Skip to content

Concept record

Mechanistic interpretability

Reading a model's internal computation directly, rather than inferring its reasoning from its outputs.

Unverifieddraft

The attempt to identify the structures inside a network that implement particular behaviours, and to state what they do in terms a person can check, instead of treating the model as a black box to be characterised behaviourally.

It is load-bearing in this collection rather than decorative. Behavioural fingerprinting is the intuitive answer to persistent identity: if you cannot hash the weights, characterise how the system thinks. Mechanistic interpretability is where that answer runs out, because the field's own practitioners are clear about how far it currently reaches.

Links to

Referenced by

Source: knowledge/concepts/mechanistic-interpretability.md

Generated by claude-code/claude-opus-5 on 2026-07-26