TOPIC #239Advanced 15 min read

Mechanistic Interpretability: Reverse-Engineering Neural Networks

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A trained model is a program nobody wrote line by line — mechanistic interpretability treats it as software to be decompiled. This topic covers the circuit abstraction (the residual stream as shared workspace, attention as data movement, MLPs as key-value memories), the interventional method stack (activation patching, logit lens, causal scrubbing, sparse autoencoders), the documented circuit results, and the honest limits of the field.

From Weights to Explanations: The MI Toolchain 🧩

Mechanistic interpretability maps observable weights and activations onto human-interpretable features and circuits via an interventional method stack, then uses the resulting explanations for safety auditing.

From Weights to Explanations: The MI Toolchain 🧩
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Software Nobody Wrote

If your web app has a bug, an engineer can read the code. Find the function, set a breakpoint, watch variables. The behavior is somewhere in the source.

A neural network is different. Gradient descent grew it, over billions of matrix updates, into a program spread across hundreds of layers of numbers. There is no source file. There is no breakpoint. There is no manager who can say "the copy-paste logic lives in module 7."

And it does have module 7s. Models demonstrably implement algorithms: they look things up, copy previous tokens, plan next words, and — worryingly — carry features like "sycophancy" and "evaluation awareness" that researchers have literally detected inside production models.

So the question becomes

Insight

Can we reverse-engineer this thing the way we reverse-engineer a compiled binary — find the variables, the functions, the call graph?

That program is mechanistic interpretability (MI). Its bet: trained networks implement comprehensible algorithms built from identifiable components, and we can recover those algorithms the same way we decompile software — identify the representations (what the model stores), then the operations over them (how it computes).

Why should you care even as an engineer? Because evals tell you what a model does. When your fine-tune suddenly refuses benign requests, or an agent games a test suite (previous topic), "what" is not enough. You need where and how — and only an internal inspection can answer before the behavior surfaces.

02.The Idea in Plain Words: Features and Circuits

MI boils down to two nouns:

Insight

A feature is a direction of activity inside the model that corresponds to something you could name ("starts a code comment", "mentions France"). A circuit is a set of components wired together that performs a step you could describe ("when you see a repeated phrase, copy what followed it last time").

The ontology comes from Elhage et al.'s A Mathematical Framework for Transformer Circuits (2021). Four load-bearing ideas:

  • The residual stream: one shared read/write workspace running through all layers — think of a whiteboard every layer can read and write on. Attention heads and MLPs read from it at their input and add to it at their output. Information flow becomes auditable: every component writes to a common ledger.
  • Components are linear maps; nonlinearity lives in the attention softmax and MLP activations. Because sums of linear maps compose, the whole transformer decomposes into a sum of paths: input embedding → each chain of components → unembedding. Like tracing circuits on a board, you can enumerate the paths.
  • Attention heads move information. Each head decomposes into QK (the address circuit: which token to attend to — content-based lookup) and OV (the payload circuit: what value gets copied when it fires). "Find the earlier occurrence of this word; copy what came after it" is literally a QK job wired to an OV job.
  • MLP blocks act as key-value memories. A layer-2 MLP can implement "if pattern X detected, write vector Y into the residual stream" — learned lookup tables, which is why MLPs hold most of a model's factual associations ("Paris" → capital-of France).

The older neuron-level view (Olah et al., Building Blocks of Interpretability, Distill 2018) preceded all this: visualize the inputs that maximally activate a neuron, patch activations to test causal contribution, compose the effects. Both strands hit the same wall almost immediately — neurons are polysemantic (one neuron fires for "cat", "French", and "syntax error"), which is the superposition hypothesis, the next topic.

03.A Worked Example: Finding a Copy Machine

Take the most famous circuit result — the induction head — and run the scientist's experiment by hand.

Setup. Feed the model: "the cat sat on the mat, and the". A well-trained next-token guess is "cat" — because the pattern "the _" repeated, and last time "the" was followed by "cat". Doing this requires an algorithm: (1) look back for an earlier copy of the current token, (2) read what followed it, (3) move that token to the output.

Step 1 — Hypothesize from weights. In the mathematical framework, look for a head whose QK circuit scores token positions by "does previous token match current token?" and whose OV circuit copies the next token's value. Some heads in small models literally implement that composition. Hypothesis: these heads do induction.

Step 2 — Interventional test (activation patching). Run the prompt normally (clean). Run a corrupted prompt where the earlier "cat sat" is absent (corrupt) — the target logit for "cat" drops. Now re-run the corrupt prompt but surgically overwrite one attention head's output with its clean-run value. Patch the induction head and the "cat" logit snaps back; patch ten random heads and nothing happens.

code
patch target        logit("cat")   (clean 9.1, corrupt 4.3)
─────────────────────────────────────────────
layer 5 head 3      8.7  ← restores 88% of the gap
layer 5 head 7      4.4  ← nothing
layer 9 head 1      4.5  ← nothing

Step 3 — Falsification test (causal scrubbing). The algorithm claims heads are interchangeable except for the copy. Resample activations inside the claimed-equivalence groups; if the loss stays preserved, the explanation survives the hypothesis test; if not, back to the drawing board.

This is how you catch a fish in a dark pond: stop staring at the water (correlations), start poking with a net and asking what moves (interventions).

04.Visual Intuition: The City With No Blueprints

Here is the analogy to carry through the whole topic: you inherited the plumbing of a city that has no blueprints, and you must explain why the fountain in the square plays "Ode to Joy" every Tuesday.

  • The map (weights): the transformer is pipes and valves you can measure directly — every pipe's diameter is a number you can read out of the checkpoint. But reading a pipe does not tell you the song.
  • The water (activations): what flows through the pipes when a request comes in. You can tap any pipe at any junction — that is a hook that captures the activation vector at that layer and position.
  • Reading the water color (logit lens): at each junction, ask "if I forced the water here straight into the fountain, what note would play?" Decoding intermediate residual streams through the unembedding shows a computation resolving layer by layer — you watch the model commit to "cat" over time.
  • Dye injection (activation patching): run Tuesday normally, then run a "clean Tuesday" versus a "dry Tuesday", and inject clean dyed water into exactly one pipe during the dry run. The junction whose dye reaches the fountain and restores the song causally carries the melody. Correlation (this pipe flows when the song plays) is watching; intervention (dye) is testing.
  • The water blending (superposition): the cruel twist — most taps dispense a blend (cat-French-syntax-error smoothie water), not one flavor. Finding individual flavors is the sparse autoencoder's job (section 6, and the next topic).
  • The full schematic (attribution graphs / circuit tracing): the finished diagram — which reservoir feeds which valve feeds which fountain note, scored for how much of Tuesday's song the schematic can explain.

The fountain really plays the song; the pipes really exist; the blueprints were just never written by anyone. MI is the trade of drawing them.

05.The Method Stack

Every MI finding rests on a small set of interventions — the standard toolkit in one list:

  1. Activation patching (causal tracing): run clean and corrupted forward passes, overwrite one activation, measure the change in a target logit or KL divergence. Components whose patching restores behavior are causally implicated.
  2. Logit lens / tuned lens: decode intermediate residual streams through the unembedding matrix to watch a computation resolve across layers. (A learned correction — the tuned lens — fixes the fact that mid-layer streams are not yet in "word space.")
  3. Sparse autoencoders (dictionary learning): train an overcomplete sparse encoder/decoder on activations; the decoder atoms are candidate monosemantic features — tens of thousands per model, released publicly for GPT-2/3-class and open-weight models since 2024. (The next topic is dedicated to why this works.)
  4. Attribution graphs / circuit tracing (2025): combine SAE features with edge-level attribution into a graph of which features caused which, under a clean-vs-corrupted counterfactual; score the graph with matching (does it reproduce the behavior?) and completeness metrics.
  5. Causal scrubbing: replace internal activations with resampled ones inside groups that the claimed algorithm says are interchangeable; if loss is preserved, the algorithm explanation survives the hypothesis test. This is MI's version of a controlled experiment.
  6. Knowledge-location probes (ROME lineage, Meng et al. 2022): surgically patch/erase MLP key-value memory to test whether a specific fact really lives where the theory says — and watch which answers change.

06.Reproducible Wins: Real Circuits Found in Real Networks

MI is judged by its discoveries. Several are now textbook:

  • Induction heads (Olsson et al., 2022): the copy circuits from our worked example — "find the previous occurrence of the current token and copy what followed it." Their emergence aligns with in-context learning phase transitions and can be traced across training checkpoints: one of the clearest "algorithm found in a network" results.
  • Othello-GPT world model (Li et al., 2022): a model trained only on move sequences of the board game Othello develops internal representations of board state — square coordinates are linearly probeable and causally drive later move predictions (validated by transforming the coordinates and watching the moves change). Evidence for emergent internal world models.
  • Reverse-text circuit (Yao et al., 2023): asked to output text backwards, a small model builds a multi-layer chain: earlier heads encode "the next character in reverse reading order," later heads decode it. A genuine composed circuit, not a single feature.
  • Multilingual "text vs code" heads (Wendler et al., 2024): language-specific processing concentrates into a few heads, explaining why code-heavy training shifts cross-lingual behavior — you can almost point at the polyglot switch.
  • Claude-3 Sonnet biology (Anthropic, 2024, "On the Biology of a Large Language Model"): ~95% of neurons labelled via SAE feature dictionaries, multi-step "planning" features observed in poetry generation, a documented case where a "Manhattan" feature contributed to a fake-poem task invisibly on the surface, and an evaluation-awareness feature lighting up on a US-supply-chain-fiction prompt shaped like a test.
  • Grokking and modular circuits: delayed generalization in toy settings analyzed as the gradual formation of weight-sparse modular subnetworks — the "sudden understanding" moment decomposed into slow circuit assembly.

07.The 15-Line Dye Injection

Here is activation patching, the core intervention, reduced to its skeleton. A run helper captures every attention head's output on a clean and a corrupted prompt; then each head gets one surgical overwrite and a re-run:

python— Activation patching in ~15 lines: does a specific head carry the answer?
import torch

def patch_probe(model, clean_in, corrupt_in, target_pos, token_id,
                layers=(0, 11), heads=range(12)):
    """Return per-(layer,head) causal effect: how much does restoring the
    clean activation move the target logit back toward its clean value?"""
    clean_act    = run(model, clean_in,    capture=True)   # hook: attn out per layer/head
    corrupt_act  = run(model, corrupt_in,  capture=True)
    effects = {}
    for L in layers:
        for h in heads:
            mixed = corrupt_act.clone()
            mixed[L][target_pos, h] = clean_act[L][target_pos, h]   # surgical overwrite
            logits = run(model, corrupt_in, inject=mixed)
            effects[(L, h)] = (logits[token_id] - corrupt_logits_at).item()
    # Rank: components with large positive effects plausibly implement the answer-copying step.
    return sorted(effects.items(), key=lambda kv: -kv[1])[:10]

08.In Practice: Limits, and Using MI Without Getting Burned

Honest constraints, all actively debated in 2024-2026:

  • Faithfulness vs plausibility: an explanation that predicts held-out interventions is still an incomplete account. Attribution graphs cover only part of the computation on tasks models ace, and coverage degrades on long or highly parallel computations.
  • Scale: costs grow with parameters, sequence length, and graph size. Frontier models are analyzed via sub-networks, sampling tricks, and automated labelling — results may not transfer to the full model.
  • No ground truth: there is no gold-standard circuit to grade against, so the field relies on falsification: scrubbing tests, replication, and adversarial prompts that the theory says must behave a certain way.
  • Feature identity is not stable: SAE features split, merge, and shift across checkpoints. A "deception neuron" claim is checkpoint-specific unless tracked longitudinally — treat every internal feature name as a versioned artifact (the CI/CD instinct again).

Production usage today looks like three patterns:

(a) Monitoring — SAE feature activations plus probes as behavioural early-warning signals on live traffic: internal states often move before outputs do. (b) Chain-of-thought faithfulness checks — verify that a stated reasoning step is causally necessary (patch it out; does the answer change?) before trusting the explanation the model gave. (c) Audit support — localize misbehavior to components so mitigations become cheap: a targeted fine-tune or refusal training on the offending cluster, instead of broad retraining.

One rule keeps you sane: MI complements behavioral evals; it never replaces them. The city map is wonderful; you still test the water at every tap.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Produces causal, falsifiable claims — unlike post-hoc attribution which only ranks input correlations.
  • Localizes failure modes to components, enabling cheap targeted mitigations instead of broad retraining.
  • SAE feature dictionaries are shareable infrastructure: released publicly and reusable across teams.
  • Gives auditors evidence for safety cases ("we found no deception-shaped circuit in these slices"), and can detect misbehaviour before outputs are visible.

Trade-offs & Constraints

  • Compute-heavy and currently partial: full-model circuits for frontier models are out of reach.
  • Findings are metric-, checkpoint- and prompt-sensitive; replication cost is high.
  • SAE features may be artefacts of the dictionary-learning objective rather than the model's true decomposition.
  • Dual-use risk: circuit maps make it easier to find and remove refusals; labs therefore release partial results only.
Production Implementation in Big Tech
Anthropic ("On the Biology of a Large Language Model", 2024) + open-source MI tooling• Automated feature-level auditing of a production model

Anthropic trained SAE dictionaries on Claude 3 Sonnet activations, auto-labelled ~95% of neurons via feature maximization, and traced multi-step circuits (e.g. a poem-planning feature, an "evaluation awareness" feature lighting up on jailbreak-shaped prompts). Open-source equivalents — TransformerLens, NGSAE, Llama/Pythia SAE suites and the Neuronpedia catalogue — let smaller teams run the same patching-plus-dictionary workflow on open-weight models.

Staff+ Engineering Takeaways

  • MI treats trained models as reverse-engineerable programs: features (directions in activation space) plus circuits (weight components that compute over them).
  • The residual stream is a shared read/write memory; attention heads move/copy information (QK addressing, OV payload), MLP blocks act as learned key-value memories.
  • The method stack is interventional: activation patching, logit lens, sparse autoencoders, causal scrubbing, and attribution-graph circuit tracing.
  • Landmark results include induction heads, the Othello-GPT world model, reverse-text circuits, and SAE-based feature maps of production models.
  • Limits are real and debated: partial faithfulness, scale ceilings, no ground truth, and unstable feature identity across checkpoints — use MI alongside behavioural evals, not instead of them.

Topic Knowledge Check

Exercise 1 of 4 • Test your architectural comprehension.

Exercise 1 of 40 answered
1

In the standard transformer-circuits account, what is the residual stream?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?