Probing Classifiers: What Do Models Actually Represent?
A probe is a tiny second model trained on a frozen model's activations to ask: "can this property be read off?" This topic covers how to build one, why a high probe score does NOT prove the model uses the property, and how control tasks, model-free baselines and amnesic counterfactuals turn a decoding test into causal evidence.
Probe Pipeline and the Validation Battery 🧪
A probe measures what is *decodable* from activations. Turning "decodable" into "represented and used" requires control tasks, model-free baselines, causal interventions and sensitivity sweeps.
01.The Problem: You Cannot Ask a Model What It Knows
When a language model answers a grammar question correctly, you cannot simply ask it how it did that.
Ask it directly and you get a confident story, not the truth. Models invent plausible explanations — the same confabulation you met in hallucination topics.
So interpretability researchers built something better than an interview: a test.
Say you suspect a model internally represents part-of-speech, dependency structure, sentiment, truthfulness, or the language of its input. You want one number that answers
Is this property written anywhere in the model's internal activity?
The tool is a probe (also called a classifier probe or diagnostic classifier):
A probe is a tiny second model, trained on the frozen activations of the first model, that tries to read one property off them.
The claim being tested is: "if a low-capacity probe can recover X from h_L, then X is linearly/near-linearly available in h_L."
One word to make plain: an activation h_L is just the list of numbers the model writes at layer L while processing your input. "Frozen" means the big model is not trained any further — you only run it forward and copy its notes.
Probes became the workhorse of NLP interpretability because they are cheap and transfer across models. But they come with a famous warning (Belinkov, Probing Classifiers: Promises, Shortcomings, and Advances, 2022) that powers the rest of this topic:
Accuracy alone is a weak evidentiary standard. It says something is decodable, not that the model uses it.
First the recipe. Then the trap. Then the fixes.
02.The Idea in Plain Words: The Probe Recipe
Build one properly and you touch exactly five knobs:
- Pick the representation site. One token's vector, the
[CLS]/last-token summary, or mean-pooling over the sentence? And probe every layer — results vary enormously with depth. - Choose probe capacity on purpose. A linear probe is the conservative choice: it only asks "can a straight cut separate these classes?" An MLP probe instead measures "extractable with bounded computation".
- Label your data with an automatic tool (a parser, a classifier, a rubric) — and remember that label noise sets the accuracy ceiling.
- Split by input, never by (input, label) pairs. Otherwise the probe memorizes examples and reports inflated numbers.
- Report variance across random seeds and initializations. Single-run differences under ~2 points are noise, not findings.
That is the whole mechanism. The probe itself can be as dumb as logistic regression on cached activation vectors — the training step is ordinary machine learning. The only unusual thing is where the features come from: another model's internal states.
Think of it as a reading-comprehension test applied to the model's scratch paper. If a tiny trained reader can answer questions from the scratch paper at layer 8, then something about the answer is written at layer 8.
The remaining sections are about how carefully you have to phrase that sentence.
03.A Simple Worked Example: What Does 92% Even Mean?
Let's put tiny numbers on the trap. You probe layer 8 of a language model for a syntactic relation (which word depends on which), in a binary version of the task, so pure guessing scores 50%.
Your report has three numbers:
- Probe accuracy: 92% — impressive!
- Control task (same pipeline, labels randomly permuted): 51% — good, your pipeline isn't leaking.
- Model-free baseline (an independent classifier predicts the property from the raw input text, never looking at the model): 90%.
Now stare at that last line.
If the property is that predictable from the sentence itself, the probe's 92% mostly reflects facts about English, not facts about the model's computation. The model didn't have to represent the relation in any deep sense — the answer is sitting in the input.
Change only the third number: model-free baseline = 55%. The story flips. Now the activations carry structure that is not lying around on the surface — the probe became evidence about the model.
So never report a probe score alone. Always report the triple:
- probe (decodable?)
- shuffled-label control (is my pipeline honest?)
- model-free baseline (did I learn about the model, or about the data?)
Every gap between them is a different conclusion.
04.Visual Intuition and the Analogy: The Chef and the Recipe Book
Carry one picture through the whole topic: a chef cooking from a notebook.
The model processing an input is the chef mid-recipe. The activations at layer L are page L of the notes the chef writes while cooking. A probe is a translator who reads page L and guesses what the notes are about.
codeinput: "The cat sat on the mat" | v +--------------+ page 4 +---------+ | |-------------->| probe |-- 92%: "noun-phrase here" | frozen | page 8 | (tiny | | model |-------------->| reader) | | | page 12 +---------+ +--------------+--------------------------> ... | v final output (the finished dish)
If the translator reads "cumin: 2g" written on page 8, they have proven one thing only:
cumin is written down.
But a kitchen supports three different claims, and papers constantly confuse them:
- Cumin appears in the notebook. — decodable
- The chef keeps that note in a form they can act on. — represented in usable form
- The cumin actually went into the dish. — causally used
And here is the elegant trick: to test claim 3, erase the cumin line before the chef reaches that step and see whether the dish changes. That is precisely what amnesic probing does — remove the property from the activations and measure whether behaviour degrades.
The rest of this topic is about which of the three claims each experiment can actually support.
05.Cause vs Correlate: Why Accuracy Alone Is Weak Evidence
A probe can score high for reasons that have nothing to do with the model's computation. Four documented traps:
- Prediction vs use. If activations encode the previous token, a probe can predict syntactic properties the model never actually uses for its task. The honest comparison is the model-free baseline: how well can
Xbe predicted from the raw input by an independent classifier? If that matches the probe, the probe reveals little about the model. - Bottleneck artefacts. LayerNorm, attention averaging, and positional encodings leak surface statistics (length, position, character n-grams) into every layer — producing high probe scores for trivially available properties.
- Over-probing. Very high accuracy often reflects the probe's own capacity and label leakage, not representation quality.
- Under-probing. Failure to decode does not prove absence either: the property may sit nonlinearly, in a different subspace, or in a format a simple probe can't reach. Probes give upper bounds on what is decodable, not lower bounds on what is represented.
Two remedies became standard:
Control tasks (Vig & Belinkov; Hewitt & Manning): run the identical pipeline on randomly permuted labels. A real pipeline must collapse to chance on them; if the probe still scores well, your plumbing is leaking.
Amnesic probing (Elazar et al., 2021): take the direction the probe found and project it out of the representation — surgically erase the cumin line — then re-measure the downstream task. If removing X degrades behaviour, X was causally used. This reframes probing from "decode and report" to "decode, ablate, and measure behaviour", a genuinely stronger design — though the erasure can also damage unrelated information sharing the same subspace, so interpretation still takes care.
06.Why AI Cares: Probing for Safety in 2024-2026
Probes stopped being only a linguistics tool. They are now practical monitoring and audit instruments:
- Truthfulness / confidence probes. Linear probes on mid-layer activations predict factual correctness and self-consistency — sometimes better than the model's own verbalized confidence. Used as cheap abstention signals in RAG pipelines (RAG = retrieval-augmented generation, the pattern of feeding retrieved documents into the prompt).
- Emotion & reward-direction probes. Anthropic (2025) extracted named "emotion vectors" from Claude via contrast prompts and showed they causally shift behaviour on vignette tasks (a despair vector raised reward-hacking rate by 40%+). That is exactly probe-plus-intervention used as an alignment audit.
- Deception / evaluation-awareness monitors. Probes over concept directions are candidate signals for "the model believes it is being tested" — though Anthropic's own work found monitor activation was often driven by surface cues of test-shaped text rather than the underlying construct (a model-free baseline problem, in the wild).
- Refusal and safety-relevance detectors at the embedding layer (Llama Guard-style classifiers) are production probes: fast, cheap, small, and auditable as a separate artefact from the model.
- Language and modality probes debug multilingual regressions: they reveal whether a fine-tune collapsed language-specific directions.
The engineering guidance: a probe monitor is a model artefact under test. It needs its own eval set, its own false-positive rate, drift monitoring against model updates, and a documented mapping from "probe fires" to "operator action".
import numpy as np, torch
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GroupKFold
def concept_direction(h_pos, h_neg): # h_*: (n_tokens, d) activations
v = h_pos.mean(0) - h_neg.mean(0)
return v / v.norm() # unit direction, e.g. "harmfulness"
def train_probe(X, y, groups):
clf = LogisticRegression(max_iter=2000)
cv = GroupKFold(n_splits=5) # groups = distinct inputs!
scores = cross_val_score(clf, X, y, cv=cv, scoring="f1_macro")
return scores.mean(), scores.std()
acc, sd = train_probe(H, labels, input_ids)
ctrl, _ = train_probe(H, rng.permutation(labels), input_ids) # CONTROL TASK
print(f"probe={acc:.3f}±{sd:.3f} control={ctrl:.3f}")
# Evidence requires: acc >> control AND acc > model-free baseline from raw text.
# Then test causal use: amnesic projection removing v, re-measure task loss.07.In Practice: Reporting Standards That Survive Review
If a probe result will inform a release decision or an audit, meet this bar — each item closes one of the traps above:
- Multi-layer, multi-seed sweeps reported as a curve, not one number; say which layer you chose and why (quoting only the best layer inflates results — report the whole profile).
- Capacity ablation: linear vs 1-hidden-layer vs high-capacity probe, plus the model-free control classifier on raw text.
- Label-quality check: inter-annotator agreement or parser accuracy for the probing task; state the ceiling your labels impose.
- Statistical treatment: paired bootstrap over inputs; single-run differences are not findings.
- Downstream validation: either amnesic/causal ablation, or behavioural tests — does the probe's signal predict failure modes on held-out adversarial prompts?
- Superposition caveat: under superposition (models packing more concepts than dimensions), the "true" concept may be spread over several directions, so one linear probe underestimates availability; SAE-based features are the complementary instrument.
The translator in the kitchen can tell you what is written down. Only the erased-line experiment tells you whether the dish depends on it. Keeping those two sentences apart — in your methods section and in your safety case — is the entire discipline of modern probing.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Extremely cheap: one forward pass with hooks, then logistic regression on cached activations.
- Model-agnostic and transferable across architectures, layers and tasks.
- With control tasks and amnesic interventions, produces evidence about causal use, not just correlation.
- Naturally becomes a runtime monitor (probe = tiny classifier) for content safety, abstention and drift.
Trade-offs & Constraints
- Decodable ≠ used; naive accuracy claims are routinely overturned by model-free baselines.
- Nonlinear/superposed encoding means negative results are weak evidence.
- Probe findings are checkpoint-specific; they decay silently when the base model is updated.
- Label noise, pooling choice, and best-layer selection create easy self-deception.
Safety teams train small concept probes (harmfulness, evaluation-awareness, emotion/urgency directions) on activations of the deployed model and use them as fast pre-filters ahead of an LLM guardrail, with contrast-pair directions also supporting steering experiments. The same pattern is exposed as standalone products: Llama Guard and prompt-injection classifiers are effectively production probing pipelines with their own versioning and false-positive budgets.
Staff+ Engineering Takeaways
- A probe measures what is *decodable* from a frozen representation; converting that into "the model uses it" requires intervention.
- Standard validation stack: control tasks (permuted labels), model-free baselines, multi-layer/multi-seed sweeps, and downstream behavioural checks.
- Amnesic probing (project the found direction out and measure behaviour) and activation patching are the causal upgrades to classical probes.
- Negative results are weak: nonlinear or superposed encoding can hide from simple probes.
- Probes double as cheap production monitors — but they are models themselves and need eval sets, drift monitoring, and version pinning.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
A probe predicts dependency relations at 91% accuracy, and a model-free classifier on the raw input reaches 90%. What is the right conclusion?
How clear and actionable was this distributed systems breakdown?