TOPIC #240Advanced 13 min read

The Superposition Hypothesis

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Networks store far more features than they have neurons by overlapping them at near-perpendicular angles — that is the superposition hypothesis. It explains why one neuron fires for "cat", "French", and "syntax errors" (polysemanticity), why penalizing activity makes neurons specialize (monosemanticity, shown in Anthropic's Toy Models), and why sparse autoencoders — a change of basis — are the right tool to unpack the mixture.

Superposition: More Features Than Neurons 📐

A network with n neurons can encode m >> n features by assigning each feature a direction and relying on near-orthogonality to limit interference. Sparsity pressure (explicit or via SAEs) encourages a decomposition into monosemantic directions.

Superposition: More Features Than Neurons 📐
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: More Concepts Than Room

Do a quick, depressing count.

A small language model has a residual stream of, say, 1024 dimensions — 1024 "neuron slots" per token position (the previous topic: that stream is the shared workspace). How many distinct concepts does language contain? Product names alone number in the millions, and every one is useful for next-token prediction — rare French verbs, the syntax of a Python comment, the month of January.

Pretraining pays for modelling rare concepts too. So the model faces a choice for each rare feature:

  • Drop it → pay permanent loss on every rare occurrence.
  • Give it a dedicated neuron → there are not enough neurons; frequent features would starve.
  • Share → store several features in the same neurons, overlapping, like double-parked bicycles leaning against the same railing.
Insight

The network chose to share. That choice is the superposition hypothesis.

This hypothesis, named and demonstrated by Anthropic's Toy Models of Superposition (2023), also explains the single most confusing observation in all of neuron-level interpretability (previous topic): why one neuron fires for cats, French text, and syntax errors. If you expected each neuron to be "about" one concept, superposition says you were reading the wrong coordinate system.

02.The Idea in Plain Words: Overlap, Quietly

The hypothesis, formally:

Insight

A trained network represents more distinct features than it has directions in its activation space by overlapping them — assigning each feature a (learned) linear direction so that simultaneously active features interfere as little as possible.

Unpack the geometry. A "feature" (cat-ness, French-ness) gets a direction — a unit vector d in the n-dimensional activation space. "The cat feature is active" means the activation vector slides some distance along d_cat.

The cruel and beautiful part is how sharing works: if two features are encoded along directions d_i and d_j, the reconstruction error from activating them simultaneously scales with the cosine of the angle between them:

interference ∝ cos θ_ij

  • Same direction (θ = 0°, cos = 1): total mix-up — the two features are indistinguishable.
  • Perpendicular (θ = 90°, cos = 0): zero interference — both can be fully active and the network can still read each one out.
  • Near-perpendicular: small but nonzero cost.

So the network packs m features into n dimensions by choosing near-perpendicular angles for everyone. The cost math explains the allocation strategy: a rare feature (active 1 in 10,000 tokens) tolerates a shared, slightly-tilted direction because it almost never collides with anything; a frequent feature pays interference constantly, so it gets pushed toward its own dedicated direction. Rarity buys tolerance for crowding.

03.A Worked Example: Two Dimensions, Three Features

Shrink the space until you can draw it: 2 dimensions (neurons x, y), 3 features to encode.

  • Feature A → (1, 0) — pure x.
  • Feature B → (0, 1) — pure y. Perfectly orthogonal to A: cos = 0.
  • Feature C — no room left! Tilt it: (cos 75°, sin 75°) ≈ (0.26, 0.97).

Read the consequences like a neuron would. Neuron x measures coordinate 1. Its value is now:

x = 1·A + 0·B + 0.26·C

So x fires for A and (weakly) for C. Neuron y:

y = 0·A + 1·B + 0.97·C

y fires for B and, more strongly, for C. Both neurons are polysemantic — "cat, French, errors" in miniature — and yet all three features are recoverable, because C's direction differs from A's and B's: given many examples, the network (or you, with linear algebra) can solve the mixture. The interference cost of C's tilt is only cos 75° ≈ 0.26 per co-activation, small if C is rare.

Now the Toy Models experiment, which is exactly this picture with a learning algorithm: a 1-layer ReLU network with 50 neurons trained to reconstruct inputs carrying 100 sparse features. Result: it works — via superposition — and its neurons look polysemantic. Then Anthropic added an activation penalty (making activity expensive pushes toward sparser codes): the same task reorganizes toward near-monosemantic neurons. Sparsity pressure → feature splitting. The hypothesis made a prediction; the prediction held in the lab before it held at scale.

04.Visual Intuition: Draw the Directions

Picture the 2-D example from above as arrows in a plane:

code
   y ▲            B↑
     │      C ↗  (75° — the rare tenant,
     │    ╱  ╱     squeezed in at a tilt)
     │  ╱  ╱
     │╱ ╱
     ●───────────▶ x
      A→ (1,0)

  angle(A,C)=75° → cos≈0.26  small interference
  angle(A,B)=90° → cos=0     zero interference

The x-axis neuron is a flashlight casting shadows: it only sees each arrow's shadow on the x-axis — A and C both cast a shadow. Rotate the flashlight to C's angle and you read C alone. A neuron is one fixed flashlight direction; a feature is an arrow; superposition is a room full of arrows with few flashlights.

The coat-closet version of the same picture, to carry through: a hook (neuron) with scarves (features) hanging at slightly different angles. A single glance from the door (one coordinate) blurs them into one silhouette — polysemantic. Step back, put on different glasses that look along a scarf's own angle (a change of basis), and each scarf separates cleanly. The closet never had more hooks; you changed where you stood.

05.The Neighbouring Idea: The Linear Representation Hypothesis

Superposition is usually paired with its optimistic sibling:

Insight

The linear representation hypothesis: high-level concepts are represented as linear directions (or small subspaces), and relationships between concepts as linear maps.

Canonical evidence: gender bias directions in word embeddings (Bolukbasi et al., 2016), where we − man and they − woman share a component; "sentiment" / "truth" / "language" directions probeable in middle layers; arithmetic and even spatial coordinates recovered in small models.

The reconciliation — the key insight of this whole topic — is that both statements hold simultaneously:

Insight

Features are linear directions, but they are not axis-aligned with neurons, and they are not orthogonal to each other.

A neuron is one measurement of one direction in a space where many features partially live. That is exactly why single-unit inspection (staring at one flashlight's shadow) was so confusing for a decade, and why the field moved to basis-change methods.

Where the geometry becomes practical:

  • Concept vectors: directions like "harmfulness", "sycophancy", "evaluation-awareness", estimated from contrast pairs (answers before/after the concept appears). Readable at runtime — project the activation, detect the state — and steerable — add or subtract a small multiple of the direction to nudge behaviour.
  • Refusal directions (Arditi et al., 2024): refusal behaviour rides on a small number of directions; ablating them can remove refusals. Striking technically — and striking for safety, because the same map tells an attacker where to cut (dual-use).
  • Feature interference explains fine-tuning fragility: many tasks sharing directions means updating one behaviour can degrade unrelated ones — your safety fine-tune "forgets" politeness because they hung on adjacent scarves.

06.Sparse Autoencoders: Buying a Bigger Rack

If the model's neurons are the wrong flashlight set, change basis: decompose activations against an overcomplete dictionary — more atoms than dimensions — with pressure that only a few atoms fire per input. That is precisely a sparse autoencoder (SAE): encode an activation x into a much wider latent z, penalize `z's' activity, decode back; the decoder columns become candidate feature directions.

The practical pipeline and its knobs:

  1. Collect activations at a chosen site (residual stream after layer L, or an MLP output — layer choice trades interpretability against downstream re-mixing).
  2. Train the objective min ||x − Dec(Enc(x))||² + λ·L1(z): reconstruction fidelity plus sparsity; modern variants use TopK or gating to control the active-atom count directly.
  3. Label features automatically: retrieve the activating examples for each atom, hand them to an LLM, ask for a name, and score labelling faithfulness (automated interpretability — how the field scaled past human inspection).
  4. Evaluate: reconstruction error, shrinkage / explained variance, replacement/ablation studies (does deleting the SAE hurt less than deleting equivalent neurons?), and feature-splitting behaviour.

Results so far: SAEs on GPT-2/Pythia/Llama-3/Mixtral-class models ship dictionaries of 4k–1M features per layer; Anthropic's scaling effort labelled ~95% of neurons via features and found cross-code, in-progress-plan, and evaluation-awareness features. Crosscoders (shared dictionaries across layers or models) attack the same problem with a coupled basis.

Counterpoints to carry into any design doc: SAE features reconstruct loss well but ablations sometimes show only modest behavioural effect; "feature absorption"/"splitting" artefacts appear; and one concept can be spread over several atoms — a single-atom monitor has silent false negatives.

See it run on live traffic — a tiny monitor over chosen atoms:

python— Reading a feature's activation on live traffic (SAE monitor sketch)
import torch

class SaeMonitor:
    def __init__(self, sae, feature_ids, layer=-1):
        self.sae, self.feature_ids, self.layer = sae, feature_ids, layer

    @torch.no_grad()
    def score(self, residual_stream):        # (..., d_model) at chosen site
        z = self.sae.encode(residual_stream)  # sparse latent, d_dict >> d_model
        return z[..., self.feature_ids]       # activation of monitored features

    def alert(self, stream, threshold=0.35):
        s = self.score(stream).amax(dim=-1)   # max over tokens in the window
        return (s > threshold).any().item()

# Caveat: monitor multiple atoms per concept — features split across atoms,
# so a single id gives unreliable recall. Track per-checkpoint stability too.

07.In Practice: What This Buys You and What It Costs

Why engineers should care. Superposition turns "the model is inscrutable" into a concrete, tractable statement: the model is written in a basis you were not given. Interpretability becomes linear algebra with a shopping list — find better directions (SAEs), read projections (probes), steer along them (concept vectors), monitor them in production (the code block above). It also explains real operational phenomena for free: why one neuron is many things, why fine-tunes interfere, and why safety behaviours can be cut out with a rank-one edit — which is precisely why frontier labs treat this knowledge as dual-use.

What it does not (yet) buy you. The open questions are load-bearing:

  • Do SAE atoms recover the features the model causally uses, or one convenient basis among many?
  • How do you handle concepts that split across atoms or absorb into each other?
  • Feature directions are checkpoint-dependent — claims must be re-verified per model version, the same versioning discipline as the alignment topic's "version everything".

The honest summary sentence: the network double-parks ideas in a space too small for them, near-perpendicular angles keep the fender-benders rare, and sparse autoencoders are the tow trucks that sort the parking lot — powerful infrastructure, with the caveat that nobody has proven the towed cars are the real cars the model drives.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Explains polysemantic neurons and fine-tuning interference as capacity phenomena rather than bugs.
  • Turns interpretability into a tractable optimization problem: find a better basis, not a magic neuron.
  • Enables runtime monitors and interventions (probing/steering along feature directions) on production models.
  • Predicted correctly: sparsity pressure producing monosemanticity was demonstrated in Toy Models before large-scale SAE work.

Trade-offs & Constraints

  • SAE dictionaries are expensive to train, and layer/site choice changes the feature set.
  • Reconstruction quality is not proof that atoms correspond to causally used features.
  • Concepts spread over many atoms: monitors built on single features have silent false negatives.
  • Feature directions are checkpoint-dependent — claims do not automatically transfer across model versions.
Production Implementation in Big Tech
Anthropic + OpenAI/MATS + open SAE ecosystem• Feature dictionaries as shared interpretability infrastructure

Anthropic trained SAE stacks across Claude 3 Haiku/Sonnet layers and used the resulting feature space to explain multi-step behaviour. OpenAI, with the Mechanistic Interpretability Technical Substrate, released open SAEs for GPT-3.5-class models, and community suites (SAELens, Neuronpedia, Llama/Pythia dictionaries) made per-layer feature bases a reusable artefact — effectively a "disassembler symbol table" that other teams can query instead of retraining dictionaries.

Staff+ Engineering Takeaways

  • Superposition: networks encode more features than neurons by using near-orthogonal directions; interference grows with cosine similarity between them.
  • It explains polysemantic neurons as projections onto shared directions, and predicts that sparsity pressure yields monosemantic units (shown in Anthropic's Toy Models).
  • Paired with the linear representation hypothesis, it reframes interpretability as changing to a better basis rather than finding "the concept neuron".
  • Sparse autoencoders implement that basis change as dictionary learning; released dictionaries label the great majority of neurons in some production-scale models.
  • Open questions: do SAE atoms reflect causally used features, how to handle split/absorbed features, and how to make claims robust across checkpoints.

Topic Knowledge Check

Exercise 1 of 4 • Test your architectural comprehension.

Exercise 1 of 40 answered
1

Under the superposition hypothesis, why do neurons appear polysemantic?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?