Emergent Capabilities: Real Phase Transitions or Measurement Artifacts?
Some AI abilities look like they switch on overnight once a model is big enough — like water suddenly boiling. A famous 2023 critique says the "switch" is often just your grading rule: the skill was growing smoothly the whole time, but your test only reports all-or-nothing scores. This topic walks through both sides and what you should actually do when planning products.
The Emergence Debate 🌊
The same curves read as "phase transition" under exact-match scoring and as "smooth improvement crossing a usability threshold" under continuous metrics.
01.The Problem: A Skill That Appears Overnight?
Imagine you test three AI models on 3-digit addition.
Small model: 2% correct. Medium model: 4% correct. Huge model: 82% correct.
Nothing in between.
It looks like the ability switched on.
Did the big model suddenly gain a skill the small ones did not have?
That is the puzzle of emergent capabilities.
Most model improvements are boring and smooth.
As you add parameters, the loss drops along a gentle curve.
Predictable. Boring.
But a handful of tasks — multi-digit arithmetic, chain-of-thought reasoning, translating languages the model never saw — look different.
They sit flat near zero for a long time.
Then they jump.
So the question the whole field argued about:
Is the cliff real, or is it an illusion created by how we measure?
02.The Idea in Plain Words: Two Camps, One Graph
Camp 1 — the emergence claim.
Wei et al. (2022, "Emergent Abilities of Large Language Models") define an ability as emergent if:
- it is absent in small models, and
- it is unpredictable from their scaling curves, then
- it appears sharply past a parameter/FLOP threshold.
Their examples from the paper and follow-ups:
- Unprompted multi-digit arithmetic: near 0% for 3-digit addition until roughly 100B parameters, then it jumps.
- Chain-of-thought prompting (Topic 160 — the trick of adding "Let's think step by step" style worked examples to the prompt): reasoning gains were negligible below ~62B parameters and materialized above.
- Analogies, word unscrambling, few-shot translation of unseen languages — similar cliff shapes.
Wei et al. also separated two ideas:
- Quantitative emergence: a sudden jump in a discrete metric. The curve changes shape.
- Qualitative novelty: a genuinely new kind of ability. Much rarer claim.
They noted emergent tasks typically require many composed steps, where each individual step is near chance on its own.
Camp 2 — the mirage critique.
Schaeffer, Mirzasoleiman & Shah (2023, "Are Emergent Abilities of Large Language Models a Mirage?", arXiv 2304.15004) argued:
The cliff is not in the model. The cliff is in your metric.
The model's underlying competence improves smoothly and predictably.
But the evaluation metric — exact-match accuracy, SQuAD-style EM, F1 with rounding — is discontinuous.
It squashes a smooth sigmoid of true ability into a visual cliff of reported score.
Their supporting evidence:
- Shrinking a "cliff" task's metric turns the curve smooth. Example: report token-level edit distance (how many edits away from the right answer) instead of exact match, and many "emergent" curves become gentle slopes.
- ICL variance: just adding a handful of in-context demonstrations shifts scores across the threshold — so the "jump" partly reflects prompt artifacts, not model physics.
- Cramming experiments (Ruan et al.): capabilities that the famous 2022 curve said would "only appear" at ~10^25 FLOPs can be forced out at ~10^20 FLOPs by intensive training. Emergence partly measured training under-deployment, not a law of nature.
The honest 2026 synthesis:
- Emergence as a literal phase transition (like water boiling) is contested.
- But usability thresholds are real — a capability that exists yet runs at 60% accuracy is functionally absent in production. Nobody ships a calculator that is wrong 4 times in 10.
03.A Tiny Worked Example: How a Grading Rule Builds a Cliff
Same student, two report cards.
Suppose a model's skill at 3-digit addition improves in per-digit accuracy, smoothly, as it grows:
codemodel size per-digit accuracy 1B 30% 10B 50% 60B 70% 100B 85% 300B 97%
Now grade it two ways on the FULL problem (3 digits, all must be right):
Probability the whole answer is exactly right ≈ (per-digit accuracy)³:
code1B: 0.30³ ≈ 3% 10B: 0.50³ ≈ 12% 60B: 0.70³ ≈ 34% 100B: 0.85³ ≈ 61% 300B: 0.97³ ≈ 91%
Notice what happened.
- The underlying skill is a smooth staircase: 30 → 50 → 70 → 85 → 97.
- The all-or-nothing score is steeper, and if you round anything below ~5% to "the model cannot do this", you get: 0%, 0%, ~0%, then 61%, 91% — a cliff.
Same model. Same training. Only the tape measure changed.
That is the entire mirage argument in three lines of arithmetic.
04.Visual Intuition: One Curve, Two Stories
Plot score (vertical) against model size (horizontal).
The true competence line keeps bending upward, quietly:
codescore ▲ │ ● 91% exact-match │ ● 61% / │ ╱╱╱╱╱╱╱╱╱╱╱╱╱╱╱ ← true skill, smooth │ ╱╱╱╱ │╱╱╱ 0 └────┴──────┴──────┴─────┴─────► model size 1B 10B 100B 300B exact-match metric "reads" this as: 0% ──── 0% ──── 0% ─── ⚡JUMP⚡ ─── 91% (anything tiny rounds to "absent" until it doesn't)
The emergence camp sees the ⚡ as a phase transition — water boiling at 100°C.
The mirage camp sees the ⚡ as a ruler with only two markings.
Here is the key picture in one line:
codesmooth ability → [discontinuous metric] → cliff-shaped score (exact match, EM, F1)
Swap the box for a continuous one (edit distance, per-step credit, partial points) and the cliff flattens back into a slope.
05.The Analogy: The All-or-Nothing Driving Test
Carry one story through everything here: a driving school that grades badly.
A student practices parallel parking for months.
Her skill improves smoothly — 30 cm off, then 20 cm, then 10 cm.
But the examiner only checks one box:
"Did the car go in with ZERO corrections? Yes/No."
Every student fails for the first year. Pass rate: 0%.
Then one week, suddenly 80% pass.
The school announces: "Parallel parking ability EMERGED at hour 40! It is a phase transition!"
An outsider checks the practice logs and says:
"Her steering error shrank smoothly the whole time. The cliff is in your pass mark."
Both statements are useful:
- The metric cliff is an artifact (mirage camp).
- But a license is genuinely binary — a driver who parked OK 60% of the time cannot be hired to park ambulances (usability-threshold camp).
In AI: per-digit accuracy / edit distance is the practice log. Exact-match score is the pass/fail box. Your production users only experience the pass/fail box.
06.Why AI Cares: This Is a Budget Decision, Not Philosophy
The emergence debate dictates how you build with LLMs. Four concrete rules:
- Never assume a capability is absent because a small-model eval showed ~0% exact-match. Probe with continuous metrics, few-shot prompts, and partial-credit scoring before writing a model size off. The skill may be there at 34%, just rounded to zero by your metric.
- Conversely, never assume scale will rescue a task. Compositional, multi-step reliability historically arrives late — and it burns money to wait. RL-trained reasoning models from 2024–2025 shifted this again: reliability became something you can buy at test time with thinking tokens, not only at train time with more parameters.
- Buy capability per-token, not per-parameter: frontier labs now report capability-vs-inference-FLOPs curves, because that is the axis your product actually rides — your bill arrives per generated token, not per parameter.
- Red-team the metric: if your roadmap depends on a benchmark cliff, check whether a smoother metric (or verifier-based grading, where a program checks each step) removes it.
import math
# pretend per-digit accuracy as model size grows
per_digit = {1: 0.30, 10: 0.50, 60: 0.70, 100: 0.85, 300: 0.97}
for size, p in per_digit.items():
exact_match = round(p ** 3, 2) # all-or-nothing: 3 digits all right
edits = round(3 * (1 - p), 2) # continuous: expected wrong digits
print(f"{size:>4}B exact-match={exact_match:<5} mean-edits={edits}")
# exact-match: 0.03, 0.12, 0.34, 0.61, 0.91 <- 'cliff' past 60B
# mean-edits: 2.1, 1.5, 0.9, 0.45, 0.09 <- smooth the whole wayArchitectural Trade-offs & Production Realities
Architectural Advantages
- The emergence lens usefully flags that compositionally hard tasks behave differently from recall tasks.
- Usability thresholds are a legitimate engineering concept: 60%-reliable tools genuinely are absent for production.
- The mirage work forced better evaluation hygiene across the field.
Trade-offs & Constraints
- Cliff-waiting is expensive: betting that scale will deliver a capability is an unbounded budget gamble.
- Discontinuous metrics hide signal below the threshold and cause both false alarms and false confidence.
- Test-time reasoning compute partially decouples capability from model size, complicating every scaling story.
Gemini's development emphasized capability prediction across sizes, and later labs' "benchmark-forecasting" work (fitting curves on small probes to predict pass@k on SWE-bench-class tasks) is applied emergence science: deciding which cliff shapes are real and predictable before a model ships.
Staff+ Engineering Takeaways
- Emergent = ability appears abruptly past a scale threshold and is unpredictable from smaller models (Wei et al. 2022).
- Schaeffer et al. 2023: many "cliffs" come from discontinuous metrics (exact match, accuracy) applied to smoothly improving competence.
- Practical truth: hidden capability may exist below thresholds, and small gains can cross real usability cliffs.
- Choose continuous, per-step metrics and few-shot probes before committing to model sizes.
- 2024–2026: reasoning via test-time compute adds a fourth axis to the size-vs-capability story.
Topic Knowledge Check
Exercise 1 of 2 • Test your architectural comprehension.
According to the "mirage" critique, the apparent sudden emergence of capabilities is mostly explained by:
How clear and actionable was this distributed systems breakdown?