Extended Thinking and Test-Time Compute
For twenty years, AI capability grew by spending compute during training. From 2024-2026, a second axis matters just as much: compute spent AT INFERENCE — thinking longer (sequential) or trying more candidates and verifying them (parallel). This topic covers the scaling laws behind both, why verifier quality is the ceiling, how difficulty-based allocation lets a small model beat a 14x larger one, and why every API now sells "thinking effort" as a priced knob.
01.The Problem: Training Compute Hit the Price Wall
Here is the old religion, in one plain sentence: bigger models trained on more data get better, following smooth scaling laws (Kaplan 2020, Chinchilla 2022 — the "parameters × data × FLOPs" curves).
By 2023 that religion had a problem:
- Frontier training runs cost tens of millions of dollars and the next doubling costs roughly the same again.
- And the curve is smooth — no amount of the same kind of training produces a qualitative jump.
Meanwhile, a different observation kept repeating in labs: the model was underperforming its own ability. Ask a model the same problem ten times and it solves it sometimes; the single-answer score was understating what the weights already contain.
So the question became
Instead of buying a better model... can we just use the one we have harder?
Spend extra compute when answering — think longer, or try more answers and check them. Capability that was latent, unlocked by a bill instead of a training run.
That is test-time compute (TTC) — called TTT, inference scaling, or "thinking budget" depending on the vendor. It became the second scaling axis of 2024–2026, the engine behind every "reasoning effort" dial in every API, and the reason small open models can sometimes beat giants.
Two Inference Scaling Axes and the Effort Knob 🎛️
Two Inference Scaling Axes and the Effort Knob 🎛️
Test-time compute splits into sequential (longer deliberation) and parallel (more samples, better verification); allocation by prompt difficulty decides whether the budget buys capability or waste.
Unlock Topic #219: Extended Thinking and Test-Time Compute
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?