Scaling Laws and Chinchilla: How to Spend a Training Budget
Training a model costs a fixed pile of compute — how do you split it between parameters and data? Scaling laws say loss falls in smooth, predictable power laws, so you can measure small and forecast the $100M run. Chinchilla added the correction: spend on both equally, roughly 20 tokens per parameter — and production models then overtrain small ones on purpose because serving, not training, dominates lifetime cost.
Spending One Compute Budget Three Ways ⚖️
Chinchilla showed loss is minimized when parameters and tokens grow together. Production teams then deliberately overtrain small models because serving — not training — dominates lifetime cost.
01.The Problem: A $100M Run Is a Budget-Splitting Decision
Picture the moment a lab commits to a frontier pretraining run (Topic 141).
They have a fixed pile of compute — call it one budget. It can be spent in wildly different ways:
- a huge model (hundreds of billions of parameters) fed a few hundred billion tokens,
- or a modest model fed trillions of tokens,
- or something in between.
Which split gives the smartest final model?
Before 2020, that was answered by taste, papers copied between labs, and a lot of anxiety. Training was alchemy: nobody could tell you, before spending, whether the run would be good.
Then two results in a row turned the question into arithmetic.
- Kaplan et al. (2020): loss falls in a smooth, predictable curve as you grow size or data — so you can measure the curve cheaply and forecast the expensive run.
- Chinchilla (Hoffmann et al., 2022): the forecast exposes how to split the budget — and it said the giants of 2020 (GPT-3 175B, Gopher 280B) had split it badly.
This topic is those two results, the 20-token-per-parameter rule, and why production models deliberately break it.
02.The Idea in Plain Words: Loss Follows a Power Law
A scaling law is simply
A smooth formula that says: multiply your compute, and loss drops by a predictable amount — every time, across orders of magnitude.
Kaplan et al. (2020) discovered that Transformer language-modeling loss decreases smoothly and predictably as any of three factors grows:
- N — model size (parameters),
- D — dataset size (tokens),
- C — compute.
Fitting all three jointly with GPT-3-era scaling:
L(N, D, C) ≈ 3.7 + \frac{1.5}{N^{0.23}} + \frac{1.5}{D^{0.24}} \quad (power-law terms, N in the billions)
Unpack the plain consequences of that formula:
- Everything is a curve with no surprises. No magic size, no threshold where "understanding" appears — loss glides down as N or D rise (power laws: multiply the input by 10, lose a fixed amount of loss).
- You can extrapolate. Train tiny probe models on ~1% of the budget, fit the curve, and confidently predict the final loss of a $100M run before spending it.
- Budgets become plans. This turned frontier training from alchemy into capacity planning — and explains why labs raced straight up the curve to GPT-3 (175B) and Gopher (280B), assuming bigger N was the answer.
That last assumption is exactly what Section 4 takes apart.
03.A Simple Worked Example: Predicting a Run Before Buying It
Suppose your probes fit a simplified law: every 10x of compute drops loss by about 0.5 nats, in a straight line on a log scale.
codecompute (log scale) predicted loss 10^18 FLOPs 5.0 <- probe model (cheap, hours) 10^19 4.5 <- probe 10^20 4.0 <- probe 10^22 (the real run) 3.0 <- FORECAST before spending $100M
You spent pennies to pre-purchase confidence in a number. That is the whole economic magic of power laws.
Now the budget-splitting arithmetic, Chinchilla-style. Two ways to spend roughly the same training compute (compute ≈ proportional to N × D, parameters times tokens seen):
codeparameters tokens N x D (relative work) Gopher: 280B 300B 280 x 300 = 84,000 Chinchilla: 70B 1.4T 70 x 1400 = 98,000 (4x smaller) (~4.7x more data) ~same total compute
Same pile of money. Different split. Chinchilla — the 4x-smaller model trained on 5x more data — beat Gopher on 67 of 70 evals.
The lesson in one sentence: when the budget is fixed, parameters and data must grow together, not parameters alone.
04.The Chinchilla Correction: Balance N and D
In 2022, Hoffmann et al. at DeepMind ("Training Compute-Optimal Large Language Models", arXiv 2203.15556) re-derived the laws with three different estimation methods (FLOP-matched, step-wise, parametric — the result was triangulated, not a lucky fit) and found Kaplan had a critical flaw: the 2020 fits could not separate the N and D exponents, biasing labs toward over-sized models trained on too little data.
The corrected result: for a fixed training compute budget C, loss is minimized when model size and data size grow at the same rate — roughly N \propto C^{0.5} and D \propto C^{0.5}, i.e., about 20 training tokens per parameter.
In plain words: if you can afford to double one, you should double the other too. A model of 70B parameters "deserves" ~1.4T tokens; a 175B model needs ~3.5T.
GPT-3 (175B on 300B tokens ≈ 1.7 tokens/parameter) and Gopher (280B on 300B ≈ 1 token/param) were, by this standard, roughly 4x over-parameterized — giants starved of reading, when a smaller brain with five times the library wins the same fight for the same money.
The empirical punchline was Section 3's matchup: Chinchilla (70B parameters, 1.4T tokens) beat Gopher (280B parameters, 300B tokens) on 67/70 evals despite using the same training compute.
05.Visual Intuition: One Budget, Three Splits
A fixed compute budget is a rectangle: width = parameters N, height = tokens D. Same area = same compute. Chinchilla says the square-ish shape is best for loss.
codeD (tokens) 1.4T | +-----+ | | CHIN| <- 70B x 1.4T: balanced, lowest loss 300B| +----+ + +--------+ | |GOPHER| | |OVERTRND| <- e.g. 8B x 15T: slightly | +------+ | +--------+ worse loss, cheap to serve +------------+-----------------+---------► N (params) 70B 280B 8B
- Tall-thin Gopher (huge N, few D): pays for neurons it never fed.
- Square-ish Chinchilla: loss-optimal at that compute.
- Flat-wide overtrained (tiny N, oceans of D): not loss-optimal — and that is the point; see next section.
The dependency chain to memorize:
fixed C → choose split of N and D → loss minimized near N ≈ D balance → but serving cost depends on N alone → shift the split on purpose
06.The Analogy: Educating One Employee on a Fixed Scholarship
Carry one analogy: you are funding the education of a single employee, then paying their salary forever.
- Parameters N = the size of the brain you buy (bigger brain = higher salary per word they speak).
- Tokens D = the number of books you make them read.
- Compute C = the one-time scholarship that must be split between the two.
Kaplan-era labs spent nearly the whole scholarship on a giant brain and handed it three books. The genius, poorly read, bills per hour — and its ideas are shaky anyway.
Chinchilla noticed the square-root balance: the giant brain and the well-read midbrain came out at the same price, and the well-read one outperformed on 67 of 70 exams.
Then production adds the final twist: after hiring, you pay the employee per word spoken (inference), and they will speak trillions of words serving users. Suddenly the smart money hires the smaller, ferociously well-read person — cheaper per word, nearly as smart, because extra reading was almost free.
That last move is Section 7: deliberate overtraining.
07.Training Compute vs. Inference Compute: Overtraining Is Rational
Chinchilla optimizes loss at training time only. But most frontier models generate trillions of tokens over their serving life, and each inference token costs about 2x-forward-pass-FLOPs per parameter — salary per word, in the analogy.
If expected served tokens are large, the lifetime-optimal model is therefore smaller and more data-fed than Chinchilla-optimal — deliberately "overtrained" (more data than Chinchilla says the parameters deserve).
This is the regime all production 2024–2026 models live in:
- Llama 3 8B was trained on ~15T tokens ≈ 1,900 tokens per parameter — two orders of magnitude past Chinchilla's 20 — because it serves hundreds of millions of users, each call paying per-parameter salary (Topic 141's recipe).
- GPT-4o / Claude / Gemini-class deployments favor small fast models with huge context for the same reason.
But what if the library runs dry? High-quality unique text is finite. The Muennighoff et al. work on data-constrained scaling formalizes the limit: repeating data across 4–16 epochs retains most of its value, then decays — re-read books cost little but teach less each pass.
Rule of thumb for system design interviews:
Training budget chooses what you can afford to learn; traffic volume chooses how small you should have made it.
08.In Practice: Using Scaling Laws in Your Own Runs
A lab's scaling workflow today, step by step:
- Run probes: train a ladder of small models (10M–1B) at several N/D splits with shared tokenizer and data order — keep everything else fixed so the curve measures what you think it measures.
- Fit the law: estimate exponents for your architecture and data mixture — exponents shift across architectures, so never copy constants blindly (Kaplan's numbers already mislead once).
- Extrapolate and validate: predict frontier loss, then sanity-check with eval-transfer — small-model evals predict large-model evals only for smooth, non-emergent metrics (see Topic 143; an emergent capability can stay flat on the probe ladder and appear only at scale).
- Watch the frontier: loss vs. FLOPs has been closing a power-law gradient for years; the 2024–2026 era added reasoning-time compute (test-time scaling) as a fourth axis, where extra tokens generated at inference buy capability once bought only at training (Concept 137's o-series).
The three-axis map to leave with:
codeparameters N ──┐ data D ──┼──► training budget (Chinchilla balances these two) compute C ──┘ test-time tokens ──► inference budget (the new axis: think longer)
Scaling laws did not make training cheap. They made it a purchase you can price before buying it — and, once priced, spendable in the way your traffic, not just your loss curve, demands.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Loss is predictable across orders of magnitude — budgets can be planned, not gambled.
- Chinchilla-style balance extracts maximum capability per dollar of training compute.
- Overtrained small models slash serving cost and latency with modest loss sacrifice.
Trade-offs & Constraints
- Fitted laws are architecture/mixture-specific; extrapolating others' constants misleads.
- Low loss ≠ capability on emergent tasks (multi-step reasoning may not transfer from probes).
- High-quality unique data is finite; data-constrained regimes distort the classic laws.
Llama 3 trained comparatively small models (8B, 70B) on ~15T tokens — orders of magnitude beyond Chinchilla-optimal — explicitly trading training loss for drastically cheaper inference at hundreds-of-millions-of-users scale, a decision the team described as "beyond Chinchilla scaling".
Staff+ Engineering Takeaways
- LM loss follows smooth power laws in parameters N, data D, and compute C — training runs become extrapolatable experiments.
- Chinchilla (2022): compute-optimal means N and D scale together, about 20 tokens per parameter; earlier fits over-weighted parameters.
- 70B trained on 1.4T tokens beat 280B trained on 300B tokens at equal compute.
- Serving cost dominates lifetime cost, so production models are deliberately overtrained (small + huge data).
- Test-time compute (reasoning tokens at inference) is the newest scaling axis beyond N and D.
Topic Knowledge Check
Exercise 1 of 2 • Test your architectural comprehension.
Under the Chinchilla compute-optimal rule with fixed training compute C, loss is minimized when:
How clear and actionable was this distributed systems breakdown?