TOPIC #142Advanced 14 min read

Scaling Laws and Chinchilla: How to Spend a Training Budget

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Training a model costs a fixed pile of compute — how do you split it between parameters and data? Scaling laws say loss falls in smooth, predictable power laws, so you can measure small and forecast the $100M run. Chinchilla added the correction: spend on both equally, roughly 20 tokens per parameter — and production models then overtrain small ones on purpose because serving, not training, dominates lifetime cost.

Spending One Compute Budget Three Ways ⚖️

Chinchilla showed loss is minimized when parameters and tokens grow together. Production teams then deliberately overtrain small models because serving — not training — dominates lifetime cost.

Spending One Compute Budget Three Ways ⚖️
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A $100M Run Is a Budget-Splitting Decision

Picture the moment a lab commits to a frontier pretraining run (Topic 141).

They have a fixed pile of compute — call it one budget. It can be spent in wildly different ways:

  • a huge model (hundreds of billions of parameters) fed a few hundred billion tokens,
  • or a modest model fed trillions of tokens,
  • or something in between.
Insight

Which split gives the smartest final model?

Before 2020, that was answered by taste, papers copied between labs, and a lot of anxiety. Training was alchemy: nobody could tell you, before spending, whether the run would be good.

Then two results in a row turned the question into arithmetic.

  1. Kaplan et al. (2020): loss falls in a smooth, predictable curve as you grow size or data — so you can measure the curve cheaply and forecast the expensive run.
  2. Chinchilla (Hoffmann et al., 2022): the forecast exposes how to split the budget — and it said the giants of 2020 (GPT-3 175B, Gopher 280B) had split it badly.

This topic is those two results, the 20-token-per-parameter rule, and why production models deliberately break it.

02.The Idea in Plain Words: Loss Follows a Power Law

A scaling law is simply

Insight

A smooth formula that says: multiply your compute, and loss drops by a predictable amount — every time, across orders of magnitude.

Kaplan et al. (2020) discovered that Transformer language-modeling loss decreases smoothly and predictably as any of three factors grows:

  • N — model size (parameters),
  • D — dataset size (tokens),
  • C — compute.

Fitting all three jointly with GPT-3-era scaling:

L(N, D, C) ≈ 3.7 + \frac{1.5}{N^{0.23}} + \frac{1.5}{D^{0.24}} \quad (power-law terms, N in the billions)

Unpack the plain consequences of that formula:

  • Everything is a curve with no surprises. No magic size, no threshold where "understanding" appears — loss glides down as N or D rise (power laws: multiply the input by 10, lose a fixed amount of loss).
  • You can extrapolate. Train tiny probe models on ~1% of the budget, fit the curve, and confidently predict the final loss of a $100M run before spending it.
  • Budgets become plans. This turned frontier training from alchemy into capacity planning — and explains why labs raced straight up the curve to GPT-3 (175B) and Gopher (280B), assuming bigger N was the answer.

That last assumption is exactly what Section 4 takes apart.

03.A Simple Worked Example: Predicting a Run Before Buying It

Suppose your probes fit a simplified law: every 10x of compute drops loss by about 0.5 nats, in a straight line on a log scale.

code
 compute (log scale)   predicted loss
 10^18  FLOPs              5.0      <- probe model (cheap, hours)
 10^19                      4.5      <- probe
 10^20                      4.0      <- probe
 10^22  (the real run)      3.0      <- FORECAST before spending $100M

You spent pennies to pre-purchase confidence in a number. That is the whole economic magic of power laws.

Now the budget-splitting arithmetic, Chinchilla-style. Two ways to spend roughly the same training compute (compute ≈ proportional to N × D, parameters times tokens seen):

code
              parameters    tokens        N x D (relative work)
 Gopher:      280B          300B          280 x 300 = 84,000
 Chinchilla:  70B           1.4T          70 x 1400 = 98,000
             (4x smaller)  (~4.7x more data)   ~same total compute

Same pile of money. Different split. Chinchilla — the 4x-smaller model trained on 5x more data — beat Gopher on 67 of 70 evals.

The lesson in one sentence: when the budget is fixed, parameters and data must grow together, not parameters alone.

04.The Chinchilla Correction: Balance N and D

In 2022, Hoffmann et al. at DeepMind ("Training Compute-Optimal Large Language Models", arXiv 2203.15556) re-derived the laws with three different estimation methods (FLOP-matched, step-wise, parametric — the result was triangulated, not a lucky fit) and found Kaplan had a critical flaw: the 2020 fits could not separate the N and D exponents, biasing labs toward over-sized models trained on too little data.

The corrected result: for a fixed training compute budget C, loss is minimized when model size and data size grow at the same rate — roughly N \propto C^{0.5} and D \propto C^{0.5}, i.e., about 20 training tokens per parameter.

In plain words: if you can afford to double one, you should double the other too. A model of 70B parameters "deserves" ~1.4T tokens; a 175B model needs ~3.5T.

GPT-3 (175B on 300B tokens ≈ 1.7 tokens/parameter) and Gopher (280B on 300B ≈ 1 token/param) were, by this standard, roughly 4x over-parameterized — giants starved of reading, when a smaller brain with five times the library wins the same fight for the same money.

The empirical punchline was Section 3's matchup: Chinchilla (70B parameters, 1.4T tokens) beat Gopher (280B parameters, 300B tokens) on 67/70 evals despite using the same training compute.

05.Visual Intuition: One Budget, Three Splits

A fixed compute budget is a rectangle: width = parameters N, height = tokens D. Same area = same compute. Chinchilla says the square-ish shape is best for loss.

code
  D (tokens)
   1.4T |      +-----+
        |      | CHIN|      <- 70B x 1.4T: balanced, lowest loss
    300B| +----+     +  +--------+
        | |GOPHER|   |  |OVERTRND|  <- e.g. 8B x 15T: slightly
        | +------+   |  +--------+     worse loss, cheap to serve
        +------------+-----------------+---------► N (params)
             70B      280B        8B
  • Tall-thin Gopher (huge N, few D): pays for neurons it never fed.
  • Square-ish Chinchilla: loss-optimal at that compute.
  • Flat-wide overtrained (tiny N, oceans of D): not loss-optimal — and that is the point; see next section.

The dependency chain to memorize:

fixed C → choose split of N and D → loss minimized near N ≈ D balance → but serving cost depends on N alone → shift the split on purpose

06.The Analogy: Educating One Employee on a Fixed Scholarship

Carry one analogy: you are funding the education of a single employee, then paying their salary forever.

  • Parameters N = the size of the brain you buy (bigger brain = higher salary per word they speak).
  • Tokens D = the number of books you make them read.
  • Compute C = the one-time scholarship that must be split between the two.

Kaplan-era labs spent nearly the whole scholarship on a giant brain and handed it three books. The genius, poorly read, bills per hour — and its ideas are shaky anyway.

Chinchilla noticed the square-root balance: the giant brain and the well-read midbrain came out at the same price, and the well-read one outperformed on 67 of 70 exams.

Then production adds the final twist: after hiring, you pay the employee per word spoken (inference), and they will speak trillions of words serving users. Suddenly the smart money hires the smaller, ferociously well-read person — cheaper per word, nearly as smart, because extra reading was almost free.

That last move is Section 7: deliberate overtraining.

07.Training Compute vs. Inference Compute: Overtraining Is Rational

Chinchilla optimizes loss at training time only. But most frontier models generate trillions of tokens over their serving life, and each inference token costs about 2x-forward-pass-FLOPs per parameter — salary per word, in the analogy.

If expected served tokens are large, the lifetime-optimal model is therefore smaller and more data-fed than Chinchilla-optimal — deliberately "overtrained" (more data than Chinchilla says the parameters deserve).

This is the regime all production 2024–2026 models live in:

  • Llama 3 8B was trained on ~15T tokens ≈ 1,900 tokens per parameter — two orders of magnitude past Chinchilla's 20 — because it serves hundreds of millions of users, each call paying per-parameter salary (Topic 141's recipe).
  • GPT-4o / Claude / Gemini-class deployments favor small fast models with huge context for the same reason.

But what if the library runs dry? High-quality unique text is finite. The Muennighoff et al. work on data-constrained scaling formalizes the limit: repeating data across 4–16 epochs retains most of its value, then decays — re-read books cost little but teach less each pass.

Rule of thumb for system design interviews:

Insight

Training budget chooses what you can afford to learn; traffic volume chooses how small you should have made it.

08.In Practice: Using Scaling Laws in Your Own Runs

A lab's scaling workflow today, step by step:

  1. Run probes: train a ladder of small models (10M–1B) at several N/D splits with shared tokenizer and data order — keep everything else fixed so the curve measures what you think it measures.
  2. Fit the law: estimate exponents for your architecture and data mixture — exponents shift across architectures, so never copy constants blindly (Kaplan's numbers already mislead once).
  3. Extrapolate and validate: predict frontier loss, then sanity-check with eval-transfer — small-model evals predict large-model evals only for smooth, non-emergent metrics (see Topic 143; an emergent capability can stay flat on the probe ladder and appear only at scale).
  4. Watch the frontier: loss vs. FLOPs has been closing a power-law gradient for years; the 2024–2026 era added reasoning-time compute (test-time scaling) as a fourth axis, where extra tokens generated at inference buy capability once bought only at training (Concept 137's o-series).

The three-axis map to leave with:

code
 parameters N ──┐
 data D       ──┼──► training budget (Chinchilla balances these two)
 compute C     ──┘
 test-time tokens ──► inference budget (the new axis: think longer)

Scaling laws did not make training cheap. They made it a purchase you can price before buying it — and, once priced, spendable in the way your traffic, not just your loss curve, demands.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Loss is predictable across orders of magnitude — budgets can be planned, not gambled.
  • Chinchilla-style balance extracts maximum capability per dollar of training compute.
  • Overtrained small models slash serving cost and latency with modest loss sacrifice.

Trade-offs & Constraints

  • Fitted laws are architecture/mixture-specific; extrapolating others' constants misleads.
  • Low loss ≠ capability on emergent tasks (multi-step reasoning may not transfer from probes).
  • High-quality unique data is finite; data-constrained regimes distort the classic laws.
Production Implementation in Big Tech
Meta — Llama 3• Choosing the 8B/70B Point on the Curve

Llama 3 trained comparatively small models (8B, 70B) on ~15T tokens — orders of magnitude beyond Chinchilla-optimal — explicitly trading training loss for drastically cheaper inference at hundreds-of-millions-of-users scale, a decision the team described as "beyond Chinchilla scaling".

Staff+ Engineering Takeaways

  • LM loss follows smooth power laws in parameters N, data D, and compute C — training runs become extrapolatable experiments.
  • Chinchilla (2022): compute-optimal means N and D scale together, about 20 tokens per parameter; earlier fits over-weighted parameters.
  • 70B trained on 1.4T tokens beat 280B trained on 300B tokens at equal compute.
  • Serving cost dominates lifetime cost, so production models are deliberately overtrained (small + huge data).
  • Test-time compute (reasoning tokens at inference) is the newest scaling axis beyond N and D.

Topic Knowledge Check

Exercise 1 of 2 • Test your architectural comprehension.

Exercise 1 of 20 answered
1

Under the Chinchilla compute-optimal rule with fixed training compute C, loss is minimized when:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum