TOPIC #141Advanced 14 min read

LLM Pretraining: Next-Token Prediction at Internet Scale

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

How base language models are built: one self-supervised game — guess the next token — replayed on trillions of tokens of curated web text. This topic covers the objective, the tokenizer and corpus engineering that bound what a fixed budget can learn, the training dynamics of thousand-GPU runs, and why the model that comes out still cannot follow instructions.

From Raw Text to Base Model 🔍

Pretraining is a data pipeline wrapped around one simple objective: predict the next token. The model that emerges can continue text but has never been trained to be helpful.

From Raw Text to Base Model 🔍
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: You Cannot Label the Internet

Remember how classic machine learning works: someone pays annotators to label data. Cats in pictures. Sentiment in reviews.

Now imagine you want a model that knows everything about language: grammar, history, law, code, jokes, physics.

Insight

How many annotators would that take?

Infinite. And infinitely expensive. There is no labeled dataset for "all human knowledge in text" — and any you tried to build would be a rounding error next to what actually exists.

But here is the observation that unlocked the entire LLM era:

Insight

The text already contains its own answer key.

Take any sentence. Hide the next word. The sentence itself tells you what it was:

"The capital of France is ___" → "Paris"

No human labeled anything. The label was simply the token that happens to come next. Every book, every forum post, every line of open-source code is simultaneously the question and the answer.

That single trick — using tomorrow's word as today's label — is self-supervised pretraining, and it is the only reason models trained on trillions of tokens exist at all.

02.The Idea in Plain Words: Compress the Internet by Predicting the Next Token

Pretraining optimizes a single self-supervised loss: given a prefix of tokens, maximize the probability of the next token. For a sequence (x_1, ..., x_T), the model parameterized by \theta minimizes the average negative log-likelihood:

L(\theta) = -\frac{1}{T} \sum_{t=1}^{T} log p_\theta(x_t \mid x_1, ..., x_{t-1})

In plain words, line by line:

  • Read a chunk of text left to right (the causal rule from Concept 138).
  • At every position, the model outputs a probability over its whole vocabulary for "what comes next."
  • Reward it by how much probability it put on the word that actually came next — the log p term.
  • Average over the sequence, flip the sign (we minimize negative likelihood), and train with gradient descent.

No human labels are required — the text itself is the supervision signal, which is what makes internet-scale training feasible.

And here is the surprise that defines this era. To minimize this humble prediction error, the model is forced to internalize:

  • grammar (wrong agreement raises the loss),
  • facts ("Paris" really does follow "capital of France is"),
  • world structure, code syntax, style, even crude reasoning patterns — because all of them reduce prediction error on real text.

This is the core bet behind the modern LLM era, articulated by Rich Sutton as "The Bitter Lesson": general methods that leverage computation and data ultimately beat hand-engineered linguistic structure. A decoder-only Transformer (the GPT lineage, Concept 137) trained on trillions of tokens becomes a universal lossless compressor of human text — and compression, as Ilya Sutskever puts it, is closely related to intelligence: to predict a file well, you must understand it.

03.A Simple Worked Example: One Sentence, Four Exercises

Take the tiny sequence of 5 tokens: I / love / hot / coffee (pretend single-word tokens).

From no labels, you get 4 training exercises for free — just by sliding a window:

code
 prefix seen        exercise               truth
 [I]                next = ?               love
 [I, love]          next = ?               hot
 [I, love, hot]     next = ?               coffee

(4 exercises from 5 tokens — one per boundary. In real training, the whole sequence is processed in a single parallel pass, all positions at once — Concept 138.)

Suppose a half-trained model says:

code
 after [I, love]:        tea 0.4 | coffee 0.3 | pizza 0.1 | ...
 after [I, love, hot]:   water 0.5 | coffee 0.2 | ...

Loss = average −log P(truth):

−log 0.3 = 1.20 and −log 0.2 = 1.61 → mean ≈ 1.4 nats → perplexity ≈ e^1.4 ≈ 4.1 (Concept 140: the pretraining loss and the eval metric are literally the same number, exponentiated).

Training nudges probability from tea/water toward coffee. Now scale that one line of arithmetic by:

  • 500 bytes per average web page → millions of pages,
  • trillions of tokens of text,
  • a few billion parameters to hold what it learns.

That is pretraining. Nothing else.

04.Visual Intuition: The Pipeline Around One Small Loop

The objective is tiny; the machinery around it is the topic. The whole operation:

code
 RAW TEXT          trillions of tokens from web/books/code
     |
     v
 TOKENIZER         chop text into 32k-256k subword pieces
     |              (the alphabet the model will think in)
     v
 CURATION          dedup, filter, scrub PII, mix domains
     |              (garbage in = garbage baked in forever)
     v
 +-----------------------------+
 |  THE LOOP (the small part)  |
 |  predict next token         |
 |  compute loss               |
 |  gradient update (AdamW)    |  x 10^5-10^6 steps
 +-----------------------------+   on 1000s of GPUs
     |
     v
 BASE MODEL        completes text, not instructions
     |
     v
 POST-TRAINING     SFT / RLHF -> assistant (later topics)

Two things to see in this picture:

  • Everything before the loop decides what the loop can learn. A bad tokenizer or a polluted corpus cannot be fixed by more steps.
  • The loop itself is embarrassingly simple — that is the point. The complexity lives in scale, not in the math.

05.The Analogy: The Library Intern Nobody Ever Taught to Answer

Carry one analogy through the rest of this topic: pretraining produces a library intern.

  • The intern spends years reading every shelf — novels, forums, manuals, code repos (trillions of tokens).
  • Her only exercise: at every page turn, guess the next word, over and over. No teacher. No exams about understanding. Just reading, forever.
  • She absorbs enormous knowledge — grammar, facts, style, argument patterns — because all of it helps her guess faster.
  • Then one day you walk in and ask: "What is the capital of France?"

And she does something maddening. She doesn't answer. She continues:

"...What is the capital of France?" asked the quizmaster, sweating under the lights.

Because in all her reading, a question like yours was almost always the start of a quiz transcript, not a request for help. Nobody ever taught her to be an assistant.

That is the two-part lesson of this topic:

  1. Pretraining installs knowledge (Sections 6–7: how, and at what cost).
  2. Post-training installs behavior (Section 8; Topics 144–149) — the day the intern gets coached to actually help.

06.Before the Loop: Tokenizers and Corpus Engineering

Before any GPU runs, two unglamorous decisions dominate final quality:

  • Tokenizer. Models train on subword tokens (BPE or SentencePiece), typically 32k–256k vocabulary. A poorly trained tokenizer inflates token counts for multilingual or code text — the same article takes more positions, so effectively the model sees less knowledge per FLOP. Byte-level tokenizers (GPT-family) never fail on unknown characters; Llama 3 moved to a 128k vocabulary partly to cut tokenization overhead for non-English and code.
  • Corpus curation. Raw Common Crawl is repetitive and polluted. Pipelines apply fuzzy and exact deduplication (MinHash), quality classifiers, PII scrubbing, malware/CSAM filtering, and deliberate domain mixing ratios (web, code, math, books). Modern Llama-3-class recipes report trillions of raw tokens compressed to roughly 15T high-quality training tokens.

Why "unglamorous" is the wrong word — do the arithmetic on the intern:

  • Budget = fixed number of pages she can read (compute).
  • A bloated tokenizer means each book costs more pages. Fewer books, same money.
  • Duplicates mean re-reading the same chapter. Memorizing a duplicate teaches nothing new.

Data quality is not a footnote: deduplication interacts directly with scaling laws (Topic 142) because repeated data has diminishing marginal value.

07.Running the Loop at Scale: Training Dynamics

A pretraining run is a months-long, multi-thousand-GPU job whose failure modes are statistical, not symbolic:

  • Optimization: AdamW with a warmup-then-decay (usually cosine) learning-rate schedule; batch sizes from ~1M tokens (GPT-3) to tens of millions of tokens per step in 2020s frontier runs.
  • Loss spikes: sudden training-loss spikes (bad data batches or LR instabilities) require checkpoint rollback and sometimes data excision — a major operational concern. (Think of it as the intern re-reading a corrupted chapter, then quarantining it forever.)
  • Schedule length matters more than final model size: training far "beyond Chinchilla" (many tokens per parameter) yields models that are smaller, cheaper to serve, and often better than oversized models trained on less data (see Topic 142).
  • Observability: practitioners watch not just loss curves but downstream evals (HellaSwag, MMLU, HumanEval) every few thousand steps, because validation loss alone hides regressions.

The entire pretraining objective, in four lines — the smallness is the lesson:

python— The entire pretraining objective in pseudocode
# logits: (batch, seq_len, vocab) from the Transformer
logits = model(input_ids)
targets = roll_left(input_ids)          # each token predicts the NEXT token
loss = cross_entropy(logits.view(-1, V),
                       targets.view(-1),
                       ignore_index=PAD)
loss.backward(); optimizer.step()       # repeat ~10^18 FLOPs away

08.Why It Still Fails You: What a Base Model Cannot Do

The output of pretraining is a base model: it samples continuations of a prompt, with no notion of "assistant".

Ask it "What is the capital of France?" and it may answer with a quiz-show transcript instead of "Paris", because that is statistically similar to its training text — the intern, doing exactly what her one exercise taught.

Base models also have no safety training: they will happily continue harmful requests, because harmful text exists in the corpus and the loss never distinguished it.

The split to memorize, because it organizes the entire rest of this phase:

pretraining installs knowledge → post-training installs behavior

Turning a base model into a usable assistant requires the post-training stack covered later in this phase: instruction tuning / SFT (Topic 144), then preference optimization via RLHF, DPO, or Constitutional AI (Topics 145–149).

In production practice, the decision follows from the intern metaphor's price tag:

  • Pretraining is the one-time, hundreds-of-millions-of-dollars read (frontier runs exceed $100M).
  • Fine-tuning or prompting an existing model is 100–1000x cheaper — which is why Section 6's data choices and Section 7's run health are other people's problems for almost every team building on LLMs, and a genuine strategic choice only for exotic domains (see trade-offs below).

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Zero labeling cost: supervision comes free from raw text.
  • One pretraining run yields a general-purpose model adaptable to every downstream task.
  • Capabilities scale predictably with compute and data (scaling laws).

Trade-offs & Constraints

  • Enormous capital cost (frontier runs exceed $100M) and long iteration cycles.
  • Base models are not helpful by default; a full post-training pipeline is still required.
  • Data curation mistakes are baked in for the life of the model.
Production Implementation in Big Tech
Meta (Llama 3)• Open-Weights Pretraining at 15T Tokens

Llama 3 405B was pretrained on over 15 trillion curated tokens with a 128k-tokenizer, multi-stage data filtering and dedup, and careful domain mixing of web, code, and math — then post-trained with SFT, rejection sampling, and DPO to produce the instruct variants.

Staff+ Engineering Takeaways

  • Pretraining = self-supervised next-token prediction; the text itself supplies every label.
  • Tokenizer and data curation choices materially bound what a fixed compute budget can learn.
  • Training long (high tokens-per-parameter) beats training big, and interacts directly with scaling laws.
  • Base models complete text; instruction following and safety come only from post-training.
  • The split "pretraining installs knowledge, post-training installs behavior" organizes the entire LLM pipeline.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Why does LLM pretraining require no human-annotated labels?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?