TOPIC #144Advanced 12 min read

Instruction Tuning / SFT: Teaching a Base Model to Be an Assistant

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A base model can finish any text, but it will not answer your question — it has never been shown what an "answer" looks like. Supervised fine-tuning (SFT) fixes that with thousands of (instruction, ideal-response) examples, training only on the response part. This topic covers what actually changes, why a small high-quality dataset (LIMA) rivals a huge one, and the limits imitation leaves behind.

SFT in the Alignment Pipeline 🎯

SFT reuses the pretraining loss but applies it only to the assistant turns, so the model learns "given this instruction in this chat format, produce this response".

SFT in the Alignment Pipeline 🎯
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A Genius That Only Finishes Sentences

You train a model on half the internet.

It becomes brilliant at one thing:

Insight

Predict the next token.

So you ask it:

User: What is the capital of France?

And it replies:

code
User: What is the population?
User: Who are you?

Not wrong, exactly. It is completing the pattern of a Q&A page, not answering you.

The knowledge is in there. The manners are not:

  • answer questions instead of generating more questions,
  • stop when the answer is done,
  • follow the format you asked for.

Instruction tuning / SFT (Supervised Fine-Tuning) is the training stage that installs those manners.

02.The Idea in Plain Words: Show, Don't Just Read

SFT is:

Insight

Fine-tuning on datasets of (instruction, ideal response) pairs — same next-token objective as pretraining, with two changes.

Change 1 — the text is dialogue-shaped.

Prompts are rendered with a chat template: special tokens mark the user role, the assistant role, and where each turn begins and ends. The model sees conversations, formatted exactly how conversations will arrive at serving time.

Change 2 — the loss is computed only on assistant tokens.

Prompt tokens are masked — the model is graded only on the reply.

Unpack the pieces:

  • "Next-token objective" = the same cross-entropy training you already know from pretraining (predict the following token; get punished for surprise). Nothing new mathematically.
  • "Masked" = those positions are excluded from the loss entirely. The model is not trained to reproduce the user's question — it already "knows" text; it must learn to generate good answers.
  • "Ideal response" = a human-written (or distilled) best-practice answer: the thing you wish the model had said.

And the distribution changes too: instead of raw web text, the model now sees questions, rewrites, code requests, refusals, and their model answers.

The lineage, in plain words:

  • T0 and FLAN (2021–2022): fine-tune a model on hundreds of NLP tasks formatted as instructions → it generalizes to unseen tasks. The surprise that started it all.
  • InstructGPT (2022): its 1.3B SFT+RLHF model beat GPT-3 175B on human preference for instruction following — behavior beats size.
  • FLAN-PaLM 540B ("Scaling Instruction-Finetuned Language Models", 2022): cemented that instruction tuning multiplies few-shot ability across diverse task mixtures.

03.A Tiny Worked Example: What the Model Gets Graded On

One training pair:

code
instruction: "Name a prime number."
response:    "7 is prime."

Rendered with a simple chat template, the token sequence looks like:

code
<|user|> Name a prime number. <|end|> <|assistant|> 7 is prime. <|end|>
|____________ masked (no loss) ____________|^^^^^^^^^^^ loss here ^^^^^^^|
  • On the masked span, the model still reads every token (that is the input).
  • On the graded span, at each position it must predict the next token: after "<|assistant|>" it had better put probability mass on "7", not on "Name a prime number."
  • Loss = average surprise over just the answer tokens.

So SFT literally teaches: when an assistant turn starts, produce this kind of text, then stop.

04.Visual Intuition: One Loss, Two Jobs

code
PRETRAINING (read the whole page)
  text:  [████████████████████████]   ← loss on everything
  learns: language, facts, style — but no "answer mode"

SFT (only grade the reply)
  turn:  user     [░░░░░░░░░░]        ← masked, input only
         assistant[██████]           ← loss here
  learns: WHEN asked X, ANSWER with Y, in chat format

Read the arrows:

pretraining → fills the library (knowledge). SFT → installs the reference desk (how to serve the knowledge).

Same model weights, same gradient machinery — only the data shape and the mask change.

05.The Analogy: The Well-Read Intern Given a Script Manual

Picture a base model as an intern who has read every book in the library — including every question-and-answer page, but never trained to be the answerer.

Ask them "what is the capital of France?" and they might recite more of the encyclopedia entry, because entries is all they have ever produced.

SFT is the first week on the job at the front desk:

  • The manager hands them thousands of sample exchanges: guest asked X, ideal staffer answered Y.
  • The intern is graded only on their lines, never on repeating the guest's question (that is the loss mask).
  • After a few thousand scripted exchanges, the format clicks: greet, answer, stop.

Key insight from LIMA (below): the intern did not need to be re-educated. The books had already been read. The script manual just showed them how to act.

python— The loss mask that defines SFT (pseudo-code)
rendered = chat_template([{"role":"user","content":inst},
                          {"role":"assistant","content":resp}])
logits = model(input_ids)
labels = input_ids.clone()
labels[assistant_span == 0] = -100   # mask prompt/user tokens
loss = cross_entropy(logits, labels) # gradient only on the reply

06.Data: Quality, Diversity, and the LIMA Result

SFT data design follows one empirical law:

Insight

A few thousand exemplar responses teach the format; the breadth of coverage teaches robustness.

  • Sources: human-written (InstructGPT contractors), model-generated with heavy curation (self-instruct, distillation from stronger models), and expert-annotated subsets for code/math.
  • LIMA (2023): fine-tuning Llama-70B on just 52,000 carefully curated, high-quality examples produced responses humans preferred nearly as much as those from models trained on hundreds of thousands. Conclusion: "Less Is More for Alignment" — SFT mostly installs style and format; the knowledge is already in the base model.
  • Coverage still matters: every capability and refusal pattern you never show in SFT data tends to be missing or unstable. A model that has never seen a polite refusal example will improvise one — badly. Modern recipes blend general chat, coding, math, tool use, long-context, and safety examples, then continue with preference optimization (Topics 145–149).

07.What SFT Cannot Do: The Imitation Ceiling

Back to the intern: a script manual teaches them to copy ideal lines. It cannot teach judgment about which of two acceptable lines is better.

SFT imitates demonstrated behavior; it cannot express preferences it never saw graded:

  • No relative judgments: SFT treats every example as equally perfect. "Answer A is better than answer B" — the signal that makes models genuinely helpful, safe, and concise — requires preference training (RLHF/Topic 145, DPO/Topic 148).
  • Sycophancy and verbosity leak in: annotator-written SFT data tends to reward long, agreeable answers; studies through 2024 consistently show SFT amplifies sycophancy because that style dominates the demonstration distribution. The intern copies the manners of the scripts, including the bad ones.
  • Alignment tax / brittleness: aggressive SFT on narrow styles measurably degrades base-model capabilities — which is why labs keep SFT epochs low (often <3) and learning rates small. Repeating one script too many times makes the intern lose their book knowledge.
  • The imitation ceiling: the model can at best match the average quality of its demonstrations; pushing beyond demonstrated quality is precisely what the reward-and-optimization stages are for.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Cheap and stable: hours of compute on a curated dataset transforms a base model into a usable assistant.
  • Installs consistent formatting, tool-call schemas, and domain style that prompting cannot guarantee.
  • Quality-focused small datasets (LIMA-style) rival huge ones, lowering annotation cost.

Trade-offs & Constraints

  • Imitation-only: cannot encode pairwise preferences or trade-offs beyond the data.
  • Amplifies annotator biases — verbosity, sycophancy, refusals in the wrong places.
  • Risk of capability regression (alignment tax) if epochs/LR are too aggressive.
Production Implementation in Big Tech
OpenAI• InstructGPT (2022)

Three-stage recipe: SFT on contractor-written demonstrations of ideal assistant behavior, then a reward model from human preference rankings, then PPO. The 1.3B InstructGPT output was preferred over GPT-3 175B — proof that the SFT+preference stack changes effective behavior more than raw scale.

Staff+ Engineering Takeaways

  • SFT = the pretraining objective on (instruction, response) data, with loss only on assistant tokens.
  • It installs format, style, and instruction-following; knowledge already lives in the base model (LIMA).
  • Data quality and coverage of desired behaviors matter more than raw example count.
  • SFT is imitation: it cannot express preferences, and it inherits annotator biases like verbosity and sycophancy.
  • RLHF/DPO stages follow SFT precisely because demonstrated-average quality is a ceiling.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

In SFT, the training loss is computed on:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?