TOPIC #175Advanced 13 min read

Full Fine-tuning

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Full fine-tuning continues training on every weight of a pretrained model until it behaves like a specialist. That buys the highest quality ceiling — and costs ~16 bytes of memory per parameter, which is why FSDP/ZeRO sharding, BF16, and activation checkpointing are part of the recipe.

Full Fine-tuning Memory Anatomy 💾

Full fine-tuning must hold parameters, gradients, and optimizer states for every weight. The ~16 bytes/param rule is why multi-GPU sharding is mandatory beyond small models.

Full Fine-tuning Memory Anatomy 💾
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: When Prompting Hits a Ceiling

You have a pretrained model — one that read most of the internet and learned general language. You want it to behave like a specialist: a medical report writer, a legal-drafting assistant, a model that speaks your company's exact internal format.

Prompting gets you shockingly far (topics 150–160). But prompts have ceilings:

  • The knowledge and style must already be in the weights; a prompt can only coax out what exists.
  • A long system prompt describing your format is re-read on every single call — tokens, latency, and the model drifting anyway.
  • Deep changes — a new reasoning style, a tool-calling protocol, a genuinely different tone — are fought by everything the base model was trained to be.

So the question becomes

Insight

What if we don't coax the model — we retrain it on our own examples until the new behavior lives in its weights?

That is fine-tuning: the model already did the multi-million-dollar job of learning language; we continue its training on a much smaller dataset of our examples, so it keeps the general capability and bends toward ours. (Training itself — loss, gradients, updating weights by gradient descent — was topics 45–49; in one sentence: show the model examples, measure how wrong it is, nudge every weight a little in the direction that made it less wrong.)

Full fine-tuning (FFT) is the plain, maximal version of that idea.

02.The Idea in Plain Words: Every Knob Moves

Full fine-tuning is simply

Insight

Continue gradient descent on ALL parameters of the pretrained model, using a supervised dataset of instructions, responses, or domain text.

Read the pieces:

  • ALL parameters. Every weight in every layer is trainable. Nothing is frozen, no modules are added — the pretrained weights are directly overwritten toward the new distribution. (Contrast: PEFT/LoRA, topics 176–177, freeze the base and train a tiny slice.)
  • Supervised dataset. You need paired examples: input → the output you want. "Here is an ER intake note → here is the assessment-and-plan paragraph a specialist would write." The stage is often called SFT (supervised fine-tuning).
  • Continue, not restart. The model starts already near a good solution, so we take small steps — a learning rate roughly 10–100x smaller than pretraining used.

Why anyone pays for this: FFT is the quality ceiling of adaptation. It is the method Meta used for the Llama 2/3 chat models' supervised fine-tuning stage, and it is the default when a provider builds a domain-specialized model (medical, legal, code) intended to replace the base model rather than coexist with it.

And the two behavioral costs to know before paying:

  • Catastrophic forgetting. Aggressive updates erase general capabilities — teach 10k medical examples hard enough and the model starts mangling plain conversation. Mitigations: mix 10–30% general instruction data into every batch, keep the learning rate low, and early-stop against a held-out general benchmark, not just your task metric.
  • Overfitting. Below roughly 1k–5k examples, billions of free parameters memorize surface patterns instead of learning the pattern behind them; at that data scale PEFT or in-context learning usually generalizes better.

03.A Simple Worked Example: The 16-Byte-per-Param Bill

The number that decides whether FFT is feasible is pure arithmetic. With mixed-precision AdamW (BF16 compute, FP32 optimizer math), each of the P parameters costs:

code
BF16 weight            2 bytes   (the live copy used in forward/backward)
BF16 gradient          2 bytes   (∂L/∂θ for this weight)
FP32 master weight     4 bytes   (high-precision copy the optimizer edits)
FP32 momentum          4 bytes   (running average of gradients)
FP32 variance          4 bytes   (running average of squared gradients)
─────────────────────
                      16 bytes   per parameter,  before activations!

(One plain sentence on AdamW: the standard optimizer that gives each weight its own adaptive step size, remembering recent gradient size and direction — hence the two extra 4-byte registers per weight.)

Now plug in real models:

  • 7B model: 7 × 10⁹ × 16 bytes ≈ 112 GB of model state. One 80 GB A100? Doesn't fit. → ZeRO-2/FSDP sharding across 2–4 GPUs.
  • 70B model: 70 × 10⁹ × 16 ≈ 1.1 TB. → realistically 16–32 H100s with ZeRO-3 or FSDP, plus activation checkpointing, or a tensor/pipeline-parallel Megatron-style setup.

Plus activations — the intermediate values saved during the forward pass so the backward pass can compute gradients — which grow with batch size × sequence length × layers, and can dwarf the 16 bytes on long contexts.

Hold the contrast for interviews (topic 176 in one sentence: train only ~0.1–1% of the parameters): full fine-tuning pays AdamW's 12 optimizer bytes for every weight; PEFT pays it only for its trainable sliver. That single asymmetry is why the whole PEFT industry exists.

04.Visual Intuition: One Backpack Strap Per Knob

Picture the model as a vast piano — one key per parameter — and training as moving every key to retune it. The cost is not the keys themselves; it is the backpack each moving key must wear:

code
   per trainable weight:
   ┌──────────────────────────────┐
   │ weight (2B)   gradient (2B)  │   ← the "playing" copies
   │ master (4B)  mom (4B)  var(4B)│  ← the optimizer's memory of
   └──────────────────────────────┘     *how* this key should move

   7,000,000,000 keys × that backpack = 112 GB
   ...even though only 14 GB of it is the piano itself.

So the piano fits on one GPU — the backpacks don't. The infrastructure story of this topic is three moves on exactly this picture:

code
 1 shard the backpacks   ZeRO/FSDP: GPU₁ holds optimizer
    across GPUs           state for keys 1–2B, GPU₂ for 2–4B...
 2 shrink the backpacks   8-bit Adam / Adafactor: ~6 B/param
 3 don't save the sheet   activation checkpointing: recompute
    music                forward intermediates during backward

(Full sharding detail: ZeRO stage 2 splits optimizer state + gradients; stage 3 also splits the parameters themselves — which trades a lot of GPU-to-GPU communication for the memory back.)

05.The Analogy: Retraining the Whole Orchestra

Carry this picture through topics 175–177: a world-class orchestra you want to play your repertoire.

  • The pretrained model is the orchestra as it stands — superb sight-readers, trained on every genre.
  • Prompting is handing them detailed notes before the concert. Works surprisingly well, but the notes are re-read every performance and only steer what the players can already do.
  • Full fine-tuning is canceling a season of concerts and putting every single musician through chamber-music school. Every violinist's muscle memory changes. When they come back, they are a chamber orchestra — the deepest possible transformation, the highest ceiling.
  • And the two costs are exactly musical:
    • Catastrophic forgetting: rehearsing only your repertoire makes them slowly lose the symphonic repertoire they played for decades. (Mix in general "sight-reading practice" — the 10–30% general data.)
    • Money: a season of paid rehearsals for all 80 musicians, with private tutors for each (the optimizer state per weight), costs far more than the instruments themselves.

Topics 176–177 introduce the alternative: keep the orchestra untouched and hire a small annotate-and-section-coach layer — same concert, 1% of the rehearsal cost. But when you truly need the musicians themselves to change, FFT is the ceiling.

06.The Training Recipe: Hyperparameters That Actually Get Used

FFT hyperparameters are an order of magnitude smaller than from-scratch training, because the model is already near a good loss basin — you're nudging, not carving:

  • Learning rate: 1e-5 to 2e-5 for 7B–13B models, with cosine decay (the LR glides down instead of stopping abruptly) and 3–10% warmup (ease in, so the first big gradients don't wreck the basin). Larger models use the lower end, or layer-wise LR decay — deeper layers move less than the final ones.
  • Epochs (passes over your data): 1–3 for instruction data. More epochs sharply increase forgetting and memorization — the curve that taught everyone "stop early."
  • Batch: effective batch 128–512 sequences. Gradient accumulation substitutes for missing GPU memory: run 32 tiny micro-batches, sum their gradients, then take one optimizer step — same math, fraction of the memory.
  • Precision: BF16 compute (or FP16 with loss scaling); FP32 master weights inside the optimizer. BF16 keeps the dynamic range of FP32 at half the bytes — the reason 16-bit training is stable at all.
  • Data packing: concatenate short examples to fill the sequence window (no wasted padding tokens), and use FlashAttention/SDPA so long-context activations stay tractable.

The code below is the canonical shape: Hugging Face TRL, BF16, gradient accumulation 4×32 = effective batch 128, gradient checkpointing on, cosine schedule, packing.

python— SFT with Hugging Face TRL in BF16 + gradient accumulation
from transformers import TrainingArguments
from trl import SFTTrainer

args = TrainingArguments(
    per_device_train_batch_size=4,
    gradient_accumulation_steps=32,   # effective batch 128
    learning_rate=1.5e-5,
    num_train_epochs=2,
    bf16=True,
    optim="adamw_torch_fused",
    gradient_checkpointing=True,
    warmup_ratio=0.03,
    lr_scheduler_type="cosine",
)
trainer = SFTTrainer(model=model, args=args,
                     train_dataset=dataset,
                     max_seq_length=4096,
                     packing=True)

07.In Practice: Infrastructure, and When Full Fine-tuning Wins

The enabling stack, in one plain sentence each:

  • DeepSpeed ZeRO / PyTorch FSDP: shard optimizer state, gradients, and (ZeRO-3) parameters across GPUs, so 8 cards hold what 1 cannot. Mandatory beyond small models.
  • Activation (gradient) checkpointing: don't store forward intermediates; recompute them during the backward pass — trade ~30% extra compute for a large activation-memory win.
  • Tensor / sequence / pipeline parallelism (Megatron-style): when even sharded state exceeds one machine, split each layer's math and layers themselves across devices.
  • BF16 + FlashAttention: the compute-side defaults that make long sequences affordable.

Now the decision rule. Choose FFT when at least two of these hold:

  1. You have >5k–10k high-quality examples (below that, the parameter count memorizes).
  2. You need deep behavioral change — a new output format, a domain reasoning style, a tool-calling protocol — not just topical answers.
  3. You will serve one dedicated model (so artifact size is irrelevant — no per-tenant swapping).
  4. You have the multi-GPU budget the 16-bytes-per-param bill demands.

FFT also composes with post-training pipelines: RLHF/DPO stages on frontier models are conventionally built on top of a fully fine-tuned SFT checkpoint rather than a LoRA merge — the "official" path from base model to chat product still runs through this topic. (The Llama 2 paper is the honest case study: full fine-tuning at 7B/13B, but they switched to LoRA at 34B/70B because FFT cost more for comparable quality — a preview of topic 176.)

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Highest achievable quality ceiling — every weight can adapt to the task distribution.
  • Deployed artifact is a single plain model: no adapter plumbing, fully compatible with any inference server or quantizer.
  • Best fit for deep style/format/protocol changes and for downstream RLHF or DPO stages.

Trade-offs & Constraints

  • ~16 bytes/param of optimizer memory forces multi-GPU FSDP/ZeRO setups beyond small models.
  • Catastrophic forgetting of general capabilities without careful data mixing and early stopping.
  • Overkills small datasets: below ~5k examples PEFT or prompting often generalize better.
  • Every tenant/domain needs its own full checkpoint — storage and serving cost multiply per variant.
Production Implementation in Big Tech
Meta (Llama 2 / Llama 3 post-training)• Supervised fine-tuning of chat models

Meta fully fine-tuned 7B and 13B Llama 2 chat models on hundreds of thousands of instruction-following examples; for the 34B/70B tiers the paper explicitly switched to LoRA because full fine-tuning was more expensive with comparable quality. A production echo of the memory economics discussed above.

Staff+ Engineering Takeaways

  • Full fine-tuning updates every weight and remains the quality ceiling for deep behavioral change.
  • AdamW model state costs ~16 bytes per parameter plus activations — 7B needs multi-GPU sharding, 70B needs a small cluster.
  • Recipes use 1e-5-range learning rates, 1-3 epochs, and general-data mixing to contain catastrophic forgetting.
  • FSDP/ZeRO sharding, BF16 compute, activation checkpointing, and FlashAttention are the standard enabling infrastructure.
  • Prefer PEFT when data is small, variants are numerous, or GPUs are scarce; prefer FFT when you ship one dedicated model.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Roughly how much memory does AdamW model state require per parameter during BF16 mixed-precision full fine-tuning?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?