Parameter-Efficient Fine-tuning (PEFT)
PEFT freezes the pretrained model and trains only 0.05–5% of its parameters, because task adaptation lives in a low-dimensional corner of weight space. That tiny trainable slice turned fine-tuning into a single-GPU job and created the multi-tenant adapter-serving economy of 2024–2026.
PEFT Method Taxonomy 🧩
Parameter-efficient methods either add small trainable modules around a frozen base, select a tiny subset of existing weights to train, or reinitialize task-specific components.
01.The Problem: You Cannot Afford the Orchestra Rehearsal
Recall the last topic in one plain sentence: full fine-tuning (FFT) continues gradient descent on every weight, and with AdamW that costs ~16 bytes per parameter — about 112 GB of model state for a 7B model, ~1.1 TB for 70B, so multi-GPU sharding is mandatory.
Now put that bill in front of a real business:
- You run an API company. You have 400 customers, each wanting the same base model tuned to their voice, their JSON schemas, their domain. FFT says: 400 full checkpoints, 400 cluster jobs, 400 dedicated GPU deployments.
- You are a researcher with one 48 GB GPU and a 7B model. FFT says: doesn't fit, please come back with 4 GPUs.
- Even where FFT fits, it behaves badly on small data: millions of free parameters will happily memorize your 800 examples instead of generalizing.
So the question becomes
Does a model really need to change every one of its billions of weights to learn a new task — or just a small, well-chosen slice?
Parameter-Efficient Fine-Tuning (PEFT) is the family of answers that says: freeze the pretrained base, train a sliver — typically 0.05%–5% of parameters — and let that sliver do the adapting. It became the default mode of LLM adaptation in 2024–2026, and this topic is the map of the family: why it works, its three method families, and the serving economics it created. (The flagship member, LoRA, gets its own topic, 177.)
02.The Idea in Plain Words: Adaptation Is Low-Dimensional
PEFT rests on one empirical finding:
Pretrained networks are over-parameterized for any single task — the change a task needs lives in a tiny corner of weight space.
The evidence goes back further than LLMs. Aghajanyan et al. (2020) measured the intrinsic dimension of fine-tuning: restrict BERT/RoBERTa updates to a random subspace of just a few hundred directions, and fine-tuning still converges to full-quality GLUE results. A model with 100M+ knobs only ever turns a few hundred of them per task.
The 2024–2026 consensus extends this to LLMs: task-specific weight deltas ΔW (the difference between base weights and the fully fine-tuned ones) have small singular-value spectra — meaning ΔW is "thin": it can be well-approximated by a low-rank matrix. That is exactly the assumption LoRA encodes (topic 177 in one sentence: learn ΔW as a product of two narrow matrices).
Given that, the recipe writes itself:
- Freeze the base weights entirely (
requires_grad=False— no gradients, no optimizer state for them). - Train the sliver: either new small modules attached around the base, or a chosen subset of existing weights.
And quality follows: results that in 2022 sat a few points below FFT on classification/seq2seq benchmarks now, on modern instruction datasets, often match or exceed it for style/format adaptation — with ~1000x less optimizer memory (topic 175's 16-bytes-per-param applies to only the trainable slice), single-GPU jobs, and ~40 MB checkpoints.
One more free win: the frozen base is inherent regularization. Fewer trainable knobs means less capacity to memorize your 800 examples — forgetting of general capability is structurally limited because general capability was never touched.
03.A Simple Worked Example: Do the Math on a 7B Model
Take a 7B-parameter model, BF16. Run the two bills side by side.
Full fine-tuning (topic 175):
code16 bytes × 7,000,000,000 params ≈ 112 GB model state → 2–4× 80 GB GPUs minimum checkpoint written: ~14 GB (every weight changed a little)
LoRA at rank 16 on all linear layers (the code in section 6 prints this):
codetrainable params: 19,922,944 (~0.29% of 6.76B) 16 bytes × 19.9M ≈ 319 MB optimizer-side state checkpoint written: ~40–80 MB base still loaded: ~14 GB ← the part PEFT does NOT shrink
So: a cluster job became a laptop job. The ~40–80 MB checkpoint is the second magic number — thousands of customer variants now fit in an S3 bucket, not a model registry.
For a third feel, the smallest of the family: BitFit trains only the bias terms — about 0.08% of parameters — and on small/medium models already recovers most of full fine-tuning's gain. That is the intrinsic-dimensionality claim made concrete: sometimes even biases alone are enough, because the task really only needed a few hundred useful directions.
04.Visual Intuition: The Frozen Machine and the Sticker Panel
Picture the pretrained model as a huge, glass-covered control panel — billions of dials, all welded shut (frozen):
codeFROZEN BASE MODEL (billions of dials, welded) ┌──────────────────────────────────────────────┐ │ ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● │ │ ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● │ │ ● = frozen weight ★ = what each method trains └──────────────────────────────────────────────┘ ▲ ▲ ▲ │ ADDITIVE: │ SELECTIVE: │ RETRAIN: stick a small pick a few rip off the aftermarket dial existing dials output dials, panel ON TOP (★ = biases, mount new (LoRA, adapters, layer subset) task dials IA³, prefix tokens) (heads, embeds)
Three families = three different answers to "where do the trainable knobs come from":
- Additive — bolt on a new, tiny trainable panel next to the frozen one.
- Selective — un-weld a chosen subset of existing dials.
- Retrain — replace whole task-facing components and train only those.
And the panel-vs-sticker picture explains the callout: the giant panel still has to be powered up and read on every forward pass (that's inference memory); PEFT only made the reprogramming cheap.
05.The Analogy: Annotation Booklets for an Unchanged Orchestra
Continuing the orchestra from topic 175: full fine-tuning put every musician through chamber-music school — deep, expensive, and everyone actually changed.
PEFT is a different contract: the orchestra never rehearses a single note of its own. Instead you publish a small annotation booklet — "in bar 12, bow nearer the bridge; in the finale, lean the accents" — and hand it out. The booklet is a few pages (~0.3% of the total sheet music), cheap to author (train), cheap to store, and any concert can be re-created by pairing the same orchestra with a different booklet.
Now see the three families as three ways of authoring that booklet:
- Additive: a separate slip clipped onto the printed score (LoRA branches, adapter modules, prefix tokens) — the original pages untouched.
- Selective: you're allowed to pencil over only the dynamic markings (biases — BitFit), or only the last two pages (top layers).
- Retrain: you throw away the cover and program notes and print new ones (task heads/embeddings).
The economics write themselves — that is section 7: one orchestra on stage (the shared base in GPU memory), four hundred booklets in a drawer, swapped between songs (per-request adapters). And the honest caveat, in the same picture: the booklet steers the players but never retrains them — when a violinist fundamentally cannot play the passage, an annotation can't fix the muscle memory (the quality ceiling of section 8).
06.Under the Hood: The Three Families and Their Methods
Additive methods — insert new trainable modules, all base weights frozen:
- Adapters (Houlsby et al., 2019, arXiv:1902.00751): tiny bottleneck MLPs (down-project, nonlinearity, up-project) inserted in the residual path of each Transformer block. The original result: ~3.6% extra parameters matched full BERT fine-tuning on GLUE. Cost: they sit inside the forward graph, so inference pays for them.
- LoRA (Hu et al., 2021, arXiv:2106.09685): a low-rank update ΔW to the weight matrices themselves — full story in topic 177. The unique gift: ΔW merges into the base weights, so serving latency is unchanged.
- IA³ (arXiv:2205.05638): even cheaper — just three elementwise learned scale vectors that multiply into key/value and feed-forward activations. Tune the gains, not the weights.
- Prefix / Prompt tuning (arXiv:2101.00190, 2104.08691): prepend learned virtual tokens (or their KV states) that every layer attends to — conditioning from inside the sequence. Detailed in topic 180; note the inference tax now: those extra tokens ride along in every attention computation.
Selective methods — train a chosen subset of existing weights:
- BitFit (arXiv:2106.10199): biases only (~0.08% of parameters) — most of the gain on small/medium models.
- Layer-subset / sparse tuning: update only the top transformer blocks (or a chosen slice), freezing the rest.
Retraining methods — discard/reinitialize task-specific components and train just those (task heads, embeddings; in the LLM era this family survives mainly as embed-tuning and adapter-based head replacement).
The ranking that emerged from benchmark sweeps (GLUE, Super-Natural Instructions, code tasks), worth memorizing:
codequality per trainable parameter at LLM scale: LoRA ≥ adapters > prefix ≈ prompt tuning > BitFit + LoRA is uniquely MERGEABLE → zero added inference latency
The code block shows the whole interface in practice: five lines of config wrap any Hugging Face model and the trainable count collapses ~1000x.
from peft import LoraConfig, get_peft_model
cfg = LoraConfig(r=16, lora_alpha=32, target_modules="all-linear",
lora_dropout=0.05, task_type="CAUSAL_LM")
model = get_peft_model(base_model, cfg)
model.print_trainable_parameters()
# trainable params: 19,922,944 || all params: 6,761,267,072 || 0.29%07.The Economics That Changed Deployment
PEFT did two things: it turned fine-tuning from a cluster job into a single-GPU job (section 3's arithmetic), and — more disruptively — it changed serving architecture:
- Checkpoint size. A rank-16 LoRA on a 7B model is ~20–80 MB versus 14 GB for the full model. Thousands of task variants fit in S3 buckets, not model registries.
- Multi-tenant serving. One shared base model resident in GPU memory + per-request hot-swap of different adapters — Hugging Face PEFT
load_multiple_adapters, and vLLM/SGLang multi-LoRA serving since 2024. This is literally how Fireworks, Together, Modal, and Replicate price per-tenant fine-tunes affordably: four hundred customers, one replica, one booklet each (the orchestra picture, invoiced). - Privacy and licensing. Base weights never move; tenants ship only adapters. A lab can license a foundation model while customers own their deltas — and a customer can delete their fine-tune by deleting one small file.
Compare the deployment shapes directly:
codeFFT world: customer A → dedicated GPU + 14 GB checkpoint customer B → dedicated GPU + 14 GB checkpoint (×400) PEFT world: ┌─ one base replica (shared, in memory) ─┐ │ req #1 → + adapter_A (40 MB, hot-swap) │ │ req #2 → + adapter_B (40 MB, same batch)│ └─────────────────────────────────────────┘ (×1)
08.In Practice: Limits and Failure Modes
PEFT is not free quality. Know the edges:
- The ceiling is real. At very large target-task distributions (50k+ examples of a genuinely new capability), full fine-tuning still tends to win: additive modules bottleneck what the base can reorganize. The booklet can steer the orchestra; it cannot teach new muscle memory.
- Conditioning overhead. Prefix/prompt methods add per-layer attention cost at inference and are notoriously harder to optimize (longer effective sequences, sensitive initialization).
- Hyperparameter sensitivity is higher than FFT. A single adapter rank, alpha, and target-module set can swing results several points — so adapter recipes require ablation, not defaults (topic 177's section on r/α/targets exists for this reason).
- Merging is a one-way door. Un-meritted adapters roll back by unloading one file — but once you merge a LoRA into the base, the delta is baked in; a data-incident rollback after merging means re-exporting the whole model.
- Governance drift. Thousands of tiny adapters complicate provenance, evals, and "which variant is live?" audits — cheap artifacts still need registries, versioning, and tests.
Decision shortcut: many variants on shared infrastructure, small data, tight GPUs, reversible updates → PEFT. Fundamental capability expansion with big data and a cluster → FFT (or continued pretraining) — and for the single-GPU version of that power move, QLoRA (topic 178) combines PEFT with a quantized base.
Architectural Trade-offs & Production Realities
Architectural Advantages
- 0.05-5% trainable parameters: optimizer state shrinks ~100-1000x, enabling single-GPU fine-tuning of 7B-70B models.
- Tiny checkpoints make thousands of per-task/per-tenant variants storage-cheap and swappable at serving time.
- Inherent regularization: frozen base curbs overfitting and catastrophic forgetting on small datasets.
- One shared base GPU replica can serve many adapters concurrently (vLLM/PEFT multi-adapter support).
Trade-offs & Constraints
- The full frozen base must still load and run every forward pass — memory/bandwidth savings require quantization too.
- Quality ceiling below full fine-tuning for deep, distribution-shifting capability changes.
- Higher hyperparameter sensitivity (rank, alpha, targets, family choice); conditioning methods add inference overhead.
- Adapter sprawl complicates provenance, evals, and rollback governance at scale.
Providers keep a single replica of an open model (e.g., Llama-3-8B) resident in GPU memory and load customer-specific LoRA adapters (tens of MB each) per request, batched together. Customers pay for a fraction of a GPU instead of a dedicated 70B-class deployment — an architecture only possible because PEFT checkpoints are tiny and swappable.
Staff+ Engineering Takeaways
- PEFT freezes the base model and trains 0.05-5% of parameters, justified by the low intrinsic dimensionality of task adaptation.
- Three families: additive modules (LoRA, adapters, IA³), selective weight training (BitFit), and retraining/reinit of task components.
- LoRA dominates modern practice because it is mergeable — zero inference latency — while prefix/prompt methods add conditioning overhead.
- PEFT shrinks optimizer state and checkpoints, not base-model inference cost; combine with QLoRA for full memory collapse.
- Tiny adapters enabled the multi-tenant serving economics that define commercial LLM fine-tuning platforms in 2024-2026.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Which statement about memory during LoRA fine-tuning of a frozen 7B BF16 model is correct?
How clear and actionable was this distributed systems breakdown?