PHASE 10 CURRICULUM

Fine-tuning & Adaptation

Progress0 of 9 (0%)

Adapting pretrained models to your domain without rebuilding them:

Key Architectural Domains & Syllabus
full fine-tuning economics
parameter-efficient methods (LoRA, QLoRA, adapter modules, prefix and prompt tuning)
quantization for training and inference (INT8/INT4, FP16/BF16, GPTQ, AWQ)
knowledge distillation
structured pruning. Includes the decision framework for when to prompt
retrieve
or fine-tune
9 In-Depth Topics ~72 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 10

0 of 9 completed
#175Full Fine-tuningAdvancedFREE

Full fine-tuning continues training on every weight of a pretrained model until it behaves like a specialist. That buys the highest quality ceiling — and costs ~16 bytes of memory per parameter, which is why FSDP/ZeRO sharding, BF16, and activation checkpointing are part of the recipe.

13 min read•3 Quiz Questions

PEFT freezes the pretrained model and trains only 0.05–5% of its parameters, because task adaptation lives in a low-dimensional corner of weight space. That tiny trainable slice turned fine-tuning into a single-GPU job and created the multi-tenant adapter-serving economy of 2024–2026.

13 min read•3 Quiz Questions

LoRA freezes the pretrained weights and learns the change as a thin product of two small matrices (ΔW = (α/r)·B·A), because task updates are low-rank. Fewer trainable parameters, no inference latency after merging, and a whole 2024–2026 ecosystem of scaling fixes and multi-adapter serving built on that one equation.

14 min read•3 Quiz Questions

Fine-tuning a huge model normally needs a monster server. QLoRA stores the frozen base model in 4-bit NF4, unpacks it to BF16 just for the math, and sends every gradient into a small BF16 LoRA adapter — collapsing 65B-class training onto a single 48 GB GPU.

12 min read•3 Quiz Questions
#179Adapter ModulesAdvancedFREE

The original parameter-efficient fine-tuning method: tiny bottleneck side-paths spliced into a frozen transformer. Houlsby-style serial and parallel placement, reduction factors, multi-task routing and fusion, and why adapters still matter next to mergeable LoRA.

12 min read•3 Quiz Questions

Teach a frozen model a new task without touching a single weight: prompt tuning learns fake input tokens, prefix tuning learns per-layer attention key/value states. Here is what each one trains, why the optimization is touchy, and where they still win in 2024-2026 stacks.

12 min read•3 Quiz Questions

Every number in a model costs bytes, and bytes decide what fits on a GPU and how fast tokens come out. This topic is the map of numeric formats — FP32/FP16/BF16/FP8/INT8/INT4 — plus the post-training quantization methods (GPTQ, AWQ, SmoothQuant) and the memory-bandwidth math behind them.

13 min read•3 Quiz Questions
#182Knowledge DistillationAdvanced PRO

Compress capability, not just bytes: a big teacher model trains a small student by sharing its soft probability distributions (via temperature) or, in the LLM era, its written output as synthetic data. The Phi playbook made this the cheapest way to raise what a small model can do.

12 min read•3 Quiz Questions
#183Model PruningAdvanced PRO

Most trained weights barely matter — pruning deletes them. But only certain deletion patterns convert into real speed: unstructured zeros need special kernels, 2:4 blocks ride NVIDIA hardware for ~2x, and structured cuts shrink the model outright. Here is the ladder, the one-shot LLM methods (SparseGPT, Wanda), and the honest math of when sparsity pays.

12 min read•3 Quiz Questions