Model Merging: SLERP, TIES, DARE
Model merging combines fine-tuned checkpoints by arithmetic on weights — no training, no pooled data. This topic explains why same-lineage merges work at all (linear mode connectivity), the methods (weight-average soups, task vectors, SLERP, TIES, DARE), and why merges so often wreck calibration and instruction-following, plus the recipe that keeps them honest.
01.The Problem: Two Good Models, Zero Training Budget
You have two checkpoints. One fine-tuned on math is great at proofs. The other fine-tuned on code is great at Python. You want one model that is good at both.
The textbook answer is: pool the data and fine-tune again. Three problems with that answer:
- You may not have both datasets. Privacy, licensing, competing tenants. The math data belongs to someone else.
- You may not have training compute. A 70B fine-tune run is weeks of GPUs.
- You may not have permission to train at all. You downloaded two open checkpoints; retraining from scratch is not on the table.
So instead of training, you can do arithmetic. Take the numbers inside model A and the numbers inside model B, combine them with a formula, and write the result to disk as model AB. Minutes of GPU time. No data. No gradients.
How can just averaging numbers possibly produce a model that keeps both skills?
Because fine-tuning barely moves the weights — a trained checkpoint sits close to its starting point, and the directions it moved often live in compatible subspaces. Merging is the art of combining those movement directions without letting them collide.
The title of this section should already warn you: it often works, and when it fails it fails quietly (section 6).
Merging Is Cheap Search, Not Free Lunch 🧬
Merging Is Cheap Search, Not Free Lunch 🧬
Every practical merge converts checkpoints into task vectors, applies a density/sign/geometric rule, and then must repair and re-evaluate - because the arithmetic is trivial and the damage is subtle.
Unlock Topic #259: Model Merging: SLERP, TIES, DARE
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?