Gradient Boosting: XGBoost, LightGBM & CatBoost
Gradient boosting trains many small decision trees one after another. Each new tree only fixes the mistakes the earlier trees left behind. That simple loop is secretly gradient descent on the loss (Friedman, 2001) — and the engineering around it (XGBoost, LightGBM, CatBoost) is why these models still dominate tabular data.
01.The Problem: Your First Model Is Wrong — In a Useful Way
Imagine you predict house prices from size.
You fit one small decision tree. It's okay. It learns "small houses are cheap, big houses are expensive."
But it misses details. It says every mid-size house costs 300k, when some go for450k.
So the question becomes
Can I build a second model whose only job is to fix the mistakes of the first?
That's the seed of gradient boosting.
You're not throwing the first tree away. You keep it, measure exactly where it's wrong, and train a second tree to predict the errors themselves. Then a third to fix what the second missed. And so on.
Wait — isn't that cheating? Doesn't the model just memorize the training set?
Sometimes, yes. That's exactly why boosting has built-in brakes (shrinkage and early stopping), and why the way you fix mistakes matters so much. You'll see both below.
One contrast to anchor everything: random forests (Topic 59) grow many trees in parallel, each independent, mainly to reduce variance (noise sensitivity). Boosting grows trees in sequence, each dependent on the last, mainly to reduce bias (systematic error). Same ingredient — weak trees — opposite recipe.
Boosting as Gradient Descent in Function Space 🚀
Boosting as Gradient Descent in Function Space 🚀
Each new tree fits the gradients of the loss with respect to the current ensemble prediction; the three flagship libraries differ in how they grow and regularize those trees.
Unlock Topic #60: Gradient Boosting: XGBoost, LightGBM & CatBoost
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?