Gini Impurity
Decision trees need one number to rate how messy a node is. Gini impurity is that number: the chance you would guess a random sample's label wrong if you drew a fake label from the node's own mix. This topic computes it, weights it across a split, and shows how CART uses it to pick questions.
01.The Problem: A Tree Must Compare Questions — It Needs One Number
Remember what a decision tree does (Topic 56: a flowchart of if-then questions, built by trying many questions and keeping the best ones).
Here is the moment of truth. The tree stands at a node holding some data and must choose ONE question from dozens of candidates:
age ≤ 30?income ≤ 45k?visited website in last 7 days?
"Which question is BEST?"
To answer, the tree needs to look at the piles a question creates and rate them with one single number: how mixed are the class labels inside this pile?
Messy pile = the number is high. Pure pile = the number is 0.
That number is called the impurity.
Several impurity scores exist — entropy (Topic 57), classification error — but the CART algorithm (Classification and Regression Trees, Breiman et al., 1984) is built on one specific score, and it is the default you will meet in scikit-learn every day.
It is called Gini impurity.
How a Tree Uses Gini to Pick a Split 🔪
How a Tree Uses Gini to Pick a Split 🔪
Every candidate split produces child nodes, each with its own Gini score. The split with the lowest sample-weighted child impurity wins.
Unlock Topic #58: Gini Impurity
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?