TOPIC #58Beginner 11 min read

Gini Impurity

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Decision trees need one number to rate how messy a node is. Gini impurity is that number: the chance you would guess a random sample's label wrong if you drew a fake label from the node's own mix. This topic computes it, weights it across a split, and shows how CART uses it to pick questions.

01.The Problem: A Tree Must Compare Questions — It Needs One Number

Remember what a decision tree does (Topic 56: a flowchart of if-then questions, built by trying many questions and keeping the best ones).

Here is the moment of truth. The tree stands at a node holding some data and must choose ONE question from dozens of candidates:

  • age ≤ 30?
  • income ≤ 45k?
  • visited website in last 7 days?
Insight

"Which question is BEST?"

To answer, the tree needs to look at the piles a question creates and rate them with one single number: how mixed are the class labels inside this pile?

Messy pile = the number is high. Pure pile = the number is 0.

That number is called the impurity.

Several impurity scores exist — entropy (Topic 57), classification error — but the CART algorithm (Classification and Regression Trees, Breiman et al., 1984) is built on one specific score, and it is the default you will meet in scikit-learn every day.

It is called Gini impurity.

How a Tree Uses Gini to Pick a Split 🔪

PRO Architecture Blueprint

How a Tree Uses Gini to Pick a Split 🔪

Every candidate split produces child nodes, each with its own Gini score. The split with the lowest sample-weighted child impurity wins.

How a Tree Uses Gini to Pick a Split 🔪
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #58: Gini Impurity

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?