Train / Validation / Test Split
A student who rewrites the answer key while practicing looks perfect but learns nothing. Splitting data into train (study), validation (mock exams) and a sealed test set (one final exam) is what turns a model's self-reported score into an honest prediction of how it will do in the real world.
01.The Problem: A Student Who Grades Himself on His Own Notes
Imagine a student preparing for an exam.
At home she practices on the exact questions that will appear on the exam. At the kitchen table she scores 100%.
Is she ready?
You have no idea — she may have simply memorized the answers.
A model trained and then evaluated on the same data is that student. It looks near-perfect because it has memorized the answers. That fake confidence is overfitting (Topic 52), and without a check you cannot tell genuine signal from memorization.
What you actually care about is the opposite question:
How will this model do on data it has never seen?
That is called generalization — and the only data you will ever deploy against is unseen data. So you must hide some of your data from the model, purely for honest grading.
Three Splits, Three Purposes
Three Splits, Three Purposes
Train fits parameters, validation tunes hyperparameters and selects models, and the untouched test set gives an unbiased final estimate of generalization.
Unlock Topic #34: Train / Validation / Test Split
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?