Batch, Mini-Batch, Epoch, and Iteration: The Training-Loop Grammar
Four words you must never mix up: a batch is the handful of examples used per look, an iteration is one look, an update is one adjustment, an epoch is one full pass over the data. This topic fixes the grammar and shows why batch size is the most consequential number in any training run — it changes what the optimizer learns, not just how fast.
01.The Problem: Two Engineers, Two Meanings of "Batch"
You post online: "my run took 10 epochs with batch 2048."
Reply A: "Wait — how many iterations is that?"
Reply B: "Why epochs? This is 2026, we count steps."
Reply C: "Batch per GPU or total?"
What do all these words even MEAN?
Every training conversation assumes this vocabulary, and mixing it up causes real bugs: learning-rate schedulers that end halfway through training, evaluations that never fire, "effective batch size" claims that are 8× off.
Even worse: choosing how much data to look at before each update changes what the model learns, not just how fast it learns.
This topic is the grammar. Five words, in order.
Epochs, Iterations, and Updates
Epochs, Iterations, and Updates
One epoch = one pass over shuffled data = N/B iterations = N/B parameter updates (for simple SGD; gradient accumulation makes updates sparser than iterations).
Unlock Topic #87: Batch, Mini-Batch, Epoch, and Iteration: The Training-Loop Grammar
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?