Max & Average Pooling
Pooling shrinks each feature map on its own using a fixed little operator — max keeps the strongest value in a window, average keeps the mean. It has no learnable parameters and never mixes channels. This topic compares the two, derives output sizes, and explains why classic local pooling is being replaced by strided convolutions and Global Average Pooling.
01.The Problem: The Maps Keep Getting Too Big
A feature map from the first layers is huge — 224x224x64 is over three million numbers per image.
That is a lot of memory and a lot of multiplying for the next layer.
And honestly, a lot of it is over-precise: once we know "there's an edge around here," we do not always need to know it is at pixel (47, 102) and not (48, 103).
So we want to:
- shrink the map,
- stay tolerant to tiny shifts (a moved cat is still a cat),
- and do it cheaply, ideally with nothing to learn.
How do you squash a map down without losing the parts that matter?
The classic answer is pooling.
Max vs average pooling on one map 🟦🟨
Max vs average pooling on one map 🟦🟨
Both pool over each 2x2 block with stride 2, halving resolution. Max outputs the block maximum (keeps the peak detector response); average outputs the block mean (smooths toward background).
Unlock Topic #101: Max & Average Pooling
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?