Feature Maps and Activation Maps
A conv layer does not output a single number — it outputs a stack of 2D score sheets, one per filter, called feature maps, which together form an H x W x C activation volume. This topic defines pre- vs post-activation maps, the tensor shape (NCHW vs NHWC), and how map size shrinks while depth grows through a network.
From filters to a 3D activation volume 🗂️
A layer first computes linear feature maps (z = W*x + b), then applies ReLU to get activation maps. The C_out stacked maps form the output tensor consumed by the next layer.
01.The Problem: The Filter Slides — But What Does It Hand Back?
You now know a filter is a small learnable block of weights, and that convolution is a sliding dot product.
But a dot product is one number.
So if a 3x3 filter scores 56x56 = 3,136 window positions, what comes out of that layer? One number? A list? A shape?
This matters enormously:
- You cannot wire layers together if you don't know what shape each one produces.
- You cannot debug "size mismatch at the linear layer" errors.
- You cannot understand why a CNN can segment an image (label every pixel) instead of just classifying it.
The answer is the feature map, and this whole topic is about getting comfortable looking at it.
02.The Idea in Plain Words: One Filter, One Sheet of Scores
A feature map (a.k.a. activation map) is simply
The 2D grid of scores one filter writes as it slides across the input.
Each cell of that grid holds the dot-product answer for one window position.
- A big positive value → "the pattern this filter loves is strongly present right here."
- A zero / negative → "nothing matching here."
Because a layer holds C_out filters, it emits C_out feature maps. Stack them along the channel axis and you get a 3D volume — the layer's real output.
So the chain is: one filter → one 2D map → C_out maps → one 3D volume.
03.A Tiny Worked Example: Reading a Sheet
Imagine one 3x3 edge filter over a tiny 4x4 image (stride 1, no padding). The window fits in 2x2 places, so the filter produces a 2x2 map.
codevertical-edge map: ┌────────┐ │ 18 2 │ top-left window: strong edge here │ 0 -5 │ bottom-right: no vertical edge └────────┘
Each number answers "was there a vertical edge at this spot?" Now imagine a second filter (horizontal edges) yields another 2x2 sheet, and a third (blobs) a third.
Stack the three sheets:
codeoutput volume = 2 x 2 x 3 (H x W x C) ↑ ↑ ↑ rows cols detectors
Read it as: "a 2x2 grid of locations, each described by 3 detector readings." That little cube is the layer output, and the next conv treats its 3 channels as its C_in.
04.Visual Intuition: A Stack of Transparency Sheets
Picture the input photo flat on a light table.
- Each filter is a highlighter that, on its own transparency sheet, marks every place its pattern shows up.
- The photo's edges get a vertical-edge sheet, a horizontal-edge sheet, a blob sheet, …
C_outsheets in all. - The channel axis is literally the physical stack of sheets.
- The spatial (H, W) axes are the position on each sheet — still aligned with the original photo.
So a feature volume is "many annotated copies of the same picture, each highlighting one learned feature." Channels are not colors — they are learned detectors: one fires for horizontal lines, one for a particular fur texture, one for a dog-snout-like part.
Where the whole stack is stored in memory differs by framework:
codeoutput tensor shape = (batch, C_out, H_out, W_out) # PyTorch NCHW ("channels first") = (batch, H_out, W_out, C_out) # TensorFlow NHWC ("channels last")
Same data, different sheet order. The stack-vs-plane idea is identical.
So when someone asks "what's the shape coming out of this conv layer?", answer in two pieces: the spatial size H_out x W_out comes from the sliding arithmetic (it shrinks with pooling/stride), and the depth C_out is just "how many filters did this layer own." Never confuse the two axes — that mix-up is the classic "size mismatch at the linear layer" bug.
05.The Analogy: The Inspector's Report Pages
Return to the mural inspector (from the CNN overview) with a clipboard.
After sweeping one stencil (filter) across the whole wall, she does not say "the mural has one edge." She fills an entire report page: one box per wall region, ticked by how strongly her pattern showed there.
- One stencil → one report page (a feature map).
- A drawer of
C_outstencils →C_outpages bound together (a channel-depth volume). - The next inspector reads the bound report as her input, so her stencils must be as thick as the report has pages (
C_inagain). - Stepping back and summarizing blocks of pages = pooling (shrinks the page, keeps the gist).
This "report of where, per detector" framing is exactly why the network can later point at a region (segmentation, Grad-CAM) instead of only naming the whole picture.
06.Spatial Size vs Depth: What Changes Going Deeper
Through a well-designed CNN, two axes move in opposite directions:
- Spatial size (H, W) shrinks as you stack pooling and strided-conv layers:
224 → 112 → 56 → 28 → 14 → 7. - Depth (C_out) grows as you add more filters per stage:
64 → 128 → 256 → 512.
Why balance them? The product H * W * C — the feature-vector length that eventually gets flattened — often stays roughly constant. That is why VGG's final 7x7x512 = 25,088-long vector is a sensible bottleneck instead of an explosion.
Two rules drive every shape you'll ever trace:
- Map size at any layer follows
floor((H - K + 2P)/S) + 1. - Depth simply equals the number of filters you chose for that layer.
Trace it live — each print shows the (batch, channels, H, W) NCHW shape after every stage:
import torch, torch.nn as nn
x = torch.randn(1, 3, 224, 224)
stages = nn.Sequential(
nn.Conv2d(3, 64, 3, padding=1), nn.ReLU(), # 64 x 224 x 224
nn.MaxPool2d(2), # 64 x 112 x 112
nn.Conv2d(64, 128, 3, padding=1), nn.ReLU(), # 128 x 112 x 112
nn.MaxPool2d(2), # 128 x 56 x 56
nn.Conv2d(128, 256, 3, padding=1), nn.ReLU(), # 256 x 56 x 56
)
for i, layer in enumerate(stages):
x = layer(x)
print(f"{layer.__class__.__name__}: {tuple(x.shape)}")07.Why It Matters: The Map Remembers Where (and How to Look Inside)
Here is the superpower hidden in "one sheet per detector."
Within a single activation map, spatial position is preserved: the value at cell (14, 30) corresponds to a definite location in the original image (scaled by whatever downsampling happened). That is not true of a dense layer's output vector — position was destroyed the moment you flattened.
Two consequences engineers exploit daily:
- Dense prediction. Because the map is itself an image-sized grid of per-location scores, CNNs can do segmentation and detection — label every pixel, not just the whole photo. (Later we add Global Average Pooling, which deliberately collapses each map to one number, trading away position to gain invariance for pure classification.)
- Grad-CAM interpretability. For a classifier, the final feature map (say
7x7x2048) is where you plug Grad-CAM: back-propagating the chosen class score onto that map produces a coarse heat-map showing which region of the image the network "looked at." This only works because the deep map still carries position.
08.In Practice: Taming Channel Depth
Not all C_out channels earn their keep — some fire near-identically, wasting memory and compute. Two standard tricks manage feature-map depth:
- Grouped convolutions split
C_inintoggroups; each filter reads onlyC_in/gchannels (fewer cross-channel sums), cutting parameters for a small accuracy cost. - Bottleneck (1x1) layers first shrink map depth, run a cheap
3x3on the thin version, then expand it back — the ResNet bottleneck's whole reason for existing.
One cost to keep in view: activation memory scales with batch x C x H x W. For high-resolution inputs, storing all these maps for backprop — not the weights — is the dominant VRAM expense. That's why you sometimes downsample early, or, when a head only needs a global descriptor, end with adaptive average pooling to 1x1 instead of flattening huge maps.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Maps are interpretable and visualizable — you can literally plot any channel.
- Preserving spatial position enables dense tasks (segmentation/detection heads).
- Stacking maps into depth gives the next layer many complementary detectors.
Trade-offs & Constraints
- Large early maps (224x224) are memory-heavy — activation storage dominates training VRAM.
- Blindly increasing C_out wastes compute on redundant channels.
- Pre- vs post-activation terminology is inconsistently used across papers.
Distill.pub-style feature-visualization and Lucid optimize input images to maximize a chosen activation map, revealing which intermediate feature (a single scalar in a deep channel) corresponds to "striped window" or "dog" concepts. These maps are the currency of CNN debugging.
Staff+ Engineering Takeaways
- One filter → one 2D feature map; C_out filters → a H x W x C_out activation volume.
- Pre-activation z = W*x + b; post-activation a = ReLU(z) — the stored activations feed backprop.
- Spatial size shrinks (pooling/stride) while channel depth grows (more filters) through the network.
- A map value is a "pattern-present-here" score that preserves input position — enabling dense prediction.
- Final maps enable Grad-CAM; grouped and bottleneck convs control map depth for efficiency.
- Activation memory scales with batch x C x H x W — the dominant VRAM cost for high-res inputs.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
The channel depth of a conv layer's output feature volume equals:
How clear and actionable was this distributed systems breakdown?