The Convolution Operation
Under the hood, a convolution is just a sliding dot product: a learnable filter is laid over a small window, the overlapping numbers are multiplied and summed into one score, and the window slides on. This topic walks that arithmetic, shows how 3D filters span color channels, and clears up the cross-correlation naming nuance PyTorch actually performs.
One conv output = one windowed dot product 🧮
Each pixel of the output map is a weighted sum of a local input patch. The filter slides across the input; at every position it multiplies element-by-element, sums to a scalar, adds a bias, and writes that scalar into the output map.
01.The Problem: How Does a Filter Actually "Read" One Window?
Last topic said a filter is a small learnable stencil you slide across an image.
But "slide a stencil" is a metaphor.
What is the arithmetic inside one position?
You have a tiny square of pixels — say 3x3 of them — and a matching 3x3 grid of learned numbers.
You need one output value that answers:
"How strongly does the pattern this filter represents appear in this exact window?"
That is the whole job. Everything else — sliding, channels, many filters — is just repeating this one step.
So the question is: how do you squash 3x3 = 9 input pixels and 9 filter weights into a single "how well does this match" score?
02.The Idea in Plain Words: One Output = One Similarity Score
A convolution is simply
A dot product between a learnable filter and a local window, slid over every position.
For a 3x3 grayscale input patch and a 3x3 filter, one output value is:
y = Σ_i Σ_j ( x[i,j] * w[i,j] ) + b
Read it piece by piece:
x[i,j]— the pixel in the window at rowi, columnj.w[i,j]— the filter weight sitting on top of that pixel.x[i,j] * w[i,j]— multiply the two overlapping numbers.Σ_i Σ_j— add up all 9 products into one scalar.+ b— add one learnable bias (a threshold that decides how easy it is to "fire").
The result y is one number. Move the window one step right (or down) and do it again. Collect every position's y and you have the feature map.
03.A Tiny Worked Example: Scoring a Patch
Use a vertical-edge filter — it wants "dark on the left, light on the right":
codefilter w = patch A (an edge): patch B (flat gray): -1 0 1 0 0 7 5 5 5 -1 0 1 0 0 7 5 5 5 -1 0 1 0 0 7 5 5 5
Multiply overlapping numbers and add:
Patch A (real edge): each row contributes (-1*0) + (0*0) + (1*7) = 7. Three rows → 7 + 7 + 7 = 21. Big positive score. The edge detector fires.
Patch B (flat gray): each row contributes (-1*5) + (0*5) + (1*5) = -5 + 5 = 0. Total 0. Nothing fires.
Now add a bias b = -10:
- Patch A:
21 - 10 = 11→ still positive, fires. - Patch B:
0 - 10 = -10→ negative, suppressed.
The bias just tunes the firing threshold. The network learns both the ws and b by backprop — none of these numbers are hand-set. The filter simply becomes a good edge detector because edges help the final task.
04.Visual Intuition: A Flashlight Sweeping a Photo
Picture a 5x5 photo and a 3x3 filter as a small flashlight you press against it. The beam lights up a 3x3 block; you compute one score for that block; you slide right; repeat; wrap to the next row.
code[ ■ ■ ■ . . ] window here → score y(0,0) [ ■ ■ ■ . . ] [ ■ ■ ■ . . ] [ . . . . . ] [ . . . . . ] slide right one: [ . ■ ■ ■ . ] → y(0,1) [ . ■ ■ ■ . ] [ . ■ ■ ■ . ]
The grid of scores you filled in is the feature map. Each of its cells is literally "how much this pattern showed up at that spot." That is why a conv output is still a little image — position is preserved.
05.The Analogy: A Fill-in-the-Blanks Answer Key
Think of the filter as a grading rubric, and the window as one student's answer.
- The rubric assigns a weight to each slot: high positive for "great, mention it," negative for "penalize this," zero for "ignore this."
- Grading one answer = multiply each student response by its slot weight and total the marks. That single total is the dot product.
- A high mark = this answer strongly matches what the rubric rewards (the pattern is present).
- The bias is the pass line: the score you need before you call it a hit.
- You grade every student (every window position) with the same rubric. Same weights, everywhere — that's weight sharing.
A vertical-edge rubric grades "does this window look like a left-to-right brightness jump?" And the network writes and rewrites the rubric itself, trying to make the final class prediction correct.
06.Real Inputs Are 3D: Filters That Span Color
So far, flat grayscale. Real images have depth. A color image is H x W x 3 (red, green, blue stacked).
That means a filter cannot be flat 3x3 — it must be a 3x3x3 cube.
The rule that trips everyone up:
A filter's depth always equals the input's depth.
So one 3x3x3 filter has 3*3*3 = 27 weights + 1 bias = 28 numbers. It looks at all three color channels at once. The dot product now sums over the extra axis:
y = Σ_c Σ_i Σ_j ( x[c,i,j] * w[c,i,j] ) + b
c runs over the input channels.
Crucial: the filter is summed across the entire input depth at each window position, collapsing C_in channels into a single scalar. So one filter → one output number per position → one output map, no matter how many channels the input had.
07.Many Filters = Many Output Channels
If one filter collapses everything to one map, how do we get 64 or 256 channels?
By using a bank of filters. One filter is one pattern detector, so to learn many patterns you use many filters.
If a layer has C_out filters, each spanning the full input depth, you get C_out output feature maps stacked into an H' x W' x C_out volume.
The next layer then receives that thicker volume, so ITS filters become K x K x C_out. This chain is why a filter's depth is dictated by the previous layer's number of filters — and why parameter counts balloon with depth, motivating the small-3x3 and depthwise-separable tricks covered later.
The budget math, worth memorizing:
- Parameters per filter:
K * K * C_in + 1(the +1 is the bias). - Total layer weights:
(K*K*C_in + 1) * C_out.
import torch, torch.nn.functional as F
x = torch.randn(1, 3, 5, 5) # (batch, C_in, H, W)
w = torch.randn(8, 3, 3, 3) # (C_out, C_in, K, K)
b = torch.zeros(8)
out = F.conv2d(x, w, bias=b, stride=1, padding=0)
print(out.shape) # torch.Size([1, 8, 3, 3])
# A single 3x3x3 filter, one position:
patch = x[0, :, 0:3, 0:3] # (3,3,3)
manual = (patch * w[0]).sum() + b[0] # == out[0,0,0,0]08.Why This Is Cheaper Than Dense (and What AI Does With It)
Now count what one little filter actually does.
A 3x3x3 filter has 28 parameters, but you reuse it at every output position. On a 56x56 map that is 56*56 = 3,136 positions, each doing 27 multiply-adds — yet only 28 numbers to learn.
A dense layer wiring every input pixel to every output pixel for the same map would need millions of weights.
Convolution gets its power from two pillars of deep learning (alongside non-linearity):
- Weight sharing — one set of
K*K*C_inweights reused everywhere → orders of magnitude fewer parameters. - Sparsity of connections — each output sees only a small
K x K x C_inneighborhood, not the whole image.
Because the operation is a perfectly regular sliding dot product, hardware (cuDNN, Apple's neural engine, TPUs) implements it with matrix-multiply tricks like im2col and Winograd, hitting teraFLOPs of throughput. The same arithmetic you just did by hand at one window is what every GPU runs billions of times per second.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Locality: each output depends on a tiny neighborhood, matching visual structure.
- Weight sharing → orders of magnitude fewer parameters than dense layers.
- Parallelizable across positions — ideal for GPUs/TPUs (im2col, Winograd, cuDNN kernels).
Trade-offs & Constraints
- Fixed kernel shape cannot adapt to variable-scale objects without multi-scale tricks.
- Each filter collapses all input channels — cross-channel mixing only via many filters.
- A true convolution (with flip) differs from library cross-correlation — a naming trap.
Because convolution is a regular sliding-window dot product, hardware vendors implement it with im2col + matrix multiply, the Winograd transform, or systolic arrays. The deterministic locality of conv is exactly what lets tensor cores and neural engines achieve teraFLOPs of throughput.
Staff+ Engineering Takeaways
- A convolution output value = elementwise multiply of a local patch with the kernel, summed, plus bias.
- Filters are 3D: their depth always equals the input channel count; summing over depth yields one scalar per position.
- One filter → one output channel; a bank of C_out filters → C_out-deep feature map.
- PyTorch performs cross-correlation (no kernel flip), which is fine because kernels are learned.
- Weight sharing and local connectivity make convs far cheaper and better-regularized than dense layers.
- Layer parameters: (K * K * C_in + 1) * C_out.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
For a color (3-channel) input, a single 3x3 conv filter produces how many output channels?
How clear and actionable was this distributed systems breakdown?