Filters (Kernels)
The filter (or kernel) is the learnable heart of a convolution: a small 3D block of weights that acts as a sliding pattern detector. This topic nails down filter shape (K x K x C_in), the filter-vs-bank distinction, how to count its parameters, and why tiny learned 3x3 kernels beat big hand-written ones.
A filter bank and its output depth 🎛️
Each filter is a K x K x C_in weight tensor that scans the full input depth and yields one output map. The number of filters in the bank (C_out) sets the output channel depth.
01.The Problem: You Keep Hearing "Filter" — What Exactly Is It?
The last two topics said a convolution "slides a filter" and "scores a window."
But people use the words filter, kernel, convolution, and feature map loosely, and beginners mix them up constantly.
So before anything else, get one object crisp in your head:
What is a filter, physically? How big is it? How many numbers does it hold? And how many does a whole layer hold?
Answer those and half of all shape bugs in your first CNN vanish. That's what this topic is for.
02.The Idea in Plain Words: A Filter Is One Small Block of Weights
A filter (a.k.a. kernel) is simply
A tiny learnable block of numbers of shape
K x K x C_in, used as one sliding pattern detector.
Unpack every symbol:
Kis the spatial extent — the "kernel size." Most often 3, 5, or 7. It's how wide/tall the detector's window is.C_inis the input depth. The filter's depth must equal the input's channel count, because it sums over every input channel at each position.
Dimensionality changes the shape but not the idea:
- 1D (audio, text):
K x C_in. - 2D (images):
K x K x C_in. - 3D (video, medical volumes):
K x K x K x C_in.
Three crucial distinctions to keep straight, in order of size:
- One filter = one
K x K x C_intensor → produces one output feature map. - One convolutional layer = a bank of
C_outfilters → produces aH' x W' x C_outoutput volume. - The word "kernel" in deep learning means this weight tensor, NOT the classic signal-processing "kernel function" — in ML the two happen to coincide.
03.A Tiny Worked Example: Counting the Numbers in One Conv2d
Let's build a nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3) and count every parameter.
One filter: K*K*C_in = 3*3*3 = 27 weights, plus 1 bias = 28 numbers.
The layer has C_out = 64 such filters:
code64 * 27 weights = 1,728 64 * 1 bias = 64 total = 1,792
Compare that to a dense layer doing the same job on a 224x224x3 image: hundreds of millions. The filter is a cookie cutter you reuse everywhere, so it's tiny.
The formula behind it, worth memorizing:
parameters = (K*K*C_in + 1) * C_out = (3*3*3 + 1) * 64 = 28 * 64 = 1,792
Notice what sets what: K and C_in set each filter's size; C_out is just how many filters you own.
import torch.nn as nn
conv = nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3, padding=1)
print(conv.weight.shape) # torch.Size([64, 3, 3, 3]) -> C_out, C_in, K, K
print(conv.bias.shape) # torch.Size([64]) -> one bias per filter
# Parameter budget:
params = conv.weight.numel() + conv.bias.numel()
print(params) # 64*3*3*3 + 64 = 1,79204.Visual Intuition: A Drawer of Cookie Cutters
Think of a convolutional layer as a drawer of cookie cutters.
codeC_in = 3 layers of dough stacked (R,G,B) cutter 1 (3x3) ┌───┐ -> "vertical edge" -> one imprint map cutter 2 (3x3) ┌───┐ -> "horizontal edge" -> one imprint map cutter 3 (5x5) ┌─┐ -> "blob/corner" -> one imprint map ... ││ cutter 64 (3x3) └───┘ -> "whisker-like" -> one imprint map 64 cutters stamped across the dough = a 64-deep sheet of imprints
- Each cutter is one filter (
K x K x C_in). Its shape (the cut) is the pattern it detects. - The dough depth (
C_in) is what the cutter must press through — so the cutter's depth matches the dough. - You own a bank of cutters (
C_out). Stamping all of them across the whole dough gives theC_out-deep output. Kis the cutter's width — how big a bite of context each press takes.
05.Kernels Are Learned, Not Hand-Set
Old computer vision hand-wrote the cutters. Engineers designed operators like the Sobel filter ([[-1,0,1]]) or Gabor filters to detect edges at chosen orientations. They knew what an edge looked like and coded it in.
CNNs do the opposite. They start from random weights and learn the optimal filter values by gradient descent on the task loss. Nobody tells the network what an edge is.
Remarkably, after training on ImageNet, the first-layer filters self-organize into oriented, color-opponent edge detectors that look like Gabor filters — the network rediscovered hand-crafted vision from scratch, purely to reduce its error.
Because the kernel is learned, you almost never set its values. You choose only three things:
- its size (
K), - how many you want (
C_out), - how they slide (stride / padding — covered in the next topics).
06.Choosing K: The Small-Kernel Doctrine
Kernel size is a trade-off between three things: receptive field (how much context one press covers), parameter cost, and depth of non-linearity.
- A
5x5covers a larger area in one step, but costs(25/9) ≈ 2.8xthe parameters of a3x3. - Two stacked
3x3convs see the same5x5region as one5x5conv, but use fewer weight-multiplications (2 * 9 = 18vs25) and get one extra ReLU in between, so they model more non-linearly.
That single insight — same receptive field, fewer parameters, more non-linearity — is the VGG argument that made 3x3 the default.
A quick cheat-sheet of sizes:
- 1x1: "network in network" — mixes channels with no spatial context; a cheap projection.
- 3x3: the workhorse; good locality/parameter balance.
- 5x5 / 7x7: early layers or wide-context needs (AlexNet even opened with an 11x11 first layer).
07.Visual Intuition: Why Two 3x3 Beat One 5x5
Here is the receptive-field picture that makes the doctrine click.
One 5x5 filter, in a single press, looks at a 5x5 patch of the input:
codeone 5x5 conv: two stacked 3x3 convs: ┌─────┐ press 3x3 -> small map │ . . │ sees 5x5 press 3x3 again over THAT │ .X. │ at once .X. now reaches 5x5 of input │ . . │ (because each late cell already └─────┘ pooled a 3x3 of early cells)
Both end up "seeing" the same 5x5 stretch of the original image. But count the work:
- One 5x5:
25weight-multiplications, and zero ReLUs applied in between — it is one straight linear map over that patch. - Two 3x3:
9 + 9 = 18multiplications (cheaper) and a ReLU squeezed in the middle (more non-linear).
Same context, fewer parameters, extra non-linearity. Stack a third 3x3 and you cover 7x7 for only 27 multiplications. That is why deep-and-thin (many 3x3 layers) wins over shallow-and-wide, and why VGG made 3x3 the universal default.
08.In Practice: Multi-Scale Filters Within a Layer
Why commit to one cutter size when you can use several at once?
A single layer can run banks of different sizes in parallel — this is the heart of the Inception module, which applies 1x1, 3x3, and 5x5 convolutions simultaneously and concatenates their maps. The network captures fine detail and broad context in the same stage instead of guessing one receptive field. The output depth then equals the sum of the parallel banks' C_out values.
The other practical lever is parameter thrift. Mobile nets refuse to pay for a full K x K x C_in cutter per channel; depthwise-separable convolution uses one tiny spatial cutter per input channel, then a cheap 1x1 to mix — cutting parameters roughly 8–9x for phone inference. So in real deployments you pick K and C_out not just for accuracy but to hit a latency budget on the chip you ship to.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Tiny parameter count relative to dense layers (a 3x3x3 filter is just 28 numbers).
- Learned filters adapt automatically to whatever features the data and loss demand.
- Reusing one filter across the image gives translation equivariance for free.
Trade-offs & Constraints
- Fixed spatial size limits each filter's direct context to K x K.
- More filters = more capacity but quadratic growth in parameters with K and depth.
- Large kernels (e.g. 11x11) are parameter-heavy and rarely used outside first layers.
Inception stacks parallel 1x1/3x3/5x5 banks to span multiple receptive fields at once. MobileNet instead replaces standard K x K x C_in filters with depthwise-separable ones (one small spatial filter per input channel, then a 1x1 to mix), cutting parameters ~8-9x for mobile inference.
Staff+ Engineering Takeaways
- A filter (kernel) is a learnable K x K x C_in weight tensor; its depth always matches the input's channel count.
- One filter yields one output map; a bank of C_out filters yields a C_out-deep volume.
- Kernel values are learned by gradient descent, rediscovering edge/Gabor-like detectors from data.
- Stacked 3x3 filters are parameter-efficient and more non-linear than a single large K.
- 1x1 convolutions mix channels without spatial context; multi-scale banks (Inception) broaden receptive field.
- Layer parameters = (K * K * C_in + 1) * C_out.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
The spatial size (K) of a conv filter directly controls what?
How clear and actionable was this distributed systems breakdown?