Stride
Stride is how many pixels the filter jumps between placements. Stride 1 slides one pixel and keeps full size; stride 2 hops two pixels and roughly halves the output map — downsampling built right into the convolution. This topic covers the stride arithmetic, the stride-vs-pooling trade-off, and the checkerboard artifacts careless striding can cause.
Stride sets the slide step 📐
With stride 1 the window moves one pixel at a time and the output stays large; with stride 2 it jumps two pixels, halving the output width and height.
01.The Problem: Stopping at Every Single Pixel Is Slow
So far our filter has been sliding one pixel at a time.
On a 224x224 image that means about 50,000 window placements per filter. Doable, but wasteful when all we really need is a coarser sense of "is there an edge around here?"
Two desires pull against sliding so finely:
- We want smaller maps deeper in the network (less memory, less compute, wider context per cell).
- We want to shrink without throwing away a separate pooling stage.
So the question becomes
Can the filter take bigger steps as it slides?
Yes. That step size is stride — and it is the simplest downsampling knob you have.
02.The Idea in Plain Words: Stride = How Far the Window Jumps
Stride (S) is simply
The number of pixels the filter jumps between successive placements, horizontally and vertically.
S = 1→ slide one pixel (dense, overlapping windows, output stays big).S = 2→ hop two pixels, skipping every other position (output roughly halves).S = 3→ hop three, and so on.
Stride is one of the four convolution hyper-parameters: kernel size, padding, stride, dilation. You pick it; it is not learned.
What makes it special is that stride downsamples inside the convolution itself. You do not need a separate pooling layer to shrink the map — a stride-2 conv shrinks as it filters. Fewer layers, one fused operation.
Quick vocabulary so the next topics click:
- A conv with
S = 1is sometimes called non-strided or full-resolution. S = 2is a halving conv (each side cuts in half, so the area cuts to a quarter).- The cousin operation, dilation, also enlarges what a filter "sees" — but instead of skipping outputs like stride, it spreads the filter taps apart while still emitting at every position. Keep the two straight: stride throws away output cells; dilation stretches the window and keeps them all.
03.A Tiny Worked Example: 5x5 Input, 3x3 Filter
Take a 5x5 input, a 3x3 filter, S = 1, no padding.
S = 1: the window's top-left can sit at columns 0,1,2 →3columns →3x3output. (Formula check below.)
Now make it S = 2:
- The window sits at column 0, then jumps to column 2, then column 4? At column 4 the
3x3window needs columns 4,5,6 — but the image ends at 4. So the last placement is dropped. - Valid positions: 0 and 2 →
2columns →2x2output. Halved.
Let the formula confirm it, then we use it for the big real numbers.
One more useful habit: the same rule runs independently on height and width. A square stride applies the identical jump vertically, so a 5x5 with S = 2 becomes 2x2 on both axes — the area drops from 25 cells to 4. That area-collapse (÷S²) is why stride is such an effective compute saver: each doubling of stride removes roughly three-quarters of the downstream work.
04.The Output-Size Formula
For input size W (and H), kernel K, padding P, stride S, the output spatial size is:
codeH_out = floor((H_in - K + 2P) / S) + 1 W_out = floor((W_in - K + 2P) / S) + 1
The key thing to see: stride sits in the denominator. Doubling S roughly halves the output dimensions (exactly, when things divide evenly).
Re-doing the tiny example with it: floor((5 - 3 + 0)/2) + 1 = floor(2/2) + 1 = 1 + 1 = 2 → 2x2. Matches.
05.Visual Intuition: Stepping Stones Across a Stream
Picture the input as a row of stepping stones and the filter as your boot.
codestride 1 (every stone): [1][2][3][4][5][6][7][8] you step on: ^ ^ ^ ^ ^ ^ ^ ^ -> 8 outputs stride 2 (skip one): [1][2][3][4][5][6][7][8] you step on: ^ ^ ^ ^ -> 4 outputs skipped: 2 4 6 8 (gone)
Bigger strides = fewer, more spread-out readings. You cover the whole stream with half the steps — but you have thrown away everything that lived on the stones you skipped. That lost fine detail is exactly the resolution you pay for the speed.
06.The Analogy: A Gardener Watering in Longer Strides
Imagine watering a long garden bed with a watering can that soaks a small patch each time you set it down.
- Stride 1 — you set the can down on every spot. Perfect coverage, but you walk the whole length dozens of times. Slow.
- Stride 2 — you set it down, hop over one patch, set down again. Half the steps, twice the speed. But the skipped patches get less water — a thin sprout between your stops can be missed entirely.
Stride is that hop. It saves effort, and the price is resolution: details that fall in the gaps may not be captured. For classification you rarely care. For tasks that must point at an exact pixel (segmentation, detection), you care a lot — so you keep stride low, or recover detail elsewhere.
07.Worked Examples on Real Numbers
Take a 224 x 224 input, kernel 3, padding 1 (a standard "same-size" block):
- S = 1:
floor((224 - 3 + 2)/1) + 1 = 224→ full resolution preserved (VGG "same-res" blocks). - S = 2:
floor((224 - 3 + 2)/2) + 1 = floor(223/2)+1 = 112→ halved. This is how ResNet'sconv1(stride 2 at 7x7) and each downsampling stage cut spatial size without pooling. - A
1x1conv withS=2:floor((224-1+0)/2)+1 = 112— 1x1 convs with stride are an efficient channel-reduce + downsample combo.
import torch, torch.nn as nn
x = torch.randn(1, 64, 224, 224)
# Option A: strided convolution (downsample inside the filter)
strided = nn.Conv2d(64, 128, kernel_size=3, stride=2, padding=1)(x)
print("strided conv:", tuple(strided.shape)) # (1, 128, 112, 112)
# Option B: regular conv then max pooling
pool = nn.MaxPool2d(2)(nn.Conv2d(64, 128, kernel_size=3, padding=1, stride=1)(x))
print("conv+pool :", tuple(pool.shape)) # (1, 128, 112, 112)08.Stride > 1 vs Pooling: Why the Field Moved
Historically, downsampling was done with a dedicated 2x2 max-pool after each conv. Modern architectures increasingly fold the stride into the convolution and drop pooling entirely (ResNet, EfficientNet). Why the shift?
- Fewer layers, same compute. A stride-2 conv does filtering and downsampling in one op.
- Learnable downsampling. The conv weights choose what to keep; max-pool keeps only the maximum, which can discard complementary information.
- Better GPU utilization. One fused kernel instead of a conv plus a separate pool memory round-trip.
The catch: strided convs (and especially transposed convolutions used for upsampling) can cause aliasing / information loss, because no smoothing is applied before you start skipping positions. Pooling-with-anti-aliasing or blur-pool alternatives mitigate this in image-generation models.
09.In Practice: The Checkerboard Artifacts Pitfall
Subsampling without low-pass filtering is a textbook violation of the Nyquist theorem: high-frequency detail "folds" into false low-frequency content.
With strided convs — and most visibly with transposed convs used to upsample — those fold-over artifacts show up as grid-like checkerboard patterns in the output image: faint squares that have nothing to do with the real content.
The standard fix (Odeniya et al., 2016) is one of two moves:
- Resize with nearest-neighbor / bilinear interpolation first, then a regular stride-1 conv, instead of a transposed-conv with stride 2.
- Or choose kernel sizes divisible by the stride, so windows overlap evenly and no position is systematically favored.
This matters most for segmentation masks and generative decoders, where a grid artifact is glaring. It is the direct consequence of the gardener lesson: skip too coarsely without smoothing, and the gaps you leave become fake structure.
The one-line interview rule that ties it together:
Stride is powerful, cheap downsampling — but smooth first, then skip. Any time you reduce resolution (striding down or transposed-conv up), apply a low-pass filter around the subsample step and the checkerboards vanish. That single principle (anti-aliasing) is why modern image generators replaced transposed convs with "resize + conv" blocks.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Single op performs spatial reduction and feature extraction — fewer layers.
- Learns what to downsample rather than blindly max-pooling.
- Reduces activation memory and compute in deeper stages.
Trade-offs & Constraints
- Discards spatial resolution — bad for tasks needing exact localization.
- Naive striding can alias, producing checkerboard artifacts (especially upsampling).
- Hard to choose K/S combos cleanly; non-divisible sizes cause truncation bugs.
ResNet and EfficientNet halve spatial size at each stage via stride-2 (depthwise) convolutions instead of pooling layers. This keeps latency predictable on mobile NPUs and avoids pooling's information discarding — critical when input resolution is already small.
Staff+ Engineering Takeaways
- Stride is the pixel step between filter applications; it downsamples inside the conv itself.
- Output size = floor((W_in - K + 2P) / S) + 1; S sits in the denominator.
- Strided convs are the modern replacement for max-pooling: fewer layers, learnable, GPU-friendly.
- Without anti-aliasing, stride (esp. transposed conv) aliases into checkerboard artifacts.
- Stride trades spatial resolution for compute savings — keep it low for localization-heavy tasks.
- Stride is chosen, not learned; combine carefully with kernel size for divisible shapes.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
A 7x7 kernel, stride 2, padding 3 applied to a 224x224 image produces output of size:
How clear and actionable was this distributed systems breakdown?