Padding: Same vs Valid
Padding adds a synthetic border of made-up values around the input so the filter can cover edge pixels and you can control output size. "Same" padding keeps the spatial dimensions (at stride 1); "Valid" adds no border and lets the map shrink. This topic derives the padding amounts and the corner-handling subtleties.
Valid shrinks, Same preserves 🧱
Valid convolutions place the kernel only where it fits fully inside the input, shrinking the map. Same convolutions add a border (usually zeros) of size that, at stride 1, keeps the output equal to the input dimensions.
01.The Problem: Corners and Edges Get Ignored
A 3x3 filter is only fully-on-the-image when its center sits at least one pixel in from every border.
Slide it to a corner or an edge and part of the window hangs off into nothing.
So without special handling, a corner pixel can be a window center only once, while a central pixel gets visited K x K times.
Two ugly consequences pile up:
- Edges are under-learned. The outer
⌊K/2⌋pixels barely contribute, so the network learns less about the border — exactly where object boundaries often live. - The map shrinks every layer. A
3x3valid conv eats 2 from each side:224 → 222 → 220…. Stack enough layers and most of the image is gone.
So the question becomes
How do we let the filter treat edge pixels normally, and control how fast the map shrinks?
The answer is padding: draw a fake border around the input before sliding.
02.The Idea in Plain Words: A Border of Made-Up Pixels
Padding (P) is simply
A ring of synthetic values pasted around the input, so the filter has something to read when it reaches the edge.
Now the window can be centered on the border pixels, because their neighbors are no longer missing — they are padded values.
You choose two things:
- How thick the border is (
Ppixels per side), and - What fills it (usually
0, sometimes a reflection of the real edge).
With the right thickness, you can make the output come out the same size as the input, which is the whole point of "Same" padding. Pad with nothing and you get "Valid" padding, where the map shrinks.
03.A Tiny Worked Example: 5x5 With a 3x3 Filter
A 5x5 input, a 3x3 filter, stride 1.
No padding (Valid):
- The window center can only sit on the inner
3x3of positions. - Output =
5 - 3 + 1 = 3→ a3x3map. The four real borders were barely visited.
Pad with one ring of zeros (Same):
- Now the input is effectively
7x7. - The window center can sit on all
5x5original positions. - Output =
7 - 3 + 1 = 5→ a5x5map. Size preserved, and every original edge pixel got centered like any other.
That P = 1 for a 3x3 is the general rule P = (K-1)/2. For K = 5 you would pad 2; for K = 7, pad 3.
Notice what changed: with padding, all 25 original positions became window centers and got inspected; without it, only the inner 9 did. Padding is not adding information — it is spending a little compute on a fake border so the real edges are treated fairly and the map keeps its size.
04.Visual Intuition: Matting a Photo in a Frame
Think of hanging a photo in a frame with a cardboard mat (the border around a picture).
codeno padding (Valid): [ . . . . . ] the cutter only [ . ■ ■ ■ . ] fits fully inside pad 1 (Same): [ . . . . . ] the mat lets the 0 0 0 0 0 0 0 [ ■ ■ ■ ■ ■ ] cutter sit on the 0 [ . . . . . ] 0 [ ■ ■ ■ ■ ■ ] real edge too 0 [ . . . . . ] 0 [ ■ ■ ■ ■ ■ ] 0 [ . . . . . ] 0 [ ■ ■ ■ ■ ■ ] 0 0 0 0 0 0 0 [ ■ ■ ■ ■ ■ ] ■ = output cell 0 = fake mat
The mat (padded zeros) is not "real" image content — but it gives the filter something uniform to chew on at the boundary, so the photo's true edges are no longer special cases.
05.The Analogy: Tiling a Floor Without Hanging Off the Edge
Imagine a tile-cutter that must rest on a 3x3 patch of floor to score one tile.
- Near the middle, easy: the cutter always has floor under it.
- At the wall, the cutter would hang off into the baseboard — positions that "don't exist."
You have two strategies:
- Valid: only cut where the whole `3x3) fits inside the real floor. You skip a band along every wall. Fewer cuts, and the border tiles never get inspected.
- Same (pad): snap a cheap strip of scrap flooring around the room's edge. Now the cutter fits everywhere, including against the walls, and every real tile — even the border ones — gets its proper look.
The scrap strip is zero-padding. It is fake, but it lets the machinery run uniformly over the whole floor, edges included.
06.Valid Padding (a.k.a. "No Padding")
Valid means the filter is applied only at positions where it lies fully inside the input — the most "valid" placements. In the size formula P = 0, so with stride 1:
codeH_out = H_in - K + 1 W_out = W_in - K + 1
A 5x5 input convolved 3x3 valid → 3x3 output, exactly as the worked example showed. Each valid layer eats K-1 from each dimension, so valid-only architectures shrink rapidly.
Valid is still chosen deliberately when you cannot afford border pollution — for instance the early AlexNet/LeNet stages, or when you want to convert a fixed input into a smaller feature region.
# For odd kernels with stride 1, "same" padding keeps size:
def same_pad(kernel):
return kernel // 2 # e.g. K=3 -> 1, K=5 -> 2, K=7 -> 3
# For even kernels (e.g. K=2) the split is asymmetric: total pad = K-1.
def same_pad_even(kernel):
left = (kernel - 1) // 2
right = (kernel - 1) - left
return left, right
print(same_pad(3), same_pad(7)) # 1 3
print(same_pad_even(2)) # (0, 1) # total 107.Same Padding: Making Output Equal Input
Same padding adds enough border that, at stride 1, H_out = H_in and W_out = W_in. For an odd kernel K, the padding per side is P = (K - 1) / 2 (both sides padded equally).
Prove it against the output formula — plug P = (K-1)/2 and S = 1:
codeH_out = floor((H_in - K + 2P)/1) + 1 = H_in - K + 2*((K-1)/2) + 1 = H_in - K + (K-1) + 1 = H_in ✓
The -K and +K cancel, leaving the input size untouched. That cancellation is the reason the rule is (K-1)/2 and not some other number.
PyTorch's padding='same' does this automatically (and even handles even kernels with the asymmetric left/right split shown in the code block). When stride is 2, "same" typically means ceil(H_in/S) — an exact halving — which is what makes it the natural companion to stride-based downsampling.
08.In Practice: Choosing Between Same and Valid
Practical guidance from production backbones:
- Default to same-padding for the bulk of conv stages so intermediate maps line up spatially. This is mandatory for residual connections (the identity shortcut and the conv branch must be the same size to add elementwise) and for segmentation, where the feature grid must align pixel-for-pixel with the input grid.
- Use valid sparingly when you want intentional shrinkage with a large kernel, or when the border carries no meaning and zero-padding would inject false structure — a black frame around an all-bright medical scan is actively misleading; use
reflect/replicatethere. - Watch the corner budget. With small inputs and many layers, even same-padding means true corner pixels receive fewer real contributions; the padding only pretends they do. The genuine global context arrives later, from large receptive fields or attention — not from the border trick.
The interview one-liner: Same keeps the grid intact so things align; Valid honestly reports only fully-covered windows. Knowing which alignment you need (residuals, skip connections, dense labels) picks the answer for you.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Same padding preserves spatial size → clean residual/dense alignment and stable dimensions.
- Border pixels stop being under-visited; edges are detected like the interior.
- Pairs with stride-2 to yield exact spatial halving.
Trade-offs & Constraints
- Zero-padding injects artificial (black) borders that may mislead on high-contrast edges.
- Even kernels need asymmetric padding — a subtle framework difference.
- Same keeps more activations alive → larger memory footprint than valid.
U-Net encodes with same-padded convs so each downsampled feature map stays spatially consistent with its skip-connection counterpart from the encoder, enabling clean concatenation during upsampling. The original paper alternatively used valid convs and cropped borders to avoid zero-padding artifacts — a deliberate trade-off.
Staff+ Engineering Takeaways
- Padding surrounds the input with synthetic values so filters can cover borders and control output size.
- Valid (P=0) shrinks output by K-1 per dimension; Same preserves dimensions at stride 1.
- Same padding amount = (K-1)/2 per side for odd kernels; even kernels split the padding asymmetrically.
- Frameworks default to zero-padding; reflect/replicate avoid injecting false borders.
- Same-padding is required for residual shortcuts and dense prediction alignment.
- With stride, "same" typically means ceil(H/S), enabling exact spatial halving.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
For a 5x5 kernel with stride 1, how much zero padding per side is needed to keep a same-size output?
How clear and actionable was this distributed systems breakdown?