TOPIC #100Beginner 11 min read

Padding: Same vs Valid

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Padding adds a synthetic border of made-up values around the input so the filter can cover edge pixels and you can control output size. "Same" padding keeps the spatial dimensions (at stride 1); "Valid" adds no border and lets the map shrink. This topic derives the padding amounts and the corner-handling subtleties.

Valid shrinks, Same preserves 🧱

Valid convolutions place the kernel only where it fits fully inside the input, shrinking the map. Same convolutions add a border (usually zeros) of size that, at stride 1, keeps the output equal to the input dimensions.

Valid shrinks, Same preserves 🧱
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Corners and Edges Get Ignored

A 3x3 filter is only fully-on-the-image when its center sits at least one pixel in from every border.

Slide it to a corner or an edge and part of the window hangs off into nothing.

So without special handling, a corner pixel can be a window center only once, while a central pixel gets visited K x K times.

Two ugly consequences pile up:

  • Edges are under-learned. The outer ⌊K/2⌋ pixels barely contribute, so the network learns less about the border — exactly where object boundaries often live.
  • The map shrinks every layer. A 3x3 valid conv eats 2 from each side: 224 → 222 → 220…. Stack enough layers and most of the image is gone.

So the question becomes

Insight

How do we let the filter treat edge pixels normally, and control how fast the map shrinks?

The answer is padding: draw a fake border around the input before sliding.

02.The Idea in Plain Words: A Border of Made-Up Pixels

Padding (P) is simply

Insight

A ring of synthetic values pasted around the input, so the filter has something to read when it reaches the edge.

Now the window can be centered on the border pixels, because their neighbors are no longer missing — they are padded values.

You choose two things:

  • How thick the border is (P pixels per side), and
  • What fills it (usually 0, sometimes a reflection of the real edge).

With the right thickness, you can make the output come out the same size as the input, which is the whole point of "Same" padding. Pad with nothing and you get "Valid" padding, where the map shrinks.

03.A Tiny Worked Example: 5x5 With a 3x3 Filter

A 5x5 input, a 3x3 filter, stride 1.

No padding (Valid):

  • The window center can only sit on the inner 3x3 of positions.
  • Output = 5 - 3 + 1 = 3 → a 3x3 map. The four real borders were barely visited.

Pad with one ring of zeros (Same):

  • Now the input is effectively 7x7.
  • The window center can sit on all 5x5 original positions.
  • Output = 7 - 3 + 1 = 5 → a 5x5 map. Size preserved, and every original edge pixel got centered like any other.

That P = 1 for a 3x3 is the general rule P = (K-1)/2. For K = 5 you would pad 2; for K = 7, pad 3.

Notice what changed: with padding, all 25 original positions became window centers and got inspected; without it, only the inner 9 did. Padding is not adding information — it is spending a little compute on a fake border so the real edges are treated fairly and the map keeps its size.

04.Visual Intuition: Matting a Photo in a Frame

Think of hanging a photo in a frame with a cardboard mat (the border around a picture).

code
   no padding (Valid):    [ . . . . . ]      the cutter only
                          [ . ■ ■ ■ . ]      fits fully inside
   pad 1 (Same):          [ . . . . . ]      the mat lets the
   0 0 0 0 0 0 0          [ ■ ■ ■ ■ ■ ]      cutter sit on the
   0 [ . . . . . ] 0      [ ■ ■ ■ ■ ■ ]      real edge too
   0 [ . . . . . ] 0      [ ■ ■ ■ ■ ■ ]
   0 [ . . . . . ] 0      [ ■ ■ ■ ■ ■ ]
   0 0 0 0 0 0 0          [ ■ ■ ■ ■ ■ ]
                          ■ = output cell    0 = fake mat

The mat (padded zeros) is not "real" image content — but it gives the filter something uniform to chew on at the boundary, so the photo's true edges are no longer special cases.

05.The Analogy: Tiling a Floor Without Hanging Off the Edge

Imagine a tile-cutter that must rest on a 3x3 patch of floor to score one tile.

  • Near the middle, easy: the cutter always has floor under it.
  • At the wall, the cutter would hang off into the baseboard — positions that "don't exist."

You have two strategies:

  • Valid: only cut where the whole `3x3) fits inside the real floor. You skip a band along every wall. Fewer cuts, and the border tiles never get inspected.
  • Same (pad): snap a cheap strip of scrap flooring around the room's edge. Now the cutter fits everywhere, including against the walls, and every real tile — even the border ones — gets its proper look.

The scrap strip is zero-padding. It is fake, but it lets the machinery run uniformly over the whole floor, edges included.

06.Valid Padding (a.k.a. "No Padding")

Valid means the filter is applied only at positions where it lies fully inside the input — the most "valid" placements. In the size formula P = 0, so with stride 1:

code
H_out = H_in - K + 1
W_out = W_in - K + 1

A 5x5 input convolved 3x3 valid → 3x3 output, exactly as the worked example showed. Each valid layer eats K-1 from each dimension, so valid-only architectures shrink rapidly.

Valid is still chosen deliberately when you cannot afford border pollution — for instance the early AlexNet/LeNet stages, or when you want to convert a fixed input into a smaller feature region.

python— Computing the Same-padding amount for a kernel
# For odd kernels with stride 1, "same" padding keeps size:
def same_pad(kernel):
    return kernel // 2          # e.g. K=3 -> 1, K=5 -> 2, K=7 -> 3

# For even kernels (e.g. K=2) the split is asymmetric: total pad = K-1.
def same_pad_even(kernel):
    left  = (kernel - 1) // 2
    right = (kernel - 1) - left
    return left, right

print(same_pad(3), same_pad(7))   # 1 3
print(same_pad_even(2))            # (0, 1)  # total 1

07.Same Padding: Making Output Equal Input

Same padding adds enough border that, at stride 1, H_out = H_in and W_out = W_in. For an odd kernel K, the padding per side is P = (K - 1) / 2 (both sides padded equally).

Prove it against the output formula — plug P = (K-1)/2 and S = 1:

code
H_out = floor((H_in - K + 2P)/1) + 1
      = H_in - K + 2*((K-1)/2) + 1
      = H_in - K + (K-1) + 1 = H_in   ✓

The -K and +K cancel, leaving the input size untouched. That cancellation is the reason the rule is (K-1)/2 and not some other number.

PyTorch's padding='same' does this automatically (and even handles even kernels with the asymmetric left/right split shown in the code block). When stride is 2, "same" typically means ceil(H_in/S) — an exact halving — which is what makes it the natural companion to stride-based downsampling.

08.In Practice: Choosing Between Same and Valid

Practical guidance from production backbones:

  • Default to same-padding for the bulk of conv stages so intermediate maps line up spatially. This is mandatory for residual connections (the identity shortcut and the conv branch must be the same size to add elementwise) and for segmentation, where the feature grid must align pixel-for-pixel with the input grid.
  • Use valid sparingly when you want intentional shrinkage with a large kernel, or when the border carries no meaning and zero-padding would inject false structure — a black frame around an all-bright medical scan is actively misleading; use reflect/replicate there.
  • Watch the corner budget. With small inputs and many layers, even same-padding means true corner pixels receive fewer real contributions; the padding only pretends they do. The genuine global context arrives later, from large receptive fields or attention — not from the border trick.

The interview one-liner: Same keeps the grid intact so things align; Valid honestly reports only fully-covered windows. Knowing which alignment you need (residuals, skip connections, dense labels) picks the answer for you.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Same padding preserves spatial size → clean residual/dense alignment and stable dimensions.
  • Border pixels stop being under-visited; edges are detected like the interior.
  • Pairs with stride-2 to yield exact spatial halving.

Trade-offs & Constraints

  • Zero-padding injects artificial (black) borders that may mislead on high-contrast edges.
  • Even kernels need asymmetric padding — a subtle framework difference.
  • Same keeps more activations alive → larger memory footprint than valid.
Production Implementation in Big Tech
U-Net ( biomedical segmentation )• Keeping encoder/decoder grids aligned

U-Net encodes with same-padded convs so each downsampled feature map stays spatially consistent with its skip-connection counterpart from the encoder, enabling clean concatenation during upsampling. The original paper alternatively used valid convs and cropped borders to avoid zero-padding artifacts — a deliberate trade-off.

Staff+ Engineering Takeaways

  • Padding surrounds the input with synthetic values so filters can cover borders and control output size.
  • Valid (P=0) shrinks output by K-1 per dimension; Same preserves dimensions at stride 1.
  • Same padding amount = (K-1)/2 per side for odd kernels; even kernels split the padding asymmetrically.
  • Frameworks default to zero-padding; reflect/replicate avoid injecting false borders.
  • Same-padding is required for residual shortcuts and dense prediction alignment.
  • With stride, "same" typically means ceil(H/S), enabling exact spatial halving.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

For a 5x5 kernel with stride 1, how much zero padding per side is needed to keep a same-size output?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?

Related Concepts & Cross-References

Indexed from curriculum