TOPIC #95Beginner 12 min read

Convolutional Neural Networks (CNNs)

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Plain dense networks choke on images because they treat pixels as one giant unordered list. CNNs fix this with small filters that slide across the picture and share their weights everywhere. The result is the Convolution → ReLU → Pooling → Fully-Connected pipeline behind modern computer vision.

Anatomy of a CNN 🔬

A CNN interleaves convolution and pooling stages that progressively shrink spatial size while growing channel depth, then a small dense head maps the final feature vector to class scores.

Anatomy of a CNN 🔬
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: A Plain Network Flattens a Photo Into Soup

A normal feed-forward network (also called a fully-connected or dense layer — every input neuron wires to every output neuron) does not see a picture.

It sees one long, flat list of numbers.

A color photo is a stack of grids: 224 x 224 x 3 pixels — height, width, and three color channels (red, green, blue).

To feed it to a dense layer, you smash that stack into one flat vector of 224 * 224 * 3 = 150,528 numbers.

Now add a single hidden layer of 1,000 neurons.

Each of the 150,528 inputs wires to each of the 1,000 neurons:

150,528 * 1,000 ≈ 150M weights

...before the output layer even starts.

That is computationally brutal. But the bigger crime is that it is statistically naive. It throws away three facts that make vision easy for us and hard for machines:

  • Spatial structure matters. A cat in the top-left and a cat in the bottom-right are the same cat. A flat list has no idea which pixels are neighbors.
  • Patterns repeat. An edge detector that works in one corner should work everywhere. Dense weights are frozen to absolute pixel positions and cannot be reused.
  • Objects are built from parts. Eyes, wheels, and texture patches combine into faces and cars. A flat vector cannot express that composition.

So the question becomes

Insight

How do we look at an image as an image, not as soup?

CNNs answer it with two ideas: give each neuron a small window to look through, and share the same weights across the whole picture.

02.The Idea in Plain Words: A Tiny Window, Slid Everywhere

A CNN is simply

Insight

A small learnable patch of weights that slides across the image and fires wherever its pattern appears.

That patch is called a filter (or kernel). Let's unpack the sentence.

  • Small — a filter covers only a little neighborhood, like 3x3 pixels, not the whole image. This respects spatial structure: each neuron looks at one local spot.
  • Learnable — the numbers inside the filter are tuned by gradient descent, exactly like dense weights. Nobody hand-writes them.
  • Slides — the same filter moves left-to-right, top-to-bottom, like reading a page line by line. This gives you a whole map of answers, one per position.
  • Shares — the same weights are used at every position. This is weight sharing: detect a vertical edge once, detect it everywhere for free.

Compare the two worldviews at one spot:

  • Dense: "pixel #40,217 matters to neuron #9." (position-locked, no reuse)
  • Conv: "this 3x3 window looks like a vertical edge." (reused at every window)

Weight sharing is why a filter is a pattern detector instead of a pixel detector.

03.A Tiny Worked Example: One Filter, Four Positions

Forget 150M weights. Look at one 3x3 filter on a tiny 4x4 grayscale patch.

Suppose the learned filter is a vertical-edge detector:

code
 -1  0  1
 -1  0  1
 -1  0  1

It rewards "dark on the left, light on the right."

Now slide it across the first row of window positions and, at each spot, multiply overlapping numbers and add them up.

  • Over a plain gray region (all 7s): positive and negative columns cancel → output ≈ 0. No edge. Nothing fires.
  • Over a left-dark, right-light boundary: the -1s land on dark, the +1s on light → products add to a large positive number. The edge is present. Filter fires.
  • Over a left-light, right-dark boundary: sign flips → large negative → a reversed edge.

Result: one number per window position, arranged back into a small grid. That grid is the filter's feature map — a heat-map of "where does this pattern appear?"

One more filter (say, horizontal edges) gives a second map. A hundred filters give a hundred maps. That stack of maps is what a CNN layer outputs.

04.Visual Intuition: The Map Shrinks While Depth Grows

Picture the CNN as an hourglass that trades space for meaning.

At the start you have a big, shallow input (a photo). At the end you have a small, deep stack of abstract detectors, then a verdict.

code
  width→                              meaning ↓
  [ 224x224 x 3 ]   photo            (3 channels: RGB)
        |  conv + ReLU + pool
        v
  [ 112x112 x 64 ]  edges/colors     (64 detectors)
        |  conv + ReLU + pool
        v
  [ 56x56 x 128 ]   textures/shapes
        |  conv + ReLU + pool
        v
  [ 28x28 x 256 ]   parts (eyes/wheels)
        |
        v  flatten → fully-connected head
     [ class scores → softmax ]       "cat? dog?"

Notice the two arrows moving in opposite directions:

  • Spatial size H, W shrinks: 224 → 112 → 56 → 28 (pooling/strides chop it down).
  • Channel depth C grows: 3 → 64 → 128 → 256 (more filters per stage).

The total amount of data H * W * C stays roughly constant. The network is not losing information — it is compressing where-things-are into what-things-are.

05.The Analogy: Inspecting a Mural With Reusable Stencils

Carry one picture through the rest of the topic: you are auditing a giant museum mural with a handheld stencil light.

  • The mural is the input image — huge, full of detail.
  • A stencil is a filter: a cut-out shape that only "lights up" when it matches a local pattern (a vertical brushstroke, a spiral, an eye).
  • You slide the stencil across the whole wall, marking each spot where it fits. That marking sheet is the feature map.
  • You own many stencils in a roll (a filter bank): early ones match simple strokes; late ones match complex figures.
  • You share each stencil everywhere — you don't carve a new one per square foot. That's weight sharing.
  • Stepping back after each pass (pooling/downsampling) lets you trade fine position for a truer sense of "this region has a face."

You never memorized all 150k pixels. You just kept asking each small window one question:

Insight

"Does my pattern show up here?"

Everything later in the course — kernels, stride, padding, pooling, receptive fields — is just about which stencils you hold, how far you slide, and how often you step back.

06.The Canonical Building Blocks

Every CNN, from LeNet-5 (1998) to a ResNet-152 (2015), is assembled from a handful of layers, always in this rhythm:

  1. Convolution layer — applies N learnable filters to the input volume, producing N output feature maps. This is where spatial patterns are detected.
  2. ReLU activation — element-wise max(0, x), introducing the non-linearity that lets stacked convs model complex shapes (without it, stacking convs would collapse into one linear map).
  3. Pooling (subsampling) — typically 2x2 max pooling with stride 2, halving width and height to build invariance and save compute.
  4. Fully-connected head — after flattening, a dense layer maps the global features to class logits.

So the repeating stage is Convolution → ReLU → Pooling, and the spatial dimensions shrink layer by layer (224 → 112 → 56 → 28...) while channel depth grows (3 → 64 → 128 → 256...), keeping the feature volume roughly constant in size.

In code, the whole recipe is a few lines:

python— A minimal CNN in PyTorch — two conv blocks then a dense head
import torch.nn as nn

class MiniCNN(nn.Module):
    def __init__(self, n_classes=10):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(3, 32, kernel_size=3, padding=1), nn.ReLU(),
            nn.MaxPool2d(2),                        # 224 -> 112
            nn.Conv2d(32, 64, kernel_size=3, padding=1), nn.ReLU(),
            nn.MaxPool2d(2),                        # 112 -> 56
        )
        self.classifier = nn.Sequential(
            nn.Flatten(),
            nn.Linear(64 * 56 * 56, 256), nn.ReLU(),
            nn.Linear(256, n_classes),
        )

    def forward(self, x):
        return self.classifier(self.features(x))

07.Why AI Cares: The Hierarchy Comes For Free

Here is the payoff of stacking convolutions: automatic feature hierarchy. You never tell the network "look for edges." Visualizing early-layer activations (a tradition since the AlexNet era) shows it discovers this on its own:

  • Layer 1: oriented edges and color blobs — almost hand-designed Gabor filters, but learned purely from data.
  • Layer 2–3: textures, corners, simple shapes assembled from edge responses.
  • Deeper layers: object parts — eyes, wheels, window grids.
  • Final layers: high-level, class-specific whole-object detectors.

Why does a network whose filters are only 3x3 ever "see" a whole face?

Because each layer looks "back" through the ones below it. Layer 5's 3x3 window covers pixels that already summarize wide regions of layer 1. That growing coverage is the receptive field, and it is why depth buys context.

This is the real reason deep CNNs beat old hand-engineered vision (SIFT, HOG features): the machine learns the features themselves.

08.In Practice: The Arithmetic You Use Every Day

Two rules cover most daily CNN wiring and prevent the classic "shape mismatch at the linear layer" bug.

Rule 1 — Output spatial size. For one conv/pool layer with input size W, kernel K, padding P, and stride S:

output = floor((W - K + 2P) / S) + 1

Rule 2 — Output depth. The number of channels is simply the number of filters N in that layer, independent of the formula above.

Worked check: input W = 224, K = 3, P = 1, S = 1 (a "same-size" conv block):

output = floor((224 - 3 + 2*1) / 1) + 1 = floor(223) + 1 = 224

Then a 2x2 max-pool with stride 2 halves it to 112. Repeat, and you get the 224 → 112 → 56 ladder from the hourglass sketch. Keeping these two numbers straight — where the size comes from (the formula) and where the depth comes from (the filter count) — is the single most common CNN whiteboard question.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Weight sharing slashes parameters versus dense layers, making images trainable.
  • Translation equivariance: a learned pattern is detected wherever it appears.
  • Automatic hierarchy — no manual feature engineering (SIFT/HOG) required.

Trade-offs & Constraints

  • Locality biases struggle with long-range dependencies (motivating Transformers).
  • Pooling and striding discard fine spatial detail; dense prediction needs upsampling.
  • Deep stacks need careful normalization/residuals to train stably.
Production Implementation in Big Tech
Tesla Autopilot / Waymo• Real-time visual perception

Onboard CNN backbones (evolving HydraNet / BEV-style ResNet ensembles) extract features from multi-camera streams; the hierarchical features feed detection and occupancy networks that must run at 36+ FPS on embedded accelerators — driving the choice of compact conv blocks and aggressive downsampling.

Staff+ Engineering Takeaways

  • CNNs replace flat dense layers with local, weight-shared filters that respect image geometry.
  • The standard stage is Convolution → ReLU → Pooling; depth grows as spatial size shrinks.
  • Stacked convs learn an automatic hierarchy: edges → textures → parts → objects.
  • Output size follows floor((W - K + 2P)/S) + 1; output depth equals the number of filters.
  • Weight sharing is what makes training on high-dimensional pixels computationally feasible.
  • Locality is a strength for images but a limitation for global relationships, which Transformers address.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

What is the PRIMARY reason fully-connected layers are impractical for raw images?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?