Convolutional Neural Networks & Vision
How machines actually see:
All Topics in Phase 5
0 of 15 completedPlain dense networks choke on images because they treat pixels as one giant unordered list. CNNs fix this with small filters that slide across the picture and share their weights everywhere. The result is the Convolution → ReLU → Pooling → Fully-Connected pipeline behind modern computer vision.
Under the hood, a convolution is just a sliding dot product: a learnable filter is laid over a small window, the overlapping numbers are multiplied and summed into one score, and the window slides on. This topic walks that arithmetic, shows how 3D filters span color channels, and clears up the cross-correlation naming nuance PyTorch actually performs.
The filter (or kernel) is the learnable heart of a convolution: a small 3D block of weights that acts as a sliding pattern detector. This topic nails down filter shape (K x K x C_in), the filter-vs-bank distinction, how to count its parameters, and why tiny learned 3x3 kernels beat big hand-written ones.
A conv layer does not output a single number — it outputs a stack of 2D score sheets, one per filter, called feature maps, which together form an H x W x C activation volume. This topic defines pre- vs post-activation maps, the tensor shape (NCHW vs NHWC), and how map size shrinks while depth grows through a network.
Stride is how many pixels the filter jumps between placements. Stride 1 slides one pixel and keeps full size; stride 2 hops two pixels and roughly halves the output map — downsampling built right into the convolution. This topic covers the stride arithmetic, the stride-vs-pooling trade-off, and the checkerboard artifacts careless striding can cause.
Padding adds a synthetic border of made-up values around the input so the filter can cover edge pixels and you can control output size. "Same" padding keeps the spatial dimensions (at stride 1); "Valid" adds no border and lets the map shrink. This topic derives the padding amounts and the corner-handling subtleties.
Pooling shrinks each feature map on its own using a fixed little operator — max keeps the strongest value in a window, average keeps the mean. It has no learnable parameters and never mixes channels. This topic compares the two, derives output sizes, and explains why classic local pooling is being replaced by strided convolutions and Global Average Pooling.
A single neuron in a CNN only looks at a tiny window of the image — so how does the network ever "see" a whole bus? The answer is the receptive field: the region of the input that can influence a neuron. Stacking small convolutions widens that window layer by layer, and this topic derives the math, exposes the effective-receptive-field gap, and connects to dilated and multi-scale designs.
How did neural networks go from reading handwritten digits to beating humans at naming anything in a photo? This is the story of five architectures — LeNet-5, AlexNet, VGG, ResNet, EfficientNet — each of which won a round of the ImageNet competition by changing one big thing. We cover what each changed, why it mattered, and which design lessons still apply in 2024-2026.
Common sense said deeper networks are better — until deeper networks started getting worse. The fix was beautifully simple: let each layer ADD a small correction to its input instead of rewriting it from scratch (y = F(x) + x). That one wiring change cured the degradation problem, gave gradients a shortcut highway, and now sits inside every Transformer you have ever used.
You have 2,000 images and no hope of training a big network from scratch. The fix: start from a network that already studied 1.28 million — its early layers (edges, textures) work for any image, so you only train the part that knows YOUR classes. This topic covers feature extraction vs fine-tuning, freezing, discriminative learning rates, and modern self-supervised / vision-language pretraining.
A model trained on 5,000 exact images memorizes them and fails on the 5,001st. Augmentation fixes this by showing the network a fresh, randomly transformed copy every time — flipped, brightened, cropped, even blended with a neighbor — as long as the label stays true. This topic covers the standard operator menu, label-aware pipelines for detection, and the mix-style tricks (Cutout, Mixup, CutMix) that power modern training.
Classification says "there is a cat." A self-driving car needs "cat, 84% sure, 340 pixels right, 120 up, 90 wide, 60 tall." Object detection answers where every object is AND what it is, as a set of boxes. This topic contrasts two-stage (R-CNN family) and one-stage (YOLO/SSD) detectors, builds up proposals, anchors, IoU, confidence and Non-Maximum Suppression, and surveys the 2024-2026 anchor-free and Transformer (DETR) landscape.
A bounding box around a pedestrian still includes slices of sidewalk. Safety-critical systems need the opposite: a class label for every single pixel — road here, car there, sky above. That is semantic segmentation. This topic builds the fully-convolutional idea, the encoder-decoder U-Net with skip connections, dilation/ASPP for context at full resolution, and the shift to promptable foundation models like SAM.
Transformers ran language. In 2020, ViT asked: what if an image were just another sentence — 196 patch-words read by standard self-attention, with no convolutions at all? This topic covers patch embedding, the [CLS] token and positional embeddings, how attention replaces the CNN receptive field (instant global context, quadratic cost), the data-hungry pretraining caveat, and ViT's role as the backbone of CLIP, DINOv2, and SAM.