PHASE 5 CURRICULUM

Convolutional Neural Networks & Vision

Progress0 of 15 (0%)

How machines actually see:

Key Architectural Domains & Syllabus
convolution and pooling mechanics
receptive fields
classic architectures (LeNet, AlexNet, VGG, ResNet, EfficientNet)
transfer learning for vision
data augmentation
object detection (YOLO family)
semantic and instance segmentation
image generation foundations
the Vision Transformer that later fused vision with the transformer lineage
15 In-Depth Topics ~120 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 5

0 of 15 completed

Plain dense networks choke on images because they treat pixels as one giant unordered list. CNNs fix this with small filters that slide across the picture and share their weights everywhere. The result is the Convolution → ReLU → Pooling → Fully-Connected pipeline behind modern computer vision.

12 min read•3 Quiz Questions
#96The Convolution OperationIntermediateFREE

Under the hood, a convolution is just a sliding dot product: a learnable filter is laid over a small window, the overlapping numbers are multiplied and summed into one score, and the window slides on. This topic walks that arithmetic, shows how 3D filters span color channels, and clears up the cross-correlation naming nuance PyTorch actually performs.

11 min read•3 Quiz Questions
#97Filters (Kernels)BeginnerFREE

The filter (or kernel) is the learnable heart of a convolution: a small 3D block of weights that acts as a sliding pattern detector. This topic nails down filter shape (K x K x C_in), the filter-vs-bank distinction, how to count its parameters, and why tiny learned 3x3 kernels beat big hand-written ones.

11 min read•3 Quiz Questions

A conv layer does not output a single number — it outputs a stack of 2D score sheets, one per filter, called feature maps, which together form an H x W x C activation volume. This topic defines pre- vs post-activation maps, the tensor shape (NCHW vs NHWC), and how map size shrinks while depth grows through a network.

11 min read•3 Quiz Questions
#99StrideBeginnerFREE

Stride is how many pixels the filter jumps between placements. Stride 1 slides one pixel and keeps full size; stride 2 hops two pixels and roughly halves the output map — downsampling built right into the convolution. This topic covers the stride arithmetic, the stride-vs-pooling trade-off, and the checkerboard artifacts careless striding can cause.

11 min read•3 Quiz Questions
#100Padding: Same vs ValidBeginnerFREE

Padding adds a synthetic border of made-up values around the input so the filter can cover edge pixels and you can control output size. "Same" padding keeps the spatial dimensions (at stride 1); "Valid" adds no border and lets the map shrink. This topic derives the padding amounts and the corner-handling subtleties.

11 min read•3 Quiz Questions
#101Max & Average PoolingBeginner PRO

Pooling shrinks each feature map on its own using a fixed little operator — max keeps the strongest value in a window, average keeps the mean. It has no learnable parameters and never mixes channels. This topic compares the two, derives output sizes, and explains why classic local pooling is being replaced by strided convolutions and Global Average Pooling.

11 min read•3 Quiz Questions
#102Receptive FieldIntermediate PRO

A single neuron in a CNN only looks at a tiny window of the image — so how does the network ever "see" a whole bus? The answer is the receptive field: the region of the input that can influence a neuron. Stacking small convolutions widens that window layer by layer, and this topic derives the math, exposes the effective-receptive-field gap, and connects to dilated and multi-scale designs.

11 min read•3 Quiz Questions

How did neural networks go from reading handwritten digits to beating humans at naming anything in a photo? This is the story of five architectures — LeNet-5, AlexNet, VGG, ResNet, EfficientNet — each of which won a round of the ImageNet competition by changing one big thing. We cover what each changed, why it mattered, and which design lessons still apply in 2024-2026.

12 min read•3 Quiz Questions
#104Skip and Residual ConnectionsIntermediate PRO

Common sense said deeper networks are better — until deeper networks started getting worse. The fix was beautifully simple: let each layer ADD a small correction to its input instead of rewriting it from scratch (y = F(x) + x). That one wiring change cured the degradation problem, gave gradients a shortcut highway, and now sits inside every Transformer you have ever used.

11 min read•3 Quiz Questions
#105Transfer LearningIntermediate PRO

You have 2,000 images and no hope of training a big network from scratch. The fix: start from a network that already studied 1.28 million — its early layers (edges, textures) work for any image, so you only train the part that knows YOUR classes. This topic covers feature extraction vs fine-tuning, freezing, discriminative learning rates, and modern self-supervised / vision-language pretraining.

11 min read•3 Quiz Questions
#106Image AugmentationBeginner PRO

A model trained on 5,000 exact images memorizes them and fails on the 5,001st. Augmentation fixes this by showing the network a fresh, randomly transformed copy every time — flipped, brightened, cropped, even blended with a neighbor — as long as the label stays true. This topic covers the standard operator menu, label-aware pipelines for detection, and the mix-style tricks (Cutout, Mixup, CutMix) that power modern training.

11 min read•3 Quiz Questions

Classification says "there is a cat." A self-driving car needs "cat, 84% sure, 340 pixels right, 120 up, 90 wide, 60 tall." Object detection answers where every object is AND what it is, as a set of boxes. This topic contrasts two-stage (R-CNN family) and one-stage (YOLO/SSD) detectors, builds up proposals, anchors, IoU, confidence and Non-Maximum Suppression, and surveys the 2024-2026 anchor-free and Transformer (DETR) landscape.

12 min read•3 Quiz Questions
#108Semantic SegmentationAdvanced PRO

A bounding box around a pedestrian still includes slices of sidewalk. Safety-critical systems need the opposite: a class label for every single pixel — road here, car there, sky above. That is semantic segmentation. This topic builds the fully-convolutional idea, the encoder-decoder U-Net with skip connections, dilation/ASPP for context at full resolution, and the shift to promptable foundation models like SAM.

11 min read•3 Quiz Questions

Transformers ran language. In 2020, ViT asked: what if an image were just another sentence — 196 patch-words read by standard self-attention, with no convolutions at all? This topic covers patch embedding, the [CLS] token and positional embeddings, how attention replaces the CNN receptive field (instant global context, quadratic cost), the data-hungry pretraining caveat, and ViT's role as the backbone of CLIP, DINOv2, and SAM.

12 min read•3 Quiz Questions