Semantic Segmentation
A bounding box around a pedestrian still includes slices of sidewalk. Safety-critical systems need the opposite: a class label for every single pixel — road here, car there, sky above. That is semantic segmentation. This topic builds the fully-convolutional idea, the encoder-decoder U-Net with skip connections, dilation/ASPP for context at full resolution, and the shift to promptable foundation models like SAM.
01.The Problem: The Box Is a Lie
Your self-driving car runs a detector (see the object-detection topic) and gets a bounding box around a pedestrian.
Now look closer: the box is a rectangle. The pedestrian is a person-shaped hole in that rectangle — around their legs and between their arms, the box contains road.
So the question becomes
If I need to plan a path the car can actually drive on, don't I need to know which pixels are road?
One label per image (classification) is too little. One box per object (detection) is too coarse. Sometimes you need the full truth:
"Every pixel in this 224x224 image — what is it?"
That is semantic segmentation: turn an H x W x 3 image into an H x W map of class IDs — "road," "car," "pedestrian," "sky," "tumor." It is called dense prediction because the output is an image-sized grid of answers rather than a single vector.
And it powers real products: drivable-area masks, organ and tumor delineation in radiotherapy, photo background removal, satellite land-cover maps, robot traversability.
U-Net encoder-decoder with skip connections 🧩
U-Net encoder-decoder with skip connections 🧩
The encoder shrinks resolution while growing semantic depth; the decoder upsamples back to pixel resolution, concatenating same-size encoder feature maps (skips) to recover fine spatial detail lost during downsampling.
Unlock Topic #108: Semantic Segmentation
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?