Vision Transformers (ViT)
Transformers ran language. In 2020, ViT asked: what if an image were just another sentence — 196 patch-words read by standard self-attention, with no convolutions at all? This topic covers patch embedding, the [CLS] token and positional embeddings, how attention replaces the CNN receptive field (instant global context, quadratic cost), the data-hungry pretraining caveat, and ViT's role as the backbone of CLIP, DINOv2, and SAM.
01.The Problem: The CNN Never Truly Sees the Whole Picture
A CNN neuron's view of the input is bounded by its receptive field — the region of pixels that can influence it, which widens by only ~2 pixels per stacked 3x3 layer (the receptive-field topic has the derivation).
And even that generous geometry overstates things: measurements of the Effective Receptive Field showed trained CNNs concentrate their real attention on a Gaussian core much smaller than the theoretical window. Deep CNNs approximate global context. They never natively get it.
So when a question needs relations across the image —
Is that the same person who appears at the top-left and the bottom-right? Is the umbrella over the head, or over the stranger three meters away?
— the CNN must learn long-range wiring through many indirect hops.
Meanwhile, in language AI, a different architecture had taken over: the Transformer (2017) — a stack of layers built on self-attention, where every word directly looks at every other word. Its native unit is a sequence.
Dosovitskiy et al. (2020) asked the blunt question:
If Transformers can read any sequence... why not make the image be a sequence?
No kernels. No locality. No gradual widening. Just: chop the image into pieces, call the pieces "words," and let attention do what attention does — connect any piece to any piece, from layer one.
ViT: image → token sequence 🧩
ViT: image → token sequence 🧩
An image is chopped into fixed-size patches, each flattened and linearly projected into an embedding, positional encodings and a learnable [CLS] token are prepended, and the resulting token sequence is processed by stacked Transformer blocks.
Unlock Topic #109: Vision Transformers (ViT)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?