CLIP: Contrastive Language-Image Pre-training
CLIP trains two towers - an image encoder and a text encoder - with a matching game over 400M web image-caption pairs so that matched pairs land next to each other in one shared embedding space. That single space gives zero-shot classification (just type new labels), image/text retrieval, and the text-conditioning interface used by Stable Diffusion and today's vision-language models.
01.The Problem: A Classifier Only Knows the Labels It Was Shown
The classic image classifier is trained like this: collect photos of 1000 species, teach the network those 1000 names, deploy.
Then someone asks it to recognize "a dog riding a skateboard". Failure. The label does not exist. To add it you must gather images, annotate them, and retrain.
So the question becomes
Is there a way to teach a model the concepts themselves, so that brand-new categories can be answered without any retraining — just by typing them?
Meanwhile, the web was sitting on the biggest free teaching material ever assembled: billions of images, each with a sentence about it (alt-text, captions, filenames). Humans had already written down what the pictures show — for free, in plain language.
CLIP (Contrastive Language-Image Pre-training, Radford et al., OpenAI, 2021) used that material to do something modest on paper and enormous in practice: teach a picture-net and a sentence-net to agree on where things live in the same map.
CLIP: Two Towers, One Shared Space
CLIP: Two Towers, One Shared Space
An image encoder and a text encoder project each input into a common embedding space. The symmetric InfoNCE loss pulls the N correct image-text pairs together and pushes apart the N*N - N incorrect pairs in each batch, aligning vision and language.
Unlock Topic #190: CLIP: Contrastive Language-Image Pre-training
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?