Multimodal Embeddings
An embedding turns anything — text, image, PDF page — into a list of numbers, and a multimodal embedding puts them all in ONE shared number space so you can search pictures with words. This covers the three designs (CLIP-style dual encoders, SigLIP scaling, ColPali late interaction, LLM-based embedders) and the very real storage bills they create.
01.The Problem: Your Search Knows Words but Not Pictures
You have 10 million files: PDFs, screenshots, product photos, slide decks.
You want to type "the chart showing revenue growth in Q3" and find the picture.
Classic search (BM25, keyword match) can only find letters. It cannot see that a photo of a golden retriever matches the words "a dog on grass". It cannot tell that a scanned invoice looks different from a scanned contract.
So the question becomes
How do you search things that have no words — or whose words are trapped in pixels?
The answer: convert everything into numbers first. Not one number — a list of numbers that captures meaning. Then "searching" becomes "finding lists of numbers that are close to each other".
- A text embedding turns a sentence into that list.
- An image embedding turns a picture into the same kind of list.
- A multimodal embedding puts both lists in the same space, so a sentence and a picture can be "close" — the dog photo sits near the words "a dog on grass".
Carry this analogy through the topic: one giant library catalog for everything. Books (text), photographs, maps, even the scanned invoices — all filed in the same aisles by subject, so you can walk from a word to a picture without changing buildings. The three encoding designs in this topic are three ways of writing catalog cards.
Three Embedding Designs, Three Storage Profiles 🗂️
Three Embedding Designs, Three Storage Profiles 🗂️
Architecture choice determines index size and query cost as much as accuracy does: one vector per item, one matrix per page, or one LLM forward pass per item.
Unlock Topic #262: Multimodal Embeddings
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?