Multimodal Models (LLaVA, GPT-4o, Gemini)
How vision, audio, and language fuse into one model: the open recipe bolts a pretrained vision encoder onto an LLM through a tiny projector (LLaVA), trained first to align the modalities then to follow multimodal instructions - while GPT-4o and Gemini train one transformer over interleaved text, audio, and video end to end, buying coherence, native speech, and million-token context at enormous cost.
01.The Problem: A Genius That Can Only Read
By 2023, LLMs could write code, reason about essays, and explain physics — and could not tell you what was in a photograph. They are brains in a dark room: every input must already be text.
But most of what people actually want is not text:
- "what is wrong with this chest X-ray?",
- "summarize the meeting from this whiteboard photo",
- "click the button in this app screenshot",
- "answer while I talk to you" — voice in, voice out.
The brute-force fix — train a brand-new model from scratch on everything — is absurdly expensive and throws away the strongest pretrained models humanity has. So the question becomes
Can we give eyes and ears to an LLM we already have, instead of raising a new mind?
That is the vision-language model (VLM) project. And the open answer turned out to be embarrassingly modular: take a model that already sees (CLIP — topic 190), take a model that already reads (an LLM), and train a tiny translator between them.
Fusing Modalities into One Transformer
Fusing Modalities into One Transformer
The common pattern: pre-trained per-modality encoders produce feature tokens that a lightweight projector maps into the LLM's embedding space, where they are consumed alongside text tokens as one interleaved sequence. Truly natively-multimodal models (GPT-4o, Gemini) instead train a single transformer over all modalities end to end.
Unlock Topic #195: Multimodal Models (LLaVA, GPT-4o, Gemini)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?