Text-to-Video Generation (Sora, Lumiere)
Text-to-video is latent diffusion plus one brutal new demand: every frame must agree with every other frame. This topic covers the two tools that buy temporal consistency - spatio-temporal attention and a video VAE that compresses time as well as space - plus Sora's spacetime-patch Diffusion Transformer, Lumiere's single-stage Space-Time U-Net, and the 2025-2026 frontier.
01.The Problem: One Pretty Picture Is Easy; Two Identical Ones Are Not
Ask Stable Diffusion (topic 188) for "a red fox trotting through snow" and you get a lovely frame. Ask for the video and you have a new contract:
- frame 1 has a fox,
- frame 2 must have the same fox — same shape, same coat, same direction of travel —
- …and frame 24 × 4 seconds of it, with snow falling continuously, paws alternating, light stable.
A single frame just needs to look real. A video needs every frame to agree on object identity, layout, lighting, and physics while moving plausibly.
The naive attempt: run an image model once per frame. The result flickers and morphs — the fox changes face every 200 milliseconds, snow teleport-pop, legs melt. Why? Each frame was generated as an independent coin flip; nothing bound frame 7 to frame 6.
So the question becomes
How do we make a diffusion model that does not just denoise a picture, but denoises a spacetime volume — where every patch constantly checks what its neighbors-in-time are doing?
Carry one analogy through: a flipbook. You draw the same runner on 100 pages and flip. Any single page can be a masterpiece — but if page 57 gives the runner a different shirt, the flick shows it. The hardest part of the flipbook is not drawing; it is not losing the character between pages. Video models are flipbook artists whose 100 pages are denoised simultaneously and can talk to each other.
Latent Video Diffusion Pipeline
Latent Video Diffusion Pipeline
A video VAE compresses frames in both space and time into latents, cut into spacetime patches like a ViT. A diffusion transformer iteratively denoises those patches while cross-attending to the prompt, then a decoder renders frames. Spatio-temporal attention is what keeps subjects and motion consistent across time.
Unlock Topic #192: Text-to-Video Generation (Sora, Lumiere)
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?