Generative AI Beyond Text
The generative model zoo beyond autoregressive text:
All Topics in Phase 11
0 of 12 completedA GAN learns to create by pitting two networks against each other: a forger (generator) and a detective (discriminator). The zero-sum game has an elegant optimum — matching the data distribution exactly — but training it is notoriously unstable. Here is the math, the failure modes, the StyleGAN lineage, and where adversarial losses still run 2024-2026 products.
A VAE compresses data into a deliberately smooth probabilistic latent space, so any random point decodes to something meaningful. The enabler is the reparameterization trick (z = mu + sigma*epsilon), the objective is the ELBO (reconstruction minus KL), and the legacy is the VQ-autoencoder backbone inside Stable Diffusion.
Diffusion generates by iterative refinement: destroy data with a fixed noising schedule, then train one small neural network to undo a little noise at a time. That simple noise-prediction regression unified DDPMs, score matching, and flow matching into the dominant generative paradigm for image, video, and audio in 2024-2026.
DDPM (2020) generates images by learning to reverse a simple process: gradually bury a real photo in Gaussian noise until it is pure static, then train one small skill — "given this noisy image, tell me what noise is in it" — and undo the damage one tiny step at a time. The whole method reduces to a noise-prediction MSE, and it launched the diffusion revolution behind Stable Diffusion and Sora.
Stable Diffusion packages four separable parts — a VAE that shrinks pixels into a small latent, a CLIP text encoder that reads your prompt, a cross-attention denoising U-Net, and a scheduler that plans the steps — into an open text-to-image model that runs on one consumer GPU, and then evolved through SDXL and SD3/FLUX into an entire extensible ecosystem.
LDM is the two-stage framework behind Stable Diffusion: first train a perceptual autoencoder to shrink images into a small latent where diffusion is ~48x cheaper, then run the diffusion entirely in that latent, steering it with any condition (text, sketch, mask) injected through cross-attention plus classifier-free guidance.
CLIP trains two towers - an image encoder and a text encoder - with a matching game over 400M web image-caption pairs so that matched pairs land next to each other in one shared embedding space. That single space gives zero-shot classification (just type new labels), image/text retrieval, and the text-conditioning interface used by Stable Diffusion and today's vision-language models.
CFG makes diffusion models obey prompts by a training trick and a sampling trick: randomly drop the condition so ONE network learns both "denoise with the prompt" and "denoise without it", then amplify the difference between the two predictions at every step. The guidance scale w is the fidelity-vs-diversity dial behind every guidance_scale and negative_prompt in production - and 2024-era distillation folds its 2x cost back into one pass.
Text-to-video is latent diffusion plus one brutal new demand: every frame must agree with every other frame. This topic covers the two tools that buy temporal consistency - spatio-temporal attention and a video VAE that compresses time as well as space - plus Sora's spacetime-patch Diffusion Transformer, Lumiere's single-stage Space-Time U-Net, and the 2025-2026 frontier.
Speech modeling runs in two mirrored directions: TTS expands text down to a waveform, ASR compresses a waveform up to words. This topic traces both from hand-glued HMM modules to end-to-end neural nets - acoustic models, vocoders, CTC/Transducer/seq2seq heads, self-supervised encoders, and RVQ codec tokens - and shows how 2024-2026 native-audio LLMs fold the two pipelines into one.
Whisper is a plain seq2seq Transformer ASR that got its near-human robustness from data, not architecture: ~680k hours of messy, weakly-labeled web audio taught one model to transcribe, translate, and align about 100 languages zero-shot via a special-token prompt - and its distilled/streamed descendants (faster-whisper, Distil-Whisper, whisper.cpp) are how open transcription ships today.
How vision, audio, and language fuse into one model: the open recipe bolts a pretrained vision encoder onto an LLM through a tiny projector (LLaVA), trained first to align the modalities then to follow multimodal instructions - while GPT-4o and Gemini train one transformer over interleaved text, audio, and video end to end, buying coherence, native speech, and million-token context at enormous cost.