PHASE 11 CURRICULUM

Generative AI Beyond Text

Progress0 of 12 (0%)

The generative model zoo beyond autoregressive text:

Key Architectural Domains & Syllabus
variational autoencoders
adversarial training with GANs
the diffusion process and denoising score matching
latent diffusion and the image-generation stack (Stable Diffusion, control mechanisms)
video and audio generation
multimodal models that see and speak
the safety considerations unique to generative systems
12 In-Depth Topics ~96 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 11

0 of 12 completed

A GAN learns to create by pitting two networks against each other: a forger (generator) and a detective (discriminator). The zero-sum game has an elegant optimum — matching the data distribution exactly — but training it is notoriously unstable. Here is the math, the failure modes, the StyleGAN lineage, and where adversarial losses still run 2024-2026 products.

14 min read•3 Quiz Questions

A VAE compresses data into a deliberately smooth probabilistic latent space, so any random point decodes to something meaningful. The enabler is the reparameterization trick (z = mu + sigma*epsilon), the objective is the ELBO (reconstruction minus KL), and the legacy is the VQ-autoencoder backbone inside Stable Diffusion.

13 min read•3 Quiz Questions

Diffusion generates by iterative refinement: destroy data with a fixed noising schedule, then train one small neural network to undo a little noise at a time. That simple noise-prediction regression unified DDPMs, score matching, and flow matching into the dominant generative paradigm for image, video, and audio in 2024-2026.

14 min read•3 Quiz Questions

DDPM (2020) generates images by learning to reverse a simple process: gradually bury a real photo in Gaussian noise until it is pure static, then train one small skill — "given this noisy image, tell me what noise is in it" — and undo the damage one tiny step at a time. The whole method reduces to a noise-prediction MSE, and it launched the diffusion revolution behind Stable Diffusion and Sora.

16 min read•4 Quiz Questions

Stable Diffusion packages four separable parts — a VAE that shrinks pixels into a small latent, a CLIP text encoder that reads your prompt, a cross-attention denoising U-Net, and a scheduler that plans the steps — into an open text-to-image model that runs on one consumer GPU, and then evolved through SDXL and SD3/FLUX into an entire extensible ecosystem.

16 min read•4 Quiz Questions

LDM is the two-stage framework behind Stable Diffusion: first train a perceptual autoencoder to shrink images into a small latent where diffusion is ~48x cheaper, then run the diffusion entirely in that latent, steering it with any condition (text, sketch, mask) injected through cross-attention plus classifier-free guidance.

15 min read•4 Quiz Questions

CLIP trains two towers - an image encoder and a text encoder - with a matching game over 400M web image-caption pairs so that matched pairs land next to each other in one shared embedding space. That single space gives zero-shot classification (just type new labels), image/text retrieval, and the text-conditioning interface used by Stable Diffusion and today's vision-language models.

15 min read•4 Quiz Questions

CFG makes diffusion models obey prompts by a training trick and a sampling trick: randomly drop the condition so ONE network learns both "denoise with the prompt" and "denoise without it", then amplify the difference between the two predictions at every step. The guidance scale w is the fidelity-vs-diversity dial behind every guidance_scale and negative_prompt in production - and 2024-era distillation folds its 2x cost back into one pass.

15 min read•4 Quiz Questions

Text-to-video is latent diffusion plus one brutal new demand: every frame must agree with every other frame. This topic covers the two tools that buy temporal consistency - spatio-temporal attention and a video VAE that compresses time as well as space - plus Sora's spacetime-patch Diffusion Transformer, Lumiere's single-stage Space-Time U-Net, and the 2025-2026 frontier.

16 min read•4 Quiz Questions

Speech modeling runs in two mirrored directions: TTS expands text down to a waveform, ASR compresses a waveform up to words. This topic traces both from hand-glued HMM modules to end-to-end neural nets - acoustic models, vocoders, CTC/Transducer/seq2seq heads, self-supervised encoders, and RVQ codec tokens - and shows how 2024-2026 native-audio LLMs fold the two pipelines into one.

16 min read•4 Quiz Questions

Whisper is a plain seq2seq Transformer ASR that got its near-human robustness from data, not architecture: ~680k hours of messy, weakly-labeled web audio taught one model to transcribe, translate, and align about 100 languages zero-shot via a special-token prompt - and its distilled/streamed descendants (faster-whisper, Distil-Whisper, whisper.cpp) are how open transcription ships today.

15 min read•4 Quiz Questions

How vision, audio, and language fuse into one model: the open recipe bolts a pretrained vision encoder onto an LLM through a tiny projector (LLaVA), trained first to align the modalities then to follow multimodal instructions - while GPT-4o and Gemini train one transformer over interleaved text, audio, and video end to end, buying coherence, native speech, and million-token context at enormous cost.

16 min read•4 Quiz Questions