PHASE 16 CURRICULUM

Current Frontiers & Trends (2024-2026)

Progress0 of 18 (0%)

The live research frontier explained from first principles:

Key Architectural Domains & Syllabus
reasoning models and test-time compute scaling (o1/o3, DeepSeek-R1)
RL with verifiable rewards
small language models and on-device intelligence
agentic computer use
world models and video understanding
synthetic data at scale
continued pretraining
retrieval-native architectures
the economic and capability trade-offs between open-weight and closed frontier models
18 In-Depth Topics ~144 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 16

0 of 18 completed

Frontier AI models today learn largely from text written by other AI models. This topic covers the factory: a strong teacher model generates examples, a verifier keeps only the ones that pass a real check, and the survivors are blended with human data. Plus the risks — model collapse, contamination, licensing — and how to run it as infrastructure.

16 min read•3 Quiz Questions

AI coding tools went from autocomplete ghosts (2021) to chat-with-your-repo (2023) to agents that read files, edit code, run tests, and open pull requests (2024-2026). The surprise: the product is mostly the loop around the model — context, tools, verification, budgets, guardrails — not the model itself.

16 min read•3 Quiz Questions
#253Multimodal ReasoningAdvancedFREE

Vision-language models can now reason over pictures: a vision encoder turns images into pseudo-tokens that join the text stream, and 2025-2026 models act on what they see mid-thought — cropping, zooming, running OCR, writing measuring code. This topic covers the two architectures, the grounding mechanisms that made VLMs useful, where they still fail, and how to serve them.

16 min read•3 Quiz Questions

An agent that nails a five-minute task in 2024 can still derail on a five-hour one. The reason is math, not magic: success decays like p^N as dependent steps stack up, while context rots and the goal drifts. This topic covers the four failure modes and the architecture — durable state, exit criteria, sub-agents, verifier gates, checkpoints — that extends the reliable horizon.

17 min read•3 Quiz Questions

Computer-use agents drive a GUI the way a human does: screenshot in, click-and-type out. This topic covers how grounding turns language into a click target, why pixel loops are slow and fragile, where structured interfaces (a11y trees, Playwright, MCP, APIs) beat vision, and what OSWorld/WebArena scores actually mean.

16 min read•3 Quiz Questions

Classic RLHF pays humans to rate answers — slow, expensive, fooled by long text, and useless for judging a 300-step proof. The 2024-2026 stack swapped in four alternatives: AI-labelled preferences (RLAIF/CAI), offline pair losses (DPO/KTO/SimPO), self-play, and programmatic verification (RLVR + GRPO). This topic is how each works and when to use which.

17 min read•3 Quiz Questions

A model writes a 12-step solution and gets one number: right or wrong. Which step caused the failure? That is credit assignment, and it is the crux of training reasoning. Outcome reward models grade the final answer; process reward models grade every step. This topic covers what each buys, what each costs, and how 2024-2026 practice blends them with rule-based verifiers.

16 min read•3 Quiz Questions

Can an AI improve itself — writing its own practice problems, grading its own answers, even rewriting its own tools? From STaR and Reflexion to self-rewarding loops, zero-data self-play, and agents that rewrite their own scaffolds (Darwin Gödel Machine), this topic shows what actually bootstraps, where the ceiling is, and how to build the loop safely.

17 min read•3 Quiz Questions

Model merging combines fine-tuned checkpoints by arithmetic on weights — no training, no pooled data. This topic explains why same-lineage merges work at all (linear mode connectivity), the methods (weight-average soups, task vectors, SLERP, TIES, DARE), and why merges so often wreck calibration and instruction-following, plus the recipe that keeps them honest.

17 min read•3 Quiz Questions
#260Open vs Closed ModelsAdvanced PRO

A closed model is a service you rent; an open model is one you can download, change and run yourself. But "open" is really a ladder of five rungs, and between 2024 and 2026 open-weight models (DeepSeek, Qwen3, Llama, Gemma, Mistral) closed most of the quality gap. The real choice is cost, control, customisation and risk — not ideology.

15 min read•3 Quiz Questions

Instead of training a bigger model, spend more compute when answering: sample many attempts and rerank, revise with feedback, or let the model think longer. This topic covers the techniques, the verifier you always need, the overthinking problem, and the serving costs that make thinking budgets a product feature.

15 min read•3 Quiz Questions
#262Multimodal EmbeddingsAdvanced PRO

An embedding turns anything — text, image, PDF page — into a list of numbers, and a multimodal embedding puts them all in ONE shared number space so you can search pictures with words. This covers the three designs (CLIP-style dual encoders, SigLIP scaling, ColPali late interaction, LLM-based embedders) and the very real storage bills they create.

15 min read•3 Quiz Questions

An LLM forgets everything when the chat ends. Memory systems fix that: a small "desk" of facts always in context, big filing cabinets outside it, and rules for what to write, update, and forget. MemGPT started it by treating the context window like RAM; the hard engineering turns out to be the write policy, not the storage.

16 min read•3 Quiz Questions
#264LLM and Prompt CachingAdvanced PRO

Every LLM call re-reads its prompt from scratch — unless you cache it. Prompt caching reuses the model's "reading notes" (the KV cache) for the unchanged *beginning* of your prompt, cutting cost up to 90% and time-to-first-token by ~80% on long prompts. This topic covers what is cached, provider pricing, the prompt layout that actually hits, and the routing behind it.

15 min read•3 Quiz Questions
#265Mixture of AgentsAdvanced PRO

Mixture of Agents asks several different LLMs the same question, lets each read the others' drafts as references, and has one aggregator fuse a better final answer. It is an ensemble at inference time — not Mixture-of-Experts — and while it buys real quality on open-ended tasks, it costs 10x tokens. Its best 2026 role is offline: generating training data.

15 min read•3 Quiz Questions

Plain RAG blindly trusts whatever it retrieves. Corrective RAG adds a quality gate that grades the documents (use them, decompose them, or throw them away and search elsewhere). Speculative RAG attacks the other problem — slowness — by fetching likely contexts in parallel before the router decides. This topic covers both, plus Self-RAG and the metrics that keep them honest.

15 min read•3 Quiz Questions

Programs need values, not essays. Structured output makes an LLM fill a form instead of writing free text — and grammar-constrained decoding is the only method that *guarantees* the output matches your schema, by blocking every token that would break it. This topic covers how that works, what "strict mode" really enforces, and the semantics-versus-syntax gap retries cannot close.

15 min read•3 Quiz Questions

Open-ended answers have no auto-grader, so we hire a model to grade models. This topic covers judging protocols (pointwise, pairwise, checklists), the four documented judge biases and their fixes, how Chatbot Arena-style Elo/Bradley-Terry rankings are computed from votes, and why a judge is a sensor with known error — never ground truth.

15 min read•3 Quiz Questions