PHASE 17 CURRICULUM

AI Engineering & Agentic Development Tooling

Progress0 of 52 (0%)

The practitioner capstone:

Key Architectural Domains & Syllabus
LLM API engineering (streaming, retries, cost control)
structured outputs and function-calling schemas
prompt management as versioned code
evaluation harnesses and LLM-as-judge
guardrails in application code
orchestration frameworks (LangGraph, LlamaIndex, DSPy, OpenAI Agents SDK)
context engineering for agents
coding assistants and agentic development workflows
AI product patterns
latency/cost optimization
the roles (ML engineer, AI engineer, research engineer) with the portfolio projects that prove them
52 In-Depth Topics ~416 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 17

0 of 52 completed

Cursor (by Anysphere) is a copy of VS Code rebuilt around one idea: an AI assistant that has already read your whole codebase, can edit many files itself, runs the tests, and fixes what breaks. It pairs a repo index, a fast in-house model called Composer, a plan-edit-run-verify agent loop, and MCP tool plugins.

10 min read•3 Quiz Questions

Windsurf (formerly Codeium) is an AI-native IDE whose flagship agent, Cascade, plans and executes multi-step tasks while assembling codebase context automatically. In July 2025 it was split three ways — Google licensed the tech and hired the founders, Cognition (Devin's maker) bought the company — which makes it a case study in both agentic-IDE design and vendor risk.

10 min read•3 Quiz Questions

GitHub Copilot grew from an autocomplete popup into a whole platform: completions and chat in every IDE, an agent mode, a cloud coding agent that turns GitHub issues into draft pull requests, a CLI, MCP tools, and premium-request billing that turned model choice into a budgeting skill.

10 min read•3 Quiz Questions

Claude Code is Anthropic's coding agent that lives in your terminal: type claude, and it reads your repo, edits files, and runs shell commands — governed by a committed CLAUDE.md memory file, deterministic hooks, context-isolating subagents, MCP tools, and permission prompts instead of an IDE UI. The same engine runs unattended in CI.

10 min read•3 Quiz Questions

Devin (by Cognition) is a fully autonomous coding agent that runs in its own cloud VM with a shell, editor, and browser. You queue it a ticket; it plans, writes code, runs tests, browses docs, and reports back with a pull request. It is billed in ACUs (agent compute units), grounded by Devin's Wiki / DeepWiki, and paired with the Windsurf IDE under one company.

10 min read•3 Quiz Questions

Replit Agent is a full-stack autonomous builder living inside Replit's cloud IDE: describe an app and it plans, writes the code, provisions the database and auth, tests its own UI in a real browser, and deploys it with one click. Agent 3 (Sept 2025) pushed autonomy to sessions of up to ~200 minutes — with platform lock-in as the price.

10 min read•3 Quiz Questions

Cline and its fork Roo Code are open-source coding agents that live inside VS Code. You bring your own API key (BYOK), so nothing is hidden: every plan, prompt, file diff, and shell command is visible and human-approved step by step, any model can be plugged in — even a fully local one — and MCP servers extend its tools. They are the transparency counterweight to closed agent IDEs.

10 min read•3 Quiz Questions

Aider is the open-source terminal pair programmer with one defining habit: every accepted change becomes a clean, attributed git commit. Its tree-sitter repo map gives the model cheap whole-repo awareness, and its architect/editor two-model split cut the cost of multi-file work — many commercial tools later copied these concepts.

10 min read•3 Quiz Questions

Vibe coding — Karpathy's February 2025 term for steering AI agents almost entirely by intent, accepting changes without reading the code — is a genuine workflow and a genuine risk. The useful move is treating it as a dial: prototype-grade freedom on one end, production-grade discipline (specs, tests, review, provenance) on the other.

11 min read•3 Quiz Questions

Context engineering is the art of curating what fills a model's finite attention window at each step: instructions, retrieved code, memory, tool schemas, and message history. It explains why agent IDEs index repos, why CLAUDE.md works, and why long sessions rot.

12 min read•3 Quiz Questions

LangChain is the most-deployed LLM application framework: provider-agnostic interfaces for models, tools, and retrieval, and — since the October 2025 1.0 release — a lean agent abstraction (create_agent) built on LangGraph runtime with structured outputs and middleware.

11 min read•2 Quiz Questions

LlamaIndex is the data-side counterpart to LangChain: connectors, ingestion pipelines, and index/retriever abstractions over 300+ document and data sources, plus Workflows for event-driven agents and LlamaParse for enterprise document extraction.

11 min read•2 Quiz Questions

The Vercel AI SDK is the standard TypeScript layer for AI features: a provider-agnostic message model, streaming primitives (useChat) that bind LLM tokens to React/Svelte/Vue UI, tool calling, structured generation with Zod, and — since v5/v6 — first-class Agent abstractions.

11 min read•3 Quiz Questions

Semantic Kernel is Microsoft's enterprise SDK for embedding LLM features in .NET, Python, and Java apps — kernel + plugins + connectors with serious auth/telemetry posture. Its agent layer is converging into the Microsoft Agent Framework (AutoGen + SK, 2025).

11 min read•3 Quiz Questions

Haystack (deepset) models LLM applications as typed, inspectable pipelines of components — converters, embedders, retrievers, rankers, generators — running over pluggable document stores. Haystack 2.x rebuilt the framework around explicit, serializable DAGs for production search and RAG.

11 min read•3 Quiz Questions

AutoGen (Microsoft Research) pioneered conversation-driven multi-agent systems — agents solving tasks by talking to each other, in group chats, with code executors. v0.4 rebuilt on an actor model; in 2025-2026 it is in maintenance mode, succeeded by Microsoft Agent Framework, with AG2 as the community fork.

11 min read•3 Quiz Questions

CrewAI orchestrates agents as a team: role/goal/backstory personas executing tasks in sequential or hierarchical crews, plus Flows for event-driven deterministic pipelines. Standalone Python framework with a large course-trained community and an enterprise platform (AMP).

11 min read•3 Quiz Questions

LangGraph is the low-level agent runtime behind LangChain 1.0: model agent work as a graph of nodes over shared typed state, with checkpointed durable execution, human-in-the-loop interrupts, time-travel debugging, and LangGraph Platform deployment.

12 min read•4 Quiz Questions

LLMs answer in prose; your code needs data. Instructor wraps any LLM client so the model must fill out a Pydantic form, and failed checks trigger automatic repair retries until you get typed Python objects. Pydantic AI (1.0, September 2025) is the same team's agent framework built on one principle: typed, verifiable outputs are the interface between models and software.

12 min read•2 Quiz Questions

Guardrails AI (open-source project by Shreya Rajpal) puts a programmable security checkpoint between what an LLM says — on the way in and on the way out — and your application: a hub of composable validators (PII, toxicity, regex/quality, groundedness, malware) attached declaratively, each with an explicit on-failure policy: fix, mask, filter, or fail closed.

12 min read•2 Quiz Questions

Your agent made one bad sentence three steps after a small mistake nobody saw. LangSmith is the flight recorder plus test harness for LLM apps: hierarchical traces of every run, golden datasets with evaluators wired into CI, versioned prompts, annotation queues, and production monitoring — the LangChain ecosystem's observability layer, usable standalone via OpenTelemetry.

12 min read•2 Quiz Questions

Braintrust packages the evaluation loop into one product: datasets built from real production traces, scorers (plain functions, inspectable LLM judges, and the open Autoevals library), playgrounds that score dozens of prompt/model variants at once, experiments diffed against baselines, and CI gates via the SDK — the strongest "evals first" platform in the LLM-observability category.

12 min read•2 Quiz Questions

A RAG app answers wrong — but did it fetch the wrong pages, or write wrong words from the right pages? Ragas is the de-facto open-source evaluation library that splits those two failures: reference-free LLM-based metrics (faithfulness, answer relevancy, context precision/recall), synthetic test-set generation from your own documents, and an evaluate() API that drops into any CI or dashboard.

12 min read•2 Quiz Questions

Your ML team already logs training runs and model artifacts in Weights & Biases — do LLM apps need a second, separate vendor? W&B Weave brings LLM-app observability into the experiment-tracking house: decorate any function with @weave.op and every call is traced with lineage, plus datasets, scorers, evaluate() experiments, and live production monitors sharing the same org, auth, and dashboards as classical ML.

11 min read•2 Quiz Questions

Two poles of the 2025-2026 LLM-observability market, and both fix the same two pains. Arize Phoenix: open-source, OpenTelemetry/OTLP-native tracing and evals you can run anywhere and leave without pain. Galileo: a managed AI-reliability platform whose Luna small models score every production call cheaply in real time — replacing 1% sampling with 100% coverage — now folding into Cisco/Splunk enterprise observability.

12 min read•2 Quiz Questions

Safety evals turn policy into executable tests: refusal/over-refusal suites, prompt-injection benchmarks, jailbreak and red-team regressions, PII/policy scorers, and human-review escalation — versioned datasets that gate deploys like any other test.

8 min read•2 Quiz Questions

Prompts drift, so 2025-era teams manage them like code: immutable versions with commit messages, environments (dev/stage/prod), rollback, deployment pins, experiment lineage linking version→scores→traffic — via Git-native files or prompt registries (LangSmith Hub, PEVA/Agenta-style, portals).

7 min read•2 Quiz Questions

Instead of training your own AI model, you can rent one over the internet. The OpenAI API is the most-copied order form for renting models: pick a model by name, send messages, and get text, embeddings, or guaranteed-schema JSON back — with tool calling, prompt caching, and half-price overnight batch jobs built in.

12 min read•3 Quiz Questions

Anthropic rents out the Claude models through the Messages API, a stateless but very explicit contract built for long-running agents: budgeted thinking you can see, typed tool results, deliberate prompt-cache breakpoints, and the Model Context Protocol (MCP) — the open tool-plugging standard Anthropic created and the industry adopted.

12 min read•2 Quiz Questions

Google rents out its Gemini models through two doors: the free-to-start Gemini API for developers and enterprise-grade Vertex AI on Google Cloud. The family is natively multimodal (text, images, audio, video in one stream) with the largest context windows in the frontier class (1M-2M tokens), plus two Google-only tricks: grounding in real Google Search and deep discounts via context caching.

12 min read•2 Quiz Questions

Bedrock is the food court inside the AWS building you already secured: many vendors frontier models (Anthropic Claude, Meta Llama, Mistral, Amazon Nova and more) behind one API, inheriting your IAM, VPC endpoints, CloudTrail audit, and existing contract. On top sit managed glue — Knowledge Bases for RAG, Agents/AgentCore, and Guardrails — priced and provisioned for enterprises.

12 min read•2 Quiz Questions

Azure OpenAI serves the same OpenAI model weights under a Microsoft enterprise contract: Entra ID authentication, regional data residency, no training on your prompts, always-on content filters, and capacity you reserve as deployments and PTU units. Since Ignite 2025 it lives inside the Microsoft Foundry platform, but the API keeps the familiar OpenAI shape.

11 min read•2 Quiz Questions

Groq sells one thing: speed. Its LPU chips keep model weights in fast on-chip memory instead of slow GPU video memory, so open models (Llama, Qwen, DeepSeek distills, Gemma) generate at 300-1000+ tokens/sec through an OpenAI-compatible GroqCloud API. When a human is waiting on the answer, latency becomes the product.

11 min read•2 Quiz Questions

Open-weight models are free to use but painful to serve. Three companies — Together AI, Fireworks AI, and Replicate — rent you a production kitchen for those recipes: a broad catalog with 1M-token contexts and GPU clusters, an enterprise low-latency runtime, and a Stripe-style anything-model API. Same weights, dramatically lower serving effort and often 2-10x lower prices.

11 min read•2 Quiz Questions

Ollama made running open-weight models on your own laptop a single command: ollama run pulls a quantized GGUF file, starts a llama.cpp-based engine, and serves a localhost API that speaks both its own and the OpenAI format. No cloud, no key, nothing leaves your machine — with a registry, a desktop app, and optional cloud routing for models too big for your hardware.

11 min read•2 Quiz Questions

Hugging Face is the GitHub of AI: 1M+ model/dataset/spaces artifacts with permissive model cards and widgets, plus Inference Endpoints that promote any repo into dedicated, autoscaling, VPC-isolated production serving.

8 min read•2 Quiz Questions

For years the best AI models were locked away — you could only rent them through an API. Meta's Llama line (2023-2025) changed that by publishing the trained weights for anyone to download and run, from phones (1B) to flagships (405B and the MoE Llama 4). The lesson of the family: an AI wins not just by being smart, but by being free to host and easy to build on.

12 min read•3 Quiz Questions

Mistral AI is the European lab that made "small and efficient" fashionable: Mistral 7B (Apache 2.0) beat models twice its size, Mixtral 8x7B proved sparse Mixture-of-Experts in production (46.7B stored, only 12.8B awake per token), and the lineup later grew into a full stack — Large/Medium/Small tiers, Pixtral vision, Devstral coding, the La Plateforme API, and Le Chat.

12 min read•3 Quiz Questions

Qwen started as a solid Chinese-English bilingual model and became arguably the most-deployed open-weight family: Qwen2.5 covered every size from 0.5B to 72B and won its weight class, QwQ and the R1 era brought open reasoning, and Qwen3 (April 2025) put a "think long or answer fast" switch inside one Apache-2.0 model that spans tiny dense sizes up to a 235B MoE.

12 min read•3 Quiz Questions

DeepSeek, a Chinese hedge-fund-born lab, published two bombs: V3 — a 671B-parameter MoE (only 37B awake per token) claimed to cost about $5.6M to train, MIT-licensed — and R1, a reasoning model whose long "aha" chains of thought emerged from pure reinforcement learning against verifiable answers. Together they reset the industry's price/perception of what frontier-adjacent AI must cost.

13 min read•4 Quiz Questions

Gemma is Google's way of publishing its big-lab research in a size you can actually run: the huge Gemini "teacher" is distilled into small "student" models (1B-27B) that fit a laptop or phone. Gemma 3 added vision and 140+ languages, QAT checkpoints make 4-bit quality survive, and Gemma 3n's MatFormer nests many smaller models inside one checkpoint for tight device RAM budgets.

12 min read•3 Quiz Questions

Microsoft's Phi line bets that *what you train on* beats *how big you build*: a 1.3B model taught from curated web plus synthetic "textbook" data beat 70B-class models on code tests with ~1/50 the parameters, Phi-3 shipped MIT-licensed 3.8B/14B workhorses with 128k context, and Phi-4 taught itself to reason using RLAVF — reinforcement learning from automated AI verifiers instead of human preference labels.

12 min read•3 Quiz Questions

ElevenLabs turned speech into a product layer: text-to-speech that sounds human and emotional (~75 ms first-audio on its Flash/Turbo tiers for realtime bots), instant voice cloning with consent checks, a Voice Library marketplace, Scribe speech-to-text, Dubbing Studio, and an Agents runtime — which is why it became the default voice inside many AI products.

12 min read•3 Quiz Questions

Instead of chaining three models (listen → transcribe → think → speak), the 2024-2026 generation of realtime APIs — OpenAI's Realtime API (WebRTC/SIP, gpt-realtime) and Google's Gemini Live — uses one audio-native model that hears tone and speaks with it, hitting conversational latency. The engineering that decides whether it feels human is turn-taking, barge-in, latency budgets, and telephony — not model IQ.

13 min read•3 Quiz Questions
#313Computer-Use & Browser AgentsIntermediate PRO

Agents that operate computers the way humans do: the model looks at the screen (or the page structure) and moves the mouse and types. This topic covers the two implementation styles, the flagship products (Anthropic computer use, OpenAI Operator/CUA and Agent Mode, the open browser-agent ecosystem), and the safety and reliability engineering they demand.

11 min read•3 Quiz Questions

Real products eat messy inputs: scanned PDFs, photos, meeting recordings, video. This topic is about turning them into model-ready context — the vision/OCR/ASR/video stages, the tokenization and cost math behind each choice, and the durable two-stage pattern: normalize cheaply first, then spend expensive reasoning.

11 min read•3 Quiz Questions

The AI Engineer builds products on top of models rather than training them: LLM APIs, RAG, agents, evals, and production plumbing. Named by swyx in 2023, the role became one of the most-hired engineering titles by 2025-2026. This topic maps the job, the skill stack ranked by hiring weight, the adjacent roles it is often confused with, and the market reality.

11 min read•3 Quiz Questions

Prompt engineering exploded as a standalone role in 2022-2023 (Anthropic famously hired for it in 2024), then matured. By 2025-2026 it is less a job title and more a discipline: context engineering, eval-driven iteration, and prompt systems versioned like code inside AI product teams. This topic traces the craft, the collapse of the title, and why the skill pays better than ever.

11 min read•3 Quiz Questions

LLMOps is the platform-ops discipline for model-powered systems: a gateway for routing, keys and budgets; tracing and observability; eval pipelines wired in as CI/CD; cost and quota control; guardrails; and incident response for services that are non-deterministic by nature. This topic maps what the role owns, the toolchain, the adjacent roles, and the market.

11 min read•3 Quiz Questions

Invented at Palantir and resurrected for the AI era: a Forward-Deployed Engineer is a software engineer stationed at the customer, making platform technology actually work inside their messy environment. OpenAI, Anthropic, Mistral, Databricks and others scaled the role aggressively through 2025-2026 because AI pilots die in integration reality, not on benchmarks.

11 min read•3 Quiz Questions

AI PMs own product judgment where the core technology is probabilistic: assessing what models are genuinely good at, writing evals-as-spec instead of behavior checklists, making cost/latency/quality trade-offs, designing trust and failure UX, and pricing agent products. This topic shows how classic PM craft is being reinvented for the model era.

11 min read•3 Quiz Questions

Applied AI/Research Engineers make models production-usable: they read papers as specifications, reproduce and extend research, fine-tune and distill models, optimize inference with vLLM/quantization/speculative decoding, and build the evaluation infrastructure that defines what better means. This topic maps the seat between research science and ML engineering, its technical surface, and the market around it.

11 min read•3 Quiz Questions