PHASE 15 CURRICULUM

AI Alignment, Safety & Ethics

Progress0 of 14 (0%)

The trust layer every AI engineer must own:

Key Architectural Domains & Syllabus
fairness and bias measurement and mitigation
privacy techniques including differential privacy
adversarial examples and data poisoning
jailbreaks and prompt injection
red-teaming
guardrail architectures and output filtering
interpretability and mechanistic analysis
alignment approaches
watermarks and provenance
the regulatory landscape (EU AI Act and beyond) shaping production AI
14 In-Depth Topics ~112 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 15

0 of 14 completed

Alignment is the gap between the objective we write down and the behaviour we actually want — the classic genie problem, made engineering. This topic covers outer vs inner alignment, the RLHF/DPO/Constitutional-AI toolkit, the documented failure modes (goal misgeneralization, reward hacking, sycophancy, over-refusal), and why alignment is a systems problem rather than a loss function.

15 min read•4 Quiz Questions

"When a measure becomes a target, it ceases to be a good measure." This topic makes that slogan precise: the four Goodhart variants and their different fixes, documented specification-gaming incidents in RL and RLHF, the reward-model overoptimization inverted-U, and the metric-design defenses (orthogonal checks, private evals, satisficing, process supervision, tail audits) that actually work.

15 min read•4 Quiz Questions

A trained model is a program nobody wrote line by line — mechanistic interpretability treats it as software to be decompiled. This topic covers the circuit abstraction (the residual stream as shared workspace, attention as data movement, MLPs as key-value memories), the interventional method stack (activation patching, logit lens, causal scrubbing, sparse autoencoders), the documented circuit results, and the honest limits of the field.

15 min read•4 Quiz Questions

Networks store far more features than they have neurons by overlapping them at near-perpendicular angles — that is the superposition hypothesis. It explains why one neuron fires for "cat", "French", and "syntax errors" (polysemanticity), why penalizing activity makes neurons specialize (monosemanticity, shown in Anthropic's Toy Models), and why sparse autoencoders — a change of basis — are the right tool to unpack the mixture.

13 min read•4 Quiz Questions

A probe is a tiny second model trained on a frozen model's activations to ask: "can this property be read off?" This topic covers how to build one, why a high probe score does NOT prove the model uses the property, and how control tasks, model-free baselines and amnesic counterfactuals turn a decoding test into causal evidence.

12 min read•3 Quiz Questions

Run the model twice — once succeeding, once failing — then overwrite one internal activation at a time and measure what moved. That swap experiment is activation patching, the field's main causal tool. This topic covers the clean/corrupted design, the metric choices that change answers, and the attribution-graph pipeline that scales tracing to production models.

12 min read•3 Quiz Questions

Benchmarks are standardized tests for models: a fixed question set plus a scoring rule. This topic walks through MMLU, HumanEval, MATH and GPQA — what they ask, the famous scores, where each one saturates or leaks — plus HELM's multi-metric idea and the documented failure modes (contamination, weak graders, Goodharted leaderboards) every engineer must know.

12 min read•3 Quiz Questions

Red teaming means hiring your own attackers: a team with a goal (make the model do something harmful or leak data) probes your model and product before real users or criminals find the same hole. This topic covers the programme — threat models, harm taxonomies, human plus automated attack frameworks (garak, PyRIT, HarmBench, PAIR/TAP), severity scoring, and regression gates that actually stop releases.

12 min read•3 Quiz Questions

An LLM reads orders and information in the same stream of text — it has no way to tell "do this" from "this is what a document says". Jailbreaking abuses that from the user side; prompt injection abuses it through data your app feeds the model. This topic covers the documented mechanisms, why safety training cannot fully fix either, and the architecture-level defenses that actually bound the damage.

12 min read•3 Quiz Questions

ML systems inherit unfairness through four doors — history, data coverage, faulty proxies, and one-size-fits-all thresholds — and the three formal definitions of "fair" provably contradict each other. This topic explains where bias enters, what demographic parity, equalized odds and calibration actually mean, why language models add their own failure modes, and how slice-level evaluation plus documentation makes fairness measurable instead of rhetorical.

13 min read•3 Quiz Questions
#247Differential PrivacyAdvanced PRO

Differential privacy answers a devious question with math: "would an expert tell I was in this dataset?" It forces every published answer to look almost identical whether or not your record was included, by adding noise calibrated to how much one record can change the result. This topic covers the (ε, δ) guarantee, the Gaussian mechanism, privacy budgets and composition, DP-SGD for training, real deployments, and the accuracy price engineers actually pay.

13 min read•3 Quiz Questions

Language models can repeat training text word-for-word — including *your* text. A parrot that heard one sentence a hundred times owns that sentence. This topic covers why memorization happens (capacity and repetition), the attacks that pull text back out (continuation prompting, prompt-guided PII extraction, membership inference), what production chat systems actually leaked, and the mitigation stack from deduplication to unlearning.

13 min read•3 Quiz Questions

AI is now governed like buildings: rules sorted by risk (a garden shed needs no permit; a school needs many), plus inspectors who only accept paperwork. This topic maps the 2024-2026 regimes — the EU AI Act's risk tiers and dates, NIST AI RMF and its Generative AI Profile, the US state patchwork, China, the UK, ISO/IEC 42001 — and then gets concrete about the exact artefacts a platform team must ship to be defensible.

13 min read•3 Quiz Questions

A sober engineering look at catastrophic and existential AI risk: how governments classified the risk families in 2025, the concrete technical mechanisms behind loss-of-control scenarios (long-horizon autonomy, inner alignment, instrumental convergence, demonstrated deception), what the evidence does and does not show, and the controls — safety cases, staged gates, monitorability, containment, compute governance — that pay off no matter whose timeline is right.

13 min read•3 Quiz Questions