AI & ML MASTER REFERENCE MATRIX

Formulas, Scaling Laws & Decision Matrices for AI/ML Interviews

The quick-reference companion for the 320-topic AI & ML curriculum: ML math essentials, classical complexity tables, training defaults, transformer reference specs, LLM scaling laws, RAG numbers, fine-tuning & quantization charts, and MLOps checklists.

1. Mathematical Foundations for ML

Bayes’ theoremP(A|B) = P(B|A) P(A) / P(B) posterior = likelihood x prior / evidence
EntropyH(p) = -SUM_x p(x) log2 p(x) minimum mean bits to encode draws from p
Cross-entropyH(p,q) = -SUM_x p(x) log q(x) = H(p) + KL(p||q) the classification loss
KL divergenceKL(p||q) = SUM_x p log (p/q) >= 0, asymmetric VAE prior, distillation, PPO/KL-to-reference penalty
Gradient descenttheta <- theta - lr * grad_theta L(theta) backprop = chain rule over the graph
MLE vs MAPMLE: argmax log p(data|theta); MAP: argmax log p(data|theta) + log p(theta)
L1 vs L2 regularizationL1 (lasso) drives weights to exactly zero - sparse feature selection. L2 (ridge) shrinks smoothly toward zero. Elastic net mixes both. Typical lambda 0.001 to 1. This is why L1 = lasso and L2 = ridge.
Variance vs biasTotal error = bias^2 + variance + irreducible noise. Underfit = high bias; overfit = high variance. Regularization and data both attack variance; capacity attacks bias.

2. Data Work & Classical ML Complexity

AlgorithmTrainPredictSweet Spot
Linear regressionO(n d^2) closed form (normal eq.)O(d)Baselines + interpretability on tabular data
Logistic regressionIterative GD, ~O(n d) / epochO(d)Calibrated probabilities, huge sparse text features
Decision treeO(n d log n)O(depth)Fast prototyping, explainable splits; overfits deep
Random foresttrees x tree costtrees x O(depth)Robust tabular default; OOB error is free validation
XGBoost / LightGBMiters x O(n d log n)O(iters x depth)Still wins most Kaggle tabular competitions
SVM (RBF kernel)O(n^2 - n^3)O(#SVs x d)Small/medium clean datasets; scales features first
k-NNO(1) (stores data)O(n d)Tiny datasets, embedding similarity, k = 3-15
Naive BayesO(n d)O(d)Text-classification baseline, streaming updates
k-meansO(n k t) per run (t ~ 10-100 iters)O(k d)Spherical clusters, quantization, color reduction
DBSCANO(n log n) with spatial index-Arbitrary shapes + noise/outlier detection
PCA (SVD)O(n d^2)O(k d)2-20 components capture 50-95% variance for viz
Precision / Recall / F1P = TP/(TP+FP); R = TP/(TP+FN); F1 = 2PR/(P+R)
ROC-AUC vs PR-AUCROC survives class imbalance (few FPs); report PR-AUC when positives < 1%
Splits80/10/10 or 60/20/20; stratified for imbalance; walk-forward (never shuffle) for time series

Data leakage kills silently

Standardizers fit on the FULL dataset, target-encoded features, timestamps after the label event - all inflate validation scores and collapse in production. Fit every transform on the train split only, inside a cross-validation pipeline.

3. Neural Network Training Reference

ActivationFormulaWhere it lives
ReLUmax(0, x)CNN/MLP hidden layers; cheap, dying-ReLU risk; He init
GELUx * Phi(x)Transformers and LLMs (smooth ReLU); modern default
SiLU / Swishx * sigmoid(x)LLaMA-style FFNs, EfficientNet
Sigmoid1 / (1 + e^-x)Binary outputs + LSTM/GRU gates; saturates
Tanh2*sigmoid(2x) - 1Recurrent hidden states; zero-centered
Softmaxexp(x_i) / SUM exp(x_j)Multiclass heads; subtract max for numerical stability
OptimizerDefault LRNotes
SGD + momentum 0.90.01 - 0.1Vision from scratch; cosine decay; best final accuracy when tuned
AdamW (0.9 / 0.999, eps 1e-8, wd 0.01-0.1)1e-4 - 3e-4 pretrain; 1e-5 - 2e-5 fine-tuneDecoupled weight decay; the transformer workhorse
Adafactor1e-3 - 3e-3Near-zero optimizer memory for huge seq2seq models
BatchNorm vs LayerNormBN: normalize per-feature across the batch (CNNs; needs a batch). LN: normalize per-sample across features (transformers/RNNs; works at batch size 1).
Dropout0.1 in transformer blocks, 0.2-0.5 in overfit MLP heads; train-only (inference is the expectation).
Warmup + clippingLinear LR warmup for the first 0.5-3% of steps; clip grad-norm to 1.0 for RNNs/transformers.
Mixed precisionBF16/FP16 forward + FP32 master weights + loss scaling; ~2x throughput on tensor-core GPUs.

4. CNNs, Vision & Transfer Learning

Conv output sizeO = floor((W - K + 2P) / S) + 1 per spatial dimension
Conv parametersK x K x C_in x C_out + C_out (3x3, 64->64 filters = 3,712 params)
Stacked 3x3sTwo 3x3 layers see a 5x5 field with 8/9 the params and one extra nonlinearity
ArchitectureYearParamsWhy it matters
LeNet-5199860KFirst deployed CNN (bankcheck OCR)
AlexNet201260MGPU + ReLU + dropout ImageNet moment
VGG-162014138MDepth via repeated 3x3 blocks; simple recipe
ResNet-50201525MSkip connections: y = F(x) + x, 100+ layers trainable
EfficientNet-B020195MCompound scaling of depth x width x resolution
ViT-B/16202086MPatches-as-tokens transformer; needs data or distillation
Transfer-learning recipeFrozen backbone + new head at lr 1e-4; unfreeze top blocks at lr 1e-5. Under ~10K images, frozen features often beat full fine-tunes.
Detection familiesTwo-stage (Faster R-CNN: accurate, region proposals) vs one-stage (YOLO/SSD: real-time); anchor-free (FCOS, DETR) is the modern line.
Augmentation defaultsFlips, random resized crop, color jitter 0.2, rotation +/-15 deg, RandAugment for scale.

5. Attention & Transformer Architecture Reference

Scaled dot-productAttention(Q,K,V) = softmax(Q K^T / sqrt(d_k)) V - the 1/sqrt(d) keeps logits ~N(0,1)
ComplexityO(n^2 d) compute, O(n^2) memory in sequence length n - the quadratic tax FlashAttention attacks
Multi-headhead_dim = d_model / h; per head, Q,K,V project to d_model; concat + W_O
Decoder params rule~ 12 x L x h^2 + V x h (4h^2 attention + 8h^2 FFN per layer)
KV cache / token2 x L x kv_heads x head_dim x bytes x tokens x batch
Configd_modelLayersHeadsFFNVocab
BERT-base (encoder)76812123072 (4x)30K WordPiece
GPT-2 XL (decoder)16004825640050K BPE
7B-class LLM (LLaMA-2 archetype)40963232 (head_dim 128)11008 (~2.7x)32K SentencePiece
13B-class512040401382432K
FFN blockLinear(h, 4h) -> GELU/SwiGLU -> Linear(4h, h); two-thirds of transformer parameters live here.
Norm placementPre-norm (LN before each sublayer) trains far more stably than post-norm; RMSNorm (Llama) drops mean-centering for speed.
Positional encodingSinusoidal/learned (2017) -> learned -> RoPE (Llama and most post-2023 LLMs), ALiBi; RoPE + NTK/YaRN scaling gives long context.
Seq2seq vs decoder-onlyEncoder-decoder (T5/BART) for translation & summarization; decoder-only (GPT lineage) for open-ended generation; encoders (BERT) for classification/embeddings.

6. LLM Pretraining, Scaling Laws & Decoding

Causal LM lossL(theta) = -(1/T) SUM_t log p_theta(x_t | x_1..x_(t-1)) - next-token cross-entropy
Training computeC ~ 6 N D FLOPs (N params, D tokens, fwd+bwd). 7B x 1T tokens ~ 4.2e22 FLOPs.
Chinchilla ruleCompute-optimal D ~ 20 x N tokens; N ~ C^0.5 and D ~ C^0.5. Chinchilla 70B/1.4T beat Gopher 280B/300B at equal compute on 67/70 evals.
Overtrained servingLlama-3 8B on ~15T tokens = ~1,900 tokens/param - deliberately past 20x because per-token inference cost dominates lifetime.
PerplexityPPL = exp(cross-entropy loss); comparable only for the same tokenizer + corpus
Decoding knobRangeGuidance
Temperature0 - 1+0-0.2 extraction/code; ~0.7 chat default; >1.2 creative, drifts
Top-k40 typicalCut tail; use with temp but avoid stacking both hard filters
Top-p (nucleus)0.9 - 0.95Dynamic tail cut; the usual single knob with temperature
Beam searchbeams 2-5Offline translation/summarization; poor for open chat (repetitive)
TokenizersBPE / WordPiece / SentencePiece; vocab 32K-256K; ~0.75 English words per token; multilingual coverage trades vocab size for per-language efficiency.
Post-training stackPretrain (next token) -> SFT (instructions) -> preference alignment (RLHF / DPO) -> reasoning RL (RLVR). Each stage is a different loss on the same weights.
Context windows4K -> 128K -> 1M+ (2024-2026); nominal != effective - check needle-in-a-haystack / RULER before trusting long context.

7. RAG, Vector Retrieval & Agents

Chunking defaults256-512 token chunks (embedders trained on short spans), overlap 10-20% (tune DOWN - heavy overlap doubles index and bill), recursive/structural splitting first, small-to-big (parent-child) retrieval.
EmbeddingsDims 384 (MiniLM) - 1536 - 3072; cosine == dot-product on normalized vectors; pick by MTEB rank for your language/domain.
Dense vs sparseBM25 nails exact terms, IDs and jargon; dense nails paraphrase and synonyms; hybrid + RRF fusion beats either alone - the production default.
RerankingBi-encoder recalls top 50-100 -> cross-encoder reranks to top 5-10; biggest quality-per-latency win in the RAG stack.
Retrieval evalsrecall@k, MRR, nDCG for search; faithfulness / answer-relevance (RAGAS-style) + a golden set for end-to-end.
ANN indexIdeaTuning knobsTrade-off
HNSWLayered proximity graphM 16-48; efSearch 64-512Recall 0.95+ at ms latency; RAM-hungry, slow builds
IVF-FlatCoarse centroids, scan bucketsnlist ~ sqrt(N); nprobeFast, tunable recall/speed; dead weight without quantization
IVF-PQProduct-quantized vectorspq m / code_size10-50x RAM cut, some recall loss
SCANNAsymmetric quantization (Google)-Strong recall/speed on big normalized vectors

Agent patterns that actually ship:

  • ReAct loop: Thought -> Action (tool call) -> Observation, repeated until Final Answer. Cap iterations, add timeouts.
  • Function calling = JSON-schema tool contracts; validate arguments before execution; least-privilege the tools.
  • Planning: CoT for linear reasoning, ToT for branching search, GoT for graph-structured aggregation.
  • Memory: buffer window -> rolling summary -> vector long-term store; MemGPT-style paging for beyond-context tasks.
  • Guardrails: read/write separation, human approval on irreversible actions, sandboxed code exec, prompt-injection filtering on retrieved content.

8. Fine-tuning, PEFT & Quantization

Full-FT memory~16 bytes/param (fp16 weights + grads + fp32 Adam states) + activations; a 7B model needs ~112 GB before the batch
LoRA updateh = W0 x + (alpha/r) B A x; A: r x d_in, B: d_out x r; trainable ~ r x (d_in + d_out) per targeted layer
LoRA defaultsr = 8-64 (task-dependent), alpha = 16-32 (~2r), target q,v first - then all linear layers; merge W + BA for zero-latency serving
PrecisionBytes/param7B modelNotes
FP32428 GBTraining master weights / research baselines
FP16 / BF16214 GBStandard inference; BF16 keeps FP32 exponent range
INT817 GBSmoothQuant / llama.cpp Q8; near-lossless
INT4 (NF4)0.53.5 GBGPTQ (Hessian-aware), AWQ (activation-aware), bitsandbytes NF4 - the QLoRA base
Prompt vs RAG vs fine-tunePrompt first. RAG to inject fresh/private KNOWLEDGE. Fine-tune to change FORM - style, format, tone, tool behavior, latency cost. Never fine-tune to memorize facts.
DistillationTeacher outputs -> soft-label / CoT-trace training of a student; GPT-4 -> 4o-mini and DeepSeek-R1 -> small reasoners are the canonical wins.
Pruning60-90% sparsity + retrain historically; in the LLM era, quantization captured most of the win.

9. Diffusion, GANs, VAEs & Multimodal

FamilyObjectiveStrengthWeakness
VAEELBO: reconstruction + KL to priorSmooth latent space, cheapBlurry samples
GANMinimax generator vs discriminatorSharp, fast inferenceUnstable training, mode collapse
DiffusionIterative denoising (score matching)Best quality, stable, controllableSlow sampling (fixed by distillation)
AutoregressiveCross-entropy over tokenized spaceUnifies text/audio/image pipelinesSequential generation cost
DDPM -> fast samplingForward: T=1000 beta-scheduled noise steps. Reverse learned denoiser samples in 20-50 steps with DDIM; step-distillation (LCM/Turbo) reaches 1-4.
Latent diffusionVAE f8: 512x512x3 image -> 64x64x4 latent; UNet (~860M in SD1.5) denoises in latent space; CLIP text encoder conditions via cross-attention.
Classifier-free guidanceout = uncond + s (cond - uncond); s = 5-10 (too high burns saturation/artifacts); flow matching (SD3) replaces the UNet noise schedule with straight transport.
SpeechWhisper: log-mel -> seq2seq transformer, ~39 languages zero-shot; modern TTS is autoregressive codec LMs (VALL-E) or flow/diffusion (F5, XTTS).
Vision-language wiringViT image encoder + projection layer feeds visual tokens into the LLM context (LLaVA-style), or native interleaved training (GPT-4o/Gemini class).

10. Reinforcement Learning & RLHF

MDP + return(S, A, P, R, gamma); G_t = SUM_k gamma^k r_(t+k+1)
Bellman (optimal V)V*(s) = max_a E[ r + gamma V*(s’) ]
Q-learning updateQ(s,a) <- Q(s,a) + alpha [ r + gamma max_a’ Q(s’,a’) - Q(s,a) ]
PPO objectiveE[ min( rho A, clip(rho, 1-eps, 1+eps) A ) ]; eps = 0.2, GAE lambda ~0.95, KL-to-reference penalty
DPO loss-log sigmoid( beta [ log pi(y_w|x)/ref - log pi(y_l|x)/ref ] ); beta ~ 0.1 - preferences without a reward model or PPO
RLHF pipelineSFT -> reward model on pairwise human preferences -> PPO against the RM; ~4 model copies in memory (policy, ref, RM, value).
RLVR / GRPORewards from verifiable checks (math answers, unit tests) + group-relative advantages (no critic) - the DeepSeek-R1 / o-series reasoning recipe.
Exploration vs exploitationepsilon-greedy, UCB, entropy bonuses; offline RL exists because real interaction is expensive/dangerous.

11. Modern Inference Architectures & Serving Numbers

MoETop-2 expert routing per token. Mixtral 8x7B: 47B total / 13B active params - dense-class quality at active-param cost, total-param memory, load-balancing aux loss.
GQA / MQAShare K/V heads across query heads (Llama-2 70B: 64 Q / 8 KV) - 8x smaller KV cache for a sliver of quality.
KV-cache sizing~0.13 MB/token for an 8B-class FP16 model; 100K-token context ~13 GB per sequence x batch. The cache, not the weights, caps long-context concurrency.
FlashAttentionExact attention with IO-aware tiling; O(n) memory instead of O(n^2), 2-4x speed; table stakes in every 2024+ stack.
Speculative decodingSmall drafter proposes k tokens, big model verifies in parallel - 2-3x decode speedup, output distribution provably unchanged.
PagedAttention / vLLMOS-style paged KV blocks kill fragmentation (~zero waste) + continuous batching: 2-5x throughput vs naive HF serving.
Serving targetsChat UX: time-to-first-token < 500 ms, decode >= 30 tok/s/user; throughput side: maximize tokens/s/GPU with batching + FP8/INT4.

Pretraining compute sanity check

C ~ 6ND FLOPs; A100 ~312 TFLOPS BF16, H100 ~989. At ~40% MFU, 7B x 1T tokens ~ 4.2e22 FLOPs ~ 3,400 A100-days. Always divide by MFU, not peak - peak-only math underestimates cost 2-3x.

12. MLOps & Production Deployment Checklists

Ship-readiness checklist:

  • Every run tracked: params, metrics, data version, git SHA (MLflow / W&B) - an untracked experiment is a re-run waiting to happen.
  • Feature parity between training and serving (offline vs online feature store) is the #1 silent production bug class.
  • Deployment ladder: shadow (mirror traffic, zero user risk) -> canary 1-5% with auto-rollback -> A/B on business metrics.
  • Monitor: input drift (PSI > 0.2 = investigate), prediction distribution shift, latency p50/p95, cost per call, periodic eval quality.
  • CI/CD gates: schema tests -> one-batch overfit test -> offline eval threshold -> shadow -> staged promote from the model registry.
  • LLM ops add: prompt versioning, token/cost budgets per route, semantic caching, guardrail checks in the hot path, eval-suite regressions.
Latency vs throughputSingle inference: latency-bound. Batching: turns memory-bound decode into throughput. Dynamic batch + continuous batching is the serving lever.
Exchange formatsONNX for portability, TensorRT/OpenVINO for vendor-specific kernels; export quantized to hit the memory budget.

13. Alignment, Safety, Privacy & Fairness

Fairness criterionDefinitionTension
Demographic parityP(pred +1 | group A) = P(pred +1 | group B)Conflicts with calibration and merit-based labels
Equalized oddsEqual TPR AND equal FPR across groupsHard when base rates genuinely differ
Equal opportunityEqual TPR onlyIgnores false-positive harms
Calibration within groupPredicted p == observed rate per groupCompatible with unequal outcomes
Differential privacy(eps, delta)-DP adds calibrated noise (DP-SGD, output perturbation). eps < 1 strong and costly, 1-10 typical deployments, >10 weak. Privacy-accuracy trade is real.
Attack surfaceAdversarial examples (FGSM/PGD), data poisoning, membership inference, extraction, and for LLMs: jailbreaks + prompt injection (goal hijack via untrusted content).
Defense in depthSeparate untrusted content from instructions, output-schema validation, least-privilege tools, red-team pre- and post-launch, eval-gated rollouts.
Mechanistic interpretabilitySuperposition, probing classifiers, activation patching / causal tracing, sparse autoencoders extracting human-readable features.
GovernanceEU AI Act risk tiers (prohibited / high-risk / transparency / minimal); NIST AI RMF as the voluntary framing most teams operationalize.

14. Course Map: 17 Phases of the AI & ML Curriculum

Every formula and number above is taught interactively - with visuals, quizzes and production case studies - across these phases.

1Mathematical Foundations for Machine Learning
24 topicsBuild the linear algebra, calculus, and probability intuition that every ML algorithm secretly runs on.
2Data & Programming Foundations
15 topicsMaster NumPy, pandas, visualization, and the data hygiene habits that decide whether your models succeed or silently fail.
3Classical Machine Learning
35 topicsLinear and logistic regression, trees, ensembles, clustering, and the evaluation discipline behind trustworthy models.
4Neural Network Fundamentals
20 topicsFrom perceptrons to training loops: activations, loss, backprop, optimizers, initialization, and modern normalization.
5Convolutional Neural Networks & Vision
15 topicsConvolutions, pooling, landmark architectures, and modern vision from ViT to detection and segmentation.
6Sequence Models & Attention
16 topicsRNNs, LSTMs, embeddings, seq2seq, and the attention mechanism that reset the entire field.
7NLP & Word Representations
15 topicsTokenization, classical text pipelines, and the transformer-based NLP stack from BERT to sentence embeddings.
8Large Language Models
20 topicsThe GPT lineage: pretraining, scaling laws, in-context learning, decoding, prompting, and embeddings at scale.
9Retrieval, Memory & Agents
14 topicsRAG pipelines, vector databases, tool use, agent architectures, and the protocols wiring LLMs to the world.
10Fine-tuning & Adaptation
9 topicsLoRA, QLoRA, adapters, quantization, distillation, and pruning for adapting powerful models cheaply.
11Generative AI Beyond Text
12 topicsDiffusion, GANs, VAEs, and multimodal generation across images, audio, video, and code.
12Reinforcement Learning
13 topicsMDPs, policy and value methods, RLHF, and the RL-from-verifiable-rewards engines behind reasoning models.
13Modern Architecture Advances
11 topicsMixture-of-Experts, FlashAttention, RoPE, GQA, KV caching, speculative decoding, long context, and Mamba.
14MLOps & Deployment
17 topicsTraining frameworks, GPU serving, vLLM, experiment tracking, CI/CD for ML, monitoring, and edge inference.
15AI Alignment, Safety & Ethics
14 topicsBias, robustness, adversarial attacks, interpretability, guardrails, regulation, and responsible deployment.
16Current Frontiers & Trends (2024-2026)
18 topicsReasoning models, test-time compute, small language models, multimodal agents, and the open-vs-closed race.
17AI Engineering & Agentic Development Tooling
52 topicsThe AI engineer stack: LLM APIs, prompt systems, evals, agent frameworks, coding copilots, and career paths.

These are the summary cards — the course is the story

Every matrix above expands into full interactive lessons with diagrams, quizzes, and production case studies across 17 phases and 320 topics.

Explore the AI & ML Course