Formulas, Scaling Laws & Decision Matrices for AI/ML Interviews
The quick-reference companion for the 320-topic AI & ML curriculum: ML math essentials, classical complexity tables, training defaults, transformer reference specs, LLM scaling laws, RAG numbers, fine-tuning & quantization charts, and MLOps checklists.
1. Mathematical Foundations for ML
| Bayes’ theorem | P(A|B) = P(B|A) P(A) / P(B) posterior = likelihood x prior / evidence |
|---|---|
| Entropy | H(p) = -SUM_x p(x) log2 p(x) minimum mean bits to encode draws from p |
| Cross-entropy | H(p,q) = -SUM_x p(x) log q(x) = H(p) + KL(p||q) the classification loss |
| KL divergence | KL(p||q) = SUM_x p log (p/q) >= 0, asymmetric VAE prior, distillation, PPO/KL-to-reference penalty |
| Gradient descent | theta <- theta - lr * grad_theta L(theta) backprop = chain rule over the graph |
| MLE vs MAP | MLE: argmax log p(data|theta); MAP: argmax log p(data|theta) + log p(theta) |
| L1 vs L2 regularization | L1 (lasso) drives weights to exactly zero - sparse feature selection. L2 (ridge) shrinks smoothly toward zero. Elastic net mixes both. Typical lambda 0.001 to 1. This is why L1 = lasso and L2 = ridge. |
|---|---|
| Variance vs bias | Total error = bias^2 + variance + irreducible noise. Underfit = high bias; overfit = high variance. Regularization and data both attack variance; capacity attacks bias. |
2. Data Work & Classical ML Complexity
| Algorithm | Train | Predict | Sweet Spot |
|---|---|---|---|
| Linear regression | O(n d^2) closed form (normal eq.) | O(d) | Baselines + interpretability on tabular data |
| Logistic regression | Iterative GD, ~O(n d) / epoch | O(d) | Calibrated probabilities, huge sparse text features |
| Decision tree | O(n d log n) | O(depth) | Fast prototyping, explainable splits; overfits deep |
| Random forest | trees x tree cost | trees x O(depth) | Robust tabular default; OOB error is free validation |
| XGBoost / LightGBM | iters x O(n d log n) | O(iters x depth) | Still wins most Kaggle tabular competitions |
| SVM (RBF kernel) | O(n^2 - n^3) | O(#SVs x d) | Small/medium clean datasets; scales features first |
| k-NN | O(1) (stores data) | O(n d) | Tiny datasets, embedding similarity, k = 3-15 |
| Naive Bayes | O(n d) | O(d) | Text-classification baseline, streaming updates |
| k-means | O(n k t) per run (t ~ 10-100 iters) | O(k d) | Spherical clusters, quantization, color reduction |
| DBSCAN | O(n log n) with spatial index | - | Arbitrary shapes + noise/outlier detection |
| PCA (SVD) | O(n d^2) | O(k d) | 2-20 components capture 50-95% variance for viz |
| Precision / Recall / F1 | P = TP/(TP+FP); R = TP/(TP+FN); F1 = 2PR/(P+R) |
|---|---|
| ROC-AUC vs PR-AUC | ROC survives class imbalance (few FPs); report PR-AUC when positives < 1% |
| Splits | 80/10/10 or 60/20/20; stratified for imbalance; walk-forward (never shuffle) for time series |
Data leakage kills silently
Standardizers fit on the FULL dataset, target-encoded features, timestamps after the label event - all inflate validation scores and collapse in production. Fit every transform on the train split only, inside a cross-validation pipeline.
3. Neural Network Training Reference
| Activation | Formula | Where it lives |
|---|---|---|
| ReLU | max(0, x) | CNN/MLP hidden layers; cheap, dying-ReLU risk; He init |
| GELU | x * Phi(x) | Transformers and LLMs (smooth ReLU); modern default |
| SiLU / Swish | x * sigmoid(x) | LLaMA-style FFNs, EfficientNet |
| Sigmoid | 1 / (1 + e^-x) | Binary outputs + LSTM/GRU gates; saturates |
| Tanh | 2*sigmoid(2x) - 1 | Recurrent hidden states; zero-centered |
| Softmax | exp(x_i) / SUM exp(x_j) | Multiclass heads; subtract max for numerical stability |
| Optimizer | Default LR | Notes |
|---|---|---|
| SGD + momentum 0.9 | 0.01 - 0.1 | Vision from scratch; cosine decay; best final accuracy when tuned |
| AdamW (0.9 / 0.999, eps 1e-8, wd 0.01-0.1) | 1e-4 - 3e-4 pretrain; 1e-5 - 2e-5 fine-tune | Decoupled weight decay; the transformer workhorse |
| Adafactor | 1e-3 - 3e-3 | Near-zero optimizer memory for huge seq2seq models |
| BatchNorm vs LayerNorm | BN: normalize per-feature across the batch (CNNs; needs a batch). LN: normalize per-sample across features (transformers/RNNs; works at batch size 1). |
|---|---|
| Dropout | 0.1 in transformer blocks, 0.2-0.5 in overfit MLP heads; train-only (inference is the expectation). |
| Warmup + clipping | Linear LR warmup for the first 0.5-3% of steps; clip grad-norm to 1.0 for RNNs/transformers. |
| Mixed precision | BF16/FP16 forward + FP32 master weights + loss scaling; ~2x throughput on tensor-core GPUs. |
4. CNNs, Vision & Transfer Learning
| Conv output size | O = floor((W - K + 2P) / S) + 1 per spatial dimension |
|---|---|
| Conv parameters | K x K x C_in x C_out + C_out (3x3, 64->64 filters = 3,712 params) |
| Stacked 3x3s | Two 3x3 layers see a 5x5 field with 8/9 the params and one extra nonlinearity |
| Architecture | Year | Params | Why it matters |
|---|---|---|---|
| LeNet-5 | 1998 | 60K | First deployed CNN (bankcheck OCR) |
| AlexNet | 2012 | 60M | GPU + ReLU + dropout ImageNet moment |
| VGG-16 | 2014 | 138M | Depth via repeated 3x3 blocks; simple recipe |
| ResNet-50 | 2015 | 25M | Skip connections: y = F(x) + x, 100+ layers trainable |
| EfficientNet-B0 | 2019 | 5M | Compound scaling of depth x width x resolution |
| ViT-B/16 | 2020 | 86M | Patches-as-tokens transformer; needs data or distillation |
| Transfer-learning recipe | Frozen backbone + new head at lr 1e-4; unfreeze top blocks at lr 1e-5. Under ~10K images, frozen features often beat full fine-tunes. |
|---|---|
| Detection families | Two-stage (Faster R-CNN: accurate, region proposals) vs one-stage (YOLO/SSD: real-time); anchor-free (FCOS, DETR) is the modern line. |
| Augmentation defaults | Flips, random resized crop, color jitter 0.2, rotation +/-15 deg, RandAugment for scale. |
5. Attention & Transformer Architecture Reference
| Scaled dot-product | Attention(Q,K,V) = softmax(Q K^T / sqrt(d_k)) V - the 1/sqrt(d) keeps logits ~N(0,1) |
|---|---|
| Complexity | O(n^2 d) compute, O(n^2) memory in sequence length n - the quadratic tax FlashAttention attacks |
| Multi-head | head_dim = d_model / h; per head, Q,K,V project to d_model; concat + W_O |
| Decoder params rule | ~ 12 x L x h^2 + V x h (4h^2 attention + 8h^2 FFN per layer) |
| KV cache / token | 2 x L x kv_heads x head_dim x bytes x tokens x batch |
| Config | d_model | Layers | Heads | FFN | Vocab |
|---|---|---|---|---|---|
| BERT-base (encoder) | 768 | 12 | 12 | 3072 (4x) | 30K WordPiece |
| GPT-2 XL (decoder) | 1600 | 48 | 25 | 6400 | 50K BPE |
| 7B-class LLM (LLaMA-2 archetype) | 4096 | 32 | 32 (head_dim 128) | 11008 (~2.7x) | 32K SentencePiece |
| 13B-class | 5120 | 40 | 40 | 13824 | 32K |
| FFN block | Linear(h, 4h) -> GELU/SwiGLU -> Linear(4h, h); two-thirds of transformer parameters live here. |
|---|---|
| Norm placement | Pre-norm (LN before each sublayer) trains far more stably than post-norm; RMSNorm (Llama) drops mean-centering for speed. |
| Positional encoding | Sinusoidal/learned (2017) -> learned -> RoPE (Llama and most post-2023 LLMs), ALiBi; RoPE + NTK/YaRN scaling gives long context. |
| Seq2seq vs decoder-only | Encoder-decoder (T5/BART) for translation & summarization; decoder-only (GPT lineage) for open-ended generation; encoders (BERT) for classification/embeddings. |
6. LLM Pretraining, Scaling Laws & Decoding
| Causal LM loss | L(theta) = -(1/T) SUM_t log p_theta(x_t | x_1..x_(t-1)) - next-token cross-entropy |
|---|---|
| Training compute | C ~ 6 N D FLOPs (N params, D tokens, fwd+bwd). 7B x 1T tokens ~ 4.2e22 FLOPs. |
| Chinchilla rule | Compute-optimal D ~ 20 x N tokens; N ~ C^0.5 and D ~ C^0.5. Chinchilla 70B/1.4T beat Gopher 280B/300B at equal compute on 67/70 evals. |
| Overtrained serving | Llama-3 8B on ~15T tokens = ~1,900 tokens/param - deliberately past 20x because per-token inference cost dominates lifetime. |
| Perplexity | PPL = exp(cross-entropy loss); comparable only for the same tokenizer + corpus |
| Decoding knob | Range | Guidance |
|---|---|---|
| Temperature | 0 - 1+ | 0-0.2 extraction/code; ~0.7 chat default; >1.2 creative, drifts |
| Top-k | 40 typical | Cut tail; use with temp but avoid stacking both hard filters |
| Top-p (nucleus) | 0.9 - 0.95 | Dynamic tail cut; the usual single knob with temperature |
| Beam search | beams 2-5 | Offline translation/summarization; poor for open chat (repetitive) |
| Tokenizers | BPE / WordPiece / SentencePiece; vocab 32K-256K; ~0.75 English words per token; multilingual coverage trades vocab size for per-language efficiency. |
|---|---|
| Post-training stack | Pretrain (next token) -> SFT (instructions) -> preference alignment (RLHF / DPO) -> reasoning RL (RLVR). Each stage is a different loss on the same weights. |
| Context windows | 4K -> 128K -> 1M+ (2024-2026); nominal != effective - check needle-in-a-haystack / RULER before trusting long context. |
7. RAG, Vector Retrieval & Agents
| Chunking defaults | 256-512 token chunks (embedders trained on short spans), overlap 10-20% (tune DOWN - heavy overlap doubles index and bill), recursive/structural splitting first, small-to-big (parent-child) retrieval. |
|---|---|
| Embeddings | Dims 384 (MiniLM) - 1536 - 3072; cosine == dot-product on normalized vectors; pick by MTEB rank for your language/domain. |
| Dense vs sparse | BM25 nails exact terms, IDs and jargon; dense nails paraphrase and synonyms; hybrid + RRF fusion beats either alone - the production default. |
| Reranking | Bi-encoder recalls top 50-100 -> cross-encoder reranks to top 5-10; biggest quality-per-latency win in the RAG stack. |
| Retrieval evals | recall@k, MRR, nDCG for search; faithfulness / answer-relevance (RAGAS-style) + a golden set for end-to-end. |
| ANN index | Idea | Tuning knobs | Trade-off |
|---|---|---|---|
| HNSW | Layered proximity graph | M 16-48; efSearch 64-512 | Recall 0.95+ at ms latency; RAM-hungry, slow builds |
| IVF-Flat | Coarse centroids, scan buckets | nlist ~ sqrt(N); nprobe | Fast, tunable recall/speed; dead weight without quantization |
| IVF-PQ | Product-quantized vectors | pq m / code_size | 10-50x RAM cut, some recall loss |
| SCANN | Asymmetric quantization (Google) | - | Strong recall/speed on big normalized vectors |
Agent patterns that actually ship:
- ReAct loop: Thought -> Action (tool call) -> Observation, repeated until Final Answer. Cap iterations, add timeouts.
- Function calling = JSON-schema tool contracts; validate arguments before execution; least-privilege the tools.
- Planning: CoT for linear reasoning, ToT for branching search, GoT for graph-structured aggregation.
- Memory: buffer window -> rolling summary -> vector long-term store; MemGPT-style paging for beyond-context tasks.
- Guardrails: read/write separation, human approval on irreversible actions, sandboxed code exec, prompt-injection filtering on retrieved content.
8. Fine-tuning, PEFT & Quantization
| Full-FT memory | ~16 bytes/param (fp16 weights + grads + fp32 Adam states) + activations; a 7B model needs ~112 GB before the batch |
|---|---|
| LoRA update | h = W0 x + (alpha/r) B A x; A: r x d_in, B: d_out x r; trainable ~ r x (d_in + d_out) per targeted layer |
| LoRA defaults | r = 8-64 (task-dependent), alpha = 16-32 (~2r), target q,v first - then all linear layers; merge W + BA for zero-latency serving |
| Precision | Bytes/param | 7B model | Notes |
|---|---|---|---|
| FP32 | 4 | 28 GB | Training master weights / research baselines |
| FP16 / BF16 | 2 | 14 GB | Standard inference; BF16 keeps FP32 exponent range |
| INT8 | 1 | 7 GB | SmoothQuant / llama.cpp Q8; near-lossless |
| INT4 (NF4) | 0.5 | 3.5 GB | GPTQ (Hessian-aware), AWQ (activation-aware), bitsandbytes NF4 - the QLoRA base |
| Prompt vs RAG vs fine-tune | Prompt first. RAG to inject fresh/private KNOWLEDGE. Fine-tune to change FORM - style, format, tone, tool behavior, latency cost. Never fine-tune to memorize facts. |
|---|---|
| Distillation | Teacher outputs -> soft-label / CoT-trace training of a student; GPT-4 -> 4o-mini and DeepSeek-R1 -> small reasoners are the canonical wins. |
| Pruning | 60-90% sparsity + retrain historically; in the LLM era, quantization captured most of the win. |
9. Diffusion, GANs, VAEs & Multimodal
| Family | Objective | Strength | Weakness |
|---|---|---|---|
| VAE | ELBO: reconstruction + KL to prior | Smooth latent space, cheap | Blurry samples |
| GAN | Minimax generator vs discriminator | Sharp, fast inference | Unstable training, mode collapse |
| Diffusion | Iterative denoising (score matching) | Best quality, stable, controllable | Slow sampling (fixed by distillation) |
| Autoregressive | Cross-entropy over tokenized space | Unifies text/audio/image pipelines | Sequential generation cost |
| DDPM -> fast sampling | Forward: T=1000 beta-scheduled noise steps. Reverse learned denoiser samples in 20-50 steps with DDIM; step-distillation (LCM/Turbo) reaches 1-4. |
|---|---|
| Latent diffusion | VAE f8: 512x512x3 image -> 64x64x4 latent; UNet (~860M in SD1.5) denoises in latent space; CLIP text encoder conditions via cross-attention. |
| Classifier-free guidance | out = uncond + s (cond - uncond); s = 5-10 (too high burns saturation/artifacts); flow matching (SD3) replaces the UNet noise schedule with straight transport. |
| Speech | Whisper: log-mel -> seq2seq transformer, ~39 languages zero-shot; modern TTS is autoregressive codec LMs (VALL-E) or flow/diffusion (F5, XTTS). |
| Vision-language wiring | ViT image encoder + projection layer feeds visual tokens into the LLM context (LLaVA-style), or native interleaved training (GPT-4o/Gemini class). |
10. Reinforcement Learning & RLHF
| MDP + return | (S, A, P, R, gamma); G_t = SUM_k gamma^k r_(t+k+1) |
|---|---|
| Bellman (optimal V) | V*(s) = max_a E[ r + gamma V*(s’) ] |
| Q-learning update | Q(s,a) <- Q(s,a) + alpha [ r + gamma max_a’ Q(s’,a’) - Q(s,a) ] |
| PPO objective | E[ min( rho A, clip(rho, 1-eps, 1+eps) A ) ]; eps = 0.2, GAE lambda ~0.95, KL-to-reference penalty |
| DPO loss | -log sigmoid( beta [ log pi(y_w|x)/ref - log pi(y_l|x)/ref ] ); beta ~ 0.1 - preferences without a reward model or PPO |
| RLHF pipeline | SFT -> reward model on pairwise human preferences -> PPO against the RM; ~4 model copies in memory (policy, ref, RM, value). |
|---|---|
| RLVR / GRPO | Rewards from verifiable checks (math answers, unit tests) + group-relative advantages (no critic) - the DeepSeek-R1 / o-series reasoning recipe. |
| Exploration vs exploitation | epsilon-greedy, UCB, entropy bonuses; offline RL exists because real interaction is expensive/dangerous. |
11. Modern Inference Architectures & Serving Numbers
| MoE | Top-2 expert routing per token. Mixtral 8x7B: 47B total / 13B active params - dense-class quality at active-param cost, total-param memory, load-balancing aux loss. |
|---|---|
| GQA / MQA | Share K/V heads across query heads (Llama-2 70B: 64 Q / 8 KV) - 8x smaller KV cache for a sliver of quality. |
| KV-cache sizing | ~0.13 MB/token for an 8B-class FP16 model; 100K-token context ~13 GB per sequence x batch. The cache, not the weights, caps long-context concurrency. |
| FlashAttention | Exact attention with IO-aware tiling; O(n) memory instead of O(n^2), 2-4x speed; table stakes in every 2024+ stack. |
| Speculative decoding | Small drafter proposes k tokens, big model verifies in parallel - 2-3x decode speedup, output distribution provably unchanged. |
| PagedAttention / vLLM | OS-style paged KV blocks kill fragmentation (~zero waste) + continuous batching: 2-5x throughput vs naive HF serving. |
| Serving targets | Chat UX: time-to-first-token < 500 ms, decode >= 30 tok/s/user; throughput side: maximize tokens/s/GPU with batching + FP8/INT4. |
Pretraining compute sanity check
C ~ 6ND FLOPs; A100 ~312 TFLOPS BF16, H100 ~989. At ~40% MFU, 7B x 1T tokens ~ 4.2e22 FLOPs ~ 3,400 A100-days. Always divide by MFU, not peak - peak-only math underestimates cost 2-3x.
12. MLOps & Production Deployment Checklists
Ship-readiness checklist:
- Every run tracked: params, metrics, data version, git SHA (MLflow / W&B) - an untracked experiment is a re-run waiting to happen.
- Feature parity between training and serving (offline vs online feature store) is the #1 silent production bug class.
- Deployment ladder: shadow (mirror traffic, zero user risk) -> canary 1-5% with auto-rollback -> A/B on business metrics.
- Monitor: input drift (PSI > 0.2 = investigate), prediction distribution shift, latency p50/p95, cost per call, periodic eval quality.
- CI/CD gates: schema tests -> one-batch overfit test -> offline eval threshold -> shadow -> staged promote from the model registry.
- LLM ops add: prompt versioning, token/cost budgets per route, semantic caching, guardrail checks in the hot path, eval-suite regressions.
| Latency vs throughput | Single inference: latency-bound. Batching: turns memory-bound decode into throughput. Dynamic batch + continuous batching is the serving lever. |
|---|---|
| Exchange formats | ONNX for portability, TensorRT/OpenVINO for vendor-specific kernels; export quantized to hit the memory budget. |
13. Alignment, Safety, Privacy & Fairness
| Fairness criterion | Definition | Tension |
|---|---|---|
| Demographic parity | P(pred +1 | group A) = P(pred +1 | group B) | Conflicts with calibration and merit-based labels |
| Equalized odds | Equal TPR AND equal FPR across groups | Hard when base rates genuinely differ |
| Equal opportunity | Equal TPR only | Ignores false-positive harms |
| Calibration within group | Predicted p == observed rate per group | Compatible with unequal outcomes |
| Differential privacy | (eps, delta)-DP adds calibrated noise (DP-SGD, output perturbation). eps < 1 strong and costly, 1-10 typical deployments, >10 weak. Privacy-accuracy trade is real. |
|---|---|
| Attack surface | Adversarial examples (FGSM/PGD), data poisoning, membership inference, extraction, and for LLMs: jailbreaks + prompt injection (goal hijack via untrusted content). |
| Defense in depth | Separate untrusted content from instructions, output-schema validation, least-privilege tools, red-team pre- and post-launch, eval-gated rollouts. |
| Mechanistic interpretability | Superposition, probing classifiers, activation patching / causal tracing, sparse autoencoders extracting human-readable features. |
| Governance | EU AI Act risk tiers (prohibited / high-risk / transparency / minimal); NIST AI RMF as the voluntary framing most teams operationalize. |
14. Course Map: 17 Phases of the AI & ML Curriculum
Every formula and number above is taught interactively - with visuals, quizzes and production case studies - across these phases.
These are the summary cards — the course is the story
Every matrix above expands into full interactive lessons with diagrams, quizzes, and production case studies across 17 phases and 320 topics.
Explore the AI & ML Course