Mathematical Foundations for Machine Learning
Before touching any model, you need the language machines think in:
All Topics in Phase 1
0 of 24 completedA tensor is just a box of numbers. One number = scalar, a row = vector, a grid = matrix, a stack of grids = higher-rank tensor. The "rank" is how many coordinates you need to grab one number. In real ML, getting the shape right is most of the debugging.
Adding vectors means adding position by position. Multiplying by a scalar means stretching every position. From word arithmetic (king − man + woman ≈ queen) to residual connections and every optimizer update, these two tiny moves are the workhorses of ML.
The dot product takes two lists of numbers, multiplies matching entries, and adds them up. One number out. That single recipe gives you cosine similarity, vector-database search, every neuron, and the Q·K scores that let Transformers "attend".
Matrix multiplication is a stack of dot products: each output entry is one row meeting one column. The inner shapes must match, the order matters, and a whole neural-network forward pass is just a chain of these products — the single workload GPUs are built to run.
A transpose just swaps rows and columns: entry (i, j) becomes entry (j, i). In code it is almost free (the data stays put, only the labels change), yet it shows up everywhere — QKᵀ in attention, Wᵀ in every gradient, and half your shape errors are really missing transposes.
The identity matrix is a "do nothing" machine; the inverse is the matching "undo" machine. A matrix has an inverse only if it never squashes a direction to zero. When it does, ML practice stops fighting: you solve instead of inverting, and you add λ·I (ridge) so the undo button always works.
Every matrix hides a few lucky directions it can only stretch, never rotate. Those directions are eigenvectors, the stretch factors are eigenvalues — and this one idea powers PCA, PageRank, spectral clustering, loss-landscape curvature, and the stability of deep networks.
A norm is a formal way to ask "how big is this vector?" — and the formula you pick quietly decides everything: MSE vs MAE losses, why L1 penalties make weights exactly zero, what weight decay does, and how gradients get clipped.
The derivative answers one question: "if I nudge this input a tiny bit right here, what happens to the output?" It is the slope of the tangent line, the best local straight-line copy of a curve — and every training step in AI is built from it.
A model has millions of knobs and one score. A partial derivative answers the fundamental question of training — "if I nudge just THIS weight and freeze all the others, what happens to the loss?" Do it for every knob and you have the gradient.
A gradient is simply all the partial derivatives combined into one vector. It tells you which way is steepest uphill — so walking the opposite way (θ ← θ − lr·∇L) drops the loss fastest. That one idea powers all of deep learning.
The chain rule says rates of change multiply along nested functions; backpropagation is the chain rule executed backwards over a computation graph, one sweep for all gradients. This pair is THE algorithm that made deep learning possible — here it is from first principles, with code you can write from memory.
The gradient answers "which way is downhill?" — but not "what shape is the ground?" The Jacobian is the complete sensitivity table for vector-valued maps, and the Hessian is the curvature table that classifies flat spots as minima, maxima, or saddles and powers second-order optimization.
A convex loss surface is one smooth bowl: wherever you drop the marble, it rolls to the same bottom, so gradient descent comes with a global guarantee. Deep learning deliberately breaks convexity to buy expressiveness — and this topic explains why training still works.
A random variable is not a variable — it is a function that converts messy outcomes into numbers so math can do probability with them. Here are the three descriptors that capture its behavior — the PMF (discrete), the PDF (continuous), and the CDF (both) — and why the datasets, labels and batches in ML are random variables.
A probability distribution is a named shape that matches a data story: counting rare arrivals? Poisson. Averaged noise? Gaussian. Uncertainty about a probability itself? Beta. This tour covers the distributions every ML practitioner must know — and the deep fact that choosing an output distribution is choosing a loss function.
A distribution is a lot of information; expectation and variance squash it into two numbers: the center of mass (the probability-weighted average ML actually optimizes) and the spread (the noise that governs estimators, batch sizes, and the bias-variance tradeoff).
You usually know the probability of the evidence given a hypothesis; you want the probability of the hypothesis given the evidence. Bayes' theorem flips one into the other — prior × likelihood over evidence — and the medical-test example shows why ignoring base rates fools almost everyone.
You have data and a model with knobs. MLE says: turn the knobs to the setting that makes the data you actually saw most plausible. That one principle produces the frequency estimate, MSE, cross-entropy — and every neural-net loss — while carrying a few famous failure modes.
MAP asks: which parameters are most believable given BOTH the data and what you believed beforehand? The beautiful punchline: that "beforehand" belief is exactly a regularizer — a Gaussian prior becomes L2 weight decay, a Laplace prior becomes L1, a Beta prior becomes smoothing.
Surprisal says: rare news costs more, and the price is −log p. Entropy averages that cost over the whole distribution — and Shannon proved this "average surprise" is simultaneously your uncertainty and the best compression anyone can achieve.
Cross-entropy is the surprise bill: your model sets the odds, reality picks the outcome, and you pay −log(probability you assigned to what actually happened). Minimizing it means being honestly surprised as little as possible — and with softmax the gradient collapses to the beautiful error signal q − y.
KL divergence is the overcharge on your surprise bill: how much extra you pay because your belief q differs from reality p. It is never negative, zero only when the beliefs match, stubbornly asymmetric — and it quietly powers VAEs, distillation, and RLHF.
A computer cannot hold the product of hundreds of tiny probabilities — it silently rounds to zero. Logarithms fix this by turning those products into sums, and the log-sum-exp trick is the one maneuver that makes softmax and likelihoods overflow-proof. Almost all of ML numerics is built on this pair.