PHASE 3 CURRICULUM

Classical Machine Learning

Progress0 of 35 (0%)

Classical algorithms still win most real-world tabular problems.

Key Architectural Domains & Syllabus
This phase covers regression in depth
logistic regression as a classifier
decision trees
random forests
gradient boosting (XGBoost, LightGBM)
SVMs
k-NN
naive Bayes
k-means and DBSCAN clustering
PCA and dimensionality reduction
bias-variance
regularization
metrics (precision/recall, ROC, PR-AUC)
cross-validation
hyperparameter tuning
scikit-learn pipelines
35 In-Depth Topics ~280 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 3

0 of 35 completed
#40Supervised LearningBeginnerFREE

Supervised learning is teaching with an answer key: you show a model many examples that each have the right label, and it learns a rule. The real goal is not to ace the examples you showed, but to get NEW examples right. This topic covers the labels, the two task families, and the train/validate/test contract that makes a score trustworthy.

10 min read•3 Quiz Questions
#41Unsupervised LearningBeginnerFREE

Unsupervised learning finds structure in data that has NO answer key. It groups similar things (clustering), compresses many features into a few (dimensionality reduction), flags the weird ones (anomaly detection), and produces embeddings. The hard part is proving a result is good when there is no label to check against.

11 min read•3 Quiz Questions
#42Semi-supervised LearningIntermediateFREE

Semi-supervised learning uses a small set of human-labeled examples plus a huge unlabeled pool to cut annotation cost. The model guesses labels for the unlabeled data (pseudo-labels) and learns from them. It saves money when the data forms clean clusters, and it collapses from confirmation bias when the model confidently trusts its own wrong guesses.

11 min read•3 Quiz Questions
#43Self-supervised LearningIntermediateFREE

Self-supervised learning makes its own labels by hiding or splitting the input and asking the model to guess the missing piece — no humans needed. The invented goal (a pretext task) is not the point; the understanding the model builds to solve it (the representation / embeddings) is. This is the foundation of every modern foundation model, including LLMs.

12 min read•3 Quiz Questions

The two basic prediction jobs: regression predicts a NUMBER (how much), classification picks a CATEGORY (which one). The type of your target decides the loss, the metric, and whether a decision threshold even exists. This topic walks through both, shows how to convert a problem from one to the other, and explains the tricks (binning, probabilities, counts, ordered labels) that connect them.

12 min read•3 Quiz Questions
#45Linear RegressionBeginnerFREE

The oldest useful model in machine learning: you predict a number as base + weights x features, and fit those weights by least squares. It has an exact closed-form solve, clean diagnostics (R2, RMSE), and strict assumptions — and it fails silently on collinear features and on extrapolation.

10 min read•3 Quiz Questions
#46Logistic RegressionBeginner PRO

A linear model for probabilities. It fits the log-odds as a straight line, squashes that line into a probability with the sigmoid, and trains it on cross-entropy. There is no closed-form solve, but the objective is convex and the decision boundary stays auditable.

10 min read•3 Quiz Questions
#47The Sigmoid FunctionBeginner PRO

The S-shaped curve that turns any unbounded score into a probability between 0 and 1. Its clever derivative shortcut made it the original neural-network activation, but its flat saturated ends vanish gradients — which is why ReLU-family activations replaced it inside hidden layers while it stayed on output heads and gates.

11 min read•3 Quiz Questions
#48Loss and Cost FunctionsIntermediate PRO

What a model actually optimizes. The loss is the penalty for one mistake, the cost adds a size penalty on the weights, and the metric is just what humans read. Choosing the loss encodes what you believe about the noise, how much you fear outliers, and whether you need honest probabilities.

10 min read•3 Quiz Questions
#49Gradient DescentIntermediate PRO

The general-purpose optimizer behind almost all of machine learning. When a cost function is too big or too tangled to solve in one shot, you minimize it by repeatedly stepping opposite the gradient (w ← w − η∇J). A first-order argument guarantees a local decrease; curvature bounds the safe step size and the condition number κ = L/μ sets the convergence speed — which is why standardizing features, preconditioning, momentum, and Adam all exist.

10 min read•3 Quiz Questions
#50Stochastic Gradient DescentIntermediate PRO

Plain gradient descent wants the exact average slope over every training example before each step — impossibly slow with millions of examples. Stochastic Gradient Descent (SGD) instead stirs the pot and tastes one spoonful (or a small cup of m spoonfuls): the direction is noisy, the average direction is exactly right, and every step gets n times cheaper. The bonus: that noise often helps the model generalize.

12 min read•3 Quiz Questions
#51Learning RateIntermediate PRO

The gradient tells the model which way to walk, but not how far. That distance is the learning rate η — the single most sensitive knob in machine learning. Too small and training crawls; too large and it oscillates, then explodes to NaN. This topic covers why, what each failure looks like on a loss curve, the schedules (warmup + decay) that fix both, and cheap ways to find the right value.

12 min read•3 Quiz Questions
#52OverfittingIntermediate PRO

A model overfits when it memorizes the training set instead of learning the process behind it: it aces the data it studied and fails on data it has never seen. The symptom is the train/validation gap. This topic gives you the diagnostic (learning curves), the four mechanisms that cause the gap (capacity, data, leakage, evaluation abuse), and the ranked list of remedies that actually work.

12 min read•3 Quiz Questions
#53UnderfittingIntermediate PRO

Underfitting is the failure mode teams under-diagnose: the model is too simple (or under-trained, or starved of signal) to even fit its own training data. The tell is a poor TRAINING score with a small train/validation gap, and flat parallel learning curves. The fixes go the opposite way from overfitting: add features, capacity, and training — remove regularization — in that order.

12 min read•3 Quiz Questions
#54Bias-Variance TradeoffAdvanced PRO

Every prediction error at a point splits exactly into three parts: Bias² (how far the average model misses the truth), Variance (how much the model thrashes when you retrain it on different data), and irreducible noise σ² (the jitter in y itself). Model complexity trades the first two against each other, making total error U-shaped. This decomposition is the equation behind why Topic 52 (overfitting) and Topic 53 (underfitting) happen — and what actually fixes each.

13 min read•3 Quiz Questions
#55RegularizationAdvanced PRO

A model that tries too hard memorizes noise. Regularization fixes that by charging the model a "simplicity tax" on its weights: L2 shrinks everything smoothly, L1 zeroes coefficients outright, and early stopping, dropout, and augmentation apply the same trade in different ways.

12 min read•3 Quiz Questions
#56Decision TreesIntermediate PRO

A decision tree is a flowchart of if-then questions the computer builds for itself by playing 20 Questions with the data. This topic covers how CART picks each question, why one tree is wildly unstable, and how pruning and ensembles fix that.

11 min read•3 Quiz Questions
#57Entropy & Information GainIntermediate PRO

Entropy counts how surprising a node's labels are, in bits. Information gain is how much a candidate question removes that surprise — the score that picks every split in ID3/C4.5 trees — plus the gain-ratio fix for the trap where high-cardinality columns "win" by memorizing.

11 min read•3 Quiz Questions
#58Gini ImpurityBeginner PRO

Decision trees need one number to rate how messy a node is. Gini impurity is that number: the chance you would guess a random sample's label wrong if you drew a fake label from the node's own mix. This topic computes it, weights it across a split, and shows how CART uses it to pick questions.

11 min read•3 Quiz Questions
#59Random ForestsIntermediate PRO

One decision tree is a moody expert: ask it twice on slightly different data and it changes its mind. A random forest grows hundreds of deliberately different trees and lets them vote — averaging cancels their individual mood swings and yields one of the most robust general-purpose tabular learners.

11 min read•3 Quiz Questions

Gradient boosting trains many small decision trees one after another. Each new tree only fixes the mistakes the earlier trees left behind. That simple loop is secretly gradient descent on the loss (Friedman, 2001) — and the engineering around it (XGBoost, LightGBM, CatBoost) is why these models still dominate tabular data.

12 min read•3 Quiz Questions
#61Support Vector MachinesIntermediate PRO

A Support Vector Machine is a line (or hyperplane) drawn to leave the widest possible gap between two classes. You will see why a fat gap means safer predictions, how the C knob prices mistakes, why the hinge loss is the same idea in disguise, and why only a few "support vectors" are the whole model.

11 min read•3 Quiz Questions
#62The Kernel TrickAdvanced PRO

The kernel trick lets a straight-line algorithm solve curved problems without ever computing the curved coordinates: replace every dot product by a similarity function k(x, z). You will see why this is exact (not an approximation), what makes a kernel valid (Mercer/PSD), how the RBF gamma dial controls wiggle, and which other algorithms can be kernelized.

11 min read•3 Quiz Questions
#63K-Nearest NeighborsBeginner PRO

K-Nearest Neighbors never trains. It stores the labeled data and answers any question by looking at the k most similar past examples: they vote (classification) or average (regression). You will see how to pick k, why distances demand scaling, which metric fits which data, and why "nearest" quietly breaks in high dimensions.

11 min read•3 Quiz Questions
#64Naive BayesBeginner PRO

Naive Bayes asks: how would each class *generate* this email? It counts word frequencies per class, plugs them into Bayes' theorem, and pretends words are independent given the class — a gloriously false shortcut that turns an impossible counting problem into two tally tables, solved in a linear pass.

11 min read•3 Quiz Questions
#65K-Means ClusteringBeginner PRO

K-means finds k groups in unlabeled data by repeating two simple steps: assign every point to its nearest center, then move each center to the average of its points. The loop always converges — but only to a local minimum of inertia — so careful seeding (k-means++) and restarts are what make it trustworthy.

13 min read•3 Quiz Questions
#66DBSCANIntermediate PRO

DBSCAN groups points that live in crowded neighborhoods, connected by chains of dense areas — so clusters can have any shape, the number of clusters is something the data tells you, and lonely leftover points get an honest "outlier" label instead of a forced home.

12 min read•3 Quiz Questions

PCA rotates your data so the directions with the most spread become the first axes. Keep those, drop the flat ones, and you compress, denoise, and visualize — with the exact linear algebra (eigenvectors of the covariance, or SVD of the centered matrix) doing the finding.

13 min read•3 Quiz Questions
#68t-SNE & UMAPAdvanced PRO

t-SNE and UMAP turn high-dimensional data into a 2-D picture by keeping one promise: points that are neighbors up high stay neighbors on the page. That promise is why the plots are gorgeous — and why distances BETWEEN the blobs you see are not real.

13 min read•3 Quiz Questions

One train/test split can lie — an unlucky draw moves your score several points. k-fold cross-validation rotates the test role so every row is held out exactly once, and the matching splitters (stratified, grouped, temporal) make the rotation mirror real deployment instead of faking it.

13 min read•3 Quiz Questions
#70The Confusion MatrixBeginner PRO

The confusion matrix is the 2x2 table underneath every classification metric: TP/FP/FN/TN counts, the metric derivations, why accuracy lies on imbalanced data, and how the two kinds of error become a business decision.

9 min read•3 Quiz Questions
#71Precision, Recall & F1Beginner PRO

Precision and recall are the two error rates that survive imbalance: precision asks "can I trust an alarm?", recall asks "did I catch them all?". F1 is their honest harmonic compromise, and the precision-recall curve is how you pick the operating point the business actually wants.

9 min read•3 Quiz Questions
#72ROC Curve & AUCIntermediate PRO

The ROC curve shows what your model would do at EVERY threshold, not just one: plot catch-rate against false-alarm-rate as the cut-off slides, and the area under it (AUC) becomes the probability that a random positive outscores a random negative. Threshold-free and comparable across base rates — which is also exactly what it hides.

10 min read•3 Quiz Questions

Models come with dials you set before training — learning rate, depth, regularization. Tuning is searching those dials for the best validation score, and each try costs a full training run. This topic covers where to look (the search space), how to pick candidates (grid, random, Bayesian/TPE), and how to stop bad candidates early without fooling yourself.

10 min read•3 Quiz Questions

Combining several mediocre models can beat one carefully tuned model — if their mistakes differ. Bagging averages independent peers to cancel noise (variance), boosting chains correctors to chase the residual signal (bias), and stacking trains a small meta-model to learn which base model to trust where.

11 min read•3 Quiz Questions