PHASE 4 CURRICULUM

Neural Network Fundamentals

Progress0 of 20 (0%)

The core deep-learning loop, built from first principles.

Key Architectural Domains & Syllabus
You will construct neurons and MLPs
choose activations
define loss functions
implement forward and backward passes with gradient descent variants (SGD, momentum, Adam)
handle learning-rate schedules
weight initialization
batch normalization and layer norm
dropout and regularization
learn to debug training curves like an engineer rather than guess at hyperparameters
20 In-Depth Topics ~160 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 4

0 of 20 completed

A perceptron is the smallest neural unit you can build: multiply each input by a weight, add them up with a bias, and answer 1 only if the score clears a bar. This page walks the math with tiny numbers, shows how it learns from mistakes, and explains why its XOR failure ignited the first AI winter.

12 min read•3 Quiz Questions

One perceptron draws one straight line and gets stuck on XOR (topic 75). Stack perceptrons into layers with nonlinearities between them, and each layer re-draws the map until the classes fall apart along a straight cut. This page shows the composition math, solves XOR with two hidden units, and explains what the universal approximation theorem does — and does not — promise.

12 min read•4 Quiz Questions

Between every matrix multiply in a neural network sits a tiny per-number rule called an activation. That rule decides whether your "deep" network is a universal function approximator or an expensive straight line. This page builds the selection criteria (saturation, zero-centering, sparsity, smoothness), tours the family from sigmoid to SwiGLU, and names the defaults used in 2026 production models.

12 min read•4 Quiz Questions

ReLU is one line: keep positive scores, zero out negatives. Topic 77 showed saturating sigmoids were strangling gradients in deep 1990s networks — ReLU's unbounded positive side, sparsity, and near-free compute trained the first very deep networks and helped ignite 2012. This page walks the function with numbers, the dying-ReLU failure mode, and the Leaky/PReLU/ELU patches it spawned.

12 min read•4 Quiz Questions

Sigmoid squashes any score into (0,1) — perfect for "what is the probability?" — and tanh is the same S-curve re-centered on zero. Both carried neural networks from 1986 to 2010, and both share one fatal habit: for extreme inputs they flatten out, and a flat function passes on zero gradient. This page walks their formulas, the saturation arithmetic that capped depth, the zig-zag from non-centered outputs, and the three places they still run production AI today.

12 min read•4 Quiz Questions

Your model spits out raw scores. Softmax turns any list of scores into probabilities that add up to 1 — the output layer of almost every classifier and the heart of attention. Learn its shift-invariance, the temperature knob, the subtract-the-max stability trick, and its famously clean cross-entropy gradient (p − y).

11 min read•3 Quiz Questions

ReLU is a light switch: full on or dead off, with a sharp kink at zero. GELU and SiLU replaced the switch with a dimmer — a smooth probability gate that keeps gradients informative near zero. That one smoothing choice powers BERT, GPT, ViT, LLaMA and Mistral.

11 min read•3 Quiz Questions

You have weights — now what? Forward propagation is the answer: layer by layer, each neuron computes a weighted sum plus bias, then passes it through a gate. Trace the pipeline with tiny numbers, see what gets cached for training, and count what one pass costs in FLOPs and memory.

11 min read•3 Quiz Questions

A network has 100 million weights; training needs a gradient for every one of them. Brute force would take 100 million forward passes. Backprop gets all of them for about twice the cost of one — by running the chain rule backward through the cached forward graph. Build the δ recursion, split the two gradients, and see what autodiff actually is.

12 min read•3 Quiz Questions

Training outcomes are decided before the first gradient flows: the distribution you draw initial weights from controls whether signals survive 50 layers. Zero init never breaks symmetry; too-small init fades the signal; too-large init explodes it. Xavier and He compute the exact variance that keeps every layer's gain at 1.

11 min read•3 Quiz Questions

Deep networks in the 2000s simply refused to learn: the learning signal is born at the last layer and shrinks a little more at every layer it passes back through, until early layers get nothing. This topic shows why that happens (gradients are multiplied, not added) and the toolkit that fixed it — ReLU, good init, normalization, residual connections, and gates.

12 min read•3 Quiz Questions

The mirror image of the vanishing gradient (Topic 85): if each layer multiplies the learning signal by more than 1, gradients grow exponentially until weights leap across space, numbers overflow, and your loss turns into NaN. Learn the fingerprint of a blowing-up run and the defenses: gradient clipping, normalization, careful init, and loss scaling.

12 min read•3 Quiz Questions

Four words you must never mix up: a batch is the handful of examples used per look, an iteration is one look, an update is one adjustment, an epoch is one full pass over the data. This topic fixes the grammar and shows why batch size is the most consequential number in any training run — it changes what the optimizer learns, not just how fast.

12 min read•3 Quiz Questions

Plain SGD forgets everything after each step, so it zig-zags uselessly in narrow valleys. Momentum gives the optimizer memory: it keeps a running velocity that averages recent gradients, which cancels oscillation and multiplies steady progress by up to 10×. This topic covers the heavy-ball formulas, why β = 0.9 means "remember ~10 steps", and Nesterov's look-ahead upgrade.

12 min read•3 Quiz Questions

Adam is the default optimizer of the deep-learning era because it answers a question momentum never could: how big should THIS knob's step be, given how this knob has behaved in the past? It does so with two running averages — of the gradients (momentum) and of the squared gradients (a per-knob size gauge) — plus a small correction for starting the averages at zero. This topic walks every symbol, the AdamW weight-decay fix, and the memory cost that shapes GPU clusters.

13 min read•3 Quiz Questions

The learning rate is not a single number — it is a curve you draw over training. You start small (warmup), climb to a peak, then ease down (decay). This topic explains why that shape became the default, with the formulas, the schedules, and the real 2026 settings.

12 min read•3 Quiz Questions

Batch Normalization standardizes each channel using the current batch's mean and variance while training, then switches to a memorized running average at inference. This is the layer that let CNNs go deeper and learning rates go 10x higher — and where a huge share of training bugs live.

12 min read•3 Quiz Questions

LayerNorm answers BatchNorm's batch-coupling problem: standardize each sample's own features, so no batch statistics and no train/eval switch are needed. This topic shows the math, why LayerNorm became the normalization of transformers, and what pre-LN vs post-LN costs.

12 min read•3 Quiz Questions

Dropout randomly switches off a fraction of neurons at every training step, so no neuron can lean on any single partner. It works like training a huge ensemble that shares weights. This topic covers the mechanism, inverted scaling, where to place it, and its modern decline in LLM pretraining.

12 min read•3 Quiz Questions

Weight decay gently pulls every weight toward zero at each step, favoring small, smooth solutions. For SGD it is exactly L2 regularization; for Adam it is not — AdamW fixes that by decaying the weights directly. This topic covers the math, the AdamW correction, and the no-decay-on-norms convention.

12 min read•3 Quiz Questions