Derivatives: Instantaneous Rates and Local Linear Maps
The derivative answers one question: "if I nudge this input a tiny bit right here, what happens to the output?" It is the slope of the tangent line, the best local straight-line copy of a curve — and every training step in AI is built from it.
Training as Repeated Tangent-Following
Every gradient-descent iteration: approximate the loss by its tangent line, walk a small step downhill along that line, then re-measure. Derivatives make the whole loop definable.
01.The Problem: Averages Hide What Happens *Right Now*
You drive 100 km in 2 hours.
Average speed: 50 km/h.
But were you ever going faster than that? Slower? The average can't say. Maybe you crawled through a city and flew down a highway.
A function is the same. Say the loss is L(w). You can ask:
"If I move the weight from 2 to 3, how much did the loss change on average?"
That's just a slope between two points — the average rate: [L(3) − L(2)] / 1.
But training doesn't move weights from 2 to 3. It nudges them a tiny amount, from here, right now. And it needs to know:
"At exactly w = 2, which way is downhill, and how steep is it under my feet?"
The average-over-a-stretch can't answer that. A city crawl + highway fly can average to "flat" while every single moment was anything but.
You need a slope-measuring device that works at a single point. That device is the derivative.
02.The Idea in Plain Words: The Slope Right Here
The derivative of f at x is what's left when you shrink the "stretch" to nothing:
f'(x) = lim_(h→0) [f(x + h) − f(x)] / h
Read it slowly: take the average rate over a tiny step h... then let h shrink toward zero. Whatever number the ratio settles on is the derivative.
You should hold three pictures of f'(x) in your head at once:
- Instantaneous rate of change.
f'(x) = 2means "output gains ≈ 2 units per unit of x, right here" — your speedometer reading, not your trip average. - Slope of the tangent line. The best straight-line copy of
fnear x:
f(x+h) ≈ f(x) + f'(x)·h
This is called local linearization, and it is the deep reason first-order optimizers work: zoom in enough on any smooth function and it looks linear. Up close, every curve is a staircase of straight pieces.
- A flatness detector.
f'(x) = 0marks a flat point — a local minimum, local maximum, or (in higher dimensions) a saddle (Topic 7). If the slope is zero, a tiny step changes nothing to first order.
The third picture has fine print: flat ≠ best. A plateau is also flat. Telling minima from maxima from saddles needs the second derivative — kept for Section 7 and Topic 13.
def f(w):
return w**2 + 3*w
def f_prime(w): # analytic: 2w + 3
return 2*w + 3
def finite_diff(f, w, h=1e-5): # numerical estimate
return (f(w + h) - f(w - h)) / (2*h)
print(f_prime(2.0)) # 7.0
print(finite_diff(f, 2.0)) # 7.0000000...
# h too small -> catastrophic cancellation; too big -> truncation bias.
# Autodiff (Topic 12) gives exact derivatives at scale instead.03.A Simple Worked Example: Shrink h Until You See 7
Use f(w) = w² + 3w at w = 2. First, f(2) = 4 + 6 = 10.
Now compute the average rate over a step h, and keep shrinking h:
codeh = 1.00 : [f(3) − f(2)] / 1 = (18 − 10) / 1 = 8.00 h = 0.10 : [f(2.1) − f(2)] / 0.1 = (11.51 − 10) / .1 = 7.10 h = 0.01 : [f(2.01) − f(2)] / .01 = 7.01 h → 0 : settles on 7
The numbers march toward 7. So f'(2) = 7: at exactly w = 2, nudging w upward by a tiny h raises f by about 7h.
You can also get it algebraically. For this f the secant slope works out to exactly 7 + h (try it: f(2+h) − f(2) = (4+4h+h² + 6+3h) − 10 = 7h + h², and dividing by h gives 7 + h). Let h → 0 and the +h dust disappears. Seven.
And the general formula, from the rules in Section 6: f'(w) = 2w + 3, so f'(2) = 4 + 3 = 7. ✓ Three roads — secants, algebra, rules — same number.
One more check of the local-linearization claim: predict f(2.1) with the tangent line: f(2) + 7·0.1 = 10.7. True value: f(2.1) = 4.41 + 6.3 = 10.71. Off by 0.01. For tiny nudges, the straight line is essentially the function.
Notice what went wrong at the start: h = 1 overestimated (8, not 7). Big steps lie. Remember that — it is why learning rates exist.
04.Visual Intuition: Zoom In and the Curve Goes Straight
Picture the curve, the secant line through two points, and the tangent that survives when the second point slides in:
codesecant (h = 1) tangent (h → 0) ╲ ╱ ╲ ╱╲ ╱╱╱╱ ← tangent: touches, ╲╱ ╲ ╲ ╱ slope 7 at w = 2 ╱ ●2 ╲ curve ╲●╱ ● ────┴──────┴────── ─────┴────── w=2 w=3 w=2
And the zoom view of local linearization:
codezoom 1× zoom 4× zoom 16× zoom 64× ╱╲ ╱─╲ ╱——╲ ▬▬▬▬▬ ╱ ╲ ╱ ╲ ╱ ╲ straight!
Smooth curves are liars at full zoom and honest up close. f'(2) = 7 is a promise that is true near 2 and gradually false as you walk away.
That single fact generates half of ML's hyperparameter obsession:
- the step must be small (else the straight-line model lies) → learning rates,
- or you re-measure the slope often → mini-batches, epochs,
- or you test how far the promise reaches → line search, trust regions (Section 8).
05.The Analogy: A Speedometer on a Winding Mountain Road
Carry one analogy through the rest of the topic: you're driving a mountain road at night, and the only instrument you trust is the speedometer.
- The odometer log (average speed over each 10 km stretch) is the secant slope. Useful for planning, useless for the next second.
- The speedometer needle is the derivative: the rate right now, at this exact meter of road.
- The needle tells you direction and urgency: reading +7 means "the altitude climbs 7 meters per kilometer of road right here." To lose altitude fastest in the immediate next moment, drive backwards at exactly that rate.
- If the needle reads 0, you're on flat ground — could be a valley floor, a summit, or a mountain pass (a saddle). The speedometer alone can't tell you which. You'd need to feel whether the road curves up or down around you — that's the second derivative (Section 7).
- If you floor it based on one needle reading, the road will bend and betray you. The reading is local. Drive at a speed the curve deserves.
Every optimizer you'll ever meet is a driver doing exactly this: read the needle, nudge, re-read, nudge — trusting the needle only for the next few meters. Neural network training is driving a mountain road using nothing but a speedometer, billions of times per second.
06.The Rules: Differentiation Is an Algorithm, Not a Miracle
Computing derivatives from the limit definition every time would be agony. Fortunately there are rules — and each one is a small algorithm template that software can execute:
- Linearity:
(αf + βg)' = αf' + βg'— sums and scalings pass through; element-wise sums scale gradients component-wise. - Power / polynomial:
d/dx xⁿ = n·xⁿ⁻¹— the workhorse behindw² → 2win our example. - Product rule:
(fg)' = f'g + fg'— each factor takes a turn differentiating while the other sits still. This is the derivation kernel of backprop through multiplicative interactions: attention's Q·K, gating in LSTMs/GLU, LoRA's B·A. - Quotient rule (product rule with a reciprocal) and exponential/log:
d/dx eˣ = eˣ,d/dx ln x = 1/x— logs matter because they convert likelihood products into sums, so gradients stay numerically sane. - Chain rule:
(f∘g)' = f'(g(x))·g'(x)— derivatives of nested functions multiply. It gets its own topic (12) because it is backpropagation.
Why does an ML person care about rules this much? Because each operation in a training graph (add, multiply, matmul, exp, divide) has a hand-written local derivative, and autodiff just composes those locals using these rules. That is the code-level meaning of "the layer knows how to backprop itself."
The derivative of f(w) = w² + 3w you computed in Section 3? Linearity + power rule: 2w + 3. Same arithmetic, one line.
07.Why AI Cares: The Derivative of Each Activation Function Is a Trainability Verdict
A neural network is a chain of linear layer → activation blocks. Backprop (Topic 12) will multiply the activation derivatives along that chain. So: an activation's value shapes what the network can express; its derivative shapes whether it can be trained at all.
The vanishing-gradient catalog, in one glance:
- Sigmoid σ(x): a smooth S from 0 to 1; peak derivative 0.25 at x = 0. Stacked through L layers, gradients multiply by ≤ 0.25 each step → vanish ~exponentially (
0.25ᴸ). Historically fatal for deep nets. - Tanh: derivative ≤ 1 (peak at 0) — better, but still vanishing in the saturation regions (large |x| ⇒ slope ≈ 0, the needle stuck at zero).
- ReLU max(0, x): derivative 1 for x > 0, 0 for x < 0. Identity slope preserves the signal backwards for active units — the key to training 2012+ deep nets. The "dying ReLU" failure is the flip side: units stuck negative pass zero gradient forever.
- GELU / SiLU/Swish (default in 2024–2026 LLMs): smooth versions of ReLU, derivative near 1 in the active region with small negative dips — designed for exactly the same slope-vs-stability tradeoff.
- Softmax + cross-entropy: the combined derivative simplifies to
predicted − target— the famously clean gradient behind classification training loops (you'll meet it again in Topic 11).
Speedometer translation: sigmoid in saturation = a needle reading 0 on a road that's actually steep — the layer claims flatness and starves its own upstream layers.
08.In Practice: From One Slope to Every Training Loop
Everything above compresses into one line. For a 1-D loss, gradient descent is literally
w ← w − lr·f'(w)
Take a step proportional to minus the slope of the tangent model. Downhill at the speedometer's advice. Everything later in this course generalizes exactly one idea — replace f' with the gradient vector (Topic 11) computed via the chain rule (Topic 12).
Practical corollaries interviewers probe:
- Derivatives are local: a good linear model of the loss at w0 says nothing far away — hence learning rates, line search, trust regions. Our h = 1 estimate of 7 gave 8; the error grows with distance.
f' = 0is necessary but not sufficient for a minimum — flat ground includes maxima and passes. Check the second derivative (or, in many dimensions, eigenvalues — Topic 7) to classify the stationary point.- Non-differentiable kinks (ReLU at 0, abs at 0) are handled by subgradients: frameworks pick a convention (often 0) and training tolerates the measure-zero ambiguity — one wrong meter reading at one exact point doesn't ruin the drive.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Tangent-line approximation makes almost any smooth optimization problem locally trivial to move on.
- Closed-form derivatives of elementary ops let autodiff compose exact gradients efficiently.
- Sign of f′ gives direction for free (increase/decrease), zero of f′ locates candidates for optima.
Trade-offs & Constraints
- First derivatives alone mislead far from the point (nonlinearity, saddle points).
- Finite-difference estimates are noisy, biased by h, and cost one full evaluation per input dimension.
- Saturating activation derivatives (sigmoid/tanh) cause vanishing gradients in deep stacks.
GPT-class models replaced ReLU with GELU/SiLU because the smooth derivative profile avoids hard zero-gating of gradients while keeping slope ≈ 1 in the active region. The entire "why this activation" design discussion is applied single-variable derivative analysis at billion-parameter scale.
Staff+ Engineering Takeaways
- f′(x) is simultaneously the instantaneous rate, the tangent slope, and the coefficient of the best local linear approximation.
- Gradient descent for 1-D losses is literally stepping opposite the tangent slope; everything else generalizes that.
- Differentiation rules (product, chain, log) are the op-level templates autodiff systems stitch together.
- Activation trainability is derivative design: sigmoid ≤ 0.25 vanishes, ReLU = 1 when active, GELU smooths the kink.
- Stationary points need second-derivative/curvature checks (Topic 7) to classify as minima, maxima, or saddles.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
The best local straight-line approximation of a smooth f near x0 is:
How clear and actionable was this distributed systems breakdown?