TOPIC #10Intermediate 10 min read

Partial Derivatives: Sensitivity in Multivariable Worlds

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A model has millions of knobs and one score. A partial derivative answers the fundamental question of training — "if I nudge just THIS weight and freeze all the others, what happens to the loss?" Do it for every knob and you have the gradient.

Partial Derivatives as Slices

Hold every other variable constant; the function becomes 1-D along that axis; take an ordinary derivative. Do it once per parameter and you have the gradient.

Partial Derivatives as Slices
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: One Score, Millions of Knobs

Topic 9 gave you the derivative: nudge the one input of a function, watch the output move.

But a trained model isn't like that.

A model has millions of inputs — its weights — and exactly one output you care about during training: the loss.

So when the loss reads 2.47, one weight is to blame... partly. Which ones pushed the score up? Which pulled it down? By how much?

Insight

"If I nudge only weight #4,031 and freeze every other weight, what happens to the loss?"

That question, asked once per weight, is the question every training step answers. Its name is the partial derivative.

Why "partial"? Because you get a partial picture of the function: the honest slope along one axis, with everything else held still. A million partials later, you've rebuilt the full picture (that stack is the gradient — Topic 11, right next door).

02.The Idea in Plain Words: Freeze Everything But One

A partial derivative is simply

Insight

An ordinary derivative where every other variable is pretending to be a constant.

Formally: for f(x₁, x₂, …, xₙ), the partial derivative with respect to xᵢ, written ∂f/∂xᵢ (or f_xᵢ), is the derivative of the slice of f where all other variables are frozen.

The computation recipe is suspiciously boring:

  1. Pick one variable.
  2. Look at every other symbol through squinted eyes — treat it as a plain number.
  3. Apply the Topic 9 rules (power, product, chain, exp/log) and differentiate.

That's it. No new rules. The partial derivative machine reuses everything you already know about single-variable slopes.

Two housekeeping facts worth knowing:

  • The funny symbol ∂ is just a decorated d, shouting "other variables exist, but I'm ignoring them for now."
  • For smooth functions the order of operations doesn't matter: differentiate by x then y, or y then x — same answer (Clairaut's theorem: mixed partials commute). Autograd's double-backward passes silently rely on this.
python— Partial derivatives by hand vs PyTorch autograd
import torch

x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)
f = x**2 * y + torch.sin(y) + torch.exp(x)

f.backward()                    # computes ALL partials in one sweep
print(x.grad)   # 2xy + e^x = 12 + 7.389 = 19.389
print(y.grad)   # x^2 + cos(y) = 4 - 0.990 = 3.010

# Numerical check: perturb only x, keep y fixed
def fn(a, b):
    return a**2 * b + torch.sin(torch.tensor(b)) + torch.exp(torch.tensor(a))
h = 1e-5
print((fn(2.0 + h, 3.0) - fn(2.0 - h, 3.0)) / (2 * h))  # ~19.389

03.A Simple Worked Example: Three Terms, Two Squints

Take f(x, y) = x²y + sin(y) + eˣ and stand at x = 2, y = 3.

Partial in x — squint; y and sin(y) are now furniture:

  • x²y: y is a constant multiplier → derivative 2x·y
  • sin(y): pure constant → derivative 0 (it vanishes)
  • eˣ: → eˣ

∂f/∂x = 2xy + eˣ

At (2, 3): 2·2·3 + e² = 12 + 7.389 = 19.389.

Partial in y — squint differently; now x is furniture:

  • x²y: x² is the constant multiplier → derivative x²·1 = x²
  • sin(y): → cos(y)
  • eˣ: constant → 0 (vanishes)

∂f/∂y = x² + cos(y)

At (2, 3): 4 + cos(3) = 4 − 0.990 = 3.010.

Same function, two different squints, two different answers. And both are honest:

  • 19.389 says: at this setting, nudging x by +0.01 (leaving y frozen) lifts f by about 0.194.
  • 3.010 says: the same nudge on y lifts f by only about 0.030.

x is roughly 6× more sensitive than y here. If f were your loss, x is the knob that demands attention first. The PyTorch block above reproduces both numbers to the decimal — f.backward() filled x.grad and y.grad automatically.

04.Visual Intuition: Cutting Slices Out of a Landscape

Picture f(x, y) as a hilly landscape: two horizontal axes, height = output.

The partial derivative is what you get when you slice the landscape with a vertical plane:

code
   full surface f(x,y)          slice at y = 3 (freeze y!)
        ▄▄▄╱╱╱╱                      ▄▄▄▄
      ▄▄╱╱╱╱   ╱╱╱                  ╱╱    ╲        ← a 1-D curve!
    ╱╱╱╱    ╱╱╱        ╲____╱╱╲____╱   ●  ← slope of
    ╱   ╲╱╱╱                            the slice at x=2  = ∂f/∂x
                                    = 19.389 here
  • Freeze y at 3 → you cut a curve out of the surface → take its ordinary slope → that's ∂f/∂x.
  • Freeze x at 2 → cut a curve the other way → its slope is ∂f/∂y.
  • A partial derivative is never "the slope of the landscape" — it's the slope of one cut through it. Walk east and you measure ∂f/∂x; walk north and you measure ∂f/∂y; the hill between the two has no single slope.

And the Hessian, later? Slice twice: how does the east-slope change as you walk north. That's ∂²f/∂x∂y — a curvature of slopes.

Do all the slices — every axis, one frozen world at a time — and stack the results into a vector. Congratulations, you've built the gradient (Topic 11 does exactly this).

05.The Analogy: The Sound Engineer and the Million-Slider Console

Carry this analogy through the rest of the topic: a mixing console with a million sliders, feeding one speaker.

  • Each slider is a weight. The speaker plays one number: the loss. A high volume = bad prediction.
  • The engineer needs the mix quieter. She can't guess. So she runs experiments:
Insight

First partial derivative: "Solo one slider — freeze every other slider — and note how much louder it gets when I slide it up a bit."

  • ∂L/∂θᵢ = +19.4 → this slider is a shouty one: moving it down quiets the mix a LOT.
  • ∂L/∂θⱼ = +0.003 → this slider barely matters right now. Leave it alone.
  • A negative value → pulling this slider up actually quiets things. Counterintuitive, but the meter doesn't lie.

Stack a million solo readings into one list and the producer has the full sensitivity report — the gradient.

The mixed part — literally mixing — gives you the second-order story:

Insight

Second partial derivative: "When I raise slider A, does slider B's effect get stronger or weaker?"

Because on a real console, boosting the vocal fader changes how much the reverb fader matters. Sliders interact. The Hessian (Section 7) is the giant table of all pairwise interactions, and the Jacobian (Section 8) is what you build when the console has a million speakers instead of one.

One warning the engineer knows well: each solo reading is local. At slider position (2, 3), ∂f/∂x = 19.389 — move far and the reading changes. Partials are point statements, not laws.

06.Why AI Cares: Two Partials That Became the Hello-World of Training

The smallest real ML example: fitting a line ŷ = wx + b with mean squared error over n samples:

L(w, b) = (1/n) Σᵢ (w·xᵢ + b − yᵢ)²

Two knobs (w and b), so exactly two partial derivatives — and both come straight from the Topic 9 rules (the chain rule: square → 2·(residual); the inner residual → xᵢ or 1):

  • ∂L/∂w = (2/n) Σᵢ (wxᵢ + b − yᵢ)·xᵢ
  • ∂L/∂b = (2/n) Σᵢ (wxᵢ + b − yᵢ)·1

Read the first one in plain words: "average, over all samples, (how wrong you were) × (how much x contributed)", times 2. If you overshot on samples where x is large, ∂L/∂w > 0 → shrink w. Every optimizer in every framework is just accumulating these slices for every scalar in every tensor: w ← w − lr·∂L/∂w, b ← b − lr·∂L/∂b. Textbook SGD, born from two partials.

Scale check: a 70B-parameter model has 70 billion partial derivatives — one per weight — computed together in roughly two forward passes' worth of work, thanks to the machinery of Topics 11–12.

So the interpretation to memorize: each ∂L/∂θᵢ answers "by how much does L change per unit increase of θᵢ, right now, others fixed?" — the per-weight sensitivity score that decides whether that weight should grow or shrink, and by roughly how much.

07.Going Deeper: The Hessian — a Table of Slider Interactions

Differentiate the partials again and you get second-order partial derivatives ∂²L/∂θᵢ∂θⱼ: for each pair of knobs, "how does knob i's sensitivity change as knob j moves?"

With p parameters, these assemble into the Hessian matrix H ∈ R^{p×p}:

  • Diagonal ∂²L/∂θᵢ²: curvature along each parameter axis — bounds on the usable learning rate (a stiff knob needs a gentle hand).
  • Off-diagonals: interaction curvature — whether raising one weight makes another want to rise or fall (feature/parameter correlation; the "vocal fader changes the reverb's effect" reading).
  • H is symmetric for smooth f (Clairaut again), so its eigendecomposition is well-behaved (Topic 7): the eigenvalues are the principal curvatures of the loss surface, and the eigenvectors its stiff/soft directions.

Second-order methods use this table to take better steps than "follow the local slope" — Newton, K-FAC, and 2024–2026-era optimizers like Sophia and Shampoo. But storing a full p×p Hessian for p = 70×10⁹ needs ~10¹⁹ numbers — impossible. Hence the industry of approximations: low-rank, block-diagonal, and Hutchinson-probe estimates (Topic 13).

08.When the Output Is Not One Number: the Jacobian

Everything above assumed one output (a scalar loss). But a layer usually maps a vector in, vector out: F: Rⁿ → Rᵐ.

Then every input affects every output, so there are m·n sensitivities ∂Fᵢ/∂xⱼ. They organize into the Jacobian J ∈ R^{m×n} with

Jᵢⱼ = ∂Fᵢ/∂xⱼ

  • Rows of J = the gradient of each output component (m gradients, one per "speaker").
  • Columns = how one input perturbs all outputs.

Two practical truths about Jacobians in deep learning:

  1. Backprop never forms the full Jacobian for giant layers. It computes vector–Jacobian products vᵀJ — contracting the Jacobian with the incoming gradient stream — one row of Jacobian information per pass. That's why autograd costs O(#ops), not O(#parameters²).
  2. "Forward mode autodiff" computes Jacobian-vector products (Jv) instead, winning when outputs ≪ inputs — niche uses in sensitivity analysis and physics-informed ML. Reverse mode (VJPs, i.e. backprop) wins the scalar-loss regime of deep learning (the mechanism is Topic 12).

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Reduces a mountain-dimensional problem to independent 1-D slopes — exactly what first-order optimizers consume.
  • One reverse sweep of autodiff yields all partials simultaneously at ~2× forward cost.
  • Second-order partials (H) unlock curvature-aware methods and stability analysis when affordable.

Trade-offs & Constraints

  • Each partial ignores interactions: "others constant" is only valid infinitesimally; far-apart coordinate moves couple.
  • Full Jacobians/Hessians scale as O(mn) / O(p²) — prohibitive at model scale without approximation.
  • Non-smooth models (ReLU kinks) make some partials undefined; conventions silently choose a value.
Production Implementation in Big Tech
PyTorch autograd engine• Computing ∂loss/∂θ for every parameter, every step

torch.autograd records every op producing the loss, then one backward() call executes hand-written per-op partial-derivative rules composed by the chain rule. model.parameters() gradients are literally dictionaries of partial derivatives ∂L/∂θᵢ — enabling Adam, clipping, and distributed all-reduce of millions of sensitivities per millisecond.

Staff+ Engineering Takeaways

  • A partial derivative is the slope of f along one coordinate axis with all other variables frozen.
  • Compute ML partials with ordinary differentiation rules; ∂L/∂θᵢ is the per-weight sensitivity that steers updates.
  • The gradient vector stacks first partials; the Jacobian stacks first partials of vector-valued maps; the Hessian stacks second partials.
  • Backprop computes all partials in one reverse sweep (VJP products) instead of n separate finite-difference probes.
  • Linear-regression partials (∂L/∂w, ∂L/∂b) are the smallest complete example of what every optimizer consumes.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

For f(x, y) = x²y³, the partial derivative ∂f/∂x equals:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?