TOPIC #13Intermediate 12 min read

Jacobian & Hessian

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

The gradient answers "which way is downhill?" — but not "what shape is the ground?" The Jacobian is the complete sensitivity table for vector-valued maps, and the Hessian is the curvature table that classifies flat spots as minima, maxima, or saddles and powers second-order optimization.

Classifying Critical Points with the Hessian

At any point where the gradient vanishes, the eigenvalues of the Hessian reveal the local geometry: all positive means a local minimum, all negative a local maximum, mixed signs a saddle point.

Classifying Critical Points with the Hessian
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: The Gradient Tells You Where to Step, Not What the Ground Looks Like

Two topics ago (Topic 11): the gradient ∇f collects every partial derivative (Topic 10) into one vector. It says "here's the steepest uphill direction," so gradient descent walks the other way, −∇f.

Almost every optimizer in machine learning — SGD, Adam, AdaGrad — consumes nothing but gradients. First-derivative numbers. Full stop.

But gradients have two blind spots:

Insight

"The slope is zero here. Great... is that a valley floor, a mountaintop, or a mountain pass?"

The gradient shrugs. Zero slope looks identical at all three.

Insight

"The slope is 0.1 to the left and 0.1 to the right — is the ground gently curving down, or is it a razor-tight canyon?"

The gradient shrugs again. Two places can have the same slope and wildly different shapes — and shape decides how big a step is safe.

To fix the blind spots you need to finish two sentences:

  • "How does each of my m outputs change with each of my n inputs?" → a table of first derivatives → the Jacobian.
  • "How does each slope itself change as I move?" → a table of second derivatives → the Hessian.

That's this whole topic: two matrices that turn "which way?" into "what shape?"

02.The Idea in Plain Words: Two Completed Tables

The Jacobian is simply

Insight

The full grid of first derivatives, when a vector in becomes a vector out.

For f: Rⁿ → Rᵐ (n inputs, m outputs), build an m × n matrix J:

Jᵢⱼ = ∂fᵢ / ∂xⱼ

One row per output, one column per input. Entry (i, j): "if I poke input j, how much does output i wiggle?" With m·n outputs×inputs pairs, that's m·n sensitivities — all of them, in one table.

Sanity anchors:

  • A scalar function (m = 1) has a Jacobian that's just the gradient as a row vector. Gradients are the special case.
  • A linear layer y = Wx has Jacobian W: the map is its own derivative.
  • The softmax of a vector has a dense n × n Jacobian, because nudging any input logit shifts every output probability. (Section 6 computes it with numbers.)

The Hessian is

Insight

The full grid of second derivatives, when you have one scalar output.

For f: Rⁿ → R, the n × n matrix H with

Hᵢⱼ = ∂²f / ∂xᵢ ∂xⱼ

"Take the slope along axis j. Now measure how that slope changes as you move along axis i." When f is smooth, mixed partials commute (Schwarz's theorem — the same fact Topic 10 leaned on), so H is symmetric. Symmetric means real eigenvalues and orthogonal eigenvectors (Topic 7) — and that is what makes Hessian eigenvalues so useful: they are the curvatures along the surface's principal bending directions, one number per direction.

Dimension contract worth memorizing: Jacobian of Rⁿ → Rᵐ is m × n; Hessian of a scalar function on Rⁿ is n × n.

03.A Simple Worked Example: A Bowl, a Dome, and a Pass

Three tiny two-variable functions. In each, the gradient is zero at the origin — the first-derivative test gives identical verdicts. Watch the Hessian tell them apart.

Bowl: f(x, y) = x² + 3y²

  • Gradient: (2x, 6y) → (0, 0) at the origin. Slope: dead flat. Hmm.
  • Hessian: [[2, 0], [0, 6]] — already diagonal, so eigenvalues are λ = 2 and λ = 6, both positive.
  • Verdict: curves up in every direction → local minimum. And the eigenvalue sizes say the y-direction is 3× stiffer: along x you can use lr up to ≈ 2/2 = 1, along y only ≈ 2/6 ≈ 0.33 before you'd bounce out of the canyon (Topic 7's rule).

Dome: g(x, y) = −x² − 3y²

  • Same flat origin. Hessian [[-2, 0], [0, -6]]: eigenvalues both negative → local maximum.

Mountain pass: s(x, y) = x² − y²

  • Same flat origin. Hessian [[2, 0], [0, -2]]: eigenvalues +2 and −2 — mixed signs.
  • Verdict: saddle point. A bowl if you walk east-west, an upside-down bowl north-south.

That's the classification rule in one line: at a zero-gradient point, all Hessian eigenvalues > 0 → minimum, all < 0 → maximum, mixed → saddle, any near 0 → flat direction.

High-dimensional fine print: real loss surfaces have millions of eigenvalues per point. For all of them to be positive by luck is astronomically unlikely — so in practice most "flat spots" you get stuck near are saddles or flat valleys, not bad minima. That observation (Dauphin et al., 2014) reshaped how researchers think about training.

04.Visual Intuition: The Shape Under Your Feet

Draw the three examples as ground surfaces, all perfectly flat at the marked center point:

code
      BOWL (min)             DOME (max)            PASS (saddle)
        ╲_╱╲_╱                 ╱‾╲╱‾╲              ╲_╱ │ ╲_╱
       ╱ ● ╲                 ╲ ● ╱               up ╱  ●  ╲ down
      along x: valley        along x: peak        along x:  valley
      along y: valley        along y: peak        along y:  peak  ← disagrees!

And the Hessian eigenvalues, plotted as arrows at the flat point:

code
   min:     ●   λ₁=+2 ↗ curves up      both bend UP     → stay, you're in a pit
   max:     ●   λ₁=−2 ↘ curves down    both bend DOWN   → you're on a summit
   saddle:  ●   +2 one way, −2 the other → a pass: two valleys, two ridges
   flat:    ●   λ ≈ 0 → a long shallow trench: no information, slow crawling

Read the pass (saddle) picture like terrain: from a mountain pass you can walk down into two valleys and up onto two ridges — the ground is flat exactly at the pass, yet it is neither pit nor peak. Mixed-sign eigenvalues are the algebra of that scene.

The near-zero case matters most in training: flat directions (λ ≈ 0) are long, shallow valleys where the gradient barely moves — gradient descent crawls along them, and this is one core reason learning rates are hard to tune.

05.The Analogy: A Surveyor Mapping Flat Ground

Carry this through the rest of the topic: you're a land surveyor, paid to find where the lowest ground is, on a fog-bound estate with one instrument: a tiny spirit level.

  • The gradient is a big, accurate ruler tilted one way: it tells you which direction descends right now. Roll your ball that way.
  • But at a flat spot the ruler reads zero everywhere — and zero is ambiguous. Before buying land at a flat spot, you want the shape: is this a basin, a hillock, or a natural pass between hills?
  • So you do the second-derivative survey: tilt your little board along every independent direction and record how fast the slope itself changes. A basin: tilting any way sends the edge up. Hillock: any way sends it down. Pass: up along the trail, down along the ridgeline. Mixed answers = a pass.
  • Your survey sheet — every direction × every direction, how the tilt changes — is the Hessian. The eigenvalues are the two "principal directions" of the terrain: the directions along which the ground bends purely up or purely down, with nothing mixed.
  • A shrewd developer's rule: the steeper the bend (bigger λ), the shorter your step from that spot must be. λ = 6 is a tighter canyon than λ = 2 — hence learning rate ≈ 2/λ ceilings, and hence the condition number κ = λmax/λmin: if the estate's tightest canyon is 1000× tighter than its widest valley (κ = 1000), one global step size can't serve both. You zig-zag across the ravine instead of marching down it.
  • Rescuing a fog-bound gradient-descent walker: the pass has downhill escape routes the ball can roll out of (that's SGD noise escaping saddles); a basin has none — which is exactly why you'd rather converge to a flat basin than keep falling into passes you stall at.

Newton's method (Section 7) is the surveyor becoming the builder: "I know the exact shape now, so I'll place my next step at the bottom of the fitted bowl, in one stride."

06.The Jacobian in Practice: Backprop Never Builds One

The Jacobian is conceptually central and operationally avoided — a tension every ML engineer should be able to state.

Conceptually: every layer is a Rⁿ → Rᵐ map, so every layer has an m × n Jacobian. For the softmax with 100K vocabulary logits, that's a 10⁵ × 10⁵ dense table — 10 billion sensitivities for one layer.

Operationally: backpropagation does not literally build giant Jacobians. Autograd computes vector-Jacobian products (VJPs; the forward-mode twin is the Jacobian-vector product): it contracts each layer's Jacobian with the incoming gradient stream in one pass — vᵀJ instead of J. Since the loss is a scalar, there is exactly one such stream, and each backward step harvests one row of Jacobian information per op. That is why the chain rule (Topic 12) scales to networks with millions of layers while remaining cheap per layer: cost is O(#ops), not O(m·n).

The code below brute-forces the softmax Jacobian with finite differences so you can see the pattern the VJP hides — and spot the signature property: every row sums to ~0. That's the probabilities-still-add-to-1 constraint in derivative form: poke any logit and the probability mass just moves between classes, never grows.

python— Finite-difference Jacobian of a vector-valued map (softmax)
import numpy as np

def softmax(z):
    e = np.exp(z - z.max())      # shifted for numerical stability
    return e / e.sum()

def jacobian(f, x, eps=1e-6):
    y0 = f(x); m, n = y0.size, x.size
    J = np.zeros((m, n))
    for j in range(n):
        xp = x.copy(); xp[j] += eps
        J[:, j] = (f(xp) - y0) / (2 * eps)
    return J

z = np.array([2.0, 0.0, -1.0])
J = jacobian(softmax, z)
print(J.round(3))
# rows: how each output probability reacts to each input logit
print("row sums ~ 0:", J.sum(axis=1).round(6))

07.Why AI Cares: Newton's Dream and the O(n²) Nightmare

If the Hessian tells you the shape, why not use it for the step itself?

Newton's method does exactly that. It solves ∇f(x + δ) = 0 under the quadratic model, giving the update:

x ← x − H⁻¹∇f

Translation in surveyor language: fit the local bowl, then teleport to its bottom — instead of guessing a step size off a slope, take the geometrically correct step. Two superpowers follow:

  • Near a well-behaved minimum it converges in a few steps (quadratic convergence).
  • It is scale-invariant: reparameterize the loss — stretch or squash any axis — and the trajectory is unchanged. The H⁻¹ cancels the distortion. Vanilla gradient descent has no such immunity; that's ravine zig-zag in the fog.

The nightmare is arithmetic. For n = 10⁷ parameters:

  • storing H needs O(n²) memory — 10¹⁴ entries;
  • inverting it costs O(n³).

Completely infeasible. Practical deep learning therefore buys partial curvature with approximations, each a different slice of the survey sheet:

  1. Diagonal Hessian (AdaHessian): keep only H_ii, preconditioning each parameter individually — the curvature of each axis, ignoring cross-terms.
  2. Hutchinson estimation: estimate curvature with random probe vectors and Hessian-vector products, never materializing H.
  3. Quasi-Newton (L-BFGS): build a cheap curvature model from gradient differences over recent steps — common in classical ML and differentiable rendering.
  4. Fisher information / natural gradient / K-FAC: use the expected Hessian of the log-likelihood (the Fisher matrix); block-factorized variants (K-FAC) scale to large networks.
  5. Adam and RMSProp: implicit, diagonal curvature estimates from running gradient statistics — a pragmatic surrogate for second-order scaling, which is why they behave better than SGD on badly conditioned problems.

08.In Practice: Reading the Landscape Like an Operator

In day-to-day training, nobody diagonalizes anything — but the Hessian's ideas show up as operations:

  • Saddle-point thinking: "SGD noise escapes saddles but not minima" (Topic 7) is pure Section 3 geometry — a pass always has downhill directions for the noise to find; a basin has none. So getting stuck at a bad local minimum is rare in high dimensions; getting stuck near flat directions and passes is the everyday slowdown.
  • Condition-number diagnostics: the ratio λmax/λmin of the local Hessian predicts whether one global learning rate can work at all — and motivates per-layer normalization, per-coordinate scaling (Adam), and preconditioners (K-FAC, Shampoo, Sophia).
  • Sharpness-aware training: flat basins (small ‖H‖) generalize better than sharp pits; SAM-style optimizers deliberately hunt for them. The mermaid diagram above is the classification flow they all share: gradient zero? → read H's eigenvalues → min / max / saddle / flat.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Newton-type updates converge quadratically near the minimum and need no hand-tuned learning rate.
  • Curvature eigenvalues diagnose the loss landscape: minima, saddles, and flat directions.
  • Second-order preconditioning makes optimization invariant to parameter scaling.

Trade-offs & Constraints

  • Full Hessian storage is O(n²) and inversion is O(n³): impossible for deep networks.
  • Mini-batch Hessian estimates are extremely noisy compared to batch gradients.
  • Indefinite curvature (saddles) makes naive Newton steps unstable far from minima.
Production Implementation in Big Tech
Google DeepMind• K-FAC: scaling second-order optimization

Kronecker-Factored Approximate Curvature (K-FAC) approximates the Fisher information matrix (the expected Hessian of the log-likelihood) for each layer as a Kronecker product of two much smaller matrices, enabling Newton-like updates on networks with tens of millions of parameters and measurably faster training on convolutional models.

Staff+ Engineering Takeaways

  • The Jacobian is the complete first-derivative map of f: R^n → R^m, with shape m × n; gradients are the special case m = 1.
  • The Hessian is the symmetric matrix of second partial derivatives; its eigenvalues classify critical points as minima, maxima, or saddle points.
  • Backpropagation is a chain of vector-Jacobian products, which is why autograd scales without ever building full Jacobians.
  • The condition number (ratio of extreme Hessian eigenvalues) explains ravine-like losses and motivates adaptive optimizers.
  • Newton's method is fast but O(n²)-O(n³) in cost, so deep learning uses diagonal, low-rank, quasi-Newton, or Fisher-based curvature approximations.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

A function f maps R^5 → R^3. What are the shapes of its Jacobian and, for a scalar loss g: R^5 → R, its Hessian?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?