Logistic Regression
A linear model for probabilities. It fits the log-odds as a straight line, squashes that line into a probability with the sigmoid, and trains it on cross-entropy. There is no closed-form solve, but the objective is convex and the decision boundary stays auditable.
01.The Problem: A Straight Line Gives Silly Probabilities
Picture a loan officer deciding whether to approve you.
The inputs are your income, your debt, and your credit history.
What she wants back is one number: the probability you will repay.
Can't we just run linear regression and read off that number?
No. Try it and three things break:
- Predictions leave [0,1] — a "probability" of 1.3 is nonsense.
- The errors are unequal across x by construction (for a 0/1 target, Var(y|x) = p(1−p)), so squared-error assumptions collapse.
- One extreme row tilts the whole line, because squared loss punishes big errors quadratically.
Worse, a straight line treats the jump from p=0.01 to p=0.05 as identical to the jump from p=0.51 to p=0.55 — even though the real-world odds change enormously there (about 4x versus barely at all).
So what quantity is plausibly a straight line?
That single question is the whole insight behind logistic regression.
From Linear Score to Decision ⚖️
From Linear Score to Decision ⚖️
Logistic regression fits a linear function on the log-odds scale, transforms it with the sigmoid into a probability, and optimizes it with cross-entropy — the threshold is applied only at the end.
Unlock Topic #46: Logistic Regression
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?