Linear Regression
The oldest useful model in machine learning: you predict a number as base + weights x features, and fit those weights by least squares. It has an exact closed-form solve, clean diagnostics (R2, RMSE), and strict assumptions — and it fails silently on collinear features and on extrapolation.
Fitting a Line, Then Auditing It 📈
Least squares has an exact solution, so the modeling work is not the fit — it is the residual diagnostics and the collinearity audit that decide whether coefficients are trustworthy.
01.The Problem: Predicting One Number From Others
Imagine you run a small taxi company.
For every past ride you know two things: how far it went (km) and how long it took (min).
You also know the final fare.
Now a new ride comes in. You want to guess its fare from its km and minutes.
How do you turn some inputs into one predicted number?
The oldest, most honest answer is a straight-line rule:
- a fixed base you pay just for getting in,
- plus a rate per km times the distance,
- plus a rate per minute times the time.
That rule is linear regression.
In one plain sentence:
Linear regression predicts a number by multiplying each input by a learned weight and adding the results.
The weights (base, rate-per-km, rate-per-min) are learned from your old rides.
Learning means: find the weights whose guesses land closest to the real fares you collected.
Sounds obvious. But three things are actually tricky:
- How do we measure "closest"?
- How do we find those weights, fast?
- When does a straight line quietly lie to us?
Everything below answers those three.
02.The Model in Plain Words: Base + Weights × Features
The model says the target is a weighted sum of features plus some noise:
y = β₀ + β₁x₁ + … + βₚxₚ + ε
Read it slowly:
yis the number you predict (the fare).x₁ … xₚare the inputs (km, minutes, and any others).β₀is the base — the intercept.β₁ … βₚare the rates — the coefficients (or weights).εis everything your features don't explain — the noise.
In matrix form the same line is written y = Xβ + ε, where X stacks all rows.
One word matters: linear.
It means linear in the parameters (the β's), not in the raw input x.
So you can bend the line by transforming features first:
y = β₀ + β₁x + β₂x²
This is still a linear regression — it is linear in β₁ and β₂.
Polynomial and Fourier features let one straight-line machine fit curved data.
Non-linearity is handled by feature engineering, not by changing the algorithm.
03.How Do We Measure "Close"? Least Squares
We need one number for "how wrong" the current weights are.
OLS (ordinary least squares) uses the residual sum of squares:
RSS(β) = ‖y − Xβ‖²₂ = Σᵢ (yᵢ − ŷᵢ)²
A residual is just real − predicted for one row. We square each residual, then add them up.
Why square? Two reasons:
- Big misses get punished extra hard, so a wild error stands out.
- Squared error is the maximum-likelihood choice exactly when the noise is symmetric and bell-shaped: ε ~ N(0, σ²).
That second point is also a warning: squared error is not universal. If the errors are lopsided or have extreme tails, it is the wrong ruler (see the Loss Functions topic).
A tiny ride dataset makes this concrete. Suppose one input, km, and three rides:
coderide km real fare guess with β0=10, β1=6 1 1 18 16 -> residual 2 -> squared 4 2 2 30 22 -> residual 8 -> squared 64 3 3 44 28 -> residual 16 -> squared 256
RSS = 4 + 64 + 256 = 324 for these weights.
OLS asks: which β makes RSS the smallest?
Because RSS is a convex quadratic in β (one smooth upward bowl), there is exactly one best answer — no lucky restarts, no local traps.
04.Finding the Weights: The Normal Equations
Because RSS is a convex bowl, calculus finds its bottom by setting the slope to zero. That gives the normal equations:
XᵀXβ = Xᵀy ⟹ β̂ = (XᵀX)⁻¹Xᵀy
Three consequences follow immediately — they are why OLS is both beloved and dangerous.
1. Exact, non-iterative.
No learning rate, no epochs, no local minima — one matrix solve.
With a Cholesky factorization it costs O(np² + p³), which is why linear regression trains in milliseconds on millions of rows.
2. Geometric meaning.
Xβ̂ is the orthogonal projection of y onto the flat space spanned by X's columns. The residual sticks straight out, perpendicular to every feature column:
Σᵢ eᵢxᵢⱼ = 0 for all j
That is also why, with an intercept, the residuals sum to zero. A picture of the projection:
codey (real fares) | |\ residual e (perpendicular) | \ | Xβ̂ <- closest point on the feature "flat" | / +---------------- column space of X
3. Maximum likelihood under Gaussian noise.
Minimizing squared error equals maximizing likelihood when ε ~ N(0, σ²) — the same ruler point from §3.
05.Is the Line Any Good? R², Adjusted R², and RMSE
Once fitted, how do you report quality?
R² = 1 − RSS/TSS is the fraction of the target's variance the model explains, where TSS = Σ(yᵢ − ȳ)² (ȳ is the mean fare).
It is scale-free, easy to say out loud — and easy to abuse:
- R² never decreases when you add a feature, so 200 random noise columns can post a proud in-sample R². Adjusted R² penalizes p relative to n; AIC and BIC do the same.
- Training R² is not a performance claim. Report out-of-sample R² from a held-out split or CV fold.
- Two baselines keep it honest: predicting the mean gives R² = 0 exactly, and a negative R² means your model is worse than the mean — a real and common outcome on new data.
- For lopsided or heavy-tailed targets, report MAE and a percentile of absolute error alongside RMSE.
Coefficients are not causal effects.
βⱼ reads: "the expected change in y per one-unit increase in xⱼ, holding the other included features fixed."
That holding-fixed clause is doing enormous work:
- It only means something if the other features genuinely capture the confounders.
- It changes entirely when you add or drop a feature.
- If one variable's effect depends on another, you must add an interaction term (a product like
x₁·x₂); once an interaction is present, plain main-effect reading stops.
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score, mean_squared_error
from sklearn.model_selection import cross_val_score
import numpy as np
model = LinearRegression().fit(X_train, y_train)
resid = y_val - model.predict(X_val)
print("coef:", np.round(model.coef_, 4))
print("intercept:", round(model.intercept_, 4))
print("in-sample R2:", round(model.score(X_train, y_train), 4))
print("out-of-sample R2:", round(r2_score(y_val, model.predict(X_val)), 4))
print("RMSE:", round(np.sqrt(mean_squared_error(y_val, model.predict(X_val))), 4))
# Honest generalization estimate, not the optimistic train score
print("cv R2:", cross_val_score(model, X, y, cv=5, scoring="r2").mean().round(4))
# For standard errors and p-values use statsmodels (adds constant for intercept)
# import statsmodels.api as sm
# sm.OLS(y_train, sm.add_constant(X_train)).fit().summary()06.The Assumptions and Their Failure Modes
OLS is unbiased, consistent, and (under the classical conditions) the best linear unbiased estimator. Each guarantee has a matching assumption, and each violation leaves a signature in the residuals.
- Linearity in parameters / correct features. Left-out curvature shows as a U-shape in residuals-versus-fitted plots. Cure: add polynomial, spline, binning, or interaction terms.
- Exogeneity: E[ε|x] = 0. If a predictor correlates with the noise — omitted variables, measurement error in
x, or choosingxafter observingy(reverse causation) — coefficients are biased, and no amount of data fixes it. Measurement error pulls the affected coefficient toward zero; omitted-variable bias hides the effect inside correlated features. - Homoscedasticity (constant error variance) and no autocorrelation. A funnel shape in residual plots signals heteroscedasticity: it does not bias β̂ but makes standard errors and confidence intervals wrong. Cure: weighted least squares, a log-transform of
y, or robust (sandwich/HC) standard errors. Serially correlated residuals in time-series data invalidate naive p-values entirely. - Normal errors — only required for exact confidence intervals and hypothesis tests with small n, not for prediction. With large samples, the central limit theorem largely rescues inference.
Multicollinearity is the fifth practical problem.
When features are near-linear combinations of each other, XᵀX becomes ill-conditioned and its inverse unstable. So β̂ swings wildly with tiny data changes, even though predictions stay fine.
- Detect it with VIF (VIF > 5-10 is concerning; > 30 is severe) or the condition number of
XᵀX. - Symptoms: coefficient flips between retrains, huge standard errors, nonsense signs.
- Remedies: drop or combine redundant features, use ridge (topic 55), or use PCA.
Also beware high-dimensional p > n, where XᵀX is singular and OLS exactly interpolates the training data with zero residual and no generalization. Regularized or reduced-rank methods become mandatory there.
07.Where Linear Regression Still Wins in Production
Back to the taxi: your straight-line fare rule is the first model you always fit, and often the final one. The reasons are economic as much as statistical:
- Interpretability and auditability. Coefficients are reportable to regulators, finance teams, and customers. Scorecards in credit and insurance still lean on linear/logistic forms for this reason.
- Training and serving cost. One matrix solve offline, a dot product online — sub-millisecond at any QPS, no GPU, no feature-store round trips beyond the feature vector.
- Sample efficiency with few signal-rich features. When n is small relative to p, a linear model often beats a flexible one simply because it cannot overfit elaborate structure.
- Composability. Residual boosting is initialized from a linear fit; generalized additive models sum learned linear pieces; embeddings plus a linear head is a standard production pattern.
When linear regression is the wrong tool: strongly non-monotonic or saturating relationships you cannot feature-engineer away, targets with hard bounds or count/ordinal structure, heavy-tailed noise where squared loss is hijacked by a handful of rows, and interactions across many variables. The classical escalations are polynomial/spline features (staying linear), trees and gradient boosting (topic 56), or a generalized linear model with a link function (topic 46).
The professional habit is always the same sequence:
fit the linear baseline → measure validation RMSE → plot residuals → justify each added unit of complexity by the error it removes.
Architectural Trade-offs & Production Realities
Architectural Advantages
- Closed-form exact fit in one solve — no hyperparameters, no convergence tuning, trains on billions of rows via streaming least squares.
- Coefficients are directly interpretable and reportable, with standard errors and confidence intervals available.
- Strong baseline that most complex models must beat; frequently the best choice with small datasets.
- Cheap to serve (a dot product) and composable with splines, interactions, GAMs, and residual boosting.
Trade-offs & Constraints
- Assumes additive linear effects; misses interactions and curvature unless you engineer them explicitly.
- Squared loss is dominated by outliers, and heavy-tailed targets distort β̂ badly.
- Multicollinearity makes coefficients unstable and uninterpretable even when predictions are fine.
- Extrapolates blindly outside the observed feature range — a silent production failure mode.
- Biased coefficients under endogeneity, omitted variables, or measurement error, which more data cannot cure.
Home price prediction has historically been built on least-squares foundations: square footage, bedroom/bathroom counts, lot size, and location dummies enter as additive linear effects with interaction and polynomial terms. The model is trained fast, refreshed continuously, and its coefficients provide the human-auditable "why" that price-explanation products require, with tree ensembles layered on top for residual structure.
Staff+ Engineering Takeaways
- OLS minimizes ‖y − Xβ‖² and has the closed-form solution β̂ = (XᵀX)⁻¹Xᵀy — convex, exact, no learning rate.
- Squared-error fitting is maximum likelihood under Gaussian noise, so outliers and heavy tails distort β̂.
- Fitted values are the orthogonal projection of y onto the column space of X, and coefficients are conditional associations — not causal effects — so endogeneity biases them at any sample size.
- Multicollinearity leaves predictions intact but makes coefficients unstable and uninterpretable — detect with VIF and treat with ridge, PCA, or feature removal; when p > n, OLS is undefined or exactly interpolating.
- Report out-of-sample RMSE and R² against a predict-the-mean baseline; in-sample R² always rewards pointless complexity.
- Linear models extrapolate blindly, so production use needs feature-range monitoring and documented scoring envelopes.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Adding a feature that is an exact linear combination of existing features causes:
How clear and actionable was this distributed systems breakdown?