TOPIC #45Beginner 10 min read

Linear Regression

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

The oldest useful model in machine learning: you predict a number as base + weights x features, and fit those weights by least squares. It has an exact closed-form solve, clean diagnostics (R2, RMSE), and strict assumptions — and it fails silently on collinear features and on extrapolation.

Fitting a Line, Then Auditing It 📈

Least squares has an exact solution, so the modeling work is not the fit — it is the residual diagnostics and the collinearity audit that decide whether coefficients are trustworthy.

Fitting a Line, Then Auditing It 📈
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Predicting One Number From Others

Imagine you run a small taxi company.

For every past ride you know two things: how far it went (km) and how long it took (min).

You also know the final fare.

Now a new ride comes in. You want to guess its fare from its km and minutes.

Insight

How do you turn some inputs into one predicted number?

The oldest, most honest answer is a straight-line rule:

  • a fixed base you pay just for getting in,
  • plus a rate per km times the distance,
  • plus a rate per minute times the time.

That rule is linear regression.

In one plain sentence:

Insight

Linear regression predicts a number by multiplying each input by a learned weight and adding the results.

The weights (base, rate-per-km, rate-per-min) are learned from your old rides.

Learning means: find the weights whose guesses land closest to the real fares you collected.

Sounds obvious. But three things are actually tricky:

  • How do we measure "closest"?
  • How do we find those weights, fast?
  • When does a straight line quietly lie to us?

Everything below answers those three.

02.The Model in Plain Words: Base + Weights × Features

The model says the target is a weighted sum of features plus some noise:

y = β₀ + β₁x₁ + … + βₚxₚ + ε

Read it slowly:

  • y is the number you predict (the fare).
  • x₁ … xₚ are the inputs (km, minutes, and any others).
  • β₀ is the base — the intercept.
  • β₁ … βₚ are the rates — the coefficients (or weights).
  • ε is everything your features don't explain — the noise.

In matrix form the same line is written y = Xβ + ε, where X stacks all rows.

One word matters: linear.

It means linear in the parameters (the β's), not in the raw input x.

So you can bend the line by transforming features first:

y = β₀ + β₁x + β₂x²

This is still a linear regression — it is linear in β₁ and β₂.

Polynomial and Fourier features let one straight-line machine fit curved data.

Insight

Non-linearity is handled by feature engineering, not by changing the algorithm.

03.How Do We Measure "Close"? Least Squares

We need one number for "how wrong" the current weights are.

OLS (ordinary least squares) uses the residual sum of squares:

RSS(β) = ‖y − Xβ‖²₂ = Σᵢ (yᵢ − ŷᵢ)²

A residual is just real − predicted for one row. We square each residual, then add them up.

Why square? Two reasons:

  • Big misses get punished extra hard, so a wild error stands out.
  • Squared error is the maximum-likelihood choice exactly when the noise is symmetric and bell-shaped: ε ~ N(0, σ²).

That second point is also a warning: squared error is not universal. If the errors are lopsided or have extreme tails, it is the wrong ruler (see the Loss Functions topic).

A tiny ride dataset makes this concrete. Suppose one input, km, and three rides:

code
ride   km   real fare   guess with β0=10, β1=6
1      1    18          16      -> residual 2  -> squared 4
2      2    30          22      -> residual 8  -> squared 64
3      3    44          28      -> residual 16 -> squared 256

RSS = 4 + 64 + 256 = 324 for these weights.

OLS asks: which β makes RSS the smallest?

Because RSS is a convex quadratic in β (one smooth upward bowl), there is exactly one best answer — no lucky restarts, no local traps.

04.Finding the Weights: The Normal Equations

Because RSS is a convex bowl, calculus finds its bottom by setting the slope to zero. That gives the normal equations:

XᵀXβ = Xᵀy ⟹ β̂ = (XᵀX)⁻¹Xᵀy

Three consequences follow immediately — they are why OLS is both beloved and dangerous.

1. Exact, non-iterative.

No learning rate, no epochs, no local minima — one matrix solve.

With a Cholesky factorization it costs O(np² + p³), which is why linear regression trains in milliseconds on millions of rows.

2. Geometric meaning.

Xβ̂ is the orthogonal projection of y onto the flat space spanned by X's columns. The residual sticks straight out, perpendicular to every feature column:

Σᵢ eᵢxᵢⱼ = 0 for all j

That is also why, with an intercept, the residuals sum to zero. A picture of the projection:

code
        y (real fares)
        |
        |\  residual e (perpendicular)
        |  \
        |    Xβ̂  <- closest point on the feature "flat"
        |  /
        +----------------  column space of X

3. Maximum likelihood under Gaussian noise.

Minimizing squared error equals maximizing likelihood when ε ~ N(0, σ²) — the same ruler point from §3.

05.Is the Line Any Good? R², Adjusted R², and RMSE

Once fitted, how do you report quality?

R² = 1 − RSS/TSS is the fraction of the target's variance the model explains, where TSS = Σ(yᵢ − ȳ)² (ȳ is the mean fare).

It is scale-free, easy to say out loud — and easy to abuse:

  • R² never decreases when you add a feature, so 200 random noise columns can post a proud in-sample R². Adjusted R² penalizes p relative to n; AIC and BIC do the same.
  • Training R² is not a performance claim. Report out-of-sample R² from a held-out split or CV fold.
  • Two baselines keep it honest: predicting the mean gives R² = 0 exactly, and a negative R² means your model is worse than the mean — a real and common outcome on new data.
  • For lopsided or heavy-tailed targets, report MAE and a percentile of absolute error alongside RMSE.
Insight

Coefficients are not causal effects.

βⱼ reads: "the expected change in y per one-unit increase in xⱼ, holding the other included features fixed."

That holding-fixed clause is doing enormous work:

  • It only means something if the other features genuinely capture the confounders.
  • It changes entirely when you add or drop a feature.
  • If one variable's effect depends on another, you must add an interaction term (a product like x₁·x₂); once an interaction is present, plain main-effect reading stops.
python— OLS with diagnostics in scikit-learn and statsmodels
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score, mean_squared_error
from sklearn.model_selection import cross_val_score
import numpy as np

model = LinearRegression().fit(X_train, y_train)
resid = y_val - model.predict(X_val)

print("coef:", np.round(model.coef_, 4))
print("intercept:", round(model.intercept_, 4))
print("in-sample R2:", round(model.score(X_train, y_train), 4))
print("out-of-sample R2:", round(r2_score(y_val, model.predict(X_val)), 4))
print("RMSE:", round(np.sqrt(mean_squared_error(y_val, model.predict(X_val))), 4))

# Honest generalization estimate, not the optimistic train score
print("cv R2:", cross_val_score(model, X, y, cv=5, scoring="r2").mean().round(4))

# For standard errors and p-values use statsmodels (adds constant for intercept)
# import statsmodels.api as sm
# sm.OLS(y_train, sm.add_constant(X_train)).fit().summary()

06.The Assumptions and Their Failure Modes

OLS is unbiased, consistent, and (under the classical conditions) the best linear unbiased estimator. Each guarantee has a matching assumption, and each violation leaves a signature in the residuals.

  1. Linearity in parameters / correct features. Left-out curvature shows as a U-shape in residuals-versus-fitted plots. Cure: add polynomial, spline, binning, or interaction terms.
  2. Exogeneity: E[ε|x] = 0. If a predictor correlates with the noise — omitted variables, measurement error in x, or choosing x after observing y (reverse causation) — coefficients are biased, and no amount of data fixes it. Measurement error pulls the affected coefficient toward zero; omitted-variable bias hides the effect inside correlated features.
  3. Homoscedasticity (constant error variance) and no autocorrelation. A funnel shape in residual plots signals heteroscedasticity: it does not bias β̂ but makes standard errors and confidence intervals wrong. Cure: weighted least squares, a log-transform of y, or robust (sandwich/HC) standard errors. Serially correlated residuals in time-series data invalidate naive p-values entirely.
  4. Normal errors — only required for exact confidence intervals and hypothesis tests with small n, not for prediction. With large samples, the central limit theorem largely rescues inference.
Insight

Multicollinearity is the fifth practical problem.

When features are near-linear combinations of each other, XᵀX becomes ill-conditioned and its inverse unstable. So β̂ swings wildly with tiny data changes, even though predictions stay fine.

  • Detect it with VIF (VIF > 5-10 is concerning; > 30 is severe) or the condition number of XᵀX.
  • Symptoms: coefficient flips between retrains, huge standard errors, nonsense signs.
  • Remedies: drop or combine redundant features, use ridge (topic 55), or use PCA.

Also beware high-dimensional p > n, where XᵀX is singular and OLS exactly interpolates the training data with zero residual and no generalization. Regularized or reduced-rank methods become mandatory there.

07.Where Linear Regression Still Wins in Production

Back to the taxi: your straight-line fare rule is the first model you always fit, and often the final one. The reasons are economic as much as statistical:

  • Interpretability and auditability. Coefficients are reportable to regulators, finance teams, and customers. Scorecards in credit and insurance still lean on linear/logistic forms for this reason.
  • Training and serving cost. One matrix solve offline, a dot product online — sub-millisecond at any QPS, no GPU, no feature-store round trips beyond the feature vector.
  • Sample efficiency with few signal-rich features. When n is small relative to p, a linear model often beats a flexible one simply because it cannot overfit elaborate structure.
  • Composability. Residual boosting is initialized from a linear fit; generalized additive models sum learned linear pieces; embeddings plus a linear head is a standard production pattern.

When linear regression is the wrong tool: strongly non-monotonic or saturating relationships you cannot feature-engineer away, targets with hard bounds or count/ordinal structure, heavy-tailed noise where squared loss is hijacked by a handful of rows, and interactions across many variables. The classical escalations are polynomial/spline features (staying linear), trees and gradient boosting (topic 56), or a generalized linear model with a link function (topic 46).

The professional habit is always the same sequence:

Insight

fit the linear baseline → measure validation RMSE → plot residuals → justify each added unit of complexity by the error it removes.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Closed-form exact fit in one solve — no hyperparameters, no convergence tuning, trains on billions of rows via streaming least squares.
  • Coefficients are directly interpretable and reportable, with standard errors and confidence intervals available.
  • Strong baseline that most complex models must beat; frequently the best choice with small datasets.
  • Cheap to serve (a dot product) and composable with splines, interactions, GAMs, and residual boosting.

Trade-offs & Constraints

  • Assumes additive linear effects; misses interactions and curvature unless you engineer them explicitly.
  • Squared loss is dominated by outliers, and heavy-tailed targets distort β̂ badly.
  • Multicollinearity makes coefficients unstable and uninterpretable even when predictions are fine.
  • Extrapolates blindly outside the observed feature range — a silent production failure mode.
  • Biased coefficients under endogeneity, omitted variables, or measurement error, which more data cannot cure.
Production Implementation in Big Tech
Zillow (early Zestimate) and most real-estate valuation models• Housing price estimation from structured attributes

Home price prediction has historically been built on least-squares foundations: square footage, bedroom/bathroom counts, lot size, and location dummies enter as additive linear effects with interaction and polynomial terms. The model is trained fast, refreshed continuously, and its coefficients provide the human-auditable "why" that price-explanation products require, with tree ensembles layered on top for residual structure.

Staff+ Engineering Takeaways

  • OLS minimizes ‖y − Xβ‖² and has the closed-form solution β̂ = (XᵀX)⁻¹Xᵀy — convex, exact, no learning rate.
  • Squared-error fitting is maximum likelihood under Gaussian noise, so outliers and heavy tails distort β̂.
  • Fitted values are the orthogonal projection of y onto the column space of X, and coefficients are conditional associations — not causal effects — so endogeneity biases them at any sample size.
  • Multicollinearity leaves predictions intact but makes coefficients unstable and uninterpretable — detect with VIF and treat with ridge, PCA, or feature removal; when p > n, OLS is undefined or exactly interpolating.
  • Report out-of-sample RMSE and R² against a predict-the-mean baseline; in-sample R² always rewards pointless complexity.
  • Linear models extrapolate blindly, so production use needs feature-range monitoring and documented scoring envelopes.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Adding a feature that is an exact linear combination of existing features causes:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?