What Is a Feature?
A feature is one measurable input column fed to a model; stack them and you get the design matrix X. Learn the feature types, the (n_samples, n_features) convention, and why feature quality — not model fancy-ness — sets the accuracy ceiling.
From Columns to the Design Matrix
Selected and transformed columns become features, arranged as columns of a design matrix X whose rows are samples and which the model maps to a target.
01.The Problem: A Model Cannot "See" a House
You want a program to price houses.
A model cannot look at brick, mortar or a garden. It only crunches numbers (from topic 25: at bottom, everything is a numeric array).
So the question becomes
How do I describe "this house" in numbers a model can digest?
You answer that question with features.
02.The Idea in Plain Words: One Measurable Property Per Column
A feature is
one measurable property of a sample, used as an input to a model.
- A sample is one thing you want to predict about — one house, one email, one patient.
- A feature is one number describing it — area, bedrooms, age.
Collect the features for every sample as columns and you get the design matrix X:
rows are samples (observations), columns are features.
A model learns a mapping from this matrix to a target value (the answer you want, e.g., the price).
In scikit-learn the convention is fixed and worth memorizing:
Xis 2-D with shape(n_samples, n_features)- the target
yis 1-D with shape(n_samples,)
Features are numeric in the matrix even when the original data was text or categorical — turning them numeric is exactly what encoding does (a later topic).
import numpy as np
# 3 houses; features = [area_sqm, bedrooms, age_years]
X = np.array([
[70, 2, 10],
[95, 3, 5],
[120, 3, 2],
])
y = np.array([250, 340, 410]) # price in $k
print(X.shape) # (3, 3) -> (n_samples, n_features)03.A Simple Worked Example: Reading the Matrix
Take the X from the code above and read it out loud, like a table:
codearea bedrooms age -> price (y) house 1 70 2 10 -> 250 house 2 95 3 5 -> 340 house 3 120 3 2 -> 410 ^ feature columns ^ one row = one sample
- Row 1 says: "a 70 m², 2-bedroom, 10-year-old house sold for $250k."
- The model's job: learn how the three columns combine to predict the price.
- Sanity peek: as
areagrows (70 → 95 → 120), price grows (250 → 340 → 410). Area carries signal about price. That is what makes it a good feature.
Notice one house is described by three numbers, and ten houses would be ten rows. Same idea, any size.
04.Types of Features
Features differ by what they measure, and each type has preferred handling:
- Numerical / continuous — age, price, temperature; any real value in a range.
- Numerical / discrete (count) — number of clicks, rooms; whole numbers.
- Categorical (nominal) — color, city; unordered labels. "Red" is not greater than "blue".
- Ordinal — size S/M/L; ordered labels with meaningful rank. M really is between S and L.
- Binary — yes/no, 0/1 flags.
- Datetime — timestamps, often broken into year, month, day-of-week, hour.
05.Features vs Raw Columns
Columns and features are not the same thing.
- Not every column should be a feature. An ID column is usually dropped (no signal).
- One column can spawn several features. A single timestamp can yield hour-of-day, is-weekend and days-since-signup.
Carry this analogy through the topic: your model is a chef, and the features are the ingredients.
A world-class chef cannot cook a great meal from rotten tomatoes.
- Great ingredients (informative features) → even a simple recipe (a linear model) shines.
- Bad ingredients (useless columns) → no recipe saves the dish.
Good ML is less about exotic models and more about representing information through the right features.
06.Feature Quality Determines the Ceiling — and Why AI Cares
A model can only exploit signal that is present in its inputs.
What if the features simply do not correlate with the target?
Then no amount of tuning helps — this is why feature engineering and feature selection (later topics) have outsized impact.
The failure modes to remember:
- Irrelevant features add noise and slow training.
- Redundant features (the same information twice) do the same, and can destabilize some models.
- Leakage-prone features (information that would not exist at prediction time — a separate topic) make evaluation dishonest: the model looks brilliant in testing and fails in production.
Why AI cares: every model in this course — regression, trees, neural nets, LLM training data — consumes features. Before you ever choose an algorithm, you decide what to feed it. That decision is the first and often the biggest lever on accuracy.
Architectural Trade-offs & Production Realities
Architectural Advantages
- More informative features raise the achievable accuracy ceiling.
- Typed handling (scaling, encoding) makes models converge faster and fairer.
- Feature-centric thinking keeps you focused on data, not just algorithms.
Trade-offs & Constraints
- Many features risk redundancy, noise and the curse of dimensionality.
- Poorly chosen features can leak target information and inflate metrics.
- Feature handling adds pipeline complexity and versioning needs.
A regression model uses area, bedrooms and age as features but drops the listing ID; area and age correlate strongly with price while the ID does not, illustrating that informative columns — not all columns — drive predictions.
Staff+ Engineering Takeaways
- A feature is one measurable input column; together they form design matrix X.
- scikit-learn expects X of shape (n_samples, n_features) and 1-D y.
- Features are numeric in the matrix even when the source data is categorical or datetime.
- Feature type (numerical, ordinal, nominal, binary, temporal) drives how you preprocess it.
- The informativeness of features, not model fancy-ness, sets the accuracy ceiling.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
In the scikit-learn design matrix convention, X has shape:
How clear and actionable was this distributed systems breakdown?