Data & Programming Foundations
Real ML is 80% data work.
All Topics in Phase 2
0 of 15 completedPython loops are slow because they check every number one by one. NumPy fixes this by storing numbers in one tidy block of same-typed memory, so whole-array math (vectorization) and shape-stretching math (broadcasting) run in compiled code — the numerical backbone of almost every AI/ML library.
Real data arrives as messy tables — mixed types, missing values, several files. A pandas DataFrame is a labeled, spreadsheet-like table that lets you load, clean, join and reshape that data until it becomes the numeric array a model can eat.
A wall of numbers hides its patterns; a picture reveals them. Matplotlib is the full-control canvas and brushes; Seaborn is the pre-stocked paint-by-numbers kit built on top of it. Together they turn data into EDA insight and model-report figures.
A feature is one measurable input column fed to a model; stack them and you get the design matrix X. Learn the feature types, the (n_samples, n_features) convention, and why feature quality — not model fancy-ness — sets the accuracy ceiling.
The label (or target) is the answer key a supervised model is graded against while it learns. Learn the terminology, how classification vs regression targets pick your metrics, and why defining the target well shapes the whole project.
Some data sits in tidy rows and columns (tables); some is text, images, audio and video with no schema at all. This topic contrasts the two families and shows how the data type dictates storage, preprocessing, model choice and hardware.
Real data arrives messy: gaps, typos, dates as text, fake 999 readings. Preprocessing is the cooking step — wash, peel, chop — that turns it into a clean numeric matrix a model can actually read, without ever peeking at the test set.
Age runs 0–100, income runs 0–1,000,000 — on one ruler, income shouts and age whispers. Scaling fixes that. Min-Max stretches every value into [0,1]; Z-score standardization recenters to mean 0 and spread 1. This topic shows both formulas, tiny worked examples, and when to pick each.
Models read numbers, not words like "red" or "Delhi". Encoding converts categories into numbers without smuggling in fake relationships: one-hot for unordered labels, ordinal for genuinely ranked ones, and target/embedding tricks for columns with thousands of categories.
A student who rewrites the answer key while practicing looks perfect but learns nothing. Splitting data into train (study), validation (mock exams) and a sealed test set (one final exam) is what turns a model's self-reported score into an honest prediction of how it will do in the real world.
Data leakage is reading tomorrow's newspaper to answer today's quiz: the model secretly gets information it will not have when it must actually predict. Offline scores look amazing, production collapses. This topic catalogs the four ways it sneaks in and the discipline that keeps it out.
Models only see the columns you hand them. Feature engineering is reshaping raw data so the signal becomes visible — ratios, interactions, date parts, rolling aggregates — and on tabular problems it beats model choice as the biggest accuracy lever.
A packed suitcase with 40 clothes for a 3-day trip makes you slower, not readier. Feature engineering (Topic 36) adds columns; selection subtracts the useless ones — via filters (cheap statistics), wrappers (try-subsets) and embedded methods (Lasso/tree importances) — to cut noise, overfitting and cost.
A smoke detector that never beeps is right 99.9% of the time and worthless 100% of the time. When fraud, disease or churn is rare, accuracy rewards the model for ignoring exactly the class you care about. This topic covers better metrics, resampling, and cost-sensitive loss.