PHASE 2 CURRICULUM

Data & Programming Foundations

Progress0 of 15 (0%)

Real ML is 80% data work.

Key Architectural Domains & Syllabus
You will learn NumPy arrays and vectorization
pandas for tabular wrangling
Matplotlib and Seaborn for exploratory plots
features vs labels
preprocessing
normalization vs standardization
categorical encoding
train/validation/test splits
data leakage
feature engineering and selection
class imbalance
EDA workflows used daily by practicing ML engineers
15 In-Depth Topics ~120 Minutes Reading Time Interactive Quizzes & Assessments

All Topics in Phase 2

0 of 15 completed

Python loops are slow because they check every number one by one. NumPy fixes this by storing numbers in one tidy block of same-typed memory, so whole-array math (vectorization) and shape-stretching math (broadcasting) run in compiled code — the numerical backbone of almost every AI/ML library.

10 min read•3 Quiz Questions

Real data arrives as messy tables — mixed types, missing values, several files. A pandas DataFrame is a labeled, spreadsheet-like table that lets you load, clean, join and reshape that data until it becomes the numeric array a model can eat.

11 min read•3 Quiz Questions

A wall of numbers hides its patterns; a picture reveals them. Matplotlib is the full-control canvas and brushes; Seaborn is the pre-stocked paint-by-numbers kit built on top of it. Together they turn data into EDA insight and model-report figures.

10 min read•3 Quiz Questions
#28What Is a Feature?BeginnerFREE

A feature is one measurable input column fed to a model; stack them and you get the design matrix X. Learn the feature types, the (n_samples, n_features) convention, and why feature quality — not model fancy-ness — sets the accuracy ceiling.

9 min read•3 Quiz Questions

The label (or target) is the answer key a supervised model is graded against while it learns. Learn the terminology, how classification vs regression targets pick your metrics, and why defining the target well shapes the whole project.

9 min read•3 Quiz Questions

Some data sits in tidy rows and columns (tables); some is text, images, audio and video with no schema at all. This topic contrasts the two families and shows how the data type dictates storage, preprocessing, model choice and hardware.

9 min read•3 Quiz Questions

Real data arrives messy: gaps, typos, dates as text, fake 999 readings. Preprocessing is the cooking step — wash, peel, chop — that turns it into a clean numeric matrix a model can actually read, without ever peeking at the test set.

10 min read•3 Quiz Questions

Age runs 0–100, income runs 0–1,000,000 — on one ruler, income shouts and age whispers. Scaling fixes that. Min-Max stretches every value into [0,1]; Z-score standardization recenters to mean 0 and spread 1. This topic shows both formulas, tiny worked examples, and when to pick each.

9 min read•3 Quiz Questions

Models read numbers, not words like "red" or "Delhi". Encoding converts categories into numbers without smuggling in fake relationships: one-hot for unordered labels, ordinal for genuinely ranked ones, and target/embedding tricks for columns with thousands of categories.

9 min read•3 Quiz Questions

A student who rewrites the answer key while practicing looks perfect but learns nothing. Splitting data into train (study), validation (mock exams) and a sealed test set (one final exam) is what turns a model's self-reported score into an honest prediction of how it will do in the real world.

9 min read•3 Quiz Questions
#35Data LeakageIntermediate PRO

Data leakage is reading tomorrow's newspaper to answer today's quiz: the model secretly gets information it will not have when it must actually predict. Offline scores look amazing, production collapses. This topic catalogs the four ways it sneaks in and the discipline that keeps it out.

9 min read•3 Quiz Questions
#36Feature EngineeringIntermediate PRO

Models only see the columns you hand them. Feature engineering is reshaping raw data so the signal becomes visible — ratios, interactions, date parts, rolling aggregates — and on tabular problems it beats model choice as the biggest accuracy lever.

10 min read•3 Quiz Questions
#37Feature SelectionIntermediate PRO

A packed suitcase with 40 clothes for a 3-day trip makes you slower, not readier. Feature engineering (Topic 36) adds columns; selection subtracts the useless ones — via filters (cheap statistics), wrappers (try-subsets) and embedded methods (Lasso/tree importances) — to cut noise, overfitting and cost.

9 min read•3 Quiz Questions
#38Class ImbalanceIntermediate PRO

A smoke detector that never beeps is right 99.9% of the time and worthless 100% of the time. When fraud, disease or churn is rare, accuracy rewards the model for ignoring exactly the class you care about. This topic covers better metrics, resampling, and cost-sensitive loss.

10 min read•3 Quiz Questions

EDA is the home inspection before you buy the house: a structured look at distributions, missing values, outliers and relationships before any model is trained. It turns preprocessing and feature choices from guesses into evidence, and catches leaks and disasters early.

10 min read•3 Quiz Questions