/ ML, DL & GenAI / Phase 1 Phase 1 Days 1-24
ML Foundations: From Data Analyst to First Real Models Mental model, generalization, just-enough math, linear regression and classification from scratch, validation discipline, capstone
Phase Goal Turn a competent data analyst into someone who can frame an ML problem, prove generalization, derive and code linear models from scratch, evaluate them honestly, and ship a defensible capstone — before touching trees, boosters, or neural nets.
Day 1: What Machine Learning Actually Is Inspect the descriptions and schemas of three real datasets (Titanic, Ames housing, 20 Newsgroups); record example, features, target, prediction timing and task family. Build a tiny learned rule in Python and compare its arithmetic with NumPy. No imports of sklearn. Use only tools already taught: explain an early library call fully, then rebuild its mechanism once prerequisites are ready; offer an executable small-data or CPU route Add shape, dtype, and range assertions so silent bugs become loud bugs Compare with a taught reference under matched data, objective, dtype, and justified tolerance; do not demand identical stochastic runs
Evaluation & Failure Modes Decide when ML is the wrong tool: use direct rules for specified policies and analytics for calculations on known records. Build a "do I even need ML?" checklist. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 2: Your First Model in 30 Lines
Evaluation & Failure Modes Compare train and held-out predictions without claiming perfect accuracy always means leakage. Tomorrow investigates generalization and why a split must represent intended use. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 3: The Generalization Problem
Evaluation & Failure Modes Use the gap as a clue with competing explanations. Diagnose target, contamination, temporal and preprocessing leakage; demonstrate group/time boundaries and a training-only fitted transform. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 4: Your Second Model: Linear Regression with Scikit-Learn
Evaluation & Failure Modes When linear regression fails ugly: non-linear relationships, outliers that drag the line, and the residual plot that tells you so. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 5: Your Third Model: Logistic Regression on Titanic
Evaluation & Failure Modes The three things every beginner does wrong on Titanic: leaks the test set during EDA, ignores the base rate, and trusts accuracy on an imbalanced problem. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 6: End-to-End Mini-Project: Ship a Notebook Lab: pick a Kaggle or UCI dataset, write a 6-section notebook (problem, data, baseline, model, evaluation, what I would do next), publish it to GitHub. Use only tools already taught: explain an early library call fully, then rebuild its mechanism once prerequisites are ready; offer an executable small-data or CPU route Add shape, dtype, and range assertions so silent bugs become loud bugs Compare with a taught reference under matched data, objective, dtype, and justified tolerance; do not demand identical stochastic runs
Evaluation & Failure Modes The four readability sins of beginner notebooks: no headings, no baseline, no test set, no written conclusion. Self-grade your notebook against this rubric. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 7: How a Model Actually Works: Vectors, Matrices, Shapes
Evaluation & Failure Modes The three shape bugs that break every junior ML engineer: silent broadcasting, transposed weights, and label/prediction shape mismatch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 8: Calculus and Gradients for Learning
Evaluation & Failure Modes Finite-difference gradient checks: the 10-line sanity test that catches silent bugs that destroy weeks of training time. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 9: Probability and Likelihood for ML
Evaluation & Failure Modes Base-rate failures: why "the model is 95% accurate" almost always means nothing. The cancer-test problem worked end to end. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 10: Loss Functions: How Models Know They Are Wrong
Evaluation & Failure Modes Choosing the loss before the model. How asymmetric business cost (false positives cheap, false negatives ruinous) maps directly to weighted loss. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 11: Gradient Descent From Scratch
Evaluation & Failure Modes Reading a loss curve like a doctor reads an ECG: divergence, oscillation, plateau, and the slow-burn convergence that looks fine but is actually stuck. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 12: Rebuild Linear Regression From Scratch
Evaluation & Failure Modes Multicollinearity and ill-conditioning: when the matrix you must invert is almost singular, and what it does to your coefficients. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 13: Linear Regression With Gradient Descent at Scale
Evaluation & Failure Modes Why your loss diverges and the four most common fixes: scale features, lower the LR, add a bias, switch to mini-batch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 14: Regularization: Ridge, Lasso, and Elastic Net
Evaluation & Failure Modes When not to regularize. Why standardization is non-optional with L1/L2. The "best lambda" trap on tiny validation sets. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 15: Logistic Regression: From Line to Probability
Evaluation & Failure Modes The failure mode that makes beginners pick MSE for classification, and the loss surface that punishes them with stuck gradients. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 16: Rebuild Logistic Regression From Scratch
Evaluation & Failure Modes Silent failure: probabilities that look right but are wildly miscalibrated. The reliability diagram introduced as a teaser for the calibration day. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 17: Softmax and Multiclass Classification
Evaluation & Failure Modes Per-class error analysis. Why one accuracy number for a 10-class problem hides everything important. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 18: Classification Metrics That Actually Matter
Evaluation & Failure Modes Imbalanced data: how 99% accuracy can mean your model is useless. The fraud-detection worked example. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 19: Calibration, Thresholds, and Business Cost
Evaluation & Failure Modes Cost-blind models in production: how a "great" model loses money because nobody put dollars on the metric. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 20: Bias, Variance, and Learning Curves
Evaluation & Failure Modes Prescribing the next experiment from a learning curve in 30 seconds: more data, more capacity, more regularization, or different features. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 21: Cross-Validation Done Right
Evaluation & Failure Modes Validation strategies that quietly cheat: preprocessing before splitting, target encoding before splitting, and time-series with random folds. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 22: Data Audits and the Leakage Hunt
Evaluation & Failure Modes Data cards as a release artifact. The patterns that make data trustworthy enough to model on. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 23: Experiment Discipline: Configs, Seeds, Model Cards
Evaluation & Failure Modes Reading and writing the model card an honest engineer can defend in a review meeting. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 24: Capstone: Linear Models, End to End Final lab: pick one real tabular dataset (UCI, OpenML, or domain-specific). Deliver a scratch model, a sklearn parity model, and a written report. Use only tools already taught: explain an early library call fully, then rebuild its mechanism once prerequisites are ready; offer an executable small-data or CPU route Add shape, dtype, and range assertions so silent bugs become loud bugs Compare with a taught reference under matched data, objective, dtype, and justified tolerance; do not demand identical stochastic runs
Evaluation & Failure Modes The readiness report that decides whether the project should ship, get more data, or be killed. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Capstone: Linear Models, End to End On one real tabular dataset, deliver a data audit, a from-scratch linear model, a sklearn parity model, regularization and calibration sweeps, cross-validated metrics with uncertainty, and a one-page model card that an interviewer or manager could read in two minutes.
CS229-style derivation notes for the loss and gradient NumPy scratch implementation matched against sklearn Data card, model card, notebook, and READMEAfter this phase, you'll be able to:
Explain ML in one sentence and decide when to use it Reason about generalization, leakage, and the train-test gap Derive and implement linear and logistic regression from scratch with gradient descent Evaluate models with the right metric, calibration, and CV, and document them in a model card You can deliver a linear-models case study end to end — audit, scratch model, sklearn parity, calibration, cross-validated intervals, and a written model card. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.