/ ML, DL & GenAI / Phase 2 Phase 2 Days 25-44
Classical Supervised ML in Depth KNN, Naive Bayes, trees, forests, gradient boosting, XGBoost/LightGBM/CatBoost, SVMs, pipelines, tuning, calibration, interpretation, packaging
Phase Goal Become dangerous on tabular data: derive and tune every classical supervised algorithm, build leakage-safe sklearn pipelines, calibrate and interpret models, and ship a packaged tuned booster with a written model card.
Day 25: KNN and Distance-Based Learning
Evaluation & Failure Modes When KNN silently fails: irrelevant features, mixed scales, high-dimensional data, and unbalanced classes. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 26: Naive Bayes and Generative Classification
Evaluation & Failure Modes When the independence lie still works (text, high-d sparse) and when it spectacularly does not (correlated tabular features). Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 27: Generative vs Discriminative Models
Evaluation & Failure Modes Detecting assumption mismatch in generative models with quantile and covariance diagnostics. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 28: Decision Trees: Splits, Impurity, Pruning
Evaluation & Failure Modes Why a single tree always overfits and the four knobs that tame it: depth, min-samples-leaf, min-impurity-decrease, ccp_alpha. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 29: Bagging and Random Forests
Evaluation & Failure Modes Feature-importance traps: how correlated features distort attribution, and the permutation-importance fix. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 30: Boosting Intuition: AdaBoost and Additive Models
Evaluation & Failure Modes Why boosting overfits if you let it run forever. Early stopping as a real, used-in-production lever. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 31: Gradient Boosting From Scratch
Evaluation & Failure Modes The tuning order for boosters: which knob to turn first, second, and which one you should leave alone. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 32: XGBoost, LightGBM, and CatBoost
Evaluation & Failure Modes Tabular competition traps: leakage in CV folds, target encoding without folds, feature explosion, eval-set overfitting. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 33: SVMs: Margins, Hinge Loss, Soft Margins
Evaluation & Failure Modes When SVMs beat logistic regression and when they are just slower with no upside. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 34: Kernel SVMs and the Kernel Trick
Evaluation & Failure Modes Why kernel SVMs do not scale and why boosted trees beat them on tabular data today. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 35: Pipelines and Leakage-Safe Preprocessing
Evaluation & Failure Modes Serving skew: when train-time and inference-time preprocessing silently diverge in production. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 36: Categorical Encoding Deep Dive
Evaluation & Failure Modes Category drift in production. Which encoders survive it and which silently corrupt predictions. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 37: Imbalanced Classification
Evaluation & Failure Modes Imbalance pitfalls: SMOTE-before-CV leakage, metric mismatch, and the "great recall" trap. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 38: Hyperparameter Tuning: Grid, Random, Bayesian
Evaluation & Failure Modes Tuning luck: when the "best" hyperparameters are within noise of the previous run, and how to know. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 39: Interpretation: Permutation, PDP, ICE, SHAP
Evaluation & Failure Modes The four ways feature importance lies: correlation, leakage, scale, and interactions. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 40: Causal Caution: Prediction vs Intervention
Evaluation & Failure Modes Overclaiming causality: the sentence patterns that should make you stop and the language to use instead. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 41: Time Series Baselines and Walk-Forward CV
Evaluation & Failure Modes Future leakage: the feature you accidentally engineered from tomorrow, and the checklist that catches it. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 42: Calibration and Probability Repair
Evaluation & Failure Modes When calibration silently breaks under distribution shift and how to monitor for it. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 43: Packaging and Serving Classical Models
Evaluation & Failure Modes Inference readiness: latency, memory, and the requirements.txt that bites you in six months. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 44: Capstone: Classical ML Case Study Final lab: deliver notebook, tuned pipeline, calibration plots, SHAP report, model card, and a deployable FastAPI artifact. Use only tools already taught: explain an early library call fully, then rebuild its mechanism once prerequisites are ready; offer an executable small-data or CPU route Add shape, dtype, and range assertions so silent bugs become loud bugs Compare with a taught reference under matched data, objective, dtype, and justified tolerance; do not demand identical stochastic runs
Evaluation & Failure Modes The case-study writeup an interviewer can use to evaluate your judgment without re-running your code. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Capstone: Classical ML Case Study Build a supervised tabular ML case study end to end: data audit, baseline, ensemble, tuned booster, calibration, cross-validated metrics with intervals, SHAP-based explanation, model card, and a packaged FastAPI demo.
ISL with Python-style workflow CS229/CMU-style derivations for the core algorithms used Pipeline artifact, model card, notebook, FastAPI demo, and READMEAfter this phase, you'll be able to:
Derive, code, and compare KNN, Naive Bayes, trees, forests, boosters, and SVMs Build leakage-safe sklearn pipelines for mixed numeric/categorical data Tune, calibrate, and interpret models with SHAP and reliability curves Package a tuned tabular model behind FastAPI with a model card You can place on a competitive tabular task and explain every modeling choice in plain English. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.