/ ML, DL & GenAI / Phase 6 Phase 6 Days 109-128
NLP Foundations and Classic Neural NLP Tokenization, n-grams, TF-IDF, word2vec, GloVe, sequence labeling, parsing, RNNs, LSTMs, seq2seq, attention, QA, summarization
Phase Goal Make NLP deep before transformers: understand language tasks, classic baselines, embeddings, sequence models, attention, evaluation, and linguistic error analysis.
Day 109: NLP Map and Linguistic Foundations
Evaluation & Failure Modes Language task quality: what makes a "good" NLP dataset and what makes one secretly broken. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 110: Text Normalization and Tokenization Basics
Evaluation & Failure Modes Preprocessing damage: the silent information loss that is hard to detect after the fact. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 111: N-Gram Language Models
Evaluation & Failure Modes LM baseline quality. Why a smart n-gram model is hard to beat on tiny datasets. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 112: Naive Bayes for Text
Evaluation & Failure Modes Text classification errors: the patterns that exposed every chain email of the 2000s. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 113: TF-IDF and Sparse Linear Models
Evaluation & Failure Modes Feature interpretation: what the top weights mean and when they lie. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 114: Word2Vec Skip-Gram
Evaluation & Failure Modes Embedding quality: the analogy tests and their limits. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 115: GloVe and Co-Occurrence Embeddings
Evaluation & Failure Modes Semantic bias as encoded in any large word-embedding model. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 116: Embedding Evaluation and Bias
Evaluation & Failure Modes Embedding misuse in downstream tasks and the patterns to test for before shipping. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 117: Sequence Labelling and NER
Evaluation & Failure Modes Span error analysis. The categories that matter beyond a single F1. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 118: Dependency Parsing and Syntax
Evaluation & Failure Modes Syntax errors and what they reveal about your tokenizer. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 119: RNNs for Language
Evaluation & Failure Modes Sequence memory limits. Why RNNs struggle past a few hundred tokens. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 120: LSTM and GRU Mechanics
Evaluation & Failure Modes Long-sequence behaviour and the bias toward recent tokens. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 121: Seq2Seq Models
Evaluation & Failure Modes Generation errors: repetition, premature stopping, and the empty-output failure. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 122: Attention Before Transformers
Evaluation & Failure Modes Alignment failures and what they tell you about the dataset. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 123: Machine Translation and Summarization
Evaluation & Failure Modes Generation metric limits and why every paper needs a human-eval section. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 124: Question Answering and Information Extraction
Evaluation & Failure Modes Answer grounding: the abstention behaviour that makes a QA system trustworthy. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 125: Hugging Face Datasets and Tokenizers
Evaluation & Failure Modes Data pipeline pitfalls: silent truncation, padding bias, and collator mismatch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 126: Classic NLP Project Design
Evaluation & Failure Modes Project feasibility. The questions to answer before kicking off. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 127: NLP Error Analysis Workshop
Evaluation & Failure Modes NLP robustness: the slices that should ship in every model card. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 128: Capstone: Classic NLP System Final lab: deliver a classic NLP system (classification, NER, QA, summarization, or retrieval). Use only tools already taught: explain an early library call fully, then rebuild its mechanism once prerequisites are ready; offer an executable small-data or CPU route Add shape, dtype, and range assertions so silent bugs become loud bugs Compare with a taught reference under matched data, objective, dtype, and justified tolerance; do not demand identical stochastic runs
Evaluation & Failure Modes NLP capstone quality. The structure that gets you an NLP-engineer interview. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Capstone: Classic NLP System Build a text classification, NER, QA, summarization, or retrieval baseline system using classic NLP and early neural methods before transformers.
CS224N word vectors and neural NLP foundations Jurafsky and Martin-style task framing Dataset audit, model comparison, error report, and READMEAfter this phase, you'll be able to:
Build classic NLP and sparse text baselines Train and evaluate word embeddings Understand RNN, LSTM, seq2seq, and attention mechanics Analyze NLP errors by language patterns and data issues You can approach transformers with real NLP foundations instead of only copying Hugging Face examples. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.