Phase 6Days 109-128

NLP Foundations and Classic Neural NLP

Tokenization, n-grams, TF-IDF, word2vec, GloVe, sequence labeling, parsing, RNNs, LSTMs, seq2seq, attention, QA, summarization

Phase Goal

Make NLP deep before transformers: understand language tasks, classic baselines, embeddings, sequence models, attention, evaluation, and linguistic error analysis.

Progress

Day 109: NLP Map and Linguistic Foundations

Day 110: Text Normalization and Tokenization Basics

Day 111: N-Gram Language Models

Day 112: Naive Bayes for Text

Day 113: TF-IDF and Sparse Linear Models

Day 114: Word2Vec Skip-Gram

Day 115: GloVe and Co-Occurrence Embeddings

Day 116: Embedding Evaluation and Bias

Day 117: Sequence Labelling and NER

Day 118: Dependency Parsing and Syntax

Day 119: RNNs for Language

Day 120: LSTM and GRU Mechanics

Day 121: Seq2Seq Models

Day 122: Attention Before Transformers

Day 123: Machine Translation and Summarization

Day 124: Question Answering and Information Extraction

Day 125: Hugging Face Datasets and Tokenizers

Day 126: Classic NLP Project Design

Day 127: NLP Error Analysis Workshop

Day 128: Capstone: Classic NLP System

Capstone
Capstone: Classic NLP System

Build a text classification, NER, QA, summarization, or retrieval baseline system using classic NLP and early neural methods before transformers.

  • CS224N word vectors and neural NLP foundations
  • Jurafsky and Martin-style task framing
  • Dataset audit, model comparison, error report, and README

Phase Complete!

After this phase, you'll be able to:

  • Build classic NLP and sparse text baselines
  • Train and evaluate word embeddings
  • Understand RNN, LSTM, seq2seq, and attention mechanics
  • Analyze NLP errors by language patterns and data issues

You can approach transformers with real NLP foundations instead of only copying Hugging Face examples. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.