Phase 8Days 149-176

Reinforcement Learning, LLM Post-Training, and Evaluation

Bandits, MDPs, value learning, policy gradients and PPO before RLHF; SFT, DPO, calibrated decisions, evaluation and serving

Phase Goal

Treat LLMs like serious systems: adapt them, evaluate them, test safety and factuality, understand reasoning and post-training, and make release decisions with evidence.

Progress

Day 149: Bandits and Learning from Feedback

Day 150: States, Actions, Returns, and Markov Decision Processes

Day 151: Bellman Equations and Dynamic Programming for RL

Day 152: Monte Carlo, Temporal Difference, and Q-Learning

Day 153: Deep Q-Learning and Approximation Failures

Day 154: Policy Gradients and REINFORCE

Day 155: Actor-Critic, Advantages, and PPO

Day 156: Reward Design, Offline Data, and the RL Readiness Gate

Day 157: LLM Lifecycle and Product Fit

Day 158: Instruction Tuning and Chat Templates

Day 159: Supervised Fine-Tuning Data Quality

Day 160: Preference Data and Reward Modeling

Day 161: RLHF Conceptual Pipeline

Day 162: DPO and Modern Preference Optimization

Day 163: Reasoning Models and Test-Time Compute

Day 164: Multilinguality and Tokenization Fairness

Day 165: LLM Evaluation Design

Day 166: LLM-as-Judge and Human Evaluation

Day 167: Safety, Refusal, Policy Evaluation

Day 168: Hallucination and Factuality

Day 169: Interpretability and Mechanistic Probing

Day 170: Model Compression and Quantization

Day 171: Serving Open Models Locally

Day 172: Prompt Security and Adversarial Robustness

Day 173: Data Governance for LLMs

Day 174: Model Selection and Leaderboards

Day 175: LLM Research Reading Sprint

Day 176: Capstone: Fine-Tuned and Evaluated LLM

Capstone
Capstone: Fine-Tuned and Evaluated LLM

Fine-tune or adapt an open model with LoRA or an SFT-style workflow, then evaluate usefulness, safety, factuality, cost, and release risk.

  • CS224N post-training and benchmarking topics
  • DeepLearning.AI LLM lifecycle framing
  • Adapter, eval harness, safety notes, and README

Phase Complete!

After this phase, you'll be able to:

  • Curate SFT and preference datasets
  • Understand RLHF, DPO, reasoning, and PEFT tradeoffs
  • Build LLM eval, safety, factuality, and judge workflows
  • Serve and select models under real constraints

You can adapt and evaluate LLMs with rigor instead of relying on vibes. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.