Phase 9Days 177-196

RAG, Agents, and Production LLM Applications

RAG architecture, ingestion, chunking, embeddings, vector stores, hybrid search, reranking, grounded generation, tool use, agents, observability, security

Phase Goal

Build the most important 2026 AI product skills in depth: reliable RAG, tool use, agents, traces, evals, security, cost controls, and deployed LLM app backends.

Progress

Day 177: RAG Architecture and Failure Modes

Day 178: Document Loading and Parsing

Day 179: Chunking Strategy Deep Dive

Day 180: Embedding Model Selection

Day 181: Vector Stores and Index Operations

Day 182: Hybrid Retrieval with BM25 and Dense Search

Day 183: Reranking and Context Compression

Day 184: Grounded Generation and Citations

Day 185: RAG Evaluation and Observability

Day 186: Advanced RAG Patterns

Day 187: Tool Calling and Function Schemas

Day 188: Agent Loops and ReAct

Day 189: Agent Memory and State

Day 190: Planning, Reflection, Multi-Agent Patterns

Day 191: LangGraph, LlamaIndex, smolagents

Day 192: Agent Evaluation Benchmarks

Day 193: LLM App Backend Engineering

Day 194: Cost, Latency, UX for LLM Apps

Day 195: LLM Security and Production Guardrails

Day 196: Capstone: Production RAG or Agent App

Capstone
Capstone: Production RAG or Agent App

Build a serious RAG or agent product with evals, citations or tool traces, security tests, cost controls, observability notes, and a deployed demo.

  • CS224N agents/RAG/evaluation topics
  • Hugging Face Agents Course frameworks
  • FSDL LLMOps mindset
  • Deployed app, eval harness, traces, and ops README

Phase Complete!

After this phase, you'll be able to:

  • Build and evaluate robust RAG pipelines
  • Implement tool-calling and agent loops
  • Use agent frameworks without losing control of traces and state
  • Ship LLM apps with security, latency, cost, and observability discipline

You can build GenAI products that are testable systems, not only impressive demos. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.