Day 197: Generative Model Taxonomy
Evaluation & Failure Modes Generation quality. The metrics each family looks good on for the wrong reasons. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 198: VAEs and Latent Variable Models
Evaluation & Failure Modes Latent quality. The signs that posterior collapse is silently eating your model. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 199: GANs and Adversarial Training
Evaluation & Failure Modes GAN failure modes: mode collapse, discriminator overpower, and the loss curves that say everything is fine when it is not. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 200: Diffusion Intuition and Noise Schedules
Evaluation & Failure Modes Denoising behaviour. The schedule choices that change everything. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 201: DDPM From Scratch
Evaluation & Failure Modes Sample quality. What changes when you tweak the schedule, the loss, or the U-Net. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 202: Guidance and Conditional Generation
Evaluation & Failure Modes Control failures. The cases where guidance breaks fidelity. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 203: Latent Diffusion and Stable Diffusion
Evaluation & Failure Modes Pipeline assumptions that are not obvious from the API. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 204: Fine-Tuning Diffusion: LoRA and DreamBooth
Evaluation & Failure Modes Style and identity overfit. The visible patterns that warn you. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 205: Continuous Dynamics and Numerical Sampling
Evaluation & Failure Modes Separate training error from integration error; vary the number of steps and report accuracy and cost. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 206: Flow Matching From Scratch
Evaluation & Failure Modes Inspect sample coverage, path choice and integration error; do not infer image-scale superiority from a toy experiment. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 207: Video Generation: Time, Latents, and Conditioning
Evaluation & Failure Modes Test identity drift, flicker, motion collapse and sensitivity to frame rate using explicit short-clip comparisons. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 208: Video Generation: Controlled Experiments and Evaluation
Evaluation & Failure Modes Log failures as well as successes, keep evaluation clips separate, and distinguish controllable motion from plausible-looking motion. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 209: Cameras, Projection, and Coordinate Frames
Evaluation & Failure Modes Catch handedness, matrix-order, degree/radian and world-to-camera mistakes with a known synthetic scene. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 210: Multi-View Geometry and 3D Representations Correspondences, triangulation, camera calibration, point clouds, meshes, voxels, signed-distance fields and visibility, with small geometric examples. Track every shape, dtype, and assumption through the derivation, not just the final equation Work one tiny numeric example by hand before touching code Name the assumptions that have to hold for this to be useful and where they break
Evaluation & Failure Modes Test bad camera poses, weak baselines and occlusion; low reprojection error alone does not certify all scene geometry. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 211: Neural Radiance Fields and Differentiable Rendering
Evaluation & Failure Modes Inspect unseen views and pose errors; distinguish scene reconstruction, novel-view rendering and generating a new scene. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 212: 3D Gaussian Splatting and Representation Tradeoffs
Evaluation & Failure Modes Separate a pedagogical toy renderer from a full optimized implementation; inspect floaters, overfitting and poorly observed surfaces. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 213: Text and Image to 3D: Priors, Consistency, and Assets
Evaluation & Failure Modes Measure multi-view inconsistencies, duplicate faces, geometry defects, editability and provenance; explicitly separate executed training from checkpoint inspection. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 214: World Models: Representation, Prediction, and Planning
Evaluation & Failure Modes Test unseen action sequences, compounding rollout errors and partial observability; a plausible generated video is not proof of accurate controllable dynamics. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 215: Research Studio: Alternative Explanations and Ablations
Evaluation & Failure Modes Report negative results and changed beliefs; separate novelty from a reproduction and avoid general claims from one selected run. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 216: Research Studio: Reproduction and Product Decisions
Evaluation & Failure Modes Deliver code, environment, data description, result table, limitations and next experiment; explain which claims remain untested. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 217: Image, Video, Audio Generation
Evaluation & Failure Modes Modality failures specific to video and audio. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 218: Generative AI Evaluation and Safety
Evaluation & Failure Modes Safety risk specific to generative models — the categories every launch must cover. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 219: RL Transfer: Bandits, Control, and Product Decisions
Evaluation & Failure Modes Evaluate reward misspecification, feedback bias and uncertainty; document when deploying an RL policy would be unjustified. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 220: Data and Model Versioning
Evaluation & Failure Modes Lineage gaps that bite three months after launch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 221: Testing ML and LLM Systems
Evaluation & Failure Modes Silent breakage. The categories that traditional tests will miss every time. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 222: Deployment Patterns for AI Systems
Evaluation & Failure Modes Deployment tradeoffs and the four serving patterns you should know cold. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 223: Monitoring Drift, Quality, and Cost
Evaluation & Failure Modes Operational risk. The dashboards that catch issues before customers do. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 224: Privacy, Licensing, Governance
Evaluation & Failure Modes Release readiness. The questions that block, or unblock, a launch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 225: ML System Design Interviews
Evaluation & Failure Modes Interview gaps. The four sections candidates always under-deliver on. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 226: Portfolio Strategy and Case Studies
Evaluation & Failure Modes Portfolio weakness. The patterns that recruiters skip past in five seconds. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 227: Final Capstone Build Sprint
Evaluation & Failure Modes Capstone execution. The signs you should ship vs polish for one more day. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 228: Final Launch and Career Readiness
Evaluation & Failure Modes Career readiness. The artefacts a hiring manager wants in their inbox. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Capstone: 2026 AI Product Launch Launch a portfolio-grade AI product that can include ML, DL, LLMs, RAG, agents, diffusion, or multimodal features, with evaluation, deployment, monitoring, governance, and a case-study writeup.
Hugging Face Diffusion Course concepts Full Stack Deep Learning production arc Public demo, eval report, model/system card, monitoring plan, and portfolio case studyAfter this phase, you'll be able to:
Understand major generative model families and diffusion internals Fine-tune and evaluate generative AI workflows Test, deploy, monitor, and govern AI systems Launch a portfolio-grade AI product and explain it in interviews You can present yourself as a modern applied AI engineer with depth across ML, DL, LLMs, RAG, agents, and production systems. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.