/ ML, DL & GenAI / Phase 5 Phase 5 Days 85-108
Deep Learning Architectures, Vision, and Multimodal Perception OpenCV, image processing and geometry, CNNs, attention before ViTs, detection, segmentation, self-supervised and multimodal perception
Phase Goal Go beyond image-classifier demos: understand modern perception architectures, multimodal embeddings, qualitative error analysis, safety, and deployment constraints.
Day 85: Images as Arrays and OpenCV Foundations
Evaluation & Failure Modes Diagnose wrong colors, clipped ranges, distorted aspect ratios, and train/serve preprocessing mismatch with assertions and a visual test grid. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 86: Classical Image Processing and a Non-Neural Baseline
Evaluation & Failure Modes Vary illumination, noise, touching objects and scale; document where the classical approach is sufficient and where its assumptions fail. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 87: Image Geometry, Features, and Motion
Evaluation & Failure Modes Measure alignment error and failure under occlusion, poor texture and non-planar scenes; a 2D warp is not a complete 3D reconstruction. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 88: Convolution From Scratch
Evaluation & Failure Modes Spatial shape bugs: the four ways your output tensor ends up the wrong size. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 89: CNN Training and Regularization
Evaluation & Failure Modes Vision baseline quality: when "more augmentation" stops helping and architecture changes start. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 90: ResNet and Modern CNN Design
Evaluation & Failure Modes Deep CNN stability: vanishing gradients without residuals, and how to spot them. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 91: Transfer Learning and Fine-Tuning
Evaluation & Failure Modes Transfer failure: when pretraining hurts and you are better off training from scratch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 92: Attention for Vision: From Weighted Averages to QKV
Evaluation & Failure Modes Distinguish attention weights from causal explanations; diagnose the wrong softmax axis and missing position information before tomorrow uses image patches. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 93: Vision Transformers
Evaluation & Failure Modes ViT tradeoffs: when you should not pick ViT. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 94: Object Detection Geometry
Evaluation & Failure Modes Annotation quality. Why the dataset usually limits you, not the architecture. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 95: YOLO, Faster R-CNN, Detection Workflows
Evaluation & Failure Modes Detection failures: small objects, occlusion, and label-noise effects. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 96: Semantic and Instance Segmentation
Evaluation & Failure Modes Mask quality. Why boundary-IoU matters more than overall IoU for thin objects. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 97: Self-Supervised Vision
Evaluation & Failure Modes Representation quality. How to evaluate embeddings without a final task. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 98: CLIP and Image-Text Embeddings
Evaluation & Failure Modes Prompt and bias failures. The categories CLIP silently misclassifies and why. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 99: Document AI and OCR
Evaluation & Failure Modes Field-level accuracy. The metric that matters in production, not page-level. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 100: Audio and Speech Foundations
Evaluation & Failure Modes Speech model limits: accents, noise, domain shift, and the bias they encode. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 101: Video Understanding
Evaluation & Failure Modes Temporal failures. Why static-image labelling does not transfer to video. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 102: Vision-Language Models
Evaluation & Failure Modes Visual hallucinations. The patterns to flag in any VLM-powered product. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 103: Multimodal Retrieval
Evaluation & Failure Modes Retrieval failures specific to multimodal: scale mismatch, modality bias, and the cross-modal "false friends". Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 104: Synthetic Data for Perception
Evaluation & Failure Modes Domain gap. Why a model trained on synthetic data often fails specifically on the easy cases. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 105: Perception Model Deployment
Evaluation & Failure Modes Deployment readiness: the four checks before exposing a vision endpoint. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 106: Safety, Privacy, Bias in Perception
Evaluation & Failure Modes Responsible release. The questions that should kill a model launch. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 107: CS231n-Style Vision Project Sprint
Evaluation & Failure Modes Project evidence: the artefacts a reviewer expects. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Day 108: Capstone: Multimodal Perception System
Evaluation & Failure Modes Perception capstone quality. The interview story that fits in three minutes. Inspect individual errors, slices, prompts, or traces — never trust a single average score Catalogue leakage, overfitting, bias, hallucination, latency, cost, or operational risks for this technique Write the next experiment from the failure mode you observed Close the notebook and reconstruct one mechanism; log help used and schedule a delayed changed-task check, initially around 1/3/7/14/30 days Keep ordinary review within 20 minutes; if weak prerequisites accumulate, pause new material and repair one dependency Record two competing explanations and the smallest experiment that distinguishes them; distinguish a negative result from an implementation failure
Capstone: Multimodal Perception System Build a computer vision or multimodal perception system with dataset auditing, model comparison, qualitative error review, deployment, and safety notes.
CS231n-style visual reasoning CLIP or VLM workflow Demo, model card, evaluation notebook, and READMEAfter this phase, you'll be able to:
Implement and fine-tune CNN and ViT workflows Evaluate detection, segmentation, and VLM systems Build multimodal retrieval and perception demos Document safety, bias, latency, and deployment limits You can build and critique vision or multimodal systems beyond screenshots and accuracy numbers. Gate: explain without notes, build one independent artifact, diagnose a deliberate failure, and repeat a changed task after a delay. Record help and repair missing prerequisites.