Video: Evals for AI Agents: How Product Builders Get the Most Out of Every New Model | Duration: 3632s | Summary: Evals for AI Agents: How Product Builders Get the Most Out of Every New Model | Chapters: Introduction and Welcome (2.72s), Evaluation Challenges (163.22s), Agent Evaluation Evolution (355.93s), Evaluation Harness (582.705s), Building Eval Sets (864.465s), Eval Loop Architecture (1019.245s), Test Suite Design (1232.93s), Test Types & Design (1425.705s), Reducing Eval Costs (3178.563s), Adversarial Testing (3276.97s), Model Migration Strategy (3370.975s), LLM Judge Optimization (3494.055s), Closing Remarks (3606.28s)