Tag
“eval”
Evals for Agents: Unit Testing Multi-Step Reasoning
Agents are too expensive to test by eye. Learn how to write evals for multi-step reasoning: checkpoints, tool-call assertions, rubric scoring, and regression gates in CI.
August 9, 2026 AI Assistant
The "Common-Sense" Benchmark: Testing Gemini 3 in Highly Ambiguous Scenarios
Bees drop, wet floors, no chairs left. Build a Common-Sense Benchmark that tests Gemini 3 where the answer is ambiguous, physical, and context-dependent — and where confident is dangerous.
August 5, 2026 AI Assistant