Evals
Sep 18, 2026
8 min
How we evaluate agents before they ship
Benchmarks you write yourself beat leaderboards you read about.

Leaderboards tell you how a model does on someone else’s problem. Your own benchmarks tell you how it does on yours.
Write the test first
Collect fifty real tasks, define a pass, and run every model and prompt change against them before it ships.
Keep reading
More notes.

Strategy
Sep 30, 2026
Your AI strategy is a data strategy
Models change every quarter. Clean, governed data keeps compounding.

Evals
Sep 18, 2026
How we evaluate agents before they ship
Benchmarks you write yourself beat leaderboards you read about.

Governance
Sep 04, 2026
Governance without the slowdown
Policy as code lets teams move fast and stay inside the lines.