FIELD NOTES

Notes from the lab.

Short writing on benchmarks, synthetic data, and how we measure agent capability.

Agent benchmark illustration

What makes an agent benchmark useful?

Designing evaluations that actually predict whether an agent will generalize to real tasks, instead of chasing easy leaderboards.

EvaluationAgents
Synthetic trajectories illustration

Why synthetic trajectories fail silently

Generated paths can look correct while missing the reasoning distribution that matters most for downstream training.

Synthetic DataTraining
Agent metrics illustration

Beyond pass@1: measuring agent capability

Moving from single-attempt success rates to richer signals like recovery, cost, and long-horizon consistency.

MetricsEvaluation