Resources
Evaluation Science Briefs
Concise, source-linked articles on how to design, interpret, and use evidence from AI agent evaluations.Explore AES research and tools →Featured / Start here
What’s Wrong With AI Eval “Science”?
AI evaluation produces many scores, but the scientific object behind the number is often underspecified.
Read brief- 01
How Close Is an Agent Evaluation to the Real Workflow?
When does benchmark performance support a claim about performance in the workflow we actually care about?
- 02
What Does an Agent Evaluation Actually Support?
Observed runs, evaluation results, target-setting performance, and decisions require different evidence.
- 03
How Do You Know Your Agent Grader Works?
Tests, rubrics, parsers, and LLM judges need validation before their scores can support comparisons between agents.
- 04
Is This Agent Result Reproducible?
Repetition under one configuration cannot show whether a result survives changes in tasks, prompts, tools, environments, implementations, or time.
- 05
Is This Agent Good Enough for the Decision?
A score supports a decision only when uncertainty, thresholds, failure costs, alternatives, and the structure of errors are explicit.
- 06
What Should an Agent Evaluation Report So Someone Else Can Use It?
A reusable evaluation report identifies the system, task frame, run policy, grader, uncertainty, target claim, decision context, and limits of the evidence.
- 07
Agent Evaluation Requires Methods Beyond Psychometrics
Agent evaluation problems can break at different inferential steps, and each requires different evidence.
- 08
Beyond Building the Most Difficult Benchmarks
Harder benchmarks can measure frontier capability, while evidence for everyday agent use also requires measures of reliability, stability, cost, and failure.