Agent systems
Results depend on models, scaffolds, tools, prompts, and environments.
About AES
Agent Evaluation Science (AES) is an independent nonprofit initiative advancing the scientific evaluation of AI agents. We develop public methods and evidence for benchmark design and validation, and build evaluation systems that support valid, representative, and decision-relevant claims.
AES develops the scientific methods and evidence needed for agent evaluation, drawing on measurement, experimental design, reproducible infrastructure, and evidence synthesis. As agents move from benchmark tasks into real work, evaluation must connect a clear question to observations produced under stated conditions.
A valid evaluation begins with a defined design. The ability, decision, population, setting, and conditions of interest are defined first; tasks, samples, comparisons, and measures are then selected to support the intended inference, with attention to bias, uncertainty, and representativeness. Collecting all available data does not by itself produce valid evidence.
A leaderboard number becomes evidence only when its meaning, conditions, uncertainty, and intended decision are clear.
Results depend on models, scaffolds, tools, prompts, and environments.
Stochastic, path-dependent behavior requires repetition and uncertainty estimates.
Cost, actions, recoverable errors, and risk disappear when only success is reported.
Written data, sent messages, spent money, and changed state are part of the outcome.
Sandboxes, tools, and task definitions can change the conclusion.
Agent evaluation draws on established scientific fields, while agent systems also introduce conditions that require these methods to be adapted.
| Agent-evaluation problem | Established scientific foundation |
|---|---|
| Are we measuring the intended thing? | Measurement theory / psychometrics |
| Does it represent the work and setting of interest? | Epidemiology / sampling science |
| Are comparisons interpretable? | Experimental design / statistics |
| What caused the difference? | Causal inference |
| Does it reflect real human interaction and workflow? | HCI / human factors |
| Are system failures captured beyond final score? | Systems / reliability & safety engineering |
| What can we conclude across evaluations? | Evidence synthesis / meta-science |
Develop public methods, field guidance, and reporting practices for designing and validating agent benchmarks and evaluations.
Build, audit, replicate, compare, and study benchmarks to test what evaluation results can support.
Develop reusable research resources and the technical systems needed to turn validated designs into executable evaluations.
Connect evaluation researchers with domain experts through collaborations, working groups, and scientific events.
Collaboration can take one of three forms.
Study evaluation methods, benchmark validity, existing evidence, or new research questions with AES.
Work with AES to turn domain workflows, cases, or proprietary data into a scientifically designed benchmark, including task definition, sampling, scoring, validation, and execution design.
Share research opportunities, connect researchers and domain experts, co-host events, and help relevant work reach the broader evaluation community.
AES publishes the scientific methodology behind its benchmark work. Proprietary data, private test sets, customer-specific implementation, and internal benchmark-production infrastructure remain private when required. Collaborations can start with a single research question, benchmark, dataset, workflow, resource, or event, with scope and data rights agreed for the project.