Subscribe

About AES

Evaluation is evidence generation

Agent Evaluation Science (AES) is an independent nonprofit initiative advancing the scientific evaluation of AI agents. We develop public methods and evidence for benchmark design and validation, and build evaluation systems that support valid, representative, and decision-relevant claims.

IWhy we exist

AES develops the scientific methods and evidence needed for agent evaluation, drawing on measurement, experimental design, reproducible infrastructure, and evidence synthesis. As agents move from benchmark tasks into real work, evaluation must connect a clear question to observations produced under stated conditions.

A valid evaluation begins with a defined design. The ability, decision, population, setting, and conditions of interest are defined first; tasks, samples, comparisons, and measures are then selected to support the intended inference, with attention to bias, uncertainty, and representativeness. Collecting all available data does not by itself produce valid evidence.

A leaderboard number becomes evidence only when its meaning, conditions, uncertainty, and intended decision are clear.
IIWhat makes agents different
A

Agent systems

Results depend on models, scaffolds, tools, prompts, and environments.

B

Runs are samples

Stochastic, path-dependent behavior requires repetition and uncertainty estimates.

C

Trajectories matter

Cost, actions, recoverable errors, and risk disappear when only success is reported.

D

Actions have consequences

Written data, sent messages, spent money, and changed state are part of the outcome.

E

Environments are instruments

Sandboxes, tools, and task definitions can change the conclusion.

IIIThe scientific foundations

Agent evaluation draws on established scientific fields, while agent systems also introduce conditions that require these methods to be adapted.

Agent-evaluation problemEstablished scientific foundation
Are we measuring the intended thing?Measurement theory / psychometrics
Does it represent the work and setting of interest?Epidemiology / sampling science
Are comparisons interpretable?Experimental design / statistics
What caused the difference?Causal inference
Does it reflect real human interaction and workflow?HCI / human factors
Are system failures captured beyond final score?Systems / reliability & safety engineering
What can we conclude across evaluations?Evidence synthesis / meta-science
IVWhat we are building
  1. AMethods & guidance

    Develop public methods, field guidance, and reporting practices for designing and validating agent benchmarks and evaluations.

  2. BBenchmark evidence

    Build, audit, replicate, compare, and study benchmarks to test what evaluation results can support.

  3. CEvaluation infrastructure

    Develop reusable research resources and the technical systems needed to turn validated designs into executable evaluations.

  4. DResearch community

    Connect evaluation researchers with domain experts through collaborations, working groups, and scientific events.

VWork with AES

Collaboration can take one of three forms.

A

Research collaboration

Study evaluation methods, benchmark validity, existing evidence, or new research questions with AES.

B

Benchmark development

Work with AES to turn domain workflows, cases, or proprietary data into a scientifically designed benchmark, including task definition, sampling, scoring, validation, and execution design.

C

Network & community partnership

Share research opportunities, connect researchers and domain experts, co-host events, and help relevant work reach the broader evaluation community.

AES publishes the scientific methodology behind its benchmark work. Proprietary data, private test sets, customer-specific implementation, and internal benchmark-production infrastructure remain private when required. Collaborations can start with a single research question, benchmark, dataset, workflow, resource, or event, with scope and data rights agreed for the project.