Subscribe

Resources

Research

AES benchmark methodology, papers, benchmark studies, audits, evidence reviews, and work in development.
Evaluation Science BriefsRead source-linked articles on methods and evidence for agent evaluation

Public research

Same activity: analysis + record-keeping

Answer-scored task

  1. Fixed corpus
  2. Final numerical answer
  3. Checked

Reviewable work product

  1. Sources
  2. Extracted values
  3. Calculation
  4. Assumptions
  5. Checked as a package

What is evaluated determines what the score can support.

Method · Benchmark design · Preprint

Designing Benchmarks for Knowledge Work

Preprint

Specifies what knowledge-work agent benchmarks represent: the work activity, tested setting, required work product, and evaluated result. The accompanying reference derives an 18-activity inventory from O*NET task statements and applies the framework to existing agent benchmarks.

Without state contract

Parsed
v1
Native
v2
Submission relationship unclear

StagedWorkspace

Parsed
Cₜ
Workspace
Wₜ
Native
Wₜ
Review
Δₜ
Submit
Wₜ
edit → hash change → stalerefreshed / current
Experimental study · Evaluation infrastructure · Preprint

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

Preprint

A study of workspace state as an experimental variable in knowledge-work agent evaluation. StagedWorkspace keeps parsed views, native files, review diffs, and submitted artifacts tied to explicit workspace versions, and tests how access and review conditions change measured agent performance.

8.3–12.1 ppOfficeQA improvement with dual parsed/native access vs. single view
4.7–9.2 ptsAPEX rubric-score improvement vs. single view
57 tasksPaired review-axis experiment on the effect of visible diffs

Research in development

Method · Field guide · In development

AES Benchmark Methodology

Intended
use
Task
population
Benchmark
design
Validation

A public field guide for designing and validating agent benchmarks, covering the target work and task population, sampling, environments, references and graders, repeated execution, uncertainty, validity evidence, versioning, and reporting.

The methodology describes the evidence a defensible benchmark should provide; proprietary datasets, hidden test sets, customer-specific implementation, and internal production infrastructure are outside the public field guide.

Evidence study · In development

EvidenceChecker

EvaluationRationaleDecision

documented linkage?

A study of how evaluation evidence is documented, interpreted, and connected to engineering decisions in AI-agent development, including whether those decisions can be traced to the evidence used to support them.

Benchmark · In development

Evidence Generation Agents

Paper-worldAgent runWork
products
Artifact
grading

An executable benchmark for evaluating agents on evidence-generating scientific workflows using paper-grounded environments, isolated execution, and artifact-based grading.

The current implementation begins with epidemiologic research workflows and will expand.