Method · Field guide · In development
AES Benchmark Methodology
Intended
use→Task
population→Benchmark
design→Validation
A public field guide for designing and validating agent benchmarks, covering the target work and task population, sampling, environments, references and graders, repeated execution, uncertainty, validity evidence, versioning, and reporting.
The methodology describes the evidence a defensible benchmark should provide; proprietary datasets, hidden test sets, customer-specific implementation, and internal production infrastructure are outside the public field guide.
Evidence study · In development
EvidenceChecker
Evaluation→Rationale→Decision
documented linkage?
A study of how evaluation evidence is documented, interpreted, and connected to engineering decisions in AI-agent development, including whether those decisions can be traced to the evidence used to support them.
Benchmark · In development
Evidence Generation Agents
Paper-world→Agent run→Work
products→Artifact
grading
An executable benchmark for evaluating agents on evidence-generating scientific workflows using paper-grounded environments, isolated execution, and artifact-based grading.
The current implementation begins with epidemiologic research workflows and will expand.