Subscribe

Resources

Evaluation Science Briefs

Concise, source-linked articles on how to design, interpret, and use evidence from AI agent evaluations.Explore AES research and tools
  1. 01

    How Close Is an Agent Evaluation to the Real Workflow?

    When does benchmark performance support a claim about performance in the workflow we actually care about?

    5 min read

  2. 02

    What Does an Agent Evaluation Actually Support?

    Observed runs, evaluation results, target-setting performance, and decisions require different evidence.

    6 min read

  3. 03

    How Do You Know Your Agent Grader Works?

    Tests, rubrics, parsers, and LLM judges need validation before their scores can support comparisons between agents.

    6 min read

  4. 04

    Is This Agent Result Reproducible?

    Repetition under one configuration cannot show whether a result survives changes in tasks, prompts, tools, environments, implementations, or time.

    6 min read

  5. 05

    Is This Agent Good Enough for the Decision?

    A score supports a decision only when uncertainty, thresholds, failure costs, alternatives, and the structure of errors are explicit.

    6 min read

  6. 06

    What Should an Agent Evaluation Report So Someone Else Can Use It?

    A reusable evaluation report identifies the system, task frame, run policy, grader, uncertainty, target claim, decision context, and limits of the evidence.

    6 min read

  7. 07

    Agent Evaluation Requires Methods Beyond Psychometrics

    Agent evaluation problems can break at different inferential steps, and each requires different evidence.

    7 min read

  8. 08

    Beyond Building the Most Difficult Benchmarks

    Harder benchmarks can measure frontier capability, while evidence for everyday agent use also requires measures of reliability, stability, cost, and failure.

    7 min read