Fall 2026 · New York City

Agent Evaluation Science

A one-day symposium on the methods, measures, systems, and evidence used to evaluate AI agents.

Date
Friday, November 20, 2026
Location
Downtown Manhattan · New York City
Submissions
Short papers and 2-page extended abstracts via OpenReview

Attendance is free. Paper submission is not required to register.

Explore the symposium
Scientific evaluationConstruct validityMeasurement reliabilityBenchmark integrityAgent safetyGrader designFailure analysisEvaluation infrastructureReal-world evidenceReproducible systems

Why this symposium

Scientific methods and evidence for agent evaluation

AI agents act across tools, environments, and extended workflows. Evaluating them requires evidence beyond a single score or benchmark.

The symposium brings together researchers and practitioners working on measurement, benchmark design, evaluation infrastructure, reliability, safety, and real-world performance.

1focused day
2submission formats
4evaluation stages

Evaluation science agenda

Four stages organize the symposium

01

Diagnose

Identify what current evaluations measure, where they fail, and which gaps limit reliable conclusions.

02

Measure

Develop frameworks, constructs, metrics, graders, and evidence for reliability and validity.

03

Operationalize

Build runnable benchmarks, environments, harnesses, and reproducible evaluation infrastructure.

04

Apply

Evaluate agents in realistic settings using failure evidence and real-world performance.

Call for presentations

Submit work on agent evaluation

Submission details →

Submit evaluation studies, audits, negative results, failure analyses, tools, datasets, production lessons, and real-world evaluation cases. Ongoing and previously published work can also be presented through the extended abstract track.

Industry and real-world submissions are assessed on documented evidence, production lessons, and findings relevant beyond a single system or setting. A new benchmark or algorithm is not required.

01

Submissions open

August 252026
02

Abstract registration deadline

October 202026
03

Full submission deadline

October 252026
04

Notification of acceptance

November 52026
05

Symposium

November 202026

Affiliations represented

Speakers and committee members are affiliated with universities, research institutes, and industry organizations.

  • Agent Evaluation Science
  • Harvard University
  • Massachusetts General Hospital
  • BRIDGE GenAI Lab
  • Princeton University
  • Cornell University · Cornell Tech
  • The University of Texas at Austin
  • University of Technology Sydney
  • Australian Artificial Intelligence Institute
  • University of Alberta
  • Amii
  • Meta Superintelligence Labs
  • Sony AI
  • Raycaster
  • Google DeepMind
  • Massachusetts Institute of Technology
  • MIT Media Lab
  • Columbia University
  • DAPLab
  • New York University
  • Cohere Labs
  • J.P. Morgan
  • IBM Research
  • Microsoft Research
  • Mercor
  • University of Illinois Urbana-Champaign
  • Dartmouth College

Affiliations describe speakers’ and committee members’ institutional or organizational connections. They do not indicate sponsorship.

Agent Evaluation Science Fall 2026Friday, November 20 · Downtown Manhattan, New York City
Meet the committee