Diagnose
Find what current evaluations measure, where they fail, and which gaps limit reliable conclusions.
Fall 2026 · New York City
Symposium on the science of AI agent evaluation.
Evaluation science
Find what current evaluations measure, where they fail, and which gaps limit reliable conclusions.
Develop frameworks, constructs, metrics, graders, and validity evidence.
Create benchmarks, environments, harnesses, and reproducible evaluation infrastructure.
Evaluate agents in realistic and consequential settings using real-world evidence.
About the symposium
AI Agent Eval Sci Fall 2026 invites research and practice contributions on the science of evaluating AI agents. We welcome work that improves how agent capabilities, failures, reliability, safety, and real-world performance are measured, as well as evidence from teams building and deploying agent systems.
The meeting connects evaluation methodology with deployment experience across research, industry, and high-stakes practice.
Program
The detailed schedule is being finalized. The symposium will include the following components.
Approximately 10 contributed oral presentations
Approximately 20 poster presentations
Invited academic and industry talks
Real-world agent evaluation case studies
Poster and reception session
Structured discussion connecting evaluation methodology with deployment experience
Presentation plans include 12-minute contributed talks with 3 minutes for discussion. Program details remain subject to change.
Speakers
Invited speakers and participating organizations will be added as the program is confirmed.
Call for contributed presentations
Authors may submit through either an archival research paper track or a non-archival presentation track. Archival status and presentation format are independent.
Track 01
For original research authors wish to publish in the symposium proceedings.
A proceedings volume in the EPiC Series in Computing is being pursued. Publication is planned and remains subject to formal EasyChair/EPiC approval.
Track 02
For work authors wish to present and discuss without publication in archival proceedings.
This track supports later archival publication eligibility and applied teams presenting real-world evaluation evidence.
Oral presentations
Selected from both tracks. Priority is given to important evaluation questions, strong evidence, lessons beyond one model or application, important failures or real-world challenges, and distinctive perspectives.
Industry and real-world case studies are fully eligible. A new benchmark or algorithm is not required.
Poster presentations
Posters may present complete studies, work in progress, tools, datasets, infrastructure, audits, replications, negative findings, real-world failures, published work, or early ideas.
Selection is based on relevance and discussion value, not lower scientific importance.
Benchmark and leaderboard audits, systematic reviews, replication studies, negative results, contamination, grader instability, harness effects, coverage gaps, reproducibility failures, and mismatches between scores and real-world behavior.
Evaluation frameworks, construct and capability definitions, metrics, rubrics, human evaluation, automated graders, executable verifiers, trajectory-level measures, work-product assessment, uncertainty, reliability, validity, efficiency, cost, and safety measures.
Benchmarks, realistic environments, evaluation harnesses, tool and state management, trace collection, repeated-run evaluation, versioning, observability, benchmark integrity, scalable grading, reproducibility infrastructure, and action or trajectory audits.
Coding, research, scientific, healthcare, enterprise, workplace, web, computer-use, multimodal, and other tool-using or autonomous agents. Concrete failure cases, production traces, deployment evaluations, and evidence linking formal evaluation to practical failure are encouraged.
Industry submissions are assessed on concrete failure evidence, production lessons, evaluation methodology, and generalizable insight. Academic novelty alone is not required.
Submission portal
The submission link will be posted when the symposium installation is open.
Important dates
All deadlines are 11:59 PM Anywhere on Earth (AoE).
Venue
The symposium is designed primarily as an in-person meeting. Venue and travel details will be announced after arrangements are finalized.
Limited remote presentation may be approved when necessary.