Fall 2026 · New York City

Call for Presentations

Short papers and extended abstracts on the science and practice of AI agent evaluation.

Submission options

Choose your submission

Short paper

Presentation-only
Length
4–6 pages, excluding references
Best for
Completed studies, substantial tools or datasets, and detailed empirical analyses

Extended abstract

Presentation-only
Length
2 pages, excluding references
Best for
Ongoing or published work, tools, demonstrations, datasets, failures, and industry cases

Presentation preference

  • Oral preferred
  • Poster preferred
  • Either

Final format is determined by the program committee.

Scope

What we are looking for

01

Diagnose

Audits, replications, negative results, contamination, grader instability, and evaluation gaps.

02

Measure

Constructs, metrics, rubrics, human evaluation, automated graders, reliability, validity, cost, and safety.

03

Operationalize

Benchmarks, environments, harnesses, traces, repeated runs, versioning, observability, and scalable grading.

04

Apply

Coding, research, science, healthcare, enterprise, web, computer-use, multimodal, and other tool-using agents.

Selection

How submissions are evaluated

  1. 01A clear evaluation target and appropriately bounded claims
  2. 02Relevant evidence supported by sound methods and reporting
  3. 03Findings relevant to evaluation research or practice
  4. 04Relevance to the symposium scope and scientific discussion

Industry and real-world submissions are assessed on documented evidence, production lessons, and findings relevant beyond a single system or setting. A new benchmark or algorithm is not required.

Author information

Submission requirements

Submission system
Submit a single PDF through OpenReview. OpenReview is the only submission system for this symposium.
Formatting
Use a readable PDF with a title, author names and affiliations, and references. Page limits exclude references. No conference template is required.
OpenReview selections
Select Short Paper or Extended Abstract and a presentation preference.
Attendance
At least one author of every accepted contribution must register and present.
Review
Submissions are reviewed through OpenReview. Reviews and decisions are managed in the OpenReview venue.
Author identity
Submissions are not anonymous. Include author names and affiliations in the manuscript and OpenReview record.
Prior publication
Ongoing work, previously published work, tools, datasets, negative results, failure analyses, and real-world cases are all eligible. Identify and cite any prior publication.

FAQ

Common questions

Can I submit published work?

Yes. Published work, ongoing work, tools, datasets, negative results, and failure analyses are eligible. Identify and cite any prior publication.

Can I present remotely?

The meeting is primarily in person in New York City. Limited remote presentation may be approved when necessary.

Who decides whether I give an oral or poster?

Authors indicate Oral preferred, Poster preferred, or Either. The program committee makes the final assignment.

Questions about the call?contact@evalscience.org