On September 30, 2026, researchers from MIT, Harvard, Northeastern, Brown, UMass Amherst, and industry met at MIT for the first AES Boston Working Group. Participants work on agent benchmarks, data systems, RL environments, behavioral evaluation, cognitive science, multi-agent systems, and deployed agents.
We asked everyone three questions: what would make agent evaluation more scientific, what is the main bottleneck in their own work, and what they are doing about it. This note summarizes what we heard. We hope readers take away two points. First, building trustworthy evaluations at scale is now harder than building any single good benchmark, and the most common way to scale, having models write and grade tasks, has failure modes that are easy to miss. Second, many of the methods needed to address this already exist in other fields, and bringing those fields together is a large part of the work.
Task generation and review create linked bottlenecks
This came up more than anything else. Writing hard tasks takes a long time, and the time grows as models improve. Participants who use models to generate tasks described several recurring problems:
- Models have poor judgment about which problems are worth asking. Even with rich context, generated questions tend to be broken, too easy, or difficult only in trivial ways.
- A separate judge model is often not independent of the generator. In practice, the judge tends to grade the generator's stated reasoning, so some unsolved problems get recorded as solved.
- New fields have no experts to hire, and a human review step at the end of a large pipeline becomes the bottleneck. One participant argued that expert input is more useful earlier, when deciding what a good problem in the field looks like.
- RL environments are quick to build and slow to harden. One participant estimated that making an environment robust to reward hacking takes about twice as long as building it, and that frontier labs reportedly spend $2,000 to $10,000 per environment on quality checks. We are not aware of a public equivalent.
Running evaluations is also getting expensive. Long-horizon tasks can run for days, and multi-agent studies at the scale labs run are out of reach for most academic groups.
Hard tasks can obscure the capability being measured
As models improve, well-specified tasks saturate quickly. Making tasks harder often adds dependencies that make it unclear what is being measured. One example from the discussion: a benchmark meant to test whether an agent can use a new software library is not measuring that skill if a strong model can solve its tasks without using the library. Several participants suspected that when frontier models score in the low single digits on a benchmark, flawed questions are often part of the reason.
End-to-end work involves retrieval, data preparation, missing values, method choice, changing information, and dependencies between steps. Many different pipelines can reach the same correct answer, so scoring against a single reference solution works poorly. One group instead checks for the building blocks that any correct pipeline must contain. Researchers also lack access to the tasks and queries that deployed systems actually receive, so benchmark designers have to guess at that distribution.
Diagnosis requires attention to the whole system
Failures can come from the model, the harness, context, memory, tools, the environment, or the grader. In long trajectories, which can run to millions of tokens, finding the cause is hard, and naive LLM judges often blame the wrong component. Models asked to diagnose a failure also tend to stick with their first explanation.
That last observation connected to a different project in the room. A linguist studying multi-turn sycophancy, where models abandon correct answers after a user pushes back, was looking at the other side of the same behavior: how much weight a model gives to an earlier turn when new information arrives. Connections like this were one of the more useful results of putting people from different fields in one room.
The discussion drew on methods from epidemiology, psychology, cognitive science, linguistics, social theory, and manufacturing quality control.
Questions for the group
A few questions came up that nobody in the room could answer:
- Can self-improving or auto-research systems get past human-level benchmarks, or do loops without external grounding drift and collapse first?
- Can research taste be defined well enough to measure? Citation counts are a weak proxy.
- For controlled studies of agent behavior, is prompting rigorous enough, or do we need to intervene on model internals?
- Does an evaluation still measure the same capability as agents get stronger and discover new shortcuts?
- Can we verify that task difficulty comes from the target capability rather than artifacts or unrelated complexity?
- Where should scarce domain expertise enter generation, filtering, reference construction, and auditing?
- When generators, judges, source evidence, or diagnostic tools share models or assumptions, how correlated are their errors?
- What population of real work is a benchmark meant to represent, and how should tasks be sampled or weighted?
- For meta-agents, research agents, and multi-agent systems, what should be measured beyond final task success?
Continuing the work
We plan to continue the Boston Working Group as a recurring forum, with participants proposing topics and presenting work in progress. We will feature participants' work on evalscience.org and use the sessions to identify collaborations and shared research questions. Several participants also want to develop longer-term shared outputs, including a joint paper across multiple sessions and a public resource organizing evaluation methods, benchmarks, and known pitfalls.
AES is running similar sessions in other research hubs, as well as the Agent Evaluation Science Symposium in New York on November 20.
We are also starting to plan an AES New England Agent Evaluation Science Summit. A draft proposal is ready, and we are looking for people who want to help co-organize the event and shape the program, partnerships, and logistics.
If you work on any of these problems, would like to join AES or the Boston Working Group, or want to help organize the New England summit, we would like to hear from you through evalscience.org. We would especially like to hear from practitioners who can share realistic workloads or failure data, and from institutions that can host sessions or provide compute.