Evaluation Science Brief
Is This Agent Good Enough for the Decision?
Suppose an evaluation reports:
Agent A: 74% success.
A procurement requirement says:
Minimum acceptable performance: 70%.
It is tempting to conclude that the agent passed. That conclusion requires more information than the two numbers provide. What is the uncertainty around 74%? What kinds of failures make up the remaining 26%? Is 70% a threshold derived from consequences or simply a convenient round number? What is the status quo? Does the agent reduce cost or review burden? What happens if a system that is actually below the requirement is accepted? These are decision questions.
Separate the evidence from the action
GRADE rates certainty of evidence separately from the strength of a recommendation, and its Evidence-to-Decision framework adds considerations such as benefits and harms, resources, feasibility, and acceptability [1,2]. For agents, an evaluation can provide strong evidence that a system achieves a specified level of performance under specified conditions, while deployment still depends on the decision context. Evaluation reporting should keep the result, broader claim, and deployment recommendation separate.
Uncertainty changes whether a threshold is cleared
Conformity assessment deals directly with measurements near a specification limit. JCGM 106 describes decision rules for taking measurement uncertainty into account when declaring conformity [3]. One common approach is guard banding: the acceptance boundary is set so that the risk of falsely accepting a nonconforming item is controlled. For an agent evaluation, suppose the nominal threshold is 70% and the estimate is 74%, but the uncertainty around the estimate is large enough to include values below 70%. The point estimate alone does not justify a clear “pass.” Whether the system passes depends on the pre-specified decision rule and the tolerated probability of false acceptance. The exact guard band should not be invented after seeing the result. The rule belongs in the evaluation plan.

The threshold should reflect consequences
“Is 74% good enough?” has no general answer. The relevant quantities depend on the workflow. An undetected research error, a failed code change, and an unnecessary request for human review have different consequences. A system used 10 times per month and a system used 100,000 times per day can tolerate different failure structures even at the same average success rate. Decision-curve analysis in medicine formalizes a related idea: a model is evaluated in relation to the consequences attached to different decision thresholds, instead of treating predictive accuracy as the final objective [4]. Although the exact clinical formula does not transfer mechanically to agent deployment, the threshold should similarly be tied to the relative cost of wrong actions and missed opportunities. Relevant inputs for an agent workflow can include:
False acceptance cost: what happens if an inadequate agent is accepted or an incorrect output is trusted?
False rejection cost: what is lost if a useful agent or useful output is rejected?
Review cost: how much human checking is required, and at what volume?
Failure detectability: are the important failures obvious or silent?
Alternatives: what do the current process and competing systems achieve?
Keep reliability and failure structure in the decision
Two systems with the same mean can support different decisions. One may fail consistently on a known subset of tasks that can be routed elsewhere. Another may fail unpredictably across the full task distribution. A third may have rare but severe failures that dominate the expected harm. A decision threshold therefore should not depend on the mean alone. The evidence package should preserve relevant reliability distributions, severity, and detectability of failure modes. This is also where the target workflow matters. The same agent may be acceptable for a low-consequence drafting task and unacceptable for an action that changes production data.
Decide whether more evaluation is worth doing
When a result is close to the decision boundary, value-of-information methods ask whether reducing uncertainty is worth the cost of collecting more evidence. A full formal analysis is not necessary for every agent evaluation. If plausible new evidence could change the decision, the remaining evaluation can target the decision-relevant uncertainty; otherwise, more benchmarking may have little value.
Decision Threshold Card
| Field | Record |
|---|---|
| Decision | What action is being considered? |
| Target workflow | Where and how will the agent be used? |
| Candidate and comparator | Agent, status quo, and relevant alternatives |
| Measured outcome | What result is directly observed? |
| Estimate and uncertainty | Point estimate plus relevant uncertainty |
| Decision threshold | What counts as acceptable, and why? |
| False acceptance cost | Consequence of accepting an inadequate system/output |
| False rejection cost | Consequence of rejecting an adequate system/output |
| Reliability / failure structure | Repeatability, tail failures, severity, detectability |
| Cost and operational constraints | Review burden, latency, resources, volume |
| Decision rule | How uncertainty is handled at the threshold |
| Additional evidence | What could plausibly change the decision? |
The card records the threshold and its assumptions for review.
References
- Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924–926.
- Alonso-Coello P, Schünemann HJ, Moberg J, et al. GRADE Evidence to Decision (EtD) frameworks: a systematic and transparent approach to making well informed healthcare choices. BMJ. 2016;353:i2016.
- JCGM. Evaluation of measurement data — The role of measurement uncertainty in conformity assessment (JCGM 106:2012). 2012. doi:10.59161/JCGM106-2012.
- Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Medical Decision Making. 2006;26(6):565–574. doi:10.1177/0272989X06295361.
- Kane MT. Validating the Interpretations and Uses of Test Scores. Journal of Educational Measurement. 2013;50(1):1–73.
- Kapoor S, Stroebl B, Siegel ZS, Nadgir N, Narayanan A. AI Agents That Matter. Transactions on Machine Learning Research. 2025.