Subscribe
Evaluation Science Briefs

Evaluation Science Brief

Is This Agent Good Enough for the Decision?

By Yining Hua, MSc.6 min read
Evidence path
Evaluation evidence
Decision rule
Action

Suppose an evaluation reports:

Agent A: 74% success.

A procurement requirement says:

Minimum acceptable performance: 70%.

It is tempting to conclude that the agent passed. That conclusion requires more information than the two numbers provide. What is the uncertainty around 74%? What kinds of failures make up the remaining 26%? Is 70% a threshold derived from consequences or simply a convenient round number? What is the status quo? Does the agent reduce cost or review burden? What happens if a system that is actually below the requirement is accepted? These are decision questions.

Separate the evidence from the action

GRADE rates certainty of evidence separately from the strength of a recommendation, and its Evidence-to-Decision framework adds considerations such as benefits and harms, resources, feasibility, and acceptability [1,2]. For agents, an evaluation can provide strong evidence that a system achieves a specified level of performance under specified conditions, while deployment still depends on the decision context. Evaluation reporting should keep the result, broader claim, and deployment recommendation separate.

Uncertainty changes whether a threshold is cleared

Conformity assessment deals directly with measurements near a specification limit. JCGM 106 describes decision rules for taking measurement uncertainty into account when declaring conformity [3]. One common approach is guard banding: the acceptance boundary is set so that the risk of falsely accepting a nonconforming item is controlled. For an agent evaluation, suppose the nominal threshold is 70% and the estimate is 74%, but the uncertainty around the estimate is large enough to include values below 70%. The point estimate alone does not justify a clear “pass.” Whether the system passes depends on the pre-specified decision rule and the tolerated probability of false acceptance. The exact guard band should not be invented after seeing the result. The rule belongs in the evaluation plan.

Decision threshold diagram comparing estimates clearly below, overlapping, and clearly above a threshold, with uncertainty intervals and point estimates.

The threshold should reflect consequences

“Is 74% good enough?” has no general answer. The relevant quantities depend on the workflow. An undetected research error, a failed code change, and an unnecessary request for human review have different consequences. A system used 10 times per month and a system used 100,000 times per day can tolerate different failure structures even at the same average success rate. Decision-curve analysis in medicine formalizes a related idea: a model is evaluated in relation to the consequences attached to different decision thresholds, instead of treating predictive accuracy as the final objective [4]. Although the exact clinical formula does not transfer mechanically to agent deployment, the threshold should similarly be tied to the relative cost of wrong actions and missed opportunities. Relevant inputs for an agent workflow can include:

False acceptance cost: what happens if an inadequate agent is accepted or an incorrect output is trusted?
False rejection cost: what is lost if a useful agent or useful output is rejected?
Review cost: how much human checking is required, and at what volume?
Failure detectability: are the important failures obvious or silent?
Alternatives: what do the current process and competing systems achieve?

Keep reliability and failure structure in the decision

Two systems with the same mean can support different decisions. One may fail consistently on a known subset of tasks that can be routed elsewhere. Another may fail unpredictably across the full task distribution. A third may have rare but severe failures that dominate the expected harm. A decision threshold therefore should not depend on the mean alone. The evidence package should preserve relevant reliability distributions, severity, and detectability of failure modes. This is also where the target workflow matters. The same agent may be acceptable for a low-consequence drafting task and unacceptable for an action that changes production data.

Decide whether more evaluation is worth doing

When a result is close to the decision boundary, value-of-information methods ask whether reducing uncertainty is worth the cost of collecting more evidence. A full formal analysis is not necessary for every agent evaluation. If plausible new evidence could change the decision, the remaining evaluation can target the decision-relevant uncertainty; otherwise, more benchmarking may have little value.

Decision Threshold Card

FieldRecord
DecisionWhat action is being considered?
Target workflowWhere and how will the agent be used?
Candidate and comparatorAgent, status quo, and relevant alternatives
Measured outcomeWhat result is directly observed?
Estimate and uncertaintyPoint estimate plus relevant uncertainty
Decision thresholdWhat counts as acceptable, and why?
False acceptance costConsequence of accepting an inadequate system/output
False rejection costConsequence of rejecting an adequate system/output
Reliability / failure structureRepeatability, tail failures, severity, detectability
Cost and operational constraintsReview burden, latency, resources, volume
Decision ruleHow uncertainty is handled at the threshold
Additional evidenceWhat could plausibly change the decision?

The card records the threshold and its assumptions for review.

References

  1. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924–926.
  2. Alonso-Coello P, Schünemann HJ, Moberg J, et al. GRADE Evidence to Decision (EtD) frameworks: a systematic and transparent approach to making well informed healthcare choices. BMJ. 2016;353:i2016.
  3. JCGM. Evaluation of measurement data — The role of measurement uncertainty in conformity assessment (JCGM 106:2012). 2012. doi:10.59161/JCGM106-2012.
  4. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Medical Decision Making. 2006;26(6):565–574. doi:10.1177/0272989X06295361.
  5. Kane MT. Validating the Interpretations and Uses of Test Scores. Journal of Educational Measurement. 2013;50(1):1–73.
  6. Kapoor S, Stroebl B, Siegel ZS, Nadgir N, Narayanan A. AI Agents That Matter. Transactions on Machine Learning Research. 2025.