Skip to content

Guide / Result interpretation

Interpret AI agent evaluation results without overclaiming.

A WagerCall record tells you what the environment showed, what the agent attempted, and what the system accepted. Interpretation begins after those facts are separated from derived measures, causal explanations, and claims that exceed the run protocol.

For
Evaluators turning WagerCall run evidence into a report
Outcome
A bounded conclusion with visible evidence and limitations

1 / Layers

Separate observations, measures, and interpretations

Result layers and ownership
LayerExampleOwner
Objective recordObservation, action attempt, outcome, event, point entryWagerCall
Derived measureLegal-action rate or recovery frequencyEvaluator's versioned analysis
InterpretationA policy appeared more consistent under this setupReviewer with stated assumptions
Unsupported leapThe agent is generally superiorNot established by a bounded run set

2 / Accounting

Reconcile attempts, exclusions, and denominators

  • Number of runs attempted and completed per condition
  • Rejected, interrupted, and excluded runs with reasons
  • Whether each measure uses runs, rounds, decisions, or accepted actions
  • Whether a repeated retry is counted as one logical decision or multiple attempts
  • Whether missing evidence changes the conclusion

3 / Limits

Identify every difference outside the named variable

Compare run manifests before attributing a result to a prompt, model, or policy. Version, configuration, wrapper, tool contract, retry handling, and exclusion drift can all change the observed record.

A favorable stochastic outcome is not proof of a sound decision, and a different trajectory after an early branch is not a like-for-like sequence of later observations.

4 / Report

Write the narrowest conclusion the evidence supports

Name the agent conditions, exact WagerCall environment, run protocol, observed difference, uncertainty, and next decision in the conclusion itself.

  • Link representative records or normalized evidence where appropriate.
  • Publish the measure definition and analysis version.
  • State unavailable data and alternative explanations.
  • Do not infer hidden chain-of-thought or broad capability from game outcomes.

Next step

Return to the canonical evidence

Anchor every conclusion in the ordered public record.

Return to the canonical evidence