1 / Layers
Separate observations, measures, and interpretations
| Layer | Example | Owner |
|---|---|---|
| Objective record | Observation, action attempt, outcome, event, point entry | WagerCall |
| Derived measure | Legal-action rate or recovery frequency | Evaluator's versioned analysis |
| Interpretation | A policy appeared more consistent under this setup | Reviewer with stated assumptions |
| Unsupported leap | The agent is generally superior | Not established by a bounded run set |
2 / Accounting
Reconcile attempts, exclusions, and denominators
- Number of runs attempted and completed per condition
- Rejected, interrupted, and excluded runs with reasons
- Whether each measure uses runs, rounds, decisions, or accepted actions
- Whether a repeated retry is counted as one logical decision or multiple attempts
- Whether missing evidence changes the conclusion
3 / Limits
Identify every difference outside the named variable
Compare run manifests before attributing a result to a prompt, model, or policy. Version, configuration, wrapper, tool contract, retry handling, and exclusion drift can all change the observed record.
A favorable stochastic outcome is not proof of a sound decision, and a different trajectory after an early branch is not a like-for-like sequence of later observations.
4 / Report
Write the narrowest conclusion the evidence supports
Name the agent conditions, exact WagerCall environment, run protocol, observed difference, uncertainty, and next decision in the conclusion itself.
- Link representative records or normalized evidence where appropriate.
- Publish the measure definition and analysis version.
- State unavailable data and alternative explanations.
- Do not infer hidden chain-of-thought or broad capability from game outcomes.