Platform
Evaluation
Conventional classification models lend themselves readily to measurement. Agent systems are considerably harder to assess, because a given task may have several acceptable answers, a substantial proportion of cases fall into an ambiguous middle, and the route taken to an answer matters as much as the answer itself.
What we measure
| Measure | Definition | Why it is on the list |
|---|---|---|
| Outcome quality | Rubric score against an adjudicated set | The headline claim. Meaningless without the agreement rate beside it. |
| Inter-rater agreement | Agreement between human adjudicators | Bounds every other claim. If reviewers agree 0.8 of the time, a two-point difference is noise. |
| Tool-call precision | Calls made that a competent reviewer would have made | Catches the agent that queries five systems when one would do. |
| Tool-call recall | Calls a competent reviewer would have made that were made | Catches the agent that answers confidently without checking anything. |
| Escalation calibration | False escalations and, harder, missed escalations | The second is where real risk sits and almost nobody instruments it. |
| Cost per resolved task | Total cost divided by cases actually closed | Per-request cost flatters systems with a high retry or handoff rate. |
| Cost variance (p95/p50) | Spread of cost across cases | Determines whether the system can be budgeted at all. |
| Replay fidelity | Traces that reproduce the same terminal state | What a supervisory review actually depends on. |
Building the dataset
A few hundred carefully adjudicated cases will generally outperform several thousand loosely labelled ones. Our datasets combine three groups in deliberate proportion: routine cases that establish a baseline, known-difficult cases gathered from the staff currently performing the task, and adversarial cases constructed to test specific assumptions. The third group is the smallest and accounts for most of the diagnostic value.
Each case is accompanied by documented reasoning for its inclusion. A test case whose purpose is no longer understood is usually removed the first time it fails, taking the requirement it encoded with it.
Using a model as judge
Model-based judging is effective within defined limits, and those limits should be treated as mandatory. The judge must be calibrated against human review on a held-out sample before it is relied upon, the agreement rate should be reported alongside every result it produces, the same model should never act as both judge and subject, and a permanently human-reviewed subset should be retained, since drift in the judge is not otherwise visible.
What an evaluation report contains
- 01 · The claim being tested, expressed as a proposition capable of being disproved.
- 02 · Dataset provenance, covering the source of cases, who adjudicated them and what was excluded.
- 03 · Inter-rater agreement, reported before any headline result.
- 04 · Results expressed with uncertainty ranges rather than point estimates.
- 05 · Judge calibration figures, where a model judge was used.
- 06 · Limitations and matters the evaluation does not address, written by the person who conducted it.
- 07 · The conditions or changes that would invalidate the result.
Review of an existing evaluation suite
A frequent first engagement is a review of the evaluation suite you already have, resulting in a written assessment of the conclusions it can and cannot legitimately support. The findings are not always comfortable, but they are considerably less expensive to act on before go-live than afterwards.