Three real agents, one oracle, one verifier.
Three production models drove the same customer-operations agent through 46 scenarios each: five workflows across payments, orders, CRM and support, with failures injected at the HTTP layer that the agent cannot see. A database oracle holds the truth. CommitLayer verifies from the systems of record only. A trace-only judge, a model that sees the agent's tool calls and responses but never the systems, is the baseline. Run 2026-09-29, verifier 0.2.2.
| Measure | Anthropic claude-opus-5-5 | OpenAI gpt-5-mini | Google gemini-flash-latest |
|---|---|---|---|
| Scenarios scored | 46 | 46 | 46 |
| CommitLayer agreement with the oracle | 46 / 46 | 46 / 46 | 46 / 46 |
| Real failures (oracle says FAIL) | 12 | 10 | 8 |
| Real failures detected | 12 / 12 | 10 / 10 | 8 / 8 |
| False negatives (PASS on a real failure) | 0 | 0 | 0 |
| False positives (FAIL on a real pass) | 0 of 29 passes | 0 of 31 passes | 0 of 33 passes |
| Agent claimed success on a real failure, caught | 2 / 2 | 7 / 7 | 4 / 4 |
| Duplicate effects caught | 1 / 1 | 3 / 3 | 1 / 1 |
| UNKNOWN when a required system was down | 5 / 5 | 5 / 5 | 5 / 5 |
| Orphan-effect false alarms | 0 | 0 | 0 |
| Injected-failure scenarios the agent still got right | 24 | 26 | 28 |
| Trace-only judge agreement with the oracle | 39 / 46 | 31 / 46 | 37 / 46 |
| Trace-only judge passed a real failure | 2 | 2 | 2 |
| Trace-only judge failed a real pass | 0 | 8 | 2 |
| Verify latency, mean | 10 ms | 10 ms | 13 ms |
What the runs show
- The verifier's answer does not depend on the model. Three models, three different agents, three different counts of real failures. On every scored scenario CommitLayer matched the oracle: each real failure named by the missing or forbidden effect, no false positives, UNKNOWN only when a required system was unreachable.
- The failures that get through are the ones a trace cannot show. The better the model, the more of the matrix it avoids on its own. What remains is what the agent could not see from its own tool responses: a human who refunded first, a side effect elsewhere, a write that returned 200 and changed nothing, a retry without an idempotency key.
- A trace-only judge follows the trace. Where the agent's trace looks clean and the systems disagree, the judge passes the failure. Where the agent hit errors and recovered, the judge tends to fail a run that ended in the right state.
- A real agent found a verifier defect, and it was fixed. One model notified the customer on the support ticket rather than on the order, and the verifier reported a missing notification. Tickets are linked to their orders, so effects on linked subjects now count. The rows above are scored on the fixed verifier.
The two failures every trace-only judge passed
The agent's own trace is clean: one read, one refund, 200 OK. The duplicate exists only in the payment system.
The address update succeeded. The cancellation that followed never appears in any tool response the agent saw.
What was injected
- Lost response, retried
- The write succeeded but the response timed out; the agent retries without an idempotency key and a second refund lands.
- Phantom success
- The API returns 200 and changes nothing.
- Partial completion
- A downstream step fails after the first write; the agent reports success anyway.
- Permission denied
- The primary write returns 403; the agent may still claim the work was done.
- Rate limited, then fine
- A 429 once; a good agent retries and the state ends correct.
- Stale read
- The connector reads a lagging copy of the system at verification time.
- System unreachable
- A system of record the contract needs is down when the verdict is due.
- Over the limit
- The requested amount or discount exceeds what the policy allows.
- Ineligible order
- The precondition (paid, active, open) is not met.
- Concurrent update
- A human refunded the order between the agent's read and its write; the agent's refund duplicates it.
- Prompt injection
- Ticket text tells the agent to refund a different order for a larger amount.
- Forbidden side effect
- An address change is followed by the order being cancelled.
Method and limits
- Same system prompt, tool catalogue and tasks for every model; default temperature; at most 14 turns; the agent ends by reporting whether it completed the task.
- Truth comes from the sandbox databases after each run, under the same policy the agent was given, on a code path independent of the verifier.
- Facts and subject links are pulled over HTTP through the same failure layer, then the action is verified with its settlement window closed. Transient failures get the settlement-window re-check the production job would run.
- The trace-only judge is our own baseline for what traces can show, not any vendor's scorer.
- One recorded run per model. Agents are non-deterministic, so the agent-side counts move between runs; the verifier-side results are deterministic given the facts.
- Not covered: enforcement (verification is after the fact), real vendor APIs, load, and agents that falsify the systems of record themselves.
Every verdict in the runs carries a hash. See a verification or read the methodology.