Prove what your agents did in your systems.
Action contracts for agents that act in your systems. Define what a valid action looks like, record what the agent did, and get a verdict from the systems of record, independent of what the agent says.
Free for developers, no card. Read-only connectors. No agent changes required.
The failures a transcript cannot show.
Wrong customer
Executed twice
Exceeded permissions
Downstream state never reached
Define → Record → Observe → Verify
- 1
Define
An action contract in YAML: the allowed action, required inputs, expected effects, limits, and forbidden effects.
contract: refund_workflow action: refund_order required_inputs: [amount, currency] limits: amount: { lte: 250 } expected_effects: - stripe.refund_created - shopify.order_cancelled forbidden_effects: - duplicate: stripe.refund_created - 2
Record
The agent's action, via one API call or an OpenTelemetry span. No change to the agent's logic.
POST /v1/actions { "contract_id": "refund_workflow", "action_type": "refund_order", "subject_id": "order_10231", "params": { "amount": 180 } } - 3
Observe
Read-only facts from the systems of record. Never the agent's own account of what it did.
shopify order_status 10:41:30 stripe refund_created 10:42:07 crm note_added 10:42:09 messaging customer_notified 10:42:11
- 4
Verify
PASS / FAIL / UNKNOWN per lens with plain-language reasons, hashed and chained to the previous verdict.
outcome: FAIL policy: FAIL reason: expected_effect_missing: shopify.order_cancelled hash: sha256:3f9a2c… prev: sha256:b21c7e…
One refund. Four expected effects. One that never happened.
- stripe.refund_created observed within 15m of the actionobserved 10:42:07
- shopify.order_cancelled observed within 15m of the actionNOT OBSERVED
- crm.note_added observed within 15m of the actionobserved 10:42:09
- messaging.customer_notified observed within 15m of the actionobserved 10:42:11
Three verdicts, plain-language reasons.
| Reason code | Label | In plain language |
|---|---|---|
| precondition_unmet | Precondition not met | A fact the contract requires before the action may run was not observed in the system of record, or did not have the required value. |
| limit_exceeded | Limit exceeded | A parameter of the action was outside the limit set in the contract, for example a refund amount above the cap. |
| expected_effect_missing | Expected effect not observed | The contract expects a change in a system of record after the action, and no matching fact was observed inside the settlement window. |
| forbidden_effect_observed | Forbidden effect observed | A fact the contract forbids, for example a second refund on the same order, was observed inside the window. |
| duplicate_effect | Duplicate effect | The same effect was observed more than once for the subject, which suggests the action was executed twice. |
| effects_pending | Effects still within settlement window | Not every expected effect has been observed yet, and the settlement window has not closed. The verdict is UNKNOWN until it does. |
| verified_condition_not_met | Verified condition not met | None of the conditions that would verify the claim were satisfied. |
| duplicate_claim_for_subject | Duplicate claim for subject | Another claim for the same subject exists in the same billing period; only the earliest counts. |
| missing_fact_source | Required fact source not connected | A fact source the definition depends on is not connected, so there are no facts to verify against. |
| identity_unmatched | Subject not found in fact sources | The claim's subject could not be found in any connected fact source, so nothing can be verified either way. |
Everything from the Developer plan up.
YAML with preconditions, limits, expected effects, forbidden effects and a settlement window. Versioned, dry-runnable, compiled to deterministic rules.
One call records the action. Facts arrive from read-only connectors, CSV, the facts API or OpenTelemetry spans from the agent runtime.
PASS / FAIL / UNKNOWN per lens, a reason code that names the missing or forbidden effect, the facts that decided it, and a hash chained to the previous verdict.
A provisional verdict while effects are still landing, re-checked automatically when the window closes. The re-check is not metered.
Effects observed in your systems with no recorded action behind them: the wrong-customer and injected-instruction cases.
action.verified on every settled verdict, and a public page where anyone can confirm a verdict hash without an account.
Connectors to more systems of record, the orphan-effect report and OpenTelemetry ingestion come with plan size. Pricing.
Tested against three real agents.
Claude, GPT and Gemini each drove the same agent through 46 scenarios with injected failures. CommitLayer matched the database oracle on every scenario for every model.