Prove what your AI agents actually did.
CommitLayer verifies that every action an agent takes in your business systems was valid, within limits, and produced the intended outcome. Independent of what the agent says.
Free for developers, no card. Read-only connectors. No agent changes required.
Agents no longer just answer. They act.
A support agent now takes a message on WhatsApp, updates the CRM, refunds the order and closes the ticket. Its own report says it worked. The systems of record are the only place to find out whether it did.
- Agent
- Salesforce
- Shopify
- Zendesk
Wrong customer
Executed twice
Exceeded permissions
Downstream state never reached
One refund. Four expected effects. One that never happened.
The agent reported success. The payment was refunded, the CRM was updated and the customer was told. The order in Shopify was never cancelled, so the warehouse still ships it.
- stripe.refund_created observed within 15m of the actionobserved 10:42:07
- shopify.order_cancelled observed within 15m of the actionNOT OBSERVED
- crm.note_added observed within 15m of the actionobserved 10:42:09
- messaging.customer_notified observed within 15m of the actionobserved 10:42:11
Define → Record → Observe → Verify
- 1
Define
An action contract in YAML: the allowed action, required inputs, expected effects, limits, and forbidden effects.
contract: refund_workflow action: refund_order required_inputs: [amount, currency] limits: amount: { lte: 250 } expected_effects: - stripe.refund_created - shopify.order_cancelled forbidden_effects: - duplicate: stripe.refund_created - 2
Record
The agent's action, via one API call or an OpenTelemetry span. No change to the agent's logic.
POST /v1/actions { "contract_id": "refund_workflow", "action_type": "refund_order", "subject_id": "order_10231", "params": { "amount": 180 } } - 3
Observe
Read-only facts from the systems of record. Never the agent's own account of what it did.
shopify order_status 10:41:30 stripe refund_created 10:42:07 crm note_added 10:42:09 messaging customer_notified 10:42:11
- 4
Verify
PASS / FAIL / UNKNOWN per lens with plain-language reasons, hashed and chained to the previous verdict.
outcome: FAIL policy: FAIL reason: expected_effect_missing: shopify.order_cancelled hash: sha256:3f9a2c… prev: sha256:b21c7e…
Three verdicts, plain-language reasons.
Every action gets exactly one verdict per lens: outcome (did the intended change land?) and policy (was the action valid and within limits?). Reason codes explain why, and UNKNOWN is reported on its own so a gap in the evidence is never counted as a failure.
| Reason code | Label | In plain language |
|---|---|---|
| precondition_unmet | Precondition not met | A fact the contract requires before the action may run was not observed in the system of record, or did not have the required value. |
| limit_exceeded | Limit exceeded | A parameter of the action was outside the limit set in the contract, for example a refund amount above the cap. |
| expected_effect_missing | Expected effect not observed | The contract expects a change in a system of record after the action, and no matching fact was observed inside the settlement window. |
| forbidden_effect_observed | Forbidden effect observed | A fact the contract forbids, for example a second refund on the same order, was observed inside the window. |
| duplicate_effect | Duplicate effect | The same effect was observed more than once for the subject, which suggests the action was executed twice. |
| effects_pending | Effects still within settlement window | Not every expected effect has been observed yet, and the settlement window has not closed. The verdict is UNKNOWN until it does. |
| verified_condition_not_met | Verified condition not met | None of the conditions that would verify the claim were satisfied. |
| duplicate_claim_for_subject | Duplicate claim for subject | Another claim for the same subject exists in the same billing period; only the earliest counts. |
| missing_fact_source | Required fact source not connected | A fact source the definition depends on is not connected, so there are no facts to verify against. |
| identity_unmatched | Subject not found in fact sources | The claim's subject could not be found in any connected fact source, so nothing can be verified either way. |
Tested against three real agents.
Claude, GPT and Gemini each drove the same agent through 46 scenarios with failures injected where the agent cannot see them. A database oracle held the truth. CommitLayer matched it on every scenario, for every model; a judge that only reads the agent's trace did not. Run 2026-09-29.
Every model was fooled by the same two cases, a human who refunded first and a side effect in another system, and the trace-only judge passed both every time. The systems of record did not.
One verification layer, two products.
The same engine, contracts, connectors, facts, verdicts and hash chain, applied to the actions your agents take and to the outcomes your vendors bill for.
CommitLayer Verify
Action contracts for agents that act in your systems.
- Contracts: preconditions, limits, expected and forbidden effects
- Record with one API call or an OpenTelemetry span
- PASS / FAIL / UNKNOWN from the systems of record, hashed and chained
CommitLayer Reconcile
Outcome-priced vendor invoices, verified against your systems of record. Monthly statements, invoice reconciliation, dispute packages.
- Vendor and strict lenses on every billed outcome
- Monthly statements, invoice reconciliation, dispute packages
- Definitions transcribed from public documentation, cited
Built to be trusted by the people who sign off on agent actions.
Connectors request read-only, least-privilege scopes, listed in the Trust Center with the exact fields pulled. Action contracts observe; they never write to your systems.
Every verdict under an action contract is hashed and chained to the previous one. Verdict, statement and dispute-package hashes are publicly verifiable.
CommitLayer is paid by customers only. No payments, partnerships or data deals with the agent platforms or vendors it verifies.
Start with one contract.
Developer is free with no time limit: write an action contract, record one action, and read the verdict from your own systems. Pricing and answers to the questions operators ask first are one click away.