Skip to content

Prove what your AI agents actually did.

CommitLayer verifies that every action an agent takes in your business systems was valid, within limits, and produced the intended outcome. Independent of what the agent says.

Free for developers, no card. Read-only connectors. No agent changes required.

The problem

Agents no longer just answer. They act.

A support agent now takes a message on WhatsApp, updates the CRM, refunds the order and closes the ticket. Its own report says it worked. The systems of record are the only place to find out whether it did.

  1. WhatsApp
  2. Agent
  3. Salesforce
  4. Shopify
  5. Zendesk

Wrong customer

The refund went to a different order than the one the customer wrote in about. Every step succeeded; the outcome is still wrong.

Executed twice

A retry after a timeout created a second refund. The agent reports one; the payment system holds two.

Exceeded permissions

The amount was above the cap the workflow allows, or a parameter the contract does not permit was set.

Downstream state never reached

The payment was refunded but the order stayed active, so the warehouse still ships it. Nothing in the transcript says so.
Sample verification

One refund. Four expected effects. One that never happened.

The agent reported success. The payment was refunded, the CRM was updated and the customer was told. The order in Shopify was never cancelled, so the warehouse still ships it.

Synthetic demo data.
refund_order · order_10231FAIL
Contract:Refund workflow (internal agent)
Precondition:latest shopify.order_status before the action had status=eligible — met
Limit:amount above 250 — within limit (180)
Expected effects:4
  • stripe.refund_created observed within 15m of the actionobserved 10:42:07
  • shopify.order_cancelled observed within 15m of the actionNOT OBSERVED
  • crm.note_added observed within 15m of the actionobserved 10:42:09
  • messaging.customer_notified observed within 15m of the actionobserved 10:42:11
Forbidden:more than 1 stripe.refund_created effect(s) observed within 15m — none observed
Forbidden:stripe.refund_created observed with amount above 250 — none observed
Result:FAIL
Reason:Shopify order remained active despite a successful payment refund
Evidence hash:sha256:d1b0a88bbf37381f… Verify publicly
How it works

Define → Record → Observe → Verify

  1. 1

    Define

    An action contract in YAML: the allowed action, required inputs, expected effects, limits, and forbidden effects.

    contract: refund_workflow
    action: refund_order
    required_inputs: [amount, currency]
    limits:
      amount: { lte: 250 }
    expected_effects:
      - stripe.refund_created
      - shopify.order_cancelled
    forbidden_effects:
      - duplicate: stripe.refund_created
  2. 2

    Record

    The agent's action, via one API call or an OpenTelemetry span. No change to the agent's logic.

    POST /v1/actions
    {
      "contract_id": "refund_workflow",
      "action_type": "refund_order",
      "subject_id": "order_10231",
      "params": { "amount": 180 }
    }
  3. 3

    Observe

    Read-only facts from the systems of record. Never the agent's own account of what it did.

    shopify    order_status       10:41:30
    stripe     refund_created     10:42:07
    crm        note_added         10:42:09
    messaging  customer_notified  10:42:11
  4. 4

    Verify

    PASS / FAIL / UNKNOWN per lens with plain-language reasons, hashed and chained to the previous verdict.

    outcome: FAIL
    policy:  FAIL
    reason:  expected_effect_missing:
             shopify.order_cancelled
    hash:    sha256:3f9a2c…
    prev:    sha256:b21c7e…
Verdicts

Three verdicts, plain-language reasons.

Every action gets exactly one verdict per lens: outcome (did the intended change land?) and policy (was the action valid and within limits?). Reason codes explain why, and UNKNOWN is reported on its own so a gap in the evidence is never counted as a failure.

PASS
Every precondition held, every parameter was within its limit, no forbidden effect was observed, and every expected effect landed in the systems of record inside the window. Each fact is pointed to by source, record and sha256.
FAIL
At least one check failed on the facts: a precondition was not met, a limit was exceeded, a forbidden effect (such as a duplicate refund) was observed, or an expected effect never reached the system of record.
UNKNOWN
The facts needed to decide are not available yet, or not at all: the settlement window is still open, no connected source covers the subject, or the identity could not be matched. Reported separately; never counted as a FAIL.
Reason codes and their plain-language meaning
Reason codeLabelIn plain language
precondition_unmetPrecondition not metA fact the contract requires before the action may run was not observed in the system of record, or did not have the required value.
limit_exceededLimit exceededA parameter of the action was outside the limit set in the contract, for example a refund amount above the cap.
expected_effect_missingExpected effect not observedThe contract expects a change in a system of record after the action, and no matching fact was observed inside the settlement window.
forbidden_effect_observedForbidden effect observedA fact the contract forbids, for example a second refund on the same order, was observed inside the window.
duplicate_effectDuplicate effectThe same effect was observed more than once for the subject, which suggests the action was executed twice.
effects_pendingEffects still within settlement windowNot every expected effect has been observed yet, and the settlement window has not closed. The verdict is UNKNOWN until it does.
verified_condition_not_metVerified condition not metNone of the conditions that would verify the claim were satisfied.
duplicate_claim_for_subjectDuplicate claim for subjectAnother claim for the same subject exists in the same billing period; only the earliest counts.
missing_fact_sourceRequired fact source not connectedA fact source the definition depends on is not connected, so there are no facts to verify against.
identity_unmatchedSubject not found in fact sourcesThe claim's subject could not be found in any connected fact source, so nothing can be verified either way.
Benchmark

Tested against three real agents.

Claude, GPT and Gemini each drove the same agent through 46 scenarios with failures injected where the agent cannot see them. A database oracle held the truth. CommitLayer matched it on every scenario, for every model; a judge that only reads the agent's trace did not. Run 2026-09-29.

Anthropic
claude-opus-5-5
46 / 46
verdicts matched the oracle · 12 real failures, all named
Trace-only judge: 39 / 46
OpenAI
gpt-5-mini
46 / 46
verdicts matched the oracle · 10 real failures, all named
Trace-only judge: 31 / 46
Google
gemini-flash-latest
46 / 46
verdicts matched the oracle · 8 real failures, all named
Trace-only judge: 37 / 46

Every model was fooled by the same two cases, a human who refunded first and a side effect in another system, and the trace-only judge passed both every time. The systems of record did not.

Products

One verification layer, two products.

The same engine, contracts, connectors, facts, verdicts and hash chain, applied to the actions your agents take and to the outcomes your vendors bill for.

CommitLayer Verify

Action contracts for agents that act in your systems.

  • Contracts: preconditions, limits, expected and forbidden effects
  • Record with one API call or an OpenTelemetry span
  • PASS / FAIL / UNKNOWN from the systems of record, hashed and chained

CommitLayer Reconcile

Outcome-priced vendor invoices, verified against your systems of record. Monthly statements, invoice reconciliation, dispute packages.

  • Vendor and strict lenses on every billed outcome
  • Monthly statements, invoice reconciliation, dispute packages
  • Definitions transcribed from public documentation, cited
Security

Built to be trusted by the people who sign off on agent actions.

Read-only by design

Connectors request read-only, least-privilege scopes, listed in the Trust Center with the exact fields pulled. Action contracts observe; they never write to your systems.

Evidence integrity

Every verdict under an action contract is hashed and chained to the previous one. Verdict, statement and dispute-package hashes are publicly verifiable.

Independence policy

CommitLayer is paid by customers only. No payments, partnerships or data deals with the agent platforms or vendors it verifies.

Start with one contract.

Developer is free with no time limit: write an action contract, record one action, and read the verdict from your own systems. Pricing and answers to the questions operators ask first are one click away.

    CommitLayer — The verification layer for enterprise AI actions.