Skip to content
CommitLayer Verify

Prove what your agents did in your systems.

Action contracts for agents that act in your systems. Define what a valid action looks like, record what the agent did, and get a verdict from the systems of record, independent of what the agent says.

Free for developers, no card. Read-only connectors. No agent changes required.

What it catches

The failures a transcript cannot show.

Wrong customer

The refund went to a different order than the one the customer wrote in about. Every step succeeded; the outcome is still wrong.

Executed twice

A retry after a timeout created a second refund. The agent reports one; the payment system holds two.

Exceeded permissions

The amount was above the cap the workflow allows, or a parameter the contract does not permit was set.

Downstream state never reached

The payment was refunded but the order stayed active, so the warehouse still ships it. Nothing in the transcript says so.
How it works

Define → Record → Observe → Verify

  1. 1

    Define

    An action contract in YAML: the allowed action, required inputs, expected effects, limits, and forbidden effects.

    contract: refund_workflow
    action: refund_order
    required_inputs: [amount, currency]
    limits:
      amount: { lte: 250 }
    expected_effects:
      - stripe.refund_created
      - shopify.order_cancelled
    forbidden_effects:
      - duplicate: stripe.refund_created
  2. 2

    Record

    The agent's action, via one API call or an OpenTelemetry span. No change to the agent's logic.

    POST /v1/actions
    {
      "contract_id": "refund_workflow",
      "action_type": "refund_order",
      "subject_id": "order_10231",
      "params": { "amount": 180 }
    }
  3. 3

    Observe

    Read-only facts from the systems of record. Never the agent's own account of what it did.

    shopify    order_status       10:41:30
    stripe     refund_created     10:42:07
    crm        note_added         10:42:09
    messaging  customer_notified  10:42:11
  4. 4

    Verify

    PASS / FAIL / UNKNOWN per lens with plain-language reasons, hashed and chained to the previous verdict.

    outcome: FAIL
    policy:  FAIL
    reason:  expected_effect_missing:
             shopify.order_cancelled
    hash:    sha256:3f9a2c…
    prev:    sha256:b21c7e…
Sample verification

One refund. Four expected effects. One that never happened.

Synthetic demo data.
refund_order · order_10231FAIL
Contract:Refund workflow (internal agent)
Precondition:latest shopify.order_status before the action had status=eligible — met
Limit:amount above 250 — within limit (180)
Expected effects:4
  • stripe.refund_created observed within 15m of the actionobserved 10:42:07
  • shopify.order_cancelled observed within 15m of the actionNOT OBSERVED
  • crm.note_added observed within 15m of the actionobserved 10:42:09
  • messaging.customer_notified observed within 15m of the actionobserved 10:42:11
Forbidden:more than 1 stripe.refund_created effect(s) observed within 15m — none observed
Forbidden:stripe.refund_created observed with amount above 250 — none observed
Result:FAIL
Reason:Shopify order remained active despite a successful payment refund
Evidence hash:sha256:d1b0a88bbf37381f… Verify publicly
Verdicts

Three verdicts, plain-language reasons.

PASS
Every precondition held, every parameter was within its limit, no forbidden effect was observed, and every expected effect landed in the systems of record inside the window. Each fact is pointed to by source, record and sha256.
FAIL
At least one check failed on the facts: a precondition was not met, a limit was exceeded, a forbidden effect (such as a duplicate refund) was observed, or an expected effect never reached the system of record.
UNKNOWN
The facts needed to decide are not available yet, or not at all: the settlement window is still open, no connected source covers the subject, or the identity could not be matched. Reported separately; never counted as a FAIL.
Reason codes and their plain-language meaning
Reason codeLabelIn plain language
precondition_unmetPrecondition not metA fact the contract requires before the action may run was not observed in the system of record, or did not have the required value.
limit_exceededLimit exceededA parameter of the action was outside the limit set in the contract, for example a refund amount above the cap.
expected_effect_missingExpected effect not observedThe contract expects a change in a system of record after the action, and no matching fact was observed inside the settlement window.
forbidden_effect_observedForbidden effect observedA fact the contract forbids, for example a second refund on the same order, was observed inside the window.
duplicate_effectDuplicate effectThe same effect was observed more than once for the subject, which suggests the action was executed twice.
effects_pendingEffects still within settlement windowNot every expected effect has been observed yet, and the settlement window has not closed. The verdict is UNKNOWN until it does.
verified_condition_not_metVerified condition not metNone of the conditions that would verify the claim were satisfied.
duplicate_claim_for_subjectDuplicate claim for subjectAnother claim for the same subject exists in the same billing period; only the earliest counts.
missing_fact_sourceRequired fact source not connectedA fact source the definition depends on is not connected, so there are no facts to verify against.
identity_unmatchedSubject not found in fact sourcesThe claim's subject could not be found in any connected fact source, so nothing can be verified either way.
What is included

Everything from the Developer plan up.

Action contracts

YAML with preconditions, limits, expected effects, forbidden effects and a settlement window. Versioned, dry-runnable, compiled to deterministic rules.

Record and facts APIs

One call records the action. Facts arrive from read-only connectors, CSV, the facts API or OpenTelemetry spans from the agent runtime.

Verdicts with evidence

PASS / FAIL / UNKNOWN per lens, a reason code that names the missing or forbidden effect, the facts that decided it, and a hash chained to the previous verdict.

Settlement window

A provisional verdict while effects are still landing, re-checked automatically when the window closes. The re-check is not metered.

Orphan-effect report

Effects observed in your systems with no recorded action behind them: the wrong-customer and injected-instruction cases.

Webhook and public verification

action.verified on every settled verdict, and a public page where anyone can confirm a verdict hash without an account.

Connectors to more systems of record, the orphan-effect report and OpenTelemetry ingestion come with plan size. Pricing.

Benchmark

Tested against three real agents.

Claude, GPT and Gemini each drove the same agent through 46 scenarios with injected failures. CommitLayer matched the database oracle on every scenario for every model.

claude-opus-5-5
46 / 46
gpt-5-mini
46 / 46
gemini-flash-latest
46 / 46
    CommitLayer Verify · CommitLayer