done mean verified. A code task that mutated files can reach done only when two independent checks agree: a deterministic gate (the task’s own build/test/lint command, judged by its exit code) passes, and, on a green gate, for a risky or large change, a read-only critic (an independent reviewer that cannot push) raises no blocking finding. The principle underneath is builder≠judge: the agent that did the work never certifies its own work. A generator self-grading is a known failure mode; Clawboo’s only signals into done are a machine truth (an exit code) and a structurally-independent reviewer.
This page explains what verification is and isn’t, the two-layer model and the typed verdict it produces, the swe-af severity taxonomy that rations blocking, the bounded fix loop and its completed_with_debt exit, and the intrinsic board gate that makes the rule un-bypassable.
The verification gate is intrinsic to the board state machine, not an opt-in step. Any transition to
done is rejected when the task carries a non-promotable verdict, by any caller, including the generic board PATCH route, the Tasks-MCP update_status tool, and the orchestrator. The only escape is an explicit, audited humanOverride. See The intrinsic board gate.What it is, and what it isn’t
Verification is a gate on→ done for file-mutating work, not a continuous test runner and not a quality score. It runs once, at completion, when a task’s worktree carries a non-empty diff. It produces a typed verdict that the board state machine reads mechanically, never prose a model can argue with.
A few things verification is not:
- Not a quality opinion. The deterministic gate is an exit code, not a judgement; the critic emits structured findings against a fixed severity taxonomy, not a free-form review.
- Not run on every change. A read-only task, a research task, an empty diff, or a small low-risk diff is not sent to the critic; a model is spent only where it earns its keep (see When the critic runs).
- Not a deadlock for the task, and not a rescue for its chain. When the attempt budget is spent, the verdict is recorded as
completed_with_debtand the task landsblockedwith its delegator notified, rather than sitting in a column that implies someone is still working on it. Its dependents keep waiting: the ready query requires every dependency to bedone. The move buys visibility, not recovery (see The bounded fix loop). - Not applicable to non-worktree work. A task with no isolated worktree (an OpenClaw-substrate run, a non-file-mutating native run) has no diff to verify, so it carries no verdict; and a task with no verdict is unverified, not failing.
→ done transition and a single verification cell on the task row); the worktree subsystem owns the isolated checkout the gate runs in; the verification subsystem owns the gate-and-critic composition. Verification is downstream of delegation: the orchestrator routes a structured failure verdict back to the team leader as a real, actionable fix request, not the string “FAIL”.
The model
The completion path for a file-mutating task: land the work in review, run the deterministic gate, run the critic only if the gate is green and the diff is risky, then map the verdict to a terminal status. The board gate then enforces that mapping.The two layers
The deterministic gate
The deterministic gate is the strongest signal because it reads a truth, not an opinion. It runs the task’s configured verify command, the build/test/lint command from the worktree’s system-of-record, inside the isolated worktree and records the result as a typedDeterministicResult (command, exitCode, passed, scrubbed stdout/stderr tails, durationMs, timedOut).
The command is resolved without scraping rendered output. The scaffold writes VERIFY_CMD='<shell-quoted>' into init.sh and a - Verify: `<cmd>` line into VERIFICATION.md; the gate parses that structured line (init.sh first, then VERIFICATION.md) and reverses the bash single-quote escaping. passed is true only when the command exits 0 and did not time out. A timeout kills the whole process tree (the shell wrapper plus the real test runner it spawned). The command + a scrubbed output tail are appended to VERIFICATION.md as evidence; the typed verdict on the task is the source of truth.
A missing or placeholder verify command is a structured fail, not a skip. done requires real evidence, so the absence of a check cannot certify completion; an unconfigured task fails the gate with (none configured) until someone sets VERIFY_CMD.
The deterministic gate uses
shell: true because the verify command is a free-form shell string (unlike the runtime drivers, which spawn a known binary by argv). It sets windowsHide to suppress the console popup on Windows, matching the repo’s spawn convention.The critic
When the gate is green and the change is non-trivial, an independent reviewer runs second. The critic’s independence is structural: it provisions a detached review worktree, checked out at the work’s committed SHA with no branch, so the reviewer literally has nothing to push and cannot mutate the builder’s branch. It then drives a reviewer adapter with a structured-output instruction and parses a typedCriticVerdict.
Independence is layered. At minimum it is context-level: a fresh session, a detached push-less checkout, and no builder home directory; the reviewer never shares the builder’s persisted native memory. When an operator sets a distinct reviewer model via CLAWBOO_REVIEWER_MODEL, independence becomes model-level too. The stored verdict records the reviewerModel and reviewerRuntime, so a same-model review’s bias caveat stays visible rather than hidden.
The critic is asked to emit only a single JSON object of findings, each a typed Finding with a severity, a title, an optional body / file path / start line, and a confidence. The output is parsed through the same structured-judge drive the eval grader uses (@clawboo/obs): valid JSON yields findings, an empty result yields no findings (a valid, good outcome), and anything unparseable becomes a single non-blocking other finding rather than a crash. A malformed or failed critic must never block the deterministic verdict; the gate is the hard signal; the critic is the second opinion.
When the critic runs
The critic is rationed. It fires only when at least one of these is true:
A small, undelegated, unflagged change skips the critic entirely (
ran: false), no review worktree, no model spend, and its verdict comes from the deterministic gate alone. The thresholds (files, lines) are overridable.
Composing one verdict
The two layers fold into a single attempt status, with the deterministic gate as the hard authority:- A red gate is always
fail, full stop, the critic never even runs. - A green gate plus a critic with a blocking finding is
fail, route the fix back. - A green gate with the critic not run, or only non-blocking findings, is
pass.
This is deliberate: a style nit or a perf observation must not deadlock a task, but a security hole, a crash, data loss, a wrong algorithm, or unmet acceptance criteria must. Rationing blocking to genuine defects keeps the fix loop from churning on cosmetics.
A failing attempt does not return the string “FAIL”. It carries a structured
{ what, why, howToFix }, for a red gate, the failing command and its scrubbed output tail; for blocking critic findings, the list of [severity] title (file:line), so the leader routes a concrete fix, not a guess.
The bounded fix loop and completed_with_debt
The independent evaluator is permanent; only the retry budget is bounded. Each verify attempt is recorded in an attempts[] array on the task, that array is the loop history, which is why the → done gate is a single-row read.
After a failing attempt, the loop decides whether to retry or stop. One budget governs both halves of that decision, verifyMaxAttempts(), which reads CLAWBOO_MAX_FIX_CYCLES as a count of fix cycles and returns the total attempts, one more. The default is 2 attempts: the initial verify plus one fix cycle. While attempts remain, the task goes back to in_progress with the structured fix note and the specialist tries again. On the last attempt the verdict is marked completed_with_debt and the open issues are recorded as debtNotes (each a { criterion, severity, justification }), and the task itself routes to blocked for a human.
completed_with_debt is not an unconditional pass. Its promotability is the load-bearing rule:
completed_with_debtover a green deterministic gate is still promotable by the board’s→ donegate, but the completion path no longer promotes it. A green gate with a failing verdict means the critic raised a blocking finding (security, crash, data loss, wrong algorithm, missing acceptance criteria), so the task routes toblockedfor review rather than landingdone.completed_with_debtover a red deterministic gate (the build/test gate is still failing after the loop exhausts) → not promotable. A red gate is the canonical blocking case; it routes toblockedfor a human rather than silently shipping.
pass lands done. A completed_with_debt verdict, and any fail that has spent the attempt budget, lands blocked with a system comment recording how many attempts failed and an inbox notice to the delegator, or to the team leader when the task has no delegator. A fail with attempts left reverts to in_progress so the next fix cycle can re-dispatch it.
The intrinsic board gate
The rule above is enforced in the board state machine itself, so it cannot be skipped by reachingdone through a different door. Every transition to done runs a single shared check, isVerdictPromotable, against the freshly-read task row inside a BEGIN IMMEDIATE transaction:
pass→ promotable.completed_with_debt→ promotable only if the latest attempt’s deterministic gate was green.- anything else (
fail, or an unparseable verdict that is present and non-promotable) → not promotable. - no stored verdict → the task is unverified, not failing, and lands
donenormally. The gate blocks known-failing verdicts, not un-run verification. (An unparseable-but-present cell is treated leniently, if promotability can’t be determined, the gate doesn’t block.)
verification_required, which the board REST layer maps to a 409. The state machine is the backstop for any caller that reaches done another way; the worktree completion path is stricter still and promotes only a clean pass, so “done means verified” holds at every entry point.
The one escape is humanOverride, a human deciding to ship despite a non-promotable verdict. It is the only way a task with a known-failing verdict can reach done, and the caller must audit it: the PATCH /api/board/:taskId handler writes a verification audit row ({ override: true, route: 'board_patch', priorStatus, to: 'done' }) whenever an override-to-done succeeds. The override is never silent.
Moving a task back to
todo clears the stored verification verdict, on the in_progress → todo re-claim path and on the blocked → todo re-queue a human uses after an exhausted fix loop. This is intentional: a release is a cross-runtime rebind boundary, and a previous runtime’s failing verdict must not gate a fresh runtime’s legitimate completion. The next runtime re-verifies from scratch.Design rationale and trade-offs
Verification exists because a self-grading generator is unreliable, and because “the agent said it’s done” is not evidence. Makingdone mean verified buys a real completion guarantee, at the cost of a verify command per task and, on risky changes, a second model run.
The two-layer split is deliberate. The deterministic gate is cheap, objective, and non-negotiable: a red gate is always a failure, with no model in the loop to be talked out of it. The critic adds judgement that an exit code can’t capture (a security hole that still compiles, a wrong algorithm that still passes thin tests); but judgement is expensive and a same-model self-review is biased, so it is rationed to risky surfaces and made structurally independent (detached, push-less, no shared home).
completed_with_debt is the record, not a shortcut. Bounding the retry budget (never the evaluator) keeps a task from churning forever, and the debt notes say what was still open when the budget ran out; but nothing ships on that record. Only a clean pass lands done, so a red gate and an unresolved blocking finding both end in front of a human. Non-blocking findings never reach this path at all: a green gate with only style / perf / other findings is a pass. The trade-off is that a task nobody revisits stays blocked indefinitely, and its dependents stay unready with it.
Making the gate intrinsic to the state machine, rather than a step the orchestrator is trusted to call, is what makes the rule un-bypassable. The cost is a small inline read of the verification cell on every → done transition; the benefit is that no caller (MCP tool, REST route, future orchestrator) can route around it without the audited override.
Boundaries and non-goals
- Not a CI system. Verification runs the task’s own verify command once at completion; it does not provide a pipeline, scheduling, caching, or matrix builds.
- Only worktree-backed work is verified. Tasks without an isolated worktree carry no verdict and land
doneun-gated. Verification is for file-mutating work that has a diff to check. - The critic’s model independence is opt-in. Without
CLAWBOO_REVIEWER_MODEL, the critic reuses the run’s own adapter factory; independence is context-level (fresh session, detached checkout, no builder home), and the verdict records the reviewer model so the same-model caveat is visible. - The deterministic gate trusts the configured command. It reads an exit code; it does not validate that the command actually exercises the change. A weak
VERIFY_CMDyields a weak gate.
These docs describe Clawboo v0.3.1, the current release.
See also
- The board, the state machine the
→ donegate lives in - Governance, budgets, circuit breakers, caps, and approvals that bound a run
- Worktrees and handoff, the isolated checkout the gate runs in and the system-of-record
VERIFY_CMDlives in - Delegation and orchestration, how a structured fix verdict routes back to the leader
- Board API, the REST surface, including the
409 verification_requiredand the auditedhumanOverride - Glossary, canonical term definitions