packages/db/src/board/: four modules: repository.ts (the data access), contention.ts (the write recipe), schemas.ts (the zod shapes for REST bodies and raw-CTE results), and verification.ts (verdict persistence — setTaskVerification — feeding the intrinsic → done gate). The pure transition rules sit one package over, in @clawboo/board-core, and are re-exported through @clawboo/db unchanged. The connection-level pragmas they depend on are set in packages/db/src/db.ts. If you want the why behind the board’s existence, narration-vs-authority, the worktree linkage, read the concept page first; this page assumes it.
What it is, and what it isn’t
The board data-access layer is the only place in the codebase that touches the board tables.apps/web never reaches a board table with raw Drizzle; it calls a repository function. That is deliberate: it keeps query construction out of the app, and it is the single seam a future SQLite→Postgres or multi-tenant swap would target. Every read that could be tenant-scoped takes an optional scope argument carrying a tenantId; the column exists on every board table but no filtering is active in v0.3.1.
It is not an ORM-over-everything layer. It is a hand-written set of typed functions over five tables (tasks, task_deps, task_comments, workspaces, execution_processes), each function doing exactly one board operation. The pure transition rules live apart from the data access so they are unit-testable without a database; the contention helpers live apart so the rest of the repository reads as plain Drizzle.
The repository imports one cross-package symbol:
isVerdictPromotable from @clawboo/governance. That is the shared builder-≠-judge rule, kept in one place so the board state machine and the worktree completion path enforce “done means verified” identically. See Verification.The state machine
@clawboo/board-core’s state-machine.ts is pure: a TaskStatus union of seven values, a LEGAL transition map, and the accessors over it. There is no I/O here — and no imports at all, which is what lets the browser bundle the same file the server enforces.
It lives outside @clawboo/db deliberately. Three layers need these rules and only one of them can touch a database: the repository enforces them inside the write transaction, @clawboo/team-orchestration types its BoardClient with TaskStatus, and the board UI derives its columns and its status editor from TASK_STATUSES + legalTargets. When each declared its own copy they agreed, but nothing linked them: a newly-legal transition on the server would have left the UI hiding a move it now accepts, with green CI.
canTransition(from, to) returns true when from === to (same-status is an idempotent no-op, so re-emitting a transition is harmless) and otherwise checks membership in LEGAL[from]. legalTargets(from) returns that row as a fresh array, which is how a UI enumerates the moves it may offer without copying the table. done and cancelled have empty target lists; they are terminal (a test pins isTerminal(s) ⇔ legalTargets(s).length === 0, so the two can never disagree). isLocked(status) returns true for in_progress and in_review (a locked task is actively owned and must not have its assignee reassigned). isTerminal(status) returns true for done and cancelled.
The enforcement happens in updateStatus in the repository, not in the REST layer:
The row is read inside the BEGIN IMMEDIATE transaction, so the transition is checked against the freshly read status; two concurrent callers cannot both pass a stale pre-check. Any pre-check the REST layer does first is purely fast-fail ergonomics; the transactional check is the real gate. An illegal transition surfaces as 409 from the REST handler.
Two side effects fire inside that same write:
- Terminal states set
completedAt. Reachingdoneorcancelledstamps the completion timestamp. → todois the release path, and it unassigns. Moving a task back totodoclearsassigneeAgentId,assigneeRuntime, and the storedverificationverdict. The assignee clear is what lets the atomic claim re-acquire the task (the claim guards onassignee IS NULL); without it, every re-fire of a released task would409. The verdict clear is the cross-runtime rebind boundary: a previous runtime’s failing verdict must not gate a fresh runtime’s legitimate completion.releaseTaskperforms the same three-field clear directly for thein_progress → todocase.
The intrinsic verification gate
The→ done gate is the board enforcing builder-≠-judge in the substrate itself, not in a caller:
verdictCellPromotable is a lightweight read of the persisted JSON cell. A full zod parse happens on write (setTaskVerification in verification.ts); here a tiny inline JSON.parse plus the shared isVerdictPromotable rule avoids importing board/verification.ts, which would create an import cycle with getTask. The semantics:
The gate is un-bypassable by any caller except an explicit
opts.humanOverride: true. The REST PATCH /api/board/:taskId handler accepts humanOverride in the request body and records the bypass in the governance audit log; the override is auditable, never silent. The autonomous path always writes a verdict via verifyTask before this transition, so the gate is exercised on genuine completions, not as a routine bypass.A subtlety for contributors:
isVerdictPromotable(null) in @clawboo/governance returns false, but verdictCellPromotable(null) in the repository returns true. They are not contradictory; verdictCellPromotable only calls isVerdictPromotable when a cell exists. A null cell short-circuits to “promotable” because a missing verdict means “unverified,” which the gate deliberately does not block. Do not “fix” this by passing the null straight through.The atomic claim
A task is worked by exactly one assignee. The board guarantees that with a single conditional UPDATE, not a read-then-write:status='todo' AND assignee IS NULL AND dropped=0) is part of the same atomic statement that performs the write, at most one concurrent caller can win. The winner gets the updated row back via RETURNING; every loser gets a zero-row result.
That zero-row result is data, not an error. claimTask distinguishes the two failure modes with a follow-up existence check and returns { ok: false, reason: 'conflict' } when the row exists (someone else claimed it) or { ok: false, reason: 'not_found' } when it does not. The REST layer maps conflict → 409 and not_found → 404. The rule the whole system follows is never retry a 409; a conflict means someone legitimately owns the work, so a retry would either no-op or fight for a task already taken.
Stale re-claim of a dead, abandoned in_progress task is deliberately not the claim’s job. The claim’s WHERE only matches todo, so an abandoned in_progress task is invisible to it. Liveness lives in one place; reconciliation releases such a task back to todo, after which an ordinary claim re-acquires it. Keeping the claim a pure mutex and liveness a separate pass is what makes both correct.
Dependencies and the recursive CTEs
Tasks form a blocks / blocked-by graph intask_deps, with a composite primary key on (task_id, depends_on_task_id) that prevents duplicate edges. linkDep writes the whole edge inside immediateWrite: it rejects a self-link outright, then probes reachability with a recursive CTE and throws TaskDependencyCycleError when taskId is already upstream of dependsOnTaskId — that is, when the new edge would close a direct or transitive cycle. Only after the probe does it insert, still with onConflictDoNothing, so re-linking an existing edge remains a harmless no-op. Probe and insert share one BEGIN IMMEDIATE transaction, so two concurrent links cannot race a cycle into the graph.
A cycle is refused at write time because nothing downstream can recover from one:
getReadyTasks requires every dependency to be done, so every task in a cycle is permanently un-ready and no error surfaces anywhere — the ready-pump simply has nothing to fire. TaskDependencyCycleError carries code: 'task_dependency_cycle'; the link_task MCP tool maps it to a tool-error, and the REST dependency route maps it to a 409 with that code in error — a client mistake, not a server fault.todo, not dropped, and every one of its dependencies is done. getReadyTasks expresses that with a NOT EXISTS subquery and orders the results by priority descending, then updatedAt descending:
getDependents walks the dependency graph downstream, every task transitively blocked by a given task:
blocked or otherwise can’t reach done, its downstream chain can never become ready and would otherwise sit forever as ghost todo cards. cancelDependents walks getDependents, cancels only the still-pending (todo / backlog) members via updateStatus, and returns the cancelled rows so the orchestrator can report the stalled plan to the team leader. Tasks already in_progress, done, or cancelled are left untouched.
getAncestors walks the parent chain via the self-referential parent_task_id, returning a minimal { id, parent_task_id, title, status } row per ancestor. Its raw rows are parsed with ancestorRowsSchema from schemas.ts before being returned. It is the depth cap’s counter at two call sites, both reading the same DEFAULT_MAX_DEPTH: the team orchestrator refuses to spawn when the SOURCE task is already at the ceiling (ancestors.length >= max, so the child would be one past it), and the executor runner refuses to dispatch a task that is itself past it (ancestors.length > max). The >=/> split is what makes the deepest creatable task also dispatchable. The creation-time cap enforces the orchestrator’s rule but does not use this helper — see createCappedSubtask below.
createCappedSubtask is the board’s own bounded create, and the one the Tasks MCP create tools call. It does the parent lookup, an indexed COUNT(*) of the parent’s non-dropped children, a step-bounded parent-chain walk, and the insert inside a single BEGIN IMMEDIATE transaction, returning { ok: false, reason: 'parent_not_found' | 'child_cap' | 'depth_cap' } rather than throwing. The transaction is the point: the Tasks MCP ships as a stdio bin, one OS process per attached runtime on the shared database file, so a count-then-insert window would let two agents both land the N+1th child. It walks the chain itself instead of calling getAncestors because that helper takes a ClawbooDb rather than a transaction handle, and because a hard step bound cannot spin on corrupt data while a write lock is held. createCappedRootTask is its parentless sibling: same one-transaction shape, but the guard is a rolling-window RATE on root creation rather than a total, because a per-parent ceiling has no subject on a root and a lifetime ceiling would eventually jam a long-lived board. createSubtask remains the uncapped primitive for the orchestrator, the REST board, and the evals harness. Every count comes from durable board state, never from the model.
Worktree linkage and execution rows
The board owns pointers into the worktree subsystem, not the worktree itself.setTaskWorkspaceRefs records worktreeRef and branchRef on the task; createWorkspace / updateWorkspaceStatus track a workspaces row through active → archived (cleanup) or active → stale (garbage collection). Provisioning a worktree for a read-only or research task is refused by the REST layer with 422 no_isolation; the worktree is only for the concurrency boundary that file-mutating work needs. See Worktrees and handoff.
Each spawned run is an execution_processes row, a per-run ledger for any executor. createExecutionProcess opens it (started running); completeExecutionProcess closes it with an outcome (succeeded, failed, timed_out, or cancelled) plus an optional token/cost ledger and git before/after commit checkpoints. The close is guarded on status = 'running', so a terminal row is immutable: the first terminal wins, and a losing writer gets null back and appends no second execution_completed. That guard is load-bearing because the ledger is the dispatch pump’s fire-policy input, where cancelled means user Stop; a late abort overwriting an already-swept timed_out row would reclassify infra death as user intent and park refireable work for good. By convention an exec row is opened only after a successful claim, because this ledger is exactly what crash recovery reads on restart.
The contention recipe
Clawboo is team-first: many agents may write one SQLite file, and out-of-process consumers (the MCP stdio bins an external runtime spawns) open the same file the Express server serves. Without care, concurrent writers hit SQLite’s single-writer lock and degrade into a “convoy.” The board’s answer is a layered recipe: connection-level pragmas plus two application-level pieces. Every connection sets these pragmas at open (openDb, packages/db/src/db.ts). The Express server opens exactly one, at boot; each out-of-process MCP stdio bin opens one of its own:
On top of those,
contention.ts adds two application-level mechanisms:
withWriteRetry(fn)runs a synchronous write and retries only transient lock errors,SQLITE_BUSY,SQLITE_BUSY_SNAPSHOT,SQLITE_LOCKED(the setisBusyErrorrecognizes), with a jittered backoff (20–150ms) until a wall-clock budget expires (1500ms by default, tunable withCLAWBOO_DB_WRITE_BUDGET_MS, with a24-attempt backstop). The budget is a deadline rather than an attempt count because the retry sleep blocks the calling thread, which in the server is the event loop: an attempt count does not bound how long one contended write can freeze the server, a deadline does. Worst case is the budget plus one finalbusy_timeout, about1.75s. On exhaustion it throwsWriteBudgetExhaustedError, which keepscode = 'SQLITE_BUSY'soisBusyErrorstill recognizes it and existing callers are unaffected. A 0-row result, like a lost claim race, is not an exception; it is returned to the caller unretried, which is what lets callers honor the never-retry-a-409 rule. Any non-lock error propagates immediately.immediateWrite(db, cb)runscbinside adb.transaction(cb, { behavior: 'immediate' });BEGIN IMMEDIATEacquires the write lock up front rather than escalating mid-transaction, avoiding lock-escalation deadlocks. The whole transaction re-runs from scratch on a transient lock error becauseimmediateWriteitself is wrapped inwithWriteRetry.updateStatus,reconcileOrphans, andreconcileStaleInProgressall run through it.
The retry sleep is synchronous and non-busy-spinning: it uses
Atomics.wait on a throwaway SharedArrayBuffer. better-sqlite3 is fully synchronous, so making every repository method async just to await a backoff would be a needless cost; the synchronous block on a transient lock is the correct shape.wal_autocheckpoint=50 PASSIVE checkpoint. It is the SQLite-blessed mechanism and needs no access to the raw handle.
Orphan and stale reconciliation
Two recovery passes keep the board from accumulating stuck work, one for crashes, one for abandonment. Both run in a singleimmediateWrite transaction and never block boot.
reconcileOrphans runs once at startup, gated on the heartbeat. It takes opts.staleAfterMs (default 6 * TASK_HEARTBEAT_MS, and the server passes the same CLAWBOO_BOARD_STALE_TTL_MS the sweep uses) and reaps a running, un-tombstoned execution_processes row only when its task row is missing or its updatedAt predates that cutoff. A row whose task is still beating belongs to a live process and must survive boot. Each reaped execution is marked failed, gets recovery_tombstone=1 so a second pass is a no-op (no infinite auto-resume), and its task is released back to todo if the task is currently in_progress or in_review. The release clears the assignee, the assignee runtime, and the verification verdict, the same unassign-on-release discipline as updateStatus(→todo). The server’s boot sequence calls it best-effort, logging and continuing on failure.
reconcileStaleInProgress runs at boot and on an interval. It releases an in_progress task whose owner has stopped proving it is alive, whether that owner died with the server or wedged inside a live one. The release publishes a task_released lifecycle event, which detaches the task from any resident orchestrator still mapping it to a session and wakes that team’s dispatch pump. A task that is in_progress, not dropped, whose updatedAt predates a TTL cutoff, and whose execution is still running is timed out (the exec → timed_out) and released to todo. The server wires it as one immediate sweep plus a setInterval(...).unref() loop.
tasks.updatedAt is a liveness signal. Every drain that claims a task also beats it: startTaskHeartbeat bumps updatedAt every TASK_HEARTBEAT_MS (30 seconds) for as long as the run owns the task, and the beat is a timer, not a side effect of events, so a silent-but-alive run keeps beating. That turns the sweep from an hour-scale guess into a minutes-scale fact, so the TTL is 3 minutes by default (six missed beats), tunable via CLAWBOO_BOARD_STALE_TTL_MS; the sweep interval is CLAWBOO_BOARD_STALE_SWEEP_MS, default 60 seconds.The beat is what makes the short TTL safe, so a drain that claims a task without beating it will have its live work swept. There are three: the executor runner, the team orchestrator’s serverDeliver, and the routines dispatcher. Any new one must beat too.The orchestrator’s own idle watchdog is still the first responder for a delegate that goes quiet: the per-team server orchestrator sweeps every 30 seconds and fails any delegate silent past DELEGATION_IDLE_TIMEOUT_MS (8 minutes), or past OPEN_TOOL_CALL_TIMEOUT_MS (24 minutes) while the delegate is still inside an unresolved tool call, since a build or a test run is legitimately silent. Because that watchdog is server-side, it keeps running whether or not a browser is open. The board’s stale TTL plus its sweep interval must stay under that 8-minute window: the sweep is what publishes task_released, and that release is what detaches a stale session before the watchdog can reach it.The no-migration-ladder model
There is no migration ladder. TheCREATE TABLE IF NOT EXISTS block in ensureSchema (packages/db/src/schemaBootstrap.ts) is the sole schema-creation source: it declares every table and column on a fresh database outright. The codebase carries no hand-written forward-only ALTERs either; the one in-place upgrade step, reconcileSchema, is derived from that same DDL, adding whatever columns an existing database is missing (see Database schema). A change that rewrites or removes an existing column is still a hard reset of the local DB. The package no longer ships db:migrate or db:generate scripts; only db:studio remains, and there is no drizzle/ directory on disk.
schema.ts is the Drizzle type layer over the same tables, used for typed queries and $inferSelect / $inferInsert types, never to apply migrations. Because nothing keeps the two descriptions in sync automatically, schemaSource.test.ts is the guard: it builds a database via the real createDb(), reads the live { table → column-name set } via PRAGMA table_info, and asserts it matches the same map derived from the schema.ts type layer (and vice versa). It also pins the posture, asserting the npm files array excludes drizzle, that no db:migrate / db:generate scripts exist, and that no migration-ladder directory is present.
schemaSource.test.ts compares column names only; column type, NOT NULL, DEFAULT, primary key, foreign key, and index drift are not checked, because the Drizzle-column → SQLite-PRAGMA mapping is lossy and would produce false drift. The test catches the drift that matters most (a column or table added to one source but not the other); deeper shape verification is deferred until a real schema change. The FTS5 virtual table and its shadow tables are excluded; they are raw DDL in schemaBootstrap.ts that Drizzle cannot model.Design rationale and trade-offs
The board is a hand-written repository, not a generic ORM facade, for the same reason the claim is a conditional UPDATE rather than a transaction-wrapped read-modify-write: the correctness properties (single-assignee, legal transitions, crash recovery) are easier to see and to test as explicit, single-purpose functions over five tables than as generated query plans. The pure state machine and the contention helpers are split out precisely so each can be reasoned about in isolation, transition legality with no database, lock handling with no business logic. SQLite with the WAL recipe is the deliberate concurrency choice: it ships in-process, needs no server, and tolerates the team-scale write contention Clawboo needs without a database to administer. The single data-access layer is the seam where a future SQLite→Postgres or multi-tenant move would land; the dormanttenant_id column on every board table is the placeholder for it.
The cost is a second persistence layer beside each runtime’s own session state, and a recovery surface (the two reconciliation passes plus the stale-sweep TTL) whose correctness rests on every claiming drain beating its task row. The board buys durability, race-freedom, and recoverability; the bill is that liveness is only as good as the beat, so a drain that claims without beating has its live work swept.
Boundaries and non-goals
- Not the agent or session registry. The board references agents, runtimes, and sessions by id (soft refs, no foreign key) and owns no agent identity, agent files, or live session state. Those belong to the registry of record and the runtime. Only the internal
parent_task_idself-reference is FK-enforced. - Not a general-purpose tracker. The statuses, transitions, and recovery passes are tuned for runtime-driven agent execution, not for a human-facing issue tracker.
- Single implicit tenant today. Every board table carries a
tenant_idcolumn, but it is a dormant seam; no per-tenant filtering is active in v0.3.1. Multi-tenant scoping is a future seam, not a shipped feature.
These docs describe Clawboo v0.3.1, the current release.
See also
- The board, the concept-level model this page implements
- The executor runner, the consumer that drives claim → run → verify → handoff
- Verification, the builder-≠-judge gate the
→ donerule enforces - Worktrees and handoff, the isolated world a task’s work happens in
- Database schema, the full table and column definitions
- Board API, the REST surface over these primitives
- Glossary, canonical term definitions