Domain model
4. Domain model
| Object | Responsibility | Lifetime |
|---|---|---|
| Task | Issue, instructions, decisions, progress, results, PR link | Until completed, cancelled or failed |
| Workspace | A folder on the host or an external SSD: its agent clone, the agents’ worktrees, the integration branch, tool config and caches (D42) | Created and removed by the human; outlives tasks, runs and environments |
| Agent | A named role in a workspace (for example docs, runtime): its worktree and branch agent/<role>, instructions, permission profile and session (D42) | Until the human removes it; works on many tasks in turn |
| Run | One execution of an agent in an environment | Start, pause, resume, terminate |
Environment (env) | Container or VM backing a workspace; it runs the runs of the workspace’s agents at once (D42). A console is an environment without an agent (D43) | Stopped or recreated independently of workspace data |
| Decision | A question, approval or review request raised to the human | Until answered |
| ReviewCandidate | Task → branch → commit SHA → PR → CI results, keyed by SHA | One per prepared revision: created when cleanup pins the SHA, before the push (§4.5) |
| Event | Append-only record of instructions, observations, decisions, actions | Permanent (audit trail and UI feed) |
4.1 State machines
Task, run and environment each get their own small FSM with explicit legal transitions; these are specified before coding.
Task (D13):
From To queuedrunning,cancelledrunningawaiting_guidance,ready_for_review,failed,cancelledawaiting_guidancerunning,failed,cancelledready_for_reviewrunning(rework),completed,cancelledcompleted,cancelled,failednone (terminal) stateDiagram-v2 [*] --> queued queued --> running queued --> cancelled running --> awaiting_guidance running --> ready_for_review running --> failed running --> cancelled awaiting_guidance --> running awaiting_guidance --> failed awaiting_guidance --> cancelled ready_for_review --> running: rework ready_for_review --> completed ready_for_review --> cancelled completed --> [*] cancelled --> [*] failed --> [*]awaiting_guidanceis for blocking Decisions raised by a run: a question or a tool approval from a live run, orauth_expiredorquota_exhaustedfor a paused one (D23). The review Decisions ofready_for_review(“Ready to push?”, §4.5) leave the task inready_for_review.- Rework (
ready_for_review → running) happens when the push is declined or the PR needs changes. It starts a new run on the same workspace and topic. completedis set when the pushed PR is merged on the forge.failedis terminal and maps to exit code 10 (§9.2). A failed run does not fail its task: it opens a blocking Decision (retry or cancel). A task fails only when a hard limit ends it (time or cost budget, a lost workspace) or the human answers that Decision with “give up”. Both happen while the task isrunningorawaiting_guidance. Inready_for_reviewa lost workspace opens a review Decision instead (rework or cancel), so no other state needs a way intofailed(D21).- A run paused by the human leaves its task
running: pause is a run state. A run paused byauth_expiredorquota_exhaustedopens a blocking Decision, which moves the task toawaiting_guidancelike any blocking Decision raised by the run (D21).
Run (issue #15):
From To startingrunning,stopped,failed,interruptedrunningpaused,stopped,failed,interruptedpausedstarting(relaunch),running(cooperative pause),stopped,interrupted(cooperative pause only)interruptedstarting(resume),stopped,failedstopped,failednone (terminal) stateDiagram-v2 [*] --> starting starting --> running starting --> stopped starting --> failed starting --> interrupted running --> paused running --> stopped running --> failed running --> interrupted paused --> starting: relaunch paused --> running: cooperative pause paused --> stopped paused --> interrupted: cooperative pause lost interrupted --> starting: resume interrupted --> stopped interrupted --> failed stopped --> [*] failed --> [*]- Pause is a run state. Pausing never changes the environment (§4.3).
- Resuming a paused run. None of the measured agents has a cooperative pause (D11), so pause is a hard interrupt: the agent process is gone while the run is paused, and resume relaunches it from the session. That is
paused → starting, like a resume frominterrupted, so a failed relaunch can end infailed.paused → runningis for an agent that reports cooperative pause, whose process stays alive. A pause survives an environment stop (issues #201, #221): once its agent is stopped, a hard-paused run has no process to lose, so it stayspausedacross an environment stop, a supervisor or host restart and a lost environment, with its open login or quota question; only the human’s resume, an answer that resumes or a due “resume at reset” moves it on, never the reconciler’s recovery. Its environment is started, if it is stopped, when the run is resumed. Only an agent with a cooperative pause, whose process stays alive, goes frompausedtointerruptedwhen that process is lost. Every resume after a pause, a cancel or an interruption starts with a briefing from the supervisor (D27). stoppedis a normal end: the agent finished, or the run was cancelled.failedis an agent crash, a failed start or a failed resume.- Into
interrupted: the reconciler (§5.3) marks astartingorrunningrun (or a cooperatively paused one, above)interruptedwhen its process or environment is gone: a supervisor or host restart, or a lost environment. Nothing else sets it, and a pending approval is raised again on resume (§4.2). - Out of
interrupted:startingwhen the reconciler resumes the agent from its session in a running environment (§4.3);stoppedwhen the task is cancelled;failedwhen resuming is impossible or its attempts are used up. - One active run per environment counts
interrupted(issues #201, #216). Until several agents at once are built (#94), an environment holds at most one run that isstarting,running,pausedorinterrupted(§4.3): an interrupted run still owns its environment, because it is resumed there. An environment whose agent stop is still under way counts as busy too, whatever state its run was saved in, until that stop and its fallback have ended (the fifth path below; #238): a cancelled or failed run is terminal once saved, yet its agent may still run. A new run’s start and every path that moves a run tostartingcheck it, the run itself excepted. - One live agent per run (issues #201, #202). Every path that moves a run to
starting, the human’s resume and the reconciler’s recovery alike, holds that run’s lock, reads the run again after taking it, and refuses when an agent session is already attached to it; it never replaces a session that is attached. The path takes the run’s session slot before it savesstartingand releases the lock once that write is saved, so from then until the launch ends the slot guards the run and a second path is refused as already attached (#225). A path that loses this race leaves the live session as it is and changes nothing else. So, within one supervisor process, cancel and pause always reach the agent that runs, and the reconciler never takes a live agent for a lost one. - No surviving agent before a relaunch (issues #201, #216). A supervisor that ends without stopping its agent (a crash, a kill, or a crash between saving a pause and stopping the agent) can leave the agent process running in the guest: ending the runtime’s exec client does not end it (spike #7, Case 4, verified on Apple Container 1.5.0), and a restarted supervisor cannot find it. So no agent is launched, a new run’s included, in an environment the current supervisor process did not start without stopping and starting that environment first. Stopping an environment is how an agent left over from an earlier process is ended; that a stop ends every process in the environment is unverified (seen in passing in the Apple Container spike, not measured; #216’s live test measures it). Five paths apply it. The first four act only on an environment the current supervisor process did not start, and the fifth only on the environment of a run whose own agent stop failed; because every launch goes through one of them, an environment is never stopped under an agent this process launched for a run that is to go on running:
- The reconciler interrupts a
startingorrunningrun that has no attached session and then stops that run’s environment. The stop interrupts everystartingorrunningrun in that environment, attached or not, as any observed stop does; a hard-paused run stayspaused(Intointerrupted, above); each resumes with the briefing (D27). Recovery then starts the environment as for any stopped one. On every pass (issues #201, #221) the reconciler first stops, before it observes anything, each running workspace environment that its store records and the current process did not start, whatever runs it holds (the console’s environment, where no agent is launched, is left alone): a paused run in it stayspaused, a lost run is interrupted as above, and a recorded environment with no active run is stopped too. The owner label is shared by every supervisor of the host, so an environment the store does not record is left alone: it may hold another supervisor’s live agent, and no agent of this supervisor was ever launched there, because every launch saves its environment first. A stop that fails is tried again on the next pass. A new process has started no environment, so its first pass ends whatever an earlier process left behind, a pause saved before a crash stopped its agent and a cancel whose stop failed included; while the supervisor is down, such an agent runs unsupervised until that pass. - Every path that moves a run to
starting, frompausedorinterrupted(the reconciler’s recovery, a quota resume that is due, the human’s resume, and an answer that resumes) first stops and starts again such an environment; when that stop fails, the run keeps its state, nothing is launched, and for aninterruptedrun the failure counts an attempt. - A new run’s start in such an environment (a task started on an agent, a retry, a rework) first stops and starts it; when that stop fails, nothing is launched and nothing is saved: the start returns the error, as an environment that cannot be brought up does before the run is saved (§5.3, Start), so a retry’s or a rework’s answer stays open.
- Cancelling a
pausedorinterruptedrun in such an environment also stops it, so a cancel reaches a leftover agent. When that stop fails, the cancel stands, the human is told that the environment was not stopped, and the reconciler’s pass stops it (above). - Any agent stop that fails (issues #221, #222, #238): a pause or a suspension (§4.2), a cancel of a
startingorrunningrun, a budget stop (§5.7, D48) or kill-all (§7, item 7) stops the run’s environment instead. The run already has the state its path gave it (paused,stoppedorfailed), which the coupling rules allow. A paused run still owns its environment; for a stopped or failed run, the environment counts as busy from before the run’s state is saved until the agent stop and this fallback have ended (One active run per environment, above; #238), so no other run is launched there in between. When that stop fails too, the run keeps that state, the human is told that its agent may still run (the error a human’s own pause or cancel returns; for a suspension, a budget stop or kill-all, which no human call waits on, and for a stop that ends after the human’s call has returned, as a cancel of a run that is stillstartingdoes, whose agent the start stops once it is attached: the eventagent.may_runon the task, which the API’s task view andwhr showshow with the run (to be built, #238), and the noticeagent_may_runin the inbox, §4.2; not thewhr servelog alone; kill-all’s result lists each task whose stops failed, and each whose agent stop is still pending in a starting run asagent_stop_pending(to be built, #238), without waiting for it, so the kill switch never waits on a start), and the current process forgets that it started the environment: it then counts as one this process did not start, so the reconciler’s next pass tries the stop again and no agent is launched there without a stop and start first. A pidfile in the guest is not proof, because the agent can change it.
- The reconciler interrupts a
- Terminal runs are never reused. A retry or rework starts a new run on the same workspace and topic (§4.1 task rules).
Environment (issue #15):
From To provisioningstopped(created),deleted(provisioning failed or abandoned)stoppedrunning,deletedrunningstoppeddeletednone (terminal) stateDiagram-v2 [*] --> provisioning provisioning --> stopped: created provisioning --> deleted: failed or abandoned stopped --> running stopped --> deleted running --> stopped deleted --> [*]- Start and stop move between
stoppedandrunning. After a service or host restart every environment isstopped(spike #2, §5.3), so the reconciler observesrunning → stoppedand starts the ones that should run. - Delete only from
stopped. A running environment is stopped first. - Recycling (§4.3) deletes an environment and provisions a new one, so it is a new Environment with a new ID. An environment is never reprovisioned.
- Start and stop move between
Coupling rules (issue #16). The three machines are not independent; a task aggregate (the task with its runs, their environments and its ReviewCandidates) checks these guards before a change is made:
- One live run. A task has at most one run that is not
stoppedorfailed, and a run is only created or started (frompausedorinterrupted) in an environment that isrunning. - A new run is new.
StartRuntakes a run that has not started (no state) and a non-empty ID that no run of the task has; a stopped or failed run is never passed back in. It is refused unless the task isqueued,runningorready_for_review(rework, §4.1); acompleted,cancelled,failedorawaiting_guidancetask starts no run. Once every guard has passed,StartRunmoves aqueuedorready_for_reviewtask torunningin the same change, so a rework run never stops under aready_for_reviewtask; a refused start moves nothing. - Pause never stops the environment. Pausing a run changes the run only. An environment is not stopped while a run in it is
startingorrunning; a run is stopped or interrupted first. ready_for_reviewneeds a stopped run and a pinned SHA. The task’s latest run isstopped(afailedrun does not qualify), and a ReviewCandidate exists whose SHA is pinned (§4.5). The current revision is the most recent ReviewCandidate, and it records the run that produced it: it must be the latest run. Only the latest run may pin a revision, so a late pin from an older run cannot become the current one. After rework, a new run that stops without pinning a new commit leaves the task not ready, so the old SHA and its old CI result never carry over.- CI belongs to a SHA. A pipeline result is recorded on the ReviewCandidate of its own commit. A pass on an earlier SHA never counts for the current revision: where CI is required,
ready_for_reviewis refused until the current SHA has passed. - Violations are conflicts. A broken coupling rule and an illegal transition are reported as a conflict, exit code 5 (§9.2). An unknown run, environment or SHA is “not found”, exit code 3.
- One live run. A task has at most one run that is not
4.2 Decision object
Fields: ID, task, run (empty for a review Decision, which no live run raised), kind (question | approval | review), blocking flag, subject, input (untrusted, capped), commit SHA, options, status, created and answered timestamps, deadline, answer, reason and answering actor (issue #17). The inbox, whr inbox, notifications and the audit trail hang off it. “Awaiting guidance” is the state a task enters while a blocking Decision raised by a live run is open; review Decisions belong to ready_for_review (§4.1).
Approvals are live and blocking. Spike #1 showed the pattern with Claude Code: the agent’s permission prompt is routed to the supervisor, which opens an approval Decision carrying the tool name and a capped copy of its input. The agent stays blocked until a human answers allow or deny, with an optional reason that is passed back to the agent. Rules:
- Fail closed. Deny on timeout (the spike used 10 minutes) and when the supervisor is unreachable.
- What is shown is what is allowed, or it says so. The input is capped at 2000 characters, and an approval whose input was cut anywhere on the way (in the adapter or the Decision) is marked
input_truncated: the inbox,whr inboxand the API say that the allow covers more than is shown, so a human can deny what they cannot read. The tool name is the subject: a prompt whose name is not a tool name (letters, digits,_ . : -, at most 128 characters) is denied without asking. At most 8 prompts of one session wait for the human at once; one more is denied without asking, so an agent cannot flood the inbox. - Plan approval. In plan mode the agent’s
ExitPlanModearrives as an approval whose subject is the plan. - Capped input. Tool inputs in a Decision are capped (the spike used 2,000 characters); the full input stays with the agent. All of it is untrusted data.
- Status.
openbecomesanswered,expired(the deadline passed) orsuperseded(the supervisor restarted or the run was paused); each is terminal. Onlyansweredwithallowever permits anything: an open, expired or superseded approval is a denial. - Approval channel transport (D26). Approvals travel over the agent’s stdio control protocol (
--permission-prompt-tool stdio,stream-jsonin and out) on the runtime’s interactive exec channel (container exec -i):control_request(can_use_tool) on stdout,control_response(allow or deny) on stdin, matched byrequest_id. An answer for an unknownrequest_idis denied, and open requests are denied when they are superseded (D23). A closed stdin fails closed without running the tool, and a crashed supervisor end or an unanswered request leaves the guest agent waiting with the tool unrun (verified, spike #7), so when the channel is lost the supervisor stops the agent withwhr-shim(D25) and records the request as denied. No guest-to-host network path, no listener and no approval token in the guest. - Deadline. An approval always has one (default 10 minutes). An answer that arrives after it is refused and the Decision expires, so a late “allow” does not count. An approval that has none (a lost value, a damaged row) counts as past it: an answer expires it, and
Allowsis false. The store refuses to write such an approval and refuses to load one. Only a question or a review Decision may have no deadline. - A refusal can still change the Decision. A late answer expires the Decision, and an allow for another commit is recorded as a denial, while both return an error. The store’s
RespondDecisionloads, answers and saves in one transaction and saves the change even when the answer is refused, so a refusal is never lost to a rollback. The caller still receives the error (exit code 5). - Bad input is a usage error. A missing actor or time, an answer that is not offered and a malformed new Decision are not conflicts: they exit with code 2 (§9.2). A conflict (code 5) depends on the Decision’s state; a usage error does not.
- Commit SHA. A review Decision such as “Ready to push?” carries the pinned SHA of its ReviewCandidate (§4.5, §6). An allow is tied to that SHA: an allow given for a different SHA is recorded as a denial, and the supervisor asks
Allows(sha)against the commit it is about to push, so an approval for an earlier revision never covers a later one. - Pause and restart (D23). A Decision raised by a live run (a question or an approval) is
supersededwhen the supervisor restarts or the run is paused: pause is a hard interrupt (D11), so the agent process that asked is gone and could not receive the answer. The reconciler, or the resume, raises a new one when the agent asks again (§5.3). An answer sent to a superseded Decision is refused as a conflict (§9.2, exit code 5). A review Decision belongs to no run, so it survives both, and so do the login and quota questions below: they are raised for a run that is already paused, and nothing waits on them. This is the one restart rule; §4.1 and §5.3 refer to it. - Expired auth and quota (D23).
auth_expiredandquota_exhaustedare blockingquestionDecisions raised for the paused run, not approvals, because nothing is being permitted. Their options are fixed: forauth_expired, “Signed in again, resume” or “Cancel”; forquota_exhausted, “Resume now”, “Resume at reset” (with the reset time when the agent reports it) or “Cancel”. The run stays paused and the task moves toawaiting_guidance(§4.1, D21). Suspending is a pause (issue #222): when the agent reports either while it runs, the run is savedpausedwith its question first, then its agent process is stopped as for the human’s pause (D11, D25), so nothing runs while the question waits; a stop that fails is handled as for the human’s pause (§4.1, “No surviving agent before a relaunch”). They have no deadline: waiting is safe, because nothing runs until they are answered. - Rebase conflict (D42). A conflict found when the supervisor rebases an agent’s branch onto the integration branch, before the export (§4.5), is a blocking
questionDecision raised for the stopped run whose branch it is, because no live run can be asked and the run ended there. The rebase is aborted inside the environment first, so the branch is as the agent left it. The options are fixed: “Rework” (a new run on the same agent, whose briefing names the conflicting paths, D27), “Resolved in the console, retry” (the human fixed it in the console, D43, and the supervisor rebases again) and “Cancel”. The input lists the conflicting paths as untrusted data. A rebase outside a task, such as one asked for by the human, returns the conflict to the caller and raises no Decision. - Refused prepare (D51, #253). A prepare that is refused for any other reason (a failed check, a refused commit message, a tree or topic cap, a rewritten history, no check configured, a forge error) is a blocking
questionDecision raised for the stopped run, causeprepare_failed, for the same reason as a rebase conflict. The options are fixed: “Rework” (a new run on the same agent, whose briefing names the reason, D27), “Retry” (prepare again, for a fault outside the agent’s work) and “Cancel”. The input is the reason and the check’s capped output, as untrusted data. - Refused publish (D51, #253). A publish step that the Guard or the forge refuses for a reason a retry cannot change (§4.5, “An outstanding publish”) is a blocking
questionDecision for the task, causepublish_failed. The options are fixed: “Retry” (complete the publish again under the same approval of the same SHA, for a fault fixed outsidewhr; the Guard checks everything again), “Rework” (a new run on the same agent, whose briefing names the reason, D27) and “Cancel”. The input is the refusal, as untrusted data. It belongs to no run, so, like the review Decision, it leaves the task inready_for_review. - Agent may still run (issue #238; the inbox entry and its push to be built, #238). When an agent stop and the environment stop after it both fail on a path no human call waits on (a suspension, a budget stop, kill-all, a stop that ends after the human’s call returned; §4.1, fifth path), the supervisor raises a non-blocking
questionDecision for that run, causeagent_may_run, beside the task event, after the task’s state is saved: the inbox lists it with the task, a terminal task included, and nothing supersedes it (not the cancel kill-all makes next, not a restart, not the task’s end), so only the human’s answer closes it; and it is pushed (§9.4), because an agent that may run unwatched is what the human must hear about without looking. Its input names the path and the error, as untrusted data; its one option is “Seen”, and answering it records who saw it and changes nothing else. Being non-blocking, it leaves the task’s state alone. The inbox gets no other entry kind: everything it lists is a Decision. - Egress requests (D38). A host that a repository’s
devcontainer.jsonrequests or a lockfile suggests is opened only by the human’s allow, stored per repository on the supervisor’s side; nothing in the repository grants it. The question is asked once, at the first start of a run that would use the environment: the run is created and staysstarting, and one blockingapprovalDecision per unanswered host is raised for it, with causeegress_requestand the host in its own field (a host name the reader validated, never parsed from the subject or input, which only describe it and its source). The task moves toawaiting_guidanceas for any blocking approval. Allow adds the host to the repository’s allowlist and deny records it as refused; either answer is kept, so later runs and other workspaces of the repository are not asked again. The deadline and the restart rule are an approval’s: an expired or superseded request is a denial for this run only and is asked again at the next start. The environment exists already (it is made with the workspace, D42) and starts with the hosts already allowed. The agent process is started only once every request of the run is answered; if an answer allowed a new host, the supervisor first recreates the environment’s proxy sidecar with the new allowlist, which is safe because no agent process runs yet. A denied host leaves the run to start without it. The allowlist never changes while an agent process runs. A host added to the repository later is asked at the next start. Changing an answer is a policy change (§6), outside the run. - Feature sources (D38, issue #127). A devcontainer feature from a source outside
ghcr.io/devcontainers/features/is asked like an egress host: one blocking approval per reference, causefeature_source, the reference in its own field, the answer kept per repository and answered only with a passkey step-up on the web (D45). An allow takes effect when the environment is next built: a new workspace, orwhr ws rebuild <workspace>, which the human runs while no agent of that workspace runs and which keeps the home and build volumes (issue #128); nothing rebuilds by itself. - Card queue (D30, D40, issue #71). A card the human moves to an agent queue on the board raises a blocking question, cause
board_queue, “Accept this task?” with the options start and cancel; it never starts a run by itself, once per card state. The project board does not say who moved a card, so the mover is treated as untrusted and only the issue author’s trust tier is applied; the human’s answer is what starts the run. - Task state (D13, D23). Only a blocking Decision raised by a run moves its task to
awaiting_guidance: a question or approval from a live run, a login or quota question for a paused one, or a rebase conflict or a refused prepare for the stopped run whose branch it is. A review Decision belongs to no run and leaves the task inready_for_review.
4.3 Lifecycle rules
- IDE/SSH disconnect does not pause the agent or stop the environment.
- Detached sessions. The agent process runs inside the environment under the supervisor with no client attached. Chat is attach and detach over a persistent session, never the owner of the process. Session state (transcript and agent session ID) lives in the workspace, so it survives a supervisor restart and a reattach from another client. The reconciler (§5.3) resumes from that session after a restart.
- A long run can outlive its credentials. Expired auth or an exhausted usage window pauses the run and opens a blocking Decision (§5.2) instead of failing it.
- Pausing lets a human inspect or edit without concurrent agent changes. Human takeover holds an explicit, visible workspace lock (agent vs human); on resume the agent is re-synced with the human’s changes.
- Stopping an environment preserves files, logs and agent session state.
- Agent-level checkpoint/resume is distinct from VM suspend and must work on backends without suspend.
- Environments awaiting review may be stopped.
- Worker recycling (recreate the environment, keep the workspace) is a first-class lifecycle action because freed guest memory is not returned to macOS.
- Only validated runner × backend pairs may claim resumability.
- Workspaces and agents (D42, issue #90). A
Workspacehas a unique name (lowercase letters, digits and-, 1 to 40 characters, starting with a letter), a host folder, the repository it was seeded from (owner/name), its integration branch (any branch name the configuration accepts: a valid ref name outsideagent/, never chosen by an agent or an issue; issue #229). The agents’ namespace is compared without case (issue #254): a name whose first path component isagentin any case (agent,Agent/x,AGENT/main) is refused by every check that keeps a branch out of the namespace (the configuration, the workspace, the forge Guard’starget(), andhostgit’s fetch and prepare of a target, which also catch a record stored before this rule; to be built, #262), because the supervisor’s repositories live on a case-insensitive disk whererefs/heads/Agent/xandrefs/heads/agent/xare one file, and git cannot holdrefs/heads/agentbesiderefs/heads/agent/<role>. The checks that admit only an agent’s branch point the other way and stay an exact, lowercaseagent/match (the Guard’sagentBranchfor a push,hostgit’sPush), so a case variant is never accepted as an agent’s branch. A workspace also has its environment. AnAgenthas a role (the same character rules, at most 30 characters), unique within its workspace, and gets the branchagent/<role>and the worktree/ws/wt/<role>inside the environment from it; neither is chosen freely, so a name from a repository’s issue or an agent’s own output never becomes a path or a ref by itself. A task is assigned to one agent (Task.AgentID) and its runs carry the same agent. The model allows several agents per environment, but the service allows one active run (starting,running,pausedorinterrupted, §4.1) per environment and refuses a second with the conflictenvironment busyuntil several agents at once are built (#94). Workspaces and agents are registry records the human creates and removes; they are not part of a task’s aggregate and are not changed by task events, so each change is recorded as an audit event in the streamworkspace:<id>(the event log is keyed by a task or by such a stream). Creating a workspace checks the folder (empty, below a workspace root, not inside a git repository; §5.6), seeds the agent clone withhostgit(the only time the host runs git there), records the workspace, provisions and starts its environment with the folder mounted read-write at/ws, and adds the first agent:git worktree addruns inside the environment, never on the host. A failed step takes back what the earlier ones made. Starting a task on an agent records the task and its run and checks the one-run rule in one transaction, so two starts cannot both pass; it then starts the agent in its worktree. If the agent cannot be started the run fails into a retry-or-cancel Decision and the environment is free again. One environment serves its workspace’s tasks one after the other, so an environment ID can be in several tasks. Rebasing an agent’s branch onto the integration branch runsgit rebasein the agent’s worktree inside the environment, and is refused while the agent has a starting or running run. A conflict aborts the rebase and is returned with git’s report as the conflictrebase-conflict; the export (#91) turns it into a Decision on the task. The repository cache and per-task clones ofhostgitstay until that export replaces the fetch from a stopped environment (§4.5).
4.4 Persistence semantics
Decided from spike #2 (Apple Container, issue #2; confirm on other backends):
- Workspaces live on the host or an external SSD (D42), outside containers and volumes, and are mounted into their environment read-write. A workspace holds its own agent clone and the agents’ worktrees; the human’s own repositories are never mounted. Workspaces survive any environment, and one environment serves all agents of a workspace. This replaces per-task clones of a repository cache (D17); the agent clone stays in the workspace folder, and dependency directories follow D39 per worktree. Whether the agents’ worktrees move onto a volume is open (D39, issue #119); if they do, it is one volume per workspace environment with a directory per agent, kept when the environment is recreated.
- The agent home (auth directory, session, caches) is one writable named volume per environment. A volume survives stop, start, delete and rebuild.
- A writable volume is exclusive. While one container has it read-write, no other container can attach it, not even read-only (the second start fails with “The storage device attachment is invalid”). A rebuild stops the old container before the new one starts. A volume can be shared read-only by several containers.
- A bind-mounted checkout is slower but usable. For 5000 files with 300 edits, the first
git statustook 0.3 to 1.1 s on a bind mount against 0.12 s in a volume, later runs about 100 to 130 ms against 75 to 118 ms,git add363 ms against 120 ms, andcommit667 ms against 334 ms. Git tuning (untracked cache,feature.manyFiles, preloaded index) changed steady-statestatusby almost nothing, because the cost is the cold first stat of every file. Very large repositories were not measured; the cold pass grows with the file count. - Dependency directories stay off the bind mount (D39). Spike #40 measured file-metadata work on a bind mount at 20 to 100 times the cost on a volume for large trees (153 000 files: warm
git status8 s against 0.1 s, copying the tree 128 s against 1.5 s; 40 000 small files of anode_modulesshape: 1.8 s against 0.04 s), while a CPU-bound build took the same time on both. So the untracked trees that hold most of a repository’s files are moved to a per-environment volume where the tool reads its location from the environment (CARGO_TARGET_DIR,UV_PROJECT_ENVIRONMENT), and covered by a volume mounted over the checkout only where it cannot be moved (node_modules). A mount point inside the bind mount cannot be removed or renamed in the guest (EBUSY), shows the volume’slost+foundand hides host files at that path (spikevolume-over-bind), which is why relocation comes first. The checkout itself stays on the host. Results: spike #40. - Caches and build output stay on the volume, outside the checkout. A real build (this repository’s
make check, including golangci-lint and editorconfig-checker compiled from source, 4 CPUs) took 45 s cold and 2.3 s warm with everything on a volume; with only the checkout on a bind mount it was 41 s and 2.7 s, no measurable cost; with the Go caches (about 25,700 module files and 7,800 build-cache files) on the bind mount as well it was 56 s and 4.7 s, about twice as slow warm. The adapter setsGOMODCACHE,GOCACHEand the package-manager and build-directory equivalents to paths on the environment’s volume. A tree with many small files of its own, such asnode_modules, will behave like the cache case and was not measured. - Bind mounts sync both ways at once, and files the guest root writes appear owned by the host user.
- The root filesystem is disposable. It survives stop and start and is lost on delete, so nothing that matters lives there.
- Mounts are rejected by the adapter, not the runtime. The runtime accepts any host path. The adapter calls
runtime.CheckMountbefore every mount: it resolves symlinks first, then rejects$HOMEand its parents,~/.ssh, other secrets directories and runtime sockets (§7.4), and the conformance suite checks it. - Tools are not part of the workspace. Agent CLIs come from a shared read-only store (§5.6).
Still open: retention and garbage collection of completed tasks, topics and workspaces on the host disk (§8), quotas (a named volume is a sparse image with a virtual size of 512 GiB), and repositories that track more than about 50 000 files, whose git status stays slow on a bind mount even with dependencies on a volume (D39, §12), and with them whether the agents’ worktrees move onto a volume (issue #119, waiting on its measurement).
4.5 Topics, checkouts and cleanup before push
A topic is one line of work: one branch (agent/<topic>) with its own checkout on the host. Several topics are in flight at once, each with its own environment and agent, and the developer can open any checkout in their editor at the same time. That is the worktree idea: parallel topics that are merged and tidied locally before anything leaves the machine.
Since D42 a topic is an agent’s branch in a workspace: each named agent works in its own worktree of the workspace’s agent clone, rebases onto the integration branch, and its commits leave as a git bundle exported from the running environment (issue #91), not through a fetch from a stopped environment. The import is a fetch of the bundle into a supervisor-owned repository that lacks its objects: git bundle verify alone passed a bundle truncated to half its size, while the fetch refused it and left no ref (spike #89). The per-task clone of D17 is gone (issue #95); its cleanup and push rules (prepare, per-SHA approval, push, no rewriting of pushed commits) carry over unchanged.
The export as built (issue #91). Request.Agent names the agent whose branch is exported. The supervisor’s forge mirror (the repository cache of D17 without its per-task clones, D42) stays: a bare repository per forge repository, fetched only from the forge, never from a workspace and never mounted into an environment. Prepare first refreshes the target from it into the supervisor’s repository, because the bundle’s prerequisite is the branch’s merge base with the integration branch. For a first round it then runs git rebase in the agent’s worktree inside the environment; a conflict aborts the rebase and raises the blocking question of §4.2 for the stopped run, with the conflicting paths as untrusted input (a later round is a fast-forward and is not rebased: pushed commits are never rewritten). Then git merge-base and git bundle create - <base>..<branch> run in the guest, through the same exec path (and its launcher) as the agent, under a timeout (default 5 minutes). Their output is streamed into Repo.ImportBundle, which caps it (default 512 MiB), fetches it into a scratch repository that borrows only the supervisor’s objects, checks that the branch arrived whole and only then fetches it into the supervisor’s repository; the scratch is removed on every path. The environment keeps running, and no git runs in the workspace on the host. A rewrite of commits already pushed is refused as before, by hostgit.Prepare against the revision’s recorded source. OpenCopy exports the same way, so the editor copy is fresh while the environment runs, and is cloned from the supervisor’s repository, never from the workspace.
Checkout layout (spike #2, item 9):
- Each workspace has its own full agent clone (D42), made once by
SeedAgentClonefrom the forge or from a repository path: an independent clone with no alternates, no hardlinks, no template and an origin that is a plain https URL, never a host path. The agent works in a worktree of it inside the environment. The supervisor keeps no per-task clones and mounts no object cache into an environment. - Clone depth is a per-repository setting,
clone_depth, default 0 (full history), of the forge mirror. A positive value fetches the mirror with--depthfor very large repositories; refreshing it then needs a forced refspec and a target is taken withfetch --update-shallow; once the target moves further than the depth past a branch’s fork point there is no merge base. The supervisor therefore checks for a merge base before a rebase and deepens the mirror and its own copy (--deepen) until it finds one. A partial clone (--filter=blob:none) keeps the full history and may be the better option; it was not tested. - Worktrees of one repository (
git worktree) are the developer’s tool on the host for their own parallel topics. They are not handed to agents. A worktree’s.gitfile points at the shared repository by host path: mounting only the worktree fails (not a git repository), and mounting the shared repository exposes every branch, the shared hooks and config, and the other worktrees’ metadata.git worktree add --relative-pathsmakes the pointer work if the mounts keep the relative layout, but the exposure is the same.
Cleanup before push:
- Agents do not push. The agent commits in its topic’s checkout. Nothing leaves the host until a cleanup step has run and been approved.
- Prepare for push. The supervisor, on its own copy of the topic (
hostgit, below), rebases it onto its target, folds attempts into one commit per finished change (fixup and autosquash of unpushed commits only), checks the commit messages (conventional commits and the repository’s commit linter), signs the commits with the bot key (§7.7) and runs the repository’s checks. Several finished topics can be merged into one integration branch first. - Approval. The result is a ReviewCandidate (a pinned commit SHA) and a Decision, “Ready to push?”, that shows the commit list and the diff stat. Approval is per commit SHA, as for any review (§6).
- Push and PR. On approval the supervisor pushes the prepared
agent/*branch with a token minted for that push alone (D51, §7.3) and opens or updates the PR. Merging stays human, on the forge. - Pushed commits are never rewritten. After a push a topic is only extended; rewriting pushed history needs an explicit request. Each revision records the agent’s own tip it was prepared from; a follow-up round rebases only the agent’s newer commits onto the pushed revision, so the next push is a fast-forward. An agent that rewrote commits it had already handed in is refused rather than guessed at. Publishing needs the task in
ready_for_reviewand an approval of its current revision, so a cancelled task or an approval of an earlier revision publishes nothing, and the repository’s checks are required. - Done or cancelled topics. The checkout and its branch are removed by the retention rules (§4.4, issue #54) after the push is merged or the topic is cancelled. Unpushed work is kept until the owner discards it.
Prepare and push in code (issue #27). The steps above run in this order, on the supervisor’s own copy: the target comes from the mirror, the agent’s branch is imported from a bundle (above), EnsureMergeBase finds a base, and Prepare rebases the branch onto the target in a temporary worktree with --autosquash (hooks off, the bot as committer, commits signed with the bot key), then lints every commit message and returns the new tip. A topic with no commit after its target is an error, never a revision pinned for review with 0 commits (hostgit.ErrNoCommits, issue #254), so a tip nothing signed or linted is never offered for approval. Running the repository’s own checks is not done by hostgit: they are the repository’s code, so a Checker runs them in an environment, never on the host. The tip is pinned (PinRevision) and a review Decision “Ready to push?” is raised for that SHA. Push sends one agent/* branch, fast-forward only, as the exact approved commit (<sha>:refs/heads/<branch>), never with force and never tags.
sequenceDiagram
participant A as Agent (environment)
participant S as Supervisor (hostgit)
participant C as Checker (environment)
participant D as Developer
participant F as Forge
A->>A: commits on agent/<topic>
A-->>S: the run stops
S->>A: git bundle create - (the environment keeps running)
S->>S: ImportBundle into its own copy
S->>S: EnsureMergeBase with the cached target
S->>S: Prepare: rebase --autosquash, sign, lint messages
S->>C: run the repository's checks
C-->>S: result
S->>S: PinRevision (ReviewCandidate for the SHA)
S->>D: Decision "Ready to push?" (commits, diff stat)
alt approved for this SHA
D-->>S: allow
S->>F: push <sha>:refs/heads/agent/<topic>, fast-forward only
S->>F: open or update the PR
F-->>D: review and merge, by the human
else declined, or the PR needs changes
D-->>S: deny
S->>A: rework: a new run on the same topic
end
Publishing in production (D51, issue #244; built: the check and its configuration in #251, the push credential in #252, the trigger, the answer, the completion and the wiring in whr serve in #253). The parts above are joined into one path:
- Trigger. When a task’s latest run ends
stoppedbecause its agent finished (not a cancel), the service prepares the topic: the export,Prepare, the check and the pin, then “Ready to push?” for the pinned SHA. A restart in between is repeated: the reconciler prepares again arunningtask whose latest run isstopped, with no revision pinned for that run and no open Decision; preparing is repeatable, because it starts again from the bundle. From the moment the run is savedstoppedfor a prepare until the prepare has pinned a revision or been refused, the environment and the agent count as busy: a run start in that environment and a new task on that agent are refused withenvironment busy(One active run per environment, §4.1), so no other run’s commits reach the export or run beside the check; the reconciler’s repeat holds them the same way. - The check. The command is the repository’s
checkin the supervisor’s configuration, elsecustomizations.workharbor.checkread at the default branch’s commit (D38), elsepre-commit run --all-fileswhen the default branch has.pre-commit-config.yaml. With none of them the prepare fails closed (ErrNoChecks): nothing is published without a check. TheCheckerruns it in the task’s workspace environment, started first if it is stopped, as the agent’s user, in a check worktree detached at the prepared SHA that is no agent’s worktree and is removed afterwards from inside the guest. The prepared commits reach the guest as a bundle from the supervisor’s repository, with prerequisites the agent’s clone holds, streamed through the exec’s stdin; the host writes nothing into the workspace and runs no git there. The check gets no credential ofwhr’s (no forge token, no API key); a subscription login in the environment stays readable to it, as it is to the agent (accepted, §7.3). It has only the environment’s egress allowlist, a timeout (default 30 minutes) and an output capped to its last 64 KiB, which is untrusted input (§7.1). The check counts as running, and the environment as busy, until its process group is gone: after every check, a normal exit included, the checker ends the check’s process group (D25) and looks for it, and a process still there stops the environment as on the fifth path of §4.1 (under the start lock; when that stop fails too, the process forgets that it started the environment), with the busy mark held until that has ended. The check leaves a receipt (issue #259, built in #253): an event on the task bound to the prepared SHA, naming the command and where it came from (the configuration,devcontainer.jsonorpre-commit), the exit status, the duration and the capped output, which “Ready to push?”,prepare_failedandwhr showshow, so the human sees which SHA was checked, by what and with what result. A process that leaves the group (setsid, a double fork) escapes, as it does for an agent: accepted, because it stays inside the workspace’s one trust domain (D42) and the environment’s egress allowlist. Exit status 0 passes. The check is a quality gate, not a security control: the agent writes the tree it checks, including what the command runs, so the per-SHA approval, the forge’s CI and its ruleset stay the controls. - A refused prepare (a failed check, a refused commit message, a cap, a rewritten history, no check configured, a forge error) raises the blocking question
prepare_failedfor the stopped run (§4.2). - Identity and lint. The committer is the GitHub App’s bot identity, its name (
<slug>[bot]) and noreply address read from the App (D15), never typed into the configuration; commits are signed with the bot key (§7.7). The host lints messages with a built-in linter the repository’s configuration names (commit_lint:conventional, the default, orworkharbor, the rules of this repository’sinternal/commitlint); a repository’s own linter configuration is repository code and runs only as part of its check, in the environment. - Answering “Ready to push?” is one service call, whichever way it comes: the JSON API’s
POST /v1/decisions/{id}/answeron the host socket,whr approve <decision> --sha <commit>over the same API, or the web UI only after a passkey step-up naming the Decision and its SHA (§7.5, D45). It records the answer first; an allow then starts the publish, a deny publishes nothing. No handler and no CLI command callsPublishitself. - An outstanding publish is completed, never repeated. An allowed review Decision of a task that is still
ready_for_review, whose SHA is the current revision and is not recorded as pushed, and which has no openpublish_failedquestion, is an outstanding publish; answering that question “Retry” makes it outstanding again, while “Rework” and “Cancel” move the task out ofready_for_review. The reconciler’s pass completes it, so a restart between the answer, the push and the record of the push loses nothing. Every forge step is idempotent for the approved SHA: pushing a branch that is already at it succeeds, a fast-forward whose target already points at it succeeds, and opening a PR first looks up the open PR whose head is the repository’s own branch (<owner>:agent/<role>, never a fork’s branch of the same name) and whose base is the expected target, and updates that one, so a retry never opens a second PR. Completing it needs no new approval, because the human allowed exactly that SHA, and the Guard checks it again at every step. A step that fails for a transport fault (a network error, a timeout, a server error, a rate limit) leaves the publish outstanding, is retried on a later pass with backoff, and each failure is an event on the task thatwhr showshows. A refusal that a retry cannot change (not a fast-forward, a target the Guard refuses, no approval for the SHA, a forbidden request or a missing installation) ends the outstanding publish with the blocking questionpublish_failed(§4.2), and so does any failure not known to be a transport fault: an unknown error fails closed, and the human’s “Retry” is the way back. Git reports a failed push only as text, sohostgittells a transport fault by markers in git’s own messages under the C locale; a marker found in a line the server sends (remote:), or transport faults that go on for 24 hours after the first failed attempt (DefaultPublishMaxAge, counted from the task’spublish.attemptevents, so a restart keeps it), end the publish the same way (#270). A cancelled task, or a newer revision, ends it.
The editor copy (issue #59). The developer’s editor is never pointed at an agent’s checkout: opening it runs planted repository config on the host (VS Code and its git extension, direnv, .vscode/tasks.json; threat model T15). EditorCopy clones the supervisor’s own bare copy of the topic (which the bundle import filled) into a directory outside the workspace root, with an empty template and none of the agent’s config, hooks or alternates, and refreshes it later with a fast-forward only, so the developer’s edits are never overwritten. It refuses a destination inside the workspace root, so the agent’s checkout cannot be handed out by mistake. The commits themselves are the agent’s work and may contain files an editor acts on (.vscode/tasks.json, .envrc, .devcontainer), so the copy lists them and the UI warns before the folder is trusted.
The forge contract (§10) separates Git transport from the forge API. The adapter is wrapped in a Guard that holds the autonomy table: OpenPR and UpdatePR take an Approval (the Decision and the commit SHA) and are refused unless the Decision allows that SHA and the branch at the forge points at it; Push needs the same proof and is the only caller of the adapter’s Pusher, which holds the push credential (D51, §7.3); Merge, Tag, Release and Deploy are refused whatever the table says (the table’s ceilings), and the inner adapter is never called for them. An Issue carries the author association for trust tiers (#53), and VerifyWebhook takes the request’s http.Header.
The host treats every agent-writable checkout as hostile (§7.4), and since D42 runs no git in one at all: an agent’s commits arrive as a bundle read as a stream, never as a fetch from its .git.
hostgit (issue #19) is the only way the host runs git on anything an agent can write:
- No git in a workspace on the host. The only time the host runs git in a workspace’s clone is when it does not yet exist:
SeedAgentClonemakes it. After that the agent works inside the environment, and its commits leave as a bundle (ImportBundle): read as a stream with a size cap, fetched into a scratch repository that borrows only the supervisor’s objects, checked as a whole (git bundle verifyalone passed a bundle cut in half, spike #89) and only then fetched into the supervisor’s repository, with object checks (fsckObjects) on. Cleanup (rebase, fold, sign) and push run only on that supervisor-owned repository. An earlier design let the host read an agent’s checkout with an allowlist of read-only plumbing after a check of its.git; that code is removed. - The expanded tree is capped before any checkout. Only a bundle’s compressed size was capped, and a few kilobytes of tree objects can name billions of files, or billions of trees with no file at all (a tree that lists the same subtree twice, nested).
Repo.CheckTreestreamsgit ls-tree -r -t -l -zof a commit, counting trees as well as files, and stops git as soon as it passesMaxTreeEntries(200,000 entries) orMaxTreeBytes(2 GiB of blobs).Preparereplays every commit of the topic into a host worktree, soCheckCommitschecks the tree of each commit after the target (or after the pushed source in a follow-up round), once per tree id, for at mostMaxTopicCommits(2,000) commits: a bomb that a later commit deletes is refused too.--autosquashbuilds a tree no single commit has (the union of what a chain of fixups added), soCheckCommitsalso caps what the topic adds as a whole: the entries and blob bytes each commit introduces over its parent, summed, atMaxTopicEntries(400,000) andMaxTopicBytes(4 GiB). The branch is resolved to a commit ID once, and the check and the checkout use that ID, so a branch that moves in between cannot slip past. The check,PrepareandEditorCopyrun under hard deadlines (CheckTimeout,PrepareTimeout,EditorCopyTimeout). A refusal isErrTreeTooLarge. Tests build a 2^30-file bomb from 31 objects, a 2^40-tree bomb with no file from 41, and a bomb in one commit deleted by the next, and require a quick refusal. - A hardened floor on every host git command. Git starts with an empty environment (no
GIT_*from the host, a minimalPATH, an emptyHOME, no system or global config, no terminal prompt, no optional locks, no replace objects) and-coverrides that disable hooks (core.hooksPath), the file-system monitor, the pager and editor, credential helpers, signing, submodule recursion, and every transport exceptfilewhere a fetch from a local repository needs it. - The forge mirror (issue #45, D42).
Cacheis the bare repository per forge repository, with the repository’sclone_depth(default 0).Refreshfetches one branch from the forge with the forced refspec+refs/heads/<b>:refs/heads/<b>, no tags, objects checked, and--depthfor a positive depth (an existing shallow mirror is made complete with--unshallowwhen the depth is 0). The source is an absolute path or anhttps://URL without credentials; every other scheme (ssh,ext::,file://) and anything that starts with a dash is refused, and an authenticated fetch comes with the forge adapter (issue #27). It is fed only from the forge, never from a workspace, and never mounted into an environment. - One writer per cache (issue #67). Every operation on a cache (refresh, deepen, clone, merge-base search, maintenance) runs under a per-cache lock: a mutex shared by every handle of the same path and a file lock (
whr.lock) for other processes, waited for with the caller’s context. Git’s own lock files (shallow.lock,packed-refs.lock, ref locks) are never left to block the next refresh: while the cache lock is held no other supervisor process runs git on it, so a lock file found then is stale and is removed first. Three parallel shallow refreshes had failed 40 times in 60 onshallow.lockwithout it. - Names map to directories by a strict rule.
CachePath(repo)accepts onlyowner/namemade of ASCII letters, digits,.,_and-(each part starting with a letter or digit, at most 100 characters), lowercases it, and appends a short hash, soFoo/xandfoo/xare one cache on a case-insensitive disk and no name, path separator or Unicode form can leave the cache root. A cache path must be absolute and inside the cache root when one is set (WithCacheRoot), a relative path is refused instead of being created under/, and a local source inside a workspace root (agent-writable) is refused. - Merge base before a rebase.
EnsureMergeBaselooks for a merge base between the target and the topic in the supervisor’s copy. While there is none and either side is shallow it deepens the cache (fetch --deepen) and then the copy (fetch --update-shallow --deepen) by a step, up to a limit, and reportsErrNoMergeBasewhen the limit is reached or both sides are complete, so the caller opens a Decision instead of rebasing blindly. - Tested. Tests plant a hook,
core.fsmonitorand a clean filter in a workspace’s.git; the bundle import never runs them, and a control run of host git in the clone shows each one fires (spike #89). Truncated, corrupted, oversize and garbage bundles are refused and leave no ref.