Skip to content

Domain model

4. Domain model

ObjectResponsibilityLifetime
TaskIssue, instructions, decisions, progress, results, PR linkUntil completed, cancelled or failed
WorkspaceA folder on the host or an external SSD: its agent clone, the agents’ worktrees, the integration branch, tool config and caches (D42)Created and removed by the human; outlives tasks, runs and environments
AgentA named role in a workspace (for example docs, runtime): its worktree and branch agent/<role>, instructions, permission profile and session (D42)Until the human removes it; works on many tasks in turn
RunOne execution of an agent in an environmentStart, pause, resume, terminate
Environment (env)Container or VM backing a workspace; it runs the runs of the workspace’s agents at once (D42). A console is an environment without an agent (D43)Stopped or recreated independently of workspace data
DecisionA question, approval or review request raised to the humanUntil answered
ReviewCandidateTask → branch → commit SHA → PR → CI results, keyed by SHAOne per prepared revision: created when cleanup pins the SHA, before the push (§4.5)
EventAppend-only record of instructions, observations, decisions, actionsPermanent (audit trail and UI feed)

4.1 State machines

Task, run and environment each get their own small FSM with explicit legal transitions; these are specified before coding.

  • Task (D13):

    FromTo
    queuedrunning, cancelled
    runningawaiting_guidance, ready_for_review, failed, cancelled
    awaiting_guidancerunning, failed, cancelled
    ready_for_reviewrunning (rework), completed, cancelled
    completed, cancelled, failednone (terminal)
        stateDiagram-v2
        [*] --> queued
        queued --> running
        queued --> cancelled
        running --> awaiting_guidance
        running --> ready_for_review
        running --> failed
        running --> cancelled
        awaiting_guidance --> running
        awaiting_guidance --> failed
        awaiting_guidance --> cancelled
        ready_for_review --> running: rework
        ready_for_review --> completed
        ready_for_review --> cancelled
        completed --> [*]
        cancelled --> [*]
        failed --> [*]
      
    • awaiting_guidance is for blocking Decisions raised by a run: a question or a tool approval from a live run, or auth_expired or quota_exhausted for a paused one (D23). The review Decisions of ready_for_review (“Ready to push?”, §4.5) leave the task in ready_for_review.
    • Rework (ready_for_review → running) happens when the push is declined or the PR needs changes. It starts a new run on the same workspace and topic.
    • completed is set when the pushed PR is merged on the forge.
    • failed is terminal and maps to exit code 10 (§9.2). A failed run does not fail its task: it opens a blocking Decision (retry or cancel). A task fails only when a hard limit ends it (time or cost budget, a lost workspace) or the human answers that Decision with “give up”. Both happen while the task is running or awaiting_guidance. In ready_for_review a lost workspace opens a review Decision instead (rework or cancel), so no other state needs a way into failed (D21).
    • A run paused by the human leaves its task running: pause is a run state. A run paused by auth_expired or quota_exhausted opens a blocking Decision, which moves the task to awaiting_guidance like any blocking Decision raised by the run (D21).
  • Run (issue #15):

    FromTo
    startingrunning, stopped, failed, interrupted
    runningpaused, stopped, failed, interrupted
    pausedstarting (relaunch), running (cooperative pause), stopped, interrupted (cooperative pause only)
    interruptedstarting (resume), stopped, failed
    stopped, failednone (terminal)
        stateDiagram-v2
        [*] --> starting
        starting --> running
        starting --> stopped
        starting --> failed
        starting --> interrupted
        running --> paused
        running --> stopped
        running --> failed
        running --> interrupted
        paused --> starting: relaunch
        paused --> running: cooperative pause
        paused --> stopped
        paused --> interrupted: cooperative pause lost
        interrupted --> starting: resume
        interrupted --> stopped
        interrupted --> failed
        stopped --> [*]
        failed --> [*]
      
    • Pause is a run state. Pausing never changes the environment (§4.3).
    • Resuming a paused run. None of the measured agents has a cooperative pause (D11), so pause is a hard interrupt: the agent process is gone while the run is paused, and resume relaunches it from the session. That is paused → starting, like a resume from interrupted, so a failed relaunch can end in failed. paused → running is for an agent that reports cooperative pause, whose process stays alive. A pause survives an environment stop (issues #201, #221): once its agent is stopped, a hard-paused run has no process to lose, so it stays paused across an environment stop, a supervisor or host restart and a lost environment, with its open login or quota question; only the human’s resume, an answer that resumes or a due “resume at reset” moves it on, never the reconciler’s recovery. Its environment is started, if it is stopped, when the run is resumed. Only an agent with a cooperative pause, whose process stays alive, goes from paused to interrupted when that process is lost. Every resume after a pause, a cancel or an interruption starts with a briefing from the supervisor (D27).
    • stopped is a normal end: the agent finished, or the run was cancelled. failed is an agent crash, a failed start or a failed resume.
    • Into interrupted: the reconciler (§5.3) marks a starting or running run (or a cooperatively paused one, above) interrupted when its process or environment is gone: a supervisor or host restart, or a lost environment. Nothing else sets it, and a pending approval is raised again on resume (§4.2).
    • Out of interrupted: starting when the reconciler resumes the agent from its session in a running environment (§4.3); stopped when the task is cancelled; failed when resuming is impossible or its attempts are used up.
    • One active run per environment counts interrupted (issues #201, #216). Until several agents at once are built (#94), an environment holds at most one run that is starting, running, paused or interrupted (§4.3): an interrupted run still owns its environment, because it is resumed there. An environment whose agent stop is still under way counts as busy too, whatever state its run was saved in, until that stop and its fallback have ended (the fifth path below; #238): a cancelled or failed run is terminal once saved, yet its agent may still run. A new run’s start and every path that moves a run to starting check it, the run itself excepted.
    • One live agent per run (issues #201, #202). Every path that moves a run to starting, the human’s resume and the reconciler’s recovery alike, holds that run’s lock, reads the run again after taking it, and refuses when an agent session is already attached to it; it never replaces a session that is attached. The path takes the run’s session slot before it saves starting and releases the lock once that write is saved, so from then until the launch ends the slot guards the run and a second path is refused as already attached (#225). A path that loses this race leaves the live session as it is and changes nothing else. So, within one supervisor process, cancel and pause always reach the agent that runs, and the reconciler never takes a live agent for a lost one.
    • No surviving agent before a relaunch (issues #201, #216). A supervisor that ends without stopping its agent (a crash, a kill, or a crash between saving a pause and stopping the agent) can leave the agent process running in the guest: ending the runtime’s exec client does not end it (spike #7, Case 4, verified on Apple Container 1.5.0), and a restarted supervisor cannot find it. So no agent is launched, a new run’s included, in an environment the current supervisor process did not start without stopping and starting that environment first. Stopping an environment is how an agent left over from an earlier process is ended; that a stop ends every process in the environment is unverified (seen in passing in the Apple Container spike, not measured; #216’s live test measures it). Five paths apply it. The first four act only on an environment the current supervisor process did not start, and the fifth only on the environment of a run whose own agent stop failed; because every launch goes through one of them, an environment is never stopped under an agent this process launched for a run that is to go on running:
      • The reconciler interrupts a starting or running run that has no attached session and then stops that run’s environment. The stop interrupts every starting or running run in that environment, attached or not, as any observed stop does; a hard-paused run stays paused (Into interrupted, above); each resumes with the briefing (D27). Recovery then starts the environment as for any stopped one. On every pass (issues #201, #221) the reconciler first stops, before it observes anything, each running workspace environment that its store records and the current process did not start, whatever runs it holds (the console’s environment, where no agent is launched, is left alone): a paused run in it stays paused, a lost run is interrupted as above, and a recorded environment with no active run is stopped too. The owner label is shared by every supervisor of the host, so an environment the store does not record is left alone: it may hold another supervisor’s live agent, and no agent of this supervisor was ever launched there, because every launch saves its environment first. A stop that fails is tried again on the next pass. A new process has started no environment, so its first pass ends whatever an earlier process left behind, a pause saved before a crash stopped its agent and a cancel whose stop failed included; while the supervisor is down, such an agent runs unsupervised until that pass.
      • Every path that moves a run to starting, from paused or interrupted (the reconciler’s recovery, a quota resume that is due, the human’s resume, and an answer that resumes) first stops and starts again such an environment; when that stop fails, the run keeps its state, nothing is launched, and for an interrupted run the failure counts an attempt.
      • A new run’s start in such an environment (a task started on an agent, a retry, a rework) first stops and starts it; when that stop fails, nothing is launched and nothing is saved: the start returns the error, as an environment that cannot be brought up does before the run is saved (§5.3, Start), so a retry’s or a rework’s answer stays open.
      • Cancelling a paused or interrupted run in such an environment also stops it, so a cancel reaches a leftover agent. When that stop fails, the cancel stands, the human is told that the environment was not stopped, and the reconciler’s pass stops it (above).
      • Any agent stop that fails (issues #221, #222, #238): a pause or a suspension (§4.2), a cancel of a starting or running run, a budget stop (§5.7, D48) or kill-all (§7, item 7) stops the run’s environment instead. The run already has the state its path gave it (paused, stopped or failed), which the coupling rules allow. A paused run still owns its environment; for a stopped or failed run, the environment counts as busy from before the run’s state is saved until the agent stop and this fallback have ended (One active run per environment, above; #238), so no other run is launched there in between. When that stop fails too, the run keeps that state, the human is told that its agent may still run (the error a human’s own pause or cancel returns; for a suspension, a budget stop or kill-all, which no human call waits on, and for a stop that ends after the human’s call has returned, as a cancel of a run that is still starting does, whose agent the start stops once it is attached: the event agent.may_run on the task, which the API’s task view and whr show show with the run (to be built, #238), and the notice agent_may_run in the inbox, §4.2; not the whr serve log alone; kill-all’s result lists each task whose stops failed, and each whose agent stop is still pending in a starting run as agent_stop_pending (to be built, #238), without waiting for it, so the kill switch never waits on a start), and the current process forgets that it started the environment: it then counts as one this process did not start, so the reconciler’s next pass tries the stop again and no agent is launched there without a stop and start first. A pidfile in the guest is not proof, because the agent can change it.
    • Terminal runs are never reused. A retry or rework starts a new run on the same workspace and topic (§4.1 task rules).
  • Environment (issue #15):

    FromTo
    provisioningstopped (created), deleted (provisioning failed or abandoned)
    stoppedrunning, deleted
    runningstopped
    deletednone (terminal)
        stateDiagram-v2
        [*] --> provisioning
        provisioning --> stopped: created
        provisioning --> deleted: failed or abandoned
        stopped --> running
        stopped --> deleted
        running --> stopped
        deleted --> [*]
      
    • Start and stop move between stopped and running. After a service or host restart every environment is stopped (spike #2, §5.3), so the reconciler observes running → stopped and starts the ones that should run.
    • Delete only from stopped. A running environment is stopped first.
    • Recycling (§4.3) deletes an environment and provisions a new one, so it is a new Environment with a new ID. An environment is never reprovisioned.
  • Coupling rules (issue #16). The three machines are not independent; a task aggregate (the task with its runs, their environments and its ReviewCandidates) checks these guards before a change is made:

    • One live run. A task has at most one run that is not stopped or failed, and a run is only created or started (from paused or interrupted) in an environment that is running.
    • A new run is new. StartRun takes a run that has not started (no state) and a non-empty ID that no run of the task has; a stopped or failed run is never passed back in. It is refused unless the task is queued, running or ready_for_review (rework, §4.1); a completed, cancelled, failed or awaiting_guidance task starts no run. Once every guard has passed, StartRun moves a queued or ready_for_review task to running in the same change, so a rework run never stops under a ready_for_review task; a refused start moves nothing.
    • Pause never stops the environment. Pausing a run changes the run only. An environment is not stopped while a run in it is starting or running; a run is stopped or interrupted first.
    • ready_for_review needs a stopped run and a pinned SHA. The task’s latest run is stopped (a failed run does not qualify), and a ReviewCandidate exists whose SHA is pinned (§4.5). The current revision is the most recent ReviewCandidate, and it records the run that produced it: it must be the latest run. Only the latest run may pin a revision, so a late pin from an older run cannot become the current one. After rework, a new run that stops without pinning a new commit leaves the task not ready, so the old SHA and its old CI result never carry over.
    • CI belongs to a SHA. A pipeline result is recorded on the ReviewCandidate of its own commit. A pass on an earlier SHA never counts for the current revision: where CI is required, ready_for_review is refused until the current SHA has passed.
    • Violations are conflicts. A broken coupling rule and an illegal transition are reported as a conflict, exit code 5 (§9.2). An unknown run, environment or SHA is “not found”, exit code 3.

4.2 Decision object

Fields: ID, task, run (empty for a review Decision, which no live run raised), kind (question | approval | review), blocking flag, subject, input (untrusted, capped), commit SHA, options, status, created and answered timestamps, deadline, answer, reason and answering actor (issue #17). The inbox, whr inbox, notifications and the audit trail hang off it. “Awaiting guidance” is the state a task enters while a blocking Decision raised by a live run is open; review Decisions belong to ready_for_review (§4.1).

Approvals are live and blocking. Spike #1 showed the pattern with Claude Code: the agent’s permission prompt is routed to the supervisor, which opens an approval Decision carrying the tool name and a capped copy of its input. The agent stays blocked until a human answers allow or deny, with an optional reason that is passed back to the agent. Rules:

  • Fail closed. Deny on timeout (the spike used 10 minutes) and when the supervisor is unreachable.
  • What is shown is what is allowed, or it says so. The input is capped at 2000 characters, and an approval whose input was cut anywhere on the way (in the adapter or the Decision) is marked input_truncated: the inbox, whr inbox and the API say that the allow covers more than is shown, so a human can deny what they cannot read. The tool name is the subject: a prompt whose name is not a tool name (letters, digits, _ . : -, at most 128 characters) is denied without asking. At most 8 prompts of one session wait for the human at once; one more is denied without asking, so an agent cannot flood the inbox.
  • Plan approval. In plan mode the agent’s ExitPlanMode arrives as an approval whose subject is the plan.
  • Capped input. Tool inputs in a Decision are capped (the spike used 2,000 characters); the full input stays with the agent. All of it is untrusted data.
  • Status. open becomes answered, expired (the deadline passed) or superseded (the supervisor restarted or the run was paused); each is terminal. Only answered with allow ever permits anything: an open, expired or superseded approval is a denial.
  • Approval channel transport (D26). Approvals travel over the agent’s stdio control protocol (--permission-prompt-tool stdio, stream-json in and out) on the runtime’s interactive exec channel (container exec -i): control_request (can_use_tool) on stdout, control_response (allow or deny) on stdin, matched by request_id. An answer for an unknown request_id is denied, and open requests are denied when they are superseded (D23). A closed stdin fails closed without running the tool, and a crashed supervisor end or an unanswered request leaves the guest agent waiting with the tool unrun (verified, spike #7), so when the channel is lost the supervisor stops the agent with whr-shim (D25) and records the request as denied. No guest-to-host network path, no listener and no approval token in the guest.
  • Deadline. An approval always has one (default 10 minutes). An answer that arrives after it is refused and the Decision expires, so a late “allow” does not count. An approval that has none (a lost value, a damaged row) counts as past it: an answer expires it, and Allows is false. The store refuses to write such an approval and refuses to load one. Only a question or a review Decision may have no deadline.
  • A refusal can still change the Decision. A late answer expires the Decision, and an allow for another commit is recorded as a denial, while both return an error. The store’s RespondDecision loads, answers and saves in one transaction and saves the change even when the answer is refused, so a refusal is never lost to a rollback. The caller still receives the error (exit code 5).
  • Bad input is a usage error. A missing actor or time, an answer that is not offered and a malformed new Decision are not conflicts: they exit with code 2 (§9.2). A conflict (code 5) depends on the Decision’s state; a usage error does not.
  • Commit SHA. A review Decision such as “Ready to push?” carries the pinned SHA of its ReviewCandidate (§4.5, §6). An allow is tied to that SHA: an allow given for a different SHA is recorded as a denial, and the supervisor asks Allows(sha) against the commit it is about to push, so an approval for an earlier revision never covers a later one.
  • Pause and restart (D23). A Decision raised by a live run (a question or an approval) is superseded when the supervisor restarts or the run is paused: pause is a hard interrupt (D11), so the agent process that asked is gone and could not receive the answer. The reconciler, or the resume, raises a new one when the agent asks again (§5.3). An answer sent to a superseded Decision is refused as a conflict (§9.2, exit code 5). A review Decision belongs to no run, so it survives both, and so do the login and quota questions below: they are raised for a run that is already paused, and nothing waits on them. This is the one restart rule; §4.1 and §5.3 refer to it.
  • Expired auth and quota (D23). auth_expired and quota_exhausted are blocking question Decisions raised for the paused run, not approvals, because nothing is being permitted. Their options are fixed: for auth_expired, “Signed in again, resume” or “Cancel”; for quota_exhausted, “Resume now”, “Resume at reset” (with the reset time when the agent reports it) or “Cancel”. The run stays paused and the task moves to awaiting_guidance (§4.1, D21). Suspending is a pause (issue #222): when the agent reports either while it runs, the run is saved paused with its question first, then its agent process is stopped as for the human’s pause (D11, D25), so nothing runs while the question waits; a stop that fails is handled as for the human’s pause (§4.1, “No surviving agent before a relaunch”). They have no deadline: waiting is safe, because nothing runs until they are answered.
  • Rebase conflict (D42). A conflict found when the supervisor rebases an agent’s branch onto the integration branch, before the export (§4.5), is a blocking question Decision raised for the stopped run whose branch it is, because no live run can be asked and the run ended there. The rebase is aborted inside the environment first, so the branch is as the agent left it. The options are fixed: “Rework” (a new run on the same agent, whose briefing names the conflicting paths, D27), “Resolved in the console, retry” (the human fixed it in the console, D43, and the supervisor rebases again) and “Cancel”. The input lists the conflicting paths as untrusted data. A rebase outside a task, such as one asked for by the human, returns the conflict to the caller and raises no Decision.
  • Refused prepare (D51, #253). A prepare that is refused for any other reason (a failed check, a refused commit message, a tree or topic cap, a rewritten history, no check configured, a forge error) is a blocking question Decision raised for the stopped run, cause prepare_failed, for the same reason as a rebase conflict. The options are fixed: “Rework” (a new run on the same agent, whose briefing names the reason, D27), “Retry” (prepare again, for a fault outside the agent’s work) and “Cancel”. The input is the reason and the check’s capped output, as untrusted data.
  • Refused publish (D51, #253). A publish step that the Guard or the forge refuses for a reason a retry cannot change (§4.5, “An outstanding publish”) is a blocking question Decision for the task, cause publish_failed. The options are fixed: “Retry” (complete the publish again under the same approval of the same SHA, for a fault fixed outside whr; the Guard checks everything again), “Rework” (a new run on the same agent, whose briefing names the reason, D27) and “Cancel”. The input is the refusal, as untrusted data. It belongs to no run, so, like the review Decision, it leaves the task in ready_for_review.
  • Agent may still run (issue #238; the inbox entry and its push to be built, #238). When an agent stop and the environment stop after it both fail on a path no human call waits on (a suspension, a budget stop, kill-all, a stop that ends after the human’s call returned; §4.1, fifth path), the supervisor raises a non-blocking question Decision for that run, cause agent_may_run, beside the task event, after the task’s state is saved: the inbox lists it with the task, a terminal task included, and nothing supersedes it (not the cancel kill-all makes next, not a restart, not the task’s end), so only the human’s answer closes it; and it is pushed (§9.4), because an agent that may run unwatched is what the human must hear about without looking. Its input names the path and the error, as untrusted data; its one option is “Seen”, and answering it records who saw it and changes nothing else. Being non-blocking, it leaves the task’s state alone. The inbox gets no other entry kind: everything it lists is a Decision.
  • Egress requests (D38). A host that a repository’s devcontainer.json requests or a lockfile suggests is opened only by the human’s allow, stored per repository on the supervisor’s side; nothing in the repository grants it. The question is asked once, at the first start of a run that would use the environment: the run is created and stays starting, and one blocking approval Decision per unanswered host is raised for it, with cause egress_request and the host in its own field (a host name the reader validated, never parsed from the subject or input, which only describe it and its source). The task moves to awaiting_guidance as for any blocking approval. Allow adds the host to the repository’s allowlist and deny records it as refused; either answer is kept, so later runs and other workspaces of the repository are not asked again. The deadline and the restart rule are an approval’s: an expired or superseded request is a denial for this run only and is asked again at the next start. The environment exists already (it is made with the workspace, D42) and starts with the hosts already allowed. The agent process is started only once every request of the run is answered; if an answer allowed a new host, the supervisor first recreates the environment’s proxy sidecar with the new allowlist, which is safe because no agent process runs yet. A denied host leaves the run to start without it. The allowlist never changes while an agent process runs. A host added to the repository later is asked at the next start. Changing an answer is a policy change (§6), outside the run.
  • Feature sources (D38, issue #127). A devcontainer feature from a source outside ghcr.io/devcontainers/features/ is asked like an egress host: one blocking approval per reference, cause feature_source, the reference in its own field, the answer kept per repository and answered only with a passkey step-up on the web (D45). An allow takes effect when the environment is next built: a new workspace, or whr ws rebuild <workspace>, which the human runs while no agent of that workspace runs and which keeps the home and build volumes (issue #128); nothing rebuilds by itself.
  • Card queue (D30, D40, issue #71). A card the human moves to an agent queue on the board raises a blocking question, cause board_queue, “Accept this task?” with the options start and cancel; it never starts a run by itself, once per card state. The project board does not say who moved a card, so the mover is treated as untrusted and only the issue author’s trust tier is applied; the human’s answer is what starts the run.
  • Task state (D13, D23). Only a blocking Decision raised by a run moves its task to awaiting_guidance: a question or approval from a live run, a login or quota question for a paused one, or a rebase conflict or a refused prepare for the stopped run whose branch it is. A review Decision belongs to no run and leaves the task in ready_for_review.

4.3 Lifecycle rules

  • IDE/SSH disconnect does not pause the agent or stop the environment.
  • Detached sessions. The agent process runs inside the environment under the supervisor with no client attached. Chat is attach and detach over a persistent session, never the owner of the process. Session state (transcript and agent session ID) lives in the workspace, so it survives a supervisor restart and a reattach from another client. The reconciler (§5.3) resumes from that session after a restart.
  • A long run can outlive its credentials. Expired auth or an exhausted usage window pauses the run and opens a blocking Decision (§5.2) instead of failing it.
  • Pausing lets a human inspect or edit without concurrent agent changes. Human takeover holds an explicit, visible workspace lock (agent vs human); on resume the agent is re-synced with the human’s changes.
  • Stopping an environment preserves files, logs and agent session state.
  • Agent-level checkpoint/resume is distinct from VM suspend and must work on backends without suspend.
  • Environments awaiting review may be stopped.
  • Worker recycling (recreate the environment, keep the workspace) is a first-class lifecycle action because freed guest memory is not returned to macOS.
  • Only validated runner × backend pairs may claim resumability.
  • Workspaces and agents (D42, issue #90). A Workspace has a unique name (lowercase letters, digits and -, 1 to 40 characters, starting with a letter), a host folder, the repository it was seeded from (owner/name), its integration branch (any branch name the configuration accepts: a valid ref name outside agent/, never chosen by an agent or an issue; issue #229). The agents’ namespace is compared without case (issue #254): a name whose first path component is agent in any case (agent, Agent/x, AGENT/main) is refused by every check that keeps a branch out of the namespace (the configuration, the workspace, the forge Guard’s target(), and hostgit’s fetch and prepare of a target, which also catch a record stored before this rule; to be built, #262), because the supervisor’s repositories live on a case-insensitive disk where refs/heads/Agent/x and refs/heads/agent/x are one file, and git cannot hold refs/heads/agent beside refs/heads/agent/<role>. The checks that admit only an agent’s branch point the other way and stay an exact, lowercase agent/ match (the Guard’s agentBranch for a push, hostgit’s Push), so a case variant is never accepted as an agent’s branch. A workspace also has its environment. An Agent has a role (the same character rules, at most 30 characters), unique within its workspace, and gets the branch agent/<role> and the worktree /ws/wt/<role> inside the environment from it; neither is chosen freely, so a name from a repository’s issue or an agent’s own output never becomes a path or a ref by itself. A task is assigned to one agent (Task.AgentID) and its runs carry the same agent. The model allows several agents per environment, but the service allows one active run (starting, running, paused or interrupted, §4.1) per environment and refuses a second with the conflict environment busy until several agents at once are built (#94). Workspaces and agents are registry records the human creates and removes; they are not part of a task’s aggregate and are not changed by task events, so each change is recorded as an audit event in the stream workspace:<id> (the event log is keyed by a task or by such a stream). Creating a workspace checks the folder (empty, below a workspace root, not inside a git repository; §5.6), seeds the agent clone with hostgit (the only time the host runs git there), records the workspace, provisions and starts its environment with the folder mounted read-write at /ws, and adds the first agent: git worktree add runs inside the environment, never on the host. A failed step takes back what the earlier ones made. Starting a task on an agent records the task and its run and checks the one-run rule in one transaction, so two starts cannot both pass; it then starts the agent in its worktree. If the agent cannot be started the run fails into a retry-or-cancel Decision and the environment is free again. One environment serves its workspace’s tasks one after the other, so an environment ID can be in several tasks. Rebasing an agent’s branch onto the integration branch runs git rebase in the agent’s worktree inside the environment, and is refused while the agent has a starting or running run. A conflict aborts the rebase and is returned with git’s report as the conflict rebase-conflict; the export (#91) turns it into a Decision on the task. The repository cache and per-task clones of hostgit stay until that export replaces the fetch from a stopped environment (§4.5).

4.4 Persistence semantics

Decided from spike #2 (Apple Container, issue #2; confirm on other backends):

  • Workspaces live on the host or an external SSD (D42), outside containers and volumes, and are mounted into their environment read-write. A workspace holds its own agent clone and the agents’ worktrees; the human’s own repositories are never mounted. Workspaces survive any environment, and one environment serves all agents of a workspace. This replaces per-task clones of a repository cache (D17); the agent clone stays in the workspace folder, and dependency directories follow D39 per worktree. Whether the agents’ worktrees move onto a volume is open (D39, issue #119); if they do, it is one volume per workspace environment with a directory per agent, kept when the environment is recreated.
  • The agent home (auth directory, session, caches) is one writable named volume per environment. A volume survives stop, start, delete and rebuild.
  • A writable volume is exclusive. While one container has it read-write, no other container can attach it, not even read-only (the second start fails with “The storage device attachment is invalid”). A rebuild stops the old container before the new one starts. A volume can be shared read-only by several containers.
  • A bind-mounted checkout is slower but usable. For 5000 files with 300 edits, the first git status took 0.3 to 1.1 s on a bind mount against 0.12 s in a volume, later runs about 100 to 130 ms against 75 to 118 ms, git add 363 ms against 120 ms, and commit 667 ms against 334 ms. Git tuning (untracked cache, feature.manyFiles, preloaded index) changed steady-state status by almost nothing, because the cost is the cold first stat of every file. Very large repositories were not measured; the cold pass grows with the file count.
  • Dependency directories stay off the bind mount (D39). Spike #40 measured file-metadata work on a bind mount at 20 to 100 times the cost on a volume for large trees (153 000 files: warm git status 8 s against 0.1 s, copying the tree 128 s against 1.5 s; 40 000 small files of a node_modules shape: 1.8 s against 0.04 s), while a CPU-bound build took the same time on both. So the untracked trees that hold most of a repository’s files are moved to a per-environment volume where the tool reads its location from the environment (CARGO_TARGET_DIR, UV_PROJECT_ENVIRONMENT), and covered by a volume mounted over the checkout only where it cannot be moved (node_modules). A mount point inside the bind mount cannot be removed or renamed in the guest (EBUSY), shows the volume’s lost+found and hides host files at that path (spike volume-over-bind), which is why relocation comes first. The checkout itself stays on the host. Results: spike #40.
  • Caches and build output stay on the volume, outside the checkout. A real build (this repository’s make check, including golangci-lint and editorconfig-checker compiled from source, 4 CPUs) took 45 s cold and 2.3 s warm with everything on a volume; with only the checkout on a bind mount it was 41 s and 2.7 s, no measurable cost; with the Go caches (about 25,700 module files and 7,800 build-cache files) on the bind mount as well it was 56 s and 4.7 s, about twice as slow warm. The adapter sets GOMODCACHE, GOCACHE and the package-manager and build-directory equivalents to paths on the environment’s volume. A tree with many small files of its own, such as node_modules, will behave like the cache case and was not measured.
  • Bind mounts sync both ways at once, and files the guest root writes appear owned by the host user.
  • The root filesystem is disposable. It survives stop and start and is lost on delete, so nothing that matters lives there.
  • Mounts are rejected by the adapter, not the runtime. The runtime accepts any host path. The adapter calls runtime.CheckMount before every mount: it resolves symlinks first, then rejects $HOME and its parents, ~/.ssh, other secrets directories and runtime sockets (§7.4), and the conformance suite checks it.
  • Tools are not part of the workspace. Agent CLIs come from a shared read-only store (§5.6).

Still open: retention and garbage collection of completed tasks, topics and workspaces on the host disk (§8), quotas (a named volume is a sparse image with a virtual size of 512 GiB), and repositories that track more than about 50 000 files, whose git status stays slow on a bind mount even with dependencies on a volume (D39, §12), and with them whether the agents’ worktrees move onto a volume (issue #119, waiting on its measurement).

4.5 Topics, checkouts and cleanup before push

A topic is one line of work: one branch (agent/<topic>) with its own checkout on the host. Several topics are in flight at once, each with its own environment and agent, and the developer can open any checkout in their editor at the same time. That is the worktree idea: parallel topics that are merged and tidied locally before anything leaves the machine.

Since D42 a topic is an agent’s branch in a workspace: each named agent works in its own worktree of the workspace’s agent clone, rebases onto the integration branch, and its commits leave as a git bundle exported from the running environment (issue #91), not through a fetch from a stopped environment. The import is a fetch of the bundle into a supervisor-owned repository that lacks its objects: git bundle verify alone passed a bundle truncated to half its size, while the fetch refused it and left no ref (spike #89). The per-task clone of D17 is gone (issue #95); its cleanup and push rules (prepare, per-SHA approval, push, no rewriting of pushed commits) carry over unchanged.

The export as built (issue #91). Request.Agent names the agent whose branch is exported. The supervisor’s forge mirror (the repository cache of D17 without its per-task clones, D42) stays: a bare repository per forge repository, fetched only from the forge, never from a workspace and never mounted into an environment. Prepare first refreshes the target from it into the supervisor’s repository, because the bundle’s prerequisite is the branch’s merge base with the integration branch. For a first round it then runs git rebase in the agent’s worktree inside the environment; a conflict aborts the rebase and raises the blocking question of §4.2 for the stopped run, with the conflicting paths as untrusted input (a later round is a fast-forward and is not rebased: pushed commits are never rewritten). Then git merge-base and git bundle create - <base>..<branch> run in the guest, through the same exec path (and its launcher) as the agent, under a timeout (default 5 minutes). Their output is streamed into Repo.ImportBundle, which caps it (default 512 MiB), fetches it into a scratch repository that borrows only the supervisor’s objects, checks that the branch arrived whole and only then fetches it into the supervisor’s repository; the scratch is removed on every path. The environment keeps running, and no git runs in the workspace on the host. A rewrite of commits already pushed is refused as before, by hostgit.Prepare against the revision’s recorded source. OpenCopy exports the same way, so the editor copy is fresh while the environment runs, and is cloned from the supervisor’s repository, never from the workspace.

Checkout layout (spike #2, item 9):

  • Each workspace has its own full agent clone (D42), made once by SeedAgentClone from the forge or from a repository path: an independent clone with no alternates, no hardlinks, no template and an origin that is a plain https URL, never a host path. The agent works in a worktree of it inside the environment. The supervisor keeps no per-task clones and mounts no object cache into an environment.
  • Clone depth is a per-repository setting, clone_depth, default 0 (full history), of the forge mirror. A positive value fetches the mirror with --depth for very large repositories; refreshing it then needs a forced refspec and a target is taken with fetch --update-shallow; once the target moves further than the depth past a branch’s fork point there is no merge base. The supervisor therefore checks for a merge base before a rebase and deepens the mirror and its own copy (--deepen) until it finds one. A partial clone (--filter=blob:none) keeps the full history and may be the better option; it was not tested.
  • Worktrees of one repository (git worktree) are the developer’s tool on the host for their own parallel topics. They are not handed to agents. A worktree’s .git file points at the shared repository by host path: mounting only the worktree fails (not a git repository), and mounting the shared repository exposes every branch, the shared hooks and config, and the other worktrees’ metadata. git worktree add --relative-paths makes the pointer work if the mounts keep the relative layout, but the exposure is the same.

Cleanup before push:

  • Agents do not push. The agent commits in its topic’s checkout. Nothing leaves the host until a cleanup step has run and been approved.
  • Prepare for push. The supervisor, on its own copy of the topic (hostgit, below), rebases it onto its target, folds attempts into one commit per finished change (fixup and autosquash of unpushed commits only), checks the commit messages (conventional commits and the repository’s commit linter), signs the commits with the bot key (§7.7) and runs the repository’s checks. Several finished topics can be merged into one integration branch first.
  • Approval. The result is a ReviewCandidate (a pinned commit SHA) and a Decision, “Ready to push?”, that shows the commit list and the diff stat. Approval is per commit SHA, as for any review (§6).
  • Push and PR. On approval the supervisor pushes the prepared agent/* branch with a token minted for that push alone (D51, §7.3) and opens or updates the PR. Merging stays human, on the forge.
  • Pushed commits are never rewritten. After a push a topic is only extended; rewriting pushed history needs an explicit request. Each revision records the agent’s own tip it was prepared from; a follow-up round rebases only the agent’s newer commits onto the pushed revision, so the next push is a fast-forward. An agent that rewrote commits it had already handed in is refused rather than guessed at. Publishing needs the task in ready_for_review and an approval of its current revision, so a cancelled task or an approval of an earlier revision publishes nothing, and the repository’s checks are required.
  • Done or cancelled topics. The checkout and its branch are removed by the retention rules (§4.4, issue #54) after the push is merged or the topic is cancelled. Unpushed work is kept until the owner discards it.

Prepare and push in code (issue #27). The steps above run in this order, on the supervisor’s own copy: the target comes from the mirror, the agent’s branch is imported from a bundle (above), EnsureMergeBase finds a base, and Prepare rebases the branch onto the target in a temporary worktree with --autosquash (hooks off, the bot as committer, commits signed with the bot key), then lints every commit message and returns the new tip. A topic with no commit after its target is an error, never a revision pinned for review with 0 commits (hostgit.ErrNoCommits, issue #254), so a tip nothing signed or linted is never offered for approval. Running the repository’s own checks is not done by hostgit: they are the repository’s code, so a Checker runs them in an environment, never on the host. The tip is pinned (PinRevision) and a review Decision “Ready to push?” is raised for that SHA. Push sends one agent/* branch, fast-forward only, as the exact approved commit (<sha>:refs/heads/<branch>), never with force and never tags.

    sequenceDiagram
    participant A as Agent (environment)
    participant S as Supervisor (hostgit)
    participant C as Checker (environment)
    participant D as Developer
    participant F as Forge
    A->>A: commits on agent/<topic>
    A-->>S: the run stops
    S->>A: git bundle create - (the environment keeps running)
    S->>S: ImportBundle into its own copy
    S->>S: EnsureMergeBase with the cached target
    S->>S: Prepare: rebase --autosquash, sign, lint messages
    S->>C: run the repository's checks
    C-->>S: result
    S->>S: PinRevision (ReviewCandidate for the SHA)
    S->>D: Decision "Ready to push?" (commits, diff stat)
    alt approved for this SHA
        D-->>S: allow
        S->>F: push <sha>:refs/heads/agent/<topic>, fast-forward only
        S->>F: open or update the PR
        F-->>D: review and merge, by the human
    else declined, or the PR needs changes
        D-->>S: deny
        S->>A: rework: a new run on the same topic
    end
  

Publishing in production (D51, issue #244; built: the check and its configuration in #251, the push credential in #252, the trigger, the answer, the completion and the wiring in whr serve in #253). The parts above are joined into one path:

  • Trigger. When a task’s latest run ends stopped because its agent finished (not a cancel), the service prepares the topic: the export, Prepare, the check and the pin, then “Ready to push?” for the pinned SHA. A restart in between is repeated: the reconciler prepares again a running task whose latest run is stopped, with no revision pinned for that run and no open Decision; preparing is repeatable, because it starts again from the bundle. From the moment the run is saved stopped for a prepare until the prepare has pinned a revision or been refused, the environment and the agent count as busy: a run start in that environment and a new task on that agent are refused with environment busy (One active run per environment, §4.1), so no other run’s commits reach the export or run beside the check; the reconciler’s repeat holds them the same way.
  • The check. The command is the repository’s check in the supervisor’s configuration, else customizations.workharbor.check read at the default branch’s commit (D38), else pre-commit run --all-files when the default branch has .pre-commit-config.yaml. With none of them the prepare fails closed (ErrNoChecks): nothing is published without a check. The Checker runs it in the task’s workspace environment, started first if it is stopped, as the agent’s user, in a check worktree detached at the prepared SHA that is no agent’s worktree and is removed afterwards from inside the guest. The prepared commits reach the guest as a bundle from the supervisor’s repository, with prerequisites the agent’s clone holds, streamed through the exec’s stdin; the host writes nothing into the workspace and runs no git there. The check gets no credential of whr’s (no forge token, no API key); a subscription login in the environment stays readable to it, as it is to the agent (accepted, §7.3). It has only the environment’s egress allowlist, a timeout (default 30 minutes) and an output capped to its last 64 KiB, which is untrusted input (§7.1). The check counts as running, and the environment as busy, until its process group is gone: after every check, a normal exit included, the checker ends the check’s process group (D25) and looks for it, and a process still there stops the environment as on the fifth path of §4.1 (under the start lock; when that stop fails too, the process forgets that it started the environment), with the busy mark held until that has ended. The check leaves a receipt (issue #259, built in #253): an event on the task bound to the prepared SHA, naming the command and where it came from (the configuration, devcontainer.json or pre-commit), the exit status, the duration and the capped output, which “Ready to push?”, prepare_failed and whr show show, so the human sees which SHA was checked, by what and with what result. A process that leaves the group (setsid, a double fork) escapes, as it does for an agent: accepted, because it stays inside the workspace’s one trust domain (D42) and the environment’s egress allowlist. Exit status 0 passes. The check is a quality gate, not a security control: the agent writes the tree it checks, including what the command runs, so the per-SHA approval, the forge’s CI and its ruleset stay the controls.
  • A refused prepare (a failed check, a refused commit message, a cap, a rewritten history, no check configured, a forge error) raises the blocking question prepare_failed for the stopped run (§4.2).
  • Identity and lint. The committer is the GitHub App’s bot identity, its name (<slug>[bot]) and noreply address read from the App (D15), never typed into the configuration; commits are signed with the bot key (§7.7). The host lints messages with a built-in linter the repository’s configuration names (commit_lint: conventional, the default, or workharbor, the rules of this repository’s internal/commitlint); a repository’s own linter configuration is repository code and runs only as part of its check, in the environment.
  • Answering “Ready to push?” is one service call, whichever way it comes: the JSON API’s POST /v1/decisions/{id}/answer on the host socket, whr approve <decision> --sha <commit> over the same API, or the web UI only after a passkey step-up naming the Decision and its SHA (§7.5, D45). It records the answer first; an allow then starts the publish, a deny publishes nothing. No handler and no CLI command calls Publish itself.
  • An outstanding publish is completed, never repeated. An allowed review Decision of a task that is still ready_for_review, whose SHA is the current revision and is not recorded as pushed, and which has no open publish_failed question, is an outstanding publish; answering that question “Retry” makes it outstanding again, while “Rework” and “Cancel” move the task out of ready_for_review. The reconciler’s pass completes it, so a restart between the answer, the push and the record of the push loses nothing. Every forge step is idempotent for the approved SHA: pushing a branch that is already at it succeeds, a fast-forward whose target already points at it succeeds, and opening a PR first looks up the open PR whose head is the repository’s own branch (<owner>:agent/<role>, never a fork’s branch of the same name) and whose base is the expected target, and updates that one, so a retry never opens a second PR. Completing it needs no new approval, because the human allowed exactly that SHA, and the Guard checks it again at every step. A step that fails for a transport fault (a network error, a timeout, a server error, a rate limit) leaves the publish outstanding, is retried on a later pass with backoff, and each failure is an event on the task that whr show shows. A refusal that a retry cannot change (not a fast-forward, a target the Guard refuses, no approval for the SHA, a forbidden request or a missing installation) ends the outstanding publish with the blocking question publish_failed (§4.2), and so does any failure not known to be a transport fault: an unknown error fails closed, and the human’s “Retry” is the way back. Git reports a failed push only as text, so hostgit tells a transport fault by markers in git’s own messages under the C locale; a marker found in a line the server sends (remote:), or transport faults that go on for 24 hours after the first failed attempt (DefaultPublishMaxAge, counted from the task’s publish.attempt events, so a restart keeps it), end the publish the same way (#270). A cancelled task, or a newer revision, ends it.

The editor copy (issue #59). The developer’s editor is never pointed at an agent’s checkout: opening it runs planted repository config on the host (VS Code and its git extension, direnv, .vscode/tasks.json; threat model T15). EditorCopy clones the supervisor’s own bare copy of the topic (which the bundle import filled) into a directory outside the workspace root, with an empty template and none of the agent’s config, hooks or alternates, and refreshes it later with a fast-forward only, so the developer’s edits are never overwritten. It refuses a destination inside the workspace root, so the agent’s checkout cannot be handed out by mistake. The commits themselves are the agent’s work and may contain files an editor acts on (.vscode/tasks.json, .envrc, .devcontainer), so the copy lists them and the UI warns before the folder is trusted.

The forge contract (§10) separates Git transport from the forge API. The adapter is wrapped in a Guard that holds the autonomy table: OpenPR and UpdatePR take an Approval (the Decision and the commit SHA) and are refused unless the Decision allows that SHA and the branch at the forge points at it; Push needs the same proof and is the only caller of the adapter’s Pusher, which holds the push credential (D51, §7.3); Merge, Tag, Release and Deploy are refused whatever the table says (the table’s ceilings), and the inner adapter is never called for them. An Issue carries the author association for trust tiers (#53), and VerifyWebhook takes the request’s http.Header.

The host treats every agent-writable checkout as hostile (§7.4), and since D42 runs no git in one at all: an agent’s commits arrive as a bundle read as a stream, never as a fetch from its .git.

hostgit (issue #19) is the only way the host runs git on anything an agent can write:

  • No git in a workspace on the host. The only time the host runs git in a workspace’s clone is when it does not yet exist: SeedAgentClone makes it. After that the agent works inside the environment, and its commits leave as a bundle (ImportBundle): read as a stream with a size cap, fetched into a scratch repository that borrows only the supervisor’s objects, checked as a whole (git bundle verify alone passed a bundle cut in half, spike #89) and only then fetched into the supervisor’s repository, with object checks (fsckObjects) on. Cleanup (rebase, fold, sign) and push run only on that supervisor-owned repository. An earlier design let the host read an agent’s checkout with an allowlist of read-only plumbing after a check of its .git; that code is removed.
  • The expanded tree is capped before any checkout. Only a bundle’s compressed size was capped, and a few kilobytes of tree objects can name billions of files, or billions of trees with no file at all (a tree that lists the same subtree twice, nested). Repo.CheckTree streams git ls-tree -r -t -l -z of a commit, counting trees as well as files, and stops git as soon as it passes MaxTreeEntries (200,000 entries) or MaxTreeBytes (2 GiB of blobs). Prepare replays every commit of the topic into a host worktree, so CheckCommits checks the tree of each commit after the target (or after the pushed source in a follow-up round), once per tree id, for at most MaxTopicCommits (2,000) commits: a bomb that a later commit deletes is refused too. --autosquash builds a tree no single commit has (the union of what a chain of fixups added), so CheckCommits also caps what the topic adds as a whole: the entries and blob bytes each commit introduces over its parent, summed, at MaxTopicEntries (400,000) and MaxTopicBytes (4 GiB). The branch is resolved to a commit ID once, and the check and the checkout use that ID, so a branch that moves in between cannot slip past. The check, Prepare and EditorCopy run under hard deadlines (CheckTimeout, PrepareTimeout, EditorCopyTimeout). A refusal is ErrTreeTooLarge. Tests build a 2^30-file bomb from 31 objects, a 2^40-tree bomb with no file from 41, and a bomb in one commit deleted by the next, and require a quick refusal.
  • A hardened floor on every host git command. Git starts with an empty environment (no GIT_* from the host, a minimal PATH, an empty HOME, no system or global config, no terminal prompt, no optional locks, no replace objects) and -c overrides that disable hooks (core.hooksPath), the file-system monitor, the pager and editor, credential helpers, signing, submodule recursion, and every transport except file where a fetch from a local repository needs it.
  • The forge mirror (issue #45, D42). Cache is the bare repository per forge repository, with the repository’s clone_depth (default 0). Refresh fetches one branch from the forge with the forced refspec +refs/heads/<b>:refs/heads/<b>, no tags, objects checked, and --depth for a positive depth (an existing shallow mirror is made complete with --unshallow when the depth is 0). The source is an absolute path or an https:// URL without credentials; every other scheme (ssh, ext::, file://) and anything that starts with a dash is refused, and an authenticated fetch comes with the forge adapter (issue #27). It is fed only from the forge, never from a workspace, and never mounted into an environment.
  • One writer per cache (issue #67). Every operation on a cache (refresh, deepen, clone, merge-base search, maintenance) runs under a per-cache lock: a mutex shared by every handle of the same path and a file lock (whr.lock) for other processes, waited for with the caller’s context. Git’s own lock files (shallow.lock, packed-refs.lock, ref locks) are never left to block the next refresh: while the cache lock is held no other supervisor process runs git on it, so a lock file found then is stale and is removed first. Three parallel shallow refreshes had failed 40 times in 60 on shallow.lock without it.
  • Names map to directories by a strict rule. CachePath(repo) accepts only owner/name made of ASCII letters, digits, ., _ and - (each part starting with a letter or digit, at most 100 characters), lowercases it, and appends a short hash, so Foo/x and foo/x are one cache on a case-insensitive disk and no name, path separator or Unicode form can leave the cache root. A cache path must be absolute and inside the cache root when one is set (WithCacheRoot), a relative path is refused instead of being created under /, and a local source inside a workspace root (agent-writable) is refused.
  • Merge base before a rebase. EnsureMergeBase looks for a merge base between the target and the topic in the supervisor’s copy. While there is none and either side is shallow it deepens the cache (fetch --deepen) and then the copy (fetch --update-shallow --deepen) by a step, up to a limit, and reports ErrNoMergeBase when the limit is reached or both sides are complete, so the caller opens a Decision instead of rebasing blindly.
  • Tested. Tests plant a hook, core.fsmonitor and a clean filter in a workspace’s .git; the bundle import never runs them, and a control run of host git in the clone shows each one fires (spike #89). Truncated, corrupted, oversize and garbage bundles are refused and leave no ref.