Skip to content

Architecture

5. Architecture

Logical components live in one Go binary on the Mac; boundaries are package interfaces, not microservices.

ComponentResponsibility
Control planeTasks, runs, workspaces, decisions, policies, integration config, event log, reconciler
Web UIv0: task list, live transcript and inbox; writes are answering Decisions, sending messages, starting tasks, pause/resume/cancel and transcript purge (§9.3). Server-rendered (D8)
CLI (whr)Same operations through the shared API
Host workerExecutes a fixed set of authorized lifecycle operations beside the runtime
Agent adapterStart, observe, instruct, pause, resume a coding agent
Runtime adapterProvision and manage environments
Forge adapterIssues, PRs, reviews, metadata, webhooks; enforces the policy table
CI adapterInterface only in release 1
Credential serviceEncrypted store (macOS Keychain holds the master key); issues per-run credentials

Host worker and credential service are package boundaries on one host. Until the encrypted store exists, release 1’s credential service is a set of secret files (0600, owned by the supervisor’s user, one link, outside every root, read with config.ReadSecret, never logged): every secret file the configuration names: the GitHub App key (github.key_file), the supervisor’s API token (api_token_file), the bot’s commit-signing key (bot_signing_key_file, D51), the console’s SSH authority (console.ssh_ca_key_file), the ntfy topic and token (ntfy.topic_file, ntfy.token_file) and, in the api-key auth mode, the agent’s key (agent_api_key_env_file) (issue #177). Keep the contract remote-capable so a remote worker can be added later, without building inter-process auth now.

5.1 Runtime adapter

Covers provision, start/stop/delete, inspect, resource limits, logs, exec, storage and endpoint discovery. The Apple Container behaviour below was measured in the Apple Container and host reachability spikes. Capabilities are explicit and never assumed:

CapabilityPurpose
Isolation boundaryShared-kernel container, guest kernel or full VM
CPU architectureImage, toolchain and IDE compatibility
Persistent storageWhat survives stop, rebuild, delete
NetworkingReachability and supported isolation controls
Suspend/checkpointReported support only
SSH/browser accessIntervention endpoints

The Go contract (issue #20, internal/runtime):

  • Spec carries what the hardened environment needs: image, owner and labels, CPUs, memory and a disk quota, the network (default or --internal, by name), the user (never root), a read-only root, cap-drop ALL, --init, tmpfs mounts and the mounts. Validate rejects a spec that is not hardened: a root or empty user, no cap-drop ALL, no --init, no owner, a relative or duplicate target. Mounts are bind (a host path, checked by CheckMount, §7.4) or volume (a name); a bind mount of a forbidden path never reaches the runtime.
  • Owner label. Every environment carries the supervisor’s owner label. List(owner) returns only environments with that label, and an adapter refuses to start, stop or delete one it does not own. Removal is by exact ID, never a pattern (rm --all deletes every container on the machine).
  • Inspect returns a typed Info: the ID, owner, labels, state (provisioning, running, stopped, deleted as in §4.1) and the current address, which is empty unless the environment is running and is never stored. An unknown environment is ErrNotFound.
  • Exec streams: separate stdout and stderr readers, Wait for the exit code, and cancellation through the context. Exec in an environment that is not running is ErrNotRunning. ExecRequest.Stdin is an optional reader the adapter copies into the process and closes at its end; with a pipe the caller keeps writing while the command runs, which is how the Claude Code adapter sends stream-json input and mid-run messages (§5.2). Without a reader the command’s stdin is closed. container exec carries stdin through to the guest: the runtime conformance suite’s stdin checks pass against the Apple Container adapter (go test -tags applecontainer ./internal/runtime/apple, container 1.5.0, issue #26).
  • Cancel kills the process in the guest. Cancelling the context ends the stream, and the process inside the environment must be gone, not only the client (spike #2: SIGINT was not forwarded, and the agent kept running). The suite proves it with a second exec that looks for the process, so a backend that only returns from Wait fails.
  • Lifecycle is idempotent. Start of a running and stop of a stopped environment succeed, so a reconciler can retry (§5.3).
  • Hardening is required, not possible (issue #26). Spec.Validate is not enough on its own: a spec with the default network, a writable root, another task’s volume and a bind of ~/.ssh validated. runtime.Prepare is the one checked step. It runs Validate, requires Network.Internal and ReadOnlyRoot, checks every bind mount with CheckMount (and CheckMountsWithin the workspace roots), checks that every named volume belongs to this environment (Owns), and returns a PreparedSpec with unexported fields. Provision takes only a PreparedSpec, so nothing unchecked can reach the runtime. CheckMount returns the resolved path and Prepare puts that path in the prepared spec, so the adapter mounts exactly what was checked: a source swapped for a symlink after the check is not followed.
  • The contract owns the surroundings. Provisioning creates the per-environment --internal network, the agent-home volume and the egress sidecar (Spec.Egress: the proxy’s image, its allowlist and the host binary of the proxy), and deleting the environment removes all three, by exact name. Resources reports them, so the conformance suite can check that they exist and that they are gone. A writable volume is exclusive (§4.4): starting a second environment whose volume another running environment holds read-write is refused with ErrVolumeBusy.
  • Cache objects are read-only. For a --shared topic, Spec.Alternates lists the cache’s objects directory (issue #45), and Prepare turns each into a read-only bind mount at the same host path, because the clone’s alternates file names that path. They are checked like any bind mount and must lie under the cache root.
  • A terminal is not a pair of pipes (D43, issue #92). runtime.TerminalAdapter.Terminal(env, request) runs a command with a terminal and returns a runtime.Terminal: one stream for the command’s output and one for its input, Resize(cols, rows), Wait for the exit code and Close, which hangs the terminal up. The Apple adapter makes a pseudo-terminal of its own (on macOS and Linux, golang.org/x/sys/unix, which was already a dependency of golang.org/x/term), runs container exec -t -i with it as the client’s controlling terminal in a session of its own, and holds the master: the guest command has a real terminal with a size, a resize reaches it, and closing the master sends the client SIGHUP, which ends the guest’s terminal and so the shell. The environment variables travel through the 0600 env file like the ones of Exec. verified on container 1.5.0: TestTerminalLive (the size asked for, TERM, a resize, input, the exit code) and TestTerminalCloseEndsTheGuestCommand (a sleep of the guest is gone within seconds of the close), and the conformance suite’s terminal check (output, exit code, input, resize, Close, an unknown and a stopped environment).
  • The egress allowlist changes only before an agent starts (§4.2). runtime.EgressUpdater.UpdateEgress(env, prepared spec) replaces the environment’s proxy sidecar with one that has the spec’s allowlist: the old sidecar is stopped and deleted, the new one is created on the same networks and, if the environment runs, started. The spec must be prepared, carry an Egress and name the environment’s own network. The sidecar’s address changes, so a process that already runs keeps the old one: the supervisor calls it only while no agent process runs. Inspect reports the sidecar’s allowlist (Info.EgressAllow, sorted; on Apple Container from the sidecar’s workharbor.egress label), so the supervisor changes the sidecar only when the allowlist differs. verified on container 1.5.0: the conformance suite checks the update on a stopped and on a running environment, that one sidecar remains and the refusals, and TestEgressThroughTheSidecar checks that a newly allowed host answers CONNECT 200 and a dropped one 403 through the new proxy.
  • Fakes and conformance. runtimetest has an in-memory fake that can simulate a service restart (every environment becomes stopped, as measured below) and the conformance suite every backend must pass; the real Apple Container adapter (#26) runs the same suite on the Mac. internal/runtime/apple runs it behind the applecontainer build tag (go test -tags applecontainer ./internal/runtime/apple, container 1.5.0 with a local fedora image, about 30 s). The same tag runs the egress conformance: a direct connection fails, an allowlisted host answers through the sidecar (CONNECT 200), a raw IP and an unlisted host get 403, and the guest resolves no names. The adapter passes mounts with -v (its value is not split at commas, so a path cannot add an option; a : in a path is refused in Prepare), exec environment through --env-file /dev/fd/3, a pipe the supervisor writes and the CLI reads (no value on the command line, and no file: the 0600 temp file this replaced stayed for the life of the exec and a crash left it behind, issue #122; whr serve sweeps such files an earlier run left, verified for the CLI reading a pipe on container 1.5.0 by the live exec and terminal tests), reserves the workharbor. label prefix for its own labels, never --ssh, --rm or a published port, finds a volume’s name in the mount type (its source is the image file path), and acts on exact names and the owner label only. An environment’s Addr and Proxy are read from container list on every call and never stored.
  • Volumes belong to the environment’s user. A new volume is an empty ext4 filesystem owned by root, so an unprivileged agent could not write its home. At creation the adapter runs one short container as root with CAP_CHOWN only, the volume and the read-only tool store mounted, on the environment’s internal network, and whr-shim chown from the tool store as its entrypoint: it gives the volume’s top directory to the spec’s numeric user without following a link, and the container is deleted (found by the serve integration run, issue #28). No program of the environment’s image runs as root, because since D38 that image may be built from the repository, and the helper has no way out. It is the only thing the adapter runs as root.
  • Environment from the repository (D38, issues #77 and #76). internal/devcontainer resolves the environment of a repository at its default branch: Resolve turns the ref into one commit and every later read, the file, the toolchain files, the lockfiles and the build context, uses that commit through git ls-tree and git cat-file, never a working tree, so a topic’s devcontainer.json, an edit not committed and a branch that moves between two reads change nothing. The order is a devcontainer.json (.devcontainer/devcontainer.json or .devcontainer.json, JSON with comments and trailing commas), else a Dockerfile or Containerfile at the root, else a default image chosen from the toolchain files (go.mod, .tool-versions, mise.toml, .nvmrc, .python-version, in that order; the version must be digits and dots, and the image comes from the supervisor’s table, the official language images by default, or its base image when there is none, D43). The file’s image or build.dockerfile, build.context and build.args, containerEnv, postCreateCommand (a string or an array, run inside the environment by sh -c or directly, never on the host), forwardPorts (plain port numbers, the preview ports of D33) and customizations.workharbor are read. customizations.workharbor holds egress (host names only, each a request), and the hints check, previewPorts, agent and tools, which suggest and decide nothing. features are fetched, checked and applied to an image environment, or on top of the image a build.dockerfile builds (issues #108 and #127, below). It refuses the whole file, naming every key, when it has initializeCommand, mounts, workspaceMount, runArgs, privileged, capAdd, securityOpt, a remoteUser or containerUser that is root (by name or any numeric spelling of uid 0), a containerEnv or build.args variable the supervisor sets or the agent reads (the proxy variables, PATH, HOME, LD_PRELOAD, LD_LIBRARY_PATH, and names starting with WHR_, CLAUDE_, ANTHROPIC_, OPENAI_, CODEX_, GEMINI_, GOOGLE_ or BUILDKIT_, runtime.ReservedEnv; the proxy variables are predefined build arguments, so they would steer the build’s traffic), or a build.dockerfile or build.context that is absolute or leads outside the repository. The file must be a regular blob of at most 1 MiB, not a symbolic link, the ref may not look like an option or a range, and a repository that cannot be read is an error rather than “no devcontainer”; every other key is ignored with a note. Build. Environment.Stage writes the build context and the Dockerfile from the commit into a directory the supervisor owns (Export: regular files only, a symbolic link or a submodule is skipped and listed, no .git, no attributes or filters, at most 20000 files and 512 MiB), and a runtime.Builder builds it (runtime.BuildSpec, checked like a Spec: a clean absolute context, a valid tag, no reserved build argument). The tag carries a digest of the commit, the Dockerfile, the context and the arguments, so the same commit reuses its image and a new commit builds a new one. The Apple adapter runs container build --progress plain --tag … --file … --build-arg … -- <context> (verified on container 1.5.0: TestBuildLive, go test -tags applecontainer); it never passes --ssh, --secret or --output. The build runs in the builder VM, which the egress allowlist does not cover: a RUN step can reach any host, and what that means for a reviewed Dockerfile is a question for §7.2 (open, proposed on #76). Spec. Environment.Spec applies the environment to the supervisor’s hardened Spec: only the image and Spec.Env (the containerEnv, refused again by Spec.Validate for a reserved name) come from the repository; the user, the network, the mounts, the capabilities and the limits stay the supervisor’s, and remoteUser, a named user, is not mapped to a uid. In whr serve. serve.Environment refreshes the supervisor’s own mirror of the integration branch, resolves the environment of that commit there and, for a repository with a devcontainer or a Dockerfile, gives Workspaces.provision the means to make its image: once per tag, from a context exported from the commit into a directory of the supervisor’s own, through the runtime’s builder. provision then runs that image with Environment.Spec (the containerEnv; everything else of the spec stays the supervisor’s) and the usual git check. A repository that cannot be read gets the supervisor’s default environment and the failure is reported; one whose image cannot be built is refused, with the builder’s output, and nothing is left behind. postCreateCommand runs at a run’s start, after the egress allowlist is up to date and before the agent starts (so what it fetches goes through the hosts the human allowed), in the agent’s worktree with the agent’s environment, as the agent user, once per environment and command set (a marker in the agent home); a command that fails ends the run as failed, with the tail of its output. Egress. A host the file requests and a host a lockfile suggests (go.sum, package-lock.json, yarn.lock, pnpm-lock.yaml, poetry.lock, uv.lock, Pipfile.lock, Cargo.lock, Gemfile.lock) are HostRequests, minus the hosts already answered for the repository; nothing is allowed by being listed. The answers, allow or deny, are stored per repository in the supervisor’s database (egress_hosts, Store.SetEgressHost), never in the repository. Service.PendingEgress lists the hosts of an environment that the repository has not answered yet, and Service.RequestEgress raises one blocking approval per host for a starting or running run (TaskAggregate.RaiseEgressRequest: cause egress_request, the validated host in the Decision’s own hostname column, all or nothing). AnswerDecision keeps the answer for the task’s repository, needs no waiting agent (none exists before the run starts), and Service.EgressAllow returns the allowed hosts for the environment’s allowlist. An expired or superseded request keeps nothing: it is a denial for its run, and the host is asked again at the next start. The start flow (§4.2). Workspaces.launch saves the run, then gateEgress reads the repository’s environment from the supervisor’s own mirror of the integration branch (Config.Environment, in whr serve a refresh of the mirror and devcontainer.Resolve of that commit, never a workspace) and, if it has hosts not yet answered, raises the requests and leaves the run starting with its slot held, so the reconciler does not take it for lost. AnswerDecision starts the agent when the last request of the run is no longer open, and so does the reconciler when the last one expired. Before the agent process starts, applyEgress compares the sidecar’s reported allowlist with the supervisor’s hosts plus those allowed for the repository and, if they differ, replaces the sidecar through runtime.EgressUpdater (§5.1), which is safe because no agent process runs yet; the allowlist never changes while one runs. A new workspace of the repository is provisioned with the allowed hosts at once. A repository that cannot be read (no network, a private repository before #27) starts the run with the allowed hosts and reports the failure: what could not be read allows nothing. Cancelling a waiting run drops its wait, and an answer to a superseded request is refused. The start is detached. The agent start of a run (applyEgress, postCreate, the agent itself) runs in a goroutine of its own on a context that is not the request’s (Service.startDetached): a phone whose request drops must not fail a run whose answer was stored, and a slow postCreate must not hold the reconciler. At most one start runs per run, it is tracked so Shutdown cancels and waits for it, and Cancel stops one already running (a postCreate command ends with its context, and an agent that had just started is stopped); a cancelled start does not fail the run, which the human stopped. StartTask waits up to 30 seconds for the start, so a quick one returns with the agent running and a start that fails in that time returns its error, while a longer one returns with the run starting and a failure is reported. postCreate is bounded by environment.post_create_timeout (a Go duration, default 10m): a command still running then fails the run with that reason. This repository’s own file builds .devcontainer/Dockerfile (the Go version of go.mod, a non-root agent user with a passwd entry, which the tests’ ssh-keygen needs) and asks for proxy.golang.org and sum.golang.org. A repository may not name an image under whr.invalid/ (rule 4a): Parse refuses the whole file naming image (the host compared without regard to case, quoting or white space) and naming a build.args value that mentions whr.invalid, and Stage refuses a Dockerfile that mentions whr.invalid anywhere (RefuseBuiltFrom), so a FROM, a COPY --from, a RUN --mount=...,from=, an ARG default or a continued, quoted or # escape= spelling cannot build on another repository’s built image. The check is a scan of the file’s text and of build.args after lower-casing and removing backslashes, backticks, quotes and white space; it parses no Dockerfile. It is defence in depth, not a closed door: a name assembled from ARG pieces (ARG H=whr, ARG D=invalid, FROM ${H}.${D}/x) is not in the text and is not caught (TestRefuseBuiltFromKnownGapNameAssembledFromArgs). The exposure is a repository naming another repository’s locally built image, whose content came from that repository’s default branch, within one supervisor and one human. Measured (verified, spike/builder-store, a local branch until it is pushed): container build resolves a FROM of a locally tagged whr.invalid/ image from the host’s image store when it is run without --pull, so the scan above is the only guard, and with --pull it fails. Stage also refuses a syntax parser directive that names anything but docker/dockerfile, with an optional tag and digest (RefuseSyntaxDirective), because a custom frontend is an image the builder pulls and runs: the marker may be # or //, with any Unicode white space around the key (as BuildKit’s own detection), and a file that starts with { and has a syntax key is refused as JSON, which BuildKit also reads; a #! first line is discarded first, as BuildKit does, so a shebang does not hide either. Only the leading comment block counts, and when unsure the check refuses, so an odd first comment line can cost a Dockerfile its build. BUILDKIT_ is a reserved prefix, so build.args cannot set BUILDKIT_SYNTAX. Gaps: an ONBUILD instruction in a base image the repository names has no static fix, and COPY --from= and RUN --mount=type=bind,from= of a local whr.invalid/ image resolve in the builder like a FROM does, without --pull (verified, spike/builder-store), so the text scan is what catches all three. The builder-side rule, so that the builder cannot resolve a locally tagged whr.invalid/ image at all, is pending wh/design; the gap is accepted for one supervisor and one human.

Do not pretend backends share Docker semantics. One runtime conformance suite (the §12 checklist, automated) must pass for every backend; it turns capability flags into verified claims.

Measured on Apple Container 1.5.0 (macOS 26.6.2, spike #2, issue #2):

CapabilityObserved
Isolation boundaryA lightweight VM per container: its own Linux kernel and one host runtime process each
CPU architecturearm64 guests; a Rosetta flag exists and was not tested
Persistent storageNamed volumes (ext4 image files, exclusive while writable) and bind mounts survive delete; the root filesystem does not (§4.4)
NetworkingThe default NAT network reaches the LAN, the internet, other containers and host services bound to all interfaces. --internal networks block the internet, the LAN and other containers, but not the host: an internal guest reaches host services bound to the Mac’s LAN address or to all interfaces (issue #69, branch spike/host-reachability). Only loopback-only listeners are out of reach (§7.2)
Resource limits--cpus sets the vCPU count; --memory is enforced by a cgroup inside the VM (a larger allocation is killed with exit 137 and the container survives)
Restart policyNone. After a crash or a service restart every container is stopped (§5.3)
Suspend/checkpointNone observed
SSH/browser accessNot tested; exec works

Adapter rules that follow from it:

  • Always pass --init: a stop took 145 ms with it and 5.3 s without, because PID 1 ignored SIGTERM.
  • Never store a container’s IP; it changes across recreate and restart. Read it with inspect.
  • Remove only containers by exact ID. container rm --all deletes every container on the machine, including ones the supervisor did not create.
  • Never pass --ssh, which forwards the host ssh-agent into the container.
  • A container starts in about 1.1 s and exec is ready in about 100 ms, so recycling environments (§4.3) is cheap.
  • Devcontainer features as built (D38, issue #108; internal/oci, internal/devcontainer/feature; unverified end to end: the libraries and the Resolve and Stage integration are tested against a fake registry, and the build itself is the spike’s, not run here). Resolve reads each features entry in the order of the file with its options (a string is the version, an object its options as strings, false switches it off) and, for an image environment, fetches it on the host through internal/oci: an anonymous token from the registry’s own realm only, the manifest by tag or digest, its sha256 taken as the feature’s identity, the one feature layer by digest with a size cap (16 MiB), a timeout and a streamed hash, redirects only to https and public addresses and without the token, and an extraction that refuses absolute and .. paths, links, devices, duplicates and anything past 2,000 entries (files and directories alike), 64 MiB or 12 levels; a devcontainer.json that asks for more than 20 features is refused whole. devcontainer-feature.json is then read: privileged, mounts, capAdd, securityOpt, init, entrypoint, dependsOn and every lifecycle command refuse the feature (a false or empty value asks for nothing), and containerEnv may not set the proxy variables or any *_PROXY, LD_*, HOME or a WHR_, vendor or supervisor name (PATH may change). That refusal is not a security boundary, because install.sh runs as root and can write any file in the image it builds: it keeps a feature from setting the supervisor’s own variables or the proxy by accident. The boundary is that only an allowed source runs, below, and that the build runs in the builder VM (§7.2). A refused feature, or one from outside ghcr.io/devcontainers/features/ that the human has not allowed for the repository, is left out and named in the environment’s notes (a foreign one is also listed in Environment.ForeignFeatures, as written and with the manifest digest it resolves to now, to be asked about, issue #127); a registry that cannot be reached or a digest that does not match fails the resolution. Options are checked against the feature’s own declarations and written to a shell file with single quotes escaped; features install in file order except where installsAfter says, a cycle is refused and a missing dependency is noted, never fetched. Stage writes a context that holds only the features, extracted again, and a Dockerfile that installs each as root (_REMOTE_USER=root) on the environment’s image and then writes each feature’s containerEnv as ENV, which is how its tool reaches PATH. The image tag is keyed on the base image, each manifest digest and the options, so a moved tag changes nothing until the next resolution. The build has the builder’s network, outside the egress allowlist: the accepted risk of §7.2. A source outside the allowed one (issue #127). gateEgress raises one blocking approval per unanswered foreign reference, with the cause feature_source, the reference in the Decision’s own feature column (ValidFeatureRef: registry path with a tag or digest, plain characters only, never parsed from the subject) and the manifest digest it resolved to in its own feature_digest column (ValidDigest, sha256: and 64 hex), which the question text states as a supervisor fact; to learn it a foreign reference is resolved as far as its manifest, and no blob is fetched until an allow applies, next to the egress requests of the same run; the run stays starting until none is open (the same wait, AsksBeforeStart). The answer is kept per repository and reference with the digest (feature_sources, Store.SetFeatureSource). An allow applies only while the reference resolves to the digest the human saw: serve.Environment hands feature.Resolver.Approved the allowed references with their digests, and it is a function of the reference and the digest, so a tag that was moved to other bytes is not approved and is asked again, with the new digest; an allow given before the digest was recorded never matches. A deny stays per reference, whatever it resolves to, until the human changes it. Answering needs a fresh passkey assertion bound to the reference and the digest (D45), as for an egress host. The image is made when a workspace is provisioned, so an answer applies to environments made after it: the environment of the workspace the run started in keeps the image it has (open: whether an allow should rebuild it, proposed on #127). On a Dockerfile. With a build.dockerfile the features are a second build on top of the first: Stage builds the repository’s Dockerfile under BaseTag (the tag without the features), StageFeatures then writes the features’ Dockerfile FROM that tag, tagged as Tag, which is keyed on both. Apple Container’s builder resolves a FROM of an image in the local store when run without --pull (verified, spike/builder-store).
  • Cancel via in-guest launcher (whr-shim, D25): do not cancel by signalling the host container exec client; Apple Container fails with "failed to send signal: missing signal in xpc message" and leaves the guest process running. Instead, start commands under /tools/whr-shim run -pidfile ... (in their own process group) and cancel with container exec <id> /tools/whr-shim kill -pidfile ... -grace 500ms. Measured in spike #10: cooperative processes terminate via SIGINT in 10–15 ms (125 ms total exec roundtrip), stubborn trees ignoring SIGINT are killed via SIGKILL after grace in ~514 ms (600 ms roundtrip), exit code is preserved (130 on SIGINT), and zero orphan processes remain.

5.2 Agent adapter

Specified as explicitly as the runtime contract, and versioned: the contract carries a contract_version, and an adapter declares which version it implements. Release 1 targets Claude Code and Codex CLI as built-in adapters against it; as built, only Claude Code is composed in serve (internal/serve/real.go), and the Codex CLI adapter is issue #35 (§13, capability matrix). Codex CLI lacks mid-run injection and host-routed approvals in what spike #1 could test (approvals and cancel: spike #7; planted configuration: spike #68), so in release 1 it runs in the degraded mode below, labelled in the UI. Full mode needs every capability marked so; an agent without them runs degraded (D12 requires full mode only of Claude Code, the first agent). Capability flags:

  • headless / unattended operation
  • mid-run message injection (required for full mode): send a user message into a running session and report how it was delivered (injected now, or at the next turn). Without it an agent cannot be a remote-controlled assistant (§1); an agent that lacks it may only run in a degraded mode that the UI labels
  • structured event stream (required for every adapter): messages, tool calls, diffs and test results as typed events, which feed the live transcript (§9.3)
  • cooperative pause (e.g. stop after current turn); reported false by every measured agent (D11)
  • session persistence and resume
  • PR/issue tooling
  • “awaiting guidance” signal (how the agent raises a blocking Decision)
  • approval prompts routed to the host (required for full mode): the agent blocks on a permission request and the supervisor answers it with a human’s allow or deny and a reason (§4.2). An agent without it can only run with a fixed allowlist and every other action denied
  • auth modes, reported explicitly and never assumed:
    • api-key: the key stays in the host-side proxy and is issued per run (§7.3). Until that proxy exists, the key comes from a 0600 file on the supervisor’s side (the configuration’s agent_api_key_env_file) and reaches the agent through the runtime’s env file, never a command line.
    • subscription: a consumer-plan login (for example Claude or ChatGPT sign-in) kept in a dedicated per-environment auth directory. The CLI refreshes the token itself, so it cannot sit behind the proxy.
    • whr never handles a subscription credential (D40). The human signs in inside the environment through the vendor’s own flow, from a terminal attached to it, and the CLI writes the login to the agent-home volume, where it survives stop, start and rebuild (D16). whr does not read, copy, store, relay or log it, and does not record that terminal session. A subscription token found in the supervisor’s configuration is refused. Which flows work behind the egress sidecar is spike #82.
  • auth and quota blocking states: the adapter reports auth_expired and quota_exhausted (with the reset time when known). Each opens a blocking Decision and pauses the run instead of failing or retrying. Re-login is a UI action through a browser or device-code flow.

The Go contract (issue #20, internal/agent):

  • ContractVersion is a constant; Capabilities carries the version the adapter implements, the flags above, the auth modes and whether it reports quota. Mode() computes the mode from them: full needs headless, structured events, mid-run injection and host-routed approvals; an agent with events and headless only is degraded; anything less is unsupported.
  • Start and Resume take a StartSpec (environment, working directory, prompt, auth mode, permission mode, tool allowlist, approver and approval timeout) and return a Session. Events is a typed stream: message, tool call, tool result, diff, test result, usage, approval, auth_expired, quota_exhausted (with the reset time when known), result and error.
  • The session ID arrives with the session event. Session.ID() may be empty until then (spike #1: Claude Code reports it after the first message). Instruct and Stop are called after the session event, and the suite waits for it; a started session that never reports an ID fails the suite.
  • Permission mode and allowlist. manual routes every permission prompt to the host and needs an adapter with host approvals. dontAsk never asks: a tool on AllowedTools runs and every other is denied, and the approver is not consulted. A degraded agent (§5.2) runs in dontAsk; the allowlist is also what limits an agent that cannot be asked. An empty mode means manual. An adapter that cannot honour the mode refuses to start with ErrUnsupported. bypassPermissions and auto are not in the contract (§6).
  • The approval event is the record. An approval event carries the request ID, the tool, the allow or deny and the reason, and the audit entry (§5.4) is written from it. The suite asserts on the event and not on text the agent prints.
  • Stop cancels a pending approval. An approval that waits for a human is cancelled with the session (D23), so the approver’s context ends, the approval event records a denial and the result is stopped.
  • Usage events carry a typed payload (§5.7).
  • Instruct returns the delivery: injected (now), next_turn (after the running tool, as measured for Claude Code) or resumed_turn (degraded: the message becomes a resumed turn). An agent without injection never claims the first two.
  • Approvals are a host callback, and fail closed. The adapter blocks the agent on its permission request and asks the Approver. An error, a cancelled context or no answer within the approval timeout is a denial, and the agent sees the denial.
  • Cooperative pause stays a capability flag (D11): a session implements Pauser only if the flag is true, and the conformance suite checks both ways.
  • Auth and quota end a run without failing it. The session emits auth_expired or quota_exhausted and finishes with that status and no error, so the supervisor opens a blocking Decision (§4.2) instead of retrying; the session stays resumable.
  • Stop is a hard interrupt and leaves the session resumable. agenttest has a scripted fake agent and the conformance suite; the Claude Code adapter (#25) and the Codex adapter (#35) run it too.

The Claude Code adapter (issue #25, internal/agent/claude). It runs claude -p --input-format stream-json --output-format stream-json --verbose through the runtime’s Exec with Stdin (§5.1), as spike #1 did, and turns each output line into an agent event:

  • Session. One process is one run of turns: the prompt is the first user message, a mid-run Instruct is another line on stdin (next_turn, as measured), and the adapter closes stdin after the result event so the process ends. A further turn is Resume with --resume <id>. The session event comes from system/init, which only appears after the first message.
  • Stop cancels the Exec context. The runtime contract (§5.1) requires that this ends the process in the guest, so the adapter needs no signal of its own; the session stays resumable by its ID.
  • Events. assistant text is a message, tool_use a tool call, tool_result a tool result, thinking is dropped, and result ends the run and produces one usage event with the model, the reported cost, the token counts from result.usage (input, output, cache read, cache write) and the usage windows from the last rate_limit_event. A tool result with tool_result_meta[].non_execution_kind (recorded: permission-rule for a denial) is a tool that never ran, and is recorded as denied.
  • auth_expired comes from the error code authentication_failed on an assistant event, never from subtype or apiKeySource (§5.2, spike #1). The run ends auth_expired even though the result says is_error with subtype: "success".
  • Permission modes. dontAsk runs with --allowedTools and the CLI denies the rest silently; the adapter records each decision as an approval event (an allowed tool from its allowlist, a denial from the permission_denied system event). A tool that is not on the allowlist and returns an is_error result with no denial event is recorded as denied, never allowed, because that is also how a denial can come back; each approval ID is unique, including denials that match no tool use. manual runs the CLI with --permission-mode manual --permission-prompts host --permission-prompt-tool stdio (D26, issue #75) and the same settings flags, so every permission prompt arrives as a control_request (can_use_tool) on stdout. Each one is asked of the Approver in its own goroutine, with the tool and its input capped (a Bash command, a plan, or the tool’s input as compact JSON), and the answer goes back as a control_response carrying the request’s request_id, allow or deny with the human’s reason as the message. The adapter records every answer as an approval event, and it fails closed: no answer in time is a denial sent to the agent; a request of another kind is refused with an error response at once, and a second request with an ID that is still open is denied without being asked; if the answer cannot be written, because the channel is lost while the agent still runs, the allow is recorded as the denial it is and the agent is stopped (the runtime ends the process in the guest with whr-shim, D25); and when the output ends or the session is stopped, whatever is still open is denied and recorded before the events end, so no request outlives its process. HostApprovals is true. That an allow, a deny, a closed stdin and a wrong request_id behave this way is verified by the recordings of spike #7 (internal/agent/claude/testdata/recorded/, replayed against the adapter in approval_test.go); the shape of the error response to an unsupported request kind and the flags together with the supervisor’s settings flags against the real CLI are unverified. The service turns each prompt into a blocking approval Decision of the task and returns the human’s answer (internal/service/approval.go); an answer to an approval no agent waits for, such as after a pause or a restart, is refused (D23).
  • Only the supervisor writes what the CLI reads. The agent can write its own home (CLAUDE_CONFIG_DIR) and its repository, and Claude Code reads settings, hooks, MCP servers and skills from both. Measured with Claude Code 2.1.285 (branch spike/claude-config): with no flags, a hook planted in the repository’s .claude/settings.json and settings.local.json, one in the agent home’s settings.json, a project .mcp.json server and a project skill all took effect; with --setting-sources user the agent home’s hook still fired, so pinning to user was not enough. With --setting-sources "" (no source), --strict-mcp-config and a supervisor-written --settings (a file or an inline JSON string), none of the planted hook, MCP server or skill took effect, and the supervisor’s own hook did. The adapter passes those flags on every start and every resume, plus --disable-slash-commands (skills) as defense in depth. That planted permission rules are ignored too is inferred: the same loader reads them, but it was not shown without a model call.
  • What the CLI still allows on its own (§6): the built-in read-only allows (spike #1: pwd and git status ran in manual and dontAsk), the built-in plugins, and whatever --allowedTools and the supervisor’s --settings grant. Anything else is denied in dontAsk.
  • Bounded input. stderr is kept to a capped number of lines and bytes. A result that has arrived is kept when Stop comes late, and Resume reports ErrNoSession only for an error result that says the session is missing; any other early failure is an ordinary error.
  • Unverified, taken from the spike’s field names only: quota_exhausted is raised when a usage window reaches utilization 1 and rate_limit_info.status is anything but allowed (what rate_limit_event.status reads when exhausted was not observed); a Resume of an unknown session is recognized as an error result before any init and no login failure (the CLI’s real signal was not captured). Verified on real streams (spike #7, spike/agent-approval commit 22ddc7f, internal/agent/claude/testdata/recorded): the init and result shapes, the token counts, the rate-limit windows with epoch-second reset times, the denial marker, and session_id on every event. Still constructed from the spike’s description, because no recording exists: the expired login and the exhausted quota.

Measured in spike #1 (issue #1; branch spike/transcript, RESULTS.md), with Claude Code 2.1.285, Codex CLI 0.159.2 and Antigravity agy 1.1.12 on one machine:

CapabilityClaude CodeCodex CLIAntigravity
Headless, typed eventsYes (stream-json in and out)exec --json; only start and error events seenYes (--output-format stream-json)
Mid-run messageYes, picked up at the next model stepNot found in exec (unverified)No: one prompt per run
ResumeYes (--resume), same session IDexec resume exists, untested--conversation <id>, untested
Approvals to the hostYes, on the host through an MCP prompt tool; in a container, the stdio control protocol (D26, measured in spike #7)UntestedNone found in print mode; the run ends in ERROR on a denial
Usage windowStructured: five-hour and seven-day windows with reset timeText only, reset time inside the messageNot observed
CancelHard interrupt only, session stays resumableUntestedUntested

Findings that shape the contract:

  • A user message sent mid-run is delivered at the next model step, after the running tool finishes, not by interrupting it. The UI says so.
  • The session ID only appears after the first user message, and user messages are not echoed in the output, so the supervisor logs its own.
  • There is no cooperative pause; cancel is a hard interrupt.
  • Agents without streaming input (Codex exec, agy print mode) run in the degraded mode: a message becomes a resumed turn, labelled in the UI.
  • Token-level streaming is available from Claude Code (--include-partial-messages adds text_delta chunks, and the full message still follows). Coalesce deltas (about 150 ms) and let the final message replace them; deltas are live-only and never stored (§5.4).
  • A missing or expired login is signalled by an assistant event with error: "authentication_failed", then a result with is_error: true and subtype: "success". Detect auth_expired from the error code, never from subtype, and never from apiKeySource, which reads none both for a subscription login and for no login at all. A login that expires mid-session was not reproduced.

5.3 Reconciler

Desired state lives in the database. A loop compares it with actual runtime state, marks orphaned runs interrupted, and resumes from the agent session rather than the VM. This is the answer to Apple Container’s missing restart-policy recovery.

Measured in spike #2 (issue #2):

  • No restart policy. After the host-side runtime process of a container was killed, it stayed stopped, with its volume state intact, until started again.
  • A service restart ends everything. container system stop took 0.4 s and ended every container VM at once; after container system start (0.4 s) every container was stopped, including those that had been running. Nothing came back by itself. Root filesystems, volumes, bind-mounted data, networks (also custom --internal ones) and images survived; every process inside the containers was gone.
  • After a Mac reboot nothing starts the services either: there is no LaunchAgent or LaunchDaemon plist for them on disk. The supervisor’s own launchd job must run container system start --disable-kernel-install and then reconcile. The flag skips the interactive kernel-install prompt, so the kernel must already be installed once at setup (container system kernel set --recommended); without it no container starts. A reboot itself was not triggered.
  • Recovery loop. List containers, start those that should be running, wait for exec to answer (about 100 ms after start), then resume the agent from its session. Container IPs change on every start, so they are read again each time and never stored.

The Go service (issue #23, internal/service) is the one layer the JSON API and the web UI call (D8). Its reconciler is DB-first (D6) and takes the clock and the runtime and agent adapters as parameters, so tests run it on the fakes with an injected clock. One pass, step 1 once and the rest per active task:

  1. List the environments the runtime reports for this owner (List(owner), never every container). Then, before anything is observed, stop each running workspace environment that this supervisor’s store records and the current process did not start, the console’s excepted, whatever runs it holds; the stopped ones count as seen stopped in step 2, and a failed stop is tried again on the next pass (§4.1, “No surviving agent before a relaunch”, #221). The owner label is shared by every supervisor of the host, so an environment the store does not record is left alone. This sweep runs once per pass, not per task.
  2. For each environment of the task, ObserveEnv with what the runtime says; an environment the runtime no longer knows is observed as gone. starting and running runs in an environment seen stopped or gone, and a run paused by an agent’s cooperative pause, become interrupted and their open questions and approvals are superseded; a hard-paused run stays paused with its login or quota question (§4.1, #221).
  3. A run without a session is lost. A starting or running run that has no attached agent session is interrupted, even when its environment is still up: after a supervisor restart the environment survives, and so can the agent process, which the supervisor no longer holds (spike #7, Case 4). The run is interrupted, and when the current supervisor process did not start its environment, the reconciler then stops that environment, which interrupts every starting or running run in it (§4.1, “No surviving agent before a relaunch”, #216); an environment this process started is never stopped for this. A run whose launch is in progress is not lost; the service marks it before it calls the agent. After step 1’s sweep this stop is needed only when the sweep’s stop of that environment failed (a run’s environment is always recorded, because every launch saves it first).
  4. For each interrupted run, unless it waits for a reset it chose (“resume at reset”, §4.2): check that Resume can succeed (no open login or quota question) and that no other active run owns its environment before starting anything, so a run that waits for the human costs no environment; stop its environment first when this process did not start it (§4.1, “No surviving agent before a relaunch”), start it if it is stopped, wait until exec answers (polling with the injected clock, bounded), observe the environment running, move the run to starting with Resume and relaunch the agent with Resume(sessionID). The session ID is recorded on the run when the agent reports it (the session event), because the session, not the environment, is what survives. On success the run is running. A session the agent no longer knows (ErrNoSession), or a run that never reported one, ends failed at once, which opens a blocking retry-or-cancel Decision (§4.1). Any other error returns the run to interrupted and counts an attempt; when the attempts (3 by default) are used up the run ends failed too.
  5. Every resume starts with the briefing (D27): the first message to the agent says that the process ended, that a tool call that was running may have had effects that are unknown or partial, which Decisions were superseded (their tool and input), and that the agent must check the workspace before repeating anything.
  6. Resume the runs whose resume_at_reset Decision is due.
  7. A shutdown interrupts, it does not stop. Shutdown ends the sessions, and the runs they served become interrupted, so the next start resumes them; only a stop the human asks for (cancel) or an agent that finished ends a run for good.

The sessions map is keyed by run, and a finished session removes its entry only if the entry is still that session, so a relaunch during a pass is not lost.

The slice’s use cases (issue #95, internal/service) are what the API and the CLI call. Handlers hold no logic.

  • Start (whr run <issue-url> --agent <ws>/<role>, D42) takes an issue URL and an existing agent of an existing workspace; it never clones. In order: parse the URL and refuse a repository other than the workspace’s; load the issue through the forge (everything in it is untrusted data); apply the trust tier of #53 (a hook that today allows every issue and is the place to add the tiers, not a way around them); build the first prompt from the issue, marked as untrusted; start the workspace’s environment and wait until exec answers; save the task and its first run in starting together with the one-run check, in one transaction (§4.3); start the agent in the agent’s worktree; mark the run running and attach the session. Each step that changes state is an event of the aggregate (task.state, run.started, run.state, run.session). A start interrupted after the save leaves a starting run without a session, which the reconciler takes for lost and fails into a retry-or-cancel Decision (point 3 above); one interrupted before the save left nothing.
  • Trust tiers (issue #53). The author association GitHub reports for an issue is read as a tier: OWNER, MEMBER and COLLABORATOR are trusted, everything else (CONTRIBUTOR, FIRST_TIME_CONTRIBUTOR, FIRST_TIMER, MANNEQUIN, NONE, and anything missing or new) is untrusted: fail closed. whr run on an issue by an untrusted author starts nothing: it records the task, still queued, and raises a blocking question for the human (start or cancel) that shows the author, their association and the issue text, as untrusted data. start loads the issue again and starts the run only if what the human saw is still the start of it, otherwise it refuses and the human runs it again; cancel cancels the task. The task is marked as having untrusted input, and from then on the policy decision for it takes that context (Table.DecideIn): an action the table lets run on its own (auto) asks when it has an outward effect (push, open a pull request, comment) and the input is untrusted, never looser than the table. A run that combines private data, untrusted input and outbound network needs approval (§7.1): untrusted input is held for a Decision before the run starts, and the supervisor does not know whether the repository is private, so it assumes it is and always holds untrusted input.
  • A new run on an existing task (StartRun) serves rework after changes were requested and the answers of two questions: retry of a failed run, and rework of the rebase-conflict question of §4.2. It checks the one-run rule like Start and starts the agent with the briefing of D27 plus, for a rebase conflict, the conflicting paths as untrusted data. retry of a rebase conflict only frees the task: the human runs the export again.
  • Say sends a message to the task’s live run with the session’s Instruct, records an instruction.sent event (the run, how it was delivered and the text, redacted) and returns the delivery, so a degraded agent’s resumed turn is shown as that and never as an injection. A task with no attached session is a conflict.
  • List and Show read the store: a summary of every task, and one task with its runs, open Decisions and current revision.
  • Events (Subscribe) replays durable events of a task from a sequence number out of the store and then follows live ones. The agent’s observations (messages, tool calls and results, diffs, test results, usage) are stored as transcript-tier events, so they replay and have retention (§5.4); ephemeral events (token deltas, heartbeats) are published to subscribers from memory, redacted like stored ones, and never stored. The store is the replay buffer, so a client follows a history of any length at its own pace. Publishing never waits for a slow subscriber: one that falls behind loses the ephemeral events it missed and is caught up from the store, so it misses no durable event.
  • A session outlives the request that starts it. The agent runs for hours and the API call that started it returns at once, so the supervisor starts and resumes a session with a context that does not end with the request; Shutdown stops it and leaves its run resumable (found by the serve integration run, where the session died with its request). The supervisor also gives the agent process its environment for this start in StartSpec.Env: the egress proxy’s address, read from the runtime because it changes with every start and never stored, and HOME on the agent home volume. It holds no secret.
  • Credentials (D40): in API-key mode the key comes from config.AgentAPIKey() and reaches the agent through the adapter’s Env; in subscription mode nothing is passed and the human signs in inside the environment (spike #82, stubbed until it lands). The egress sidecar runs the installed libexec/whr/whr-proxy-linux-arm64.

Only the container’s state and the agent session ID are stored; the address Info.Addr is read for use and never written. Nothing in the reconciler sets a state by itself: it calls the aggregate.

5.4 Events, idempotency and retention

Per-task append-only event log doubles as audit trail, UI feed and CLI stream. Every mutating command accepts an idempotency key.

Retention. A chat grows with every message, tool call, tool result and diff, so the log has two tiers:

  • Audit entries are never purged: state changes, Decisions and their answers, usage records (§5.7), approvals (including the tool and a capped input), commits and PR links, credential issue and revoke, policy denials, permission-mode changes, and the record of every purge.
  • Transcript content is bulk and has retention: assistant text, tool inputs and results, diffs, thinking and attachments. Streamed token deltas are never kept durably, only the final message (§5.2).
  • Limits. A size cap and an age limit per task, with the cap and limit set by policy, and a manual purge from the web UI and whr purge (§9.3). Deleting a task purges its transcript.
  • A purge records itself. It deletes transcript content and keeps one audit entry: who, when, and what was removed (event count and bytes). Audit entries refer to transcript content by hash, so a purge leaves a verifiable gap and never silently rewrites history (§7.7).
  • The kill switch writes an audit entry too (§7.7). Service.KillAll (POST /v1/kill-all, whr kill-all once the CLI exists, #97) stops every live session, cancels every unfinished task, revokes the installation tokens the forge client holds (DELETE /installation/token, one per repository; the agent’s own credentials stay with the human, D40) and then writes supervisor.kill_all into the supervisor’s own stream: who, which tasks, how many tokens and what failed. It goes on after a failure, so one task that cannot be cancelled does not leave the others running, and the entry is written whatever happened.
  • The audit entries form a hash chain, indexed by commit (§7.7). Every audit-tier event gets one row in audit_chain, written in the same transaction: a SHA-256 over the event (sequence number, task, kind, tier, time and the stored payload, each length-prefixed) and the hash of the entry before it, so changing, removing or inserting an entry breaks every hash after it. The rows are append-only like the events (triggers refuse an update or a delete) and a purge of the transcript leaves them alone. VerifyAudit recomputes the chain and fails at the first broken entry; an entry written outside Append has no row and is caught too. A chain cut off at its end is a shorter valid chain, so AuditHead (sequence number and hash of the last entry) is what the human records somewhere the supervisor cannot write, and VerifyAudit takes it to check the end. Audit entries from before the chain (migration 0010) have no row and are counted as unchained, not covered. An entry whose payload has a sha field of a full lower-case commit name (the pin, the push, the CI result, the pull request, a review Decision and its answer) also records it in the chain row, so CommitAudit returns the trail of one commit, oldest first; the SHA is part of what the chain checks.
  • The agent’s own session is separate. A purge does not touch the session the agent resumes from; shrinking the agent’s context (compaction or a new session) is a different action with its own consequence, the agent forgetting, and is not offered as a purge.
  • Redaction at ingest. Secrets are redacted before anything is written: events, Decision inputs and audit entries alike (§7.3). Audit entries are never purged, so redacting later would be too late. What remains is treated as untrusted data when shown.
  • Live-only events. Token deltas and heartbeats go to connected clients through an in-memory fan-out and are never written. --since (§9.2) replays durable events only; a client that reconnects mid-message gets the final message when it is written.

The store (issue #21, internal/store):

  • SQLite in WAL mode through the pure-Go driver modernc.org/sqlite, so the supervisor stays one static binary (D3) that cross-compiles to Linux. Foreign keys are on. Migrations are embedded SQL applied in order and recorded in schema_migrations; a database newer than the binary is refused.
  • Tables: tasks (saved together with their runs, environments and review candidates), decisions, events and idempotency keys.
  • Versions and compare-and-swap. A task aggregate and a Decision each carry an integer version. A save writes WHERE version = expected and bumps it; no row updated means another writer got there first, reported as a conflict (exit code 5). An answer and an expiry of one Decision cannot both win, and neither can two commands on one task.
  • One transaction. The domain methods record the events a change produced (TakeEvents), and the store writes the new state and those events together, so there is never a state without its audit entry or an audit entry without its state.
  • Events are append-only with a monotonic, never reused sequence number, which is what --since (§9.2) replays. Each has a tier. Triggers refuse any update and any delete of an audit row; the only way rows leave is Purge, which deletes transcript rows and appends one audit entry (who, when, how many events and bytes) in the same transaction.
  • Idempotency. A mutating command runs under its key: the response is stored in the same transaction as the changes, so a replay of the same key and request returns the stored response without doing anything again, and the same key with a different request is refused as a conflict.
  • Redaction is on by default (issue #22, internal/redact). The store redacts before it writes: event payloads (audit and transcript), the text of Decisions (subject, input, reason, answer, actor), task and candidate text, stored idempotency responses and the actor of a purge. What is held in memory is not changed, only what is persisted. Idempotency keys are opaque identifiers chosen by the client and are stored as given, so a client must not put a secret in one.
  • What the redactor knows. Exact secrets registered for a run (the scoped tokens the credential service issued, §7.3) are replaced in every form they may take in text: raw, JSON-escaped, URL-escaped and base64. Well-known token formats are replaced whether or not they were registered: GitHub, GitLab, Anthropic, OpenAI, AWS, Slack and Google keys, JWTs, private key blocks, bearer tokens, passwords in URLs and values of secret-looking names (token, password, api_key, authorization). A numeric value, such as a usage count named input_tokens, is left alone, because usage records must stay correct (§5.7). The same redactor wraps log output and slog records, so a token is kept out of logs by the same rules.
  • A floor, not a proof. A secret in a format the redactor has never seen, and not registered, passes. A canary test therefore pushes a token of every kind through every write path, closes the database and scans the files, including the write-ahead log, for it; a control run shows the scan finds a token that was not redacted.

5.5 Adapter plugins

New agents (and later runtime or forge backends) are added as out-of-process plugins, not in-process code. A plugin is a separate executable that speaks the versioned adapter contract (§5.2) over stdio or a local socket (JSON-RPC style). Go’s in-process plugin package is not used: it is fragile and would put third-party code inside the supervisor.

  • Release 1: the contract is the design; Claude Code and Codex CLI are built-in adapters against it. No loader.
  • Medium term: a plugin loader, once two built-in adapters have proved the contract. Whether an existing agent-client protocol (for example Zed’s ACP) already covers part of the contract is unverified; check it in the §12 scorecard and reuse it if it fits.
  • Conformance. A plugin declares its capabilities and must pass the same conformance suite as a built-in adapter, so a capability flag is a verified claim (§5.1).
  • Wire types. Everything that crosses the seam has a stable JSON name (snake case): capabilities, the start spec, events, results, approval requests and answers. Times are RFC 3339 and a zero time is left out. An error crosses as a string code (unsupported, unsupported_auth, no_approver, no_session, not_running, bad_spec), and agent.ErrorFor maps a code back to the sentinel, so errors.Is works on the supervisor’s side of the seam.
  • The approver is a reverse call. The approver cannot be a Go callback across a process. The plugin sends an approve request (an approval request) to the supervisor on the same connection and waits for the answer. The supervisor applies the approval timeout and the fail-closed rule (§5.2), so a lost connection, a plugin that dies or a late answer is a denial, and the plugin never decides.
  • Trust. See §7.8: plugins are installed explicitly and run isolated.

5.6 Tool store

Environments run stock images, or, for a repository without a devcontainer, the workharbor base image of D44: a first-class base plus git and CA certificates, built once by whr. Fedora and Ubuntu LTS are the first-class bases, pinned by digest and covered by the conformance suite, with Fedora the default; other bases work best-effort (D43). The agent CLIs (Claude Code, Codex CLI, later others) and the supervisor’s own helpers live once in a versioned, immutable tool store on the host and are mounted read-only into each environment, in the manner of a Nix store. This replaces installing an agent in every container, which took about 11 s and 230 MB each in spike #2.

  • The base image (D44, issue #98). internal/baseimage ships one Containerfile per first-class base in the binary (Containerfile.fedora, the default, and Containerfile.ubuntu): the stock image pinned by digest, then git-core and ca-certificates (Fedora) or git and ca-certificates (Ubuntu) and nothing else, with no user and no copied file. Its tag is whr.invalid/whr-base/<base>:<12 hex of the Containerfile's SHA-256>, so it stays the same until the pin changes. whr serve calls baseimage.Ensure at start when environment.image is unset: if the runtime has no image with that tag (HasImage, container image inspect), it builds one through the same runtime.Builder as a repository’s image (an empty context), so a later start builds nothing. environment.base (fedora or ubuntu) picks the base; environment.image replaces the base image altogether. verified on container 1.5.0 for both bases: go test -tags applecontainer -run TestBaseImageLive ./internal/baseimage builds the image, finds it the second time, and runs as uid 1000 git worktree add, git bundle create and an HTTPS git ls-remote.
  • An image without git is refused at workspace creation. After the environment starts, Workspaces.Create runs git --version in it. A non-zero exit is refused, with the image name and what git said, before any worktree step; the environment, its network and its home volume are taken back. The check applies to every image: the base image, environment.image and a repository’s own (#76).
  • The console image (D43, issue #92). internal/console ships the console’s image definitions next to the base image’s: one Containerfile per first-class base, the same stock image pinned by the same digest, with git, zsh, fish, jq, curl, ripgrep, tmux, make, less, an editor (vim-minimal or vim-tiny), util-linux with script, a non-root user whr (uid 1000) and the git wrapper as /usr/local/bin/git. It has no Docker CLI and no container engine and holds no credentials. The tag is whr.invalid/whr-console/<base>:<12 hex of a SHA-256 over the Containerfile and the wrapper>, built once through the same builder (baseimage.EnsureImage, with the wrapper as the only file of the context). verified on container 1.5.0 for both bases: go test -tags applecontainer -run TestConsoleImageLive ./internal/console. Alpine (issue #93). The console has a third image, Containerfile.alpine (Alpine 3.24.2, pinned by digest; apk for the same tools, adduser for the user, and a password field of * because an sshd built without PAM refuses a locked ! account even for a certificate). It is for the console only: console.base in the configuration picks fedora, ubuntu or alpine (default environment.base), and agent environments stay on a glibc base until the agent CLIs’ musl builds are verified. console.Alpine is not a baseimage.Distro, which a test pins. Measured (verified, container 1.5.0): the console image checks (TestConsoleImageLive with WHR_TEST_BASE=alpine: the tools are there, no engine, and the git wrapper stops the planted hook, fsmonitor, filter, textconv and alias under busybox sh and musl git), the runtime conformance suite, TestConsoleSSHLive and the three TestTerminal tests (the session kill with pkill -s) all pass on it; the suites ran on an image built by hand from the same files, since the tests that take WHR_TEST_IMAGE do not build one.
  • Images whr builds are named under whr.invalid/ and never pulled (rule 4a, issue #133). The environment, base and console tags are whr.invalid/whr-env/<owner>:<hash>, whr.invalid/whr-base/<base>:<hash> and whr.invalid/whr-console/<base>:<hash>; runtime.BuiltImageHost is the one definition of the prefix. whr.invalid is an RFC 6761 name no registry answers, so the name cannot be published elsewhere. container create has no never-pull option (only container build has --pull), so Provision asks container image inspect first and refuses a missing whr.invalid/ image with “built by whr and is not here: build it” before it creates a network, a volume or a container. Measured on container 1.5.0 (verified, TestMissingBuiltImageIsNeverFetchedLive): container create of a missing whr.invalid/ image fails after about 10 seconds with “failed to resolve either repository hostname”, while a missing bare name goes to registry-1.docker.io; container build with a missing FROM whr.invalid/... fails the same way, closed, in about 12 seconds. The egress sidecar’s image is checked the same way, in Provision and in UpdateEgress (it is the environment’s own image, or the console’s, when built), before any network contact. A repository’s own image is not rewritten: one under whr.invalid/ is refused (a scan of the file’s text and of build.args, see the D38 bullet: defence in depth, and a name assembled from ARG pieces is not caught; the builder does use a local whr.invalid/ image as a FROM, so the scan matters, verified by spike/builder-store; a # syntax= frontend other than docker/dockerfile is refused too), any other keeps the CLI’s behaviour. The fake runtime still accepts a missing image. Images under the old names (whr-env/, whr-base/, whr-console/) are orphaned: nothing deletes them and whr builds under the new tag on next use.
  • The console’s git wrapper. A workspace’s .git is written by agents, so the configuration of a repository is hostile input to whoever runs git in it (spike #89: a hook, fsmonitor, a clean filter, a textconv driver and an alias all ran). whr-git runs the real git with core.hooksPath=/dev/null, core.fsmonitor=false and protocol.ext.allow=never, and overrides by name every setting git uses to run a command: core.pager, core.editor, sequence.editor, core.sshCommand, core.askPass, core.gitProxy, pager.*, filter.*.clean, .smudge, .process and .required, diff.external, diff.*.textconv and .command, merge.*.driver, mergetool.* and difftool.*, alias.* (to a static refusal), credential.helper, gpg.program, uploadpack.packObjectsHook, remote.*.uploadpack, .receivepack and .vcs, trailer.*.cmd and the browser and man commands. The overrides go through GIT_CONFIG_COUNT, GIT_CONFIG_KEY_n and GIT_CONFIG_VALUE_n, which win over every file. A key is trusted only if the console’s own system or global configuration sets it too, and then the override is that value, so the human’s own aliases and helpers keep working; the origin of a setting is never parsed, because an include gives a repository the choice of the path. Keys read first-value-wins (remote.*.uploadpack) go into a temporary file git reads as its system configuration, which comes before the repository’s. It is POSIX sh, because Ubuntu’s /bin/sh is dash. The tests plant each kind in a real repository and run it under plain git, which must run it, and under the wrapper, which must not; pager, editor and credential.helper need a terminal or a helper, so their value is read instead, and no test runs a credential command.
  • The console environment (D43, issue #92). service.Consoles opens, reuses and closes one console: Open runs the optional image build (console.Ensure, on first use, so whr serve starts without it), prepares and provisions serve.ConsoleOptions.For’s spec and starts it. The spec is hardened like an agent’s environment (non-root uid 1000, read-only root, cap-drop ALL, --init, tmpfs /tmp and /run, an internal network of its own behind the egress sidecar), with every workspace root mounted read-only at /workspaces/<root> (two roots with one base name get a numeric suffix), each workspace requested read-write mounted over its place in its root, and the home volume <owner>-console-home at /home/whr. There is no tool store, no agent and no secret: the console reaches the hosts of console.egress_allow (the package registries of the supported toolchains and github.com for read-only fetches by default; never the model API) and nothing else. The console is found by two labels (whr.console, and whr.console.rw with the sorted IDs of its writable workspaces), not in the database, so the reconciler, which looks at the environments of tasks, never touches it. The same writable set reuses an open console (a stopped one is started); another set is a conflict, because changing the mounts of an environment under open shells pulls the floor from under them, so it is closed first. Close stops and deletes the environment, its network and its sidecar, and keeps the home volume, which holds the human’s own dotfiles and history. verified on container 1.5.0: the runtime conformance suite passes on both console images (WHR_TEST_IMAGE=<console image> go test -tags applecontainer ./internal/runtime/apple, about 100 s each). Consoles.Shell opens a login zsh in the open console as whr (uid 1000) in the workspace’s directory, with HOME, USER, SHELL, LANG, WHR_CONSOLE, the client’s TERM if it is a plain terminal name (anything else becomes xterm-256color, because it comes from the client) and the egress proxy’s variables, through runtime.TerminalAdapter, and writes the supervisor.console audit entry; at most eight shells at once, and CloseShells hangs them all up when the supervisor stops. The API route and the whr console command carry it as a terminal stream (§9). The guest’s command ends with the terminal: closing the master hangs the guest’s terminal up, which the live test shows ends a sleep within seconds.
  • Rebuilding a workspace’s environment (issue #128). An allowed feature source, a new devcontainer.json commit or a bumped base image takes effect only when an environment is next built, so Workspaces.Rebuild (whr ws rebuild <workspace>, POST /v1/workspaces/{workspace}/rebuild) replaces a long-lived workspace’s environment. It keeps what the agents have: the workspace folder, so every worktree and branch, and the home and build volumes. The container, its network and its egress sidecar are new, with the hosts the repository has been allowed, as at provision (prepareSpec and bringUp, the two halves provision is now made of). The order is the safe one: refuse while any run of the workspace is unfinished, naming it and its state (Store.UnfinishedRuns: starting, running, paused and also interrupted, which the reconciler would resume in the old environment), and mark the workspace as being rebuilt under the lock a run’s start takes (saveStartingRun), so neither can pass the other’s check and a task started meanwhile is refused; saveStartingRun reads the workspace record again under that lock and refuses a start whose record names an environment a rebuild has since replaced; every other path that would start or use the environment refuses with a conflict while the mark is set (Service.refuseWhileRebuilding: the start of a task, export and prepare, AddAgent, Rebase), and the reconciler skips a run of a workspace being rebuilt and takes it up on its next pass, so the old environment is never started during a rebuild; build the new image and check the spec first, so a repository whose image cannot be built costs nothing; stop the old environment, because a volume is held by one running environment at a time; provision the new one on a network of its own name, start it, wait for it and check it has git; switch the record and write the audit entry in one change (Store.SwapWorkspaceEnv, a compare-and-set on the old environment); and only then delete the old container, its network and its sidecar. If the new environment cannot be brought up, or the record cannot be switched, the new one is taken back and the old one is started again, as it was (it is left stopped if it was stopped). The audit entry workspace.rebuilt names both environments, both images and the digest the runtime recorded for each (runtime.Info.ImageDigest, from the container’s image descriptor). verified on container 1.5.0 for what the runtime must do (TestAnEnvironmentCanBeReplacedOnTheSameVolume: a second environment on the first’s volume while the first is stopped, the volume’s contents kept, the first removed with its network). That test found that the helper container ownVolume runs could stay behind and keep the old network from being deleted; the helper is now removed with retries and again when its environment is deleted. A kept home volume meets the same user: the user of an environment is fixed by the supervisor (1000:1000, in SpecOptions.For; TestTheUserOfAnEnvironmentIsFixedSoAKeptVolumeMeetsTheSameUID pins it), so the ownership helper, which runs only on a new volume, is not needed again; a change of that user would need the helper to run on a kept volume’s root first (it does not recurse). A path that starts or uses an environment (a task start, an export, a rebase, adding an agent) takes a lease on it under the same lock, so a rebuild is refused while one is under way and a lease is refused while a rebuild is; the rebuild reads the workspace record again inside its critical section and works from the fresh one, never from a stale environment ID. A rebuild does not re-run a repository’s postCreateCommand (its marker is in the agent home, which is kept); that is open.
  • The host’s own IPv6 prefixes are refused (issue #124, §7.2). serve.HostIPv6Prefixes reads the global-unicast IPv6 prefixes of the host’s interfaces (not loopback, link-local, unique-local or IPv4; at most 16) whenever a spec is built, runtime.Egress.DenyPrefixes carries them, runtime.Prepare refuses one that is not an IPv6 prefix or is wider than /16, and the Apple adapter passes them to whr-proxy -deny-prefixes. egress.Proxy.Deny makes the proxy refuse an address inside one with ErrNotPublic (logged as a denial), both on the resolved answer and in the dialer’s control hook; egress.Public is unchanged. The list is read when a spec is built, so it goes stale: after the ISP renumbers or the host changes networks, a running environment or console still refuses the old prefix and allows the new LAN until it is rebuilt (whr ws rebuild <workspace>, issue #128) or the console is closed and opened again (whr console --close); a refresh of a running sidecar is not built. The list fails open: when the interfaces cannot be read, or there are more prefixes than the cap, the proxy still applies the private ranges and the prefixes it got, and the supervisor logs it (with the count dropped over the cap). A reported prefix wider than /16 is dropped and logged rather than passed to runtime.Prepare, which would refuse every spec. A /128 interface mask (DHCPv6, utun) would refuse only the host’s own address and not its LAN unverified: not measured on macOS. Tested with fake prefixes and a fake interface list; not run against a sidecar with real global IPv6 unverified, because container 1.5.0 gave a container only a link-local IPv6 address.
  • The allowlist matches exactly (issue #123, §7.2). egress.New keeps two lists: exact host names and *.suffix wildcards. Allowed admits a host that is an entry, or that lies below a wildcard entry and is not the bare name; an IP address never matches; an entry with a * anywhere but a leading *. is dropped, so it allows nothing. The wildcard form is accepted in environment.egress_allow and console.egress_allow (config.validEgressEntry: a host name or *. and a host name with at least two labels, so *.com is refused) and by runtime.Prepare for the sidecar’s list; a repository’s request never is: devcontainer.json entries with a * are ignored with a note that names the wildcard, domain.ValidHost refuses one for a Decision, and the lockfile suggestions are fixed exact names. The built-in console list is six exact hosts, each with its reason in config.DefaultConsoleEgress, and the default agent list is api.anthropic.com. The sidecar also passes on the bytes a client sent with its CONNECT (issue #124).
  • The build volume (D39, issue #80). Every workspace environment has a second volume, <owner>-build-<workspace>, mounted read-write at /var/whr/build (serve.GuestBuild), owned like the agent home (Owns), created and deleted with the environment. Each agent’s tools write their output there, in /var/whr/build/<role>, and not in the bind-mounted checkout, where file metadata costs 20 to 100 times as much (spike #40): the run’s environment sets CARGO_TARGET_DIR to …/<role>/cargo-target and UV_PROJECT_ENVIRONMENT to …/<role>/venv (Workspaces.buildEnv; the post-create commands get the same), and adding an agent makes its directory (mkdir -p), for tools that do not make their parents. Removing an agent clears its directory (rm -rf of a direct child of the build directory, only; a failure is reported, not returned), so a role added again does not inherit the old output. It is per agent, because two agents of one workspace work on different branches and must not share a virtual environment; the Go caches already live on the agent home. WorkspaceConfig.BuildDir is empty in a supervisor without the volume, and then nothing is set, because the paths would not exist on a read-only root. unverified with a real cargo and uv until wh/verify runs them; the volume itself is the conformance suite’s ownership and writability checks. The size of the checkout. When a run starts, Workspaces.noteRepoSize counts the files its worktree tracks, in the environment (git ls-files -z, because the host never runs git in a workspace), and records it as a task.repo_size audit entry in the task’s events (domain.RepoSize: the run, the count, a warning and its text above 50 000 files, D39); a count that cannot be made is reported and never fails the start. The supervisor’s git commands in the guest (gitEnv: this count, the worktree, the rebase, the bundle export) run with core.hooksPath=/dev/null, core.fsmonitor=false and protocol.ext.allow=never through GIT_CONFIG_COUNT, because a planted core.fsmonitor runs on git ls-files (measured, git 2.54) and a .git written by an agent is hostile input; a test plants a hook and an fsmonitor and shows plain git runs both and the environment neither. Not built, and why (issue #80): the mount-over path for node_modules is blocked by two facts measured or read from the code. git worktree add refuses a directory that already exists and is not empty (git 2.54: fatal: '…' already exists, also with --no-checkout), and the runtime makes the mount point in the host folder when it creates the container, before the worktree exists; and a container’s mounts are fixed when it is created, so an agent added later gets none. Both need a decision (D39, D42) before it is built.
  • The console’s SSH certificate authority (issue #32). internal/sshca is the user authority the console’s sshd will trust, and nothing else: no password, no authorized_keys. Its Ed25519 (or ECDSA) key is a secret file (0600, one link, owned by the user, outside every root) read with config.ReadSecret, registered with the redactor, and never in a guest; only its public key, as an authorized_keys line, goes into the console. Generate makes the key with an exclusive create and never overwrites one. Sign turns a client’s public key into a certificate for one session: one principal (a plain user name), a life of ten minutes by default and an hour at most, started a minute in the past for clock skew, a random serial, a key ID that is cut to printable characters for sshd’s log, permit-pty only, and permit-port-forwarding only when asked for (the editors’ remote modes need it). It refuses a certificate as input, an RSA key under 3072 bits and any key type it does not know. The client keeps its private key, so the supervisor never sees one. The console’s sshd. The console images install openssh-server and ship whr-sshd and sshd_config (in the image tag’s digest). sshd runs once per connection, in inetd mode (sshd -i), as the console user, started by the supervisor with exec and carried over its API, so the console listens on no port and SSH is reachable exactly as far as the API is (nothing new on the VPN). It is not root, because the console drops every capability. The configuration trusts one thing: TrustedUserCAKeys is the authority’s public key, which whr-sshd writes from WHR_SSH_CA on every start (so a file changed inside the console is overwritten); AuthorizedPrincipalsFile allows the principal whr; AuthenticationMethods publickey only, no password, no keyboard-interactive, AuthorizedKeysFile none, no root, no agent or X11 forwarding, no tunnels, no user environment. Port forwarding, where a certificate grants it, reaches only the console’s own loopback; unix sockets are not forwarded. StrictModes is off because /tmp (mode 1777) would fail it, and the files are the console user’s own and rewritten per connection. The host key is made once in the console’s home volume (whr-sshd hostkey prints it), so the console keeps its identity and a client can pin it: the private key is made aside and linked in by one ln, so of several first connections at once one wins, and the public half is derived from the winner with ssh-keygen -y, so the pair always belongs together. Every file sshd or a login shell reads from /tmp (the authority key, the principal file, the proxy variables) is replaced by a rename, never rewritten in place. The service and the API. Consoles.SSHCertificate signs a client’s public key (one line, no certificate, a key sshca accepts) for the principal whr with a key ID of whr-<actor>-<time>, asks the console for its host key (whr-sshd hostkey, which must be one Ed25519 line) and audits the certificate (supervisor.console_ssh, which never holds a key). Consoles.SSH starts whr-sshd in the console with exec, as the console user, with WHR_SSH_CA (the authority’s public key), HOME and the egress proxy’s variables in the environment and the connection as its standard input; it returns a stream of bytes, ends the sshd on close and reports a bad exit with the tail of its log (-e). At most eight connections at once; none without an authority or without an open console. The API carries it unframed after an upgrade (§9.7), and whr ssh is its client. Measured (verified, TestConsoleSSHLive, go test -tags applecontainer, container 1.5.0, Fedora console image): a certificate of the authority for the principal whr gets a shell as whr in the read-only root with the proxy variables the supervisor passed on; a certificate of another authority, one for another principal, a bare key and a password do not get in; a certificate without the forwarding permission forwards nothing, and one with it cannot reach another host; a client that pins another host key refuses the console, and the host key is the same on every call, also for six first calls at once. The Ubuntu image is unverified until the same test runs on it.
store/<hash8>-<name>-<version>-<platform>/bin/<name>     content-addressed, never modified
profiles/<profile>/bin/<name> -> ../../../store/.../bin/<name>
  • Verified on download. Claude Code is fetched from the vendor’s release URL and checked against its SHA-256 manifest; Codex CLI is the static musl build from its GitHub release. The store hash is recorded in the run’s audit entry, so every run says exactly which tool version ran it.
  • One build per libc. The glibc build of Claude Code ran in fedora, debian and ubuntu and failed in alpine; its musl build ran only in alpine; Codex’s static build ran in all four. The adapter picks the profile from the image’s libc (the vendor installer does the same, by looking for the musl loader). A tool that needs shared libraries beyond libc would need its closure in the store, as Nix does; none of the tested tools did.
  • Immutable from inside. The store is mounted read-only; touch, rm, appending to a tool, chmod and replacing a symlink all failed, and re-hashing every entry afterwards showed no change. Writable state (the agent home, with its auth directory and session) stays on the per-environment volume (§4.4).
  • Versions are profiles. Environment A ran Claude Code 2.1.286 and environment B 2.1.285 at the same time, each resolving to its own store entry. An upgrade is a new entry and a profile change, and a rollback is a profile change.
  • Mounting. A read-only bind mount and a read-only volume both worked, with the same startup cost (about 100 to 130 ms for claude --version), and four containers shared one store at once. A volume can be attached read-only by several containers but is exclusive while writable (§4.4), so the shared store is read-only everywhere.
  • Network. The agent starts without network, so the egress allowlist (§7.2) no longer has to permit the download host.

Building the store (issue #74, internal/toolstore, whr tools build -store <dir> -shim <whr-shim>). The pins live in the repository (internal/toolstore/pins.json: name, version, platform, release URL and SHA-256), so a version change is a reviewed commit and never a download-time decision. A download must match the pin and the SHA-256 the vendor’s own manifest.json lists for the platform (the manifest of Claude Code 2.1.285 does, and its linux-arm64 entry equals the hash spike #7 measured); a mismatch in either is refused before anything is stored. Antigravity is different: its pin holds the https archive URL, the archive’s SHA-512 and the extracted file’s SHA-256, and at run time only that pin is trusted, so the build fetches no vendor manifest; the vendor’s manifest and sigstore provenance are checked only when a pin is written or updated (scripts/antigravity-pin-check.sh), never by whr tools build. The build holds Claude Code alone unless -tools names more, so Antigravity is opt-in (an unknown name, or one with no pin for the platform, builds nothing). The binary is placed content-addressed at store/<hash8>-<name>-<version>-<platform>/bin/<name> through a staging directory and a rename, and the entry, its bin and the tool are made read-only (Verify re-hashes every entry and reports a changed or writable one). whr-shim is added from the launcher built from the same commit, and a profile profiles/<tool>-<version>[-<tool>-<version>...][-musl]/ (each tool in the order given, -musl last for a musl build, so Claude Code alone is claude-<version>) links to the entries with relative links, replaced in one move, so an upgrade or a rollback is a profile change. The vendor’s manifest signature (manifestSignatureEnforcement) is not verified here: unverified.

Verifying the store (issue #125). Every entry records the tool’s full SHA-256 in a read-only sha256 file next to it, because the eight digits in the entry’s name are 32 bits and a label, not a check. Store.Verify hashes each tool, opened without following a link, and compares it with that record; a tool with a built-in pin is compared with the pin, and a directory belongs to a pin by its name, version and platform whatever its first eight digits are, so a look-alike with another hash8 is severe, which is compiled into the binary, because the record and the tool are written by the same user and whoever changes one can change both; only a tool without a pin, one added from a file, is compared with the record; with neither only the name’s digits are checked, which is reported as a problem that is not severe. It also refuses (severe) a tool that is writable, that is a link or not a regular file, an entry that is a link or not named <hash8>-<name>-<version>-<platform>, a record that is a link or does not start with the name’s hash, and a tool that has gone. It verifies the profiles as well: each link in a profile’s bin must be the relative form Profile writes (../../../store/<entry>/bin/<tool>) into an existing entry whose tool is the link’s own name; a profile and its bin must be real directories, not links, and each link must resolve (every link followed) to exactly its entry’s bin/<tool> under the store’s root; a pinned tool name may link only to a pinned entry of that name, at the pin’s own entry, so a link to an unpinned version or another hash is severe; an absolute or escaping target, a regular file in place of a link and a link to a missing entry are severe. The store and profiles directories themselves must be real directories, not links (an absolute link would resolve in the guest through the image’s rootfs); the root may be a link. whr serve runs it at start and refuses to start on a severe problem, because every environment mounts the store and runs what is in it (an unchecked entry is logged); whr doctor’s tool-store step fails on one. Verify runs at whr serve start and in whr doctor, not while serving: a change after the start is not seen until the next check. The store is mounted read-only into an environment, so this guards the host’s side: a changed file, a restore from a bad backup, another process of the host user.

The configuration file (issue #74, internal/config) is JSON, decoded strictly (an unknown or misspelt key is an error), so the standard library is enough; TOML would add a dependency for comments only. It holds the repositories (owner/name and clone_depth), the roots (one or more workspace roots, on the internal disk or an external SSD, and the tool store; the repository cache root is gone with D42), the GitHub App ID and its key file, the agent-login env file, the API token file and the listen address. Secrets are paths to files, never values: each must be a regular file, not a link, owned by the user, with mode 0600 and some content, and outside the workspace root, where an agent could read it. The listen address must be a loopback IP (D29: a guest reaches every other address of the host); the roots must exist and must not overlap, since a workspace root is where an agent writes, and a root may not lie inside a git repository. Two optional keys (issue #97): state_dir, where the supervisor’s database lives (default ~/.local/state/whr; it may not overlap a workspace root), and environment, the supervisor’s choice of what an agent environment is: the stock base image (Fedora by default, provisional until pinned by digest, D43), the host names the agent may reach through the egress proxy (api.anthropic.com by default; plain DNS names only, no IP address, wildcard or port, because the proxy refuses a raw IP) and the CPU, memory and disk limits. CheckWorkspacePath decides whether a folder may become a workspace (issue #90): an existing empty directory strictly below a workspace root, resolved and compared by identity, and not inside a git repository, because a human’s own repository is never mounted. whr serve validates all of it at start and reports every problem with its key. The full onboarding stays in #29.

Open: how new versions are discovered (a developer bumps a pin in a commit), and how the store is garbage collected.

5.7 Usage and cost

A remote for agents that run detached for an hour has to say what they used. The supervisor records usage per run and reports it; it never meters the model traffic itself.

  • Source: the agent’s own reports. Each turn becomes a usage event with a typed payload: model, token counts, the cost (reported by the agent) and the balance when reported, and the usage windows (name, utilization from 0 to 1, reset time). The token counts are optional: a missing block means the agent did not report them, which is not the same as zero, so a token budget never reads an unknown count as 0 tokens. An adapter that reports usage sets ReportsUsage, and the suite checks that the events are well formed. Claude Code’s result event carries total_cost_usd and usage with input, output, cache-read and cache-creation tokens (recorded in spike #7). Antigravity sends usage in step_update and result; Codex CLI is unverified.
  • Reported, never estimated. Every cost is the agent’s own figure (reported), and so is every other number: tokens, windows and the balance. workharbor does not price tokens from a table of its own, because such a table goes stale. A turn the agent reports without a cost has tokens and no cost, and is counted as such.
  • Auth mode decides what the number means (§5.2). In api-key mode cost is real spend. In subscription mode it is API-equivalent, not billed: the plan is paid flat, and the usage-window utilization (five-hour and seven-day, §5.2) is what limits the developer, so the UI leads with that.
  • One usage window per account (D40). Concurrent runs on one subscription share its usage window, so the UI shows the window as one account-wide figure across all running tasks, not per task, and a quota stop pauses every run on that account with one Decision each (D23).
  • Bounds. A turn’s counts are validated at ingest: no negative figure, and no token count, cost (micro-USD) or duration (ms) above 10^11, which is far beyond a real turn and keeps a column of 9 x 10^7 maximal turns below the int64 limit, so the SQL SUM over the never-purged rows cannot overflow. The Claude adapter drops an out-of-range figure as unreported, and the Go sums saturate rather than wrap.
  • Kept as audit entries. Usage rows are small and survive a transcript purge (§5.4), so totals stay correct after history is deleted. Each turn is one audit-tier event usage.recorded (domain.UsageRecorded: run, repository, agent, auth mode, model, tokens, cost and windows), which the service writes in place of the transcript copy; the totals read these events, so a purge changes none of them. The token block and the cost are optional and a turn without them is counted in turns_without_tokens and turns_without_cost, never added as zero. The latest reading of every usage window is kept per account (usage_windows, one row per account and window name, a late event never rolls a reading back): in release 1 an account is the agent (claude), because whr never sees the credential (D40).
  • Balance. An agent may report what the account has left with a turn, an optional balance in the usage report next to the token counts: the latest reading is shown (Service.Usage) and a turn without one is unknown, not zero. No agent recorded so far streams one: the Claude Code runs of spike #7 carry no balance field, so whether Claude Code does for an API key is unverified, and the Claude adapter does not fill the field yet.
  • Reports. Totals per run, task, repository and day or month: whr usage [--task <task>] [--since <time>] [--json], and per-task usage plus a usage-window meter in the web UI (§9.3). whr show <task> prints the task’s usage line (usage_line in GET /v1/tasks/{task}). Service.Usage returns the rows and the account’s windows as the JSON of whr usage --json (a golden test pins its shape): a row is one group and one auth mode, so a subscription’s API-equivalent, not billed cost is never added to an API key’s real spend. Service.UsageLine is the line whr show prints; in subscription mode it leads with the window and calls the cost API-equivalent, not billed. GET /v1/usage serves it (task, repo, since, until and by, which is all, run, task, repo, agent, model, day or month; days and months group in the supervisor’s time zone, not UTC), and whr usage [--task <ref>] [--repo <owner/name>] [--since <time>] [--until <time>] [--by <group>] [--json] prints it: stdout is the table (with --json, the API’s envelope), the usage windows come first on a subscription, a turn without tokens shows - rather than 0, and a cost is always “reported” and, on a subscription, “API-equivalent, not billed”. --since and --until take an RFC 3339 time, a date, or how long ago (36h, 7d).
  • Budgets read the same counters. The per-run and per-task token and cost budgets of §7.4 compare against these totals: a soft threshold notifies (§9.4), and a hard limit ends the task as failed (D13). They are set in the configuration (budgets.per_run, budgets.per_task: max_tokens and max_cost_usd, and soft_percent, 80 by default) and checked after every recorded turn (Service.checkBudgets, reading Store.UsageTotals). Tokens are all four counts the agent reported, and the cost is the reported cost, so a turn that reported neither adds nothing and an unknown figure never trips a limit. The soft threshold writes one budget.warned audit entry per task, scope and metric and sends one push; the hard limit stops the live session, stops the task’s runs, supersedes its open Decisions, writes a budget.exceeded audit entry naming the scope, metric, limit and used amount, and moves the task to failed (TaskAggregate.ExceedBudget). The push says the budget was exceeded, not that the run ended.

8. Resources

Starting estimates, to be replaced by measurement. Assumes API-backed agents, not local inference.

WorkloadGuest memoryTarget
Terminal agent, Git, small scripts0.75–1 GiBup to 6–8 light (stretch)
Agent, language server, moderate builds1.5–2 GiB~4
JetBrains indexing or heavy builds3–4 GiB1–2 plus light workers

Review corrections to the original budget:

  • 8–10 GiB guests + 4–6 GiB macOS leaves near-zero slack on 16 GB. Plan for 4 concurrent instances on 16 GB, about 6 on 24 GB and about 10 on 32 GB (D32).

Capacity by memory (estimates; issue #39 measures them). macOS, the supervisor and the VPN take 4–6 GiB; a moderate environment (agent and build, including VM overhead) about 2–2.5 GiB; heavy builds or JetBrains 3–4 GiB.

MemoryLeft for environmentsModerate environmentsHeavy
16 GB (budget)about 10–12 GiBabout 41–2
24 GBabout 18–20 GiBabout 63–4
32 GB (recommended)about 26–28 GiBabout 105–6

Storage for about 4 concurrent environments (estimates; issue #54 measures them): macOS, apps and Homebrew 35–50 GB; Apple Container images 2–10 GB; per environment a root filesystem of 1–3 GB and an agent-home volume with build caches of 3–15 GB; repository caches and topic clones by repository size; stopped environments awaiting review stay on disk; keep 10–20% free. That is roughly 120–250 GB in use, so 512 GB is comfortable and 256 GB is tight. A few stopped spike containers already took 8.2 GB on the development Mac.

  • Per-VM overhead sits outside the guest limit; a Linux kernel plus a Node-based agent is typically 300–500 MB, so the 0.75 GiB floor is tight.
  • Freed guest pages are not returned to macOS: recycling is policy, not an occasional fix.
  • Admission control uses host memory pressure (memory_pressure, vm_stat) plus static limits; heavy jobs are serialised. A simple admission counter ships in release 1; full scheduling is deferred.
  • Benchmark on the real Mac mini before designing the scheduler.
  • Explicitly set CPU/memory; container machine defaults to half host RAM and shares the host home.