Architecture
5. Architecture
Logical components live in one Go binary on the Mac; boundaries are package interfaces, not microservices.
| Component | Responsibility |
|---|---|
| Control plane | Tasks, runs, workspaces, decisions, policies, integration config, event log, reconciler |
| Web UI | v0: task list, live transcript and inbox; writes are answering Decisions, sending messages, starting tasks, pause/resume/cancel and transcript purge (§9.3). Server-rendered (D8) |
CLI (whr) | Same operations through the shared API |
| Host worker | Executes a fixed set of authorized lifecycle operations beside the runtime |
| Agent adapter | Start, observe, instruct, pause, resume a coding agent |
| Runtime adapter | Provision and manage environments |
| Forge adapter | Issues, PRs, reviews, metadata, webhooks; enforces the policy table |
| CI adapter | Interface only in release 1 |
| Credential service | Encrypted store (macOS Keychain holds the master key); issues per-run credentials |
Host worker and credential service are package boundaries on one host. Until the encrypted store exists, release 1’s credential service is a set of secret files (0600, owned by the supervisor’s user, one link, outside every root, read with config.ReadSecret, never logged): every secret file the configuration names: the GitHub App key (github.key_file), the supervisor’s API token (api_token_file), the bot’s commit-signing key (bot_signing_key_file, D51), the console’s SSH authority (console.ssh_ca_key_file), the ntfy topic and token (ntfy.topic_file, ntfy.token_file) and, in the api-key auth mode, the agent’s key (agent_api_key_env_file) (issue #177). Keep the contract remote-capable so a remote worker can be added later, without building inter-process auth now.
5.1 Runtime adapter
Covers provision, start/stop/delete, inspect, resource limits, logs, exec, storage and endpoint discovery. The Apple Container behaviour below was measured in the Apple Container and host reachability spikes. Capabilities are explicit and never assumed:
| Capability | Purpose |
|---|---|
| Isolation boundary | Shared-kernel container, guest kernel or full VM |
| CPU architecture | Image, toolchain and IDE compatibility |
| Persistent storage | What survives stop, rebuild, delete |
| Networking | Reachability and supported isolation controls |
| Suspend/checkpoint | Reported support only |
| SSH/browser access | Intervention endpoints |
The Go contract (issue #20, internal/runtime):
Speccarries what the hardened environment needs: image, owner and labels, CPUs, memory and a disk quota, the network (default or--internal, by name), the user (never root), a read-only root,cap-drop ALL,--init, tmpfs mounts and the mounts.Validaterejects a spec that is not hardened: a root or empty user, nocap-drop ALL, no--init, no owner, a relative or duplicate target. Mounts arebind(a host path, checked byCheckMount, §7.4) orvolume(a name); a bind mount of a forbidden path never reaches the runtime.- Owner label. Every environment carries the supervisor’s owner label.
List(owner)returns only environments with that label, and an adapter refuses to start, stop or delete one it does not own. Removal is by exact ID, never a pattern (rm --alldeletes every container on the machine). Inspectreturns a typedInfo: the ID, owner, labels, state (provisioning,running,stopped,deletedas in §4.1) and the current address, which is empty unless the environment is running and is never stored. An unknown environment isErrNotFound.Execstreams: separate stdout and stderr readers,Waitfor the exit code, and cancellation through the context. Exec in an environment that is not running isErrNotRunning.ExecRequest.Stdinis an optional reader the adapter copies into the process and closes at its end; with a pipe the caller keeps writing while the command runs, which is how the Claude Code adapter sendsstream-jsoninput and mid-run messages (§5.2). Without a reader the command’s stdin is closed.container execcarries stdin through to the guest: the runtime conformance suite’s stdin checks pass against the Apple Container adapter (go test -tags applecontainer ./internal/runtime/apple, container 1.5.0, issue #26).- Cancel kills the process in the guest. Cancelling the context ends the stream, and the process inside the environment must be gone, not only the client (spike #2: SIGINT was not forwarded, and the agent kept running). The suite proves it with a second
execthat looks for the process, so a backend that only returns fromWaitfails. - Lifecycle is idempotent. Start of a running and stop of a stopped environment succeed, so a reconciler can retry (§5.3).
- Hardening is required, not possible (issue #26).
Spec.Validateis not enough on its own: a spec with the default network, a writable root, another task’s volume and a bind of~/.sshvalidated.runtime.Prepareis the one checked step. It runsValidate, requiresNetwork.InternalandReadOnlyRoot, checks every bind mount withCheckMount(andCheckMountsWithinthe workspace roots), checks that every named volume belongs to this environment (Owns), and returns aPreparedSpecwith unexported fields.Provisiontakes only aPreparedSpec, so nothing unchecked can reach the runtime.CheckMountreturns the resolved path andPrepareputs that path in the prepared spec, so the adapter mounts exactly what was checked: a source swapped for a symlink after the check is not followed. - The contract owns the surroundings. Provisioning creates the per-environment
--internalnetwork, the agent-home volume and the egress sidecar (Spec.Egress: the proxy’s image, its allowlist and the host binary of the proxy), and deleting the environment removes all three, by exact name.Resourcesreports them, so the conformance suite can check that they exist and that they are gone. A writable volume is exclusive (§4.4): starting a second environment whose volume another running environment holds read-write is refused withErrVolumeBusy. - Cache objects are read-only. For a
--sharedtopic,Spec.Alternateslists the cache’sobjectsdirectory (issue #45), andPrepareturns each into a read-only bind mount at the same host path, because the clone’s alternates file names that path. They are checked like any bind mount and must lie under the cache root. - A terminal is not a pair of pipes (D43, issue #92).
runtime.TerminalAdapter.Terminal(env, request)runs a command with a terminal and returns aruntime.Terminal: one stream for the command’s output and one for its input,Resize(cols, rows),Waitfor the exit code andClose, which hangs the terminal up. The Apple adapter makes a pseudo-terminal of its own (on macOS and Linux,golang.org/x/sys/unix, which was already a dependency ofgolang.org/x/term), runscontainer exec -t -iwith it as the client’s controlling terminal in a session of its own, and holds the master: the guest command has a real terminal with a size, a resize reaches it, and closing the master sends the client SIGHUP, which ends the guest’s terminal and so the shell. The environment variables travel through the0600env file like the ones ofExec. verified oncontainer1.5.0:TestTerminalLive(the size asked for,TERM, a resize, input, the exit code) andTestTerminalCloseEndsTheGuestCommand(asleepof the guest is gone within seconds of the close), and the conformance suite’s terminal check (output, exit code, input, resize,Close, an unknown and a stopped environment). - The egress allowlist changes only before an agent starts (§4.2).
runtime.EgressUpdater.UpdateEgress(env, prepared spec)replaces the environment’s proxy sidecar with one that has the spec’s allowlist: the old sidecar is stopped and deleted, the new one is created on the same networks and, if the environment runs, started. The spec must be prepared, carry anEgressand name the environment’s own network. The sidecar’s address changes, so a process that already runs keeps the old one: the supervisor calls it only while no agent process runs.Inspectreports the sidecar’s allowlist (Info.EgressAllow, sorted; on Apple Container from the sidecar’sworkharbor.egresslabel), so the supervisor changes the sidecar only when the allowlist differs. verified oncontainer1.5.0: the conformance suite checks the update on a stopped and on a running environment, that one sidecar remains and the refusals, andTestEgressThroughTheSidecarchecks that a newly allowed host answersCONNECT200 and a dropped one 403 through the new proxy. - Fakes and conformance.
runtimetesthas an in-memory fake that can simulate a service restart (every environment becomesstopped, as measured below) and the conformance suite every backend must pass; the real Apple Container adapter (#26) runs the same suite on the Mac.internal/runtime/appleruns it behind theapplecontainerbuild tag (go test -tags applecontainer ./internal/runtime/apple,container1.5.0 with a localfedoraimage, about 30 s). The same tag runs the egress conformance: a direct connection fails, an allowlisted host answers through the sidecar (CONNECT200), a raw IP and an unlisted host get 403, and the guest resolves no names. The adapter passes mounts with-v(its value is not split at commas, so a path cannot add an option; a:in a path is refused inPrepare), exec environment through--env-file /dev/fd/3, a pipe the supervisor writes and the CLI reads (no value on the command line, and no file: the0600temp file this replaced stayed for the life of the exec and a crash left it behind, issue #122;whr servesweeps such files an earlier run left, verified for the CLI reading a pipe on container 1.5.0 by the live exec and terminal tests), reserves theworkharbor.label prefix for its own labels, never--ssh,--rmor a published port, finds a volume’s name in the mount type (itssourceis the image file path), and acts on exact names and the owner label only. An environment’sAddrandProxyare read fromcontainer liston every call and never stored. - Volumes belong to the environment’s user. A new volume is an empty ext4 filesystem owned by root, so an unprivileged agent could not write its home. At creation the adapter runs one short container as root with
CAP_CHOWNonly, the volume and the read-only tool store mounted, on the environment’s internal network, andwhr-shim chownfrom the tool store as its entrypoint: it gives the volume’s top directory to the spec’s numeric user without following a link, and the container is deleted (found by the serve integration run, issue #28). No program of the environment’s image runs as root, because since D38 that image may be built from the repository, and the helper has no way out. It is the only thing the adapter runs as root. - Environment from the repository (D38, issues #77 and #76).
internal/devcontainerresolves the environment of a repository at its default branch:Resolveturns the ref into one commit and every later read, the file, the toolchain files, the lockfiles and the build context, uses that commit throughgit ls-treeandgit cat-file, never a working tree, so a topic’sdevcontainer.json, an edit not committed and a branch that moves between two reads change nothing. The order is adevcontainer.json(.devcontainer/devcontainer.jsonor.devcontainer.json, JSON with comments and trailing commas), else aDockerfileorContainerfileat the root, else a default image chosen from the toolchain files (go.mod,.tool-versions,mise.toml,.nvmrc,.python-version, in that order; the version must be digits and dots, and the image comes from the supervisor’s table, the official language images by default, or its base image when there is none, D43). The file’simageorbuild.dockerfile,build.contextandbuild.args,containerEnv,postCreateCommand(a string or an array, run inside the environment bysh -cor directly, never on the host),forwardPorts(plain port numbers, the preview ports of D33) andcustomizations.workharborare read.customizations.workharborholdsegress(host names only, each a request), and the hintscheck,previewPorts,agentandtools, which suggest and decide nothing.featuresare fetched, checked and applied to animageenvironment, or on top of the image abuild.dockerfilebuilds (issues #108 and #127, below). It refuses the whole file, naming every key, when it hasinitializeCommand,mounts,workspaceMount,runArgs,privileged,capAdd,securityOpt, aremoteUserorcontainerUserthat is root (by name or any numeric spelling of uid 0), acontainerEnvorbuild.argsvariable the supervisor sets or the agent reads (the proxy variables,PATH,HOME,LD_PRELOAD,LD_LIBRARY_PATH, and names starting withWHR_,CLAUDE_,ANTHROPIC_,OPENAI_,CODEX_,GEMINI_,GOOGLE_orBUILDKIT_,runtime.ReservedEnv; the proxy variables are predefined build arguments, so they would steer the build’s traffic), or abuild.dockerfileorbuild.contextthat is absolute or leads outside the repository. The file must be a regular blob of at most 1 MiB, not a symbolic link, the ref may not look like an option or a range, and a repository that cannot be read is an error rather than “no devcontainer”; every other key is ignored with a note. Build.Environment.Stagewrites the build context and the Dockerfile from the commit into a directory the supervisor owns (Export: regular files only, a symbolic link or a submodule is skipped and listed, no.git, no attributes or filters, at most 20000 files and 512 MiB), and aruntime.Builderbuilds it (runtime.BuildSpec, checked like aSpec: a clean absolute context, a valid tag, no reserved build argument). The tag carries a digest of the commit, the Dockerfile, the context and the arguments, so the same commit reuses its image and a new commit builds a new one. The Apple adapter runscontainer build --progress plain --tag … --file … --build-arg … -- <context>(verified oncontainer1.5.0:TestBuildLive,go test -tags applecontainer); it never passes--ssh,--secretor--output. The build runs in the builder VM, which the egress allowlist does not cover: aRUNstep can reach any host, and what that means for a reviewed Dockerfile is a question for §7.2 (open, proposed on #76). Spec.Environment.Specapplies the environment to the supervisor’s hardenedSpec: only the image andSpec.Env(thecontainerEnv, refused again bySpec.Validatefor a reserved name) come from the repository; the user, the network, the mounts, the capabilities and the limits stay the supervisor’s, andremoteUser, a named user, is not mapped to a uid. Inwhr serve.serve.Environmentrefreshes the supervisor’s own mirror of the integration branch, resolves the environment of that commit there and, for a repository with a devcontainer or a Dockerfile, givesWorkspaces.provisionthe means to make its image: once per tag, from a context exported from the commit into a directory of the supervisor’s own, through the runtime’s builder.provisionthen runs that image withEnvironment.Spec(thecontainerEnv; everything else of the spec stays the supervisor’s) and the usual git check. A repository that cannot be read gets the supervisor’s default environment and the failure is reported; one whose image cannot be built is refused, with the builder’s output, and nothing is left behind.postCreateCommandruns at a run’s start, after the egress allowlist is up to date and before the agent starts (so what it fetches goes through the hosts the human allowed), in the agent’s worktree with the agent’s environment, as the agent user, once per environment and command set (a marker in the agent home); a command that fails ends the run as failed, with the tail of its output. Egress. A host the file requests and a host a lockfile suggests (go.sum,package-lock.json,yarn.lock,pnpm-lock.yaml,poetry.lock,uv.lock,Pipfile.lock,Cargo.lock,Gemfile.lock) areHostRequests, minus the hosts already answered for the repository; nothing is allowed by being listed. The answers, allow or deny, are stored per repository in the supervisor’s database (egress_hosts,Store.SetEgressHost), never in the repository.Service.PendingEgresslists the hosts of an environment that the repository has not answered yet, andService.RequestEgressraises one blocking approval per host for a starting or running run (TaskAggregate.RaiseEgressRequest: causeegress_request, the validated host in the Decision’s ownhostnamecolumn, all or nothing).AnswerDecisionkeeps the answer for the task’s repository, needs no waiting agent (none exists before the run starts), andService.EgressAllowreturns the allowed hosts for the environment’s allowlist. An expired or superseded request keeps nothing: it is a denial for its run, and the host is asked again at the next start. The start flow (§4.2).Workspaces.launchsaves the run, thengateEgressreads the repository’s environment from the supervisor’s own mirror of the integration branch (Config.Environment, inwhr servea refresh of the mirror anddevcontainer.Resolveof that commit, never a workspace) and, if it has hosts not yet answered, raises the requests and leaves the runstartingwith its slot held, so the reconciler does not take it for lost.AnswerDecisionstarts the agent when the last request of the run is no longer open, and so does the reconciler when the last one expired. Before the agent process starts,applyEgresscompares the sidecar’s reported allowlist with the supervisor’s hosts plus those allowed for the repository and, if they differ, replaces the sidecar throughruntime.EgressUpdater(§5.1), which is safe because no agent process runs yet; the allowlist never changes while one runs. A new workspace of the repository is provisioned with the allowed hosts at once. A repository that cannot be read (no network, a private repository before #27) starts the run with the allowed hosts and reports the failure: what could not be read allows nothing. Cancelling a waiting run drops its wait, and an answer to a superseded request is refused. The start is detached. The agent start of a run (applyEgress,postCreate, the agent itself) runs in a goroutine of its own on a context that is not the request’s (Service.startDetached): a phone whose request drops must not fail a run whose answer was stored, and a slowpostCreatemust not hold the reconciler. At most one start runs per run, it is tracked so Shutdown cancels and waits for it, andCancelstops one already running (apostCreatecommand ends with its context, and an agent that had just started is stopped); a cancelled start does not fail the run, which the human stopped.StartTaskwaits up to 30 seconds for the start, so a quick one returns with the agent running and a start that fails in that time returns its error, while a longer one returns with the runstartingand a failure is reported.postCreateis bounded byenvironment.post_create_timeout(a Go duration, default10m): a command still running then fails the run with that reason. This repository’s own file builds.devcontainer/Dockerfile(the Go version ofgo.mod, a non-rootagentuser with a passwd entry, which the tests’ssh-keygenneeds) and asks forproxy.golang.organdsum.golang.org. A repository may not name an image underwhr.invalid/(rule 4a):Parserefuses the whole file namingimage(the host compared without regard to case, quoting or white space) and naming abuild.argsvalue that mentionswhr.invalid, andStagerefuses a Dockerfile that mentionswhr.invalidanywhere (RefuseBuiltFrom), so aFROM, aCOPY --from, aRUN --mount=...,from=, anARGdefault or a continued, quoted or# escape=spelling cannot build on another repository’s built image. The check is a scan of the file’s text and ofbuild.argsafter lower-casing and removing backslashes, backticks, quotes and white space; it parses no Dockerfile. It is defence in depth, not a closed door: a name assembled fromARGpieces (ARG H=whr,ARG D=invalid,FROM ${H}.${D}/x) is not in the text and is not caught (TestRefuseBuiltFromKnownGapNameAssembledFromArgs). The exposure is a repository naming another repository’s locally built image, whose content came from that repository’s default branch, within one supervisor and one human. Measured (verified,spike/builder-store, a local branch until it is pushed):container buildresolves aFROMof a locally taggedwhr.invalid/image from the host’s image store when it is run without--pull, so the scan above is the only guard, and with--pullit fails.Stagealso refuses asyntaxparser directive that names anything butdocker/dockerfile, with an optional tag and digest (RefuseSyntaxDirective), because a custom frontend is an image the builder pulls and runs: the marker may be#or//, with any Unicode white space around the key (as BuildKit’s own detection), and a file that starts with{and has asyntaxkey is refused as JSON, which BuildKit also reads; a#!first line is discarded first, as BuildKit does, so a shebang does not hide either. Only the leading comment block counts, and when unsure the check refuses, so an odd first comment line can cost a Dockerfile its build.BUILDKIT_is a reserved prefix, sobuild.argscannot setBUILDKIT_SYNTAX. Gaps: anONBUILDinstruction in a base image the repository names has no static fix, andCOPY --from=andRUN --mount=type=bind,from=of a localwhr.invalid/image resolve in the builder like aFROMdoes, without--pull(verified,spike/builder-store), so the text scan is what catches all three. The builder-side rule, so that the builder cannot resolve a locally taggedwhr.invalid/image at all, is pending wh/design; the gap is accepted for one supervisor and one human.
Do not pretend backends share Docker semantics. One runtime conformance suite (the §12 checklist, automated) must pass for every backend; it turns capability flags into verified claims.
Measured on Apple Container 1.5.0 (macOS 26.6.2, spike #2, issue #2):
| Capability | Observed |
|---|---|
| Isolation boundary | A lightweight VM per container: its own Linux kernel and one host runtime process each |
| CPU architecture | arm64 guests; a Rosetta flag exists and was not tested |
| Persistent storage | Named volumes (ext4 image files, exclusive while writable) and bind mounts survive delete; the root filesystem does not (§4.4) |
| Networking | The default NAT network reaches the LAN, the internet, other containers and host services bound to all interfaces. --internal networks block the internet, the LAN and other containers, but not the host: an internal guest reaches host services bound to the Mac’s LAN address or to all interfaces (issue #69, branch spike/host-reachability). Only loopback-only listeners are out of reach (§7.2) |
| Resource limits | --cpus sets the vCPU count; --memory is enforced by a cgroup inside the VM (a larger allocation is killed with exit 137 and the container survives) |
| Restart policy | None. After a crash or a service restart every container is stopped (§5.3) |
| Suspend/checkpoint | None observed |
| SSH/browser access | Not tested; exec works |
Adapter rules that follow from it:
- Always pass
--init: a stop took 145 ms with it and 5.3 s without, because PID 1 ignored SIGTERM. - Never store a container’s IP; it changes across recreate and restart. Read it with
inspect. - Remove only containers by exact ID.
container rm --alldeletes every container on the machine, including ones the supervisor did not create. - Never pass
--ssh, which forwards the host ssh-agent into the container. - A container starts in about 1.1 s and
execis ready in about 100 ms, so recycling environments (§4.3) is cheap. - Devcontainer features as built (D38, issue #108;
internal/oci,internal/devcontainer/feature; unverified end to end: the libraries and the Resolve and Stage integration are tested against a fake registry, and the build itself is the spike’s, not run here).Resolvereads eachfeaturesentry in the order of the file with its options (a string is the version, an object its options as strings,falseswitches it off) and, for animageenvironment, fetches it on the host throughinternal/oci: an anonymous token from the registry’s own realm only, the manifest by tag or digest, its sha256 taken as the feature’s identity, the one feature layer by digest with a size cap (16 MiB), a timeout and a streamed hash, redirects only to https and public addresses and without the token, and an extraction that refuses absolute and..paths, links, devices, duplicates and anything past 2,000 entries (files and directories alike), 64 MiB or 12 levels; adevcontainer.jsonthat asks for more than 20 features is refused whole.devcontainer-feature.jsonis then read:privileged,mounts,capAdd,securityOpt,init,entrypoint,dependsOnand every lifecycle command refuse the feature (afalseor empty value asks for nothing), andcontainerEnvmay not set the proxy variables or any*_PROXY,LD_*,HOMEor aWHR_, vendor or supervisor name (PATHmay change). That refusal is not a security boundary, becauseinstall.shruns as root and can write any file in the image it builds: it keeps a feature from setting the supervisor’s own variables or the proxy by accident. The boundary is that only an allowed source runs, below, and that the build runs in the builder VM (§7.2). A refused feature, or one from outsideghcr.io/devcontainers/features/that the human has not allowed for the repository, is left out and named in the environment’s notes (a foreign one is also listed inEnvironment.ForeignFeatures, as written and with the manifest digest it resolves to now, to be asked about, issue #127); a registry that cannot be reached or a digest that does not match fails the resolution. Options are checked against the feature’s own declarations and written to a shell file with single quotes escaped; features install in file order except whereinstallsAftersays, a cycle is refused and a missing dependency is noted, never fetched.Stagewrites a context that holds only the features, extracted again, and a Dockerfile that installs each as root (_REMOTE_USER=root) on the environment’s image and then writes each feature’scontainerEnvasENV, which is how its tool reachesPATH. The image tag is keyed on the base image, each manifest digest and the options, so a moved tag changes nothing until the next resolution. The build has the builder’s network, outside the egress allowlist: the accepted risk of §7.2. A source outside the allowed one (issue #127).gateEgressraises one blocking approval per unanswered foreign reference, with the causefeature_source, the reference in the Decision’s ownfeaturecolumn (ValidFeatureRef: registry path with a tag or digest, plain characters only, never parsed from the subject) and the manifest digest it resolved to in its ownfeature_digestcolumn (ValidDigest,sha256:and 64 hex), which the question text states as a supervisor fact; to learn it a foreign reference is resolved as far as its manifest, and no blob is fetched until an allow applies, next to the egress requests of the same run; the run staysstartinguntil none is open (the same wait,AsksBeforeStart). The answer is kept per repository and reference with the digest (feature_sources,Store.SetFeatureSource). An allow applies only while the reference resolves to the digest the human saw:serve.Environmenthandsfeature.Resolver.Approvedthe allowed references with their digests, and it is a function of the reference and the digest, so a tag that was moved to other bytes is not approved and is asked again, with the new digest; an allow given before the digest was recorded never matches. A deny stays per reference, whatever it resolves to, until the human changes it. Answering needs a fresh passkey assertion bound to the reference and the digest (D45), as for an egress host. The image is made when a workspace is provisioned, so an answer applies to environments made after it: the environment of the workspace the run started in keeps the image it has (open: whether an allow should rebuild it, proposed on #127). On a Dockerfile. With abuild.dockerfilethe features are a second build on top of the first:Stagebuilds the repository’s Dockerfile underBaseTag(the tag without the features),StageFeaturesthen writes the features’ DockerfileFROMthat tag, tagged asTag, which is keyed on both. Apple Container’s builder resolves aFROMof an image in the local store when run without--pull(verified,spike/builder-store). - Cancel via in-guest launcher (
whr-shim, D25): do not cancel by signalling the hostcontainer execclient; Apple Container fails with"failed to send signal: missing signal in xpc message"and leaves the guest process running. Instead, start commands under/tools/whr-shim run -pidfile ...(in their own process group) and cancel withcontainer exec <id> /tools/whr-shim kill -pidfile ... -grace 500ms. Measured in spike #10: cooperative processes terminate via SIGINT in 10–15 ms (125 ms total exec roundtrip), stubborn trees ignoring SIGINT are killed via SIGKILL after grace in ~514 ms (600 ms roundtrip), exit code is preserved (130 on SIGINT), and zero orphan processes remain.
5.2 Agent adapter
Specified as explicitly as the runtime contract, and versioned: the contract carries a contract_version, and an adapter declares which version it implements. Release 1 targets Claude Code and Codex CLI as built-in adapters against it; as built, only Claude Code is composed in serve (internal/serve/real.go), and the Codex CLI adapter is issue #35 (§13, capability matrix). Codex CLI lacks mid-run injection and host-routed approvals in what spike #1 could test (approvals and cancel: spike #7; planted configuration: spike #68), so in release 1 it runs in the degraded mode below, labelled in the UI. Full mode needs every capability marked so; an agent without them runs degraded (D12 requires full mode only of Claude Code, the first agent). Capability flags:
- headless / unattended operation
- mid-run message injection (required for full mode): send a user message into a running session and report how it was delivered (injected now, or at the next turn). Without it an agent cannot be a remote-controlled assistant (§1); an agent that lacks it may only run in a degraded mode that the UI labels
- structured event stream (required for every adapter): messages, tool calls, diffs and test results as typed events, which feed the live transcript (§9.3)
- cooperative pause (e.g. stop after current turn); reported false by every measured agent (D11)
- session persistence and resume
- PR/issue tooling
- “awaiting guidance” signal (how the agent raises a blocking Decision)
- approval prompts routed to the host (required for full mode): the agent blocks on a permission request and the supervisor answers it with a human’s allow or deny and a reason (§4.2). An agent without it can only run with a fixed allowlist and every other action denied
- auth modes, reported explicitly and never assumed:
api-key: the key stays in the host-side proxy and is issued per run (§7.3). Until that proxy exists, the key comes from a0600file on the supervisor’s side (the configuration’sagent_api_key_env_file) and reaches the agent through the runtime’s env file, never a command line.subscription: a consumer-plan login (for example Claude or ChatGPT sign-in) kept in a dedicated per-environment auth directory. The CLI refreshes the token itself, so it cannot sit behind the proxy.whrnever handles a subscription credential (D40). The human signs in inside the environment through the vendor’s own flow, from a terminal attached to it, and the CLI writes the login to the agent-home volume, where it survives stop, start and rebuild (D16).whrdoes not read, copy, store, relay or log it, and does not record that terminal session. A subscription token found in the supervisor’s configuration is refused. Which flows work behind the egress sidecar is spike #82.
- auth and quota blocking states: the adapter reports
auth_expiredandquota_exhausted(with the reset time when known). Each opens a blocking Decision and pauses the run instead of failing or retrying. Re-login is a UI action through a browser or device-code flow.
The Go contract (issue #20, internal/agent):
ContractVersionis a constant;Capabilitiescarries the version the adapter implements, the flags above, the auth modes and whether it reports quota.Mode()computes the mode from them: full needs headless, structured events, mid-run injection and host-routed approvals; an agent with events and headless only is degraded; anything less is unsupported.StartandResumetake aStartSpec(environment, working directory, prompt, auth mode, permission mode, tool allowlist, approver and approval timeout) and return aSession.Eventsis a typed stream: message, tool call, tool result, diff, test result, usage, approval,auth_expired,quota_exhausted(with the reset time when known), result and error.- The session ID arrives with the
sessionevent.Session.ID()may be empty until then (spike #1: Claude Code reports it after the first message).InstructandStopare called after thesessionevent, and the suite waits for it; a started session that never reports an ID fails the suite. - Permission mode and allowlist.
manualroutes every permission prompt to the host and needs an adapter with host approvals.dontAsknever asks: a tool onAllowedToolsruns and every other is denied, and the approver is not consulted. A degraded agent (§5.2) runs indontAsk; the allowlist is also what limits an agent that cannot be asked. An empty mode meansmanual. An adapter that cannot honour the mode refuses to start withErrUnsupported.bypassPermissionsandautoare not in the contract (§6). - The approval event is the record. An
approvalevent carries the request ID, the tool, the allow or deny and the reason, and the audit entry (§5.4) is written from it. The suite asserts on the event and not on text the agent prints. - Stop cancels a pending approval. An approval that waits for a human is cancelled with the session (D23), so the approver’s context ends, the approval event records a denial and the result is
stopped. - Usage events carry a typed payload (§5.7).
Instructreturns the delivery:injected(now),next_turn(after the running tool, as measured for Claude Code) orresumed_turn(degraded: the message becomes a resumed turn). An agent without injection never claims the first two.- Approvals are a host callback, and fail closed. The adapter blocks the agent on its permission request and asks the
Approver. An error, a cancelled context or no answer within the approval timeout is a denial, and the agent sees the denial. - Cooperative pause stays a capability flag (D11): a session implements
Pauseronly if the flag is true, and the conformance suite checks both ways. - Auth and quota end a run without failing it. The session emits
auth_expiredorquota_exhaustedand finishes with that status and no error, so the supervisor opens a blocking Decision (§4.2) instead of retrying; the session stays resumable. - Stop is a hard interrupt and leaves the session resumable.
agenttesthas a scripted fake agent and the conformance suite; the Claude Code adapter (#25) and the Codex adapter (#35) run it too.
The Claude Code adapter (issue #25, internal/agent/claude). It runs claude -p --input-format stream-json --output-format stream-json --verbose through the runtime’s Exec with Stdin (§5.1), as spike #1 did, and turns each output line into an agent event:
- Session. One process is one run of turns: the prompt is the first user message, a mid-run
Instructis another line on stdin (next_turn, as measured), and the adapter closes stdin after theresultevent so the process ends. A further turn isResumewith--resume <id>. Thesessionevent comes fromsystem/init, which only appears after the first message. - Stop cancels the
Execcontext. The runtime contract (§5.1) requires that this ends the process in the guest, so the adapter needs no signal of its own; the session stays resumable by its ID. - Events.
assistanttext is a message,tool_usea tool call,tool_resulta tool result,thinkingis dropped, andresultends the run and produces oneusageevent with the model, the reported cost, the token counts fromresult.usage(input, output, cache read, cache write) and the usage windows from the lastrate_limit_event. A tool result withtool_result_meta[].non_execution_kind(recorded:permission-rulefor a denial) is a tool that never ran, and is recorded as denied. auth_expiredcomes from theerrorcodeauthentication_failedon anassistantevent, never fromsubtypeorapiKeySource(§5.2, spike #1). The run endsauth_expiredeven though theresultsaysis_errorwithsubtype: "success".- Permission modes.
dontAskruns with--allowedToolsand the CLI denies the rest silently; the adapter records each decision as an approval event (an allowed tool from its allowlist, a denial from thepermission_deniedsystem event). A tool that is not on the allowlist and returns anis_errorresult with no denial event is recorded as denied, never allowed, because that is also how a denial can come back; each approval ID is unique, including denials that match no tool use.manualruns the CLI with--permission-mode manual --permission-prompts host --permission-prompt-tool stdio(D26, issue #75) and the same settings flags, so every permission prompt arrives as acontrol_request(can_use_tool) on stdout. Each one is asked of theApproverin its own goroutine, with the tool and its input capped (a Bash command, a plan, or the tool’s input as compact JSON), and the answer goes back as acontrol_responsecarrying the request’srequest_id,allowordenywith the human’s reason as the message. The adapter records every answer as an approval event, and it fails closed: no answer in time is a denial sent to the agent; a request of another kind is refused with an error response at once, and a second request with an ID that is still open is denied without being asked; if the answer cannot be written, because the channel is lost while the agent still runs, the allow is recorded as the denial it is and the agent is stopped (the runtime ends the process in the guest withwhr-shim, D25); and when the output ends or the session is stopped, whatever is still open is denied and recorded before the events end, so no request outlives its process.HostApprovalsis true. That an allow, a deny, a closed stdin and a wrongrequest_idbehave this way is verified by the recordings of spike #7 (internal/agent/claude/testdata/recorded/, replayed against the adapter inapproval_test.go); the shape of the error response to an unsupported request kind and the flags together with the supervisor’s settings flags against the real CLI are unverified. The service turns each prompt into a blocking approval Decision of the task and returns the human’s answer (internal/service/approval.go); an answer to an approval no agent waits for, such as after a pause or a restart, is refused (D23). - Only the supervisor writes what the CLI reads. The agent can write its own home (
CLAUDE_CONFIG_DIR) and its repository, and Claude Code reads settings, hooks, MCP servers and skills from both. Measured with Claude Code 2.1.285 (branchspike/claude-config): with no flags, a hook planted in the repository’s.claude/settings.jsonandsettings.local.json, one in the agent home’ssettings.json, a project.mcp.jsonserver and a project skill all took effect; with--setting-sources userthe agent home’s hook still fired, so pinning touserwas not enough. With--setting-sources ""(no source),--strict-mcp-configand a supervisor-written--settings(a file or an inline JSON string), none of the planted hook, MCP server or skill took effect, and the supervisor’s own hook did. The adapter passes those flags on every start and every resume, plus--disable-slash-commands(skills) as defense in depth. That planted permission rules are ignored too is inferred: the same loader reads them, but it was not shown without a model call. - What the CLI still allows on its own (§6): the built-in read-only allows (spike #1:
pwdandgit statusran inmanualanddontAsk), the built-in plugins, and whatever--allowedToolsand the supervisor’s--settingsgrant. Anything else is denied indontAsk. - Bounded input. stderr is kept to a capped number of lines and bytes. A result that has arrived is kept when
Stopcomes late, andResumereportsErrNoSessiononly for an error result that says the session is missing; any other early failure is an ordinary error. - Unverified, taken from the spike’s field names only:
quota_exhaustedis raised when a usage window reaches utilization 1 andrate_limit_info.statusis anything butallowed(whatrate_limit_event.statusreads when exhausted was not observed); aResumeof an unknown session is recognized as an errorresultbefore anyinitand no login failure (the CLI’s real signal was not captured). Verified on real streams (spike #7,spike/agent-approvalcommit22ddc7f,internal/agent/claude/testdata/recorded): theinitandresultshapes, the token counts, the rate-limit windows with epoch-second reset times, the denial marker, andsession_idon every event. Still constructed from the spike’s description, because no recording exists: the expired login and the exhausted quota.
Measured in spike #1 (issue #1; branch spike/transcript, RESULTS.md), with Claude Code 2.1.285, Codex CLI 0.159.2 and Antigravity agy 1.1.12 on one machine:
| Capability | Claude Code | Codex CLI | Antigravity |
|---|---|---|---|
| Headless, typed events | Yes (stream-json in and out) | exec --json; only start and error events seen | Yes (--output-format stream-json) |
| Mid-run message | Yes, picked up at the next model step | Not found in exec (unverified) | No: one prompt per run |
| Resume | Yes (--resume), same session ID | exec resume exists, untested | --conversation <id>, untested |
| Approvals to the host | Yes, on the host through an MCP prompt tool; in a container, the stdio control protocol (D26, measured in spike #7) | Untested | None found in print mode; the run ends in ERROR on a denial |
| Usage window | Structured: five-hour and seven-day windows with reset time | Text only, reset time inside the message | Not observed |
| Cancel | Hard interrupt only, session stays resumable | Untested | Untested |
Findings that shape the contract:
- A user message sent mid-run is delivered at the next model step, after the running tool finishes, not by interrupting it. The UI says so.
- The session ID only appears after the first user message, and user messages are not echoed in the output, so the supervisor logs its own.
- There is no cooperative pause; cancel is a hard interrupt.
- Agents without streaming input (Codex
exec,agyprint mode) run in the degraded mode: a message becomes a resumed turn, labelled in the UI. - Token-level streaming is available from Claude Code (
--include-partial-messagesaddstext_deltachunks, and the full message still follows). Coalesce deltas (about 150 ms) and let the final message replace them; deltas are live-only and never stored (§5.4). - A missing or expired login is signalled by an
assistantevent witherror: "authentication_failed", then aresultwithis_error: trueandsubtype: "success". Detectauth_expiredfrom the error code, never fromsubtype, and never fromapiKeySource, which readsnoneboth for a subscription login and for no login at all. A login that expires mid-session was not reproduced.
5.3 Reconciler
Desired state lives in the database. A loop compares it with actual runtime state, marks orphaned runs interrupted, and resumes from the agent session rather than the VM. This is the answer to Apple Container’s missing restart-policy recovery.
Measured in spike #2 (issue #2):
- No restart policy. After the host-side runtime process of a container was killed, it stayed
stopped, with its volume state intact, until started again. - A service restart ends everything.
container system stoptook 0.4 s and ended every container VM at once; aftercontainer system start(0.4 s) every container wasstopped, including those that had been running. Nothing came back by itself. Root filesystems, volumes, bind-mounted data, networks (also custom--internalones) and images survived; every process inside the containers was gone. - After a Mac reboot nothing starts the services either: there is no LaunchAgent or LaunchDaemon plist for them on disk. The supervisor’s own launchd job must run
container system start --disable-kernel-installand then reconcile. The flag skips the interactive kernel-install prompt, so the kernel must already be installed once at setup (container system kernel set --recommended); without it no container starts. A reboot itself was not triggered. - Recovery loop. List containers, start those that should be running, wait for
execto answer (about 100 ms after start), then resume the agent from its session. Container IPs change on every start, so they are read again each time and never stored.
The Go service (issue #23, internal/service) is the one layer the JSON API and the web UI call (D8). Its reconciler is DB-first (D6) and takes the clock and the runtime and agent adapters as parameters, so tests run it on the fakes with an injected clock. One pass, step 1 once and the rest per active task:
- List the environments the runtime reports for this owner (
List(owner), never every container). Then, before anything is observed, stop each running workspace environment that this supervisor’s store records and the current process did not start, the console’s excepted, whatever runs it holds; the stopped ones count as seen stopped in step 2, and a failed stop is tried again on the next pass (§4.1, “No surviving agent before a relaunch”, #221). The owner label is shared by every supervisor of the host, so an environment the store does not record is left alone. This sweep runs once per pass, not per task. - For each environment of the task,
ObserveEnvwith what the runtime says; an environment the runtime no longer knows is observed as gone.startingandrunningruns in an environment seen stopped or gone, and a run paused by an agent’s cooperative pause, becomeinterruptedand their open questions and approvals are superseded; a hard-paused run stayspausedwith its login or quota question (§4.1, #221). - A run without a session is lost. A
startingorrunningrun that has no attached agent session isinterrupted, even when its environment is still up: after a supervisor restart the environment survives, and so can the agent process, which the supervisor no longer holds (spike #7, Case 4). The run is interrupted, and when the current supervisor process did not start its environment, the reconciler then stops that environment, which interrupts everystartingorrunningrun in it (§4.1, “No surviving agent before a relaunch”, #216); an environment this process started is never stopped for this. A run whose launch is in progress is not lost; the service marks it before it calls the agent. After step 1’s sweep this stop is needed only when the sweep’s stop of that environment failed (a run’s environment is always recorded, because every launch saves it first). - For each
interruptedrun, unless it waits for a reset it chose (“resume at reset”, §4.2): check thatResumecan succeed (no open login or quota question) and that no other active run owns its environment before starting anything, so a run that waits for the human costs no environment; stop its environment first when this process did not start it (§4.1, “No surviving agent before a relaunch”), start it if it is stopped, wait untilexecanswers (polling with the injected clock, bounded), observe the environmentrunning, move the run tostartingwithResumeand relaunch the agent withResume(sessionID). The session ID is recorded on the run when the agent reports it (thesessionevent), because the session, not the environment, is what survives. On success the run isrunning. A session the agent no longer knows (ErrNoSession), or a run that never reported one, endsfailedat once, which opens a blocking retry-or-cancel Decision (§4.1). Any other error returns the run tointerruptedand counts an attempt; when the attempts (3 by default) are used up the run endsfailedtoo. - Every resume starts with the briefing (D27): the first message to the agent says that the process ended, that a tool call that was running may have had effects that are unknown or partial, which Decisions were superseded (their tool and input), and that the agent must check the workspace before repeating anything.
- Resume the runs whose
resume_at_resetDecision is due. - A shutdown interrupts, it does not stop.
Shutdownends the sessions, and the runs they served becomeinterrupted, so the next start resumes them; only a stop the human asks for (cancel) or an agent that finished ends a run for good.
The sessions map is keyed by run, and a finished session removes its entry only if the entry is still that session, so a relaunch during a pass is not lost.
The slice’s use cases (issue #95, internal/service) are what the API and the CLI call. Handlers hold no logic.
- Start (
whr run <issue-url> --agent <ws>/<role>, D42) takes an issue URL and an existing agent of an existing workspace; it never clones. In order: parse the URL and refuse a repository other than the workspace’s; load the issue through the forge (everything in it is untrusted data); apply the trust tier of #53 (a hook that today allows every issue and is the place to add the tiers, not a way around them); build the first prompt from the issue, marked as untrusted; start the workspace’s environment and wait until exec answers; save the task and its first run instartingtogether with the one-run check, in one transaction (§4.3); start the agent in the agent’s worktree; mark the runrunningand attach the session. Each step that changes state is an event of the aggregate (task.state,run.started,run.state,run.session). A start interrupted after the save leaves astartingrun without a session, which the reconciler takes for lost and fails into a retry-or-cancel Decision (point 3 above); one interrupted before the save left nothing. - Trust tiers (issue #53). The author association GitHub reports for an issue is read as a tier:
OWNER,MEMBERandCOLLABORATORare trusted, everything else (CONTRIBUTOR,FIRST_TIME_CONTRIBUTOR,FIRST_TIMER,MANNEQUIN,NONE, and anything missing or new) is untrusted: fail closed.whr runon an issue by an untrusted author starts nothing: it records the task, stillqueued, and raises a blocking question for the human (startorcancel) that shows the author, their association and the issue text, as untrusted data.startloads the issue again and starts the run only if what the human saw is still the start of it, otherwise it refuses and the human runs it again;cancelcancels the task. The task is marked as having untrusted input, and from then on the policy decision for it takes that context (Table.DecideIn): an action the table lets run on its own (auto) asks when it has an outward effect (push, open a pull request, comment) and the input is untrusted, never looser than the table. A run that combines private data, untrusted input and outbound network needs approval (§7.1): untrusted input is held for a Decision before the run starts, and the supervisor does not know whether the repository is private, so it assumes it is and always holds untrusted input. - A new run on an existing task (
StartRun) serves rework after changes were requested and the answers of two questions:retryof a failed run, andreworkof the rebase-conflict question of §4.2. It checks the one-run rule like Start and starts the agent with the briefing of D27 plus, for a rebase conflict, the conflicting paths as untrusted data.retryof a rebase conflict only frees the task: the human runs the export again. - Say sends a message to the task’s live run with the session’s
Instruct, records aninstruction.sentevent (the run, how it was delivered and the text, redacted) and returns the delivery, so a degraded agent’s resumed turn is shown as that and never as an injection. A task with no attached session is a conflict. - List and Show read the store: a summary of every task, and one task with its runs, open Decisions and current revision.
- Events (
Subscribe) replays durable events of a task from a sequence number out of the store and then follows live ones. The agent’s observations (messages, tool calls and results, diffs, test results, usage) are stored as transcript-tier events, so they replay and have retention (§5.4); ephemeral events (token deltas, heartbeats) are published to subscribers from memory, redacted like stored ones, and never stored. The store is the replay buffer, so a client follows a history of any length at its own pace. Publishing never waits for a slow subscriber: one that falls behind loses the ephemeral events it missed and is caught up from the store, so it misses no durable event. - A session outlives the request that starts it. The agent runs for hours and the API call that started it returns at once, so the supervisor starts and resumes a session with a context that does not end with the request;
Shutdownstops it and leaves its run resumable (found by the serve integration run, where the session died with its request). The supervisor also gives the agent process its environment for this start inStartSpec.Env: the egress proxy’s address, read from the runtime because it changes with every start and never stored, andHOMEon the agent home volume. It holds no secret. - Credentials (D40): in API-key mode the key comes from
config.AgentAPIKey()and reaches the agent through the adapter’sEnv; in subscription mode nothing is passed and the human signs in inside the environment (spike #82, stubbed until it lands). The egress sidecar runs the installedlibexec/whr/whr-proxy-linux-arm64.
Only the container’s state and the agent session ID are stored; the address Info.Addr is read for use and never written. Nothing in the reconciler sets a state by itself: it calls the aggregate.
5.4 Events, idempotency and retention
Per-task append-only event log doubles as audit trail, UI feed and CLI stream. Every mutating command accepts an idempotency key.
Retention. A chat grows with every message, tool call, tool result and diff, so the log has two tiers:
- Audit entries are never purged: state changes, Decisions and their answers, usage records (§5.7), approvals (including the tool and a capped input), commits and PR links, credential issue and revoke, policy denials, permission-mode changes, and the record of every purge.
- Transcript content is bulk and has retention: assistant text, tool inputs and results, diffs, thinking and attachments. Streamed token deltas are never kept durably, only the final message (§5.2).
- Limits. A size cap and an age limit per task, with the cap and limit set by policy, and a manual purge from the web UI and
whr purge(§9.3). Deleting a task purges its transcript. - A purge records itself. It deletes transcript content and keeps one audit entry: who, when, and what was removed (event count and bytes). Audit entries refer to transcript content by hash, so a purge leaves a verifiable gap and never silently rewrites history (§7.7).
- The kill switch writes an audit entry too (§7.7).
Service.KillAll(POST /v1/kill-all,whr kill-allonce the CLI exists, #97) stops every live session, cancels every unfinished task, revokes the installation tokens the forge client holds (DELETE /installation/token, one per repository; the agent’s own credentials stay with the human, D40) and then writessupervisor.kill_allinto the supervisor’s own stream: who, which tasks, how many tokens and what failed. It goes on after a failure, so one task that cannot be cancelled does not leave the others running, and the entry is written whatever happened. - The audit entries form a hash chain, indexed by commit (§7.7). Every audit-tier event gets one row in
audit_chain, written in the same transaction: a SHA-256 over the event (sequence number, task, kind, tier, time and the stored payload, each length-prefixed) and the hash of the entry before it, so changing, removing or inserting an entry breaks every hash after it. The rows are append-only like the events (triggers refuse an update or a delete) and a purge of the transcript leaves them alone.VerifyAuditrecomputes the chain and fails at the first broken entry; an entry written outsideAppendhas no row and is caught too. A chain cut off at its end is a shorter valid chain, soAuditHead(sequence number and hash of the last entry) is what the human records somewhere the supervisor cannot write, andVerifyAudittakes it to check the end. Audit entries from before the chain (migration 0010) have no row and are counted asunchained, not covered. An entry whose payload has ashafield of a full lower-case commit name (the pin, the push, the CI result, the pull request, a review Decision and its answer) also records it in the chain row, soCommitAuditreturns the trail of one commit, oldest first; the SHA is part of what the chain checks. - The agent’s own session is separate. A purge does not touch the session the agent resumes from; shrinking the agent’s context (compaction or a new session) is a different action with its own consequence, the agent forgetting, and is not offered as a purge.
- Redaction at ingest. Secrets are redacted before anything is written: events, Decision inputs and audit entries alike (§7.3). Audit entries are never purged, so redacting later would be too late. What remains is treated as untrusted data when shown.
- Live-only events. Token deltas and heartbeats go to connected clients through an in-memory fan-out and are never written.
--since(§9.2) replays durable events only; a client that reconnects mid-message gets the final message when it is written.
The store (issue #21, internal/store):
- SQLite in WAL mode through the pure-Go driver
modernc.org/sqlite, so the supervisor stays one static binary (D3) that cross-compiles to Linux. Foreign keys are on. Migrations are embedded SQL applied in order and recorded inschema_migrations; a database newer than the binary is refused. - Tables: tasks (saved together with their runs, environments and review candidates), decisions, events and idempotency keys.
- Versions and compare-and-swap. A task aggregate and a Decision each carry an integer version. A save writes
WHERE version = expectedand bumps it; no row updated means another writer got there first, reported as a conflict (exit code 5). An answer and an expiry of one Decision cannot both win, and neither can two commands on one task. - One transaction. The domain methods record the events a change produced (
TakeEvents), and the store writes the new state and those events together, so there is never a state without its audit entry or an audit entry without its state. - Events are append-only with a monotonic, never reused sequence number, which is what
--since(§9.2) replays. Each has a tier. Triggers refuse any update and any delete of an audit row; the only way rows leave isPurge, which deletes transcript rows and appends one audit entry (who, when, how many events and bytes) in the same transaction. - Idempotency. A mutating command runs under its key: the response is stored in the same transaction as the changes, so a replay of the same key and request returns the stored response without doing anything again, and the same key with a different request is refused as a conflict.
- Redaction is on by default (issue #22,
internal/redact). The store redacts before it writes: event payloads (audit and transcript), the text of Decisions (subject, input, reason, answer, actor), task and candidate text, stored idempotency responses and the actor of a purge. What is held in memory is not changed, only what is persisted. Idempotency keys are opaque identifiers chosen by the client and are stored as given, so a client must not put a secret in one. - What the redactor knows. Exact secrets registered for a run (the scoped tokens the credential service issued, §7.3) are replaced in every form they may take in text: raw, JSON-escaped, URL-escaped and base64. Well-known token formats are replaced whether or not they were registered: GitHub, GitLab, Anthropic, OpenAI, AWS, Slack and Google keys, JWTs, private key blocks, bearer tokens, passwords in URLs and values of secret-looking names (
token,password,api_key,authorization). A numeric value, such as a usage count namedinput_tokens, is left alone, because usage records must stay correct (§5.7). The same redactor wraps log output andslogrecords, so a token is kept out of logs by the same rules. - A floor, not a proof. A secret in a format the redactor has never seen, and not registered, passes. A canary test therefore pushes a token of every kind through every write path, closes the database and scans the files, including the write-ahead log, for it; a control run shows the scan finds a token that was not redacted.
5.5 Adapter plugins
New agents (and later runtime or forge backends) are added as out-of-process plugins, not in-process code. A plugin is a separate executable that speaks the versioned adapter contract (§5.2) over stdio or a local socket (JSON-RPC style). Go’s in-process plugin package is not used: it is fragile and would put third-party code inside the supervisor.
- Release 1: the contract is the design; Claude Code and Codex CLI are built-in adapters against it. No loader.
- Medium term: a plugin loader, once two built-in adapters have proved the contract. Whether an existing agent-client protocol (for example Zed’s ACP) already covers part of the contract is unverified; check it in the §12 scorecard and reuse it if it fits.
- Conformance. A plugin declares its capabilities and must pass the same conformance suite as a built-in adapter, so a capability flag is a verified claim (§5.1).
- Wire types. Everything that crosses the seam has a stable JSON name (snake case): capabilities, the start spec, events, results, approval requests and answers. Times are RFC 3339 and a zero time is left out. An error crosses as a string code (
unsupported,unsupported_auth,no_approver,no_session,not_running,bad_spec), andagent.ErrorFormaps a code back to the sentinel, soerrors.Isworks on the supervisor’s side of the seam. - The approver is a reverse call. The approver cannot be a Go callback across a process. The plugin sends an
approverequest (an approval request) to the supervisor on the same connection and waits for the answer. The supervisor applies the approval timeout and the fail-closed rule (§5.2), so a lost connection, a plugin that dies or a late answer is a denial, and the plugin never decides. - Trust. See §7.8: plugins are installed explicitly and run isolated.
5.6 Tool store
Environments run stock images, or, for a repository without a devcontainer, the workharbor base image of D44: a first-class base plus git and CA certificates, built once by whr. Fedora and Ubuntu LTS are the first-class bases, pinned by digest and covered by the conformance suite, with Fedora the default; other bases work best-effort (D43). The agent CLIs (Claude Code, Codex CLI, later others) and the supervisor’s own helpers live once in a versioned, immutable tool store on the host and are mounted read-only into each environment, in the manner of a Nix store. This replaces installing an agent in every container, which took about 11 s and 230 MB each in spike #2.
- The base image (D44, issue #98).
internal/baseimageships one Containerfile per first-class base in the binary (Containerfile.fedora, the default, andContainerfile.ubuntu): the stock image pinned by digest, thengit-coreandca-certificates(Fedora) orgitandca-certificates(Ubuntu) and nothing else, with no user and no copied file. Its tag iswhr.invalid/whr-base/<base>:<12 hex of the Containerfile's SHA-256>, so it stays the same until the pin changes.whr servecallsbaseimage.Ensureat start whenenvironment.imageis unset: if the runtime has no image with that tag (HasImage,container image inspect), it builds one through the sameruntime.Builderas a repository’s image (an empty context), so a later start builds nothing.environment.base(fedoraorubuntu) picks the base;environment.imagereplaces the base image altogether. verified oncontainer1.5.0 for both bases:go test -tags applecontainer -run TestBaseImageLive ./internal/baseimagebuilds the image, finds it the second time, and runs as uid 1000git worktree add,git bundle createand an HTTPSgit ls-remote. - An image without git is refused at workspace creation. After the environment starts,
Workspaces.Createrunsgit --versionin it. A non-zero exit is refused, with the image name and what git said, before any worktree step; the environment, its network and its home volume are taken back. The check applies to every image: the base image,environment.imageand a repository’s own (#76). - The console image (D43, issue #92).
internal/consoleships the console’s image definitions next to the base image’s: one Containerfile per first-class base, the same stock image pinned by the same digest, withgit, zsh, fish,jq,curl,ripgrep,tmux,make,less, an editor (vim-minimalorvim-tiny),util-linuxwithscript, a non-root userwhr(uid 1000) and the git wrapper as/usr/local/bin/git. It has no Docker CLI and no container engine and holds no credentials. The tag iswhr.invalid/whr-console/<base>:<12 hex of a SHA-256 over the Containerfile and the wrapper>, built once through the same builder (baseimage.EnsureImage, with the wrapper as the only file of the context). verified oncontainer1.5.0 for both bases:go test -tags applecontainer -run TestConsoleImageLive ./internal/console. Alpine (issue #93). The console has a third image,Containerfile.alpine(Alpine 3.24.2, pinned by digest;apkfor the same tools,adduserfor the user, and a password field of*because an sshd built without PAM refuses a locked!account even for a certificate). It is for the console only:console.basein the configuration picks fedora, ubuntu or alpine (defaultenvironment.base), and agent environments stay on a glibc base until the agent CLIs’ musl builds are verified.console.Alpineis not abaseimage.Distro, which a test pins. Measured (verified,container1.5.0): the console image checks (TestConsoleImageLivewithWHR_TEST_BASE=alpine: the tools are there, no engine, and the git wrapper stops the planted hook, fsmonitor, filter, textconv and alias under busyboxshand musl git), the runtime conformance suite,TestConsoleSSHLiveand the threeTestTerminaltests (the session kill withpkill -s) all pass on it; the suites ran on an image built by hand from the same files, since the tests that takeWHR_TEST_IMAGEdo not build one. - Images
whrbuilds are named underwhr.invalid/and never pulled (rule 4a, issue #133). The environment, base and console tags arewhr.invalid/whr-env/<owner>:<hash>,whr.invalid/whr-base/<base>:<hash>andwhr.invalid/whr-console/<base>:<hash>;runtime.BuiltImageHostis the one definition of the prefix.whr.invalidis an RFC 6761 name no registry answers, so the name cannot be published elsewhere.container createhas no never-pull option (onlycontainer buildhas--pull), soProvisionaskscontainer image inspectfirst and refuses a missingwhr.invalid/image with “built by whr and is not here: build it” before it creates a network, a volume or a container. Measured on container 1.5.0 (verified,TestMissingBuiltImageIsNeverFetchedLive):container createof a missingwhr.invalid/image fails after about 10 seconds with “failed to resolve either repository hostname”, while a missing bare name goes toregistry-1.docker.io;container buildwith a missingFROM whr.invalid/...fails the same way, closed, in about 12 seconds. The egress sidecar’s image is checked the same way, inProvisionand inUpdateEgress(it is the environment’s own image, or the console’s, when built), before any network contact. A repository’s ownimageis not rewritten: one underwhr.invalid/is refused (a scan of the file’s text and ofbuild.args, see the D38 bullet: defence in depth, and a name assembled fromARGpieces is not caught; the builder does use a localwhr.invalid/image as aFROM, so the scan matters, verified byspike/builder-store; a# syntax=frontend other thandocker/dockerfileis refused too), any other keeps the CLI’s behaviour. The fake runtime still accepts a missing image. Images under the old names (whr-env/,whr-base/,whr-console/) are orphaned: nothing deletes them andwhrbuilds under the new tag on next use. - The console’s git wrapper. A workspace’s
.gitis written by agents, so the configuration of a repository is hostile input to whoever runs git in it (spike #89: a hook, fsmonitor, a clean filter, a textconv driver and an alias all ran).whr-gitruns the real git withcore.hooksPath=/dev/null,core.fsmonitor=falseandprotocol.ext.allow=never, and overrides by name every setting git uses to run a command:core.pager,core.editor,sequence.editor,core.sshCommand,core.askPass,core.gitProxy,pager.*,filter.*.clean,.smudge,.processand.required,diff.external,diff.*.textconvand.command,merge.*.driver,mergetool.*anddifftool.*,alias.*(to a static refusal),credential.helper,gpg.program,uploadpack.packObjectsHook,remote.*.uploadpack,.receivepackand.vcs,trailer.*.cmdand the browser and man commands. The overrides go throughGIT_CONFIG_COUNT,GIT_CONFIG_KEY_nandGIT_CONFIG_VALUE_n, which win over every file. A key is trusted only if the console’s own system or global configuration sets it too, and then the override is that value, so the human’s own aliases and helpers keep working; the origin of a setting is never parsed, because an include gives a repository the choice of the path. Keys read first-value-wins (remote.*.uploadpack) go into a temporary file git reads as its system configuration, which comes before the repository’s. It is POSIX sh, because Ubuntu’s/bin/shis dash. The tests plant each kind in a real repository and run it under plain git, which must run it, and under the wrapper, which must not;pager,editorandcredential.helperneed a terminal or a helper, so their value is read instead, and no test runs a credential command. - The console environment (D43, issue #92).
service.Consolesopens, reuses and closes one console:Openruns the optional image build (console.Ensure, on first use, sowhr servestarts without it), prepares and provisionsserve.ConsoleOptions.For’s spec and starts it. The spec is hardened like an agent’s environment (non-root uid 1000, read-only root,cap-drop ALL,--init, tmpfs/tmpand/run, an internal network of its own behind the egress sidecar), with every workspace root mounted read-only at/workspaces/<root>(two roots with one base name get a numeric suffix), each workspace requested read-write mounted over its place in its root, and the home volume<owner>-console-homeat/home/whr. There is no tool store, no agent and no secret: the console reaches the hosts ofconsole.egress_allow(the package registries of the supported toolchains andgithub.comfor read-only fetches by default; never the model API) and nothing else. The console is found by two labels (whr.console, andwhr.console.rwwith the sorted IDs of its writable workspaces), not in the database, so the reconciler, which looks at the environments of tasks, never touches it. The same writable set reuses an open console (a stopped one is started); another set is a conflict, because changing the mounts of an environment under open shells pulls the floor from under them, so it is closed first.Closestops and deletes the environment, its network and its sidecar, and keeps the home volume, which holds the human’s own dotfiles and history. verified oncontainer1.5.0: the runtime conformance suite passes on both console images (WHR_TEST_IMAGE=<console image> go test -tags applecontainer ./internal/runtime/apple, about 100 s each).Consoles.Shellopens a loginzshin the open console aswhr(uid 1000) in the workspace’s directory, withHOME,USER,SHELL,LANG,WHR_CONSOLE, the client’sTERMif it is a plain terminal name (anything else becomesxterm-256color, because it comes from the client) and the egress proxy’s variables, throughruntime.TerminalAdapter, and writes thesupervisor.consoleaudit entry; at most eight shells at once, andCloseShellshangs them all up when the supervisor stops. The API route and thewhr consolecommand carry it as a terminal stream (§9). The guest’s command ends with the terminal: closing the master hangs the guest’s terminal up, which the live test shows ends asleepwithin seconds. - Rebuilding a workspace’s environment (issue #128). An allowed feature source, a new
devcontainer.jsoncommit or a bumped base image takes effect only when an environment is next built, soWorkspaces.Rebuild(whr ws rebuild <workspace>,POST /v1/workspaces/{workspace}/rebuild) replaces a long-lived workspace’s environment. It keeps what the agents have: the workspace folder, so every worktree and branch, and the home and build volumes. The container, its network and its egress sidecar are new, with the hosts the repository has been allowed, as at provision (prepareSpecandbringUp, the two halvesprovisionis now made of). The order is the safe one: refuse while any run of the workspace is unfinished, naming it and its state (Store.UnfinishedRuns: starting, running, paused and also interrupted, which the reconciler would resume in the old environment), and mark the workspace as being rebuilt under the lock a run’s start takes (saveStartingRun), so neither can pass the other’s check and a task started meanwhile is refused;saveStartingRunreads the workspace record again under that lock and refuses a start whose record names an environment a rebuild has since replaced; every other path that would start or use the environment refuses with a conflict while the mark is set (Service.refuseWhileRebuilding: the start of a task, export and prepare,AddAgent,Rebase), and the reconciler skips a run of a workspace being rebuilt and takes it up on its next pass, so the old environment is never started during a rebuild; build the new image and check the spec first, so a repository whose image cannot be built costs nothing; stop the old environment, because a volume is held by one running environment at a time; provision the new one on a network of its own name, start it, wait for it and check it has git; switch the record and write the audit entry in one change (Store.SwapWorkspaceEnv, a compare-and-set on the old environment); and only then delete the old container, its network and its sidecar. If the new environment cannot be brought up, or the record cannot be switched, the new one is taken back and the old one is started again, as it was (it is left stopped if it was stopped). The audit entryworkspace.rebuiltnames both environments, both images and the digest the runtime recorded for each (runtime.Info.ImageDigest, from the container’s image descriptor). verified oncontainer1.5.0 for what the runtime must do (TestAnEnvironmentCanBeReplacedOnTheSameVolume: a second environment on the first’s volume while the first is stopped, the volume’s contents kept, the first removed with its network). That test found that the helper containerownVolumeruns could stay behind and keep the old network from being deleted; the helper is now removed with retries and again when its environment is deleted. A kept home volume meets the same user: the user of an environment is fixed by the supervisor (1000:1000, inSpecOptions.For;TestTheUserOfAnEnvironmentIsFixedSoAKeptVolumeMeetsTheSameUIDpins it), so the ownership helper, which runs only on a new volume, is not needed again; a change of that user would need the helper to run on a kept volume’s root first (it does not recurse). A path that starts or uses an environment (a task start, an export, a rebase, adding an agent) takes a lease on it under the same lock, so a rebuild is refused while one is under way and a lease is refused while a rebuild is; the rebuild reads the workspace record again inside its critical section and works from the fresh one, never from a stale environment ID. A rebuild does not re-run a repository’spostCreateCommand(its marker is in the agent home, which is kept); that is open. - The host’s own IPv6 prefixes are refused (issue #124, §7.2).
serve.HostIPv6Prefixesreads the global-unicast IPv6 prefixes of the host’s interfaces (not loopback, link-local, unique-local or IPv4; at most 16) whenever a spec is built,runtime.Egress.DenyPrefixescarries them,runtime.Preparerefuses one that is not an IPv6 prefix or is wider than /16, and the Apple adapter passes them towhr-proxy -deny-prefixes.egress.Proxy.Denymakes the proxy refuse an address inside one withErrNotPublic(logged as a denial), both on the resolved answer and in the dialer’s control hook;egress.Publicis unchanged. The list is read when a spec is built, so it goes stale: after the ISP renumbers or the host changes networks, a running environment or console still refuses the old prefix and allows the new LAN until it is rebuilt (whr ws rebuild <workspace>, issue #128) or the console is closed and opened again (whr console --close); a refresh of a running sidecar is not built. The list fails open: when the interfaces cannot be read, or there are more prefixes than the cap, the proxy still applies the private ranges and the prefixes it got, and the supervisor logs it (with the count dropped over the cap). A reported prefix wider than /16 is dropped and logged rather than passed toruntime.Prepare, which would refuse every spec. A /128 interface mask (DHCPv6, utun) would refuse only the host’s own address and not its LAN unverified: not measured on macOS. Tested with fake prefixes and a fake interface list; not run against a sidecar with real global IPv6 unverified, because container 1.5.0 gave a container only a link-local IPv6 address. - The allowlist matches exactly (issue #123, §7.2).
egress.Newkeeps two lists: exact host names and*.suffixwildcards.Allowedadmits a host that is an entry, or that lies below a wildcard entry and is not the bare name; an IP address never matches; an entry with a*anywhere but a leading*.is dropped, so it allows nothing. The wildcard form is accepted inenvironment.egress_allowandconsole.egress_allow(config.validEgressEntry: a host name or*.and a host name with at least two labels, so*.comis refused) and byruntime.Preparefor the sidecar’s list; a repository’s request never is:devcontainer.jsonentries with a*are ignored with a note that names the wildcard,domain.ValidHostrefuses one for a Decision, and the lockfile suggestions are fixed exact names. The built-in console list is six exact hosts, each with its reason inconfig.DefaultConsoleEgress, and the default agent list isapi.anthropic.com. The sidecar also passes on the bytes a client sent with its CONNECT (issue #124). - The build volume (D39, issue #80). Every workspace environment has a second volume,
<owner>-build-<workspace>, mounted read-write at/var/whr/build(serve.GuestBuild), owned like the agent home (Owns), created and deleted with the environment. Each agent’s tools write their output there, in/var/whr/build/<role>, and not in the bind-mounted checkout, where file metadata costs 20 to 100 times as much (spike #40): the run’s environment setsCARGO_TARGET_DIRto…/<role>/cargo-targetandUV_PROJECT_ENVIRONMENTto…/<role>/venv(Workspaces.buildEnv; the post-create commands get the same), and adding an agent makes its directory (mkdir -p), for tools that do not make their parents. Removing an agent clears its directory (rm -rfof a direct child of the build directory, only; a failure is reported, not returned), so a role added again does not inherit the old output. It is per agent, because two agents of one workspace work on different branches and must not share a virtual environment; the Go caches already live on the agent home.WorkspaceConfig.BuildDiris empty in a supervisor without the volume, and then nothing is set, because the paths would not exist on a read-only root. unverified with a realcargoanduvuntilwh/verifyruns them; the volume itself is the conformance suite’s ownership and writability checks. The size of the checkout. When a run starts,Workspaces.noteRepoSizecounts the files its worktree tracks, in the environment (git ls-files -z, because the host never runs git in a workspace), and records it as atask.repo_sizeaudit entry in the task’s events (domain.RepoSize: the run, the count, a warning and its text above 50 000 files, D39); a count that cannot be made is reported and never fails the start. The supervisor’s git commands in the guest (gitEnv: this count, the worktree, the rebase, the bundle export) run withcore.hooksPath=/dev/null,core.fsmonitor=falseandprotocol.ext.allow=neverthroughGIT_CONFIG_COUNT, because a plantedcore.fsmonitorruns ongit ls-files(measured, git 2.54) and a.gitwritten by an agent is hostile input; a test plants a hook and an fsmonitor and shows plain git runs both and the environment neither. Not built, and why (issue #80): the mount-over path fornode_modulesis blocked by two facts measured or read from the code.git worktree addrefuses a directory that already exists and is not empty (git 2.54:fatal: '…' already exists, also with--no-checkout), and the runtime makes the mount point in the host folder when it creates the container, before the worktree exists; and a container’s mounts are fixed when it is created, so an agent added later gets none. Both need a decision (D39, D42) before it is built. - The console’s SSH certificate authority (issue #32).
internal/sshcais the user authority the console’s sshd will trust, and nothing else: no password, noauthorized_keys. Its Ed25519 (or ECDSA) key is a secret file (0600, one link, owned by the user, outside every root) read withconfig.ReadSecret, registered with the redactor, and never in a guest; only its public key, as anauthorized_keysline, goes into the console.Generatemakes the key with an exclusive create and never overwrites one.Signturns a client’s public key into a certificate for one session: one principal (a plain user name), a life of ten minutes by default and an hour at most, started a minute in the past for clock skew, a random serial, a key ID that is cut to printable characters for sshd’s log,permit-ptyonly, andpermit-port-forwardingonly when asked for (the editors’ remote modes need it). It refuses a certificate as input, an RSA key under 3072 bits and any key type it does not know. The client keeps its private key, so the supervisor never sees one. The console’s sshd. The console images installopenssh-serverand shipwhr-sshdandsshd_config(in the image tag’s digest). sshd runs once per connection, in inetd mode (sshd -i), as the console user, started by the supervisor withexecand carried over its API, so the console listens on no port and SSH is reachable exactly as far as the API is (nothing new on the VPN). It is not root, because the console drops every capability. The configuration trusts one thing:TrustedUserCAKeysis the authority’s public key, whichwhr-sshdwrites fromWHR_SSH_CAon every start (so a file changed inside the console is overwritten);AuthorizedPrincipalsFileallows the principalwhr;AuthenticationMethods publickeyonly, no password, no keyboard-interactive,AuthorizedKeysFile none, no root, no agent or X11 forwarding, no tunnels, no user environment. Port forwarding, where a certificate grants it, reaches only the console’s own loopback; unix sockets are not forwarded.StrictModesis off because/tmp(mode 1777) would fail it, and the files are the console user’s own and rewritten per connection. The host key is made once in the console’s home volume (whr-sshd hostkeyprints it), so the console keeps its identity and a client can pin it: the private key is made aside and linked in by oneln, so of several first connections at once one wins, and the public half is derived from the winner withssh-keygen -y, so the pair always belongs together. Every file sshd or a login shell reads from/tmp(the authority key, the principal file, the proxy variables) is replaced by a rename, never rewritten in place. The service and the API.Consoles.SSHCertificatesigns a client’s public key (one line, no certificate, a keysshcaaccepts) for the principalwhrwith a key ID ofwhr-<actor>-<time>, asks the console for its host key (whr-sshd hostkey, which must be one Ed25519 line) and audits the certificate (supervisor.console_ssh, which never holds a key).Consoles.SSHstartswhr-sshdin the console withexec, as the console user, withWHR_SSH_CA(the authority’s public key),HOMEand the egress proxy’s variables in the environment and the connection as its standard input; it returns a stream of bytes, ends the sshd on close and reports a bad exit with the tail of its log (-e). At most eight connections at once; none without an authority or without an open console. The API carries it unframed after an upgrade (§9.7), andwhr sshis its client. Measured (verified,TestConsoleSSHLive,go test -tags applecontainer,container1.5.0, Fedora console image): a certificate of the authority for the principalwhrgets a shell aswhrin the read-only root with the proxy variables the supervisor passed on; a certificate of another authority, one for another principal, a bare key and a password do not get in; a certificate without the forwarding permission forwards nothing, and one with it cannot reach another host; a client that pins another host key refuses the console, and the host key is the same on every call, also for six first calls at once. The Ubuntu image is unverified until the same test runs on it.
store/<hash8>-<name>-<version>-<platform>/bin/<name> content-addressed, never modified
profiles/<profile>/bin/<name> -> ../../../store/.../bin/<name>- Verified on download. Claude Code is fetched from the vendor’s release URL and checked against its SHA-256 manifest; Codex CLI is the static musl build from its GitHub release. The store hash is recorded in the run’s audit entry, so every run says exactly which tool version ran it.
- One build per libc. The glibc build of Claude Code ran in fedora, debian and ubuntu and failed in alpine; its musl build ran only in alpine; Codex’s static build ran in all four. The adapter picks the profile from the image’s libc (the vendor installer does the same, by looking for the musl loader). A tool that needs shared libraries beyond libc would need its closure in the store, as Nix does; none of the tested tools did.
- Immutable from inside. The store is mounted read-only;
touch,rm, appending to a tool,chmodand replacing a symlink all failed, and re-hashing every entry afterwards showed no change. Writable state (the agent home, with its auth directory and session) stays on the per-environment volume (§4.4). - Versions are profiles. Environment A ran Claude Code 2.1.286 and environment B 2.1.285 at the same time, each resolving to its own store entry. An upgrade is a new entry and a profile change, and a rollback is a profile change.
- Mounting. A read-only bind mount and a read-only volume both worked, with the same startup cost (about 100 to 130 ms for
claude --version), and four containers shared one store at once. A volume can be attached read-only by several containers but is exclusive while writable (§4.4), so the shared store is read-only everywhere. - Network. The agent starts without network, so the egress allowlist (§7.2) no longer has to permit the download host.
Building the store (issue #74, internal/toolstore, whr tools build -store <dir> -shim <whr-shim>). The pins live in the repository (internal/toolstore/pins.json: name, version, platform, release URL and SHA-256), so a version change is a reviewed commit and never a download-time decision. A download must match the pin and the SHA-256 the vendor’s own manifest.json lists for the platform (the manifest of Claude Code 2.1.285 does, and its linux-arm64 entry equals the hash spike #7 measured); a mismatch in either is refused before anything is stored. Antigravity is different: its pin holds the https archive URL, the archive’s SHA-512 and the extracted file’s SHA-256, and at run time only that pin is trusted, so the build fetches no vendor manifest; the vendor’s manifest and sigstore provenance are checked only when a pin is written or updated (scripts/antigravity-pin-check.sh), never by whr tools build. The build holds Claude Code alone unless -tools names more, so Antigravity is opt-in (an unknown name, or one with no pin for the platform, builds nothing). The binary is placed content-addressed at store/<hash8>-<name>-<version>-<platform>/bin/<name> through a staging directory and a rename, and the entry, its bin and the tool are made read-only (Verify re-hashes every entry and reports a changed or writable one). whr-shim is added from the launcher built from the same commit, and a profile profiles/<tool>-<version>[-<tool>-<version>...][-musl]/ (each tool in the order given, -musl last for a musl build, so Claude Code alone is claude-<version>) links to the entries with relative links, replaced in one move, so an upgrade or a rollback is a profile change. The vendor’s manifest signature (manifestSignatureEnforcement) is not verified here: unverified.
Verifying the store (issue #125). Every entry records the tool’s full SHA-256 in a read-only sha256 file next to it, because the eight digits in the entry’s name are 32 bits and a label, not a check. Store.Verify hashes each tool, opened without following a link, and compares it with that record; a tool with a built-in pin is compared with the pin, and a directory belongs to a pin by its name, version and platform whatever its first eight digits are, so a look-alike with another hash8 is severe, which is compiled into the binary, because the record and the tool are written by the same user and whoever changes one can change both; only a tool without a pin, one added from a file, is compared with the record; with neither only the name’s digits are checked, which is reported as a problem that is not severe. It also refuses (severe) a tool that is writable, that is a link or not a regular file, an entry that is a link or not named <hash8>-<name>-<version>-<platform>, a record that is a link or does not start with the name’s hash, and a tool that has gone. It verifies the profiles as well: each link in a profile’s bin must be the relative form Profile writes (../../../store/<entry>/bin/<tool>) into an existing entry whose tool is the link’s own name; a profile and its bin must be real directories, not links, and each link must resolve (every link followed) to exactly its entry’s bin/<tool> under the store’s root; a pinned tool name may link only to a pinned entry of that name, at the pin’s own entry, so a link to an unpinned version or another hash is severe; an absolute or escaping target, a regular file in place of a link and a link to a missing entry are severe. The store and profiles directories themselves must be real directories, not links (an absolute link would resolve in the guest through the image’s rootfs); the root may be a link. whr serve runs it at start and refuses to start on a severe problem, because every environment mounts the store and runs what is in it (an unchecked entry is logged); whr doctor’s tool-store step fails on one. Verify runs at whr serve start and in whr doctor, not while serving: a change after the start is not seen until the next check. The store is mounted read-only into an environment, so this guards the host’s side: a changed file, a restore from a bad backup, another process of the host user.
The configuration file (issue #74, internal/config) is JSON, decoded strictly (an unknown or misspelt key is an error), so the standard library is enough; TOML would add a dependency for comments only. It holds the repositories (owner/name and clone_depth), the roots (one or more workspace roots, on the internal disk or an external SSD, and the tool store; the repository cache root is gone with D42), the GitHub App ID and its key file, the agent-login env file, the API token file and the listen address. Secrets are paths to files, never values: each must be a regular file, not a link, owned by the user, with mode 0600 and some content, and outside the workspace root, where an agent could read it. The listen address must be a loopback IP (D29: a guest reaches every other address of the host); the roots must exist and must not overlap, since a workspace root is where an agent writes, and a root may not lie inside a git repository. Two optional keys (issue #97): state_dir, where the supervisor’s database lives (default ~/.local/state/whr; it may not overlap a workspace root), and environment, the supervisor’s choice of what an agent environment is: the stock base image (Fedora by default, provisional until pinned by digest, D43), the host names the agent may reach through the egress proxy (api.anthropic.com by default; plain DNS names only, no IP address, wildcard or port, because the proxy refuses a raw IP) and the CPU, memory and disk limits. CheckWorkspacePath decides whether a folder may become a workspace (issue #90): an existing empty directory strictly below a workspace root, resolved and compared by identity, and not inside a git repository, because a human’s own repository is never mounted. whr serve validates all of it at start and reports every problem with its key. The full onboarding stays in #29.
Open: how new versions are discovered (a developer bumps a pin in a commit), and how the store is garbage collected.
5.7 Usage and cost
A remote for agents that run detached for an hour has to say what they used. The supervisor records usage per run and reports it; it never meters the model traffic itself.
- Source: the agent’s own reports. Each turn becomes a
usageevent with a typed payload: model, token counts, the cost (reported by the agent) and the balance when reported, and the usage windows (name, utilization from 0 to 1, reset time). The token counts are optional: a missing block means the agent did not report them, which is not the same as zero, so a token budget never reads an unknown count as 0 tokens. An adapter that reports usage setsReportsUsage, and the suite checks that the events are well formed. Claude Code’sresultevent carriestotal_cost_usdandusagewith input, output, cache-read and cache-creation tokens (recorded in spike #7). Antigravity sendsusageinstep_updateandresult; Codex CLI is unverified. - Reported, never estimated. Every cost is the agent’s own figure (
reported), and so is every other number: tokens, windows and the balance. workharbor does not price tokens from a table of its own, because such a table goes stale. A turn the agent reports without a cost has tokens and no cost, and is counted as such. - Auth mode decides what the number means (§5.2). In
api-keymode cost is real spend. Insubscriptionmode it is API-equivalent, not billed: the plan is paid flat, and the usage-window utilization (five-hour and seven-day, §5.2) is what limits the developer, so the UI leads with that. - One usage window per account (D40). Concurrent runs on one subscription share its usage window, so the UI shows the window as one account-wide figure across all running tasks, not per task, and a quota stop pauses every run on that account with one Decision each (D23).
- Bounds. A turn’s counts are validated at ingest: no negative figure, and no token count, cost (micro-USD) or duration (ms) above 10^11, which is far beyond a real turn and keeps a column of 9 x 10^7 maximal turns below the int64 limit, so the SQL
SUMover the never-purged rows cannot overflow. The Claude adapter drops an out-of-range figure as unreported, and the Go sums saturate rather than wrap. - Kept as audit entries. Usage rows are small and survive a transcript purge (§5.4), so totals stay correct after history is deleted. Each turn is one audit-tier event
usage.recorded(domain.UsageRecorded: run, repository, agent, auth mode, model, tokens, cost and windows), which the service writes in place of the transcript copy; the totals read these events, so a purge changes none of them. The token block and the cost are optional and a turn without them is counted inturns_without_tokensandturns_without_cost, never added as zero. The latest reading of every usage window is kept per account (usage_windows, one row per account and window name, a late event never rolls a reading back): in release 1 an account is the agent (claude), becausewhrnever sees the credential (D40). - Balance. An agent may report what the account has left with a turn, an optional
balancein the usage report next to the token counts: the latest reading is shown (Service.Usage) and a turn without one is unknown, not zero. No agent recorded so far streams one: the Claude Code runs of spike #7 carry no balance field, so whether Claude Code does for an API key is unverified, and the Claude adapter does not fill the field yet. - Reports. Totals per run, task, repository and day or month:
whr usage [--task <task>] [--since <time>] [--json], and per-task usage plus a usage-window meter in the web UI (§9.3).whr show <task>prints the task’s usage line (usage_lineinGET /v1/tasks/{task}).Service.Usagereturns the rows and the account’s windows as the JSON ofwhr usage --json(a golden test pins its shape): a row is one group and one auth mode, so a subscription’s API-equivalent, not billed cost is never added to an API key’s real spend.Service.UsageLineis the linewhr showprints; in subscription mode it leads with the window and calls the cost API-equivalent, not billed.GET /v1/usageserves it (task,repo,since,untilandby, which isall,run,task,repo,agent,model,dayormonth; days and months group in the supervisor’s time zone, not UTC), andwhr usage [--task <ref>] [--repo <owner/name>] [--since <time>] [--until <time>] [--by <group>] [--json]prints it: stdout is the table (with--json, the API’s envelope), the usage windows come first on a subscription, a turn without tokens shows-rather than 0, and a cost is always “reported” and, on a subscription, “API-equivalent, not billed”.--sinceand--untiltake an RFC 3339 time, a date, or how long ago (36h,7d). - Budgets read the same counters. The per-run and per-task token and cost budgets of §7.4 compare against these totals: a soft threshold notifies (§9.4), and a hard limit ends the task as
failed(D13). They are set in the configuration (budgets.per_run,budgets.per_task:max_tokensandmax_cost_usd, andsoft_percent, 80 by default) and checked after every recorded turn (Service.checkBudgets, readingStore.UsageTotals). Tokens are all four counts the agent reported, and the cost is the reported cost, so a turn that reported neither adds nothing and an unknown figure never trips a limit. The soft threshold writes onebudget.warnedaudit entry per task, scope and metric and sends one push; the hard limit stops the live session, stops the task’s runs, supersedes its open Decisions, writes abudget.exceededaudit entry naming the scope, metric, limit and used amount, and moves the task tofailed(TaskAggregate.ExceedBudget). The push says the budget was exceeded, not that the run ended.
8. Resources
Starting estimates, to be replaced by measurement. Assumes API-backed agents, not local inference.
| Workload | Guest memory | Target |
|---|---|---|
| Terminal agent, Git, small scripts | 0.75–1 GiB | up to 6–8 light (stretch) |
| Agent, language server, moderate builds | 1.5–2 GiB | ~4 |
| JetBrains indexing or heavy builds | 3–4 GiB | 1–2 plus light workers |
Review corrections to the original budget:
- 8–10 GiB guests + 4–6 GiB macOS leaves near-zero slack on 16 GB. Plan for 4 concurrent instances on 16 GB, about 6 on 24 GB and about 10 on 32 GB (D32).
Capacity by memory (estimates; issue #39 measures them). macOS, the supervisor and the VPN take 4–6 GiB; a moderate environment (agent and build, including VM overhead) about 2–2.5 GiB; heavy builds or JetBrains 3–4 GiB.
| Memory | Left for environments | Moderate environments | Heavy |
|---|---|---|---|
| 16 GB (budget) | about 10–12 GiB | about 4 | 1–2 |
| 24 GB | about 18–20 GiB | about 6 | 3–4 |
| 32 GB (recommended) | about 26–28 GiB | about 10 | 5–6 |
Storage for about 4 concurrent environments (estimates; issue #54 measures them): macOS, apps and Homebrew 35–50 GB; Apple Container images 2–10 GB; per environment a root filesystem of 1–3 GB and an agent-home volume with build caches of 3–15 GB; repository caches and topic clones by repository size; stopped environments awaiting review stay on disk; keep 10–20% free. That is roughly 120–250 GB in use, so 512 GB is comfortable and 256 GB is tight. A few stopped spike containers already took 8.2 GB on the development Mac.
- Per-VM overhead sits outside the guest limit; a Linux kernel plus a Node-based agent is typically 300–500 MB, so the 0.75 GiB floor is tight.
- Freed guest pages are not returned to macOS: recycling is policy, not an occasional fix.
- Admission control uses host memory pressure (
memory_pressure,vm_stat) plus static limits; heavy jobs are serialised. A simple admission counter ships in release 1; full scheduling is deferred. - Benchmark on the real Mac mini before designing the scheduler.
- Explicitly set CPU/memory;
container machinedefaults to half host RAM and shares the host home.