Decisions
3. Key decisions
| # | Decision | Rationale |
|---|---|---|
| D1 | Reuse runtimes, agents and IDE connections; build a thin supervision layer | Original 9/10 direction, kept |
| D2 | Native Apple Container path via the container CLI behind the runtime adapter. Portainer/Socktainer demoted to an optional compatibility shim (rating 7 → ~4–5) | Three layers, partial compatibility, known exec and restart-recovery gaps; an adapter is needed anyway |
| D3 | Go, single static binary (server, host worker and whr as subcommands), SQLite in WAL mode, embedded web UI (see D8), OpenAPI as the source for the JSON API and the whr client types, SSE for live events | One host, one user; easy launchd packaging, later Linux cross-compile |
| D4 | Agent runner chosen first, by scorecard (§12), preferring one with a structured headless protocol | Constrains the whole model |
| D5 | Approval boundaries are data (a policy table), enforced at the forge adapter, never by prompts | Prompt rules are not a security boundary |
| D6 | Supervisor is a DB-first reconciler, not a process tree | Only realistic answer to reboot/restart gaps |
| D7 | Harbor metaphor is for branding and UI section names only; CLI and API use plain nouns | Guessable, searchable commands |
| D8 | Web UI: server-rendered Go with templ templates, htmx and SSE, embedded in the binary. No Node toolchain and no CSS framework in release 1. The JSON API and the HTML handlers call the same service layer, so nothing is implemented twice | Release 1 is a small remote-control UI: live transcript, messages, start and pause/cancel, and answering Decisions (§9.3). One language, one binary, fewer dependencies and a smaller attack surface on the same origin. Revisit (Svelte) if the UI needs rich client-side state such as inline diff review or a takeover panel |
| D9 | Documentation site: Hugo with the Hextra theme (Go module, pinned version), deployed to GitHub Pages, dark by default with a light toggle | Go toolchain only, no Ruby or Node. Fast builds, built-in search and dark mode. Replaces the earlier Jekyll setup |
| D10 | Decisions are rows in this table. Each is proposed and settled in a GitHub issue labelled decision, then recorded here with its rationale. Separate decision-record pages come only if the design is split into several pages | One place to look; CONTRIBUTING already names this table as the record |
| D11 | Cooperative pause stays an agent capability flag, reported per adapter and never assumed. None of the measured agents has it (spike #1), so their adapters report it false; pause then means a hard interrupt followed by a resume from the agent session, and the UI says so | Keeps the contract ready for an agent that can stop after its current turn, without pretending the current ones can |
| D12 | Release 1 starts with a CLI-only vertical slice: whr run <issue-url>, then whr logs -f, whr say and whr cancel, and whr approve pushes the prepared agent/* branch and opens the PR, with Claude Code in Apple Container on one forge. The web UI, PWA, SSH, notifications and the Codex CLI adapter follow in the rest of release 1 (the Codex adapter is a target, not yet built: #35) | Proves the service layer, adapters and policy end to end before any UI, and gives D8 a working API to check against |
| D13 | Task state machine (§4.1): a terminal failed state, rework from ready_for_review back to running, and awaiting_guidance only for blocking Decisions raised by a live run, so “Ready to push?” leaves the task in ready_for_review | Gives exit code 10 a state, makes the rework path explicit, and keeps the review gate from looking like a stalled run in the inbox |
| D14 | CLI framework: spf13/cobra, without viper (issue #46). The root command silences cobra’s own error and usage output, sets stdout and stderr explicitly, and main maps errors to internal/exitcode (§9.2) | The design needs generated shell completion with dynamic task and workspace IDs and a noun-verb grammar with aliases (§9.1, §9.2); cobra provides both and can generate the CLI reference. Rated above kong, urfave/cli and the standard library flag |
| D15 | Release 1 forge: GitHub, through a GitHub App installation (issue #6). The App is the bot identity: its installation tokens last about an hour and are scoped to the repositories it is installed on and to the permissions it asks for (contents and pull requests write, issues write, metadata read). Only when a board is configured (D30) does it also ask for the organization permission organization_projects: write; the manifest, the installation token and whr doctor then all expect it, and without a board none of them does. Its private key lives in the credential service. A ruleset on the default branch requires a human review and lists no bypass for the App, so it cannot merge (§6); today the active ruleset main blocks deletion, merge commits and force pushes with no bypass actor, and the required review is added with the App’s push flow (open, #27, #73). Gitea, Forgejo and GitLab follow behind the same forge adapter | The repository and its CI already live on GitHub, so the limits can be checked against a real ruleset at once. Short-lived, repo-scoped tokens match §7.3 better than a long-lived PAT, and an App is a separate identity that commits and audit entries can name. That an App’s installation token cannot bypass a ruleset without being listed as a bypass actor is unverified until #27 tests it |
| D16 | Persistence (§4.4): repositories live on the host and are mounted into environments; the agent home (auth directory, session, caches) is one writable named volume per environment; build caches stay on that volume outside the checkout; the root filesystem is disposable | Measured in spike #2: volumes and bind mounts survive delete, the root filesystem does not, and caches on a bind mount roughly double warm builds |
| D17 | Replaced by D42. Agent checkouts are per-task clones of a bare cache whose objects are mounted read-only (§4.5); git worktree stays the developer’s own tool and is never handed to an agent | A worktree’s .git exposes the shared repository’s branches, hooks and config (spike #2, item 9) |
| D18 | Agents never push (§4.5, §6). The agent commits in its checkout; the supervisor rebases, folds and checks the topic, opens a “Ready to push?” Decision per commit SHA, and pushes and opens the PR only after approval | Nothing leaves the host unreviewed, and policy is enforced outside the agent (D5) |
| D19 | Stock images plus a shared read-only tool store (§5.6): agent CLIs live once, content-addressed, on the host and are mounted read-only; versions are profiles | Spike #2: installing per container cost about 11 s and 230 MB; the store is immutable from inside and shared by several containers |
| D20 | Adapters are built in for release 1 and out-of-process plugins later (§5.5), never Go’s in-process plugin package | Two built-in adapters prove the contract first; third-party code stays out of the supervisor process |
| D21 | Task state machine, amending D13 (§4.1): a run paused by auth_expired or quota_exhausted moves its task to awaiting_guidance; a task fails only from running or awaiting_guidance, and a lost workspace in ready_for_review opens a review Decision (rework or cancel) | Settles the gaps found in the review of the D13 implementation without new transitions, so the code in internal/domain already agrees |
| D22 | Build the supervision layer; adopt none of the agent-task supervisors (issue #5, confirms D1). OpenHands, Vibe Kanban, Sculptor and Coder Agents were assessed from their docs and repositories (§11). Borrow: ACP as a candidate generic agent-adapter protocol (§5.5) and OpenHands’ confirmation states; Sculptor’s Claude control-protocol integration and editable message queue; Claude Remote Control’s phone UX as a reference and a fallback for Claude | None meets the non-negotiable parts of release 1 together: Apple Container, default-deny egress per environment, approvals routed to a human for Claude Code and Codex under subscription logins, an agent that never pushes, more than one forge. Adapting one would replace its runtime, policy and forge layers, which is most of workharbor. Vendor remotes cover one vendor and push to GitHub only. Desk research only: the claims marked unverified in §11 were not tried |
| D23 | Decisions around pauses (§4.2): auth_expired and quota_exhausted are blocking question Decisions with fixed options (re-login and resume, resume now or at the reset, cancel) and no deadline; pausing a run supersedes every open Decision the run raised, questions and approvals alike, and they are raised again when the agent asks after resuming | Nothing is permitted by a login or quota answer, so it is a question, and waiting on it is safe, so it does not fail closed. A paused agent’s process is gone (D11), so an answer could only reach a dead process or the wrong request; superseding reuses the restart rule |
| D24 | Releases are human-signed tags built by GoReleaser into draft releases, distributed through a Homebrew tap (§13, Releases). The human pushes a signed v* tag; a workflow checks it and has GoReleaser build whr, checksums, an SBOM and a build-provenance attestation into a draft GitHub release; you publish it, and only then is the tap updated. Versions are 0.x until the JSON API, the adapter contract_version and the database migrations are stable; the first release, v0.1.0, is cut when the release 1 slice demo (issue #28) passes on the Mac mini. Before it, dogfood builds are prerelease tags (v0.1.0-alpha.N) whose draft is never published: the administrator installs a draft with make install-release, which checks the checksums and the build-provenance attestation, into a prefix whr cannot write. The tap is a formula rendered by its own workflow, and only a published, non-prerelease release updates it. Signatures (issue #180): the build-provenance attestation is the release signature. The release job signs it keylessly through Sigstore (id-token: write in that job only), bound to the release workflow, the tag and the tagged commit, over the digest of every artifact listed in checksums.txt; its Sigstore bundle is also attached to the draft as a release asset, so a download can be checked offline. No GPG key and no long-lived signing key in CI, and no second cosign signature over checksums.txt, which the attestation already covers. Verification (make install-release and the manual) pins the workflow identity, not just the repository. A development installation (issue #261) is the one exception to the admin-owned prefix: make install from an approved commit into $HOME/.local or the developer’s --prefix, accepted only when whr setup and whr doctor are given --dev on that call, or when the configuration holds development_prefix (issue #276, to be built); the supervisor’s own account can then replace the binary (threat model, Accepted risks), so it is for a developer’s machine, never the dogfood or reference host. Its prefix is refused when it is /, the account’s home directory or a directory above it, the home taken both from $HOME and from the directory service (/usr/bin/dscl, run by its absolute path, never found through PATH; a failed or empty lookup fails --dev closed). Directories are compared by identity, never by spelling: the prefix is stated once and compared with os.SameFile against the home and each directory above it, because the default APFS volume ignores case and /System/Volumes/Data reaches the same directories under another name; development_prefix and its file checks follow the same rule (issue #278: the lookup is built in e664c8a, the absolute path and the identity comparison are to be built). These checks guard against a mistake, not against the account itself, which sets its own $HOME and PATH and can replace the binary it installed anyway. The remembered development installation (Werner’s decision on #276): development_prefix, a top-level key of the supervisor’s configuration holding the absolute prefix, is written by whr only on an explicit whr setup --dev, shown and confirmed like any fix, and removed by whr setup --managed or by hand; whr never writes it through an environment variable (no WORKHARBOR_DEV, none read for it), a default, whr doctor, whr serve, whr service, a repair line or a workspace’s or repository’s file. whr setup, whr doctor and whr service install read it as --dev --prefix <value>, an explicit --dev or --prefix wins, an explicit --prefix without --dev being a managed call that ignores the key, and whr serve only logs it at start. Anything running as the account can write the key too, since the file is the account’s own; what limits it is that a managed installation refuses it and that the account can already replace a development installation. It loosens nothing beyond --dev: every --dev check runs again on each read, and the key is refused, as a configuration error, when its value fails any check of the --prefix flag (absolute, no control, bidirectional or separator characters) or is a managed prefix, when whr runs from a managed prefix, or when the configuration file is not a regular single-link file of the running account or root, closed to group and other writers, outside every workspace root and git working tree, checked on the opened descriptor (O_NOFOLLOW, then fstat) as the secret files are, never by path and then opened. The workspace roots come from that same file, so that check catches a stray file, not a crafted one; the managed-prefix refusal is the control. whr doctor reports it as warn on every run, naming the key, the file and whr setup --managed, so the weaker mode is never silent | Tags and releases are already human-only (§6, AGENTS.md), and a signed tag is the trust anchor; the tag is also the Go module version, so there is no version file. GoReleaser covers cross-builds, checksums, SBOMs, signing and the tap in one pinned tool. The draft is where you check the assets before anyone can install them, so the tap must not point at a draft. A tap avoids notarizing a downloaded binary for now. A formula, not a cask: Homebrew does not quarantine a formula’s download, while a cask of an unsigned binary would need its quarantine attribute removed (both unverified until #62 installs it). A draft is downloadable only by a repository writer, which suits dogfooding: the binary the supervisor runs was built by CI from a signed tag on main and attested, not on the host, and an admin-owned prefix means nothing running as whr, an agent’s escape included, can replace it. Keyless signing leaves no key to store, rotate or steal, and its transparency-log entry names the workflow and commit, which a GPG signature made from a CI secret would not; the attestation already writes that entry today, so attaching its bundle publishes nothing new. That OpenSSF Scorecard counts the attached bundle as provenance is unverified until a release carries it (#180) |
| D25 | Launcher and host-initiated cancel (§5.1, D19): agent and command processes run under whr-shim, a small static binary in the shared tool store. The launcher starts the command in its own process group (setpgid) and writes the PID to a file. Cancellation is triggered from the host via container exec <id> /tools/whr-shim kill -pidfile <path> -grace <duration>, which sends SIGINT to the whole process group and falls back to SIGKILL after the grace period | Signalling the container exec client fails in Apple Container (missing signal in xpc message) and leaves processes running in the guest (spike #2; issue #10, whose cancel of a real agent is pending in #7). whr-shim terminates the entire process tree reliably from inside the guest without leaving orphan processes (measured in spike #10: cooperative cancel exits in ~10 ms, stubborn trees killed after 500 ms grace in ~514 ms, 0 orphans) |
| D26 | Approval channel: the stdio control protocol (§4.2, §5.2, §12): Claude Code runs with --permission-prompt-tool stdio and stream-json in and out. Permission requests arrive as control_request (can_use_tool) messages on the exec stdout, and the supervisor answers with a control_response (allow or deny) on stdin. Whatever ends the channel, the run ends with a denial: the supervisor stops the agent with whr-shim (D25) and never lets a request outlive it | No network listener on the host or sidecar, no path from the internal network to the supervisor, and no approval token the guest could read. Evidence (verified, spike #7, spike/agent-approval at 567b5ad, scripts and raw streams committed): in a container on Apple Container 1.5.0, allow and deny round-trip and the tool runs or does not; a closed stdin fails closed without running the tool; a killed container exec client leaves the guest agent alive with its request pending and the tool unrun, and whr-shim reaps it in about 0.6 s with no orphans; an unanswered request holds the tool until the deadline stop; a response with an unknown request_id is ignored. Hence the whr-shim stop whatever ends the channel |
| D27 | Resume briefing (§4.1, §4.2): when a run resumes after a pause, a cancel or an interruption, the supervisor’s first message to the agent says what it knows: which tool call was running or waiting for approval when the process ended, that its effects are unknown and may be partial, and which open Decisions were superseded. The agent is told to check the workspace before repeating anything | Measured in spike #7, case 7 (branch spike/agent-approval): after a loop was killed 3 s into a Bash call, Claude Code’s resumed session told the model “The command was never executed … rejected before it could run”, which was false. Left to the agent’s own account, a resumed run may skip or repeat work |
| D28 | Host software through Homebrew; agent CLIs only through the tool store (manual, host setup). The Mac mini gets Apple Container, whr (the tap of D24) and a VPN client from a Brewfile installed with brew bundle, with HOMEBREW_NO_AUTO_UPDATE=1 so nothing upgrades unasked. Claude Code, Codex CLI and other agent CLIs are never installed on the host for workharbor: they come from the verified, versioned tool store (D19). Nix (nix-darwin) remains possible for a host that already uses it | The host needs few packages, and D24 already distributes whr through a tap. Nix would be reproducible but heavy for one appliance, and it duplicates what the tool store does for agents. Whether Apple Container is packaged in nixpkgs is unverified |
| D29 | The web UI listens on loopback only and a forwarder carries remote access to it; the JSON API is on a host-only socket; the token is always required (§7.5). The web UI binds to 127.0.0.1, which no guest reaches (issue #69), and only it is forwarded. The JSON API is never forwarded: it is served on a unix socket in whr’s state directory (directory 0700, socket 0600), reached only by the host CLI as the whr user, and the forwarded listener serves no /v1 route, so a leaked API token cannot answer a review, allow an egress host or enrol a passkey from the phone network (D45; review of #101). The phone reaches them through a forwarder: tailscale serve (the default), or, for a VPN that ends on the router such as a FRITZ!Box with WireGuard, a small proxy on the Mac’s LAN address admitted by a pf rule to the router’s VPN clients only. Every request needs the API token. A pf rule blocks the container subnets from the host’s own addresses, and host services that listen on all interfaces are turned off or hardened | Issue #69 (branch spike/host-reachability, run.out) measured that every guest, --internal included, reaches host listeners on the LAN address or on all interfaces, and none reaches a loopback-only listener; spike #2’s “--internal blocks the host” was wrong. A forwarder on any non-loopback address is reachable by guests too, so the token, not the address, is the guard there, and pf narrows it. Whether a guest reaches the Tailscale address, the pf rules themselves, and the Application Firewall’s effect are unverified (issue #69) |
| D30 | The forge board mirrors task state; the supervisor writes it (issue #70). Through the forge adapter, as an optional capability, the supervisor keeps a project board current: a task awaiting guidance moves its card to “Needs you” (first), running → In progress, ready for review → Ready to push, completed → Done; Session names the agent, and the card links to the task in the web UI; a failed task moves its card to Needs you, because a human has to look, and a cancelled one back to Todo, so no card stays In progress for a task that has stopped; queued changes no card. The write is the supervisor’s own update_board action (auto in §6), never an agent’s. Whether an App’s installation token can write a user-owned project at all is unverified: the permission of D15 covers organization projects, so a user-owned board may need an organization-owned board instead; that and the real board’s field names wait for the check with the real App (#73). Agents never write to the board. Starting a task by moving its card to an agent queue comes later, only through an “Accept this task?” Decision and the trust tiers (issue #71) | The GitHub board is a good planning dashboard and the web UI the control surface; mirroring keeps one current view without rebuilding a board in workharbor. Whether an App installation token can write fields on a user-owned project, and whether card-move webhooks exist for one, is unverified: issue #70 tests it first, and moving the repository and project into an organization is the fallback the human decides on |
| D31 | GitHub is reached through its API from Go, with the App’s installation token, never through gh (§10). The forge adapter has a typed client for the REST API (issues, pull requests, rulesets) and the GraphQL API (Projects v2 fields), mints installation tokens from a JWT signed with the App key using the standard library, and handles rate limits and errors as typed values. gh stays a developer and agent-session tool for this repository, not part of the product (issue #27) | gh would carry the user’s own broadly scoped token, the wrong identity for a bot, add a host dependency against D28, and leave errors and rate limits to output parsing. Libraries such as go-github and githubv4 are added only if the hand-written client grows large enough to justify them (AGENTS.md) |
| D32 | Recommended host: Mac mini M6 with 32 GB memory and 512 GB storage; on a budget 16 GB (§2, §8), with 512 GB, or with 256 GB plus an external SSD for repositories, workspaces and backups; 24 GB / 512 GB sits in between. Spend on memory before storage, and add an external SSD rather than paying for 1 TB internal. The M5 Pro is not worth its premium for API-backed agents | US Apple Store prices on 1 October 2026: M6 16 GB / 256 GB $899, 16 / 512 $1,099, 24 / 512 $1,299, 32 / 512 $1,499, 24 GB / 1 TB $1,599, 32 GB / 1 TB $1,799; M5 Pro 24 / 512 $1,699 (Germany: M6 from €1,049). Each memory step costs $200 and buys about four more concurrent environments; memory cannot be upgraded later, while storage can be added externally. That makes 32 GB about $150 per environment against about $215 for 24 GB and $275 for 16 / 512. Environment counts are estimates until issue #39; whether Apple Container’s storage can move to an external SSD is unverified (issue #54) |
| D33 | Web app previews go through a preview proxy in whr (§9.3, issue #72). When an agent runs a dev server in its environment, whr proxies a preview of one declared port through the proxy sidecar, the only container on both networks, and tailscale serve (or the router-VPN forwarder of D29) carries it to the developer. Each preview has its own origin, never the web UI’s; it needs a per-preview token, lives only while the environment runs, forwards only to that port, and passes WebSocket upgrades for hot reload | The developer’s phone or laptop cannot reach an --internal environment, and should not; the supervisor already knows which task, environment and port belong together. A preview serves untrusted, agent-written code in the developer’s browser, so sharing the UI’s origin would let it read the session and answer Decisions. Port publishing straight to the host, Traefik or Caddy, Tailscale inside each guest, and Tailscale Funnel were rejected: they bypass the supervisor, need routing data it already has, put a key in the guest, or publish unreviewed code. That the sidecar can relay inbound traffic to the internal network is unverified (issues #69, #72) |
| D34 | Dogfood first: workharbor develops workharbor as early as possible (§13). A Dogfood milestone holds the smallest set that runs one real workharbor issue through whr end to end: whr serve and the core commands (#24), the Claude Code adapter in degraded mode (#25, dontAsk with a fixed allowlist; host approvals follow with #7), the Apple Container adapter (#26), push after approval (#27), the reconciler fixes (#66) and the adapter’s permission fix (#68), on a Mac mini M4 with 16 GB (#73). The supervisor always runs an installed binary built from an approved commit on main (a dogfood draft release installed with make install-release until the first release, then the tap, D24; make install stays for a developer’s own machine), never a topic’s working tree. From the first green run, new issues start with whr run, and each manual workaround becomes an issue labelled dogfood | Today the human supervises three agent sessions by hand: relaying messages, pushing, ticking criteria, keeping the board, and catching duplicated work and a leaked token, which are all workharbor features. Degraded mode works now and takes #7 off the critical path; the push stays human-approved (D18). Agents working on workharbor edit the code that constrains them, including the policy, so the running supervisor must come from reviewed code. A host process runtime was rejected: it would be faster but would normalise unisolated agents |
| D35 | The phone and a 12-inch tablet are the primary clients (§9.6). The phone serves short, urgent interactions (answer, approve a tool, stop a run, glance at the harbor); the tablet replaces the laptop for reviewing a topic before push, supervising several tasks and planning with an agent. One server-rendered UI with a phone layout and a two-pane tablet layout, installed as a PWA. Approving “Ready to push?” asks for a passkey on any device | The developer detaches while agents work and returns when one needs them (§1), which happens away from a desk; a 12-inch tablet with a keyboard covers the review that the phone’s screen cannot. Publishing code is the one irreversible step a lost or unlocked phone could take, so it alone needs a fresh check of who is approving. That a passkey prompt works in an installed PWA on both devices is unverified |
| D36 | The autonomy table’s defaults and fixed floor (§6, issue #9). Commit in the topic’s checkout: auto; push an agent/* branch: ask, carried out by the supervisor after the “Ready to push?” approval; open or update a PR and comment on the issue: auto, after the push; merge, tag, release and deploy: forbid. Whatever a repository’s table says, merge, tag, release and deploy stay forbid and push stays at most ask; an override may only tighten; an unknown action or mode is forbid | Implemented and tested in internal/policy (issues #4, #51); this row records it as decided. Sensitive actions triggered by untrusted input asking (§6) follow with the trust tiers (issue #53) |
| D37 | The CLI grammar of the dogfood slice is stable (§9.1, issue #9): whr serve, run, ls, logs -f, say, cancel, inbox, approve, reject and answer <decision> <option> (for a question’s fixed options, D23), with the scripting contract of §9.2. The other commands stay provisional until they are built | These are the commands the first dogfood run uses (D34); fixing their names now lets #24, scripts and the manual (#65) rely on them. answer is separate from approve and reject because a question has its own options, not allow or deny |
| D38 | Repository configuration uses established files, read from the default branch only; security settings live only on the supervisor’s side (§4.5, §5.1, issues #76, #77). The environment comes from .devcontainer/devcontainer.json (or a Dockerfile/Containerfile), of which only a safe subset is honoured: image, build, features, containerEnv, postCreateCommand inside the environment, forwardPorts as preview ports (D33) and customizations.workharbor; initializeCommand (it runs on the host), mounts, runArgs, privileged, capAdd, securityOpt and a root user are refused. Agent instructions come from AGENTS.md only, not CLAUDE.md. Checks before push come from the repository’s own check command or .pre-commit-config.yaml, commit rules from its commit linter, toolchain versions from go.mod, .tool-versions/mise.toml and the like, and suggested registry hosts from lockfiles. No .workharbor file: the few hints workharbor needs go in customizations.workharbor. The autonomy table, push and merge rules, mounts, the agent’s permission mode and any egress beyond what the human confirmed come only from the supervisor’s config; a repository may request an egress host, which becomes a Decision whose answer the supervisor stores Features (#108): the supervisor fetches each feature on the host as an OCI artifact, with a size cap, a timeout, a digest check against the manifest and safe extraction (no absolute or .. paths, links or devices); the reference is resolved to its manifest digest when the environment is read from the default branch’s commit, the image tag is keyed on that digest and the digest is recorded, so a moved tag changes nothing until the next resolution. Features from ghcr.io/devcontainers/features/ are allowed; any other source needs a Decision, kept per repository like an egress host. A feature with privileged, mounts, capAdd, securityOpt, init, entrypoint or a lifecycle command is refused; its containerEnv becomes ENV in the image, where PATH may change but the proxy variables, LD_PRELOAD, LD_LIBRARY_PATH, HOME and the supervisor’s and vendors’ prefixes are refused; dependencies apply in the file’s order, a cycle is refused and a missing one is noted, never fetched. | Established files work for humans and other tools too, and the devcontainer spec has an official extension point. An agent can edit every file in its topic, including configuration, so reading only the reviewed default branch keeps an agent from configuring its own environment or permissions; even reviewed files may only request, never grant, security-relevant settings. Claude Code v2.1.277 or later reads AGENTS.md by default when no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md exists in the working directory or above it (Claude Code documentation, “How Claude remembers your project”, read 1 October 2026; the pinned 2.1.285 qualifies), and Codex CLI reads AGENTS.md, so one file serves both. So the supervisor keeps every CLAUDE.md-family file out of the directories above the workspace and passes the instruction-file setting (pluginConfigs of the built-in agents-md plugin) in its own --settings file, which project and local settings cannot override (#68). Whether devcontainer features build on Apple Container is unverified |
| D39 | Dependency and build directories stay off the bind-mounted checkout: relocated to a per-environment volume where the tool allows it, mounted over the checkout only where it does not; the checkout stays on the host (§4.4, issue #80; refines D16). Where a tool reads its output location from the environment, that location points to a per-environment volume outside the checkout (CARGO_TARGET_DIR, UV_PROJECT_ENVIRONMENT; the Go caches already live on the agent home), so the checkout has no mount point. Only a directory a tool cannot move (node_modules for npm) gets a volume mounted over it; it is listed per repository in volume_dirs, with defaults by lockfile, and owned like the agent home. A tracked path is never relocated or covered. The supervisor warns when a checkout tracks more than 50 000 files; putting a whole checkout on a volume for such repositories is open (§12) Amended (#80): in release 1 node_modules stays on the bind mount, because git worktree add refuses a non-empty directory the runtime would have to create first, and an environment’s mounts are fixed at creation, so an agent added later would get none; the real fix is the worktree itself on a volume (open, issue #119). Amended (#119), the shape: if worktrees move, they move onto one worktree volume per workspace environment, with a directory per agent worktree, never one volume per agent, because a container’s mounts are fixed at creation and an agent added later must get its worktree without recreating the environment. The volume belongs to the workspace, not the environment: recreating or rebuilding the environment keeps it, and only the workspace’s retention (issue #54) deletes it. The agent clone and its object store stay in the workspace folder on the host, which stays the human’s view; commits still leave only as a bundle exported from the running environment (D42), and the host still never runs git in a workspace. Whether worktrees move at all waits on the measurement in issue #119 (git status, install and build in a worktree on a volume whose git directory stays on the bind mount, against spike #40, and how the editor copy and the bundle export reach it) open. | Spike #40 (branch spike/bind-mount-perf, results.txt): work that touches many files is 20 to 100 times slower on a virtiofs bind mount than on a volume (40 000 small files: warm git status 1.8 s against 0.04 s, copy 35 s against 0.4 s; 153 000 files: 8 s against 0.1 s, copy 128 s against 1.5 s), while a CPU-bound build costs the same. Dependency trees hold most of the files in a typical repository, so this removes most of the cost while hostgit, the editor copy and the shared object store keep working on the host (D16, D17). A volume mounted over a subdirectory of the bind mount works on Apple Container 1.5.0 (branch spike/volume-over-bind, results.txt): the guest sees it, the host sees an empty directory, and the contents survive stop, start and a rebuild. But the mount point cannot be removed or renamed in the guest (rm -rf, rmdir and mv fail with EBUSY), a fresh volume shows lost+found, and the volume hides any host files at that path; relocation has none of these problems, so it comes first. Whether npm ci and npm install work over such a mount point is unverified (issue #80) |
| D40 | A subscription login stays within the vendor’s terms: whr never touches the credential, and only the human starts runs (§5.2, §6, §7.3, issue #81). whr never reads, stores, relays or logs a subscription credential: the human signs in inside the environment through the vendor’s own flow, and the CLI keeps the login on that environment’s agent-home volume (D16); only an API key may be configured on the supervisor’s side. One human per supervisor, and that human starts every run (whr run, the UI, or accepting a card’s Decision, #71); forge events and webhooks only fill the inbox, there is no scheduled run or unattended backlog, and queued means a started run that waits for host resources. Only the unmodified vendor binaries run (D19), with none of their sign-in methods removed. A quota stop pauses the run with a Decision (D23), with no retry and no account rotation. Concurrency is not capped for licence reasons: every run draws on the account’s one usage window, shown as one figure (§5.7), and admission is by host resources (§8) | Anthropic’s Claude Code terms (legal and compliance page, read 1 October 2026): subscription sign-in serves ordinary use of Claude Code and Anthropic’s own apps; developers may not collect, store or intermediate Claude.ai credentials or session tokens, and sign-in completes through Anthropic’s own flow; an end user may sign in to the unmodified binary with their own subscription, also where a platform hosts it; limits assume ordinary, individual use. For Codex the ChatGPT Terms of Use apply, and OpenAI recommends an API key for CI/CD. The earlier plan, a claude setup-token token kept in a host file and passed to every environment (spike #2, issue #73), had whr store and pass the credential. A one-run cap was rejected (the terms regulate through usage limits, and parallel topics are the point of §1), and so were mandatory host approvals in subscription mode (a security setting, D36, that would block degraded-mode dogfooding, D34). How sign-in works inside a hardened environment, and whether it can be done from the phone, is open (spike #82); the reading of the terms is unverified legal advice: it is ours, not the vendors' |
| D41 | Auth modes are phased: subscription logins with guardrails first, API keys in the medium term, an enterprise offering in the long term (§5.2, §13, issues #81, #83). Release 1 supports subscription logins under D40’s guardrails (the human signs in inside the environment, starts every run, and a quota stop pauses; the shared usage window, #48; the agent’s permission mode, §6), and an API key only through the configuration’s key file. In the medium term API-key mode becomes a full peer: a host-side key proxy keeps the key out of the environment (§7.3), with spend budgets and per-key usage (#83). In the long term an enterprise offering serves several developers under the vendors’ commercial terms; that changes D40’s one-human rule and needs its own decision | The human’s product direction (1 October 2026). Subscriptions are what the developer already pays for and what dogfooding uses (D34); API keys are the vendors’ path for automation and products; several developers on one supervisor are outside consumer terms (D40), so they wait for commercial terms and a per-user identity model |
| D42 | Workspaces with named agents replace per-task clones (§4, §4.4, §4.5, issues #88, #90, #91; replaces D17). A workspace is a folder on the host or an external SSD, created by the human and mounted into one long-lived environment. It holds its own agent clone, seeded from the forge or the human’s repository, and never a worktree of the human’s own repository: that repository’s .git is never mounted into a container. A workspace hosts several named agents (for example docs, runtime, review), one worktree and branch each (agent/<role>), each with its own process group, session, approval channel, instructions and permission profile; tasks are assigned to an agent, which persists while tasks come and go. Agents keep their branches rebased on the workspace’s integration branch (the one its workflow preset names: the configured integration_branch, develop, or the forge’s default branch; any valid ref name outside agent/, issue #229), and a conflict becomes a Decision. The host never runs git in a workspace: commits leave as a git bundle streamed out of the running environment, verified, and then prepared and pushed after a per-SHA “Ready to push?” as before (D18). workharbor gives agents no shared cache and no per-task clones. The supervisor keeps one bare mirror per forge repository for itself, the old cache without its clones: fed only from the forge, never from a workspace, never mounted into an environment, and the source of the target branch, its history for the merge base, and the base the imported bundles build on. The human may use worktrees in their own repositories, which workharbor never touches. Accepted: the agents of one workspace are one trust domain and can change each other’s worktrees. It arrives in two steps: dogfooding runs one named agent per workspace (#90, #91), and several agents in one environment at once follow in release 1 (#94) | The human’s direction (1 October 2026), matching how workharbor itself is developed: one worktree per long-lived, role-named session, rebasing onto main. Several agents in one environment save a VM each, which matters on the 16 GB host (D32). Spike #2, item 9, measured why the human’s own repository stays out: a hook and a core.fsmonitor command planted in a mounted shared .git ran on the host, and a shared .git exposes every branch; a separate agent clone keeps both inside the workspace. Spike #89 measured the export (branch spike/workspaces): a bundle streamed from a running environment took 108 ms for 3 commits and 1.8 s for 401 commits with a 60 MB blob, and nothing planted in the clone’s .git ran on the host; git bundle verify passed a bundle truncated to half its size, so the import is the fetch into a repository without the objects, which refused both a truncated and a corrupted bundle. It also confirmed that one agent can edit another’s worktree and move its branch. Several agent sessions on one sign-in (#82) and a workspace on an external SSD are open |
| D43 | A console environment for the human’s shell work; only the admin logs in to the host (§7.4, §7.6, issues #88, #92). A console is an environment without an agent that mounts the workspaces root, read-only by default and read-write per workspace on request, for browsing, editing, git and tests across workspaces. Fedora and Ubuntu LTS are the first-class base images, for the console and for agent environments alike: pinned by digest and covered by the runtime conformance suite, with Fedora the default; other bases work best-effort. The console image adds zsh and fish, git, jq, curl, ripgrep, tmux, an editor and make; it has no Docker CLI and no container engine, and the human can swap or extend the image within D38’s safe subset. It holds no forge or agent credentials, sits behind the egress sidecar, and runs git with hooks and fsmonitor off through GIT_CONFIG_* environment variables, because a workspace’s .git is agent-written. whr console and SSH (#32) land in the console; logging in to the host is for the admin (OS and whr upgrades, recovery) | Daily shell work on the host would run the agents’ planted repository config there and needs a host account per user; inside the console it runs in a VM with no host home, secrets or runtime socket. Fedora is the base every spike measured under Apple Container 1.5.0, its repositories carry the tooling, and it uses glibc like the tool store’s builds (spike #2); Ubuntu LTS is first-class too because many devcontainer base images (D38) are Debian or Ubuntu; spike #89 built console images on Fedora 44 and Ubuntu 24.04.5 LTS, and the runtime conformance suite passed on both. Alpine’s musl and busybox differ from what the agents’ toolchains expect, so it follows later: first for the console (#93), then for agent environments. All three agent CLIs publish linux-arm64 musl builds (spike musl-cli, #152; checked at the vendors’ sources, not run). What still blocks it is a run of each CLI and the adapter’s conformance suite on Alpine in Apple Container (#161) and musl pins in the tool store with the build picked by the base’s libc (#162) open; Nix on a base image is a later option for pinned tooling. Spike #89 measured the git settings: GIT_CONFIG_COUNT with core.hooksPath=/dev/null and core.fsmonitor=false stops hooks and fsmonitor, but filter drivers, textconv and aliases set in a workspace’s config run unless each is overridden by name, so the console’s git wrapper overrides every filter.*, diff.*.textconv and alias.* it finds, and what remains is an accepted risk inside the console VM |
| D44 | Without a devcontainer, environments run a workharbor base image with git (§5.6, issue #98). When a repository has no devcontainer.json (D38), the environment’s image is one whr builds once from a Containerfile it ships: a first-class base (D43; Fedora by default, Ubuntu LTS as the alternative) pinned by digest, plus git and CA certificates, built through the same builder as a repository’s image (#76) and rebuilt when the pinned base changes. A repository’s own image must provide git too; one without it is refused when the workspace is created, with a message that says so, rather than failing later in git worktree add. Git is not put into the tool store | The environment needs git for the agents’ worktrees and for the bundle export (D42), and the stock Fedora image has none (the serve integration run, issue #28 groundwork). The tool store holds the agent CLIs and the supervisor’s own static helpers (D19), and there is no pinned, verifiable static git build to put there; relying on each repository’s devcontainer helps only repositories that have one. A small built base keeps D19’s point, one copy of the tools instead of an install per container, while git, a system package, comes from the base’s own package manager |
| D45 | The human proves who they are with a passkey on the phone; sensitive answers need a fresh one, bound to the Decision (§7.5, issues #100, #101). The web app on the phone and tablet (D35) signs in with a passkey (WebAuthn) that requires user verification (Face ID or a fingerprint), bound to whr’s HTTPS name behind the forwarder (D29); whr stores only its public key. Answering a sensitive Decision (“Ready to push?”, allowing an egress host, a policy change, an operation on a secret) needs a fresh user-verified assertion whose challenge names the Decision and, for a review, its commit SHA, so the approval covers exactly that commit (D18). The static API token stays for the local CLI on the host. No TOTP; recovery is the admin enrolling a new passkey on the host (D43). Unlocking whr’s secrets after a restart with a passkey (WebAuthn PRF) is a later option (#140) | workharbor runs headless and is operated remotely, so a prompt on the host (a passphrase, Touch ID, a hardware-key touch) cannot be answered, and whr needs its own secrets unattended: those stay with the whr user (D40, host setup). The human’s identity is therefore checked on the phone. A passkey leaves no shared secret on the host and cannot be phished to another site; a TOTP seed would be one more host secret, its codes can be relayed, and a code does not bind to the commit being approved. Two leaks in this project’s own development (a token in a message, a key in a committed script) showed that what matters most is that values never pass through a session, a script or a message |
| D46 | First-time setup is a CLI wizard over the doctor’s checks: whr setup host as the administrator, whr setup as whr (§9.5, manual host setup, issue #104). Every whr doctor check is a setup step with an optional fix: the exact commands it runs, shown before they run, or a guided text for what a CLI cannot do. whr doctor runs every check, both phases’ included, read-only, and names on each failing or not verified line the setup command that fixes it (whr setup host --only <step>, whr setup --only <step>; #153); whr setup checks, fixes after a confirmation and checks again, so the two cannot drift. whr setup host (the administrator’s part: the whr user, power, firewall, SSH keys only, FileVault status, the Brewfile, the admin-owned prefix of D24, or with --dev a development prefix the running account owns, issue #261) never runs as root itself and refuses to run as root or as a standard whr (an administrator account may run both parts, D49): each privileged command goes through sudo as its own argv. whr setup (the whr part, in its desktop session: the container system, secrets, the configuration, the GitHub App, the tool store, the launchd job) writes secrets only as generated or typed-without-echo 0600 files and never overwrites one. Steps macOS or a third party keeps for the human (Screen Sharing under TCC, enabling FileVault, Tailscale sign-in, the App’s confirm click, the ruleset check) are guided and then checked. Every step is idempotent; --dry-run changes nothing. Host tuning for a headless Mac (#150) follows the same rule: a setting with a command-line tool (Power Nap in the power step’s pmset -a, mdutil -i off on a workspace volume) is a fix shown and confirmed like any other, a setting without one (Apple Intelligence, on-device analysis) is guided and checked, and nothing the wizard runs deletes Apple’s files or stops Apple’s processes (a mediaanalysisd cache cleared or the daemon killed is a stopgap the OS undoes, so it stays manual-only text). Built in Go on Apple’s command-line tools, not on Swift frameworks or configuration profiles | Nearly every host step is a command-line tool (sysadminctl, pmset, socketfilterfw, fdesetup, launchctl, brew, container), so the manual’s checklist can be run and rerun rather than typed, and the reference Mac mini (#73) becomes reproducible. Apple’s frameworks add nothing here: OpenDirectory duplicates sysadminctl, SMAppService needs a signed app bundle, privileged helpers need signing, and configuration profiles need the human’s approval in System Settings or an MDM; TCC permissions cannot be scripted at all (all unverified on macOS 26). A wizard that ran as root, or as whr, would let the supervisor’s account change the host; per-command sudo keeps every privileged action visible and confirmed. The web wizard of §9.5 stays medium term and will drive the same steps |
| D47 | Each repository runs under a workflow preset: prototype, integration (the default) or published (§6, issue #105). A preset is a named setting of the policy table, where approved commits go, the PR rule, the agent permission mode and the ruleset whr doctor expects. prototype fast-forwards an integration branch to the approved SHA with no PR; integration opens a PR into an integration branch that the human merges and promotes; published opens a PR into the default branch, runs the agent in manual mode and asks for egress again when its source changes. The floor is the same in every preset: a per-SHA approval before anything leaves, no merge, tag, release or deploy by an agent, and secrets, isolation and egress unchanged. The preset is set in the supervisor’s configuration, never by the repository; changing it is a policy change, audited and, once D45 exists, confirmed with a passkey | One flow for every repository makes a prototype pay for a PR and a review per change, while a free per-action configuration makes unsafe combinations easy and audits hard. Presets give the speed where the stakes are low and keep every guarantee that matters, because the floor is not part of a preset. The fast-forward of prototype is the human’s action, carried out by the supervisor after approving the exact SHA; it is not an agent merge (D18). A preset in a repository file was rejected: the agent can edit the repository, so it could choose its own preset (D38). Whether a ruleset can name the App as the only writer of a branch besides the human is unverified |
| D48 | In api-key mode the key lives in a model-API gateway in the environment’s sidecar; the agent holds only a per-run token (§7.3, issue #83; medium term, D41). A small whr gateway runs next to the egress proxy in the sidecar, which runs no agent code. The agent’s CLI is pointed at it by the vendor’s base-URL setting over the environment’s internal network; the gateway swaps the per-run token for the real key and forwards over TLS to a vendor host and endpoint list fixed in the code, with no redirects and a size cap. There is no TLS interception and no CA in the guest. The per-run token (32 random bytes) is made at run start, reaches the gateway over stdin, is registered with the redactor and revoked when the run ends, when the sidecar stops and at supervisor start. The agent’s allowlist drops the vendor host, so the gateway is the only path to the model API. The gateway counts the vendor’s token usage per run; budgets.per_day joins the per-run and per-task budgets in api-key mode, and the supervisor can stop a token at once. A CLI that cannot take a base URL is refused in api-key mode unless the human opts in for that agent, an audited policy change, and whr doctor then says the key is in the environment | The key leaving the agent’s environment is the point of D41’s medium term: today a compromised agent can copy the key and spend after the run. A listener on the host is ruled out because guests reach host addresses but not loopback (D29, #69); the sidecar already holds egress and is separate from the agent’s VM. Residual risks: spend up to the budgets while the run lives, prompt content reaching the vendor (as today), and a compromised sidecar leaking the key. Which CLIs accept a base URL, and whether they need the vendor host for anything else, is unverified until the spike that gates the build (#109) |
| D49 | workharbor’s account is a recommendation, not a requirement: any account but root, and the doctor warns instead of refusing (§7, manual host setup, issues #154 and #157; Werner’s decision of 3 Oct 2026). The supervisor may run as a dedicated standard user (recommended), a dedicated administrator, or the developer’s own account on a dual-use Mac; root stays refused everywhere. whr doctor gets a warn status (measured, working, weaker than recommended; it does not fail the exit code). Its account check is ok for a dedicated standard account, warn for an administrator or the developer’s own account (the configuration’s `account: dedicated | shared) without remote access, and failfor the same with remote access configured: a web UI forwarder,whr ssh and the console over the VPN, or the host's own Remote Login (sshd) or Screen Sharing; the supervisor runs either way, and whr setupasks one explicit confirmation that names the risk. An administrator account may runwhr setup hostandwhr setupitself, one after the other; the last, optional stepdrop-adminremoves it from theadmin group (dseditgroup, shown, confirmed, then sudo -k) and tells the human to log out or restart, since sessions and the LaunchAgent started before keep their group membership until then (unverified); it refuses while no other administrator exists, and is never offered for a sharedaccount. What a weaker account gives up is named, not hidden: on asharedaccount every process the developer runs (an editor extension, a build script) can read the API token and use the socket, so D29's "thewhruser only" and D43's "only the admin logs in to the host" do not hold there; and an administrator orsharedaccount can write the Homebrew prefix and replace thewhrbinary, so D24's admin-owned prefix protects only a standard account. The doctor'swarnandfail` messages say so |
| D51 | Publishing runs in production as one path: prepare when a run stops, the check in the workspace’s environment, the App as committer with a signing key of its own, a token per push through a pipe, and a publish the reconciler completes (§4.2, §4.5, §7.3, §7.7, issue #244; built in #251, #252 and #253). When a task’s latest run ends stopped because its agent finished, the service prepares the topic and raises “Ready to push?” for the exact SHA; a refused prepare raises the blocking question prepare_failed. The check command is the repository’s check in the supervisor’s configuration, else customizations.workharbor.check at the default branch’s commit (D38), else pre-commit run --all-files when the default branch has .pre-commit-config.yaml; without one the prepare fails closed. It runs in the task’s workspace environment as the agent’s user, in a check worktree detached at the prepared SHA, with no credential, under a timeout and with its output capped as untrusted. The committer is the GitHub App’s bot identity, read from the App (D15); commits are signed with an SSH key in a 0600 file named by bot_signing_key_file, and a missing key fails the prepare, never an unsigned commit. Messages are linted on the host by a built-in linter the configuration names (commit_lint); a repository’s own linter is repository code and runs only inside its check. Each push gets an installation token for the one repository with contents: write, minted by the forge adapter’s Pusher, revoked when the push returns, and handed to git only through a credential helper fed by a pipe, never argv, environment, URL, file or log; only the Guard calls the Pusher. Answering a review Decision is one service call for the JSON API, whr approve and the web UI (after a passkey step-up, D45): the answer is recorded first and an allow starts the publish; an allowed, current, unpushed SHA is an outstanding publish that the reconciler’s pass completes, every forge step idempotent for that SHA, and a refusal that a retry cannot change ends it with the blocking question publish_failed. From the run’s stop until the prepare has pinned a revision or been refused, the environment and the agent are busy, and the check runs until its process group is gone | The parts of #27 existed and nothing joined them, so #28 could not open a PR. The check runs where the agent’s own code already runs, so it adds no trust domain and needs no second environment per check; it is a quality gate, not a security control, because the agent writes the tree it checks, and the per-SHA approval, the forge’s CI and its ruleset stay the controls. The App is already the bot identity (D15). A token per push with only contents: write and revoked at once is the narrowest credential the forge offers (§7.3), and a pipe leaves no copy on disk, in a process list or in a transcript, stricter than the 0600 file the Hard rules allow (D48 hands the API key over the same way). Recording the answer before the push and repeating only idempotent steps makes a restart lose nothing and open no second PR; completing it needs no new approval, because the human allowed exactly that SHA and the Guard checks it again. That the pipe reaches the helper through git’s own subprocesses up to an https push to the forge, and whether GitHub marks commits signed with a key no account holds as verified, are unverified until #28’s live run (#252 measured the pipe through git’s own subprocesses over loopback http) |