Skip to content

Interfaces

9. Interfaces

9.1 CLI grammar

Noun-verb with short aliases. Nouns: task, ws, run, env, inbox. whr task create is the canonical long form; common verbs are hoisted.

whr serve                                  # the supervisor: JSON API, web UI, reconciler
whr login --server <url>
whr run <issue-url> [--agent <ws>/<role>] # create + start on a named agent (D42); the core demo
whr ls [--json]
whr show <task>                            # review card: diff stat, CI for current SHA, open decisions, usage line (tests and agent notes: not available yet)
whr logs <task> -f
whr watch                                  # live event stream
whr say <task> "msg"                       # or -f guidance.md, or - for stdin
whr pause|resume|cancel <task>
whr purge <task> [--yes]                   # delete the transcript after saying what goes; keeps audit entries, usage and Decisions (age and size limits come later)
whr usage [--task <task>] [--since <time>]       # tokens and cost per run, task, repo and period (§5.7)
whr inbox [--watch]
whr approve|reject <decision>
whr answer <decision> <option> [--text "..."]  # a question's fixed option, or a free answer (D23)
whr diff <task>
whr ws add <path> --from <repo-url|path>   # a workspace with its agent clone (D42)
whr agent add <ws> <role>                  # a named agent with its worktree agent/<role>
whr ws rebuild <ws>                       # recreate the environment from the image the repository resolves to now (#128)
whr console [<ws>] [--write] [--status] [--close]   # a shell in the console environment (D43)
whr ssh [--forward] [-- <command>]         # SSH into the console with a certificate that lasts minutes (#32)
whr ssh --config [--forward]               # emit a ~/.ssh/config block (ProxyCommand through whr) for VS Code/JetBrains
whr ssh --takeover <ws>                    # later: pause the agent and take the lock
whr open <ws> --editor vscode
whr wait <task> --for decision|done
whr kill-all
whr doctor                                 # server, auth, runtime capabilities, SSH config

Avoid review as a verb (ambiguous). Task refs accept a short ID, a prefix or repo#42. Stable (D37): serve, run, ls, logs, say, cancel, inbox, approve, reject, answer. The other names are provisional until they are built.

9.2 Scripting contract

  • Exit codes: 0 ok · 1 error · 2 usage · 3 not found · 4 auth · 5 conflict/wrong state · 6 needs human input (whr wait) · 7 timeout · 10 task failed.
  • --json returns a stable envelope with schema_version; --jsonl for streams.
  • Streams via SSE, resumable with --since <event-id>.
  • Idempotency-Key on mutations; --dry-run for destructive actions.
  • Stdout is data, stderr is human text; honour NO_COLOR and TTY detection.
  • Shell completion generated by the CLI framework (cobra, D14), with dynamic task/workspace ID completion.

9.3 Web UI

v0 shows task/issue, repo, branch, PR, recent actions, test results, pending Decisions and resource use, with a single summary card per task. Sections “Harbor” (overview) and “Inbox”.

Because the app is a remote for coding agents (§1), v0 also carries the core remote-control loop:

  • Live transcript. A structured, streamed view of the agent session (messages, tool calls, diffs, test results) over SSE. Not a raw terminal mirror, which reads badly on a phone.
  • Send a message to the running agent (mid-run instruction injection, §5.2). Delivery is reported honestly: injected now, delivered at the next turn, or, for an agent in degraded mode, as a resumed turn.
  • Start a task from an issue or a repo, choosing the agent.
  • Pause, resume and cancel a run.
  • Answer Decisions, as before.
  • History controls. Clear view hides older events and loads them on request, without deleting anything. Purge transcript deletes the stored transcript content after a confirmation that states what goes (event count and size), what stays (the audit entries and a record of the purge) and that it cannot be undone. A chat is paged and virtualized, so a very long one stays usable on a phone.

Usage on the dashboard (issue #111). The Harbor page summarises what every agent used over a period: total cost, API and wall time, code changes in approved commits and the number of runs, then the same broken down by agent (<workspace>/<role>) and by model (input, output, cache-read and cache-write tokens, cost, cache share). Every cost is labelled reported or estimated; in subscription mode it is API-equivalent and not billed, and the account’s usage window leads (§5.7, D40); in api-key mode it is real spend, shown against the daily budget (D48). whr usage --by agent|model|day gives the same numbers from the same service method.

Usage on the dashboard as built (issue #111; internal/service/usage.go, internal/store/usage.go, internal/web/usage.go, GET /v1/usage/summary; unverified on a real run: the totals are tested against recorded and synthetic usage events). Service.UsageSummary(period) is the one method behind the Harbor card and whr usage --by agent|model|day. A period is today, 7 days, 30 days or all, and a day begins at the supervisor’s midnight in its time zone (service.Config.Location, the host’s by default), not UTC’s: a turn belongs to the day it happened on, so a run that spans midnight is split between two days and counts as a run on both. The card shows, for the period, one total row per auth mode (turns, runs, API time, wall time, cost; a subscription’s notional cost is never added to an API key’s spend), then a breakdown by agent (<workspace>/<role>; a row links to that agent’s tasks) and by model, with input, output, cache-read and cache-write tokens, the cache share (cache read over all input tokens) and the cost, each sortable by a column link. Time is what the agent reported on its result event (api_ms, wall_ms). Every cost is labelled: reported (the agent’s own figure), estimated (kept in its own column and never added to a reported one; nothing records an estimate yet, #48), mixed, or none; on a subscription the usage window (five-hour, seven-day) leads the card and the cost reads “API-equivalent, not billed”; with an API key it reads “spend”. The card reads only usage audit entries, so a transcript purge changes nothing. Code changes: Prepare measures the diff stat of a revision against what it was rebased onto (git diff --numstat: files, lines added and removed; a binary file counts as a file and no lines), the review candidate keeps it, “Ready to push?” shows it in its subject, and approving writes a review.approved audit entry with the commit, who approved it and that size; the card sums those entries over the period (a denial adds nothing). Not built yet: the daily budget of D48 against API-key spend (the budget is not built, #83).

Writes in v0 are therefore: answer Decisions, send messages, start tasks, pause/resume/cancel, and purge a transcript. Editor launch and takeover (whr ssh --takeover) come after v0.

Previews (D33). A web app the agent runs is previewed through whr’s preview proxy on its own origin, never the UI’s, opened on request and closed with its environment (issue #72).

Stack (D8). templ templates rendered by the Go server, htmx for partial updates and form posts, and the htmx SSE extension for live event and inbox updates. htmx is vendored and version-pinned; styling is plain CSS with design tokens shared with the documentation site (navy and teal, light and dark). Pages are semantic HTML first, so they work without JavaScript for reading. Handlers stay thin: they call the same service layer as the JSON API. Diffs are server-rendered (or use a small library such as diff2html); an interactive terminal (xterm.js) is out of scope for v0.

Web UI v0 as built (issue #30, internal/web). The UI is served on the loopback listen address, behind the forwarder (D29), by whr serve, and that listener has no /v1 route: the JSON API is served only on a unix socket (§9.7). Pages are templ templates compiled into the binary; htmx and its SSE extension are vendored, version-pinned and served from /static with the stylesheet. Every page works as plain HTML: each action is a form with a POST and a redirect, and htmx (hx-boost) only makes it a partial update. The screens: Usage (issue #48). The harbor page leads with a Usage panel (internal/web/usage.go): the account’s usage windows as meters (five hour, seven day, with the reset time), because on a subscription they are the figure that matters (D40); below them the totals of the last 24 hours (turns, tokens in and out, the agent’s reported cost, called notional on a subscription) and the balance the agent last reported. The panel is left out when there is no usage and when it cannot be read (the failure is reported, the page still shows the tasks). A task page shows its own usage line, the one service.FormatUsageLine writes for the CLI. Both read Backend.Usage, so they show what whr usage shows.

  • Sign in (/login). The credential is the static API token until a passkey is enrolled (D45, below); a correct token starts a session (a random ID kept in memory, so a restart signs everyone out) and the cookie is HttpOnly, SameSite=Strict, Secure behind the HTTPS forwarder (a request over TLS or by public_url’s host name; X-Forwarded-Proto is never believed), with a path of /. Failed attempts are slowed down. Everything but /login and /static needs the session, SSE included; a request without one is redirected to /login. Authentication sits behind one small interface (web.Auth: the session of a request, sign in, sign out) and a passkey sign-in starts the same session without touching a handler.
  • Harbor (/). Counts, the tasks that need you first (a task awaiting guidance), then the rest, and the form to start a task: an issue URL and the agent to run it, chosen from the workspaces’ agents (POST /tasks, the service’s Run; an untrusted author’s issue is held and says so, §7.1).
  • Inbox (/inbox). Every open Decision with its subject and input as untrusted text. A tool approval has Allow and Deny with an optional reason; a question has its fixed options; an “Accept this task?” hold has Start and Cancel. A “Ready to push?” review and an egress host have no plain button: they are answered with a passkey (see The board’s agent queue as built (D30, D40, issue #71; internal/service/queue.go, forge.QueueReader, board.queue_status; unverified against a real board: tested against a fake GraphQL API). Setting board.queue_status to a Status option (for example “Agent queue”) makes the supervisor read the project board every minute, by polling, since a webhook needs an address the forwarder does not publish (D29). A card in that column never starts a run: it raises a blocking “Accept this task?” Decision (cause board_queue, options start and cancel) on a new queued task with no run, showing the issue’s author and association and its text as untrusted data and a hash of that text, so a start refuses text that changed after the question; whoever moved the card is shown as unknown, because a project item names no actor, and is therefore treated as untrusted. Only the supervisor’s one human answers it, in the web UI or with whr answer; start starts one run now, which waits only for host resources, and cancel cancels the task, so a card nobody accepts never runs and the column is no unattended backlog. A non-trusted issue author marks the task as having untrusted input, as in #53. The agent is the card’s Session field when it names an agent of the issue’s repository, else the repository’s only agent, else the card is reported once and nothing is asked. The supervisor remembers each card’s last-seen update time (board_queue): a card asks once per state, so declining it does not ask again until it is moved again, and an issue that already has an unfinished task is not queued twice. Cards of repositories the supervisor does not work on, drafts and pull requests are ignored.

Passkeys as built), and with none enrolled the page says how to enrol one on the host or to answer with whr approve. The form POST refuses them whatever the session, so there is no weaker path.

  • Task (/tasks/{id}). The state, repository, issue, branch and the run, the task’s open Decisions, the live transcript (the recent events, then new ones over SSE with Last-Event-ID, so a dropped connection resumes without a gap or a repeat), a box to send a message to the running agent with the delivery reported as the service reports it (injected, next turn, or resumed turn), Make an editor copy, which shows the supervisor’s own copy of the agent’s branch (OpenCopy, never the agent’s checkout) and lists the files in it that an editor may run by itself, and Cancel behind a confirmation page.

Rules every handler follows (security rules of the UI, decided with issue #30; a change to them is a design-owner change like §6 and §7). A handler parses the request, calls the same service method as the JSON API and renders; it holds no business logic. Every POST carries the session’s CSRF token (a hidden field, compared in constant time) and is refused when Origin or Sec-Fetch-Site says another site sent it. Every write carries an idempotency key made when the form was rendered, kept through the store like the API’s, so a double tap or a retry answers once and the repeat shows the same result. Text from an issue, a transcript, a tool input or an error is escaped by the template and never inserted as HTML; no template uses a raw-HTML escape hatch. The page sets Content-Security-Policy: default-src 'none'; script-src 'self'; style-src 'self'; img-src 'self'; connect-src 'self'; form-action 'self'; base-uri 'none'; frame-ancestors 'none', so there is no inline script or style, plus X-Content-Type-Options, Referrer-Policy: no-referrer and Cache-Control: no-store. A phone gets one column with controls in thumb reach and targets of at least 44 pixels, and a wide screen gets the task list beside the task; nothing depends on hover. Until the passkey step-up of D45 exists, the UI refuses to answer a review Decision (“Ready to push?”) on the server side, not only by hiding the button: pushing stays with the CLI on the host. The session cookie is HttpOnly, SameSite=Strict, Secure behind the HTTPS forwarder, with no Domain attribute and, over HTTPS, the __Host- prefix (__Host-whr_session: a browser takes it only when it is Secure, has no Domain and Path=/, and a cookie another port of the same host name sets under the plain name, such as a preview’s, is not read as the session; plain loopback HTTP keeps whr_session, since __Host- needs Secure; the CSRF token is held server-side with the session and is no cookie), and lives at most 12 hours. The one script of the UI’s own beyond the vendored htmx (app.js, which keeps Last-Event-ID across reconnects) is served from the same origin under the same CSP; it adds no client-side state or routing, so it stays within D8.

Passkeys as built (D45, issue #101; internal/passkey, internal/web/passkeys.go, internal/web/static/passkey.js; provisional unverified until a real phone has done it). WebAuthn through the go-webauthn library, kept behind web.Passkeys and api.Passkeys. They need public_url in the configuration (board.public_url is used when it is absent): the relying-party ID is its host name and the origin is its origin, so a passkey made on one name does not work on another; without it passkeys are off and the token signs in. Credentials are resident (discoverable) and user verification is required; whr keeps the public key, a sign counter and the backup flags, never a secret. Enrolment and revocation are only from the host: whr passkey add [name] prints a one-time link (single use, 5 minutes, only its hash is kept) that the human opens on the phone, whr passkey ls and whr passkey rm <id> list and revoke, through POST /v1/passkeys/enrolments, GET /v1/passkeys and DELETE /v1/passkeys/{id} with the API token; the web UI has no route that adds or removes one. Once one is enrolled the token no longer signs in to the web UI (no fallback; the token stays for the CLI), and revoking the last one turns it back on. Sign-in refuses an assertion without user verification, and one whose counter did not move forward or whose clone-warning is set. A step-up answers a sensitive Decision (today a review and an egress host): the challenge is a hash of the Decision ID, the commit SHA (or host:<name>) and a random nonce, held server-side for 2 minutes, usable once and only by the session that began it; the answer is recorded only when the assertion names exactly the Decision and the SHA or host it has now, and a review answer carries the SHA so the service checks it again. Policy changes and secret operations on the web (issue #107; internal/store/change.go, internal/web/changes.go, /changes). The same step-up, with a change in place of a Decision. A workflow change (D47) is still made in the configuration, never in the web UI. When whr serve starts and a repository’s preset or integration branch there differs from what was recorded, it refuses to start, as before, unless the host passes --accept-workflow-change; with a passkey enrolled it instead starts, keeps that repository on the recorded workflow (the held one, which tasks started meanwhile record) and raises a pending change, which /changes shows as from-and-to. Confirming it takes a step-up whose challenge names the change’s ID and its text (workflow <repo> <from> on <branch> to <to> on <branch>); that records the change in the audit log of workflow changes as confirmed by web+passkey, and it takes effect when whr serve is restarted, so a running supervisor never changes its own rules under a task. A pending change is withdrawn when the configuration no longer asks for it, and it is refused as stale when the recorded workflow is no longer the one it was raised against. There is no way to reject one on the web: the host edits the configuration back. A secret operation is one that touches the credentials whr holds: today the revocation of the forge tokens (the token half of whr kill-all, §7.7), a button on /changes that needs a step-up naming the operation and writes a supervisor.tokens_revoked audit entry (who, how many). No secret value is ever shown, entered or stored through the web UI; rotating or setting one stays a file the host writes (0600). Both need a passkey: without one the web offers nothing and says to use the host. The host path stays (--accept-workflow-change, whr kill-all), and nothing here lets a web session alone do either. Sessions follow their passkey: a web session remembers the passkey it was started with; enrolling the first passkey ends every session the API token started, and whr passkey rm ends the sessions of the revoked passkey and their live streams (web.TokenAuth.EndSessions, wired in whr serve). If the store cannot say whether a passkey is enrolled, the token does not sign in (fail closed). Limits: the challenges held in memory are in two pools: sign-in ceremonies (256), where a full pool evicts its oldest and never refuses a new begin, so a flood cannot keep the human out, and step-up and enrolment ceremonies (64), which refuse when full and which no sign-in can fill; beginning a sign-in is guarded for the whole supervisor at 600 a minute (a CPU guard, counted for the whole supervisor because the clients arrive through a forwarder that hides their address); a refused assertion is not counted and there is no lockout, since an assertion cannot be guessed and a lockout would only be a lever for anyone who reaches the forwarder; and the web answers 429 then. No TOTP and no password.

Previews as built (D33, issue #72; internal/preview, internal/service/preview.go, whr preview, /v1/previews; unverified live: it is tested against a fake runtime, and whether the sidecar can relay inbound traffic is issue #69). The configuration’s preview.first_port and preview.last_port (at most 50 ports, not the UI’s) turn it on. The human opens a preview of a port that the repository’s environment declares on its default branch (forwardPorts, customizations.workharbor.previewPorts) while that environment runs: from the task page, or with whr preview open <task> <port>. The agent cannot ask for one. Each preview listens on its own loopback port, which the forwarder maps (tailscale serve --https=<port>), so it has an origin of its own; the link is https://<public_url host>:<port>/?whr_preview=<grant>. The grant works once for 5 minutes and is traded for a per-preview HttpOnly, SameSite=Strict cookie (named with the __Host- prefix over https, so a browser takes only one that is Secure, has no Domain and Path=/; a cookie the server did not issue opens nothing); every other request needs that cookie. The proxy forwards to the one environment and port the preview was opened for, through runtime.Previewer.DialPreview, whatever a request names (absolute URIs, other ports and other hosts are ignored); it answers 421 to a Host that is not the forwarder’s name or loopback, rewrites Host to localhost:<port>, passes WebSocket upgrades, and never passes whr_-prefixed cookies to the app nor lets the app set them. Because a port is reused by the next preview, a request that registers a service worker is refused, the grant exchange sends Clear-Site-Data: "storage", "cache" (never "cookies"), and the app’s own Strict-Transport-Security and Clear-Site-Data are dropped. Every proxied response carries the proxy’s own Content-Security-Policy (default-src 'self'; script-src 'self' 'unsafe-inline' 'unsafe-eval'; style-src 'self' 'unsafe-inline'; connect-src 'self'; img-src 'self' data:; form-action 'self'; base-uri 'none'; frame-ancestors 'none': the app can load and call only itself, and inline script and eval stay allowed for dev servers) in place of the app’s; leaving by navigation is accepted in the threat model. A preview opened from the web ends with the sessions that opened it (sign-out, revoke, sweep, and expiry or idling, which a sweep once a minute ends even if the cookie never comes back); one opened from the CLI or API does not. Whether a request is HTTPS is decided from public_url (its host name gives Secure and __Host- cookies) and never from X-Forwarded-Proto. A preview ends when its environment stops (checked on every request and every 10 seconds), when it is closed, after 12 hours and when the supervisor restarts; opening and closing are audit entries (supervisor.preview_opened, supervisor.preview_closed) without the link. A browser does not keep cookies apart by port, so an agent-written preview can overwrite whr’s session cookie and sign the human out of the UI; the UI’s cross-origin POST checks still hold. A runtime without Previewer (Apple Container today) answers that previews are unavailable.

The installed app as built (issue #33; internal/web/static/{manifest.webmanifest,sw.js,pwa.js,offline.html,icons/}, internal/web/pwa_test.go; unverified on a real phone and tablet: only the files and the rules below are tested). /manifest.webmanifest, /sw.js and /offline are served without a session, because the browser fetches them without cookies, and hold nothing private; the CSP gains manifest-src 'self'. Installing needs HTTPS, so it works over the forwarder’s name (D29) and on localhost. The service worker caches the app shell only: the files of /static/ and the offline page, named in one list in sw.js. A page, a transcript, a diff, a Decision, an event stream or any POST is never cached and never handled by the worker; a page fetched while offline gets the offline page, which says that nothing is kept on the device. A test reads the list and fails if a path outside /static/ and /offline is cached or the worker gains a second place that writes the cache. Devices: each browser that signs in (token or passkey) has its own session, and /devices lists them with a label made of fixed words from the User-Agent (“Phone, Safari”), when each signed in and was last used; any device can end another’s session, or its own, so a lost phone is signed out from the tablet. A phone’s session ends after 15 minutes without a request, a tablet’s or a computer’s after 8 hours, and every session after 12 hours whatever its use. This is the rule (D35, D45): 15 minutes idle on a phone, 8 hours elsewhere, and a fresh passkey step-up for anything sensitive whatever the age of the session. A live stream ends with its session: sign-out, a revoke from another device and expiry cancel it at once, and every heartbeat checks the session again without counting as use, so an open stream neither outlives a revoked session nor keeps an idle one alive (web.Streams). These are sessions, not tokens: the API token for the host’s CLI stays one secret. The phone class is guessed from the User-Agent, which is the client’s own claim, so it only sets a convenience, never a right.

Pause, resume and purge as built (issue #106; internal/service/pause.go; provisional unverified against a real agent, tested with the fake agent). Pause (Service.Pause, POST /tasks/{task}/pause, whr pause, the task page’s button) is the hard interrupt of D11: the aggregate moves the run to paused, supersedes the questions and approvals it raised (D23, §4.2) and frees the task from guidance it waited on, then the session is stopped through the runtime, so whr-shim cancels the process group (D25). The run is saved as paused first, so the session’s end is not taken for a loss; the environment keeps running (§4.3). Only a running run pauses. Resume (Service.Resume, POST /tasks/{task}/resume, whr resume) relaunches the agent from its session with the briefing of D27, which names the superseded requests, so the agent asks them again. It works on a paused or an interrupted run, is refused while a login or quota question of the run is open (answer it, or cancel) and while the run is already running, and a forgotten session or used-up attempts fail the run and open the retry-or-cancel question, as an answer that resumes does. Purge (Service.PurgeTranscript, GET /tasks/{task}/transcript for the size, POST /tasks/{task}/purge, whr purge, the task page’s “Delete the transcript”) deletes all transcript-tier events of the task through Store.Purge (§5.4) and writes one audit entry (store.purged: who, events, bytes and a digest of what went; who is web:<device ID> for a web session, the ID /devices lists, and api for the host’s token). The audit entries and their chain, usage rows and Decisions stay, and so does the agent’s own session. The confirmation says how many events and bytes go and what stays: the CLI asks you to type purge (or takes --yes), the web page shows it, and the API needs {"confirm": true}. A purge is refused while a run of the task is running or starting, because the transcript is still being written: pause it first. Age and size limits (§5.4) are not built.

The two layouts and the keyboard as built (issue #30; internal/web/static/{app.css,app.js}, internal/web/layout_test.go; unverified in a browser and on a real phone and tablet: only the markup and the stylesheet are tested over HTTP, and the shortcuts’ behaviour is not run). The task page carries the task list in an aside.tasklist next to the task in one server-rendered page; the stylesheet hides it by default (the phone’s single column) and shows it as a left pane in @media (min-width: 56rem) and (orientation: landscape) (the tablet’s two panes, T2), so switching tasks keeps the page. Touch targets are at least 44 points (inputs, buttons, tabs and cards, tested), and no :hover rule shows or hides anything. Shortcuts (app.js, only when no field has focus): g then h, i, c or d go to the harbor, inbox, changes or devices; j and k move through the task list; m or / focus the message box; Ctrl or ⌘ with Enter sends the message being typed; ? shows the list that every signed-in page carries in a details element; Escape drops the focus. They only move focus and follow links: no shortcut answers a Decision, and the script does no network request or HTML insertion (tested). The phone and tablet walkthroughs wait for the live demo (#28, #73).

9.4 Notifications

Push when a blocking Decision stops a task: the value of a supervisor is not having to watch it. This includes auth_expired and quota_exhausted (§5.2), which are the most likely reasons a detached run stalls.

Default channel: ntfy (phone and desktop apps; self-hosted or ntfy.sh). Generic webhook and macOS notification are secondary channels; Pushover, Telegram and Web Push are later options. ntfy’s iOS delivery through a self-hosted server is unverified.

  • Events: a new blocking Decision (question, approval, review), auth_expired, quota_exhausted, the notice agent_may_run (§4.2, issue #238), and a run that ended or failed. Deduplicate and rate-limit per task so a stalled run does not notify repeatedly: one notify.Throttle shared by every kind sits in front of the channel (whr serve, issue #177), drops a message equal to one sent within the hour (same task, kind, Decision and run), and lets a task send at most five an hour. The cap counts every kind, so a blocking Decision can be dropped after five earlier pushes for the same task; the inbox still shows it.
  • Generic payload. Task ID, event kind and a link only. Never issue text, code, logs, transcripts or tokens: the message leaves the host, may pass a public relay, and issue text is untrusted input.
  • Link, not action. The notification opens the task in the web UI behind the supervisor login. No approve or answer buttons in the push. Approvals stay per commit SHA (§6).
  • Reachability. The web UI listens on loopback, behind the forwarder (D29), so the link uses the forwarder’s name and the phone needs the VPN to open it.
  • Topic protection. A long random topic, or an access token on a self-hosted server. Keep the topic and token in the credential service, never in the repo, the configuration file or logs. In release 1 that is a secret file each (ntfy.topic_file, ntfy.token_file: 0600, owned by whr, one link, outside every root, read with config.ReadSecret), as for the GitHub App key and the SSH authority (§5); whr serve refuses a topic under 20 characters (issue #177).
  • Best effort. The inbox stays the source of truth; a push can arrive late or be lost.

9.5 Onboarding

First run is a guided sequence of six steps. The steps are the contract; the surface differs by phase. Release 1 delivers them through whr login, whr doctor and the setup wizard of D46, because the v0 web UI is scoped to remote control of running tasks (§9.3). A web wizard over the same steps is a medium-term item (§13). Each step can be skipped and re-run later.

  1. Sign in. Server URL (reached over the VPN, never public) and the single static access token, stored encrypted. OAuth sign-in comes later (§10).
  2. Connect the forge. GitHub in release 1 (D15): install the workharbor GitHub App on the chosen repositories; Gitea, Forgejo and GitLab later. Verify the limits the forge enforces, not prompts (§6): the bot can push agent/* branches and open PRs, branch protection requires a human review, the bot cannot bypass it, and merge, tag, release and deploy stay forbidden.
  3. Choose the agent login. subscription (the human signs in inside each environment through the vendor’s own flow; whr never sees the credential, D40; how is spike #82) or api-key (kept on the supervisor’s side), per §5.2. The subscription option states the accepted risk of §7.3 and links the vendor terms page of the manual.
  4. Check the host. The checks of whr doctor: server and token, container runtime, forbidden mounts rejected, default-deny egress, agent session surviving a reboot, capacity (plan for 4 concurrent environments, §8). A check that has not been verified is reported as not verified, never as passed (spike #2 measured Apple Container isolation and egress; reboot survival is still unverified, §12).
  5. Set up phone notifications. ntfy provider (self-hosted or ntfy.sh), a random topic stored in the credential service, and a test push that carries the generic payload of §9.4. Remind that the link needs the VPN. In release 1 the human generates the topic into its secret file and adds the ntfy block to the configuration (host setup, step 10); neither whr setup nor whr doctor generates it or sends a test push yet.
  6. Ready. Summary of what was configured and what is not yet verified, then the first command: whr run <issue-url>.

The setup wizard (D46, issue #104). A setup step is a doctor check plus an optional fix, one value, so whr doctor and whr setup cannot disagree. A fix is a list of commands (argv, never a shell string) shown before it runs, or a guided text with the System Settings pane or link to open; after a fix the check runs again. whr setup host covers the administrator’s part of the host setup and whr setup the whr user’s, in its desktop session; --dry-run prints every command, --only and --from pick steps, and a step whose check passes does nothing. The command set and their behaviour on macOS 26 are unverified until the wizard has set up the reference Mac mini (#73).

Creating the App (step 2). whr github app create (the name is provisional) makes the operator’s own GitHub App through GitHub’s manifest flow, so nobody fills in the creation form or downloads a .pem; the manual steps stay as the fallback (host setup, step 11). Every operator needs their own private App: the key mints tokens for every installation, so one cannot be shared.

  • The manifest asks for a private App with no active webhook, no events and exactly Contents and Issues and Pull requests write plus Metadata read (github.AppPermissions, the same set an installation token is minted with). --org targets an organization’s endpoint instead of the operator’s account.
  • The command listens on the configuration’s loopback listen address for the redirect, which the forwarder of D29 sends to whr’s HTTPS name (--public-url), and so whr serve must be stopped meanwhile: the supervisor needs the App, so it cannot be up before the App exists. It prints a link on that name; the page behind it posts the manifest to GitHub when the human presses a button, so it works from the phone.
  • The setup server serves two pages and nothing else (it has no API and no token), and both need a state that this command minted: unguessable, valid for ten minutes, compared in constant time against the live ones, and spent by the redirect, so a forged, expired or replayed value gets a 400 and leaves no file and no request to GitHub. The redirect’s code is exchanged with POST /app-manifests/{code}/conversions, which needs no credential.
  • The private key goes straight to a 0600 file github-app-<id>.pem in the configuration’s directory (exclusive create, no symbolic link followed, never overwriting) and is registered with the redactor, as are the client secret and webhook secret, which are then dropped. Nothing secret is printed, logged or shown on the page. The command prints the two lines github.app_id and github.key_file for the configuration instead of editing the file, which stays the operator’s, and the install link; the operator installs the App on selected repositories and checks that the ruleset of main does not list it as a bypass actor (D15).
  • whr doctor (forge-app) asks GitHub (GET /app, GET /repos/{repo}/installation) whether the App is installed on every configured repository with exactly those permissions, and reports not_verified when GitHub cannot be reached. Whether GitHub accepts the manifest as sent is unverified until it runs against github.com.

whr doctor runs the checks of these steps, and (D46, issue #153) the checks of both setup phases after them, read-only: the host steps, then the user steps, in the wizard’s order, then the shared checks. It prints one tab-separated line per check on stdout: ok, fail, not_verified or skipped, then the check’s name, what it found and, as a fourth column, the setup command that fixes it (empty when it passes or no step fixes it); --json carries the same as phase and fix on each check. On stderr each failing or not verified line with a fix is repeated as → whr setup host --only power or → whr setup --only config-base; a shared check that cannot run without the configuration says → whr setup. The user phase’s checks describe the account that runs them: as another account than --user (default whr) they are not_verified with “run whr doctor as whr”, and their fix says (run as whr). doctor never fixes, asks or runs sudo; its runner only reads. warn (D49) means measured, working and weaker than recommended, as an administrator whr or a shared account: it prints, is in --json and is ignored by the exit code. Only fail is a non-zero exit; a check nobody has measured says not_verified, never ok, and --skip <check> leaves one out until the next run. The shared checks: config, account, server, forge-key, forge-app, forge-board, forge-limits, agent-login, runtime, mounts, egress, reboot, capacity, lane-agents and notifications. lane-agents is a dogfood-phase check (D34, issue #160): inside a checkout with .claude/agents/wh-*.md it reads permissions.allow of Claude Code’s settings files (.claude/settings.json, .claude/settings.local.json of the main checkout, ~/.claude/settings.json) and warns, never fails, when a lane agent has no Agent(<name>) rule (a bare Agent allows all); it never writes. In release 1 the live measurements (default-deny egress, reboot survival, capacity, the forge’s limits) are not_verified: the command does not run them yet. ok on forge-key, agent-login, runtime and notifications (configured and valid; no test push is sent) means the file or binary is in order, not that GitHub, the vendor or the container service accepted it.

The setup wizard as built (D46, issue #104; internal/setup, the steps in internal/doctor). A step is a doctor.Check with a phase (host or user), a title and an optional Fix: argument vectors (Cmd, never a shell string, with a Sudo flag so sudo is its own argv), an in-process action for a secret or a configuration, a Build that asks for what a command needs after the confirmation, or a guide with the pane or link to open. whr doctor runs every check, whr setup host and whr setup run the steps of their phase through one Host interface (commands, opening a URL, confirmation, the prompts), and a test passes a fake, so no test runs sudo or changes a system. The host steps are whr-user, power, firewall, ssh-keys-only, filevault, homebrew, brew-packages, brew-pin, prefix, autologout, workspace-volume, media-analysis and spotlight and the optional screen-sharing and tailscale (media-analysis fails when the busiest mediaanalysisd of any account is at or over 50 % of a core, and reads the cache only when run as whr; spotlight is not_verified unless mdutil -s says enabled or disabled, and for a root on the internal disk); the user steps run in this order, with no cycle: config-dir, api-token, agent-key (optional), ssh-ca (optional: the console’s SSH authority key, made in the configuration directory, never overwritten), container-start, container-kernel (its fix needs the running container system: a step names a service it needs with Needs and one its fix brings up with Provides, and a test holds the order to them; issue #265), standard-user-check, config-base (enough for whr github app create), github-app, config-github (adds the App to the file as a diff, after a y, written atomically, keeping keys it does not know), tool-store, service-install, drop-admin (D49, issue #157). Rules: the host part refuses root and a standard whr user (an administrator whr account may run it, D49), the user part refuses root, any other user and a shell outside the desktop session (the same check as whr service install), both refuse a whr that is not in an admin-owned prefix (unless --dev names a development installation the running account owns, D24, issue #261) or is in a git working tree (a dry run says so and goes on); without a terminal only --dry-run runs; sudo -v runs once after the first confirmation and is never refreshed; a root-owned file is written to a 0600 file in a 0700 temporary directory and installed with sudo install -m 0644 -o root -g wheel; a secret is generated or typed without echo into a 0600 file with an exclusive create and is never overwritten or printed; --dry-run runs the read-only checks for real and prints the fixes without running, opening or asking anything. The step names complete in the shell. The tools’ output formats (pmset -g, socketfilterfw, fdesetup status, dscl, container system status, ps -axo, du -sk, mdutil -s, df -P) were read from documentation and checked read-only on the developer’s own Mac only for some of them; none of it has set up a fresh machine, so all of it is unverified until issue #73.

Workflow presets as built (D47, issue #105). repositories[].workflow (prototype, integration by default, published; an unknown value fails the configuration check) and integration_branch are read from the supervisor’s configuration only (policy.Preset, config.Repository.Preset and Target). A prototype repository must name an integration_branch, or the configuration check refuses it, and whr doctor fails when that branch is the default branch, which the configuration cannot know. A preset’s table is checked like any override: the ceilings of §6 lower anything that would loosen the floor, which internal/policy tests for all three. Guard.WithTable decides with the stricter of each row of the configured table and the preset’s (policy.Table.Stricter; a row only one lists keeps its mode), so a preset can tighten what the supervisor was configured with but never loosen it. Publisher.Publish pushes the approved agent branch and then goes by the task’s preset: prototype calls Guard.FastForward, which needs the per-SHA approval and the table’s fast_forward_branch action (ask in prototype, forbidden in the others), requires the agent’s branch on the forge to be at the approved commit, refuses an agent branch as target, refuses the repository’s default branch (it asks the forge for it through forge.DefaultBrancher and fails closed when the adapter cannot say), and the GitHub client moves the branch with force: false, so a branch that moved is ErrNotFastForward and the topic is rebased, never forced; integration opens the pull request into the integration branch (OpenPRInto), published into the default branch. Changing a repository’s integration_branch is a policy change like changing its preset: whr serve records both (repo_workflows, migration 0016), refuses to start until --accept-workflow-change confirms a difference, and appends it to workflow_changes, with the old and new branch. Repository names are matched without regard to case there, so a name written in another case cannot slip past the check. A task records the integration branch it started with (tasks.integration_branch, migration 0015) and publishes to it, under the stricter of its own preset and the repository’s current one (policy.Stricter: published over integration over prototype), so a looser configuration never applies to a task already started and a stricter one applies at once. One review Decision still covers a topic’s series of commits and names the head SHA. The agent runs in manual mode for published and dontAsk otherwise unless agent_permission_mode overrides it. A task records the preset it started under (tasks.workflow), and that one decides its mode and its publish, so a looser preset never applies to a run already started; a change of a repository’s preset is refused at whr serve start unless --accept-workflow-change confirms it on the host, and is appended to workflow_changes (the audit of policy changes); a fresh passkey assertion replaces the host confirmation once D45 exists. whr doctor (forge-workflow) reads the rules of the branch the preset writes to (GET /repos/{repo}/rules/branches/{branch}) and the bypass list of their rulesets and compares them with the table above: a mismatch fails, and anything it cannot read (the rules, the bypass list, who else may write to a prototype’s branch) is not verified. Egress under published (Service.PendingEgress, migration 0014): an answer remembers what the host was requested by, a SHA-256 of the devcontainer.json at the commit the environment is built from (Config.SourceDigest), “lockfile” for a suggestion, or “unknown” for an answer from before this. Under published an answer holds only while that digest is the current one: otherwise it is forgotten (its row is deleted), so the host counts as unanswered for the repository, neither allowed nor denied, and is asked again; the next applyEgress then drops the stale allow from the proxy’s allowlist. Lockfile suggestions are ignored, and a repository with no devcontainer.json keeps no published allow. The digest covers the whole file, so any change to it asks again about every host: that errs toward asking, and a finer digest of just the requests is possible later. The other presets ask once per repository as before, and the preset used is the one the task started under. All of it is tested against a fake forge only; whether a ruleset can name the App as the only other writer, what permission the rules read needs, and the fast-forward under a real ruleset are unverified until #73.

Publishing settings as built (D51, issue #251). repositories[].check (the check command, one line; it wins over the repository’s own), repositories[].commit_lint (conventional, the default, or workharbor) and the top-level bot_signing_key_file are read from the supervisor’s configuration only. The key file is a secret file like the others (0600, owned by the supervisor’s user, outside every root) and whr doctor’s bot-key check warns when it is not set and fails when it is not an unencrypted SSH ed25519 key; the mode and the owner are refused by the configuration check. The check is service.RepoChecker (issue #251): the prepared commits go to the environment as a bundle on the exec’s stdin (hostgit.Repo.StreamBundle), a script there adds a check worktree under /tmp/whr-check-<sha> and runs the command in it, and the worktree is removed from inside the guest. Service.HoldEnvironment marks the environment busy (a run start is refused with environment busy); RepoChecker holds it while the check runs, and the publish path (#253) holds it from the run’s stop until the revision is pinned or the prepare is refused. After every check, a normal exit included, the checker kills the check’s process group (a process the check left in the background does not outlive it) and looks for it, and a process still there stops the environment (under the start lock, and the supervisor forgets that it started the environment, so a failed stop makes the next launch there stop and start it first). A process that starts a new session or process group (setsid, a double fork) escapes, as it does for an agent: accepted. A hold and a run start never both pass because the store has a single connection (TestTheStoreKeepsOneConnection). Check output and the guest’s text in an error have control characters, bidirectional controls and line separators removed or escaped (§7.1). That a real runtime’s cancel ends the check’s process group through whr-shim is unverified until a live run (spike #7 measured the launcher for an agent, not for this script). The check has no credential of the supervisor’s, but it runs as the agent’s user, so it can read what that user can read, a subscription login in the environment included: it is a quality gate and not a security control.

9.6 Mobile clients: phone and 12-inch tablet

The phone and a 12-inch tablet are the primary clients (D35); a laptop browser is a larger tablet. Both run the installed PWA (§13) over the VPN and its forwarder (D29). The phone is for short, urgent interactions; the tablet replaces the laptop for reviewing and longer supervision.

Phone (about 6 inches, one hand, seconds to a few minutes):

#Use caseNeedsWhen
P1A push says a task needs you; open it straight in the inboxntfy link to the Decision (§9.4), deep links, fast cold startR1
P2Answer a question: pick a fixed option or type or dictate a short answerOption buttons, a text field, idempotent submit (§9.2)R1
P3Approve or deny a tool request before its deadlineTool name, capped input, countdown to the fail-closed deadline (§4.2), one tap eachR1
P4Handle an expired login or exhausted quota“Signed in again, resume”, “Resume at reset” or “Cancel” (D23); the vendor’s sign-in link or device codeR1
P5Check the harbor at a glanceCounts by state, tasks needing you first, usage-window meter (§5.7)R1
P6Send the running agent a short instructionMessage box, delivery shown as injected, next turn or resumed turnR1
P7Stop a runaway runPause or cancel per task; kill-all behind a confirmationR1
P8Start a task from an issue seen elsewhereShare an issue link to the PWA (Web Share Target), pick the agentLater
P9Follow the live transcript for a minuteTail mode, the last events only, collapsed tool outputR1

12-inch tablet (landscape, often with a keyboard, minutes to an hour):

#Use caseNeedsWhen
T1Review a topic and answer “Ready to push?”Commit list, side-by-side diff in landscape, checks and CI for the pinned SHA (§4.5)R1
T2Supervise several tasks at onceTwo panes: task list and transcript; switch tasks without losing scrollR1
T3Plan with the agent, approve its planLong messages with a hardware keyboard, plan approval (§6)R1
T4Preview the web app the agent builds next to its transcriptThe preview on its own origin (D33) in a second window or Split ViewR1 Complete
T5Start tasks deliberatelyPick an issue, the agent, the permission mode and instructionsR1
T6Watch usage and budgets, and the planning boardUsage per task and window (§5.7), a link to the forge board (D30)R1 Complete
T7Audit and housekeepingEvent log, purge a transcript with its confirmation (§5.4)R1 Complete
T8Take over or edit by handEditor launch on the supervisor’s copy (§4.5)Later

What follows for the web UI and the PWA:

  • Two layouts from one server-rendered page (D8): a single column with actions in thumb reach on the phone; two panes in landscape on the tablet, with keyboard shortcuts and pointer hover. No separate mobile app (§13).
  • Touch first: targets of at least 44 points, no hover-only controls, readable with large text settings, light and dark.
  • Flaky networks: live views resume with Last-Event-ID (§9.2), and every answer carries an idempotency key, so a double tap or a retry never answers twice.
  • Approving a push needs a fresh check of who you are: on any device, “Ready to push?” asks for a passkey (WebAuthn) before it accepts the approval, because a phone left unlocked must not be able to publish code. Tool approvals and questions do not, as their deadlines are short and they publish nothing.
  • Per-device tokens, revocable from the other device (§13), and a short idle timeout on the phone.
  • Nothing sensitive leaves the server: notifications stay generic (§9.4); diffs and transcripts are rendered on demand and not cached for offline use by the service worker.

9.7 The JSON API

whr serve serves a JSON API (issue #96, internal/api) that the CLI, the web UI and scripts call. The contract is the OpenAPI 3.1 document internal/api/openapi.json, which GET /v1/openapi.json also serves; handlers are written by hand and a test fails when a route in the document has no handler or the other way round. Handlers are thin: they decode, call the service (internal/service) and encode, and hold no logic.

  • Binding. net/http from the standard library. The JSON API is served only on a unix socket, api.sock in state_dir (api.ListenSocket, D29, §7.5): the directory must be a real directory owned by the supervisor’s user and closed to everyone else (0700; one that is open to others, a link to a directory or one owned by another user is refused, by the configuration check and again at start, and a missing one is created 0700), a lock file in it (api.sock.lock, flock, held for as long as whr serve runs) lets only one supervisor start, so a second one cannot remove the socket the first serves, the socket is created under a umask that leaves it 0600 and checked afterwards, a socket that is already served is not taken over, a stale file is replaced, and a regular file in the way is an error. A unix socket path is at most 100 bytes here (macOS allows 104), so a longer state_dir is refused with a message. The web UI is on listen, a loopback address only (api.Listen: the server refuses anything else and checks the address it actually bound), and has no /v1 route, so a leaked API token cannot be used from the forwarded network. The CLI (cli.NewClient), whr doctor’s server check and whr setup connect to the socket and send the token to nothing else; the one thing still on listen is whr github app create’s setup server.
  • Authentication. One static bearer token from the configuration’s api_token_file (config.ReadSecret), sent as Authorization: Bearer <token>. It is compared in constant time (digests of both sides, so the length is not leaked), is required on every route including the event stream and /v1/openapi.json, and is never logged, echoed or put in an error. A missing or wrong token is 401 with exit code 4.
  • Envelope. Every response is {"schema_version": 1, "ok": true, "data": ...} or {"schema_version": 1, "ok": false, "error": {"code", "exit_code", "message"}}. The exit code is the CLI’s (§9.2) and comes from internal/exitcode; the HTTP status follows it (2 → 400, 3 → 404, 4 → 401, 5 → 409, 6 → 409, 7 → 504, 1 and 10 → 500). Golden tests pin both.
  • Routes (all under /v1): GET /tasks (?active=true), POST /tasks (start from an issue URL on an agent), GET /tasks/{task}, POST /tasks/{task}/say, POST /tasks/{task}/cancel, POST /tasks/{task}/pause and /resume, GET /tasks/{task}/transcript (what a purge would delete) and POST /tasks/{task}/purge (body {"confirm": true} or it is refused), GET /console, POST /console (open it, with read_write naming the workspaces to mount read-write; an open console with another set is a conflict), DELETE /console, GET /console/shell (a shell as a terminal stream, below), POST /console/ssh/certificate (sign a client’s public key for one SSH session) and GET /console/ssh (one SSH connection as a byte stream, below), POST /kill-all (the kill switch: body {"confirm": true} or it is refused; it stops every run, cancels every unfinished task and revokes the forge tokens, goes on after a failure and lists what failed), GET /tasks/{task}/events, GET /inbox, POST /decisions/{decision}/answer (an approval is allow or deny with the commit it was shown), GET /workspaces, POST /workspaces (create one: name, an empty folder below a workspace root, the repository, the integration branch, the source of the agent clone and the first agent’s role; it seeds the clone, provisions the environment and makes the worktree, which takes a while, and a failure takes back what it made), DELETE /workspaces/{workspace} (only without agents: it removes the environment and its home volume and leaves the folder to the human), POST /workspaces/{workspace}/agents (add a named agent), POST /workspaces/{workspace}/rebuild (replace the environment with one made from the image the repository resolves to now: refused with a conflict that names the run while one of the workspace is live; the answer is the old and new environment and image, by reference and digest), DELETE /workspaces/{workspace}/agents/{role} (only without an unfinished task; the worktree and branch stay in the clone, so no unpushed work is lost), GET /health and GET /openapi.json. Tasks carry their agent as <workspace>/<role>.
  • The console SSH routes (issue #32). POST /v1/console/ssh/certificate takes {"public_key": "<one line>", "forwarding": false} and returns the certificate (one authorized_keys line), the console’s host key (one line, to pin), the principal whr and the expiry; it needs the console open and a configured authority (console.ssh_ca_key_file), and every certificate is an audit entry (supervisor.console_ssh: who, key ID, serial, forwarding, expiry; never a key). GET /v1/console/ssh with Connection: Upgrade and Upgrade: whr-ssh/1 answers 101 and then carries the bytes of one SSH connection, unframed, to and from an sshd -i that the supervisor starts for it in the console; at most eight at once, and closing the stream ends the sshd. Both routes are on the API’s private socket, behind the token. whr ssh makes a key on the machine it runs on ($WHR_SSH_DIR, else $XDG_STATE_HOME/whr/ssh, else ~/.local/state/whr/ssh; the private key never leaves it), gets a certificate and the host key before every connection, writes them next to the key and runs ssh with -F /dev/null, only that certificate, only the pinned host key, no password and no agent, and itself as the ProxyCommand. whr ssh --config prints the same as a Host whr-console block for ~/.ssh/config, led by a Match host whr-console exec "whr ssh --refresh" line: ssh runs it before it reads the certificate, and it renews the certificate and the pinned host key only when the certificate on disk has less than four minutes left, is not for this key or lacks the forwarding permission asked for (the ProxyCommand, whr ssh --proxy, does the same, so one signing serves both and a connection never signs twice), so VS Code Remote-SSH and JetBrains need nothing renewed by hand. The block also says IdentityAgent none, quotes paths that have a space, and doubles % in the Match command; a path with a quote, a backslash or a line break is refused. On the macOS ssh measured by hand, ssh runs the Match command, then the ProxyCommand, then loads the identity and certificate files; another OpenSSH may order them otherwise, which is why the Match line is there: unverified on the editors’ own ssh until wh/verify runs them. --forward asks for port forwarding, which the certificate then carries and sshd restricts to the console’s own loopback.
  • The console shell stream. GET /v1/console/shell?workspace=&term=&cols=&rows= with Connection: Upgrade and Upgrade: whr-terminal/1 turns the connection into a terminal stream after an answer of 101 Switching Protocols; everything that can fail does so as an ordinary error before that. The stream carries frames both ways (internal/termproto): one type byte, a four-byte big-endian length and a payload of at most 64 KiB. The client sends data (D, the terminal’s input) and its size (R, two big-endian uint16, columns then rows); the server sends data (D, the terminal’s output) and, last, the exit code (X, a big-endian int32). A frame the client may not send, or a client that goes away, hangs the terminal up. The route is on the API’s private socket like the others, so it is behind the token and not reachable through the web UI or from the phone; the console serves at most eight shells at once, and each one is an audit entry (supervisor.console: who, the starting workspace, the writable workspaces). whr console puts the user’s terminal in raw mode, sends its size and every change of it, and returns the shell’s exit code as its own.
  • Events are server-sent events. Durable events carry id: <seq>, so a client reconnects with Last-Event-ID (or ?since=) and the server replays them from the store before following live ones. Ephemeral events (token deltas) and the heartbeat comments are sent from memory and have no id; they are never stored (§5.4).
  • Idempotency. A POST may carry Idempotency-Key. The first request with a key runs and its successful response is stored under the key and a hash of the request; the same key and request returns the stored response (Idempotent-Replay: true) without running again, and the same key with another request is a conflict. Requests with the same key are serialised in the process. The effect of a POST (a runtime and an agent) cannot be made part of one database transaction, so the response is stored after the effect: a crash between the two lets a retry run again. The domain rules turn such a rerun into a conflict (one live run per environment, a Decision is answered once), not a duplicate. A failed request is not stored.

9.8 The CLI as built

whr (issue #97, internal/cli, cmd/whr) is built on spf13/cobra (D14), without viper: the command tree, flag errors, help and shell completion are what cobra is for, and the standard library has none of them. Cobra’s own error and usage printing is silenced; main maps every error to internal/exitcode, prints one line to stderr, and a command line cobra refuses (unknown command, flag or argument count) is exit code 2.

  • Workspaces and agents (issue #99, provisional until D37 names them): whr ws add <name> --path <folder> --repo <owner/name> --role <role> [--from <path|url>] [--branch <name>], whr ws ls, whr ws rm <name>, whr agent add <workspace> <role>, whr agent ls [workspace] and whr agent rm <workspace>/<role>. They call the routes above; the server checks the folder (CheckWorkspacePath) and the source, not the client. --branch only restates the workspace’s branch, which is the repository’s publication target (the configuration’s integration_branch, develop for the integration preset, or the forge’s default branch for a published repository); a different one is refused, and so is a name that is not a valid branch name outside agent/. whr ls prints an agent as <workspace>/<role>.
  • The job (issue #38; provisional): whr service install|uninstall|status runs whr serve as a LaunchAgent in the user’s Aqua session, the only place Apple Container’s services (gui/<uid>) can be reached. The plist (internal/launchd) runs container system start --disable-kernel-install and then exec whr serve --config <file> through /bin/sh, with the three paths as arguments, never text of the script; KeepAlive with ThrottleInterval 30, LimitLoadToSessionType Aqua, PATH and nothing else in the environment (no secret: the configuration holds paths), and its output in ~/Library/Logs/whr. Uninstall and status need only the label and the home directory, so they work when whr or container is gone; root is refused, a sudo shell can report Aqua; the log directory is made 0700 even when it exists, and a reinstall tries the load again for a moment after the unload. It is written 0644 by the user, binary path as resolved from the installed whr (/opt/whr or $(brew --prefix)/opt/whr, D24; with whr setup --dev, the development installation it passes as --whr <binary>, issue #261) and refused inside a git working tree (D34), after launchctl managername says Aqua and the configuration loads. A golden file and plutil -lint check the plist; install, restart on kill -9 and uninstall were run live (verified). The job’s logs are not rotated, a standard-user whr and a real reboot are unverified (issue #73).
  • Opening a topic (issue #24, #59; provisional): whr open <workspace>[/<role>] (POST /v1/workspaces/{workspace}/open) exports the agent’s branch out of its running environment into the supervisor’s own repository, clones it into <state_dir>/open/<workspace>.<role> (a supervisor-owned copy with no hooks and none of the agent’s config, never the agent’s checkout, refreshed by fast-forward and never over your edits) and prints the path on stdout, so code "$(whr open docs)" works; the files in the copy an editor may run by itself are listed on stderr. The role may be left out when the workspace has one agent. The mirror is fetched from the repository’s public https address, so a private repository fails at the fetch until the authenticated fetch of the push flow (#27) lands (unverified against a real repository). --editor of the sketch above is not built: the command prints a path.
  • The stable set (D37): serve, run, ls, logs [-f], say, cancel, inbox, approve, reject, answer; version and tools build stay. The other names of §9.1 are provisional and not built.
  • A client of the API. Every command except serve, version and tools calls the JSON API (§9.7). The server address is listen and the token is read with config.ReadSecret from api_token_file, both from the configuration file (--config, then $WHR_CONFIG, then ~/.config/whr/config.json). A token never comes from a command line or the environment, and a token file other users can read is refused.
  • Output. Stdout is data and nothing else: a table, or with --json the API’s envelope exactly as received (schema_version, ok, data), or with logs -f --json one event per line. Everything for a human (errors, reconnect notices) goes to stderr as one line. The exit code is the server’s (§9.2). Text that came from an agent, an issue or the forge is shown with control characters replaced by ?, so it cannot move the cursor or rewrite the terminal.
  • References. A task is named by its ID, a unique prefix of it, or repo#42 / #42 (the newest unfinished task of that issue); a decision by its ID or a unique prefix. An ambiguous reference is a usage error that lists the candidates; an unknown one is exit code 3.
  • Approving. A review Decision is answered for one commit, so approve and answer on a review need --sha <commit>, the commit whr inbox showed; without it the command refuses before it sends anything.
  • Mutations send a fresh Idempotency-Key (--idempotency-key reuses one), so a retry by a script is safe (§9.7).
  • Following (logs -f) reconnects after a dropped connection with Last-Event-ID set to the last event it printed, so nothing is lost or printed twice; Ctrl-C ends it with exit code 0. Without -f, logs prints the stored events (GET /v1/tasks/{task}/log) and returns.
  • Completion is cobra’s, with dynamic task IDs, decision IDs and their options, workspace names and <workspace>/<role> for --agent and whr agent rm. whr completion bash|zsh|fish|powershell prints the script on stdout (for example source <(whr completion zsh)); it needs no configuration or server, and the dynamic values come from the API when the user presses tab.
  • whr serve (internal/serve) loads and validates the whole configuration, opens the database in state_dir, builds the Apple Container runtime (with whr-shim from the tool store as the launcher), the Claude Code adapter, hostgit and the environment spec, reconciles once before it accepts a request (D6) and then every 30 seconds, and serves the API (§9.7) on the unix socket and the web UI on the loopback listen address until it is stopped; on the way out the sessions are stopped and their runs stay resumable. An environment is hardened (read-only root, all capabilities dropped, a non-root user, an internal network of its own), mounts the tool store read-only at /tools and the workspace folder at /ws, keeps the agent home on a named volume of its own, and reaches the network only through the egress proxy installed by make install. agent_permission_mode chooses how the agent’s permission prompts are handled: dontAsk (the default) runs only the tools in agent_allowed_tools, which is then required: the supervisor refuses to start without it rather than invent a policy; manual sends every prompt to the inbox as an approval Decision (D26, issue #75) and refuses an allowlist. The forge is the GitHub App client of §10.

10. Forge, CI and identity integrations

The GitHub client as built (issue #27, internal/forge/github, standard library only). It authenticates as the App: an RS256 JWT (backdated a minute, valid nine, iss the App ID) signed with the key from github.key_file (read with config.ReadSecret, PKCS#1 or PKCS#8, registered with the redactor) is exchanged for an installation token scoped to the one repository the call is about, with the permissions the supervisor needs and no more (issues and pull_requests write for comments and PRs, contents write for the host’s push, metadata read). The token is cached until a minute before it expires, minted again after a 401, registered with the redactor the moment it exists, held in memory only, and never reaches an environment (D18). A call about a repository that is not configured is refused before a request is made. It implements forge.Adapter: GetIssue (title, body, author and the author association the trust tiers read; a pull request is not an issue), CommentIssue, BranchSHA, OpenPR and UpdatePR (each checks that the branch is at the approved commit and fails with ErrStale otherwise, on top of the Guard, which also refuses any head that is not an agent/* branch before it calls the forge) and VerifyWebhook (HMAC-SHA256, constant time, refused when no secret is configured). Failures are typed: ErrNotFound, ErrAuth and ErrRateLimited (with the time to retry from Retry-After or X-RateLimit-Reset), all also an *APIError. The base URL must be https, or http to a loopback address for a test, and an error never carries a credential. GraphQL, for the project board (D30), is not built yet. Verified so far against a fake API; the real-ruleset check waits for the App of issue #73. The board mirror as built (issue #70, D30; forge.Board, internal/service/board.go, internal/forge/github/board.go). With board in the configuration, the supervisor writes the project board of the issue a task started from. A worker applies each change in order and never holds up or fails a task: a write that fails is reported and the next one still runs. Every saved task-state event is mapped: awaiting_guidance to “Needs you”, running to “In progress”, ready_for_review to “Ready to push”, completed to “Done”, failed to “Needs you” and cancelled to “Todo” (a queued task changes no card), and the last change of a batch wins. The card’s Session is the agent as <workspace>/<role> (a text field, or the option of that name in a single-select field; an option the board does not have is left alone) and its Task text field links the task on board.public_url. Nothing written comes from the issue’s text, so untrusted input cannot steer a card, and the write goes through the Guard under the table action update_board (auto by default, never an agent’s action). The GitHub client uses the GraphQL API (D31): it reads the project and its fields once, finds the issue’s card or adds the issue to the project, and sets each field with updateProjectV2ItemFieldValue; GraphQL errors, a missing project, field or option are errors (ErrBoard), and the project is read again after one. When GitHub refuses the App the board (a missing permission, or a project its token cannot see, which a user-owned project may be), the error is also ErrBoardNotWritable, not a generic 403, and whr doctor (forge-board) reports “board not writable” for it; the same check reads the project and its four Status options without writing and reports not_verified when they are all there, because only a real write proves access. Only the board’s own installation token asks for the extra permission organization_projects: write (the manifest and whr doctor expect it when a board is configured); every other token keeps the base permissions, so an installation that has not accepted it, or a user-owned board, costs the board and nothing else (ErrBoardPermission, which whr doctor names). All of this is tested against a fake GraphQL API only: whether an App installation token can write fields on a user-owned project, whether card-move webhooks exist for one, and the field names of the real board are unverified until issue #70’s spike runs with the real App (#73); for a user-owned project the permission does not apply and the fallback is the human’s decision (D30).

Keep Git transport separate from forge API operations. The forge API is called from Go with the App’s installation token, never through gh (D31), and a project board can mirror task state (D30). Release 1 ships one forge, GitHub (D15), through a GitHub App installation, with no manual-handoff half-state.

ProviderLoginAutomation
GiteaOAuth2, configurable instance URLScoped bot credentials
Forgejo / CodebergOAuth2; Codeberg as Forgejo presetForgejo adapter
GitLabOAuth2, hosted or self-managedGitLab adapter
GitHubOAuth2GitHub App installation credentials

Result chain: Task → branch → commit SHA → PR → CI results (ReviewCandidate). CI adapter (Drone, later Woodpecker) covers state, links, logs/artifacts, retry and cancel; it is an interface only in release 1. CLI login is browser-based with supervisor-issued credentials; a supervisor-owned device flow supports headless terminals. No separate identity service for a personal deployment.

The account as built (D49, issue #157). The configuration key account is dedicated (the default) or shared and is the human’s statement; whr setup asks once in config-base. The shared check account (any account may run it) reads dseditgroup -o checkmember -m <user> admin and the configuration as plain JSON, so it works before the file is complete, and says ok for a dedicated standard account, warn for an administrator or a shared account, and fail for those with remote access, which the check reads as a public_url (or board.public_url) or the console’s ssh_ca_key_file, or the Mac’s own Remote Login or Screen Sharing being on (launchctl print system/com.openssh.sshd and system/com.apple.screensharing print something; a job it cannot print counts as off, so the reading is unverified); the tailscale step cannot be read and does not count. The messages name what is given up: on a shared account every local process can read the API token and drive the supervisor, and an administrator or shared account can replace the whr binary in the prefix. After drop-admin the step tells the human to log out or restart. whr-user says warn for an administrator whr, the prefix check accepts a prefix owned by an administrator whr, and whr setup asks one y naming the risk when account would fail. drop-admin is Skipped (“refused”) when dscl . -read /Groups/admin GroupMembership lists nobody but root and the account, or cannot be read, and “not offered” for account: shared; otherwise it is warn and its fix is sudo dseditgroup -o edit -d <user> -t user admin then sudo -k, with the refusals checked again after the confirmation. Those two output formats and the effect of sudo -k on a running session are unverified until wh/verify measures them (issue #157).