Agent Workspaces
Superseded (2026-07-31): the workspace now runs as a k3s pod, not a host service. The shared-host Remote Control server described below was retired —
agent-ws-agent.serviceandagent-repos-init.serviceare gone fromnixos/ringtail/(the file is nowagent-heph-spoke.nix, which keeps only theagentuser + heph spoke). The agent ran containerized with its owntag:agentTailscale identity; see agent-containerization for that design. Theagent-wspod is now retired too, superseded by talos — its image, k8s manifests, and ArgoCD app are gone. This page is retained for the isolation-model rationale and history — the §Isolation & security and §The heph spoke sections still describe real, current concerns.
A persistent Claude Code Remote Control server on
ringtail — a single home-base session environment rooted in the agents
repo — that Erich steers on demand from the Claude mobile app or claude.ai/code.
The server spawns isolated per-session git worktrees of the home-base repo, so
several sessions can run concurrently; every other repo is a sibling checkout
the session cds into.
This is the infrastructure landing of a prototype researched in the research
repo (2026/July/Remote-Agent-Sessions-On-Ringtail/). BlumeOps is the platform;
the agent workflows are built upon it.
Design
Claude mobile app / claude.ai/code
│ Anthropic relay (outbound-only from ringtail; no inbound ports)
▼
ringtail — user `agent` (isolated, non-wheel)
systemd: agent-ws-agent.service (the single home-base workspace)
└─ claude remote-control --spawn worktree --name ringtail-agent
cwd = ~/code/personal/agents
PATH prepends the transparent `op` shim + ~/.cargo/bin (heph)
+ the report toolchain (mise/uv/pandoc/typst/weasyprint)
systemd: agent-repos-init.service (oneshot, clones/updates repos first)
Checkouts live at ~/code/personal/<repo> (not under a per-workspace dir) so
the paths agents read in the repo docs — every AGENTS.md/CLAUDE.md assumes
~/code/personal/… — actually resolve on the agent box.
The home base: the agents repo
The session root is the agents repo:
a small repo whose AGENTS.md (with CLAUDE.md symlinked to it) carries the
base instructions every remote session wakes up with — the repo map
(blumeops, hephaestus, research, …), the shared toolbox (heph, mise, uv, op,
tea, …), and the execution environments (gilbert, ringtail,
ringtail-agent). Per-repo details stay in each repo’s own
AGENTS.md/CLAUDE.md, which the session reads after cding in.
Repos cloned by agent-repos-init:
| Repo | Role |
|---|---|
agents | Primary — session cwd, worktree-isolated, base instructions |
hephaestus, hephaestus.nvim | Sibling checkouts |
research | Sibling checkout |
timberborn-parsimony | Sibling checkout |
blumeops | Pool-only, author-only — see §blumeops |
Historical list.
agent-repos-initis gone; the live pool is the clone loop in the talos entrypoint (default.nix in the eblume/talos repo), driven by blumeopsrepos.jsonmounted as a ConfigMap (rolls the pod on change) — see agent-containerization §“The repo pool”.
Concurrency caveat. Only the primary repo gets per-session worktree isolation (that is what
--spawn worktreeoperates on). Sibling checkouts are shared between concurrent sessions — the base instructions tell agents to work siblings on a session-named branch (or a manualgit worktree addinto the session’s own worktree) so two sessions never fight over a checkout.Resolved in the containerized model (now in talos): the workspace tool makes those per-session worktrees automatically from a
SessionStarthook, so the convention is now enforced rather than requested. See agent-containerization §“Per-session worktrees”.
Why one home-base server
A Remote Control server is rooted in its start directory; new sessions spawn as
worktrees of that repo and inherit its CLAUDE.md/tooling. The home-base shape
exploits that once: a session wakes up with the base instructions loaded and
cds to whatever repo the task needs. One environment in the app, one service
to operate, one trust seed, one place to keep cross-repo instructions — and
adding a repo to the fleet is a one-line clone entry plus a paragraph in the
agents repo, not a new server. Idle servers cost nothing either way (no
inference until a session is active).
Superseded decision: per-repo servers (2026-07-08 → 2026-07-11). The first landing ran one server per repo (
ringtail-hephaestus,ringtail-research,ringtail-playground, brieflyringtail-parsimony) so each session woke up inside its repo with instructions preloaded. In practice the per-repo environment picker added friction (N environments, N services, N trust seeds, duplicated base instructions) for little gain — sessions regularly needed to cross repos anyway. Replaced by the singleagentshome base; the per-repo trick (cwd sets the instruction context) still does the work, just once. Theplaygroundworkspace was dropped outright — a home-base worktree is already a safe scratch space.
blumeops: author-only, not a server
blumeops is cloned into the pool at ~/code/personal/blumeops (agent-owned, so
git works) but is not a workspace — there is no ringtail-blumeops Remote
Control server. Agents cd into it from any session to author changes and
open PRs as the bot. The bot has only read on the
canonical repo and works through its fork agents/blumeops: the clone’s
origin is the fork (push), upstream is eblume/blumeops (fetch) — branch off
upstream/main, and PRs are cross-repo. This is safe because blumeops is a
public, secret-free repo: the gate has never been the code, it is the
blumeops 1Password vault, cluster access, and CI execution — and an
agent holds none of them.
- No deploy via ansible.
provision-{indri,ringtail}and mostmisetasksop readthe blumeops vault inpre_tasks(so even--check --diffdies at the lookup, not just apply). An agent authenticates only as theagents-vault service account →403. Vault access is per-vault, all-or-nothing (no item-level whitelist), and biometricopcan’t be routed to a headless worker anyway — so there is no least-privilege subset to grant. That is by design: the vault is the operational-secret blast-radius boundary (argocd break-glass, every ansible secret). - No deploy via k8s. The k3s admin kubeconfig is
0600root-only (--write-kubeconfig-mode=600), so the agent can neitherkubectlnor read theargocd-*secrets; it is non-wheelwith no sudo. (It was0644until 2026-07-10 — a hole that did hand the agent cluster-admin and an ArgoCD-admin deploy path; see §Isolation.) - No deploy via CI. blumeops’ deploy workflows carry Actions secrets
(
ARGOCD_AUTH_TOKEN,FLY_DEPLOY_TOKEN,ZOT_CI_API_KEY,MAIN_PUSH_TOKEN).workflow_dispatchis write-gated and the bot has only read, so it cannot run a workflow — not even a modified one on a branch — to read those secrets. Without this, write access would let the bot dispatch a branch workflow that exfiltrates the deploy tokens (Forgejo has no per-run approval gate for write users), which is exactly why the bot is read-only + fork here.
Net: an agent can edit code/docs and open a blumeops PR, but a human
provision/sync on gilbert (biometric op) always sits between that PR and
production, and main is branch-protected (push + merge whitelisted to
eblume) on top. The agent-owned pool clone is distinct from the root-owned
deploy checkout at /etc/blumeops that nixos-rebuild builds from (the agent
can read its files but not git it — dubious-ownership guard).
Superseded decision. blumeops was dropped entirely on 2026-07-08 on the reasoning that author-only was too thin to bother standing up. Reinstated 2026-07-10 as a pool-only clone: authoring blumeops from other workflows turned out to be valuable context worth having, and the deploy gate holds without a server. (
adelaide-baby-shower-app, once cloned alongside, can be added the same way if wanted — no vault dependency.project-templatewas added on 2026-08-07 aswrite+canonical.)
Isolation & security
The boundaries, weakest-first:
- Vault boundary — agents authenticate to 1Password only as the
agents-ringtail-rwservice account (read/write, theagentsvault only). Everything outside that vault is unreachable. This is the hard secrets wall. - User boundary — the
agentuser is not inwheel, has its own home, and never sees Erich’s~/.claudecredentials, repo checkouts, or shell history. Also (co-tenant on the k3s host): the k3s admin kubeconfig is0600root-only (--write-kubeconfig-mode=600), so the agent cannotkubectlthe cluster or read k8s secrets. This was0644(world-readable) until 2026-07-10 — a real hole that gave the agent cluster-admin and, viaargocd-initial-admin-secret, an ArgoCD-admin deploy path entirely bypassing the vault gate below. Locking it is what makes the k8s deploy surface actually inaccessible to agents. - Credential handling — the service-account token lives at
/etc/agents/op-token(owned byagent, mode 0400) and is injected by a transparentopshim (see below) so it never appears in a session’s environment (agents dumpenvwhile debugging, and transcripts get archived). - Git-mediated writes — agents push as a dedicated Forgejo bot
(agents-forgejo-bot) and open PRs rather than committing to
main. For the repos the bot has write on (hephaestus, …) this is convention backed by human PR review. Forblumeopsit is enforced: the bot has only read on the canonical repo and pushes to its fork (agents/blumeops), opening cross-repo PRs. Read-not-write is deliberate —workflow_dispatchis write-gated, so a read-only bot cannot run blumeops CI and therefore cannot reach the deploy-credentialed Actions secrets (ARGOCD_AUTH_TOKEN,FLY_DEPLOY_TOKEN,ZOT_CI_API_KEY,MAIN_PUSH_TOKEN);mainis also push+merge-whitelisted toeblume. This is what actually keeps deploy credentials out of agent reach — layered over the fact that the bot holds neither the blumeops vault nor cluster access, so even a hypotheticalmaincommit deploys nothing until a human provisions. PRs are opened withtea, authenticated by an agents-ownedwrite:repository+write:issuePAT (agents-forgejo-tokenin the agents vault) that the launcher exports asFORGEJO_TOKENand seeds into tea’s config — no blumeops-vault dependency, and the PAT is self-bounded to the bot’s own repos (agents-forgejo-bot). - Network (in migration) — the bare-metal phase has ambient tailnet
reach: because Tailscale ACL identity is per-device, the
agentuser inherits ringtail’stag:homelabtrust and can reach indri directly (ssh erichblume@indri). The containerized phase gives the pod its owntag:agenttailnet identity behind an egress allowlist (Claude relay, 1Password,forge.ops.eblu.me, the heph hub), which is the only thing that lets the ACL fence the agent apart from the host — the identity foundation for that has landed inpulumi/tailscale/policy.hujson.
Residual risk (bare-metal phase): a same-uid agent can read
/etc/agents/op-token and the agent user’s own OAuth credential file — the
shim prevents accidental leakage, not a determined read. The vault scoping
(1) bounds the blast radius; the container phase closes the uid gap. The main
live threat is prompt injection driving malicious commits or vault-secret
exfiltration; because the bot can push to main (see 4), the backstop is
that no credential the agent can reach triggers a deploy on its own — a human
provision/sync step always sits between a commit and production.
The transparent op shim
Existing scripts call plain op and must work inside and outside agent
environments. The shim (pkgs.writeShellScriptBin "op") is prepended to the
workspace PATH: if OP_SERVICE_ACCOUNT_TOKEN is unset it reads
/etc/agents/op-token, then execs the real op. Outside agent services the
shim is not on PATH, so normal op (desktop/biometric) is unchanged.
op gotchas (learned the hard way): Claude Code’s Bash tool does not persist environment between calls, so
export OP_SERVICE_ACCOUNT_TOKEN=…in one call is gone by the next — the shim exists precisely so nothing needs to export it. Also,opparses non-TTY stdin as a JSON item template, so inside scripts call it with</dev/nullor writes fail with a misleading “invalid JSON” error that looks like a permission denial.
Writing secrets into the vault
When an agent (or a human) stores a sensitive value in the agents vault,
put it in a concealed field, not a plain-text one, so the 1Password GUI
masks it when the item is examined: op item edit <item> --vault agents "api-key[password]=…" (the [password] field type is concealed;
[text]/STRING is shown in the clear). Only truly public values (e.g. an SSH
public key) belong in a [text] field. This is defence-in-depth behind an
already-unlocked vault — minor, but free.
The heph spoke (deliberately in-boundary)
heph is the counterpoint to the vault isolation above: it is
intentionally inside the agent trust boundary, because it is the substrate
agentic workflows run on (task discovery, logging, canonical-context docs). So
the agent user runs a real hephd spoke synced to the indri hub and can
read/write the owner’s real tasks — this is the point, not a leak. (The blumeops
1Password vault stays isolated; heph is a separate, deliberately-included
substrate.)
Shape (all source-controlled — see agent-workspaces.nix; the version pin,
hub/OIDC endpoints, and install/timer/path machinery are shared with Erich’s
own eblume-heph-spoke via heph-common.nix — see ringtail §Heph
Spokes):
agent-heph-install.service— a oneshot thatcargo installsheph+hephdat a pinned tag (hephTag) using a mise-resolved Rust toolchain. nixpkgs’rustclags heph’s fast-moving floor (it shipped 1.91 vs the ≥1.92 a heph dep needs), so the toolchain comes frommise x rust@stableover thenix-ld+ build-deps setup the host already provides for mise runtimes. Idempotent: it version-checks and only recompiles on a tag bump.agent-heph-spoke.service— runshephd --mode local --hub-url http://indri…:8787(spoke sync is HTTP-only) authenticating via OIDC.- The spoke’s token lives in the agents vault (
op://agents/heph-spoke-token), not a file, via hephd’s command token store (--token-load-cmd 'op read …'/--token-save-cmd heph-token-save) — no plaintext token at rest. Theheph-token-savewrapper writes refreshes back to the vault without ever putting the token in argv (/proc/<pid>/cmdlineis world-readable) using op template files +jq --rawfile.
Identity & revocation. The spoke authenticates as a dedicated
heph-agents Authentik user in a heph-scoped group (not admins — that
would grant every admin-gated app), so it is independently revocable from the
human login. The hub admits it as a co-owner via hephd --authorized-sub <sub>
(the sub is a hashed_user_id, kept in the blumeops vault and templated into the
indri unit). Two independent kill switches, neither touching your own logins:
disable the heph-agents Authentik user, or drop its sub from the hub’s
--authorized-sub and restart (the vault token goes inert even if unexpired).
Bound the refresh-token lifetime on the Authentik provider as the third lever.
See bootstrap-agent-workspaces for the one-time seeding steps.
Report toolchain
The workspace PATH is a deliberately minimal curated set (op shim, git, ssh, tea,
coreutils, ~/.cargo/bin, ~/.local/bin) — it does not include
/run/current-system/sw/bin, so system-wide tools there are invisible to agent
sessions. The research repo’s compile-report / save-session tasks
(mise run …, each a uv run --script program) therefore need their toolchain
added explicitly. agent-workspaces.nix puts mise, uv, pandoc, typst, and
weasyprint (all nixpkgs builds) on the session PATH.
Alongside the report toolchain the launcher also adds a general CLI toolbox
(cliTools: gawk, jq, curl, python3) — the staples a plain shell agent
reaches for to munge text/JSON and write quick scripts. Without them sessions hit
command not found on awk/jq pipelines and fall back to slower workarounds
(WebFetch for curl, a full uv run --script for a python3 one-liner). uv
is already present for uv run --script; python3 is a bare interpreter for
one-off snippets and stdlib. All are nixpkgs builds on the curated PATH, so no
/run/current-system/sw/bin exposure is implied.
WeasyPrint is the wrinkle: compile-report pip-installs it into uv’s ephemeral
venv, and that venv can only render a PDF if WeasyPrint’s native libraries
(Pango, HarfBuzz, fontconfig, …) are discoverable. nix store libs live in no
default loader path, so the launcher exports LD_LIBRARY_PATH over
pkgs.{pango,glib,harfbuzz,fontconfig,freetype,gdk-pixbuf,cairo,libffi} — the
Linux counterpart of the repo’s macOS Brewfile. (pkgs.weasyprint is on PATH so
weasyprint resolves as a CLI, but the render path is the uv venv, not that
binary.) MISE_TRUSTED_CONFIG_PATHS=~/code/personal is set so mise run … never
stalls on an interactive trust prompt.
A global
mise.tomlinstalling these tools via mise backends was considered and rejected: mise’s prebuilt binaries fight NixOS’s dynamic linker (they neednix-ldat best), whereas nixpkgs builds are deterministic and already patched. The task’s goal — the report tasks run unchanged — is met by the nix PATH, not a mise config.
Build toolchain
The curated PATH also omitted a C toolchain, so cargo build — and anything
that links a native binary — failed in a session with linker `cc` not found.
The heph install oneshot sidesteps this with its own gcc/pkg-config env
(hephBuildDeps), but interactive sessions inherited none of it, so agents could
not compile or verify Rust they authored.
The launcher now adds buildTools (gcc, binutils, pkg-config,
gnumake) to the session PATH and exports CC=gcc. Rust itself still comes from
mise (rust@stable) — nixpkgs’ rustc lags, exactly as for the heph install
— so this change supplies only the linker and pkg-config that mise’s toolchain
needs to actually build.
For the gamedev Bevy project specifically, the launcher also exposes Bevy’s
Linux native deps the same way the report toolchain exposes WeasyPrint’s: alsa
and udev (pkg-config-probed at build) go on PKG_CONFIG_PATH via
gameBuildDeps, and the windowing/graphics libs that Bevy dlopens at run time
(vulkan-loader, libxkbcommon, wayland, libGL, and the core xorg
client libs) go on LD_LIBRARY_PATH via gameLibs.
This lets agents
cargo build/cargo checkto verify their work. It does not make the headless box run a windowed Bevy app — there is no GPU or display — so actually playtesting a build stays a human job on gilbert/ringtail, consistent with the timberborn-parsimony playtest rule.
Authentication
Remote Control requires a full-scope claude.ai subscription OAuth login —
the credential an interactive claude auth login writes to
~/.claude/.credentials.json, which lives on the PVC because it is refreshed
in place on use. Nothing else works: not an API key, and not a long-lived
claude setup-token credential (CLAUDE_CODE_OAUTH_TOKEN) — those are
inference-only, and claude exits at startup with Remote Control requires a full-scope login token … Run claude auth login. v0.16.0 shipped exactly
that token and crash-looped on the startup probe until v0.17.0 reverted it
(2026-08-07). An inference call succeeding with a setup-token proves nothing
about Remote Control — that’s the test that let the regression ship.
The login credential does not survive unattended operation: its refresh token
expires ~29 days after login (access tokens last ~8h and refresh in
place), and Claude Code refresh tokens are single-use, so concurrent sessions
in one pod can invalidate each other’s and end it sooner
(claude-code#24317).
Worse, its death is silent — see Known warts. The expiry was
managed operationally: a recurring 21-day heph chore (“Rotate Claude OAuth
login”, Blumeops project) re-ran the login with ~8 days of buffer. (The rotation
runbook retired with the agent-ws pod.)
The ~29 days is measured, not documented — a 2026-08-07T18:01Z login wrote
refreshTokenExpiresAtof 2026-09-05T19:54Z. Anthropic publishes no figure. These docs briefly claimed ~7 days, from mistaking the PVC’s creation date for the login date; the credential had actually been carried onto the PVC from the pre-container host login 22 days earlier. The runbook reads the real deadline off the credential at each rotation rather than trusting the number here.
Terms of use
This deployment is unsupported (Anthropic won’t treat breakage as a bug; see Known warts) but not prohibited. It stays within subscription terms as ordinary, individual usage: one human (Erich), one subscription, native Claude Code via OAuth, steered on demand. Claude Code’s legal terms peg Pro/Max limits to “ordinary, individual usage” and reserve OAuth for “ordinary use of Claude Code and other native Anthropic applications” — so the line to respect is attended, bounded use, not a constantly-running autonomous fleet. Two guardrails keep it clearly inside: (1) one human, one subscription, on demand — never multi-user or serving others (that would require API-key auth via the Agent SDK/Console instead); (2) cap or human-gate anything scheduled, since the token-wasteful always-on pattern is also the fair-use risk.
Operations
- Deploy:
mise run provision-ringtail(writes/etc/agents/*secrets via ansible, thennixos-rebuild switch). First-ever deploy needs the one-time bootstrap-agent-workspaces steps (OAuth login, trust + consent seeding). - The
agent-wshost services and pod are both retired now; live operations for the current agent run under talos (talos-design). The host model above is retained for history.
Known warts
claudefor theagentuser is installed by the official self-updating installer during bootstrap (imperative), not from nixpkgs — nixpkgs lags the npm channel and Remote Control is fast-moving. The serviceExecStartpoints at~agent/.local/bin/claude.- The service uses
util-linux’sscriptto allocate a PTY (Remote Control renders a status TUI and has no--headlessflag yet: anthropics/claude-code#30447). - This whole deployment shape (Remote Control server, headless, multi-session) is unsupported by Anthropic — treat upgrades as capable of breaking it.
- An auth failure here is invisible to every health signal we have. When the
PVC OAuth credential expired on 2026-08-07, remote-control kept its gateway
websocket up and kept printing
✔︎ Connected, so the pod health probe (which only asserts that a claude process holds an ESTABLISHED TCP connection) stayed green — while every session start failedFailed to authenticate: OAuth session expired and could not be refreshed. The 21-day rotation chore kept the expiry from being hit, but nothing detected a bad credential:claude auth status --jsoninside the pod was the check.
Related
- agent-containerization — the migration off this shared-host model
- talos-design — the agent service that supersedes the retired
agent-wspod - bootstrap-agent-workspaces — one-time setup runbook
- agents-forgejo-bot — the bot account and its key
- security-model — service accounts and the
agentsvault - ringtail — the host