# 9agents design 9agents is the coding-agents fs daemon: a long-running 9P2000 server that replaces the `zmxify` rc script's introspection half with a unified, file-shaped, always-fresh view of every AI-agent harness's state on the machine — Claude Code (`~/.claude`), Codex (`~/.codex`), omp (`~/.omp`), hermes (`~/.hermes`) and dsh (`~/.dsh`), plus their skills. Everything is read-only and read live from disk at request time; nothing is cached, so a session transcript grows as its harness writes it and a new session appears as soon as its file lands. One binary, hosted Linux only (it posts through `cloud9.post`). The backend is `src/tree.zig`, a `cloud9.fs.Server` backend run by `cloud9.serve.Runner`; the CLI is `src/main.zig`. Like its siblings (9proc, 9ns) the build is a fragment in `build.zig` wired into the root behind `-D9agents` (steps `9agents`, `9agents-test`, `9agents-itest`; installed by the plain `zig build` next to the other programs). ## The tree (v1 contract) ``` / read-only root (dr-xr-xr-x) /README the tree explained in place: what each entry is, the /active fields, and the exclusions. Static text, so it is the one fact needing no buffer. /pid the daemon's pid, one line /uptime seconds since start, one line /claude/ projects/ mirror of ~/.claude/projects//: one dir per project, its *.jsonl session transcripts as plain readable files, memory/ subdirs included history ~/.claude/history.jsonl (global prompt history) skills/ mirror of ~/.claude/skills /codex/ sessions/ mirror of ~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl session-index ~/.codex/session_index.jsonl history ~/.codex/history.jsonl /omp/ mirror of ~/.omp/agent (history.db, agent.db, models.db, ... served as raw readable blobs; sqlite parsing is a later step — v1 is files) /hermes/ mirror of ~/.hermes (state.db raw, logs/) /dsh/ mirror of ~/.dsh (profiles/, storages/, ... whatever it holds) /skills/ union view: plain dirs claude/, codex/, omp/ mirroring each harness's skills dir; a harness without one simply has no entry ``` Mirroring is lazy: nothing is walked at startup. A lookup, getattr or readdir stats and lists the real files at request time. Unknown files and directories appear as themselves, readable. A file that vanishes between lookup and read answers ENOENT cleanly; a file that appears is visible on the next request. Writes, creates, removes and setattrs answer EPERM (create/remove/wstat are not declared as backend features, so the engine refuses them itself; the write and setattr ops answer `E.PERM`, as does an open for write). Modes are reported read-only regardless of the real bits: directories `0o555`, files `0o444`. Sizes and mtimes are real. The engine's known fidelity gap applies: directory records carry fixed modes and length 0, so `9p ls -l` shows placeholders; `9p stat` is the accurate one. ## The exclusions (the security boundary) The daemon must never become a credential reader for anything that mounts it. The rule is absolute — `excluded()` in tree.zig, unit-tested — and applies to every path component of every subtree, by lookup *and* by readdir, so an excluded name neither resolves nor lists: - names containing `credentials`, `token`, `auth` or `secret` (case-insensitive substrings; "auth" also hides "author-notes" — the price of never guessing wrong); - names ending `.key`, `.pem` or `.env`; - the config/settings files of the mirrored harnesses, which may embed API keys: `settings.json`, `settings.yaml`, `settings.local.json`, `config.yml`, `config.yaml`, `config.toml`, `models.yml`, `models.yaml` (Claude Code's settings.json is not in the tree anyway: /claude serves only projects/, skills/ and history.jsonl); - `.ssh`, and `id_rsa`/`id_dsa`/`id_ecdsa`/`id_ed25519` prefixes. Symlinks are never served, at any depth: they are an escape hatch around the root pinning (a symlink out of `~/.claude` would make the daemon read outside the five roots). A symlink answers ENOENT and is not listed. Path resolution is the mechanism that makes that true, and it is not string composition. `openIn` walks from the pinned base one component at a time, each opened with `O_NOFOLLOW`: the intermediate components with `O_PATH|O_DIRECTORY` (so a component swapped for a symlink fails with ENOTDIR rather than redirecting the walk) and the last with the caller's flags. Every stat, read and readdir goes through it. Composing `/` and opening that in one call is what the first version did, and it was wrong in a way a test could not see: only the final component's symlink was refused, so a name replaced between the walk and the read — the plain TOCTOU — served bytes from outside every root, and a directory above it could redirect the whole path at any time. `O_PATH` also means a fifo or device left in a harness root is stat'd without its open ever blocking; a read refuses anything that is not a regular file. The base itself comes from $HOME or `--root NAME=PATH` at startup and is resolved normally (it is configuration, and may legitimately be a link). Relative paths are composed only from remembered (root, rel) pairs plus single-segment, engine-vetted names. Client bytes never form a path: names with `/` or NUL are refused before the filesystem is touched, `.` and `..` are refused as components inside the walk, and `..` as a lookup resolves through the backend's own parent map, never the kernel's. ## Qids, and what the protocol says they mean Read out of the Plan 9 source (`~/05-genizah/principia-softwarica`), not recalled — the three rules below were all being broken. `intro(5)`: "The qid represents the server's unique identification for the file being accessed: two files on the same server hierarchy are the same if and only if their qids are the same. ... The path is an integer unique among all files in the hierarchy. If a file is deleted and recreated with the same name in the same directory, the old and new path components of the qids should be different. The version is a version number for a file; typically, it is incremented every time the file is modified." * **The qid path is the file's identity, not the server's bookkeeping.** The backend's node id is a slot in the path table, and a slot freed by its last release returns with a new generation — the guard that stops a forged node id resolving. Reporting that as the qid path meant `9p stat` of one unchanged file answered a *different* path every time (`...50000104`, `...70000104`, `...90000104`), which by the rule above says "a different file" on every walk. The qid path now comes from the file itself, the way u9fs builds one — `qid.vers = st->st_mtime ^ (st->st_size << 8)` and a path from the inode with the device mixed in (`u9fs.c`) — while the engine goes on resolving by node id. `fs.Attr.path` exists for exactly this split and the engine uses it for nothing else. The device is folded in because an inode number is only unique within its filesystem and the five roots need not share one. * **The qid version has to move when the file does.** It was 0 everywhere, which tells every client that nothing this server holds ever changes — precisely wrong for a tree whose whole point is transcripts that grow while you read them. It is now `mtime ^ (size << 8)` for a file on disk (u9fs's formula: the shift is what makes an append within one second still change it) and a hash of the text for the fields this server derives rather than reads, where `status` flipping `busy` to `idle` keeps its size and only the content can carry the change. * **A directory's length is 0 on purpose.** `stat(5)`: "Directories and most files representing devices have a conventional length of 0." That part of the engine's "fidelity gap" note below is the convention, not a shortfall. ## Errors are strings `error(5)` is `Rerror tag[2] ename[s]`: 9P2000 has no numeric error, and acme names its refusals in words (`Eperm[] = "permission denied"`, `Eexist[] = "file does not exist"`, `editors/acme/fsys.c`). A server that answers only bare errnos throws away the one channel the protocol gives it for explaining itself, and a move has seven different reasons to refuse that all used to arrive as one undifferentiated `EPERM`: moves not allowed, illegal name, no session resolved, no cwd, dsh has no resume form, nothing to resume, the name is taken. Each says so now, and so does the over-cap listing, which used to claim "Too many open files in system" — a true sentence about the wrong resource. The errno is still carried beside the text, because the reader may be a Linux mount rather than a person: 9ns maps the *text* back to an errno by case-insensitive substring, first match wins (`enameToErrno`, 9ns/src/nine.zig). So each message keeps the phrase that carries its errno — "not permitted" for EPERM, "invalid" for EINVAL, "exists" for EEXIST, "no such" for ENOENT — and avoids the ones that would silently retarget it: "permission" and "denied" map to EACCES, "read-only" to EROFS, and the bare substring "fid" to EBADF. That is why the read-only refusal reads "9agents serves a view, not a store: writing is not permitted" and not the more obvious wording. Opens refused by the mode bits never reach the backend at all: the engine checks `perm` first and answers `e_perm`, which is already acme's spelling. ## Backend shape `Harness` (tree.zig) is the backend of `fs.Server(Harness, opts)` on `serve.Runner` (main.zig): allocation-free request paths, comptime caps, `Io` passed explicitly, static memory in .bss (the path table and the per-connection buffers; ~3 MB of tables, untouched pages cost nothing). - **Node ids** pack `Node{idx: u8, kind: u8, serial: u48}` in the u64, zmx-style. Static skeleton nodes (facts, the harness dirs, skills) are `kind = .top`; every mirrored file is `kind = .path` with `serial = (slot, generation)` of its table entry. Two walks of the same file find the same entry, so their node ids — and qid paths — are equal; an entry freed by its last release takes a new generation, so a stale forged node id never resolves. - **The path table** remembers at most 4096 remembered (root, rel) pairs, each with a Wyhash of the pair for dedupe and a reference count. The backend declares `features = .{ .references = true }`, so the engine asks lookups for `.` too and pays every reference back with exactly one `release`; an entry's refcount reaching zero frees it. Readdir records name entries without references (a listing does not pin), so a later walk of the same name finds the same entry and the same node id. A full table answers `E.NFILE` — on a lookup, and on a listing that cannot name all its entries (see Readdir). - **Fresh reads**: a read opens the file, `readPositionalAll`s the window at the request's offset, and closes it — every read is against the live file, and the engine handles offsets, so `cat` of a 5 MB transcript through a mntgen mount works. - **Readdir** collects a directory's surviving names, sorts them for cross-read cursor stability, and stages the engine's records (`node:u64le dir:u8 len:u8 name`), skipping `req.off` records. The comptime caps (1024 names, 128 KiB of name bytes, the path table) are **loud**: a directory that would not fit answers `E.NFILE` instead of a short listing, because a short listing cannot be told apart from a small directory and is therefore a wrong answer, not a limitation. The staging buffer filling is different and normal — it ends that read at the last record that fit, and the next read continues from there; staging stops at the first record that does not fit rather than packing a later, shorter one behind it, which would drop that entry from the listing entirely. A directory that changes between the reads of one listing may shift its cursor — the fresh-tree trade, accepted in v1. - **Locking**: one mutex spans `handle` and the reply (the answer's bytes point into shared staging buffers), so the backend is serialized across connections. v1 accepts this: it is an introspection fs, not a throughput service. ## CLI and process model ``` usage: 9agents [--unix PATH | --tcp IP:PORT | --fd N] [--no-post] [--name NAME] [--root NAME=PATH]... ``` - Default: post itself under the name `agents` via `serve.Runner.listenPosted`, so it appears at `$XDG_RUNTIME_DIR/9p/agents` and every interactive fish (self-wrapped in `9ns --mntgen`) sees it at `/mnt/9p/agents` with zero configuration. `stop()` unposts; the stale-socket protocol covers a killed daemon. - `--unix`/`--tcp` listen forms (beside the post, or with `--no-post`), `--name` to post under another name, `--root NAME=PATH` to pin any of the five roots elsewhere (defaults are `$HOME/.`; the tests use fake homes and never the live roots). - `--fd N` serves exactly one 9P session over the connected stream on descriptor N (socket activation, `9ns --spawn` handoffs), driving the engine directly; it posts nothing. - SIGTERM/SIGINT stop cleanly (unpost); SIGPIPE is ignored. - No daemonization in v1: the daemon never forks or detaches, so it is run under something that supervises it. On this machine that is a systemd user service (`~/.config/systemd/user/9agents.service`, `Restart=always`, lingering on), which is what keeps `/mnt/9p/agents` there across logouts and reboots; a restart reclaims the posted name through the stale-socket protocol. For a throwaway instance the zmx way still works — `zmx run agents -d 9agents`. ## Integration test plan (`test/e2e.sh`) Part A runs entirely on fixtures (a fake home pinned by `--root`, a scratch `XDG_RUNTIME_DIR` registry, plan9port's `9p` as the client): 1. the posted name is listed in the registry; 2. `9p ls /` shows the roots; a fixture transcript reads back byte-identical (`cmp` against the source file); the skills union mirrors the harness tree; 3. credentials-shaped files are unreachable by both listing and reading, at several depths, in every root; 4. a file appended after the daemon started is visible immediately, and a new session directory appears in the next listing; 5. the pid fact answers the daemon's pid; writes answer EPERM; 6. the `--unix` and `--fd` listen forms serve; 7. SIGTERM unposts the name. Part B is the money shot, read-only against the real roots and the real `$XDG_RUNTIME_DIR/9p` — only when the `agents` name is free (a live daemon owning the name skips it, not fails): the posted name in the real registry, `9p` walks of the live `~/.claude`, a real transcript byte-identical through the tree, the live `~/.claude/.credentials.json` unreachable, the `9ns --mntgen` mount of `/mnt/9p/agents` with `head` of `/claude/history`, and the stop unposting the real name. Unit tests (tree.zig) cover the exclusions (including the predicate itself, spelled out), path-composition safety (slashes, NULs, symlinks, the `..` chain, NOTDIR), node id packing (same file equal, distinct files differ, the union identity, id recycling through release), the vanishing file (ENOENT on lookup, getattr and read), EPERM on write/setattr/open-for-write, and live visibility — all against a fake HOME under a test tmp dir, never the real harness roots. ## Verification - `zig build` (installs `9agents` beside 9ns, 9proc-demo, 9web). - `zig build test` — the library's 80 tests stay green. - `zig build 9agents-test` — the tree's unit tests. - `zig build 9agents-itest` — `test/e2e.sh` (fixture + real registry). - `zig build 9proc-check-freestanding` — the daemon is hosted-only but must not break the root module. - `zig build programs-test` / `programs-itest` — the umbrella steps, which include 9agents's. ## /active: the derived view (v2) v1 mirrors files. `/active` is the first *semantic* layer: one directory per live agent, normalized across harnesses, so `ls /active` answers "what is running right now" whatever wrote it. It is the discovery half of `~/.local/bin/zmxify`, whose resolution ladder it ports. ``` /active/ claude/ 345104/ pid 345104 ppid 344980 started 1790013029 cwd /home/goblin name goblin-e5 title the session's title session 1f9a74f7-... model claude-opus-5[1m] via registry zmx harness (read-write, below) status busy only where the harness publishes one transcript the live .jsonl, growing as it writes agents/ /{model,transcript} omp/ 155574/ ... ``` A harness directory holds one directory per live process of it, named by pid. `/proc` spells it that way and so does this: a compound `claude-345104` would make you parse a name to recover `harness`, which is already the directory above it, and `/active/omp/*` would not glob. There is no `updated` file. The last write to a session is the mtime of `transcript`, which `stat` already carries; serving it again as its own file would be the same fact twice, and the copy is the one that goes stale. These are zmxify's picker columns — pid, harness, dir, age, zmx, via, session, title — as files, which is the point: the script stops scanning `/proc`, reading fds and querying sqlite itself and just reads the tree. **`status` is a promise only some harnesses make.** It is the harness's own word for what it is doing, and only Claude Code publishes one (`busy`, in `sessions/.json`); for every other harness the file is simply absent, the way a `skills` mount with no target is absent. Deriving one from the process's `/proc` state would answer a different question (`S` means "not on a CPU this instant", not "waiting for you"). For every other harness the honest activity signal is the mtime of `transcript`. `ls /active/*/*/status` tells you who publishes the stronger answer. An entry is named `-`: unique, stable for the process's life, and it sorts by harness. The pid is the identity because it is what `/proc` and every harness's own registry agree on. ### Finding the agents `/proc` is scanned for a process whose `argv[0]` basename is `omp`, `claude`, `codex`, `hermes` or `dsh`, or a python running `hermes_cli.main`. A command line carrying one of the daemon words (`gateway`, `dashboard`, `mcp`, `mcp-server`, `app-server`, `exec-server`, `serve`, `daemon`, `acp`, `ps`, `render`, `export`, `__omp_worker_daemon_broker`) is a server or a helper, never a session. The daemon's own ancestry is excluded, so 9agents can never list or act on the process tree it lives in. Liveness is `/proc/` existing **and** its `stat` field 22 (starttime) matching the one remembered for the slot. A pid that has been reused is a different process and does not list; a corpse never lists. ### Resolving the session Each harness is asked in its own terms — the `fd -> dir -> db` ladder zmxify worked out, with the route named in `via` so a wrong guess is visible rather than silent: | harness | route | `via` | |---|---|---| | claude | `~/.claude/sessions/.json`, which the harness maintains itself: sessionId, cwd, name, status, version, and `procStart` as a pid-reuse guard | `registry` | | omp | the transcript it holds open under `~/.omp/agent/sessions/`, newest first; else the store directory named after its cwd with `/` becoming `-` | `fd`, `dir` | | dsh | the `session-/session.jsonl.zstd` it holds open; zstd, so there is no title | `fd` | | codex | the rollout it holds open — `sessions/YYYY/MM/DD/rollout--.jsonl`, whose *name* carries the session id, so its sqlite is never opened | `fd` | | hermes | nothing to resolve: see below | `none` | A resolved session is checked against the process's cwd before it is believed (the head of an omp transcript carries `"cwd"`). A process whose session does not resolve still appears, with everything `/proc` knows and an empty `session`: the view never pretends to know what it does not, and a move on it is refused. codex's id is the last 36 bytes of the rollout's name, not what splitting on `-` gives — the timestamp in front of it holds dashes too. **hermes is the one harness with no answer here, and it needs none.** Its sessions live only in `state.db`, with no per-session file to find; but the only hermes processes that run are the gateway and the dashboard, and both are daemon-shaped, so neither is a session anything should move. zmxify resolved a hermes id from sqlite and then declined to touch the process holding it, for the same reason. If an interactive hermes ever exists, this is the gap, and it is the one place sqlite would buy something. ### Freshness A slot remembers only what identifies the agent: harness, pid, starttime, cwd, session id and transcript path. Every *field* is re-derived from disk when it is read, so `status` is never stale and a transcript grows under `cat`. `/proc` is rescanned on a readdir of `/active` and on a lookup that misses, not on every read. ### `zmx`: the write path, and why it is not a ctl Writing a zmx session name into `/active//zmx` moves that agent into a zmx session of that name: the zmxify action half, as a file. Reading `zmx` gives the `ZMX_SESSION` of the process, empty when it runs outside zmx. Writing sets it. The file means *which zmx session this agent lives in*, and writing a name makes that true — state, not a verb channel. A `ctl` taking words would be the ordinary Plan 9 spelling (`/proc/n/ctl`), and an executable `zmxify` script served in the tree would be the zmx `attach` spelling, but a script that shells out to a local binary is a lie over a remote mount: it would run against a session that is not on the client's machine. A write is served where the authority is, so it survives being mounted from anywhere. What a write does, in order, refusing before it destroys anything: 1. The name must be zmx's label charset (`[A-Za-z0-9._-]`) and unused by a live session, else `EEXIST`. 2. The agent's session must have resolved, else `EPERM`. Nothing is killed that has nowhere to come back to — zmxify's invariant. 3. `SIGTERM`, then `SIGKILL` after 5s, giving up at 12s with `EIO`. 4. `fork`, `chdir` to the agent's cwd, `exec zmx run -d` with a **fixed argv per harness** (`claude --resume `, `omp --resume `, `codex resume `, `hermes --resume `). No client byte ever reaches `exec`: the only thing the client supplies is the session name, and it is validated first. 5. The write returns once `$XDG_RUNTIME_DIR/9p/zmx/` appears (zmx self-posts), or `EIO` on timeout. The agent comes back under a new pid, so the old `/active/` is gone and a new entry takes its place with `zmx` reading the new name. The window between the kill and the exec is the same exposure zmxify has always had, and it is why step 2 comes first. **This is the tree's first write path, and it is an escalation**: a client that can write this file can kill the user's agents and cause a process to be spawned. It is therefore off unless `--allow-move` is given, and the read-only daemon stays the default. Everything else in the tree still answers `EPERM` to writes. ### Where this is not Plan 9 Named, because they are choices rather than oversights: * **`via` describes how the server found out**, not something true of the process. That is diagnostics in the interface. It stays because a resolution that guesses wrong silently is worse than a wart — it is why zmxify's picker showed the route too. * **A write to `zmx` does not change the object, it replaces it.** The process is killed and another starts under a new pid, so the entry written to disappears. `/proc/n/ctl` accepting `kill` is at least honest about being a verb with a consequence. The file still beats an executable served in the tree, which would lie outright over a remote mount, so it stays — with this paragraph. * **A move blocks the whole daemon** for as long as the kill and the restart take, because one mutex spans `handle` and the reply. A Plan 9 server keeps answering other fids meanwhile. The engine already has parked replies and Tflush for exactly this; the move does not use them yet, and that is the first thing to revisit. * **`/active` lives inside the mirror** rather than being its own posted service. "What files exist" and "what is running" are different jobs, and a stricter reading would separate them; they share the pinned roots and the parsing glue, so they share a daemon. ### Replacing zmxify There is no replacement script, because the mount is the replacement: ``` ls /mnt/9p/agents/active/*/* what is running cat /mnt/9p/agents/active/claude/345104/session what it is echo mine > /mnt/9p/agents/active/claude/345104/zmx zmx attach mine ``` `~/.local/bin/zmxify` was 307 lines of rc because there was no filesystem to ask: it scanned `/proc`, classified command lines, read fd tables and queried two sqlite databases to work out what the four lines above now read. All of that moved into the daemon, where it is tested. Writing a wrapper back on top would only re-derive what the tree already says, and would become a second, worse interface beside the real one. Two behaviours the old script had are worth knowing, because they are now the caller's to arrange and both are one line: pick a session name nothing holds (the registry, `$XDG_RUNTIME_DIR/9p/zmx`, says which are taken) and `zmx attach` afterwards. A third is gone rather than moved. The old script refused to act on its own ancestry, because it did the killing itself: kill your own parent and the script dies before it can re-exec, losing the session it was rescuing. The daemon kills, re-execs and waits for the new session to post, and completes all of it whether or not the client is still connected — verified by hanging up immediately after sending the write and watching the move land anyway. So moving the terminal you are sitting in works, which is the common case; the shell drops, and you reattach with the name you wrote. ### Testability `--proc DIR` overrides `/proc` the way `--root NAME=PATH` overrides a harness root, so the scan, the liveness guard, the resolution ladder and the exclusions are unit-tested against a fixture process tree and never against the live machine. The move itself is exercised end-to-end against a fake harness in `test/e2e.sh`. ## Out of scope for v1 (documented, not hidden) - Write paths (create/write/setattr stay EPERM), resume/attach automation (the zmxify re-exec half), sqlite parsing (omp/hermes sessions stay raw blobs), cross-session search, any caching or invalidation layer. - Streaming/parked reads (a read answers at once; the transcripts are plain files), and union-directory semantics beyond the plain /skills/ dirs.