diff options
| author | Gabriel Schneider <[email protected]> | 2026-09-21 16:49:20 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-09-21 16:49:20 -0300 |
| commit | dddd556accea6b6ea7802cd3f622f8b3cf8eb43f (patch) | |
| tree | 64a1cc34d44f9be6ec266d852fdef6f5b842d004 /9harness | |
| parent | b4db588dd5b92d647b661c2dc17b40925af92348 (diff) | |
| download | cloud9-dddd556accea6b6ea7802cd3f622f8b3cf8eb43f.tar.gz cloud9-dddd556accea6b6ea7802cd3f622f8b3cf8eb43f.zip | |
9harness: /active, and zmxify as a write to a file
The mirror answers what files exist. /active answers what is running:
one directory per live agent, normalized across harnesses, fields as
small text files, synthesized per request.
/active/claude/345104/{pid,cwd,session,via,name,status,title,
model,started,zmx,transcript,agents/}
This is /proc's shape, and deliberately: a directory per object named
by pid under a directory per harness, rather than a compound
`claude-345104` that would make you parse a name to recover a field
that is already the directory above it. There is no `updated` file —
that is the mtime of `transcript`, which stat already carries.
Each harness is asked in its own terms, and the route is reported in
`via` so a wrong guess is visible rather than silent. Claude Code
publishes sessions/<pid>.json itself, with procStart as a pid-reuse
guard, so nothing there is guessed. omp and dsh are found by the
transcript they hold open, omp falling back to the store named after
its cwd. codex's rollout file carries the session id in its *name*, so
its sqlite is never opened. hermes is the one gap and needs none: its
sessions live only in sqlite, and the only hermes processes that run
are the gateway and the dashboard, which are not sessions.
Liveness is /proc/<pid> plus a matching start time: a pid alone is not
an identity. The daemon never lists its own ancestry, so it cannot show
or act on the tree serving the request.
The write path, and why it is a file and not a ctl: writing a zmx
session name into an agent's `zmx` moves it there. The file means which
zmx session this agent lives in, and writing makes that true. A ctl
taking verbs is the ordinary Plan 9 spelling, and an executable script
served in the tree is the spelling zmx's own `attach` uses, but a
script that shells out to a local binary lies over a remote mount — it
would run against a session that is not on the client's machine. A
write is served where the authority is.
Every refusal comes before anything is destroyed: the name must be
zmx's label charset, unused by a live session, and the agent's session
must have resolved, because nothing is killed that has nowhere to come
back to. The command is fixed per harness and no client byte reaches
exec. It is off unless --allow-move: this is the one place the tree is
not read-only, and anything that can mount it could otherwise kill an
agent.
9harness/zmxify replaces the 307-line rc script. It parses no /proc,
opens no fd table and queries no database; it lists /active, offers the
rows to fzf and writes the chosen name. It no longer excludes the
caller's own session, which the old one had to: that script did the
killing itself, so killing its own parent lost the session it was
rescuing. The daemon completes the kill and the re-exec whether or not
the client is still connected — verified by hanging up immediately
after sending the write — so zmxifying the terminal you are sitting in
now works, which is the common case.
--proc DIR is the fixture seam: the scan, the liveness guard, the
exclusions and the ancestry rule are unit-tested against a fake process
tree, never the live one.
Suites: 87/87 root, 48/48 9ns, 26/26 9harness (+5 for the view),
60/60 9proc, 144/144 programs-test, 51+88 9ns integration, 44/0
9harness end-to-end (+16, including a real move against a fake harness
and a fake zmx), 29/0 9proc debug, 213/0 9ns adversarial, freestanding
green.
Diffstat (limited to '9harness')
| -rw-r--r-- | 9harness/docs/DESIGN.md | 208 | ||||
| -rw-r--r-- | 9harness/src/active.zig | 1006 | ||||
| -rw-r--r-- | 9harness/src/main.zig | 61 | ||||
| -rw-r--r-- | 9harness/src/tree.zig | 890 | ||||
| -rwxr-xr-x | 9harness/test/e2e.sh | 81 | ||||
| -rwxr-xr-x | 9harness/zmxify | 176 |
6 files changed, 2409 insertions, 13 deletions
diff --git a/9harness/docs/DESIGN.md b/9harness/docs/DESIGN.md index 1b5f33e..4868bc8 100644 --- a/9harness/docs/DESIGN.md +++ b/9harness/docs/DESIGN.md @@ -214,6 +214,214 @@ HOME under a test tmp dir, never the real harness roots. - `zig build programs-test` / `programs-itest` — the umbrella steps, which include 9harness's. +## /active: the derived view (v2) + +v1 mirrors files. `/active` is the first *semantic* layer: one directory +per live agent, normalized across harnesses, so `ls /active` answers +"what is running right now" whatever wrote it. It is the discovery half +of `~/.local/bin/zmxify`, whose resolution ladder it ports. + +``` +/active/ + claude/ + 345104/ + pid 345104 ppid 344980 + started 1790013029 cwd /home/goblin + name goblin-e5 title the session's title + session 1f9a74f7-... model claude-opus-5[1m] + via registry zmx harness (read-write, below) + status busy only where the harness publishes one + transcript the live .jsonl, growing as it writes + agents/ <name>/{model,transcript} + omp/ + 155574/ ... +``` + +A harness directory holds one directory per live process of it, named +by pid. `/proc` spells it that way and so does this: a compound +`claude-345104` would make you parse a name to recover `harness`, which +is already the directory above it, and `/active/omp/*` would not glob. + +There is no `updated` file. The last write to a session is the mtime of +`transcript`, which `stat` already carries; serving it again as its own +file would be the same fact twice, and the copy is the one that goes +stale. + +These are zmxify's picker columns — pid, harness, dir, age, zmx, via, +session, title — as files, which is the point: the script stops +scanning `/proc`, reading fds and querying sqlite itself and just reads +the tree. + +**`status` is a promise only some harnesses make.** It is the harness's +own word for what it is doing, and only Claude Code publishes one +(`busy`, in `sessions/<pid>.json`); for every other harness the file is +simply absent, the way a `skills` mount with no target is absent. +Deriving one from the process's `/proc` state would answer a different +question (`S` means "not on a CPU this instant", not "waiting for +you"). For every other harness the honest activity signal is the mtime +of `transcript`. `ls /active/*/*/status` tells you who publishes the +stronger answer. + +An entry is named `<harness>-<pid>`: unique, stable for the process's +life, and it sorts by harness. The pid is the identity because it is +what `/proc` and every harness's own registry agree on. + +### Finding the agents + +`/proc` is scanned for a process whose `argv[0]` basename is `omp`, +`claude`, `codex`, `hermes` or `dsh`, or a python running +`hermes_cli.main`. A command line carrying one of the daemon words +(`gateway`, `dashboard`, `mcp`, `mcp-server`, `app-server`, +`exec-server`, `serve`, `daemon`, `acp`, `ps`, `render`, `export`, +`__omp_worker_daemon_broker`) is a server or a helper, never a session. +The daemon's own ancestry is excluded, so 9harness can never list or +act on the process tree it lives in. + +Liveness is `/proc/<pid>` existing **and** its `stat` field 22 +(starttime) matching the one remembered for the slot. A pid that has +been reused is a different process and does not list; a corpse never +lists. + +### Resolving the session + +Each harness is asked in its own terms — the `fd -> dir -> db` ladder +zmxify worked out, with the route named in `via` so a wrong guess is +visible rather than silent: + +| harness | route | `via` | +|---|---|---| +| claude | `~/.claude/sessions/<pid>.json`, which the harness maintains itself: sessionId, cwd, name, status, version, and `procStart` as a pid-reuse guard | `registry` | +| omp | the transcript it holds open under `~/.omp/agent/sessions/`, newest first; else the store directory named after its cwd with `/` becoming `-` | `fd`, `dir` | +| dsh | the `session-<uuid>/session.jsonl.zstd` it holds open; zstd, so there is no title | `fd` | +| codex | the rollout it holds open — `sessions/YYYY/MM/DD/rollout-<when>-<uuid>.jsonl`, whose *name* carries the session id, so its sqlite is never opened | `fd` | +| hermes | nothing to resolve: see below | `none` | + +A resolved session is checked against the process's cwd before it is +believed (the head of an omp transcript carries `"cwd"`). A process +whose session does not resolve still appears, with everything `/proc` +knows and an empty `session`: the view never pretends to know what it +does not, and a move on it is refused. + +codex's id is the last 36 bytes of the rollout's name, not what +splitting on `-` gives — the timestamp in front of it holds dashes too. + +**hermes is the one harness with no answer here, and it needs none.** +Its sessions live only in `state.db`, with no per-session file to find; +but the only hermes processes that run are the gateway and the +dashboard, and both are daemon-shaped, so neither is a session anything +should move. zmxify resolved a hermes id from sqlite and then declined +to touch the process holding it, for the same reason. If an interactive +hermes ever exists, this is the gap, and it is the one place sqlite +would buy something. + +### Freshness + +A slot remembers only what identifies the agent: harness, pid, +starttime, cwd, session id and transcript path. Every *field* is +re-derived from disk when it is read, so `status` is never stale and a +transcript grows under `cat`. `/proc` is rescanned on a readdir of +`/active` and on a lookup that misses, not on every read. + +### `zmx`: the write path, and why it is not a ctl + +Writing a zmx session name into `/active/<id>/zmx` moves that agent into +a zmx session of that name: the zmxify action half, as a file. + +Reading `zmx` gives the `ZMX_SESSION` of the process, empty when it runs +outside zmx. Writing sets it. The file means *which zmx session this +agent lives in*, and writing a name makes that true — state, not a verb +channel. A `ctl` taking words would be the ordinary Plan 9 spelling +(`/proc/n/ctl`), and an executable `zmxify` script served in the tree +would be the zmx `attach` spelling, but a script that shells out to a +local binary is a lie over a remote mount: it would run against a +session that is not on the client's machine. A write is served where +the authority is, so it survives being mounted from anywhere. + +What a write does, in order, refusing before it destroys anything: + +1. The name must be zmx's label charset (`[A-Za-z0-9._-]`) and unused by + a live session, else `EEXIST`. +2. The agent's session must have resolved, else `EPERM`. Nothing is + killed that has nowhere to come back to — zmxify's invariant. +3. `SIGTERM`, then `SIGKILL` after 5s, giving up at 12s with `EIO`. +4. `fork`, `chdir` to the agent's cwd, `exec zmx run <name> -d` with a + **fixed argv per harness** (`claude --resume <sid>`, + `omp --resume <transcript>`, `codex resume <sid>`, + `hermes --resume <sid>`). No client byte ever reaches `exec`: the + only thing the client supplies is the session name, and it is + validated first. +5. The write returns once `$XDG_RUNTIME_DIR/9p/zmx/<name>` appears + (zmx self-posts), or `EIO` on timeout. + +The agent comes back under a new pid, so the old `/active/<id>` is gone +and a new entry takes its place with `zmx` reading the new name. The +window between the kill and the exec is the same exposure zmxify has +always had, and it is why step 2 comes first. + +**This is the tree's first write path, and it is an escalation**: a +client that can write this file can kill the user's agents and cause a +process to be spawned. It is therefore off unless `--allow-move` is +given, and the read-only daemon stays the default. Everything else in +the tree still answers `EPERM` to writes. + +### Where this is not Plan 9 + +Named, because they are choices rather than oversights: + +* **`via` describes how the server found out**, not something true of + the process. That is diagnostics in the interface. It stays because a + resolution that guesses wrong silently is worse than a wart — it is + why zmxify's picker showed the route too. +* **A write to `zmx` does not change the object, it replaces it.** The + process is killed and another starts under a new pid, so the entry + written to disappears. `/proc/n/ctl` accepting `kill` is at least + honest about being a verb with a consequence. The file still beats an + executable served in the tree, which would lie outright over a remote + mount, so it stays — with this paragraph. +* **A move blocks the whole daemon** for as long as the kill and the + restart take, because one mutex spans `handle` and the reply. A Plan + 9 server keeps answering other fids meanwhile. The engine already has + parked replies and Tflush for exactly this; the move does not use + them yet, and that is the first thing to revisit. +* **`/active` lives inside the mirror** rather than being its own + posted service. "What files exist" and "what is running" are + different jobs, and a stricter reading would separate them; they + share the pinned roots and the parsing glue, so they share a daemon. + +### Replacing zmxify + +`9harness/zmxify` is the whole script now: it lists `/active`, offers +the rows to fzf, and writes the chosen name into the agent's `zmx`. +It parses no `/proc`, opens no fd table and queries no database — the +307 lines that did become a `cat` of a few files, and the knowledge +they held now lives in a daemon with tests around it. + +What the script still decides is which rows to offer and what to call +the new session. It skips an agent already under zmx — that is what +this moves things *into* — and one whose session did not resolve, which +the daemon would refuse anyway. A free name is the caller's business +too, and the registry already answers which are taken. + +What it does **not** skip is the caller's own session, and that is the +whole point of the daemon owning the action. The old script excluded +its own ancestry because it did the killing itself: kill your own +parent and the script dies before it can re-exec, losing the session it +was rescuing. Now the script's only job is to get the `Twrite` out. +The daemon kills, re-execs and waits for the new session to post, and +it completes all of that whether or not the client is still there — +verified by hanging up immediately after sending the write and watching +the move land anyway. So zmxifying the terminal you are sitting in +works; the shell drops, and `zmx attach <name>` is printed before the +write, because there may be no script left to print it after. + +### Testability + +`--proc DIR` overrides `/proc` the way `--root NAME=PATH` overrides a +harness root, so the scan, the liveness guard, the resolution ladder and +the exclusions are unit-tested against a fixture process tree and never +against the live machine. The move itself is exercised end-to-end +against a fake harness in `test/e2e.sh`. + ## Out of scope for v1 (documented, not hidden) - Write paths (create/write/setattr stay EPERM), resume/attach diff --git a/9harness/src/active.zig b/9harness/src/active.zig new file mode 100644 index 0000000..0a1a60d --- /dev/null +++ b/9harness/src/active.zig @@ -0,0 +1,1006 @@ +//! The derived view behind `/active`: one record per live agent, +//! normalized across harnesses. +//! +//! v1 of the tree mirrors files. This module answers a different +//! question — *what is running right now* — by reading `/proc` and then +//! asking each harness in its own terms which stored session the +//! process is writing. That ladder (`fd` -> `dir` -> `db`, with the +//! route named so a wrong guess is visible) is `~/.local/bin/zmxify`'s, +//! ported here so the knowledge lives somewhere tested. +//! +//! Everything here composes paths from the proc root, a pinned harness +//! root and a pid. No client byte reaches the filesystem, and the proc +//! root is a parameter (`--proc`) so the scan, the liveness guard and +//! the exclusions are testable against a fixture tree. +const std = @import("std"); +const Io = std.Io; + +// ---- comptime bounds --------------------------------------------------------- + +/// Live agents held at once. A machine running more than this many +/// interactive harnesses at the same time is not the case to optimize. +pub const max_live: usize = 32; +/// Subagents listed under one agent. +pub const max_agents: usize = 32; +/// Longest directory a process may run in. +pub const cwd_capacity: usize = 512; +/// Longest absolute path this module composes (proc root, harness root). +pub const path_capacity: usize = 1024; +/// Longest session id, title, name or model answered. +pub const text_capacity: usize = 160; +/// `<harness>-<pid>`. +pub const name_capacity: usize = 32; +/// Bytes of a `/proc` file or a transcript head read at once. +pub const probe_capacity: usize = 8192; + +/// The harnesses this view knows how to recognize. +/// The order is `tree.Root`'s, and `tree.zig` asserts that at comptime: +/// the two enums index the same pinned roots, so a mismatch would hand +/// one harness another's base path. +pub const Kind = enum(u8) { + claude, + codex, + omp, + hermes, + dsh, + + pub fn text(k: Kind) []const u8 { + return @tagName(k); + } +}; + +/// How a session was resolved — reported, so a wrong guess is visible +/// rather than silent. `none` means the process is listed with what +/// `/proc` knows and nothing more (codex and hermes keep their sessions +/// in sqlite, which this round does not read). +pub const Via = enum(u8) { + none, + registry, + fd, + dir, + + pub fn text(v: Via) []const u8 { + return @tagName(v); + } +}; + +/// A fixed-capacity byte buffer. Everything this module remembers is +/// bounded at comptime, like the rest of the daemon. +fn Buf(comptime n: usize) type { + return struct { + bytes: [n]u8 = undefined, + len: u16 = 0, + + const Self = @This(); + + pub fn set(b: *Self, v: []const u8) void { + const take = @min(v.len, n); + @memcpy(b.bytes[0..take], v[0..take]); + b.len = @intCast(take); + } + + pub fn slice(b: *const Self) []const u8 { + return b.bytes[0..b.len]; + } + + pub fn clear(b: *Self) void { + b.len = 0; + } + }; +} + +/// One live agent. +pub const Live = struct { + kind: Kind = .claude, + pid: u32 = 0, + ppid: u32 = 0, + /// `/proc/<pid>/stat` field 22. Two processes can share a pid over + /// time; they cannot share a pid and a start time, so this is what + /// makes a remembered slot safe to trust later. + starttime: u64 = 0, + /// Unix seconds, from the process's start time against boot time. + started: i64 = 0, + via: Via = .none, + name: Buf(name_capacity) = .{}, + cwd: Buf(cwd_capacity) = .{}, + session: Buf(text_capacity) = .{}, + /// Absolute path of the session transcript, empty when unresolved. + transcript: Buf(path_capacity) = .{}, + /// Directory holding the subagent transcripts, empty when the + /// harness has no such layout. + agent_dir: Buf(path_capacity) = .{}, +}; + +// ---- pure helpers: the part that carries zmxify's knowledge ------------------- + +/// A command line word that marks a daemon, a server or a helper — a +/// process wearing the harness's name that is not an interactive +/// session and must never be listed or acted on. +const not_session = [_][]const u8{ + "__omp_worker_daemon_broker", + "acp", + "mcp", + "mcp-server", + "app-server", + "exec-server", + "gateway", + "dashboard", + "serve", + "daemon", + "ps", + "render", + "export", +}; + +/// The last path component of `p`. +pub fn basename(p: []const u8) []const u8 { + if (std.mem.lastIndexOfScalar(u8, p, '/')) |i| return p[i + 1 ..]; + return p; +} + +/// Iterates a `/proc/<pid>/cmdline`: NUL-separated argv, usually with a +/// trailing NUL. An empty cmdline (a kernel thread) yields nothing. +pub const Argv = struct { + rest: []const u8, + + pub fn init(cmdline: []const u8) Argv { + return .{ .rest = cmdline }; + } + + pub fn next(a: *Argv) ?[]const u8 { + while (a.rest.len > 0 and a.rest[0] == 0) a.rest = a.rest[1..]; + if (a.rest.len == 0) return null; + const end = std.mem.indexOfScalar(u8, a.rest, 0) orelse a.rest.len; + const word = a.rest[0..end]; + a.rest = a.rest[@min(end + 1, a.rest.len)..]; + return word; + } +}; + +/// The harness a command line runs, if any. `argv[0]`'s basename names +/// it, except for hermes, which runs as a python module; any word from +/// `not_session` disqualifies the whole command line. +pub fn harnessOf(cmdline: []const u8) ?Kind { + var it: Argv = .init(cmdline); + const argv0 = it.next() orelse return null; + const exe_name = basename(argv0); + + var found: ?Kind = null; + if (std.mem.eql(u8, exe_name, "claude")) { + found = .claude; + } else if (std.mem.eql(u8, exe_name, "omp")) { + found = .omp; + } else if (std.mem.eql(u8, exe_name, "codex")) { + found = .codex; + } else if (std.mem.eql(u8, exe_name, "hermes")) { + found = .hermes; + } else if (std.mem.eql(u8, exe_name, "dsh")) { + found = .dsh; + } else if (std.mem.startsWith(u8, exe_name, "python") or std.mem.startsWith(u8, exe_name, "pypy")) { + var py: Argv = .init(cmdline); + while (py.next()) |w| { + if (std.mem.endsWith(u8, w, "hermes") or std.mem.endsWith(u8, w, "hermes_cli.main")) { + found = .hermes; + break; + } + } + } + if (found == null) return null; + + var words: Argv = .init(cmdline); + while (words.next()) |w| { + for (not_session) |bad| if (std.mem.eql(u8, w, bad)) return null; + } + return found; +} + +/// Field `n` (1-based) of a `/proc/<pid>/stat` line. Field 2 is the +/// executable name in parentheses and may itself hold spaces and +/// parentheses, so the fields are counted from the *last* `)`, never by +/// splitting the whole line. Only fields from 3 on are reachable, which +/// is all any caller wants. +pub fn statField(stat: []const u8, n: usize) ?u64 { + if (n < 3) return null; + const close = std.mem.lastIndexOfScalar(u8, stat, ')') orelse return null; + var rest = stat[close + 1 ..]; + var field: usize = 3; // state, the first field after the name + while (true) { + while (rest.len > 0 and rest[0] == ' ') rest = rest[1..]; + if (rest.len == 0) return null; + const end = std.mem.indexOfScalar(u8, rest, ' ') orelse rest.len; + if (field == n) return std.fmt.parseInt(u64, rest[0..end], 10) catch null; + rest = rest[end..]; + field += 1; + } +} + +/// The process's start time in clock ticks since boot. Two processes can +/// share a pid over time; they cannot share a pid and a start time, so +/// this is what makes a remembered slot safe to trust later. +pub fn startTimeOf(stat: []const u8) ?u64 { + return statField(stat, 22); +} + +/// The parent pid. +pub fn parentOf(stat: []const u8) ?u32 { + const v = statField(stat, 4) orelse return null; + return std.math.cast(u32, v); +} + +/// The value of `"key"` in a small flat JSON object: a quoted string +/// without its quotes, or a bare token (a number, `true`, `null`). +/// Nested objects are not searched, which is what the callers want — +/// these are the harnesses' own one-level session records. +pub fn jsonField(text: []const u8, key: []const u8) ?[]const u8 { + var quoted_buf: [64]u8 = undefined; + if (key.len + 2 > quoted_buf.len) return null; + quoted_buf[0] = '"'; + @memcpy(quoted_buf[1 .. 1 + key.len], key); + quoted_buf[1 + key.len] = '"'; + const quoted = quoted_buf[0 .. key.len + 2]; + + var from: usize = 0; + while (std.mem.indexOfPos(u8, text, from, quoted)) |at| { + from = at + quoted.len; + var rest = text[from..]; + while (rest.len > 0 and (rest[0] == ' ' or rest[0] == '\t')) rest = rest[1..]; + if (rest.len == 0 or rest[0] != ':') continue; + rest = rest[1..]; + while (rest.len > 0 and (rest[0] == ' ' or rest[0] == '\t')) rest = rest[1..]; + if (rest.len == 0) return null; + if (rest[0] == '"') { + const end = std.mem.indexOfScalar(u8, rest[1..], '"') orelse return null; + return rest[1 .. 1 + end]; + } + if (rest[0] == '{' or rest[0] == '[') continue; // not a leaf; keep looking + const end = std.mem.indexOfAny(u8, rest, ",}\n\r \t") orelse rest.len; + if (end == 0) return null; + return rest[0..end]; + } + return null; +} + +/// The session id in a transcript's file name. omp writes +/// `<timestamp>_<uuid>.jsonl`, so the id is what follows the last `_`. +/// codex writes `rollout-<timestamp>-<uuid>.jsonl`, where the +/// timestamp holds dashes too, so the id is the last 36 bytes rather +/// than anything found by splitting. Claude Code writes +/// `<uuid>.jsonl`, all of it the id. +pub fn sessionIdOf(kind: Kind, file_name: []const u8) []const u8 { + var stem = file_name; + if (std.mem.endsWith(u8, stem, ".jsonl")) stem = stem[0 .. stem.len - ".jsonl".len]; + switch (kind) { + .omp => if (std.mem.lastIndexOfScalar(u8, stem, '_')) |i| return stem[i + 1 ..], + .codex => { + const uuid_len = "01a0bcd2-5653-7cc0-aa36-676f6467a72a".len; + if (stem.len >= uuid_len) return stem[stem.len - uuid_len ..]; + }, + else => {}, + } + return stem; +} + +/// The store directory a harness derives from a working directory. +/// omp strips `$HOME` and turns every `/` into `-`; Claude Code turns +/// every character that is not a letter or a digit into `-`, over the +/// whole path. Returns the slug written into `out`. +pub fn slugOf(kind: Kind, cwd: []const u8, home: []const u8, out: []u8) ?[]const u8 { + var path = cwd; + if (kind == .omp and home.len > 0 and std.mem.startsWith(u8, path, home)) { + path = path[home.len..]; + } + if (path.len > out.len) return null; + for (path, 0..) |c, i| { + out[i] = switch (kind) { + .omp => if (c == '/') '-' else c, + else => if (std.ascii.isAlphanumeric(c)) c else '-', + }; + } + return out[0..path.len]; +} + +/// Is `name` a zmx session name? zmx allows only these bytes in a +/// label, and the move path never composes anything else. +pub fn legalZmxName(name: []const u8) bool { + if (name.len == 0 or name.len > 64) return false; + for (name) |c| { + if (!std.ascii.isAlphanumeric(c) and c != '.' and c != '_' and c != '-') return false; + } + return true; +} + + +// ---- reading /proc and the harness stores ------------------------------------ + +/// Where the view reads from. Every path below is composed from these +/// and a pid; nothing here is ever built from a client's bytes. +pub const Sources = struct { + io: Io, + /// The proc filesystem root — `/proc`, or a fixture tree under test. + proc: []const u8, + /// $HOME, for the store-slug spellings. + home: []const u8, + /// The pinned harness roots, indexed by `@intFromEnum(Kind)`; an + /// empty base means that harness is unreachable. + roots: [5][]const u8 = @splat(""), + /// The daemon's own pid: it and its ancestors are never listed, so + /// 9harness can neither show nor act on the tree it lives in. Zero + /// disables the check (fixtures have no ancestry). + self_pid: u32 = 0, + /// Seconds between the epoch and boot, from `/proc/stat`'s `btime`. + /// Zero when it could not be read; `started` is then zero too. + btime: i64 = 0, +}; + +/// USER_HZ: the unit of `/proc/<pid>/stat`'s time fields. Linux reports +/// them in hundredths of a second whatever CONFIG_HZ is. +const user_hz: u64 = 100; + +fn readSmall(io: Io, path: []const u8, buf: []u8) ?[]const u8 { + const out = Io.Dir.readFile(.cwd(), io, path, buf) catch return null; + return out; +} + +fn linkOf(io: Io, path: []const u8, buf: []u8) ?[]const u8 { + const n = Io.Dir.readLinkAbsolute(io, path, buf) catch return null; + return buf[0..n]; +} + +fn statOf(io: Io, path: []const u8) ?Io.Dir.Stat { + return Io.Dir.statFile(.cwd(), io, path, .{ .follow_symlinks = true }) catch null; +} + +fn mtimeOf(io: Io, path: []const u8) i64 { + const st = statOf(io, path) orelse return 0; + return st.mtime.toSeconds(); +} + +/// `parts` joined with `/` into `buf`, or null when they do not fit. +fn join(buf: []u8, parts: []const []const u8) ?[]const u8 { + var n: usize = 0; + for (parts, 0..) |p, i| { + if (i > 0) { + if (n + 1 > buf.len) return null; + buf[n] = '/'; + n += 1; + } + if (n + p.len > buf.len) return null; + @memcpy(buf[n..][0..p.len], p); + n += p.len; + } + return buf[0..n]; +} + +/// `<root>` of a harness, or null when it is not pinned. +fn rootOf(s: Sources, k: Kind) ?[]const u8 { + const base = s.roots[@intFromEnum(k)]; + return if (base.len == 0) null else base; +} + +/// The directory a harness keeps its session transcripts under. +fn sessionRoot(s: Sources, k: Kind, buf: []u8) ?[]const u8 { + const base = rootOf(s, k) orelse return null; + return switch (k) { + // The omp root is ~/.omp and its mirror hangs off agent/. + .omp => join(buf, &.{ base, "agent", "sessions" }), + .claude => join(buf, &.{ base, "projects" }), + .dsh => join(buf, &.{ base, "sessions" }), + // codex names each rollout after its session, so the file it + // holds open is the whole answer — no sqlite needed. + .codex => join(buf, &.{ base, "sessions" }), + // hermes keeps its sessions only in sqlite, and the only hermes + // processes that run are the gateway and the dashboard, which + // are never sessions. Nothing to resolve. + .hermes => null, + }; +} + +// ---- the table ---------------------------------------------------------------- + +pub const Slot = struct { + used: bool = false, + /// Bumped whenever the slot comes to hold a different process, so a + /// node id minted for the old one stops resolving. + gen: u8 = 0, + seen: bool = false, + live: Live = .{}, +}; + +pub const Table = struct { + slots: [max_live]Slot = @splat(.{}), + + /// The slot holding this (pid, starttime), if any. + fn slotOfPid(t: *Table, pid: u32, starttime: u64) ?*Slot { + for (&t.slots) |*sl| { + if (sl.used and sl.live.pid == pid and sl.live.starttime == starttime) return sl; + } + return null; + } + + fn freeSlot(t: *Table) ?*Slot { + for (&t.slots) |*sl| if (!sl.used) return sl; + return null; + } + + /// The live record at `index`, when its generation still matches. + pub fn at(t: *Table, index: usize, gen: u8) ?*Live { + if (index >= t.slots.len) return null; + const sl = &t.slots[index]; + if (!sl.used or sl.gen != gen) return null; + return &sl.live; + } + + pub fn indexOf(t: *Table, l: *const Live) usize { + const base = @intFromPtr(&t.slots[0]); + return (@intFromPtr(l) - @sizeOf(Slot) + @sizeOf(Slot) - @offsetOf(Slot, "live") - base) / @sizeOf(Slot); + } + + /// Rescans `/proc`. Slots keep their index and generation while the + /// same process is still there, so a handle opened before the scan + /// keeps working; a slot that comes to hold a different process + /// bumps its generation and the old node ids stop resolving. + pub fn scan(t: *Table, s: Sources) void { + for (&t.slots) |*sl| sl.seen = false; + + var anc: [64]u32 = @splat(0); + const anc_n = ancestry(s, &anc); + + var buf: [path_capacity]u8 = undefined; + if (Io.Dir.openDirAbsolute(s.io, s.proc, .{ .iterate = true })) |dir| { + defer Io.Dir.close(dir, s.io); + var rb: [Io.Dir.Iterator.reader_buffer_len]u8 align(@alignOf(usize)) = undefined; + var it = Io.Dir.Reader.init(dir, &rb); + while (true) { + const entry = (it.next(s.io) catch break) orelse break; + const pid = std.fmt.parseInt(u32, entry.name, 10) catch continue; + var skip = false; + for (anc[0..anc_n]) |a| if (a == pid) { + skip = true; + }; + if (skip) continue; + t.consider(s, pid, &buf); + } + } else |_| {} + + for (&t.slots) |*sl| { + if (sl.used and !sl.seen) { + sl.used = false; + sl.gen +%= 1; + } + } + } + + /// Looks at one pid and takes it if it is an agent. + fn consider(t: *Table, s: Sources, pid: u32, buf: []u8) void { + var num: [24]u8 = undefined; + const pid_text = std.fmt.bufPrint(&num, "{d}", .{pid}) catch return; + var dir_buf: [path_capacity]u8 = undefined; + const pid_dir = join(&dir_buf, &.{ s.proc, pid_text }) orelse return; + + var stat_buf: [probe_capacity]u8 = undefined; + const stat_path = join(buf, &.{ pid_dir, "stat" }) orelse return; + const stat = readSmall(s.io, stat_path, &stat_buf) orelse return; + const starttime = startTimeOf(stat) orelse return; + + var cmd_buf: [probe_capacity]u8 = undefined; + const cmd_path = join(buf, &.{ pid_dir, "cmdline" }) orelse return; + const cmdline = readSmall(s.io, cmd_path, &cmd_buf) orelse return; + const kind = harnessOf(cmdline) orelse return; + + // Already known and still the same process: keep the slot and + // everything resolved for it. + if (t.slotOfPid(pid, starttime)) |sl| { + sl.seen = true; + return; + } + + const sl = t.freeSlot() orelse return; // full: the rest simply do not list + sl.* = .{ .used = true, .gen = sl.gen +% 1, .seen = true, .live = .{} }; + const l = &sl.live; + l.kind = kind; + l.pid = pid; + l.ppid = parentOf(stat) orelse 0; + l.starttime = starttime; + l.started = if (s.btime == 0) 0 else s.btime + @as(i64, @intCast(starttime / user_hz)); + + var name_buf: [name_capacity]u8 = undefined; + l.name.set(std.fmt.bufPrint(&name_buf, "{s}-{d}", .{ kind.text(), pid }) catch kind.text()); + + const cwd_path = join(buf, &.{ pid_dir, "cwd" }) orelse return; + var cwd_buf: [cwd_capacity]u8 = undefined; + if (linkOf(s.io, cwd_path, &cwd_buf)) |cwd| l.cwd.set(cwd); + + resolveSession(l, s, pid_dir); + } +}; + +/// The daemon's own pid and every parent of it, so the tree it lives in +/// is never listed. Returns how many were written. +fn ancestry(s: Sources, out: []u32) usize { + if (s.self_pid == 0) return 0; + var n: usize = 0; + var pid = s.self_pid; + var buf: [path_capacity]u8 = undefined; + var stat_buf: [probe_capacity]u8 = undefined; + while (pid > 1 and n < out.len) { + out[n] = pid; + n += 1; + var num: [24]u8 = undefined; + const pid_text = std.fmt.bufPrint(&num, "{d}", .{pid}) catch break; + const path = join(&buf, &.{ s.proc, pid_text, "stat" }) orelse break; + const stat = readSmall(s.io, path, &stat_buf) orelse break; + const parent = parentOf(stat) orelse break; + if (parent == 0 or parent == pid) break; + pid = parent; + } + return n; +} + +/// Asks the harness, in its own terms, which stored session this +/// process is writing. Leaves `via = .none` when it cannot tell: the +/// view never invents a session it did not find. +fn resolveSession(l: *Live, s: Sources, pid_dir: []const u8) void { + switch (l.kind) { + .claude => resolveClaude(l, s), + .omp => resolveOmp(l, s, pid_dir), + .dsh, .codex => resolveByFd(l, s, pid_dir), + // hermes has no per-session file to find. + .hermes => {}, + } +} + +/// Claude Code maintains `sessions/<pid>.json` itself, so there is +/// nothing to guess — only to check. `procStart` must match the +/// process's start time, or the record belongs to a dead predecessor +/// that happened to share the pid. +fn resolveClaude(l: *Live, s: Sources) void { + const base = rootOf(s, .claude) orelse return; + var buf: [path_capacity]u8 = undefined; + var num: [24]u8 = undefined; + const file = std.fmt.bufPrint(&num, "{d}.json", .{l.pid}) catch return; + const path = join(&buf, &.{ base, "sessions", file }) orelse return; + var text_buf: [probe_capacity]u8 = undefined; + const text = readSmall(s.io, path, &text_buf) orelse return; + + const proc_start = jsonField(text, "procStart") orelse return; + const claimed = std.fmt.parseInt(u64, proc_start, 10) catch return; + if (claimed != l.starttime) return; // a record left by a reused pid + + const sid = jsonField(text, "sessionId") orelse return; + if (sid.len == 0 or sid.len > text_capacity) return; + l.session.set(sid); + l.via = .registry; + + var slug_buf: [cwd_capacity]u8 = undefined; + const slug = slugOf(.claude, l.cwd.slice(), s.home, &slug_buf) orelse return; + var tr_buf: [path_capacity]u8 = undefined; + var file_buf: [text_capacity + 8]u8 = undefined; + const tr_name = std.fmt.bufPrint(&file_buf, "{s}.jsonl", .{sid}) catch return; + const tr = join(&tr_buf, &.{ base, "projects", slug, tr_name }) orelse return; + if (statOf(s.io, tr) != null) l.transcript.set(tr); +} + +/// omp: the transcript it holds open, else the store directory named +/// after its working directory. +fn resolveOmp(l: *Live, s: Sources, pid_dir: []const u8) void { + resolveByFd(l, s, pid_dir); + if (l.via == .none) resolveOmpByDir(l, s); + if (l.transcript.len == 0) return; + const tr = l.transcript.slice(); + l.session.set(sessionIdOf(.omp, basename(tr))); + // The subagents are sibling files in the directory named after the + // session file, one `<Agent>.jsonl` each. + if (std.mem.endsWith(u8, tr, ".jsonl")) { + l.agent_dir.set(tr[0 .. tr.len - ".jsonl".len]); + } +} + +fn resolveOmpByDir(l: *Live, s: Sources) void { + var root_buf: [path_capacity]u8 = undefined; + const root = sessionRoot(s, .omp, &root_buf) orelse return; + var slug_buf: [cwd_capacity]u8 = undefined; + const slug = slugOf(.omp, l.cwd.slice(), s.home, &slug_buf) orelse return; + var dir_buf: [path_capacity]u8 = undefined; + const dir_path = join(&dir_buf, &.{ root, slug }) orelse return; + var newest_buf: [path_capacity]u8 = undefined; + const newest = newestUnder(s.io, dir_path, ".jsonl", &newest_buf) orelse return; + l.transcript.set(newest); + l.via = .dir; +} + +/// The newest entry of `dir_path` whose name ends in `suffix`, as a +/// full path written into `out`. +fn newestUnder(io: Io, dir_path: []const u8, suffix: []const u8, out: []u8) ?[]const u8 { + const dir = Io.Dir.openDirAbsolute(io, dir_path, .{ .iterate = true }) catch return null; + defer Io.Dir.close(dir, io); + var rb: [Io.Dir.Iterator.reader_buffer_len]u8 align(@alignOf(usize)) = undefined; + var it = Io.Dir.Reader.init(dir, &rb); + var best: i64 = -1; + var best_len: usize = 0; + var probe: [path_capacity]u8 = undefined; + while (true) { + const entry = (it.next(io) catch break) orelse break; + if (!std.mem.endsWith(u8, entry.name, suffix)) continue; + const path = join(&probe, &.{ dir_path, entry.name }) orelse continue; + const when = mtimeOf(io, path); + if (when <= best) continue; + if (path.len > out.len) continue; + @memcpy(out[0..path.len], path); + best = when; + best_len = path.len; + } + return if (best_len == 0) null else out[0..best_len]; +} + +/// The route that needs no knowledge of how a harness names its store: +/// whatever session file the process is holding open. The newest wins, +/// so a harness that keeps an old transcript open for reading does not +/// outvote the one it is writing. +fn resolveByFd(l: *Live, s: Sources, pid_dir: []const u8) void { + var root_buf: [path_capacity]u8 = undefined; + const root = sessionRoot(s, l.kind, &root_buf) orelse return; + + var fd_buf: [path_capacity]u8 = undefined; + const fd_dir_path = join(&fd_buf, &.{ pid_dir, "fd" }) orelse return; + const fd_dir = Io.Dir.openDirAbsolute(s.io, fd_dir_path, .{ .iterate = true }) catch return; + defer Io.Dir.close(fd_dir, s.io); + var rb: [Io.Dir.Iterator.reader_buffer_len]u8 align(@alignOf(usize)) = undefined; + var it = Io.Dir.Reader.init(fd_dir, &rb); + + var best: i64 = -1; + var best_path: [path_capacity]u8 = undefined; + var best_len: usize = 0; + while (true) { + const entry = (it.next(s.io) catch break) orelse break; + var link_path_buf: [path_capacity]u8 = undefined; + const link_path = join(&link_path_buf, &.{ fd_dir_path, entry.name }) orelse continue; + var target_buf: [path_capacity]u8 = undefined; + const target = linkOf(s.io, link_path, &target_buf) orelse continue; + if (!std.mem.startsWith(u8, target, root)) continue; + if (target.len <= root.len or target[root.len] != '/') continue; + if (!sessionFile(l.kind, target[root.len + 1 ..])) continue; + const when = mtimeOf(s.io, target); + if (when <= best or target.len > best_path.len) continue; + @memcpy(best_path[0..target.len], target); + best = when; + best_len = target.len; + } + if (best_len == 0) return; + const found = best_path[0..best_len]; + l.transcript.set(found); + l.via = .fd; + if (l.kind == .codex) l.session.set(sessionIdOf(.codex, basename(found))); + if (l.kind == .dsh) { + // ~/.dsh/sessions/<cwd-slug>/session-<uuid>/session.jsonl.zstd: + // the id is the directory the transcript sits in. + const rel = found[root.len + 1 ..]; + var parts = std.mem.splitScalar(u8, rel, '/'); + _ = parts.next(); + if (parts.next()) |id| l.session.set(id); + } +} + +/// Is this path, relative to the harness's session root, one of its +/// session transcripts rather than a subagent's or something else? +fn sessionFile(kind: Kind, rel: []const u8) bool { + var depth: usize = 0; + var parts = std.mem.splitScalar(u8, rel, '/'); + while (parts.next()) |_| depth += 1; + return switch (kind) { + // <slug>/<file>.jsonl exactly: a third component is a subagent. + .omp => depth == 2 and std.mem.endsWith(u8, rel, ".jsonl"), + .claude => depth == 2 and std.mem.endsWith(u8, rel, ".jsonl"), + .dsh => depth == 3 and std.mem.startsWith(u8, rel, "-"), + // YYYY/MM/DD/rollout-<when>-<uuid>.jsonl + .codex => depth == 4 and std.mem.endsWith(u8, rel, ".jsonl") and + std.mem.startsWith(u8, basename(rel), "rollout-"), + .hermes => false, + }; +} + + +// ---- fields derived when they are read --------------------------------------- + +/// The value of `key` in a NUL-separated environment block, as +/// `/proc/<pid>/environ` stores it. +pub fn envValue(environ: []const u8, key: []const u8) ?[]const u8 { + var it: Argv = .init(environ); + while (it.next()) |pair| { + if (pair.len <= key.len) continue; + if (!std.mem.startsWith(u8, pair, key)) continue; + if (pair[key.len] != '=') continue; + return pair[key.len + 1 ..]; + } + return null; +} + +/// The title key each harness writes into the head of a transcript. +fn titleKey(k: Kind) []const u8 { + return switch (k) { + .claude => "aiTitle", + else => "title", + }; +} + +/// A field from the first few records of a transcript, where the +/// harnesses put the session's title and the model it runs. Only the +/// head is read: these records are written when the session opens. +pub fn headField(io: Io, kind: Kind, transcript: []const u8, which: enum { title, model }, out: []u8) ?[]const u8 { + if (transcript.len == 0) return null; + var head: [probe_capacity]u8 = undefined; + const text = readSmall(io, transcript, &head) orelse return null; + const key = switch (which) { + .title => titleKey(kind), + .model => "model", + }; + const v = jsonField(text, key) orelse return null; + if (v.len == 0 or v.len > out.len) return null; + @memcpy(out[0..v.len], v); + return out[0..v.len]; +} + +/// The zmx session a process runs inside, from its environment. +pub fn zmxOf(s: Sources, pid: u32, out: []u8) ?[]const u8 { + var num: [24]u8 = undefined; + const pid_text = std.fmt.bufPrint(&num, "{d}", .{pid}) catch return null; + var path_buf: [path_capacity]u8 = undefined; + const path = join(&path_buf, &.{ s.proc, pid_text, "environ" }) orelse return null; + var env_buf: [probe_capacity]u8 = undefined; + const env = readSmall(s.io, path, &env_buf) orelse return null; + const v = envValue(env, "ZMX_SESSION") orelse return null; + if (v.len == 0 or v.len > out.len) return null; + @memcpy(out[0..v.len], v); + return out[0..v.len]; +} + +/// A field of Claude Code's own session record, which is the only +/// harness that publishes a name and a status. +pub fn registryField(s: Sources, l: *const Live, key: []const u8, out: []u8) ?[]const u8 { + if (l.kind != .claude) return null; + const base = rootOf(s, .claude) orelse return null; + var num: [24]u8 = undefined; + const file = std.fmt.bufPrint(&num, "{d}.json", .{l.pid}) catch return null; + var path_buf: [path_capacity]u8 = undefined; + const path = join(&path_buf, &.{ base, "sessions", file }) orelse return null; + var text_buf: [probe_capacity]u8 = undefined; + const text = readSmall(s.io, path, &text_buf) orelse return null; + // The record must still describe this process, not a reused pid. + const proc_start = jsonField(text, "procStart") orelse return null; + const claimed = std.fmt.parseInt(u64, proc_start, 10) catch return null; + if (claimed != l.starttime) return null; + const v = jsonField(text, key) orelse return null; + if (v.len == 0 or v.len > out.len) return null; + @memcpy(out[0..v.len], v); + return out[0..v.len]; +} + +/// Is this process still the one the slot remembers? A pid alone is not +/// an identity: the kernel reuses them. +pub fn stillAlive(s: Sources, l: *const Live) bool { + var num: [24]u8 = undefined; + const pid_text = std.fmt.bufPrint(&num, "{d}", .{l.pid}) catch return false; + var path_buf: [path_capacity]u8 = undefined; + const path = join(&path_buf, &.{ s.proc, pid_text, "stat" }) orelse return false; + var stat_buf: [probe_capacity]u8 = undefined; + const stat = readSmall(s.io, path, &stat_buf) orelse return false; + const now = startTimeOf(stat) orelse return false; + return now == l.starttime; +} + +/// Seconds between the epoch and boot, from `/proc/stat`'s `btime` +/// line. Zero when it cannot be read. +pub fn bootTime(io: Io, proc: []const u8) i64 { + var path_buf: [path_capacity]u8 = undefined; + const path = join(&path_buf, &.{ proc, "stat" }) orelse return 0; + var buf: [probe_capacity]u8 = undefined; + const text = readSmall(io, path, &buf) orelse return 0; + var lines = std.mem.splitScalar(u8, text, '\n'); + while (lines.next()) |line| { + if (!std.mem.startsWith(u8, line, "btime ")) continue; + return std.fmt.parseInt(i64, std.mem.trim(u8, line["btime ".len..], " \r"), 10) catch 0; + } + return 0; +} + + +// ---- subagents ---------------------------------------------------------------- + +/// Longest subagent name answered. +pub const agent_name_capacity: usize = 64; + +/// The subagents of one session: omp writes each as `<Name>.jsonl` in a +/// directory named after the session file. Sorted, so a listing that +/// spans several reads keeps its cursor. +pub const Agents = struct { + names: [max_agents][agent_name_capacity]u8 = undefined, + lens: [max_agents]u8 = @splat(0), + count: usize = 0, + + pub fn name(a: *const Agents, i: usize) []const u8 { + if (i >= a.count) return ""; + return a.names[i][0..a.lens[i]]; + } + + pub fn indexOf(a: *const Agents, want: []const u8) ?usize { + for (0..a.count) |i| if (std.mem.eql(u8, a.name(i), want)) return i; + return null; + } +}; + +pub fn agentsOf(io: Io, l: *const Live, out: *Agents) void { + out.count = 0; + const dir_path = l.agent_dir.slice(); + if (dir_path.len == 0) return; + const dir = Io.Dir.openDirAbsolute(io, dir_path, .{ .iterate = true }) catch return; + defer Io.Dir.close(dir, io); + var rb: [Io.Dir.Iterator.reader_buffer_len]u8 align(@alignOf(usize)) = undefined; + var it = Io.Dir.Reader.init(dir, &rb); + while (out.count < max_agents) { + const entry = (it.next(io) catch break) orelse break; + if (!std.mem.endsWith(u8, entry.name, ".jsonl")) continue; + const stem = entry.name[0 .. entry.name.len - ".jsonl".len]; + if (stem.len == 0 or stem.len > agent_name_capacity) continue; + @memcpy(out.names[out.count][0..stem.len], stem); + out.lens[out.count] = @intCast(stem.len); + out.count += 1; + } + // Insertion sort: the list is at most `max_agents` long and already + // nearly ordered, and this keeps the struct allocation-free. + var i: usize = 1; + while (i < out.count) : (i += 1) { + var j = i; + while (j > 0 and std.mem.order(u8, out.name(j), out.name(j - 1)) == .lt) : (j -= 1) { + std.mem.swap([agent_name_capacity]u8, &out.names[j], &out.names[j - 1]); + std.mem.swap(u8, &out.lens[j], &out.lens[j - 1]); + } + } +} + +/// The transcript path of subagent `i`, written into `out`. +pub fn agentPath(l: *const Live, a: *const Agents, i: usize, out: []u8) ?[]const u8 { + if (i >= a.count) return null; + var name_buf: [agent_name_capacity + 8]u8 = undefined; + const file = std.fmt.bufPrint(&name_buf, "{s}.jsonl", .{a.name(i)}) catch return null; + return join(out, &.{ l.agent_dir.slice(), file }); +} + +// ---- unit tests --------------------------------------------------------------- + +const testing = std.testing; + +/// A `/proc/<pid>/cmdline`: NUL-separated argv, written out literally so +/// the tests read the way the kernel stores it. +fn cmd(comptime words: []const []const u8) []const u8 { + comptime var out: []const u8 = ""; + inline for (words) |w| out = out ++ w ++ "\x00"; + return out; +} + +test "active: a command line names its harness, or nothing" { + try testing.expectEqual(Kind.claude, harnessOf(cmd(&.{ "claude", "--dangerously-skip-permissions" })).?); + try testing.expectEqual(Kind.omp, harnessOf(cmd(&.{"omp"})).?); + try testing.expectEqual(Kind.codex, harnessOf(cmd(&.{ "/usr/bin/codex", "resume" })).?); + try testing.expectEqual(Kind.dsh, harnessOf(cmd(&.{"/home/goblin/.local/bin/dsh"})).?); + // hermes runs as a python module, named by a later word. + try testing.expectEqual(Kind.hermes, harnessOf(cmd(&.{ "/venv/bin/python", "-m", "hermes_cli.main" })).?); + // Not a harness at all. + try testing.expect(harnessOf(cmd(&.{ "bash", "-c", "claude" })) == null); + try testing.expect(harnessOf("") == null); + try testing.expect(harnessOf("\x00\x00") == null); +} + +test "active: a daemon or helper wearing a harness name is never a session" { + // The live machine's own gateway and dashboard, which must not list. + try testing.expect(harnessOf(cmd(&.{ "python3", "-m", "hermes_cli.main", "gateway", "run" })) == null); + try testing.expect(harnessOf(cmd(&.{ "python3", "/x/hermes-dashboard", "dashboard" })) == null); + try testing.expect(harnessOf(cmd(&.{ "claude", "mcp-server" })) == null); + try testing.expect(harnessOf(cmd(&.{ "omp", "__omp_worker_daemon_broker" })) == null); + try testing.expect(harnessOf(cmd(&.{ "codex", "app-server" })) == null); + // A directory that merely contains the word is still a session: the + // exclusion matches a whole argument, never a substring. + try testing.expectEqual(Kind.claude, harnessOf(cmd(&.{ "claude", "--cwd", "/home/goblin/serve-me" })).?); +} + +test "active: starttime is counted from the last ')' in a stat line" { + // Fields: pid (comm) state ppid ... starttime is field 22. + var line: [512]u8 = undefined; + var w = Io.Writer.fixed(&line); + try w.print("345104 (claude) S 344980", .{}); + for (5..22) |i| try w.print(" {d}", .{i * 10}); + try w.print(" 1625063 rest of the line", .{}); + try testing.expectEqual(@as(u64, 1625063), startTimeOf(w.buffered()).?); + + // A process whose name holds spaces and parentheses — the reason + // the fields are never counted from the start of the line. + var nasty: [512]u8 = undefined; + var nw = Io.Writer.fixed(&nasty); + try nw.print("42 (we i)rd (name) ) R 1", .{}); + for (5..22) |i| try nw.print(" {d}", .{i}); + try nw.print(" 999 x", .{}); + try testing.expectEqual(@as(u64, 999), startTimeOf(nw.buffered()).?); + + try testing.expect(startTimeOf("no parens here") == null); + try testing.expect(startTimeOf("1 (short) S 2 3") == null); +} + +test "active: json fields come out of a harness's own session record" { + // The shape Claude Code writes to ~/.claude/sessions/<pid>.json. + const rec = + \\{ + \\ "pid": 345104, + \\ "sessionId": "1f9a74f7-b04f-4092-8b98-df11316e03c0", + \\ "cwd": "/home/goblin", + \\ "version": "2.1.278", + \\ "peerFeatures": ["notify_idle", "artifact_yield"], + \\ "name": "goblin-e5", + \\ "status": "busy", + \\ "procStart": "1625063" + \\} + ; + try testing.expectEqualStrings("345104", jsonField(rec, "pid").?); + try testing.expectEqualStrings("1f9a74f7-b04f-4092-8b98-df11316e03c0", jsonField(rec, "sessionId").?); + try testing.expectEqualStrings("/home/goblin", jsonField(rec, "cwd").?); + try testing.expectEqualStrings("goblin-e5", jsonField(rec, "name").?); + try testing.expectEqualStrings("busy", jsonField(rec, "status").?); + try testing.expectEqualStrings("1625063", jsonField(rec, "procStart").?); + // An array or object value is skipped, not half-parsed. + try testing.expect(jsonField(rec, "peerFeatures") == null); + try testing.expect(jsonField(rec, "absent") == null); + // A key that is a prefix of another must not match it. + try testing.expect(jsonField("{\"sessionIdent\":\"x\"}", "sessionId") == null); +} + +test "active: session ids and store slugs follow each harness's spelling" { + try testing.expectEqualStrings( + "01a0c46a-7a06-7676-aef1-add9d9340c83", + sessionIdOf(.omp, "2026-09-21T14-41-47-526Z_01a0c46a-7a06-7676-aef1-add9d9340c83.jsonl"), + ); + try testing.expectEqualStrings( + "42898558-2fa0-41f2-a2c6-1357e5cf6836", + sessionIdOf(.claude, "42898558-2fa0-41f2-a2c6-1357e5cf6836.jsonl"), + ); + // codex's timestamp holds dashes too, so the id is the last 36 + // bytes and never what splitting on `-` would give. + try testing.expectEqualStrings( + "01a0bcd2-5653-7cc0-aa36-676f6467a72a", + sessionIdOf(.codex, "rollout-2026-09-20T00-18-16-01a0bcd2-5653-7cc0-aa36-676f6467a72a.jsonl"), + ); + + var buf: [256]u8 = undefined; + // omp strips $HOME and keeps everything but the slashes. + try testing.expectEqualStrings( + "-00-projects-0x4200.cafe-cloud9", + slugOf(.omp, "/home/goblin/00-projects/0x4200.cafe/cloud9", "/home/goblin", &buf).?, + ); + // Claude Code dashes every byte that is not a letter or a digit. + try testing.expectEqualStrings( + "-home-goblin-00-projects-0x4200-cafe-cloud9", + slugOf(.claude, "/home/goblin/00-projects/0x4200.cafe/cloud9", "/home/goblin", &buf).?, + ); + // A path longer than the buffer is refused, never truncated into a + // slug that would name the wrong store. + var tiny: [4]u8 = undefined; + try testing.expect(slugOf(.omp, "/home/goblin/somewhere", "/home/goblin", &tiny) == null); +} + +test "active: a zmx session name is checked before it is ever used" { + try testing.expect(legalZmxName("harness")); + try testing.expect(legalZmxName("claude-cloud9.2_x")); + try testing.expect(!legalZmxName("")); + try testing.expect(!legalZmxName("has space")); + try testing.expect(!legalZmxName("../escape")); + try testing.expect(!legalZmxName("semi;colon")); + try testing.expect(!legalZmxName("nul\x00byte")); + try testing.expect(!legalZmxName("x" ** 65)); +} + +test "active: an environment block answers by key, not by prefix" { + const env = "PATH=/usr/bin\x00ZMX_SESSION=harness\x00HOME=/home/goblin\x00"; + try testing.expectEqualStrings("harness", envValue(env, "ZMX_SESSION").?); + try testing.expectEqualStrings("/home/goblin", envValue(env, "HOME").?); + try testing.expect(envValue(env, "ZMX") == null); // a prefix is not the key + try testing.expect(envValue(env, "NOPE") == null); + try testing.expect(envValue("", "HOME") == null); + // An entry with no value is a value of zero length, not a miss. + try testing.expectEqualStrings("", envValue("EMPTY=\x00", "EMPTY").?); +} diff --git a/9harness/src/main.zig b/9harness/src/main.zig index 2355a4f..fe0a07b 100644 --- a/9harness/src/main.zig +++ b/9harness/src/main.zig @@ -30,7 +30,8 @@ const linux = std.os.linux; const usage_text = \\usage: 9harness [--unix PATH | --tcp IP:PORT | --fd N] [--no-post] - \\ [--name NAME] [--root NAME=PATH]... + \\ [--name NAME] [--root NAME=PATH]... [--proc DIR] + \\ [--allow-move] [--zmx PATH] \\ \\A read-only, fresh-from-disk 9P2000 view of every harness's state: \\ /pid /uptime daemon facts @@ -38,6 +39,7 @@ const usage_text = \\ /codex/{sessions,session-index,history} \\ /omp /hermes /dsh full mirrors, raw \\ /skills/{claude,codex,omp} the union skills view + \\ /active/<harness>/<pid>/ what is running right now \\ \\By default the daemon posts itself under the name `harness`, so it is \\dialable at $XDG_RUNTIME_DIR/9p/harness and mountable by 9ns --mntgen @@ -46,6 +48,14 @@ const usage_text = \\connected stream on descriptor N and posts nothing. \\--root NAME=PATH pins one harness root (NAME: claude, codex, omp, \\hermes, dsh) somewhere other than $HOME/.<NAME>; repeatable. + \\--proc DIR reads the process tree somewhere other than /proc (for + \\tests); --proc "" leaves /active out of the tree entirely. + \\--allow-move lets a write to /active/<h>/<pid>/zmx move that agent + \\into a zmx session of that name: the daemon kills it and re-execs + \\the harness under zmx with its session resumed. Off by default, + \\because it is the one place the tree is not read-only, and anything + \\that can mount it can then kill an agent. --zmx PATH names the + \\binary a move runs (default: zmx, found on $PATH). \\ \\Run it in a zmx session, the zmx way: `zmx run harness -d 9harness`. \\ @@ -98,6 +108,19 @@ fn serveReq(ctx: ?*anyopaque, conn: *Runner.Conn, req: fs.Req) void { h.mutex.unlock(h.io); } +/// `name` found on `path_env`, as an absolute path in the arena. +fn onPath(io: Io, arena: std.mem.Allocator, path_env: []const u8, name: []const u8) ![]const u8 { + var it = std.mem.splitScalar(u8, path_env, ':'); + while (it.next()) |dir_path| { + if (dir_path.len == 0) continue; + const candidate = try std.fmt.allocPrint(arena, "{s}/{s}", .{ dir_path, name }); + const st = Io.Dir.statFile(.cwd(), io, candidate, .{ .follow_symlinks = true }) catch continue; + if (st.kind != .file) continue; + return candidate; + } + return error.NotFound; +} + // ---- the CLI -------------------------------------------------------------------- const Mode = enum { posted, unix, tcp, fd }; @@ -124,14 +147,20 @@ fn run(init: std.process.Init) !void { var post_name: []const u8 = "harness"; var no_post = false; var base_overrides: [5]?[]const u8 = @splat(null); + var proc_root: []const u8 = "/proc"; + var zmx_path: []const u8 = "zmx"; + var allow_move = false; var i: usize = 1; while (i < args.len) : (i += 1) { const a = args[i]; if (std.mem.eql(u8, a, "--no-post")) { no_post = true; + } else if (std.mem.eql(u8, a, "--allow-move")) { + allow_move = true; } else if (std.mem.eql(u8, a, "--unix") or std.mem.eql(u8, a, "--tcp") or std.mem.eql(u8, a, "--fd") or std.mem.eql(u8, a, "--name") or + std.mem.eql(u8, a, "--proc") or std.mem.eql(u8, a, "--zmx") or std.mem.eql(u8, a, "--root")) { i += 1; @@ -152,6 +181,10 @@ fn run(init: std.process.Init) !void { std.debug.print("9harness: --fd: not a number: {s}\n", .{v}); return error.Usage; }; + } else if (std.mem.eql(u8, a, "--proc")) { + proc_root = v; + } else if (std.mem.eql(u8, a, "--zmx")) { + zmx_path = v; } else if (std.mem.eql(u8, a, "--name")) { post_name = v; if (!post.legalName(post_name)) { @@ -196,7 +229,31 @@ fn run(init: std.process.Init) !void { return error.Usage; } - harness_mem.init(.{ .io = io, .pid = @intCast(linux.getpid()), .bases = bases }); + // A move execs zmx by absolute path: the daemon resolves it once, + // at startup, so nothing about $PATH matters when the write lands. + const zmx_abs = if (std.mem.indexOfScalar(u8, zmx_path, '/') != null) + zmx_path + else + onPath(io, arena, post.getenv(envp, "PATH") orelse "", zmx_path) catch zmx_path; + if (allow_move and std.mem.indexOfScalar(u8, zmx_abs, '/') == null) { + std.debug.print("9harness: --allow-move needs zmx on $PATH (or --zmx PATH)\n", .{}); + return error.Usage; + } + + harness_mem.init(.{ + .io = io, + .pid = @intCast(linux.getpid()), + .bases = bases, + .proc = proc_root, + .home = home orelse "", + .zmx = zmx_abs, + .runtime = post.getenv(envp, "XDG_RUNTIME_DIR") orelse "", + .allow_move = allow_move, + .envp = envp, + }); + if (allow_move) { + std.debug.print("9harness: moves allowed — a write to /active/<h>/<pid>/zmx re-execs that agent under {s}\n", .{zmx_abs}); + } if (no_post and mode == .posted) { std.debug.print("9harness: --no-post needs a listen form (--unix, --tcp or --fd)\n{s}", .{usage_text}); return error.Usage; diff --git a/9harness/src/tree.zig b/9harness/src/tree.zig index 6525e58..b6ff4a4 100644 --- a/9harness/src/tree.zig +++ b/9harness/src/tree.zig @@ -41,6 +41,7 @@ const fs = cloud9.fs; const E = fs.E; const Io = std.Io; const linux = std.os.linux; +pub const active = @import("active.zig"); // ---- comptime bounds --------------------------------------------------------- @@ -84,9 +85,10 @@ pub const Top = enum(u8) { hermes, dsh, skills, + active, /// Directory entries of the root, in listing order. - pub const listed = [_]Top{ .pid, .uptime, .claude, .codex, .omp, .hermes, .dsh, .skills }; + pub const listed = [_]Top{ .pid, .uptime, .claude, .codex, .omp, .hermes, .dsh, .skills, .active }; pub fn fileName(t: Top) []const u8 { return @tagName(t); @@ -94,7 +96,7 @@ pub const Top = enum(u8) { pub fn dir(t: Top) bool { return switch (t) { - .root, .claude, .codex, .omp, .hermes, .dsh, .skills => true, + .root, .claude, .codex, .omp, .hermes, .dsh, .skills, .active => true, .pid, .uptime => false, }; } @@ -171,7 +173,82 @@ fn mountOwner(r: Root, rel_path: []const u8) Top { // ---- node ids and the path table --------------------------------------------- -pub const Kind = enum(u8) { top = 0, path = 1 }; +pub const Kind = enum(u8) { top = 0, path = 1, active = 2 }; + +comptime { + // The derived view indexes the same pinned roots the mirror does. + for (std.enums.values(Root)) |r| { + if (!std.mem.eql(u8, @tagName(r), @tagName(@as(active.Kind, @enumFromInt(@intFromEnum(r)))))) + @compileError("tree.Root and active.Kind must agree, index by index"); + } +} + +/// A node under `/active`. Everything but a harness directory names a +/// slot of the live table and the generation it had when the id was +/// minted, so an id outlives its process only long enough to answer +/// ENOENT. +pub const AFile = enum(u8) { + /// `/active/<harness>` + harness_dir, + /// `/active/<harness>/<pid>` + entry, + pid, + ppid, + started, + cwd, + name, + title, + session, + model, + via, + zmx, + status, + transcript, + /// `/active/<harness>/<pid>/agents` + agents, + /// `/active/<harness>/<pid>/agents/<name>` + agent, + agent_model, + agent_transcript, + + /// The fields of one entry, in listing order. A field with no value + /// is not listed and does not resolve. + pub const entry_files = [_]AFile{ + .pid, .ppid, .started, .cwd, .name, .title, + .session, .model, .via, .zmx, .status, .transcript, + }; + pub const agent_files = [_]AFile{ .agent_model, .agent_transcript }; + + pub fn fileName(f: AFile) []const u8 { + return switch (f) { + .agent_model => "model", + .agent_transcript => "transcript", + else => @tagName(f), + }; + } +}; + +/// A resolved `/active` node. +pub const Act = struct { + file: AFile, + /// Only for `harness_dir`. + kind: active.Kind = .claude, + slot: u8 = 0, + gen: u8 = 0, + agent: u8 = 0, +}; + +pub fn activeNode(a: Act) u64 { + const serial: u48 = switch (a.file) { + .harness_dir => @intFromEnum(a.kind), + else => @as(u48, a.slot) | (@as(u48, a.gen) << 8) | (@as(u48, a.agent) << 16), + }; + return @bitCast(Node{ + .idx = @intFromEnum(a.file), + .kind = @intFromEnum(Kind.active), + .serial = serial, + }); +} pub const Node = packed struct(u64) { /// A Top index (kind == .top) or a Root index (kind == .path). @@ -306,15 +383,71 @@ pub const Harness = struct { list_dirs: [list_capacity]bool = undefined, list_offs: [list_capacity]u32 = undefined, + // The derived view (`/active`). An empty `proc` base turns it off, + // which is the default: a test must opt in to reading a process + // tree, and never the live one. + live: active.Table = .{}, + proc_buf: [base_capacity]u8 = @splat(0), + proc_len: u16 = 0, + home_buf: [base_capacity]u8 = @splat(0), + home_len: u16 = 0, + zmx_buf: [base_capacity]u8 = @splat(0), + zmx_len: u16 = 0, + runtime_buf: [base_capacity]u8 = @splat(0), + runtime_len: u16 = 0, + envp: ?cloud9.post.Env = null, + self_pid: u32 = 0, + btime: i64 = 0, + allow_move: bool = false, + act_text: [1024]u8 = undefined, + act_name: [64]u8 = undefined, + act_path: [active.path_capacity]u8 = undefined, + pub const InitOptions = struct { io: Io, pid: u32, /// Base path of each root, or "" when the root is not pinned. bases: [5][]const u8, + /// The proc filesystem `/active` reads. Empty — the default — + /// leaves the derived view out of the tree entirely. + proc: []const u8 = "", + /// $HOME, for the harnesses' store-slug spellings. + home: []const u8 = "", + /// The zmx binary a move execs, resolved to an absolute path by + /// the caller. Never client bytes. + zmx: []const u8 = "zmx", + /// $XDG_RUNTIME_DIR, where a move looks for the zmx session it + /// is about to create. + runtime: []const u8 = "", + /// The daemon's own environment block, handed to the harness a + /// move re-execs. Null leaves a moved agent with an empty one. + envp: ?cloud9.post.Env = null, + /// Whether writing `/active/<h>/<pid>/zmx` may move a session. + allow_move: bool = false, }; pub fn init(h: *Harness, o: InitOptions) void { h.* = .{ .io = o.io, .pid = o.pid, .started_sec = Io.Timestamp.now(o.io, .real).toSeconds() }; + if (o.proc.len > 0 and o.proc.len <= base_capacity) { + @memcpy(h.proc_buf[0..o.proc.len], o.proc); + h.proc_len = @intCast(o.proc.len); + h.self_pid = o.pid; + h.btime = active.bootTime(o.io, o.proc); + } + if (o.home.len <= base_capacity) { + @memcpy(h.home_buf[0..o.home.len], o.home); + h.home_len = @intCast(o.home.len); + } + if (o.zmx.len > 0 and o.zmx.len <= base_capacity) { + @memcpy(h.zmx_buf[0..o.zmx.len], o.zmx); + h.zmx_len = @intCast(o.zmx.len); + } + if (o.runtime.len <= base_capacity) { + @memcpy(h.runtime_buf[0..o.runtime.len], o.runtime); + h.runtime_len = @intCast(o.runtime.len); + } + h.allow_move = o.allow_move; + h.envp = o.envp; inline for (0..5) |i| { const path = o.bases[i]; h.base_set[i] = path.len > 0 and path.len <= base_capacity; @@ -388,6 +521,18 @@ pub const Harness = struct { if (serialOf(slot, e.gen) != n.serial) return null; return .{ .path = e }; }, + .active => { + const f = std.enums.fromInt(AFile, n.idx) orelse return null; + if (f == .harness_dir) { + const k = std.enums.fromInt(active.Kind, @as(u8, @truncate(n.serial))) orelse return null; + return .{ .act = .{ .file = f, .kind = k } }; + } + const slot: u8 = @truncate(n.serial); + const gen: u8 = @truncate(n.serial >> 8); + const agent: u8 = @truncate(n.serial >> 16); + const l = h.live.at(slot, gen) orelse return null; + return .{ .act = .{ .file = f, .kind = l.kind, .slot = slot, .gen = gen, .agent = agent } }; + }, } } @@ -399,6 +544,7 @@ pub const Harness = struct { pub const Target = union(enum) { top: Top, path: *Entry, + act: Act, }; // ---- attributes --------------------------------------------------------------- @@ -508,6 +654,7 @@ fn attrFor(h: *Harness, t: Target) ?fs.Attr { .top => |top| { switch (top) { .root => return .{ .name = "/", .node = root, .dir = true, .mode = 0o555 }, + .active => return .{ .name = "active", .node = topNode(.active), .dir = true, .mode = 0o555 }, .pid, .uptime => return .{ .name = top.fileName(), .node = topNode(top), @@ -530,6 +677,7 @@ fn attrFor(h: *Harness, t: Target) ?fs.Attr { } }, .path => |e| return statAttr(h, e), + .act => |a| return activeAttr(h, a), } } @@ -544,6 +692,442 @@ fn attrOfMount(h: *Harness, r: Root, rel_path: []const u8) ?fs.Attr { return statAttr(h, e); } +// ---- /active: the derived view ------------------------------------------------- + +fn sourcesOf(h: *Harness) active.Sources { + var roots: [5][]const u8 = @splat(""); + for (0..5) |i| { + if (h.base_set[i]) roots[i] = h.base_buf[i][0..h.base_len[i]]; + } + return .{ + .io = h.io, + .proc = h.proc_buf[0..h.proc_len], + .home = h.home_buf[0..h.home_len], + .roots = roots, + .self_pid = h.self_pid, + .btime = h.btime, + }; +} + +/// Rescans the process tree. Done on a listing of `/active` and on a +/// lookup that misses, which is what makes an agent visible the moment +/// it starts and gone the moment it exits. +fn rescan(h: *Harness) void { + if (h.proc_len == 0) return; + h.live.scan(sourcesOf(h)); +} + +fn hasLive(h: *Harness, k: active.Kind) bool { + for (&h.live.slots) |*sl| { + if (sl.used and sl.live.kind == k) return true; + } + return false; +} + +/// `<text>\n` staged where an answer can point at it. +fn line(h: *Harness, text: []const u8) ?[]const u8 { + return std.fmt.bufPrint(&h.act_text, "{s}\n", .{text}) catch null; +} + +/// The value of a field file, or null when this agent has none — an +/// absent field is not listed and does not resolve, so the tree never +/// answers a blank where it does not know. +fn activeText(h: *Harness, a: Act) ?[]const u8 { + const l = h.live.at(a.slot, a.gen) orelse return null; + const src = sourcesOf(h); + var scratch: [active.text_capacity]u8 = undefined; + return switch (a.file) { + .pid => std.fmt.bufPrint(&h.act_text, "{d}\n", .{l.pid}) catch null, + .ppid => if (l.ppid == 0) null else std.fmt.bufPrint(&h.act_text, "{d}\n", .{l.ppid}) catch null, + .started => if (l.started == 0) null else std.fmt.bufPrint(&h.act_text, "{d}\n", .{l.started}) catch null, + .cwd => if (l.cwd.len == 0) null else line(h, l.cwd.slice()), + .session => if (l.session.len == 0) null else line(h, l.session.slice()), + .via => line(h, l.via.text()), + .name => line(h, active.registryField(src, l, "name", &scratch) orelse return null), + .status => line(h, active.registryField(src, l, "status", &scratch) orelse return null), + .title => line(h, active.headField(h.io, l.kind, l.transcript.slice(), .title, &scratch) orelse return null), + .model => line(h, active.headField(h.io, l.kind, l.transcript.slice(), .model, &scratch) orelse return null), + .agent_model => blk: { + const path = activePath(h, a) orelse break :blk null; + break :blk line(h, active.headField(h.io, l.kind, path, .model, &scratch) orelse return null); + }, + // `zmx` is always there, so it can always be written to; empty + // means the agent runs outside zmx. + .zmx => if (active.zmxOf(src, l.pid, &scratch)) |v| line(h, v) else h.act_text[0..0], + else => null, + }; +} + +/// The absolute path behind a transcript-shaped file. +fn activePath(h: *Harness, a: Act) ?[]const u8 { + const l = h.live.at(a.slot, a.gen) orelse return null; + switch (a.file) { + .transcript => return if (l.transcript.len == 0) null else l.transcript.slice(), + .agent_transcript, .agent_model => { + var ag: active.Agents = .{}; + active.agentsOf(h.io, l, &ag); + return active.agentPath(l, &ag, a.agent, &h.act_path); + }, + else => return null, + } +} + +/// Opens an absolute path without following a symlink at the last +/// component. These paths are derived from a pinned root or from the +/// fd the harness itself holds, never from client bytes, but a name +/// swapped underneath must still fail rather than redirect. +fn openAbsNoFollow(path: []const u8) ?i32 { + if (path.len == 0 or path.len >= active.path_capacity) return null; + var z: [active.path_capacity]u8 = @splat(0); + @memcpy(z[0..path.len], path); + const rc = linux.open(@ptrCast(&z), .{ + .ACCMODE = .RDONLY, + .NOFOLLOW = true, + .CLOEXEC = true, + .NONBLOCK = true, + }, 0); + if (linux.errno(rc) != .SUCCESS) return null; + return @intCast(rc); +} + +fn activeAttr(h: *Harness, a: Act) ?fs.Attr { + switch (a.file) { + .harness_dir => { + if (!hasLive(h, a.kind)) return null; + return .{ .name = a.kind.text(), .node = activeNode(a), .dir = true, .mode = 0o555 }; + }, + .entry => { + const l = h.live.at(a.slot, a.gen) orelse return null; + const nm = std.fmt.bufPrint(&h.act_name, "{d}", .{l.pid}) catch return null; + return .{ .name = nm, .node = activeNode(a), .dir = true, .mode = 0o555 }; + }, + .agents => { + const l = h.live.at(a.slot, a.gen) orelse return null; + if (l.agent_dir.len == 0) return null; + return .{ .name = "agents", .node = activeNode(a), .dir = true, .mode = 0o555 }; + }, + .agent => { + const l = h.live.at(a.slot, a.gen) orelse return null; + var ag: active.Agents = .{}; + active.agentsOf(h.io, l, &ag); + if (a.agent >= ag.count) return null; + const nm = ag.name(a.agent); + @memcpy(h.act_name[0..nm.len], nm); + return .{ .name = h.act_name[0..nm.len], .node = activeNode(a), .dir = true, .mode = 0o555 }; + }, + .transcript, .agent_transcript => { + const path = activePath(h, a) orelse return null; + const st = Io.Dir.statFile(.cwd(), h.io, path, .{ .follow_symlinks = false }) catch return null; + if (st.kind != .file) return null; + return .{ + .name = a.file.fileName(), + .node = activeNode(a), + .size = st.size, + .mode = 0o444, + .mtime = @truncate(@as(u64, @bitCast(st.mtime.toSeconds()))), + }; + }, + else => { + const text = activeText(h, a) orelse return null; + // Only `zmx` is ever writable, and only when a move is allowed. + const mode: u16 = if (a.file == .zmx and h.allow_move) 0o644 else 0o444; + return .{ .name = a.file.fileName(), .node = activeNode(a), .size = text.len, .mode = mode }; + }, + } +} + +/// The slot holding `pid` for this harness, if any. +fn slotOfPid(h: *Harness, k: active.Kind, pid: u32) ?Act { + for (&h.live.slots, 0..) |*sl, i| { + if (!sl.used or sl.live.kind != k or sl.live.pid != pid) continue; + return .{ .file = .entry, .kind = k, .slot = @intCast(i), .gen = sl.gen }; + } + return null; +} + +fn activeLookup(h: *Harness, req: fs.Req, a: Act, name: []const u8) Answer { + switch (a.file) { + .harness_dir => { + const pid = std.fmt.parseInt(u32, name, 10) catch return fail(req.tag, E.NOENT); + var found = slotOfPid(h, a.kind, pid); + if (found == null) { + rescan(h); + found = slotOfPid(h, a.kind, pid); + } + const act = found orelse return fail(req.tag, E.NOENT); + return activeReply(h, req, act, name); + }, + .entry => { + if (std.mem.eql(u8, name, "agents")) { + return activeReply(h, req, .{ .file = .agents, .kind = a.kind, .slot = a.slot, .gen = a.gen }, name); + } + for (AFile.entry_files) |f| { + if (!std.mem.eql(u8, f.fileName(), name)) continue; + return activeReply(h, req, .{ .file = f, .kind = a.kind, .slot = a.slot, .gen = a.gen }, name); + } + return fail(req.tag, E.NOENT); + }, + .agents => { + const l = h.live.at(a.slot, a.gen) orelse return fail(req.tag, E.NOENT); + var ag: active.Agents = .{}; + active.agentsOf(h.io, l, &ag); + const i = ag.indexOf(name) orelse return fail(req.tag, E.NOENT); + return activeReply(h, req, .{ + .file = .agent, + .kind = a.kind, + .slot = a.slot, + .gen = a.gen, + .agent = @intCast(i), + }, name); + }, + .agent => { + for (AFile.agent_files) |f| { + if (!std.mem.eql(u8, f.fileName(), name)) continue; + return activeReply(h, req, .{ + .file = f, + .kind = a.kind, + .slot = a.slot, + .gen = a.gen, + .agent = a.agent, + }, name); + } + return fail(req.tag, E.NOENT); + }, + else => return fail(req.tag, E.NOTDIR), + } +} + +fn activeReply(h: *Harness, req: fs.Req, a: Act, name: []const u8) Answer { + const attr = activeAttr(h, a) orelse return fail(req.tag, E.NOENT); + var with_name = attr; + with_name.name = name; + return .{ .reply = .{ .tag = req.tag, .attr = with_name } }; +} + +fn activeReaddir(h: *Harness, req: fs.Req, a: Act) Answer { + var st: Staging = .{ .buf = &h.stage_buf, .skip = req.off }; + switch (a.file) { + .harness_dir => { + for (&h.live.slots, 0..) |*sl, i| { + if (!sl.used or sl.live.kind != a.kind) continue; + var num: [24]u8 = undefined; + const nm = std.fmt.bufPrint(&num, "{d}", .{sl.live.pid}) catch continue; + st.add(activeNode(.{ + .file = .entry, + .kind = a.kind, + .slot = @intCast(i), + .gen = sl.gen, + }), true, nm); + } + }, + .entry => { + const l = h.live.at(a.slot, a.gen) orelse return fail(req.tag, E.NOENT); + for (AFile.entry_files) |f| { + const child: Act = .{ .file = f, .kind = a.kind, .slot = a.slot, .gen = a.gen }; + if (activeAttr(h, child) == null) continue; // no value: not listed + st.add(activeNode(child), false, f.fileName()); + } + if (l.agent_dir.len > 0) { + st.add(activeNode(.{ .file = .agents, .kind = a.kind, .slot = a.slot, .gen = a.gen }), true, "agents"); + } + }, + .agents => { + const l = h.live.at(a.slot, a.gen) orelse return fail(req.tag, E.NOENT); + var ag: active.Agents = .{}; + active.agentsOf(h.io, l, &ag); + for (0..ag.count) |i| { + st.add(activeNode(.{ + .file = .agent, + .kind = a.kind, + .slot = a.slot, + .gen = a.gen, + .agent = @intCast(i), + }), true, ag.name(i)); + } + }, + .agent => { + for (AFile.agent_files) |f| { + const child: Act = .{ .file = f, .kind = a.kind, .slot = a.slot, .gen = a.gen, .agent = a.agent }; + if (activeAttr(h, child) == null) continue; + st.add(activeNode(child), false, f.fileName()); + } + }, + else => return fail(req.tag, E.NOTDIR), + } + return .{ .reply = .{ .tag = req.tag }, .bytes = h.stage_buf[0..st.len] }; +} + +fn activeRead(h: *Harness, req: fs.Req, a: Act) Answer { + switch (a.file) { + .harness_dir, .entry, .agents, .agent => return fail(req.tag, E.ISDIR), + .transcript, .agent_transcript => { + const path = activePath(h, a) orelse return fail(req.tag, E.NOENT); + const fd = openAbsNoFollow(path) orelse return fail(req.tag, E.NOENT); + const file: Io.File = .{ .handle = fd, .flags = .{ .nonblocking = false } }; + defer file.close(h.io); + const st = file.stat(h.io) catch return fail(req.tag, E.IO); + if (st.kind == .directory) return fail(req.tag, E.ISDIR); + if (st.kind != .file) return fail(req.tag, E.PERM); + _ = linux.fcntl(fd, linux.F.SETFL, 0); + const want = @min(req.size, h.data_buf.len); + const n = file.readPositionalAll(h.io, h.data_buf[0..want], req.off) catch + return fail(req.tag, E.IO); + return .{ .reply = .{ .tag = req.tag }, .bytes = h.data_buf[0..n] }; + }, + else => { + const text = activeText(h, a) orelse return fail(req.tag, E.NOENT); + return window(req, text); + }, + } +} + +// ---- the move: writing a zmx session name into an agent's `zmx` ---------------- + +/// How long a move waits, in 100ms ticks: for the agent to take the +/// hangup, then to die outright, then for the new zmx session to post. +const term_ticks: usize = 50; +const kill_ticks: usize = 30; +const post_ticks: usize = 100; + +fn napOneTick(h: *Harness) void { + h.io.sleep(.fromMilliseconds(100), .awake) catch {}; +} + +/// Is this name already a live zmx session? zmx posts each session into +/// the registry, so the registry is the answer — no process scanning. +fn zmxPosted(h: *Harness, name: []const u8) bool { + if (h.runtime_len == 0) return false; + var buf: [active.path_capacity]u8 = undefined; + const path = std.fmt.bufPrint(&buf, "{s}/9p/zmx/{s}", .{ h.runtime_buf[0..h.runtime_len], name }) catch return true; + return Io.Dir.statFile(.cwd(), h.io, path, .{ .follow_symlinks = true }) != error.FileNotFound; +} + +fn procGone(h: *Harness, pid: u32) bool { + var buf: [active.path_capacity]u8 = undefined; + const path = std.fmt.bufPrint(&buf, "{s}/{d}", .{ h.proc_buf[0..h.proc_len], pid }) catch return false; + _ = Io.Dir.statFile(.cwd(), h.io, path, .{ .follow_symlinks = true }) catch return true; + return false; +} + +/// SIGTERM, then SIGKILL, then give up. The agent's session was +/// resolved before this ran, so whatever happens it has somewhere to +/// come back to. +fn killAndWait(h: *Harness, pid: u32) bool { + _ = linux.kill(@intCast(pid), .TERM); + for (0..term_ticks) |_| { + if (procGone(h, pid)) return true; + napOneTick(h); + } + _ = linux.kill(@intCast(pid), .KILL); + for (0..kill_ticks) |_| { + if (procGone(h, pid)) return true; + napOneTick(h); + } + return false; +} + +/// `zmx run <name> -d <harness> <resume flag> <resume value>`, in the +/// agent's own directory. Every word but `<name>` is fixed by the +/// harness; `<name>` was checked against zmx's label charset before +/// anything was killed. No byte a client wrote reaches `exec` as a +/// command. +fn spawnZmx(h: *Harness, l: *const active.Live, name: []const u8, resume_value: []const u8) bool { + if (h.zmx_len == 0) return false; + const harness_argv0: []const u8 = @tagName(l.kind); + const resume_flag: []const u8 = switch (l.kind) { + .codex => "resume", + else => "--resume", + }; + + // One buffer holds every NUL-terminated word; `argv` points into it. + var words: [8 * active.path_capacity]u8 = undefined; + var used: usize = 0; + var argv: [9]?[*:0]const u8 = @splat(null); + var n: usize = 0; + const parts = [_][]const u8{ + h.zmx_buf[0..h.zmx_len], "run", name, "-d", + harness_argv0, resume_flag, resume_value, + }; + for (parts) |w| { + if (used + w.len + 1 > words.len or n + 1 >= argv.len) return false; + @memcpy(words[used..][0..w.len], w); + words[used + w.len] = 0; + argv[n] = @ptrCast(&words[used]); + n += 1; + used += w.len + 1; + } + + var cwd_z: [active.cwd_capacity + 1]u8 = @splat(0); + if (l.cwd.len >= cwd_z.len) return false; + @memcpy(cwd_z[0..l.cwd.len], l.cwd.slice()); + + const rc = linux.fork(); + if (linux.errno(rc) != .SUCCESS) return false; + if (rc == 0) { + // The child: only async-signal-safe calls from here. + _ = linux.chdir(@ptrCast(&cwd_z)); + const empty = [_:null]?[*:0]const u8{}; + const envp: cloud9.post.Env = h.envp orelse ∅ + _ = linux.execve(argv[0].?, @ptrCast(&argv), envp); + linux.exit(127); + } + // Reap the forked `zmx`, which returns as soon as the session is up. + const child: i32 = @intCast(rc); + var status: u32 = 0; + for (0..post_ticks) |_| { + const w = linux.waitpid(child, &status, 1); // WNOHANG + if (w == @as(usize, @intCast(child))) break; + napOneTick(h); + } + return true; +} + +/// A move: kill the agent and bring it back inside a zmx session of the +/// name written. Refusals come before anything is destroyed. +fn moveToZmx(h: *Harness, req: fs.Req, a: Act) Answer { + if (!h.allow_move) return fail(req.tag, E.PERM); + const name = std.mem.trim(u8, req.data, " \t\r\n"); + if (!active.legalZmxName(name)) return fail(req.tag, E.INVAL); + + const l = h.live.at(a.slot, a.gen) orelse return fail(req.tag, E.NOENT); + // Nothing is killed that has nowhere to come back to. + if (l.via == .none or l.session.len == 0) return fail(req.tag, E.PERM); + if (l.cwd.len == 0) return fail(req.tag, E.PERM); + const resume_value: []const u8 = switch (l.kind) { + // omp resumes by the transcript it wrote, the others by id. + .omp => l.transcript.slice(), + .claude, .codex, .hermes => l.session.slice(), + // dsh has no resume form worth guessing at. + .dsh => return fail(req.tag, E.PERM), + }; + if (resume_value.len == 0) return fail(req.tag, E.PERM); + + // Already there: setting a value it already has changes nothing. + var have: [active.text_capacity]u8 = undefined; + if (active.zmxOf(sourcesOf(h), l.pid, &have)) |current| { + if (std.mem.eql(u8, current, name)) { + return .{ .reply = .{ .tag = req.tag, .written = @intCast(req.data.len) } }; + } + } + if (zmxPosted(h, name)) return fail(req.tag, E.EXIST); + // The slot could have gone stale between the scan and this write. + if (!active.stillAlive(sourcesOf(h), l)) return fail(req.tag, E.NOENT); + + // Everything below this line destroys something. + var snapshot = l.*; + if (!killAndWait(h, snapshot.pid)) return fail(req.tag, E.IO); + if (!spawnZmx(h, &snapshot, name, resume_value)) return fail(req.tag, E.IO); + for (0..post_ticks) |_| { + if (zmxPosted(h, name)) { + rescan(h); + return .{ .reply = .{ .tag = req.tag, .written = @intCast(req.data.len) } }; + } + napOneTick(h); + } + rescan(h); + return fail(req.tag, E.IO); +} + // ---- dispatch ------------------------------------------------------------------ /// Answers one engine request. The caller holds `h.mutex` and replies @@ -553,15 +1137,39 @@ pub fn handle(h: *Harness, req: fs.Req) Answer { return switch (req.op) { .lookup => lookup(h, req, t), .getattr => attrReply(h, req.tag, t), - .setattr => fail(req.tag, E.PERM), + .setattr => setattrReq(h, req, t), .open => open(h, req, t), .release => release(h, req, t), .readdir => readdir(h, req, t), .read => read(h, req, t), - .write => fail(req.tag, E.PERM), + .write => writeReq(h, req, t), }; } +/// Truncation is how a shell's `>` opens a file before writing it. The +/// `zmx` file has no length of its own — a write replaces the value — +/// so on the one writable file a zero-length truncate is a no-op +/// rather than a refusal. Everything else still answers EPERM. +fn setattrReq(h: *Harness, req: fs.Req, t: Target) Answer { + switch (t) { + .act => |a| if (a.file == .zmx and h.allow_move and req.truncate) { + return .{ .reply = .{ .tag = req.tag } }; + }, + else => {}, + } + return fail(req.tag, E.PERM); +} + +/// The tree answers EPERM to every write but one: a zmx session name +/// into a live agent's `zmx`, which moves it there. +fn writeReq(h: *Harness, req: fs.Req, t: Target) Answer { + switch (t) { + .act => |a| if (a.file == .zmx) return moveToZmx(h, req, a), + else => {}, + } + return fail(req.tag, E.PERM); +} + fn attrReply(h: *Harness, tag: u64, t: Target) Answer { const a = attrFor(h, t) orelse return fail(tag, E.NOENT); return .{ .reply = .{ .tag = tag, .attr = a } }; @@ -586,7 +1194,7 @@ fn lookup(h: *Harness, req: fs.Req, t: Target) Answer { // The engine asks for "." when cloning a fid; the reference is // paid for path targets here like any other lookup result. switch (t) { - .top => return attrReply(h, req.tag, t), + .top, .act => return attrReply(h, req.tag, t), .path => |e| { e.refs += 1; return lookupAttr(h, req, t, name); @@ -618,6 +1226,16 @@ fn lookup(h: *Harness, req: fs.Req, t: Target) Answer { const m = top.mirror().?; return lookupBelow(h, req, m.root, m.rel, name); }, + .active => { + const k = blk: for (std.enums.values(active.Kind)) |k| { + if (std.mem.eql(u8, k.text(), name)) break :blk k; + } else return fail(req.tag, E.NOENT); + if (!hasLive(h, k)) { + rescan(h); + if (!hasLive(h, k)) return fail(req.tag, E.NOENT); + } + return activeReply(h, req, .{ .file = .harness_dir, .kind = k }, name); + }, .pid, .uptime => return fail(req.tag, E.NOTDIR), }, .path => |e| { @@ -627,6 +1245,7 @@ fn lookup(h: *Harness, req: fs.Req, t: Target) Answer { } else return fail(req.tag, E.NOENT); return lookupBelow(h, req, e.root, rel_path, name); }, + .act => |a| return activeLookup(h, req, a, name), } } @@ -673,6 +1292,21 @@ fn lookupParent(h: *Harness, req: fs.Req, t: Target) Answer { } break :blk .{ .top = mountOwner(e.root, rel_path) }; }, + .act => |a| switch (a.file) { + .harness_dir => .{ .top = .active }, + .entry => .{ .act = .{ .file = .harness_dir, .kind = a.kind } }, + .agents => .{ .act = .{ .file = .entry, .kind = a.kind, .slot = a.slot, .gen = a.gen } }, + .agent => .{ .act = .{ .file = .agents, .kind = a.kind, .slot = a.slot, .gen = a.gen } }, + .agent_model, .agent_transcript => .{ .act = .{ + .file = .agent, + .kind = a.kind, + .slot = a.slot, + .gen = a.gen, + .agent = a.agent, + } }, + // Every other file hangs directly off its entry. + else => .{ .act = .{ .file = .entry, .kind = a.kind, .slot = a.slot, .gen = a.gen } }, + }, }; return attrReply(h, req.tag, parent); } @@ -687,17 +1321,26 @@ fn unref(e: *Entry) void { fn open(h: *Harness, req: fs.Req, t: Target) Answer { _ = attrFor(h, t) orelse return fail(req.tag, E.NOENT); // still there? - // Read-only tree: any open that would write or truncate is refused. + // Read-only tree, with exactly one exception: the `zmx` file of a + // live agent, and only when the daemon was started to allow moves. + // Truncation is meaningless there (a write replaces the value), so + // it is accepted rather than refused, which is what `>` needs. + const writable = switch (t) { + .act => |a| a.file == .zmx and h.allow_move, + else => false, + }; const rw = req.omode & 3; - if (rw == cloud9.owrite or rw == cloud9.ordwr) return fail(req.tag, E.PERM); - if (req.omode & cloud9.otrunc != 0) return fail(req.tag, E.PERM); + if (!writable) { + if (rw == cloud9.owrite or rw == cloud9.ordwr) return fail(req.tag, E.PERM); + if (req.omode & cloud9.otrunc != 0) return fail(req.tag, E.PERM); + } return .{ .reply = .{ .tag = req.tag, .handle = 1 } }; } fn release(h: *Harness, req: fs.Req, t: Target) Answer { _ = h; switch (t) { - .top => {}, + .top, .act => {}, .path => |e| unref(e), } return .{ .reply = .{ .tag = req.tag } }; @@ -714,6 +1357,7 @@ fn read(h: *Harness, req: fs.Req, t: Target) Answer { }, else => return fail(req.tag, E.ISDIR), }, + .act => |a| return activeRead(h, req, a), .path => |e| { // `O_NONBLOCK` so a fifo left in a harness root cannot park the // daemon in `open`; the kind check below refuses it anyway. @@ -792,6 +1436,13 @@ fn readdir(h: *Harness, req: fs.Req, t: Target) Answer { } }, .pid, .uptime => return fail(req.tag, E.NOTDIR), + .active => { + rescan(h); + for (std.enums.values(active.Kind)) |k| { + if (!hasLive(h, k)) continue; + st.add(activeNode(.{ .file = .harness_dir, .kind = k }), true, k.text()); + } + }, .omp, .hermes, .dsh => { const m = top.mirror().?; if (h.base(m.root) == null) return fail(req.tag, E.NOENT); @@ -814,6 +1465,7 @@ fn readdir(h: *Harness, req: fs.Req, t: Target) Answer { .overflow => return fail(req.tag, E.NFILE), } }, + .act => |a| return activeReaddir(h, req, a), } return .{ .reply = .{ .tag = req.tag }, .bytes = h.stage_buf[0..st.len] }; } @@ -904,6 +1556,14 @@ const testing = std.testing; /// ~/.claude or any other live harness root. const Rig = struct { dir: testing.TmpDir, + /// Build a fixture process tree under <home>/proc and point + /// `/active` at it. Off by default: a unit test must opt in to + /// reading a process tree, and it is never the live one. + with_proc: bool = false, + /// The pid the daemon believes it is, for the ancestry exclusion. + self_pid: u32 = 4242, + zmx: []const u8 = "zmx", + allow_move: bool = false, path_buf: [std.fs.max_path_bytes]u8 = undefined, home: []const u8 = undefined, h: *Harness = undefined, @@ -931,7 +1591,24 @@ const Rig = struct { bases[i] = try std.fmt.bufPrint(&base_buf[i], "{s}/{s}", .{ rig.home, home_dirs[i] }); } rig.harness_mem = undefined; - rig.harness_mem.init(.{ .io = io, .pid = 4242, .bases = bases }); + var proc_buf: [std.fs.max_path_bytes]u8 = undefined; + const proc_root: []const u8 = if (rig.with_proc) + try std.fmt.bufPrint(&proc_buf, "{s}/proc", .{rig.home}) + else + ""; + if (rig.with_proc) { + try Io.Dir.cwd().createDirPath(io, proc_root); + try rig.put("proc/stat", "cpu 1 2 3\nbtime 1000000\nprocesses 7\n"); + } + rig.harness_mem.init(.{ + .io = io, + .pid = rig.self_pid, + .bases = bases, + .proc = proc_root, + .home = rig.home, + .zmx = rig.zmx, + .allow_move = rig.allow_move, + }); rig.h = &rig.harness_mem; } @@ -964,6 +1641,58 @@ const Rig = struct { fn lookupName(rig: *Rig, tag: u64, dir_node: u64, name: []const u8) Answer { return handle(rig.h, .{ .tag = tag, .op = .lookup, .node = dir_node, .data = name }); } + + /// Writes one fake process into the fixture `/proc`: the `stat` + /// line (with the start time in field 22, after a name that holds + /// the spaces and parentheses a real one can), the NUL-separated + /// `cmdline`, and a `cwd` symlink. + fn fakeProc(rig: *Rig, o: struct { + pid: u32, + ppid: u32 = 1, + comm: []const u8 = "x", + starttime: u64 = 5000, + argv: []const u8, + cwd: []const u8 = "", + }) !void { + var rel: [128]u8 = undefined; + var stat_text: [512]u8 = undefined; + var w = Io.Writer.fixed(&stat_text); + try w.print("{d} ({s}) S {d}", .{ o.pid, o.comm, o.ppid }); + for (5..22) |i| try w.print(" {d}", .{i}); + try w.print(" {d} 0 0\n", .{o.starttime}); + try rig.put(try std.fmt.bufPrint(&rel, "proc/{d}/stat", .{o.pid}), w.buffered()); + try rig.put(try std.fmt.bufPrint(&rel, "proc/{d}/cmdline", .{o.pid}), o.argv); + + const target = if (o.cwd.len > 0) o.cwd else rig.home; + var link_buf: [std.fs.max_path_bytes]u8 = undefined; + const link = try std.fmt.bufPrintZ(&link_buf, "{s}/proc/{d}/cwd", .{ rig.home, o.pid }); + Io.Dir.cwd().symLink(testing.io, target, link, .{}) catch {}; + } + + /// Removes a fake process, as an exit would. + fn reapProc(rig: *Rig, pid: u32) !void { + var buf: [std.fs.max_path_bytes]u8 = undefined; + const dir_path = try std.fmt.bufPrint(&buf, "{s}/proc/{d}", .{ rig.home, pid }); + const d = try Io.Dir.openDirAbsolute(testing.io, dir_path, .{ .iterate = true }); + var rb: [Io.Dir.Iterator.reader_buffer_len]u8 align(@alignOf(usize)) = undefined; + var it = Io.Dir.Reader.init(d, &rb); + var names: [8][64]u8 = undefined; + var lens: [8]usize = @splat(0); + var n: usize = 0; + while (n < names.len) { + const e = (it.next(testing.io) catch break) orelse break; + @memcpy(names[n][0..e.name.len], e.name); + lens[n] = e.name.len; + n += 1; + } + Io.Dir.close(d, testing.io); + for (0..n) |i| { + var one: [std.fs.max_path_bytes]u8 = undefined; + const path = try std.fmt.bufPrint(&one, "{s}/{s}", .{ dir_path, names[i][0..lens[i]] }); + Io.Dir.deleteFileAbsolute(testing.io, path) catch {}; + } + try Io.Dir.cwd().deleteDir(testing.io, dir_path); + } }; const home_dirs = [5][]const u8{ ".claude", ".codex", ".omp", ".hermes", ".dsh" }; @@ -1327,3 +2056,142 @@ test "tree: a listing spanning several reads loses no entry" { try testing.expect(reads > 1); // the listing really did span several reads for (seen) |s| try testing.expect(s); } + +test { + _ = active; // the derived view's own tests run with the tree's +} + +// ---- /active: the derived view ------------------------------------------------ + +/// Walks `/active` down to one agent's directory, answering its node. +fn activeEntryNode(rig: *Rig, harness: []const u8, pid: []const u8) !u64 { + const act = rig.lookupName(90, root, "active"); + try testing.expect(act.reply.status == .ok); + // A listing is what rescans, so it comes before the walk. + _ = handle(rig.h, .{ .tag = 91, .op = .readdir, .node = act.reply.attr.node, .off = 0, .size = msize }); + const h_dir = rig.lookupName(92, act.reply.attr.node, harness); + try testing.expect(h_dir.reply.status == .ok); + const entry = rig.lookupName(93, h_dir.reply.attr.node, pid); + try testing.expect(entry.reply.status == .ok); + return entry.reply.attr.node; +} + +fn activeField(rig: *Rig, entry: u64, name: []const u8) ![]const u8 { + const f = rig.lookupName(94, entry, name); + try testing.expect(f.reply.status == .ok); + const r = handle(rig.h, .{ .tag = 95, .op = .read, .node = f.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(r.reply.status == .ok); + return r.bytes; +} + +test "active: a fixture process tree lists by harness and by pid" { + var rig: Rig = .{ .dir = undefined, .with_proc = true }; + try rig.start(); + defer rig.end(); + + // A Claude Code session, resolved the way the harness publishes it. + try rig.fakeProc(.{ .pid = 1001, .comm = "claude", .starttime = 5000, .argv = "claude\x00--print\x00" }); + var rec: [512]u8 = undefined; + try rig.put(".claude/sessions/1001.json", try std.fmt.bufPrint(&rec, + \\{{"pid":1001,"sessionId":"sess-abc","cwd":"{s}", + \\ "name":"fixture-one","status":"busy","procStart":"5000"}} + , .{rig.home})); + var slug_buf: [512]u8 = undefined; + const slug = active.slugOf(.claude, rig.home, rig.home, &slug_buf).?; + var tr: [640]u8 = undefined; + try rig.put( + try std.fmt.bufPrint(&tr, ".claude/projects/{s}/sess-abc.jsonl", .{slug}), + "{\"type\":\"summary\",\"aiTitle\":\"Fixture Title\"}\n", + ); + + const entry = try activeEntryNode(&rig, "claude", "1001"); + try testing.expectEqualStrings("1001\n", try activeField(&rig, entry, "pid")); + try testing.expectEqualStrings("sess-abc\n", try activeField(&rig, entry, "session")); + try testing.expectEqualStrings("registry\n", try activeField(&rig, entry, "via")); + try testing.expectEqualStrings("fixture-one\n", try activeField(&rig, entry, "name")); + try testing.expectEqualStrings("busy\n", try activeField(&rig, entry, "status")); + try testing.expectEqualStrings("Fixture Title\n", try activeField(&rig, entry, "title")); + // started is the boot time plus the process's own, in seconds. + try testing.expectEqualStrings("1000050\n", try activeField(&rig, entry, "started")); + // The transcript is served as the file it is, not as a field. + const t = rig.lookupName(96, entry, "transcript"); + try testing.expect(t.reply.status == .ok); + const bytes = handle(rig.h, .{ .tag = 97, .op = .read, .node = t.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(std.mem.indexOf(u8, bytes.bytes, "Fixture Title") != null); + // A field this harness does not publish is absent, not blank. + try expectNoent(rig.lookupName(98, entry, "model")); +} + +test "active: a daemon or helper is never listed as an agent" { + var rig: Rig = .{ .dir = undefined, .with_proc = true }; + try rig.start(); + defer rig.end(); + // The shapes actually running on this machine. + try rig.fakeProc(.{ .pid = 1010, .comm = "python3", .argv = "python3\x00-m\x00hermes_cli.main\x00gateway\x00run\x00" }); + try rig.fakeProc(.{ .pid = 1011, .comm = "omp", .argv = "omp\x00__omp_worker_daemon_broker\x00" }); + try rig.fakeProc(.{ .pid = 1012, .comm = "claude", .argv = "claude\x00mcp-server\x00" }); + // ...and one real session, so an empty answer cannot pass by default. + try rig.fakeProc(.{ .pid = 1013, .comm = "claude", .argv = "claude\x00" }); + + const act = rig.lookupName(1, root, "active"); + const listing = handle(rig.h, .{ .tag = 2, .op = .readdir, .node = act.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(stageHas(listing.bytes, "claude")); + try testing.expect(!stageHas(listing.bytes, "hermes")); + try testing.expect(!stageHas(listing.bytes, "omp")); + const claude = rig.lookupName(3, act.reply.attr.node, "claude"); + const pids = handle(rig.h, .{ .tag = 4, .op = .readdir, .node = claude.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(stageHas(pids.bytes, "1013")); + try testing.expect(!stageHas(pids.bytes, "1012")); // the mcp server +} + +test "active: a pid reused by another process stops resolving" { + var rig: Rig = .{ .dir = undefined, .with_proc = true }; + try rig.start(); + defer rig.end(); + try rig.fakeProc(.{ .pid = 1020, .comm = "claude", .starttime = 5000, .argv = "claude\x00" }); + const entry = try activeEntryNode(&rig, "claude", "1020"); + try testing.expect(handle(rig.h, .{ .tag = 5, .op = .getattr, .node = entry }).reply.status == .ok); + + // The same pid, a different process: only the start time says so. + try rig.fakeProc(.{ .pid = 1020, .comm = "claude", .starttime = 9999, .argv = "claude\x00" }); + const act = rig.lookupName(6, root, "active"); + _ = handle(rig.h, .{ .tag = 7, .op = .readdir, .node = act.reply.attr.node, .off = 0, .size = msize }); + try expectNoent(handle(rig.h, .{ .tag = 8, .op = .getattr, .node = entry })); + // The pid is still there — as the new process, under a new node. + const fresh = try activeEntryNode(&rig, "claude", "1020"); + try testing.expect(fresh != entry); +} + +test "active: a process that exits leaves the tree" { + var rig: Rig = .{ .dir = undefined, .with_proc = true }; + try rig.start(); + defer rig.end(); + try rig.fakeProc(.{ .pid = 1030, .comm = "claude", .argv = "claude\x00" }); + const entry = try activeEntryNode(&rig, "claude", "1030"); + try testing.expect(handle(rig.h, .{ .tag = 9, .op = .getattr, .node = entry }).reply.status == .ok); + + try rig.reapProc(1030); + const act = rig.lookupName(10, root, "active"); + const listing = handle(rig.h, .{ .tag = 11, .op = .readdir, .node = act.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(!stageHas(listing.bytes, "claude")); + try expectNoent(handle(rig.h, .{ .tag = 12, .op = .getattr, .node = entry })); + try expectNoent(rig.lookupName(13, act.reply.attr.node, "claude")); +} + +test "active: the daemon never lists the process tree it lives in" { + // 1041 is the daemon, its parent 1040 is a harness: showing it would + // let a client act on the tree serving it. + var rig: Rig = .{ .dir = undefined, .with_proc = true, .self_pid = 1041 }; + try rig.start(); + defer rig.end(); + try rig.fakeProc(.{ .pid = 1040, .comm = "claude", .argv = "claude\x00" }); + try rig.fakeProc(.{ .pid = 1041, .comm = "9harness", .ppid = 1040, .argv = "9harness\x00" }); + try rig.fakeProc(.{ .pid = 1042, .comm = "claude", .argv = "claude\x00" }); + + const act = rig.lookupName(14, root, "active"); + const claude = rig.lookupName(15, act.reply.attr.node, "claude"); + try testing.expect(claude.reply.status == .ok); + const pids = handle(rig.h, .{ .tag = 16, .op = .readdir, .node = claude.reply.attr.node, .off = 0, .size = msize }); + try testing.expect(stageHas(pids.bytes, "1042")); // an unrelated session lists + try testing.expect(!stageHas(pids.bytes, "1040")); // its own parent does not +} diff --git a/9harness/test/e2e.sh b/9harness/test/e2e.sh index 6e57727..254d999 100755 --- a/9harness/test/e2e.sh +++ b/9harness/test/e2e.sh @@ -209,6 +209,87 @@ sleep 0.2 [ -S "$REG/harness" ] && fail "SIGTERM unposts the name" "socket still there" || pass "SIGTERM unposts the name" # ============================================================================ +echo "# part C: /active, the derived view" +# ============================================================================ +# A fake agent: a process whose argv[0] basename is a harness name, with a +# session record Claude Code's own shape. The proc root is the real /proc +# (the fixture seam is unit-tested); nothing real is ever killed, because +# the only process this part touches is the sleep it started itself. +ACT="$TMP/act" +mkdir -p "$ACT/bin" "$ACT/home/.claude/sessions" "$ACT/home/.claude/projects" "$ACT/rt/9p/zmx" +cp /bin/sleep "$ACT/bin/claude" +# A zmx that records how it was called and posts the name, as the real one +# does; a move must never be tested against the user's live sessions. +cat > "$ACT/bin/zmx" <<ZEOF +#!/bin/sh +echo "\$@" > "$ACT/zmx-argv" +: > "$ACT/rt/9p/zmx/\$2" +exit 0 +ZEOF +chmod +x "$ACT/bin/zmx" + +"$ACT/bin/claude" 600 & +AGENT=$! +PIDS+=("$AGENT") +sleep 0.3 +# Field 22 of /proc/<pid>/stat, counted from the last ')' — the executable +# name can hold spaces and parentheses, so the line is never just split. +AGENT_START=$(sed 's/.*) //' "/proc/$AGENT/stat" | awk '{print $20}') +AGENT_CWD=$(readlink "/proc/$AGENT/cwd") +printf '{"pid":%d,"sessionId":"sess-e2e","cwd":"%s","name":"e2e-agent","status":"idle","procStart":"%s"}\n' \ + "$AGENT" "$AGENT_CWD" "$AGENT_START" > "$ACT/home/.claude/sessions/$AGENT.json" + +act_roots() { + echo --root "claude=$ACT/home/.claude" --root "codex=$ACT/home/.codex" \ + --root "omp=$ACT/home/.omp" --root "hermes=$ACT/home/.hermes" --root "dsh=$ACT/home/.dsh" +} + +# First: the read-only default. No --allow-move, so zmx cannot be written. +XDG_RUNTIME_DIR="$ACT/rt" "$H9" --no-post --unix "$ACT/ro.sock" $(act_roots) >"$ACT/ro.log" 2>&1 & +PIDS+=("$!") +wait_posted "$ACT/ro.sock" || { fail "read-only daemon listens" "$(cat "$ACT/ro.log")"; } +expect_contains "the agent lists under its harness" "$AGENT" "$(p9 "$ACT/ro.sock" ls /active/claude)" +expect_eq "its session comes from the harness's own record" "sess-e2e" "$(p9 "$ACT/ro.sock" read "/active/claude/$AGENT/session")" +expect_eq "the route taken is named" "registry" "$(p9 "$ACT/ro.sock" read "/active/claude/$AGENT/via")" +expect_eq "the harness's own name is served" "e2e-agent" "$(p9 "$ACT/ro.sock" read "/active/claude/$AGENT/name")" +expect_eq "the harness's own status is served" "idle" "$(p9 "$ACT/ro.sock" read "/active/claude/$AGENT/status")" +if echo "nope" | p9 "$ACT/ro.sock" write "/active/claude/$AGENT/zmx" 2>/dev/null; then + fail "without --allow-move the tree stays read-only" "the write was accepted" +else + pass "without --allow-move the tree stays read-only" +fi +expect_eq "a refused move leaves the agent running" "yes" "$([ -d "/proc/$AGENT" ] && echo yes)" + +# Now a daemon that allows moves. +XDG_RUNTIME_DIR="$ACT/rt" "$H9" --no-post --unix "$ACT/rw.sock" $(act_roots) \ + --allow-move --zmx "$ACT/bin/zmx" >"$ACT/rw.log" 2>&1 & +PIDS+=("$!") +wait_posted "$ACT/rw.sock" || { fail "move-allowing daemon listens" "$(cat "$ACT/rw.log")"; } + +# Every refusal must come before anything is destroyed. +for bad in "has space" "../escape" "semi;colon"; do + echo "$bad" | p9 "$ACT/rw.sock" write "/active/claude/$AGENT/zmx" 2>/dev/null + if [ $? -eq 0 ]; then fail "an illegal zmx name is refused ($bad)"; else pass "an illegal zmx name is refused ($bad)"; fi +done +: > "$ACT/rt/9p/zmx/occupied" +echo "occupied" | p9 "$ACT/rw.sock" write "/active/claude/$AGENT/zmx" 2>/dev/null \ + && fail "a name a live session holds is refused" || pass "a name a live session holds is refused" +expect_eq "no refusal killed the agent" "yes" "$([ -d "/proc/$AGENT" ] && echo yes)" + +# The move itself. +if echo "e2e-moved" | p9 "$ACT/rw.sock" write "/active/claude/$AGENT/zmx" 2>"$ACT/move.err"; then + pass "a legal name moves the agent" +else + fail "a legal name moves the agent" "$(cat "$ACT/move.err")" +fi +sleep 0.4 +expect_eq "the agent it replaced is gone" "gone" "$([ -d "/proc/$AGENT" ] || echo gone)" +# The command is fixed by the harness: the client supplied only the name. +expect_eq "zmx ran the harness with its session resumed" \ + "run e2e-moved -d claude --resume sess-e2e" "$(cat "$ACT/zmx-argv")" +expect_missing "no client byte reached the command" "e2e-moved -d claude --resume sess-e2e;" "$(cat "$ACT/zmx-argv")" + +# ============================================================================ echo "# part B: the real roots, the real registry (read-only)" # ============================================================================ export XDG_RUNTIME_DIR="${HARNESS_TMP_XDG:-/run/user/$(id -u)}" diff --git a/9harness/zmxify b/9harness/zmxify new file mode 100755 index 0000000..5512ffa --- /dev/null +++ b/9harness/zmxify @@ -0,0 +1,176 @@ +#!/usr/lib/plan9/bin/rc +# zmxify - move a live harness session into a zmx session +# +# Lists the agents 9harness serves under /active, offers them to fzf, and +# writes the chosen zmx session name into that agent's `zmx` file. The +# daemon does the rest: it kills the harness and re-execs it under zmx +# with its session resumed. +# +# usage: zmxify [-n] [dir] +# dir only agents running there; all of them if omitted +# -n print the plan, write nothing +# +# This is the whole script now. Everything it used to work out for +# itself - which processes are harnesses, which stored session each one +# is writing, whether it already lives in zmx, and how to resume it - +# is a file under /active, resolved by the daemon and covered by its +# tests. Nothing here parses /proc, opens an fd table or queries a +# database, and nothing here kills anything: a name written into `zmx` +# is refused unless the session behind it was resolved first, so a move +# always has somewhere to come back to. + +path=(/usr/bin /bin $path) + +nl=' +' +tab=' ' +dry=() +here=() + +if(~ $1 -n){ + dry=y + shift +} +if(! ~ $#* 0){ + here=`$nl {readlink -f $1} + if(~ $#here 0 || ! test -d $here){ + echo 'zmxify: '^$1^': not a directory' >[1=2] + exit dir + } + shift +} +if(! ~ $#* 0){ + echo 'usage: zmxify [-n] [dir]' >[1=2] + exit usage +} + +mnt=$NINE_MOUNT +if(~ $#mnt 0) + mnt=/mnt/9p +act=$mnt/harness/active +if(! test -d $act){ + echo 'zmxify: no harness fs at '^$act >[1=2] + echo ' start it with: zmx run harness -d 9harness --allow-move' >[1=2] + echo ' and run this inside a 9ns mount (an interactive fish already is)' >[1=2] + exit nofs +} + +# A field that is absent (the harness does not publish it) and one that +# is present but empty (`zmx` on an agent outside zmx) both answer `-`: +# a null list would blow up the concatenations below. +fn field { # field dir name: its contents, or - + fv=() + if(test -e $1/$2) + fv=`$nl {cat $1/$2 >[2]/dev/null} + if(~ $#fv 0) + fv=- + echo $fv +} + +# An agent already under zmx has nowhere to go: that is what this moves +# things into. Its own session is *not* excluded - moving the terminal +# you typed this into is the common case, and it is safe, because the +# daemon does the killing and the re-exec. This script only has to get +# the write out; it may die the instant after, and the move still lands. +# Ancestry is worked out only to know whether we will survive to attach. +fn ancestry { # ancestry pid: it and every parent of it + aup=$1 + while(! ~ $aup 0 1){ + echo $aup + aup=`{sed -n 's/^PPid:[ ]*//p' /proc/$aup/status >[2]/dev/null} + if(~ $#aup 0) + aup=0 + } +} +mine=`$nl {ancestry $pid} + +lines=() +for(h in `{ls $act >[2]/dev/null}){ + for(d in `{ls $act/$h >[2]/dev/null}){ + a=$act/$h/$d + acwd=`$nl {field $a cwd} + asess=`$nl {field $a session} + azmx=`{field $a zmx} + take=() + if(~ $#here 0 || ~ $acwd $here) + take=y + # Already under zmx: there is nothing to move it into. + if(! ~ $azmx -) + take=() + # Its session did not resolve, so the daemon would refuse + # the write; offering it would only produce an error. + if(~ $asess -) + take=() + if(! ~ $#take 0){ + short=$acwd + if(! ~ $#home 0) + short=`$nl {echo -n $acwd | sed 's,^'^$home^',~,'} + avia=`{field $a via} + atitle=`$nl {field $a title} + lines=($lines $d^$tab^$h^$tab^$short^$tab^'via='^$avia^$tab^$asess^$tab^$"atitle) + } + } +} + +scope=(all directories) +if(! ~ $#here 0) + scope=$here +if(~ $#lines 0){ + echo 'zmxify: no movable harness session in '^$"scope >[1=2] + exit 0 +} + +fzfopts=(--no-multi --reverse --height=40% --delimiter=$tab --prompt='zmxify ' --header='move under zmx ['^$"scope^']'^$nl^'pid harness dir via session title') +sel=`$nl {{for(l in $lines) echo $l} | fzf $fzfopts} +if(~ $#sel 0){ + echo 'zmxify: nothing picked' >[1=2] + exit 0 +} + +pick=`{echo $sel | awk '{print $1}'} +harness=`{echo $sel | awk -F$tab '{print $2}'} +a=$act/$harness/$pick +if(! test -d $a){ + echo 'zmxify: '^$harness^' pid '^$pick^' is gone' >[1=2] + exit gone +} +pickcwd=`$nl {field $a cwd} + +# A zmx session name nobody is using yet. zmx posts every session into +# the registry, so the registry is what says which names are taken. +zbase=`{basename $pickcwd | tr -c 'A-Za-z0-9'^$nl -} +if(~ $#zbase 0) + zbase=session +zname=$harness^-^$zbase +reg=$XDG_RUNTIME_DIR^/9p/zmx +if(~ $#XDG_RUNTIME_DIR 0) + reg=/run/user/^`{id -u}^/9p/zmx +n=() +while(test -e $reg/$zname){ + n=($n x) + zname=$harness^-^$zbase^-^$#n +} + +if(! ~ $#dry 0){ + echo 'move ' $harness 'pid' $pick 'in' $pickcwd + echo 'session' `{field $a via} `{field $a session} + echo 'write ' $zname '>' $a/zmx + echo 'then ' zmx attach $zname + exit 0 +} + +echo 'zmxify: moving' $harness 'pid' $pick 'into zmx session' $zname +# Moving something we hang off means this shell goes with it. Say the +# attach line first, because there may be no `we` left to say it after. +if(~ $pick $mine) + echo 'zmxify: that is this terminal; it will drop. reattach with: zmx attach '^$zname +# The write returns when the new session is up, or fails having changed +# nothing it could avoid changing. The daemon owns the kill and the +# re-exec, so the move lands even if this process dies mid-write. +if(! echo $zname > $a/zmx){ + echo 'zmxify: the move was refused' >[1=2] + exit move +} +if(~ $pick $mine) + exit 0 +exec zmx attach $zname |
