summaryrefslogtreecommitdiff
path: root/9agents/docs/DESIGN.md
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-09-25 16:40:49 -0300
committerGabriel Schneider <[email protected]>2026-09-25 17:52:33 -0300
commit1c1b192d4ef59199a7196229d56901b1f6512678 (patch)
tree9f6e4c1efcf877d1231aa58975ccd923ad4f2f2e /9agents/docs/DESIGN.md
parent73602127d15d10a1932b6fe916bd18a608054980 (diff)
downloadcloud9-main.tar.gz
cloud9-main.zip
9agents: /active/<h>/<pid>/chat, the conversation as one file per messageHEADmain
Each message of a claude or codex agent's transcript is a file named <id>-<kind>: the id is its position in the transcript from 0, as 8 digits so the names sort; the kind is user, assistant, thinking, system or agent, and a tool call is named after its tool (00000002-bash) and its result after the call (00000003-bash-result). A file is the message as text (a tool call: its name, then its input); its mtime is when it was said. chat.zig indexes a transcript incrementally, keyed by (dev, ino) and checked by birth time and a hash of its first and last indexed bytes, so a file rewritten in place is indexed again. It holds offsets, never bytes: every read reads the file, and a read resumes where the last one stopped so a long message is not decoded from its start each time. The indexes and the runner live in lazily backed mappings (NORESERVE, NOHUGEPAGE), and the fid table is 32768 so a kernel mount can hold every message of a long chat. A 9agents started inside a user namespace (from a 9ns-wrapped terminal) cannot read /proc/<pid>/fd of processes outside it, so the fd route never resolved codex there; /active now also finds a codex session through the thread writer locks it holds (/proc/locks), via = lock. Fixes to /active found on the way: a fifo in place of a transcript or record no longer blocks the daemon; a transcript is opened from its pinned root a component at a time, so a symlinked directory or a sessionId with .. cannot reach outside it; slot generations are 24 bits, so a stale node id cannot come to name another agent; listing a harness rescans; an agent that had not resolved yet is asked again. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Diffstat (limited to '9agents/docs/DESIGN.md')
-rw-r--r--9agents/docs/DESIGN.md142
1 files changed, 139 insertions, 3 deletions
diff --git a/9agents/docs/DESIGN.md b/9agents/docs/DESIGN.md
index 8d399d3..1530153 100644
--- a/9agents/docs/DESIGN.md
+++ b/9agents/docs/DESIGN.md
@@ -174,8 +174,17 @@ checks `perm` first and answers `e_perm`, which is already acme's spelling.
`Harness` (tree.zig) is the backend of `fs.Server(Harness, opts)` on
`serve.Runner` (main.zig): allocation-free request paths, comptime caps,
-`Io` passed explicitly, static memory in .bss (the path table and the
-per-connection buffers; ~3 MB of tables, untouched pages cost nothing).
+`Io` passed explicitly, static memory: the harness (the path table,
+~3 MB) in .bss, and the two large tables in anonymous mappings of their
+own, `MAP_NORESERVE` and `MADV_NOHUGEPAGE`, so they are address space
+until used — the runner (~160 MB: 16 connections, each with a fid table
+of 32768) and the chat indexes (~100 MB). A kernel mount holds a fid for
+every inode the kernel caches, and `ls -l` of one long chat caches
+thousands, so the fid table has to be large; the engine sets itself up in
+place (`initIn`) and writes a fid slot only when a client first uses it.
+Huge pages are refused because each connection's and each index's header
+sits in its own 2 MB stretch, and the first write to one would back all
+of it: with them, the idle daemon was 79 MB; without, 11 MB.
- **Node ids** pack `Node{idx: u8, kind: u8, serial: u48}` in the u64,
zmx-style. Static skeleton nodes (facts, the harness dirs, skills) are
@@ -307,6 +316,7 @@ of `~/.local/bin/zmxify`, whose resolution ladder it ports.
status busy only where the harness publishes one
transcript the live .jsonl, growing as it writes
agents/ <name>/{model,transcript}
+ chat/ 00000000-user 00000001-assistant 00000002-bash ... (claude, codex; below)
omp/
155574/ ...
```
@@ -367,7 +377,7 @@ visible rather than silent:
| claude | `~/.claude/sessions/<pid>.json`, which the harness maintains itself: sessionId, cwd, name, status, version, and `procStart` as a pid-reuse guard | `registry` |
| omp | the transcript it holds open under `~/.omp/agent/sessions/`, newest first; else the store directory named after its cwd with `/` becoming `-` | `fd`, `dir` |
| dsh | the `session-<uuid>/session.jsonl.zstd` it holds open; zstd, so there is no title | `fd` |
-| codex | the rollout it holds open — `sessions/YYYY/MM/DD/rollout-<when>-<uuid>.jsonl`, whose *name* carries the session id, so its sqlite is never opened | `fd` |
+| codex | the rollout it holds open — `sessions/YYYY/MM/DD/rollout-<when>-<uuid>.jsonl`, whose *name* carries the session id, so its sqlite is never opened; else the thread whose writer lock it holds (below) | `fd`, `lock` |
| hermes | nothing to resolve: see below | `none` |
A resolved session is checked against the process's cwd before it is
@@ -496,6 +506,121 @@ and watching the move land anyway. So moving the terminal you are
sitting in works, which is the common case; the shell drops, and you
reattach with the name you wrote.
+**When the fds are closed, codex's locks name its session.** A daemon
+started inside a user namespace — anything launched from a terminal that
+`9ns` wraps (`unshare -U`, mapping only the user's uid) — cannot read
+`/proc/<pid>/fd`, `cwd` or `environ` of a process outside that namespace:
+the kernel's ptrace access check fails across it, whatever the uid. The
+systemd service runs in the initial namespace and the fd route works
+there; `zmx run agents -d 9agents` from a terminal does not, and the fd
+route finds nothing. What stays visible is the locks:
+codex holds a write `flock` on `~/.codex/thread-writer-locks/<thread>.lock`
+for every thread it is writing, and `/proc/locks` is world-readable and
+names the holder's pid and the locked file's inode. So the `lock` route
+reads the pid's `FLOCK WRITE` lines, matches their inodes against the lock
+directory, and turns each thread id into its rollout: the id is a UUIDv7,
+whose first 48 bits are the milliseconds it was minted at, so the
+`YYYY/MM/DD` directory is the UTC day or one either side of it (the
+rollouts are named in local time). A process holds its own thread and its
+subagents' (multi-agent v2 runs them in-process); its own is the one whose
+first record says `"thread_source":"user"`, and a process holding several
+(a `/new`) resolves to the newest.
+
+The inodes are compared without the device. On btrfs `/proc/locks` prints
+the superblock's device (`00:37`) and `stat` the subvolume's anonymous one
+(`0:56`), so the pair never agrees; every candidate is in one directory,
+where the inode alone is unique.
+
+### chat/: the conversation, one file per message
+
+```
+/active/claude/345104/chat/
+ 00000000-user 00000001-assistant 00000002-bash 00000003-bash-result
+ 00000004-thinking 00000005-read 00000006-read-result 00000007-system ...
+```
+
+`chat/` is the agent's transcript read as messages: one file per message,
+named `<id>-<kind>`. The id is the message's position in the transcript,
+counted from 0 and written as 8 digits, so a plain `ls` lists them in the
+order they were said. Transcripts are append-only, so an id never moves
+and a message written later takes the next one. The kind says who is
+speaking, normalized across harnesses:
+
+| kind | claude | codex |
+|---|---|---|
+| `user` | a `user` record's text; a prompt typed while the model worked (`queued_command`, origin human) | `message`, role user |
+| `assistant` | an `assistant` text block | `message`, role assistant |
+| `thinking` | a thinking block, when it is not redacted to an empty string | a `reasoning` summary, when there is one |
+| *tool* (`bash`, `read`, `shell`, `apply_patch`, ...) | `tool_use`: name, then input | `function_call`, `custom_tool_call`, `*_call` (`web_search`, `local_shell`): name, then arguments |
+| *tool*`-result` | `tool_result`, matched to its call by `tool_use_id` | `*_output`, matched by `call_id` |
+| `tool-call`, `tool-result` | a call whose tool has no usable name, and a result whose call is not known | the same |
+| `system` | `isMeta` or compact-summary user text, `system` records, queued prompts not from a person | `message`, role developer or system; user-role text codex writes itself (AGENTS.md, `<environment_context>`, goals, plugin lists) |
+| `agent` | — | `agent_message`: the sending agent's path, then its text |
+
+A tool call is named after its tool: lowercased, with anything but
+letters, digits, `_`, `.` and `-` turned into `-`, and cut at 48 bytes. A
+tool whose name would read as another kind (`user`, `tool-result`,
+anything ending `-result`) is named `tool-call` instead, so a name never
+lies about who spoke. Its result is found through the call's id, among the
+last 256 calls; results follow their calls closely, even when a model runs
+calls in parallel. The index keeps each chat's distinct tool names (up to
+255) and gives every message a one-byte reference to one.
+
+Everything else in a transcript is the harness's bookkeeping (attachments,
+mode changes, titles, token counts, snapshots, and codex's `event_msg`
+copies of what the `response_item`s already say) and is not a message.
+Claude Code writes each content block of a reply as its own record, and
+each is its own message here too. A subagent's thread (`isSidechain`) is
+not this chat.
+
+The file is the message as text: JSON strings unescaped, each part on its
+own line, a newline added where the text lacks one, and an image as
+`[image]`. A tool call is the tool's name and then its input, which stays
+JSON — it is data the model wrote, not prose. A block of a kind this does
+not know is served as its JSON rather than dropped. The mtime of a message
+is the timestamp its record carries, so `ls -lt` shows when it was said.
+
+A name has one spelling: the id is always 8 digits, so `1-user`,
+`000000001-user`, `00000001-assistant` for a user message, or a bare `1`
+do not resolve, and every message answers to exactly one path.
+Only complete lines count: a record the harness is half-way through
+writing waits for its newline, which is what keeps ids from being handed
+out twice.
+
+**The index, and why it is not a cache.** Finding message N means knowing
+where every message before it is, which is a scan of the transcript, and
+codex rollouts reach 200 MB. So each transcript gets an index
+(`src/chat.zig`): per message, the spans of the file that render it — an
+offset, a length, and how to decode it. It holds positions, never bytes:
+every read still reads the file. It is keyed by the file's (dev, ino) and
+checked against it on every request; a transcript that grew is indexed
+from where the index stopped. Anything else is another file and is
+indexed again from the start: one that shrank, one born again under a
+reused inode (its birth time differs), and one rewritten in place to the
+same size or longer, which keeps its inode and only its bytes can give
+away — so the index also keeps a hash of the first and the last 64 bytes
+it covered, and checks both. Eight transcripts are indexed at once,
+reused least recently used first, in a pool of its own mapped
+`MAP_NORESERVE`: ~100 MB of address space, none of it memory until an
+index writes it.
+
+A string's output offset does not map onto its input offset, so a read
+from the middle of a long message would have to decode everything before
+it, and reading a message front to back would cost the square of its
+length. The pool remembers where the last read stopped — the input offset
+of the last escape or run it began, and the output offset that produced —
+and the next read of the same message resumes there.
+
+The first read of a chat pays for the scan: a 17 MB Claude Code
+transcript takes ~150 ms in a Debug build, the largest codex rollout on
+this machine (228 MB) ~2 s in Debug and ~130 ms in ReleaseFast. That time
+is spent under the daemon's one mutex, like a move. After that, a request
+costs a stat and whatever the transcript grew by.
+
+The caps are loud, as everywhere else: an index holds 262144 messages (the
+longest rollout on this machine has 32k), and a transcript with more
+answers the listing with an error rather than a list that stops early.
+
### Testability
`--proc DIR` overrides `/proc` the way `--root NAME=PATH` overrides a
@@ -513,3 +638,14 @@ against a fake harness in `test/e2e.sh`.
- Streaming/parked reads (a read answers at once; the transcripts are
plain files), and union-directory semantics beyond the plain
/skills/<harness> dirs.
+- `chat/` for sessions that are not running: it hangs off `/active`
+ entries only. The mirrors serve a transcript as a file, and Claude
+ Code already keeps a `<session>/` directory beside each transcript for
+ its subagents, so a chat there needs a name of its own.
+- Chats of subagents (omp's `agents/`, codex's in-process threads) and of
+ omp, hermes and dsh.
+- A session changed inside a running process (`/clear` in Claude Code,
+ `/new` in codex): the `/active` entry keeps the transcript it resolved
+ when it first resolved one. (An agent that resolved nothing is asked
+ again on every scan, since a harness seen in its first instant may not
+ have written its record or taken its lock yet.)