summaryrefslogtreecommitdiff
Commit message (Collapse)AuthorAge
* 9agents: /active/<h>/<pid>/chat, the conversation as one file per messageHEADmainGabriel Schneider6 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Each message of a claude or codex agent's transcript is a file named <id>-<kind>: the id is its position in the transcript from 0, as 8 digits so the names sort; the kind is user, assistant, thinking, system or agent, and a tool call is named after its tool (00000002-bash) and its result after the call (00000003-bash-result). A file is the message as text (a tool call: its name, then its input); its mtime is when it was said. chat.zig indexes a transcript incrementally, keyed by (dev, ino) and checked by birth time and a hash of its first and last indexed bytes, so a file rewritten in place is indexed again. It holds offsets, never bytes: every read reads the file, and a read resumes where the last one stopped so a long message is not decoded from its start each time. The indexes and the runner live in lazily backed mappings (NORESERVE, NOHUGEPAGE), and the fid table is 32768 so a kernel mount can hold every message of a long chat. A 9agents started inside a user namespace (from a 9ns-wrapped terminal) cannot read /proc/<pid>/fd of processes outside it, so the fd route never resolved codex there; /active now also finds a codex session through the thread writer locks it holds (/proc/locks), via = lock. Fixes to /active found on the way: a fifo in place of a transcript or record no longer blocks the daemon; a transcript is opened from its pinned root a component at a time, so a symlinked directory or a sessionId with .. cannot reach outside it; slot generations are 24 bits, so a stale node id cannot come to name another agent; listing a harness rescans; an agent that had not resolved yet is asked again. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
* fs, serve: set an engine up in place, so a large fid table costs only the ↵Gabriel Schneider6 days
| | | | | | | | | | | | | | | fids used A kernel mount (9ns) holds a fid for every inode the kernel caches, so a server that wants `ls -l` of a directory of thousands of files to work needs a fid table of tens of thousands. Server.initIn sets an engine up in place and, with fid_index, never writes a fid slot at or past high_water; the walks over the table (references, reset, orphan) stop there. Runner.init sets each connection's small fields one by one and leaves the engine to start(), so connection slots nobody uses are never written. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
* Rename 9harness to 9agents; README, file-backed qids, worded errorsGabriel Schneider6 days
| | | | | | - 9harness/ becomes 9agents/ (build option -D9agents, package paths). - 9agents serves /README, reports qid paths from the file's (dev, ino) and qid versions that move with the file, and names its refusals.
* fs: a parked job whose fid was clunked is answered, not asked againGabriel Schneider10 days
| | | | | | | | | | | | A client that gives up on a parked open sends Tclunk for its fid without a Tflush; the fid goes, the parked job stays. On retry the job went back to the backend, which did the work -- opened a handle, made an object -- and the reply then found no fid and dropped it, handle and all. Now the retry looks for the fid first and answers "fid unknown" without asking. Found by an adversarial review of pardes's use of the engine. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* 9ns --mntgen: a server that never answers stalls only its own nameGabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it would sometimes freeze the whole session: no input reached it and nothing under /mnt/9p answered, until the shell was killed from outside. The cause was one posted server that accepted a connection and then never spoke 9P — pardes, answering its 9P from the same loop that was walking its own mount, was the one on this machine, but any wedged or half-dead server does it. Three things conspired, and each is fixed on its own: * The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion, Tattach and Tstat on the one thread that reads /dev/fuse, so while that server kept quiet no request for any name was read, and no FUSE_INTERRUPT either. Now the dispatcher makes a Mount without touching the network and queues the walk to the mount's worker, which dials while serving it. The dial is the request in flight, so an interrupt of the walk abandons it at once (`Session.abort_on_cancel`: nothing to flush before a session exists) and the walk answers EINTR; a failed dial leaves the mount undialed for the next walk to retry; a full listen backlog (the server stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of the name goes through the same queue and is answered from the remembered root attr, so the dispatcher never holds a session at all. * Once the dispatcher had read a request the process behind it was unkillable (FUSE waits out a request userspace has taken), and an INTERRUPT for a request still sitting in a mount's queue was dropped. The dispatcher now takes a queued request out and answers EINTR itself, and forwards only in-flight ones to the worker; queue and in-flight unique are read under the mount's mutex, where the worker moves a request from one to the other. A Tflush the server never answers is given 3s (`Session.flush_grace_ms`) and then the session is declared wedged: the request answers EINTR, the mount dies, the next walk makes a new one. * The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the INIT reply every LOOKUP and READDIR in a directory takes its inode lock, so one parked walk held up every other name under /mnt/9p however free the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now negotiated when the kernel offers it. What remains is the kernel's own serialization of lookups of one *name*: a second walker into the parked name waits for the first walk to end, and only then proceeds (and can be interrupted in its turn). An adversarial review of the above found three more things, fixed here: the single-connection bridge's one-slot stash stopped polling the FUSE fd while a second request was parked, so an INTERRUPT could not arrive (and parallel dirops make a second request routine) — the stash is now a queue of copies and the fd is always watched; a dead or wedged mount kept its socket open until exit, where a late-answering single-threaded server could block on it — the session is closed when the mount dies; and teardown after DESTROY or ENODEV (the child still alive, so stop_fd says nothing) could join a worker parked on a mute server forever — the sockets are shut down before the join. The flush grace is a deadline now, not a timer restarted on every wakeup. A black-box run against the binary (hostile servers: mute, garbage, close-after-accept, full backlog, 100 mute names, interrupt storms, 300 deaths of one server) found that a dead mount kept its socket, its interrupt pipe and a megabyte of buffers until exit — three descriptors per death — so `retire` now frees all of it and keeps only the slot; descriptors, threads and RSS stay flat across 400 deaths. The 4096-slot cap per process remains and is documented. Reproduced with a socket that accepts and never writes, posted beside 9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at once, the parked walker dies of its signal within milliseconds, and a server that answers the handshake but ignores reads and Tflush releases its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now check exactly that; nine.zig gains unit tests for the grace and the abort. Also in this change: the uncommitted ESTALE-on-death and FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount path here builds on. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* fs: let open, truncate, clunk and remove park on again, not only reads and ↵Gabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | writes A backend that is not ready for a request answers `again` and the engine parks the request, so the connection goes on answering everything else and a retry asks the parked request later. That only worked for a read, readdir or write; every other op answering `again` failed with EAGAIN on the spot. pardes needs the rest. Its 9P is answered on the connection's task while the editor's own thread may be out in a syscall in the middle of a step, and in that window a request that would change a pane -- an open of /pane/new, a truncating open or wstat, a remove -- has to wait, not fail, and has to wait without holding the connection up, because the editor's own request may be the next frame on that very connection (it is, when the syscall goes through a mount of the editor's own tree). Parking is exactly that. A parked open, wstat, clunk or remove goes back into the job slot on retry and its reply then goes through `jobReply` like any other, which is why the slot keeps only what an open needs (`omode`, `step`); a wstat parks only as the truncation to zero a Linux client sends for O_TRUNC, and a clunk parks before it lets its fid go, since a release that was never paid still owns its handle. A walk, attach, stat, create or renaming wstat keeps names in the input frame that parking lets go of, so those still answer EAGAIN. A parked job retried and parked again sits the round out like a retried read does -- without that the first version of this looped forever in `retry`. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Package 9harness: the published tarball was missing itGabriel Schneider10 days
| | | | | | | | | | | | | build.zig @imports each program's build fragment unconditionally, so a program absent from .paths builds fine from a checkout and fails for every consumer with zig-pkg/cloud9-.../9harness/build.zig: unable to load 'build.zig': FileNotFound 9ns and 9proc were listed; 9harness was not, from the change that added it. A pardes build against cloud9 main found it — the first consumer to try. Nothing in the suites covers "the package a dependent fetches actually builds", which is why it went unnoticed.
* 9harness: drop the zmxify script; the mount is the replacementGabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | The script listed /active, ran fzf over it and wrote the chosen name into an agent's zmx. Everything in it but the picker, a free-name loop and the attach was `ls` and `cat` against the tree — 130 lines re-deriving what four lines of shell already read: ls /mnt/9p/harness/active/*/* cat /mnt/9p/harness/active/claude/345104/session echo mine > /mnt/9p/harness/active/claude/345104/zmx zmx attach mine zmxify was 307 lines of rc because there was no filesystem to ask. The answer to that is the filesystem, not a smaller script on top of it, which would only become a second and worse interface beside the real one. DESIGN.md keeps what a caller now arranges itself (a free name, the attach) and why the old script's own-ancestry rule is gone rather than moved.
* 9harness: /active, and zmxify as a write to a fileGabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The mirror answers what files exist. /active answers what is running: one directory per live agent, normalized across harnesses, fields as small text files, synthesized per request. /active/claude/345104/{pid,cwd,session,via,name,status,title, model,started,zmx,transcript,agents/} This is /proc's shape, and deliberately: a directory per object named by pid under a directory per harness, rather than a compound `claude-345104` that would make you parse a name to recover a field that is already the directory above it. There is no `updated` file — that is the mtime of `transcript`, which stat already carries. Each harness is asked in its own terms, and the route is reported in `via` so a wrong guess is visible rather than silent. Claude Code publishes sessions/<pid>.json itself, with procStart as a pid-reuse guard, so nothing there is guessed. omp and dsh are found by the transcript they hold open, omp falling back to the store named after its cwd. codex's rollout file carries the session id in its *name*, so its sqlite is never opened. hermes is the one gap and needs none: its sessions live only in sqlite, and the only hermes processes that run are the gateway and the dashboard, which are not sessions. Liveness is /proc/<pid> plus a matching start time: a pid alone is not an identity. The daemon never lists its own ancestry, so it cannot show or act on the tree serving the request. The write path, and why it is a file and not a ctl: writing a zmx session name into an agent's `zmx` moves it there. The file means which zmx session this agent lives in, and writing makes that true. A ctl taking verbs is the ordinary Plan 9 spelling, and an executable script served in the tree is the spelling zmx's own `attach` uses, but a script that shells out to a local binary lies over a remote mount — it would run against a session that is not on the client's machine. A write is served where the authority is. Every refusal comes before anything is destroyed: the name must be zmx's label charset, unused by a live session, and the agent's session must have resolved, because nothing is killed that has nowhere to come back to. The command is fixed per harness and no client byte reaches exec. It is off unless --allow-move: this is the one place the tree is not read-only, and anything that can mount it could otherwise kill an agent. 9harness/zmxify replaces the 307-line rc script. It parses no /proc, opens no fd table and queries no database; it lists /active, offers the rows to fzf and writes the chosen name. It no longer excludes the caller's own session, which the old one had to: that script did the killing itself, so killing its own parent lost the session it was rescuing. The daemon completes the kill and the re-exec whether or not the client is still connected — verified by hanging up immediately after sending the write — so zmxifying the terminal you are sitting in now works, which is the common case. --proc DIR is the fixture seam: the scan, the liveness guard, the exclusions and the ancestry rule are unit-tested against a fake process tree, never the live one. Suites: 87/87 root, 48/48 9ns, 26/26 9harness (+5 for the view), 60/60 9proc, 144/144 programs-test, 51+88 9ns integration, 44/0 9harness end-to-end (+16, including a real move against a fake harness and a fake zmx), 29/0 9proc debug, 213/0 9ns adversarial, freestanding green.
* 9harness: the harness fs daemonGabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The last item of the 9P plan, replacing zmxify's introspection half: a read-only, fresh-from-disk 9P view of every agent harness's state on this machine, posted as `harness` like any other service, so a shell inside a 9ns --mntgen mount reads it at /mnt/9p/harness with no setup. /pid /uptime /claude/{projects,history,skills} /codex/{sessions,session-index,history} /omp /hermes /dsh the mirrors /skills/{claude,codex,omp} Nothing is cached: a lookup, getattr or readdir walks the real filesystem, so a transcript grows as its harness writes it and a new session appears as soon as its file lands. Writes answer EPERM, and no name that looks like a credential, key, token or auth store is ever answered at any depth. Three findings from the adversarial pass, each with its regression: - The read path composed <base>/<rel> and opened it in one call, which follows symlinks. A name swapped for a link between the walk and the read served bytes from outside every pinned root (proved against /etc/passwd). Every stat, read and readdir now resolves through openIn, which walks from the base one component at a time with O_NOFOLLOW, and O_PATH for the intermediates, so no component can redirect the walk. O_PATH also keeps a fifo in a root from parking the daemon in open(); a read refuses anything but a regular file. - Joining a child onto an empty relative path returned an uncopied scratch slice, so every file at the top of a mirror root (/hermes/x, /dsh/x) listed but read back uninitialized stack bytes. - A directory past the comptime caps was served short, and a short listing cannot be told from a small directory. The caps answer NFILE now. Staging also stops at the first record that does not fit instead of packing a shorter one behind it, which dropped that entry from the listing across the read boundary. Suites: 13/13 unit (fake HOME, never the live roots), 36/0 end-to-end including the mntgen money shot and the live ~/.claude/.credentials.json proved unreachable, 131/131 programs-test.
* 9ns --mntgen: registry subdirectories are mount points tooGabriel Schneider10 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A registry entry that is a directory is now served the way the root is: a synthetic directory listing the real one, dialing the sockets inside it on walk and recursing into further directories, to max_synth_depth (8) levels across max_synth_dirs (64) synthetic nodes. That is the plan9port mntgen shape and the layout zmx now posts under, so a live session reads at /mnt/9p/zmx/<name>. Before this a directory in the registry was dialed like a socket and answered EIO for good. post gains the two entry points the traversal needs: postedDir (the registry scan, against any directory) and dialPath (a dial by composed path, no name validation). Hardening, each from an attack that broke the code: - BATCH_FORGET carries entries for many owners and puts 0 in the header nodeid, so routing it by the header dropped all of them: 32 of 64 synthetic slots leaked in one close burst and the subdirectories that held them answered EIO forever. distributeForgets unpacks the body and hands each entry to its owner. - probe() and connectBlocking() copied a caller's path into the kernel address with no bound: a path past sun_path overran the 110-byte stack sockaddr (a panic in Debug, silent corruption in ReleaseFast). Both refuse it now, probe as `.live` so a claim never deletes what it could not inspect. - That bound then caught 9proc's own listener, which handed probe() the whole 108-byte sun_path array instead of the path inside it. The probe reads `.live` for anything it cannot ask about, so every stale socket became AlreadyListening and no server could ever take a dead predecessor's name back. It passes the path now. Suites: 87/87 root (+7 post/serve attack regressions), 48/48 9ns, 51+88 9ns integration (+4 traversal and slot-recycling checks), 213/0 9ns adversarial, 60/60 9proc plus its adversarial suites with a new stale-socket takeover check, freestanding green.
* post registry + 9ns --mntgen: the /srv translationGabriel Schneider11 days
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | cloud9.post: servers post their socket under a name in $XDG_RUNTIME_DIR/9p (post/unpost, posted, dial, Watch) and serve.Runner.listenPosted posts a server by name, unposting on stop. Names are budget-checked against the 108-byte socket path; a claim binds+listens at a private temp path and takes the name with atomic renames under flock (RENAME_NOREPLACE for free names, RENAME_EXCHANGE grab-verify-commit for stale ones): the registry path is never unlinked by a claim, live names refuse with AlreadyPosted, foreign files with NotSocket, and unpost removes only the caller's inode-matched entry. Watch surfaces inotify overflow and a replaced registry dir. 9ns --mntgen [--mount DIR] -- PROGRAM: one FUSE mount at /mnt/9p whose synthetic root lists the posted registry (no connection made); a walk into an unmounted name dials it and runs the existing bridge dispatch in a per-server worker thread, routed by mount index in the node id's top bits (ordinals never reused, cap 4096); a dead server answers EIO on its subtree and is re-dialed on the next walk. The dial watches stop_fd through Tversion (connectWatched). All existing 9ns forms are unchanged. 9proc's unix listener no longer blind-unlinks its path: a foreign non-socket is refused (Occupied), a live server is refused (AlreadyListening), only a refused socket is cleared, and stop() unlinks only the listener's own inode-matched socket. Hardened by adversarial review (GLM 5.3 x2 + DeepSeek V4.1 Flash, all high-thinking): double-bind races on one name (0 in 180k rounds), foreign-file TOCTOU deletions (0 in 4M flips), a 255-byte-name listing panic, inotify queue overflow silently dropped, listenPosted silently overwriting, dial-time Tversion hangs wedging the dispatcher, --debug silently ignored in mntgen, and xattr/statx probes answering EPERM on the synthetic root (broke `ls -l /mnt/9p`). Tests: root 80/80, 9ns 47/47, 9proc 60/60, integration 88/88 + mntgen 37/37, adversarial 213/0, freestanding riscv32 gate green.
* 9web: multiplexer, HTTP view of the tree, live streams, richer pageGabriel Schneider12 days
| | | | | | | | | | | | | | | One upstream 9P connection now serves any number of browser WebSocket sessions and, with --serve, plain 9P clients over TCP or Unix (a 9pserve- style frame remux: tags and fids remapped, Tflush forwarded, reconnect on upstream loss). /fs/<path> maps HTTP onto the tree: GET file or directory (JSON or HTML), Range, HEAD with 9P headers, PUT (create, truncate, append, trailing slash makes a directory), DELETE, and ?follow=1 or text/event-stream turning a blocking read into server-sent events with Tflush on disconnect. The page gains a lazy tree, stat panel, create, rename, delete, upload and follow mode. --probe embeds 9proc for self-introspection. New mux-test and http-fs-test steps; e2e still passes. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Add cloud9.serve: an std.Io runner around the file-server engineGabriel Schneider12 days
| | | | | | | | | | | | Runner(Backend, Options, Limits) listens on Unix or TCP, runs a reader and a serve task per connection in one Io.Group, pushes frames into an fs.Server, and lets the backend answer now or later from any task or thread (reply, flush, wake); Tflush, greet timeout, connection limit, close and stop with cancellation are covered by tests over real sockets. The engine and the backend contract stay Io-free, so the push/step mode for freestanding targets is unchanged. Documented as the two ways to drive the engine. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* 9proc core becomes a backend of cloud9.fs; engine gains optional featuresGabriel Schneider12 days
| | | | | | | | | | | | | | | | 9proc's own fid table, walk loop and dir-read engine are replaced by cloud9.fs.Server; the tree (static, vars, providers) is served through the engine's Req/Reply contract with node ids that keep the old qid scheme. Providers may answer later by returning error.Again (parked in the engine, retried each step, Tflush -> EINTR); no new files are exposed. Engine (backward compatible, all opt-in via Backend.features / Options): create, remove, wstat, reference accounting for backends that count handles, a salted fid index, name_capacity 0 (names from getattr), Reply.ename for backend-chosen error text, Attr.path/version/atime. Engine-level error strings and the 217-byte msize floor now apply to 9proc; tests updated accordingly. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Add the file-server engine: cloud9.fs.Server(Backend, Options)Gabriel Schneider12 days
| | | | | | | | | | | | | The asynchronous 9P file-server engine from the Pardes editor moves into the library: fid table, walks, directory cursors, a job/slot model where backend replies arrive later by tag (status again = parked), Tflush cancellation and orphaned fids on hangup. Allocation-free, no OS calls, no std.Io; comptime Options (fid, slot, park data, name and user capacities) replace the editor's constants. Backend contract types (Req, Reply/ReplyWith, Op, Status, Attr, E, error strings) live here. 29 engine tests plus the client tests that sat beside it in Pardes. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* 9ns: --name and /mnt/9p/<name> mounts, qid.path as inode number, interrupts ↵Gabriel Schneider12 days
| | | | | | | | | | | | | | | | | | | | as Tflush - --name NAME (default derived from the transport: socket basename, tcp-IP-PORT, spawned command, fdN) mounts at /mnt/9p/<name>; --mount still overrides. ensureMountpoint walks down and creates missing components, shadowing the deepest unwritable ancestor. NINE_MOUNT is the only exported variable. - The inode number reported to the kernel is the 9P qid.path for every node, root included; a server handing qid.path 1 to a file (Pardes /self) no longer collides with the root. - FUSE_INTERRUPT for the request in flight becomes Tflush; a blocked read returns EINTR when the server answers the flush, chunked transfers return short counts, other requests arriving meanwhile are stashed and served next. Servers ignoring Tflush still block until they answer. - 9ns-test now covers nine/bridge/fuse; new adv_bridge_interrupt suite (28); 9ns-itest grows to 88 checks. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Rename programs: 9player -> 9ns, introspect -> 9proc, app -> web (9web)Gabriel Schneider12 days
| | | | | | | | Directories, binaries, build options (-D9ns, -D9proc), step names, module name (9proc), thread and fs names, env var NINEPLAYER_MOUNT -> NINE_MOUNT, docs and test scripts. Browser assets move to web/static. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Add 9player and introspect as programs beside the libraryGabriel Schneider12 days
| | | | | | | | | | | | | 9player/: FUSE mount CLI that mounts a 9P2000 tree into a fresh user+mount namespace and runs a program in it (no root, no libfuse, no libc). introspect/: the 9P debug/introspection library (freestanding core, value renderers, Linux probe with threads/stacks/memory/breakpoints/panics) and its demo server. Each has its own build fragment; the root build.zig wires them behind -D9player/-Dintrospect with namespaced steps (9player-itest, introspect-check-freestanding, programs-test, ...) and exports the introspect module for dependents. This is the layout for related programs. Co-Authored-By: Claude Fable 5.1 <[email protected]>
* Add reusable HTTP transport, serial gateway, and WASM file browserGabriel Schneider2026-09-16
|
* Implement base 9P2000 sessions, shared transports, and conformance probesGabriel Schneider2026-09-14