| Commit message (Collapse) | Author | Age |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Each message of a claude or codex agent's transcript is a file named
<id>-<kind>: the id is its position in the transcript from 0, as 8
digits so the names sort; the kind is user, assistant, thinking, system
or agent, and a tool call is named after its tool (00000002-bash) and
its result after the call (00000003-bash-result). A file is the message
as text (a tool call: its name, then its input); its mtime is when it
was said.
chat.zig indexes a transcript incrementally, keyed by (dev, ino) and
checked by birth time and a hash of its first and last indexed bytes, so
a file rewritten in place is indexed again. It holds offsets, never
bytes: every read reads the file, and a read resumes where the last one
stopped so a long message is not decoded from its start each time. The
indexes and the runner live in lazily backed mappings (NORESERVE,
NOHUGEPAGE), and the fid table is 32768 so a kernel mount can hold every
message of a long chat.
A 9agents started inside a user namespace (from a 9ns-wrapped terminal)
cannot read /proc/<pid>/fd of processes outside it, so the fd route
never resolved codex there; /active now also finds a codex session
through the thread writer locks it holds (/proc/locks), via = lock.
Fixes to /active found on the way: a fifo in place of a transcript or
record no longer blocks the daemon; a transcript is opened from its
pinned root a component at a time, so a symlinked directory or a
sessionId with .. cannot reach outside it; slot generations are 24 bits,
so a stale node id cannot come to name another agent; listing a harness
rescans; an agent that had not resolved yet is asked again.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
fids used
A kernel mount (9ns) holds a fid for every inode the kernel caches, so a
server that wants `ls -l` of a directory of thousands of files to work
needs a fid table of tens of thousands. Server.initIn sets an engine up
in place and, with fid_index, never writes a fid slot at or past
high_water; the walks over the table (references, reset, orphan) stop
there. Runner.init sets each connection's small fields one by one and
leaves the engine to start(), so connection slots nobody uses are never
written.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
| |
|
|
|
|
| |
- 9harness/ becomes 9agents/ (build option -D9agents, package paths).
- 9agents serves /README, reports qid paths from the file's (dev, ino)
and qid versions that move with the file, and names its refusals.
|
| |
|
|
|
|
|
|
|
|
|
|
| |
A client that gives up on a parked open sends Tclunk for its fid without
a Tflush; the fid goes, the parked job stays. On retry the job went back
to the backend, which did the work -- opened a handle, made an object --
and the reply then found no fid and dropped it, handle and all. Now the
retry looks for the fid first and answers "fid unknown" without asking.
Found by an adversarial review of pardes's use of the engine.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it
would sometimes freeze the whole session: no input reached it and nothing
under /mnt/9p answered, until the shell was killed from outside. The cause
was one posted server that accepted a connection and then never spoke 9P —
pardes, answering its 9P from the same loop that was walking its own mount,
was the one on this machine, but any wedged or half-dead server does it.
Three things conspired, and each is fixed on its own:
* The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion,
Tattach and Tstat on the one thread that reads /dev/fuse, so while that
server kept quiet no request for any name was read, and no FUSE_INTERRUPT
either. Now the dispatcher makes a Mount without touching the network and
queues the walk to the mount's worker, which dials while serving it. The
dial is the request in flight, so an interrupt of the walk abandons it at
once (`Session.abort_on_cancel`: nothing to flush before a session
exists) and the walk answers EINTR; a failed dial leaves the mount
undialed for the next walk to retry; a full listen backlog (the server
stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of
the name goes through the same queue and is answered from the remembered
root attr, so the dispatcher never holds a session at all.
* Once the dispatcher had read a request the process behind it was
unkillable (FUSE waits out a request userspace has taken), and an
INTERRUPT for a request still sitting in a mount's queue was dropped. The
dispatcher now takes a queued request out and answers EINTR itself, and
forwards only in-flight ones to the worker; queue and in-flight unique
are read under the mount's mutex, where the worker moves a request from
one to the other. A Tflush the server never answers is given 3s
(`Session.flush_grace_ms`) and then the session is declared wedged: the
request answers EINTR, the mount dies, the next walk makes a new one.
* The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the
INIT reply every LOOKUP and READDIR in a directory takes its inode lock,
so one parked walk held up every other name under /mnt/9p however free
the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now
negotiated when the kernel offers it.
What remains is the kernel's own serialization of lookups of one *name*: a
second walker into the parked name waits for the first walk to end, and
only then proceeds (and can be interrupted in its turn).
An adversarial review of the above found three more things, fixed here:
the single-connection bridge's one-slot stash stopped polling the FUSE fd
while a second request was parked, so an INTERRUPT could not arrive (and
parallel dirops make a second request routine) — the stash is now a queue
of copies and the fd is always watched; a dead or wedged mount kept its
socket open until exit, where a late-answering single-threaded server
could block on it — the session is closed when the mount dies; and
teardown after DESTROY or ENODEV (the child still alive, so stop_fd says
nothing) could join a worker parked on a mute server forever — the
sockets are shut down before the join. The flush grace is a deadline now,
not a timer restarted on every wakeup. A black-box run against the binary
(hostile servers: mute, garbage, close-after-accept, full backlog, 100
mute names, interrupt storms, 300 deaths of one server) found that a dead
mount kept its socket, its interrupt pipe and a megabyte of buffers until
exit — three descriptors per death — so `retire` now frees all of it and
keeps only the slot; descriptors, threads and RSS stay flat across 400
deaths. The 4096-slot cap per process remains and is documented.
Reproduced with a socket that accepts and never writes, posted beside
9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked
behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at
once, the parked walker dies of its signal within milliseconds, and a
server that answers the handshake but ignores reads and Tflush releases
its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now
check exactly that; nine.zig gains unit tests for the grace and the
abort. Also in this change: the uncommitted ESTALE-on-death and
FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount
path here builds on.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
writes
A backend that is not ready for a request answers `again` and the engine
parks the request, so the connection goes on answering everything else and
a retry asks the parked request later. That only worked for a read, readdir
or write; every other op answering `again` failed with EAGAIN on the spot.
pardes needs the rest. Its 9P is answered on the connection's task while
the editor's own thread may be out in a syscall in the middle of a step, and
in that window a request that would change a pane -- an open of /pane/new, a
truncating open or wstat, a remove -- has to wait, not fail, and has to wait
without holding the connection up, because the editor's own request may be
the next frame on that very connection (it is, when the syscall goes through
a mount of the editor's own tree). Parking is exactly that.
A parked open, wstat, clunk or remove goes back into the job slot on retry
and its reply then goes through `jobReply` like any other, which is why the
slot keeps only what an open needs (`omode`, `step`); a wstat parks only as
the truncation to zero a Linux client sends for O_TRUNC, and a clunk parks
before it lets its fid go, since a release that was never paid still owns
its handle. A walk, attach, stat, create or renaming wstat keeps names in the
input frame that parking lets go of, so those still answer EAGAIN. A parked
job retried and parked again sits the round out like a retried read does --
without that the first version of this looped forever in `retry`.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
build.zig @imports each program's build fragment unconditionally, so a
program absent from .paths builds fine from a checkout and fails for
every consumer with
zig-pkg/cloud9-.../9harness/build.zig: unable to load 'build.zig': FileNotFound
9ns and 9proc were listed; 9harness was not, from the change that added
it. A pardes build against cloud9 main found it — the first consumer to
try. Nothing in the suites covers "the package a dependent fetches
actually builds", which is why it went unnoticed.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The script listed /active, ran fzf over it and wrote the chosen name
into an agent's zmx. Everything in it but the picker, a free-name loop
and the attach was `ls` and `cat` against the tree — 130 lines
re-deriving what four lines of shell already read:
ls /mnt/9p/harness/active/*/*
cat /mnt/9p/harness/active/claude/345104/session
echo mine > /mnt/9p/harness/active/claude/345104/zmx
zmx attach mine
zmxify was 307 lines of rc because there was no filesystem to ask. The
answer to that is the filesystem, not a smaller script on top of it,
which would only become a second and worse interface beside the real
one. DESIGN.md keeps what a caller now arranges itself (a free name,
the attach) and why the old script's own-ancestry rule is gone rather
than moved.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The mirror answers what files exist. /active answers what is running:
one directory per live agent, normalized across harnesses, fields as
small text files, synthesized per request.
/active/claude/345104/{pid,cwd,session,via,name,status,title,
model,started,zmx,transcript,agents/}
This is /proc's shape, and deliberately: a directory per object named
by pid under a directory per harness, rather than a compound
`claude-345104` that would make you parse a name to recover a field
that is already the directory above it. There is no `updated` file —
that is the mtime of `transcript`, which stat already carries.
Each harness is asked in its own terms, and the route is reported in
`via` so a wrong guess is visible rather than silent. Claude Code
publishes sessions/<pid>.json itself, with procStart as a pid-reuse
guard, so nothing there is guessed. omp and dsh are found by the
transcript they hold open, omp falling back to the store named after
its cwd. codex's rollout file carries the session id in its *name*, so
its sqlite is never opened. hermes is the one gap and needs none: its
sessions live only in sqlite, and the only hermes processes that run
are the gateway and the dashboard, which are not sessions.
Liveness is /proc/<pid> plus a matching start time: a pid alone is not
an identity. The daemon never lists its own ancestry, so it cannot show
or act on the tree serving the request.
The write path, and why it is a file and not a ctl: writing a zmx
session name into an agent's `zmx` moves it there. The file means which
zmx session this agent lives in, and writing makes that true. A ctl
taking verbs is the ordinary Plan 9 spelling, and an executable script
served in the tree is the spelling zmx's own `attach` uses, but a
script that shells out to a local binary lies over a remote mount — it
would run against a session that is not on the client's machine. A
write is served where the authority is.
Every refusal comes before anything is destroyed: the name must be
zmx's label charset, unused by a live session, and the agent's session
must have resolved, because nothing is killed that has nowhere to come
back to. The command is fixed per harness and no client byte reaches
exec. It is off unless --allow-move: this is the one place the tree is
not read-only, and anything that can mount it could otherwise kill an
agent.
9harness/zmxify replaces the 307-line rc script. It parses no /proc,
opens no fd table and queries no database; it lists /active, offers the
rows to fzf and writes the chosen name. It no longer excludes the
caller's own session, which the old one had to: that script did the
killing itself, so killing its own parent lost the session it was
rescuing. The daemon completes the kill and the re-exec whether or not
the client is still connected — verified by hanging up immediately
after sending the write — so zmxifying the terminal you are sitting in
now works, which is the common case.
--proc DIR is the fixture seam: the scan, the liveness guard, the
exclusions and the ancestry rule are unit-tested against a fake process
tree, never the live one.
Suites: 87/87 root, 48/48 9ns, 26/26 9harness (+5 for the view),
60/60 9proc, 144/144 programs-test, 51+88 9ns integration, 44/0
9harness end-to-end (+16, including a real move against a fake harness
and a fake zmx), 29/0 9proc debug, 213/0 9ns adversarial, freestanding
green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The last item of the 9P plan, replacing zmxify's introspection half: a
read-only, fresh-from-disk 9P view of every agent harness's state on
this machine, posted as `harness` like any other service, so a shell
inside a 9ns --mntgen mount reads it at /mnt/9p/harness with no setup.
/pid /uptime /claude/{projects,history,skills}
/codex/{sessions,session-index,history}
/omp /hermes /dsh the mirrors
/skills/{claude,codex,omp}
Nothing is cached: a lookup, getattr or readdir walks the real
filesystem, so a transcript grows as its harness writes it and a new
session appears as soon as its file lands. Writes answer EPERM, and no
name that looks like a credential, key, token or auth store is ever
answered at any depth.
Three findings from the adversarial pass, each with its regression:
- The read path composed <base>/<rel> and opened it in one call, which
follows symlinks. A name swapped for a link between the walk and the
read served bytes from outside every pinned root (proved against
/etc/passwd). Every stat, read and readdir now resolves through
openIn, which walks from the base one component at a time with
O_NOFOLLOW, and O_PATH for the intermediates, so no component can
redirect the walk. O_PATH also keeps a fifo in a root from parking the
daemon in open(); a read refuses anything but a regular file.
- Joining a child onto an empty relative path returned an uncopied
scratch slice, so every file at the top of a mirror root (/hermes/x,
/dsh/x) listed but read back uninitialized stack bytes.
- A directory past the comptime caps was served short, and a short
listing cannot be told from a small directory. The caps answer NFILE
now. Staging also stops at the first record that does not fit instead
of packing a shorter one behind it, which dropped that entry from the
listing across the read boundary.
Suites: 13/13 unit (fake HOME, never the live roots), 36/0 end-to-end
including the mntgen money shot and the live ~/.claude/.credentials.json
proved unreachable, 131/131 programs-test.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A registry entry that is a directory is now served the way the root is:
a synthetic directory listing the real one, dialing the sockets inside
it on walk and recursing into further directories, to max_synth_depth
(8) levels across max_synth_dirs (64) synthetic nodes. That is the
plan9port mntgen shape and the layout zmx now posts under, so a live
session reads at /mnt/9p/zmx/<name>. Before this a directory in the
registry was dialed like a socket and answered EIO for good.
post gains the two entry points the traversal needs: postedDir (the
registry scan, against any directory) and dialPath (a dial by composed
path, no name validation).
Hardening, each from an attack that broke the code:
- BATCH_FORGET carries entries for many owners and puts 0 in the header
nodeid, so routing it by the header dropped all of them: 32 of 64
synthetic slots leaked in one close burst and the subdirectories that
held them answered EIO forever. distributeForgets unpacks the body and
hands each entry to its owner.
- probe() and connectBlocking() copied a caller's path into the kernel
address with no bound: a path past sun_path overran the 110-byte stack
sockaddr (a panic in Debug, silent corruption in ReleaseFast). Both
refuse it now, probe as `.live` so a claim never deletes what it could
not inspect.
- That bound then caught 9proc's own listener, which handed probe() the
whole 108-byte sun_path array instead of the path inside it. The probe
reads `.live` for anything it cannot ask about, so every stale socket
became AlreadyListening and no server could ever take a dead
predecessor's name back. It passes the path now.
Suites: 87/87 root (+7 post/serve attack regressions), 48/48 9ns,
51+88 9ns integration (+4 traversal and slot-recycling checks), 213/0
9ns adversarial, 60/60 9proc plus its adversarial suites with a new
stale-socket takeover check, freestanding green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
cloud9.post: servers post their socket under a name in
$XDG_RUNTIME_DIR/9p (post/unpost, posted, dial, Watch) and
serve.Runner.listenPosted posts a server by name, unposting on stop.
Names are budget-checked against the 108-byte socket path; a claim
binds+listens at a private temp path and takes the name with atomic
renames under flock (RENAME_NOREPLACE for free names, RENAME_EXCHANGE
grab-verify-commit for stale ones): the registry path is never unlinked
by a claim, live names refuse with AlreadyPosted, foreign files with
NotSocket, and unpost removes only the caller's inode-matched entry.
Watch surfaces inotify overflow and a replaced registry dir.
9ns --mntgen [--mount DIR] -- PROGRAM: one FUSE mount at /mnt/9p whose
synthetic root lists the posted registry (no connection made); a walk
into an unmounted name dials it and runs the existing bridge dispatch
in a per-server worker thread, routed by mount index in the node id's
top bits (ordinals never reused, cap 4096); a dead server answers EIO
on its subtree and is re-dialed on the next walk. The dial watches
stop_fd through Tversion (connectWatched). All existing 9ns forms are
unchanged.
9proc's unix listener no longer blind-unlinks its path: a foreign
non-socket is refused (Occupied), a live server is refused
(AlreadyListening), only a refused socket is cleared, and stop()
unlinks only the listener's own inode-matched socket.
Hardened by adversarial review (GLM 5.3 x2 + DeepSeek V4.1 Flash, all
high-thinking): double-bind races on one name (0 in 180k rounds),
foreign-file TOCTOU deletions (0 in 4M flips), a 255-byte-name listing
panic, inotify queue overflow silently dropped, listenPosted silently
overwriting, dial-time Tversion hangs wedging the dispatcher, --debug
silently ignored in mntgen, and xattr/statx probes answering EPERM on
the synthetic root (broke `ls -l /mnt/9p`).
Tests: root 80/80, 9ns 47/47, 9proc 60/60, integration 88/88 +
mntgen 37/37, adversarial 213/0, freestanding riscv32 gate green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
One upstream 9P connection now serves any number of browser WebSocket
sessions and, with --serve, plain 9P clients over TCP or Unix (a 9pserve-
style frame remux: tags and fids remapped, Tflush forwarded, reconnect on
upstream loss). /fs/<path> maps HTTP onto the tree: GET file or directory
(JSON or HTML), Range, HEAD with 9P headers, PUT (create, truncate, append,
trailing slash makes a directory), DELETE, and ?follow=1 or text/event-stream
turning a blocking read into server-sent events with Tflush on disconnect.
The page gains a lazy tree, stat panel, create, rename, delete, upload and
follow mode. --probe embeds 9proc for self-introspection. New mux-test and
http-fs-test steps; e2e still passes.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
| |
Runner(Backend, Options, Limits) listens on Unix or TCP, runs a reader and a
serve task per connection in one Io.Group, pushes frames into an fs.Server,
and lets the backend answer now or later from any task or thread (reply,
flush, wake); Tflush, greet timeout, connection limit, close and stop with
cancellation are covered by tests over real sockets. The engine and the
backend contract stay Io-free, so the push/step mode for freestanding
targets is unchanged. Documented as the two ways to drive the engine.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
9proc's own fid table, walk loop and dir-read engine are replaced by
cloud9.fs.Server; the tree (static, vars, providers) is served through the
engine's Req/Reply contract with node ids that keep the old qid scheme.
Providers may answer later by returning error.Again (parked in the engine,
retried each step, Tflush -> EINTR); no new files are exposed.
Engine (backward compatible, all opt-in via Backend.features / Options):
create, remove, wstat, reference accounting for backends that count
handles, a salted fid index, name_capacity 0 (names from getattr),
Reply.ename for backend-chosen error text, Attr.path/version/atime.
Engine-level error strings and the 217-byte msize floor now apply to 9proc;
tests updated accordingly.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
The asynchronous 9P file-server engine from the Pardes editor moves into the
library: fid table, walks, directory cursors, a job/slot model where backend
replies arrive later by tag (status again = parked), Tflush cancellation and
orphaned fids on hangup. Allocation-free, no OS calls, no std.Io; comptime
Options (fid, slot, park data, name and user capacities) replace the
editor's constants. Backend contract types (Req, Reply/ReplyWith, Op,
Status, Attr, E, error strings) live here. 29 engine tests plus the client
tests that sat beside it in Pardes.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
as Tflush
- --name NAME (default derived from the transport: socket basename, tcp-IP-PORT,
spawned command, fdN) mounts at /mnt/9p/<name>; --mount still overrides.
ensureMountpoint walks down and creates missing components, shadowing the
deepest unwritable ancestor. NINE_MOUNT is the only exported variable.
- The inode number reported to the kernel is the 9P qid.path for every node,
root included; a server handing qid.path 1 to a file (Pardes /self) no
longer collides with the root.
- FUSE_INTERRUPT for the request in flight becomes Tflush; a blocked read
returns EINTR when the server answers the flush, chunked transfers return
short counts, other requests arriving meanwhile are stashed and served
next. Servers ignoring Tflush still block until they answer.
- 9ns-test now covers nine/bridge/fuse; new adv_bridge_interrupt suite (28);
9ns-itest grows to 88 checks.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
| |
Directories, binaries, build options (-D9ns, -D9proc), step names, module
name (9proc), thread and fs names, env var NINEPLAYER_MOUNT -> NINE_MOUNT,
docs and test scripts. Browser assets move to web/static.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
9player/: FUSE mount CLI that mounts a 9P2000 tree into a fresh user+mount
namespace and runs a program in it (no root, no libfuse, no libc).
introspect/: the 9P debug/introspection library (freestanding core, value
renderers, Linux probe with threads/stacks/memory/breakpoints/panics) and
its demo server. Each has its own build fragment; the root build.zig wires
them behind -D9player/-Dintrospect with namespaced steps (9player-itest,
introspect-check-freestanding, programs-test, ...) and exports the
introspect module for dependents. This is the layout for related programs.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
| | |
|
| |
|