<feed xmlns='http://www.w3.org/2005/Atom'>
<title>cloud9.git/9ns, branch main</title>
<subtitle>9p for zig</subtitle>
<id>https://git.0x4200.cafe/cloud9.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/cloud9.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/'/>
<updated>2026-09-22T14:39:25Z</updated>
<entry>
<title>9ns --mntgen: a server that never answers stalls only its own name</title>
<updated>2026-09-22T14:39:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-22T14:18:05Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=b7fc01550c7bde290cf14276d94193b5b4031dc8'/>
<id>urn:sha1:b7fc01550c7bde290cf14276d94193b5b4031dc8</id>
<content type='text'>
Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it
would sometimes freeze the whole session: no input reached it and nothing
under /mnt/9p answered, until the shell was killed from outside. The cause
was one posted server that accepted a connection and then never spoke 9P —
pardes, answering its 9P from the same loop that was walking its own mount,
was the one on this machine, but any wedged or half-dead server does it.

Three things conspired, and each is fixed on its own:

* The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion,
  Tattach and Tstat on the one thread that reads /dev/fuse, so while that
  server kept quiet no request for any name was read, and no FUSE_INTERRUPT
  either. Now the dispatcher makes a Mount without touching the network and
  queues the walk to the mount's worker, which dials while serving it. The
  dial is the request in flight, so an interrupt of the walk abandons it at
  once (`Session.abort_on_cancel`: nothing to flush before a session
  exists) and the walk answers EINTR; a failed dial leaves the mount
  undialed for the next walk to retry; a full listen backlog (the server
  stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of
  the name goes through the same queue and is answered from the remembered
  root attr, so the dispatcher never holds a session at all.

* Once the dispatcher had read a request the process behind it was
  unkillable (FUSE waits out a request userspace has taken), and an
  INTERRUPT for a request still sitting in a mount's queue was dropped. The
  dispatcher now takes a queued request out and answers EINTR itself, and
  forwards only in-flight ones to the worker; queue and in-flight unique
  are read under the mount's mutex, where the worker moves a request from
  one to the other. A Tflush the server never answers is given 3s
  (`Session.flush_grace_ms`) and then the session is declared wedged: the
  request answers EINTR, the mount dies, the next walk makes a new one.

* The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the
  INIT reply every LOOKUP and READDIR in a directory takes its inode lock,
  so one parked walk held up every other name under /mnt/9p however free
  the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now
  negotiated when the kernel offers it.

What remains is the kernel's own serialization of lookups of one *name*: a
second walker into the parked name waits for the first walk to end, and
only then proceeds (and can be interrupted in its turn).

An adversarial review of the above found three more things, fixed here:
the single-connection bridge's one-slot stash stopped polling the FUSE fd
while a second request was parked, so an INTERRUPT could not arrive (and
parallel dirops make a second request routine) — the stash is now a queue
of copies and the fd is always watched; a dead or wedged mount kept its
socket open until exit, where a late-answering single-threaded server
could block on it — the session is closed when the mount dies; and
teardown after DESTROY or ENODEV (the child still alive, so stop_fd says
nothing) could join a worker parked on a mute server forever — the
sockets are shut down before the join. The flush grace is a deadline now,
not a timer restarted on every wakeup. A black-box run against the binary
(hostile servers: mute, garbage, close-after-accept, full backlog, 100
mute names, interrupt storms, 300 deaths of one server) found that a dead
mount kept its socket, its interrupt pipe and a megabyte of buffers until
exit — three descriptors per death — so `retire` now frees all of it and
keeps only the slot; descriptors, threads and RSS stay flat across 400
deaths. The 4096-slot cap per process remains and is documented.

Reproduced with a socket that accepts and never writes, posted beside
9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked
behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at
once, the parked walker dies of its signal within milliseconds, and a
server that answers the handshake but ignores reads and Tflush releases
its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now
check exactly that; nine.zig gains unit tests for the grace and the
abort. Also in this change: the uncommitted ESTALE-on-death and
FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount
path here builds on.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>9ns --mntgen: registry subdirectories are mount points too</title>
<updated>2026-09-21T18:20:27Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T17:23:27Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=0d7e295efee1fca0935cf4a8bee9629c007dd2b6'/>
<id>urn:sha1:0d7e295efee1fca0935cf4a8bee9629c007dd2b6</id>
<content type='text'>
A registry entry that is a directory is now served the way the root is:
a synthetic directory listing the real one, dialing the sockets inside
it on walk and recursing into further directories, to max_synth_depth
(8) levels across max_synth_dirs (64) synthetic nodes. That is the
plan9port mntgen shape and the layout zmx now posts under, so a live
session reads at /mnt/9p/zmx/&lt;name&gt;. Before this a directory in the
registry was dialed like a socket and answered EIO for good.

post gains the two entry points the traversal needs: postedDir (the
registry scan, against any directory) and dialPath (a dial by composed
path, no name validation).

Hardening, each from an attack that broke the code:

- BATCH_FORGET carries entries for many owners and puts 0 in the header
  nodeid, so routing it by the header dropped all of them: 32 of 64
  synthetic slots leaked in one close burst and the subdirectories that
  held them answered EIO forever. distributeForgets unpacks the body and
  hands each entry to its owner.
- probe() and connectBlocking() copied a caller's path into the kernel
  address with no bound: a path past sun_path overran the 110-byte stack
  sockaddr (a panic in Debug, silent corruption in ReleaseFast). Both
  refuse it now, probe as `.live` so a claim never deletes what it could
  not inspect.
- That bound then caught 9proc's own listener, which handed probe() the
  whole 108-byte sun_path array instead of the path inside it. The probe
  reads `.live` for anything it cannot ask about, so every stale socket
  became AlreadyListening and no server could ever take a dead
  predecessor's name back. It passes the path now.

Suites: 87/87 root (+7 post/serve attack regressions), 48/48 9ns,
51+88 9ns integration (+4 traversal and slot-recycling checks), 213/0
9ns adversarial, 60/60 9proc plus its adversarial suites with a new
stale-socket takeover check, freestanding green.
</content>
</entry>
<entry>
<title>post registry + 9ns --mntgen: the /srv translation</title>
<updated>2026-09-21T17:13:43Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T17:13:43Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=3a23f6a29e47ace901bd4d82b9db4055fcc12bb9'/>
<id>urn:sha1:3a23f6a29e47ace901bd4d82b9db4055fcc12bb9</id>
<content type='text'>
cloud9.post: servers post their socket under a name in
$XDG_RUNTIME_DIR/9p (post/unpost, posted, dial, Watch) and
serve.Runner.listenPosted posts a server by name, unposting on stop.
Names are budget-checked against the 108-byte socket path; a claim
binds+listens at a private temp path and takes the name with atomic
renames under flock (RENAME_NOREPLACE for free names, RENAME_EXCHANGE
grab-verify-commit for stale ones): the registry path is never unlinked
by a claim, live names refuse with AlreadyPosted, foreign files with
NotSocket, and unpost removes only the caller's inode-matched entry.
Watch surfaces inotify overflow and a replaced registry dir.

9ns --mntgen [--mount DIR] -- PROGRAM: one FUSE mount at /mnt/9p whose
synthetic root lists the posted registry (no connection made); a walk
into an unmounted name dials it and runs the existing bridge dispatch
in a per-server worker thread, routed by mount index in the node id's
top bits (ordinals never reused, cap 4096); a dead server answers EIO
on its subtree and is re-dialed on the next walk. The dial watches
stop_fd through Tversion (connectWatched). All existing 9ns forms are
unchanged.

9proc's unix listener no longer blind-unlinks its path: a foreign
non-socket is refused (Occupied), a live server is refused
(AlreadyListening), only a refused socket is cleared, and stop()
unlinks only the listener's own inode-matched socket.

Hardened by adversarial review (GLM 5.3 x2 + DeepSeek V4.1 Flash, all
high-thinking): double-bind races on one name (0 in 180k rounds),
foreign-file TOCTOU deletions (0 in 4M flips), a 255-byte-name listing
panic, inotify queue overflow silently dropped, listenPosted silently
overwriting, dial-time Tversion hangs wedging the dispatcher, --debug
silently ignored in mntgen, and xattr/statx probes answering EPERM on
the synthetic root (broke `ls -l /mnt/9p`).

Tests: root 80/80, 9ns 47/47, 9proc 60/60, integration 88/88 +
mntgen 37/37, adversarial 213/0, freestanding riscv32 gate green.
</content>
</entry>
<entry>
<title>9ns: --name and /mnt/9p/&lt;name&gt; mounts, qid.path as inode number, interrupts as Tflush</title>
<updated>2026-09-20T02:55:47Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-20T02:55:47Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=3e9f8805f293f622bb885cf849b5ce47dc062ad1'/>
<id>urn:sha1:3e9f8805f293f622bb885cf849b5ce47dc062ad1</id>
<content type='text'>
- --name NAME (default derived from the transport: socket basename, tcp-IP-PORT,
  spawned command, fdN) mounts at /mnt/9p/&lt;name&gt;; --mount still overrides.
  ensureMountpoint walks down and creates missing components, shadowing the
  deepest unwritable ancestor. NINE_MOUNT is the only exported variable.
- The inode number reported to the kernel is the 9P qid.path for every node,
  root included; a server handing qid.path 1 to a file (Pardes /self) no
  longer collides with the root.
- FUSE_INTERRUPT for the request in flight becomes Tflush; a blocked read
  returns EINTR when the server answers the flush, chunked transfers return
  short counts, other requests arriving meanwhile are stashed and served
  next. Servers ignoring Tflush still block until they answer.
- 9ns-test now covers nine/bridge/fuse; new adv_bridge_interrupt suite (28);
  9ns-itest grows to 88 checks.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>Rename programs: 9player -&gt; 9ns, introspect -&gt; 9proc, app -&gt; web (9web)</title>
<updated>2026-09-20T02:28:22Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-20T02:28:22Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=ba996acfcad1698adbf4a1834fe50e73b1c6cab9'/>
<id>urn:sha1:ba996acfcad1698adbf4a1834fe50e73b1c6cab9</id>
<content type='text'>
Directories, binaries, build options (-D9ns, -D9proc), step names, module
name (9proc), thread and fs names, env var NINEPLAYER_MOUNT -&gt; NINE_MOUNT,
docs and test scripts. Browser assets move to web/static.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
</feed>
