diff options
| author | Gabriel Schneider <[email protected]> | 2026-09-22 11:18:05 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-09-22 11:39:25 -0300 |
| commit | b7fc01550c7bde290cf14276d94193b5b4031dc8 (patch) | |
| tree | 696e8fbf26c819b28ce5552df2e7dc943a6b2a8e /9ns/docs/DESIGN.md | |
| parent | 1f3aff78702b65c328384bc5b422c448751e809b (diff) | |
| download | cloud9-b7fc01550c7bde290cf14276d94193b5b4031dc8.tar.gz cloud9-b7fc01550c7bde290cf14276d94193b5b4031dc8.zip | |
9ns --mntgen: a server that never answers stalls only its own name
Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it
would sometimes freeze the whole session: no input reached it and nothing
under /mnt/9p answered, until the shell was killed from outside. The cause
was one posted server that accepted a connection and then never spoke 9P —
pardes, answering its 9P from the same loop that was walking its own mount,
was the one on this machine, but any wedged or half-dead server does it.
Three things conspired, and each is fixed on its own:
* The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion,
Tattach and Tstat on the one thread that reads /dev/fuse, so while that
server kept quiet no request for any name was read, and no FUSE_INTERRUPT
either. Now the dispatcher makes a Mount without touching the network and
queues the walk to the mount's worker, which dials while serving it. The
dial is the request in flight, so an interrupt of the walk abandons it at
once (`Session.abort_on_cancel`: nothing to flush before a session
exists) and the walk answers EINTR; a failed dial leaves the mount
undialed for the next walk to retry; a full listen backlog (the server
stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of
the name goes through the same queue and is answered from the remembered
root attr, so the dispatcher never holds a session at all.
* Once the dispatcher had read a request the process behind it was
unkillable (FUSE waits out a request userspace has taken), and an
INTERRUPT for a request still sitting in a mount's queue was dropped. The
dispatcher now takes a queued request out and answers EINTR itself, and
forwards only in-flight ones to the worker; queue and in-flight unique
are read under the mount's mutex, where the worker moves a request from
one to the other. A Tflush the server never answers is given 3s
(`Session.flush_grace_ms`) and then the session is declared wedged: the
request answers EINTR, the mount dies, the next walk makes a new one.
* The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the
INIT reply every LOOKUP and READDIR in a directory takes its inode lock,
so one parked walk held up every other name under /mnt/9p however free
the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now
negotiated when the kernel offers it.
What remains is the kernel's own serialization of lookups of one *name*: a
second walker into the parked name waits for the first walk to end, and
only then proceeds (and can be interrupted in its turn).
An adversarial review of the above found three more things, fixed here:
the single-connection bridge's one-slot stash stopped polling the FUSE fd
while a second request was parked, so an INTERRUPT could not arrive (and
parallel dirops make a second request routine) — the stash is now a queue
of copies and the fd is always watched; a dead or wedged mount kept its
socket open until exit, where a late-answering single-threaded server
could block on it — the session is closed when the mount dies; and
teardown after DESTROY or ENODEV (the child still alive, so stop_fd says
nothing) could join a worker parked on a mute server forever — the
sockets are shut down before the join. The flush grace is a deadline now,
not a timer restarted on every wakeup. A black-box run against the binary
(hostile servers: mute, garbage, close-after-accept, full backlog, 100
mute names, interrupt storms, 300 deaths of one server) found that a dead
mount kept its socket, its interrupt pipe and a megabyte of buffers until
exit — three descriptors per death — so `retire` now frees all of it and
keeps only the slot; descriptors, threads and RSS stay flat across 400
deaths. The 4096-slot cap per process remains and is documented.
Reproduced with a socket that accepts and never writes, posted beside
9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked
behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at
once, the parked walker dies of its signal within milliseconds, and a
server that answers the handshake but ignores reads and Tflush releases
its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now
check exactly that; nine.zig gains unit tests for the grace and the
abort. Also in this change: the uncommitted ESTALE-on-death and
FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount
path here builds on.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Diffstat (limited to '9ns/docs/DESIGN.md')
| -rw-r--r-- | 9ns/docs/DESIGN.md | 180 |
1 files changed, 110 insertions, 70 deletions
diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md index 6b1b80c..32c548c 100644 --- a/9ns/docs/DESIGN.md +++ b/9ns/docs/DESIGN.md @@ -144,8 +144,8 @@ listing that needs one meanwhile answers EIO rather than waiting. in new userns │ /mnt/9p ─FUSE─▶ kernel ─▶ │ dispatcher (main thread) alpha/ beta/ │ ├─ synthetic root (node 1): lists the registry - ...each a server │ ├─ LOOKUP(alpha) ── dial+attach+stat ─▶ worker 1 ─ bridge ─ 9P session ─ server alpha - │ └─ LOOKUP(beta) ── dial+attach+stat ─▶ worker 2 ─ bridge ─ 9P session ─ server beta + ...each a server │ ├─ LOOKUP(alpha) ─▶ worker 1 ─ dial+attach+stat, then bridge ─ 9P session ─ server alpha + │ └─ LOOKUP(beta) ─▶ worker 2 ─ dial+attach+stat, then bridge ─ 9P session ─ server beta ``` Threads and node ids: @@ -153,17 +153,30 @@ Threads and node ids: * **Dispatcher** (the main thread) is the only reader of `/dev/fuse`. It answers INIT/DESTROY, serves the **synthetic root** (node 1) itself, and routes every other request by the node id's top bits — the **mount - index** — to the owning server's mount; a FUSE_INTERRUPT is routed by - scanning the mounts for the one currently serving the interrupted unique - and dropping the target into its interrupt pipe. -* A **mount** is a dialed server: its own `nine.Session`, its own bridge - state (inode table, handles — the ordinary single-server translation, - unchanged) and a **worker thread**. The worker pops copied requests off a - queue and serves them one at a time, exactly the single-connection - contract; the dispatcher keeps reading `/dev/fuse` meanwhile, so one slow - server never blocks the other names. Replies go straight back on the FUSE - fd (one `writev` per reply; the kernel processes each write as one - message). + index** — to the owning server's mount. It waits on no server, ever: + the only blocking it does is the registry directory (a tmpfs) and the + mount queues' locks. A FUSE_INTERRUPT is routed by scanning the mounts: + a request still sitting in a mount's queue is taken out and answered + `EINTR` on the spot (the kernel sends an INTERRUPT once, and a request + that only runs later would otherwise run to the end with nobody left + wanting it); one in flight gets its unique dropped into that mount's + interrupt pipe. Queue and in-flight unique are read under the mount's + mutex, which is also where the worker moves a request from one to the + other, so an interrupt cannot fall between them. +* A **mount** is a posted name walked into: the socket to dial, its own + `nine.Session` and bridge state (inode table, handles — the ordinary + single-server translation, unchanged) once dialed, and a **worker + thread** that owns all of it. The dispatcher makes a mount without + touching the network and hands it the walk; the worker dials while + serving that first request. It pops copied requests off a queue and + serves them one at a time, exactly the single-connection contract; the + dispatcher keeps reading `/dev/fuse` meanwhile, so one slow server never + blocks the other names. Replies go straight back on the FUSE fd (one + `writev` per reply; the kernel processes each write as one message). + The queue is a plain list under a mutex and condition rather than an + `std.Io.Queue` because an interrupt has to find and remove a request by + unique from the middle of it — the same reason a 9P server keeps its + pending requests in a list a Tflush can search. * Node id layout: `nodeid = (mount_index << 32) | local`. Index 0 is the synthetic root; per-mount local ids start at 1 (the server's 9P root) and never exceed 2^32 (a bridge never reuses one). Mount indexes are @@ -172,40 +185,61 @@ Threads and node ids: inode of its replacement. Reported `st_ino` mixes the index into the qid.path (`(index+1) << 48` XOR), so two servers handing out the same qid.path (two ramfs instances) still get distinct inode numbers. -* **Lazy dial**: LOOKUP of an unmounted name checks the registry, dials, - attaches, stats the root and spawns the worker — all on the dispatcher - thread, in service of the walk that triggered it. No eager connection is - ever made: `ls` of the root reads the registry directory only (a plain - file dropped there is listed too — and yields EIO on the walk, never - deleted). A walk into a **stale** entry (socket present, connect refused) - answers EIO. The whole dial watches `stop_fd` (the session is built with - `nine.Session.connectWatched`, so the `Tversion` exchange is covered too): - when the program exits while a walk is parked in a dial, the dial fails - with `Stopped`, the pending LOOKUP answers EIO and 9ns follows the program - out. There is no dial timeout of our own (a slow server delays the walk, - like it would delay any 9P client), and a server that accepts but never - answers `Tversion` still parks the dispatcher until the program exits — - including the unkillable corner where the *blocked walk itself* is the - only thing keeping the program alive (the task sits in D state until the - filesystem answers; a same-user self-DoS, accepted with the pinned - "dispatcher dials, no concurrent dial" design). +* **Lazy dial, on the worker**: LOOKUP of an unmounted name stats the + registry entry (dispatcher), makes the mount and queues the LOOKUP to it; + the worker dials (connect, `Tversion`, `Tattach`, `Tstat` of the root) as + the first thing it does for that request, then answers it from the root + stat. Every later LOOKUP of the name is queued the same way and answered + from the remembered root attr, so the dispatcher never holds a session. + No eager connection is ever made: `ls` of the root reads the registry + directory only (a plain file dropped there is listed too — and yields EIO + on the walk, never deleted). A dial that fails answers the walk that + asked — `ENOENT` when the entry vanished, `EIO` for a stale entry + (connect refused), a full backlog or a server that will not speak 9P — + and leaves the mount undialed, so the next walk simply tries again. + The dial is the request in flight, so it ends the way any request does: + `stop_fd` (the program exited) fails it with `Stopped`; an interrupt of + the walk abandons it at once, with nothing sent (`abort_on_cancel`: there + is no session yet to flush anything out of) and the walk answers + `EINTR`. A server that accepts but never answers `Tversion` therefore + costs exactly the walks into its own name, and Ctrl-C ends those. + A server whose listen backlog is full (it stopped accepting) makes the + connect report `EAGAIN`; that is retried for 5s, polling `stop_fd` and + the interrupt pipe between tries, then answers `EIO`. What no design can + fix is the kernel side: the VFS serializes lookups of one *name*, so a + second walker into the parked name waits in `d_wait_lookup` until the + first walk ends — interrupt the first, and the second proceeds (and can + be interrupted in its turn). +* **FUSE_PARALLEL_DIROPS** is negotiated in the INIT reply. Without it the + kernel takes the directory inode's lock around every LOOKUP and READDIR + in it, so one parked walk would still hold up every other name under the + same directory — the whole registry root, for a mntgen mount — however + free the dispatcher is. * **Death and re-dial**: when a worker's session dies mid-request, dispatch - has already answered that request EIO, the mount is marked dead, and - everything further routed to that subtree answers EIO (a FORGET is - dropped). Nothing reconnects eagerly. Because synthetic-root entries are - served with zero entry-validity, the next walk into the name LOOKUPs it - again; a dead mount is skipped and the name is dialed afresh — a new - mount under a new index, so kernel-held inodes of the corpse keep - answering EIO until forgotten. `ls` still lists the dead name (listing - connects to nothing). Death is discovered lazily: the first walk after a - silent death answers EIO (it marks the mount dead), the next walk re-dials. + has already answered that request (`ESTALE` for a lost connection, so the + VFS redoes the path walk instead of failing; `EIO` for a protocol error), + the mount is marked dead, its entry is invalidated in the kernel's dentry + cache (`FUSE_NOTIFY_INVAL_ENTRY`, best effort), and everything further + routed to that subtree answers `ESTALE` (a FORGET is dropped). Nothing + reconnects eagerly. The next walk into the name LOOKUPs it again; a dead + mount is skipped and the name gets a new mount under a new index, so + kernel-held inodes of the corpse keep answering `ESTALE` until forgotten. + `ls` still lists the dead name (listing connects to nothing). Death is + discovered lazily: the first walk after a silent death takes the error + (it marks the mount dead), the next walk re-dials. * **Interrupts** work per mount: the worker's session polls the mount's interrupt pipe while a 9P reply is outstanding; the dispatcher writes the interrupted request's unique into it and the usual `Tflush` dance - (see *Interrupts*) follows. `stop_fd` (the child's death) is watched by - every session, so no worker can stay blocked on a hung server past the - program's exit; teardown wakes every worker, joins them, and tears down - their sessions and bridge state. + (see *Interrupts*) follows — with a grace: a server that answers neither + the request nor the `Tflush` within 3s (`nine.Session.flush_grace_ms`) + is declared gone, the request answers `EINTR`, the session is wedged and + the mount dies with it (the next walk makes a new one). The protocol + says a client waits for the Rflush; a server that has not managed one in + that long is not going to, and the process behind the interrupt is + unkillable until we stop waiting. `stop_fd` (the child's death) is + watched by every session, so no worker can stay blocked on a hung server + past the program's exit; teardown wakes every worker, joins them, and + tears down their sessions and bridge state. * The synthetic root is read-only (`dr-xr-xr-x`, like `/srv`): services are posted and unposted by their servers (`cloud9.post`'s `post`/`listenPosted`/`unpost`), not created and removed through files. @@ -290,7 +324,8 @@ setupmapping = 48, removemapping = 49, syncfs = 50, tmpfile = 51, statx = 52, _ Constants: `kernel_version = 7`, `kernel_minor = 31` (what we answer; the kernel adapts to the lower minor), `FOPEN_DIRECT_IO = 1`, `FOPEN_KEEP_CACHE = 2`, -`FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_MAX_PAGES = 1<<22`, +`FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_PARALLEL_DIROPS = 1<<18`, +`FUSE_MAX_PAGES = 1<<22`, `FATTR_MODE=1, FATTR_UID=2, FATTR_GID=4, FATTR_SIZE=8, FATTR_ATIME=16, FATTR_MTIME=32, FATTR_FH=64, FATTR_ATIME_NOW=128, FATTR_MTIME_NOW=256, FATTR_LOCKOWNER=512, FATTR_CTIME=1024`. `root_id = 1`. @@ -409,7 +444,7 @@ pub fn serve(gpa: std.mem.Allocator, fuse_fd: i32, nine: *nine.Session, root_fid // mntgen (see "mntgen: one mount, many servers" under Process model): pub const MntgenOptions = struct { - io: std.Io, // dispatcher-thread only: post.posted / post.dial + io: std.Io, // dispatcher-thread only: post.posted / registry stats env: post.Env, // XDG_RUNTIME_DIR names the registry uname: []const u8, aname: []const u8 = "", msize: u32 = 131072, }; @@ -449,7 +484,7 @@ Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod): | FUSE | 9P | |---|---| -| INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256 if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open | +| INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256, and FUSE_PARALLEL_DIROPS, each if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open | | LOOKUP(parent,name) | `walk(parent.fid → newfid, [name])`; `stat(newfid)`; dedupe by qid; `EntryOut` | | FORGET / BATCH_FORGET | `nlookup -= n`; at 0 `clunk` and drop (no reply) | | GETATTR | `stat(inode.fid)` → `AttrOut` | @@ -512,9 +547,10 @@ is outstanding, including the initial root stat: same errno through the ename table). An INTERRUPT for any other unique is consumed and dropped (the kernel expects no reply). Any other request (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is - stashed in a one-slot queue that `serve` dispatches, after swapping the two - buffers, before it polls again; while the slot is full `watch` returns -1, - so a second one cannot arrive. + copied into a queue that `serve` dispatches, in arrival order, before it + polls again; the fd stays watched throughout, so an INTERRUPT is never + stuck behind a parked request (a copy that cannot be allocated answers + `ENOMEM` on the spot). * The INTERRUPT applies to `cur_unique` only and is consumed when read: the clunks that unwind a half-done lookup/create/mkdir after an `Interrupted` walk or stat are ordinary rpcs and are not re-interrupted by it. An @@ -532,11 +568,14 @@ is outstanding, including the initial root stat: included) as much as for a caught one, and then waits for the reply; our `EINTR` is what finally lets the killed task die. -Limits: a server that ignores Tflush still blocks the mount until it answers -(the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and -while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives -after another process's request was parked is seen only once the blocked -request completes (a multi-slot stash would lift that). +Limits: a server that ignores Tflush is given `flush_grace_ms` (3s) and then +declared gone — flush(5) says the server "should answer the Tflush message +immediately", and a client that waits longer leaves an unkillable process +behind — so the reader gets `EINTR` and the session ends (the hostile +`never` mode; in single-connection mode that is the mount, in mntgen that +one name). A slow-but-alive server that cannot answer a flush inside its own +blocking read pays the same price; the alternative was a hang that only +SIGKILL of 9ns could end. ### `src/ns.zig` — namespace and process plumbing @@ -676,26 +715,27 @@ ReleaseSafe. (event files) stalls that server's subtree while it is outstanding (but not past the child's exit). It can be interrupted: killing or Ctrl-C-ing the reader sends `FUSE_INTERRUPT`, which becomes `Tflush`; servers that - honour it unblock immediately, servers that don't still block that - subtree until they answer. -* mntgen: the dispatcher dials on the main thread, in service of the walk - that triggered it: a server that accepts the connection but never - answers `Tversion` parks the whole mount for as long as the walk's - program keeps running (as it would delay any 9P client). The dial - watches `stop_fd`, so the program exiting ends it (`Stopped`, EIO to the - pending LOOKUP); no concurrent dial, no dial timeout of our own. The - remaining corner is a same-user self-DoS: the program blocked *in that - very walk* sits in D state until the filesystem answers and cannot be - killed to fire `stop_fd` — only killing the mute server (EOF) ends it. + honour it unblock immediately, servers that don't are given 3s and then + declared gone: the reader gets `EINTR` and the session (single + connection: the mount; mntgen: that name) ends. +* mntgen: a name's dial runs on that name's worker, inside the walk that + asked, with no timeout of its own beyond the 5s backlog retry: a server + that accepts and never answers `Tversion` parks the walks into its own + name, and only those, until each is interrupted or the program exits. + The one wait nothing here can cut short is the kernel's own + serialization of lookups of a single name (a second walker into the + parked name waits for the first walk to end, uninterruptibly). * mntgen: mount indexes are ordinals and never reused (cap 4096 dials per process); per-mount node ids cap at 2^32 lookups. `st_ino` mixes the index with the qid.path via XOR — collisions remain theoretically possible, just not the practical ones (identical qid.paths across two - servers). A dead mount keeps its slot — and its session socket and - interrupt-pipe descriptors — until process exit (early close would race - the dispatcher's INTERRUPT scan against fd reuse), so with a small - `ulimit -n` a re-dial storm exhausts descriptors before the index cap; - dials then fail cleanly with EIO. + servers). A dead mount releases its socket, its interrupt pipe and its + bridge state at once (the pipe under the mount's mutex, where the + dispatcher writes it) and keeps only its slot, so a server that dies and + comes back costs one index per death and nothing else; after 4096 of + them in one 9ns process no further name can be walked into (EIO) until + the shell is restarted. Reclaiming a slot once the kernel has forgotten + every node of the corpse is the next step, not taken here. * mntgen: no posting/unposting through the mount (the synthetic root is read-only; servers manage their registry entries through `cloud9.post`), and no per-name mount options: one set of |
