summaryrefslogtreecommitdiff
path: root/9ns/docs/DESIGN.md
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-09-22 11:18:05 -0300
committerGabriel Schneider <[email protected]>2026-09-22 11:39:25 -0300
commitb7fc01550c7bde290cf14276d94193b5b4031dc8 (patch)
tree696e8fbf26c819b28ce5552df2e7dc943a6b2a8e /9ns/docs/DESIGN.md
parent1f3aff78702b65c328384bc5b422c448751e809b (diff)
downloadcloud9-b7fc01550c7bde290cf14276d94193b5b4031dc8.tar.gz
cloud9-b7fc01550c7bde290cf14276d94193b5b4031dc8.zip
9ns --mntgen: a server that never answers stalls only its own name
Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it would sometimes freeze the whole session: no input reached it and nothing under /mnt/9p answered, until the shell was killed from outside. The cause was one posted server that accepted a connection and then never spoke 9P — pardes, answering its 9P from the same loop that was walking its own mount, was the one on this machine, but any wedged or half-dead server does it. Three things conspired, and each is fixed on its own: * The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion, Tattach and Tstat on the one thread that reads /dev/fuse, so while that server kept quiet no request for any name was read, and no FUSE_INTERRUPT either. Now the dispatcher makes a Mount without touching the network and queues the walk to the mount's worker, which dials while serving it. The dial is the request in flight, so an interrupt of the walk abandons it at once (`Session.abort_on_cancel`: nothing to flush before a session exists) and the walk answers EINTR; a failed dial leaves the mount undialed for the next walk to retry; a full listen backlog (the server stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of the name goes through the same queue and is answered from the remembered root attr, so the dispatcher never holds a session at all. * Once the dispatcher had read a request the process behind it was unkillable (FUSE waits out a request userspace has taken), and an INTERRUPT for a request still sitting in a mount's queue was dropped. The dispatcher now takes a queued request out and answers EINTR itself, and forwards only in-flight ones to the worker; queue and in-flight unique are read under the mount's mutex, where the worker moves a request from one to the other. A Tflush the server never answers is given 3s (`Session.flush_grace_ms`) and then the session is declared wedged: the request answers EINTR, the mount dies, the next walk makes a new one. * The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the INIT reply every LOOKUP and READDIR in a directory takes its inode lock, so one parked walk held up every other name under /mnt/9p however free the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now negotiated when the kernel offers it. What remains is the kernel's own serialization of lookups of one *name*: a second walker into the parked name waits for the first walk to end, and only then proceeds (and can be interrupted in its turn). An adversarial review of the above found three more things, fixed here: the single-connection bridge's one-slot stash stopped polling the FUSE fd while a second request was parked, so an INTERRUPT could not arrive (and parallel dirops make a second request routine) — the stash is now a queue of copies and the fd is always watched; a dead or wedged mount kept its socket open until exit, where a late-answering single-threaded server could block on it — the session is closed when the mount dies; and teardown after DESTROY or ENODEV (the child still alive, so stop_fd says nothing) could join a worker parked on a mute server forever — the sockets are shut down before the join. The flush grace is a deadline now, not a timer restarted on every wakeup. A black-box run against the binary (hostile servers: mute, garbage, close-after-accept, full backlog, 100 mute names, interrupt storms, 300 deaths of one server) found that a dead mount kept its socket, its interrupt pipe and a megabyte of buffers until exit — three descriptors per death — so `retire` now frees all of it and keeps only the slot; descriptors, threads and RSS stay flat across 400 deaths. The 4096-slot cap per process remains and is documented. Reproduced with a socket that accepts and never writes, posted beside 9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at once, the parked walker dies of its signal within milliseconds, and a server that answers the handshake but ignores reads and Tflush releases its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now check exactly that; nine.zig gains unit tests for the grace and the abort. Also in this change: the uncommitted ESTALE-on-death and FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount path here builds on. Co-Authored-By: Claude Fable 5.1 <[email protected]>
Diffstat (limited to '9ns/docs/DESIGN.md')
-rw-r--r--9ns/docs/DESIGN.md180
1 files changed, 110 insertions, 70 deletions
diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md
index 6b1b80c..32c548c 100644
--- a/9ns/docs/DESIGN.md
+++ b/9ns/docs/DESIGN.md
@@ -144,8 +144,8 @@ listing that needs one meanwhile answers EIO rather than waiting.
in new userns │
/mnt/9p ─FUSE─▶ kernel ─▶ │ dispatcher (main thread)
alpha/ beta/ │ ├─ synthetic root (node 1): lists the registry
- ...each a server │ ├─ LOOKUP(alpha) ── dial+attach+stat ─▶ worker 1 ─ bridge ─ 9P session ─ server alpha
- │ └─ LOOKUP(beta) ── dial+attach+stat ─▶ worker 2 ─ bridge ─ 9P session ─ server beta
+ ...each a server │ ├─ LOOKUP(alpha) ─▶ worker 1 ─ dial+attach+stat, then bridge ─ 9P session ─ server alpha
+ │ └─ LOOKUP(beta) ─▶ worker 2 ─ dial+attach+stat, then bridge ─ 9P session ─ server beta
```
Threads and node ids:
@@ -153,17 +153,30 @@ Threads and node ids:
* **Dispatcher** (the main thread) is the only reader of `/dev/fuse`. It
answers INIT/DESTROY, serves the **synthetic root** (node 1) itself, and
routes every other request by the node id's top bits — the **mount
- index** — to the owning server's mount; a FUSE_INTERRUPT is routed by
- scanning the mounts for the one currently serving the interrupted unique
- and dropping the target into its interrupt pipe.
-* A **mount** is a dialed server: its own `nine.Session`, its own bridge
- state (inode table, handles — the ordinary single-server translation,
- unchanged) and a **worker thread**. The worker pops copied requests off a
- queue and serves them one at a time, exactly the single-connection
- contract; the dispatcher keeps reading `/dev/fuse` meanwhile, so one slow
- server never blocks the other names. Replies go straight back on the FUSE
- fd (one `writev` per reply; the kernel processes each write as one
- message).
+ index** — to the owning server's mount. It waits on no server, ever:
+ the only blocking it does is the registry directory (a tmpfs) and the
+ mount queues' locks. A FUSE_INTERRUPT is routed by scanning the mounts:
+ a request still sitting in a mount's queue is taken out and answered
+ `EINTR` on the spot (the kernel sends an INTERRUPT once, and a request
+ that only runs later would otherwise run to the end with nobody left
+ wanting it); one in flight gets its unique dropped into that mount's
+ interrupt pipe. Queue and in-flight unique are read under the mount's
+ mutex, which is also where the worker moves a request from one to the
+ other, so an interrupt cannot fall between them.
+* A **mount** is a posted name walked into: the socket to dial, its own
+ `nine.Session` and bridge state (inode table, handles — the ordinary
+ single-server translation, unchanged) once dialed, and a **worker
+ thread** that owns all of it. The dispatcher makes a mount without
+ touching the network and hands it the walk; the worker dials while
+ serving that first request. It pops copied requests off a queue and
+ serves them one at a time, exactly the single-connection contract; the
+ dispatcher keeps reading `/dev/fuse` meanwhile, so one slow server never
+ blocks the other names. Replies go straight back on the FUSE fd (one
+ `writev` per reply; the kernel processes each write as one message).
+ The queue is a plain list under a mutex and condition rather than an
+ `std.Io.Queue` because an interrupt has to find and remove a request by
+ unique from the middle of it — the same reason a 9P server keeps its
+ pending requests in a list a Tflush can search.
* Node id layout: `nodeid = (mount_index << 32) | local`. Index 0 is the
synthetic root; per-mount local ids start at 1 (the server's 9P root) and
never exceed 2^32 (a bridge never reuses one). Mount indexes are
@@ -172,40 +185,61 @@ Threads and node ids:
inode of its replacement. Reported `st_ino` mixes the index into the
qid.path (`(index+1) << 48` XOR), so two servers handing out the same
qid.path (two ramfs instances) still get distinct inode numbers.
-* **Lazy dial**: LOOKUP of an unmounted name checks the registry, dials,
- attaches, stats the root and spawns the worker — all on the dispatcher
- thread, in service of the walk that triggered it. No eager connection is
- ever made: `ls` of the root reads the registry directory only (a plain
- file dropped there is listed too — and yields EIO on the walk, never
- deleted). A walk into a **stale** entry (socket present, connect refused)
- answers EIO. The whole dial watches `stop_fd` (the session is built with
- `nine.Session.connectWatched`, so the `Tversion` exchange is covered too):
- when the program exits while a walk is parked in a dial, the dial fails
- with `Stopped`, the pending LOOKUP answers EIO and 9ns follows the program
- out. There is no dial timeout of our own (a slow server delays the walk,
- like it would delay any 9P client), and a server that accepts but never
- answers `Tversion` still parks the dispatcher until the program exits —
- including the unkillable corner where the *blocked walk itself* is the
- only thing keeping the program alive (the task sits in D state until the
- filesystem answers; a same-user self-DoS, accepted with the pinned
- "dispatcher dials, no concurrent dial" design).
+* **Lazy dial, on the worker**: LOOKUP of an unmounted name stats the
+ registry entry (dispatcher), makes the mount and queues the LOOKUP to it;
+ the worker dials (connect, `Tversion`, `Tattach`, `Tstat` of the root) as
+ the first thing it does for that request, then answers it from the root
+ stat. Every later LOOKUP of the name is queued the same way and answered
+ from the remembered root attr, so the dispatcher never holds a session.
+ No eager connection is ever made: `ls` of the root reads the registry
+ directory only (a plain file dropped there is listed too — and yields EIO
+ on the walk, never deleted). A dial that fails answers the walk that
+ asked — `ENOENT` when the entry vanished, `EIO` for a stale entry
+ (connect refused), a full backlog or a server that will not speak 9P —
+ and leaves the mount undialed, so the next walk simply tries again.
+ The dial is the request in flight, so it ends the way any request does:
+ `stop_fd` (the program exited) fails it with `Stopped`; an interrupt of
+ the walk abandons it at once, with nothing sent (`abort_on_cancel`: there
+ is no session yet to flush anything out of) and the walk answers
+ `EINTR`. A server that accepts but never answers `Tversion` therefore
+ costs exactly the walks into its own name, and Ctrl-C ends those.
+ A server whose listen backlog is full (it stopped accepting) makes the
+ connect report `EAGAIN`; that is retried for 5s, polling `stop_fd` and
+ the interrupt pipe between tries, then answers `EIO`. What no design can
+ fix is the kernel side: the VFS serializes lookups of one *name*, so a
+ second walker into the parked name waits in `d_wait_lookup` until the
+ first walk ends — interrupt the first, and the second proceeds (and can
+ be interrupted in its turn).
+* **FUSE_PARALLEL_DIROPS** is negotiated in the INIT reply. Without it the
+ kernel takes the directory inode's lock around every LOOKUP and READDIR
+ in it, so one parked walk would still hold up every other name under the
+ same directory — the whole registry root, for a mntgen mount — however
+ free the dispatcher is.
* **Death and re-dial**: when a worker's session dies mid-request, dispatch
- has already answered that request EIO, the mount is marked dead, and
- everything further routed to that subtree answers EIO (a FORGET is
- dropped). Nothing reconnects eagerly. Because synthetic-root entries are
- served with zero entry-validity, the next walk into the name LOOKUPs it
- again; a dead mount is skipped and the name is dialed afresh — a new
- mount under a new index, so kernel-held inodes of the corpse keep
- answering EIO until forgotten. `ls` still lists the dead name (listing
- connects to nothing). Death is discovered lazily: the first walk after a
- silent death answers EIO (it marks the mount dead), the next walk re-dials.
+ has already answered that request (`ESTALE` for a lost connection, so the
+ VFS redoes the path walk instead of failing; `EIO` for a protocol error),
+ the mount is marked dead, its entry is invalidated in the kernel's dentry
+ cache (`FUSE_NOTIFY_INVAL_ENTRY`, best effort), and everything further
+ routed to that subtree answers `ESTALE` (a FORGET is dropped). Nothing
+ reconnects eagerly. The next walk into the name LOOKUPs it again; a dead
+ mount is skipped and the name gets a new mount under a new index, so
+ kernel-held inodes of the corpse keep answering `ESTALE` until forgotten.
+ `ls` still lists the dead name (listing connects to nothing). Death is
+ discovered lazily: the first walk after a silent death takes the error
+ (it marks the mount dead), the next walk re-dials.
* **Interrupts** work per mount: the worker's session polls the mount's
interrupt pipe while a 9P reply is outstanding; the dispatcher writes the
interrupted request's unique into it and the usual `Tflush` dance
- (see *Interrupts*) follows. `stop_fd` (the child's death) is watched by
- every session, so no worker can stay blocked on a hung server past the
- program's exit; teardown wakes every worker, joins them, and tears down
- their sessions and bridge state.
+ (see *Interrupts*) follows — with a grace: a server that answers neither
+ the request nor the `Tflush` within 3s (`nine.Session.flush_grace_ms`)
+ is declared gone, the request answers `EINTR`, the session is wedged and
+ the mount dies with it (the next walk makes a new one). The protocol
+ says a client waits for the Rflush; a server that has not managed one in
+ that long is not going to, and the process behind the interrupt is
+ unkillable until we stop waiting. `stop_fd` (the child's death) is
+ watched by every session, so no worker can stay blocked on a hung server
+ past the program's exit; teardown wakes every worker, joins them, and
+ tears down their sessions and bridge state.
* The synthetic root is read-only (`dr-xr-xr-x`, like `/srv`): services are
posted and unposted by their servers (`cloud9.post`'s
`post`/`listenPosted`/`unpost`), not created and removed through files.
@@ -290,7 +324,8 @@ setupmapping = 48, removemapping = 49, syncfs = 50, tmpfile = 51, statx = 52, _
Constants: `kernel_version = 7`, `kernel_minor = 31` (what we answer; the
kernel adapts to the lower minor), `FOPEN_DIRECT_IO = 1`, `FOPEN_KEEP_CACHE = 2`,
-`FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_MAX_PAGES = 1<<22`,
+`FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_PARALLEL_DIROPS = 1<<18`,
+`FUSE_MAX_PAGES = 1<<22`,
`FATTR_MODE=1, FATTR_UID=2, FATTR_GID=4, FATTR_SIZE=8, FATTR_ATIME=16,
FATTR_MTIME=32, FATTR_FH=64, FATTR_ATIME_NOW=128, FATTR_MTIME_NOW=256,
FATTR_LOCKOWNER=512, FATTR_CTIME=1024`. `root_id = 1`.
@@ -409,7 +444,7 @@ pub fn serve(gpa: std.mem.Allocator, fuse_fd: i32, nine: *nine.Session, root_fid
// mntgen (see "mntgen: one mount, many servers" under Process model):
pub const MntgenOptions = struct {
- io: std.Io, // dispatcher-thread only: post.posted / post.dial
+ io: std.Io, // dispatcher-thread only: post.posted / registry stats
env: post.Env, // XDG_RUNTIME_DIR names the registry
uname: []const u8, aname: []const u8 = "", msize: u32 = 131072,
};
@@ -449,7 +484,7 @@ Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod):
| FUSE | 9P |
|---|---|
-| INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256 if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open |
+| INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256, and FUSE_PARALLEL_DIROPS, each if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open |
| LOOKUP(parent,name) | `walk(parent.fid → newfid, [name])`; `stat(newfid)`; dedupe by qid; `EntryOut` |
| FORGET / BATCH_FORGET | `nlookup -= n`; at 0 `clunk` and drop (no reply) |
| GETATTR | `stat(inode.fid)` → `AttrOut` |
@@ -512,9 +547,10 @@ is outstanding, including the initial root stat:
same errno through the ename table). An INTERRUPT for any other unique is
consumed and dropped (the kernel expects no reply). Any other request
(FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is
- stashed in a one-slot queue that `serve` dispatches, after swapping the two
- buffers, before it polls again; while the slot is full `watch` returns -1,
- so a second one cannot arrive.
+ copied into a queue that `serve` dispatches, in arrival order, before it
+ polls again; the fd stays watched throughout, so an INTERRUPT is never
+ stuck behind a parked request (a copy that cannot be allocated answers
+ `ENOMEM` on the spot).
* The INTERRUPT applies to `cur_unique` only and is consumed when read: the
clunks that unwind a half-done lookup/create/mkdir after an `Interrupted`
walk or stat are ordinary rpcs and are not re-interrupted by it. An
@@ -532,11 +568,14 @@ is outstanding, including the initial root stat:
included) as much as for a caught one, and then waits for the reply; our
`EINTR` is what finally lets the killed task die.
-Limits: a server that ignores Tflush still blocks the mount until it answers
-(the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and
-while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives
-after another process's request was parked is seen only once the blocked
-request completes (a multi-slot stash would lift that).
+Limits: a server that ignores Tflush is given `flush_grace_ms` (3s) and then
+declared gone — flush(5) says the server "should answer the Tflush message
+immediately", and a client that waits longer leaves an unkillable process
+behind — so the reader gets `EINTR` and the session ends (the hostile
+`never` mode; in single-connection mode that is the mount, in mntgen that
+one name). A slow-but-alive server that cannot answer a flush inside its own
+blocking read pays the same price; the alternative was a hang that only
+SIGKILL of 9ns could end.
### `src/ns.zig` — namespace and process plumbing
@@ -676,26 +715,27 @@ ReleaseSafe.
(event files) stalls that server's subtree while it is outstanding (but
not past the child's exit). It can be interrupted: killing or Ctrl-C-ing
the reader sends `FUSE_INTERRUPT`, which becomes `Tflush`; servers that
- honour it unblock immediately, servers that don't still block that
- subtree until they answer.
-* mntgen: the dispatcher dials on the main thread, in service of the walk
- that triggered it: a server that accepts the connection but never
- answers `Tversion` parks the whole mount for as long as the walk's
- program keeps running (as it would delay any 9P client). The dial
- watches `stop_fd`, so the program exiting ends it (`Stopped`, EIO to the
- pending LOOKUP); no concurrent dial, no dial timeout of our own. The
- remaining corner is a same-user self-DoS: the program blocked *in that
- very walk* sits in D state until the filesystem answers and cannot be
- killed to fire `stop_fd` — only killing the mute server (EOF) ends it.
+ honour it unblock immediately, servers that don't are given 3s and then
+ declared gone: the reader gets `EINTR` and the session (single
+ connection: the mount; mntgen: that name) ends.
+* mntgen: a name's dial runs on that name's worker, inside the walk that
+ asked, with no timeout of its own beyond the 5s backlog retry: a server
+ that accepts and never answers `Tversion` parks the walks into its own
+ name, and only those, until each is interrupted or the program exits.
+ The one wait nothing here can cut short is the kernel's own
+ serialization of lookups of a single name (a second walker into the
+ parked name waits for the first walk to end, uninterruptibly).
* mntgen: mount indexes are ordinals and never reused (cap 4096 dials per
process); per-mount node ids cap at 2^32 lookups. `st_ino` mixes the
index with the qid.path via XOR — collisions remain theoretically
possible, just not the practical ones (identical qid.paths across two
- servers). A dead mount keeps its slot — and its session socket and
- interrupt-pipe descriptors — until process exit (early close would race
- the dispatcher's INTERRUPT scan against fd reuse), so with a small
- `ulimit -n` a re-dial storm exhausts descriptors before the index cap;
- dials then fail cleanly with EIO.
+ servers). A dead mount releases its socket, its interrupt pipe and its
+ bridge state at once (the pipe under the mount's mutex, where the
+ dispatcher writes it) and keeps only its slot, so a server that dies and
+ comes back costs one index per death and nothing else; after 4096 of
+ them in one 9ns process no further name can be walked into (EIO) until
+ the shell is restarted. Reclaiming a slot once the kernel has forgotten
+ every node of the corpse is the next step, not taken here.
* mntgen: no posting/unposting through the mount (the synthetic root is
read-only; servers manage their registry entries through
`cloud9.post`), and no per-name mount options: one set of