diff options
| author | Gabriel Schneider <[email protected]> | 2026-09-21 14:13:43 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-09-21 14:13:43 -0300 |
| commit | 3a23f6a29e47ace901bd4d82b9db4055fcc12bb9 (patch) | |
| tree | b82d6e7c3ebe108434ce00ca75db59cf037917e0 /9ns/docs/DESIGN.md | |
| parent | f1b53c1533539aecbf16ad19fd9156deae091f92 (diff) | |
| download | cloud9-3a23f6a29e47ace901bd4d82b9db4055fcc12bb9.tar.gz cloud9-3a23f6a29e47ace901bd4d82b9db4055fcc12bb9.zip | |
post registry + 9ns --mntgen: the /srv translation
cloud9.post: servers post their socket under a name in
$XDG_RUNTIME_DIR/9p (post/unpost, posted, dial, Watch) and
serve.Runner.listenPosted posts a server by name, unposting on stop.
Names are budget-checked against the 108-byte socket path; a claim
binds+listens at a private temp path and takes the name with atomic
renames under flock (RENAME_NOREPLACE for free names, RENAME_EXCHANGE
grab-verify-commit for stale ones): the registry path is never unlinked
by a claim, live names refuse with AlreadyPosted, foreign files with
NotSocket, and unpost removes only the caller's inode-matched entry.
Watch surfaces inotify overflow and a replaced registry dir.
9ns --mntgen [--mount DIR] -- PROGRAM: one FUSE mount at /mnt/9p whose
synthetic root lists the posted registry (no connection made); a walk
into an unmounted name dials it and runs the existing bridge dispatch
in a per-server worker thread, routed by mount index in the node id's
top bits (ordinals never reused, cap 4096); a dead server answers EIO
on its subtree and is re-dialed on the next walk. The dial watches
stop_fd through Tversion (connectWatched). All existing 9ns forms are
unchanged.
9proc's unix listener no longer blind-unlinks its path: a foreign
non-socket is refused (Occupied), a live server is refused
(AlreadyListening), only a refused socket is cleared, and stop()
unlinks only the listener's own inode-matched socket.
Hardened by adversarial review (GLM 5.3 x2 + DeepSeek V4.1 Flash, all
high-thinking): double-bind races on one name (0 in 180k rounds),
foreign-file TOCTOU deletions (0 in 4M flips), a 255-byte-name listing
panic, inotify queue overflow silently dropped, listenPosted silently
overwriting, dial-time Tversion hangs wedging the dispatcher, --debug
silently ignored in mntgen, and xattr/statx probes answering EPERM on
the synthetic root (broke `ls -l /mnt/9p`).
Tests: root 80/80, 9ns 47/47, 9proc 60/60, integration 88/88 +
mntgen 37/37, adversarial 213/0, freestanding riscv32 gate green.
Diffstat (limited to '9ns/docs/DESIGN.md')
| -rw-r--r-- | 9ns/docs/DESIGN.md | 187 |
1 files changed, 175 insertions, 12 deletions
diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md index f1589d0..1d66387 100644 --- a/9ns/docs/DESIGN.md +++ b/9ns/docs/DESIGN.md @@ -102,6 +102,101 @@ protocol we need directly against `/usr/include/linux/fuse.h`. The FUSE fd is shared with the child only until exec (CLOEXEC); the parent's copy keeps the connection alive. +### mntgen: one mount, many servers (`9ns --mntgen`) + +`9ns --mntgen [--mount DIR] -- PROGRAM` (default mountpoint `/mnt/9p`) is a +transport of its own, mutually exclusive with `--unix/--tcp/--fd/--spawn`; +`--name` is rejected (there is no single server to name) while `--uname`, +`--aname`, `--msize`, `--cache`, `--no-direct-io` and `--debug` apply to +every per-server dial. The mount it builds is the `/srv` view of the +posted-9P registry (`cloud9.post`, `$XDG_RUNTIME_DIR/9p`): a walk into a +posted name reaches that server's whole 9P tree, and nothing is connected +until something walks. XDG_RUNTIME_DIR unset is fatal before anything is +forked (the registry is not guessable; no `/tmp` fallback). + +``` + program 9ns parent + in new userns │ + /mnt/9p ─FUSE─▶ kernel ─▶ │ dispatcher (main thread) + alpha/ beta/ │ ├─ synthetic root (node 1): lists the registry + ...each a server │ ├─ LOOKUP(alpha) ── dial+attach+stat ─▶ worker 1 ─ bridge ─ 9P session ─ server alpha + │ └─ LOOKUP(beta) ── dial+attach+stat ─▶ worker 2 ─ bridge ─ 9P session ─ server beta +``` + +Threads and node ids: + +* **Dispatcher** (the main thread) is the only reader of `/dev/fuse`. It + answers INIT/DESTROY, serves the **synthetic root** (node 1) itself, and + routes every other request by the node id's top bits — the **mount + index** — to the owning server's mount; a FUSE_INTERRUPT is routed by + scanning the mounts for the one currently serving the interrupted unique + and dropping the target into its interrupt pipe. +* A **mount** is a dialed server: its own `nine.Session`, its own bridge + state (inode table, handles — the ordinary single-server translation, + unchanged) and a **worker thread**. The worker pops copied requests off a + queue and serves them one at a time, exactly the single-connection + contract; the dispatcher keeps reading `/dev/fuse` meanwhile, so one slow + server never blocks the other names. Replies go straight back on the FUSE + fd (one `writev` per reply; the kernel processes each write as one + message). +* Node id layout: `nodeid = (mount_index << 32) | local`. Index 0 is the + synthetic root; per-mount local ids start at 1 (the server's 9P root) and + never exceed 2^32 (a bridge never reuses one). Mount indexes are + **ordinals and are never reused** (cap 4096 per process), so a stale + kernel-side inode of a dead server can never be conflated with a fresh + inode of its replacement. Reported `st_ino` mixes the index into the + qid.path (`(index+1) << 48` XOR), so two servers handing out the same + qid.path (two ramfs instances) still get distinct inode numbers. +* **Lazy dial**: LOOKUP of an unmounted name checks the registry, dials, + attaches, stats the root and spawns the worker — all on the dispatcher + thread, in service of the walk that triggered it. No eager connection is + ever made: `ls` of the root reads the registry directory only (a plain + file dropped there is listed too — and yields EIO on the walk, never + deleted). A walk into a **stale** entry (socket present, connect refused) + answers EIO. The whole dial watches `stop_fd` (the session is built with + `nine.Session.connectWatched`, so the `Tversion` exchange is covered too): + when the program exits while a walk is parked in a dial, the dial fails + with `Stopped`, the pending LOOKUP answers EIO and 9ns follows the program + out. There is no dial timeout of our own (a slow server delays the walk, + like it would delay any 9P client), and a server that accepts but never + answers `Tversion` still parks the dispatcher until the program exits — + including the unkillable corner where the *blocked walk itself* is the + only thing keeping the program alive (the task sits in D state until the + filesystem answers; a same-user self-DoS, accepted with the pinned + "dispatcher dials, no concurrent dial" design). +* **Death and re-dial**: when a worker's session dies mid-request, dispatch + has already answered that request EIO, the mount is marked dead, and + everything further routed to that subtree answers EIO (a FORGET is + dropped). Nothing reconnects eagerly. Because synthetic-root entries are + served with zero entry-validity, the next walk into the name LOOKUPs it + again; a dead mount is skipped and the name is dialed afresh — a new + mount under a new index, so kernel-held inodes of the corpse keep + answering EIO until forgotten. `ls` still lists the dead name (listing + connects to nothing). Death is discovered lazily: the first walk after a + silent death answers EIO (it marks the mount dead), the next walk re-dials. +* **Interrupts** work per mount: the worker's session polls the mount's + interrupt pipe while a 9P reply is outstanding; the dispatcher writes the + interrupted request's unique into it and the usual `Tflush` dance + (see *Interrupts*) follows. `stop_fd` (the child's death) is watched by + every session, so no worker can stay blocked on a hung server past the + program's exit; teardown wakes every worker, joins them, and tears down + their sessions and bridge state. +* The synthetic root is read-only (`dr-xr-xr-x`, like `/srv`): services are + posted and unposted by their servers (`cloud9.post`'s + `post`/`listenPosted`/`unpost`), not created and removed through files. + Capability probes the kernel makes before trusting a file — xattr ops and + `STATX` — answer `ENOSYS` (as they do on server subtrees through the + ordinary dispatch): `EPERM` would leak into userland as + "Operation not permitted" blamed on the mount root by `ls -l` and plain + `stat`. Modifying ops (create, mkdir, rename, ...) keep `EPERM` — the + root's read-only nature — and unknown ops are refused, never fatal. + Root readdir snapshots the registry per OPENDIR (a new OPENDIR sees new + posts); because nothing the kernel can cache is ever served stale (zero + timeouts, snapshot per opendir), there is nothing to invalidate and no + watcher is needed (`post.Watch` exists in the library for future caching). +* Naming: `--mount` (default `/mnt/9p`) is the whole mount; `$NINE_MOUNT` + points at it as usual. + ### Mountpoint policy Default mountpoint: `/mnt/9p/<name>`, where the name is `--name NAME` or is @@ -286,12 +381,33 @@ pub const Options = struct { }; /// Runs until the FUSE fd reports ENODEV or `stop_fd` becomes readable. pub fn serve(gpa: std.mem.Allocator, fuse_fd: i32, nine: *nine.Session, root_fid: u32, stop_fd: i32, opts: Options) !void; + +// mntgen (see "mntgen: one mount, many servers" under Process model): +pub const MntgenOptions = struct { + io: std.Io, // dispatcher-thread only: post.posted / post.dial + env: post.Env, // XDG_RUNTIME_DIR names the registry + uname: []const u8, aname: []const u8 = "", msize: u32 = 131072, +}; +/// One FUSE mount whose root lists the posted-9P registry; servers dialed +/// lazily, one worker thread each; node ids carry the mount index in the +/// top bits (`mount_shift = 32`, indexes are ordinals from 1, never reused, +/// `max_mounts` = 4096; the registry snapshot buffer is 8 KiB, `max_root_dirs` +/// = 64 concurrent OPENDIRs). Runs until ENODEV, DESTROY or `stop_fd`. +pub fn serveMntgen(gpa: std.mem.Allocator, fuse_fd: i32, stop_fd: i32, mo: MntgenOptions, opts: Options) !void; +pub fn mountNode(index: u32, local: u64) u64; // (index << 32) | local +pub fn mountIndex(nodeid: u64) u32; // nodeid >> 32 +pub fn nameIno(name: []const u8) u64; // FNV-1a of a name: synthetic-root dirent inos ``` State: * `inodes: AutoHashMap(u64 /*nodeid*/, Inode{ fid: u32, qid: Qid, nlookup: u64 })`. - Node 1 is the root (`root_fid`, never forgotten). + Node 1 is the root (`root_fid`, never forgotten). In mntgen mode the + same machinery runs once per mount with node ids that already carry the + mount index; the root node id is the `Bridge.root_id` field (`fuse.root_id` + single-connection) and reported inode numbers are `qid.path ^ ino_xor` + (`ino_xor` 0 single-connection). `cur_unique` is atomic so the mntgen + dispatcher can scan it to route FUSE_INTERRUPTs. * `by_qid: AutoHashMap(u64 /*qid.path*/, u64 /*nodeid*/)` so that repeated lookups of the same file map to the same inode (the old fid is clunked and the fresh one kept). Dedupe only merges when the qid type (dir bit) also @@ -431,10 +547,16 @@ Transport (exactly one): --tcp IP:PORT TCP (IPv4/IPv6 literal) --fd N already-connected inherited descriptor --spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout + --mntgen mount the posted-9P registry ($XDG_RUNTIME_DIR/9p): one + mount whose root lists the posted names; walking into a + name dials that server (mutually exclusive with the rest) Options: --name NAME mount name: the tree appears at /mnt/9p/NAME (one path - component; default derived from the transport, see below) - --mount PATH mountpoint inside the new namespace (overrides --name) + component; default derived from the transport, see below; + not with --mntgen) + --mount PATH mountpoint inside the new namespace (overrides --name; + with --mntgen the mount is the registry view itself, + default /mnt/9p) --uname NAME 9P user name (default $USER, else "none") --aname NAME 9P tree to attach (default "") --msize BYTES maximum 9P message size to request (default 131072, max 16 MiB) @@ -446,11 +568,14 @@ PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_M Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket; --tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its first word; --fd N -> fdN; 9p when nothing usable comes out of that. +--mntgen: no per-server name; the registry mount goes to --mount (default /mnt/9p). ``` `--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's status; 125 for 9ns's own failures (usage including a bad `--name`, connect, -mount); 126/127 as usual for exec failures. +mount, and for `--mntgen` an unset XDG_RUNTIME_DIR); 126/127 as usual for +exec failures. + ### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`) @@ -491,12 +616,26 @@ Everything under a temp dir. Skips (exit 0 with a notice) when is untouched). 6. Kill tests: 9ns exits when the child exits; server death during use yields `EIO`, not a hang. +7. `test/mntgen.sh` (also `9ns-itest`): `--mntgen` against a scratch + registry (`XDG_RUNTIME_DIR` = temp dir, never the real one): two servers + posted under two names (9proc-demo's socket created inside the registry + directory — a socket at `$XDG_RUNTIME_DIR/9p/<name>` is a posted name — + plus plan9port `ramfs`, an independent 9P2000 implementation, posting + with `NAMESPACE=$REG`), the synthetic root listing without dialing + (including a non-socket file, which is listed, yields EIO on the walk, + and is never removed), lazy dial and per-server routing in one program + run, a post appearing after mount, server death → EIO on the subtree + with the name still listed and a re-dial after re-post, parallel reads on + both mounts, `$NINE_MOUNT` = `/mnt/9p` by default, and the usage errors + (`--mntgen` + any transport, `--name`, `--mntgen=x`, unset + `XDG_RUNTIME_DIR`). ## Verification -`zig build 9ns-test` (unit), `zig build 9ns-itest` (88 end-to-end -checks against 9proc-demo over unix/tcp/socketpair and against plan9port's -`ramfs`) and `zig build 9ns-adv` (adversarial suites: a scriptable +`zig build 9ns-test` (unit), `zig build 9ns-itest` (integration.sh + the +mntgen suite: end-to-end checks against 9proc-demo over unix/tcp/socketpair, +plan9port's `ramfs`, and the posted-registry multi-server mode) and +`zig build 9ns-adv` (adversarial suites: a scriptable hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against its `never_flush`/`never` modes with 28 checks, FUSE semantics through the bridge, process/namespace/signal edge cases with 51 checks, and stress). The @@ -507,11 +646,35 @@ ReleaseSafe. ## Out of scope for v1 (documented, not hidden) -* One 9P request in flight at a time: a 9P read that blocks (event files) - stalls the whole mount while it is outstanding (but not past the child's - exit). It can be interrupted: killing or Ctrl-C-ing the reader sends - `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock - immediately, servers that don't still block the mount until they answer. +* One 9P request in flight at a time, per server in mntgen mode (across + servers they proceed in parallel, one worker each): a 9P read that blocks + (event files) stalls that server's subtree while it is outstanding (but + not past the child's exit). It can be interrupted: killing or Ctrl-C-ing + the reader sends `FUSE_INTERRUPT`, which becomes `Tflush`; servers that + honour it unblock immediately, servers that don't still block that + subtree until they answer. +* mntgen: the dispatcher dials on the main thread, in service of the walk + that triggered it: a server that accepts the connection but never + answers `Tversion` parks the whole mount for as long as the walk's + program keeps running (as it would delay any 9P client). The dial + watches `stop_fd`, so the program exiting ends it (`Stopped`, EIO to the + pending LOOKUP); no concurrent dial, no dial timeout of our own. The + remaining corner is a same-user self-DoS: the program blocked *in that + very walk* sits in D state until the filesystem answers and cannot be + killed to fire `stop_fd` — only killing the mute server (EOF) ends it. +* mntgen: mount indexes are ordinals and never reused (cap 4096 dials per + process); per-mount node ids cap at 2^32 lookups. `st_ino` mixes the + index with the qid.path via XOR — collisions remain theoretically + possible, just not the practical ones (identical qid.paths across two + servers). A dead mount keeps its slot — and its session socket and + interrupt-pipe descriptors — until process exit (early close would race + the dispatcher's INTERRUPT scan against fd reuse), so with a small + `ulimit -n` a re-dial storm exhausts descriptors before the index cap; + dials then fail cleanly with EIO. +* mntgen: no posting/unposting through the mount (the synthetic root is + read-only; servers manage their registry entries through + `cloud9.post`), and no per-name mount options: one set of + `--uname/--aname/--msize/--cache` applies to every dial. * No 9P2000.u/.L: no symlinks, ownership, or extended attributes. * No PID namespace, no `/proc` remount. `--tcp` needs an IP literal. * Cross-directory rename returns `EXDEV` (9P2000 cannot move files). |
