From 3e9f8805f293f622bb885cf849b5ce47dc062ad1 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Sat, 19 Sep 2026 23:55:47 -0300 Subject: 9ns: --name and /mnt/9p/ mounts, qid.path as inode number, interrupts as Tflush - --name NAME (default derived from the transport: socket basename, tcp-IP-PORT, spawned command, fdN) mounts at /mnt/9p/; --mount still overrides. ensureMountpoint walks down and creates missing components, shadowing the deepest unwritable ancestor. NINE_MOUNT is the only exported variable. - The inode number reported to the kernel is the 9P qid.path for every node, root included; a server handing qid.path 1 to a file (Pardes /self) no longer collides with the root. - FUSE_INTERRUPT for the request in flight becomes Tflush; a blocked read returns EINTR when the server answers the flush, chunked transfers return short counts, other requests arriving meanwhile are stashed and served next. Servers ignoring Tflush still block until they answer. - 9ns-test now covers nine/bridge/fuse; new adv_bridge_interrupt suite (28); 9ns-itest grows to 88 checks. Co-Authored-By: Claude Fable 5.1 --- 9ns/docs/DESIGN.md | 172 +++++++++++++++++++++++++++++++++++++++++++---------- 1 file changed, 142 insertions(+), 30 deletions(-) (limited to '9ns/docs') diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md index 7f943a8..f1589d0 100644 --- a/9ns/docs/DESIGN.md +++ b/9ns/docs/DESIGN.md @@ -15,7 +15,7 @@ So 9ns is a tiny FUSE server that speaks 9P2000 to the real server: ``` program (fish/bash/claude) 9ns (parent) 9P server in new user+mount namespace │ (9proc-demo, - /mnt/9p ─── FUSE ───▶ kernel ──▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) + /mnt/9p/ ─FUSE─▶ kernel ─▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) │ (framing) (translation) (cloud9 Client) ``` @@ -79,7 +79,8 @@ protocol we need directly against `/usr/include/linux/fuse.h`. 8. `statx` of the mountpoint: this forces one GETATTR, which the parent serves. Without it the kernel keeps the root inode's initial uid 0 (unmapped in the namespace) and every create in the root gets `EACCES`. - 9. Sets `NINE_MOUNT=` in the environment. + 9. Sets `NINE_MOUNT=` in the environment (replacing any + inherited value; a nested 9ns overwrites it). 10. `execve` of PROGRAM with PATH search (implemented by hand; no libc). Exec failures are reported through the `CLOEXEC` status socket (errno + message); the parent prints them after the serve loop ends. @@ -90,7 +91,10 @@ protocol we need directly against `/usr/include/linux/fuse.h`. signalled). The self-pipe is also watched by the 9P session while a reply is outstanding (`Session.stop_fd` → `error.Stopped`), so a server that never answers cannot keep 9ns alive after the child is gone; a 3 s - watchdog armed from the SIGCHLD handler is the last resort. + watchdog armed from the SIGCHLD handler is the last resort. The FUSE fd + is watched during that wait too (`Session.interrupt`): a `FUSE_INTERRUPT` + for the request being served becomes a `Tflush` (see *Interrupts* under + `src/bridge.zig`). 4. Signals in the parent: `SIGINT`/`SIGQUIT` ignored (the child owns the tty and gets them itself); `SIGTERM`/`SIGHUP` forwarded to the child; `SIGPIPE` ignored; `SIGCHLD` → self-pipe. @@ -100,17 +104,43 @@ copy keeps the connection alive. ### Mountpoint policy -Default mountpoint: `/mnt/9p`. A relative `--mount` is resolved against cwd. - -* If the path is a directory: use it. -* Else try `mkdir`. If that fails with `EACCES`/`EPERM`/`EROFS` (the normal - case for `/mnt/9p` as a plain user), **shadow the parent directory**: - open an fd to the parent, mount a `tmpfs` over it, then recreate every - existing entry inside the tmpfs: directories → `mkdir` + bind mount from - `/proc/self/fd//`; symlinks → `readlinkat` + `symlink`; anything - else → empty regular file + bind mount. Then `mkdir` the target inside. - Refuse (with a clear message) if the parent has more than 4096 entries or - is `/`. This only affects the new namespace. +Default mountpoint: `/mnt/9p/`, where the name is `--name NAME` or is +derived from the transport (`main.zig`, `defaultName`): + +| transport | default name | +|---|---| +| `--unix PATH` | basename of PATH with one trailing `.sock`/`.9p`/`.socket` removed (`/tmp/9debug.sock` → `9debug`) | +| `--tcp IP:PORT` | `tcp-IP-PORT` with `:` → `-` (brackets are already gone: `[::1]:564` → `tcp---1-564`) | +| `--spawn CMD` | basename of the first whitespace-separated word of CMD | +| `--fd N` | `fdN` | + +A name is one path component: non-empty, no `/`, no NUL, not `.`/`..`. An +invalid `--name` is a usage error (125); a derived name that is not valid +(empty basename, `..`) falls back to `9p`. `--mount PATH` overrides all of +this (the name is then PATH's last component and is not used for anything); +a relative `--mount` is resolved against cwd; `/` is rejected. + +`ns.ensureMountpoint` then makes the path a directory inside the new +namespace: + +* If the path is a directory: use it (a nested 9ns with the same name + therefore mounts over the outer mount at that path). +* Else walk up to the deepest existing ancestor (which must be a directory; + a dangling symlink anywhere is an error) and `mkdir` the missing + components under it one by one. If the first of those fails with + `EACCES`/`EPERM`/`EROFS` (the normal case for `/mnt/9p/` as a plain + user), **shadow that ancestor**: open an fd to it, mount a `tmpfs` over + it, then recreate every existing entry inside the tmpfs: directories → + `mkdir` + bind mount from `/proc/self/fd//`; symlinks → + `readlinkat` + `symlink`; anything else → empty regular file + bind mount. + Then create the missing components inside. Refuse (with a clear message) + if the ancestor is `/` (also through `/proc/self/root`), is `/proc` or + below it, or has more than 4096 entries. This only affects the new + namespace. So on a host without `/mnt/9p` the shadow goes over `/mnt` and + `9p/` is created inside; with a root-owned `/mnt/9p` it goes over + `/mnt/9p`; inside a 9ns namespace `/mnt/9p` is a directory of the outer + tmpfs owned by our uid, so a nested 9ns just creates `` next to the + outer mount and both are visible. * Else fail with the errno and a hint to pass `--mount` an existing dir. ## Module contracts @@ -151,6 +181,10 @@ I/O helpers (blocking fd, no allocation beyond the caller's buffer): pub const Request = struct { header: InHeader, body: []const u8 }; /// One kernel request. Returns null on ENODEV (unmounted). Retries EINTR/EAGAIN/ENOENT. pub fn readRequest(fd: i32, buf: []u8) !?Request; +/// Same, one read(2) only: EINTR/EAGAIN/ENOENT → error.Retry (for reads that follow a poll). +pub fn readRequestOnce(fd: i32, buf: []u8) !?Request; +/// O_NONBLOCK on the device, so a request withdrawn between poll and read cannot block us. +pub fn setNonblocking(fd: i32) !void; /// Success reply: header + concatenated payload slices, single writev. pub fn reply(fd: i32, unique: u64, payloads: []const []const u8) !void; /// Error reply: negative errno. @@ -172,11 +206,20 @@ single-threaded). Fids are allocated from a free list. ```zig pub const Address = union(enum) { unix: []const u8, tcp: struct { host: []const u8, port: u16 }, fd: i32 }; +/// Owner's hook into the reply wait (the bridge's FUSE fd); see "Interrupts" under bridge.zig. +pub const Interrupt = struct { + ctx: *anyopaque, + watch: *const fn (ctx) i32, // fd to poll alongside the socket, or -1 + onReadable: *const fn (ctx) Session.Error!bool, // consume it; true = flush the request in flight + armed: *const fn (ctx) bool, // a flush was asked for the operation in progress +}; pub const Session = struct { - pub const Error = error{ Nine, Protocol, Io, Closed, TooLarge, OutOfMemory }; + pub const Error = error{ Nine, Protocol, Io, Closed, Stopped, Interrupted, TooLarge, OutOfMemory }; /// After error.Nine, `ename` holds the server's Rerror text (copied, bounded). ename: [256]u8, ename_len: usize, msize: u32, + stop_fd: i32 = -1, // readable → pending rpc fails with error.Stopped + interrupt: ?Interrupt = null, pub fn connect(gpa: std.mem.Allocator, address: Address, msize: u32) !Session; // socket+connect, version pub fn deinit(s: *Session) void; @@ -209,12 +252,20 @@ negotiated msize is `result.version.msize`; if the server answered `"unknown"`, fail with `error.Protocol`. `rpc`: submit, write all of `client.output()` (calling `wrote`), then loop: -`take()`; if null, `read` from the fd into a temp buffer and `push` (push -returns how much fit; the frame is at most msize so it always fits after a -`take`). If the fd returns 0 → `error.Closed`. If the client dies → -`error.Protocol`. A `.fail` result copies the ename and returns `error.Nine`. +`take()`; if null, poll the fd together with `stop_fd` and the interrupt +source's descriptor; when the fd is readable, `read` into a temp buffer and +`push` (push returns how much fit; the frame is at most msize so it always +fits after a `take`). If the fd returns 0 → `error.Closed`. If the client +dies → `error.Protocol`. `stop_fd` readable → `error.Stopped`. When the +interrupt source asks for it, submit `.flush = .{ .oldtag = tag }` and keep +waiting: the original reply → returned normally (the Rflush that follows is +skipped by a later call); the Rflush first → `error.Interrupted`. A `.fail` +result copies the ename and returns `error.Nine`. `read`/`write` chunk loops +stop between chunks once the source reports `armed`, and turn an +`Interrupted` chunk into a short count when earlier chunks moved data. Rerror text → errno mapping (case-insensitive substring, in this order): +`"interrupt"` → `EINTR` (a server answering a flushed request with an error); `"not exist"`, `"not found"`, `"no such"` → `ENOENT`; `"exists"` → `EEXIST`; `"not empty"` → `ENOTEMPTY`; `"not a dir"` → `ENOTDIR`; `"is a dir"` → `EISDIR`; `"permission"`, `"denied"` → `EACCES`; @@ -276,7 +327,7 @@ Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod): | STATFS | constant `Kstatfs{ bsize = 4096, namelen = 255, frsize = 4096 }` | | ACCESS | `ENOSYS` (kernel stops asking; the server enforces permissions on open) | | READLINK, SYMLINK, LINK, MKNOD, *XATTR, *LK, IOCTL, POLL, BMAP, FALLOCATE, LSEEK, COPY_FILE_RANGE, TMPFILE, STATX | `ENOSYS` | -| INTERRUPT | ignored (reply nothing) | +| INTERRUPT | read while a 9P reply is outstanding: for the request in flight → `Tflush` (see *Interrupts*); otherwise ignored (reply nothing) | | DESTROY | return from `serve` | Attr mapping from `cloud9.Stat`: `mode = (S_IFDIR if DMDIR else S_IFREG) | @@ -301,6 +352,51 @@ rather than poisoning the whole READDIR reply; `length` near 2^64 is clamped to `i64` max; the errno of a failing 9P call is latched before any cleanup clunk overwrites the session's ename. +#### Interrupts (FUSE_INTERRUPT → Tflush) + +One request at a time, but not deaf. `serve` puts the FUSE fd in `O_NONBLOCK` +mode (every read follows a poll) and installs a `nine.Interrupt` source that +`Session.rpc` polls together with the socket and `stop_fd` whenever a reply +is outstanding, including the initial root stat: + +* `onReadable` reads the request the kernel has ready into a second + 8-aligned buffer of `request_buf_len` bytes (`spare_buf`). A + `FUSE_INTERRUPT` whose `InterruptIn.unique` names the request being served + (`cur_unique`) sets `interrupted` and returns true: `rpc` sends + `Tflush(oldtag)` and keeps waiting until either the original reply arrives + (the interrupt raced it: the result is returned as if nothing happened and + the Rflush that follows is swallowed by a later call) or the Rflush does + (`error.Interrupted` → `EINTR`; a server that instead answers the flushed + request with an Rerror containing "interrupt", as Pardes does, lands on the + same errno through the ename table). An INTERRUPT for any other unique is + consumed and dropped (the kernel expects no reply). Any other request + (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is + stashed in a one-slot queue that `serve` dispatches, after swapping the two + buffers, before it polls again; while the slot is full `watch` returns -1, + so a second one cannot arrive. +* The INTERRUPT applies to `cur_unique` only and is consumed when read: the + clunks that unwind a half-done lookup/create/mkdir after an `Interrupted` + walk or stat are ordinary rpcs and are not re-interrupted by it. An + interrupted walk clunks its new fid (the Rflush alone does not say whether + the server bound it); an Rerror'd walk only frees it locally. +* Multi-step operations fail with `EINTR` at whichever rpc was flushed and + release the fids they had allocated (`--debug` prints `fids=N` per + request; the interrupt suite checks it and the server-side count). The + chunk loops (`Session.read/write`, `loadDir`) also stop between chunks + while `armed`, so a reply that won the race cannot lead into another + blocking chunk: a read or write that already moved data returns the partial + count like read(2), one that moved nothing and a directory listing return + `EINTR`. +* The kernel sends one INTERRUPT per request, for a fatal signal (SIGKILL + included) as much as for a caught one, and then waits for the reply; our + `EINTR` is what finally lets the killed task die. + +Limits: a server that ignores Tflush still blocks the mount until it answers +(the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and +while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives +after another process's request was parked is seen only once the blocked +request completes (a multi-slot stash would lift that). + ### `src/ns.zig` — namespace and process plumbing ```zig @@ -316,7 +412,7 @@ pub const Child = struct { pid: i32 }; /// fork; the child sets up the namespace, mounts, and execs. Returns once exec succeeded /// (status pipe closed) or fails with the child's error (message on stderr). pub fn spawn(gpa: std.mem.Allocator, s: Spawn) !Child; -pub fn ensureMountpoint(path: [:0]const u8) !void; // the shadowing logic, testable alone +pub fn ensureMountpoint(gpa: Allocator, path: [:0]const u8) !void; // walk down, mkdir -p, shadow; testable alone pub fn resolveMountpoint(gpa, path: []const u8) ![:0]u8; // absolute, no trailing slash pub fn findInPath(gpa, envp, name) ![:0]u8; ``` @@ -336,7 +432,9 @@ Transport (exactly one): --fd N already-connected inherited descriptor --spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout Options: - --mount PATH mountpoint inside the new namespace (default /mnt/9p) + --name NAME mount name: the tree appears at /mnt/9p/NAME (one path + component; default derived from the transport, see below) + --mount PATH mountpoint inside the new namespace (overrides --name) --uname NAME 9P user name (default $USER, else "none") --aname NAME 9P tree to attach (default "") --msize BYTES maximum 9P message size to request (default 131072, max 16 MiB) @@ -345,9 +443,13 @@ Options: --debug trace FUSE and 9P operations on stderr --help, --version PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_MOUNT. +Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket; +--tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its +first word; --fd N -> fdN; 9p when nothing usable comes out of that. ``` -Exit codes: child's status; 125 for 9ns's own failures (usage, connect, +`--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's +status; 125 for 9ns's own failures (usage including a bad `--name`, connect, mount); 126/127 as usual for exec failures. ### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`) @@ -373,23 +475,30 @@ Everything under a temp dir. Skips (exit 0 with a notice) when `mv` across dirs fails with `EXDEV`-ish message, `rm`, `rmdir`, 1 MiB random file round trip compared with `sha256sum`, `dd` with odd block sizes, many small files, `find`, exit-status propagation (`exit 7` → 7), - `$NINE_MOUNT` set, nested `9ns` inside `9ns`. + `$NINE_MOUNT` set, nested `9ns` inside `9ns`. The suites pin the + mountpoint with `--mount /mnt/9p`; the naming section then checks the + default `/mnt/9p/` (unix socket basename, `--name`, `--name=`, + invalid names → 125, nested runs with two names both visible under + `/mnt/9p`, the same name twice mounting over, `--spawn` and `--tcp` + derived names, `--mount` beating `--name`). 2. `--spawn "<9proc-demo> --stdio"` variant. 3. `--tcp 127.0.0.1:` variant. 4. If `/usr/lib/plan9/bin/ramfs` exists: `NAMESPACE=$tmp ramfs -s ramfs` creates `$tmp/ramfs`; run the scratch battery against it. -5. `--mount` with an existing dir, with a relative path, and the default - `/mnt/9p` (exercises parent shadowing; verify `/mnt`'s other entries are - still visible inside). +5. `--mount` with an existing dir, with a relative path, `/mnt/9p` and the + default `/mnt/9p/` (both exercise the shadowing of `/mnt`; verify + `/mnt`'s other entries are still visible inside and the host mount table + is untouched). 6. Kill tests: 9ns exits when the child exits; server death during use yields `EIO`, not a hang. ## Verification -`zig build 9ns-test` (unit), `zig build 9ns-itest` (74 end-to-end +`zig build 9ns-test` (unit), `zig build 9ns-itest` (88 end-to-end checks against 9proc-demo over unix/tcp/socketpair and against plan9port's `ramfs`) and `zig build 9ns-adv` (adversarial suites: a scriptable -hostile 9P server with ~30 misbehaviour modes, FUSE semantics through the +hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against +its `never_flush`/`never` modes with 28 checks, FUSE semantics through the bridge, process/namespace/signal edge cases with 51 checks, and stress). The suites that attack the 9proc server itself (a hostile raw-9P client with 181 checks, the core, the Linux layer) moved with it to @@ -399,7 +508,10 @@ ReleaseSafe. ## Out of scope for v1 (documented, not hidden) * One 9P request in flight at a time: a 9P read that blocks (event files) - stalls the whole mount until it returns (but not past the child's exit). + stalls the whole mount while it is outstanding (but not past the child's + exit). It can be interrupted: killing or Ctrl-C-ing the reader sends + `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock + immediately, servers that don't still block the mount until they answer. * No 9P2000.u/.L: no symlinks, ownership, or extended attributes. * No PID namespace, no `/proc` remount. `--tcp` needs an IP literal. * Cross-directory rename returns `EXDEV` (9P2000 cannot move files). -- cgit v1.3