diff options
Diffstat (limited to '9ns/docs')
| -rw-r--r-- | 9ns/docs/DESIGN.md | 170 |
1 files changed, 141 insertions, 29 deletions
diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md index 7f943a8..f1589d0 100644 --- a/9ns/docs/DESIGN.md +++ b/9ns/docs/DESIGN.md @@ -15,7 +15,7 @@ So 9ns is a tiny FUSE server that speaks 9P2000 to the real server: ``` program (fish/bash/claude) 9ns (parent) 9P server in new user+mount namespace │ (9proc-demo, - /mnt/9p ─── FUSE ───▶ kernel ──▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) + /mnt/9p/<name> ─FUSE─▶ kernel ─▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) │ (framing) (translation) (cloud9 Client) ``` @@ -79,7 +79,8 @@ protocol we need directly against `/usr/include/linux/fuse.h`. 8. `statx` of the mountpoint: this forces one GETATTR, which the parent serves. Without it the kernel keeps the root inode's initial uid 0 (unmapped in the namespace) and every create in the root gets `EACCES`. - 9. Sets `NINE_MOUNT=<mountpoint>` in the environment. + 9. Sets `NINE_MOUNT=<mountpoint>` in the environment (replacing any + inherited value; a nested 9ns overwrites it). 10. `execve` of PROGRAM with PATH search (implemented by hand; no libc). Exec failures are reported through the `CLOEXEC` status socket (errno + message); the parent prints them after the serve loop ends. @@ -90,7 +91,10 @@ protocol we need directly against `/usr/include/linux/fuse.h`. signalled). The self-pipe is also watched by the 9P session while a reply is outstanding (`Session.stop_fd` → `error.Stopped`), so a server that never answers cannot keep 9ns alive after the child is gone; a 3 s - watchdog armed from the SIGCHLD handler is the last resort. + watchdog armed from the SIGCHLD handler is the last resort. The FUSE fd + is watched during that wait too (`Session.interrupt`): a `FUSE_INTERRUPT` + for the request being served becomes a `Tflush` (see *Interrupts* under + `src/bridge.zig`). 4. Signals in the parent: `SIGINT`/`SIGQUIT` ignored (the child owns the tty and gets them itself); `SIGTERM`/`SIGHUP` forwarded to the child; `SIGPIPE` ignored; `SIGCHLD` → self-pipe. @@ -100,17 +104,43 @@ copy keeps the connection alive. ### Mountpoint policy -Default mountpoint: `/mnt/9p`. A relative `--mount` is resolved against cwd. +Default mountpoint: `/mnt/9p/<name>`, where the name is `--name NAME` or is +derived from the transport (`main.zig`, `defaultName`): -* If the path is a directory: use it. -* Else try `mkdir`. If that fails with `EACCES`/`EPERM`/`EROFS` (the normal - case for `/mnt/9p` as a plain user), **shadow the parent directory**: - open an fd to the parent, mount a `tmpfs` over it, then recreate every - existing entry inside the tmpfs: directories → `mkdir` + bind mount from - `/proc/self/fd/<fd>/<name>`; symlinks → `readlinkat` + `symlink`; anything - else → empty regular file + bind mount. Then `mkdir` the target inside. - Refuse (with a clear message) if the parent has more than 4096 entries or - is `/`. This only affects the new namespace. +| transport | default name | +|---|---| +| `--unix PATH` | basename of PATH with one trailing `.sock`/`.9p`/`.socket` removed (`/tmp/9debug.sock` → `9debug`) | +| `--tcp IP:PORT` | `tcp-IP-PORT` with `:` → `-` (brackets are already gone: `[::1]:564` → `tcp---1-564`) | +| `--spawn CMD` | basename of the first whitespace-separated word of CMD | +| `--fd N` | `fdN` | + +A name is one path component: non-empty, no `/`, no NUL, not `.`/`..`. An +invalid `--name` is a usage error (125); a derived name that is not valid +(empty basename, `..`) falls back to `9p`. `--mount PATH` overrides all of +this (the name is then PATH's last component and is not used for anything); +a relative `--mount` is resolved against cwd; `/` is rejected. + +`ns.ensureMountpoint` then makes the path a directory inside the new +namespace: + +* If the path is a directory: use it (a nested 9ns with the same name + therefore mounts over the outer mount at that path). +* Else walk up to the deepest existing ancestor (which must be a directory; + a dangling symlink anywhere is an error) and `mkdir` the missing + components under it one by one. If the first of those fails with + `EACCES`/`EPERM`/`EROFS` (the normal case for `/mnt/9p/<name>` as a plain + user), **shadow that ancestor**: open an fd to it, mount a `tmpfs` over + it, then recreate every existing entry inside the tmpfs: directories → + `mkdir` + bind mount from `/proc/self/fd/<fd>/<name>`; symlinks → + `readlinkat` + `symlink`; anything else → empty regular file + bind mount. + Then create the missing components inside. Refuse (with a clear message) + if the ancestor is `/` (also through `/proc/self/root`), is `/proc` or + below it, or has more than 4096 entries. This only affects the new + namespace. So on a host without `/mnt/9p` the shadow goes over `/mnt` and + `9p/<name>` is created inside; with a root-owned `/mnt/9p` it goes over + `/mnt/9p`; inside a 9ns namespace `/mnt/9p` is a directory of the outer + tmpfs owned by our uid, so a nested 9ns just creates `<name>` next to the + outer mount and both are visible. * Else fail with the errno and a hint to pass `--mount` an existing dir. ## Module contracts @@ -151,6 +181,10 @@ I/O helpers (blocking fd, no allocation beyond the caller's buffer): pub const Request = struct { header: InHeader, body: []const u8 }; /// One kernel request. Returns null on ENODEV (unmounted). Retries EINTR/EAGAIN/ENOENT. pub fn readRequest(fd: i32, buf: []u8) !?Request; +/// Same, one read(2) only: EINTR/EAGAIN/ENOENT → error.Retry (for reads that follow a poll). +pub fn readRequestOnce(fd: i32, buf: []u8) !?Request; +/// O_NONBLOCK on the device, so a request withdrawn between poll and read cannot block us. +pub fn setNonblocking(fd: i32) !void; /// Success reply: header + concatenated payload slices, single writev. pub fn reply(fd: i32, unique: u64, payloads: []const []const u8) !void; /// Error reply: negative errno. @@ -172,11 +206,20 @@ single-threaded). Fids are allocated from a free list. ```zig pub const Address = union(enum) { unix: []const u8, tcp: struct { host: []const u8, port: u16 }, fd: i32 }; +/// Owner's hook into the reply wait (the bridge's FUSE fd); see "Interrupts" under bridge.zig. +pub const Interrupt = struct { + ctx: *anyopaque, + watch: *const fn (ctx) i32, // fd to poll alongside the socket, or -1 + onReadable: *const fn (ctx) Session.Error!bool, // consume it; true = flush the request in flight + armed: *const fn (ctx) bool, // a flush was asked for the operation in progress +}; pub const Session = struct { - pub const Error = error{ Nine, Protocol, Io, Closed, TooLarge, OutOfMemory }; + pub const Error = error{ Nine, Protocol, Io, Closed, Stopped, Interrupted, TooLarge, OutOfMemory }; /// After error.Nine, `ename` holds the server's Rerror text (copied, bounded). ename: [256]u8, ename_len: usize, msize: u32, + stop_fd: i32 = -1, // readable → pending rpc fails with error.Stopped + interrupt: ?Interrupt = null, pub fn connect(gpa: std.mem.Allocator, address: Address, msize: u32) !Session; // socket+connect, version pub fn deinit(s: *Session) void; @@ -209,12 +252,20 @@ negotiated msize is `result.version.msize`; if the server answered `"unknown"`, fail with `error.Protocol`. `rpc`: submit, write all of `client.output()` (calling `wrote`), then loop: -`take()`; if null, `read` from the fd into a temp buffer and `push` (push -returns how much fit; the frame is at most msize so it always fits after a -`take`). If the fd returns 0 → `error.Closed`. If the client dies → -`error.Protocol`. A `.fail` result copies the ename and returns `error.Nine`. +`take()`; if null, poll the fd together with `stop_fd` and the interrupt +source's descriptor; when the fd is readable, `read` into a temp buffer and +`push` (push returns how much fit; the frame is at most msize so it always +fits after a `take`). If the fd returns 0 → `error.Closed`. If the client +dies → `error.Protocol`. `stop_fd` readable → `error.Stopped`. When the +interrupt source asks for it, submit `.flush = .{ .oldtag = tag }` and keep +waiting: the original reply → returned normally (the Rflush that follows is +skipped by a later call); the Rflush first → `error.Interrupted`. A `.fail` +result copies the ename and returns `error.Nine`. `read`/`write` chunk loops +stop between chunks once the source reports `armed`, and turn an +`Interrupted` chunk into a short count when earlier chunks moved data. Rerror text → errno mapping (case-insensitive substring, in this order): +`"interrupt"` → `EINTR` (a server answering a flushed request with an error); `"not exist"`, `"not found"`, `"no such"` → `ENOENT`; `"exists"` → `EEXIST`; `"not empty"` → `ENOTEMPTY`; `"not a dir"` → `ENOTDIR`; `"is a dir"` → `EISDIR`; `"permission"`, `"denied"` → `EACCES`; @@ -276,7 +327,7 @@ Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod): | STATFS | constant `Kstatfs{ bsize = 4096, namelen = 255, frsize = 4096 }` | | ACCESS | `ENOSYS` (kernel stops asking; the server enforces permissions on open) | | READLINK, SYMLINK, LINK, MKNOD, *XATTR, *LK, IOCTL, POLL, BMAP, FALLOCATE, LSEEK, COPY_FILE_RANGE, TMPFILE, STATX | `ENOSYS` | -| INTERRUPT | ignored (reply nothing) | +| INTERRUPT | read while a 9P reply is outstanding: for the request in flight → `Tflush` (see *Interrupts*); otherwise ignored (reply nothing) | | DESTROY | return from `serve` | Attr mapping from `cloud9.Stat`: `mode = (S_IFDIR if DMDIR else S_IFREG) | @@ -301,6 +352,51 @@ rather than poisoning the whole READDIR reply; `length` near 2^64 is clamped to `i64` max; the errno of a failing 9P call is latched before any cleanup clunk overwrites the session's ename. +#### Interrupts (FUSE_INTERRUPT → Tflush) + +One request at a time, but not deaf. `serve` puts the FUSE fd in `O_NONBLOCK` +mode (every read follows a poll) and installs a `nine.Interrupt` source that +`Session.rpc` polls together with the socket and `stop_fd` whenever a reply +is outstanding, including the initial root stat: + +* `onReadable` reads the request the kernel has ready into a second + 8-aligned buffer of `request_buf_len` bytes (`spare_buf`). A + `FUSE_INTERRUPT` whose `InterruptIn.unique` names the request being served + (`cur_unique`) sets `interrupted` and returns true: `rpc` sends + `Tflush(oldtag)` and keeps waiting until either the original reply arrives + (the interrupt raced it: the result is returned as if nothing happened and + the Rflush that follows is swallowed by a later call) or the Rflush does + (`error.Interrupted` → `EINTR`; a server that instead answers the flushed + request with an Rerror containing "interrupt", as Pardes does, lands on the + same errno through the ename table). An INTERRUPT for any other unique is + consumed and dropped (the kernel expects no reply). Any other request + (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is + stashed in a one-slot queue that `serve` dispatches, after swapping the two + buffers, before it polls again; while the slot is full `watch` returns -1, + so a second one cannot arrive. +* The INTERRUPT applies to `cur_unique` only and is consumed when read: the + clunks that unwind a half-done lookup/create/mkdir after an `Interrupted` + walk or stat are ordinary rpcs and are not re-interrupted by it. An + interrupted walk clunks its new fid (the Rflush alone does not say whether + the server bound it); an Rerror'd walk only frees it locally. +* Multi-step operations fail with `EINTR` at whichever rpc was flushed and + release the fids they had allocated (`--debug` prints `fids=N` per + request; the interrupt suite checks it and the server-side count). The + chunk loops (`Session.read/write`, `loadDir`) also stop between chunks + while `armed`, so a reply that won the race cannot lead into another + blocking chunk: a read or write that already moved data returns the partial + count like read(2), one that moved nothing and a directory listing return + `EINTR`. +* The kernel sends one INTERRUPT per request, for a fatal signal (SIGKILL + included) as much as for a caught one, and then waits for the reply; our + `EINTR` is what finally lets the killed task die. + +Limits: a server that ignores Tflush still blocks the mount until it answers +(the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and +while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives +after another process's request was parked is seen only once the blocked +request completes (a multi-slot stash would lift that). + ### `src/ns.zig` — namespace and process plumbing ```zig @@ -316,7 +412,7 @@ pub const Child = struct { pid: i32 }; /// fork; the child sets up the namespace, mounts, and execs. Returns once exec succeeded /// (status pipe closed) or fails with the child's error (message on stderr). pub fn spawn(gpa: std.mem.Allocator, s: Spawn) !Child; -pub fn ensureMountpoint(path: [:0]const u8) !void; // the shadowing logic, testable alone +pub fn ensureMountpoint(gpa: Allocator, path: [:0]const u8) !void; // walk down, mkdir -p, shadow; testable alone pub fn resolveMountpoint(gpa, path: []const u8) ![:0]u8; // absolute, no trailing slash pub fn findInPath(gpa, envp, name) ![:0]u8; ``` @@ -336,7 +432,9 @@ Transport (exactly one): --fd N already-connected inherited descriptor --spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout Options: - --mount PATH mountpoint inside the new namespace (default /mnt/9p) + --name NAME mount name: the tree appears at /mnt/9p/NAME (one path + component; default derived from the transport, see below) + --mount PATH mountpoint inside the new namespace (overrides --name) --uname NAME 9P user name (default $USER, else "none") --aname NAME 9P tree to attach (default "") --msize BYTES maximum 9P message size to request (default 131072, max 16 MiB) @@ -345,9 +443,13 @@ Options: --debug trace FUSE and 9P operations on stderr --help, --version PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_MOUNT. +Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket; +--tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its +first word; --fd N -> fdN; 9p when nothing usable comes out of that. ``` -Exit codes: child's status; 125 for 9ns's own failures (usage, connect, +`--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's +status; 125 for 9ns's own failures (usage including a bad `--name`, connect, mount); 126/127 as usual for exec failures. ### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`) @@ -373,23 +475,30 @@ Everything under a temp dir. Skips (exit 0 with a notice) when `mv` across dirs fails with `EXDEV`-ish message, `rm`, `rmdir`, 1 MiB random file round trip compared with `sha256sum`, `dd` with odd block sizes, many small files, `find`, exit-status propagation (`exit 7` → 7), - `$NINE_MOUNT` set, nested `9ns` inside `9ns`. + `$NINE_MOUNT` set, nested `9ns` inside `9ns`. The suites pin the + mountpoint with `--mount /mnt/9p`; the naming section then checks the + default `/mnt/9p/<name>` (unix socket basename, `--name`, `--name=`, + invalid names → 125, nested runs with two names both visible under + `/mnt/9p`, the same name twice mounting over, `--spawn` and `--tcp` + derived names, `--mount` beating `--name`). 2. `--spawn "<9proc-demo> --stdio"` variant. 3. `--tcp 127.0.0.1:<port>` variant. 4. If `/usr/lib/plan9/bin/ramfs` exists: `NAMESPACE=$tmp ramfs -s ramfs` creates `$tmp/ramfs`; run the scratch battery against it. -5. `--mount` with an existing dir, with a relative path, and the default - `/mnt/9p` (exercises parent shadowing; verify `/mnt`'s other entries are - still visible inside). +5. `--mount` with an existing dir, with a relative path, `/mnt/9p` and the + default `/mnt/9p/<name>` (both exercise the shadowing of `/mnt`; verify + `/mnt`'s other entries are still visible inside and the host mount table + is untouched). 6. Kill tests: 9ns exits when the child exits; server death during use yields `EIO`, not a hang. ## Verification -`zig build 9ns-test` (unit), `zig build 9ns-itest` (74 end-to-end +`zig build 9ns-test` (unit), `zig build 9ns-itest` (88 end-to-end checks against 9proc-demo over unix/tcp/socketpair and against plan9port's `ramfs`) and `zig build 9ns-adv` (adversarial suites: a scriptable -hostile 9P server with ~30 misbehaviour modes, FUSE semantics through the +hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against +its `never_flush`/`never` modes with 28 checks, FUSE semantics through the bridge, process/namespace/signal edge cases with 51 checks, and stress). The suites that attack the 9proc server itself (a hostile raw-9P client with 181 checks, the core, the Linux layer) moved with it to @@ -399,7 +508,10 @@ ReleaseSafe. ## Out of scope for v1 (documented, not hidden) * One 9P request in flight at a time: a 9P read that blocks (event files) - stalls the whole mount until it returns (but not past the child's exit). + stalls the whole mount while it is outstanding (but not past the child's + exit). It can be interrupted: killing or Ctrl-C-ing the reader sends + `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock + immediately, servers that don't still block the mount until they answer. * No 9P2000.u/.L: no symlinks, ownership, or extended attributes. * No PID namespace, no `/proc` remount. `--tcp` needs an IP literal. * Cross-directory rename returns `EXDEV` (9P2000 cannot move files). |
