summaryrefslogtreecommitdiff
path: root/9ns/docs
diff options
context:
space:
mode:
Diffstat (limited to '9ns/docs')
-rw-r--r--9ns/docs/DESIGN.md170
1 files changed, 141 insertions, 29 deletions
diff --git a/9ns/docs/DESIGN.md b/9ns/docs/DESIGN.md
index 7f943a8..f1589d0 100644
--- a/9ns/docs/DESIGN.md
+++ b/9ns/docs/DESIGN.md
@@ -15,7 +15,7 @@ So 9ns is a tiny FUSE server that speaks 9P2000 to the real server:
```
program (fish/bash/claude) 9ns (parent) 9P server
in new user+mount namespace │ (9proc-demo,
- /mnt/9p ─── FUSE ───▶ kernel ──▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...)
+ /mnt/9p/<name> ─FUSE─▶ kernel ─▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...)
│ (framing) (translation) (cloud9 Client)
```
@@ -79,7 +79,8 @@ protocol we need directly against `/usr/include/linux/fuse.h`.
8. `statx` of the mountpoint: this forces one GETATTR, which the parent
serves. Without it the kernel keeps the root inode's initial uid 0
(unmapped in the namespace) and every create in the root gets `EACCES`.
- 9. Sets `NINE_MOUNT=<mountpoint>` in the environment.
+ 9. Sets `NINE_MOUNT=<mountpoint>` in the environment (replacing any
+ inherited value; a nested 9ns overwrites it).
10. `execve` of PROGRAM with PATH search (implemented by hand; no libc).
Exec failures are reported through the `CLOEXEC` status socket
(errno + message); the parent prints them after the serve loop ends.
@@ -90,7 +91,10 @@ protocol we need directly against `/usr/include/linux/fuse.h`.
signalled). The self-pipe is also watched by the 9P session while a reply
is outstanding (`Session.stop_fd` → `error.Stopped`), so a server that
never answers cannot keep 9ns alive after the child is gone; a 3 s
- watchdog armed from the SIGCHLD handler is the last resort.
+ watchdog armed from the SIGCHLD handler is the last resort. The FUSE fd
+ is watched during that wait too (`Session.interrupt`): a `FUSE_INTERRUPT`
+ for the request being served becomes a `Tflush` (see *Interrupts* under
+ `src/bridge.zig`).
4. Signals in the parent: `SIGINT`/`SIGQUIT` ignored (the child owns the tty
and gets them itself); `SIGTERM`/`SIGHUP` forwarded to the child;
`SIGPIPE` ignored; `SIGCHLD` → self-pipe.
@@ -100,17 +104,43 @@ copy keeps the connection alive.
### Mountpoint policy
-Default mountpoint: `/mnt/9p`. A relative `--mount` is resolved against cwd.
+Default mountpoint: `/mnt/9p/<name>`, where the name is `--name NAME` or is
+derived from the transport (`main.zig`, `defaultName`):
-* If the path is a directory: use it.
-* Else try `mkdir`. If that fails with `EACCES`/`EPERM`/`EROFS` (the normal
- case for `/mnt/9p` as a plain user), **shadow the parent directory**:
- open an fd to the parent, mount a `tmpfs` over it, then recreate every
- existing entry inside the tmpfs: directories → `mkdir` + bind mount from
- `/proc/self/fd/<fd>/<name>`; symlinks → `readlinkat` + `symlink`; anything
- else → empty regular file + bind mount. Then `mkdir` the target inside.
- Refuse (with a clear message) if the parent has more than 4096 entries or
- is `/`. This only affects the new namespace.
+| transport | default name |
+|---|---|
+| `--unix PATH` | basename of PATH with one trailing `.sock`/`.9p`/`.socket` removed (`/tmp/9debug.sock` → `9debug`) |
+| `--tcp IP:PORT` | `tcp-IP-PORT` with `:` → `-` (brackets are already gone: `[::1]:564` → `tcp---1-564`) |
+| `--spawn CMD` | basename of the first whitespace-separated word of CMD |
+| `--fd N` | `fdN` |
+
+A name is one path component: non-empty, no `/`, no NUL, not `.`/`..`. An
+invalid `--name` is a usage error (125); a derived name that is not valid
+(empty basename, `..`) falls back to `9p`. `--mount PATH` overrides all of
+this (the name is then PATH's last component and is not used for anything);
+a relative `--mount` is resolved against cwd; `/` is rejected.
+
+`ns.ensureMountpoint` then makes the path a directory inside the new
+namespace:
+
+* If the path is a directory: use it (a nested 9ns with the same name
+ therefore mounts over the outer mount at that path).
+* Else walk up to the deepest existing ancestor (which must be a directory;
+ a dangling symlink anywhere is an error) and `mkdir` the missing
+ components under it one by one. If the first of those fails with
+ `EACCES`/`EPERM`/`EROFS` (the normal case for `/mnt/9p/<name>` as a plain
+ user), **shadow that ancestor**: open an fd to it, mount a `tmpfs` over
+ it, then recreate every existing entry inside the tmpfs: directories →
+ `mkdir` + bind mount from `/proc/self/fd/<fd>/<name>`; symlinks →
+ `readlinkat` + `symlink`; anything else → empty regular file + bind mount.
+ Then create the missing components inside. Refuse (with a clear message)
+ if the ancestor is `/` (also through `/proc/self/root`), is `/proc` or
+ below it, or has more than 4096 entries. This only affects the new
+ namespace. So on a host without `/mnt/9p` the shadow goes over `/mnt` and
+ `9p/<name>` is created inside; with a root-owned `/mnt/9p` it goes over
+ `/mnt/9p`; inside a 9ns namespace `/mnt/9p` is a directory of the outer
+ tmpfs owned by our uid, so a nested 9ns just creates `<name>` next to the
+ outer mount and both are visible.
* Else fail with the errno and a hint to pass `--mount` an existing dir.
## Module contracts
@@ -151,6 +181,10 @@ I/O helpers (blocking fd, no allocation beyond the caller's buffer):
pub const Request = struct { header: InHeader, body: []const u8 };
/// One kernel request. Returns null on ENODEV (unmounted). Retries EINTR/EAGAIN/ENOENT.
pub fn readRequest(fd: i32, buf: []u8) !?Request;
+/// Same, one read(2) only: EINTR/EAGAIN/ENOENT → error.Retry (for reads that follow a poll).
+pub fn readRequestOnce(fd: i32, buf: []u8) !?Request;
+/// O_NONBLOCK on the device, so a request withdrawn between poll and read cannot block us.
+pub fn setNonblocking(fd: i32) !void;
/// Success reply: header + concatenated payload slices, single writev.
pub fn reply(fd: i32, unique: u64, payloads: []const []const u8) !void;
/// Error reply: negative errno.
@@ -172,11 +206,20 @@ single-threaded). Fids are allocated from a free list.
```zig
pub const Address = union(enum) { unix: []const u8, tcp: struct { host: []const u8, port: u16 }, fd: i32 };
+/// Owner's hook into the reply wait (the bridge's FUSE fd); see "Interrupts" under bridge.zig.
+pub const Interrupt = struct {
+ ctx: *anyopaque,
+ watch: *const fn (ctx) i32, // fd to poll alongside the socket, or -1
+ onReadable: *const fn (ctx) Session.Error!bool, // consume it; true = flush the request in flight
+ armed: *const fn (ctx) bool, // a flush was asked for the operation in progress
+};
pub const Session = struct {
- pub const Error = error{ Nine, Protocol, Io, Closed, TooLarge, OutOfMemory };
+ pub const Error = error{ Nine, Protocol, Io, Closed, Stopped, Interrupted, TooLarge, OutOfMemory };
/// After error.Nine, `ename` holds the server's Rerror text (copied, bounded).
ename: [256]u8, ename_len: usize,
msize: u32,
+ stop_fd: i32 = -1, // readable → pending rpc fails with error.Stopped
+ interrupt: ?Interrupt = null,
pub fn connect(gpa: std.mem.Allocator, address: Address, msize: u32) !Session; // socket+connect, version
pub fn deinit(s: *Session) void;
@@ -209,12 +252,20 @@ negotiated msize is `result.version.msize`; if the server answered
`"unknown"`, fail with `error.Protocol`.
`rpc`: submit, write all of `client.output()` (calling `wrote`), then loop:
-`take()`; if null, `read` from the fd into a temp buffer and `push` (push
-returns how much fit; the frame is at most msize so it always fits after a
-`take`). If the fd returns 0 → `error.Closed`. If the client dies →
-`error.Protocol`. A `.fail` result copies the ename and returns `error.Nine`.
+`take()`; if null, poll the fd together with `stop_fd` and the interrupt
+source's descriptor; when the fd is readable, `read` into a temp buffer and
+`push` (push returns how much fit; the frame is at most msize so it always
+fits after a `take`). If the fd returns 0 → `error.Closed`. If the client
+dies → `error.Protocol`. `stop_fd` readable → `error.Stopped`. When the
+interrupt source asks for it, submit `.flush = .{ .oldtag = tag }` and keep
+waiting: the original reply → returned normally (the Rflush that follows is
+skipped by a later call); the Rflush first → `error.Interrupted`. A `.fail`
+result copies the ename and returns `error.Nine`. `read`/`write` chunk loops
+stop between chunks once the source reports `armed`, and turn an
+`Interrupted` chunk into a short count when earlier chunks moved data.
Rerror text → errno mapping (case-insensitive substring, in this order):
+`"interrupt"` → `EINTR` (a server answering a flushed request with an error);
`"not exist"`, `"not found"`, `"no such"` → `ENOENT`; `"exists"` → `EEXIST`;
`"not empty"` → `ENOTEMPTY`; `"not a dir"` → `ENOTDIR`;
`"is a dir"` → `EISDIR`; `"permission"`, `"denied"` → `EACCES`;
@@ -276,7 +327,7 @@ Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod):
| STATFS | constant `Kstatfs{ bsize = 4096, namelen = 255, frsize = 4096 }` |
| ACCESS | `ENOSYS` (kernel stops asking; the server enforces permissions on open) |
| READLINK, SYMLINK, LINK, MKNOD, *XATTR, *LK, IOCTL, POLL, BMAP, FALLOCATE, LSEEK, COPY_FILE_RANGE, TMPFILE, STATX | `ENOSYS` |
-| INTERRUPT | ignored (reply nothing) |
+| INTERRUPT | read while a 9P reply is outstanding: for the request in flight → `Tflush` (see *Interrupts*); otherwise ignored (reply nothing) |
| DESTROY | return from `serve` |
Attr mapping from `cloud9.Stat`: `mode = (S_IFDIR if DMDIR else S_IFREG) |
@@ -301,6 +352,51 @@ rather than poisoning the whole READDIR reply; `length` near 2^64 is clamped
to `i64` max; the errno of a failing 9P call is latched before any cleanup
clunk overwrites the session's ename.
+#### Interrupts (FUSE_INTERRUPT → Tflush)
+
+One request at a time, but not deaf. `serve` puts the FUSE fd in `O_NONBLOCK`
+mode (every read follows a poll) and installs a `nine.Interrupt` source that
+`Session.rpc` polls together with the socket and `stop_fd` whenever a reply
+is outstanding, including the initial root stat:
+
+* `onReadable` reads the request the kernel has ready into a second
+ 8-aligned buffer of `request_buf_len` bytes (`spare_buf`). A
+ `FUSE_INTERRUPT` whose `InterruptIn.unique` names the request being served
+ (`cur_unique`) sets `interrupted` and returns true: `rpc` sends
+ `Tflush(oldtag)` and keeps waiting until either the original reply arrives
+ (the interrupt raced it: the result is returned as if nothing happened and
+ the Rflush that follows is swallowed by a later call) or the Rflush does
+ (`error.Interrupted` → `EINTR`; a server that instead answers the flushed
+ request with an Rerror containing "interrupt", as Pardes does, lands on the
+ same errno through the ename table). An INTERRUPT for any other unique is
+ consumed and dropped (the kernel expects no reply). Any other request
+ (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is
+ stashed in a one-slot queue that `serve` dispatches, after swapping the two
+ buffers, before it polls again; while the slot is full `watch` returns -1,
+ so a second one cannot arrive.
+* The INTERRUPT applies to `cur_unique` only and is consumed when read: the
+ clunks that unwind a half-done lookup/create/mkdir after an `Interrupted`
+ walk or stat are ordinary rpcs and are not re-interrupted by it. An
+ interrupted walk clunks its new fid (the Rflush alone does not say whether
+ the server bound it); an Rerror'd walk only frees it locally.
+* Multi-step operations fail with `EINTR` at whichever rpc was flushed and
+ release the fids they had allocated (`--debug` prints `fids=N` per
+ request; the interrupt suite checks it and the server-side count). The
+ chunk loops (`Session.read/write`, `loadDir`) also stop between chunks
+ while `armed`, so a reply that won the race cannot lead into another
+ blocking chunk: a read or write that already moved data returns the partial
+ count like read(2), one that moved nothing and a directory listing return
+ `EINTR`.
+* The kernel sends one INTERRUPT per request, for a fatal signal (SIGKILL
+ included) as much as for a caught one, and then waits for the reply; our
+ `EINTR` is what finally lets the killed task die.
+
+Limits: a server that ignores Tflush still blocks the mount until it answers
+(the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and
+while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives
+after another process's request was parked is seen only once the blocked
+request completes (a multi-slot stash would lift that).
+
### `src/ns.zig` — namespace and process plumbing
```zig
@@ -316,7 +412,7 @@ pub const Child = struct { pid: i32 };
/// fork; the child sets up the namespace, mounts, and execs. Returns once exec succeeded
/// (status pipe closed) or fails with the child's error (message on stderr).
pub fn spawn(gpa: std.mem.Allocator, s: Spawn) !Child;
-pub fn ensureMountpoint(path: [:0]const u8) !void; // the shadowing logic, testable alone
+pub fn ensureMountpoint(gpa: Allocator, path: [:0]const u8) !void; // walk down, mkdir -p, shadow; testable alone
pub fn resolveMountpoint(gpa, path: []const u8) ![:0]u8; // absolute, no trailing slash
pub fn findInPath(gpa, envp, name) ![:0]u8;
```
@@ -336,7 +432,9 @@ Transport (exactly one):
--fd N already-connected inherited descriptor
--spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout
Options:
- --mount PATH mountpoint inside the new namespace (default /mnt/9p)
+ --name NAME mount name: the tree appears at /mnt/9p/NAME (one path
+ component; default derived from the transport, see below)
+ --mount PATH mountpoint inside the new namespace (overrides --name)
--uname NAME 9P user name (default $USER, else "none")
--aname NAME 9P tree to attach (default "")
--msize BYTES maximum 9P message size to request (default 131072, max 16 MiB)
@@ -345,9 +443,13 @@ Options:
--debug trace FUSE and 9P operations on stderr
--help, --version
PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_MOUNT.
+Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket;
+--tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its
+first word; --fd N -> fdN; 9p when nothing usable comes out of that.
```
-Exit codes: child's status; 125 for 9ns's own failures (usage, connect,
+`--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's
+status; 125 for 9ns's own failures (usage including a bad `--name`, connect,
mount); 126/127 as usual for exec failures.
### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`)
@@ -373,23 +475,30 @@ Everything under a temp dir. Skips (exit 0 with a notice) when
`mv` across dirs fails with `EXDEV`-ish message, `rm`, `rmdir`, 1 MiB
random file round trip compared with `sha256sum`, `dd` with odd block
sizes, many small files, `find`, exit-status propagation (`exit 7` → 7),
- `$NINE_MOUNT` set, nested `9ns` inside `9ns`.
+ `$NINE_MOUNT` set, nested `9ns` inside `9ns`. The suites pin the
+ mountpoint with `--mount /mnt/9p`; the naming section then checks the
+ default `/mnt/9p/<name>` (unix socket basename, `--name`, `--name=`,
+ invalid names → 125, nested runs with two names both visible under
+ `/mnt/9p`, the same name twice mounting over, `--spawn` and `--tcp`
+ derived names, `--mount` beating `--name`).
2. `--spawn "<9proc-demo> --stdio"` variant.
3. `--tcp 127.0.0.1:<port>` variant.
4. If `/usr/lib/plan9/bin/ramfs` exists: `NAMESPACE=$tmp ramfs -s ramfs`
creates `$tmp/ramfs`; run the scratch battery against it.
-5. `--mount` with an existing dir, with a relative path, and the default
- `/mnt/9p` (exercises parent shadowing; verify `/mnt`'s other entries are
- still visible inside).
+5. `--mount` with an existing dir, with a relative path, `/mnt/9p` and the
+ default `/mnt/9p/<name>` (both exercise the shadowing of `/mnt`; verify
+ `/mnt`'s other entries are still visible inside and the host mount table
+ is untouched).
6. Kill tests: 9ns exits when the child exits; server death during use
yields `EIO`, not a hang.
## Verification
-`zig build 9ns-test` (unit), `zig build 9ns-itest` (74 end-to-end
+`zig build 9ns-test` (unit), `zig build 9ns-itest` (88 end-to-end
checks against 9proc-demo over unix/tcp/socketpair and against plan9port's
`ramfs`) and `zig build 9ns-adv` (adversarial suites: a scriptable
-hostile 9P server with ~30 misbehaviour modes, FUSE semantics through the
+hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against
+its `never_flush`/`never` modes with 28 checks, FUSE semantics through the
bridge, process/namespace/signal edge cases with 51 checks, and stress). The
suites that attack the 9proc server itself (a hostile raw-9P client with
181 checks, the core, the Linux layer) moved with it to
@@ -399,7 +508,10 @@ ReleaseSafe.
## Out of scope for v1 (documented, not hidden)
* One 9P request in flight at a time: a 9P read that blocks (event files)
- stalls the whole mount until it returns (but not past the child's exit).
+ stalls the whole mount while it is outstanding (but not past the child's
+ exit). It can be interrupted: killing or Ctrl-C-ing the reader sends
+ `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock
+ immediately, servers that don't still block the mount until they answer.
* No 9P2000.u/.L: no symlinks, ownership, or extended attributes.
* No PID namespace, no `/proc` remount. `--tcp` needs an IP literal.
* Cross-directory rename returns `EXDEV` (9P2000 cannot move files).