# 9ns design `9ns` mounts a 9P2000 file tree served over a Unix or TCP stream socket into a **fresh mount namespace** and runs a program inside it. The program (fish, bash, `claude`, anything) sees the 9P tree as ordinary files, without root and without touching the host's mount table. ## Why FUSE The kernel's own `9p` filesystem is not mountable inside an unprivileged user namespace (it lacks `FS_USERNS_MOUNT`) and loading it needs root. FUSE has been user-namespace mountable since Linux 4.18, and `/dev/fuse` is world read/write. So 9ns is a tiny FUSE server that speaks 9P2000 to the real server: ``` program (fish/bash/claude) 9ns (parent) 9P server in new user+mount namespace │ (9proc-demo, /mnt/9p/ ─FUSE─▶ kernel ─▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) │ (framing) (translation) (cloud9 Client) ``` No libfuse: `src/fuse.zig` implements the small subset of the kernel FUSE protocol we need directly against `/usr/include/linux/fuse.h`. ## Toolchain facts (Zig 0.16) * Zig 0.16.0 at `/usr/bin/zig`, std at `/usr/lib/zig/std`. **Grep the std tree before assuming an API exists**; 0.16 moved a lot of process/fs code behind `std.Io`. Raw Linux syscalls in `std.os.linux` (`fork`, `execve`, `mount`, `unshare`, `waitpid`, `pipe2`, `socketpair`, `poll`, `read`, `write`, `open`, `openat`, `getdents64`, `sigaction`, `kill`, `readlinkat`, `mkdirat`, `symlinkat`) are the intended low-level path. They return `usize`; decode with `std.os.linux.errno(rc)` (an `E` enum, `.SUCCESS` when ok). * `std.posix.poll`, `std.posix.sigaction`, `std.posix.read`, `std.posix.kill` exist. `std.posix.fork/execve/waitpid/pipe2/socketpair` do **not**. * `pub fn main() !void` and `pub fn main(init: std.process.Init) !void` are both supported. Prefer `main(init: std.process.Init)`; `init.gpa` is a general purpose allocator, `init.arena` an arena, `init.minimal.args` the argv (`toSlice(allocator)`), `init.minimal.environ.block` the envp block. * No libc is linked. Do not use `std.c.*`. Hostname lookups are therefore out of scope: `--tcp` takes IP literals only. * 9ns lives in the cloud9 repository as `cloud9/9ns/` and is built by the root `build.zig` through the fragment `9ns/build.zig` (steps `9ns`, `9ns-test`, `9ns-itest`, `9ns-adv`; toggle `-D9ns`). cloud9 itself is imported as module `cloud9` (`@import("cloud9")`). Read `../src/client.zig`, `Server.zig`, `wire.zig` and `../docs/design.md`. Its core is allocation-free and caller-driven: you push bytes in, take results out. The demo 9P server the tests mount is the sibling program `../9proc` (`zig build 9proc`). * Standalone module tests while other files are missing (from the cloud9 root): `zig test --dep cloud9 -Mroot=9ns/src/.zig -Mcloud9=src/root.zig`. * Format everything with `zig fmt`. ## Process model ``` 9ns [options] -- PROGRAM [ARGS...] ``` 1. Parent parses args, probes that `/dev/fuse` exists, connects to the 9P server, negotiates `version` and `attach`es (fid 0 = root). Connection failures are reported before anything is forked. 2. Parent forks with a `socketpair` status channel. **Child**: 1. `unshare(CLONE_NEWUSER | CLONE_NEWNS)`. 2. Writes `/proc/self/setgroups` = `deny`, `/proc/self/uid_map` = `" 1"`, `/proc/self/gid_map` = `" 1"` (same ids inside as outside; the child creating the namespace holds full capabilities in it until exec). 3. `mount(NULL, "/", NULL, MS_REC|MS_PRIVATE, NULL)` so nothing propagates. 4. Ensures the mountpoint exists (see below). 5. Opens `/dev/fuse` (`O_RDWR|O_CLOEXEC`). The kernel refuses to mount a fuse descriptor opened from a different user namespace than the mount ("wrong user namespace for fuse device"), so this must happen here, not in the parent. 6. `mount("9ns", mountpoint, "fuse", MS_NOSUID|MS_NODEV, "fd=,rootmode=40000,user_id=,group_id=,max_read=")`. 7. Sends the fuse fd to the parent over the status socket (`SCM_RIGHTS`). 8. `statx` of the mountpoint: this forces one GETATTR, which the parent serves. Without it the kernel keeps the root inode's initial uid 0 (unmapped in the namespace) and every create in the root gets `EACCES`. 9. Sets `NINE_MOUNT=` in the environment (replacing any inherited value; a nested 9ns overwrites it). 10. `execve` of PROGRAM with PATH search (implemented by hand; no libc). Exec failures are reported through the `CLOEXEC` status socket (errno + message); the parent prints them after the serve loop ends. 3. **Parent** receives the fuse fd, then runs the FUSE loop (`bridge.serve`) until either the child exits (SIGCHLD via self-pipe) or the FUSE fd reports `ENODEV` (last process in the namespace gone, mount destroyed). It then closes the fuse fd and exits with the child's status (`128+sig` if signalled). The self-pipe is also watched by the 9P session while a reply is outstanding (`Session.stop_fd` → `error.Stopped`), so a server that never answers cannot keep 9ns alive after the child is gone; a 3 s watchdog armed from the SIGCHLD handler is the last resort. The FUSE fd is watched during that wait too (`Session.interrupt`): a `FUSE_INTERRUPT` for the request being served becomes a `Tflush` (see *Interrupts* under `src/bridge.zig`). 4. Signals in the parent: `SIGINT`/`SIGQUIT` ignored (the child owns the tty and gets them itself); `SIGTERM`/`SIGHUP` forwarded to the child; `SIGPIPE` ignored; `SIGCHLD` → self-pipe. The FUSE fd is shared with the child only until exec (CLOEXEC); the parent's copy keeps the connection alive. ### mntgen: one mount, many servers (`9ns --mntgen`) `9ns --mntgen [--mount DIR] -- PROGRAM` (default mountpoint `/mnt/9p`) is a transport of its own, mutually exclusive with `--unix/--tcp/--fd/--spawn`; `--name` is rejected (there is no single server to name) while `--uname`, `--aname`, `--msize`, `--cache`, `--no-direct-io` and `--debug` apply to every per-server dial. The mount it builds is the `/srv` view of the posted-9P registry (`cloud9.post`, `$XDG_RUNTIME_DIR/9p`): a walk into a posted name reaches that server's whole 9P tree, and nothing is connected until something walks. XDG_RUNTIME_DIR unset is fatal before anything is forked (the registry is not guessable; no `/tmp` fallback). A registry entry that is a **directory** is served the same way the root is: a synthetic directory (its node id under the reserved index `synth_index`, the slot in the low bits) listing the real directory's entries, dialing the sockets found inside on walk and recursing into further directories — up to `max_synth_depth` (8) levels, bounded by `max_synth_dirs` (64) synthetic nodes per 9ns process. This is how multi-service providers organize themselves (zmx posts its sessions under `zmx/`), the plan9port `mntgen` shape: one tree, many mounts, each entry a mount point. Non-socket, non-directory entries inside a directory answer EIO on walk, exactly like a plain file in the registry itself; a directory's slots are freed when the kernel forgets the dentry. A synthetic node's slot comes back through FORGET, and the kernel sends most of them as `BATCH_FORGET`, whose header `nodeid` is 0 and whose body carries one `(nodeid, nlookup)` per forgotten node — for any mix of owners. The dispatcher therefore cannot route a batch by its header the way it routes every other request: `distributeForgets` unpacks the body and hands each entry to its owner (synthetic root, synthetic subdirectory or per-mount bridge). Routing the batch whole instead loses every entry in it, so the 64 slots leak and a subdirectory served once answers EIO forever. The forgets themselves arrive on the kernel's schedule, not at `close`, so a slot may take a moment to return; a listing that needs one meanwhile answers EIO rather than waiting. ``` program 9ns parent in new userns │ /mnt/9p ─FUSE─▶ kernel ─▶ │ dispatcher (main thread) alpha/ beta/ │ ├─ synthetic root (node 1): lists the registry ...each a server │ ├─ LOOKUP(alpha) ── dial+attach+stat ─▶ worker 1 ─ bridge ─ 9P session ─ server alpha │ └─ LOOKUP(beta) ── dial+attach+stat ─▶ worker 2 ─ bridge ─ 9P session ─ server beta ``` Threads and node ids: * **Dispatcher** (the main thread) is the only reader of `/dev/fuse`. It answers INIT/DESTROY, serves the **synthetic root** (node 1) itself, and routes every other request by the node id's top bits — the **mount index** — to the owning server's mount; a FUSE_INTERRUPT is routed by scanning the mounts for the one currently serving the interrupted unique and dropping the target into its interrupt pipe. * A **mount** is a dialed server: its own `nine.Session`, its own bridge state (inode table, handles — the ordinary single-server translation, unchanged) and a **worker thread**. The worker pops copied requests off a queue and serves them one at a time, exactly the single-connection contract; the dispatcher keeps reading `/dev/fuse` meanwhile, so one slow server never blocks the other names. Replies go straight back on the FUSE fd (one `writev` per reply; the kernel processes each write as one message). * Node id layout: `nodeid = (mount_index << 32) | local`. Index 0 is the synthetic root; per-mount local ids start at 1 (the server's 9P root) and never exceed 2^32 (a bridge never reuses one). Mount indexes are **ordinals and are never reused** (cap 4096 per process), so a stale kernel-side inode of a dead server can never be conflated with a fresh inode of its replacement. Reported `st_ino` mixes the index into the qid.path (`(index+1) << 48` XOR), so two servers handing out the same qid.path (two ramfs instances) still get distinct inode numbers. * **Lazy dial**: LOOKUP of an unmounted name checks the registry, dials, attaches, stats the root and spawns the worker — all on the dispatcher thread, in service of the walk that triggered it. No eager connection is ever made: `ls` of the root reads the registry directory only (a plain file dropped there is listed too — and yields EIO on the walk, never deleted). A walk into a **stale** entry (socket present, connect refused) answers EIO. The whole dial watches `stop_fd` (the session is built with `nine.Session.connectWatched`, so the `Tversion` exchange is covered too): when the program exits while a walk is parked in a dial, the dial fails with `Stopped`, the pending LOOKUP answers EIO and 9ns follows the program out. There is no dial timeout of our own (a slow server delays the walk, like it would delay any 9P client), and a server that accepts but never answers `Tversion` still parks the dispatcher until the program exits — including the unkillable corner where the *blocked walk itself* is the only thing keeping the program alive (the task sits in D state until the filesystem answers; a same-user self-DoS, accepted with the pinned "dispatcher dials, no concurrent dial" design). * **Death and re-dial**: when a worker's session dies mid-request, dispatch has already answered that request EIO, the mount is marked dead, and everything further routed to that subtree answers EIO (a FORGET is dropped). Nothing reconnects eagerly. Because synthetic-root entries are served with zero entry-validity, the next walk into the name LOOKUPs it again; a dead mount is skipped and the name is dialed afresh — a new mount under a new index, so kernel-held inodes of the corpse keep answering EIO until forgotten. `ls` still lists the dead name (listing connects to nothing). Death is discovered lazily: the first walk after a silent death answers EIO (it marks the mount dead), the next walk re-dials. * **Interrupts** work per mount: the worker's session polls the mount's interrupt pipe while a 9P reply is outstanding; the dispatcher writes the interrupted request's unique into it and the usual `Tflush` dance (see *Interrupts*) follows. `stop_fd` (the child's death) is watched by every session, so no worker can stay blocked on a hung server past the program's exit; teardown wakes every worker, joins them, and tears down their sessions and bridge state. * The synthetic root is read-only (`dr-xr-xr-x`, like `/srv`): services are posted and unposted by their servers (`cloud9.post`'s `post`/`listenPosted`/`unpost`), not created and removed through files. Capability probes the kernel makes before trusting a file — xattr ops and `STATX` — answer `ENOSYS` (as they do on server subtrees through the ordinary dispatch): `EPERM` would leak into userland as "Operation not permitted" blamed on the mount root by `ls -l` and plain `stat`. Modifying ops (create, mkdir, rename, ...) keep `EPERM` — the root's read-only nature — and unknown ops are refused, never fatal. Root readdir snapshots the registry per OPENDIR (a new OPENDIR sees new posts); because nothing the kernel can cache is ever served stale (zero timeouts, snapshot per opendir), there is nothing to invalidate and no watcher is needed (`post.Watch` exists in the library for future caching). * Naming: `--mount` (default `/mnt/9p`) is the whole mount; `$NINE_MOUNT` points at it as usual. ### Mountpoint policy Default mountpoint: `/mnt/9p/`, where the name is `--name NAME` or is derived from the transport (`main.zig`, `defaultName`): | transport | default name | |---|---| | `--unix PATH` | basename of PATH with one trailing `.sock`/`.9p`/`.socket` removed (`/tmp/9debug.sock` → `9debug`) | | `--tcp IP:PORT` | `tcp-IP-PORT` with `:` → `-` (brackets are already gone: `[::1]:564` → `tcp---1-564`) | | `--spawn CMD` | basename of the first whitespace-separated word of CMD | | `--fd N` | `fdN` | A name is one path component: non-empty, no `/`, no NUL, not `.`/`..`. An invalid `--name` is a usage error (125); a derived name that is not valid (empty basename, `..`) falls back to `9p`. `--mount PATH` overrides all of this (the name is then PATH's last component and is not used for anything); a relative `--mount` is resolved against cwd; `/` is rejected. `ns.ensureMountpoint` then makes the path a directory inside the new namespace: * If the path is a directory: use it (a nested 9ns with the same name therefore mounts over the outer mount at that path). * Else walk up to the deepest existing ancestor (which must be a directory; a dangling symlink anywhere is an error) and `mkdir` the missing components under it one by one. If the first of those fails with `EACCES`/`EPERM`/`EROFS` (the normal case for `/mnt/9p/` as a plain user), **shadow that ancestor**: open an fd to it, mount a `tmpfs` over it, then recreate every existing entry inside the tmpfs: directories → `mkdir` + bind mount from `/proc/self/fd//`; symlinks → `readlinkat` + `symlink`; anything else → empty regular file + bind mount. Then create the missing components inside. Refuse (with a clear message) if the ancestor is `/` (also through `/proc/self/root`), is `/proc` or below it, or has more than 4096 entries. This only affects the new namespace. So on a host without `/mnt/9p` the shadow goes over `/mnt` and `9p/` is created inside; with a root-owned `/mnt/9p` it goes over `/mnt/9p`; inside a 9ns namespace `/mnt/9p` is a directory of the outer tmpfs owned by our uid, so a nested 9ns just creates `` next to the outer mount and both are visible. * Else fail with the errno and a hint to pass `--mount` an existing dir. ## Module contracts ### `src/fuse.zig` — kernel FUSE protocol (no policy) Extern structs mirroring `linux/fuse.h`, with `comptime` size asserts: `InHeader` (40), `OutHeader` (16), `Attr` (88), `EntryOut` (128), `AttrOut` (104), `GetattrIn` (16), `SetattrIn` (88), `OpenIn` (8), `OpenOut` (16), `ReleaseIn` (24), `FlushIn` (24), `ReadIn` (40), `WriteIn` (40), `WriteOut` (8), `CreateIn` (16), `MkdirIn` (8), `RenameIn` (8), `Rename2In` (16), `ForgetIn` (8), `BatchForgetIn` (8), `ForgetOne` (16), `FsyncIn` (16), `AccessIn` (8), `InterruptIn` (8), `Kstatfs` (80), `StatfsOut` (80), `InitIn` (64), `InitOut` (64), `Dirent` (24 header, name padded to 8), `LseekIn` (24). `pub const Opcode = enum(u32) { lookup = 1, forget = 2, getattr = 3, setattr = 4, readlink = 5, symlink = 6, mknod = 8, mkdir = 9, unlink = 10, rmdir = 11, rename = 12, link = 13, open = 14, read = 15, write = 16, statfs = 17, release = 18, fsync = 20, setxattr = 21, getxattr = 22, listxattr = 23, removexattr = 24, flush = 25, init = 26, opendir = 27, readdir = 28, releasedir = 29, fsyncdir = 30, getlk = 31, setlk = 32, setlkw = 33, access = 34, create = 35, interrupt = 36, bmap = 37, destroy = 38, ioctl = 39, poll = 40, notify_reply = 41, batch_forget = 42, fallocate = 43, readdirplus = 44, rename2 = 45, lseek = 46, copy_file_range = 47, setupmapping = 48, removemapping = 49, syncfs = 50, tmpfile = 51, statx = 52, _ }` Constants: `kernel_version = 7`, `kernel_minor = 31` (what we answer; the kernel adapts to the lower minor), `FOPEN_DIRECT_IO = 1`, `FOPEN_KEEP_CACHE = 2`, `FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_MAX_PAGES = 1<<22`, `FATTR_MODE=1, FATTR_UID=2, FATTR_GID=4, FATTR_SIZE=8, FATTR_ATIME=16, FATTR_MTIME=32, FATTR_FH=64, FATTR_ATIME_NOW=128, FATTR_MTIME_NOW=256, FATTR_LOCKOWNER=512, FATTR_CTIME=1024`. `root_id = 1`. I/O helpers (blocking fd, no allocation beyond the caller's buffer): ```zig pub const Request = struct { header: InHeader, body: []const u8 }; /// One kernel request. Returns null on ENODEV (unmounted). Retries EINTR/EAGAIN/ENOENT. pub fn readRequest(fd: i32, buf: []u8) !?Request; /// Same, one read(2) only: EINTR/EAGAIN/ENOENT → error.Retry (for reads that follow a poll). pub fn readRequestOnce(fd: i32, buf: []u8) !?Request; /// O_NONBLOCK on the device, so a request withdrawn between poll and read cannot block us. pub fn setNonblocking(fd: i32) !void; /// Success reply: header + concatenated payload slices, single writev. pub fn reply(fd: i32, unique: u64, payloads: []const []const u8) !void; /// Error reply: negative errno. pub fn replyError(fd: i32, unique: u64, err: std.os.linux.E) !void; /// Append a fuse_dirent (8-byte padded) to `buf`; returns false if it doesn't fit. pub fn addDirent(buf: []u8, used: *usize, ino: u64, off: u64, dtype: u32, name: []const u8) bool; pub fn body(comptime T: type, req: Request) !*const T; // aligned copy-free view, checks size pub fn nameAfter(comptime T: type, req: Request) ![]const u8; // NUL-terminated name after a struct ``` The request buffer must be ≥ `max_write + 4096`; 9ns uses 1 MiB + 4 KiB. Requests with an unknown/unsupported opcode get `ENOSYS`. ### `src/nine.zig` — synchronous 9P2000 session on a blocking fd Thin, synchronous RPC layer over `cloud9.Client` (which is push/take, non-blocking-agnostic). One outstanding request at a time (the FUSE loop is single-threaded). Fids are allocated from a free list. ```zig pub const Address = union(enum) { unix: []const u8, tcp: struct { host: []const u8, port: u16 }, fd: i32 }; /// Owner's hook into the reply wait (the bridge's FUSE fd); see "Interrupts" under bridge.zig. pub const Interrupt = struct { ctx: *anyopaque, watch: *const fn (ctx) i32, // fd to poll alongside the socket, or -1 onReadable: *const fn (ctx) Session.Error!bool, // consume it; true = flush the request in flight armed: *const fn (ctx) bool, // a flush was asked for the operation in progress }; pub const Session = struct { pub const Error = error{ Nine, Protocol, Io, Closed, Stopped, Interrupted, TooLarge, OutOfMemory }; /// After error.Nine, `ename` holds the server's Rerror text (copied, bounded). ename: [256]u8, ename_len: usize, msize: u32, stop_fd: i32 = -1, // readable → pending rpc fails with error.Stopped interrupt: ?Interrupt = null, pub fn connect(gpa: std.mem.Allocator, address: Address, msize: u32) !Session; // socket+connect, version pub fn deinit(s: *Session) void; pub fn attach(s: *Session, fid: u32, uname: []const u8, aname: []const u8) Error!cloud9.Qid; pub fn allocFid(s: *Session) u32; pub fn freeFid(s: *Session, fid: u32) void; /// Generic RPC. Result slices borrow the input buffer until the next call. pub fn rpc(s: *Session, req: cloud9.Client.Request) Error!cloud9.Client.Result; // Conveniences (all built on rpc): pub fn walk(s, fid: u32, newfid: u32, names: []const []const u8) Error!Walk; // Walk = { nwqid, wqid[16] }; partial walk → error.Nine with ename "not found"-ish pub fn clone(s, fid: u32) Error!u32; // allocFid + walk with no names pub fn open(s, fid: u32, mode: u8) Error!Open; // Open = { qid, iounit } pub fn create(s, fid: u32, name: []const u8, perm: u32, mode: u8) Error!Open; pub fn read(s, fid: u32, offset: u64, buf: []u8) Error!usize; // chunks by maxRead/iounit; stops at short read pub fn write(s, fid: u32, offset: u64, data: []const u8) Error!usize; // chunks; stops at short write pub fn stat(s, fid: u32) Error!cloud9.Stat; // strings borrow the input buffer pub fn wstat(s, fid: u32, st: cloud9.Stat) Error!void; pub fn clunk(s, fid: u32) Error!void; // frees the fid even on error pub fn remove(s, fid: u32) Error!void; // frees the fid even on error pub fn errno(s: *const Session) std.os.linux.E; // map ename → errno (see below) }; pub const dontcare = cloud9.Stat{ .type = 0xFFFF, .dev = 0xFFFF_FFFF, .qid = .{ .type = 0xFF, .version = 0xFFFF_FFFF, .path = 0xFFFF_FFFF_FFFF_FFFF }, .mode = 0xFFFF_FFFF, .atime = 0xFFFF_FFFF, .mtime = 0xFFFF_FFFF, .length = 0xFFFF_FFFF_FFFF_FFFF, .name = "", .uid = "", .gid = "", .muid = "" }; ``` `connect`: for `.unix` and `.tcp` create a blocking `SOCK_STREAM|SOCK_CLOEXEC` socket and connect (`TCP_NODELAY` on TCP); for `.fd` adopt it. Buffers of `msize` bytes for in/out are heap allocated. Then submit `.version`, drain output to the socket, read until `take()` yields the version result. The negotiated msize is `result.version.msize`; if the server answered `"unknown"`, fail with `error.Protocol`. `rpc`: submit, write all of `client.output()` (calling `wrote`), then loop: `take()`; if null, poll the fd together with `stop_fd` and the interrupt source's descriptor; when the fd is readable, `read` into a temp buffer and `push` (push returns how much fit; the frame is at most msize so it always fits after a `take`). If the fd returns 0 → `error.Closed`. If the client dies → `error.Protocol`. `stop_fd` readable → `error.Stopped`. When the interrupt source asks for it, submit `.flush = .{ .oldtag = tag }` and keep waiting: the original reply → returned normally (the Rflush that follows is skipped by a later call); the Rflush first → `error.Interrupted`. A `.fail` result copies the ename and returns `error.Nine`. `read`/`write` chunk loops stop between chunks once the source reports `armed`, and turn an `Interrupted` chunk into a short count when earlier chunks moved data. Rerror text → errno mapping (case-insensitive substring, in this order): `"interrupt"` → `EINTR` (a server answering a flushed request with an error); `"not exist"`, `"not found"`, `"no such"` → `ENOENT`; `"exists"` → `EEXIST`; `"not empty"` → `ENOTEMPTY`; `"not a dir"` → `ENOTDIR`; `"is a dir"` → `EISDIR`; `"permission"`, `"denied"` → `EACCES`; `"read-only"`, `"read only"`, `"readonly"` → `EROFS`; `"no space"` → `ENOSPC`; `"not allowed"`, `"not permitted"`, `"cannot"` → `EPERM`; `"fid"` → `EBADF`; `"bad offset"`, `"invalid"`, `"bad "` → `EINVAL`; `"busy"`, `"in use"` → `EBUSY`; `"too long"` → `ENAMETOOLONG`; `"not supported"`, `"unsupported"` → `ENOTSUP`; otherwise `EIO`. ### `src/bridge.zig` — FUSE ↔ 9P translation ```zig pub const Options = struct { uid: u32, gid: u32, // reported owner of every file attr_timeout_ns: u64 = 1e9, // attr/entry cache validity (0 = none) direct_io: bool = true, // FOPEN_DIRECT_IO on every regular file debug: bool = false, // trace to stderr }; /// Runs until the FUSE fd reports ENODEV or `stop_fd` becomes readable. pub fn serve(gpa: std.mem.Allocator, fuse_fd: i32, nine: *nine.Session, root_fid: u32, stop_fd: i32, opts: Options) !void; // mntgen (see "mntgen: one mount, many servers" under Process model): pub const MntgenOptions = struct { io: std.Io, // dispatcher-thread only: post.posted / post.dial env: post.Env, // XDG_RUNTIME_DIR names the registry uname: []const u8, aname: []const u8 = "", msize: u32 = 131072, }; /// One FUSE mount whose root lists the posted-9P registry; servers dialed /// lazily, one worker thread each; node ids carry the mount index in the /// top bits (`mount_shift = 32`, indexes are ordinals from 1, never reused, /// `max_mounts` = 4096; the registry snapshot buffer is 8 KiB, `max_root_dirs` /// = 64 concurrent OPENDIRs). Runs until ENODEV, DESTROY or `stop_fd`. pub fn serveMntgen(gpa: std.mem.Allocator, fuse_fd: i32, stop_fd: i32, mo: MntgenOptions, opts: Options) !void; pub fn mountNode(index: u32, local: u64) u64; // (index << 32) | local pub fn mountIndex(nodeid: u64) u32; // nodeid >> 32 pub fn nameIno(name: []const u8) u64; // FNV-1a of a name: synthetic-root dirent inos ``` State: * `inodes: AutoHashMap(u64 /*nodeid*/, Inode{ fid: u32, qid: Qid, nlookup: u64 })`. Node 1 is the root (`root_fid`, never forgotten). In mntgen mode the same machinery runs once per mount with node ids that already carry the mount index; the root node id is the `Bridge.root_id` field (`fuse.root_id` single-connection) and reported inode numbers are `qid.path ^ ino_xor` (`ino_xor` 0 single-connection). `cur_unique` is atomic so the mntgen dispatcher can scan it to route FUSE_INTERRUPTs. * `by_qid: AutoHashMap(u64 /*qid.path*/, u64 /*nodeid*/)` so that repeated lookups of the same file map to the same inode (the old fid is clunked and the fresh one kept). Dedupe only merges when the qid type (dir bit) also matches, so a server reusing a path across a file and a directory cannot poison an inode. `ino` in attrs is `qid.path` (root, or anything carrying the root's path: 1). * `handles: AutoHashMap(u64 /*fh*/, Handle{ fid: u32, dir: ?DirList })`. `DirList` is the entire directory read at first `READDIR` offset 0: `[]Entry{ name: []u8, ino: u64, dtype: u32 }` with synthetic `.` and `..` first. `READDIR` offsets are indices into that list; a `READDIR` at offset 0 re-reads the directory (rewinddir). Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod): | FUSE | 9P | |---|---| | INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256 if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open | | LOOKUP(parent,name) | `walk(parent.fid → newfid, [name])`; `stat(newfid)`; dedupe by qid; `EntryOut` | | FORGET / BATCH_FORGET | `nlookup -= n`; at 0 `clunk` and drop (no reply) | | GETATTR | `stat(inode.fid)` → `AttrOut` | | SETATTR | `stat` then `wstat` with a *dontcare* Stat: `FATTR_SIZE`→length; `FATTR_MODE`→`(old.mode & ~0o777) \| (mode & 0o777)`; `FATTR_MTIME`→mtime (`FATTR_MTIME_NOW` → now); `FATTR_ATIME` ignored; `FATTR_UID/GID` → `EPERM` unless unchanged; then `stat` again for the reply | | OPEN | `clone(inode.fid)` then `open(newfid, mode)`; mode from `O_ACCMODE` (`oread/owrite/ordwr`), `O_TRUNC` → `otrunc`; reply `OpenOut{ fh, open_flags = FOPEN_DIRECT_IO }`; on failure clunk | | OPENDIR | same with `oread`; `fh` with `dir = null` | | READ | `read(fh.fid, offset, buf[0..min(size, 1 MiB)])`; reply data | | WRITE | `write(fh.fid, offset, data)`; `WriteOut{ size = n }` | | READDIR | fill `Dirent`s from the `DirList` starting at `offset`, up to `size` bytes | | RELEASE / RELEASEDIR | `clunk(fh.fid)`; free DirList | | FLUSH / FSYNC / FSYNCDIR | ok (no-op) | | CREATE(parent,name,flags,mode) | `clone(parent)`; `create(fid, name, mode & 0o777, openmode)` → this fid is the **open** file; then `walk(parent → fid2, [name])` + `stat(fid2)` for the inode; reply `EntryOut ++ OpenOut` | | MKDIR | `clone(parent)`; `create(fid, name, DMDIR \| (mode & 0o777), oread)`; `clunk`; then lookup as above | | UNLINK / RMDIR | `walk(parent → tmp, [name])`; `remove(tmp)` | | RENAME / RENAME2 | if `newdir != parent` → `EXDEV`; else `walk(parent → tmp, [oldname])`, `wstat(tmp, dontcare with .name = newname)`, `clunk`. 9P rename never replaces, POSIX does: when the target exists (and `RENAME_NOREPLACE` is not set) a directory target is removed first; a file target is parked under a temporary name, the rename retried, and the parked file removed only after success (restored on failure) | | STATFS | constant `Kstatfs{ bsize = 4096, namelen = 255, frsize = 4096 }` | | ACCESS | `ENOSYS` (kernel stops asking; the server enforces permissions on open) | | READLINK, SYMLINK, LINK, MKNOD, *XATTR, *LK, IOCTL, POLL, BMAP, FALLOCATE, LSEEK, COPY_FILE_RANGE, TMPFILE, STATX | `ENOSYS` | | INTERRUPT | read while a 9P reply is outstanding: for the request in flight → `Tflush` (see *Interrupts*); otherwise ignored (reply nothing) | | DESTROY | return from `serve` | Attr mapping from `cloud9.Stat`: `mode = (S_IFDIR if DMDIR else S_IFREG) | (st.mode & 0o777)`; `nlink = 1`; `size = length`; `blocks = (length+511)/512`; `blksize = 4096`; `atime/mtime/ctime = st.atime/st.mtime/st.mtime`; `uid/gid = opts.uid/gid`. `Dirent.type` = `DT_DIR` (4) / `DT_REG` (8). Errors: `nine.Session.Error.Nine` → `nine.errno()`; `Closed`/`Protocol`/`Io` → `EIO` and, since the session is dead, `serve` returns `error.Closed` after replying so 9ns can report "9P server went away". With `direct_io` the kernel never trusts `length` for reads: synthetic files that report length 0 (very common in 9P) still `cat` correctly, and reads run until the server returns a short read. With `--no-direct-io` the bridge forces `attr_timeout_ns = 0`, because a cached stale size truncates reads (observed data loss on a 4 MiB copy otherwise). Hostile-server rules: directory listings are capped at 64 MiB (a server that ignores read offsets otherwise loops forever); directory records with names containing `/`, NUL, empty, `.`/`..` or longer than `FUSE_NAME_MAX` are dropped rather than poisoning the whole READDIR reply; `length` near 2^64 is clamped to `i64` max; the errno of a failing 9P call is latched before any cleanup clunk overwrites the session's ename. #### Interrupts (FUSE_INTERRUPT → Tflush) One request at a time, but not deaf. `serve` puts the FUSE fd in `O_NONBLOCK` mode (every read follows a poll) and installs a `nine.Interrupt` source that `Session.rpc` polls together with the socket and `stop_fd` whenever a reply is outstanding, including the initial root stat: * `onReadable` reads the request the kernel has ready into a second 8-aligned buffer of `request_buf_len` bytes (`spare_buf`). A `FUSE_INTERRUPT` whose `InterruptIn.unique` names the request being served (`cur_unique`) sets `interrupted` and returns true: `rpc` sends `Tflush(oldtag)` and keeps waiting until either the original reply arrives (the interrupt raced it: the result is returned as if nothing happened and the Rflush that follows is swallowed by a later call) or the Rflush does (`error.Interrupted` → `EINTR`; a server that instead answers the flushed request with an Rerror containing "interrupt", as Pardes does, lands on the same errno through the ename table). An INTERRUPT for any other unique is consumed and dropped (the kernel expects no reply). Any other request (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is stashed in a one-slot queue that `serve` dispatches, after swapping the two buffers, before it polls again; while the slot is full `watch` returns -1, so a second one cannot arrive. * The INTERRUPT applies to `cur_unique` only and is consumed when read: the clunks that unwind a half-done lookup/create/mkdir after an `Interrupted` walk or stat are ordinary rpcs and are not re-interrupted by it. An interrupted walk clunks its new fid (the Rflush alone does not say whether the server bound it); an Rerror'd walk only frees it locally. * Multi-step operations fail with `EINTR` at whichever rpc was flushed and release the fids they had allocated (`--debug` prints `fids=N` per request; the interrupt suite checks it and the server-side count). The chunk loops (`Session.read/write`, `loadDir`) also stop between chunks while `armed`, so a reply that won the race cannot lead into another blocking chunk: a read or write that already moved data returns the partial count like read(2), one that moved nothing and a directory listing return `EINTR`. * The kernel sends one INTERRUPT per request, for a fatal signal (SIGKILL included) as much as for a caught one, and then waits for the reply; our `EINTR` is what finally lets the killed task die. Limits: a server that ignores Tflush still blocks the mount until it answers (the hostile `never` mode; `SIGTERM` to 9ns ends the session as before), and while the stash is full the FUSE fd is not read, so an INTERRUPT that arrives after another process's request was parked is seen only once the blocked request completes (a multi-slot stash would lift that). ### `src/ns.zig` — namespace and process plumbing ```zig pub const Spawn = struct { argv: []const []const u8, // argv[0] is PATH-searched unless it contains '/' envp: [*:null]const ?[*:0]const u8, // inherited environment mountpoint: []const u8, // absolute fuse_fd: i32, uid: u32, gid: u32, max_read: u32, }; pub const Child = struct { pid: i32 }; /// fork; the child sets up the namespace, mounts, and execs. Returns once exec succeeded /// (status pipe closed) or fails with the child's error (message on stderr). pub fn spawn(gpa: std.mem.Allocator, s: Spawn) !Child; pub fn ensureMountpoint(gpa: Allocator, path: [:0]const u8) !void; // walk down, mkdir -p, shadow; testable alone pub fn resolveMountpoint(gpa, path: []const u8) ![:0]u8; // absolute, no trailing slash pub fn findInPath(gpa, envp, name) ![:0]u8; ``` Also exports the signal plumbing used by `main.zig`: `installSignals(child_pid_ptr: *i32) !i32` returning the SIGCHLD self-pipe read end (used as `stop_fd` for `bridge.serve`), and `waitChild(pid) !u8` → exit status (`128+sig` on signal death). ### `src/main.zig` — CLI ``` Usage: 9ns [options] -- PROGRAM [ARGS...] Transport (exactly one): --unix PATH Unix stream socket --tcp IP:PORT TCP (IPv4/IPv6 literal) --fd N already-connected inherited descriptor --spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout --mntgen mount the posted-9P registry ($XDG_RUNTIME_DIR/9p): one mount whose root lists the posted names; walking into a name dials that server (mutually exclusive with the rest) Options: --name NAME mount name: the tree appears at /mnt/9p/NAME (one path component; default derived from the transport, see below; not with --mntgen) --mount PATH mountpoint inside the new namespace (overrides --name; with --mntgen the mount is the registry view itself, default /mnt/9p) --uname NAME 9P user name (default $USER, else "none") --aname NAME 9P tree to attach (default "") --msize BYTES maximum 9P message size to request (default 131072, max 16 MiB) --cache SECONDS attr/entry cache validity, may be fractional (default 1) --no-direct-io let the kernel cache file pages (trusts stat length) --debug trace FUSE and 9P operations on stderr --help, --version PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_MOUNT. Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket; --tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its first word; --fd N -> fdN; 9p when nothing usable comes out of that. --mntgen: no per-server name; the registry mount goes to --mount (default /mnt/9p). ``` `--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's status; 125 for 9ns's own failures (usage including a bad `--name`, connect, mount, and for `--mntgen` an unset XDG_RUNTIME_DIR); 126/127 as usual for exec failures. ### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`) The demo server is a separate program in this repository, built on the 9proc library; see `../9proc/docs/LIBRARY.md` for the library contract (freestanding core, value renderers, Linux debug probe). The tree it serves keeps the paths the integration tests read (`/build/*`, `/comptime/types//*`, `/comptime/decls`, `/runtime/fn/*`, `/runtime/ctl`, `/runtime/{pid,ppid,uptime,argv,cwd,env,clients}`, `/scratch/`) and adds `/vars`, `/threads`, `/addr`, `/mem`, `/hex`, `/breakpoints` and `/panic`. ## Integration test plan (`test/integration.sh`) Run by `zig build 9ns-itest`; args: path to `9ns`, path to `9proc-demo`. Everything under a temp dir. Skips (exit 0 with a notice) when `unshare -Urm true` fails or `/dev/fuse` is missing. 1. 9proc-demo on a Unix socket; `9ns --unix … -- sh -c` scripts: `cat /mnt/9p/build/zig_version` == `zig version`; `ls` listings; `stat` sizes; `/runtime/fn/now` is numeric; `ctl` round trip; `/scratch`: create, append (`>>`), overwrite, truncate, `mkdir -p a/b/c`, rename within dir, `mv` across dirs fails with `EXDEV`-ish message, `rm`, `rmdir`, 1 MiB random file round trip compared with `sha256sum`, `dd` with odd block sizes, many small files, `find`, exit-status propagation (`exit 7` → 7), `$NINE_MOUNT` set, nested `9ns` inside `9ns`. The suites pin the mountpoint with `--mount /mnt/9p`; the naming section then checks the default `/mnt/9p/` (unix socket basename, `--name`, `--name=`, invalid names → 125, nested runs with two names both visible under `/mnt/9p`, the same name twice mounting over, `--spawn` and `--tcp` derived names, `--mount` beating `--name`). 2. `--spawn "<9proc-demo> --stdio"` variant. 3. `--tcp 127.0.0.1:` variant. 4. If `/usr/lib/plan9/bin/ramfs` exists: `NAMESPACE=$tmp ramfs -s ramfs` creates `$tmp/ramfs`; run the scratch battery against it. 5. `--mount` with an existing dir, with a relative path, `/mnt/9p` and the default `/mnt/9p/` (both exercise the shadowing of `/mnt`; verify `/mnt`'s other entries are still visible inside and the host mount table is untouched). 6. Kill tests: 9ns exits when the child exits; server death during use yields `EIO`, not a hang. 7. `test/mntgen.sh` (also `9ns-itest`): `--mntgen` against a scratch registry (`XDG_RUNTIME_DIR` = temp dir, never the real one): two servers posted under two names (9proc-demo's socket created inside the registry directory — a socket at `$XDG_RUNTIME_DIR/9p/` is a posted name — plus plan9port `ramfs`, an independent 9P2000 implementation, posting with `NAMESPACE=$REG`), the synthetic root listing without dialing (including a non-socket file, which is listed, yields EIO on the walk, and is never removed), lazy dial and per-server routing in one program run, a post appearing after mount, server death → EIO on the subtree with the name still listed and a re-dial after re-post, parallel reads on both mounts, `$NINE_MOUNT` = `/mnt/9p` by default, and the usage errors (`--mntgen` + any transport, `--name`, `--mntgen=x`, unset `XDG_RUNTIME_DIR`). ## Verification `zig build 9ns-test` (unit), `zig build 9ns-itest` (integration.sh + the mntgen suite: end-to-end checks against 9proc-demo over unix/tcp/socketpair, plan9port's `ramfs`, and the posted-registry multi-server mode) and `zig build 9ns-adv` (adversarial suites: a scriptable hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against its `never_flush`/`never` modes with 28 checks, FUSE semantics through the bridge, process/namespace/signal edge cases with 51 checks, and stress). The suites that attack the 9proc server itself (a hostile raw-9P client with 181 checks, the core, the Linux layer) moved with it to `../9proc/test` (`zig build 9proc-adv`). All pass in Debug and ReleaseSafe. ## Out of scope for v1 (documented, not hidden) * One 9P request in flight at a time, per server in mntgen mode (across servers they proceed in parallel, one worker each): a 9P read that blocks (event files) stalls that server's subtree while it is outstanding (but not past the child's exit). It can be interrupted: killing or Ctrl-C-ing the reader sends `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock immediately, servers that don't still block that subtree until they answer. * mntgen: the dispatcher dials on the main thread, in service of the walk that triggered it: a server that accepts the connection but never answers `Tversion` parks the whole mount for as long as the walk's program keeps running (as it would delay any 9P client). The dial watches `stop_fd`, so the program exiting ends it (`Stopped`, EIO to the pending LOOKUP); no concurrent dial, no dial timeout of our own. The remaining corner is a same-user self-DoS: the program blocked *in that very walk* sits in D state until the filesystem answers and cannot be killed to fire `stop_fd` — only killing the mute server (EOF) ends it. * mntgen: mount indexes are ordinals and never reused (cap 4096 dials per process); per-mount node ids cap at 2^32 lookups. `st_ino` mixes the index with the qid.path via XOR — collisions remain theoretically possible, just not the practical ones (identical qid.paths across two servers). A dead mount keeps its slot — and its session socket and interrupt-pipe descriptors — until process exit (early close would race the dispatcher's INTERRUPT scan against fd reuse), so with a small `ulimit -n` a re-dial storm exhausts descriptors before the index cap; dials then fail cleanly with EIO. * mntgen: no posting/unposting through the mount (the synthetic root is read-only; servers manage their registry entries through `cloud9.post`), and no per-name mount options: one set of `--uname/--aname/--msize/--cache` applies to every dial. * No 9P2000.u/.L: no symlinks, ownership, or extended attributes. * No PID namespace, no `/proc` remount. `--tcp` needs an IP literal. * Cross-directory rename returns `EXDEV` (9P2000 cannot move files).