# 9ns design `9ns` mounts a 9P2000 file tree served over a Unix or TCP stream socket into a **fresh mount namespace** and runs a program inside it. The program (fish, bash, `claude`, anything) sees the 9P tree as ordinary files, without root and without touching the host's mount table. ## Why FUSE The kernel's own `9p` filesystem is not mountable inside an unprivileged user namespace (it lacks `FS_USERNS_MOUNT`) and loading it needs root. FUSE has been user-namespace mountable since Linux 4.18, and `/dev/fuse` is world read/write. So 9ns is a tiny FUSE server that speaks 9P2000 to the real server: ``` program (fish/bash/claude) 9ns (parent) 9P server in new user+mount namespace │ (9proc-demo, /mnt/9p/ ─FUSE─▶ kernel ─▶│ fuse.zig ──▶ bridge.zig ──▶ nine.zig ──▶ ramfs, ...) │ (framing) (translation) (cloud9 Client) ``` No libfuse: `src/fuse.zig` implements the small subset of the kernel FUSE protocol we need directly against `/usr/include/linux/fuse.h`. ## Toolchain facts (Zig 0.16) * Zig 0.16.0 at `/usr/bin/zig`, std at `/usr/lib/zig/std`. **Grep the std tree before assuming an API exists**; 0.16 moved a lot of process/fs code behind `std.Io`. Raw Linux syscalls in `std.os.linux` (`fork`, `execve`, `mount`, `unshare`, `waitpid`, `pipe2`, `socketpair`, `poll`, `read`, `write`, `open`, `openat`, `getdents64`, `sigaction`, `kill`, `readlinkat`, `mkdirat`, `symlinkat`) are the intended low-level path. They return `usize`; decode with `std.os.linux.errno(rc)` (an `E` enum, `.SUCCESS` when ok). * `std.posix.poll`, `std.posix.sigaction`, `std.posix.read`, `std.posix.kill` exist. `std.posix.fork/execve/waitpid/pipe2/socketpair` do **not**. * `pub fn main() !void` and `pub fn main(init: std.process.Init) !void` are both supported. Prefer `main(init: std.process.Init)`; `init.gpa` is a general purpose allocator, `init.arena` an arena, `init.minimal.args` the argv (`toSlice(allocator)`), `init.minimal.environ.block` the envp block. * No libc is linked. Do not use `std.c.*`. Hostname lookups are therefore out of scope: `--tcp` takes IP literals only. * 9ns lives in the cloud9 repository as `cloud9/9ns/` and is built by the root `build.zig` through the fragment `9ns/build.zig` (steps `9ns`, `9ns-test`, `9ns-itest`, `9ns-adv`; toggle `-D9ns`). cloud9 itself is imported as module `cloud9` (`@import("cloud9")`). Read `../src/client.zig`, `Server.zig`, `wire.zig` and `../docs/design.md`. Its core is allocation-free and caller-driven: you push bytes in, take results out. The demo 9P server the tests mount is the sibling program `../9proc` (`zig build 9proc`). * Standalone module tests while other files are missing (from the cloud9 root): `zig test --dep cloud9 -Mroot=9ns/src/.zig -Mcloud9=src/root.zig`. * Format everything with `zig fmt`. ## Process model ``` 9ns [options] -- PROGRAM [ARGS...] ``` 1. Parent parses args, probes that `/dev/fuse` exists, connects to the 9P server, negotiates `version` and `attach`es (fid 0 = root). Connection failures are reported before anything is forked. 2. Parent forks with a `socketpair` status channel. **Child**: 1. `unshare(CLONE_NEWUSER | CLONE_NEWNS)`. 2. Writes `/proc/self/setgroups` = `deny`, `/proc/self/uid_map` = `" 1"`, `/proc/self/gid_map` = `" 1"` (same ids inside as outside; the child creating the namespace holds full capabilities in it until exec). 3. `mount(NULL, "/", NULL, MS_REC|MS_PRIVATE, NULL)` so nothing propagates. 4. Ensures the mountpoint exists (see below). 5. Opens `/dev/fuse` (`O_RDWR|O_CLOEXEC`). The kernel refuses to mount a fuse descriptor opened from a different user namespace than the mount ("wrong user namespace for fuse device"), so this must happen here, not in the parent. 6. `mount("9ns", mountpoint, "fuse", MS_NOSUID|MS_NODEV, "fd=,rootmode=40000,user_id=,group_id=,max_read=")`. 7. Sends the fuse fd to the parent over the status socket (`SCM_RIGHTS`). 8. `statx` of the mountpoint: this forces one GETATTR, which the parent serves. Without it the kernel keeps the root inode's initial uid 0 (unmapped in the namespace) and every create in the root gets `EACCES`. 9. Sets `NINE_MOUNT=` in the environment (replacing any inherited value; a nested 9ns overwrites it). 10. `execve` of PROGRAM with PATH search (implemented by hand; no libc). Exec failures are reported through the `CLOEXEC` status socket (errno + message); the parent prints them after the serve loop ends. 3. **Parent** receives the fuse fd, then runs the FUSE loop (`bridge.serve`) until either the child exits (SIGCHLD via self-pipe) or the FUSE fd reports `ENODEV` (last process in the namespace gone, mount destroyed). It then closes the fuse fd and exits with the child's status (`128+sig` if signalled). The self-pipe is also watched by the 9P session while a reply is outstanding (`Session.stop_fd` → `error.Stopped`), so a server that never answers cannot keep 9ns alive after the child is gone; a 3 s watchdog armed from the SIGCHLD handler is the last resort. The FUSE fd is watched during that wait too (`Session.interrupt`): a `FUSE_INTERRUPT` for the request being served becomes a `Tflush` (see *Interrupts* under `src/bridge.zig`). 4. Signals in the parent: `SIGINT`/`SIGQUIT` ignored (the child owns the tty and gets them itself); `SIGTERM`/`SIGHUP` forwarded to the child; `SIGPIPE` ignored; `SIGCHLD` → self-pipe. The FUSE fd is shared with the child only until exec (CLOEXEC); the parent's copy keeps the connection alive. ### mntgen: one mount, many servers (`9ns --mntgen`) `9ns --mntgen [--mount DIR] -- PROGRAM` (default mountpoint `/mnt/9p`) is a transport of its own, mutually exclusive with `--unix/--tcp/--fd/--spawn`; `--name` is rejected (there is no single server to name) while `--uname`, `--aname`, `--msize`, `--cache`, `--no-direct-io` and `--debug` apply to every per-server dial. The mount it builds is the `/srv` view of the posted-9P registry (`cloud9.post`, `$XDG_RUNTIME_DIR/9p`): a walk into a posted name reaches that server's whole 9P tree, and nothing is connected until something walks. XDG_RUNTIME_DIR unset is fatal before anything is forked (the registry is not guessable; no `/tmp` fallback). A registry entry that is a **directory** is served the same way the root is: a synthetic directory (its node id under the reserved index `synth_index`, the slot in the low bits) listing the real directory's entries, dialing the sockets found inside on walk and recursing into further directories — up to `max_synth_depth` (8) levels, bounded by `max_synth_dirs` (64) synthetic nodes per 9ns process. This is how multi-service providers organize themselves (zmx posts its sessions under `zmx/`), the plan9port `mntgen` shape: one tree, many mounts, each entry a mount point. Non-socket, non-directory entries inside a directory answer EIO on walk, exactly like a plain file in the registry itself; a directory's slots are freed when the kernel forgets the dentry. A synthetic node's slot comes back through FORGET, and the kernel sends most of them as `BATCH_FORGET`, whose header `nodeid` is 0 and whose body carries one `(nodeid, nlookup)` per forgotten node — for any mix of owners. The dispatcher therefore cannot route a batch by its header the way it routes every other request: `distributeForgets` unpacks the body and hands each entry to its owner (synthetic root, synthetic subdirectory or per-mount bridge). Routing the batch whole instead loses every entry in it, so the 64 slots leak and a subdirectory served once answers EIO forever. The forgets themselves arrive on the kernel's schedule, not at `close`, so a slot may take a moment to return; a listing that needs one meanwhile answers EIO rather than waiting. ``` program 9ns parent in new userns │ /mnt/9p ─FUSE─▶ kernel ─▶ │ dispatcher (main thread) alpha/ beta/ │ ├─ synthetic root (node 1): lists the registry ...each a server │ ├─ LOOKUP(alpha) ─▶ worker 1 ─ dial+attach+stat, then bridge ─ 9P session ─ server alpha │ └─ LOOKUP(beta) ─▶ worker 2 ─ dial+attach+stat, then bridge ─ 9P session ─ server beta ``` Threads and node ids: * **Dispatcher** (the main thread) is the only reader of `/dev/fuse`. It answers INIT/DESTROY, serves the **synthetic root** (node 1) itself, and routes every other request by the node id's top bits — the **mount index** — to the owning server's mount. It waits on no server, ever: the only blocking it does is the registry directory (a tmpfs) and the mount queues' locks. A FUSE_INTERRUPT is routed by scanning the mounts: a request still sitting in a mount's queue is taken out and answered `EINTR` on the spot (the kernel sends an INTERRUPT once, and a request that only runs later would otherwise run to the end with nobody left wanting it); one in flight gets its unique dropped into that mount's interrupt pipe. Queue and in-flight unique are read under the mount's mutex, which is also where the worker moves a request from one to the other, so an interrupt cannot fall between them. * A **mount** is a posted name walked into: the socket to dial, its own `nine.Session` and bridge state (inode table, handles — the ordinary single-server translation, unchanged) once dialed, and a **worker thread** that owns all of it. The dispatcher makes a mount without touching the network and hands it the walk; the worker dials while serving that first request. It pops copied requests off a queue and serves them one at a time, exactly the single-connection contract; the dispatcher keeps reading `/dev/fuse` meanwhile, so one slow server never blocks the other names. Replies go straight back on the FUSE fd (one `writev` per reply; the kernel processes each write as one message). The queue is a plain list under a mutex and condition rather than an `std.Io.Queue` because an interrupt has to find and remove a request by unique from the middle of it — the same reason a 9P server keeps its pending requests in a list a Tflush can search. * Node id layout: `nodeid = (mount_index << 32) | local`. Index 0 is the synthetic root; per-mount local ids start at 1 (the server's 9P root) and never exceed 2^32 (a bridge never reuses one). Mount indexes are **ordinals and are never reused** (cap 4096 per process), so a stale kernel-side inode of a dead server can never be conflated with a fresh inode of its replacement. Reported `st_ino` mixes the index into the qid.path (`(index+1) << 48` XOR), so two servers handing out the same qid.path (two ramfs instances) still get distinct inode numbers. * **Lazy dial, on the worker**: LOOKUP of an unmounted name stats the registry entry (dispatcher), makes the mount and queues the LOOKUP to it; the worker dials (connect, `Tversion`, `Tattach`, `Tstat` of the root) as the first thing it does for that request, then answers it from the root stat. Every later LOOKUP of the name is queued the same way and answered from the remembered root attr, so the dispatcher never holds a session. No eager connection is ever made: `ls` of the root reads the registry directory only (a plain file dropped there is listed too — and yields EIO on the walk, never deleted). A dial that fails answers the walk that asked — `ENOENT` when the entry vanished, `EIO` for a stale entry (connect refused), a full backlog or a server that will not speak 9P — and leaves the mount undialed, so the next walk simply tries again. The dial is the request in flight, so it ends the way any request does: `stop_fd` (the program exited) fails it with `Stopped`; an interrupt of the walk abandons it at once, with nothing sent (`abort_on_cancel`: there is no session yet to flush anything out of) and the walk answers `EINTR`. A server that accepts but never answers `Tversion` therefore costs exactly the walks into its own name, and Ctrl-C ends those. A server whose listen backlog is full (it stopped accepting) makes the connect report `EAGAIN`; that is retried for 5s, polling `stop_fd` and the interrupt pipe between tries, then answers `EIO`. What no design can fix is the kernel side: the VFS serializes lookups of one *name*, so a second walker into the parked name waits in `d_wait_lookup` until the first walk ends — interrupt the first, and the second proceeds (and can be interrupted in its turn). * **FUSE_PARALLEL_DIROPS** is negotiated in the INIT reply. Without it the kernel takes the directory inode's lock around every LOOKUP and READDIR in it, so one parked walk would still hold up every other name under the same directory — the whole registry root, for a mntgen mount — however free the dispatcher is. * **Death and re-dial**: when a worker's session dies mid-request, dispatch has already answered that request (`ESTALE` for a lost connection, so the VFS redoes the path walk instead of failing; `EIO` for a protocol error), the mount is marked dead, its entry is invalidated in the kernel's dentry cache (`FUSE_NOTIFY_INVAL_ENTRY`, best effort), and everything further routed to that subtree answers `ESTALE` (a FORGET is dropped). Nothing reconnects eagerly. The next walk into the name LOOKUPs it again; a dead mount is skipped and the name gets a new mount under a new index, so kernel-held inodes of the corpse keep answering `ESTALE` until forgotten. `ls` still lists the dead name (listing connects to nothing). Death is discovered lazily: the first walk after a silent death takes the error (it marks the mount dead), the next walk re-dials. * **Interrupts** work per mount: the worker's session polls the mount's interrupt pipe while a 9P reply is outstanding; the dispatcher writes the interrupted request's unique into it and the usual `Tflush` dance (see *Interrupts*) follows — with a grace: a server that answers neither the request nor the `Tflush` within 3s (`nine.Session.flush_grace_ms`) is declared gone, the request answers `EINTR`, the session is wedged and the mount dies with it (the next walk makes a new one). The protocol says a client waits for the Rflush; a server that has not managed one in that long is not going to, and the process behind the interrupt is unkillable until we stop waiting. `stop_fd` (the child's death) is watched by every session, so no worker can stay blocked on a hung server past the program's exit; teardown wakes every worker, joins them, and tears down their sessions and bridge state. * The synthetic root is read-only (`dr-xr-xr-x`, like `/srv`): services are posted and unposted by their servers (`cloud9.post`'s `post`/`listenPosted`/`unpost`), not created and removed through files. Capability probes the kernel makes before trusting a file — xattr ops and `STATX` — answer `ENOSYS` (as they do on server subtrees through the ordinary dispatch): `EPERM` would leak into userland as "Operation not permitted" blamed on the mount root by `ls -l` and plain `stat`. Modifying ops (create, mkdir, rename, ...) keep `EPERM` — the root's read-only nature — and unknown ops are refused, never fatal. Root readdir snapshots the registry per OPENDIR (a new OPENDIR sees new posts); because nothing the kernel can cache is ever served stale (zero timeouts, snapshot per opendir), there is nothing to invalidate and no watcher is needed (`post.Watch` exists in the library for future caching). * Naming: `--mount` (default `/mnt/9p`) is the whole mount; `$NINE_MOUNT` points at it as usual. ### Mountpoint policy Default mountpoint: `/mnt/9p/`, where the name is `--name NAME` or is derived from the transport (`main.zig`, `defaultName`): | transport | default name | |---|---| | `--unix PATH` | basename of PATH with one trailing `.sock`/`.9p`/`.socket` removed (`/tmp/9debug.sock` → `9debug`) | | `--tcp IP:PORT` | `tcp-IP-PORT` with `:` → `-` (brackets are already gone: `[::1]:564` → `tcp---1-564`) | | `--spawn CMD` | basename of the first whitespace-separated word of CMD | | `--fd N` | `fdN` | A name is one path component: non-empty, no `/`, no NUL, not `.`/`..`. An invalid `--name` is a usage error (125); a derived name that is not valid (empty basename, `..`) falls back to `9p`. `--mount PATH` overrides all of this (the name is then PATH's last component and is not used for anything); a relative `--mount` is resolved against cwd; `/` is rejected. `ns.ensureMountpoint` then makes the path a directory inside the new namespace: * If the path is a directory: use it (a nested 9ns with the same name therefore mounts over the outer mount at that path). * Else walk up to the deepest existing ancestor (which must be a directory; a dangling symlink anywhere is an error) and `mkdir` the missing components under it one by one. If the first of those fails with `EACCES`/`EPERM`/`EROFS` (the normal case for `/mnt/9p/` as a plain user), **shadow that ancestor**: open an fd to it, mount a `tmpfs` over it, then recreate every existing entry inside the tmpfs: directories → `mkdir` + bind mount from `/proc/self/fd//`; symlinks → `readlinkat` + `symlink`; anything else → empty regular file + bind mount. Then create the missing components inside. Refuse (with a clear message) if the ancestor is `/` (also through `/proc/self/root`), is `/proc` or below it, or has more than 4096 entries. This only affects the new namespace. So on a host without `/mnt/9p` the shadow goes over `/mnt` and `9p/` is created inside; with a root-owned `/mnt/9p` it goes over `/mnt/9p`; inside a 9ns namespace `/mnt/9p` is a directory of the outer tmpfs owned by our uid, so a nested 9ns just creates `` next to the outer mount and both are visible. * Else fail with the errno and a hint to pass `--mount` an existing dir. ## Module contracts ### `src/fuse.zig` — kernel FUSE protocol (no policy) Extern structs mirroring `linux/fuse.h`, with `comptime` size asserts: `InHeader` (40), `OutHeader` (16), `Attr` (88), `EntryOut` (128), `AttrOut` (104), `GetattrIn` (16), `SetattrIn` (88), `OpenIn` (8), `OpenOut` (16), `ReleaseIn` (24), `FlushIn` (24), `ReadIn` (40), `WriteIn` (40), `WriteOut` (8), `CreateIn` (16), `MkdirIn` (8), `RenameIn` (8), `Rename2In` (16), `ForgetIn` (8), `BatchForgetIn` (8), `ForgetOne` (16), `FsyncIn` (16), `AccessIn` (8), `InterruptIn` (8), `Kstatfs` (80), `StatfsOut` (80), `InitIn` (64), `InitOut` (64), `Dirent` (24 header, name padded to 8), `LseekIn` (24). `pub const Opcode = enum(u32) { lookup = 1, forget = 2, getattr = 3, setattr = 4, readlink = 5, symlink = 6, mknod = 8, mkdir = 9, unlink = 10, rmdir = 11, rename = 12, link = 13, open = 14, read = 15, write = 16, statfs = 17, release = 18, fsync = 20, setxattr = 21, getxattr = 22, listxattr = 23, removexattr = 24, flush = 25, init = 26, opendir = 27, readdir = 28, releasedir = 29, fsyncdir = 30, getlk = 31, setlk = 32, setlkw = 33, access = 34, create = 35, interrupt = 36, bmap = 37, destroy = 38, ioctl = 39, poll = 40, notify_reply = 41, batch_forget = 42, fallocate = 43, readdirplus = 44, rename2 = 45, lseek = 46, copy_file_range = 47, setupmapping = 48, removemapping = 49, syncfs = 50, tmpfile = 51, statx = 52, _ }` Constants: `kernel_version = 7`, `kernel_minor = 31` (what we answer; the kernel adapts to the lower minor), `FOPEN_DIRECT_IO = 1`, `FOPEN_KEEP_CACHE = 2`, `FOPEN_NONSEEKABLE = 4`, `FUSE_ASYNC_READ = 1`, `FUSE_PARALLEL_DIROPS = 1<<18`, `FUSE_MAX_PAGES = 1<<22`, `FATTR_MODE=1, FATTR_UID=2, FATTR_GID=4, FATTR_SIZE=8, FATTR_ATIME=16, FATTR_MTIME=32, FATTR_FH=64, FATTR_ATIME_NOW=128, FATTR_MTIME_NOW=256, FATTR_LOCKOWNER=512, FATTR_CTIME=1024`. `root_id = 1`. I/O helpers (blocking fd, no allocation beyond the caller's buffer): ```zig pub const Request = struct { header: InHeader, body: []const u8 }; /// One kernel request. Returns null on ENODEV (unmounted). Retries EINTR/EAGAIN/ENOENT. pub fn readRequest(fd: i32, buf: []u8) !?Request; /// Same, one read(2) only: EINTR/EAGAIN/ENOENT → error.Retry (for reads that follow a poll). pub fn readRequestOnce(fd: i32, buf: []u8) !?Request; /// O_NONBLOCK on the device, so a request withdrawn between poll and read cannot block us. pub fn setNonblocking(fd: i32) !void; /// Success reply: header + concatenated payload slices, single writev. pub fn reply(fd: i32, unique: u64, payloads: []const []const u8) !void; /// Error reply: negative errno. pub fn replyError(fd: i32, unique: u64, err: std.os.linux.E) !void; /// Append a fuse_dirent (8-byte padded) to `buf`; returns false if it doesn't fit. pub fn addDirent(buf: []u8, used: *usize, ino: u64, off: u64, dtype: u32, name: []const u8) bool; pub fn body(comptime T: type, req: Request) !*const T; // aligned copy-free view, checks size pub fn nameAfter(comptime T: type, req: Request) ![]const u8; // NUL-terminated name after a struct ``` The request buffer must be ≥ `max_write + 4096`; 9ns uses 1 MiB + 4 KiB. Requests with an unknown/unsupported opcode get `ENOSYS`. ### `src/nine.zig` — synchronous 9P2000 session on a blocking fd Thin, synchronous RPC layer over `cloud9.Client` (which is push/take, non-blocking-agnostic). One outstanding request at a time (the FUSE loop is single-threaded). Fids are allocated from a free list. ```zig pub const Address = union(enum) { unix: []const u8, tcp: struct { host: []const u8, port: u16 }, fd: i32 }; /// Owner's hook into the reply wait (the bridge's FUSE fd); see "Interrupts" under bridge.zig. pub const Interrupt = struct { ctx: *anyopaque, watch: *const fn (ctx) i32, // fd to poll alongside the socket, or -1 onReadable: *const fn (ctx) Session.Error!bool, // consume it; true = flush the request in flight armed: *const fn (ctx) bool, // a flush was asked for the operation in progress }; pub const Session = struct { pub const Error = error{ Nine, Protocol, Io, Closed, Stopped, Interrupted, TooLarge, OutOfMemory }; /// After error.Nine, `ename` holds the server's Rerror text (copied, bounded). ename: [256]u8, ename_len: usize, msize: u32, stop_fd: i32 = -1, // readable → pending rpc fails with error.Stopped interrupt: ?Interrupt = null, pub fn connect(gpa: std.mem.Allocator, address: Address, msize: u32) !Session; // socket+connect, version pub fn deinit(s: *Session) void; pub fn attach(s: *Session, fid: u32, uname: []const u8, aname: []const u8) Error!cloud9.Qid; pub fn allocFid(s: *Session) u32; pub fn freeFid(s: *Session, fid: u32) void; /// Generic RPC. Result slices borrow the input buffer until the next call. pub fn rpc(s: *Session, req: cloud9.Client.Request) Error!cloud9.Client.Result; // Conveniences (all built on rpc): pub fn walk(s, fid: u32, newfid: u32, names: []const []const u8) Error!Walk; // Walk = { nwqid, wqid[16] }; partial walk → error.Nine with ename "not found"-ish pub fn clone(s, fid: u32) Error!u32; // allocFid + walk with no names pub fn open(s, fid: u32, mode: u8) Error!Open; // Open = { qid, iounit } pub fn create(s, fid: u32, name: []const u8, perm: u32, mode: u8) Error!Open; pub fn read(s, fid: u32, offset: u64, buf: []u8) Error!usize; // chunks by maxRead/iounit; stops at short read pub fn write(s, fid: u32, offset: u64, data: []const u8) Error!usize; // chunks; stops at short write pub fn stat(s, fid: u32) Error!cloud9.Stat; // strings borrow the input buffer pub fn wstat(s, fid: u32, st: cloud9.Stat) Error!void; pub fn clunk(s, fid: u32) Error!void; // frees the fid even on error pub fn remove(s, fid: u32) Error!void; // frees the fid even on error pub fn errno(s: *const Session) std.os.linux.E; // map ename → errno (see below) }; pub const dontcare = cloud9.Stat{ .type = 0xFFFF, .dev = 0xFFFF_FFFF, .qid = .{ .type = 0xFF, .version = 0xFFFF_FFFF, .path = 0xFFFF_FFFF_FFFF_FFFF }, .mode = 0xFFFF_FFFF, .atime = 0xFFFF_FFFF, .mtime = 0xFFFF_FFFF, .length = 0xFFFF_FFFF_FFFF_FFFF, .name = "", .uid = "", .gid = "", .muid = "" }; ``` `connect`: for `.unix` and `.tcp` create a blocking `SOCK_STREAM|SOCK_CLOEXEC` socket and connect (`TCP_NODELAY` on TCP); for `.fd` adopt it. Buffers of `msize` bytes for in/out are heap allocated. Then submit `.version`, drain output to the socket, read until `take()` yields the version result. The negotiated msize is `result.version.msize`; if the server answered `"unknown"`, fail with `error.Protocol`. `rpc`: submit, write all of `client.output()` (calling `wrote`), then loop: `take()`; if null, poll the fd together with `stop_fd` and the interrupt source's descriptor; when the fd is readable, `read` into a temp buffer and `push` (push returns how much fit; the frame is at most msize so it always fits after a `take`). If the fd returns 0 → `error.Closed`. If the client dies → `error.Protocol`. `stop_fd` readable → `error.Stopped`. When the interrupt source asks for it, submit `.flush = .{ .oldtag = tag }` and keep waiting: the original reply → returned normally (the Rflush that follows is skipped by a later call); the Rflush first → `error.Interrupted`. A `.fail` result copies the ename and returns `error.Nine`. `read`/`write` chunk loops stop between chunks once the source reports `armed`, and turn an `Interrupted` chunk into a short count when earlier chunks moved data. Rerror text → errno mapping (case-insensitive substring, in this order): `"interrupt"` → `EINTR` (a server answering a flushed request with an error); `"not exist"`, `"not found"`, `"no such"` → `ENOENT`; `"exists"` → `EEXIST`; `"not empty"` → `ENOTEMPTY`; `"not a dir"` → `ENOTDIR`; `"is a dir"` → `EISDIR`; `"permission"`, `"denied"` → `EACCES`; `"read-only"`, `"read only"`, `"readonly"` → `EROFS`; `"no space"` → `ENOSPC`; `"not allowed"`, `"not permitted"`, `"cannot"` → `EPERM`; `"fid"` → `EBADF`; `"bad offset"`, `"invalid"`, `"bad "` → `EINVAL`; `"busy"`, `"in use"` → `EBUSY`; `"too long"` → `ENAMETOOLONG`; `"not supported"`, `"unsupported"` → `ENOTSUP`; otherwise `EIO`. ### `src/bridge.zig` — FUSE ↔ 9P translation ```zig pub const Options = struct { uid: u32, gid: u32, // reported owner of every file attr_timeout_ns: u64 = 1e9, // attr/entry cache validity (0 = none) direct_io: bool = true, // FOPEN_DIRECT_IO on every regular file debug: bool = false, // trace to stderr }; /// Runs until the FUSE fd reports ENODEV or `stop_fd` becomes readable. pub fn serve(gpa: std.mem.Allocator, fuse_fd: i32, nine: *nine.Session, root_fid: u32, stop_fd: i32, opts: Options) !void; // mntgen (see "mntgen: one mount, many servers" under Process model): pub const MntgenOptions = struct { io: std.Io, // dispatcher-thread only: post.posted / registry stats env: post.Env, // XDG_RUNTIME_DIR names the registry uname: []const u8, aname: []const u8 = "", msize: u32 = 131072, }; /// One FUSE mount whose root lists the posted-9P registry; servers dialed /// lazily, one worker thread each; node ids carry the mount index in the /// top bits (`mount_shift = 32`, indexes are ordinals from 1, never reused, /// `max_mounts` = 4096; the registry snapshot buffer is 8 KiB, `max_root_dirs` /// = 64 concurrent OPENDIRs). Runs until ENODEV, DESTROY or `stop_fd`. pub fn serveMntgen(gpa: std.mem.Allocator, fuse_fd: i32, stop_fd: i32, mo: MntgenOptions, opts: Options) !void; pub fn mountNode(index: u32, local: u64) u64; // (index << 32) | local pub fn mountIndex(nodeid: u64) u32; // nodeid >> 32 pub fn nameIno(name: []const u8) u64; // FNV-1a of a name: synthetic-root dirent inos ``` State: * `inodes: AutoHashMap(u64 /*nodeid*/, Inode{ fid: u32, qid: Qid, nlookup: u64 })`. Node 1 is the root (`root_fid`, never forgotten). In mntgen mode the same machinery runs once per mount with node ids that already carry the mount index; the root node id is the `Bridge.root_id` field (`fuse.root_id` single-connection) and reported inode numbers are `qid.path ^ ino_xor` (`ino_xor` 0 single-connection). `cur_unique` is atomic so the mntgen dispatcher can scan it to route FUSE_INTERRUPTs. * `by_qid: AutoHashMap(u64 /*qid.path*/, u64 /*nodeid*/)` so that repeated lookups of the same file map to the same inode (the old fid is clunked and the fresh one kept). Dedupe only merges when the qid type (dir bit) also matches, so a server reusing a path across a file and a directory cannot poison an inode. `ino` in attrs is `qid.path` (root, or anything carrying the root's path: 1). * `handles: AutoHashMap(u64 /*fh*/, Handle{ fid: u32, dir: ?DirList })`. `DirList` is the entire directory read at first `READDIR` offset 0: `[]Entry{ name: []u8, ino: u64, dtype: u32 }` with synthetic `.` and `..` first. `READDIR` offsets are indices into that list; a `READDIR` at offset 0 re-reads the directory (rewinddir). Op mapping (9P2000 has no symlinks, links, xattrs, locks, mknod): | FUSE | 9P | |---|---| | INIT | reply `InitOut{ major=7, minor=31, max_readahead=in.max_readahead, flags = FUSE_ASYNC_READ \| FUSE_ATOMIC_O_TRUNC \| FUSE_AUTO_INVAL_DATA \| FUSE_BIG_WRITES (plus FUSE_MAX_PAGES with max_pages=256, and FUSE_PARALLEL_DIROPS, each if offered), max_background=16, congestion_threshold=12, max_write=1 MiB, time_gran=1 }`. Atomic O_TRUNC matters: without it the kernel truncates via a separate SETATTR(size=0) that synthetic control files reject; with it `O_TRUNC` becomes 9P `OTRUNC` inside the open | | LOOKUP(parent,name) | `walk(parent.fid → newfid, [name])`; `stat(newfid)`; dedupe by qid; `EntryOut` | | FORGET / BATCH_FORGET | `nlookup -= n`; at 0 `clunk` and drop (no reply) | | GETATTR | `stat(inode.fid)` → `AttrOut` | | SETATTR | `stat` then `wstat` with a *dontcare* Stat: `FATTR_SIZE`→length; `FATTR_MODE`→`(old.mode & ~0o777) \| (mode & 0o777)`; `FATTR_MTIME`→mtime (`FATTR_MTIME_NOW` → now); `FATTR_ATIME` ignored; `FATTR_UID/GID` → `EPERM` unless unchanged; then `stat` again for the reply | | OPEN | `clone(inode.fid)` then `open(newfid, mode)`; mode from `O_ACCMODE` (`oread/owrite/ordwr`), `O_TRUNC` → `otrunc`; reply `OpenOut{ fh, open_flags = FOPEN_DIRECT_IO }`; on failure clunk | | OPENDIR | same with `oread`; `fh` with `dir = null` | | READ | `read(fh.fid, offset, buf[0..min(size, 1 MiB)])`; reply data | | WRITE | `write(fh.fid, offset, data)`; `WriteOut{ size = n }` | | READDIR | fill `Dirent`s from the `DirList` starting at `offset`, up to `size` bytes | | RELEASE / RELEASEDIR | `clunk(fh.fid)`; free DirList | | FLUSH / FSYNC / FSYNCDIR | ok (no-op) | | CREATE(parent,name,flags,mode) | `clone(parent)`; `create(fid, name, mode & 0o777, openmode)` → this fid is the **open** file; then `walk(parent → fid2, [name])` + `stat(fid2)` for the inode; reply `EntryOut ++ OpenOut` | | MKDIR | `clone(parent)`; `create(fid, name, DMDIR \| (mode & 0o777), oread)`; `clunk`; then lookup as above | | UNLINK / RMDIR | `walk(parent → tmp, [name])`; `remove(tmp)` | | RENAME / RENAME2 | if `newdir != parent` → `EXDEV`; else `walk(parent → tmp, [oldname])`, `wstat(tmp, dontcare with .name = newname)`, `clunk`. 9P rename never replaces, POSIX does: when the target exists (and `RENAME_NOREPLACE` is not set) a directory target is removed first; a file target is parked under a temporary name, the rename retried, and the parked file removed only after success (restored on failure) | | STATFS | constant `Kstatfs{ bsize = 4096, namelen = 255, frsize = 4096 }` | | ACCESS | `ENOSYS` (kernel stops asking; the server enforces permissions on open) | | READLINK, SYMLINK, LINK, MKNOD, *XATTR, *LK, IOCTL, POLL, BMAP, FALLOCATE, LSEEK, COPY_FILE_RANGE, TMPFILE, STATX | `ENOSYS` | | INTERRUPT | read while a 9P reply is outstanding: for the request in flight → `Tflush` (see *Interrupts*); otherwise ignored (reply nothing) | | DESTROY | return from `serve` | Attr mapping from `cloud9.Stat`: `mode = (S_IFDIR if DMDIR else S_IFREG) | (st.mode & 0o777)`; `nlink = 1`; `size = length`; `blocks = (length+511)/512`; `blksize = 4096`; `atime/mtime/ctime = st.atime/st.mtime/st.mtime`; `uid/gid = opts.uid/gid`. `Dirent.type` = `DT_DIR` (4) / `DT_REG` (8). Errors: `nine.Session.Error.Nine` → `nine.errno()`; `Closed`/`Protocol`/`Io` → `EIO` and, since the session is dead, `serve` returns `error.Closed` after replying so 9ns can report "9P server went away". With `direct_io` the kernel never trusts `length` for reads: synthetic files that report length 0 (very common in 9P) still `cat` correctly, and reads run until the server returns a short read. With `--no-direct-io` the bridge forces `attr_timeout_ns = 0`, because a cached stale size truncates reads (observed data loss on a 4 MiB copy otherwise). Hostile-server rules: directory listings are capped at 64 MiB (a server that ignores read offsets otherwise loops forever); directory records with names containing `/`, NUL, empty, `.`/`..` or longer than `FUSE_NAME_MAX` are dropped rather than poisoning the whole READDIR reply; `length` near 2^64 is clamped to `i64` max; the errno of a failing 9P call is latched before any cleanup clunk overwrites the session's ename. #### Interrupts (FUSE_INTERRUPT → Tflush) One request at a time, but not deaf. `serve` puts the FUSE fd in `O_NONBLOCK` mode (every read follows a poll) and installs a `nine.Interrupt` source that `Session.rpc` polls together with the socket and `stop_fd` whenever a reply is outstanding, including the initial root stat: * `onReadable` reads the request the kernel has ready into a second 8-aligned buffer of `request_buf_len` bytes (`spare_buf`). A `FUSE_INTERRUPT` whose `InterruptIn.unique` names the request being served (`cur_unique`) sets `interrupted` and returns true: `rpc` sends `Tflush(oldtag)` and keeps waiting until either the original reply arrives (the interrupt raced it: the result is returned as if nothing happened and the Rflush that follows is swallowed by a later call) or the Rflush does (`error.Interrupted` → `EINTR`; a server that instead answers the flushed request with an Rerror containing "interrupt", as Pardes does, lands on the same errno through the ename table). An INTERRUPT for any other unique is consumed and dropped (the kernel expects no reply). Any other request (FORGET, RELEASE, INIT during the root stat, a second process's LOOKUP) is copied into a queue that `serve` dispatches, in arrival order, before it polls again; the fd stays watched throughout, so an INTERRUPT is never stuck behind a parked request (a copy that cannot be allocated answers `ENOMEM` on the spot). * The INTERRUPT applies to `cur_unique` only and is consumed when read: the clunks that unwind a half-done lookup/create/mkdir after an `Interrupted` walk or stat are ordinary rpcs and are not re-interrupted by it. An interrupted walk clunks its new fid (the Rflush alone does not say whether the server bound it); an Rerror'd walk only frees it locally. * Multi-step operations fail with `EINTR` at whichever rpc was flushed and release the fids they had allocated (`--debug` prints `fids=N` per request; the interrupt suite checks it and the server-side count). The chunk loops (`Session.read/write`, `loadDir`) also stop between chunks while `armed`, so a reply that won the race cannot lead into another blocking chunk: a read or write that already moved data returns the partial count like read(2), one that moved nothing and a directory listing return `EINTR`. * The kernel sends one INTERRUPT per request, for a fatal signal (SIGKILL included) as much as for a caught one, and then waits for the reply; our `EINTR` is what finally lets the killed task die. Limits: a server that ignores Tflush is given `flush_grace_ms` (3s) and then declared gone — flush(5) says the server "should answer the Tflush message immediately", and a client that waits longer leaves an unkillable process behind — so the reader gets `EINTR` and the session ends (the hostile `never` mode; in single-connection mode that is the mount, in mntgen that one name). A slow-but-alive server that cannot answer a flush inside its own blocking read pays the same price; the alternative was a hang that only SIGKILL of 9ns could end. ### `src/ns.zig` — namespace and process plumbing ```zig pub const Spawn = struct { argv: []const []const u8, // argv[0] is PATH-searched unless it contains '/' envp: [*:null]const ?[*:0]const u8, // inherited environment mountpoint: []const u8, // absolute fuse_fd: i32, uid: u32, gid: u32, max_read: u32, }; pub const Child = struct { pid: i32 }; /// fork; the child sets up the namespace, mounts, and execs. Returns once exec succeeded /// (status pipe closed) or fails with the child's error (message on stderr). pub fn spawn(gpa: std.mem.Allocator, s: Spawn) !Child; pub fn ensureMountpoint(gpa: Allocator, path: [:0]const u8) !void; // walk down, mkdir -p, shadow; testable alone pub fn resolveMountpoint(gpa, path: []const u8) ![:0]u8; // absolute, no trailing slash pub fn findInPath(gpa, envp, name) ![:0]u8; ``` Also exports the signal plumbing used by `main.zig`: `installSignals(child_pid_ptr: *i32) !i32` returning the SIGCHLD self-pipe read end (used as `stop_fd` for `bridge.serve`), and `waitChild(pid) !u8` → exit status (`128+sig` on signal death). ### `src/main.zig` — CLI ``` Usage: 9ns [options] -- PROGRAM [ARGS...] Transport (exactly one): --unix PATH Unix stream socket --tcp IP:PORT TCP (IPv4/IPv6 literal) --fd N already-connected inherited descriptor --spawn CMD run CMD (via /bin/sh -c) with a socketpair on its stdin/stdout --mntgen mount the posted-9P registry ($XDG_RUNTIME_DIR/9p): one mount whose root lists the posted names; walking into a name dials that server (mutually exclusive with the rest) Options: --name NAME mount name: the tree appears at /mnt/9p/NAME (one path component; default derived from the transport, see below; not with --mntgen) --mount PATH mountpoint inside the new namespace (overrides --name; with --mntgen the mount is the registry view itself, default /mnt/9p) --uname NAME 9P user name (default $USER, else "none") --aname NAME 9P tree to attach (default "") --msize BYTES maximum 9P message size to request (default 131072, max 16 MiB) --cache SECONDS attr/entry cache validity, may be fractional (default 1) --no-direct-io let the kernel cache file pages (trusts stat length) --debug trace FUSE and 9P operations on stderr --help, --version PROGRAM defaults to $SHELL (else /bin/sh). The mountpoint is exported as $NINE_MOUNT. Default name: --unix PATH -> basename of PATH without .sock/.9p/.socket; --tcp IP:PORT -> tcp-IP-PORT (':' becomes '-'); --spawn CMD -> basename of its first word; --fd N -> fdN; 9p when nothing usable comes out of that. --mntgen: no per-server name; the registry mount goes to --mount (default /mnt/9p). ``` `--name` and `--mount` may both be given; `--mount` wins. Exit codes: child's status; 125 for 9ns's own failures (usage including a bad `--name`, connect, mount, and for `--mntgen` an unset XDG_RUNTIME_DIR); 126/127 as usual for exec failures. ### `../9proc/demo/main.zig` — demo 9P2000 server (binary `9proc-demo`) The demo server is a separate program in this repository, built on the 9proc library; see `../9proc/docs/LIBRARY.md` for the library contract (freestanding core, value renderers, Linux debug probe). The tree it serves keeps the paths the integration tests read (`/build/*`, `/comptime/types//*`, `/comptime/decls`, `/runtime/fn/*`, `/runtime/ctl`, `/runtime/{pid,ppid,uptime,argv,cwd,env,clients}`, `/scratch/`) and adds `/vars`, `/threads`, `/addr`, `/mem`, `/hex`, `/breakpoints` and `/panic`. ## Integration test plan (`test/integration.sh`) Run by `zig build 9ns-itest`; args: path to `9ns`, path to `9proc-demo`. Everything under a temp dir. Skips (exit 0 with a notice) when `unshare -Urm true` fails or `/dev/fuse` is missing. 1. 9proc-demo on a Unix socket; `9ns --unix … -- sh -c` scripts: `cat /mnt/9p/build/zig_version` == `zig version`; `ls` listings; `stat` sizes; `/runtime/fn/now` is numeric; `ctl` round trip; `/scratch`: create, append (`>>`), overwrite, truncate, `mkdir -p a/b/c`, rename within dir, `mv` across dirs fails with `EXDEV`-ish message, `rm`, `rmdir`, 1 MiB random file round trip compared with `sha256sum`, `dd` with odd block sizes, many small files, `find`, exit-status propagation (`exit 7` → 7), `$NINE_MOUNT` set, nested `9ns` inside `9ns`. The suites pin the mountpoint with `--mount /mnt/9p`; the naming section then checks the default `/mnt/9p/` (unix socket basename, `--name`, `--name=`, invalid names → 125, nested runs with two names both visible under `/mnt/9p`, the same name twice mounting over, `--spawn` and `--tcp` derived names, `--mount` beating `--name`). 2. `--spawn "<9proc-demo> --stdio"` variant. 3. `--tcp 127.0.0.1:` variant. 4. If `/usr/lib/plan9/bin/ramfs` exists: `NAMESPACE=$tmp ramfs -s ramfs` creates `$tmp/ramfs`; run the scratch battery against it. 5. `--mount` with an existing dir, with a relative path, `/mnt/9p` and the default `/mnt/9p/` (both exercise the shadowing of `/mnt`; verify `/mnt`'s other entries are still visible inside and the host mount table is untouched). 6. Kill tests: 9ns exits when the child exits; server death during use yields `EIO`, not a hang. 7. `test/mntgen.sh` (also `9ns-itest`): `--mntgen` against a scratch registry (`XDG_RUNTIME_DIR` = temp dir, never the real one): two servers posted under two names (9proc-demo's socket created inside the registry directory — a socket at `$XDG_RUNTIME_DIR/9p/` is a posted name — plus plan9port `ramfs`, an independent 9P2000 implementation, posting with `NAMESPACE=$REG`), the synthetic root listing without dialing (including a non-socket file, which is listed, yields EIO on the walk, and is never removed), lazy dial and per-server routing in one program run, a post appearing after mount, server death → EIO on the subtree with the name still listed and a re-dial after re-post, parallel reads on both mounts, `$NINE_MOUNT` = `/mnt/9p` by default, and the usage errors (`--mntgen` + any transport, `--name`, `--mntgen=x`, unset `XDG_RUNTIME_DIR`). ## Verification `zig build 9ns-test` (unit), `zig build 9ns-itest` (integration.sh + the mntgen suite: end-to-end checks against 9proc-demo over unix/tcp/socketpair, plan9port's `ramfs`, and the posted-registry multi-server mode) and `zig build 9ns-adv` (adversarial suites: a scriptable hostile 9P server with ~30 misbehaviour modes, interrupt forwarding against its `never_flush`/`never` modes with 28 checks, FUSE semantics through the bridge, process/namespace/signal edge cases with 51 checks, and stress). The suites that attack the 9proc server itself (a hostile raw-9P client with 181 checks, the core, the Linux layer) moved with it to `../9proc/test` (`zig build 9proc-adv`). All pass in Debug and ReleaseSafe. ## Out of scope for v1 (documented, not hidden) * One 9P request in flight at a time, per server in mntgen mode (across servers they proceed in parallel, one worker each): a 9P read that blocks (event files) stalls that server's subtree while it is outstanding (but not past the child's exit). It can be interrupted: killing or Ctrl-C-ing the reader sends `FUSE_INTERRUPT`, which becomes `Tflush`; servers that honour it unblock immediately, servers that don't are given 3s and then declared gone: the reader gets `EINTR` and the session (single connection: the mount; mntgen: that name) ends. * mntgen: a name's dial runs on that name's worker, inside the walk that asked, with no timeout of its own beyond the 5s backlog retry: a server that accepts and never answers `Tversion` parks the walks into its own name, and only those, until each is interrupted or the program exits. The one wait nothing here can cut short is the kernel's own serialization of lookups of a single name (a second walker into the parked name waits for the first walk to end, uninterruptibly). * mntgen: mount indexes are ordinals and never reused (cap 4096 dials per process); per-mount node ids cap at 2^32 lookups. `st_ino` mixes the index with the qid.path via XOR — collisions remain theoretically possible, just not the practical ones (identical qid.paths across two servers). A dead mount releases its socket, its interrupt pipe and its bridge state at once (the pipe under the mount's mutex, where the dispatcher writes it) and keeps only its slot, so a server that dies and comes back costs one index per death and nothing else; after 4096 of them in one 9ns process no further name can be walked into (EIO) until the shell is restarted. Reclaiming a slot once the kernel has forgotten every node of the corpse is the next step, not taken here. * mntgen: no posting/unposting through the mount (the synthetic root is read-only; servers manage their registry entries through `cloud9.post`), and no per-name mount options: one set of `--uname/--aname/--msize/--cache` applies to every dial. * No 9P2000.u/.L: no symlinks, ownership, or extended attributes. * No PID namespace, no `/proc` remount. `--tcp` needs an IP literal. * Cross-directory rename returns `EXDEV` (9P2000 cannot move files).