# 9proc: a 9P debug/introspection server as a library The demo server that 9ns's tests use grows into a library any Zig program can embed: a debugger-shaped interface where the protocol is just files. Anything that can read a filesystem (a shell, an agent, an editor, `9p`, 9ns) can inspect a running process: build facts, comptime type layouts, live values, threads and their stacks, memory, breakpoints, panics. Design rules (non-negotiable, they mirror cloud9): 1. **The core is freestanding.** No allocator, no OS, no threads, no `std.Io`. Caller-owned buffers, fixed-capacity tables sized at comptime. It must compile for `riscv32-freestanding-none` (the ESP32-P4 firmware target, `../05-zig-p4`), and `zig build 9proc-check-freestanding` proves it. 2. **Every dependency on a runtime is an explicit argument.** Features that truly need an `Allocator` or an `std.Io` take them in their `init`; nothing reaches for `std.heap.page_allocator` or a global `Io`. Where memory is needed it is preferably a caller-provided `[]u8` or a comptime-sized `Storage` struct the caller places in static memory. 3. **All allocation happens up front**, at init, from what the caller passed. Steady-state operation does not allocate. 4. **Platform layers are separate modules** (`9proc.linux`) and are the only places that touch sockets, threads, signals, `/proc` or `std.debug`. ``` 9proc/src/root.zig pub const core, vars, scratch, linux (linux only), Server(cfg) 9proc/src/core.zig the tree (static nodes, vars, providers) as a backend of cloud9.fs.Server 9proc/src/vars.zig comptime value renderers (@typeInfo) for /vars 9proc/src/scratch.zig in-memory read/write tree provider (takes an Allocator) 9proc/src/linux/probe.zig background thread + poll loop + unix/tcp/fd listeners 9proc/src/linux/debug.zig threads, stacks, registers, addr→source, memory, breakpoints, panic 9proc/demo/main.zig the `9proc-demo` binary: embeds everything, worker thread, exposed vars (paths from the cloud9 root; the library is the module `9proc` that cloud9's `build.zig` exports next to `cloud9`, wired by `9proc/build.zig`) ``` ## Core (`core.zig`) ```zig pub const Config = struct { name: []const u8 = "9proc", // /README and Stat uid/gid types: []const type = &.{}, // /comptime/types//... decls_of: ?type = null, // /comptime/decls lists this type's pub decls fns: type = struct {}, // /runtime/fn/: pub fn (ctx: *anyopaque, w: *std.Io.Writer) anyerror!void ctl: ?*const fn (ctx: *anyopaque, cmd: []const u8, out: *std.Io.Writer) anyerror!void = null, // /ctl max_fids: u16 = 64, max_providers: u8 = 8, max_vars: u8 = 32, /// Dynamic file contents are generated at open time into per-fid snapshot /// slots so that reads at arbitrary offsets are consistent. snapshot_slots: u8 = 8, snapshot_bytes: u32 = 16 * 1024, max_parked: u8 = 8, // reads a provider has parked with error.Again, per connection }; pub fn Server(comptime cfg: Config) type { return struct { pub const Storage = struct { // caller places this in static memory in: [msize]u8, out: [msize]u8, data: [msize]u8, snapshots: [cfg.snapshot_slots][cfg.snapshot_bytes]u8, }; pub const Shared = struct { // state common to all connections (providers, vars) pub fn init(name_ctx: *anyopaque) Shared; pub fn addProvider(s: *Shared, p: Provider) error{Full}!void; pub fn expose(s: *Shared, name: []const u8, ptr: anytype) error{Full}!void; // typed value → /vars/ }; pub const Backend = struct { ... }; // what cloud9.fs.Server knows of the tree: Req, Reply, features pub const Engine = cloud9.fs.Server(Backend, .{ .fid_capacity = cfg.max_fids, .slot_capacity = cfg.max_parked, ... }); pub const Conn = struct { // one 9P connection: an Engine plus the tree's per-connection state pub fn init(shared: *Shared, storage: *Storage, msize: u32) Conn; pub fn push(c: *Conn, bytes: []const u8) usize; // feed transport bytes pub fn step(c: *Conn) error{Protocol}!bool; // serve ≤ 1 engine request; false = nothing to do pub fn output(c: *const Conn) []const u8; // bytes to send pub fn wrote(c: *Conn, n: usize) void; pub fn hangup(c: *Conn) void; // drop fids, tell providers pub fn fidCount(c: *const Conn) usize; }; }; } ``` The core is a **backend of `cloud9.fs.Server`**, the library's file-server engine (cloud9's `docs/design.md`, "File server engine"): one engine for every 9P server built on cloud9. The engine owns the fid table (indexed, `Options.fid_index`, so thousands of fids cost O(1) per lookup), walks, permission checks, directory cursors, Tflush, and the releases a dropped connection owes; `Conn` wraps one engine instance and answers its requests (`lookup`, `getattr`, `setattr`, `open`, `read`, `write`, `release`, `readdir`) from the tree: the static part (comptime-generated from `cfg`: `/README`, `/build/*` via a `build_options`-like struct passed in `cfg.build`, `/comptime/types/*`, `/comptime/decls`, `/runtime/fn/*`, `/ctl`, `/vars/*`) is served directly, **providers** through their vtable. The backend declares every optional engine feature (`fs.Features`: create, remove, wstat, references), which is how Tcreate/Tremove/Twstat and the one-clunk-per-handle contract below reach providers. `step` answers the engine's `retry()` and `next()` requests synchronously, one `next()` per call; a provider that cannot answer yet parks instead (see "Answering later"). A provider is a runtime vtable mounted at a top-level name. It owns a subtree with its own naming (dynamic directories such as `/threads/` or `/addr/` cannot be enumerated at comptime): ```zig pub const Provider = struct { name: []const u8, ctx: *anyopaque, vtable: *const VTable, pub const Handle = u64; // provider-defined node id; 0 = provider root pub const VTable = struct { walk: *const fn (ctx, parent: Handle, name: []const u8) Error!Handle, stat: *const fn (ctx, h: Handle, out: *NodeStat) Error!void, // kind (dir/file), mode, length, mtime list: *const fn (ctx, dir: Handle, index: usize, out: *NodeStat) Error!bool, // nth entry; false when done open: *const fn (ctx, h: Handle, mode: u8) Error!void, read: *const fn (ctx, h: Handle, offset: u64, buf: []u8) Error!usize, write: *const fn (ctx, h: Handle, offset: u64, data: []const u8) Error!usize, create: ?*const fn (ctx, dir: Handle, name: []const u8, perm: u32, mode: u8) Error!Handle, remove: ?*const fn (ctx, h: Handle) Error!void, wstat: ?*const fn (ctx, h: Handle, st: *const cloud9.Stat) Error!void, clunk: *const fn (ctx, h: Handle) void, // fid released (also on hangup) }; pub const Error = error{ NotFound, Exists, Perm, NotDir, IsDir, NotEmpty, BadOffset, NoSpace, Io, Unsupported }; }; ``` Error → Rerror text: what the tree or a provider refuses is answered in `ename`'s Plan 9 strings, which 9ns's bridge understands (`file does not exist`, `permission denied`, `file already exists`, `directory not empty`, `not a directory`, `is a directory`, `bad offset`, `no space`, `i/o error`, `not supported`, `bad command`, `bad value`, `too many open dynamic files`); what the engine refuses on its own carries cloud9.fs's strings, the ones Linux v9fs maps back to errnos (`fid unknown or out of range`, `fid already in use`, `Too many open files in system`, `bad use of fid` for I/O on an unopened fid or a walk or clone from an open one, `file already open for I/O`, `bad offset in directory read`, `permission denied` for an open the walked mode forbids (a directory for writing, OEXEC, a static file for writing), `wstat prohibited` for any of the fields the engine owns (type, dev, qid, atime, the owner names) or a length on a directory, `illegal name`, `Invalid argument` for an Rstat that cannot fit the msize or a directory read whose count holds no whole record). The smallest msize is the engine's `fs.msize_min` (217: one full Rwalk); a Tversion below it ends the connection. Directory reads follow the 9P rule (offset 0 or previous offset+count, never split a record); the engine encodes the entries from the tree's records, with uid/gid/muid the attach uname, the modes `dirent_dir_perm`/ `dirent_file_perm` and length 0 (a stat of the entry gives the real ones). Dynamic file reads: on open the content is generated once into a snapshot slot (`open` runs the generator; `read` serves the slot; a read at offset 0 regenerates); no free slot → Rerror `too many open dynamic files`. Stats of dynamic files report length 0. Qids: static nodes get comptime paths; provider nodes get `(provider index << 56) | handle` (or `NodeStat.path` in place of the handle). The engine's node ids are the same numbers, except that provider ids are offset by one in the top byte, since the engine reads node 0 as "the node asked about". **Answering later.** `Provider.read` may return `error.Again`: the engine parks the request (`Config.max_parked` per connection; a further one fails with EAGAIN), the connection goes on serving, and every later `Conn.step` asks the provider again with the same handle and offset until it answers, whether at once or on a later step. The fid stays open; a Tflush of a parked read answers it with `Interrupted system call` before the Rflush; hangup and Tversion drop it and close the file as usual. Nothing wakes a connection by itself: the platform layer steps a connection when its transport moves, so a provider that becomes ready has to make that happen (an event stream's job, not the core's). Writes cannot park. Static memory: `Server(cfg).Storage` per connection, `Shared` once. No heap. The core has unit tests driven through `cloud9.Client` in memory, including a provider that parks. ## Value renderers (`vars.zig`) `expose(name, ptr: anytype)` builds at comptime a `VTable` for `@TypeOf(ptr.*)`: ``` /vars//value rendered text (structs: "field: value" lines, nested indented; unions: tag + payload; optionals: "null" or the value; enums: tag; ints/floats/bools; []const u8 and [*:0]const u8 as quoted strings (≤ 256 bytes); other pointers as 0x… never followed; arrays/slices ≤ 64 elements) /vars//type @typeName /vars//size @sizeOf /vars//addr 0x… /vars//raw the bytes (length = @sizeOf) /vars//f//... same layout recursively for struct fields (depth ≤ 4), leaves writable: writing text to a scalar's `value` parses and stores it (ints: decimal/0x, bools, floats, enums by tag) ``` Rendering is by a comptime-generated function table; no allocation. Writes to scalars are plain stores (not atomic; documented). ## Scratch provider (`scratch.zig`) The in-memory read/write tree from the current server, as a provider, with `init(allocator, budget_bytes)`; the only core-level component that takes an allocator, and it is optional. ## Linux layer (`linux/probe.zig`) ```zig pub const Probe = struct { pub const Options = struct { io: std.Io, // for std.debug symbolization listen: union(enum) { unix: []const u8, tcp: []const u8, fd: i32 }, max_clients: u8 = 8, msize: u32 = 64 * 1024, hold_on_panic: bool = true, capture_signal: u8 = SIGRTMIN + 3, // used to snapshot other threads breakpoints: bool = true, // install the SIGTRAP handler }; pub fn Storage(comptime max_clients: u8) type; // static: per-client Server.Storage + poll table pub fn init(p: *Probe, shared: *Server.Shared, storage: *Storage, opts: Options) !void; // listens, registers the debug provider pub fn start(p: *Probe) !void; // spawns ONE background thread running a poll loop over listener + clients pub fn stop(p: *Probe) void; // closes, joins }; ``` One thread, `poll()` over the listener and every connection; each connection is a core `Conn` fed with `push`/`step`/`output`. No per-connection threads. Symbolization uses `std.debug.getSelfDebugInfo()` with the `io` passed in and a caller-provided fixed buffer as the text arena. ## Debug provider (`linux/debug.zig`) Mounted as `/threads`, `/addr`, `/mem`, `/hex`, `/breakpoints`, `/panic`. ``` /threads/ one directory per tid, enumerated from /proc/self/task at list time /threads//name comm /threads//stat state letter + a few fields from /proc/self/task//stat /threads//stack "#n 0x in (::)" per frame /threads//regs " 0x" per general register, from the captured cpu context /addr/ dynamic dir: walk of any hex address yields a file "fn\nfile:line:col\nmodule\n" /mem/maps /proc/self/maps served by pread at the requested offset (any size) /mem/ raw bytes at address+offset via process_vm_readv/writev (never faults); writable /hex/ hexdump text of 256 bytes at address (+offset), like std.debug.dumpHex /breakpoints/ directory of tids currently stopped in @breakpoint() /breakpoints//stack, regs as above /breakpoints//ctl write "continue" (or "step"? no: continue only) to resume /panic/message the panic message, empty before any panic /panic/stack frames of the panicking thread /panic/ctl write "continue" to let the default panic handler run (abort) ``` **Capturing another thread** (`stack`, `regs`): the server thread `tgkill`s the target with `capture_signal`. The handler (SA_SIGINFO, async-signal-safe: no allocation, no locks) copies the `cpu_context.Native` obtained through `std.debug.cpu_context.fromPosixSignalContext` into a slot and futex-waits. The server thread unwinds with `std.debug.StackIterator.init(&ctx)` while the target is parked, symbolizes, then releases the slot; the target resumes. The server's own thread unwinds itself directly. Timeout 250 ms → Rerror `thread did not respond`. Threads blocked in uninterruptible syscalls simply time out. A target parked while holding std.debug's `SelfInfo` lock (it was printing a stack trace itself) cannot be unwound without deadlocking; the probe detects that with `tryLock`, releases the target and answers `i/o error` (registers still work). **Breakpoints**: `@breakpoint()` raises SIGTRAP on the executing thread only. The installed handler stores the context in a slot, marks the thread paused, and futex-waits until `/breakpoints//ctl` receives `continue`. On x86_64 the saved PC already points past `int3`; on aarch64 the handler advances PC by 4 (`brk`) before returning, but only for kernel-generated traps (`si_code > 0`); a user-sent `SIGTRAP` (`kill -TRAP`, `tgkill`) parks the thread exactly where it was, which makes it a usable "pause this thread" request. Other threads keep running; a slot table (`max_paused`, default 16) bounds simultaneous pauses and, when it is full, the trapping thread simply steps over the breakpoint (`traps_skipped` counts these). The probe's own serving thread is never parked or held: a trap or panic on it goes straight to the default behaviour, since nobody could write its `ctl` files. **Panics**: `pub const panic = proc9.linux.panic;` in the root module (built with `std.debug.FullPanic`). The first panic records message and a stack capture (`captureCurrentStackTrace` with `first_address`), publishes them, and, if `hold_on_panic` and the probe is running, futex-waits until `/panic/ctl` says `continue`; then `std.debug.defaultPanic` runs (prints the trace and aborts). A nested or second panic goes straight to the default. Signal handlers are installed by `Probe.init` (breakpoints optional) and restored by `stop`. ## Demo (`demo/main.zig`, binary `9proc-demo`) Keeps every path the existing tests read (`/build/*`, `/comptime/types/Qid/*`, `/comptime/decls`, `/runtime/fn/now|hostname|…`, `/runtime/ctl` with `add|echo|fib|sleep-ms`, `/runtime/pid|ppid|uptime|argv|cwd|env|clients`, `/scratch`), served by the library. Adds: * a worker thread running `workerLoop` that increments an exposed `State { ticks: u64, phase: enum, last_job: Job }` (`/vars/state/...`); * `/runtime/ctl` commands `trap` (the worker executes `@breakpoint()` on its next tick) and `panic` (the worker panics with a message); * `--stdio | --unix PATH | --tcp IP:PORT` as today, `--no-hold` to disable panic holding. `main` passes `init.io` and an explicit allocator to the pieces that need one; the demo's Storage is a global. ## Verification * Unit tests: core (in-memory client drives every op incl. providers, snapshots and a parked read), vars (render/set for every category), scratch, debug (capture own thread and a helper thread; breakpoint pause/continue on a helper thread; panic record path without holding). * `zig build 9proc-check-freestanding`: compiles `core.zig` + `vars.zig` for `riscv32-freestanding-none` with a tiny freestanding root that instantiates `Server(cfg)` with static Storage. * `zig build 9proc-test` (library and demo unit tests) and `zig build 9proc-debug-test` (linux/debug.zig). * 9ns's `test/integration.sh` unchanged and passing; `test/debug.sh` (`zig build 9proc-debug-itest`) through 9ns: read the worker's stack (contains `workerLoop` and `demo/main.zig:`), resolve a frame through `/addr`, dump `/hex` of the exposed state, read and write `/vars/state/f/ticks/value`, trap → `/breakpoints` lists the worker, its stack shows `workerLoop`, `continue` resumes (ticks keep increasing), panic → `/panic/message`, `/panic/stack`, `continue` → server exits non-zero. * Adversarial pass afterwards (`zig build 9proc-adv`: hostile client against the server and the core, signal races, memory reads of unmapped addresses, panic while a capture is in flight; `test/adversarial.sh` runs the same suites by hand). ## Known upstream issue (Zig 0.16 std.debug) `std.debug.SelfInfo` for ELF (`std/debug/SelfInfo/Elf.zig`, `findModule`) rebuilds its module list whenever it is asked about an address outside every known module. That frees each module's `Dwarf.Unwind` and CIE list but leaves `unwind_cache` entries pointing into the freed memory, so later unwinds read freed data: empty traces, "unwind info invalid", or segfaults once the arena reuses the block. `linux/debug.zig` records the executable's `PT_LOAD` ranges at init and refuses to hand std an address outside them (`knownCode`), which is why `/addr/` of a bogus address renders `?` instead of poisoning the process. Worth reporting upstream; the guard can go once std clears the cache.