1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
|
# 9proc: a 9P debug/introspection server as a library
The demo server that 9ns's tests use grows into a library any Zig program
can embed: a debugger-shaped interface where the protocol is just files.
Anything that can read a filesystem (a shell, an agent, an editor, `9p`,
9ns) can inspect a running process: build facts, comptime type layouts,
live values, threads and their stacks, memory, breakpoints, panics.
Design rules (non-negotiable, they mirror cloud9):
1. **The core is freestanding.** No allocator, no OS, no threads, no `std.Io`.
Caller-owned buffers, fixed-capacity tables sized at comptime. It must
compile for `riscv32-freestanding-none` (the ESP32-P4 firmware target,
`../05-zig-p4`), and `zig build 9proc-check-freestanding` proves it.
2. **Every dependency on a runtime is an explicit argument.** Features that
truly need an `Allocator` or an `std.Io` take them in their `init`; nothing
reaches for `std.heap.page_allocator` or a global `Io`. Where memory is
needed it is preferably a caller-provided `[]u8` or a comptime-sized
`Storage` struct the caller places in static memory.
3. **All allocation happens up front**, at init, from what the caller passed.
Steady-state operation does not allocate.
4. **Platform layers are separate modules** (`9proc.linux`) and are the
only places that touch sockets, threads, signals, `/proc` or `std.debug`.
```
9proc/src/root.zig pub const core, vars, scratch, linux (linux only), Server(cfg)
9proc/src/core.zig the tree (static nodes, vars, providers) as a backend of cloud9.fs.Server
9proc/src/vars.zig comptime value renderers (@typeInfo) for /vars
9proc/src/scratch.zig in-memory read/write tree provider (takes an Allocator)
9proc/src/linux/probe.zig background thread + poll loop + unix/tcp/fd listeners
9proc/src/linux/debug.zig threads, stacks, registers, addr→source, memory, breakpoints, panic
9proc/demo/main.zig the `9proc-demo` binary: embeds everything, worker thread, exposed vars
(paths from the cloud9 root; the library is the module `9proc` that
cloud9's `build.zig` exports next to `cloud9`, wired by `9proc/build.zig`)
```
## Core (`core.zig`)
```zig
pub const Config = struct {
name: []const u8 = "9proc", // /README and Stat uid/gid
types: []const type = &.{}, // /comptime/types/<short name>/...
decls_of: ?type = null, // /comptime/decls lists this type's pub decls
fns: type = struct {}, // /runtime/fn/<name>: pub fn (ctx: *anyopaque, w: *std.Io.Writer) anyerror!void
ctl: ?*const fn (ctx: *anyopaque, cmd: []const u8, out: *std.Io.Writer) anyerror!void = null, // /ctl
max_fids: u16 = 64,
max_providers: u8 = 8,
max_vars: u8 = 32,
/// Dynamic file contents are generated at open time into per-fid snapshot
/// slots so that reads at arbitrary offsets are consistent.
snapshot_slots: u8 = 8,
snapshot_bytes: u32 = 16 * 1024,
max_parked: u8 = 8, // reads a provider has parked with error.Again, per connection
};
pub fn Server(comptime cfg: Config) type {
return struct {
pub const Storage = struct { // caller places this in static memory
in: [msize]u8, out: [msize]u8, data: [msize]u8, snapshots: [cfg.snapshot_slots][cfg.snapshot_bytes]u8,
};
pub const Shared = struct { // state common to all connections (providers, vars)
pub fn init(name_ctx: *anyopaque) Shared;
pub fn addProvider(s: *Shared, p: Provider) error{Full}!void;
pub fn expose(s: *Shared, name: []const u8, ptr: anytype) error{Full}!void; // typed value → /vars/<name>
};
pub const Backend = struct { ... }; // what cloud9.fs.Server knows of the tree: Req, Reply, features
pub const Engine = cloud9.fs.Server(Backend, .{ .fid_capacity = cfg.max_fids, .slot_capacity = cfg.max_parked, ... });
pub const Conn = struct { // one 9P connection: an Engine plus the tree's per-connection state
pub fn init(shared: *Shared, storage: *Storage, msize: u32) Conn;
pub fn push(c: *Conn, bytes: []const u8) usize; // feed transport bytes
pub fn step(c: *Conn) error{Protocol}!bool; // serve ≤ 1 engine request; false = nothing to do
pub fn output(c: *const Conn) []const u8; // bytes to send
pub fn wrote(c: *Conn, n: usize) void;
pub fn hangup(c: *Conn) void; // drop fids, tell providers
pub fn fidCount(c: *const Conn) usize;
};
};
}
```
The core is a **backend of `cloud9.fs.Server`**, the library's file-server
engine (cloud9's `docs/design.md`, "File server engine"): one engine for
every 9P server built on cloud9. The engine owns the fid table (indexed,
`Options.fid_index`, so thousands of fids cost O(1) per lookup), walks,
permission checks, directory cursors, Tflush, and the releases a dropped
connection owes; `Conn` wraps one engine instance and answers its
requests (`lookup`, `getattr`, `setattr`, `open`, `read`, `write`,
`release`, `readdir`) from the tree: the static part (comptime-generated
from `cfg`: `/README`, `/build/*` via a `build_options`-like struct passed
in `cfg.build`, `/comptime/types/*`, `/comptime/decls`, `/runtime/fn/*`,
`/ctl`, `/vars/*`) is served directly, **providers** through their vtable.
The backend declares every optional engine feature (`fs.Features`:
create, remove, wstat, references), which is how Tcreate/Tremove/Twstat
and the one-clunk-per-handle contract below reach providers. `step`
answers the engine's `retry()` and `next()` requests synchronously, one
`next()` per call; a provider that cannot answer yet parks instead (see
"Answering later").
A provider is a runtime vtable mounted at a top-level name. It owns a subtree
with its own naming (dynamic directories such as `/threads/<tid>` or
`/addr/<hex>` cannot be enumerated at comptime):
```zig
pub const Provider = struct {
name: []const u8,
ctx: *anyopaque,
vtable: *const VTable,
pub const Handle = u64; // provider-defined node id; 0 = provider root
pub const VTable = struct {
walk: *const fn (ctx, parent: Handle, name: []const u8) Error!Handle,
stat: *const fn (ctx, h: Handle, out: *NodeStat) Error!void, // kind (dir/file), mode, length, mtime
list: *const fn (ctx, dir: Handle, index: usize, out: *NodeStat) Error!bool, // nth entry; false when done
open: *const fn (ctx, h: Handle, mode: u8) Error!void,
read: *const fn (ctx, h: Handle, offset: u64, buf: []u8) Error!usize,
write: *const fn (ctx, h: Handle, offset: u64, data: []const u8) Error!usize,
create: ?*const fn (ctx, dir: Handle, name: []const u8, perm: u32, mode: u8) Error!Handle,
remove: ?*const fn (ctx, h: Handle) Error!void,
wstat: ?*const fn (ctx, h: Handle, st: *const cloud9.Stat) Error!void,
clunk: *const fn (ctx, h: Handle) void, // fid released (also on hangup)
};
pub const Error = error{ NotFound, Exists, Perm, NotDir, IsDir, NotEmpty, BadOffset, NoSpace, Io, Unsupported };
};
```
Error → Rerror text: what the tree or a provider refuses is answered in
`ename`'s Plan 9 strings, which 9ns's bridge understands (`file does not
exist`, `permission denied`, `file already exists`, `directory not empty`,
`not a directory`, `is a directory`, `bad offset`, `no space`, `i/o error`,
`not supported`, `bad command`, `bad value`, `too many open dynamic
files`); what the engine refuses on its own carries cloud9.fs's strings,
the ones Linux v9fs maps back to errnos (`fid unknown or out of range`,
`fid already in use`, `Too many open files in system`, `bad use of fid`
for I/O on an unopened fid or a walk or clone from an open one, `file
already open for I/O`, `bad offset in directory read`, `permission denied`
for an open the walked mode forbids (a directory for writing, OEXEC, a
static file for writing), `wstat prohibited` for any of the fields the
engine owns (type, dev, qid, atime, the owner names) or a length on a
directory, `illegal name`, `Invalid argument` for an Rstat that cannot fit
the msize or a directory read whose count holds no whole record). The
smallest msize is the engine's `fs.msize_min` (217: one full Rwalk); a
Tversion below it ends the connection.
Directory reads follow the 9P rule (offset 0 or previous offset+count, never
split a record); the engine encodes the entries from the tree's records,
with uid/gid/muid the attach uname, the modes `dirent_dir_perm`/
`dirent_file_perm` and length 0 (a stat of the entry gives the real
ones). Dynamic file reads: on open the content is generated once into a
snapshot slot (`open` runs the generator; `read` serves the slot; a read
at offset 0 regenerates); no free slot → Rerror `too many open dynamic
files`. Stats of dynamic files report length 0.
Qids: static nodes get comptime paths; provider nodes get
`(provider index << 56) | handle` (or `NodeStat.path` in place of the
handle). The engine's node ids are the same numbers, except that
provider ids are offset by one in the top byte, since the engine reads
node 0 as "the node asked about".
**Answering later.** `Provider.read` may return `error.Again`: the engine
parks the request (`Config.max_parked` per connection; a further one fails
with EAGAIN), the connection goes on serving, and every later `Conn.step`
asks the provider again with the same handle and offset until it answers,
whether at once or on a later step. The fid stays open; a Tflush of a
parked read answers it with `Interrupted system call` before the Rflush;
hangup and Tversion drop it and close the file as usual. Nothing wakes a
connection by itself: the platform layer steps a connection when its
transport moves, so a provider that becomes ready has to make that happen
(an event stream's job, not the core's). Writes cannot park.
Static memory: `Server(cfg).Storage` per connection, `Shared` once. No heap.
The core has unit tests driven through `cloud9.Client` in memory,
including a provider that parks.
## Value renderers (`vars.zig`)
`expose(name, ptr: anytype)` builds at comptime a `VTable` for
`@TypeOf(ptr.*)`:
```
/vars/<name>/value rendered text (structs: "field: value" lines, nested indented; unions: tag + payload;
optionals: "null" or the value; enums: tag; ints/floats/bools; []const u8 and [*:0]const u8
as quoted strings (≤ 256 bytes); other pointers as 0x… never followed; arrays/slices ≤ 64 elements)
/vars/<name>/type @typeName
/vars/<name>/size @sizeOf
/vars/<name>/addr 0x…
/vars/<name>/raw the bytes (length = @sizeOf)
/vars/<name>/f/<field>/... same layout recursively for struct fields (depth ≤ 4), leaves writable:
writing text to a scalar's `value` parses and stores it (ints: decimal/0x, bools, floats, enums by tag)
```
Rendering is by a comptime-generated function table; no allocation.
Writes to scalars are plain stores (not atomic; documented).
## Scratch provider (`scratch.zig`)
The in-memory read/write tree from the current server, as a provider, with
`init(allocator, budget_bytes)`; the only core-level component that takes an
allocator, and it is optional.
## Linux layer (`linux/probe.zig`)
```zig
pub const Probe = struct {
pub const Options = struct {
io: std.Io, // for std.debug symbolization
listen: union(enum) { unix: []const u8, tcp: []const u8, fd: i32 },
max_clients: u8 = 8,
msize: u32 = 64 * 1024,
hold_on_panic: bool = true,
capture_signal: u8 = SIGRTMIN + 3, // used to snapshot other threads
breakpoints: bool = true, // install the SIGTRAP handler
};
pub fn Storage(comptime max_clients: u8) type; // static: per-client Server.Storage + poll table
pub fn init(p: *Probe, shared: *Server.Shared, storage: *Storage, opts: Options) !void; // listens, registers the debug provider
pub fn start(p: *Probe) !void; // spawns ONE background thread running a poll loop over listener + clients
pub fn stop(p: *Probe) void; // closes, joins
};
```
One thread, `poll()` over the listener and every connection; each connection
is a core `Conn` fed with `push`/`step`/`output`. No per-connection threads.
Symbolization uses `std.debug.getSelfDebugInfo()` with the `io` passed in and
a caller-provided fixed buffer as the text arena.
## Debug provider (`linux/debug.zig`)
Mounted as `/threads`, `/addr`, `/mem`, `/hex`, `/breakpoints`, `/panic`.
```
/threads/ one directory per tid, enumerated from /proc/self/task at list time
/threads/<tid>/name comm
/threads/<tid>/stat state letter + a few fields from /proc/self/task/<tid>/stat
/threads/<tid>/stack "#n 0x<addr> in <fn> (<file>:<line>:<col>)" per frame
/threads/<tid>/regs "<reg> 0x<value>" per general register, from the captured cpu context
/addr/<hex> dynamic dir: walk of any hex address yields a file "fn\nfile:line:col\nmodule\n"
/mem/maps /proc/self/maps served by pread at the requested offset (any size)
/mem/<hex> raw bytes at address+offset via process_vm_readv/writev (never faults); writable
/hex/<hex> hexdump text of 256 bytes at address (+offset), like std.debug.dumpHex
/breakpoints/ directory of tids currently stopped in @breakpoint()
/breakpoints/<tid>/stack, regs as above
/breakpoints/<tid>/ctl write "continue" (or "step"? no: continue only) to resume
/panic/message the panic message, empty before any panic
/panic/stack frames of the panicking thread
/panic/ctl write "continue" to let the default panic handler run (abort)
```
**Capturing another thread** (`stack`, `regs`): the server thread `tgkill`s
the target with `capture_signal`. The handler (SA_SIGINFO, async-signal-safe:
no allocation, no locks) copies the `cpu_context.Native` obtained through
`std.debug.cpu_context.fromPosixSignalContext` into a slot and futex-waits.
The server thread unwinds with `std.debug.StackIterator.init(&ctx)` while the
target is parked, symbolizes, then releases the slot; the target resumes. The
server's own thread unwinds itself directly. Timeout 250 ms → Rerror
`thread did not respond`. Threads blocked in uninterruptible syscalls simply
time out. A target parked while holding std.debug's `SelfInfo` lock (it was
printing a stack trace itself) cannot be unwound without deadlocking; the
probe detects that with `tryLock`, releases the target and answers
`i/o error` (registers still work).
**Breakpoints**: `@breakpoint()` raises SIGTRAP on the executing thread only.
The installed handler stores the context in a slot, marks the thread paused,
and futex-waits until `/breakpoints/<tid>/ctl` receives `continue`. On x86_64
the saved PC already points past `int3`; on aarch64 the handler advances PC by
4 (`brk`) before returning, but only for kernel-generated traps
(`si_code > 0`); a user-sent `SIGTRAP` (`kill -TRAP`, `tgkill`) parks the
thread exactly where it was, which makes it a usable "pause this thread"
request. Other threads keep running; a slot table (`max_paused`, default 16)
bounds simultaneous pauses and, when it is full, the trapping thread simply
steps over the breakpoint (`traps_skipped` counts these). The probe's own
serving thread is never parked or held: a trap or panic on it goes straight
to the default behaviour, since nobody could write its `ctl` files.
**Panics**: `pub const panic = proc9.linux.panic;` in the root module
(built with `std.debug.FullPanic`). The first panic records message and a
stack capture (`captureCurrentStackTrace` with `first_address`), publishes
them, and, if `hold_on_panic` and the probe is running, futex-waits until
`/panic/ctl` says `continue`; then `std.debug.defaultPanic` runs (prints the
trace and aborts). A nested or second panic goes straight to the default.
Signal handlers are installed by `Probe.init` (breakpoints optional) and
restored by `stop`.
## Demo (`demo/main.zig`, binary `9proc-demo`)
Keeps every path the existing tests read (`/build/*`, `/comptime/types/Qid/*`,
`/comptime/decls`, `/runtime/fn/now|hostname|…`, `/runtime/ctl` with
`add|echo|fib|sleep-ms`, `/runtime/pid|ppid|uptime|argv|cwd|env|clients`,
`/scratch`), served by the library. Adds:
* a worker thread running `workerLoop` that increments an exposed
`State { ticks: u64, phase: enum, last_job: Job }` (`/vars/state/...`);
* `/runtime/ctl` commands `trap` (the worker executes `@breakpoint()` on its
next tick) and `panic` (the worker panics with a message);
* `--stdio | --unix PATH | --tcp IP:PORT` as today, `--no-hold` to disable
panic holding.
`main` passes `init.io` and an explicit allocator to the pieces that need one;
the demo's Storage is a global.
## Verification
* Unit tests: core (in-memory client drives every op incl. providers,
snapshots and a parked read), vars (render/set for every category), scratch, debug (capture
own thread and a helper thread; breakpoint pause/continue on a helper
thread; panic record path without holding).
* `zig build 9proc-check-freestanding`: compiles `core.zig` + `vars.zig`
for `riscv32-freestanding-none` with a tiny freestanding root that
instantiates `Server(cfg)` with static Storage.
* `zig build 9proc-test` (library and demo unit tests) and
`zig build 9proc-debug-test` (linux/debug.zig).
* 9ns's `test/integration.sh` unchanged and passing; `test/debug.sh`
(`zig build 9proc-debug-itest`) through 9ns: read the worker's
stack (contains `workerLoop` and `demo/main.zig:`),
resolve a frame through `/addr`, dump `/hex` of the exposed state, read and
write `/vars/state/f/ticks/value`, trap → `/breakpoints` lists the worker,
its stack shows `workerLoop`, `continue` resumes (ticks keep increasing),
panic → `/panic/message`, `/panic/stack`, `continue` → server exits
non-zero.
* Adversarial pass afterwards (`zig build 9proc-adv`: hostile client
against the server and the core, signal races, memory reads of unmapped
addresses, panic while a capture is in flight; `test/adversarial.sh` runs
the same suites by hand).
## Known upstream issue (Zig 0.16 std.debug)
`std.debug.SelfInfo` for ELF (`std/debug/SelfInfo/Elf.zig`, `findModule`)
rebuilds its module list whenever it is asked about an address outside every
known module. That frees each module's `Dwarf.Unwind` and CIE list but leaves
`unwind_cache` entries pointing into the freed memory, so later unwinds read
freed data: empty traces, "unwind info invalid", or segfaults once the arena
reuses the block. `linux/debug.zig` records the executable's `PT_LOAD` ranges
at init and refuses to hand std an address outside them (`knownCode`), which
is why `/addr/<hex>` of a bogus address renders `?` instead of poisoning the
process. Worth reporting upstream; the guard can go once std clears the cache.
|