From 147ebd4a36ec7199074ba05bcfb79d4a656c0b74 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Thu, 27 Aug 2026 16:42:15 -0300 Subject: 9p: the client half, and a board that serves its own tree over the UART Step 5 of the 9P chain (docs/9p.typ 12.5, docs/registry.typ 9P-22, 9P-11, BOARD-1). THE CLIENT. `Client` in src/9p.zig is the mirror of `Server` and the same shape: sans-io, no allocator, no threads, no descriptor, caller-owned buffers, and it builds freestanding. 152 bytes of struct against the server's 9,488, because a client owns neither a fid table nor a park table -- the far end does. The API is submit / push+output+wrote / take. Completion is a PULL: a callback would fire inside push, inside the transport's read, inside the host's poll dispatch, which is exactly where fs9_service says filesystem work must not happen. `take()` returns the next completed operation or null, which is `Server.next()`'s loop-until-null contract read from the other side. Tags are a fixed 16-entry table indexed BY the tag, so an out-of-order reply -- which 9P allows and both reference clients rely on -- costs one bounds check. The reply's TYPE is checked against the request's op, because a tag is only as good as the table behind it. A `Done` borrows the input buffer and is valid until the next call; `take()` releases the previous frame on entry, so the rule is mechanical rather than remembered, and read data and error strings are zero-copy. And one real caller, so this is not a library with no user: the `9p` word takes a dial and a path, walks another instance's tree, and opens the bytes in a pane like any other `Look`. THE BOARD. A SECOND image, not a second role: the console runtime keeps UART0 bidirectionally and is behaviourally untouched. On the new one the UART carries 9P AND NOTHING ELSE -- no ANSI, no vaxis, no allocator, no heap module. The loop is uart.read -> push / retry+next -> handle -> reply / output -> writeSome -> wrote. `writeSome` is new and additive: `write`'s bounded spin DROPS bytes on a stalled transmitter, which on a protocol stream truncates a reply mid-message and desynchronises for good, where a short count cannot. BOARD-1's one divider write raises the line to 921600. 88,000 B text, 49,424 B bss, an 88,080-byte image -- 5.7% of the 1,536,000 B partition, against the console image's 809,536 B. THE COMPTIME BRIDGE, which is the part worth reading. `board9p.caps` is the ONLY place the GPIO tree is described; node ids, parents, names, permissions, handlers, buffer size and the per-pin directories are all derived from it, and `fan.dirs` makes `gpio//value` one table entry serving eleven pins. Modes are derived from which handlers a file has rather than declared. A second capability is a table entry, not new tree code. JP1 became a real table in the new leaf `src/board_pins.zig`, with the ASCII drawing RENDERED from it at comptime and the pin list COLLECTED from it -- the 9P image links no core and so cannot import board_memory.zig, and copying the table was not acceptable. A golden test pins the drawing byte for byte, the console's own shape test still passes, and the identical bytes are present in all three artifacts. PROVED. Two daemons: B read A's `/1/body` through the `9p` word into a pane, byte-identical to plan9port's `9p read` of the same path. Both board images build. No hardware was attached, so nothing about the board is claimed beyond what builds and what the host tests cover. zig build unit-test 585/585. fs-bench unchanged and still zero allocations on every read row. --- REVIEW FIXES FOLDED IN. Steps 3, 4 and 5 were verified on the happy path and then adversarially reviewed by three agents; eight defects, six fixed here, five of them reproduced with measurements before and after. Full writeup in docs/registry.typ `9P-27`. In brief: * a remote crash of the WHOLE daemon: one `size[4]` of zero plus one byte hit `unreachable` in `fs9_service.fill`. Also 99.7% of a core when the stuck buffer made `room == 0` return without reading. Now `srv.dead` is a hangup, checked before the room guard. * the editor froze 177 s on a dial: `connect(2)` ran on a still-BLOCKING socket before the deadline existed, and a full accept backlog waits forever. Now non-blocking with the wait spent against the budget. After: 2.03 s. * a 64 KiB pty read is exactly `queue_cap` and wiped every unread byte AND dropped itself. `notePtyOutput` splits at half the cap. Deterministic. * four silent sockets denied `--fs9` forever; connections now expire on the same five-second rule the frontend transport already had. * EMFILE spun a core; the listener pauses and leaves the poll set, as the frontend listener does. * `max_fids = 32` made `find` over `9pfuse` fail with 57 consecutive `Rerror`s -- refuting this step's own acceptance clause. 256 for a host, `board_fids` 32 for the microcontroller. Found clean and worth recording: `sig` reaches the foreground process group; the two-namespace pty lookup is right over both transports; `PaneFile`'s u4 wall is guarded; reader counts release on every abrupt-death path; `fs_origin` routing and the reply arithmetic hold under probing. --- src/9p.zig | 1237 ++++++++++++++++++++++++++++++++++++++++++++++- src/acmefs.zig | 23 +- src/board9p.zig | 874 +++++++++++++++++++++++++++++++++ src/board_memory.zig | 46 +- src/board_pins.zig | 186 +++++++ src/builtins.zig | 40 ++ src/config.zig | 8 + src/detached/server.zig | 17 +- src/esp32p4/uart.zig | 50 +- src/esp32p4_9p.zig | 454 +++++++++++++++++ src/fs9_client.zig | 695 ++++++++++++++++++++++++++ src/fs9_service.zig | 132 ++++- src/main.zig | 7 + 13 files changed, 3715 insertions(+), 54 deletions(-) create mode 100644 src/board9p.zig create mode 100644 src/board_pins.zig create mode 100644 src/esp32p4_9p.zig create mode 100644 src/fs9_client.zig (limited to 'src') diff --git a/src/9p.zig b/src/9p.zig index 8f78ed6b..068d0740 100644 --- a/src/9p.zig +++ b/src/9p.zig @@ -1834,12 +1834,31 @@ pub const orclose: u8 = 64; // --------------------------------------------------------------------------- /// Fids one connection may hold at once. A FIXED ARRAY and not a map, costed -/// in `docs/registry.typ` `9P-11`: a linear scan of thirty-two is about 0.4 µs -/// and the wire is slower than that by orders of magnitude, so the map would -/// buy nothing and cost an allocator this file does not have. +/// in `docs/registry.typ` `9P-11`: a linear scan is far cheaper than the wire, +/// so the map would buy nothing and cost an allocator this file does not have. +/// +/// 256 AND NOT 32, which is what it was, and the difference is a MOUNT. A +/// script that opens one file at a time never needs more than a handful; a +/// mounting client keeps one fid per cached inode, and this tree is three +/// top-level entries plus fourteen files per pane, so seven panes already pass +/// thirty-two and a full sixteen-pane session wants over two hundred. At the +/// old number `find` over a `9pfuse` mount failed with fifty-seven consecutive +/// `Rerror`s once the table filled — and `9pfuse` is the proof clause +/// `docs/9p.typ` §12.4 sets for this step, so the number was refuting its own +/// acceptance test. +/// +/// A `Fid` is about 64 bytes, so this is ≈16 KiB per connection against the +/// ≈34 KiB `fs9_service.zig` already budgets for one. The BOARD keeps thirty-two +/// by passing its own value: see `board_fids`, and `9P-11`'s RAM line, which is +/// costed for a microcontroller serving its own small tree and nothing else. /// /// Overflow is a refusal (`e_too_many_fids`), not a queue. -pub const max_fids: usize = 32; +pub const max_fids: usize = 256; + +/// What a microcontroller uses instead. Named here rather than spelled at the +/// call site so that the two numbers, and the reason they differ, stay next to +/// each other. +pub const board_fids: usize = 32; /// Requests that may be outstanding at once — in flight, or parked because the /// core answered `.again`. In practice this counts BLOCKED READERS: one slot @@ -4339,12 +4358,1212 @@ test "9p server: the fid table and the park table are what the board was costed // not, plus the open handle and the two-coordinate directory cursor. The // numbers are asserted here so that a change to `Fid` shows up as a diff // in the board's budget rather than as a surprise on the board. + // + // TWO budgets, because there are now two numbers. `max_fids` is the desktop + // one and it is sized for a MOUNT, which keeps a fid per cached inode; + // `board_fids` is what a microcontroller serving its own small tree uses, + // and it is the one `9P-11` costed. const S = Server(StubFs); - const fids = @sizeOf(S.Fid) * max_fids; + const entry = @sizeOf(S.Fid); const slots = @sizeOf(S.Slot) * max_slots; - try testing.expect(fids <= 3 * 1024); try testing.expect(slots <= 8 * 1024); - // Two buffers at a 4,096-byte msize — `in` and `out`'s two — plus the two - // tables, against 336 KB of free heap on the P4. - try testing.expect(3 * 4096 + fids + slots <= 24 * 1024); + + // The board: two buffers at a 4,096-byte msize — `in` and `out`'s two — + // plus the two tables, against 336 KB of free heap on the P4. + const board = entry * board_fids; + try testing.expect(board <= 3 * 1024); + try testing.expect(3 * 4096 + board + slots <= 24 * 1024); + + // The desktop, against the ≈34 KiB per connection `fs9_service` budgets. + // Eight times the fids is ≈16 KiB, and it is the price of `find` working + // over a `9pfuse` mount — see `max_fids`. + try testing.expect(entry * max_fids <= 24 * 1024); +} + +// --------------------------------------------------------------------------- +// the client +// --------------------------------------------------------------------------- + +/// Tags one client may have outstanding at once. +/// +/// SIXTEEN, and the reasoning is `max_fids`': a fixed array with no allocator, +/// scanned rather than mapped, and a refusal rather than a queue when it fills. +/// The protocol's tag space is 0..0xFFFE — `notag` is 0xFFFF and belongs to the +/// handshake — so this uses the bottom sixteen of sixty-five thousand and never +/// a number above them. THE TAG IS ITS OWN INDEX, which is what makes matching +/// a reply O(1) with no search and no bookkeeping: see `tags`. +/// +/// MANY OUTSTANDING REQUESTS ARE LEGAL and the prior art imposes no bound at +/// all: Linux's client takes a tag per request out of an IDR +/// (`net/9p/client.c:194-199`) and Plan 9's devmnt keeps an `Mntrpc` per +/// request on a free list (`devmnt.c:783-800`), because on both the outstanding +/// count is «one per process blocked in an I/O», which the kernel already +/// bounds elsewhere. Here the count is "one per thing pardes is fetching", and +/// a screen does not hold sixteen remote panes. Overflow is `error.NoTags` at +/// the moment of asking — the caller collects an answer and asks again — and +/// never a silent wait, because a client that blocks is the one thing this +/// design does not have anywhere to put. +pub const max_tags: usize = 16; + +/// `Twrite`'s own header: `size[4] type[1] tag[2] fid[4] offset[8] count[4]`. +/// What `maxWrite` subtracts from the msize. +/// +/// NOT `iohdrsz`. That number (24) is the slack a SERVER quotes in `iounit` and +/// is deliberately larger than any real header; a client sizing its own request +/// against it leaves a byte on the table on every write forever. +const twrite_header: usize = header_len + 4 + 8 + 4; + +/// `Rread`'s header: `size[4] type[1] tag[2] count[4]`. What `maxRead` +/// subtracts, and the reason a client's `count` is not simply the msize — the +/// reply has to carry a header too, and a `count` of msize is a reply eleven +/// bytes too long for the connection that asked for it. +const rread_header: usize = header_len + 4; + +comptime { + assert(twrite_header == 23); + assert(rread_header == 11); + // The tag IS the index, so the table's length is the tag space in use, and + // the handshake's `notag` must fall outside it or a `Rversion` would land + // on somebody's slot. + assert(max_tags <= notag); + // At the smallest msize this file will agree to, a read and a write must + // both still be able to carry a byte, or a connection could be negotiated + // that cannot do any I/O at all. + assert(msize_min > rread_header); + assert(msize_min > twrite_header); +} + +/// Every way `submit` can refuse, and each is a different thing for the caller +/// to do about it. +pub const ClientError = error{ + /// All `max_tags` are outstanding. Collect an answer and ask again. + NoTags, + /// The out queue has no room for this request. Write `output()` out and + /// ask again. THE ONLY BACK-PRESSURE a sans-io client has. + NoSpace, + /// The request cannot fit the negotiated msize: a `Twrite` past + /// `maxWrite()`, a `Tread` asking past `maxRead()`, a sixteen-element walk + /// of long names. REFUSED AND NOT CLAMPED, because `submit` answers with a + /// tag and nothing else — a silent clamp would leave the caller to guess + /// how much of its buffer went, and guess wrong about the offset to + /// continue from. `maxRead` and `maxWrite` are how a caller chunks first. + TooLarge, + /// A request before `Rversion` has landed, or a second `Tversion` while + /// requests are outstanding. + Handshake, + /// The stream is not 9P any more and this connection is finished. See + /// `dead`. + Dead, + /// A request no encoding of 9P admits: `nofid` as a fid, an empty walk + /// element or one with a separator in it, more than `max_welem` elements. + /// A bug in the caller, caught here rather than spent as a round trip. + BadRequest, +}; + +/// A 9P2000 client for one connection. +/// +/// THE MIRROR OF `Server`, and deliberately the same shape: caller-owned `in` +/// and `out` buffers, no allocator, no threads, no descriptor, `std` for +/// `readInt`/`writeInt` and nothing else. So it compiles for the board's +/// `riscv32-freestanding` and runs over `src/esp32p4/uart.zig`'s non-blocking +/// receive and bounded-spin transmit exactly as it runs over a unix socket in +/// `src/fs9_client.zig` — which is the whole reason for the shape, because a +/// client that owned its descriptor would be a client that could not. +/// +/// THE API IS A STATE MACHINE AND NOT `fn read() []u8`, because the core is +/// single-threaded and never blocks (docs/9p.typ §12.5). The three moving parts +/// are: +/// +/// 1. `submit(Request)` ENCODES a T-message into the out queue and hands +/// back its tag. It never waits and never touches a descriptor; a full +/// queue or a full tag table is a refusal the caller can act on. +/// 2. `push`/`output`/`wrote` move bytes, in whatever sizes the transport +/// manages, in whatever order they arrive. +/// 3. `take()` answers with the next COMPLETED operation, or null when there +/// is not a whole reply buffered yet. A caller's frame is +/// `while (client.take()) |done| ...`, which is `Server.next()`'s own +/// loop-until-null contract read from the other side. +/// +/// WHY COMPLETION IS A PULL AND NOT A CALLBACK: a callback would run inside +/// `push`, which is inside the transport's read, which is inside the host's +/// poll dispatch — and `src/fs9_service.zig` already states why filesystem +/// work must not happen there. Pulling puts the caller's own code back on the +/// caller's own stack. +/// +/// WHY THERE IS NO PER-TAG RESULT QUEUE: one frame completes exactly one +/// operation, and `take` returns it immediately, so there is never a completed +/// answer nobody has collected. That is what keeps a tag slot two bytes wide +/// instead of an msize wide, and it is why `Done` may borrow `in` (see there). +/// +/// REPLIES MAY ARRIVE IN ANY ORDER and this client does not care: the tag is +/// its own index into `tags`, so attribution is one bounds check and one +/// array read, with no assumption about arrival order anywhere in the file. +/// The reply's TYPE is checked against the request's `Op` as well, because a +/// tag is only as good as the table behind it. +/// +/// MEMORY: the two buffers, and `@sizeOf(Client)` for everything else — a +/// sixteen-entry tag table of eight-byte entries plus nine scalars, asserted at +/// the bottom of this file. Nothing here grows and nothing here is allocated. +pub const Client = struct { + /// Reply bytes the caller has pushed. `in[0..frame]` is the reply most + /// recently returned by `take`, and every slice a `Done` holds points into + /// it — which is why nothing compacts this buffer until the next `take`. + in: []u8, + /// Encoded requests, oldest first, as a byte FIFO. Every 9P message + /// carries its own length, so the queue needs no side table. + out: []u8, + + in_len: usize = 0, + frame: u32 = 0, + out_len: usize = 0, + out_off: usize = 0, + + /// Negotiated by the handshake; ZERO means "not on a protocol yet", and + /// nothing but `version` may be submitted in that state. It is also zero + /// after an `Rversion` of "unknown", which is a completed handshake with + /// no dialect in common. + msize: u32 = 0, + /// What our own `Tversion` offered, kept only so that `Rversion` can be + /// checked against it: «the server responds with its own maximum, which + /// must be less than or equal to the client's». + asked: u32 = 0, + /// A `Tversion` is outstanding. Its tag is `notag`, so it cannot live in + /// the table below — and it does not need to, because the protocol allows + /// nothing else to be outstanding beside it. + versioning: bool = false, + /// The stream is not 9P and there is no resynchronising from it. Write-once, + /// like `Server.dead`: a reply that cannot be attributed is worse than a + /// closed connection, because the caller would wait on it forever. + dead: bool = false, + + /// THE TAG TABLE, indexed BY THE TAG. `tags[t].op` is null when tag `t` is + /// free, which makes claiming a tag a scan of sixteen and matching a reply + /// a single index — and it means a caller may keep its own per-request + /// state in a plain sixteen-entry array of its own, keyed the same way, + /// with no map on either side. + tags: [max_tags]Slot = @splat(.{}), + + /// What is remembered about one outstanding request, which is as little as + /// the protocol lets us get away with: what it was, and — for a read — + /// what it asked for, because `read(5)` bounds the reply by it and a + /// server that ignores that bound is handing back bytes at offsets we + /// never asked about. + const Slot = struct { + op: ?Op = null, + count: u32 = 0, + }; + + /// What a client asked for. The tag names of `Request` and of `Result`'s + /// answers are these, so nothing maps one to the other by hand. + /// + /// EIGHT OPERATIONS AND NOT THIRTEEN, and the five absences are decisions: + /// + /// * `Tauth`: there is no authentication in this design and the server half + /// refuses it by name (`e_no_auth`). The socket's permissions are the + /// protection. + /// * `Tcreate`/`Tremove`: the server refuses both, because the shape of the + /// tree follows the pane list. Walking into `new/` is how a client creates + /// a pane, and that is a `walk`. + /// * `Twstat`: the one wstat the tree honours is a truncate, and a client + /// that wants to empty a file opens it `OTRUNC` in the same round trip. + /// * `Tflush`: nothing here has a cancel button. A flush costs a second tag + /// and brings a reply-ORDER rule with it — «the Rflush must come after the + /// original reply» — which is a rule nobody exercises if no caller can + /// change its mind, and an unexercised ordering rule in a protocol client + /// is a bug waiting for its first user. + pub const Op = enum { version, attach, walk, open, read, write, clunk, stat }; + + /// One request, as its caller states it. A `union(Op)` rather than eight + /// functions so that `submit` is one entry point with one refusal path: every + /// bound this client has — the tag table, the out queue, the msize — applies to + /// all eight identically, and a ninth operation cannot forget one of them. + /// + /// NO TAG FIELD: the tag is what `submit` HANDS BACK. A caller that chose its + /// own tags would be maintaining the table this file already maintains. + pub const Request = union(Op) { + /// The handshake. `msize` is the largest message this client will send or + /// accept, and ZERO means "as much as my buffers hold", which is the + /// answer a caller with no opinion wants. Clamped to the buffers either + /// way; see `beginVersion`. + version: struct { msize: u32 = 0 }, + /// `afid` is not a parameter: it is always `nofid`, because this client + /// never sends `Tauth`. + attach: struct { fid: u32, uname: []const u8, aname: []const u8 = "" }, + /// The path elements, already split. A SLICE OF SLICES rather than the + /// codec's fixed `[max_welem]` array, because a caller has a path and not + /// an array: the copy into the fixed array happens once, in `submit`, + /// where the `nwname` bound is checked anyway. An empty list is the legal + /// zero-element walk, which clones `fid` onto `newfid`. + walk: struct { fid: u32, newfid: u32, names: []const []const u8 }, + open: struct { fid: u32, mode: u8 }, + /// `count` is refused rather than clamped above `maxRead()`; see there. + read: struct { fid: u32, offset: u64, count: u32 }, + /// `data` is COPIED into the out queue by `submit` and is not borrowed + /// afterwards, which is what lets a caller write out of a buffer it is + /// about to reuse. + write: struct { fid: u32, offset: u64, data: []const u8 }, + clunk: struct { fid: u32 }, + stat: struct { fid: u32 }, + }; + + /// What one request came to. The answer's SHAPE, which is what a caller acts + /// on; `Done.op` says which request it belongs to and `Done.tag` says which + /// one of several. + pub const Result = union(enum) { + /// The server said no: `Rerror`'s string, and the only variant that can + /// answer ANY of the eight. Borrows the input buffer — see `Done`. + fail: []const u8, + /// `version` is "9P2000", or the literal "unknown", which is a SUCCESSFUL + /// reply meaning no dialect in common. `Client.msize` is nonzero only in + /// the first case, so the second leaves a connection on which nothing can + /// be submitted and the caller hangs up. + version: struct { msize: u32, version: []const u8 }, + attach: Qid, + /// `nwqid` may be SHORTER than the walk's element count: a partial walk is + /// a success with fewer qids, and only a failure on the FIRST element is + /// an `Rerror`. So a caller MUST compare `nwqid` against what it asked for + /// before believing its fid landed anywhere. + /// + /// The whole array is carried rather than only the last qid, because the + /// last one is the only thing THIS tree's clients want and the + /// intermediate ones are what a caching client caches against + /// (`Qid.version`). Two hundred and eight bytes, on a value the caller + /// consumes and drops. + walk: struct { nwqid: u16, wqid: [max_welem]Qid }, + open: struct { qid: Qid, iounit: u32 }, + /// The bytes, borrowing the input buffer — see `Done`. SHORTER than the + /// requested count is normal and is not the end of the file; ZERO bytes is + /// the end of the file. + read: []const u8, + /// The count actually written, which may be short — the caller advances + /// its offset by this and not by what it asked. + write: u32, + clunk: void, + /// Borrows the input buffer for its four strings — see `Done`. + stat: Stat, + }; + + /// One completed operation. + /// + /// BORROWS THE INPUT BUFFER, and this is the whole lifetime rule: a `Done` is + /// valid until the next call to anything on the `Client` that produced it. The + /// `fail` string, the `read` bytes and the `stat` strings all point into + /// `Client.in`, exactly as `decode`'s do and for the same reason — the + /// alternative is a per-tag copy of every payload, which on the board is + /// sixteen msizes of static RAM to save a caller one `@memcpy` it may not even + /// want. `take` releases the previous answer's frame on entry, so the rule is + /// enforced by construction rather than by hope: a caller that keeps a `Done` + /// across a second `take` is reading bytes the next reply has been decoded + /// into. + pub const Done = struct { + /// The tag `submit` handed out, or `notag` for the handshake. FREE again + /// the moment this is returned, so a caller that indexes its own + /// sixteen-entry table by tag must read this entry out before submitting + /// anything else. + tag: u16, + /// Which of the eight this answers. Needed beside `result` because + /// `Rerror` answers all of them and carries no hint of which. + op: Op, + result: Result, + }; + + pub const Options = struct { + /// Room for one whole reply. Caps the msize with `out`. + in: []u8, + /// Room for one whole request, at least. MORE room is what buys + /// pipelining: sixteen outstanding `Tread`s are sixteen small messages + /// that all have to fit here at once, and `submit` answers + /// `error.NoSpace` rather than blocking when they do not. + out: []u8, + }; + + /// The buffers are the caller's, which is what "no allocator" means from + /// this side: the board hands over two static arrays, a host hands over + /// two heap slices, and this file cannot tell the difference. The msize + /// follows from them and from the server's `Rversion`. + pub fn init(opts: Options) Client { + assert(opts.in.len >= msize_min); + assert(opts.out.len >= msize_min); + return .{ .in = opts.in, .out = opts.out }; + } + + /// The connection went away, or the caller is done with it. Unlike + /// `Server.hangup` there is no debt to pay: a client owes the far end + /// nothing on the way out — its fids are the server's to clean up when the + /// stream closes, which is exactly what `Server.hangup` is for. + pub fn hangup(c: *Client) void { + c.dead = true; + c.tags = @splat(.{}); + c.versioning = false; + c.msize = 0; + c.asked = 0; + c.in_len = 0; + c.frame = 0; + c.out_len = 0; + c.out_off = 0; + } + + // -- bytes in, bytes out --------------------------------------------- + // + // The four `Server` has, written out again rather than shared. They look + // identical and they are not the same three lines: `Server.push` refuses + // once dead, and `Server.hasRoom` reserves a whole msize before a request + // is handed to the core so that no reply can fail to be written. A client + // reserves nothing — it refuses at `submit`, where the caller is standing + // right there — so a shared FIFO would be one struct with two callers and + // two exceptions, which is more to read than this is. + + /// Take as much of `bytes` as there is room for, and answer how much. A + /// short answer is not a loss: it is back-pressure, and the caller + /// re-offers the tail after `take`ing what it can. Bytes are APPENDED, so + /// the reply currently being borrowed by a `Done` does not move. + pub fn push(c: *Client, bytes: []const u8) usize { + if (c.dead) return 0; + const n = @min(bytes.len, c.in.len - c.in_len); + @memcpy(c.in[c.in_len..][0..n], bytes[0..n]); + c.in_len += n; + return n; + } + + /// The requests waiting to go, oldest first, as one contiguous run of + /// whole 9P messages. Valid until the next call to anything else here. + pub fn output(c: *const Client) []const u8 { + return c.out[c.out_off..c.out_len]; + } + + /// How many of `output()`'s bytes actually left. A partial write is normal + /// on a UART and on a full socket, and the remainder stays put. + pub fn wrote(c: *Client, n: usize) void { + assert(n <= c.out_len - c.out_off); + c.out_off += n; + if (c.out_off == c.out_len) { + c.out_off = 0; + c.out_len = 0; + } + } + + /// Slide the unwritten tail down. Called only when room is wanted, so the + /// common case — a fully written queue, reset to empty by `wrote` — never + /// moves a byte. + fn compact(c: *Client) void { + assert(c.out_off <= c.out_len); + const n = c.out_len - c.out_off; + std.mem.copyForwards(u8, c.out[0..n], c.out[c.out_off..c.out_len]); + c.out_off = 0; + c.out_len = n; + } + + /// Release the reply `take` last returned and slide the rest of the input + /// down. One message-long move per message; a ring buffer would let a + /// decoded reply straddle the wrap and stop being one slice. + fn dropFrame(c: *Client) void { + assert(c.frame != 0); + assert(c.frame <= c.in_len); + const n = c.frame; + std.mem.copyForwards(u8, c.in[0 .. c.in_len - n], c.in[n..c.in_len]); + c.in_len -= n; + c.frame = 0; + } + + // -- what a caller may ask for --------------------------------------- + + /// The largest `Tread.count` this connection can answer, which is the + /// msize less `Rread`'s own header. Zero before the handshake. + pub fn maxRead(c: *const Client) u32 { + if (c.msize == 0) return 0; + return c.msize - @as(u32, @intCast(rread_header)); + } + + /// The most bytes one `Twrite` can carry, which is the msize less + /// `Twrite`'s own header. Zero before the handshake. A caller with more + /// than this chunks; see `ClientError.TooLarge` for why it is not clamped. + pub fn maxWrite(c: *const Client) u32 { + if (c.msize == 0) return 0; + return c.msize - @as(u32, @intCast(twrite_header)); + } + + /// Requests outstanding, the handshake included. What a caller's loop + /// tests to know whether there is anything left to wait for. + pub fn pending(c: *const Client) usize { + var n: usize = @intFromBool(c.versioning); + for (c.tags) |t| n += @intFromBool(t.op != null); + return n; + } + + // -- asking ------------------------------------------------------------ + + /// Encode one request into the out queue and hand back its tag. Never + /// blocks, never waits, never touches a descriptor. + /// + /// The refusals are in one order on purpose: what is wrong with the + /// REQUEST first, then what is wrong with this client's tables, so a + /// caller's bad argument never costs a tag and never half-fills the queue. + pub fn submit(c: *Client, req: Request) ClientError!u16 { + if (c.dead) return error.Dead; + if (req == .version) return c.beginVersion(req.version.msize); + // «The client must communicate the version before any other messages» + // — and until `Rversion` has landed there is no msize to bound + // anything by, which is the same gate `Server.startFrame` applies from + // the other side. + if (c.msize == 0 or c.versioning) return error.Handshake; + + const msg: Msg = switch (req) { + .version => unreachable, // handled above + .attach => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + break :blk .{ .tattach = .{ + .fid = m.fid, + .afid = nofid, + .uname = m.uname, + .aname = m.aname, + } }; + }, + .walk => |m| blk: { + if (m.fid == nofid or m.newfid == nofid) return error.BadRequest; + // `MAXWELEM` is a hard protocol bound and not a buffer size: + // every implementation refuses a seventeen-element walk, so a + // caller with a deeper path splits it into two walks. + if (m.names.len > max_welem) return error.BadRequest; + var w: [max_welem][]const u8 = @splat(""); + for (m.names, 0..) |n, i| { + // An empty element, or one with a separator in it, is a + // caller that has not split its path. `Twalk` has no + // encoding for either and a server answers the first one + // `illegal name` — a round trip spent on a bug that was + // visible from here. + if (n.len == 0) return error.BadRequest; + if (std.mem.indexOfAny(u8, n, "/\x00") != null) return error.BadRequest; + w[i] = n; + } + break :blk .{ .twalk = .{ + .fid = m.fid, + .newfid = m.newfid, + .nwname = @intCast(m.names.len), + .wname = w, + } }; + }, + .open => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + break :blk .{ .topen = .{ .fid = m.fid, .mode = m.mode } }; + }, + .read => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + // The one bound the request's own length does not express: + // what comes BACK has to fit the connection too. + if (m.count > c.maxRead()) return error.TooLarge; + break :blk .{ .tread = .{ .fid = m.fid, .offset = m.offset, .count = m.count } }; + }, + .write => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + break :blk .{ .twrite = .{ .fid = m.fid, .offset = m.offset, .data = m.data } }; + }, + .clunk => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + break :blk .{ .tclunk = .{ .fid = m.fid } }; + }, + .stat => |m| blk: { + if (m.fid == nofid) return error.BadRequest; + break :blk .{ .tstat = .{ .fid = m.fid } }; + }, + }; + // ONE ceiling for every request, which is what makes a `Twrite` and a + // sixteen-element `Twalk` obey the same rule: the msize is «the + // maximum length, in bytes, ... including the size field», and a + // client that sends more is a client the server closes on. + const need = totalLen(msg) catch return error.TooLarge; + if (need > c.msize) return error.TooLarge; + + const op = std.meta.activeTag(req); + const tag = c.claim(op) orelse return error.NoTags; + errdefer c.tags[tag] = .{}; + try c.emit(tag, msg); + if (op == .read) c.tags[tag].count = req.read.count; + return tag; + } + + /// `Tversion`, which is the one exchange with no tag and no msize behind + /// it. + /// + /// A SECOND ONE IS A CONNECTION RESET — «all fids are clunked and any + /// outstanding I/O is abandoned» (`version(5)`) — and abandoning somebody + /// else's request is not this function's decision to make. So it is + /// refused while anything is outstanding, and a caller that means to reset + /// collects its answers or hangs up first. + fn beginVersion(c: *Client, want: u32) ClientError!u16 { + if (c.pending() != 0) return error.Handshake; + // Two ceilings and the smaller wins: what one reply buffer holds, and + // what one request buffer holds. A caller with no opinion passes zero + // and gets both. + const cap: u32 = @intCast(@min(c.in.len, c.out.len, std.math.maxInt(u32))); + const m = @min(if (want == 0) cap else want, cap); + // Below the floor there is a connection that cannot carry an `Rwalk`, + // which is to say no connection at all. `init` asserts the buffers + // clear it, so this can only be a `want` the caller chose. + if (m < msize_min) return error.BadRequest; + try c.emit(notag, .{ .tversion = .{ .msize = m, .version = "9P2000" } }); + c.msize = 0; + c.asked = m; + c.versioning = true; + return notag; + } + + /// The lowest free tag, marked used. Lowest rather than round-robin so + /// that a client with one request outstanding always uses tag 0, which + /// makes a wire trace readable by eye. + fn claim(c: *Client, op: Op) ?u16 { + for (&c.tags, 0..) |*t, i| { + if (t.op != null) continue; + t.* = .{ .op = op }; + return @intCast(i); + } + return null; + } + + /// Queue one request. The only way this fails is room: `totalLen` has + /// already refused every other way `encode` can, which is why the error + /// set collapses to one value here. + fn emit(c: *Client, tag: u16, msg: Msg) ClientError!void { + if (c.out_off != 0) c.compact(); + const bytes = encode(msg, tag, c.out[c.out_len..]) catch return error.NoSpace; + c.out_len += bytes.len; + } + + // -- collecting -------------------------------------------------------- + + /// The next completed operation, or null when there is not a whole reply + /// buffered yet. Call in a loop until null, once per frame. + /// + /// A `Done` BORROWS the input buffer and is valid until the next call + /// here: the previous reply's frame is released on entry, which is what + /// makes that rule mechanical instead of a note somebody has to remember. + pub fn take(c: *Client) ?Done { + if (c.frame != 0) c.dropFrame(); + if (c.dead) return null; + const len = frameLen(c.in[0..c.in_len]) orelse return null; + // A `size` no encoder produced, or one this connection could never + // buffer: either way the stream is not 9P and waiting for the rest of + // it is waiting forever. + if (len < header_len or len > c.in.len) return c.die(); + // And a server that sends past the msize it agreed to has stopped + // speaking the protocol it agreed to. + if (c.msize != 0 and len > c.msize) return c.die(); + if (len > c.in_len) return null; + c.frame = len; + // A body this codec refuses is not a message we can attribute to a + // tag, so there is nobody to report it to. The connection ends. + const got = decode(c.in[0..len]) catch return c.die(); + return c.consume(got); + } + + /// The stream is finished. Returns null so that every refusal in `take` + /// and `consume` is one expression. + fn die(c: *Client) ?Done { + c.dead = true; + return null; + } + + /// One decoded reply onto the request it answers. + fn consume(c: *Client, got: Decoded) ?Done { + // A CLIENT READS R-MESSAGES. A T-message here is the other end of the + // connection talking, or the double-role link docs/9p.typ §7 tells us + // not to build — the exact mirror of `Server.startFrame`'s refusal, + // and the encoding's own parity does the work in both directions. + if (isT(got.msg.msgType())) return c.die(); + if (got.msg == .rversion) return c.version(got); + // Nothing may arrive before a `Tversion` has been answered, and + // nothing but the `Rversion` while one is outstanding. + if (c.versioning or c.msize == 0) return c.die(); + // The tag is its own index, so this bounds check IS the lookup. + if (got.tag >= max_tags) return c.die(); + const slot = &c.tags[got.tag]; + const op = slot.op orelse return c.die(); + const result: Result = switch (got.msg) { + // `Rerror` answers ANY of the eight, which is exactly why `Done` + // reports the op beside it: the string does not say what failed. + .rerror => |m| .{ .fail = m.ename }, + .rattach => |m| if (op != .attach) return c.die() else .{ .attach = m.qid }, + .rwalk => |m| if (op != .walk) return c.die() else .{ + .walk = .{ .nwqid = m.nwqid, .wqid = m.wqid }, + }, + .ropen => |m| if (op != .open) return c.die() else .{ + .open = .{ .qid = m.qid, .iounit = m.iounit }, + }, + .rread => |m| blk: { + if (op != .read) return c.die(); + // «count ... indicates the number of bytes returned», and + // read(5) makes it no more than what was asked. Linux calls a + // longer one a hard `-EIO` (`net/9p/client.c:1475-1479`); here + // it ends the connection, because the byte after the ones we + // asked for is a byte we have no offset to put anywhere. + if (m.data.len > slot.count) return c.die(); + break :blk .{ .read = m.data }; + }, + .rwrite => |m| if (op != .write) return c.die() else .{ .write = m.count }, + .rclunk => if (op != .clunk) return c.die() else .clunk, + .rstat => |m| if (op != .stat) return c.die() else .{ .stat = m.stat }, + // The rest are replies to requests this client does not send — + // `Rauth`, `Rcreate`, `Rremove`, `Rwstat`, `Rflush` — and one is + // not an answer at all (`Rerror`'s illegal twin `Terror`, already + // refused by parity above). A reply to a request nobody made means + // the tag space is not what we think it is. + else => return c.die(), + }; + slot.* = .{}; + return .{ .tag = got.tag, .op = op, .result = result }; + } + + /// `Rversion`: the msize handshake, from the client's side. + fn version(c: *Client, got: Decoded) ?Done { + if (!c.versioning) return c.die(); + // «Rversion ... carries the same tag», and that tag is `notag`, + // because tags do not mean anything yet. + if (got.tag != notag) return c.die(); + const m = got.msg.rversion; + // «The server responds with its own maximum, which must be less than + // or equal to the client's» — `version(5)`. A larger one is a message + // we cannot buffer, and Linux refuses it for that reason + // (`net/9p/client.c:840-843`). + if (m.msize > c.asked or m.msize < msize_min) return c.die(); + c.versioning = false; + if (std.mem.eql(u8, m.version, "9P2000")) { + c.msize = m.msize; + } else if (!std.mem.eql(u8, m.version, "unknown")) { + // The reply must be a version the client offered, or "unknown". + // Anything else — "9P2000.u", "9P2000.L", a typo — is a server + // answering a question we did not ask, and agreeing to a dialect + // this file does not implement is how a client sends a `Tattach` + // whose layout the other end reads differently. + return c.die(); + } + // "unknown" leaves `msize` at zero: a completed handshake with no + // dialect in common, on which nothing can be submitted. The CALLER + // decides whether that is worth hanging up over, which is the honest + // place for it — a fallback ladder of dialects is a policy and this is + // a codec. + return .{ .tag = notag, .op = .version, .result = .{ + .version = .{ .msize = m.msize, .version = m.version }, + } }; + } +}; + +// --------------------------------------------------------------------------- +// client tests +// --------------------------------------------------------------------------- +// +// Driven against the SERVER IN THIS FILE, in process, over two pairs of +// buffers. That is the strongest test available here and it needs no socket: +// every byte the client encodes is a byte the server decodes and vice versa, +// so a disagreement about a layout, a length or a tag fails a test rather than +// waiting for a live daemon. The stub filesystem is the server tests' own, so +// the tree the client walks is the tree those tests already pin down. + +/// One connection with a client at each end of it. Four buffers, because each +/// side owns its own two and neither may see the other's. +/// +/// `srv.out` is twice the msize because `Server` requires it; `cli.out` is not, +/// because a client reserves nothing — see `Client.Options`. +const Pair = struct { + srv_in: [4096]u8 = undefined, + srv_out: [8192]u8 = undefined, + cli_in: [4096]u8 = undefined, + cli_out: [4096]u8 = undefined, + fsys: StubFs = .{}, + srv: Srv = undefined, + cli: Client = undefined, + + /// The buffers are fields, so neither end can be built until the pair has + /// an address. + fn start(p: *Pair) void { + p.srv = Srv.init(.{ .in = &p.srv_in, .out = &p.srv_out, .root = 1 }); + p.cli = Client.init(.{ .in = &p.cli_in, .out = &p.cli_out }); + } + + fn answer(p: *Pair, req: StubFs.Req) void { + const a = p.fsys.handle(req); + p.srv.reply(&a.reply, a.bytes); + } + + /// THE WIRE: every byte both ways, and each side given every chance to + /// work, until nothing moves. A real transport does this a chunk at a time + /// in a poll loop; the tests that care about that drip bytes by hand. + fn wire(p: *Pair) void { + var moved = true; + while (moved) { + moved = false; + while (p.cli.output().len != 0) { + const n = p.srv.push(p.cli.output()); + if (n == 0) break; + p.cli.wrote(n); + moved = true; + } + while (p.srv.retry()) |req| { + p.answer(req); + moved = true; + } + while (p.srv.next()) |req| { + p.answer(req); + moved = true; + } + while (p.srv.output().len != 0) { + const n = p.cli.push(p.srv.output()); + if (n == 0) break; + p.srv.wrote(n); + moved = true; + } + } + } + + /// Submit one request, run the wire, and collect the one answer it + /// produced. The tag and the op are checked here so that no test below has + /// to repeat it. + fn one(p: *Pair, req: Client.Request) !Client.Done { + const tag = try p.cli.submit(req); + p.wire(); + const done = p.cli.take() orelse return error.NoReply; + try testing.expectEqual(tag, done.tag); + try testing.expectEqual(std.meta.activeTag(req), done.op); + // One request, one reply, and nothing left outstanding: the invariant + // that makes `pending()` usable as a loop condition. + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + return done; + } + + /// `Tversion` and `Tattach`, leaving the root on fid 0 — the client-side + /// twin of `Harness.handshake`. + fn handshake(p: *Pair) !void { + p.start(); + const v = try p.one(.{ .version = .{} }); + try testing.expectEqualStrings("9P2000", v.result.version.version); + try testing.expectEqual(@as(u16, notag), v.tag); + const a = try p.one(.{ .attach = .{ .fid = 0, .uname = "goblin" } }); + try testing.expectEqual(@as(u64, 1), a.result.attach.path); + try testing.expectEqual(qtdir, a.result.attach.type); + } +}; + +test "9p client: a whole session against the server in this file" { + var p: Pair = .{}; + try p.handshake(); + // Both ends agreed the same number, and it came off the buffers rather + // than out of the air. + try testing.expectEqual(@as(u32, 4096), p.cli.msize); + try testing.expectEqual(@as(u32, 4096 - 11), p.cli.maxRead()); + try testing.expectEqual(@as(u32, 4096 - 23), p.cli.maxWrite()); + + // A two-element walk onto pane 1's `body`. `nwqid` equals what was asked, + // which is the only thing that says the fid landed where we wanted. + const w = try p.one(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{ "1", "body" } } }); + try testing.expectEqual(@as(u16, 2), w.result.walk.nwqid); + try testing.expectEqual(@as(u64, 16), w.result.walk.wqid[0].path); + try testing.expectEqual(@as(u64, 18), w.result.walk.wqid[1].path); + try testing.expectEqual(qtfile, w.result.walk.wqid[1].type); + + const o = try p.one(.{ .open = .{ .fid = 1, .mode = ordwr } }); + try testing.expectEqual(@as(u64, 18), o.result.open.qid.path); + try testing.expectEqual(@as(u32, 4096 - iohdrsz), o.result.open.iounit); + + const r = try p.one(.{ .read = .{ .fid = 1, .offset = 0, .count = 64 } }); + try testing.expectEqualStrings("hello, body\n", r.result.read); + // Past the end is zero bytes and not an error: 9P has no EOF flag, and a + // short read is how a client learns it is done. + const eof = try p.one(.{ .read = .{ .fid = 1, .offset = 12, .count = 64 } }); + try testing.expectEqual(@as(usize, 0), eof.result.read.len); + + const wr = try p.one(.{ .write = .{ .fid = 1, .offset = 0, .data = "abc" } }); + try testing.expectEqual(@as(u32, 3), wr.result.write); + try testing.expectEqualStrings("abc", p.fsys.writes[0..p.fsys.writes_len]); + + const st = try p.one(.{ .stat = .{ .fid = 1 } }); + try testing.expectEqualStrings("body", st.result.stat.name); + try testing.expectEqual(@as(u64, 12), st.result.stat.length); + try testing.expectEqualStrings("goblin", st.result.stat.uid); + + _ = try p.one(.{ .clunk = .{ .fid = 1 } }); + // The clunk paid the core its release, which is the half of a clunk a + // client cannot see and the server tests pin down from the other side. + try testing.expectEqual(@as(u32, 1), p.fsys.releases); + // Nothing outstanding, nothing buffered, nothing owed. + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + try testing.expectEqual(@as(usize, 0), p.cli.output().len); + try testing.expect(p.cli.take() == null); + try testing.expect(!p.cli.dead); +} + +/// Move the server's queued replies to the client LAST FIRST. 9P permits it — +/// nothing in the protocol orders replies against each other — and both +/// reference clients allocate a tag per outstanding request with no in-order +/// assumption anywhere (`linux/net/9p/client.c:194-199`, +/// `plan9/devmnt.c:783-800`). A client that quietly relies on order works +/// until the day the server answers a cached stat before a blocked read. +fn deliverReversed(p: *Pair) !void { + var scratch: [4096]u8 = undefined; + const out = p.srv.output(); + try testing.expect(out.len <= scratch.len); + @memcpy(scratch[0..out.len], out); + const total = out.len; + p.srv.wrote(total); + + var at: [max_tags]usize = undefined; + var lens: [max_tags]u32 = undefined; + var count: usize = 0; + var i: usize = 0; + while (i < total) { + const len = frameLen(scratch[i..total]) orelse return error.ShortReply; + at[count] = i; + lens[count] = len; + count += 1; + i += len; + } + try testing.expect(count >= 2); + var k = count; + while (k > 0) { + k -= 1; + const f = scratch[at[k]..][0..lens[k]]; + try testing.expectEqual(f.len, p.cli.push(f)); + } +} + +test "9p client: replies out of order are matched by tag and not by arrival" { + var p: Pair = .{}; + try p.handshake(); + const w = try p.one(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{"index"} } }); + try testing.expectEqual(@as(u16, 1), w.result.walk.nwqid); + + // Two stats outstanding at once, on two different files. + const root_tag = try p.cli.submit(.{ .stat = .{ .fid = 0 } }); + const index_tag = try p.cli.submit(.{ .stat = .{ .fid = 1 } }); + try testing.expectEqual(@as(u16, 0), root_tag); + try testing.expectEqual(@as(u16, 1), index_tag); + try testing.expectEqual(@as(usize, 2), p.cli.pending()); + + // Both requests to the server, both replies produced, then handed back in + // the wrong order. + while (p.cli.output().len != 0) { + const n = p.srv.push(p.cli.output()); + p.cli.wrote(n); + } + while (p.srv.next()) |req| p.answer(req); + try deliverReversed(&p); + + // The SECOND request answers first, and it is recognised by its tag. + const first = p.cli.take() orelse return error.NoReply; + try testing.expectEqual(index_tag, first.tag); + try testing.expectEqualStrings("index", first.result.stat.name); + const second = p.cli.take() orelse return error.NoReply; + try testing.expectEqual(root_tag, second.tag); + try testing.expectEqualStrings("/", second.result.stat.name); + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + try testing.expect(!p.cli.dead); +} + +test "9p client: an Rerror answers one operation and the session carries on" { + var p: Pair = .{}; + try p.handshake(); + + // A walk failing on its FIRST element is an `Rerror` rather than a short + // `Rwalk` — the one asymmetry in walk(5), and the reason `Done` reports + // the op beside the string. + const bad = try p.one(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{"nope"} } }); + try testing.expectEqual(Client.Op.walk, bad.op); + try testing.expectEqualStrings(errString(2), bad.result.fail); + + // The failed tag is free again and the connection is untouched: an error + // is an answer, not a fault. + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + try testing.expect(!p.cli.dead); + const st = try p.one(.{ .stat = .{ .fid = 0 } }); + try testing.expectEqualStrings("/", st.result.stat.name); + + // A refusal that comes from the server's own table rather than the core's, + // spelled the way Linux's error table holds it. + const stale = try p.one(.{ .stat = .{ .fid = 9 } }); + try testing.expectEqualStrings(e_unknown_fid, stale.result.fail); + try testing.expect(!p.cli.dead); +} + +test "9p client: a reply arriving a byte at a time is taken when its last byte lands" { + var p: Pair = .{}; + try p.handshake(); + const tag = try p.cli.submit(.{ .stat = .{ .fid = 0 } }); + + // The request out, the reply produced, and then held on this side of the + // wire so it can be dripped in. + while (p.cli.output().len != 0) { + const n = p.srv.push(p.cli.output()); + p.cli.wrote(n); + } + while (p.srv.next()) |req| p.answer(req); + var scratch: [512]u8 = undefined; + const out = p.srv.output(); + try testing.expect(out.len > 4 and out.len <= scratch.len); + @memcpy(scratch[0..out.len], out); + const reply = scratch[0..out.len]; + p.srv.wrote(reply.len); + + // Every byte but the last leaves nothing to collect — including the first + // four, where `frameLen` becomes readable and still says "wait". + for (reply[0 .. reply.len - 1]) |b| { + try testing.expectEqual(@as(usize, 1), p.cli.push(&.{b})); + try testing.expect(p.cli.take() == null); + try testing.expect(!p.cli.dead); + } + try testing.expectEqual(@as(usize, 1), p.cli.push(reply[reply.len - 1 ..])); + const done = p.cli.take() orelse return error.NoReply; + try testing.expectEqual(tag, done.tag); + try testing.expectEqualStrings("/", done.result.stat.name); +} + +test "9p client: sixteen tags outstanding, and the seventeenth is refused" { + var p: Pair = .{}; + try p.handshake(); + + // Nothing is wired, so nothing is answered and every tag stays out. + var tags: [max_tags]u16 = undefined; + for (&tags, 0..) |*t, i| { + t.* = try p.cli.submit(.{ .stat = .{ .fid = 0 } }); + // Lowest free tag first, which is what makes a trace readable. + try testing.expectEqual(@as(u16, @intCast(i)), t.*); + } + try testing.expectEqual(max_tags, p.cli.pending()); + try testing.expectError(error.NoTags, p.cli.submit(.{ .stat = .{ .fid = 0 } })); + // A refused submit costs nothing: no tag, and not a byte in the queue. + const owed = p.cli.output().len; + try testing.expectError(error.NoTags, p.cli.submit(.{ .clunk = .{ .fid = 0 } })); + try testing.expectEqual(owed, p.cli.output().len); + + // Drained, every tag comes back, and the seventeenth request now fits. + p.wire(); + var seen: [max_tags]bool = @splat(false); + for (0..max_tags) |_| { + const done = p.cli.take() orelse return error.NoReply; + try testing.expectEqual(Client.Op.stat, done.op); + try testing.expect(!seen[done.tag]); + seen[done.tag] = true; + } + for (seen) |s| try testing.expect(s); + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + _ = try p.one(.{ .stat = .{ .fid = 0 } }); +} + +test "9p client: what a caller may not ask for is refused before a tag is spent" { + var p: Pair = .{}; + p.start(); + + // Nothing before the handshake, and `Tversion` is the only exception. + try testing.expectError(error.Handshake, p.cli.submit(.{ .stat = .{ .fid = 0 } })); + try p.handshake(); + + // `NOFID` is not a fid a client may name. + try testing.expectError(error.BadRequest, p.cli.submit(.{ .stat = .{ .fid = nofid } })); + try testing.expectError(error.BadRequest, p.cli.submit(.{ .clunk = .{ .fid = nofid } })); + try testing.expectError(error.BadRequest, p.cli.submit(.{ .walk = .{ .fid = 0, .newfid = nofid, .names = &.{} } })); + + // A path that has not been split, and one longer than the protocol admits. + try testing.expectError(error.BadRequest, p.cli.submit(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{"1/body"} } })); + try testing.expectError(error.BadRequest, p.cli.submit(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{""} } })); + const seventeen: [max_welem + 1][]const u8 = @splat("x"); + try testing.expectError(error.BadRequest, p.cli.submit(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &seventeen } })); + + // Both I/O bounds, each one byte past what the msize can carry. + try testing.expectError(error.TooLarge, p.cli.submit(.{ .read = .{ + .fid = 0, + .offset = 0, + .count = p.cli.maxRead() + 1, + } })); + var big: [4096]u8 = @splat('x'); + try testing.expectError(error.TooLarge, p.cli.submit(.{ .write = .{ + .fid = 0, + .offset = 0, + .data = big[0 .. p.cli.maxWrite() + 1], + } })); + // And exactly at the bound, both fit — a cap that is off by one is a cap + // that costs a round trip on every large transfer. Each one on an EMPTY + // queue, which is what `Options.out` means by "room for one whole + // request": a maximum-size `Twrite` IS the msize, so it fits beside + // nothing at all. + p.cli.wrote(p.cli.output().len); + _ = try p.cli.submit(.{ .read = .{ .fid = 0, .offset = 0, .count = p.cli.maxRead() } }); + p.cli.wrote(p.cli.output().len); + _ = try p.cli.submit(.{ .write = .{ .fid = 0, .offset = 0, .data = big[0..p.cli.maxWrite()] } }); + + // A second `Tversion` resets the connection, so it is refused while + // anything is outstanding rather than abandoning it. + try testing.expectError(error.Handshake, p.cli.submit(.{ .version = .{} })); + // And with the queue full of that one write there is nowhere to put even + // an eleven-byte `Tstat`: the out queue is the only back-pressure a + // sans-io client has, and it lands at `submit` where the caller is + // standing right there. + try testing.expectError(error.NoSpace, p.cli.submit(.{ .stat = .{ .fid = 0 } })); +} + +test "9p client: an msize below the floor, and one the server tried to raise" { + var in: [512]u8 = undefined; + var out: [512]u8 = undefined; + var buf: [64]u8 = undefined; + + // A caller asking for less than an `Rwalk` is asking for a connection that + // cannot be served. + var c = Client.init(.{ .in = &in, .out = &out }); + try testing.expectError(error.BadRequest, c.submit(.{ .version = .{ .msize = msize_min - 1 } })); + // Zero means "whatever the buffers hold", which is the smaller of the two. + _ = try c.submit(.{ .version = .{} }); + try testing.expectEqual(@as(u32, 512), c.asked); + + // A server answering with MORE than the client offered is a server whose + // next message will not fit the buffer that has to hold it. + _ = c.push(try encode(.{ .rversion = .{ .msize = 1024, .version = "9P2000" } }, notag, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + + // "unknown" is a SUCCESSFUL reply with no dialect in common: the handshake + // completes, `msize` stays zero, and nothing more can be submitted. + var c2 = Client.init(.{ .in = &in, .out = &out }); + _ = try c2.submit(.{ .version = .{} }); + _ = c2.push(try encode(.{ .rversion = .{ .msize = 512, .version = "unknown" } }, notag, &buf)); + const done = c2.take() orelse return error.NoReply; + try testing.expectEqualStrings("unknown", done.result.version.version); + try testing.expect(!c2.dead); + try testing.expectEqual(@as(u32, 0), c2.msize); + try testing.expectError(error.Handshake, c2.submit(.{ .stat = .{ .fid = 0 } })); + + // A dialect we never offered is neither: agreeing to it would be agreeing + // to a layout this file does not implement. + var c3 = Client.init(.{ .in = &in, .out = &out }); + _ = try c3.submit(.{ .version = .{} }); + _ = c3.push(try encode(.{ .rversion = .{ .msize = 512, .version = "9P2000.u" } }, notag, &buf)); + try testing.expect(c3.take() == null); + try testing.expect(c3.dead); +} + +test "9p client: what is not an answer to one of our requests ends the connection" { + var buf: [64]u8 = undefined; + // Each case gets a fresh connection past the handshake, because every one + // of them is fatal by design. + const Case = struct { + fn armed(in: []u8, out: []u8, scratch: []u8) !Client { + var c = Client.init(.{ .in = in, .out = out }); + _ = try c.submit(.{ .version = .{} }); + c.wrote(c.output().len); + _ = c.push(try encode(.{ .rversion = .{ .msize = 512, .version = "9P2000" } }, notag, scratch)); + _ = c.take() orelse return error.NoReply; + _ = try c.submit(.{ .stat = .{ .fid = 0 } }); + c.wrote(c.output().len); + return c; + } + }; + var in: [512]u8 = undefined; + var out: [512]u8 = undefined; + + // A T-message. A client reads R-messages, and the parity says so with no + // table: this is `Server.startFrame`'s refusal read from the other end. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(try encode(.{ .tstat = .{ .fid = 0 } }, 0, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // A reply on a tag nobody claimed. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(try encode(.rclunk, 3, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // A tag outside the table entirely, which no reply to us can carry. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(try encode(.rclunk, 900, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // The right tag and the WRONG SHAPE: an `Rclunk` where an `Rstat` was + // asked for. A tag is only as good as the table behind it. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(try encode(.rclunk, 0, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // A reply to a request this client never sends. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(try encode(.rwstat, 0, &buf)); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // A `size` no encoder produced, and one past the negotiated msize. Both + // are streams that will never resynchronise. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(&.{ 3, 0, 0, 0 }); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(&.{ 0, 4, 0, 0 }); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } + // And a body the codec refuses: the type byte is fine, the payload is not. + { + var c = try Case.armed(&in, &out, &buf); + _ = c.push(&.{ 8, 0, 0, 0, @intFromEnum(Type.rstat), 0, 0, 0 }); + try testing.expect(c.take() == null); + try testing.expect(c.dead); + } +} + +test "9p client: an Rread longer than the Tread asked for is refused" { + // The one bound a client cannot check from the frame alone, which is why + // `Slot` keeps the count: a server handing back more than was asked has + // given us bytes at offsets we never named. Linux calls it `-EIO`. + var p: Pair = .{}; + try p.handshake(); + _ = try p.one(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{ "1", "body" } } }); + _ = try p.one(.{ .open = .{ .fid = 1, .mode = oread } }); + + const tag = try p.cli.submit(.{ .read = .{ .fid = 1, .offset = 0, .count = 4 } }); + p.cli.wrote(p.cli.output().len); + var buf: [64]u8 = undefined; + _ = p.cli.push(try encode(.{ .rread = .{ .data = "hello, body\n" } }, tag, &buf)); + try testing.expect(p.cli.take() == null); + try testing.expect(p.cli.dead); + + // Exactly the count asked for is fine, and so is anything shorter. + var q: Pair = .{}; + try q.handshake(); + _ = try q.one(.{ .walk = .{ .fid = 0, .newfid = 1, .names = &.{ "1", "body" } } }); + _ = try q.one(.{ .open = .{ .fid = 1, .mode = oread } }); + const short = try q.one(.{ .read = .{ .fid = 1, .offset = 0, .count = 4 } }); + try testing.expectEqualStrings("hell", short.result.read); +} + +test "9p client: hangup and a dead connection refuse everything after" { + var p: Pair = .{}; + try p.handshake(); + p.cli.hangup(); + try testing.expectEqual(@as(usize, 0), p.cli.pending()); + try testing.expectEqual(@as(usize, 0), p.cli.output().len); + try testing.expectEqual(@as(usize, 0), p.cli.push("anything")); + try testing.expect(p.cli.take() == null); + try testing.expectError(error.Dead, p.cli.submit(.{ .stat = .{ .fid = 0 } })); + try testing.expectError(error.Dead, p.cli.submit(.{ .version = .{} })); +} + +test "9p client: one session is a hundred and change bytes plus its buffers" { + // The number the board is costed against, and the whole reason the client + // is shaped the way it is: sixteen eight-byte tag slots and nine scalars, + // with every payload borrowed out of the input buffer rather than copied + // into a per-tag one. A `Server` on the same connection is 9,488 B because + // it owns a fid table and a park table; a client owns neither, because the + // far end does. + try testing.expect(@sizeOf(Client.Slot) <= 8); + try testing.expect(@sizeOf(Client) <= 256); + // Two buffers at the 8,192-byte msize `src/fs9_service.zig` serves, plus + // the client itself: what one `9p` word costs while it is running. + try testing.expect(2 * 8192 + @sizeOf(Client) <= 17 * 1024); + // And at the protocol floor, which is what a board would negotiate: two + // buffers of 217 bytes each is a 9P client in under 700 bytes of RAM. + try testing.expect(2 * msize_min + @sizeOf(Client) <= 700); } diff --git a/src/acmefs.zig b/src/acmefs.zig index 21beb7f6..dc2780fe 100644 --- a/src/acmefs.zig +++ b/src/acmefs.zig @@ -690,7 +690,28 @@ pub fn notePtyOutput(p: *Pardes, id: usize, bytes: []const u8) void { if (id >= MAX_PANES or bytes.len == 0) return; const pf = &p.fs.panes[id]; if (pf.pty_readers == 0) return; - pf.pty_out.push(p.gpa, bytes); + // SPLIT, because a record larger than `queue_cap - 4` can never be + // admitted: `Queue.push`'s eviction loop pops until `peek()` is null — + // destroying every unread byte the script was still owed — and then drops + // the new record too, silently. + // + // That is not a theoretical size. Every host reads a pty master with a + // 64 KiB buffer (`pty_chunk` in detached/server.zig, `[0x10000]u8` in + // tty.zig, gui.zig and macos.zig) and a single read really does return + // 65536 on Linux — measured. So a pane running a build or a `cat` of + // anything large produces exactly the record that empties the queue, + // repeatedly, for as long as a script holds `pty/data` open. + // + // The `event` queue never met this because its records are a few dozen + // bytes; `pty/data` inherited the cap without inheriting that property. + // Half the cap per record, so a full queue is at least two records and the + // eviction loop always has something to evict. + var off: usize = 0; + while (off < bytes.len) { + const n = @min(bytes.len - off, queue_cap / 2); + pf.pty_out.push(p.gpa, bytes[off..][0..n]); + off += n; + } } // ============================================================================ diff --git a/src/board9p.zig b/src/board9p.zig new file mode 100644 index 00000000..b5b18bf2 --- /dev/null +++ b/src/board9p.zig @@ -0,0 +1,874 @@ +//! THE BOARD AS A FILESYSTEM, generated from a comptime table of what the board can do. +//! +//! The ESP32-P4 already exposes its pads and its address space — by TYPING A WORD into a tag. +//! `Gpio 20` flips a pin and prints `GPIO 20: 0->1`, `Gpio` alone draws JP1, `Peek`, `Poke` and +//! `Hexdump` reach all 2³² addresses (`src/board_memory.zig`), and every one of the four caps at +//! 4,096 bytes because the answer has to fit down a 115200-baud console. Nothing about that is +//! machine-readable and nothing about it is remote: the answer lands in an output pane, for a person +//! to read (`docs/registry.typ` `9P-11`, review note). +//! +//! This file is that same capability WITH NAMES INSTEAD OF VERBS. `cat gpio/pinout` is `Gpio`; +//! `echo 1 > gpio/20/value` is `Gpio 20`, except that it says which level it wants instead of asking +//! for whichever one it is not. A shell pipeline can do it, a script on a laptop can do it over the +//! UART, and neither needs a terminal emulator or a pane. +//! +//! WHY A TABLE, which is the whole design and not a flourish. A hand-written tree is a `Node` +//! packing, a `lookup`, a `getattr`, a `readdir`, a `read` and a `write` — six places that have to +//! agree about what exists — and the cost of adding `uptime` to it is an edit to all six plus a new +//! node id nobody else is using. The board's capabilities are a LIST, they will grow, and the entry +//! that describes one should be the only place it is described. So `caps` below is the tree: the +//! directories, the files, the per-pin fan-out, the permissions, the handlers and even the size of +//! the answer buffer are all derived from it at comptime, and the six functions at the bottom read +//! the derived table and know nothing about GPIO at all. +//! +//! WHAT A SECOND CAPABILITY COSTS, entry by entry, because "extensible" is a claim and this is the +//! evidence for it. Not implemented here — none of them is needed to serve a pin — but each is one +//! `Cap` and its handlers, and NO tree code: +//! +//! * `mem/` — `peek` and `poke` over `board_memory.readWord`/`writeWord` +//! (`src/board_memory.zig:136-144`), which are four lines of `*allowzero volatile` and already +//! compile for this target. `poke` is a WRITE handler that parses ` `, so +//! it needs `Fault.Malformed` and nothing else; `peek` needs an address to read, which a +//! stateless file cannot carry, so it is either a write-then-read pair (`echo 4ff40000 > addr; +//! cat word`, one more file and one `u32` of state) or a fan over a comptime list of interesting +//! registers. The second is free: `fan` below already generates a directory per key. +//! * `hexdump` — the same, with `scratch = 4096`: the one field that makes the shared answer +//! buffer grow, and the reason that field is in the table rather than a constant at the top. +//! * `prof` — three cycle counts from `pardes_esp32p4_frame_prof`, which the editor object already +//! exports (`src/esp32p4.zig:985`). It is the one capability that is NOT available in this +//! image: that symbol lives in the pardes object and the 9P image links none, so serving it +//! would mean either linking the editor or moving the counters. Worth saying out loud rather +//! than listing it as cheap. +//! * `uptime` and `heap` — `hal.systimer` and the heap's own free count, both of which the +//! runtime (`src/esp32p4_9p.zig`) can reach today. Two read handlers, `scratch = 24`, one +//! `Cap` each. These are the cheapest of the four and the reason the table's `board` parameter +//! is a TYPE rather than a pair of function pointers: adding `board.uptimeMs()` to the seam +//! adds a capability without changing anything here but the table. +//! +//! THE ABI IS `acmefs`'s, VERBATIM — `Op`, `Status`, `Req`, `Reply`, `Reply.Attr` with the same +//! fields and the same meanings — so `src/9p.zig`'s `Server` serves this tree with no translation +//! layer, exactly as it serves the editor's. That is the point of `Server` being a generic over the +//! filesystem rather than an importer of one (`src/9p.zig:1994-2010`), and it is what makes a board +//! image possible at all: `acmefs.zig` reaches `pardes.zig` and the whole core, and this file +//! reaches `std` and one leaf table. +//! +//! NO ALLOCATOR, NO OS, ONE REQUEST AT A TIME. Same rules as `acmefs`: `handle(req) -> Answer` is a +//! pure transaction, the answer's bytes are either `.rodata` or the one shared buffer, and they are +//! borrowed until the next call. Nothing here blocks, so `Status.again` never appears — the board +//! has no `event` file and no reader to park. +const std = @import("std"); +const board_pins = @import("board_pins.zig"); + +/// The errno values this tree returns. `acmefs.E`'s subset — the four a tree with no panes, no +/// blocking and no allocation can produce — with the same numbers, because they are Linux's and a +/// second spelling would be a second thing to check against `9p.errString`. +pub const E = struct { + pub const NOENT: u16 = 2; + pub const IO: u16 = 5; + pub const NOTDIR: u16 = 20; + pub const INVAL: u16 = 22; +}; + +/// How a HANDLER refuses, as against how the tree refuses. The tree answers ENOENT and ENOTDIR +/// itself, out of the table, before any handler runs; this is the set of things only the handler can +/// know. +/// +/// One variant today, and it is the honest count: a pad takes `0` or `1` and nothing else. A +/// capability that can refuse for a second reason adds a variant here and a prong to `errnoOf`, +/// which is the whole of what "another kind of no" costs. +pub const Fault = error{ + /// the bytes offered are not a value this file takes + Malformed, +}; + +/// The one place a `Fault` becomes a number. +fn errnoOf(f: Fault) u16 { + return switch (f) { + error.Malformed => E.INVAL, + }; +} + +/// A file's two halves, as POINTERS rather than function types: the derived table below is an +/// ordinary runtime array, and a struct holding a bare `fn` is comptime-only. +/// +/// `key` says which pad, address or counter the call is about, and `out` is the slice of the shared +/// answer buffer this file's table entry declared — exactly `scratch` bytes, so a handler cannot +/// write past its own budget. A read may also ignore `out` entirely and answer out of `.rodata`, +/// which is what the JP1 drawing does. +const ReadFn = *const fn (key: u16, out: []u8) Fault![]const u8; +const WriteFn = *const fn (key: u16, bytes: []const u8) Fault!u32; + +/// One FILE in the table. `key` is not here: it comes from the directory the file is generated +/// into, which is what makes one entry serve eleven pins. +/// +/// The MODE is derived, never declared: a file with both handlers is 0o600, a read handler alone is +/// 0o400, a write handler alone is 0o200, and neither is a compile error. A declared mode is a +/// fourth thing that can disagree with the three that decide it. +pub const FileSpec = struct { + name: []const u8, + read: ?ReadFn = null, + write: ?WriteFn = null, + /// Bytes of the shared answer buffer this file's read needs. ZERO when the read answers out of + /// `.rodata` and copies nothing, which is what `gpio/pinout` does — the JP1 drawing is 468 + /// bytes of static text and there is no reason to stage it. The largest `scratch` in the table + /// is one of the two numbers that size `Tree.out`. + scratch: u32 = 0, +}; + +/// One generated subdirectory of a capability, and its KEY: the pad, address or counter every file +/// inside it is about. `gpio/20/value` is `key = 20`. +pub const FanDir = struct { key: u16, name: []const u8 }; + +/// A capability's fan-out: one directory per key, each holding the same files. THE REASON the tree +/// has exactly the pins this board has — the dirs are collected from `board_pins.gpio_pins`, which +/// is collected from the JP1 rows, which are the schematic. +pub const Fan = struct { dirs: []const FanDir, files: []const FileSpec }; + +/// ONE CAPABILITY = ONE DIRECTORY under the root. Always a directory, even for a capability with a +/// single file: a flat root would put every capability's names in one u4 (see `block` below) and +/// would make `ls /` a list of files whose grouping a reader has to infer. `ls /` here is the list +/// of things this board can do. +pub const Cap = struct { + name: []const u8, + files: []const FileSpec = &.{}, + fan: ?Fan = null, +}; + +/// One node of the derived tree. Flat, because a table of fifteen entries scanned linearly is +/// faster than any structure with pointers in it and is the same shape `src/9p.zig`'s own test stub +/// uses — and because a scan cannot disagree with itself about what the tree contains. +const Entry = struct { + node: u64, + /// Where `..` goes. See `block`: this is also the value `src/9p.zig`'s `parentOf` derives from + /// the node id, for every entry but a fan leaf, and the test at the bottom asserts it. + parent: u64, + name: []const u8, + dir: bool, + mode: u16, + /// the pad this file is about, or zero + key: u16 = 0, + read: ?ReadFn = null, + write: ?WriteFn = null, + scratch: u32 = 0, +}; + +/// THE NODE ID PACKING, and it is not ours: it is `acmefs.Node`'s, `{ file: u4, serial: u60 }`, +/// because `src/9p.zig:1955` `parentOf` READS node ids to answer `..` and has that packing built in. +/// A tree that numbered its nodes freely would get a wrong answer to `cd ..` and no diagnostic. +/// +/// The rule, restated as arithmetic: a node's parent is the node rounded down to a multiple of 16, +/// except that a node already at a multiple of 16 — or below 16 — is a child of the root. +/// +/// * the root is 1: serial 0, so `..` is itself, which is POSIX's rule and `intro(5)`'s. +/// * a capability directory is its own BLOCK BASE, `(index + 1) * 16`, so its `..` is the root. +/// * everything inside a capability — its files AND its fan directories — is a member of that +/// block, `base + 1 .. base + 15`, so their `..` is the capability directory. Correct, which is +/// what matters for the one `..` a client actually performs: `cd /gpio/20; cd ..`. +/// * a fan LEAF (`gpio/20/value`) cannot be expressed. Its parent is a block member, and +/// `parentOf` can only produce block bases. So leaves get blocks of their own, above every +/// capability's, and `..` from one lands on an unallocated block base, which this tree answers +/// ENOENT. That is the honest failure: a walk that cannot be expressed is refused rather than +/// silently landing on a different file. No client does it — `..` from a file requires having +/// walked INTO a file, and a file is not a directory — and the fix, if one is ever wanted, is a +/// `parent` hook on `Server` so a filesystem deeper than two levels answers `..` itself. That +/// is exactly the wall `parentOf`'s own doc comment says it is (`src/9p.zig:1950-1954`), and +/// `acmefs`'s `pty/` subtree stands on the same side of it today. +const block: u64 = 16; + +/// The root, and the value the runtime hands `Server.init` as `Options.root`. One, for the same +/// reason `acmefs.TopFile.root` is one: node 0 is `{ file: 0, serial: 0 }` and cannot be a root +/// (`src/9p.zig:2190-2192`). +pub const root: u64 = 1; + +/// Every key the GPIO fan generates, re-exported for the RUNTIME's benefit: `src/esp32p4_9p.zig` +/// checks at comptime that each one is a pad `hal.gpio` will accept, which is the one thing this +/// file cannot check for itself — `max_pin` is a property of the chip package and lives in the +/// toolchain repository, and importing it here would make the tree unbuildable on a host. +pub const pins = board_pins.gpio_pins; + +/// The board's own tree, over a `board` seam the runtime supplies. +/// +/// GENERIC over the board for exactly the reason `Server` is generic over the filesystem: the pads +/// are four register files behind `hal.gpio` in the toolchain package, which exists only for +/// riscv32, and a tree that imported it could not be tested on a host at all. The seam is two +/// functions, both about the level the board is DRIVING: +/// +/// * `board.level(pin: u8) u1` +/// * `board.drive(pin: u8, level: u1) void` +/// +/// `src/esp32p4_9p.zig` implements them over `hal.gpio`, in the same four calls +/// `src/esp32p4/app.zig:200-209` uses for the `Gpio` word — the same seam, a second caller, not a +/// second copy of the register sequence. The tests below implement them over a recording stub, the +/// way `src/9p.zig`'s server tests implement a filesystem. +pub fn Tree(comptime board: type) type { + return struct { + const Self = @This(); + + // -- the ABI, which is `acmefs`'s ------------------------------------ + // + // A MIRROR, not a redefinition: `Server(acmefs)` is the instantiation that proves the + // shape, and a field that drifts from it is a compile error the moment `Server(Tree(...))` + // is built — which the tests at the bottom do. + + pub const Op = enum(u8) { lookup, getattr, setattr, open, read, write, release, readdir, statfs }; + + /// `again` is here because the ABI has it, and it never occurs: nothing on this board + /// blocks. The board's answer to "what is this pin at" is a register read. + pub const Status = enum(u8) { ok, again, err }; + + pub const Req = struct { + tag: u64, + op: Op, + node: u64, + handle: u32 = 0, + off: u64 = 0, + size: u32 = 0, + data: []const u8 = &.{}, + truncate: bool = false, + }; + + pub const Reply = struct { + tag: u64, + status: Status = .ok, + errno: u16 = 0, + attr: Attr = .{}, + handle: u32 = 0, + written: u32 = 0, + + pub const Attr = struct { + node: u64 = 0, + dir: bool = false, + size: u64 = 0, + mode: u16 = 0o600, + }; + + /// `acmefs.Reply.fail`'s twin, so a refusal is one expression here as it is there. + pub fn fail(tag: u64, e: u16) Reply { + return .{ .tag = tag, .status = .err, .errno = e }; + } + }; + + /// A reply and the bytes it points at, borrowed until the next `handle`. `Server.reply` + /// takes exactly this pair. + pub const Answer = struct { reply: Reply, bytes: []const u8 = "" }; + + // -- the handlers ---------------------------------------------------- + + /// `gpio/pinout` — JP1, as the `Gpio` word draws it, TO THE BYTE. The same + /// `board_pins.jp1_text` the word prints (`src/board_memory.zig:364`), returned out of + /// `.rodata` rather than staged, so this read costs no buffer and no copy. + fn readPinout(_: u16, _: []u8) Fault![]const u8 { + return board_pins.jp1_text; + } + + /// `gpio//value` — the level this board is DRIVING on pad `n`, as `0` or `1` and a + /// newline. + /// + /// THE DRIVEN LEVEL and not the pad's, for the reason `src/esp32p4/app.zig:196-199` gives: + /// the pad's own level is what the outside world says, and on an unconnected header pin that + /// is noise. The driven level is defined for every pin, which is what a file that a script + /// reads in a loop needs. + /// + /// The trailing newline is not decoration: `cat gpio/20/value` in a terminal and `$(cat + /// ...)` in a script both want it, and the write side accepts it back, so `cp` of one pin's + /// value onto another's is a legal round trip. + fn readValue(key: u16, out: []u8) Fault![]const u8 { + out[0] = '0' + @as(u8, board.level(@intCast(key))); + out[1] = '\n'; + return out[0..2]; + } + + /// `gpio//value` — drive pad `n` to `0` or `1`. + /// + /// WRITING THE OPPOSITE OF THE CURRENT LEVEL IS THE `Gpio` WORD'S TOGGLE, through the same + /// seam; writing the level it is already at is not a no-op, because the FIRST write to a pad + /// is what makes it an output at all (`hal.gpio.configureOutput`, four register files). So + /// this always drives, and `echo 0 > value` on a fresh boot is a meaningful command: it + /// takes the pad off whatever the IO MUX had it pointed at and holds it low. + /// + /// `0`, `1`, `0\n` and `1\n` are the whole language. Anything else is EINVAL, including + /// `true`, `high`, `01` and the empty write — a file whose only two values are one character + /// each has no room for a spelling debate, and guessing at `on` would be the beginning of + /// one. + fn writeValue(key: u16, bytes: []const u8) Fault!u32 { + const want = try oneBit(bytes); + board.drive(@intCast(key), want); + // The whole write is consumed, trailing newline included: a short count would make + // `echo` retry the tail and drive the pin a second time. + return @intCast(bytes.len); + } + + /// `0` or `1`, with at most one trailing newline (and the `\r` a Windows-ish client may put + /// in front of it). Nothing else. + fn oneBit(bytes: []const u8) Fault!u1 { + var end = bytes.len; + while (end > 0 and (bytes[end - 1] == '\n' or bytes[end - 1] == '\r')) end -= 1; + if (end != 1) return error.Malformed; + return switch (bytes[0]) { + '0' => 0, + '1' => 1, + else => error.Malformed, + }; + } + + // -- THE TABLE ------------------------------------------------------- + + /// The pin directories, one per P4 GPIO the header brings out, named by the pin number in + /// DECIMAL — the number the schematic, the silkscreen and the datasheet all use, and the one + /// literal in `board_memory.zig` that is not hex (`:394-399`). Generated from + /// `board_pins.gpio_pins`, so this list cannot contain a pin JP1 does not have. + const gpio_dirs = dirs: { + var out: [board_pins.gpio_pins.len]FanDir = undefined; + for (board_pins.gpio_pins, 0..) |pin, i| out[i] = .{ + .key = pin, + .name = std.fmt.comptimePrint("{d}", .{pin}), + }; + break :dirs out; + }; + + /// EVERYTHING THIS BOARD OFFERS, and the only place any of it is described. The tree, the + /// permissions, the handlers, the node ids and the answer buffer all come out of here. + const caps = [_]Cap{ + .{ + .name = "gpio", + .files = &.{ + .{ .name = "pinout", .read = readPinout }, + }, + .fan = .{ + .dirs = &gpio_dirs, + .files = &.{ + .{ .name = "value", .read = readValue, .write = writeValue, .scratch = 2 }, + }, + }, + }, + }; + + /// How many nodes the table generates, counted separately because it is an array length. + const node_count = count: { + var n: usize = 1; // the root + for (caps) |c| { + n += 1 + c.files.len; + if (c.fan) |f| n += f.dirs.len * (1 + f.files.len); + } + break :count n; + }; + + /// THE DERIVED TREE. Built once at comptime and `const`, so it lands in `.rodata` and costs + /// the image its bytes and the board's RAM nothing. + const table: [node_count]Entry = build: { + var out: [node_count]Entry = undefined; + out[0] = .{ .node = root, .parent = root, .name = "/", .dir = true, .mode = 0o500 }; + var at: usize = 1; + // Blocks 1..caps.len are the capability directories; fan leaves take the ones above, + // which is what keeps a leaf's unexpressible parent from landing on a real node. + var next_block: u64 = caps.len + 1; + for (caps, 0..) |c, ci| { + const dir_node = (ci + 1) * block; + out[at] = .{ .node = dir_node, .parent = root, .name = c.name, .dir = true, .mode = 0o500 }; + at += 1; + // The u4 in the node id, spent one per name inside this capability. Directories and + // files come out of the same fifteen, which is the wall `acmefs.PaneFile`'s doc + // comment describes from the other side. + var slot: u64 = 1; + for (c.files) |f| { + out[at] = fileEntry(dir_node + slot, dir_node, f, 0); + at += 1; + slot += 1; + } + if (c.fan) |fan| for (fan.dirs) |d| { + const fan_node = dir_node + slot; + slot += 1; + out[at] = .{ .node = fan_node, .parent = dir_node, .name = d.name, .dir = true, .mode = 0o500 }; + at += 1; + const leaf_base = next_block * block; + next_block += 1; + for (fan.files, 0..) |f, l| { + out[at] = fileEntry(leaf_base + 1 + l, fan_node, f, d.key); + at += 1; + } + }; + if (slot >= block) @compileError( + "capability '" ++ c.name ++ + "' has more than 15 names in it, and a node id has four bits for them:" ++ + " `acmefs.Node.file` is a u4 and `9p.parentOf` reads it. Split it into two" ++ + " capabilities, or widen the packing in acmefs.zig, 9p.zig and here at once.", + ); + } + break :build out; + }; + + /// One file's entry, with the mode derived from which handlers it has. + fn fileEntry(node: u64, parent: u64, f: FileSpec, key: u16) Entry { + const mode: u16 = if (f.read != null and f.write != null) + 0o600 + else if (f.read != null) + 0o400 + else if (f.write != null) + 0o200 + else + @compileError("file '" ++ f.name ++ "' has no read and no write, so it is a name and not a file"); + return .{ + .node = node, + .parent = parent, + .name = f.name, + .dir = false, + .mode = mode, + .key = key, + .read = f.read, + .write = f.write, + .scratch = f.scratch, + }; + } + + /// THE ONE BUFFER, and both numbers that size it come out of the table: the largest + /// `scratch` any read declares, and the widest directory's worth of staged entries. Never + /// both at once — one request is in flight at a time — so one buffer serves both, and the + /// board pays for the larger. + const out_max = size: { + var most: usize = 0; + for (table) |e| most = @max(most, e.scratch); + for (table) |d| { + if (!d.dir) continue; + var n: usize = 0; + for (table) |e| if (e.parent == d.node and e.node != d.node) { + n += dirent_fixed + e.name.len; + }; + most = @max(most, n); + } + break :size most; + }; + + /// `node[8] dir[1] namelen[1]` — `acmefs`'s staging format for a readdir + /// (`acmefs.zig:942-957`), which is what `Server` decodes. Ten bytes and then the name. + const dirent_fixed = 8 + 1 + 1; + + /// Formatted answers and staged directory entries. Valid until the next `handle`, which is + /// the borrow window `Server.reply` documents. + out: [out_max]u8 = undefined, + + /// Every request the board has been asked, for the runtime's own diagnostics. Not a + /// protocol counter — `Server` keeps those — and not a statistic anybody has to read: it is + /// the one number that distinguishes "nothing is arriving" from "everything is being + /// refused" on a board with no second console to ask. + calls: u32 = 0, + + fn find(node: u64) ?*const Entry { + for (&table) |*e| if (e.node == node) return e; + return null; + } + + fn attrOf(t: *Self, e: *const Entry) Reply.Attr { + return .{ .node = e.node, .dir = e.dir, .mode = e.mode, .size = t.sizeOf(e) }; + } + + /// A file's size is WHAT ITS READ ANSWERS, asked rather than declared. That means a + /// `getattr` of `gpio/20/value` reads the pad's output register, which is a load from a + /// peripheral and nothing more; the alternative is a second declaration in the table that + /// can disagree with the handler, on a tree whose whole claim is that there is one place per + /// fact. A write-only file has no size and reports zero, which is what `acmefs` reports for + /// every file it cannot cheaply measure. + fn sizeOf(t: *Self, e: *const Entry) u64 { + const read = e.read orelse return 0; + const bytes = read(e.key, t.out[0..e.scratch]) catch return 0; + return bytes.len; + } + + /// ONE OPERATION, and the whole of what this filesystem is. Pure: no allocation, no + /// blocking, no state but `out` and the counter. + pub fn handle(t: *Self, req: Req) Answer { + t.calls += 1; + const e = find(req.node) orelse return .{ .reply = .fail(req.tag, E.NOENT) }; + switch (req.op) { + .lookup => { + if (!e.dir) return .{ .reply = .fail(req.tag, E.NOTDIR) }; + for (&table) |*c| { + if (c.parent != req.node or c.node == req.node) continue; + if (!std.mem.eql(u8, c.name, req.data)) continue; + return .{ .reply = .{ .tag = req.tag, .attr = t.attrOf(c) } }; + } + return .{ .reply = .fail(req.tag, E.NOENT) }; + }, + .getattr => return .{ .reply = .{ .tag = req.tag, .attr = t.attrOf(e) } }, + // The only `setattr` that reaches here is a truncate, from `Topen` with `OTRUNC` + // (`src/9p.zig:2816-2822`) — which is what `echo 1 > gpio/20/value` opens with. + // Every file here is a fixed-length register view, so there is nothing to truncate + // and nothing to refuse either: answering EINVAL would make the shell's own + // redirection fail on a pin that is perfectly writable. + .setattr => { + if (e.dir) return .{ .reply = .fail(req.tag, E.INVAL) }; + return .{ .reply = .{ .tag = req.tag, .attr = t.attrOf(e) } }; + }, + // No per-open state, so one handle for every open. `Server` checks the mode against + // the fid's cached permissions before it gets here (`src/9p.zig:2075-2079`). + .open => return .{ .reply = .{ .tag = req.tag, .handle = 1 } }, + .release => return .{ .reply = .{ .tag = req.tag } }, + .read => { + if (e.dir) return .{ .reply = .fail(req.tag, E.INVAL) }; + const read = e.read orelse return .{ .reply = .fail(req.tag, E.INVAL) }; + const all = read(e.key, t.out[0..e.scratch]) catch |f| { + return .{ .reply = .fail(req.tag, errnoOf(f)) }; + }; + // Past the end is the empty read every client uses to stop, not an error. + if (req.off >= all.len) return .{ .reply = .{ .tag = req.tag } }; + const from = all[@intCast(req.off)..]; + return .{ .reply = .{ .tag = req.tag }, .bytes = from[0..@min(from.len, req.size)] }; + }, + .write => { + if (e.dir) return .{ .reply = .fail(req.tag, E.INVAL) }; + const write = e.write orelse return .{ .reply = .fail(req.tag, E.INVAL) }; + // A REGISTER IS NOT A STREAM. Every file here is one value, so the only offset + // that means anything is zero; a client that seeks and writes is describing an + // edit to a byte range this file does not have. `echo`, `9p write` and + // `cat > file` all write at zero. + if (req.off != 0) return .{ .reply = .fail(req.tag, E.INVAL) }; + const n = write(e.key, req.data) catch |f| { + return .{ .reply = .fail(req.tag, errnoOf(f)) }; + }; + return .{ .reply = .{ .tag = req.tag, .written = n } }; + }, + .readdir => { + if (!e.dir) return .{ .reply = .fail(req.tag, E.NOTDIR) }; + return .{ .reply = .{ .tag = req.tag }, .bytes = t.stage(req.node, req.off) }; + }, + // 9P2000 has no `Tstatfs` — that is a `.L` message (`src/9p.zig:24-28`) — so + // nothing reaches this. It is answered rather than `unreachable` because the ABI + // names it and a panic in a server is worse than an empty answer. + .statfs => return .{ .reply = .{ .tag = req.tag } }, + } + } + + /// A directory's children in `acmefs`'s staging format, from an ENTRY INDEX rather than a + /// byte offset — `Server` does that coordinate change and advances both cursors + /// (`src/9p.zig:2836-2849`). The whole of the widest directory fits `out` by construction, + /// so this never stages a short list for want of room; `Server` still takes only what one + /// reply holds and asks again. + fn stage(t: *Self, node: u64, skip: u64) []const u8 { + var n: usize = 0; + var seen: u64 = 0; + for (&table) |*e| { + if (e.parent != node or e.node == node) continue; + if (seen < skip) { + seen += 1; + continue; + } + std.mem.writeInt(u64, t.out[n..][0..8], e.node, .little); + t.out[n + 8] = @intFromBool(e.dir); + t.out[n + 9] = @intCast(e.name.len); + @memcpy(t.out[n + dirent_fixed ..][0..e.name.len], e.name); + n += dirent_fixed + e.name.len; + } + return t.out[0..n]; + } + }; +} + +// --------------------------------------------------------------------------- +// tests +// --------------------------------------------------------------------------- +// +// A RECORDING STUB FOR THE PADS, exactly as `src/9p.zig`'s server tests use a stub filesystem: the +// seam is two functions, so the test can hold the pads still and check what was asked of them. Every +// claim below is one a host can answer — the tree's shape, the bytes of an answer, which pin the +// seam was called with — and the one claim it cannot is stated as such: whether `hal.gpio` drives +// the pad, which only the die knows. + +const testing = std.testing; + +/// The pads, faked. `driven` is the board's output register. +const StubPads = struct { + var driven: [64]u1 = @splat(0); + var log: [16]Call = undefined; + var log_len: usize = 0; + + const Call = struct { pin: u8, level: u1 }; + + fn reset() void { + driven = @splat(0); + log_len = 0; + } + + fn level(pin: u8) u1 { + return driven[pin]; + } + + fn drive(pin: u8, want: u1) void { + driven[pin] = want; + log[log_len] = .{ .pin = pin, .level = want }; + log_len += 1; + } +}; + +const Board = Tree(StubPads); + +/// The tree, walked by name the way a client walks it: `lookup` after `lookup` from the root, which +/// is the only way to find out what the generated table actually offers. +fn walk(t: *Board, path: []const []const u8) !Board.Reply.Attr { + var at: u64 = root; + var attr: Board.Reply.Attr = .{ .node = root, .dir = true, .mode = 0o500 }; + for (path) |name| { + const a = t.handle(.{ .tag = 1, .op = .lookup, .node = at, .data = name }); + if (a.reply.status == .err) return switch (a.reply.errno) { + E.NOENT => error.NoEntry, + E.NOTDIR => error.NotDirectory, + else => error.Refused, + }; + attr = a.reply.attr; + at = attr.node; + } + return attr; +} + +fn readAll(t: *Board, node: u64) !Board.Answer { + const open = t.handle(.{ .tag = 1, .op = .open, .node = node }); + try testing.expectEqual(Board.Status.ok, open.reply.status); + return t.handle(.{ .tag = 2, .op = .read, .node = node, .handle = open.reply.handle, .size = 65535 }); +} + +test "board9p: the generated tree has exactly the header's pins, and nothing else" { + var t: Board = .{}; + + // The capability directory, and its one hand-written file. + try testing.expect((try walk(&t, &.{"gpio"})).dir); + try testing.expect(!(try walk(&t, &.{ "gpio", "pinout" })).dir); + + // Every pin JP1 brings out is a directory with a `value` in it. Eleven of them, generated. + for (board_pins.gpio_pins) |pin| { + var name: [4]u8 = undefined; + const dir = try std.fmt.bufPrint(&name, "{d}", .{pin}); + try testing.expect((try walk(&t, &.{ "gpio", dir })).dir); + const value = try walk(&t, &.{ "gpio", dir, "value" }); + try testing.expect(!value.dir); + try testing.expectEqual(@as(u16, 0o600), value.mode); + } + + // And a pin the board does not bring out is not there. 6 and 21 are real ESP32-P4 GPIOs that + // JP1 simply does not route, which is the distinction the table exists to keep: the tree has + // the pins the BOARD has, not the pins the CHIP has. + try testing.expectError(error.NoEntry, walk(&t, &.{ "gpio", "6" })); + try testing.expectError(error.NoEntry, walk(&t, &.{ "gpio", "21" })); + try testing.expectError(error.NoEntry, walk(&t, &.{ "gpio", "20", "level" })); + try testing.expectError(error.NoEntry, walk(&t, &.{"mem"})); +} + +test "board9p: a read of gpio/pinout is the bytes the Gpio word draws" { + var t: Board = .{}; + const at = try walk(&t, &.{ "gpio", "pinout" }); + // `board_memory.zig:364`'s `pinout` IS this declaration, so this is the word's own output and + // not a copy of it. The bytes themselves are pinned by `board_pins.zig`'s golden test. + const a = try readAll(&t, at.node); + try testing.expectEqualStrings(board_pins.jp1_text, a.bytes); + // The size a client is told matches what it gets, which is what makes `cat` stop in one read. + try testing.expectEqual(board_pins.jp1_text.len, at.size); + // Read-only: the drawing is the header's, not the client's. + try testing.expectEqual(@as(u16, 0o400), at.mode); + const w = t.handle(.{ .tag = 3, .op = .write, .node = at.node, .data = "x" }); + try testing.expectEqual(E.INVAL, w.reply.errno); +} + +test "board9p: writing 1 then 0 drives the pad twice, through the seam" { + StubPads.reset(); + var t: Board = .{}; + const at = try walk(&t, &.{ "gpio", "20", "value" }); + + // A fresh pad reads 0 — the level the board is DRIVING, which is defined before anybody has + // written anything. + const before = try readAll(&t, at.node); + try testing.expectEqualStrings("0\n", before.bytes); + + const one = t.handle(.{ .tag = 4, .op = .write, .node = at.node, .data = "1" }); + try testing.expectEqual(Board.Status.ok, one.reply.status); + try testing.expectEqual(@as(u32, 1), one.reply.written); + try testing.expectEqualStrings("1\n", (try readAll(&t, at.node)).bytes); + + // `echo 0 > value`, newline and all: the whole write is consumed, so the shell does not retry + // the tail and drive the pin a second time. + const zero = t.handle(.{ .tag = 5, .op = .write, .node = at.node, .data = "0\n" }); + try testing.expectEqual(@as(u32, 2), zero.reply.written); + try testing.expectEqualStrings("0\n", (try readAll(&t, at.node)).bytes); + + // TWO CALLS, the right pin, the right levels, in order. This is the whole of what the host can + // check about the seam; that `hal.gpio` then moves the pad is the die's to answer. + try testing.expectEqual(@as(usize, 2), StubPads.log_len); + try testing.expectEqual(StubPads.Call{ .pin = 20, .level = 1 }, StubPads.log[0]); + try testing.expectEqual(StubPads.Call{ .pin = 20, .level = 0 }, StubPads.log[1]); +} + +test "board9p: a pad takes 0 and 1 and refuses everything else, without touching the pads" { + StubPads.reset(); + var t: Board = .{}; + const at = try walk(&t, &.{ "gpio", "45", "value" }); + + for ([_][]const u8{ "2", "", "01", "x", "true", "high", "\n", "1 ", " 1", "10" }) |bad| { + const a = t.handle(.{ .tag = 6, .op = .write, .node = at.node, .data = bad }); + try testing.expectEqual(Board.Status.err, a.reply.status); + try testing.expectEqual(E.INVAL, a.reply.errno); + } + // A refused write is a pad that was never driven, which is the part that matters: a half-parsed + // command must not leave the board in a state nobody asked for. + try testing.expectEqual(@as(usize, 0), StubPads.log_len); + + // A register is one value, so a write at an offset is refused too, and refused before the pads. + const off = t.handle(.{ .tag = 7, .op = .write, .node = at.node, .off = 1, .data = "1" }); + try testing.expectEqual(E.INVAL, off.reply.errno); + try testing.expectEqual(@as(usize, 0), StubPads.log_len); +} + +test "board9p: every node's parent is the one 9p.parentOf derives, or an unallocated block" { + // THE ENCODING'S OWN TEST, and it defends the one thing this file cannot see: `src/9p.zig` + // answers `..` from the node id alone, by the rule restated at `block` above. A node numbered + // outside that rule would make `cd ..` land somewhere else with no diagnostic, so the rule is + // applied here to every generated node and compared against the table's own `parent`. + for (&Board.table) |*e| { + const serial = e.node >> 4; + const file = e.node & 0xF; + const derived: u64 = if (e.node == root or serial == 0 or file == 0) root else serial << 4; + if (derived == e.parent) continue; + // The one exception, and it must be exactly the one documented: a fan leaf, whose parent is + // a block MEMBER and therefore unexpressible. Its derived parent has to be a node that does + // not exist, so the walk is refused rather than landing on the wrong file. + try testing.expectEqualStrings("value", e.name); + var t: Board = .{}; + const a = t.handle(.{ .tag = 8, .op = .getattr, .node = derived }); + try testing.expectEqual(E.NOENT, a.reply.errno); + } +} + +test "board9p: a directory read lists what the table generated, in table order" { + var t: Board = .{}; + + // The root is the capability list, and today that is one name. + try testing.expectEqualStrings("gpio", (try names(&t, root, 0))[0]); + try testing.expectEqual(@as(usize, 1), (try names(&t, root, 0)).len); + + const gpio = (try walk(&t, &.{"gpio"})).node; + const listing = try names(&t, gpio, 0); + try testing.expectEqual(board_pins.gpio_pins.len + 1, listing.len); + try testing.expectEqualStrings("pinout", listing[0]); + for (board_pins.gpio_pins, 0..) |pin, i| { + var buf: [4]u8 = undefined; + try testing.expectEqualStrings(try std.fmt.bufPrint(&buf, "{d}", .{pin}), listing[i + 1]); + } + + // The cursor is an ENTRY INDEX, which is what `Server` advances between reads of a directory + // bigger than one reply. + const rest = try names(&t, gpio, 5); + try testing.expectEqual(board_pins.gpio_pins.len + 1 - 5, rest.len); + try testing.expectEqualStrings("5", rest[0]); +} + +/// The names in one staged directory read, decoded out of `acmefs`'s `node[8] dir[1] namelen[1] +/// name[]` records — the same decode `Server` does. +var name_slots: [32][]const u8 = undefined; +fn names(t: *Board, node: u64, skip: u64) ![][]const u8 { + const a = t.handle(.{ .tag = 9, .op = .readdir, .node = node, .off = skip, .size = 65535 }); + try testing.expectEqual(Board.Status.ok, a.reply.status); + var n: usize = 0; + var i: usize = 0; + while (i < a.bytes.len) { + const len = a.bytes[i + 9]; + name_slots[n] = a.bytes[i + 10 ..][0..len]; + n += 1; + i += 10 + len; + } + return name_slots[0..n]; +} + +test "board9p: the whole tree costs one buffer, and the table says how big" { + // The two numbers the board's RAM budget is quoted from. `out` is the ONLY buffer this + // filesystem has, and both of its bounds come out of the table: the widest directory's staged + // entries (gpio's twelve) and the largest read scratch (a pin's two bytes). + try testing.expectEqual(@as(usize, 143), Board.out_max); + try testing.expect(@sizeOf(Board) <= 160); + // The JP1 drawing is not in it, and that is the point of `scratch = 0`: 468 bytes of static + // text are served straight out of `.rodata`. + try testing.expect(board_pins.jp1_text.len > Board.out_max); +} + +// The proof that the ABI claim in this file's header is true, and the only place the two halves meet +// on the host: `Server` is a generic over exactly `Op`, `Status`, `Req`, `Reply` and `Reply.Attr`, +// so a field that drifts from `acmefs`'s is a compile error HERE, and a real client's bytes are what +// comes out. +// +// A PATH IMPORT, and it took two goes to get here. The first was +// `@import("9p.zig")`, which did not compile while `src/9p.zig` was also the +// ROOT of a named `ninep` module in the same link — a file belongs to exactly +// one module, and it was both. The second was a named module, declared twice in +// `build.zig`; that compiled and made this file unbuildable by anyone but +// `build.zig`, which is what broke the board image the moment the firmware +// link moved to the toolchain repository and stopped injecting modules. +// +// The path form works now because nothing declares `src/9p.zig` as a module +// root any more: `fs9_service.zig` and `fs9_client.zig` reach it by path too, +// so every link that contains it contains it once. The gain is that this file +// and `src/esp32p4_9p.zig` are self-contained — `zig test src/board9p.zig` +// works with no flags, and any builder can root an image at `nine.zig` without +// being told what modules to inject. +const ninep = @import("9p.zig"); + +test "board9p: a real 9P client reads a pin's value off this tree" { + StubPads.reset(); + const Server = ninep.Server(Board); + // The board's own buffers, at the board's own msize. See `src/esp32p4_9p.zig` for why 1024. + var in: [1024]u8 = undefined; + var out: [2048]u8 = undefined; + var fsys: Board = .{}; + var srv = Server.init(.{ .in = &in, .out = &out, .root = root }); + + var scratch: [256]u8 = undefined; + const send = struct { + fn call(s: *Server, f: *Board, buf: []u8, tag: u16, msg: ninep.Msg) !void { + const bytes = try ninep.encode(msg, tag, buf); + try testing.expectEqual(bytes.len, s.push(bytes)); + while (s.retry()) |req| { + const a = f.handle(req); + s.reply(&a.reply, a.bytes); + } + while (s.next()) |req| { + const a = f.handle(req); + s.reply(&a.reply, a.bytes); + } + } + }.call; + const reap = struct { + fn call(s: *Server) !ninep.Decoded { + const queued = s.output(); + const len = ninep.frameLen(queued) orelse return error.NoReply; + const got = try ninep.decode(queued[0..len]); + s.wrote(len); + return got; + } + }.call; + + try send(&srv, &fsys, &scratch, ninep.notag, .{ .tversion = .{ .msize = 8192, .version = "9P2000" } }); + const v = try reap(&srv); + // Clamped to what the board's buffers hold, which is the number the RAM budget was chosen for. + try testing.expectEqual(@as(u32, 1024), v.msg.rversion.msize); + + try send(&srv, &fsys, &scratch, 1, .{ .tattach = .{ .fid = 0, .afid = ninep.nofid, .uname = "goblin", .aname = "" } }); + try testing.expectEqual(root, (try reap(&srv)).msg.rattach.qid.path); + + var wname: [ninep.max_welem][]const u8 = @splat(""); + wname[0] = "gpio"; + wname[1] = "20"; + wname[2] = "value"; + try send(&srv, &fsys, &scratch, 2, .{ .twalk = .{ .fid = 0, .newfid = 1, .nwname = 3, .wname = wname } }); + try testing.expectEqual(@as(u16, 3), (try reap(&srv)).msg.rwalk.nwqid); + + try send(&srv, &fsys, &scratch, 3, .{ .topen = .{ .fid = 1, .mode = ninep.ordwr } }); + _ = try reap(&srv); + + // `echo 1 > /mnt/board/gpio/20/value`, as bytes on a wire. + try send(&srv, &fsys, &scratch, 4, .{ .twrite = .{ .fid = 1, .offset = 0, .data = "1\n" } }); + try testing.expectEqual(@as(u32, 2), (try reap(&srv)).msg.rwrite.count); + try testing.expectEqual(@as(u1, 1), StubPads.driven[20]); + + // ...and `cat` of the same file. + try send(&srv, &fsys, &scratch, 5, .{ .tread = .{ .fid = 1, .offset = 0, .count = 512 } }); + try testing.expectEqualStrings("1\n", (try reap(&srv)).msg.rread.data); + + // The refusal reaches the client as an error STRING, which is 9P's only channel for "no": EINVAL + // becomes the wording `9p.errString` gives it, and the pad is not touched. + try send(&srv, &fsys, &scratch, 6, .{ .twrite = .{ .fid = 1, .offset = 0, .data = "on" } }); + try testing.expectEqualStrings(ninep.errString(E.INVAL), (try reap(&srv)).msg.rerror.ename); + try testing.expectEqual(@as(usize, 1), StubPads.log_len); +} diff --git a/src/board_memory.zig b/src/board_memory.zig index ac567b34..4bee9e6e 100644 --- a/src/board_memory.zig +++ b/src/board_memory.zig @@ -351,43 +351,17 @@ pub fn poke(p: *Pardes, id: usize, argument: []const u8) !void { /// JP1, the 26-pin header down the left edge of the JC-ESP32P4-M3-DEV, as the board wears it: two /// columns, odd pins on the left, even on the right, pin 1 at the top. /// -/// READ OFF THE VENDOR SCHEMATIC, sheet 2 "Expand IO" -/// (`01-esp32p4-m3/docs/schematics/2_EXPAND_IO&BAT.png`), which is the only document that carries -/// this mapping - the specification PDF's "Interface Description" page is a marketing render, and -/// there is no board user guide. The sheet is a 872x1168 raster, so the assignment was taken from -/// the drawing's own geometry rather than by eye: thirteen wires leave each side of the symbol, a -/// net wire runs ~100 px to its label and a power stub ~21 px, which is what identifies pin 8 as -/// unconnected rather than as the first of the GPIO4x labels. Cross-checked against a second, -/// independent source: `05-zig-p4/build.zig` has always documented `-Dled=20` as "JP1 pin 17", and -/// GPIO20 lands on pin 17 here. +/// MOVED TO `src/board_pins.zig`, where the thirteen rows are DATA and this drawing is rendered +/// from them at comptime. Not for tidiness: the board's 9P image (`src/esp32p4_9p.zig`) links no +/// core, so it cannot import this file — this one imports `pardes.zig` — and that image serves this +/// exact drawing as `gpio/pinout` while generating its per-pin directories from the same rows. The +/// alternative was transcribing a schematic twice, which is two things to maintain and no test that +/// could say which one was wrong. The provenance moved with the rows: which sheet of which +/// schematic, how pin 8 was identified as unconnected, and what `--`, `C6_*` and `ES_I2C_*` mean. /// -/// `--` is a pin the header brings out with nothing behind it. `C6_*` are the ESP32-C6 companion's -/// pads, not the P4's, and toggling a P4 GPIO cannot reach them. `ES_I2C_*` is the audio codec's -/// bus, shared - driving either one by hand while the codec is live is a collision, which is a -/// reason to know the pin is there rather than a reason to hide it. -const pinout = - \\JP1 header - 26 pins, pin 1 top left. - \\Every number here is DECIMAL. - \\ - \\ +---------+ - \\ 3V3 | 1 | 2 | 5V - \\ 3V3 | 3 | 4 | 5V - \\ GND | 5 | 6 | GND - \\ GPIO 1 | 7 | 8 | -- - \\ GPIO 2 | 9 | 10 | GPIO 47 - \\ GPIO 3 | 11 | 12 | GPIO 46 - \\ GPIO 4 | 13 | 14 | GPIO 45 - \\ GPIO 5 | 15 | 16 | GND - \\ GPIO 20 | 17 | 18 | 3V3 - \\ GPIO 32 | 19 | 20 | C6_U0RXD - \\ GPIO 33 | 21 | 22 | C6_U0TXD - \\ES_I2C_SDA | 23 | 24 | C6_IO9 - \\ES_I2C_SCL | 25 | 26 | C6_CHIP_PU - \\ +---------+ - \\ - \\Gpio flips one: 0->1 or 1->0. - \\ -; +/// The test below is unchanged, and it is still the check that matters HERE: whoever renders this +/// drawing, the `Gpio` word's output has to fit the board's own grid in two aligned columns. +const pinout = @import("board_pins.zig").jp1_text; /// `Gpio ` flips one pad and says what it did; `Gpio` alone draws JP1. /// diff --git a/src/board_pins.zig b/src/board_pins.zig new file mode 100644 index 00000000..134cd92c --- /dev/null +++ b/src/board_pins.zig @@ -0,0 +1,186 @@ +//! JP1, the JC-ESP32P4-M3-DEV's 26-pin header, as ONE TABLE that everything else is derived from: +//! the ASCII drawing the `Gpio` word prints, and the pin directories the board's 9P tree generates. +//! +//! WHY THIS IS ITS OWN FILE, and it is the whole reason it exists. The drawing lived in +//! `src/board_memory.zig`, which imports `pardes.zig` and therefore the entire core; the board's 9P +//! image (`src/esp32p4_9p.zig`) links no core at all, so it could not have reached it. The two +//! ways out of that were a second copy of the header in the 9P tree — a table of thirteen rows +//! transcribed off a schematic, maintained twice, with no test that could tell you the day they +//! disagreed — or this: a LEAF that imports `std` and nothing else, so both sides import the same +//! thirteen rows. `board_memory.zig` keeps its `pinout` name as an alias of `jp1_text` and its own +//! shape test, so the console word's output is unchanged to the byte. +//! +//! WHY A TABLE AND NOT THE STRING. The string was the source before, and a string is fine for one +//! consumer that prints it. It is no use at all to the second, which needs to know WHICH of these +//! twenty-six pins are the P4's own GPIOs, because that is the set of directories its tree has. A +//! consumer would have to parse the drawing back out — scan for `GPIO `, take the digits, hope +//! nobody aligned a column differently — which is exactly the sort of code that works until the +//! day the drawing is edited. So the rows are data, the drawing is RENDERED from them at comptime, +//! and `gpio_pins` is COLLECTED from them at comptime. Adding a pin to the header is one row, and +//! the drawing, the pin list and the 9P tree all move together because there is only one of them. +//! +//! READ OFF THE VENDOR SCHEMATIC, sheet 2 "Expand IO" +//! (`01-esp32p4-m3/docs/schematics/2_EXPAND_IO&BAT.png`), which is the only document that carries +//! this mapping — the specification PDF's "Interface Description" page is a marketing render, and +//! there is no board user guide. The sheet is a 872x1168 raster, so the assignment was taken from +//! the drawing's own geometry rather than by eye: thirteen wires leave each side of the symbol, a +//! net wire runs ~100 px to its label and a power stub ~21 px, which is what identifies pin 8 as +//! unconnected rather than as the first of the GPIO4x labels. Cross-checked against a second, +//! independent source: `05-zig-p4/build.zig` has always documented `-Dled=20` as "JP1 pin 17", and +//! GPIO20 lands on pin 17 here. +const std = @import("std"); + +/// What is behind one header pin, and the ONE distinction that matters to both consumers: whether +/// this pad is a GPIO of the ESP32-P4 this program is running on. +/// +/// `.none` is a pin the header brings out with nothing behind it (pin 8). `.net` is a pad that is +/// not the P4's to drive as a GPIO: `3V3`, `5V` and `GND` are power, `C6_*` are the ESP32-C6 +/// companion's pins — toggling a P4 GPIO cannot reach them — and `ES_I2C_*` is the audio codec's +/// bus. The codec's two ARE P4 pads, and they are `.net` anyway, deliberately: the schematic does +/// not name their GPIO numbers, and a tree that invented one would offer a file that drives an +/// unknown pin. They stay in the drawing because a shared bus is a reason to know the pin is there. +pub const Pad = union(enum) { + none, + /// a P4 GPIO, by the number the schematic, the silkscreen and the datasheet all use + gpio: u8, + /// a named net that is not a P4 GPIO + net: []const u8, + + /// The text this pad wears in the drawing. `GPIO 47` and not `GPIO47`: the space is what the + /// header has always printed, and the shape test in `board_memory.zig` matches on it. + pub fn label(p: Pad) []const u8 { + return switch (p) { + .none => "--", + .gpio => |n| std.fmt.comptimePrint("GPIO {d}", .{n}), + .net => |s| s, + }; + } +}; + +/// One row of the header: the odd pin on the left, the even pin on its right, exactly as the board +/// wears it. The pin NUMBERS are not stored — row `i` is pins `2i+1` and `2i+2` — because a +/// hand-written number beside a row is a number that can disagree with its position. +pub const Row = struct { left: Pad, right: Pad }; + +/// JP1 itself: thirteen rows, pin 1 at the top left. THE SINGLE SOURCE for the drawing below, for +/// `gpio_pins`, and for the per-pin directories in `src/board9p.zig`. +pub const jp1 = [13]Row{ + .{ .left = .{ .net = "3V3" }, .right = .{ .net = "5V" } }, + .{ .left = .{ .net = "3V3" }, .right = .{ .net = "5V" } }, + .{ .left = .{ .net = "GND" }, .right = .{ .net = "GND" } }, + .{ .left = .{ .gpio = 1 }, .right = .none }, + .{ .left = .{ .gpio = 2 }, .right = .{ .gpio = 47 } }, + .{ .left = .{ .gpio = 3 }, .right = .{ .gpio = 46 } }, + .{ .left = .{ .gpio = 4 }, .right = .{ .gpio = 45 } }, + .{ .left = .{ .gpio = 5 }, .right = .{ .net = "GND" } }, + .{ .left = .{ .gpio = 20 }, .right = .{ .net = "3V3" } }, + .{ .left = .{ .gpio = 32 }, .right = .{ .net = "C6_U0RXD" } }, + .{ .left = .{ .gpio = 33 }, .right = .{ .net = "C6_U0TXD" } }, + .{ .left = .{ .net = "ES_I2C_SDA" }, .right = .{ .net = "C6_IO9" } }, + .{ .left = .{ .net = "ES_I2C_SCL" }, .right = .{ .net = "C6_CHIP_PU" } }, +}; + +/// The row format, and it is load-bearing rather than cosmetic: a header drawn in two columns stops +/// being a header the moment a row wraps or a column slips, and the widest row here is 34 columns +/// against the board's own 80-column grid. Ten for the left label right-aligned, two for each pin +/// number, and the three bars land under the box's own corners because the left label's field plus +/// one space is eleven characters and `+---------+` is eleven wide. +/// +/// `board_memory.zig`'s "the pinout fits the board's own grid" test is the check that this stays +/// true, and it checks the RENDERED text mechanically — every pin row's first bar in the same +/// column — rather than trusting this string. +const row_format = "{s:>10} | {d:>2} | {d:>2} | {s}\n"; + +/// The box the pin numbers sit inside. Eleven characters, indented by the left label's field width +/// plus the space before the first bar, so its corners are the bars. +const border = " +---------+\n"; + +/// JP1 as the text the `Gpio` word prints and a read of the 9P tree's `gpio/pinout` returns — the +/// SAME BYTES, which is a test in `src/board9p.zig` and not a hope. +/// +/// The trailer names the `Gpio` word, which the 9P image does not have. It is here anyway, because +/// "the same bytes" is worth more than a sentence that is true of both faces and useful to neither: +/// a person reading this table through 9P is a person who has the editor's own console in the other +/// window, and telling them the word that flips a pin is telling them something they can use. The +/// 9P equivalent — writing `0` or `1` to `gpio//value` — is documented where a 9P client will +/// look for it, which is the tree's own doc comment. +pub const jp1_text = text: { + var out: []const u8 = + \\JP1 header - 26 pins, pin 1 top left. + \\Every number here is DECIMAL. + \\ + \\ + ; + out = out ++ border; + for (jp1, 0..) |row, i| out = out ++ std.fmt.comptimePrint( + row_format, + .{ row.left.label(), 2 * i + 1, 2 * i + 2, row.right.label() }, + ); + break :text out ++ border ++ + \\ + \\Gpio flips one: 0->1 or 1->0. + \\ + ; +}; + +/// Every P4 GPIO JP1 brings out, ascending. THE SET OF PIN DIRECTORIES the board's 9P tree has, so +/// that tree has exactly the pins this board has and not a range somebody typed. +/// +/// Ascending rather than in header order, because the consumer is `ls`: the header's order puts 47 +/// between 2 and 3, and a directory listing that counts 1 2 3 4 5 20 32 33 45 46 47 is one a person +/// can scan. Nothing depends on the order — the names are the pin numbers — so it may as well be +/// the readable one. +pub const gpio_pins = pins: { + var found: [2 * jp1.len]u8 = undefined; + var n: usize = 0; + for (jp1) |row| for ([2]Pad{ row.left, row.right }) |p| switch (p) { + .gpio => |g| { + found[n] = g; + n += 1; + }, + else => {}, + }; + std.mem.sort(u8, found[0..n], {}, std.sort.asc(u8)); + break :pins found[0..n].*; +}; + +// The drawing, byte for byte, because it is the one thing here whose CORRECTNESS IS ITS SHAPE and +// because it used to be a string literal: this is the check that the renderer above reproduces what +// the console has always printed. A golden test is the right kind of duplication — the expectation +// is the thing being asserted, and if the two ever differ the diff says which byte. +test "the rendered header is the drawing the console has always printed" { + try std.testing.expectEqualStrings( + \\JP1 header - 26 pins, pin 1 top left. + \\Every number here is DECIMAL. + \\ + \\ +---------+ + \\ 3V3 | 1 | 2 | 5V + \\ 3V3 | 3 | 4 | 5V + \\ GND | 5 | 6 | GND + \\ GPIO 1 | 7 | 8 | -- + \\ GPIO 2 | 9 | 10 | GPIO 47 + \\ GPIO 3 | 11 | 12 | GPIO 46 + \\ GPIO 4 | 13 | 14 | GPIO 45 + \\ GPIO 5 | 15 | 16 | GND + \\ GPIO 20 | 17 | 18 | 3V3 + \\ GPIO 32 | 19 | 20 | C6_U0RXD + \\ GPIO 33 | 21 | 22 | C6_U0TXD + \\ES_I2C_SDA | 23 | 24 | C6_IO9 + \\ES_I2C_SCL | 25 | 26 | C6_CHIP_PU + \\ +---------+ + \\ + \\Gpio flips one: 0->1 or 1->0. + \\ + , jp1_text); +} + +// The pin list is the tree's shape, so it is asserted as a list rather than as a count: a row edited +// wrongly changes WHICH pins the board offers, and a count would not notice a 45 that became a 44. +test "the header's own GPIOs, and only those" { + try std.testing.expectEqualSlices(u8, &.{ 1, 2, 3, 4, 5, 20, 32, 33, 45, 46, 47 }, &gpio_pins); + // Pin 8 is unconnected and pin 24 is the C6's, so neither contributes a pad. Both are counted + // here rather than only drawn, because "the tree has exactly the pins the board has" is a claim + // about what is ABSENT as much as what is present. + try std.testing.expectEqual(Pad.none, jp1[3].right); + try std.testing.expectEqualStrings("C6_IO9", jp1[11].right.net); +} diff --git a/src/builtins.zig b/src/builtins.zig index 0e92f5ce..4e096c97 100644 --- a/src/builtins.zig +++ b/src/builtins.zig @@ -29,6 +29,11 @@ const image_pane = @import("image_pane.zig"); const config = @import("config.zig"); const runtime_config = @import("runtime_config.zig"); const board_memory = @import("board_memory.zig"); +/// The host half of the 9P client, for the `9p` word at the bottom. Imported +/// unconditionally and gated on `fs9_client.supported`, exactly like +/// board_memory above: nothing in it is analysed for a build whose platform +/// has no unix sockets, because the word is not registered there at all. +const fs9_client = @import("fs9_client.zig"); /// The platform's runtime-setting facilities, stated once as plain data. /// Registry generation, leader paths, Config, and EffectCode all consume this @@ -1065,3 +1070,38 @@ pub const Gpio = struct { c.p.reportError(c.id, "gpio", err); } }; + +// ---- somebody else's tree ---- + +/// `9p ` — walk to a file in ANOTHER pardes's tree, read it, and +/// open the bytes in a pane. +/// +/// THE OTHER END OF `--fs9`, and the reason the client in `src/9p.zig` is not a +/// library with no caller: one pardes serves acme's control filesystem over +/// 9P2000 on a unix socket, and this word is the second one reading it. `9p +/// work /1/body` shows you what pane 1 of the session called `work` is holding, +/// from a pane in this session, with no mount and no `plan9port` in the way. +/// +/// A DIAL IS A NAME OR A PATH: `work` resolves through the same +/// `fs9_service.socketPath` that bound it, and anything with a `/` in it is a +/// socket path taken as given. Unix sockets only for now — a 9P server across a +/// network is tunnelled (docs/9p.typ §10), and this word is not the place to +/// decide otherwise. +/// +/// It BLOCKS while it fetches, bounded by `fs9_client.budget_ms`, exactly the +/// way Look blocks on a disk read; `src/fs9_client.zig` argues that at length +/// and enforces it with a deadline rather than a promise. +pub const @"9p" = struct { + pub const takes_arg = true; + pub const enabled = fs9_client.supported; + /// Prose and not a list: the bytes are a file's, so n/N walks its words the + /// way it walks any document's, and there is nothing here to step to. + pub const output: OutputTraits = .{ .name = config.ninep_buffer, .doc = true }; + pub fn run(c: Ctx) void { + if (comptime enabled) apply(c) else unreachable; + } + fn apply(c: Ctx) void { + fs9_client.fetch(c.p, c.id, c.arg orelse "") catch |err| + c.p.reportError(c.id, "9p", err); + } +}; diff --git a/src/config.zig b/src/config.zig index 8d186435..cbe75dea 100644 --- a/src/config.zig +++ b/src/config.zig @@ -257,6 +257,10 @@ pub const leader_path = paths: { table.set(.Hexdump, null); table.set(.Gpio, null); } + // The 9P client word. Takes a dial AND a path, so it has no leader path + // for the reason the three above have none, twice over. Its gate is the + // presence of unix sockets, which is narrower than `hosted`. + if (builtins.@"9p".enabled) table.set(.@"9p", null); // The pane-local PDF commands exist only in MuPDF builds through their // explicit registry availability, so name their paths inside the same // comptime branch. PdfTint/PdfFit retain their display slots and @@ -816,6 +820,10 @@ pub const pdf_sections_buffer = "+PdfSections"; pub const hover_buffer = "+Hover"; pub const lsp_buffer = "+Lsp"; pub const changelog_buffer = "+Changelog"; +/// What `9p ` opens a remote file into. NOT the remote path: an +/// output buffer's name comes off the command that filled it, and the path is +/// the command's ARGUMENT, which is what makes two remote files two panes. +pub const ninep_buffer = "+9p"; /// The two memory windows a bare-metal build's Peek and Hexdump render. Absent /// from every hosted build along with the builtins that name them. pub const peek_buffer = "+Peek"; diff --git a/src/detached/server.zig b/src/detached/server.zig index 2d7a5238..7d286177 100644 --- a/src/detached/server.zig +++ b/src/detached/server.zig @@ -1056,6 +1056,9 @@ pub const Session = struct { // `Twrite` to `ctl` can `push_spawn`, and `dispatch` is mid-iteration // over a descriptor snapshot when it runs. if (s.ninep) |l| { + // Before the drain, so a slot held by silence is taken back on the + // same frame it expires rather than one drain later. + l.expire(); var pending = false; for (0..fs9_service.max_conns) |i| { const t = l.transport(@intCast(i)) orelse continue; @@ -1431,7 +1434,11 @@ pub const Session = struct { // a pty's — asking for it unconditionally makes every idle socket a // ready descriptor and turns the poll into a spin. if (s.ninep) |l| { - if (l.fd >= 0) { + // `accepting`, not `fd >= 0`: a listener paused after an EMFILE + // must leave the set, or the backlog it could not drain reports + // ready on every poll and spins the core. `nextDue` carries the + // moment it comes back. + if (l.accepting()) { fds[n] = .{ .fd = l.fd, .events = poll_in, .revents = 0 }; src[n] = .ninep_listener; n += 1; @@ -1600,6 +1607,14 @@ pub const Session = struct { } if (s.listener >= 0 and s.accept_paused_ms > now) due = if (due) |d| @min(d, s.accept_paused_ms) else s.accept_paused_ms; + // ...and the 9P listener's own greet deadline, for the reason its + // `expire` gives: four slots is a cheaper denial than thirty-two. + if (s.ninep) |l| { + if (l.nextDue()) |ms| { + const at9 = now + ms; + due = if (due) |d| @min(d, at9) else at9; + } + } const at = due orelse return null; return @intCast(@max(0, @min(at - now, std.math.maxInt(c_int)))); } diff --git a/src/esp32p4/uart.zig b/src/esp32p4/uart.zig index 0afd145b..53ee29df 100644 --- a/src/esp32p4/uart.zig +++ b/src/esp32p4/uart.zig @@ -30,12 +30,20 @@ //! package's `src/hal/uart.zig:52`). At 115200 the wire costs ~86 us per byte and dwarfs either //! version, so today this is merely free - and it stops being free the moment the divider is raised. //! -//! **UART0's configuration is never touched.** Not the divider, not the format, not the pad +//! **UART0's configuration is never touched, except by `setBaud`.** Not the format, not the pad //! routing, and above all not `reset()`. The second-stage bootloader configured this block, and //! `src/hal/uart.zig:195-211` records what happens if it is reset: UART_CLKDIV returns to its //! power-on value, the console turns to garbage mid-sentence, and the board takes a watchdog reset -//! with nothing readable left to explain it. Everything here touches FIFO offset 0x000 and the +//! with nothing readable left to explain it. Everything else here touches FIFO offset 0x000 and the //! status register, and nothing else. +//! +//! The DIVIDER is the one exception, and it was carved out for the second image rather than for this +//! one: `docs/registry.typ` `BOARD-1`. The editor's console is opened by a human at 115200 and the +//! firmware inherits that divider (which is why the paragraph above used to say "not the divider"); +//! the 9P image (`nine.zig`) has a program on the far end that opens the port at whatever rate the +//! image was built for, and eight times the baud is eight times less latency on every `Tread`. So +//! `setBaud` exists, this file still never calls it, and `app.zig` still never calls it — the only +//! caller is the image whose host side is opened to match. const hal = @import("hal"); const input_rescue = @import("input_rescue.zig"); @@ -71,6 +79,44 @@ pub fn write(bytes: []const u8) void { /// lying about what happened, so it is worth printing. pub var dropped: u32 = 0; +/// Push as much of `bytes` as the transmitter has room for RIGHT NOW, and answer how much. Never +/// spins, never drops, never rescues — because it never waits for anything. +/// +/// THIS IS THE SANS-IO WRITE, and it exists because `write` above is the wrong primitive for a +/// protocol server. `write` is right for a console: a frame of ANSI is atomic to the terminal on the +/// far end, half an escape sequence leaves it in the wrong colour, so blocking until the whole burst +/// is in the FIFO is real backpressure and worth the stall. A 9P reply is not atomic to anything: +/// every message carries its own length, the reader on the far end reassembles, and +/// `9p.Server.wrote(n)` exists precisely so that a partial write costs nothing but another trip +/// round the loop (`src/9p.zig:2248-2257`). So this hands over what fits and returns, and +/// `src/esp32p4_9p.zig`'s loop keeps the remainder queued in the server where it already was. +/// +/// The difference that matters is DROPPING. `write`'s bounded spin gives up after a million status +/// reads and counts the loss, which turns a stalled transmitter into a diagnosable console; the same +/// behaviour on a 9P stream would truncate a reply mid-message and desynchronise the connection for +/// good. A short count cannot desynchronise anything. +pub fn writeSome(bytes: []const u8) usize { + const n = @min(@as(usize, uart0.txFree()), bytes.len); + for (bytes[0..n]) |b| uart0.pushByte(b); + return n; +} + +/// Reprogram UART0's divider, and answer whether the rate is reachable from the clock this block is +/// running on. Nothing is touched when it is not. +/// +/// THE ONE EXCEPTION to this file's rule, and see the header for who may call it: not this file, not +/// `app.zig`, only an image whose host side is opened at the same rate. `docs/registry.typ` +/// `BOARD-1` has the numbers — 921600 is one `UART_CLKDIV_SYNC` write on the existing 40 MHz XTAL, +/// int 43 frag 6, +0.064% error, and it takes a byte from 86.8 µs to 10.85 µs. +/// +/// The hazard is worth restating where the call is: the bytes already in the FIFO go out at the OLD +/// rate, so anything written before this and not yet drained is corrupted, and a host that is not +/// reopened sees garbage from here on with no error to report. The 9P image calls this before it +/// answers its first message, which is the one moment when neither of those can have happened yet. +pub fn setBaud(baud: u32) bool { + return uart0.setBaudrate(baud, uart0.clockSource().nominalHz()); +} + /// One byte, for callers that must not touch `.rodata` to say anything - which during bring-up is /// the difference between a diagnostic and a second copy of the bug being diagnosed. pub fn writeByte(b: u8) void { diff --git a/src/esp32p4_9p.zig b/src/esp32p4_9p.zig new file mode 100644 index 00000000..dd0346ef --- /dev/null +++ b/src/esp32p4_9p.zig @@ -0,0 +1,454 @@ +//! THE BOARD AS A 9P SERVER, and nothing else: the reset entry, one UART, and a pump. +//! +//! This is the SECOND ESP32-P4 image and it is not a second role for the first one. `app.zig` is +//! the editor — a real `pardes.Pardes` core with vaxis on top, emitting ANSI down UART0 to a +//! terminal emulator on the far end. This image links none of that. Same board, same UART, same +//! flash partition, one at a time, because the editor owns UART0 bidirectionally and JP1 exposes no +//! second P4 UART (`docs/registry.typ` `9P-11`: "The board is either an editor or a filesystem at +//! any one time. Say that plainly rather than implying both"). +//! +//! THE UART CARRIES ONLY 9P. That is the whole difference from the other image and it is the point. +//! No ANSI, no vaxis, no escape sequences, no `MARK` boot markers, no `soc.rom.print` — not even +//! the heap report `app.zig:320-329` prints on every boot, which would be the single most useful +//! line here and is still not allowed, because a byte on this wire that is not part of a 9P message +//! is a byte that desynchronises whatever is parsing it. The proof that this image booted is that it +//! answers `Tversion`. +//! +//! The one thing that had to be said in some other language is a PANIC and a TRAP, and they are said +//! in 9P too: an `Rerror` carrying the message, tagged `NOTAG`. No client is waiting for that tag, +//! so `9p` reports it as an unexpected reply and prints the string — which is exactly the diagnosis +//! wanted ("the board died, here is why") delivered without putting one non-protocol byte on the +//! wire. See `panicImpl` and `trapReport`. +//! +//! ## What it serves +//! +//! `src/board9p.zig`, which is the board's own capabilities as a tree: `gpio/pinout` is the JP1 +//! drawing the editor's `Gpio` word prints, and `gpio//value` is one pad's driven level, readable +//! and writable. Both come out of a comptime table, and adding a capability to that table adds files +//! here with no code in this file changing at all. +//! +//! Deliberately NOT `src/acmefs.zig`, and the reason is the same one that makes this a second image. +//! That file is the EDITOR's control filesystem: every operation in it is about a pane, and a pane +//! only exists because a `pardes.Pardes` exists. Serving it would mean linking the editor object +//! (809,536 B of image) and instantiating the core, at which point this is `app.zig` with a +//! different output encoding rather than a 9P server. It compiles for this target — `llvm-nm` finds +//! 21,548 B of `acmefs.*` in `zig-out/pardes-esp32p4.o` — and that fact is what made this image +//! worth building, because it is what proved the filesystem layer has no host dependency. The ABI is +//! what got reused, not the tree: `src/9p.zig`'s `Server` is a generic over the filesystem, and +//! `board9p` implements `acmefs`'s `Op`/`Status`/`Req`/`Reply` verbatim, so the same server serves +//! either one and neither knows about the other. +//! +//! ## The loop +//! +//! Four lines, and every one of them is a `Server` method doing what its doc comment says: +//! +//! read bytes off the UART -> srv.push(bytes) +//! pump -> srv.retry() / srv.next() -> fsys.handle(req) -> srv.reply(...) +//! write what is queued -> srv.wrote(uart.writeSome(srv.output())) +//! +//! NOTHING BLOCKS. `uart.read` is non-blocking, `uart.writeSome` hands over what the transmit FIFO +//! has room for and answers how much, and `Server.wrote(n)` takes a partial write as an ordinary +//! answer rather than an error (`src/9p.zig:2248-2257`). So a client that stops reading cannot stall +//! this loop, and a reply larger than the 128-byte FIFO leaves over several trips round it. That is +//! the same sans-io contract `src/fs9_service.zig` gives the desktop's unix socket; the difference +//! is that there is no `poll` here and no need for one, because there is exactly one connection and +//! it is the wire. +//! +//! ## The numbers, measured rather than costed +//! +//! `.bss` IS THE WHOLE RAM BILL, because this image has no allocator: not a heap, not an arena, and +//! the 384 KiB span the editor's image hands `heapmod` is not even mapped by anything here. So the +//! board's ≈336 KB of free heap (`docs/registry.typ` `FIX-2`) is untouched at 100%, and what this +//! program spends is the 240 KiB of low L2MEM that `9P-11`'s built note names as the real binding +//! constraint. `llvm-size` on the ELF says `.bss` is 16,656 B, and every byte of it is accounted +//! for: +//! +//! 9,192 `srv` — `Server(Tree(Pads))` on riscv32. `9P-11` measured 9,488 on the +//! host; a 32-bit target's slices are half the width, and the park +//! table has thirty-two of them. +//! 1,024 `in_buf` — one msize +//! 2,048 `out_buf` — two, so no reply can fail to be queued +//! 4,108 the rescue ring — `input_rescue.Ring` inside `uart.zig`, which comes with the UART +//! 148 `fsys` — the whole tree: one 143-byte answer buffer and a counter +//! 128 `stage` +//! ------ +//! 16,648 + 8 of alignment and `uart.dropped` = 16,656 +//! +//! Add the 32,768-byte `.stack` the shared linker script gives every image built through +//! `firmware()` and the low-L2MEM total is 49,424 B, 20% of the 240 KiB — against the editor's +//! 75,236 B (20,408 `.data` + 22,060 `.bss` + the same stack). The stack is the largest single item +//! and it is inherited rather than chosen: 32 KiB is sized for the CORE's recursive layout pass +//! (`build.zig:1088-1090`), and nothing in this image recurses at all. +//! +//! FLASH: the image is 88,080 B of the 1,536,000 B partition — 5.7%, against the editor image's +//! 812,688 B (52.9%). Only 28,066 B of that is content (22,504 `.flash.text`, 5,562 B of real +//! `.flash.rodata`, 80 B of image header and checksum); the rest is the gap between the end of the +//! rodata segment and the 64 KiB-aligned origin the code segment must start on, because the ESP32 +//! flash MMU maps in 64 KiB pages and the two segments cannot share one. A tiny image pays up to +//! 64 KiB for that and there is nothing to be done about it here — it is the generated linker +//! script's arithmetic (`05-zig-p4/build.zig`), and it is why the estimate of "≈39 KiB" in `9P-11` +//! was closer to the CONTENT than to the image. +//! +//! ## Build it, flash it, talk to it +//! +//! zig build -Dplatform=esp32p4 -Desp32p4-firmware -Desp32p4-9p esp32p4-9p-flash +//! zig build -Dplatform=esp32p4 -Desp32p4-firmware -Desp32p4-9p esp32p4-9p-size # no board needed +//! +//! There is no `esp32p4-9p-attach`, and that absence is the design: what belongs on the far end of +//! this wire is a 9P client opened at `baud`, not a terminal. `9p` and `9pfuse` speak to a SOCKET, +//! so reaching this board with either means a program that copies bytes between the tty and a unix +//! socket in both directions — which is nine lines of anything and is not this file's business. +//! pardes's own client (`src/fs9_client.zig`) needs no such bridge, because a tty is already a +//! bidirectional byte stream and that is all 9P has ever asked for (`docs/registry.typ` `9P-19`). +//! +//! FLASHING THIS REPLACES THE EDITOR. Both images are written to `img.opts`'s one offset, on +//! purpose: there is one partition and the board is one thing at a time. `zig build esp32p4-flash` +//! puts the editor back. + +const std = @import("std"); +const soc = @import("soc"); +const hal = @import("hal"); +const config = @import("config"); +// PATH imports, not named modules, and that is what lets any builder root an +// image here: the toolchain repository links this file with the four platform +// modules it owns (`soc`, `hal`, `config`, `heap`) and nothing else, so a +// `@import("ninep")` here was a module only pardes's own build.zig knew to +// inject — and the image stopped building the moment that build.zig stopped +// linking it. See `src/board9p.zig`'s note on the same change. +const ninep = @import("9p.zig"); +const board9p = @import("board9p.zig"); +const uart = @import("esp32p4/uart.zig"); + +/// THE PADS, and this is the whole seam between the tree and the silicon. +/// +/// The same four `hal.gpio` calls `src/esp32p4/app.zig:200-209` makes for the editor's `Gpio` word, +/// for the reason that file gives at length: a toggle is not a write to GPIO_OUT. `configureOutput` +/// points the pad's IO MUX at the GPIO function, routes the GPIO matrix's output to it, sets the +/// drive strength and input buffer, clears the pulls and only then enables the driver — four register +/// files indexed by a per-pin table, which live in the toolchain package where `zig build diff` +/// checks their numbers against ESP-IDF's own headers. A second copy would be a second copy under no +/// test. This is a second CALLER, which is the opposite thing. +/// +/// `getDrivenLevel` and not `getLevel`: the answer is the level this board is DRIVING, which is +/// defined for every pin including one with nothing attached, where the pad's own level is whatever +/// the air says. `readback = true` enables the input buffer anyway, so a client that wants the pad +/// rather than the register has something to compare against. +/// +/// SPLIT INTO `level` AND `drive` rather than the editor's single `toggle`, because a file can say +/// which level it wants and a keystroke cannot. `Gpio 20` has one argument and has to mean "the +/// other one"; `echo 1 > gpio/20/value` says 1, which is what makes it idempotent and therefore +/// scriptable. Writing the level a pad is already at still calls `configureOutput`, and that is not +/// a wasted write: on a freshly booted board it is the call that makes the pad an output at all. +/// +/// BOTH ARE `pub` AND HAVE TO BE, for the same reason `src/esp32p4/selftest.zig:44-46` says its +/// `FakePort`'s methods are: `board9p` is a MODULE here, and duck typing across a module boundary +/// still needs the declaration to be visible from outside the file it is in. Nothing else in this +/// image is `pub`. +const Pads = struct { + pub fn level(pin: u8) u1 { + return hal.gpio.getDrivenLevel(pin); + } + + pub fn drive(pin: u8, want: u1) void { + hal.gpio.configureOutput(pin, .{ .readback = true }); + if (want == 1) hal.gpio.setHigh(pin) else hal.gpio.setLow(pin); + } +}; + +comptime { + // Every pin the tree generates has to be a pad this chip package has, and the check belongs here + // rather than in `board9p.zig`: `max_pin` is 56 on this package and lives in the toolchain + // repository, which a host-testable tree cannot import. A JP1 row edited to name GPIO 60 is a + // compile error in this image instead of an out-of-bounds register index on the die. + for (board9p.pins) |pin| { + if (pin > hal.gpio.max_pin) @compileError("JP1 names a pad this chip package does not have"); + } +} + +/// The board's tree, over the real pads. +const Fs = board9p.Tree(Pads); +const Server = ninep.Server(Fs); + +/// THE msize, and it is 1,024 rather than the 4,096 everything else in this tree assumes. +/// +/// The 4,096 floor is the LINUX KERNEL's and nobody else's: `linux/net/9p/client.c:840-843` refuses +/// to mount below it, which is why `9p.min_msize` is 4,096 and why the desktop daemon serves that. +/// Plan 9's devmnt, plan9port's `9p` and pardes's own client all accept 512 +/// (`docs/registry.typ` `9P-11`), and no Linux kernel is ever going to mount this image: the far end +/// of this wire is a serial port, and a `mount -t 9p` needs a socket or a virtio channel, neither of +/// which a CH340 is. So the floor that applies here is `9p.msize_min` — 217 bytes, DERIVED from the +/// largest reply whose size the client does not choose (`src/9p.zig:1873-1881`). +/// +/// 1,024 and not 217, because the number to size against is the widest DIRECTORY READ. `gpio/` has +/// twelve entries, a `stat` record in a directory read is 49 bytes of fixed fields plus the name plus +/// three copies of the client's `uname` (`src/9p.zig:3251-3260`), so a `goblin` reading `ls gpio/` +/// wants 12 × ~73 = ~880 bytes to get the listing in ONE round trip. At 217 it would take five, and +/// each one costs a `Tread` and an `Rread` on a wire. Everything else here is tiny: the largest file +/// in the tree is the 468-byte JP1 drawing and the largest write is two bytes. +/// +/// What it costs: `in` is one msize and `out` is two — one message going out and one being built, +/// which is what makes every reply in the server infallible — so 3,072 B for the buffers against +/// 12,288 B at a 4,096 msize. Nine kilobytes of the board's low L2MEM for a round trip nobody needs. +const msize: u32 = 1024; + +/// One whole T-message, and the ceiling on the msize this connection will agree to. +var in_buf: [msize]u8 = undefined; + +/// Two, for the reason above. `Server.hasRoom` reserves one msize before it hands any request to the +/// filesystem, which is what makes back-pressure land on `next()` returning null instead of on a +/// half-written reply. +var out_buf: [2 * msize]u8 = undefined; + +/// Bytes off the receiver on their way into the server, and the ONE buffer in this file. +/// +/// 128 is the transmit and receive FIFO depth (the toolchain package's `src/hal/uart.zig:52`), so one +/// `uart.read` can never leave more behind than one FIFO's worth, and the tail that `push` would not +/// take is re-offered next time round the loop. It is not a reassembly buffer — `Server.in` is that, +/// and it holds a whole message — it is the handover between a driver that fills a slice and a server +/// that takes what it has room for. +var stage: [128]u8 = undefined; + +/// The wire's rate, and the host must be opened to match or nothing works and nothing says so. +/// +/// 921600 rather than the 115200 the bootloader leaves behind: `docs/registry.typ` `BOARD-1`. One +/// `UART_CLKDIV_SYNC` write on the existing 40 MHz XTAL, int 43 frag 6, +0.064% error, and it takes +/// a byte from 86.8 µs to 10.85 µs — which on this loop is a warm `cat gpio/20/value` going from +/// 10.8 ms to 1.35 ms and a 1 KiB `Tread` from 89 ms to 11 ms. 2 Mbaud is representable and this +/// CH340 is unreliable there, corroborated by the flasher's own choice at `build.zig:1136-1138`. +/// +/// It is programmed before the first reply and after the input drain, which is the one moment when +/// there can be nothing in either FIFO to be corrupted by the change. +const baud: u32 = 921600; + +/// The server and the tree, both in `.bss` and both fixed for the life of the image. No allocator +/// exists in this program at all — not a heap, not an arena, not the `heapmod` the editor's image +/// hands over 384 KiB to — so `zig build esp32p4-9p-size` reporting `.bss` is reporting the whole +/// of what this server costs in RAM. +var srv: Server = undefined; +var fsys: Fs = .{}; + +export fn zig_main() noreturn { + // FIRST, before anything reads `.rodata`, exactly as `app.zig:275` does it and for the same + // reason: the JP1 drawing this image serves is 468 bytes of `.rodata` in flash, and a read of it + // through a stale cache returns whatever was there at reset. + soc.flushFlashCache(); + + // The same clock the editor's image runs at, so a latency measured on one is a latency on the + // other. A divider change that disturbs neither UART0 (XTAL) nor the flash interface (SPLL). + if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) { + 180 => .mhz180, + 360 => .mhz360, + else => .mhz90, + }); + + // The RTC watchdog is armed at reset and this loop never feeds anything. Without this the board + // resets a few seconds in, which over a wire that carries only 9P looks exactly like a client + // that cannot reach it. + _ = hal.rwdt.disable(); + + // WHAT THE BOOTLOADER LEFT ON THE WIRE, discarded before the divider changes: its own chatter + // has already been echoed at the host, and the host bridge injects a synthetic window-size + // report before this program exists. Neither is 9P, and either would be the first bytes of a + // message that never was. + _ = uart.drainInput(); + + // The rate, then. A refusal is not fatal and must not be: an unreachable divider leaves 115200 + // in place, which is a slow board rather than a silent one, and a client opened at the wrong rate + // finds out immediately because `Tversion` gets no answer it can parse. + _ = uart.setBaud(baud); + + srv = Server.init(.{ .in = &in_buf, .out = &out_buf, .root = board9p.root }); + + // THE PUMP. `stage_len` is the only state outside the server. + var stage_len: usize = 0; + while (true) { + // IN. Non-blocking, rescued bytes first (`uart.read`), and never more than the staging + // buffer's room, so a burst larger than one FIFO simply arrives over two iterations. + if (stage_len < stage.len) stage_len += uart.read(stage[stage_len..]); + if (stage_len != 0) { + // A SHORT PUSH IS NORMAL AND IS NOT A LOSS: it is the only back-pressure a sans-io + // server has (`src/9p.zig:2229-2233`). What it would not take stays here and is offered + // again after the pump has made room by finishing a message. + const took = srv.push(stage[0..stage_len]); + if (took != stage_len) std.mem.copyForwards(u8, stage[0 .. stage_len - took], stage[took..stage_len]); + stage_len -= took; + } + + // PUMP, in the order `src/fs_service.zig:209-222` requires: every parked request offered + // once, then everything the wire has, both loops to null. + // + // NOTHING ON THIS BOARD PARKS — the answer to "what level is this pad" is a register read, + // and there is no `event` file and no reader to block — so `retry()` answers null on the + // first call, every time. It is here because the contract is the contract, and because the + // first capability that does block (an interrupt-driven `gpio//edge`) needs this line to + // already exist rather than to be remembered. + while (srv.retry()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + while (srv.next()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + + // OUT. Whatever fits in the transmitter right now, and the server keeps the rest. + const queued = srv.output(); + if (queued.len != 0) srv.wrote(uart.writeSome(queued)); + + // THE STREAM WAS NOT 9P, and there is no resynchronising from that: a `size` no encoder + // could have produced, an R-message from something that thought it was the server, a + // message larger than the negotiated msize. On a socket the answer is to close the + // connection and let the client notice; on a wire that cannot be closed, the answer is to + // reset it — pay the filesystem whatever `release`s the dead fids owe it, throw away every + // byte in flight in both directions, and start a fresh connection in the same silence a + // reboot would have. A client resynchronises by sending `Tversion`, which is what a client + // does after any failure anyway. + if (srv.dead) { + srv.hangup(); + while (srv.next()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + _ = uart.drainInput(); + stage_len = 0; + srv = Server.init(.{ .in = &in_buf, .out = &out_buf, .root = board9p.root }); + } + } +} + +// --------------------------------------------------------------------------- dying in protocol + +/// A message this image is about to die with, as an `Rerror` on `NOTAG`. +/// +/// THE ONE PLACE A NON-REPLY IS SENT, and it is still a legal 9P message, which is the whole trick. +/// `NOTAG` is the tag of the `Tversion` exchange and no client has a request outstanding under it, so +/// `9p` and pardes's own client both report an unexpected reply AND PRINT THE STRING — "the board +/// panicked at 0x4000a1b8", delivered through a parser rather than past it. The alternative is what +/// the editor's image does, `MARK PARDES_PANIC` in plain text, which on this wire would be a frame +/// header of 0x4b52414d followed by garbage: an unrecoverable stream instead of a diagnosis. +/// +/// Blocking `uart.write` and not `writeSome`, because there is no loop left to come back round: this +/// is the last thing the image does, and a bounded spin that gets the whole message out is worth +/// more here than one that returns. +fn die(msg: []const u8) noreturn { + var buf: [ninep.errmax + ninep.header_len + 2]u8 = undefined; + const bytes = ninep.encode( + .{ .rerror = .{ .ename = msg[0..@min(msg.len, ninep.errmax)] } }, + ninep.notag, + &buf, + ) catch unreachable; + uart.write(bytes); + while (true) {} +} + +/// Eight hex digits into `buf`, computed arithmetically. Hand-rolled rather than `std.fmt`, for the +/// reason `uart.dumpWord` gives: this runs in a trap handler, where the less of the image it depends +/// on the more likely it is to run at all. +fn hex8(buf: *[8]u8, v: u32) void { + var shift: u5 = 28; + for (buf) |*slot| { + const nib: u8 = @intCast((v >> shift) & 0xf); + slot.* = if (nib < 10) '0' + nib else 'a' + (nib - 10); + shift -%= 4; + } +} + +/// `mtvec` is set in DIRECT mode by `_start`, so every trap and every interrupt lands here. +/// +/// A trap handler exists for the reason `app.zig:432-441` gives — the mask ROM's "Guru Meditation" +/// only prints while ITS handler is installed, and a silent fault over a serial line is +/// indistinguishable from an infinite loop — and it reports through 9P for the reason `die` gives. +export fn trapEntry() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ("j trapReport"); +} + +export fn trapReport() noreturn { + const mcause = asm volatile ("csrr %[o], mcause" + : [o] "=r" (-> u32), + ); + const mepc = asm volatile ("csrr %[o], mepc" + : [o] "=r" (-> u32), + ); + const mtval = asm volatile ("csrr %[o], mtval" + : [o] "=r" (-> u32), + ); + // The three registers that name a RISC-V fault, in the order a reader wants them: what happened, + // where, and to which address. + var msg = "trap mcause=00000000 mepc=00000000 mtval=00000000".*; + hex8(msg[12..20], mcause); + hex8(msg[26..34], mepc); + hex8(msg[41..49], mtval); + die(&msg); +} + +// --------------------------------------------------------------- the root's own duties +// +// This is a ROOT, so it owns std's configuration for this compilation unit. The editor's image has +// two of these (`app.zig` and `src/esp32p4.zig`, one per object); this image is one object and has +// one. + +/// `page_size_min`/`max`: no MMU and no pages here, but std derives alignment from them, and 4 KiB +/// is this chip's cache and DMA granularity. +/// +/// `logFn` is not cosmetic and it is not optional. std's default log implementation reaches +/// `std.debug_io`, which instantiates `std.Io.Threaded` — a thread pool, `getrandom`, `IOV_MAX`, +/// `mremap` — and one `log.warn` anywhere in the graph drags all of it into the image. This one +/// DISCARDS, which is the only honest thing it can do: there is nowhere for a log line to go on a +/// wire that carries only 9P, and a log line that went out anyway would break the connection it was +/// trying to explain. Nothing in this image's graph logs; this is the wall that keeps it that way. +pub const std_options: std.Options = .{ + .page_size_min = 4096, + .page_size_max = 4096, + .logFn = logFn, +}; + +fn logFn( + comptime _: std.log.Level, + comptime _: @EnumLiteral(), + comptime _: []const u8, + _: anytype, +) void {} + +pub const panic = std.debug.FullPanic(panicImpl); + +fn panicImpl(msg: []const u8, first_trace_addr: ?usize) noreturn { + // The address is what makes it actionable — `addr2line` against the ELF in zig-out turns it into + // a source line — so it goes in front of the message, where `errmax`'s 128-byte truncation + // cannot reach it. A panic message names a KIND of failure; the address names which one. + var buf: [ninep.errmax]u8 = undefined; + @memcpy(buf[0..7], "panic 0"); + buf[7] = 'x'; + hex8(buf[8..16], @truncate(first_trace_addr orelse 0)); + buf[16] = ' '; + const n = @min(msg.len, buf.len - 17); + @memcpy(buf[17..][0..n], msg[0..n]); + die(buf[0 .. 17 + n]); +} + +/// Reset entry, identical in shape to `app.zig:528-546` and for the identical reasons: the bootloader +/// hands over with an unspecified stack pointer and the FPU off, so enable the F extension +/// (`mstatus.FS`), establish a stack, install the trap vector, clear `.bss`, and jump into Zig. +/// +/// `.bss` MATTERS MORE HERE THAN ANYWHERE. Everything this image owns is in it — the server, its two +/// buffers, the tree, the staging buffer — so this loop is what makes the fid table empty and the +/// msize zero, and skipping it would start the server mid-connection with a client that does not +/// exist. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, trapEntry + \\ csrw mtvec, t0 + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} diff --git a/src/fs9_client.zig b/src/fs9_client.zig new file mode 100644 index 00000000..ad46dee4 --- /dev/null +++ b/src/fs9_client.zig @@ -0,0 +1,695 @@ +//! `9p `: one pardes reading a file out of another pardes's tree. +//! +//! `src/fs9_service.zig`'s MIRROR, and the other half of `9P-2`: that file is a +//! listener with connections and hands each one a `ninep.Server`, this one +//! dials a single socket and drives a `ninep.Client` over it. They share the +//! socket NAMING and nothing else — `socketPath` is imported verbatim, so a +//! bare `9p work /1/body` resolves to exactly the path a `pardes --fs9 work` +//! bound, which is the whole point of having one spelling of it. +//! +//! WHAT IS HERE, and it is the same four things any non-blocking byte stream +//! needs: connect, read into `push`, `output` out through `send` and back +//! through `wrote`, and close. Everything above that is `src/9p.zig`, which is +//! freestanding and knows about neither sockets nor panes. +//! +//! IT BLOCKS, BRIEFLY AND BOUNDED, and that is a decision rather than an +//! oversight. `src/look.zig`'s `readFile` already blocks the frame on a disk +//! read — opening a file pane is a person waiting for a file — and a remote +//! read over a unix socket on the same machine is the same wait with a context +//! switch in it. What makes it safe to say that is the BUDGET: `budget_ms` is +//! the deadline for the whole transaction, `poll` is what waits, and the +//! descriptor is non-blocking, so a peer that stops answering costs one +//! `budget_ms` pause and a message on the message row rather than a wedged +//! editor. The alternative — a request queued into the frame loop, a state +//! machine per outstanding fetch, a pane that fills in later — is an async +//! runtime, and `docs/9p.typ` §12.5 is explicit that this design does not get +//! one. +//! +//! WHY THE HOST AND NOT THE CORE. A socket is `std.c`, and `src/9p.zig` must +//! keep compiling for `wasm32-freestanding` and the board's +//! `riscv32-freestanding`; the same `ninep.Client` runs over +//! `src/esp32p4/uart.zig` with no line of this file involved. So the split is +//! the one the server half already made: protocol in the freestanding file, +//! descriptor here. +//! +//! Linux and darwin, like every other unix socket in the tree. Anywhere else +//! `supported` is false and the word is not registered at all. +const std = @import("std"); +const libc = std.c; +const nested = @import("nested.zig"); +const ninep = @import("9p.zig"); +const fs9_service = @import("fs9_service.zig"); +const pardes = @import("pardes.zig"); +const Pardes = pardes.Pardes; +const output_pane = @import("output_pane.zig"); + +/// Unix sockets, which is all this needs — `fs9_service`'s own predicate, so a +/// build that can serve 9P can dial it and one that cannot has neither. +pub const supported = fs9_service.supported; + +/// `sun_path`, from the kernel's struct. See `nested.sun_path_len`. +const sun_path_len = nested.sun_path_len; + +/// The deadline for the WHOLE transaction: connect, handshake, attach, walk, +/// open, every read, clunk. +/// +/// TWO SECONDS, and the number is about the human rather than about the wire. +/// On a local socket the whole exchange is six round trips and some memcpys — +/// microseconds — so any wait long enough to notice means the far end is not +/// answering, and the useful thing to do about that is say so. Two seconds is +/// long enough that a busy editor on the other side finishing its frame is +/// never mistaken for a dead one, and short enough that a mistyped socket name +/// on a path that happens to exist does not feel like a hang. +pub const budget_ms: i64 = 2000; + +/// The most bytes one `9p` will carry into a pane. +/// +/// A MEGABYTE, which is a quarter of `look.zig`'s cap for a virtual file +/// (`read_stream_max_bytes`) and for a sharper reason: what is on the other end +/// is a synthetic tree of live editor state, where the biggest file is one +/// pane's `body`. A megabyte of it is a large source file; ten megabytes is +/// somebody pointing this word at a `/dev/zero` equivalent, and a read loop +/// with no cap would spend the whole budget filling the heap. +pub const max_bytes: u64 = 1 << 20; + +/// The deepest path this word will walk. +/// +/// `MAXWELEM` is sixteen elements per `Twalk` and a deeper path is legal — the +/// client splits it into chunks and `transact` does — so this is not a protocol +/// bound. It is a bound on the ARGUMENT: acme's tree is two deep (`/1/body`), +/// two chunks is thirty-two, and a path with more elements than that is a typo +/// or a loop rather than a file. Refused rather than truncated, because a +/// truncated path names a different file. +pub const max_depth: usize = 2 * ninep.max_welem; + +/// What the far end reports in `Rstat`'s three name fields (`uid`, `gid`, +/// `muid`) for everything we touch, because `ninep.Server` records the +/// attach's `uname` and quotes it back. +/// +/// A CONSTANT AND NOT `$USER`: the socket's permissions are the identity here +/// (0600, in a 0700 per-user directory — docs/9p.typ §10), so this string is +/// not a credential and cannot become one. What it is for is the operator +/// reading `ls -l` on the far side, and "pardes" tells them which program +/// walked their tree, which a login name they already share with it does not. +const uname = "pardes"; + +/// Everything this word can refuse, and each one is a different thing to do +/// about it. +pub const Error = error{ + MissingDial, + MissingPath, + /// More than `max_depth` elements. + PathTooDeep, + /// The dial names no address we can form: an empty name, a name with a + /// separator or a NUL in it, a path past `sun_path`, or no runtime + /// directory to resolve a bare name against. + BadDial, + /// `socket(2)` or `connect(2)` said no: nothing is listening on that + /// socket, or its permissions are not ours. The overwhelmingly common + /// case, and it means "that pardes is not running with `--fs9`". + Dial, + /// The peer closed mid-transaction. + Hangup, + /// `budget_ms` elapsed. See there. + Timeout, + /// The stream stopped being 9P: `ninep.Client` went dead, or a reply + /// arrived whose shape does not answer the request it was tagged for. + /// Nothing can be resynchronised from here. + Botch, + /// The far end answered `Rerror`. Its own string goes on the message row — + /// see `fetch` — so this value only says "reported already". + Remote, + /// The path names a directory. A directory READ is a run of `stat` + /// records rather than text, so opening it in a pane would show a person + /// the wire format; `9p` names files. + IsDirectory, + /// The walk stopped short: some element of the path is not there. Distinct + /// from `Remote` because a partial walk is a SUCCESSFUL `Rwalk` with fewer + /// qids and carries no message to report. + NotFound, + /// Past `max_bytes`. + FileTooLarge, + /// The pane that asked went away while this was in flight. + MissingPane, +}; + +/// The far end's own words, copied out of the client's input buffer before the +/// connection is torn down and the buffer with it. `ninep.errmax` is the buffer +/// a Plan 9 client has for an error string, so it is the right size for one. +const RemoteError = struct { + buf: [ninep.errmax]u8 = undefined, + len: usize = 0, + + /// Returns the error so that every call site is `return remote.set(e)`. + fn set(r: *RemoteError, msg: []const u8) error{Remote} { + r.len = @min(msg.len, r.buf.len); + @memcpy(r.buf[0..r.len], msg[0..r.len]); + return error.Remote; + } + + fn text(r: *const RemoteError) []const u8 { + return r.buf[0..r.len]; + } +}; + +/// `9p ` — walk to a remote file, read it, and open the bytes in a +/// pane. +/// +/// The pane is an ORDINARY OUTPUT BUFFER, which is what `src/board_memory.zig` +/// puts a hexdump in and what `Grep` puts its rows in: a file pane with an +/// `output` origin, so every motion, chord, search and Look works on it for +/// free. Deliberately NOT a real file pane, even though the bytes came from +/// `look.readFile`'s own kind of read: `file_pane.open` arms a file WATCH on +/// the path it was given, and there is no local path here for `inotify` to +/// watch — the bytes live in another process's memory. An output buffer is the +/// existing answer to "text with no file behind it". +/// +/// Identified by the WHOLE argument, so `9p work /1/body` and `9p work /2/body` +/// are two panes and running either again refills its own. +pub fn fetch(p: *Pardes, id: usize, argument: []const u8) !void { + const a = std.mem.trim(u8, argument, " \t\r\n"); + const args = try parse(a); + var names: [max_depth][]const u8 = undefined; + const n = try elements(args.path, &names); + + var sock_buf: [sun_path_len]u8 = undefined; + const sock = resolve(&sock_buf, args.dial) orelse return Error.BadDial; + + var remote: RemoteError = .{}; + const content = fetchBytes(p.gpa, sock, names[0..n], &remote) catch |err| { + if (err != Error.Remote) return err; + // The far end's own wording, which is the whole error ABI in base + // 9P2000 (docs/registry.typ `9P-4`) — reported verbatim rather than + // mapped to one of ours, because it is the only thing that says which + // of the eight operations the other side objected to and why. + var buf: [ninep.errmax + 8]u8 = undefined; + p.setMessage(id, std.fmt.bufPrint(&buf, "9p: {s}", .{remote.text()}) catch "9p: refused"); + return; + }; + const pane = p.panes[id] orelse { + p.gpa.free(content); + return Error.MissingPane; + }; + const dir = if (pane.file) |f| (std.fs.path.dirname(f.path) orelse "/") else pane.cwdSlice(); + try output_pane.fillResults(p, id, dir, .{ .cmd = .@"9p" }, a, content, null); +} + +const Args = struct { dial: []const u8, path: []const u8 }; + +/// ` `, split at the FIRST run of whitespace and not tokenized. +/// +/// The path keeps its spaces, because a pane's name in acme's tree can have +/// them and a path is the last argument: `9p work /1/tag` and +/// `9p work /a name/body` both have exactly one reading. The dial cannot have +/// them, and does not need to — it is a socket name or a socket path. +fn parse(a: []const u8) Error!Args { + if (a.len == 0) return Error.MissingDial; + const cut = std.mem.indexOfAny(u8, a, " \t") orelse return Error.MissingPath; + const path = std.mem.trim(u8, a[cut..], " \t\r\n"); + if (path.len == 0) return Error.MissingPath; + return .{ .dial = a[0..cut], .path = path }; +} + +/// A path into `Twalk` elements. Separators are collapsed and a trailing one is +/// dropped, so `/1/body`, `1/body` and `//1/body/` are one file — the +/// normalisation every shell already does, done here because 9P has no +/// pathnames at all and a client that forwarded an empty element would be +/// asking for a file called "". +/// +/// ZERO ELEMENTS is the root, which is legal and is a directory; `transact` +/// refuses it there, where every other directory is refused too. +fn elements(path: []const u8, out: *[max_depth][]const u8) Error!usize { + var n: usize = 0; + var it = std.mem.tokenizeScalar(u8, path, '/'); + while (it.next()) |name| { + if (n == out.len) return Error.PathTooDeep; + out[n] = name; + n += 1; + } + return n; +} + +/// The dial, as an address. +/// +/// TWO SPELLINGS, told apart by a separator, and the distinction is the one a +/// person already makes: a NAME is what `pardes --fs9 work` was started with, +/// and it resolves through `fs9_service.socketPath` — the same function that +/// bound it, so the two can never drift. A PATH is taken as given, which is +/// what you need for a socket somewhere else entirely: a bind-mounted +/// container, a different user's runtime directory, an `ssh -L` forward. +fn resolve(buf: *[sun_path_len]u8, dial: []const u8) ?[:0]const u8 { + if (dial.len == 0) return null; + if (std.mem.indexOfScalar(u8, dial, '/') != null) { + if (std.mem.indexOfScalar(u8, dial, 0) != null) return null; + return std.fmt.bufPrintSentinel(buf, "{s}", .{dial}, 0) catch null; + } + var dir_buf: [sun_path_len:0]u8 = undefined; + const dir = nested.socketDir(&dir_buf) orelse return null; + return fs9_service.socketPath(buf, dir, dial); +} + +/// One dialled connection: the descriptor, the deadline, the client and its +/// three buffers. +/// +/// HEAP-ALLOCATED by `fetchBytes`, for `fs9_service.Listener`'s reason and one +/// more: `cl.in` and `cl.out` are slices INTO this struct, so it must never be +/// moved once `cl` is initialised, and at three msizes it is 24 KiB, which does +/// not belong on the frame's stack. +/// +/// The msize is `fs9_service.msize`, the one number the serving side is already +/// sized from. A client on the same machine reading the same tree has no reason +/// to pick a different one, and picking the same one means the handshake never +/// clamps. +const Session = struct { + fd: c_int, + /// `nowMs()` past which every wait gives up. + deadline: i64, + cl: ninep.Client = undefined, + in: [fs9_service.msize]u8 = undefined, + out: [fs9_service.msize]u8 = undefined, + /// A frame-local staging buffer rather than a read straight into the + /// client's tail: advancing `in_len` is `push`'s business, and reaching + /// past it to do it here would make this file a second author of + /// `9p.zig`'s invariants for the sake of one memcpy per 8 KiB. Exactly + /// `fs9_service.fill`'s reasoning, from the other side. + stage: [fs9_service.msize]u8 = undefined, + + /// Wait for `events` on the descriptor, or give up. THE ONLY PLACE THIS + /// FILE BLOCKS, and the only place the budget is spent. + fn wait(s: *Session, events: i16) Error!void { + while (true) { + const left = s.deadline - nowMs(); + if (left <= 0) return Error.Timeout; + var fds = [1]libc.pollfd{.{ .fd = s.fd, .events = events, .revents = 0 }}; + const ready = libc.poll(&fds, 1, @intCast(@min(left, budget_ms))); + if (ready < 0) { + if (libc.errno(ready) == .INTR) continue; + return Error.Hangup; + } + if (ready == 0) return Error.Timeout; + // What we asked for wins over HUP: a peer that wrote a reply and + // then closed reports both at once, and those bytes are ours. + if (fds[0].revents & events != 0) return; + return Error.Hangup; + } + } + + /// Push everything the client owes the wire, and nothing else. Split out of + /// `settle` so that `dropNoWait` can send a message it will never collect a + /// reply for. Bounded by the same deadline `wait` enforces. + fn flush(s: *Session) Error!void { + while (s.cl.output().len != 0) { + try s.wait(poll_out); + const bytes = s.cl.output(); + const sent = libc.send(s.fd, bytes.ptr, bytes.len, nosignal); + if (sent < 0) switch (libc.errno(sent)) { + .INTR, .AGAIN => continue, + else => return Error.Hangup, + }; + // No progress and no error: looping on it is a spin, and a + // spin in here is the editor at 100% of a core. + if (sent == 0) return Error.Hangup; + s.cl.wrote(@intCast(sent)); + } + } + + /// Drive the client until the one outstanding request answers: flush what + /// we owe, collect if a reply is already buffered, otherwise wait and read. + /// + /// LOCK-STEP, deliberately, and it is worth saying why given that + /// `ninep.Client` allows sixteen requests in flight. A `9p` word is one + /// person waiting for one file, and its round trips are strictly ordered + /// anyway — you cannot read a fid you have not opened, or open one you have + /// not walked to. The one place pipelining would pay is the read loop, and + /// on a local socket at an 8 KiB msize a megabyte is 128 round trips of a + /// few microseconds each; buying that back would cost this file a request + /// window, an out-of-order reassembly buffer and a reason for both. The + /// CLIENT is where the sixteen tags live, so the board's runtime and any + /// future caller get them without this file having spent them. + fn settle(s: *Session) Error!ninep.Client.Done { + while (true) { + try s.flush(); + if (s.cl.take()) |done| return done; + if (s.cl.dead) return Error.Botch; + try s.wait(poll_in); + const room = s.cl.in.len - s.cl.in_len; + // Cannot happen: one reply is at most one msize and the buffer is + // exactly that, so a full buffer with nothing to take would mean + // the far end sent a frame it told us it would not. + if (room == 0) return Error.Botch; + const got = libc.read(s.fd, &s.stage, @min(room, s.stage.len)); + if (got == 0) return Error.Hangup; + if (got < 0) switch (libc.errno(got)) { + .INTR, .AGAIN => continue, + else => return Error.Hangup, + }; + const n = s.cl.push(s.stage[0..@intCast(got)]); + // The read was clamped to the room, so this cannot be short; it is + // asserted rather than ignored because silently dropping wire + // bytes desynchronises the stream, which is the one failure 9P + // cannot resynchronise from. + std.debug.assert(n == @as(usize, @intCast(got))); + } + } + + /// One request, one reply, and the two answers that are not the one asked + /// for folded into errors here so that `transact` reads as a script. + fn ask(s: *Session, req: ninep.Client.Request, remote: *RemoteError) Error!ninep.Client.Result { + _ = s.cl.submit(req) catch return Error.Botch; + const done = try s.settle(); + if (done.result == .fail) return remote.set(done.result.fail); + // The client already refuses a reply whose shape does not match the + // request's op (it kills the connection), so this can only be an + // `Rerror` we have just handled. Checked anyway: a `switch` here would + // be a second copy of that table. + if (std.mem.eql(u8, @tagName(done.result), @tagName(std.meta.activeTag(req)))) return done.result; + return Error.Botch; + } + + /// Clunk a fid and WAIT for the answer, because the caller is about to + /// reuse the number. `transact`'s walk alternates between fids 1 and 2, and + /// in-order processing is the only thing that makes that safe. + /// + /// Best effort otherwise: a refused clunk still frees the fid on both sides + /// (`clunk(5)`), so there is nothing here worth failing a fetch over. + fn drop(s: *Session, fid: u32) void { + _ = s.cl.submit(.{ .clunk = .{ .fid = fid } }) catch return; + _ = s.settle() catch {}; + } + + /// Clunk a fid and do NOT wait. For the last one, where the descriptor is + /// closed on the next line and closing it frees every fid the connection + /// held — so the `Rclunk` is not merely unwanted, it is unobservable. + /// + /// Waiting for it cost the whole budget against a peer that answers + /// everything else and ignores clunks: measured at 2.005 s to deliver a + /// file that was already in hand, and 2.003 s to report an error decided + /// 1.4 ms in. The bytes still go out — a well-behaved peer gets its clunk + /// and frees the fid immediately rather than at hangup — but nothing here + /// reads the reply. + fn dropNoWait(s: *Session, fid: u32) void { + _ = s.cl.submit(.{ .clunk = .{ .fid = fid } }) catch return; + s.flush() catch {}; + } +}; + +/// The whole transaction, and the only function here that knows 9P's order of +/// operations: version, attach, walk, open, read to the end, clunk. +fn transact( + s: *Session, + names: []const []const u8, + out: *std.Io.Writer.Allocating, + remote: *RemoteError, +) !void { + // The handshake. A server that answers "unknown" has no dialect in common + // with us and leaves `msize` at zero, which is a connection nothing can be + // submitted on — reported as a botch, because there is no fallback ladder + // here to climb down. + _ = try s.ask(.{ .version = .{} }, remote); + if (s.cl.msize == 0) return Error.Botch; + + // Fid 0 is the root for the life of the connection; 1 and 2 alternate as + // the walk descends, so a chunked walk never needs a third. + const root: u32 = 0; + var here = (try s.ask(.{ .attach = .{ .fid = root, .uname = uname } }, remote)).attach; + var cur: u32 = root; + var next: u32 = 1; + + var i: usize = 0; + while (i < names.len) { + // `MAXWELEM` elements at a time, which is what makes a path deeper than + // sixteen work at all: every implementation refuses a seventeenth + // element, so a deep path is several walks with an intermediate fid. + const n = @min(ninep.max_welem, names.len - i); + const w = (try s.ask(.{ .walk = .{ + .fid = cur, + .newfid = next, + .names = names[i..][0..n], + } }, remote)).walk; + // A PARTIAL WALK IS A SUCCESS with fewer qids, and only a failure on + // the first element is an `Rerror`. So this comparison is the whole of + // "did the path exist", and skipping it is how a client ends up + // reading the wrong file. + if (w.nwqid != n) return Error.NotFound; + here = w.wqid[n - 1]; + if (cur != root) s.drop(cur); + cur = next; + next = if (next == 1) 2 else 1; + i += n; + } + defer s.dropNoWait(cur); + + // The qid the walk landed on already says what this is, so the refusal + // costs no round trip — and it catches the bare `/` too, which is zero + // elements and the root. + if (here.type & ninep.qtdir != 0) return Error.IsDirectory; + + _ = try s.ask(.{ .open = .{ .fid = cur, .mode = ninep.oread } }, remote); + + // Read to the end. 9P has no EOF flag: a reply SHORTER than the count is + // ordinary and means nothing, and a reply of ZERO bytes is the end of the + // file (`read(5)`). The offset advances by what came back and never by what + // was asked, which is the same rule a POSIX read loop follows. + var off: u64 = 0; + while (true) { + if (off >= max_bytes) return Error.FileTooLarge; + const want: u32 = @intCast(@min(@as(u64, s.cl.maxRead()), max_bytes - off)); + const data = (try s.ask(.{ .read = .{ .fid = cur, .offset = off, .count = want } }, remote)).read; + if (data.len == 0) return; + try out.writer.writeAll(data); + off += data.len; + } +} + +/// Dial, transact, and hand back the bytes — gpa-owned, the way +/// `look.readFile`'s are, so the pane adopts them with no second copy. +fn fetchBytes( + gpa: std.mem.Allocator, + sock: [:0]const u8, + names: []const []const u8, + remote: *RemoteError, +) ![]u8 { + if (comptime !supported) return Error.Dial; + // The deadline starts BEFORE the dial, because the dial is part of the + // transaction and used not to be bounded by anything at all. See `connect`. + const deadline = nowMs() +| budget_ms; + const fd = try connect(sock, deadline); + const s = gpa.create(Session) catch return error.OutOfMemory; + defer { + _ = libc.close(fd); + gpa.destroy(s); + } + s.* = .{ .fd = fd, .deadline = deadline }; + s.cl = .init(.{ .in = &s.in, .out = &s.out }); + + var out: std.Io.Writer.Allocating = .init(gpa); + errdefer out.deinit(); + try transact(s, names, &out, remote); + return out.toOwnedSlice(); +} + +/// Connect to a unix socket, non-blocking from the first moment there is +/// anything to wait for — which is the connect itself. +/// +/// This used to leave the connect BLOCKING, on the argument that a unix socket +/// either completes at once or refuses at once. That is true only while the +/// listener's accept queue has room. When it is full, Linux's +/// `unix_stream_connect` waits in `unix_wait_for_peer` for `sk_sndtimeo`, which +/// defaults to MAX_SCHEDULE_TIMEOUT — forever. The core is single-threaded, so +/// that is the whole editor: no frame, no keystroke, no filesystem request +/// served. Measured at 177 seconds against a peer that had called `listen` and +/// never `accept`, and it ended only because the peer was killed. Nothing in +/// pardes would have ended it, and the trigger needs no hostility — a peer that +/// is itself wedged does it, and a bare path names sockets pardes does not own. +/// +/// So the descriptor is non-blocking before the connect and the wait is spent +/// against the caller's deadline. Both refusals have to be handled and they are +/// different: on AF_UNIX a full backlog is EAGAIN, NOT the EINPROGRESS a TCP +/// connect would give, so EAGAIN retries until the deadline and EINPROGRESS +/// waits for POLLOUT and then asks SO_ERROR what actually happened. +fn connect(sock: [:0]const u8, deadline: i64) Error!c_int { + if (sock.len + 1 > sun_path_len) return Error.BadDial; + var addr: libc.sockaddr.un = .{ .path = @splat(0) }; + @memcpy(addr.path[0 .. sock.len + 1], sock[0 .. sock.len + 1]); + const fd = libc.socket(libc.AF.UNIX, libc.SOCK.STREAM, 0); + if (fd < 0) return Error.Dial; + nested.setCloexec(fd); + setNonblock(fd); + errdefer _ = libc.close(fd); + while (true) { + if (libc.connect(fd, @ptrCast(&addr), @sizeOf(@TypeOf(addr))) == 0) break; + switch (libc._errno().*) { + // The backlog is full. Nobody is obliged to drain it, so this is a + // poll on the clock rather than on the descriptor: there is no + // event to wait for, only room that may or may not appear. + @intFromEnum(libc.E.AGAIN), @intFromEnum(libc.E.INTR) => { + if (nowMs() >= deadline) return Error.Dial; + nap(2); + }, + // Someone is listening and the connect is under way. This one IS a + // descriptor event, so wait for it and then ask what it was. + @intFromEnum(libc.E.INPROGRESS), @intFromEnum(libc.E.ALREADY) => { + const left = deadline - nowMs(); + if (left <= 0) return Error.Dial; + var pfd: [1]libc.pollfd = .{.{ .fd = fd, .events = poll_out, .revents = 0 }}; + if (libc.poll(&pfd, 1, @intCast(@min(left, 1000))) <= 0) continue; + var err: c_int = 0; + var len: libc.socklen_t = @sizeOf(c_int); + if (libc.getsockopt(fd, libc.SOL.SOCKET, libc.SO.ERROR, @ptrCast(&err), &len) != 0) + return Error.Dial; + if (err == 0) break; + return Error.Dial; + }, + // Already connected by a previous round of this loop. + @intFromEnum(libc.E.ISCONN) => break, + else => return Error.Dial, + } + } + if (comptime nested.darwin) { + // linux says MSG_NOSIGNAL per write, darwin once per socket. A peer + // that dies mid-transaction must not take the editor down with it. + const on: c_int = 1; + _ = libc.setsockopt(fd, libc.SOL.SOCKET, libc.SO.NOSIGPIPE, &on, @sizeOf(c_int)); + } + return fd; +} + +/// The descriptor is non-blocking and `poll` does the waiting, because that is +/// the only shape in which the budget above is enforceable: a blocking `read` +/// has no deadline to give it. `fs9_service`'s own `setNonblock` is not reused +/// for its stated reason — importing a daemon into a path the tty and GUI +/// shells take would make a frontend transport a dependency of a builtin. +fn setNonblock(fd: c_int) void { + const flags = libc.fcntl(fd, libc.F.GETFL, @as(c_int, 0)); + if (flags < 0) return; + var o: libc.O = @bitCast(@as(u32, @bitCast(flags))); + o.NONBLOCK = true; + _ = libc.fcntl(fd, libc.F.SETFL, @as(c_int, @bitCast(@as(u32, @bitCast(o))))); +} + +/// Milliseconds on the MONOTONIC clock, which is the only clock a deadline may +/// be measured against: the wall clock can be stepped, and an NTP correction +/// landing mid-fetch would turn a two-second budget into a hang or into an +/// instant timeout depending on which way it went. +/// +/// A clock that will not answer is reported as THE END OF TIME, so the budget +/// expires on the first wait rather than never — the saturating `+|` at the one +/// call site that adds to it is what makes that safe. +fn nowMs() i64 { + var ts: libc.timespec = undefined; + if (libc.clock_gettime(.MONOTONIC, &ts) != 0) return std.math.maxInt(i64); + return @as(i64, ts.sec) * std.time.ms_per_s + @divTrunc(ts.nsec, std.time.ns_per_ms); +} + +const poll_in: i16 = @intCast(libc.POLL.IN); +const poll_out: i16 = @intCast(libc.POLL.OUT); + +/// Sleep a couple of milliseconds while a full accept backlog drains. There is +/// no descriptor to wait on for that — the room either appears or the deadline +/// arrives — so this is the one place here that waits on the clock. `poll` with +/// no descriptors is the portable spelling and needs no `nanosleep` import. +fn nap(ms: c_int) void { + _ = libc.poll(&[0]libc.pollfd{}, 0, ms); +} + +/// A dead peer must never kill the editor. linux says it per write, darwin once +/// per socket (see `connect`). +const nosignal: u32 = if (nested.darwin) 0 else libc.MSG.NOSIGNAL; + +const testing = std.testing; + +test "the argument is a dial and then the rest of the line" { + const a = try parse("work /1/body"); + try testing.expectEqualStrings("work", a.dial); + try testing.expectEqualStrings("/1/body", a.path); + + // A pane's name in acme's tree may contain spaces, and the path is the last + // argument, so it keeps them. + const spaced = try parse("work /a name/body"); + try testing.expectEqualStrings("work", spaced.dial); + try testing.expectEqualStrings("/a name/body", spaced.path); + + // A socket path as the dial, which is the other spelling. + const p = try parse("/run/user/1000/pardes-9p-work.sock /index"); + try testing.expectEqualStrings("/run/user/1000/pardes-9p-work.sock", p.dial); + try testing.expectEqualStrings("/index", p.path); + + try testing.expectError(Error.MissingDial, parse("")); + try testing.expectError(Error.MissingPath, parse("work")); + try testing.expectError(Error.MissingPath, parse("work ")); +} + +test "a path becomes walk elements, normalised the way a shell would" { + var out: [max_depth][]const u8 = undefined; + try testing.expectEqual(@as(usize, 2), try elements("/1/body", &out)); + try testing.expectEqualStrings("1", out[0]); + try testing.expectEqualStrings("body", out[1]); + + // Leading, trailing and doubled separators are one file, not four. + try testing.expectEqual(@as(usize, 2), try elements("1/body", &out)); + try testing.expectEqual(@as(usize, 2), try elements("//1//body//", &out)); + try testing.expectEqual(@as(usize, 1), try elements("/index", &out)); + + // Zero elements is the root, which is legal here and refused as a + // directory where every other directory is. + try testing.expectEqual(@as(usize, 0), try elements("/", &out)); + + // Deeper than two full walks is a typo, and truncating it would name a + // different file. + var deep: [8 * max_depth]u8 = @splat('/'); + for (0..max_depth + 1) |i| deep[i * 2 + 1] = 'a'; + try testing.expectError(Error.PathTooDeep, elements(deep[0 .. (max_depth + 1) * 2], &out)); +} + +test "a bare dial resolves to the socket --fs9 binds, and a path is taken as given" { + if (comptime !supported) return error.SkipZigTest; + var buf: [sun_path_len]u8 = undefined; + + // The one spelling both halves share: this must be the same name + // `fs9_service.socketPath` produces, or the friendly form dials nothing. + const named = resolve(&buf, "work").?; + try testing.expect(std.mem.endsWith(u8, named, "/pardes-9p-work.sock")); + var expect: [sun_path_len]u8 = undefined; + var dir_buf: [sun_path_len:0]u8 = undefined; + const dir = nested.socketDir(&dir_buf).?; + try testing.expectEqualStrings(fs9_service.socketPath(&expect, dir, "work").?, named); + + // A separator makes it a path, verbatim. + const path = resolve(&buf, "/tmp/somewhere.sock").?; + try testing.expectEqualStrings("/tmp/somewhere.sock", path); + + // And the refusals: nothing to dial, and a NUL that would truncate the + // address into something else entirely. + try testing.expect(resolve(&buf, "") == null); + try testing.expect(resolve(&buf, "/tmp/a\x00b") == null); +} + +test "a dial with nothing listening is one error and not a wait" { + if (comptime !supported) return error.SkipZigTest; + // The overwhelmingly common failure — "that pardes is not running with + // --fs9" — and it must be immediate: `connect` on a unix socket with no + // listener is refused by the kernel with no timeout in it, which is why + // `budget_ms` is never spent here. + var names: [max_depth][]const u8 = undefined; + const n = try elements("/1/body", &names); + var remote: RemoteError = .{}; + const before = nowMs(); + try testing.expectError( + Error.Dial, + fetchBytes(testing.allocator, "/tmp/pardes-9p-no-such-socket.sock", names[0..n], &remote), + ); + try testing.expect(nowMs() - before < budget_ms); +} + +test "one fetch costs three msize buffers and nothing that grows" { + // The number this word adds to a session WHILE IT RUNS, and nothing after: + // the `Session` is freed before `fetch` returns and only the content + // survives, adopted by the pane. Heap rather than stack for the reason + // `Session` states. + try testing.expectEqual(@as(usize, fs9_service.msize), @as(usize, (Session{ .fd = -1, .deadline = 0 }).in.len)); + try testing.expect(@sizeOf(Session) <= 3 * fs9_service.msize + 256); + // The client's own state is a rounding error beside its buffers, which is + // the whole point of borrowing payloads out of `in` instead of copying + // them per tag. + try testing.expect(@sizeOf(ninep.Client) <= 256); +} diff --git a/src/fs9_service.zig b/src/fs9_service.zig index f02e179e..eebfe7bf 100644 --- a/src/fs9_service.zig +++ b/src/fs9_service.zig @@ -109,6 +109,14 @@ const Conn = struct { /// stays occupied, with no descriptor, until `next()` runs dry. See /// `Srv.hangup` and `drainAll`. draining: bool = false, + /// When this peer connected, on the monotonic clock, or 0 when the clock + /// is unavailable. Read by `expire`: a connection that has not sent + /// `Tversion` within `greet_deadline_ms` is holding a slot by silence, + /// which with only four of them is a cheaper denial than the frontend + /// socket's thirty-two. `Server.msize == 0` is the "has not versioned yet" + /// flag, and version(5) requires `Tversion` before any other message, so + /// there is no legitimate client this can catch. + accepted_ms: i64 = 0, /// Undefined until `accept` initialises it in place, which it may only do /// through a pointer to this exact storage. srv: Srv = undefined, @@ -144,11 +152,24 @@ const Conn = struct { } }; +/// How long the 9P listener stays out of the poll set after an `accept` that +/// failed for a reason that persists — EMFILE and ENFILE above all. The same +/// number and the same argument as the frontend listener's own pause: the +/// connection is still in the backlog, `poll` is level triggered, and coming +/// straight back spins the core until some unrelated descriptor is freed. +const accept_pause_ms: i64 = 100; + /// The listening socket and its connections. Heap-allocated because a `Conn` /// holds slices into itself (see there) and because at three buffers per /// connection this is ≈100 KiB, which does not belong in a host's struct. pub const Listener = struct { fd: c_int = -1, + /// Do not accept before this moment on the monotonic clock. Set when + /// `accept(2)` fails for a reason that leaves the connection in the backlog + /// — EMFILE and ENFILE — because a level-triggered poll then reports the + /// listener ready forever and coming straight back spins the core. Zero + /// means accepting normally. + paused_ms: i64 = 0, /// The bound path, kept so teardown unlinks exactly what was created — /// guarded on the fd, like nested.zig's and detached/server.zig's /// `unlisten`. @@ -172,7 +193,23 @@ pub const Listener = struct { if (comptime !supported) return; for (0..max_conns + 1) |_| { const fd = libc.accept(l.fd, null, null); - if (fd < 0) return; // EAGAIN is this loop's ordinary exit + if (fd < 0) switch (libc.errno(fd)) { + // The ordinary exit: nothing more is queued. + .AGAIN => return, + // Retry: a signal, or a peer that gave up between the poll and + // the accept. Neither says anything about our capacity. + .INTR, .CONNABORTED => continue, + // Out of descriptors. The connection STAYS in the backlog, so a + // level-triggered poll reports the listener ready again at once + // and coming straight back spins the core until something + // unrelated frees an fd — measured at 99.8% of one, sustained. + // The frontend listener one file over solves it the same way. + else => { + l.paused_ms = nowMs() +| accept_pause_ms; + log.warn("--fs9: accept failed; pausing the listener for {d} ms", .{accept_pause_ms}); + return; + }, + }; nested.setCloexec(fd); setNonblock(fd); if (comptime nested.darwin) { @@ -193,6 +230,7 @@ pub const Listener = struct { }; c.fd = fd; c.draining = false; + c.accepted_ms = nowMs(); c.srv = .init(.{ .in = &c.in, .out = &c.out, @@ -201,6 +239,70 @@ pub const Listener = struct { } } + /// How long a connection may hold a slot without saying `Tversion`. The + /// same five seconds and the same argument as the frontend socket's + /// `greet_deadline_ms` (`detached/server.zig`): a slot held by silence is + /// the same denial as a full queue, arrived at from the other end. Cheaper + /// here, because there are four slots rather than thirty-two and no + /// handshake to fake. + pub const greet_deadline_ms: i64 = 5000; + + /// Take back any slot whose peer connected and then said nothing. Called + /// once per frame beside the drain; the host folds `nextDue` into its poll + /// timeout so the deadline is kept on an otherwise idle session rather than + /// whenever some other descriptor happens to wake it. + pub fn expire(l: *Listener) void { + if (comptime !supported) return; + const now = nowMs(); + if (now == 0) return; // no clock; see `nowMs` + for (&l.conns, 0..) |*c, i| { + if (c.fd < 0 or c.srv.msize != 0) continue; + if (now - c.accepted_ms < greet_deadline_ms) continue; + log.debug("--fs9: slot {d} never sent Tversion; taking it back", .{i}); + l.drop(@intCast(i)); + } + } + + /// Is the listener worth polling this round? False while it is paused after + /// a persistent `accept` failure — leaving it in the set is exactly the + /// spin the pause exists to stop. + pub fn accepting(l: *const Listener) bool { + if (comptime !supported) return false; + if (l.fd < 0) return false; + if (l.paused_ms == 0) return true; + const now = nowMs(); + return now == 0 or now >= l.paused_ms; + } + + /// Milliseconds until the earliest greet deadline, or null when nothing is + /// waiting on the clock. Floored at zero so a deadline already past polls + /// once without blocking instead of blocking on a negative timeout. + pub fn nextDue(l: *const Listener) ?i32 { + if (comptime !supported) return null; + const now = nowMs(); + if (now == 0) return null; + var due: ?i64 = null; + // The pause is a clock deadline like the greet ones: without it here, + // an idle session would sleep through the moment the listener is + // allowed back and only notice on the next unrelated wake. + if (l.paused_ms > now) due = l.paused_ms; + for (&l.conns) |*c| { + if (c.fd < 0 or c.srv.msize != 0) continue; + const at = c.accepted_ms + greet_deadline_ms; + due = if (due) |d| @min(d, at) else at; + } + const at = due orelse return null; + return @intCast(@max(0, at - now)); + } + + /// The monotonic clock in milliseconds, or 0 when there is none — which + /// every caller reads as "no deadlines this round" rather than as a time. + fn nowMs() i64 { + var ts: libc.timespec = undefined; + if (libc.clock_gettime(.MONOTONIC, &ts) != 0) return 0; + return @as(i64, ts.sec) * std.time.ms_per_s + @divTrunc(ts.nsec, std.time.ns_per_ms); + } + /// Read one chunk off connection `i` and hand it to the server. /// /// ONE read per connection per round, which is detached/server.zig's @@ -215,6 +317,13 @@ pub const Listener = struct { pub fn fill(l: *Listener, i: u8) void { if (comptime !supported) return; const c = &l.conns[i]; + // FIRST, and before the room guard below, which is the trap: once + // `startFrame` gives up on the framing, `in_len` is stuck at `in.len` + // for good, so `room == 0` returns without reading, `poll` is level + // triggered, the descriptor reports ready again immediately, and the + // loop never sleeps. Measured at 99.7% of a core, sustained, reachable + // by any process with the uid in one `write(2)`. + if (c.srv.dead) return l.drop(i); const room = c.srv.in.len - c.srv.in_len; if (room == 0) return; // A frame-local staging buffer rather than a read straight into the @@ -229,10 +338,23 @@ pub const Listener = struct { else => l.drop(i), }; const n = c.srv.push(buf[0..@intCast(got)]); - // Cannot happen — the read was clamped to the room — and it is - // asserted rather than ignored because silently dropping wire bytes - // desynchronises the stream, which is the one failure 9P cannot - // resynchronise from. + // The stream stopped being 9P. `Server.startFrame` sets `dead` when the + // framing is unrecoverable — a `size[4]` of zero, or one larger than the + // input buffer — and `push` then takes NOTHING, for good, because there + // is nowhere to resynchronise to in a protocol whose only frame marker + // is the length you were just lied to about. + // + // This has to be checked before the assert below, and the assert is why: + // it used to fire, and firing meant `unreachable` on the daemon's own + // thread — every pane, every attached frontend and the FUSE mount gone, + // reached by any client that sends one bad length and then one more + // byte. The socket is 0600 in a 0700 directory, but the whole point of + // `--fs9` is that other programs dial it, so a buggy one is enough. + if (c.srv.dead) return l.drop(i); + // NOW it cannot happen: the read was clamped to the room and the only + // other refusal is the one handled above. Asserted rather than ignored + // because silently dropping wire bytes desynchronises the stream, which + // is the one failure 9P cannot resynchronise from. std.debug.assert(n == @as(usize, @intCast(got))); } diff --git a/src/main.zig b/src/main.zig index 66d25816..004d59cb 100644 --- a/src/main.zig +++ b/src/main.zig @@ -441,6 +441,13 @@ test { // because this file reaches pardes.zig (`Server(acmefs)`); src/9p.zig, the // half that does NOT, has one in build.zig. _ = @import("fs9_service.zig"); + // Its client twin, and it needs its name here for the same reason with the + // same measurement behind it: builtins.zig imports it for the `9p` word, + // and naming a file does not make Zig analyse the tests of what IT + // imports. The protocol half's tests are in src/9p.zig's own b.addTest; + // these are the host's — the argument split, the two dial spellings, and + // the immediate refusal when nothing is listening. + _ = @import("fs9_client.zig"); if (comptime pardes.platform == .tty) { _ = @import("tty/tty.zig"); // tty.zig calls the compositor only from its runtime loop, so merely -- cgit v1.3