diff options
Diffstat (limited to 'src/esp32p4_9p.zig')
| -rw-r--r-- | src/esp32p4_9p.zig | 454 |
1 files changed, 454 insertions, 0 deletions
diff --git a/src/esp32p4_9p.zig b/src/esp32p4_9p.zig new file mode 100644 index 00000000..dd0346ef --- /dev/null +++ b/src/esp32p4_9p.zig @@ -0,0 +1,454 @@ +//! THE BOARD AS A 9P SERVER, and nothing else: the reset entry, one UART, and a pump. +//! +//! This is the SECOND ESP32-P4 image and it is not a second role for the first one. `app.zig` is +//! the editor — a real `pardes.Pardes` core with vaxis on top, emitting ANSI down UART0 to a +//! terminal emulator on the far end. This image links none of that. Same board, same UART, same +//! flash partition, one at a time, because the editor owns UART0 bidirectionally and JP1 exposes no +//! second P4 UART (`docs/registry.typ` `9P-11`: "The board is either an editor or a filesystem at +//! any one time. Say that plainly rather than implying both"). +//! +//! THE UART CARRIES ONLY 9P. That is the whole difference from the other image and it is the point. +//! No ANSI, no vaxis, no escape sequences, no `MARK` boot markers, no `soc.rom.print` — not even +//! the heap report `app.zig:320-329` prints on every boot, which would be the single most useful +//! line here and is still not allowed, because a byte on this wire that is not part of a 9P message +//! is a byte that desynchronises whatever is parsing it. The proof that this image booted is that it +//! answers `Tversion`. +//! +//! The one thing that had to be said in some other language is a PANIC and a TRAP, and they are said +//! in 9P too: an `Rerror` carrying the message, tagged `NOTAG`. No client is waiting for that tag, +//! so `9p` reports it as an unexpected reply and prints the string — which is exactly the diagnosis +//! wanted ("the board died, here is why") delivered without putting one non-protocol byte on the +//! wire. See `panicImpl` and `trapReport`. +//! +//! ## What it serves +//! +//! `src/board9p.zig`, which is the board's own capabilities as a tree: `gpio/pinout` is the JP1 +//! drawing the editor's `Gpio` word prints, and `gpio/<n>/value` is one pad's driven level, readable +//! and writable. Both come out of a comptime table, and adding a capability to that table adds files +//! here with no code in this file changing at all. +//! +//! Deliberately NOT `src/acmefs.zig`, and the reason is the same one that makes this a second image. +//! That file is the EDITOR's control filesystem: every operation in it is about a pane, and a pane +//! only exists because a `pardes.Pardes` exists. Serving it would mean linking the editor object +//! (809,536 B of image) and instantiating the core, at which point this is `app.zig` with a +//! different output encoding rather than a 9P server. It compiles for this target — `llvm-nm` finds +//! 21,548 B of `acmefs.*` in `zig-out/pardes-esp32p4.o` — and that fact is what made this image +//! worth building, because it is what proved the filesystem layer has no host dependency. The ABI is +//! what got reused, not the tree: `src/9p.zig`'s `Server` is a generic over the filesystem, and +//! `board9p` implements `acmefs`'s `Op`/`Status`/`Req`/`Reply` verbatim, so the same server serves +//! either one and neither knows about the other. +//! +//! ## The loop +//! +//! Four lines, and every one of them is a `Server` method doing what its doc comment says: +//! +//! read bytes off the UART -> srv.push(bytes) +//! pump -> srv.retry() / srv.next() -> fsys.handle(req) -> srv.reply(...) +//! write what is queued -> srv.wrote(uart.writeSome(srv.output())) +//! +//! NOTHING BLOCKS. `uart.read` is non-blocking, `uart.writeSome` hands over what the transmit FIFO +//! has room for and answers how much, and `Server.wrote(n)` takes a partial write as an ordinary +//! answer rather than an error (`src/9p.zig:2248-2257`). So a client that stops reading cannot stall +//! this loop, and a reply larger than the 128-byte FIFO leaves over several trips round it. That is +//! the same sans-io contract `src/fs9_service.zig` gives the desktop's unix socket; the difference +//! is that there is no `poll` here and no need for one, because there is exactly one connection and +//! it is the wire. +//! +//! ## The numbers, measured rather than costed +//! +//! `.bss` IS THE WHOLE RAM BILL, because this image has no allocator: not a heap, not an arena, and +//! the 384 KiB span the editor's image hands `heapmod` is not even mapped by anything here. So the +//! board's ≈336 KB of free heap (`docs/registry.typ` `FIX-2`) is untouched at 100%, and what this +//! program spends is the 240 KiB of low L2MEM that `9P-11`'s built note names as the real binding +//! constraint. `llvm-size` on the ELF says `.bss` is 16,656 B, and every byte of it is accounted +//! for: +//! +//! 9,192 `srv` — `Server(Tree(Pads))` on riscv32. `9P-11` measured 9,488 on the +//! host; a 32-bit target's slices are half the width, and the park +//! table has thirty-two of them. +//! 1,024 `in_buf` — one msize +//! 2,048 `out_buf` — two, so no reply can fail to be queued +//! 4,108 the rescue ring — `input_rescue.Ring` inside `uart.zig`, which comes with the UART +//! 148 `fsys` — the whole tree: one 143-byte answer buffer and a counter +//! 128 `stage` +//! ------ +//! 16,648 + 8 of alignment and `uart.dropped` = 16,656 +//! +//! Add the 32,768-byte `.stack` the shared linker script gives every image built through +//! `firmware()` and the low-L2MEM total is 49,424 B, 20% of the 240 KiB — against the editor's +//! 75,236 B (20,408 `.data` + 22,060 `.bss` + the same stack). The stack is the largest single item +//! and it is inherited rather than chosen: 32 KiB is sized for the CORE's recursive layout pass +//! (`build.zig:1088-1090`), and nothing in this image recurses at all. +//! +//! FLASH: the image is 88,080 B of the 1,536,000 B partition — 5.7%, against the editor image's +//! 812,688 B (52.9%). Only 28,066 B of that is content (22,504 `.flash.text`, 5,562 B of real +//! `.flash.rodata`, 80 B of image header and checksum); the rest is the gap between the end of the +//! rodata segment and the 64 KiB-aligned origin the code segment must start on, because the ESP32 +//! flash MMU maps in 64 KiB pages and the two segments cannot share one. A tiny image pays up to +//! 64 KiB for that and there is nothing to be done about it here — it is the generated linker +//! script's arithmetic (`05-zig-p4/build.zig`), and it is why the estimate of "≈39 KiB" in `9P-11` +//! was closer to the CONTENT than to the image. +//! +//! ## Build it, flash it, talk to it +//! +//! zig build -Dplatform=esp32p4 -Desp32p4-firmware -Desp32p4-9p esp32p4-9p-flash +//! zig build -Dplatform=esp32p4 -Desp32p4-firmware -Desp32p4-9p esp32p4-9p-size # no board needed +//! +//! There is no `esp32p4-9p-attach`, and that absence is the design: what belongs on the far end of +//! this wire is a 9P client opened at `baud`, not a terminal. `9p` and `9pfuse` speak to a SOCKET, +//! so reaching this board with either means a program that copies bytes between the tty and a unix +//! socket in both directions — which is nine lines of anything and is not this file's business. +//! pardes's own client (`src/fs9_client.zig`) needs no such bridge, because a tty is already a +//! bidirectional byte stream and that is all 9P has ever asked for (`docs/registry.typ` `9P-19`). +//! +//! FLASHING THIS REPLACES THE EDITOR. Both images are written to `img.opts`'s one offset, on +//! purpose: there is one partition and the board is one thing at a time. `zig build esp32p4-flash` +//! puts the editor back. + +const std = @import("std"); +const soc = @import("soc"); +const hal = @import("hal"); +const config = @import("config"); +// PATH imports, not named modules, and that is what lets any builder root an +// image here: the toolchain repository links this file with the four platform +// modules it owns (`soc`, `hal`, `config`, `heap`) and nothing else, so a +// `@import("ninep")` here was a module only pardes's own build.zig knew to +// inject — and the image stopped building the moment that build.zig stopped +// linking it. See `src/board9p.zig`'s note on the same change. +const ninep = @import("9p.zig"); +const board9p = @import("board9p.zig"); +const uart = @import("esp32p4/uart.zig"); + +/// THE PADS, and this is the whole seam between the tree and the silicon. +/// +/// The same four `hal.gpio` calls `src/esp32p4/app.zig:200-209` makes for the editor's `Gpio` word, +/// for the reason that file gives at length: a toggle is not a write to GPIO_OUT. `configureOutput` +/// points the pad's IO MUX at the GPIO function, routes the GPIO matrix's output to it, sets the +/// drive strength and input buffer, clears the pulls and only then enables the driver — four register +/// files indexed by a per-pin table, which live in the toolchain package where `zig build diff` +/// checks their numbers against ESP-IDF's own headers. A second copy would be a second copy under no +/// test. This is a second CALLER, which is the opposite thing. +/// +/// `getDrivenLevel` and not `getLevel`: the answer is the level this board is DRIVING, which is +/// defined for every pin including one with nothing attached, where the pad's own level is whatever +/// the air says. `readback = true` enables the input buffer anyway, so a client that wants the pad +/// rather than the register has something to compare against. +/// +/// SPLIT INTO `level` AND `drive` rather than the editor's single `toggle`, because a file can say +/// which level it wants and a keystroke cannot. `Gpio 20` has one argument and has to mean "the +/// other one"; `echo 1 > gpio/20/value` says 1, which is what makes it idempotent and therefore +/// scriptable. Writing the level a pad is already at still calls `configureOutput`, and that is not +/// a wasted write: on a freshly booted board it is the call that makes the pad an output at all. +/// +/// BOTH ARE `pub` AND HAVE TO BE, for the same reason `src/esp32p4/selftest.zig:44-46` says its +/// `FakePort`'s methods are: `board9p` is a MODULE here, and duck typing across a module boundary +/// still needs the declaration to be visible from outside the file it is in. Nothing else in this +/// image is `pub`. +const Pads = struct { + pub fn level(pin: u8) u1 { + return hal.gpio.getDrivenLevel(pin); + } + + pub fn drive(pin: u8, want: u1) void { + hal.gpio.configureOutput(pin, .{ .readback = true }); + if (want == 1) hal.gpio.setHigh(pin) else hal.gpio.setLow(pin); + } +}; + +comptime { + // Every pin the tree generates has to be a pad this chip package has, and the check belongs here + // rather than in `board9p.zig`: `max_pin` is 56 on this package and lives in the toolchain + // repository, which a host-testable tree cannot import. A JP1 row edited to name GPIO 60 is a + // compile error in this image instead of an out-of-bounds register index on the die. + for (board9p.pins) |pin| { + if (pin > hal.gpio.max_pin) @compileError("JP1 names a pad this chip package does not have"); + } +} + +/// The board's tree, over the real pads. +const Fs = board9p.Tree(Pads); +const Server = ninep.Server(Fs); + +/// THE msize, and it is 1,024 rather than the 4,096 everything else in this tree assumes. +/// +/// The 4,096 floor is the LINUX KERNEL's and nobody else's: `linux/net/9p/client.c:840-843` refuses +/// to mount below it, which is why `9p.min_msize` is 4,096 and why the desktop daemon serves that. +/// Plan 9's devmnt, plan9port's `9p` and pardes's own client all accept 512 +/// (`docs/registry.typ` `9P-11`), and no Linux kernel is ever going to mount this image: the far end +/// of this wire is a serial port, and a `mount -t 9p` needs a socket or a virtio channel, neither of +/// which a CH340 is. So the floor that applies here is `9p.msize_min` — 217 bytes, DERIVED from the +/// largest reply whose size the client does not choose (`src/9p.zig:1873-1881`). +/// +/// 1,024 and not 217, because the number to size against is the widest DIRECTORY READ. `gpio/` has +/// twelve entries, a `stat` record in a directory read is 49 bytes of fixed fields plus the name plus +/// three copies of the client's `uname` (`src/9p.zig:3251-3260`), so a `goblin` reading `ls gpio/` +/// wants 12 × ~73 = ~880 bytes to get the listing in ONE round trip. At 217 it would take five, and +/// each one costs a `Tread` and an `Rread` on a wire. Everything else here is tiny: the largest file +/// in the tree is the 468-byte JP1 drawing and the largest write is two bytes. +/// +/// What it costs: `in` is one msize and `out` is two — one message going out and one being built, +/// which is what makes every reply in the server infallible — so 3,072 B for the buffers against +/// 12,288 B at a 4,096 msize. Nine kilobytes of the board's low L2MEM for a round trip nobody needs. +const msize: u32 = 1024; + +/// One whole T-message, and the ceiling on the msize this connection will agree to. +var in_buf: [msize]u8 = undefined; + +/// Two, for the reason above. `Server.hasRoom` reserves one msize before it hands any request to the +/// filesystem, which is what makes back-pressure land on `next()` returning null instead of on a +/// half-written reply. +var out_buf: [2 * msize]u8 = undefined; + +/// Bytes off the receiver on their way into the server, and the ONE buffer in this file. +/// +/// 128 is the transmit and receive FIFO depth (the toolchain package's `src/hal/uart.zig:52`), so one +/// `uart.read` can never leave more behind than one FIFO's worth, and the tail that `push` would not +/// take is re-offered next time round the loop. It is not a reassembly buffer — `Server.in` is that, +/// and it holds a whole message — it is the handover between a driver that fills a slice and a server +/// that takes what it has room for. +var stage: [128]u8 = undefined; + +/// The wire's rate, and the host must be opened to match or nothing works and nothing says so. +/// +/// 921600 rather than the 115200 the bootloader leaves behind: `docs/registry.typ` `BOARD-1`. One +/// `UART_CLKDIV_SYNC` write on the existing 40 MHz XTAL, int 43 frag 6, +0.064% error, and it takes +/// a byte from 86.8 µs to 10.85 µs — which on this loop is a warm `cat gpio/20/value` going from +/// 10.8 ms to 1.35 ms and a 1 KiB `Tread` from 89 ms to 11 ms. 2 Mbaud is representable and this +/// CH340 is unreliable there, corroborated by the flasher's own choice at `build.zig:1136-1138`. +/// +/// It is programmed before the first reply and after the input drain, which is the one moment when +/// there can be nothing in either FIFO to be corrupted by the change. +const baud: u32 = 921600; + +/// The server and the tree, both in `.bss` and both fixed for the life of the image. No allocator +/// exists in this program at all — not a heap, not an arena, not the `heapmod` the editor's image +/// hands over 384 KiB to — so `zig build esp32p4-9p-size` reporting `.bss` is reporting the whole +/// of what this server costs in RAM. +var srv: Server = undefined; +var fsys: Fs = .{}; + +export fn zig_main() noreturn { + // FIRST, before anything reads `.rodata`, exactly as `app.zig:275` does it and for the same + // reason: the JP1 drawing this image serves is 468 bytes of `.rodata` in flash, and a read of it + // through a stale cache returns whatever was there at reset. + soc.flushFlashCache(); + + // The same clock the editor's image runs at, so a latency measured on one is a latency on the + // other. A divider change that disturbs neither UART0 (XTAL) nor the flash interface (SPLL). + if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) { + 180 => .mhz180, + 360 => .mhz360, + else => .mhz90, + }); + + // The RTC watchdog is armed at reset and this loop never feeds anything. Without this the board + // resets a few seconds in, which over a wire that carries only 9P looks exactly like a client + // that cannot reach it. + _ = hal.rwdt.disable(); + + // WHAT THE BOOTLOADER LEFT ON THE WIRE, discarded before the divider changes: its own chatter + // has already been echoed at the host, and the host bridge injects a synthetic window-size + // report before this program exists. Neither is 9P, and either would be the first bytes of a + // message that never was. + _ = uart.drainInput(); + + // The rate, then. A refusal is not fatal and must not be: an unreachable divider leaves 115200 + // in place, which is a slow board rather than a silent one, and a client opened at the wrong rate + // finds out immediately because `Tversion` gets no answer it can parse. + _ = uart.setBaud(baud); + + srv = Server.init(.{ .in = &in_buf, .out = &out_buf, .root = board9p.root }); + + // THE PUMP. `stage_len` is the only state outside the server. + var stage_len: usize = 0; + while (true) { + // IN. Non-blocking, rescued bytes first (`uart.read`), and never more than the staging + // buffer's room, so a burst larger than one FIFO simply arrives over two iterations. + if (stage_len < stage.len) stage_len += uart.read(stage[stage_len..]); + if (stage_len != 0) { + // A SHORT PUSH IS NORMAL AND IS NOT A LOSS: it is the only back-pressure a sans-io + // server has (`src/9p.zig:2229-2233`). What it would not take stays here and is offered + // again after the pump has made room by finishing a message. + const took = srv.push(stage[0..stage_len]); + if (took != stage_len) std.mem.copyForwards(u8, stage[0 .. stage_len - took], stage[took..stage_len]); + stage_len -= took; + } + + // PUMP, in the order `src/fs_service.zig:209-222` requires: every parked request offered + // once, then everything the wire has, both loops to null. + // + // NOTHING ON THIS BOARD PARKS — the answer to "what level is this pad" is a register read, + // and there is no `event` file and no reader to block — so `retry()` answers null on the + // first call, every time. It is here because the contract is the contract, and because the + // first capability that does block (an interrupt-driven `gpio/<n>/edge`) needs this line to + // already exist rather than to be remembered. + while (srv.retry()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + while (srv.next()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + + // OUT. Whatever fits in the transmitter right now, and the server keeps the rest. + const queued = srv.output(); + if (queued.len != 0) srv.wrote(uart.writeSome(queued)); + + // THE STREAM WAS NOT 9P, and there is no resynchronising from that: a `size` no encoder + // could have produced, an R-message from something that thought it was the server, a + // message larger than the negotiated msize. On a socket the answer is to close the + // connection and let the client notice; on a wire that cannot be closed, the answer is to + // reset it — pay the filesystem whatever `release`s the dead fids owe it, throw away every + // byte in flight in both directions, and start a fresh connection in the same silence a + // reboot would have. A client resynchronises by sending `Tversion`, which is what a client + // does after any failure anyway. + if (srv.dead) { + srv.hangup(); + while (srv.next()) |req| { + const a = fsys.handle(req); + srv.reply(&a.reply, a.bytes); + } + _ = uart.drainInput(); + stage_len = 0; + srv = Server.init(.{ .in = &in_buf, .out = &out_buf, .root = board9p.root }); + } + } +} + +// --------------------------------------------------------------------------- dying in protocol + +/// A message this image is about to die with, as an `Rerror` on `NOTAG`. +/// +/// THE ONE PLACE A NON-REPLY IS SENT, and it is still a legal 9P message, which is the whole trick. +/// `NOTAG` is the tag of the `Tversion` exchange and no client has a request outstanding under it, so +/// `9p` and pardes's own client both report an unexpected reply AND PRINT THE STRING — "the board +/// panicked at 0x4000a1b8", delivered through a parser rather than past it. The alternative is what +/// the editor's image does, `MARK PARDES_PANIC` in plain text, which on this wire would be a frame +/// header of 0x4b52414d followed by garbage: an unrecoverable stream instead of a diagnosis. +/// +/// Blocking `uart.write` and not `writeSome`, because there is no loop left to come back round: this +/// is the last thing the image does, and a bounded spin that gets the whole message out is worth +/// more here than one that returns. +fn die(msg: []const u8) noreturn { + var buf: [ninep.errmax + ninep.header_len + 2]u8 = undefined; + const bytes = ninep.encode( + .{ .rerror = .{ .ename = msg[0..@min(msg.len, ninep.errmax)] } }, + ninep.notag, + &buf, + ) catch unreachable; + uart.write(bytes); + while (true) {} +} + +/// Eight hex digits into `buf`, computed arithmetically. Hand-rolled rather than `std.fmt`, for the +/// reason `uart.dumpWord` gives: this runs in a trap handler, where the less of the image it depends +/// on the more likely it is to run at all. +fn hex8(buf: *[8]u8, v: u32) void { + var shift: u5 = 28; + for (buf) |*slot| { + const nib: u8 = @intCast((v >> shift) & 0xf); + slot.* = if (nib < 10) '0' + nib else 'a' + (nib - 10); + shift -%= 4; + } +} + +/// `mtvec` is set in DIRECT mode by `_start`, so every trap and every interrupt lands here. +/// +/// A trap handler exists for the reason `app.zig:432-441` gives — the mask ROM's "Guru Meditation" +/// only prints while ITS handler is installed, and a silent fault over a serial line is +/// indistinguishable from an infinite loop — and it reports through 9P for the reason `die` gives. +export fn trapEntry() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ("j trapReport"); +} + +export fn trapReport() noreturn { + const mcause = asm volatile ("csrr %[o], mcause" + : [o] "=r" (-> u32), + ); + const mepc = asm volatile ("csrr %[o], mepc" + : [o] "=r" (-> u32), + ); + const mtval = asm volatile ("csrr %[o], mtval" + : [o] "=r" (-> u32), + ); + // The three registers that name a RISC-V fault, in the order a reader wants them: what happened, + // where, and to which address. + var msg = "trap mcause=00000000 mepc=00000000 mtval=00000000".*; + hex8(msg[12..20], mcause); + hex8(msg[26..34], mepc); + hex8(msg[41..49], mtval); + die(&msg); +} + +// --------------------------------------------------------------- the root's own duties +// +// This is a ROOT, so it owns std's configuration for this compilation unit. The editor's image has +// two of these (`app.zig` and `src/esp32p4.zig`, one per object); this image is one object and has +// one. + +/// `page_size_min`/`max`: no MMU and no pages here, but std derives alignment from them, and 4 KiB +/// is this chip's cache and DMA granularity. +/// +/// `logFn` is not cosmetic and it is not optional. std's default log implementation reaches +/// `std.debug_io`, which instantiates `std.Io.Threaded` — a thread pool, `getrandom`, `IOV_MAX`, +/// `mremap` — and one `log.warn` anywhere in the graph drags all of it into the image. This one +/// DISCARDS, which is the only honest thing it can do: there is nowhere for a log line to go on a +/// wire that carries only 9P, and a log line that went out anyway would break the connection it was +/// trying to explain. Nothing in this image's graph logs; this is the wall that keeps it that way. +pub const std_options: std.Options = .{ + .page_size_min = 4096, + .page_size_max = 4096, + .logFn = logFn, +}; + +fn logFn( + comptime _: std.log.Level, + comptime _: @EnumLiteral(), + comptime _: []const u8, + _: anytype, +) void {} + +pub const panic = std.debug.FullPanic(panicImpl); + +fn panicImpl(msg: []const u8, first_trace_addr: ?usize) noreturn { + // The address is what makes it actionable — `addr2line` against the ELF in zig-out turns it into + // a source line — so it goes in front of the message, where `errmax`'s 128-byte truncation + // cannot reach it. A panic message names a KIND of failure; the address names which one. + var buf: [ninep.errmax]u8 = undefined; + @memcpy(buf[0..7], "panic 0"); + buf[7] = 'x'; + hex8(buf[8..16], @truncate(first_trace_addr orelse 0)); + buf[16] = ' '; + const n = @min(msg.len, buf.len - 17); + @memcpy(buf[17..][0..n], msg[0..n]); + die(buf[0 .. 17 + n]); +} + +/// Reset entry, identical in shape to `app.zig:528-546` and for the identical reasons: the bootloader +/// hands over with an unspecified stack pointer and the FPU off, so enable the F extension +/// (`mstatus.FS`), establish a stack, install the trap vector, clear `.bss`, and jump into Zig. +/// +/// `.bss` MATTERS MORE HERE THAN ANYWHERE. Everything this image owns is in it — the server, its two +/// buffers, the tree, the staging buffer — so this loop is what makes the fid table empty and the +/// msize zero, and skipping it would start the server mid-connection with a client that does not +/// exist. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, trapEntry + \\ csrw mtvec, t0 + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} |
