diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-26 13:27:46 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-27 09:47:39 -0300 |
| commit | 11f380f6d7222f2cad93c2cdf13701ea1f903d47 (patch) | |
| tree | 803194ee5853a6b4cda93f90a95e28d1f02e69ae /src/esp32p4 | |
| parent | fbc194068687e49a8490c85c9f1257a2f2bb9079 (diff) | |
| download | pardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.tar.gz pardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.zip | |
One core behind N frontends, the board's own runner moved in, and every board cap on one screen
## The wire is the effect stream, not a new protocol
`pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns
a terminal and a socket and nothing else. N frontends on one core all look at the same screen —
`screen -x`, not N sessions.
The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per
message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the
core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is
a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin
command line, and a command line cannot carry a frame.
ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be
riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit
little-endian fixed width and no message is a blit of a native struct. A protocol that only works
between two builds of the same compiler would have thrown away the one frontend that motivated it.
## The board comes in; its toolchain stays out
`src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over-
serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves
into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so
`zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the
board from this repo's `build.zig`.
The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes
core over a serial line. What stayed is everything a second project would also want — the HAL, the
register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its
own and its `build()` early-returns when it is not the root package, so this costs the package
graph exactly zero packages and the editor's own builds nothing at all.
## limits.zig: nine forgettable places become one budget
Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions —
they are ONE decision, how much memory this build may spend, taken nine times where no reader could
see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against
what it is measured against, derived from two booleans.
The payoff is testability on a machine that is not the board: the caps are ordinary comptime values,
so a host build can be compiled against the board's numbers and the parking, eviction and clamping
paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART.
## A bare `zig build`
`zig build` with no arguments now builds the tty and GUI binaries and installs them into
`~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built
one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and
"use it" two different commands for no reason.
Diffstat (limited to 'src/esp32p4')
| -rw-r--r-- | src/esp32p4/app.zig | 546 | ||||
| -rw-r--r-- | src/esp32p4/input_rescue.zig | 271 | ||||
| -rw-r--r-- | src/esp32p4/selftest.zig | 351 | ||||
| -rw-r--r-- | src/esp32p4/uart.zig | 165 |
4 files changed, 1333 insertions, 0 deletions
diff --git a/src/esp32p4/app.zig b/src/esp32p4/app.zig new file mode 100644 index 00000000..a9cf627d --- /dev/null +++ b/src/esp32p4/app.zig @@ -0,0 +1,546 @@ +//! pardes, as ESP32-P4 firmware: the reset entry, the heap, the clock, the trap handler and the +//! loop. +//! +//! There is no operating system under this. `_start` is the reset entry the second-stage bootloader +//! jumps to, and this file is the entire platform: a heap, a millisecond clock, and UART0. +//! +//! ## Why the firmware root is in the editor's repository +//! +//! It was written in the `05-zig-p4` toolchain repository, next to the SoC support it uses, and it +//! moved here because everything in it is a statement about the EDITOR. The heap span it hands over +//! is the number that decides how large a grid the board can drive; `input_chunk` is sized against +//! what applying one keystroke costs in `src/pardes.zig`; the loop's shape - read, chunk, tick, +//! render only when dirty - is this editor's loop and no one else's; and the `-Dprof` attribution +//! exists to answer "where did the 34 ms of a keystroke go" about this program. A firmware root that +//! specific to one application belongs beside it. +//! +//! What stayed behind is everything a second application would also want, and none of it is +//! duplicated here: the SoC and HAL, the translate-c register layer, the coalescing heap, `std.Io` +//! for this chip, the app descriptor, the generated linker script, the image builder, the flasher +//! and the interactive console. Those arrive as the `zig_p4` dependency, and this file imports +//! exactly four of its modules - `soc`, `hal`, `heap` and `config` - plus two sibling files, +//! `uart.zig` and `input_rescue.zig`, which are the editor's own. +//! +//! ## Where the editor is +//! +//! On the far side of a C ABI, still, and that is a choice rather than a leftover. `src/esp32p4.zig` in +//! this same repository is compiled as ONE freestanding object (`b.addObject`, rooted at that file) +//! and linked in beside this one; the `extern` declarations below are the near side of that seam. +//! +//! Importing `esp32p4.zig` as a module instead would be shorter to write and worse in every way that +//! matters. It would drag the core's whole module graph - vaxis, the themes, the allocator tiers - +//! into this root, which is the compilation that must stay small enough to reason about. It would +//! give the firmware two ways to reach the editor. And above all it would make the OBJECT path a +//! second arrangement, tested separately: that path is what `05-zig-p4 -Dpardes -Dpardes-obj=...` +//! builds, it is what every measurement in that repository's `experiments/` was taken through, and +//! it is a supported way to build this board. With the extern kept, both builds link the same eight +//! symbols against the same object file, so neither can drift and neither is the better-tested one. +//! The reasons the seam is a file at all - a nested `build.zig.zon` dependency broke every build in +//! the toolchain repository - are recorded in `src/esp32p4.zig:8-15` and `05-zig-p4/build.zig:238-260`. +//! +//! Who owns which symbol: `src/esp32p4.zig` exports all eight `pardes_esp32p4_*` functions and nothing else. +//! This file exports `_start`, `zig_main`, `trapEntry` and `trapReport`. `esp_app_desc` belongs to +//! neither and comes from the toolchain's own appdesc object, which the link adds unconditionally. +//! `abi_version` below is the one constant both sides spell, and its counterpart is +//! `src/esp32p4.zig:98` - one repository now, so a bump is two lines in one diff rather than two commits +//! in two trees. +//! +//! The seam is deliberately **bytes in, bytes out**. Everything that needs to know what a cell is - +//! vaxis, the ANSI encoder, the input parser, the capability handshake - lives on the far side, +//! next to the vaxis it is built against. What crosses is a byte stream in each direction, which is +//! exactly what a serial line is, so this file has no opinion about terminals at all. +//! +//! ## Where the memory is +//! +//! Measured on this die by the toolchain's `examples/memprobe.zig`, not read off a datasheet, and +//! written down once in the generated linker script (`05-zig-p4/build.zig:1572,1579,1584-1585`): +//! +//! 0x4FF00000..0x4FF3F000 252 KiB `l2mem`: .data/.bss/.stack are linked into this +//! 0x4FF3F000..0x4FF40000 4 KiB mask ROM .data/.bss - untouchable, ets_printf needs it +//! 0x4FF40000..0x4FFA0000 384 KiB `l2high`: handed to the editor as its entire heap +//! 0x4FFA0000..0x4FFC0000 128 KiB NOT memory - the L2 cache lives here +//! +//! That last line is why the heap is 384 KiB and not the 512 KiB an earlier version of this comment +//! claimed. The first probe wrote a pattern and read it back one page at a time and reported the +//! whole upper 512 KiB as RAM, because a store followed immediately by a load of the SAME address +//! returns the stored value whether the backing store is real, an address mirror, or merely a dirty +//! cache line. Writing every page before reading any page separates the three, and the top 128 KiB +//! then failed; handing them to an allocator hung the heap on its first free-list walk. ESP-IDF's +//! own arithmetic agrees exactly: SRAM_HIGH_SIZE = 0x80000 - CONFIG_CACHE_L2_CACHE_SIZE, with the +//! Kconfig default of 128 KiB. +//! +//! The span arrives as `__heap_start`/`__heap_end` from that script, so those addresses are written +//! down in exactly one place. The editor owns it outright: it is passed in at init and this file +//! never allocates from it. +//! +//! PSRAM is not used. The board has 32 MB fitted and it would make all of this comfortable, but +//! ESP-IDF's own ESP32-P4 implementation runs past a thousand lines - MPLL, MSPI clocking, pin +//! drive and DQS, CS timing, mode registers, a connectivity check, and an entire timing-calibration +//! subsystem - and the mask ROM offers only MMU mapping, no device init. Touching it untrained +//! faults and hangs the core, which `examples/memprobe.zig` demonstrates on purpose. + +const std = @import("std"); +const soc = @import("soc"); +const config = @import("config"); + +/// `-Dprof`: time the two phases of a keystroke on the board and print the cycle counts. A +/// diagnostic, not a feature - see the loop. +const prof = config.prof; + +/// Every byte this loop has taken off the UART, for `-Dprof`. Ground truth for "did the burst +/// arrive", which a screen reconstruction cannot answer: a character can be missing from the screen +/// because it never arrived, because the editor never applied it, or because the viewport does not +/// show that column. +var rx_total: u32 = 0; + +/// How many input bytes to hand the editor before draining the receiver again. Chosen against the +/// FIFO rather than against the editor: applying one keystroke was measured at 44 us on an empty +/// line and 63 us at 640 characters, so eight of them is at most ~0.5 ms in which nothing empties +/// the receiver, against a 128-byte FIFO that holds 11 ms of wire at 115200. Twenty times the margin +/// needed, and it costs nothing on the wire because one render still happens per loop iteration. +const input_chunk = 8; +const hal = @import("hal"); +const heapmod = @import("heap"); +const uart = @import("uart.zig"); + +// ------------------------------------------------------------------------------------- the ABI +// Eight functions, all `callconv(.c)`, all implemented in the linked object - `src/esp32p4.zig` in this +// repository, compiled for the same target and exporting exactly these names. This is the complete +// interface between this board and the editor, and it is deliberately bytes-and-memory only: the +// editor never learns what a UART is, and this file never learns what a cell is. +// +// The declarations below are a SECOND spelling of the signatures in `src/esp32p4.zig:100-127,259-...`, +// and that duplication is what a C ABI is: each side declares the wire independently, which is +// precisely why `abi_version` has to be checked. Sharing a Zig type between them would mean sharing +// a module, which would mean the core in this compilation - see the header. + +/// How the editor emits bytes. Called with finished runs of ANSI, many times per frame. +const WriteFn = *const fn (ctx: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void; + +/// The board's pads, offered to the editor. Optional on the wire so a firmware with nothing to +/// toggle passes null and the `Gpio` word reports that rather than the object guessing. +const GpioFn = *const fn (ctx: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool; + +/// This board's allocator, handed across as plain function pointers. `log2_align` is a log2 value, +/// which is exactly how `std.mem.Alignment` represents itself, so neither side needs a conversion +/// table. +/// +/// The memory belongs to THIS side: only the firmware knows that the heap is the 384 KiB at +/// 0x4FF40000, that the 128 KiB above it is L2 cache, and that PSRAM is untrained. The editor gets +/// an allocator, not an address range. +const Allocator = extern struct { + ctx: ?*anyopaque, + alloc: *const fn (ctx: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8, + resize: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool, + free: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void, +}; + +/// The one number both sides must agree on. Linkers do not type-check C symbols, so a signature +/// that drifts on one side of this seam links cleanly and then corrupts the stack; checking this +/// before calling anything else turns that into a refusal to boot. +const abi_version: u32 = 2; +extern fn pardes_esp32p4_abi_version() callconv(.c) u32; + +/// Hand over the allocator and the output sink, and state the initial window size. Returns 0, or a +/// small non-zero code this file can only report. +extern fn pardes_esp32p4_init( + alloc: *const Allocator, + write: WriteFn, + gpio: ?GpioFn, + ctx: ?*anyopaque, + cols: u16, + rows: u16, +) callconv(.c) u32; + +/// Raw bytes off the wire: keystrokes, capability-query replies, and the host bridge's in-band +/// resize reports. The editor parses all three; this file distinguishes none of them. +extern fn pardes_esp32p4_input(ptr: [*]const u8, len: usize) callconv(.c) void; + +/// Advance time. Separate from `input` because animations and timeouts must progress on a wire +/// where nothing is arriving. +extern fn pardes_esp32p4_tick(now_ms: u64) callconv(.c) void; + +/// Emit one frame through the write callback. Returns 0 or an error code. +extern fn pardes_esp32p4_render() callconv(.c) u32; + +/// Is there anything to draw - a dirty surface or a running animation? Asked every iteration so a +/// quiet editor costs no bytes on a 115200-baud link. +extern fn pardes_esp32p4_wants_frame() callconv(.c) bool; + +/// Has the user asked to leave? There is nowhere to go, so this only stops the loop. +extern fn pardes_esp32p4_quit() callconv(.c) bool; + +/// The last frame's three stages in CPU cycles: the copy of pardes's Surface into vaxis's grid, +/// vaxis's own diff-and-emit, and the push into the UART. Only meaningful under `-Dprof`; the +/// editor object always exports it, and it costs two CSR reads per stage. +extern fn pardes_esp32p4_frame_prof(copy: *u64, render: *u64, flush: *u64) callconv(.c) void; + +// ------------------------------------------------------------------------------------ the sink + +/// The write callback handed to `pardes_esp32p4_init`. No context is needed - there is one UART. +fn writeOut(_: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void { + uart.write(ptr[0..len]); +} + +/// Flip one pad and report the level before and after. The editor's `Gpio` word calls this; the +/// editor has no register of its own for it, deliberately. +/// +/// THIS IS WHY THE SEAM IS HERE. A toggle is not a write to GPIO_OUT: `configureOutput` points the +/// pad's IO MUX at the GPIO function, routes the GPIO matrix's output to it, sets the drive strength +/// and input buffer and clears the pulls, and only then enables the driver - four register files, +/// indexed by a per-pin table. That code already exists in the toolchain package's `src/hal/gpio.zig`, +/// it is the same call that package's `src/main.zig` blinks with, and its register numbers are +/// checked against ESP-IDF's own headers by `zig build diff` there. A second copy inside the editor +/// object would be a second copy under no test. +/// +/// `getDrivenLevel` rather than `getLevel`: the answer is the level this board is DRIVING, which is +/// defined for every pin. The pad's own level is what the outside world says, and on an unconnected +/// header pin that is noise. The input buffer is enabled anyway, so `Peek` of GPIO_IN_REG shows the +/// pad for anyone who wants to compare the two. +fn gpioToggle(_: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool { + if (pin > hal.gpio.max_pin) return false; + const p: u8 = @intCast(pin); + hal.gpio.configureOutput(p, .{ .readback = true }); + const before = hal.gpio.getDrivenLevel(p); + if (before == 1) hal.gpio.setLow(p) else hal.gpio.setHigh(p); + was.* = before; + now.* = hal.gpio.getDrivenLevel(p); + return true; +} + +// ------------------------------------------------------------------------------------- the heap + +/// The span the linker script hands over, from `l2high`'s ORIGIN and LENGTH. +/// +/// Reached with `@extern`, NOT with `extern const __heap_start: anyopaque` plus +/// `@intFromPtr`/`@ptrFromInt`. That spelling was here first and it was silently wrong: declaring a +/// linker symbol as an `anyopaque` OBJECT gives the optimiser a zero-sized object, so a pointer +/// derived from its address carries provenance for zero bytes, and ordinary (non-volatile) stores +/// through it are dead code it may drop. The toolchain's `examples/heapcheck.zig` caught it on the +/// die - the allocator's first block header read back as `size=2988759312 next=0xffffffff`-not, and +/// the free list walk never terminated. A `[*]u8` from `@extern` has no size to lose. +const heap_start = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_start" }); +const heap_end = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_end" }); + +fn heapSpan() []align(heapmod.Heap.granule) u8 { + return heap_start[0 .. @intFromPtr(heap_end) - @intFromPtr(heap_start)]; +} + +/// The one heap. A K&R coalescing free list over that span, validated on this die by the toolchain's +/// `examples/heapcheck.zig`: 512 blocks fill and free back to a single 393,216-byte block, a holed +/// arena still satisfies a 4 KiB request, and 20,000 random operations drain back to one block. +var gpa_heap: heapmod.Heap = undefined; + +// The four C forwarders the editor is handed. `log2_align` round-trips through +// `std.mem.Alignment`, whose representation IS the log2 value. + +fn cAlloc(_: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8 { + const a = gpa_heap.allocator(); + return a.vtable.alloc(a.ptr, len, @enumFromInt(log2_align), @returnAddress()); +} + +fn cResize(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool { + const a = gpa_heap.allocator(); + return a.vtable.resize(a.ptr, ptr[0..len], @enumFromInt(log2_align), new_len, @returnAddress()); +} + +fn cFree(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void { + const a = gpa_heap.allocator(); + a.vtable.free(a.ptr, ptr[0..len], @enumFromInt(log2_align), @returnAddress()); +} + +const editor_allocator: Allocator = .{ + .ctx = null, + .alloc = cAlloc, + .resize = cResize, + .free = cFree, +}; + +// ------------------------------------------------------------------------------------ the clock + +/// Milliseconds since boot, off the systimer - a 16 MHz counter (the toolchain package's +/// `src/hal/systimer.zig:31`), which is the cheapest trustworthy clock on this chip. `read` returns +/// null if the unit is not running, in which case time simply does not advance and the editor stops +/// animating; that is a better failure than a clock that jumps. +fn nowMs() u64 { + const us = hal.systimer.micros(.unit0) orelse return 0; + return us / 1000; +} + +// ------------------------------------------------------------------------------------- the loop + +export fn zig_main() noreturn { + // FIRST, before a single byte of `.rodata` is touched - which means before the marker below, + // because that marker IS a string literal in flash and would read as machine code without this. + soc.flushFlashCache(); + const heap = heapSpan(); + soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{ + @as(u32, @intFromPtr(heap.ptr)), + @as(u32, @intFromPtr(heap.ptr)) + @as(u32, @intCast(heap.len)), + @as(u32, @intCast(heap.len / 1024)), + }); + + // The CPU clock, before anything is timed against it. The bootloader leaves 90 MHz and the + // CPLL is already at 360, so this is a divider change that disturbs neither UART0 (XTAL) nor + // the systimer (XTAL/2.5) nor the flash interface (SPLL). See the toolchain package's + // `src/hal/clkrst.zig:setCpuFreq`. + if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) { + 180 => .mhz180, + 360 => .mhz360, + else => .mhz90, + }); + + const rwdt_was_armed = hal.rwdt.disable(); + hal.systimer.init(); + _ = rwdt_was_armed; + + const their_abi = pardes_esp32p4_abi_version(); + if (their_abi != abi_version) { + uart.write("MARK PARDES_ABI_MISMATCH\r\n"); + while (true) {} + } + + gpa_heap = heapmod.Heap.init(heap); + _ = uart.drainInput(); + + // Ask for more than any grid this board will ever render, so the SHELL's own ceiling is what + // governs - it clamps to `-Desp32p4-cols`/`-Desp32p4-rows` and reports the result. Naming 80x24 here made + // the firmware a second opinion about the geometry, which is one opinion too many. + const rc = pardes_esp32p4_init(&editor_allocator, writeOut, gpioToggle, null, 255, 255); + + if (rc != 0) { + soc.rom.print("MARK PARDES_INIT_FAIL rc=%u\r\n", .{rc}); + const s = gpa_heap.stats(); + soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{ + s.free, s.largest_free, s.free_blocks, + }); + while (true) {} + } + + // The HEAP, after the editor has taken what it needs. This is the number that decides how large + // a grid the board can drive, so it is printed on every boot rather than only on failure: a + // geometry that fits with 2 KB to spare and one that fits with 80 KB are not the same answer, + // and the difference is invisible from the host otherwise. + { + const s = gpa_heap.stats(); + soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{ + s.free, s.largest_free, s.free_blocks, + }); + } + + // The CPU clock, measured rather than assumed. Every cycle count this firmware reports is + // divided by it somewhere, and the toolchain's `src/io/chip.zig` records it as "a measured + // ~90 MHz" that nothing here reconfigures - so it is worth printing rather than remembering. The + // systimer is XTAL/2.5 = 16 MHz and is NOT derived from the CPU clock + // (`src/hal/systimer.zig:31`, `clk_tree_defs.h:196-198`), which is exactly what makes it a valid + // reference for measuring it. + if (prof) { + const t_start = hal.systimer.micros(.unit0) orelse 0; + const c_start = soc.cycles(); + // 50 ms is long enough that the systimer's 16 MHz granularity and the loop's own overhead + // are both noise, and short enough to be invisible in a boot. + while ((hal.systimer.micros(.unit0) orelse 0) -% t_start < 50_000) {} + const elapsed_us = (hal.systimer.micros(.unit0) orelse 0) -% t_start; + const elapsed_cy = soc.cycles() - c_start; + soc.rom.print("MARK CPU_HZ cycles=%u us=%u khz=%u\r\n", .{ + @as(u32, @intCast(elapsed_cy)), + @as(u32, @intCast(elapsed_us)), + @as(u32, @intCast(if (elapsed_us > 0) elapsed_cy * 1000 / elapsed_us else 0)), + }); + } + soc.rom.print("MARK PARDES_READY\r\n", .{}); + + var in: [256]u8 = undefined; + while (!pardes_esp32p4_quit()) { + // ATTRIBUTION. The host can time a keystroke's round trip but cannot see what the firmware + // spent it on, and the two candidates - parsing and editing, versus rendering - want + // opposite fixes. `soc.cycles()` is the unprivileged cycle counter, so this costs two CSR + // reads per phase and quantises at one cycle, which is four orders of magnitude below the + // milliseconds being attributed. Gated on `prof` so the shipping build carries none of it. + const n = uart.read(&in); + rx_total +%= @intCast(n); + + var input_cy: u64 = 0; + if (n > 0) { + const t0 = if (prof) soc.cycles() else 0; + // IN CHUNKS, rescuing the receiver between them. Applying a keystroke is not free and + // gets dearer as the line grows - measured at 44 us on an empty line and 63 us at 640 + // characters - so handing over a full 128-byte batch is up to 8 ms in which nothing + // drains the receiver, against a FIFO that holds only 11 ms of wire. A 600-byte paste + // lost 93 bytes to exactly that window even with the transmitter's own rescue in place. + // + // Splitting a burst at an arbitrary byte is safe: `pardes_esp32p4_input` keeps whatever it + // could not parse, which is how it already survives an escape sequence split across two + // UART reads. One render still happens per loop iteration, so this costs no extra wire. + var off: usize = 0; + while (off < n) { + const chunk = @min(input_chunk, n - off); + pardes_esp32p4_input(in[off..].ptr, chunk); + off += chunk; + if (off < n) uart.rescueNow(); + } + if (prof) input_cy = soc.cycles() - t0; + } + + pardes_esp32p4_tick(nowMs()); + + // Only when there is something to show. On a link this slow an unconditional repaint per + // iteration would saturate the wire and starve input. + if (pardes_esp32p4_wants_frame()) { + const t0 = if (prof) soc.cycles() else 0; + const err = pardes_esp32p4_render(); + if (err != 0) soc.rom.print("MARK PARDES_RENDER_FAIL rc=%u\r\n", .{err}); + if (prof) { + const render_cy = soc.cycles() - t0; + // A SECOND render with nothing changed since the first. It splits the cost in two: + // whatever this still costs is the price of walking and diffing the whole editor + // state, paid regardless of output, while the difference between the two is the + // price of the change itself. `wants_frame` is false now, so this only happens + // under -Dprof and never on a shipping build. + const t1 = soc.cycles(); + _ = pardes_esp32p4_render(); + const idle_cy = soc.cycles() - t1; + // Reported in cycles, not microseconds: the divisor is the CPU clock, which this + // firmware does not set and has only ever measured, so converting here would bake a + // guess into the data. The toolchain's `experiments/` divides by the clock it + // measured. + var copy_cy: u64 = 0; + var vx_cy: u64 = 0; + var flush_cy: u64 = 0; + pardes_esp32p4_frame_prof(©_cy, &vx_cy, &flush_cy); + soc.rom.print("PROF in=%u render=%u idle=%u copy=%u vaxis=%u flush=%u rx=%u rxdrop=%u txdrop=%u\r\n", .{ + @as(u32, @intCast(input_cy)), + @as(u32, @intCast(render_cy)), + @as(u32, @intCast(idle_cy)), + @as(u32, @intCast(copy_cy)), + @as(u32, @intCast(vx_cy)), + @as(u32, @intCast(flush_cy)), + rx_total, + uart.inputDropped(), + uart.dropped, + }); + } + } + } + + soc.rom.print("\r\nMARK PARDES_QUIT\r\n", .{}); + while (true) {} +} + +// ------------------------------------------------------------------------------------ the trap + +/// A trap handler, because the absence of one is why this port has been guessing. +/// +/// The mask ROM prints "Guru Meditation" for a trap only while ITS handler is still installed; +/// anything this image does that replaces or outgrows that path fails silently instead, and a silent +/// fault is indistinguishable from an infinite loop over a serial line. This one reports the three +/// registers that name the fault and then stops, using the direct-FIFO writer so it shares nothing +/// with the editor's buffered output. +/// +/// `mtvec` is set in DIRECT mode (low two bits zero), so every trap and every interrupt lands on +/// `trapEntry` regardless of cause - which is what a diagnostic wants. +export fn trapEntry() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ("j trapReport"); +} + +export fn trapReport() noreturn { + const mcause = asm volatile ("csrr %[o], mcause" + : [o] "=r" (-> u32), + ); + const mepc = asm volatile ("csrr %[o], mepc" + : [o] "=r" (-> u32), + ); + const mtval = asm volatile ("csrr %[o], mtval" + : [o] "=r" (-> u32), + ); + uart.write("\r\nMARK TRAP mcause="); + uart.dumpWord(mcause); + uart.write("MARK TRAP mepc="); + uart.dumpWord(mepc); + uart.write("MARK TRAP mtval="); + uart.dumpWord(mtval); + uart.write("MARK TRAP dropped="); + uart.dumpWord(uart.dropped); + while (true) {} +} + +// --------------------------------------------------------------------------- the root's own duties +// +// These are the FIRMWARE root's declarations, and they are not the same set as `src/esp32p4.zig`'s: that +// file is the root of its own object and carries its own `std_options` and `panic` for the core's +// half of the image. Two roots, two instantiations of std, one per compilation unit - which is +// exactly what the object seam buys, and why a panic in the core prints `PARDES_CORE_PANIC` through +// the write callback while a panic here prints `PARDES_PANIC` through the mask ROM. + +/// `page_size_min`/`max`: the board has no MMU and no pages, but std derives allocator alignment +/// from these. 4 KiB is the ESP32-P4's cache and DMA granularity. +/// +/// `logFn` is not cosmetic. std's default log implementation reaches `std.debug_io`, which +/// instantiates `std.Io.Threaded` - a thread pool, `getrandom`, `IOV_MAX`, `mremap` - none of which +/// exist here, and one `log.warn` from anywhere is enough to drag all of it into the image. +pub const std_options: std.Options = .{ + .page_size_min = 4096, + .page_size_max = 4096, + .logFn = logFn, +}; + +fn logFn( + comptime level: std.log.Level, + comptime scope: @EnumLiteral(), + comptime fmt: []const u8, + args: anytype, +) void { + var buf: [256]u8 = undefined; + const line = std.fmt.bufPrint(&buf, "\r\n[" ++ level.asText() ++ "/" ++ @tagName(scope) ++ "] " ++ fmt ++ "\r\n", args) catch + "\r\n[log overflow]\r\n"; + uart.write(line); +} + +pub const panic = std.debug.FullPanic(panicImpl); + +fn panicImpl(msg: []const u8, first_trace_addr: ?usize) noreturn { + // The fixed text goes out through the ROM deliberately: a panic may BE the console writer + // failing, and `ets_printf` shares nothing with `uart.write` except the FIFO itself. + // + // The MESSAGE does not, and that is a correction rather than a preference. `msg` is a Zig SLICE + // and `%s` reads until a NUL, so handing `msg.ptr` to printf prints the message and then + // whatever happens to sit after it in memory until a zero byte turns up. Literals get away with + // it; std's own panics do not, because they are formatted into a buffer - "index out of bounds: + // index 5, len 3" - and carry no terminator. `uart.write` takes a length. + soc.rom.print("\r\nMARK PARDES_PANIC ", .{}); + uart.write(msg); + // The address is what makes it actionable: addr2line against the ELF in zig-out turns it into a + // source line, and without it a panic message names a KIND of failure with no way to find which + // one of them happened. Zero when the caller had no return address to give. + soc.rom.print("\r\nMARK PARDES_PANIC_AT 0x%08x\r\n", .{@as(u32, @truncate(first_trace_addr orelse 0))}); + while (true) {} +} + +/// Reset entry. The bootloader hands over with an unspecified stack pointer and the FPU off, so: +/// enable the F extension (`mstatus.FS`, which ESP-IDF only ever turns on lazily from a trap handler +/// this image does not have), establish a stack, clear `.bss`, and call into Zig. +/// +/// The cache invalidate that this image also needs is the FIRST thing `zig_main` does, not something +/// done here. Hand-written `la t0, Cache_Invalidate_All` against an absolute linker symbol computed +/// a PC-relative target and jumped into nowhere (measured: PC=0x88b5d788 with the argument stranded +/// in a2); Zig generates the addressing for an `extern fn` correctly, and `zig_main` runs before any +/// `.rodata` is touched anyway. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, trapEntry + \\ csrw mtvec, t0 + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} diff --git a/src/esp32p4/input_rescue.zig b/src/esp32p4/input_rescue.zig new file mode 100644 index 00000000..01dd4820 --- /dev/null +++ b/src/esp32p4/input_rescue.zig @@ -0,0 +1,271 @@ +//! Keystrokes rescued from the receive FIFO while the transmitter is busy. +//! +//! THE BUG THIS EXISTS FOR. The firmware's loop is read, apply, render, write, and the write blocks +//! while the transmit FIFO is full - real backpressure, because dropping half an escape sequence +//! would leave the host terminal in the wrong colour for the rest of the session. But nothing +//! drained the RECEIVE FIFO during that wait, and the FIFO is 128 bytes (the toolchain package's +//! `src/hal/uart.zig:52`). A frame of 240 bytes is 21 ms of wire at 115200, and 21 ms of a host +//! sending at line rate is ~240 bytes, so everything past the 128th was silently gone. +//! +//! Measured on the die before the fix, typing a burst in one host write and counting what the editor +//! actually held: 128 bytes arrived intact, 200 bytes lost 88, 300 bytes lost all 300. From a +//! keyboard that is a keystroke that never lands, and it looks like a stuck key - the screen is +//! behind what was typed, and typing more appears to fix it because a later frame repaints the cells +//! the lost keystrokes would have changed. +//! +//! WHY THE POLICY LIVES HERE and not in `uart.zig`: the interesting part is a decision - drain the +//! receiver while spinning on the transmitter, and what to do when even that overflows - and the +//! decision is worth testing. `uart.zig` cannot be tested at all without the chip, because every +//! line of it is an MMIO access. `pump` takes the port as `anytype`, so the same code runs against +//! the real UART on the board and against a fake with a two-byte FIFO on the host. +//! +//! WHY IT IS IN THIS REPOSITORY. It was written in the toolchain repository, next to the UART it +//! spins on, and it moved here because every number in it is the EDITOR's. 4 KiB is sized against +//! the ~1.4 KB full repaint `src/esp32p4.zig` emits; "drop the newest, so what survives is a PREFIX of +//! what was typed" is a statement about documents rather than about UARTs, and it is the editor that +//! would otherwise appear to invent input. A policy whose every constant comes from one application +//! belongs beside that application. +//! +//! It expects NOTHING from the toolchain package - no module, no register, no target. `std` is the +//! whole import list, which is what makes the host tests at the bottom possible and what lets the +//! same source compile for riscv32-freestanding and for the host unchanged. +//! +//! ONE COPY, TWO BUILDS. This file is the only copy; the toolchain repository's is gone. Its build +//! reads this tree across a sibling-relative seam, and names this path twice: once as the +//! `input_rescue` module of `-Dpardes`'s application (`05-zig-p4/build.zig:230-231`, whose root is +//! `../02-pardes-code/src/esp32p4/app.zig` by the `-Dapp` default at `:168-169`) and once as the +//! same-named module of `zig build selftest`'s on-die root (`:437-438`). Both spell +//! `../02-pardes-code/src/esp32p4/input_rescue.zig`, so there is nothing to keep in step. +//! +//! Tested from here, both ways: the host checks at the bottom run as their own `addTest` under +//! `zig build unit-test` (`02-pardes-code/build.zig:1727-1732` - no `link_libc`, because `std` is +//! the whole import list), and the same source runs against UART0 on the die under +//! `zig build esp32p4-test`. + +const std = @import("std"); + +/// Capacity, sized for the worst frame this editor emits. +/// +/// A full repaint is ~1.4 KB, which is 121 ms of wire at 115200, and 121 ms of a host pasting at +/// line rate is ~1.4 KB of input. 4 KiB is that with headroom, a power of two so the wrap is a mask +/// rather than a division, and nothing at all against the board's RAM. +pub const capacity = 4096; + +/// A byte queue that drops the NEWEST byte when full. +/// +/// Dropping the newest rather than the oldest is deliberate: what survives is then a PREFIX of what +/// was typed. An editor that loses the end of a paste has done something a person can see and +/// correct; one that silently reorders keystrokes, or keeps the tail and discards the head, has +/// corrupted the document in a way that looks like the editor inventing input. +pub const Ring = struct { + buf: [capacity]u8 = undefined, + head: usize = 0, + len: usize = 0, + /// Bytes lost because even this overflowed. Nonzero means input was dropped; it is the honest + /// version of the bug rather than a cure for it. + dropped: u32 = 0, + + const mask = capacity - 1; + + comptime { + std.debug.assert(capacity & mask == 0); + } + + pub fn push(r: *Ring, b: u8) void { + if (r.len == capacity) { + r.dropped +%= 1; + return; + } + r.buf[(r.head + r.len) & mask] = b; + r.len += 1; + } + + /// Move as much as fits into `out`, oldest first. Returns the count. + pub fn pop(r: *Ring, out: []u8) usize { + const n = @min(out.len, r.len); + for (out[0..n]) |*slot| { + slot.* = r.buf[r.head]; + r.head = (r.head + 1) & mask; + } + r.len -= n; + return n; + } + + pub fn clear(r: *Ring) void { + r.head = 0; + r.len = 0; + } +}; + +/// Drain everything the port has received into `ring`, without waiting. +pub fn rescue(port: anytype, ring: *Ring) void { + var waiting = port.rxCount(); + while (waiting > 0) : (waiting -= 1) ring.push(port.popByte()); +} + +/// Push `bytes` through `port`, rescuing input whenever the transmitter has no room. Returns the +/// number of bytes abandoned because the transmitter stopped making progress altogether. +/// +/// The spin bound is why this returns a count rather than blocking forever: a UART whose core clock +/// has been gated never makes progress, and on a board with no debugger an infinite spin is +/// indistinguishable from a crash. A bounded wait turns that into visibly dropped output plus a +/// counter, which is a diagnosis instead of a mystery. +pub fn pump(port: anytype, ring: *Ring, bytes: []const u8, spin_limit: u32) u32 { + var rest = bytes; + while (rest.len > 0) { + // One status read per burst, not per byte: reading `txFree` once and pushing that many cuts + // the status reads by up to the FIFO depth. + var room = port.txFree(); + var spins: u32 = 0; + while (room == 0) { + // THE FIX. Every iteration of this wait is time the receiver is filling up, and this is + // the only place that can empty it. + rescue(port, ring); + spins += 1; + if (spins > spin_limit) return @intCast(rest.len); + room = port.txFree(); + } + const n = @min(room, rest.len); + for (rest[0..n]) |b| port.pushByte(b); + rest = rest[n..]; + } + return 0; +} + +// ------------------------------------------------------------------------------------ host tests + +test "the ring hands bytes back in order" { + var r: Ring = .{}; + for ("hello") |b| r.push(b); + var out: [8]u8 = undefined; + try std.testing.expectEqual(@as(usize, 5), r.pop(&out)); + try std.testing.expectEqualStrings("hello", out[0..5]); + try std.testing.expectEqual(@as(usize, 0), r.pop(&out)); +} + +test "the ring wraps without reordering" { + var r: Ring = .{}; + var out: [capacity]u8 = undefined; + // Push and pop most of the buffer so head sits near the end, then straddle the wrap. + for (0..capacity - 3) |i| r.push(@intCast(i & 0xff)); + _ = r.pop(out[0 .. capacity - 3]); + for ("straddle") |b| r.push(b); + const n = r.pop(&out); + try std.testing.expectEqualStrings("straddle", out[0..n]); +} + +test "a full ring drops the newest and says so" { + var r: Ring = .{}; + for (0..capacity) |i| r.push(@intCast(i & 0xff)); + try std.testing.expectEqual(@as(u32, 0), r.dropped); + r.push('!'); + r.push('!'); + try std.testing.expectEqual(@as(u32, 2), r.dropped); + // The head is intact: what survived is a prefix of what arrived. + var out: [4]u8 = undefined; + _ = r.pop(&out); + try std.testing.expectEqual(@as(u8, 0), out[0]); + try std.testing.expectEqual(@as(u8, 1), out[1]); +} + +/// A UART with a small transmit FIFO, a small RECEIVE FIFO, and a host that keeps typing into it. +/// +/// The receive FIFO is the part that matters and it is modelled the way the hardware behaves: it has +/// a fixed depth, and a byte that arrives when it is full is *gone*. That is the whole bug. +/// +/// Time advances on each transmitter status read, which is what `pump` does while it waits. The +/// transmitter frees a byte only every fourth tick while a typed byte lands on every one: the +/// transmitter therefore genuinely FILLS, which is the condition the bug needs. A fake whose FIFO +/// drains as fast as it fills never blocks, so `pump` never waits, so the rescue never runs and the +/// test proves nothing - the first version of this fake had exactly that flaw. +const FakePort = struct { + tx_cap: u32, + tx_used: u32 = 0, + sent: std.ArrayList(u8) = .empty, + gpa: std.mem.Allocator, + + incoming: []const u8, + delivered: usize = 0, + rx: [rx_depth]u8 = undefined, + rx_head: usize = 0, + rx_len: usize = 0, + /// Bytes the wire delivered into a full receive FIFO. The hardware has no counter for this, + /// which is exactly why the bug was invisible. + lost: u32 = 0, + + ticks: u32 = 0, + + const rx_depth = 8; + const tx_drain_every = 4; + + fn tick(p: *FakePort) void { + p.ticks += 1; + if (p.ticks % tx_drain_every == 0 and p.tx_used > 0) p.tx_used -= 1; + if (p.delivered < p.incoming.len) { + const b = p.incoming[p.delivered]; + p.delivered += 1; + if (p.rx_len == rx_depth) { + p.lost += 1; + } else { + p.rx[(p.rx_head + p.rx_len) % rx_depth] = b; + p.rx_len += 1; + } + } + } + + fn txFree(p: *FakePort) u32 { + p.tick(); + return p.tx_cap - p.tx_used; + } + + fn pushByte(p: *FakePort, b: u8) void { + p.sent.append(p.gpa, b) catch unreachable; + p.tx_used += 1; + } + + fn rxCount(p: *FakePort) u32 { + return @intCast(p.rx_len); + } + + fn popByte(p: *FakePort) u8 { + const b = p.rx[p.rx_head]; + p.rx_head = (p.rx_head + 1) % rx_depth; + p.rx_len -= 1; + return b; + } +}; + +test "a long transmit does not lose the input that arrives during it" { + // THE REGRESSION. Delete the `rescue` call inside `pump`'s wait and this fails: the receive FIFO + // is eight bytes deep, the typing below is far longer than that, and every byte that arrives + // into a full FIFO is gone with nothing to record it. That is the die's 88-of-200 in miniature. + const typed = "the quick brown fox jumps over the lazy dog, twice over, and then some more"; + var port: FakePort = .{ .tx_cap = 2, .incoming = typed, .gpa = std.testing.allocator }; + defer port.sent.deinit(std.testing.allocator); + var ring: Ring = .{}; + + const frame = "\x1b[1;1H" ++ "x" ** 400; + try std.testing.expectEqual(@as(u32, 0), pump(&port, &ring, frame, 1_000_000)); + + // Every output byte went out, in order. + try std.testing.expectEqualStrings(frame, port.sent.items); + // Nothing the wire delivered was dropped, by the FIFO or by the ring. + try std.testing.expectEqual(@as(u32, 0), port.lost); + try std.testing.expectEqual(@as(u32, 0), ring.dropped); + // And what was rescued, plus whatever is still sitting in the FIFO, is exactly what was typed - + // in order, which is the other half of the contract. + var got: [capacity]u8 = undefined; + var n = ring.pop(&got); + while (port.rxCount() > 0) : (n += 1) got[n] = port.popByte(); + try std.testing.expectEqualStrings(typed[0..port.delivered], got[0..n]); + try std.testing.expect(port.delivered == typed.len); +} + +test "a transmitter that never drains gives up and reports what it abandoned" { + var port: FakePort = .{ .tx_cap = 0, .incoming = "", .gpa = std.testing.allocator }; + defer port.sent.deinit(std.testing.allocator); + var ring: Ring = .{}; + // tx_cap 0 means txFree is always 0, so no byte can ever go out. + try std.testing.expectEqual(@as(u32, 5), pump(&port, &ring, "abcde", 32)); + try std.testing.expectEqual(@as(usize, 0), port.sent.items.len); +} diff --git a/src/esp32p4/selftest.zig b/src/esp32p4/selftest.zig new file mode 100644 index 00000000..0d692d98 --- /dev/null +++ b/src/esp32p4/selftest.zig @@ -0,0 +1,351 @@ +//! pardes's own test suite for the ESP32-P4, which runs ON THE DIE. +//! +//! WHY THIS EXISTS. Every serious bug this port has produced was invisible to a host test, and two +//! of them were invisible for months. `std.mem.eql` compares a byte at a time on this target and +//! several times faster on the host, so the firmware's largest read was three times slower than it +//! needed to be and nothing on a laptop could tell. A lone ESC resolves to the Escape key, which is +//! right when a kernel hands over a whole escape sequence and wrong when a 115200 line hands over +//! one byte every 87 microseconds. A full transmit FIFO stopped anything draining the receiver, and +//! the FIFO depth is a hardware number. None of those is a logic error you can reason your way to +//! from a host: they are properties of THIS chip, THIS clock and THIS wire. +//! +//! So the checks below are chosen by one rule: a check belongs here only if the die can answer it +//! and a host cannot. Anything that is pure logic - the ring's wrap-around, the mouse coalescer's +//! ordering - already has a deterministic host test in `zig build unit-test`, which is faster, needs +//! no hardware, and is where such a thing belongs. Duplicating those here would make this suite +//! longer and no stronger. +//! +//! ## Why the suite is in the editor's repository +//! +//! It was written in the `05-zig-p4` toolchain repository as `examples/selftest.zig`, and it was +//! never an example of anything. Re-read the paragraph above: every claim it makes is a claim about +//! THIS program on this board. The word-at-a-time `std.mem.eql` it measures against is the reason +//! the editor's frame diff compares rows a `u32` at a time; the lone-ESC decode and the +//! transmit-FIFO backpressure are `input_rescue.zig`, a sibling file in this directory; the heap +//! span and the allocator churn are what decide how large a grid this board can drive; and the +//! clock check is the divisor under every `-Dprof` cycle count the editor reports. A suite that +//! specific to one application belongs beside it, so that changing a source and changing the check +//! which defends it are one diff in one repository rather than two commits in two trees. +//! +//! What stayed behind in the toolchain is the half that is about the CHIP: the image builder's rules, +//! the register layer's field arithmetic, the console bridge's escape matcher. Those still run under +//! `zig build test` over there and need no board. +//! +//! ## What it expects from the toolchain package +//! +//! Four of its modules and no more: `soc` (mask-ROM printf, the cycle counter), `hal` (UART0, the +//! systimer, clock/reset), `config` (this build's `cpu_mhz`) and `heap` (the coalescing allocator). +//! They arrive as the `zig_p4` dependency, which also supplies what makes this an image at all - +//! the generated linker script, `ENTRY(_start)` and the app descriptor via `firmware(...).attach`, +//! then `ImageStep`, `FlashStep` and `SelftestStep`. +//! +//! `input_rescue` is the fifth import and is NOT that package's: it is the sibling file in this +//! directory, handed over as a named MODULE rather than imported as a path. Deliberately so - the +//! `pub` on `FakePort`'s methods below is what lets that module reach them by duck typing across the +//! boundary, and both builds, this repository's `esp32p4-test` and the toolchain's `selftest`, wire +//! the identical root the identical way. One file, one arrangement, nothing to drift. +//! +//! Not a `zig test` binary, deliberately. Zig's test runner wants an OS, and `std.testing.allocator` +//! is a debug allocator over the page allocator, which on freestanding is either a compile error or +//! a lie. A hand-rolled harness is thirty lines and answers to nobody. +//! +//! Run with: zig build esp32p4-test -Dplatform=esp32p4 -Desp32p4-firmware (from here) +//! or: zig build selftest (from ../05-zig-p4) + +const std = @import("std"); +const soc = @import("soc"); +const hal = @import("hal"); +const config = @import("config"); +const heapmod = @import("heap"); +const input_rescue = @import("input_rescue"); + +/// Reset entry. Identical in shape to `src/main.zig`'s and for the same reasons: the bootloader hands +/// over with an unspecified stack pointer and the FPU off, so set `mstatus.FS`, establish a stack, +/// clear `.bss`, and jump. +/// +/// Leaving this out is what made the first draft of this file unbuildable, and the failure said +/// nothing useful: `-fentry=_start` found no such symbol, `--gc-sections` then discarded every +/// function as unreachable, and the image builder reported `NotTwoMappedSegments` because +/// `.flash.text` had nothing left in it. `zig build layout` now prints `entry 0x0` for exactly that, +/// which is the same diagnosis in one line. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} + +pub const panic = std.debug.FullPanic(struct { + fn call(msg: []const u8, first_trace_addr: ?usize) noreturn { + // `msg` is a slice and carries no terminator - std's own panics are formatted into a buffer - + // so it goes out with a length rather than through a `%s` that would read past the end of it. + soc.rom.print("\r\nMARK SELFTEST_PANIC at 0x%08x: ", .{@as(u32, @truncate(first_trace_addr orelse 0))}); + hal.uart.Uart.init(0).write(msg); + // A panic is a FAILED RUN, and the host is watching for the summary line. Without this the + // run looks like a board that never answered, which is a different diagnosis entirely. + soc.rom.print("\r\nMARK SELFTEST DONE pass=%u fail=%u\r\n", .{ passed, failed + 1 }); + while (true) {} + } +}.call); + +var passed: u32 = 0; +var failed: u32 = 0; + +/// One claim about the silicon. Printed either way: a suite that only speaks up when it fails gives +/// no way to tell "all good" from "never ran", and on a board that difference matters. +fn check(name: [*:0]const u8, ok: bool) void { + if (ok) { + passed += 1; + soc.rom.print("MARK SELFTEST ok %s\r\n", .{name}); + } else { + failed += 1; + soc.rom.print("MARK SELFTEST FAIL %s\r\n", .{name}); + } +} + +/// Word-at-a-time equality, the same shape the editor's frame diff uses. +fn sameBytesWordwise(a: []const u8, b: []const u8) bool { + if (a.len != b.len) return false; + if ((@intFromPtr(a.ptr) | @intFromPtr(b.ptr)) & 3 == 0) { + const n = a.len / 4; + const wa: [*]align(4) const u32 = @ptrCast(@alignCast(a.ptr)); + const wb: [*]align(4) const u32 = @ptrCast(@alignCast(b.ptr)); + for (wa[0..n], wb[0..n]) |x, y| { + if (x != y) return false; + } + return std.mem.eql(u8, a[n * 4 ..], b[n * 4 ..]); + } + return std.mem.eql(u8, a, b); +} + +/// A UART with a receive FIFO that loses whatever arrives into a full one, as the hardware does. +/// `pub` on the methods because `input_rescue` is a separate module here and reaches them by duck +/// typing across it. +const FakePort = struct { + tx_cap: u32, + tx_used: u32 = 0, + ticks: u32 = 0, + sent: u32 = 0, + incoming: []const u8, + delivered: usize = 0, + rx: [8]u8 = undefined, + rx_head: usize = 0, + rx_len: usize = 0, + lost: u32 = 0, + + fn tick(p: *FakePort) void { + p.ticks += 1; + if (p.ticks % 4 == 0 and p.tx_used > 0) p.tx_used -= 1; + if (p.delivered < p.incoming.len) { + const byte = p.incoming[p.delivered]; + p.delivered += 1; + if (p.rx_len == p.rx.len) { + p.lost += 1; + } else { + p.rx[(p.rx_head + p.rx_len) % p.rx.len] = byte; + p.rx_len += 1; + } + } + } + pub fn txFree(p: *FakePort) u32 { + p.tick(); + return p.tx_cap - p.tx_used; + } + pub fn pushByte(p: *FakePort, _: u8) void { + p.sent += 1; + p.tx_used += 1; + } + pub fn rxCount(p: *FakePort) u32 { + return @intCast(p.rx_len); + } + pub fn popByte(p: *FakePort) u8 { + const byte = p.rx[p.rx_head]; + p.rx_head = (p.rx_head + 1) % p.rx.len; + p.rx_len -= 1; + return byte; + } +}; + +/// Backing store for the heap checks. Static, because the point is to exercise the allocator on real +/// L2MEM rather than to find out where a stack array happens to land. +var heap_area: [64 * 1024]u8 align(16) = undefined; + +export fn zig_main() noreturn { + // THE CLOCK RAISE, performed here rather than assumed, which is what turns the frequency check + // at the end into a test of `setCpuFreq` instead of a tautology. The first version of this file + // read `config.cpu_mhz` and compared the die against it without ever setting it, so + // `-Dcpu-mhz=360` failed with `khz=90001 want=360000` - the check was right and the expectation + // was wrong. Same call, and the same order, as `src/esp32p4/app.zig`. + if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) { + 360 => .mhz360, + else => .mhz90, + }); + hal.systimer.init(); + + soc.rom.print("\r\nMARK SELFTEST_START cpu_mhz=%u\r\n", .{@as(u32, config.cpu_mhz)}); + + // ------------------------------------------------ 1. the memory model the memory words assume + // + // `Peek`, `Poke` and `Hexdump` reach the bus through `*allowzero volatile` pointers and refuse an + // unaligned word. Both halves are claims about this core, and neither is checkable on a host. + { + const cell: *volatile u32 = @ptrCast(@alignCast(&heap_area[0])); + cell.* = 0xdeadbeef; + check("an aligned word round-trips through a volatile pointer", cell.* == 0xdeadbeef); + + // Every byte offset in a word, readable and writable: this is what `Hexdump` does, and it is + // why `Hexdump` needs no alignment while `Peek` does. + var all_offsets_ok = true; + for (0..4) |i| { + const at: *volatile u8 = @ptrCast(&heap_area[16 + i]); + at.* = @intCast(0xa0 + i); + if (at.* != 0xa0 + @as(u8, @intCast(i))) all_offsets_ok = false; + } + check("a byte at every offset in a word round-trips", all_offsets_ok); + + // THE VOLATILE PROMISE. Two reads of a running counter must be two reads. Were the optimiser + // allowed to fold them, `Peek` would print one value twice for a register that had changed, + // which is the one thing a memory word must never do. + const first = hal.systimer.micros(.unit0) orelse 0; + var spin: u32 = 0; + while (spin < 4000) : (spin += 1) asm volatile ("" ::: .{ .memory = true }); + const second = hal.systimer.micros(.unit0) orelse 0; + check("two reads of a live counter are two reads", second != first); + check("and that counter runs forwards", second > first); + } + + // --------------------------------------- 2. word-wise equality, on THIS instruction set + // + // The editor's frame diff compares rows a `u32` at a time because `std.mem.eql` compares a byte + // at a time here: 223 us against 66 for the same answer. "The same answer" is the part that has + // to hold on the target rather than on the host, so it is checked against `std.mem.eql` itself, + // at every difference position, at both alignments, and at lengths that are not multiples of 4. + { + var a: [64]u8 align(4) = undefined; + var b: [64]u8 align(4) = undefined; + for (&a, 0..) |*slot, i| slot.* = @intCast(i); + @memcpy(&b, &a); + + var agree = true; + for (0..a.len) |len| { + if (sameBytesWordwise(a[0..len], b[0..len]) != std.mem.eql(u8, a[0..len], b[0..len])) agree = false; + } + check("aligned equality agrees with std.mem.eql at every length", agree); + + agree = true; + for (0..a.len) |i| { + b[i] ^= 0xff; + if (sameBytesWordwise(&a, &b) != std.mem.eql(u8, &a, &b)) agree = false; + if (sameBytesWordwise(a[0..33], b[0..33]) != std.mem.eql(u8, a[0..33], b[0..33])) agree = false; + b[i] ^= 0xff; + } + check("a difference at any position is found, as std.mem.eql finds it", agree); + + agree = true; + // UNALIGNED, which is why `sameBytes` tests alignment at RUNTIME: `Cell` is all u8 fields, so + // whether a row starts on a word boundary belongs to the allocator and not to the type. + for (1..4) |off| { + const ua = a[off..]; + const ub = b[off..]; + if (sameBytesWordwise(ua, ub) != std.mem.eql(u8, ua, ub)) agree = false; + b[off + 5] ^= 0xff; + if (sameBytesWordwise(ua, ub) != std.mem.eql(u8, ua, ub)) agree = false; + b[off + 5] ^= 0xff; + } + check("unaligned spans fall back and still agree", agree); + } + + // ------------------------------------------------- 3. the rescue, on the real codegen + // + // The logic has a host test. What that cannot say is whether it behaves the same compiled for + // this core at this optimisation level, which is a question only a board answers. + { + const typed = "the quick brown fox jumps over the lazy dog, and then some more besides"; + var port: FakePort = .{ .tx_cap = 2, .incoming = typed }; + var ring: input_rescue.Ring = .{}; + const frame = "\x1b[1;1H" ++ "x" ** 300; + const abandoned = input_rescue.pump(&port, &ring, frame, 1_000_000); + + check("the frame went out whole", abandoned == 0 and port.sent == frame.len); + check("the receive FIFO never overflowed", port.lost == 0); + check("the ring dropped nothing", ring.dropped == 0); + + var got: [128]u8 = undefined; + var n = ring.pop(&got); + while (port.rxCount() > 0 and n < got.len) : (n += 1) got[n] = port.popByte(); + check("every rescued byte, in order", n == typed.len and std.mem.eql(u8, got[0..n], typed)); + } + + // --------------------------------------------------- 4. the allocator on real L2MEM + // + // The editor's whole geometry ceiling is an allocator question, and this is the allocator, on the + // memory it actually runs in rather than on a host's malloc. + { + var h = heapmod.Heap.init(heap_area[0..]); + const gpa = h.allocator(); + + const one = gpa.alloc(u32, 256) catch null; + check("a modest allocation succeeds", one != null); + if (one) |slice| { + check("and is aligned for its element", @intFromPtr(slice.ptr) % @alignOf(u32) == 0); + for (slice, 0..) |*slot, i| slot.* = @intCast(i * 7); + var intact = true; + for (slice, 0..) |slot, i| { + if (slot != i * 7) intact = false; + } + check("and holds what was written to it", intact); + gpa.free(slice); + } + + // FREE THEN REUSE. A heap that cannot hand the same bytes back is a heap that runs out, which + // on this board is the difference between a 40x12 grid and an 80x24 one. + var churn_ok = true; + for (0..64) |_| { + const block = gpa.alloc(u8, 1024) catch { + churn_ok = false; + break; + }; + gpa.free(block); + } + check("a kilobyte can be taken and returned repeatedly", churn_ok); + + // AND IT REFUSES CLEANLY. An allocator that returns garbage instead of an error when it is + // out is the failure mode that cost an afternoon during bring-up. + const absurd = gpa.alloc(u8, heap_area.len * 4); + check("an impossible allocation returns an error", absurd == error.OutOfMemory); + } + + // ------------------------------------------ 5. the clock, which everything divides by + // + // Every cycle count this firmware reports is divided by the configured frequency somewhere, and + // the systimer is clocked from the crystal rather than from the CPU - which is exactly what makes + // it a reference the CPU cannot flatter. The 90-to-360 MHz raise rested on this comparison. + { + const t0 = hal.systimer.micros(.unit0) orelse 0; + const c0 = soc.cycles(); + while ((hal.systimer.micros(.unit0) orelse 0) -% t0 < 20_000) {} + const us = (hal.systimer.micros(.unit0) orelse 0) -% t0; + const cy = soc.cycles() - c0; + const khz: u32 = if (us > 0) @intCast(cy * 1000 / us) else 0; + const want: u32 = @as(u32, config.cpu_mhz) * 1000; + // Two percent: far wider than either clock's error, far narrower than the 4x a wrong divider + // would produce. + const slack = want / 50; + soc.rom.print("MARK SELFTEST_CLOCK khz=%u want=%u\r\n", .{ khz, want }); + check("the cycle counter and the systimer agree on the CPU frequency", khz > want - slack and khz < want + slack); + } + + soc.rom.print("MARK SELFTEST DONE pass=%u fail=%u\r\n", .{ passed, failed }); + while (true) {} +} diff --git a/src/esp32p4/uart.zig b/src/esp32p4/uart.zig new file mode 100644 index 00000000..0afd145b --- /dev/null +++ b/src/esp32p4/uart.zig @@ -0,0 +1,165 @@ +//! UART0 as the editor's terminal: bytes out, bytes in, and nothing else. +//! +//! This is the whole of the firmware's I/O. There is no framebuffer and no keyboard; the board +//! emits ANSI and consumes ANSI, and the terminal emulator on the far end of the CH340 does the +//! rest of the work - including answering the editor's own capability queries, which travel down +//! this wire like any other bytes. +//! +//! WHY IT IS IN THIS REPOSITORY. It is not a UART driver - that is `hal.uart`, which stays in the +//! toolchain package and is checked against ESP-IDF's own headers by `zig build diff` there. This is +//! the EDITOR's use of one: which instance the console is, that the transmitter is real +//! backpressure because a truncated escape sequence corrupts the host terminal, that rescued +//! keystrokes must come out before FIFO ones, and that the first thing to do at startup is discard +//! the host bridge's synthetic resize report. Every one of those is a statement about the editor, so +//! the file moved to sit beside it. See `app.zig`'s header for the boundary in full. +//! +//! What it expects from the toolchain package is exactly one module: `hal`, for `hal.uart.Uart`. +//! Nothing else here reaches the chip. The sibling `input_rescue.zig` is a plain file, not a module, +//! because it is this repository's own policy. +//! +//! Deliberately not a `std.Io.Writer`. The ANSI encoding lives on the other side of the C ABI, next +//! to the vaxis that produces it (see `app.zig` for why the seam is there and not elsewhere), so +//! what crosses into this file is already a finished run of bytes. A writer here would be a second +//! buffer in front of one that already exists. +//! +//! Two decisions worth stating, because both are measurements rather than preferences. +//! +//! **Batched FIFO access.** The naive push is `while (txFree() == 0) {}` then `pushByte`, once per +//! byte: one MMIO read per byte at best, many while the FIFO is full. Reading `txFree` once and +//! then pushing that many cuts the status reads by up to the FIFO depth (128, the toolchain +//! package's `src/hal/uart.zig:52`). At 115200 the wire costs ~86 us per byte and dwarfs either +//! version, so today this is merely free - and it stops being free the moment the divider is raised. +//! +//! **UART0's configuration is never touched.** Not the divider, not the format, not the pad +//! routing, and above all not `reset()`. The second-stage bootloader configured this block, and +//! `src/hal/uart.zig:195-211` records what happens if it is reset: UART_CLKDIV returns to its +//! power-on value, the console turns to garbage mid-sentence, and the board takes a watchdog reset +//! with nothing readable left to explain it. Everything here touches FIFO offset 0x000 and the +//! status register, and nothing else. + +const hal = @import("hal"); +const input_rescue = @import("input_rescue.zig"); + +/// UART0: the instance the CH340 is wired to, and the one the ROM and bootloader configured. +const uart0 = hal.uart.Uart.init(0); + +/// Keystrokes taken off the receiver while the transmitter was full. See `input_rescue`: without +/// this, anything typed into a frame longer than the 128-byte FIFO was silently gone. +var rescued: input_rescue.Ring = .{}; + +/// Push `bytes` into the TX FIFO, blocking while it is full. +/// +/// The spin is normally bounded by the wire - a full 128-byte FIFO drains in 11 ms at 115200 - and +/// dropping instead of waiting would truncate an escape sequence, leaving the host terminal in the +/// wrong colour for the rest of the session. So the wait is real backpressure. +/// +/// But it is BOUNDED, for the reason the toolchain package's `src/hal/uart.zig:182-186` gives about +/// `update()`: a UART whose core clock has been gated never makes progress, and "on a board with no +/// debugger an infinite spin is indistinguishable from a crash". That is not hypothetical here - it +/// is how this port spent an afternoon: output stopped mid-boot with no panic and no watchdog (the +/// RTC watchdog having been correctly disabled), which looked like a hang in whatever code came next +/// rather than a stalled transmitter. A bounded wait turns that into visibly dropped output plus a +/// counter, which is a diagnosis instead of a mystery. +/// +/// The limit is per burst, not per call, and generous: 1,000,000 status reads is far longer than +/// any legitimate drain and still a fraction of a second. +pub fn write(bytes: []const u8) void { + dropped +%= input_rescue.pump(uart0, &rescued, bytes, 1_000_000); +} + +/// Bytes abandoned because the transmitter stopped making progress. Nonzero means the console is +/// lying about what happened, so it is worth printing. +pub var dropped: u32 = 0; + +/// One byte, for callers that must not touch `.rodata` to say anything - which during bring-up is +/// the difference between a diagnostic and a second copy of the bug being diagnosed. +pub fn writeByte(b: u8) void { + var spins: u32 = 0; + while (uart0.txFree() == 0) { + spins += 1; + if (spins > 1_000_000) { + dropped +%= 1; + return; + } + } + uart0.pushByte(b); +} + +/// Emit `n` bytes read from `addr` as two hex digits each, computing the digits arithmetically so +/// nothing here reads a lookup table. Used to answer "does a load from this address return what the +/// linker put there", which is not a question a string literal can be trusted to ask. +pub fn dumpHex(addr: u32, n: u32) void { + const p: [*]const volatile u8 = @ptrFromInt(addr); + var i: u32 = 0; + while (i < n) : (i += 1) { + const byte = p[i]; + for ([2]u8{ byte >> 4, byte & 0xf }) |nib| { + writeByte(if (nib < 10) '0' + nib else 'a' + (nib - 10)); + } + } + writeByte('\r'); + writeByte('\n'); +} + +/// A u32 as eight hex digits, reading no memory at all. +pub fn dumpWord(v: u32) void { + var shift: u5 = 28; + while (true) { + const nib: u8 = @intCast((v >> shift) & 0xf); + writeByte(if (nib < 10) '0' + nib else 'a' + (nib - 10)); + if (shift == 0) break; + shift -= 4; + } + writeByte('\r'); + writeByte('\n'); +} + +/// Move whatever the host has sent into `buf`, without waiting. Returns the count. +/// +/// Non-blocking on purpose: the loop has a frame to render and a core to pump, and the editor must +/// not stall on a keystroke that may never come. `rxCount` is read once per call and the FIFO +/// drained to that mark, so a fast typist or a pasted buffer cannot hold the loop here. +pub fn read(buf: []u8) usize { + // RESCUED BYTES FIRST. They arrived before anything still sitting in the FIFO, and an editor + // that reorders keystrokes is worse than one that drops them. + var n = rescued.pop(buf); + const waiting = @min(uart0.rxCount(), buf.len - n); + for (buf[n..][0..waiting]) |*slot| slot.* = uart0.popByte(); + n += waiting; + return n; +} + +/// Take whatever has arrived off the receiver right now, without waiting and without handing it to +/// anyone. For callers that are about to spend a while not reading: `write` does this while the +/// transmitter is full, and the loop does it between chunks of input, because applying a keystroke +/// gets more expensive as the line grows and 128 bytes of FIFO is only 11 ms at 115200. +pub fn rescueNow() void { + input_rescue.rescue(uart0, &rescued); +} + +/// Input abandoned because even the rescue buffer overflowed. Distinct from `dropped`, which is +/// OUTPUT abandoned by a stalled transmitter. +pub fn inputDropped() u32 { + return rescued.dropped; +} + +/// Discard anything already received, returning how much. Used once at startup: the host-side +/// bridge injects a window-size report before this program exists, and the bootloader's chatter has +/// already been echoed at the host. Neither is user input. +/// +/// Pops rather than calling `resetRxFifo`, which is a CONF0_SYNC read-modify-write plus two commits +/// on the console UART - see this file's header. +pub fn drainInput() u32 { + var discarded: u32 = 0; + while (uart0.rxCount() > 0) : (discarded += 1) _ = uart0.popByte(); + discarded += @intCast(rescued.len); + rescued.clear(); + return discarded; +} + +/// The rate the hardware is actually producing, by reading its dividers back. Reported rather than +/// assumed: the host has to be opened at the same rate, and a mismatch shows up as garbage on the +/// screen rather than as an error anyone can act on. +pub fn baudrate() u32 { + return uart0.baudrate(uart0.clockSource().nominalHz()); +} |
