summaryrefslogtreecommitdiff
path: root/src/esp32p4
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-26 13:27:46 -0300
committerGabriel Schneider <[email protected]>2026-08-27 09:47:39 -0300
commit11f380f6d7222f2cad93c2cdf13701ea1f903d47 (patch)
tree803194ee5853a6b4cda93f90a95e28d1f02e69ae /src/esp32p4
parentfbc194068687e49a8490c85c9f1257a2f2bb9079 (diff)
downloadpardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.tar.gz
pardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.zip
One core behind N frontends, the board's own runner moved in, and every board cap on one screen
## The wire is the effect stream, not a new protocol `pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns a terminal and a socket and nothing else. N frontends on one core all look at the same screen — `screen -x`, not N sessions. The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin command line, and a command line cannot carry a frame. ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit little-endian fixed width and no message is a blit of a native struct. A protocol that only works between two builds of the same compiler would have thrown away the one frontend that motivated it. ## The board comes in; its toolchain stays out `src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over- serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so `zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the board from this repo's `build.zig`. The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes core over a serial line. What stayed is everything a second project would also want — the HAL, the register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its own and its `build()` early-returns when it is not the root package, so this costs the package graph exactly zero packages and the editor's own builds nothing at all. ## limits.zig: nine forgettable places become one budget Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions — they are ONE decision, how much memory this build may spend, taken nine times where no reader could see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against what it is measured against, derived from two booleans. The payoff is testability on a machine that is not the board: the caps are ordinary comptime values, so a host build can be compiled against the board's numbers and the parking, eviction and clamping paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART. ## A bare `zig build` `zig build` with no arguments now builds the tty and GUI binaries and installs them into `~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and "use it" two different commands for no reason.
Diffstat (limited to 'src/esp32p4')
-rw-r--r--src/esp32p4/app.zig546
-rw-r--r--src/esp32p4/input_rescue.zig271
-rw-r--r--src/esp32p4/selftest.zig351
-rw-r--r--src/esp32p4/uart.zig165
4 files changed, 1333 insertions, 0 deletions
diff --git a/src/esp32p4/app.zig b/src/esp32p4/app.zig
new file mode 100644
index 00000000..a9cf627d
--- /dev/null
+++ b/src/esp32p4/app.zig
@@ -0,0 +1,546 @@
+//! pardes, as ESP32-P4 firmware: the reset entry, the heap, the clock, the trap handler and the
+//! loop.
+//!
+//! There is no operating system under this. `_start` is the reset entry the second-stage bootloader
+//! jumps to, and this file is the entire platform: a heap, a millisecond clock, and UART0.
+//!
+//! ## Why the firmware root is in the editor's repository
+//!
+//! It was written in the `05-zig-p4` toolchain repository, next to the SoC support it uses, and it
+//! moved here because everything in it is a statement about the EDITOR. The heap span it hands over
+//! is the number that decides how large a grid the board can drive; `input_chunk` is sized against
+//! what applying one keystroke costs in `src/pardes.zig`; the loop's shape - read, chunk, tick,
+//! render only when dirty - is this editor's loop and no one else's; and the `-Dprof` attribution
+//! exists to answer "where did the 34 ms of a keystroke go" about this program. A firmware root that
+//! specific to one application belongs beside it.
+//!
+//! What stayed behind is everything a second application would also want, and none of it is
+//! duplicated here: the SoC and HAL, the translate-c register layer, the coalescing heap, `std.Io`
+//! for this chip, the app descriptor, the generated linker script, the image builder, the flasher
+//! and the interactive console. Those arrive as the `zig_p4` dependency, and this file imports
+//! exactly four of its modules - `soc`, `hal`, `heap` and `config` - plus two sibling files,
+//! `uart.zig` and `input_rescue.zig`, which are the editor's own.
+//!
+//! ## Where the editor is
+//!
+//! On the far side of a C ABI, still, and that is a choice rather than a leftover. `src/esp32p4.zig` in
+//! this same repository is compiled as ONE freestanding object (`b.addObject`, rooted at that file)
+//! and linked in beside this one; the `extern` declarations below are the near side of that seam.
+//!
+//! Importing `esp32p4.zig` as a module instead would be shorter to write and worse in every way that
+//! matters. It would drag the core's whole module graph - vaxis, the themes, the allocator tiers -
+//! into this root, which is the compilation that must stay small enough to reason about. It would
+//! give the firmware two ways to reach the editor. And above all it would make the OBJECT path a
+//! second arrangement, tested separately: that path is what `05-zig-p4 -Dpardes -Dpardes-obj=...`
+//! builds, it is what every measurement in that repository's `experiments/` was taken through, and
+//! it is a supported way to build this board. With the extern kept, both builds link the same eight
+//! symbols against the same object file, so neither can drift and neither is the better-tested one.
+//! The reasons the seam is a file at all - a nested `build.zig.zon` dependency broke every build in
+//! the toolchain repository - are recorded in `src/esp32p4.zig:8-15` and `05-zig-p4/build.zig:238-260`.
+//!
+//! Who owns which symbol: `src/esp32p4.zig` exports all eight `pardes_esp32p4_*` functions and nothing else.
+//! This file exports `_start`, `zig_main`, `trapEntry` and `trapReport`. `esp_app_desc` belongs to
+//! neither and comes from the toolchain's own appdesc object, which the link adds unconditionally.
+//! `abi_version` below is the one constant both sides spell, and its counterpart is
+//! `src/esp32p4.zig:98` - one repository now, so a bump is two lines in one diff rather than two commits
+//! in two trees.
+//!
+//! The seam is deliberately **bytes in, bytes out**. Everything that needs to know what a cell is -
+//! vaxis, the ANSI encoder, the input parser, the capability handshake - lives on the far side,
+//! next to the vaxis it is built against. What crosses is a byte stream in each direction, which is
+//! exactly what a serial line is, so this file has no opinion about terminals at all.
+//!
+//! ## Where the memory is
+//!
+//! Measured on this die by the toolchain's `examples/memprobe.zig`, not read off a datasheet, and
+//! written down once in the generated linker script (`05-zig-p4/build.zig:1572,1579,1584-1585`):
+//!
+//! 0x4FF00000..0x4FF3F000 252 KiB `l2mem`: .data/.bss/.stack are linked into this
+//! 0x4FF3F000..0x4FF40000 4 KiB mask ROM .data/.bss - untouchable, ets_printf needs it
+//! 0x4FF40000..0x4FFA0000 384 KiB `l2high`: handed to the editor as its entire heap
+//! 0x4FFA0000..0x4FFC0000 128 KiB NOT memory - the L2 cache lives here
+//!
+//! That last line is why the heap is 384 KiB and not the 512 KiB an earlier version of this comment
+//! claimed. The first probe wrote a pattern and read it back one page at a time and reported the
+//! whole upper 512 KiB as RAM, because a store followed immediately by a load of the SAME address
+//! returns the stored value whether the backing store is real, an address mirror, or merely a dirty
+//! cache line. Writing every page before reading any page separates the three, and the top 128 KiB
+//! then failed; handing them to an allocator hung the heap on its first free-list walk. ESP-IDF's
+//! own arithmetic agrees exactly: SRAM_HIGH_SIZE = 0x80000 - CONFIG_CACHE_L2_CACHE_SIZE, with the
+//! Kconfig default of 128 KiB.
+//!
+//! The span arrives as `__heap_start`/`__heap_end` from that script, so those addresses are written
+//! down in exactly one place. The editor owns it outright: it is passed in at init and this file
+//! never allocates from it.
+//!
+//! PSRAM is not used. The board has 32 MB fitted and it would make all of this comfortable, but
+//! ESP-IDF's own ESP32-P4 implementation runs past a thousand lines - MPLL, MSPI clocking, pin
+//! drive and DQS, CS timing, mode registers, a connectivity check, and an entire timing-calibration
+//! subsystem - and the mask ROM offers only MMU mapping, no device init. Touching it untrained
+//! faults and hangs the core, which `examples/memprobe.zig` demonstrates on purpose.
+
+const std = @import("std");
+const soc = @import("soc");
+const config = @import("config");
+
+/// `-Dprof`: time the two phases of a keystroke on the board and print the cycle counts. A
+/// diagnostic, not a feature - see the loop.
+const prof = config.prof;
+
+/// Every byte this loop has taken off the UART, for `-Dprof`. Ground truth for "did the burst
+/// arrive", which a screen reconstruction cannot answer: a character can be missing from the screen
+/// because it never arrived, because the editor never applied it, or because the viewport does not
+/// show that column.
+var rx_total: u32 = 0;
+
+/// How many input bytes to hand the editor before draining the receiver again. Chosen against the
+/// FIFO rather than against the editor: applying one keystroke was measured at 44 us on an empty
+/// line and 63 us at 640 characters, so eight of them is at most ~0.5 ms in which nothing empties
+/// the receiver, against a 128-byte FIFO that holds 11 ms of wire at 115200. Twenty times the margin
+/// needed, and it costs nothing on the wire because one render still happens per loop iteration.
+const input_chunk = 8;
+const hal = @import("hal");
+const heapmod = @import("heap");
+const uart = @import("uart.zig");
+
+// ------------------------------------------------------------------------------------- the ABI
+// Eight functions, all `callconv(.c)`, all implemented in the linked object - `src/esp32p4.zig` in this
+// repository, compiled for the same target and exporting exactly these names. This is the complete
+// interface between this board and the editor, and it is deliberately bytes-and-memory only: the
+// editor never learns what a UART is, and this file never learns what a cell is.
+//
+// The declarations below are a SECOND spelling of the signatures in `src/esp32p4.zig:100-127,259-...`,
+// and that duplication is what a C ABI is: each side declares the wire independently, which is
+// precisely why `abi_version` has to be checked. Sharing a Zig type between them would mean sharing
+// a module, which would mean the core in this compilation - see the header.
+
+/// How the editor emits bytes. Called with finished runs of ANSI, many times per frame.
+const WriteFn = *const fn (ctx: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void;
+
+/// The board's pads, offered to the editor. Optional on the wire so a firmware with nothing to
+/// toggle passes null and the `Gpio` word reports that rather than the object guessing.
+const GpioFn = *const fn (ctx: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool;
+
+/// This board's allocator, handed across as plain function pointers. `log2_align` is a log2 value,
+/// which is exactly how `std.mem.Alignment` represents itself, so neither side needs a conversion
+/// table.
+///
+/// The memory belongs to THIS side: only the firmware knows that the heap is the 384 KiB at
+/// 0x4FF40000, that the 128 KiB above it is L2 cache, and that PSRAM is untrained. The editor gets
+/// an allocator, not an address range.
+const Allocator = extern struct {
+ ctx: ?*anyopaque,
+ alloc: *const fn (ctx: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8,
+ resize: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool,
+ free: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void,
+};
+
+/// The one number both sides must agree on. Linkers do not type-check C symbols, so a signature
+/// that drifts on one side of this seam links cleanly and then corrupts the stack; checking this
+/// before calling anything else turns that into a refusal to boot.
+const abi_version: u32 = 2;
+extern fn pardes_esp32p4_abi_version() callconv(.c) u32;
+
+/// Hand over the allocator and the output sink, and state the initial window size. Returns 0, or a
+/// small non-zero code this file can only report.
+extern fn pardes_esp32p4_init(
+ alloc: *const Allocator,
+ write: WriteFn,
+ gpio: ?GpioFn,
+ ctx: ?*anyopaque,
+ cols: u16,
+ rows: u16,
+) callconv(.c) u32;
+
+/// Raw bytes off the wire: keystrokes, capability-query replies, and the host bridge's in-band
+/// resize reports. The editor parses all three; this file distinguishes none of them.
+extern fn pardes_esp32p4_input(ptr: [*]const u8, len: usize) callconv(.c) void;
+
+/// Advance time. Separate from `input` because animations and timeouts must progress on a wire
+/// where nothing is arriving.
+extern fn pardes_esp32p4_tick(now_ms: u64) callconv(.c) void;
+
+/// Emit one frame through the write callback. Returns 0 or an error code.
+extern fn pardes_esp32p4_render() callconv(.c) u32;
+
+/// Is there anything to draw - a dirty surface or a running animation? Asked every iteration so a
+/// quiet editor costs no bytes on a 115200-baud link.
+extern fn pardes_esp32p4_wants_frame() callconv(.c) bool;
+
+/// Has the user asked to leave? There is nowhere to go, so this only stops the loop.
+extern fn pardes_esp32p4_quit() callconv(.c) bool;
+
+/// The last frame's three stages in CPU cycles: the copy of pardes's Surface into vaxis's grid,
+/// vaxis's own diff-and-emit, and the push into the UART. Only meaningful under `-Dprof`; the
+/// editor object always exports it, and it costs two CSR reads per stage.
+extern fn pardes_esp32p4_frame_prof(copy: *u64, render: *u64, flush: *u64) callconv(.c) void;
+
+// ------------------------------------------------------------------------------------ the sink
+
+/// The write callback handed to `pardes_esp32p4_init`. No context is needed - there is one UART.
+fn writeOut(_: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void {
+ uart.write(ptr[0..len]);
+}
+
+/// Flip one pad and report the level before and after. The editor's `Gpio` word calls this; the
+/// editor has no register of its own for it, deliberately.
+///
+/// THIS IS WHY THE SEAM IS HERE. A toggle is not a write to GPIO_OUT: `configureOutput` points the
+/// pad's IO MUX at the GPIO function, routes the GPIO matrix's output to it, sets the drive strength
+/// and input buffer and clears the pulls, and only then enables the driver - four register files,
+/// indexed by a per-pin table. That code already exists in the toolchain package's `src/hal/gpio.zig`,
+/// it is the same call that package's `src/main.zig` blinks with, and its register numbers are
+/// checked against ESP-IDF's own headers by `zig build diff` there. A second copy inside the editor
+/// object would be a second copy under no test.
+///
+/// `getDrivenLevel` rather than `getLevel`: the answer is the level this board is DRIVING, which is
+/// defined for every pin. The pad's own level is what the outside world says, and on an unconnected
+/// header pin that is noise. The input buffer is enabled anyway, so `Peek` of GPIO_IN_REG shows the
+/// pad for anyone who wants to compare the two.
+fn gpioToggle(_: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool {
+ if (pin > hal.gpio.max_pin) return false;
+ const p: u8 = @intCast(pin);
+ hal.gpio.configureOutput(p, .{ .readback = true });
+ const before = hal.gpio.getDrivenLevel(p);
+ if (before == 1) hal.gpio.setLow(p) else hal.gpio.setHigh(p);
+ was.* = before;
+ now.* = hal.gpio.getDrivenLevel(p);
+ return true;
+}
+
+// ------------------------------------------------------------------------------------- the heap
+
+/// The span the linker script hands over, from `l2high`'s ORIGIN and LENGTH.
+///
+/// Reached with `@extern`, NOT with `extern const __heap_start: anyopaque` plus
+/// `@intFromPtr`/`@ptrFromInt`. That spelling was here first and it was silently wrong: declaring a
+/// linker symbol as an `anyopaque` OBJECT gives the optimiser a zero-sized object, so a pointer
+/// derived from its address carries provenance for zero bytes, and ordinary (non-volatile) stores
+/// through it are dead code it may drop. The toolchain's `examples/heapcheck.zig` caught it on the
+/// die - the allocator's first block header read back as `size=2988759312 next=0xffffffff`-not, and
+/// the free list walk never terminated. A `[*]u8` from `@extern` has no size to lose.
+const heap_start = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_start" });
+const heap_end = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_end" });
+
+fn heapSpan() []align(heapmod.Heap.granule) u8 {
+ return heap_start[0 .. @intFromPtr(heap_end) - @intFromPtr(heap_start)];
+}
+
+/// The one heap. A K&R coalescing free list over that span, validated on this die by the toolchain's
+/// `examples/heapcheck.zig`: 512 blocks fill and free back to a single 393,216-byte block, a holed
+/// arena still satisfies a 4 KiB request, and 20,000 random operations drain back to one block.
+var gpa_heap: heapmod.Heap = undefined;
+
+// The four C forwarders the editor is handed. `log2_align` round-trips through
+// `std.mem.Alignment`, whose representation IS the log2 value.
+
+fn cAlloc(_: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8 {
+ const a = gpa_heap.allocator();
+ return a.vtable.alloc(a.ptr, len, @enumFromInt(log2_align), @returnAddress());
+}
+
+fn cResize(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool {
+ const a = gpa_heap.allocator();
+ return a.vtable.resize(a.ptr, ptr[0..len], @enumFromInt(log2_align), new_len, @returnAddress());
+}
+
+fn cFree(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void {
+ const a = gpa_heap.allocator();
+ a.vtable.free(a.ptr, ptr[0..len], @enumFromInt(log2_align), @returnAddress());
+}
+
+const editor_allocator: Allocator = .{
+ .ctx = null,
+ .alloc = cAlloc,
+ .resize = cResize,
+ .free = cFree,
+};
+
+// ------------------------------------------------------------------------------------ the clock
+
+/// Milliseconds since boot, off the systimer - a 16 MHz counter (the toolchain package's
+/// `src/hal/systimer.zig:31`), which is the cheapest trustworthy clock on this chip. `read` returns
+/// null if the unit is not running, in which case time simply does not advance and the editor stops
+/// animating; that is a better failure than a clock that jumps.
+fn nowMs() u64 {
+ const us = hal.systimer.micros(.unit0) orelse return 0;
+ return us / 1000;
+}
+
+// ------------------------------------------------------------------------------------- the loop
+
+export fn zig_main() noreturn {
+ // FIRST, before a single byte of `.rodata` is touched - which means before the marker below,
+ // because that marker IS a string literal in flash and would read as machine code without this.
+ soc.flushFlashCache();
+ const heap = heapSpan();
+ soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{
+ @as(u32, @intFromPtr(heap.ptr)),
+ @as(u32, @intFromPtr(heap.ptr)) + @as(u32, @intCast(heap.len)),
+ @as(u32, @intCast(heap.len / 1024)),
+ });
+
+ // The CPU clock, before anything is timed against it. The bootloader leaves 90 MHz and the
+ // CPLL is already at 360, so this is a divider change that disturbs neither UART0 (XTAL) nor
+ // the systimer (XTAL/2.5) nor the flash interface (SPLL). See the toolchain package's
+ // `src/hal/clkrst.zig:setCpuFreq`.
+ if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) {
+ 180 => .mhz180,
+ 360 => .mhz360,
+ else => .mhz90,
+ });
+
+ const rwdt_was_armed = hal.rwdt.disable();
+ hal.systimer.init();
+ _ = rwdt_was_armed;
+
+ const their_abi = pardes_esp32p4_abi_version();
+ if (their_abi != abi_version) {
+ uart.write("MARK PARDES_ABI_MISMATCH\r\n");
+ while (true) {}
+ }
+
+ gpa_heap = heapmod.Heap.init(heap);
+ _ = uart.drainInput();
+
+ // Ask for more than any grid this board will ever render, so the SHELL's own ceiling is what
+ // governs - it clamps to `-Desp32p4-cols`/`-Desp32p4-rows` and reports the result. Naming 80x24 here made
+ // the firmware a second opinion about the geometry, which is one opinion too many.
+ const rc = pardes_esp32p4_init(&editor_allocator, writeOut, gpioToggle, null, 255, 255);
+
+ if (rc != 0) {
+ soc.rom.print("MARK PARDES_INIT_FAIL rc=%u\r\n", .{rc});
+ const s = gpa_heap.stats();
+ soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{
+ s.free, s.largest_free, s.free_blocks,
+ });
+ while (true) {}
+ }
+
+ // The HEAP, after the editor has taken what it needs. This is the number that decides how large
+ // a grid the board can drive, so it is printed on every boot rather than only on failure: a
+ // geometry that fits with 2 KB to spare and one that fits with 80 KB are not the same answer,
+ // and the difference is invisible from the host otherwise.
+ {
+ const s = gpa_heap.stats();
+ soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{
+ s.free, s.largest_free, s.free_blocks,
+ });
+ }
+
+ // The CPU clock, measured rather than assumed. Every cycle count this firmware reports is
+ // divided by it somewhere, and the toolchain's `src/io/chip.zig` records it as "a measured
+ // ~90 MHz" that nothing here reconfigures - so it is worth printing rather than remembering. The
+ // systimer is XTAL/2.5 = 16 MHz and is NOT derived from the CPU clock
+ // (`src/hal/systimer.zig:31`, `clk_tree_defs.h:196-198`), which is exactly what makes it a valid
+ // reference for measuring it.
+ if (prof) {
+ const t_start = hal.systimer.micros(.unit0) orelse 0;
+ const c_start = soc.cycles();
+ // 50 ms is long enough that the systimer's 16 MHz granularity and the loop's own overhead
+ // are both noise, and short enough to be invisible in a boot.
+ while ((hal.systimer.micros(.unit0) orelse 0) -% t_start < 50_000) {}
+ const elapsed_us = (hal.systimer.micros(.unit0) orelse 0) -% t_start;
+ const elapsed_cy = soc.cycles() - c_start;
+ soc.rom.print("MARK CPU_HZ cycles=%u us=%u khz=%u\r\n", .{
+ @as(u32, @intCast(elapsed_cy)),
+ @as(u32, @intCast(elapsed_us)),
+ @as(u32, @intCast(if (elapsed_us > 0) elapsed_cy * 1000 / elapsed_us else 0)),
+ });
+ }
+ soc.rom.print("MARK PARDES_READY\r\n", .{});
+
+ var in: [256]u8 = undefined;
+ while (!pardes_esp32p4_quit()) {
+ // ATTRIBUTION. The host can time a keystroke's round trip but cannot see what the firmware
+ // spent it on, and the two candidates - parsing and editing, versus rendering - want
+ // opposite fixes. `soc.cycles()` is the unprivileged cycle counter, so this costs two CSR
+ // reads per phase and quantises at one cycle, which is four orders of magnitude below the
+ // milliseconds being attributed. Gated on `prof` so the shipping build carries none of it.
+ const n = uart.read(&in);
+ rx_total +%= @intCast(n);
+
+ var input_cy: u64 = 0;
+ if (n > 0) {
+ const t0 = if (prof) soc.cycles() else 0;
+ // IN CHUNKS, rescuing the receiver between them. Applying a keystroke is not free and
+ // gets dearer as the line grows - measured at 44 us on an empty line and 63 us at 640
+ // characters - so handing over a full 128-byte batch is up to 8 ms in which nothing
+ // drains the receiver, against a FIFO that holds only 11 ms of wire. A 600-byte paste
+ // lost 93 bytes to exactly that window even with the transmitter's own rescue in place.
+ //
+ // Splitting a burst at an arbitrary byte is safe: `pardes_esp32p4_input` keeps whatever it
+ // could not parse, which is how it already survives an escape sequence split across two
+ // UART reads. One render still happens per loop iteration, so this costs no extra wire.
+ var off: usize = 0;
+ while (off < n) {
+ const chunk = @min(input_chunk, n - off);
+ pardes_esp32p4_input(in[off..].ptr, chunk);
+ off += chunk;
+ if (off < n) uart.rescueNow();
+ }
+ if (prof) input_cy = soc.cycles() - t0;
+ }
+
+ pardes_esp32p4_tick(nowMs());
+
+ // Only when there is something to show. On a link this slow an unconditional repaint per
+ // iteration would saturate the wire and starve input.
+ if (pardes_esp32p4_wants_frame()) {
+ const t0 = if (prof) soc.cycles() else 0;
+ const err = pardes_esp32p4_render();
+ if (err != 0) soc.rom.print("MARK PARDES_RENDER_FAIL rc=%u\r\n", .{err});
+ if (prof) {
+ const render_cy = soc.cycles() - t0;
+ // A SECOND render with nothing changed since the first. It splits the cost in two:
+ // whatever this still costs is the price of walking and diffing the whole editor
+ // state, paid regardless of output, while the difference between the two is the
+ // price of the change itself. `wants_frame` is false now, so this only happens
+ // under -Dprof and never on a shipping build.
+ const t1 = soc.cycles();
+ _ = pardes_esp32p4_render();
+ const idle_cy = soc.cycles() - t1;
+ // Reported in cycles, not microseconds: the divisor is the CPU clock, which this
+ // firmware does not set and has only ever measured, so converting here would bake a
+ // guess into the data. The toolchain's `experiments/` divides by the clock it
+ // measured.
+ var copy_cy: u64 = 0;
+ var vx_cy: u64 = 0;
+ var flush_cy: u64 = 0;
+ pardes_esp32p4_frame_prof(&copy_cy, &vx_cy, &flush_cy);
+ soc.rom.print("PROF in=%u render=%u idle=%u copy=%u vaxis=%u flush=%u rx=%u rxdrop=%u txdrop=%u\r\n", .{
+ @as(u32, @intCast(input_cy)),
+ @as(u32, @intCast(render_cy)),
+ @as(u32, @intCast(idle_cy)),
+ @as(u32, @intCast(copy_cy)),
+ @as(u32, @intCast(vx_cy)),
+ @as(u32, @intCast(flush_cy)),
+ rx_total,
+ uart.inputDropped(),
+ uart.dropped,
+ });
+ }
+ }
+ }
+
+ soc.rom.print("\r\nMARK PARDES_QUIT\r\n", .{});
+ while (true) {}
+}
+
+// ------------------------------------------------------------------------------------ the trap
+
+/// A trap handler, because the absence of one is why this port has been guessing.
+///
+/// The mask ROM prints "Guru Meditation" for a trap only while ITS handler is still installed;
+/// anything this image does that replaces or outgrows that path fails silently instead, and a silent
+/// fault is indistinguishable from an infinite loop over a serial line. This one reports the three
+/// registers that name the fault and then stops, using the direct-FIFO writer so it shares nothing
+/// with the editor's buffered output.
+///
+/// `mtvec` is set in DIRECT mode (low two bits zero), so every trap and every interrupt lands on
+/// `trapEntry` regardless of cause - which is what a diagnostic wants.
+export fn trapEntry() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile ("j trapReport");
+}
+
+export fn trapReport() noreturn {
+ const mcause = asm volatile ("csrr %[o], mcause"
+ : [o] "=r" (-> u32),
+ );
+ const mepc = asm volatile ("csrr %[o], mepc"
+ : [o] "=r" (-> u32),
+ );
+ const mtval = asm volatile ("csrr %[o], mtval"
+ : [o] "=r" (-> u32),
+ );
+ uart.write("\r\nMARK TRAP mcause=");
+ uart.dumpWord(mcause);
+ uart.write("MARK TRAP mepc=");
+ uart.dumpWord(mepc);
+ uart.write("MARK TRAP mtval=");
+ uart.dumpWord(mtval);
+ uart.write("MARK TRAP dropped=");
+ uart.dumpWord(uart.dropped);
+ while (true) {}
+}
+
+// --------------------------------------------------------------------------- the root's own duties
+//
+// These are the FIRMWARE root's declarations, and they are not the same set as `src/esp32p4.zig`'s: that
+// file is the root of its own object and carries its own `std_options` and `panic` for the core's
+// half of the image. Two roots, two instantiations of std, one per compilation unit - which is
+// exactly what the object seam buys, and why a panic in the core prints `PARDES_CORE_PANIC` through
+// the write callback while a panic here prints `PARDES_PANIC` through the mask ROM.
+
+/// `page_size_min`/`max`: the board has no MMU and no pages, but std derives allocator alignment
+/// from these. 4 KiB is the ESP32-P4's cache and DMA granularity.
+///
+/// `logFn` is not cosmetic. std's default log implementation reaches `std.debug_io`, which
+/// instantiates `std.Io.Threaded` - a thread pool, `getrandom`, `IOV_MAX`, `mremap` - none of which
+/// exist here, and one `log.warn` from anywhere is enough to drag all of it into the image.
+pub const std_options: std.Options = .{
+ .page_size_min = 4096,
+ .page_size_max = 4096,
+ .logFn = logFn,
+};
+
+fn logFn(
+ comptime level: std.log.Level,
+ comptime scope: @EnumLiteral(),
+ comptime fmt: []const u8,
+ args: anytype,
+) void {
+ var buf: [256]u8 = undefined;
+ const line = std.fmt.bufPrint(&buf, "\r\n[" ++ level.asText() ++ "/" ++ @tagName(scope) ++ "] " ++ fmt ++ "\r\n", args) catch
+ "\r\n[log overflow]\r\n";
+ uart.write(line);
+}
+
+pub const panic = std.debug.FullPanic(panicImpl);
+
+fn panicImpl(msg: []const u8, first_trace_addr: ?usize) noreturn {
+ // The fixed text goes out through the ROM deliberately: a panic may BE the console writer
+ // failing, and `ets_printf` shares nothing with `uart.write` except the FIFO itself.
+ //
+ // The MESSAGE does not, and that is a correction rather than a preference. `msg` is a Zig SLICE
+ // and `%s` reads until a NUL, so handing `msg.ptr` to printf prints the message and then
+ // whatever happens to sit after it in memory until a zero byte turns up. Literals get away with
+ // it; std's own panics do not, because they are formatted into a buffer - "index out of bounds:
+ // index 5, len 3" - and carry no terminator. `uart.write` takes a length.
+ soc.rom.print("\r\nMARK PARDES_PANIC ", .{});
+ uart.write(msg);
+ // The address is what makes it actionable: addr2line against the ELF in zig-out turns it into a
+ // source line, and without it a panic message names a KIND of failure with no way to find which
+ // one of them happened. Zero when the caller had no return address to give.
+ soc.rom.print("\r\nMARK PARDES_PANIC_AT 0x%08x\r\n", .{@as(u32, @truncate(first_trace_addr orelse 0))});
+ while (true) {}
+}
+
+/// Reset entry. The bootloader hands over with an unspecified stack pointer and the FPU off, so:
+/// enable the F extension (`mstatus.FS`, which ESP-IDF only ever turns on lazily from a trap handler
+/// this image does not have), establish a stack, clear `.bss`, and call into Zig.
+///
+/// The cache invalidate that this image also needs is the FIRST thing `zig_main` does, not something
+/// done here. Hand-written `la t0, Cache_Invalidate_All` against an absolute linker symbol computed
+/// a PC-relative target and jumped into nowhere (measured: PC=0x88b5d788 with the argument stranded
+/// in a2); Zig generates the addressing for an `extern fn` correctly, and `zig_main` runs before any
+/// `.rodata` is touched anyway.
+export fn _start() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile (
+ \\ li t0, 1 << 13
+ \\ csrs mstatus, t0
+ \\ la sp, __stack_top
+ \\ mv fp, sp
+ \\ la t0, trapEntry
+ \\ csrw mtvec, t0
+ \\ la t0, __bss_start
+ \\ la t1, __bss_end
+ \\ bgeu t0, t1, 2f
+ \\1:
+ \\ sw zero, 0(t0)
+ \\ addi t0, t0, 4
+ \\ bltu t0, t1, 1b
+ \\2:
+ \\ j zig_main
+ );
+}
diff --git a/src/esp32p4/input_rescue.zig b/src/esp32p4/input_rescue.zig
new file mode 100644
index 00000000..01dd4820
--- /dev/null
+++ b/src/esp32p4/input_rescue.zig
@@ -0,0 +1,271 @@
+//! Keystrokes rescued from the receive FIFO while the transmitter is busy.
+//!
+//! THE BUG THIS EXISTS FOR. The firmware's loop is read, apply, render, write, and the write blocks
+//! while the transmit FIFO is full - real backpressure, because dropping half an escape sequence
+//! would leave the host terminal in the wrong colour for the rest of the session. But nothing
+//! drained the RECEIVE FIFO during that wait, and the FIFO is 128 bytes (the toolchain package's
+//! `src/hal/uart.zig:52`). A frame of 240 bytes is 21 ms of wire at 115200, and 21 ms of a host
+//! sending at line rate is ~240 bytes, so everything past the 128th was silently gone.
+//!
+//! Measured on the die before the fix, typing a burst in one host write and counting what the editor
+//! actually held: 128 bytes arrived intact, 200 bytes lost 88, 300 bytes lost all 300. From a
+//! keyboard that is a keystroke that never lands, and it looks like a stuck key - the screen is
+//! behind what was typed, and typing more appears to fix it because a later frame repaints the cells
+//! the lost keystrokes would have changed.
+//!
+//! WHY THE POLICY LIVES HERE and not in `uart.zig`: the interesting part is a decision - drain the
+//! receiver while spinning on the transmitter, and what to do when even that overflows - and the
+//! decision is worth testing. `uart.zig` cannot be tested at all without the chip, because every
+//! line of it is an MMIO access. `pump` takes the port as `anytype`, so the same code runs against
+//! the real UART on the board and against a fake with a two-byte FIFO on the host.
+//!
+//! WHY IT IS IN THIS REPOSITORY. It was written in the toolchain repository, next to the UART it
+//! spins on, and it moved here because every number in it is the EDITOR's. 4 KiB is sized against
+//! the ~1.4 KB full repaint `src/esp32p4.zig` emits; "drop the newest, so what survives is a PREFIX of
+//! what was typed" is a statement about documents rather than about UARTs, and it is the editor that
+//! would otherwise appear to invent input. A policy whose every constant comes from one application
+//! belongs beside that application.
+//!
+//! It expects NOTHING from the toolchain package - no module, no register, no target. `std` is the
+//! whole import list, which is what makes the host tests at the bottom possible and what lets the
+//! same source compile for riscv32-freestanding and for the host unchanged.
+//!
+//! ONE COPY, TWO BUILDS. This file is the only copy; the toolchain repository's is gone. Its build
+//! reads this tree across a sibling-relative seam, and names this path twice: once as the
+//! `input_rescue` module of `-Dpardes`'s application (`05-zig-p4/build.zig:230-231`, whose root is
+//! `../02-pardes-code/src/esp32p4/app.zig` by the `-Dapp` default at `:168-169`) and once as the
+//! same-named module of `zig build selftest`'s on-die root (`:437-438`). Both spell
+//! `../02-pardes-code/src/esp32p4/input_rescue.zig`, so there is nothing to keep in step.
+//!
+//! Tested from here, both ways: the host checks at the bottom run as their own `addTest` under
+//! `zig build unit-test` (`02-pardes-code/build.zig:1727-1732` - no `link_libc`, because `std` is
+//! the whole import list), and the same source runs against UART0 on the die under
+//! `zig build esp32p4-test`.
+
+const std = @import("std");
+
+/// Capacity, sized for the worst frame this editor emits.
+///
+/// A full repaint is ~1.4 KB, which is 121 ms of wire at 115200, and 121 ms of a host pasting at
+/// line rate is ~1.4 KB of input. 4 KiB is that with headroom, a power of two so the wrap is a mask
+/// rather than a division, and nothing at all against the board's RAM.
+pub const capacity = 4096;
+
+/// A byte queue that drops the NEWEST byte when full.
+///
+/// Dropping the newest rather than the oldest is deliberate: what survives is then a PREFIX of what
+/// was typed. An editor that loses the end of a paste has done something a person can see and
+/// correct; one that silently reorders keystrokes, or keeps the tail and discards the head, has
+/// corrupted the document in a way that looks like the editor inventing input.
+pub const Ring = struct {
+ buf: [capacity]u8 = undefined,
+ head: usize = 0,
+ len: usize = 0,
+ /// Bytes lost because even this overflowed. Nonzero means input was dropped; it is the honest
+ /// version of the bug rather than a cure for it.
+ dropped: u32 = 0,
+
+ const mask = capacity - 1;
+
+ comptime {
+ std.debug.assert(capacity & mask == 0);
+ }
+
+ pub fn push(r: *Ring, b: u8) void {
+ if (r.len == capacity) {
+ r.dropped +%= 1;
+ return;
+ }
+ r.buf[(r.head + r.len) & mask] = b;
+ r.len += 1;
+ }
+
+ /// Move as much as fits into `out`, oldest first. Returns the count.
+ pub fn pop(r: *Ring, out: []u8) usize {
+ const n = @min(out.len, r.len);
+ for (out[0..n]) |*slot| {
+ slot.* = r.buf[r.head];
+ r.head = (r.head + 1) & mask;
+ }
+ r.len -= n;
+ return n;
+ }
+
+ pub fn clear(r: *Ring) void {
+ r.head = 0;
+ r.len = 0;
+ }
+};
+
+/// Drain everything the port has received into `ring`, without waiting.
+pub fn rescue(port: anytype, ring: *Ring) void {
+ var waiting = port.rxCount();
+ while (waiting > 0) : (waiting -= 1) ring.push(port.popByte());
+}
+
+/// Push `bytes` through `port`, rescuing input whenever the transmitter has no room. Returns the
+/// number of bytes abandoned because the transmitter stopped making progress altogether.
+///
+/// The spin bound is why this returns a count rather than blocking forever: a UART whose core clock
+/// has been gated never makes progress, and on a board with no debugger an infinite spin is
+/// indistinguishable from a crash. A bounded wait turns that into visibly dropped output plus a
+/// counter, which is a diagnosis instead of a mystery.
+pub fn pump(port: anytype, ring: *Ring, bytes: []const u8, spin_limit: u32) u32 {
+ var rest = bytes;
+ while (rest.len > 0) {
+ // One status read per burst, not per byte: reading `txFree` once and pushing that many cuts
+ // the status reads by up to the FIFO depth.
+ var room = port.txFree();
+ var spins: u32 = 0;
+ while (room == 0) {
+ // THE FIX. Every iteration of this wait is time the receiver is filling up, and this is
+ // the only place that can empty it.
+ rescue(port, ring);
+ spins += 1;
+ if (spins > spin_limit) return @intCast(rest.len);
+ room = port.txFree();
+ }
+ const n = @min(room, rest.len);
+ for (rest[0..n]) |b| port.pushByte(b);
+ rest = rest[n..];
+ }
+ return 0;
+}
+
+// ------------------------------------------------------------------------------------ host tests
+
+test "the ring hands bytes back in order" {
+ var r: Ring = .{};
+ for ("hello") |b| r.push(b);
+ var out: [8]u8 = undefined;
+ try std.testing.expectEqual(@as(usize, 5), r.pop(&out));
+ try std.testing.expectEqualStrings("hello", out[0..5]);
+ try std.testing.expectEqual(@as(usize, 0), r.pop(&out));
+}
+
+test "the ring wraps without reordering" {
+ var r: Ring = .{};
+ var out: [capacity]u8 = undefined;
+ // Push and pop most of the buffer so head sits near the end, then straddle the wrap.
+ for (0..capacity - 3) |i| r.push(@intCast(i & 0xff));
+ _ = r.pop(out[0 .. capacity - 3]);
+ for ("straddle") |b| r.push(b);
+ const n = r.pop(&out);
+ try std.testing.expectEqualStrings("straddle", out[0..n]);
+}
+
+test "a full ring drops the newest and says so" {
+ var r: Ring = .{};
+ for (0..capacity) |i| r.push(@intCast(i & 0xff));
+ try std.testing.expectEqual(@as(u32, 0), r.dropped);
+ r.push('!');
+ r.push('!');
+ try std.testing.expectEqual(@as(u32, 2), r.dropped);
+ // The head is intact: what survived is a prefix of what arrived.
+ var out: [4]u8 = undefined;
+ _ = r.pop(&out);
+ try std.testing.expectEqual(@as(u8, 0), out[0]);
+ try std.testing.expectEqual(@as(u8, 1), out[1]);
+}
+
+/// A UART with a small transmit FIFO, a small RECEIVE FIFO, and a host that keeps typing into it.
+///
+/// The receive FIFO is the part that matters and it is modelled the way the hardware behaves: it has
+/// a fixed depth, and a byte that arrives when it is full is *gone*. That is the whole bug.
+///
+/// Time advances on each transmitter status read, which is what `pump` does while it waits. The
+/// transmitter frees a byte only every fourth tick while a typed byte lands on every one: the
+/// transmitter therefore genuinely FILLS, which is the condition the bug needs. A fake whose FIFO
+/// drains as fast as it fills never blocks, so `pump` never waits, so the rescue never runs and the
+/// test proves nothing - the first version of this fake had exactly that flaw.
+const FakePort = struct {
+ tx_cap: u32,
+ tx_used: u32 = 0,
+ sent: std.ArrayList(u8) = .empty,
+ gpa: std.mem.Allocator,
+
+ incoming: []const u8,
+ delivered: usize = 0,
+ rx: [rx_depth]u8 = undefined,
+ rx_head: usize = 0,
+ rx_len: usize = 0,
+ /// Bytes the wire delivered into a full receive FIFO. The hardware has no counter for this,
+ /// which is exactly why the bug was invisible.
+ lost: u32 = 0,
+
+ ticks: u32 = 0,
+
+ const rx_depth = 8;
+ const tx_drain_every = 4;
+
+ fn tick(p: *FakePort) void {
+ p.ticks += 1;
+ if (p.ticks % tx_drain_every == 0 and p.tx_used > 0) p.tx_used -= 1;
+ if (p.delivered < p.incoming.len) {
+ const b = p.incoming[p.delivered];
+ p.delivered += 1;
+ if (p.rx_len == rx_depth) {
+ p.lost += 1;
+ } else {
+ p.rx[(p.rx_head + p.rx_len) % rx_depth] = b;
+ p.rx_len += 1;
+ }
+ }
+ }
+
+ fn txFree(p: *FakePort) u32 {
+ p.tick();
+ return p.tx_cap - p.tx_used;
+ }
+
+ fn pushByte(p: *FakePort, b: u8) void {
+ p.sent.append(p.gpa, b) catch unreachable;
+ p.tx_used += 1;
+ }
+
+ fn rxCount(p: *FakePort) u32 {
+ return @intCast(p.rx_len);
+ }
+
+ fn popByte(p: *FakePort) u8 {
+ const b = p.rx[p.rx_head];
+ p.rx_head = (p.rx_head + 1) % rx_depth;
+ p.rx_len -= 1;
+ return b;
+ }
+};
+
+test "a long transmit does not lose the input that arrives during it" {
+ // THE REGRESSION. Delete the `rescue` call inside `pump`'s wait and this fails: the receive FIFO
+ // is eight bytes deep, the typing below is far longer than that, and every byte that arrives
+ // into a full FIFO is gone with nothing to record it. That is the die's 88-of-200 in miniature.
+ const typed = "the quick brown fox jumps over the lazy dog, twice over, and then some more";
+ var port: FakePort = .{ .tx_cap = 2, .incoming = typed, .gpa = std.testing.allocator };
+ defer port.sent.deinit(std.testing.allocator);
+ var ring: Ring = .{};
+
+ const frame = "\x1b[1;1H" ++ "x" ** 400;
+ try std.testing.expectEqual(@as(u32, 0), pump(&port, &ring, frame, 1_000_000));
+
+ // Every output byte went out, in order.
+ try std.testing.expectEqualStrings(frame, port.sent.items);
+ // Nothing the wire delivered was dropped, by the FIFO or by the ring.
+ try std.testing.expectEqual(@as(u32, 0), port.lost);
+ try std.testing.expectEqual(@as(u32, 0), ring.dropped);
+ // And what was rescued, plus whatever is still sitting in the FIFO, is exactly what was typed -
+ // in order, which is the other half of the contract.
+ var got: [capacity]u8 = undefined;
+ var n = ring.pop(&got);
+ while (port.rxCount() > 0) : (n += 1) got[n] = port.popByte();
+ try std.testing.expectEqualStrings(typed[0..port.delivered], got[0..n]);
+ try std.testing.expect(port.delivered == typed.len);
+}
+
+test "a transmitter that never drains gives up and reports what it abandoned" {
+ var port: FakePort = .{ .tx_cap = 0, .incoming = "", .gpa = std.testing.allocator };
+ defer port.sent.deinit(std.testing.allocator);
+ var ring: Ring = .{};
+ // tx_cap 0 means txFree is always 0, so no byte can ever go out.
+ try std.testing.expectEqual(@as(u32, 5), pump(&port, &ring, "abcde", 32));
+ try std.testing.expectEqual(@as(usize, 0), port.sent.items.len);
+}
diff --git a/src/esp32p4/selftest.zig b/src/esp32p4/selftest.zig
new file mode 100644
index 00000000..0d692d98
--- /dev/null
+++ b/src/esp32p4/selftest.zig
@@ -0,0 +1,351 @@
+//! pardes's own test suite for the ESP32-P4, which runs ON THE DIE.
+//!
+//! WHY THIS EXISTS. Every serious bug this port has produced was invisible to a host test, and two
+//! of them were invisible for months. `std.mem.eql` compares a byte at a time on this target and
+//! several times faster on the host, so the firmware's largest read was three times slower than it
+//! needed to be and nothing on a laptop could tell. A lone ESC resolves to the Escape key, which is
+//! right when a kernel hands over a whole escape sequence and wrong when a 115200 line hands over
+//! one byte every 87 microseconds. A full transmit FIFO stopped anything draining the receiver, and
+//! the FIFO depth is a hardware number. None of those is a logic error you can reason your way to
+//! from a host: they are properties of THIS chip, THIS clock and THIS wire.
+//!
+//! So the checks below are chosen by one rule: a check belongs here only if the die can answer it
+//! and a host cannot. Anything that is pure logic - the ring's wrap-around, the mouse coalescer's
+//! ordering - already has a deterministic host test in `zig build unit-test`, which is faster, needs
+//! no hardware, and is where such a thing belongs. Duplicating those here would make this suite
+//! longer and no stronger.
+//!
+//! ## Why the suite is in the editor's repository
+//!
+//! It was written in the `05-zig-p4` toolchain repository as `examples/selftest.zig`, and it was
+//! never an example of anything. Re-read the paragraph above: every claim it makes is a claim about
+//! THIS program on this board. The word-at-a-time `std.mem.eql` it measures against is the reason
+//! the editor's frame diff compares rows a `u32` at a time; the lone-ESC decode and the
+//! transmit-FIFO backpressure are `input_rescue.zig`, a sibling file in this directory; the heap
+//! span and the allocator churn are what decide how large a grid this board can drive; and the
+//! clock check is the divisor under every `-Dprof` cycle count the editor reports. A suite that
+//! specific to one application belongs beside it, so that changing a source and changing the check
+//! which defends it are one diff in one repository rather than two commits in two trees.
+//!
+//! What stayed behind in the toolchain is the half that is about the CHIP: the image builder's rules,
+//! the register layer's field arithmetic, the console bridge's escape matcher. Those still run under
+//! `zig build test` over there and need no board.
+//!
+//! ## What it expects from the toolchain package
+//!
+//! Four of its modules and no more: `soc` (mask-ROM printf, the cycle counter), `hal` (UART0, the
+//! systimer, clock/reset), `config` (this build's `cpu_mhz`) and `heap` (the coalescing allocator).
+//! They arrive as the `zig_p4` dependency, which also supplies what makes this an image at all -
+//! the generated linker script, `ENTRY(_start)` and the app descriptor via `firmware(...).attach`,
+//! then `ImageStep`, `FlashStep` and `SelftestStep`.
+//!
+//! `input_rescue` is the fifth import and is NOT that package's: it is the sibling file in this
+//! directory, handed over as a named MODULE rather than imported as a path. Deliberately so - the
+//! `pub` on `FakePort`'s methods below is what lets that module reach them by duck typing across the
+//! boundary, and both builds, this repository's `esp32p4-test` and the toolchain's `selftest`, wire
+//! the identical root the identical way. One file, one arrangement, nothing to drift.
+//!
+//! Not a `zig test` binary, deliberately. Zig's test runner wants an OS, and `std.testing.allocator`
+//! is a debug allocator over the page allocator, which on freestanding is either a compile error or
+//! a lie. A hand-rolled harness is thirty lines and answers to nobody.
+//!
+//! Run with: zig build esp32p4-test -Dplatform=esp32p4 -Desp32p4-firmware (from here)
+//! or: zig build selftest (from ../05-zig-p4)
+
+const std = @import("std");
+const soc = @import("soc");
+const hal = @import("hal");
+const config = @import("config");
+const heapmod = @import("heap");
+const input_rescue = @import("input_rescue");
+
+/// Reset entry. Identical in shape to `src/main.zig`'s and for the same reasons: the bootloader hands
+/// over with an unspecified stack pointer and the FPU off, so set `mstatus.FS`, establish a stack,
+/// clear `.bss`, and jump.
+///
+/// Leaving this out is what made the first draft of this file unbuildable, and the failure said
+/// nothing useful: `-fentry=_start` found no such symbol, `--gc-sections` then discarded every
+/// function as unreachable, and the image builder reported `NotTwoMappedSegments` because
+/// `.flash.text` had nothing left in it. `zig build layout` now prints `entry 0x0` for exactly that,
+/// which is the same diagnosis in one line.
+export fn _start() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile (
+ \\ li t0, 1 << 13
+ \\ csrs mstatus, t0
+ \\ la sp, __stack_top
+ \\ mv fp, sp
+ \\ la t0, __bss_start
+ \\ la t1, __bss_end
+ \\ bgeu t0, t1, 2f
+ \\1:
+ \\ sw zero, 0(t0)
+ \\ addi t0, t0, 4
+ \\ bltu t0, t1, 1b
+ \\2:
+ \\ j zig_main
+ );
+}
+
+pub const panic = std.debug.FullPanic(struct {
+ fn call(msg: []const u8, first_trace_addr: ?usize) noreturn {
+ // `msg` is a slice and carries no terminator - std's own panics are formatted into a buffer -
+ // so it goes out with a length rather than through a `%s` that would read past the end of it.
+ soc.rom.print("\r\nMARK SELFTEST_PANIC at 0x%08x: ", .{@as(u32, @truncate(first_trace_addr orelse 0))});
+ hal.uart.Uart.init(0).write(msg);
+ // A panic is a FAILED RUN, and the host is watching for the summary line. Without this the
+ // run looks like a board that never answered, which is a different diagnosis entirely.
+ soc.rom.print("\r\nMARK SELFTEST DONE pass=%u fail=%u\r\n", .{ passed, failed + 1 });
+ while (true) {}
+ }
+}.call);
+
+var passed: u32 = 0;
+var failed: u32 = 0;
+
+/// One claim about the silicon. Printed either way: a suite that only speaks up when it fails gives
+/// no way to tell "all good" from "never ran", and on a board that difference matters.
+fn check(name: [*:0]const u8, ok: bool) void {
+ if (ok) {
+ passed += 1;
+ soc.rom.print("MARK SELFTEST ok %s\r\n", .{name});
+ } else {
+ failed += 1;
+ soc.rom.print("MARK SELFTEST FAIL %s\r\n", .{name});
+ }
+}
+
+/// Word-at-a-time equality, the same shape the editor's frame diff uses.
+fn sameBytesWordwise(a: []const u8, b: []const u8) bool {
+ if (a.len != b.len) return false;
+ if ((@intFromPtr(a.ptr) | @intFromPtr(b.ptr)) & 3 == 0) {
+ const n = a.len / 4;
+ const wa: [*]align(4) const u32 = @ptrCast(@alignCast(a.ptr));
+ const wb: [*]align(4) const u32 = @ptrCast(@alignCast(b.ptr));
+ for (wa[0..n], wb[0..n]) |x, y| {
+ if (x != y) return false;
+ }
+ return std.mem.eql(u8, a[n * 4 ..], b[n * 4 ..]);
+ }
+ return std.mem.eql(u8, a, b);
+}
+
+/// A UART with a receive FIFO that loses whatever arrives into a full one, as the hardware does.
+/// `pub` on the methods because `input_rescue` is a separate module here and reaches them by duck
+/// typing across it.
+const FakePort = struct {
+ tx_cap: u32,
+ tx_used: u32 = 0,
+ ticks: u32 = 0,
+ sent: u32 = 0,
+ incoming: []const u8,
+ delivered: usize = 0,
+ rx: [8]u8 = undefined,
+ rx_head: usize = 0,
+ rx_len: usize = 0,
+ lost: u32 = 0,
+
+ fn tick(p: *FakePort) void {
+ p.ticks += 1;
+ if (p.ticks % 4 == 0 and p.tx_used > 0) p.tx_used -= 1;
+ if (p.delivered < p.incoming.len) {
+ const byte = p.incoming[p.delivered];
+ p.delivered += 1;
+ if (p.rx_len == p.rx.len) {
+ p.lost += 1;
+ } else {
+ p.rx[(p.rx_head + p.rx_len) % p.rx.len] = byte;
+ p.rx_len += 1;
+ }
+ }
+ }
+ pub fn txFree(p: *FakePort) u32 {
+ p.tick();
+ return p.tx_cap - p.tx_used;
+ }
+ pub fn pushByte(p: *FakePort, _: u8) void {
+ p.sent += 1;
+ p.tx_used += 1;
+ }
+ pub fn rxCount(p: *FakePort) u32 {
+ return @intCast(p.rx_len);
+ }
+ pub fn popByte(p: *FakePort) u8 {
+ const byte = p.rx[p.rx_head];
+ p.rx_head = (p.rx_head + 1) % p.rx.len;
+ p.rx_len -= 1;
+ return byte;
+ }
+};
+
+/// Backing store for the heap checks. Static, because the point is to exercise the allocator on real
+/// L2MEM rather than to find out where a stack array happens to land.
+var heap_area: [64 * 1024]u8 align(16) = undefined;
+
+export fn zig_main() noreturn {
+ // THE CLOCK RAISE, performed here rather than assumed, which is what turns the frequency check
+ // at the end into a test of `setCpuFreq` instead of a tautology. The first version of this file
+ // read `config.cpu_mhz` and compared the die against it without ever setting it, so
+ // `-Dcpu-mhz=360` failed with `khz=90001 want=360000` - the check was right and the expectation
+ // was wrong. Same call, and the same order, as `src/esp32p4/app.zig`.
+ if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) {
+ 360 => .mhz360,
+ else => .mhz90,
+ });
+ hal.systimer.init();
+
+ soc.rom.print("\r\nMARK SELFTEST_START cpu_mhz=%u\r\n", .{@as(u32, config.cpu_mhz)});
+
+ // ------------------------------------------------ 1. the memory model the memory words assume
+ //
+ // `Peek`, `Poke` and `Hexdump` reach the bus through `*allowzero volatile` pointers and refuse an
+ // unaligned word. Both halves are claims about this core, and neither is checkable on a host.
+ {
+ const cell: *volatile u32 = @ptrCast(@alignCast(&heap_area[0]));
+ cell.* = 0xdeadbeef;
+ check("an aligned word round-trips through a volatile pointer", cell.* == 0xdeadbeef);
+
+ // Every byte offset in a word, readable and writable: this is what `Hexdump` does, and it is
+ // why `Hexdump` needs no alignment while `Peek` does.
+ var all_offsets_ok = true;
+ for (0..4) |i| {
+ const at: *volatile u8 = @ptrCast(&heap_area[16 + i]);
+ at.* = @intCast(0xa0 + i);
+ if (at.* != 0xa0 + @as(u8, @intCast(i))) all_offsets_ok = false;
+ }
+ check("a byte at every offset in a word round-trips", all_offsets_ok);
+
+ // THE VOLATILE PROMISE. Two reads of a running counter must be two reads. Were the optimiser
+ // allowed to fold them, `Peek` would print one value twice for a register that had changed,
+ // which is the one thing a memory word must never do.
+ const first = hal.systimer.micros(.unit0) orelse 0;
+ var spin: u32 = 0;
+ while (spin < 4000) : (spin += 1) asm volatile ("" ::: .{ .memory = true });
+ const second = hal.systimer.micros(.unit0) orelse 0;
+ check("two reads of a live counter are two reads", second != first);
+ check("and that counter runs forwards", second > first);
+ }
+
+ // --------------------------------------- 2. word-wise equality, on THIS instruction set
+ //
+ // The editor's frame diff compares rows a `u32` at a time because `std.mem.eql` compares a byte
+ // at a time here: 223 us against 66 for the same answer. "The same answer" is the part that has
+ // to hold on the target rather than on the host, so it is checked against `std.mem.eql` itself,
+ // at every difference position, at both alignments, and at lengths that are not multiples of 4.
+ {
+ var a: [64]u8 align(4) = undefined;
+ var b: [64]u8 align(4) = undefined;
+ for (&a, 0..) |*slot, i| slot.* = @intCast(i);
+ @memcpy(&b, &a);
+
+ var agree = true;
+ for (0..a.len) |len| {
+ if (sameBytesWordwise(a[0..len], b[0..len]) != std.mem.eql(u8, a[0..len], b[0..len])) agree = false;
+ }
+ check("aligned equality agrees with std.mem.eql at every length", agree);
+
+ agree = true;
+ for (0..a.len) |i| {
+ b[i] ^= 0xff;
+ if (sameBytesWordwise(&a, &b) != std.mem.eql(u8, &a, &b)) agree = false;
+ if (sameBytesWordwise(a[0..33], b[0..33]) != std.mem.eql(u8, a[0..33], b[0..33])) agree = false;
+ b[i] ^= 0xff;
+ }
+ check("a difference at any position is found, as std.mem.eql finds it", agree);
+
+ agree = true;
+ // UNALIGNED, which is why `sameBytes` tests alignment at RUNTIME: `Cell` is all u8 fields, so
+ // whether a row starts on a word boundary belongs to the allocator and not to the type.
+ for (1..4) |off| {
+ const ua = a[off..];
+ const ub = b[off..];
+ if (sameBytesWordwise(ua, ub) != std.mem.eql(u8, ua, ub)) agree = false;
+ b[off + 5] ^= 0xff;
+ if (sameBytesWordwise(ua, ub) != std.mem.eql(u8, ua, ub)) agree = false;
+ b[off + 5] ^= 0xff;
+ }
+ check("unaligned spans fall back and still agree", agree);
+ }
+
+ // ------------------------------------------------- 3. the rescue, on the real codegen
+ //
+ // The logic has a host test. What that cannot say is whether it behaves the same compiled for
+ // this core at this optimisation level, which is a question only a board answers.
+ {
+ const typed = "the quick brown fox jumps over the lazy dog, and then some more besides";
+ var port: FakePort = .{ .tx_cap = 2, .incoming = typed };
+ var ring: input_rescue.Ring = .{};
+ const frame = "\x1b[1;1H" ++ "x" ** 300;
+ const abandoned = input_rescue.pump(&port, &ring, frame, 1_000_000);
+
+ check("the frame went out whole", abandoned == 0 and port.sent == frame.len);
+ check("the receive FIFO never overflowed", port.lost == 0);
+ check("the ring dropped nothing", ring.dropped == 0);
+
+ var got: [128]u8 = undefined;
+ var n = ring.pop(&got);
+ while (port.rxCount() > 0 and n < got.len) : (n += 1) got[n] = port.popByte();
+ check("every rescued byte, in order", n == typed.len and std.mem.eql(u8, got[0..n], typed));
+ }
+
+ // --------------------------------------------------- 4. the allocator on real L2MEM
+ //
+ // The editor's whole geometry ceiling is an allocator question, and this is the allocator, on the
+ // memory it actually runs in rather than on a host's malloc.
+ {
+ var h = heapmod.Heap.init(heap_area[0..]);
+ const gpa = h.allocator();
+
+ const one = gpa.alloc(u32, 256) catch null;
+ check("a modest allocation succeeds", one != null);
+ if (one) |slice| {
+ check("and is aligned for its element", @intFromPtr(slice.ptr) % @alignOf(u32) == 0);
+ for (slice, 0..) |*slot, i| slot.* = @intCast(i * 7);
+ var intact = true;
+ for (slice, 0..) |slot, i| {
+ if (slot != i * 7) intact = false;
+ }
+ check("and holds what was written to it", intact);
+ gpa.free(slice);
+ }
+
+ // FREE THEN REUSE. A heap that cannot hand the same bytes back is a heap that runs out, which
+ // on this board is the difference between a 40x12 grid and an 80x24 one.
+ var churn_ok = true;
+ for (0..64) |_| {
+ const block = gpa.alloc(u8, 1024) catch {
+ churn_ok = false;
+ break;
+ };
+ gpa.free(block);
+ }
+ check("a kilobyte can be taken and returned repeatedly", churn_ok);
+
+ // AND IT REFUSES CLEANLY. An allocator that returns garbage instead of an error when it is
+ // out is the failure mode that cost an afternoon during bring-up.
+ const absurd = gpa.alloc(u8, heap_area.len * 4);
+ check("an impossible allocation returns an error", absurd == error.OutOfMemory);
+ }
+
+ // ------------------------------------------ 5. the clock, which everything divides by
+ //
+ // Every cycle count this firmware reports is divided by the configured frequency somewhere, and
+ // the systimer is clocked from the crystal rather than from the CPU - which is exactly what makes
+ // it a reference the CPU cannot flatter. The 90-to-360 MHz raise rested on this comparison.
+ {
+ const t0 = hal.systimer.micros(.unit0) orelse 0;
+ const c0 = soc.cycles();
+ while ((hal.systimer.micros(.unit0) orelse 0) -% t0 < 20_000) {}
+ const us = (hal.systimer.micros(.unit0) orelse 0) -% t0;
+ const cy = soc.cycles() - c0;
+ const khz: u32 = if (us > 0) @intCast(cy * 1000 / us) else 0;
+ const want: u32 = @as(u32, config.cpu_mhz) * 1000;
+ // Two percent: far wider than either clock's error, far narrower than the 4x a wrong divider
+ // would produce.
+ const slack = want / 50;
+ soc.rom.print("MARK SELFTEST_CLOCK khz=%u want=%u\r\n", .{ khz, want });
+ check("the cycle counter and the systimer agree on the CPU frequency", khz > want - slack and khz < want + slack);
+ }
+
+ soc.rom.print("MARK SELFTEST DONE pass=%u fail=%u\r\n", .{ passed, failed });
+ while (true) {}
+}
diff --git a/src/esp32p4/uart.zig b/src/esp32p4/uart.zig
new file mode 100644
index 00000000..0afd145b
--- /dev/null
+++ b/src/esp32p4/uart.zig
@@ -0,0 +1,165 @@
+//! UART0 as the editor's terminal: bytes out, bytes in, and nothing else.
+//!
+//! This is the whole of the firmware's I/O. There is no framebuffer and no keyboard; the board
+//! emits ANSI and consumes ANSI, and the terminal emulator on the far end of the CH340 does the
+//! rest of the work - including answering the editor's own capability queries, which travel down
+//! this wire like any other bytes.
+//!
+//! WHY IT IS IN THIS REPOSITORY. It is not a UART driver - that is `hal.uart`, which stays in the
+//! toolchain package and is checked against ESP-IDF's own headers by `zig build diff` there. This is
+//! the EDITOR's use of one: which instance the console is, that the transmitter is real
+//! backpressure because a truncated escape sequence corrupts the host terminal, that rescued
+//! keystrokes must come out before FIFO ones, and that the first thing to do at startup is discard
+//! the host bridge's synthetic resize report. Every one of those is a statement about the editor, so
+//! the file moved to sit beside it. See `app.zig`'s header for the boundary in full.
+//!
+//! What it expects from the toolchain package is exactly one module: `hal`, for `hal.uart.Uart`.
+//! Nothing else here reaches the chip. The sibling `input_rescue.zig` is a plain file, not a module,
+//! because it is this repository's own policy.
+//!
+//! Deliberately not a `std.Io.Writer`. The ANSI encoding lives on the other side of the C ABI, next
+//! to the vaxis that produces it (see `app.zig` for why the seam is there and not elsewhere), so
+//! what crosses into this file is already a finished run of bytes. A writer here would be a second
+//! buffer in front of one that already exists.
+//!
+//! Two decisions worth stating, because both are measurements rather than preferences.
+//!
+//! **Batched FIFO access.** The naive push is `while (txFree() == 0) {}` then `pushByte`, once per
+//! byte: one MMIO read per byte at best, many while the FIFO is full. Reading `txFree` once and
+//! then pushing that many cuts the status reads by up to the FIFO depth (128, the toolchain
+//! package's `src/hal/uart.zig:52`). At 115200 the wire costs ~86 us per byte and dwarfs either
+//! version, so today this is merely free - and it stops being free the moment the divider is raised.
+//!
+//! **UART0's configuration is never touched.** Not the divider, not the format, not the pad
+//! routing, and above all not `reset()`. The second-stage bootloader configured this block, and
+//! `src/hal/uart.zig:195-211` records what happens if it is reset: UART_CLKDIV returns to its
+//! power-on value, the console turns to garbage mid-sentence, and the board takes a watchdog reset
+//! with nothing readable left to explain it. Everything here touches FIFO offset 0x000 and the
+//! status register, and nothing else.
+
+const hal = @import("hal");
+const input_rescue = @import("input_rescue.zig");
+
+/// UART0: the instance the CH340 is wired to, and the one the ROM and bootloader configured.
+const uart0 = hal.uart.Uart.init(0);
+
+/// Keystrokes taken off the receiver while the transmitter was full. See `input_rescue`: without
+/// this, anything typed into a frame longer than the 128-byte FIFO was silently gone.
+var rescued: input_rescue.Ring = .{};
+
+/// Push `bytes` into the TX FIFO, blocking while it is full.
+///
+/// The spin is normally bounded by the wire - a full 128-byte FIFO drains in 11 ms at 115200 - and
+/// dropping instead of waiting would truncate an escape sequence, leaving the host terminal in the
+/// wrong colour for the rest of the session. So the wait is real backpressure.
+///
+/// But it is BOUNDED, for the reason the toolchain package's `src/hal/uart.zig:182-186` gives about
+/// `update()`: a UART whose core clock has been gated never makes progress, and "on a board with no
+/// debugger an infinite spin is indistinguishable from a crash". That is not hypothetical here - it
+/// is how this port spent an afternoon: output stopped mid-boot with no panic and no watchdog (the
+/// RTC watchdog having been correctly disabled), which looked like a hang in whatever code came next
+/// rather than a stalled transmitter. A bounded wait turns that into visibly dropped output plus a
+/// counter, which is a diagnosis instead of a mystery.
+///
+/// The limit is per burst, not per call, and generous: 1,000,000 status reads is far longer than
+/// any legitimate drain and still a fraction of a second.
+pub fn write(bytes: []const u8) void {
+ dropped +%= input_rescue.pump(uart0, &rescued, bytes, 1_000_000);
+}
+
+/// Bytes abandoned because the transmitter stopped making progress. Nonzero means the console is
+/// lying about what happened, so it is worth printing.
+pub var dropped: u32 = 0;
+
+/// One byte, for callers that must not touch `.rodata` to say anything - which during bring-up is
+/// the difference between a diagnostic and a second copy of the bug being diagnosed.
+pub fn writeByte(b: u8) void {
+ var spins: u32 = 0;
+ while (uart0.txFree() == 0) {
+ spins += 1;
+ if (spins > 1_000_000) {
+ dropped +%= 1;
+ return;
+ }
+ }
+ uart0.pushByte(b);
+}
+
+/// Emit `n` bytes read from `addr` as two hex digits each, computing the digits arithmetically so
+/// nothing here reads a lookup table. Used to answer "does a load from this address return what the
+/// linker put there", which is not a question a string literal can be trusted to ask.
+pub fn dumpHex(addr: u32, n: u32) void {
+ const p: [*]const volatile u8 = @ptrFromInt(addr);
+ var i: u32 = 0;
+ while (i < n) : (i += 1) {
+ const byte = p[i];
+ for ([2]u8{ byte >> 4, byte & 0xf }) |nib| {
+ writeByte(if (nib < 10) '0' + nib else 'a' + (nib - 10));
+ }
+ }
+ writeByte('\r');
+ writeByte('\n');
+}
+
+/// A u32 as eight hex digits, reading no memory at all.
+pub fn dumpWord(v: u32) void {
+ var shift: u5 = 28;
+ while (true) {
+ const nib: u8 = @intCast((v >> shift) & 0xf);
+ writeByte(if (nib < 10) '0' + nib else 'a' + (nib - 10));
+ if (shift == 0) break;
+ shift -= 4;
+ }
+ writeByte('\r');
+ writeByte('\n');
+}
+
+/// Move whatever the host has sent into `buf`, without waiting. Returns the count.
+///
+/// Non-blocking on purpose: the loop has a frame to render and a core to pump, and the editor must
+/// not stall on a keystroke that may never come. `rxCount` is read once per call and the FIFO
+/// drained to that mark, so a fast typist or a pasted buffer cannot hold the loop here.
+pub fn read(buf: []u8) usize {
+ // RESCUED BYTES FIRST. They arrived before anything still sitting in the FIFO, and an editor
+ // that reorders keystrokes is worse than one that drops them.
+ var n = rescued.pop(buf);
+ const waiting = @min(uart0.rxCount(), buf.len - n);
+ for (buf[n..][0..waiting]) |*slot| slot.* = uart0.popByte();
+ n += waiting;
+ return n;
+}
+
+/// Take whatever has arrived off the receiver right now, without waiting and without handing it to
+/// anyone. For callers that are about to spend a while not reading: `write` does this while the
+/// transmitter is full, and the loop does it between chunks of input, because applying a keystroke
+/// gets more expensive as the line grows and 128 bytes of FIFO is only 11 ms at 115200.
+pub fn rescueNow() void {
+ input_rescue.rescue(uart0, &rescued);
+}
+
+/// Input abandoned because even the rescue buffer overflowed. Distinct from `dropped`, which is
+/// OUTPUT abandoned by a stalled transmitter.
+pub fn inputDropped() u32 {
+ return rescued.dropped;
+}
+
+/// Discard anything already received, returning how much. Used once at startup: the host-side
+/// bridge injects a window-size report before this program exists, and the bootloader's chatter has
+/// already been echoed at the host. Neither is user input.
+///
+/// Pops rather than calling `resetRxFifo`, which is a CONF0_SYNC read-modify-write plus two commits
+/// on the console UART - see this file's header.
+pub fn drainInput() u32 {
+ var discarded: u32 = 0;
+ while (uart0.rxCount() > 0) : (discarded += 1) _ = uart0.popByte();
+ discarded += @intCast(rescued.len);
+ rescued.clear();
+ return discarded;
+}
+
+/// The rate the hardware is actually producing, by reading its dividers back. Reported rather than
+/// assumed: the host has to be opened at the same rate, and a mismatch shows up as garbage on the
+/// screen rather than as an error anyone can act on.
+pub fn baudrate() u32 {
+ return uart0.baudrate(uart0.clockSource().nominalHz());
+}