summaryrefslogtreecommitdiff
path: root/src/esp32p4/app.zig
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-26 13:27:46 -0300
committerGabriel Schneider <[email protected]>2026-08-27 09:47:39 -0300
commit11f380f6d7222f2cad93c2cdf13701ea1f903d47 (patch)
tree803194ee5853a6b4cda93f90a95e28d1f02e69ae /src/esp32p4/app.zig
parentfbc194068687e49a8490c85c9f1257a2f2bb9079 (diff)
downloadpardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.tar.gz
pardes-11f380f6d7222f2cad93c2cdf13701ea1f903d47.zip
One core behind N frontends, the board's own runner moved in, and every board cap on one screen
## The wire is the effect stream, not a new protocol `pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns a terminal and a socket and nothing else. N frontends on one core all look at the same screen — `screen -x`, not N sessions. The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin command line, and a command line cannot carry a frame. ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit little-endian fixed width and no message is a blit of a native struct. A protocol that only works between two builds of the same compiler would have thrown away the one frontend that motivated it. ## The board comes in; its toolchain stays out `src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over- serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so `zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the board from this repo's `build.zig`. The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes core over a serial line. What stayed is everything a second project would also want — the HAL, the register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its own and its `build()` early-returns when it is not the root package, so this costs the package graph exactly zero packages and the editor's own builds nothing at all. ## limits.zig: nine forgettable places become one budget Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions — they are ONE decision, how much memory this build may spend, taken nine times where no reader could see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against what it is measured against, derived from two booleans. The payoff is testability on a machine that is not the board: the caps are ordinary comptime values, so a host build can be compiled against the board's numbers and the parking, eviction and clamping paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART. ## A bare `zig build` `zig build` with no arguments now builds the tty and GUI binaries and installs them into `~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and "use it" two different commands for no reason.
Diffstat (limited to 'src/esp32p4/app.zig')
-rw-r--r--src/esp32p4/app.zig546
1 files changed, 546 insertions, 0 deletions
diff --git a/src/esp32p4/app.zig b/src/esp32p4/app.zig
new file mode 100644
index 00000000..a9cf627d
--- /dev/null
+++ b/src/esp32p4/app.zig
@@ -0,0 +1,546 @@
+//! pardes, as ESP32-P4 firmware: the reset entry, the heap, the clock, the trap handler and the
+//! loop.
+//!
+//! There is no operating system under this. `_start` is the reset entry the second-stage bootloader
+//! jumps to, and this file is the entire platform: a heap, a millisecond clock, and UART0.
+//!
+//! ## Why the firmware root is in the editor's repository
+//!
+//! It was written in the `05-zig-p4` toolchain repository, next to the SoC support it uses, and it
+//! moved here because everything in it is a statement about the EDITOR. The heap span it hands over
+//! is the number that decides how large a grid the board can drive; `input_chunk` is sized against
+//! what applying one keystroke costs in `src/pardes.zig`; the loop's shape - read, chunk, tick,
+//! render only when dirty - is this editor's loop and no one else's; and the `-Dprof` attribution
+//! exists to answer "where did the 34 ms of a keystroke go" about this program. A firmware root that
+//! specific to one application belongs beside it.
+//!
+//! What stayed behind is everything a second application would also want, and none of it is
+//! duplicated here: the SoC and HAL, the translate-c register layer, the coalescing heap, `std.Io`
+//! for this chip, the app descriptor, the generated linker script, the image builder, the flasher
+//! and the interactive console. Those arrive as the `zig_p4` dependency, and this file imports
+//! exactly four of its modules - `soc`, `hal`, `heap` and `config` - plus two sibling files,
+//! `uart.zig` and `input_rescue.zig`, which are the editor's own.
+//!
+//! ## Where the editor is
+//!
+//! On the far side of a C ABI, still, and that is a choice rather than a leftover. `src/esp32p4.zig` in
+//! this same repository is compiled as ONE freestanding object (`b.addObject`, rooted at that file)
+//! and linked in beside this one; the `extern` declarations below are the near side of that seam.
+//!
+//! Importing `esp32p4.zig` as a module instead would be shorter to write and worse in every way that
+//! matters. It would drag the core's whole module graph - vaxis, the themes, the allocator tiers -
+//! into this root, which is the compilation that must stay small enough to reason about. It would
+//! give the firmware two ways to reach the editor. And above all it would make the OBJECT path a
+//! second arrangement, tested separately: that path is what `05-zig-p4 -Dpardes -Dpardes-obj=...`
+//! builds, it is what every measurement in that repository's `experiments/` was taken through, and
+//! it is a supported way to build this board. With the extern kept, both builds link the same eight
+//! symbols against the same object file, so neither can drift and neither is the better-tested one.
+//! The reasons the seam is a file at all - a nested `build.zig.zon` dependency broke every build in
+//! the toolchain repository - are recorded in `src/esp32p4.zig:8-15` and `05-zig-p4/build.zig:238-260`.
+//!
+//! Who owns which symbol: `src/esp32p4.zig` exports all eight `pardes_esp32p4_*` functions and nothing else.
+//! This file exports `_start`, `zig_main`, `trapEntry` and `trapReport`. `esp_app_desc` belongs to
+//! neither and comes from the toolchain's own appdesc object, which the link adds unconditionally.
+//! `abi_version` below is the one constant both sides spell, and its counterpart is
+//! `src/esp32p4.zig:98` - one repository now, so a bump is two lines in one diff rather than two commits
+//! in two trees.
+//!
+//! The seam is deliberately **bytes in, bytes out**. Everything that needs to know what a cell is -
+//! vaxis, the ANSI encoder, the input parser, the capability handshake - lives on the far side,
+//! next to the vaxis it is built against. What crosses is a byte stream in each direction, which is
+//! exactly what a serial line is, so this file has no opinion about terminals at all.
+//!
+//! ## Where the memory is
+//!
+//! Measured on this die by the toolchain's `examples/memprobe.zig`, not read off a datasheet, and
+//! written down once in the generated linker script (`05-zig-p4/build.zig:1572,1579,1584-1585`):
+//!
+//! 0x4FF00000..0x4FF3F000 252 KiB `l2mem`: .data/.bss/.stack are linked into this
+//! 0x4FF3F000..0x4FF40000 4 KiB mask ROM .data/.bss - untouchable, ets_printf needs it
+//! 0x4FF40000..0x4FFA0000 384 KiB `l2high`: handed to the editor as its entire heap
+//! 0x4FFA0000..0x4FFC0000 128 KiB NOT memory - the L2 cache lives here
+//!
+//! That last line is why the heap is 384 KiB and not the 512 KiB an earlier version of this comment
+//! claimed. The first probe wrote a pattern and read it back one page at a time and reported the
+//! whole upper 512 KiB as RAM, because a store followed immediately by a load of the SAME address
+//! returns the stored value whether the backing store is real, an address mirror, or merely a dirty
+//! cache line. Writing every page before reading any page separates the three, and the top 128 KiB
+//! then failed; handing them to an allocator hung the heap on its first free-list walk. ESP-IDF's
+//! own arithmetic agrees exactly: SRAM_HIGH_SIZE = 0x80000 - CONFIG_CACHE_L2_CACHE_SIZE, with the
+//! Kconfig default of 128 KiB.
+//!
+//! The span arrives as `__heap_start`/`__heap_end` from that script, so those addresses are written
+//! down in exactly one place. The editor owns it outright: it is passed in at init and this file
+//! never allocates from it.
+//!
+//! PSRAM is not used. The board has 32 MB fitted and it would make all of this comfortable, but
+//! ESP-IDF's own ESP32-P4 implementation runs past a thousand lines - MPLL, MSPI clocking, pin
+//! drive and DQS, CS timing, mode registers, a connectivity check, and an entire timing-calibration
+//! subsystem - and the mask ROM offers only MMU mapping, no device init. Touching it untrained
+//! faults and hangs the core, which `examples/memprobe.zig` demonstrates on purpose.
+
+const std = @import("std");
+const soc = @import("soc");
+const config = @import("config");
+
+/// `-Dprof`: time the two phases of a keystroke on the board and print the cycle counts. A
+/// diagnostic, not a feature - see the loop.
+const prof = config.prof;
+
+/// Every byte this loop has taken off the UART, for `-Dprof`. Ground truth for "did the burst
+/// arrive", which a screen reconstruction cannot answer: a character can be missing from the screen
+/// because it never arrived, because the editor never applied it, or because the viewport does not
+/// show that column.
+var rx_total: u32 = 0;
+
+/// How many input bytes to hand the editor before draining the receiver again. Chosen against the
+/// FIFO rather than against the editor: applying one keystroke was measured at 44 us on an empty
+/// line and 63 us at 640 characters, so eight of them is at most ~0.5 ms in which nothing empties
+/// the receiver, against a 128-byte FIFO that holds 11 ms of wire at 115200. Twenty times the margin
+/// needed, and it costs nothing on the wire because one render still happens per loop iteration.
+const input_chunk = 8;
+const hal = @import("hal");
+const heapmod = @import("heap");
+const uart = @import("uart.zig");
+
+// ------------------------------------------------------------------------------------- the ABI
+// Eight functions, all `callconv(.c)`, all implemented in the linked object - `src/esp32p4.zig` in this
+// repository, compiled for the same target and exporting exactly these names. This is the complete
+// interface between this board and the editor, and it is deliberately bytes-and-memory only: the
+// editor never learns what a UART is, and this file never learns what a cell is.
+//
+// The declarations below are a SECOND spelling of the signatures in `src/esp32p4.zig:100-127,259-...`,
+// and that duplication is what a C ABI is: each side declares the wire independently, which is
+// precisely why `abi_version` has to be checked. Sharing a Zig type between them would mean sharing
+// a module, which would mean the core in this compilation - see the header.
+
+/// How the editor emits bytes. Called with finished runs of ANSI, many times per frame.
+const WriteFn = *const fn (ctx: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void;
+
+/// The board's pads, offered to the editor. Optional on the wire so a firmware with nothing to
+/// toggle passes null and the `Gpio` word reports that rather than the object guessing.
+const GpioFn = *const fn (ctx: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool;
+
+/// This board's allocator, handed across as plain function pointers. `log2_align` is a log2 value,
+/// which is exactly how `std.mem.Alignment` represents itself, so neither side needs a conversion
+/// table.
+///
+/// The memory belongs to THIS side: only the firmware knows that the heap is the 384 KiB at
+/// 0x4FF40000, that the 128 KiB above it is L2 cache, and that PSRAM is untrained. The editor gets
+/// an allocator, not an address range.
+const Allocator = extern struct {
+ ctx: ?*anyopaque,
+ alloc: *const fn (ctx: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8,
+ resize: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool,
+ free: *const fn (ctx: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void,
+};
+
+/// The one number both sides must agree on. Linkers do not type-check C symbols, so a signature
+/// that drifts on one side of this seam links cleanly and then corrupts the stack; checking this
+/// before calling anything else turns that into a refusal to boot.
+const abi_version: u32 = 2;
+extern fn pardes_esp32p4_abi_version() callconv(.c) u32;
+
+/// Hand over the allocator and the output sink, and state the initial window size. Returns 0, or a
+/// small non-zero code this file can only report.
+extern fn pardes_esp32p4_init(
+ alloc: *const Allocator,
+ write: WriteFn,
+ gpio: ?GpioFn,
+ ctx: ?*anyopaque,
+ cols: u16,
+ rows: u16,
+) callconv(.c) u32;
+
+/// Raw bytes off the wire: keystrokes, capability-query replies, and the host bridge's in-band
+/// resize reports. The editor parses all three; this file distinguishes none of them.
+extern fn pardes_esp32p4_input(ptr: [*]const u8, len: usize) callconv(.c) void;
+
+/// Advance time. Separate from `input` because animations and timeouts must progress on a wire
+/// where nothing is arriving.
+extern fn pardes_esp32p4_tick(now_ms: u64) callconv(.c) void;
+
+/// Emit one frame through the write callback. Returns 0 or an error code.
+extern fn pardes_esp32p4_render() callconv(.c) u32;
+
+/// Is there anything to draw - a dirty surface or a running animation? Asked every iteration so a
+/// quiet editor costs no bytes on a 115200-baud link.
+extern fn pardes_esp32p4_wants_frame() callconv(.c) bool;
+
+/// Has the user asked to leave? There is nowhere to go, so this only stops the loop.
+extern fn pardes_esp32p4_quit() callconv(.c) bool;
+
+/// The last frame's three stages in CPU cycles: the copy of pardes's Surface into vaxis's grid,
+/// vaxis's own diff-and-emit, and the push into the UART. Only meaningful under `-Dprof`; the
+/// editor object always exports it, and it costs two CSR reads per stage.
+extern fn pardes_esp32p4_frame_prof(copy: *u64, render: *u64, flush: *u64) callconv(.c) void;
+
+// ------------------------------------------------------------------------------------ the sink
+
+/// The write callback handed to `pardes_esp32p4_init`. No context is needed - there is one UART.
+fn writeOut(_: ?*anyopaque, ptr: [*]const u8, len: usize) callconv(.c) void {
+ uart.write(ptr[0..len]);
+}
+
+/// Flip one pad and report the level before and after. The editor's `Gpio` word calls this; the
+/// editor has no register of its own for it, deliberately.
+///
+/// THIS IS WHY THE SEAM IS HERE. A toggle is not a write to GPIO_OUT: `configureOutput` points the
+/// pad's IO MUX at the GPIO function, routes the GPIO matrix's output to it, sets the drive strength
+/// and input buffer and clears the pulls, and only then enables the driver - four register files,
+/// indexed by a per-pin table. That code already exists in the toolchain package's `src/hal/gpio.zig`,
+/// it is the same call that package's `src/main.zig` blinks with, and its register numbers are
+/// checked against ESP-IDF's own headers by `zig build diff` there. A second copy inside the editor
+/// object would be a second copy under no test.
+///
+/// `getDrivenLevel` rather than `getLevel`: the answer is the level this board is DRIVING, which is
+/// defined for every pin. The pad's own level is what the outside world says, and on an unconnected
+/// header pin that is noise. The input buffer is enabled anyway, so `Peek` of GPIO_IN_REG shows the
+/// pad for anyone who wants to compare the two.
+fn gpioToggle(_: ?*anyopaque, pin: u16, was: *u8, now: *u8) callconv(.c) bool {
+ if (pin > hal.gpio.max_pin) return false;
+ const p: u8 = @intCast(pin);
+ hal.gpio.configureOutput(p, .{ .readback = true });
+ const before = hal.gpio.getDrivenLevel(p);
+ if (before == 1) hal.gpio.setLow(p) else hal.gpio.setHigh(p);
+ was.* = before;
+ now.* = hal.gpio.getDrivenLevel(p);
+ return true;
+}
+
+// ------------------------------------------------------------------------------------- the heap
+
+/// The span the linker script hands over, from `l2high`'s ORIGIN and LENGTH.
+///
+/// Reached with `@extern`, NOT with `extern const __heap_start: anyopaque` plus
+/// `@intFromPtr`/`@ptrFromInt`. That spelling was here first and it was silently wrong: declaring a
+/// linker symbol as an `anyopaque` OBJECT gives the optimiser a zero-sized object, so a pointer
+/// derived from its address carries provenance for zero bytes, and ordinary (non-volatile) stores
+/// through it are dead code it may drop. The toolchain's `examples/heapcheck.zig` caught it on the
+/// die - the allocator's first block header read back as `size=2988759312 next=0xffffffff`-not, and
+/// the free list walk never terminated. A `[*]u8` from `@extern` has no size to lose.
+const heap_start = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_start" });
+const heap_end = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_end" });
+
+fn heapSpan() []align(heapmod.Heap.granule) u8 {
+ return heap_start[0 .. @intFromPtr(heap_end) - @intFromPtr(heap_start)];
+}
+
+/// The one heap. A K&R coalescing free list over that span, validated on this die by the toolchain's
+/// `examples/heapcheck.zig`: 512 blocks fill and free back to a single 393,216-byte block, a holed
+/// arena still satisfies a 4 KiB request, and 20,000 random operations drain back to one block.
+var gpa_heap: heapmod.Heap = undefined;
+
+// The four C forwarders the editor is handed. `log2_align` round-trips through
+// `std.mem.Alignment`, whose representation IS the log2 value.
+
+fn cAlloc(_: ?*anyopaque, len: usize, log2_align: u8) callconv(.c) ?[*]u8 {
+ const a = gpa_heap.allocator();
+ return a.vtable.alloc(a.ptr, len, @enumFromInt(log2_align), @returnAddress());
+}
+
+fn cResize(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8, new_len: usize) callconv(.c) bool {
+ const a = gpa_heap.allocator();
+ return a.vtable.resize(a.ptr, ptr[0..len], @enumFromInt(log2_align), new_len, @returnAddress());
+}
+
+fn cFree(_: ?*anyopaque, ptr: [*]u8, len: usize, log2_align: u8) callconv(.c) void {
+ const a = gpa_heap.allocator();
+ a.vtable.free(a.ptr, ptr[0..len], @enumFromInt(log2_align), @returnAddress());
+}
+
+const editor_allocator: Allocator = .{
+ .ctx = null,
+ .alloc = cAlloc,
+ .resize = cResize,
+ .free = cFree,
+};
+
+// ------------------------------------------------------------------------------------ the clock
+
+/// Milliseconds since boot, off the systimer - a 16 MHz counter (the toolchain package's
+/// `src/hal/systimer.zig:31`), which is the cheapest trustworthy clock on this chip. `read` returns
+/// null if the unit is not running, in which case time simply does not advance and the editor stops
+/// animating; that is a better failure than a clock that jumps.
+fn nowMs() u64 {
+ const us = hal.systimer.micros(.unit0) orelse return 0;
+ return us / 1000;
+}
+
+// ------------------------------------------------------------------------------------- the loop
+
+export fn zig_main() noreturn {
+ // FIRST, before a single byte of `.rodata` is touched - which means before the marker below,
+ // because that marker IS a string literal in flash and would read as machine code without this.
+ soc.flushFlashCache();
+ const heap = heapSpan();
+ soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{
+ @as(u32, @intFromPtr(heap.ptr)),
+ @as(u32, @intFromPtr(heap.ptr)) + @as(u32, @intCast(heap.len)),
+ @as(u32, @intCast(heap.len / 1024)),
+ });
+
+ // The CPU clock, before anything is timed against it. The bootloader leaves 90 MHz and the
+ // CPLL is already at 360, so this is a divider change that disturbs neither UART0 (XTAL) nor
+ // the systimer (XTAL/2.5) nor the flash interface (SPLL). See the toolchain package's
+ // `src/hal/clkrst.zig:setCpuFreq`.
+ if (config.cpu_mhz != 90) hal.clkrst.setCpuFreq(switch (config.cpu_mhz) {
+ 180 => .mhz180,
+ 360 => .mhz360,
+ else => .mhz90,
+ });
+
+ const rwdt_was_armed = hal.rwdt.disable();
+ hal.systimer.init();
+ _ = rwdt_was_armed;
+
+ const their_abi = pardes_esp32p4_abi_version();
+ if (their_abi != abi_version) {
+ uart.write("MARK PARDES_ABI_MISMATCH\r\n");
+ while (true) {}
+ }
+
+ gpa_heap = heapmod.Heap.init(heap);
+ _ = uart.drainInput();
+
+ // Ask for more than any grid this board will ever render, so the SHELL's own ceiling is what
+ // governs - it clamps to `-Desp32p4-cols`/`-Desp32p4-rows` and reports the result. Naming 80x24 here made
+ // the firmware a second opinion about the geometry, which is one opinion too many.
+ const rc = pardes_esp32p4_init(&editor_allocator, writeOut, gpioToggle, null, 255, 255);
+
+ if (rc != 0) {
+ soc.rom.print("MARK PARDES_INIT_FAIL rc=%u\r\n", .{rc});
+ const s = gpa_heap.stats();
+ soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{
+ s.free, s.largest_free, s.free_blocks,
+ });
+ while (true) {}
+ }
+
+ // The HEAP, after the editor has taken what it needs. This is the number that decides how large
+ // a grid the board can drive, so it is printed on every boot rather than only on failure: a
+ // geometry that fits with 2 KB to spare and one that fits with 80 KB are not the same answer,
+ // and the difference is invisible from the host otherwise.
+ {
+ const s = gpa_heap.stats();
+ soc.rom.print("MARK PARDES_HEAP free=%u largest=%u blocks=%u\r\n", .{
+ s.free, s.largest_free, s.free_blocks,
+ });
+ }
+
+ // The CPU clock, measured rather than assumed. Every cycle count this firmware reports is
+ // divided by it somewhere, and the toolchain's `src/io/chip.zig` records it as "a measured
+ // ~90 MHz" that nothing here reconfigures - so it is worth printing rather than remembering. The
+ // systimer is XTAL/2.5 = 16 MHz and is NOT derived from the CPU clock
+ // (`src/hal/systimer.zig:31`, `clk_tree_defs.h:196-198`), which is exactly what makes it a valid
+ // reference for measuring it.
+ if (prof) {
+ const t_start = hal.systimer.micros(.unit0) orelse 0;
+ const c_start = soc.cycles();
+ // 50 ms is long enough that the systimer's 16 MHz granularity and the loop's own overhead
+ // are both noise, and short enough to be invisible in a boot.
+ while ((hal.systimer.micros(.unit0) orelse 0) -% t_start < 50_000) {}
+ const elapsed_us = (hal.systimer.micros(.unit0) orelse 0) -% t_start;
+ const elapsed_cy = soc.cycles() - c_start;
+ soc.rom.print("MARK CPU_HZ cycles=%u us=%u khz=%u\r\n", .{
+ @as(u32, @intCast(elapsed_cy)),
+ @as(u32, @intCast(elapsed_us)),
+ @as(u32, @intCast(if (elapsed_us > 0) elapsed_cy * 1000 / elapsed_us else 0)),
+ });
+ }
+ soc.rom.print("MARK PARDES_READY\r\n", .{});
+
+ var in: [256]u8 = undefined;
+ while (!pardes_esp32p4_quit()) {
+ // ATTRIBUTION. The host can time a keystroke's round trip but cannot see what the firmware
+ // spent it on, and the two candidates - parsing and editing, versus rendering - want
+ // opposite fixes. `soc.cycles()` is the unprivileged cycle counter, so this costs two CSR
+ // reads per phase and quantises at one cycle, which is four orders of magnitude below the
+ // milliseconds being attributed. Gated on `prof` so the shipping build carries none of it.
+ const n = uart.read(&in);
+ rx_total +%= @intCast(n);
+
+ var input_cy: u64 = 0;
+ if (n > 0) {
+ const t0 = if (prof) soc.cycles() else 0;
+ // IN CHUNKS, rescuing the receiver between them. Applying a keystroke is not free and
+ // gets dearer as the line grows - measured at 44 us on an empty line and 63 us at 640
+ // characters - so handing over a full 128-byte batch is up to 8 ms in which nothing
+ // drains the receiver, against a FIFO that holds only 11 ms of wire. A 600-byte paste
+ // lost 93 bytes to exactly that window even with the transmitter's own rescue in place.
+ //
+ // Splitting a burst at an arbitrary byte is safe: `pardes_esp32p4_input` keeps whatever it
+ // could not parse, which is how it already survives an escape sequence split across two
+ // UART reads. One render still happens per loop iteration, so this costs no extra wire.
+ var off: usize = 0;
+ while (off < n) {
+ const chunk = @min(input_chunk, n - off);
+ pardes_esp32p4_input(in[off..].ptr, chunk);
+ off += chunk;
+ if (off < n) uart.rescueNow();
+ }
+ if (prof) input_cy = soc.cycles() - t0;
+ }
+
+ pardes_esp32p4_tick(nowMs());
+
+ // Only when there is something to show. On a link this slow an unconditional repaint per
+ // iteration would saturate the wire and starve input.
+ if (pardes_esp32p4_wants_frame()) {
+ const t0 = if (prof) soc.cycles() else 0;
+ const err = pardes_esp32p4_render();
+ if (err != 0) soc.rom.print("MARK PARDES_RENDER_FAIL rc=%u\r\n", .{err});
+ if (prof) {
+ const render_cy = soc.cycles() - t0;
+ // A SECOND render with nothing changed since the first. It splits the cost in two:
+ // whatever this still costs is the price of walking and diffing the whole editor
+ // state, paid regardless of output, while the difference between the two is the
+ // price of the change itself. `wants_frame` is false now, so this only happens
+ // under -Dprof and never on a shipping build.
+ const t1 = soc.cycles();
+ _ = pardes_esp32p4_render();
+ const idle_cy = soc.cycles() - t1;
+ // Reported in cycles, not microseconds: the divisor is the CPU clock, which this
+ // firmware does not set and has only ever measured, so converting here would bake a
+ // guess into the data. The toolchain's `experiments/` divides by the clock it
+ // measured.
+ var copy_cy: u64 = 0;
+ var vx_cy: u64 = 0;
+ var flush_cy: u64 = 0;
+ pardes_esp32p4_frame_prof(&copy_cy, &vx_cy, &flush_cy);
+ soc.rom.print("PROF in=%u render=%u idle=%u copy=%u vaxis=%u flush=%u rx=%u rxdrop=%u txdrop=%u\r\n", .{
+ @as(u32, @intCast(input_cy)),
+ @as(u32, @intCast(render_cy)),
+ @as(u32, @intCast(idle_cy)),
+ @as(u32, @intCast(copy_cy)),
+ @as(u32, @intCast(vx_cy)),
+ @as(u32, @intCast(flush_cy)),
+ rx_total,
+ uart.inputDropped(),
+ uart.dropped,
+ });
+ }
+ }
+ }
+
+ soc.rom.print("\r\nMARK PARDES_QUIT\r\n", .{});
+ while (true) {}
+}
+
+// ------------------------------------------------------------------------------------ the trap
+
+/// A trap handler, because the absence of one is why this port has been guessing.
+///
+/// The mask ROM prints "Guru Meditation" for a trap only while ITS handler is still installed;
+/// anything this image does that replaces or outgrows that path fails silently instead, and a silent
+/// fault is indistinguishable from an infinite loop over a serial line. This one reports the three
+/// registers that name the fault and then stops, using the direct-FIFO writer so it shares nothing
+/// with the editor's buffered output.
+///
+/// `mtvec` is set in DIRECT mode (low two bits zero), so every trap and every interrupt lands on
+/// `trapEntry` regardless of cause - which is what a diagnostic wants.
+export fn trapEntry() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile ("j trapReport");
+}
+
+export fn trapReport() noreturn {
+ const mcause = asm volatile ("csrr %[o], mcause"
+ : [o] "=r" (-> u32),
+ );
+ const mepc = asm volatile ("csrr %[o], mepc"
+ : [o] "=r" (-> u32),
+ );
+ const mtval = asm volatile ("csrr %[o], mtval"
+ : [o] "=r" (-> u32),
+ );
+ uart.write("\r\nMARK TRAP mcause=");
+ uart.dumpWord(mcause);
+ uart.write("MARK TRAP mepc=");
+ uart.dumpWord(mepc);
+ uart.write("MARK TRAP mtval=");
+ uart.dumpWord(mtval);
+ uart.write("MARK TRAP dropped=");
+ uart.dumpWord(uart.dropped);
+ while (true) {}
+}
+
+// --------------------------------------------------------------------------- the root's own duties
+//
+// These are the FIRMWARE root's declarations, and they are not the same set as `src/esp32p4.zig`'s: that
+// file is the root of its own object and carries its own `std_options` and `panic` for the core's
+// half of the image. Two roots, two instantiations of std, one per compilation unit - which is
+// exactly what the object seam buys, and why a panic in the core prints `PARDES_CORE_PANIC` through
+// the write callback while a panic here prints `PARDES_PANIC` through the mask ROM.
+
+/// `page_size_min`/`max`: the board has no MMU and no pages, but std derives allocator alignment
+/// from these. 4 KiB is the ESP32-P4's cache and DMA granularity.
+///
+/// `logFn` is not cosmetic. std's default log implementation reaches `std.debug_io`, which
+/// instantiates `std.Io.Threaded` - a thread pool, `getrandom`, `IOV_MAX`, `mremap` - none of which
+/// exist here, and one `log.warn` from anywhere is enough to drag all of it into the image.
+pub const std_options: std.Options = .{
+ .page_size_min = 4096,
+ .page_size_max = 4096,
+ .logFn = logFn,
+};
+
+fn logFn(
+ comptime level: std.log.Level,
+ comptime scope: @EnumLiteral(),
+ comptime fmt: []const u8,
+ args: anytype,
+) void {
+ var buf: [256]u8 = undefined;
+ const line = std.fmt.bufPrint(&buf, "\r\n[" ++ level.asText() ++ "/" ++ @tagName(scope) ++ "] " ++ fmt ++ "\r\n", args) catch
+ "\r\n[log overflow]\r\n";
+ uart.write(line);
+}
+
+pub const panic = std.debug.FullPanic(panicImpl);
+
+fn panicImpl(msg: []const u8, first_trace_addr: ?usize) noreturn {
+ // The fixed text goes out through the ROM deliberately: a panic may BE the console writer
+ // failing, and `ets_printf` shares nothing with `uart.write` except the FIFO itself.
+ //
+ // The MESSAGE does not, and that is a correction rather than a preference. `msg` is a Zig SLICE
+ // and `%s` reads until a NUL, so handing `msg.ptr` to printf prints the message and then
+ // whatever happens to sit after it in memory until a zero byte turns up. Literals get away with
+ // it; std's own panics do not, because they are formatted into a buffer - "index out of bounds:
+ // index 5, len 3" - and carry no terminator. `uart.write` takes a length.
+ soc.rom.print("\r\nMARK PARDES_PANIC ", .{});
+ uart.write(msg);
+ // The address is what makes it actionable: addr2line against the ELF in zig-out turns it into a
+ // source line, and without it a panic message names a KIND of failure with no way to find which
+ // one of them happened. Zero when the caller had no return address to give.
+ soc.rom.print("\r\nMARK PARDES_PANIC_AT 0x%08x\r\n", .{@as(u32, @truncate(first_trace_addr orelse 0))});
+ while (true) {}
+}
+
+/// Reset entry. The bootloader hands over with an unspecified stack pointer and the FPU off, so:
+/// enable the F extension (`mstatus.FS`, which ESP-IDF only ever turns on lazily from a trap handler
+/// this image does not have), establish a stack, clear `.bss`, and call into Zig.
+///
+/// The cache invalidate that this image also needs is the FIRST thing `zig_main` does, not something
+/// done here. Hand-written `la t0, Cache_Invalidate_All` against an absolute linker symbol computed
+/// a PC-relative target and jumped into nowhere (measured: PC=0x88b5d788 with the argument stranded
+/// in a2); Zig generates the addressing for an `extern fn` correctly, and `zig_main` runs before any
+/// `.rodata` is touched anyway.
+export fn _start() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile (
+ \\ li t0, 1 << 13
+ \\ csrs mstatus, t0
+ \\ la sp, __stack_top
+ \\ mv fp, sp
+ \\ la t0, trapEntry
+ \\ csrw mtvec, t0
+ \\ la t0, __bss_start
+ \\ la t1, __bss_end
+ \\ bgeu t0, t1, 2f
+ \\1:
+ \\ sw zero, 0(t0)
+ \\ addi t0, t0, 4
+ \\ bltu t0, t1, 1b
+ \\2:
+ \\ j zig_main
+ );
+}