summaryrefslogtreecommitdiff
path: root/tools
Commit message (Collapse)AuthorAge
* Make the toolchain a package another build can drive, and move the editor's ↵HEADmainGabriel Schneider2026-09-16
| | | | glue to the editor
* Offer the editor the board's pads, and check one flips on the dieGabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | ## The board half of Gpio `pardes_p4_init` gained a `GpioFn` and the ABI version went to 2, which is what turns a mixed pair of builds into a refusal to boot rather than five arguments read as six. `gpioToggle` is four lines over `hal.gpio`: configure the pad as a readable output, read the level it is driving, drive the other one, read it again. It is deliberately here and not in the editor. A toggle is not a write to GPIO_OUT - `configureOutput` sets the IO MUX function, the GPIO matrix route, the drive strength, the input buffer and the pulls, then the output enable, indexed by a per-pin table - and that code is already in this repo, already the call `src/main.zig` blinks with, and already checked against ESP-IDF's headers by `zig build diff`. The editor object gets a function pointer instead of a second copy nobody tests. `getDrivenLevel`, not `getLevel`: the answer is the level the board is driving, which is defined for every pin. The pad's own level is what the outside world says, and on an unconnected header pin that is noise. The input buffer is enabled anyway so `Peek` of GPIO_IN_REG can be compared to it. ## A fifth hardware check `p4-bench --check` runs `Gpio 33` three times and requires 0->1, 1->0, 0->1. The alternation is the oracle, not either answer. `0->1` alone is what a firmware printing a hardcoded string would also say; two runs that disagree can only come from a level that was stored and read again. Three, so the third rules out an ordering coincidence. GPIO33 because `src/oracle/ledc_cases.zig` already documents it as a free pin on this board's JP1 header - pin 21 on the diagram the editor now draws. GPIO20 is the blink demo's pin and may have a wire on it. This is the only check here that crosses the whole seam: editor word, C ABI, HAL, pad, and the level back out through the message row. Nothing smaller exercises the ABI at all. ## -Dtheme-animation forwarded Same path as the geometry, for the same reason: baked into the object, wanted from here. All five checks pass on the die; host tests green; the board is flashed with md5 932898f34ca1dc31.
* Set the grid from the firmware build, and check Poke against Peek on the dieGabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | ## -Dcols / -Drows The geometry is baked into pardes's object, and that object is built by the other repository - so changing the grid was two commands in two directories, and the second one silently linked whatever the first had left behind. That is the shape of mistake that ends with a firmware whose grid is not the grid you think you flashed. `-Dcols` and `-Drows` now drive the editor's own build before linking it. A nested `zig build`, not a package dependency: the seam between these repositories stays a file, for the reasons the comment above it has always given. This only reaches across to ask for the file to be made a particular way, and it stands aside entirely when `-Dpardes-obj` names an object explicitly, because then the caller has already said which one they want. Verified by the segment table: `-Dcols=80 -Drows=24` takes the loaded segment from 20,408 to 49,944 bytes, which is the shadow grid going from 784 cells to 1,920. ## Poke and Peek, checked against each other Two more checks in `p4-bench --check`, and they are checked against each other because that is the only way to check either one without a second debugger: a Peek alone cannot tell a correct read from a stuck one, and a Poke alone cannot tell a write from a no-op. Both oracles were wrong before they were right, and both failures are worth recording. The Escape check asked whether an `x` came back on the wire. That was true until the boot buffer gained a line beginning "x selects a line", at which point a passing check began failing for a reason that had nothing to do with Escape. It now reads the CURSOR COLUMN after `0`, which the document cannot spell: 8 in normal mode, three further right if the `0` had been inserted instead. The Poke and Peek checks searched the raw byte stream for `deadbeef` and could not find it - because this renderer emits only the cells that CHANGED, jumping between runs with absolute cursor positioning, so a word on screen is frequently not a word on the wire. `stripAnsi` removes the escapes first. The Poke assertion then had to be loosened from one phrase to three facts, because the message row is 49 columns here and a wrapped line puts a row boundary inside whichever phrase straddles it. Four checks, all passing on the die: lone Escape still leaves insert mode ok (cursor col 8) a click split byte-by-byte lands at 18 ok (cursor 3,18) Poke writes a word and reads it back ok and a later Peek still finds it ok
* Fix the console: std.time.milliTimestamp does not exist, and zig build test ↵Gabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | could not see it `zig build interact` did not compile. The mouse coalescer dated its held report with `std.time.milliTimestamp`, and this Zig's `std.time` has constants and `epoch` and no clock at all: the repo's own idiom is `std.Io.Timestamp`, already wrapped as `serial.Port.nowMs`, which is what the three call sites now use. The reason it survived every check I ran is worth writing down, because the same hole will swallow the next one. `zig build test` compiles `tools/console.zig` as a test module, and Zig only analyses what is REACHABLE: no test calls `attach`, so the loop containing the bad call was never looked at. Seven passing MouseFilter tests and a green `zig build test` said nothing whatsoever about whether the file compiles as part of an executable. Nor does plain `zig build` - p4-console is not in the default install step - so the only thing that would have caught it is building or running the named step, which is also the documented way to use this repo: `zig build interact -Dpardes`. Verified the way it should have been the first time: `zig build interact -Dpardes -Dcpu-mhz=360` with piped stdin flashes, attaches, takes the keystrokes and detaches at EOF. Also compiled every step's artifacts - elf, console, bench - and re-ran the matrix that matters for the editor side: default, Debug, ReleaseFast, ReleaseSafe, ReleaseSmall, and -Dplatform=p4 at Debug and ReleaseSmall, plus -Dp4-cols/-Dp4-rows at 40x12 and 80x24. Note for anyone reading a failing `zig build console` or `bench`: those RUN their tool, so they fail with DeviceBusy when something else is attached to the port. That is not a build error, and it is how this one hid in plain sight in the middle of the output.
* Send mouse events over the wire, thinned, and check the two things only ↵Gabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | hardware can ## The console thins mouse reports The board asks for DEC 1002, so the terminal reports presses, releases and motion while a button is held. A press is one report; a DRAG is one report per cell crossed, each about a dozen bytes. A hand crosses forty cells in a tenth of a second, which is ~500 bytes, which is 43 ms of a 115200 line - and every one of those bytes is input the board must parse while it is trying to paint the result of the previous one. Unthinned, a drag makes the editor unusable for as long as the drag lasts and for a while after. `MouseFilter` holds the newest motion report and drops the ones it supersedes, on a 40 ms window: ~25 reports a second, about 2.6% of the line. Newest-wins is right for motion and only for motion - where the pointer PASSED THROUGH is not information the editor can use, since a selection is defined by where the drag began and where it is now, so an intermediate report already stale by the time it reaches the wire is pure cost. Presses, releases and wheel notches are never held: each one means something different and dropping one loses a click. Ordering is the part worth testing. A held report is released before any non-motion byte that follows it, so a release cannot overtake the motion it ends, and a drag that stops moving still delivers its final position when the window expires. The seven tests cover those, a report split across a read boundary, a non-mouse escape sequence passing through untouched, and - the one that would hurt most - a lone ESC not being swallowed, because that is how you leave insert mode. ## Two hardware checks: p4-bench --check Both are regressions that no host test can see and no latency number can show. A lone Escape still leaves insert mode. The shell now holds a solitary ESC for 10 ms because on this wire the first byte of every sequence arrives alone; if that hold ever stops expiring, Escape stops working and the editor is unusable. A click split byte-by-byte lands at the column clicked. This is the bug that made the mouse look unimplemented, and it only appears when the bytes arrive separately - which the wire does anyway, 87 us apart. The check sends them as ten separate writes with no gap, because a gap longer than the hold would expire it and the check would be exercising nothing. Two things this file deliberately does NOT check, both because a check that cannot fail honestly is worse than no check. Input loss during transmit belongs to the deterministic host test in `src/pardes/input_rescue.zig`, which loses 67 bytes with the fix removed and needs no board. And every hardware oracle for it that was tried here was worse: the cursor stops being reported past 160 characters because the wrapped line outgrows the viewport, and a screen reconstruction cannot be rebuilt mid-session because the board only sends what changed. Both false starts are recorded in the file so the next person does not repeat them. The check also found its own bugs before it found any of the firmware's: 12 ms gaps between the click's bytes expired the very hold it meant to test, `$` produces no frame when the cursor is already at the end of the line, and reading the cursor after a press alone measures the revert rather than the click.
* Run the CPU at 360 MHz: -Dcpu-mhz, and a keystroke lands at 4.37 msGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The board was executing at 90 MHz because the stock second-stage bootloader is built with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured here against the systimer, which is XTAL/2.5 and therefore an independent reference: 4,500,367 cycles in 50,004 us = exactly 90 MHz. The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse). `hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order - APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own comment says the other order passes through a state where APB or MEM violates its timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every `ets_delay_us` in the image is wrong by exactly the frequency ratio. Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only the divisor in experiments/ had to move. ## What it bought step fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x + byte compare 8.37 ms 7.1 us 9.48 ms 0.37x + 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x Compute went 6.42 -> 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte, so that stage is bounded by L2MEM bandwidth and does not care how fast the core is. The prediction that this would happen was made before the measurement and held. ## The goal was 4 ms and this is 4.37 Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency that no firmware change touches (measured independently against the protocol responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large enough to close the gap alone, and it is redundant work - `present` already computes exactly which cells moved - so emitting ANSI straight from the shadow grid would do it. I did not, because it is a from-scratch renderer and the honest verification for it needs more than the harness currently proves. ## Verification, and a bug in my own instrument Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the A/B compares screens across clocks as well as across the shadow-grid flag. The first attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw SGR parameters applied to each cell, which is history-dependent, and a faster board splits the same keystrokes across different frames. Decoding SGR into actual state - resolved foreground, background and attribute set per cell - it is identical: same characters and same style everywhere, both across clocks and across the flag. Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked like it would quantise every sample. It does not - poll(2) returns when data arrives, not when the timeout expires - and switching to a non-blocking spin moved the measured round trip by 0 us. The comment now says so, since the next reader will wonder too. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is opt-in because every number in experiments/ up to this commit was taken at 90.
* Bound the edit path's document scans, and find out they were never the problemGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | The board half: the -Dprof attribution that overturned the conclusion, its data, and the report correction. `-Dprof` adds two cycle-counter reads around `pardes_p4_input` and `pardes_p4_render` and prints both. Off by default: it puts a line on the wire per frame, which is the very resource being measured, so it answers "where did the 15 ms go" and not "how fast is it". It answered. Input is flat at ~220 us regardless of document size - 1.5% of a keystroke - and the entire ~15 ms floor plus every microsecond of the per-character slope live in `render`. The edit-path fix that the source reading implied (committed next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5 conditions x 7 trials, exactly 0% on this board. `experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's mechanism claim was wrong, says so, and carries the disproof beside it. The ranked recommendations are reordered with the renderer at #1. Also here: `--sweep position` in p4-bench, which holds the document fixed at one 320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28 bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on this board - the signature of a walk from the start of a line.
* A measuring instrument, and what it says about where the latency goesGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | "Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result.
* Move the console into its own binary; the step was fighting the progress displayGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | `zig build interact` shredded the editor's screen with fragments of `[11/13] steps +- console`. The cause is structural, not cosmetic: an interactive step lives for as long as the human does, and `std.Progress` redraws the build runner's step tree on stderr every 80 ms for all of it. std already solves this, and the solution is a lock rather than a flag. `Step.Run` with `stdio == .inherit` holds `io.lockStderr()` for the entire lifetime of the child (std/Build/Step/Run.zig:1588-1592) - the same lock `Progress` must take to draw. A hand-rolled step gets none of that unless it takes the lock itself, and `ConsoleStep` did not. So the console is now a real host program, `tools/console_main.zig`, and `console`/`interact` are `Run` steps on it. `--color off` is not needed and is no longer suggested anywhere. Measured under a real pty (a pipe hides the bug, because progress only draws to a terminal - which is why every earlier end-to-end test here looked clean): zero bytes of progress output across an 11-second session, 3 frames, the editor's `^` modified-marker landing at row 2 after two keystrokes. The second reason for a program is that "just connect" should not imply a build. Once the firmware is in flash the board runs it across resets, so the common case during use is to open the port and nothing else: zig-out/bin/p4-console # the board is already programmed zig-out/bin/p4-console --no-reset # ...and leave a live session running `--no-reset` is the interesting one and it is proven on the die: attaching to the running editor produced 309 bytes with no ESP-ROM banner and zero `boot:` lines, then a `^` frame in response to typing. The session survived a detach and reattach with no repaint, which is exactly what a 11.9 KB/s link wants. `zig build console` is 4 steps and builds no image, no app and no object. It installs `p4-console` itself rather than going through `installArtifact`, so `zig build` alone still lands exactly one file in `zig-out` - verified from an empty tree. Also here: * `tools/console.zig`'s `ModeWatch` tests were reachable by nothing. `zig build test` runs them now; the matcher has to resynchronise when the byte that broke a match is the next match's ESC, which is worth a regression test. * The port-permission advice existed twice, in `failPort` and in the new program. It now lives once, in `tools/serial.zig` beside the code that opens ports, and says nothing about the path so each caller can name its own on the first line.
* pardes as P4 firmware: the seam, and a flash-mapping bug in this toolchainGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The editor arrives as one freestanding OBJECT exporting a seven-function C ABI (src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not as a package dependency. A build.zig.zon path dependency was built first and reverted: merely DECLARING it nested pardes's ~30-package graph under this one and broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy. The seam is bytes in and bytes out, which is what a serial line is anyway: the editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the clock, and neither names the other's types. It is versioned, because linkers do not type-check C symbols and a drifted signature would link cleanly and then corrupt the stack. THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's anchor makes two mapped segments share an MMU page safely - as long as rodata does not END inside the page where text BEGINS. With a 578 KB image it does. A volatile read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared page read as code, so the first thing the firmware tried to print was machine code and it died on an instruction access fault. .flash.text is now ALIGN(0x10000), making the segments page-disjoint. The packing trick this project opened with only ever mattered when the alternative was 64 KiB of zeros in a 1 KB image. Two more findings, both recorded in README.md: * A linker symbol declared as an anyopaque OBJECT gives the optimiser a zero-sized object, so ordinary stores through a pointer derived from its address are dead code it may drop - and did, silently. The allocator's first block header read back as size=2988759312 next=0x14284684 and the free-list walk never terminated. @extern with a many-pointer has no size to lose. examples/memprobe.zig could not have caught it: it writes through a volatile pointer, which the optimiser must leave alone. * The RTC watchdog is armed at handover. Every example here had been resetting on a ten-second cycle, invisibly, because no run had ever lasted eight seconds. State, honestly: the firmware boots, clears .bss, brings up the console, disables the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB heap and calls into the editor, which sets up its sink and its environment. It then faults inside pardes_p4_init on the first allocation. The cause is measured but not fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000 higher - exactly the vaddr distance between the rodata and text segments - while pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in place, deliberately, because the next step needs them. --- correction, measured after the above was written --- Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig was right for a reason I initially got wrong and then measured. I first read bootloader_utility.c's `#else` branch, which classifies segments by address window with two independent ifs - and since the P4's DROM and IROM windows are the identical range (soc.h:146-149), I concluded the last mapped segment wins both roles and the first is never mapped. That branch does not run on this chip. The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851), whose own comment says it: "On chips with shared D/I external vaddr, we don't divide them into either D or I, as essentially they are the same." It collects mapped segments POSITIONALLY into rom_addr[2] and ends with assert(rom_index == 2); Shipping a one-segment image proved it, on the board: Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2) So the split stays, image.zig keeps enforcing exactly two - turning that boot-time abort into a build-time error - and both are now documented with the branch that actually runs and the assert that actually fires. What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may share a 64 KiB MMU page only if they also share a flash page, which the image builder's anchor guarantees - and that holds right up until an application is large enough for rodata to END inside the page where text BEGINS. With a 578 KB image it does. Measured on the die: a volatile read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own .flash.text. Every literal in that shared page read as code, so the first thing the firmware tried to print was machine code, and it died on an instruction access fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint. It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing trick this project opened with only mattered when the alternative was 64 KiB of zeros in a 1 KB image. With that fixed the firmware gets much further: entry, .bss cleared, console up, watchdog disabled, systimer running, ABI version checked, the 384 KiB heap initialised, into the editor, its sink and environment ready - and the literal at 0x4004a1d1 now reads back correctly. Still open, and characterised rather than guessed: pardes_p4_init faults on its first allocation. The allocator struct crosses the seam intact (its function pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it. Ruled out with measurements: the ELF and the image agree at that address, the flash is MD5-verified against the image, the wrong bytes are identical across three resets and two reflashes (so not a stale cache), the corruption is a contiguous run rather than 64-byte lines, and mmu_hal_map_region's arithmetic (page_num = ceil(len/page), entry from vaddr) is correct for the segments as now laid out. The next measurement is the one that settles it: read the MMU entry registers from the running application and print vaddr -> flash for every page. The register model in src/soc.zig can do that; the bisect markers are left in place for it.
* zig-p4: pure-Zig ESP32-P4 toolchainGabriel Schneider2026-08-25
build.zig generates the linker script and drives Zig's own LLD; tools/image.zig turns the ELF into a flashable image and tools/{rom,serial}.zig speak the mask ROM loader over the UART. No CMake, ninja, idf.py, esptool, or external linker. src/soc.zig is a comptime register model over ESP-IDF's own *_reg.h headers; src/hal/ adds peripheral sequences; src/io/ implements std.Io for the chip; src/oracle/ diffs this HAL against ESP-IDF's on the die.