# zig-p4 — an ESP32-P4 toolchain that is just Zig ``` zig build # compile, link, and emit a flashable image zig build flash # ...then write it to the chip and run it zig build run # flash, then print the console (ordered; `flash monitor` is not) zig build monitor # reset the board and print its console zig build console # attach a terminal to whatever is already on the board (Ctrl-] detaches) zig build interact # flash, then attach that terminal (ordered, like `run`) zig build reset # just pulse the reset line zig build size # where every byte of the image went zig build bench # measure the serial link and the editor, verified with a checksum zig build test # host tests: image builder, and the register layer's field arithmetic zig build diff # the hardware oracle: this HAL vs ESP-IDF's, on the die (needs -Doracle) zig build elf # stop at the ELF, for disassembly ``` No CMake, no ninja, no `idf.py`, no `esptool`, no external linker. Zig's own LLD does the link against a linker script this `build.zig` generates; the image builder and the serial flasher are ordinary Zig code in `tools/`, imported straight into `build.zig`, so they leave no artefacts of their own. What lands in `zig-out` is one file: the image. One exception, and it earns it: `zig build console` also installs `zig-out/bin/p4-console`, a host binary that opens the port and nothing else. Run it directly and there is no build runner in the picture — which matters, because an interactive step lives for as long as the human does, and `std.Progress` would otherwise redraw the build tree over the screen every 80 ms. std solves that for child processes by holding `io.lockStderr()` for the child's whole lifetime (`std/Build/Step/Run.zig:1588-1592`), which is the same lock `Progress` needs to draw, so both spellings are clean; the binary is simply the one that assumes the board is already flashed. ``` zig-out/bin/p4-console # the board is already programmed; just connect zig-out/bin/p4-console --no-reset # ...and do not pulse reset, so a live session survives zig-out/bin/p4-console --port /dev/ttyUSB1 --baud 115200 ``` ## Why it exists | | ESP-IDF blink | earlier Zig proof-of-concept | this | |---|---|---|---| | flashed image | 184,112 B | 66,176 B | **1,216 B** (432 B for `examples/minimal.zig`) | | build, cold | 23 s, 1,072 ninja edges | 1.3 s | **1.37 s** | | build, warm | ~1 s | 0.07 s | **0.088 s** (cache hit, no work) | | build, one file edited | ~1 s | 0.07 s | **0.121 s** | | flash + verify + run | ~2 s (esptool + stub) | ~2 s | **0.277 s** | | host dependencies | ESP-IDF 663 MB + toolchain 3.4 GB + Python | Zig + system LLD + esptool | **Zig** (the two pardes steps also want the sibling editor checkout — see Requirements) | | artefacts per build | ~1,100 files | 2 | **1** | The 66 KB → 1 KB step is the interesting one. `esptool` refuses to put two flash-mapped segments inside the same 64 KiB MMU window (`bin_image.py:832-838`) and pads the image out to the next window, which for a small program is 64 KiB of zeros. The chip only requires that each mapped segment satisfies ``` (offset of its data within the image) % 64 KiB == (its load address) % 64 KiB ``` (`esp_image_format.c:903-909`), and the MMU is perfectly happy to point two entries — or the same entry twice — at one flash page; stock ESP-IDF already aliases a page that way on every boot. `tools/image.zig` therefore packs both segments into one window and the pad disappears. Verified on ESP32-P4 rev v1.3 silicon. ## Requirements * Zig 0.16.0, plus a network fetch on the first build. `build.zig.zon` now pins exactly one package, `cloud9` — the base 9P2000 implementation — and only the GPIO 9P application reaches for it (`zig build -Dapp=../02-pardes-code/src/esp32p4_9p.zig`). Zig fetches it into `zig-pkg/`, it has no dependencies of its own, and nothing else here compiles it: `zig build`, everything under `examples/`, the host tests and the harnesses want nothing beyond Zig and this checkout. * Two steps are the exception, because they build somebody else's program: `zig build -Dpardes` and `zig build selftest` read source across a sibling-relative path from the pardes editor's checkout at `../02-pardes-code/` (the application root and `input_rescue.zig` for the former, the on-die suite for the latter), and `-Dpardes` also links `-Dpardes-obj`, which that checkout's own `zig build -Dplatform=esp32p4` emits into its `zig-out/`. Still not a package dependency — the seam is files on disk — but without that checkout beside this one those two steps cannot run. * Membership of whatever group owns the serial port (`uucp` on Arch, `dialout` on Debian). * An ESP-IDF second-stage bootloader and partition table already in flash at `0x2000` and `0x8000`. This toolchain builds and flashes *applications*; the bootloader is still Espressif's. See "What is still ESP-IDF" below. ## Options Everything is a `b.option`, so `zig build -h` lists them all. | option | default | meaning | |---|---|---| | `-Dapp=` | `src/main.zig` | application root source | | `-Dport=` | `/dev/ttyUSB0` | serial port | | `-Dbaud=` | `b921600` | flashing baud. `b2000000` does not work on this board's CH340 | | `-Dled=` | `20` | GPIO the demo blinks (20 = JP1 pin 17) | | `-Doffset=` | `0x10000` | flash offset of the app partition | | `-Dflash-size=` | `16MB` | fitted flash, written into the image header | | `-Dmin-rev`/`-Dmax-rev` | `100`/`199` | silicon revision window. The pre-v3 P4 needs 100..199 | | `-Ddescriptor=` | `minimal` | 184-byte descriptor, or `full` for the 256-byte one `esptool image-info` can parse | | `-Dstack=` | `8192` | stack size; the generated linker script follows. `32768` under `-Dpardes` | | `-Dverify=` | `true` | ask the ROM for an MD5 of what it stored and compare | | `-Delf=` | `false` | also install the ELF | | `-Doptimize=` | `ReleaseSmall` | firmware default, not Debug (Debug costs ~780 B here) | | `-Dseconds=` | `5` | how long `monitor` listens | | `-Dconsole-baud=` | `b115200` | the interactive console's rate: what the bootloader leaves UART0 at. Distinct from `-Dbaud`, which the ROM loader auto-detects | | `-Dpardes=` | `false` | build the pardes editor as the application. Needs the object below | | `-Dpardes-obj=` | `../02-pardes-code/zig-out/pardes-esp32p4.o` | the editor, compiled freestanding by its own build and linked here | ### Driving the editor over the wire `zig build interact -Dpardes` boots the editor with **one empty output buffer** and no shell pane — a shell is not a layout preference on bare metal but an impossibility, since there is nothing to fork and no pty to give a terminal pane. Booting one anyway produced exactly that: a pane with no gutter, no buffer, and every keystroke vanishing into the Fallback's silent pty. Three words exist only here, because with no OS there is no MMU and no supervisor, so all 2^32 addresses are legitimately this program's (`../02-pardes-code/src/board_memory.zig`): | word | does | |---|---| | `Peek [count]` | `count` 32-bit words, one `addr: value` row each | | `Poke ` | one 32-bit store, then a load back — the read-back is the point, since MMIO rarely returns what you wrote | | `Hexdump [len]` | `len` bytes, 16 to a row, hex columns and an ASCII gutter | There is no mouse on a serial line, so execution is the keyboard chord: write the command into the buffer, select it, execute the selection (`pardes.zig:7669`). ``` i Peek 0x501101a4 ESC type it into the buffer, then back to normal mode x select the line TAB execute the selection; `o` opens a fresh line for the next one ``` Measured on ESP32-P4 rev v1.3 silicon: two `Peek`s of the RNG register at `0x501101a4` returned `0x0e63ce71` then `0xaeaa6919`, so the `*allowzero volatile` reads are genuinely not folded; `Poke 0x5011002c 0xdeadbeef` into `LP_STORE0` read back as `0xdeadbeef` from a later `Peek`; and `Peek 0x50110001` answered `peek: MisalignedAddress` on the message row rather than taking the session down with an unhandled trap, which is the one fault that file exists to prevent. ### Measuring it `p4-bench` exists so that optimising this port is not a matter of opinion. It has two halves, because there are two different questions. ``` zig build flash -Dapp=examples/uartperf.zig # the ceiling: link + driver, nothing else zig-out/bin/p4-bench --link zig build flash -Dpardes # how much of that ceiling the editor uses zig-out/bin/p4-bench --editor zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall ``` Every `--link` number is checksummed. `tools/perfproto.zig` is a framed protocol - `"P4"`, op, length, CRC-32, payload - shared *verbatim* by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. That matters because RX overrun on this UART is undetected in hardware and uncounted in the driver: a byte that never arrived is indistinguishable from a late one, and an unchecksummed throughput figure is a guess about how fast data was corrupted. `tools/rtt.zig` holds the two timing functions everything is built on. Round trip is measured to the **first** response byte, not the last: a renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate, one that thinks for 130 ms then draws in 8 ms feels broken, and waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds throughout, because at 115200 one byte is 87 us and a millisecond clock would quantise the answer into buckets eleven bytes wide. `experiments/` holds the raw per-trial CSVs and `report.typ`, which reads them and computes its own figures - so a re-run changes the document rather than contradicting it. What it establishes on this die: | | | |---|---| | link, both directions | 99.8% of the 11,520 B/s wire, CRC verified over 32,768 B | | typing, 6 to 100 keys/s | nothing lost, wire never above 9% | | one keystroke | 81 B, round trip 17.0 ms | | a motion | 40 B, round trip 16.7 ms | | cost per character already in the line | **54.3 us, per keystroke** | The first two rows say the wire is not the problem. The last three say why: a 40-byte operation and a 206-byte one cost the same round trip, so latency is compute per event and not transmission - and it grows with the document, because the edit path copies the whole buffer every keystroke. Building the editor object `ReleaseFast` instead of `ReleaseSmall` cuts the fixed cost 13% and the per-character cost 36% for 35% more flash, which is the best ratio measured here. ## Layout ``` build.zig the toolchain: target, generated linker script, and four custom steps tools/image.zig ELF -> ESP image. Header, segments, congruence filler, checksum, SHA-256 tools/image_test.zig 8 host tests, one per rule the ROM bootloader enforces tools/rom.zig SLIP framing + the ROM loader protocol. No software stub tools/serial.zig termios2 raw mode, arbitrary baud, DTR/RTS reset dance tools/console.zig the interactive bridge: raw stdin <-> UART, and the window-size handshake src/soc.zig comptime register model: GPIO, IOMUX, mask-ROM entry points, cycle counter src/appdesc.zig esp_app_desc_t, linked as its own object so it cannot be optimised away src/main.zig demo: prints what it can prove, then blinks -Dpardes app root ../02-pardes-code/src/esp32p4/ — entry, heap, UART, input rescue, on-die suite examples/minimal.zig the floor: 432 B, blinks and nothing else examples/echo.zig UART0 duplex echo: the proof that receive works on the die examples/memprobe.zig what RAM this board actually has, measured rather than assumed examples/heapcheck.zig the allocator under an editor's workload, on the die ``` ## What RAM this board has Measured by `examples/memprobe.zig`, on the die, because it cannot be read off ESP-IDF's linker fragments. The relevant one (`esp_system/ld/esp32p4/memory.ld.in:18-33`) is parameterised on `CONFIG_CACHE_L2_CACHE_SIZE`, and that Kconfig's own help text says the size is set "on application startup" — by an application this is not. ``` 0x4FF03000..0x4FF3F000 240 KiB RAM .data/.bss/.stack live at the bottom of this 0x4FF3F000..0x4FF40000 4 KiB ROM the mask ROM's .data/.bss; ets_printf needs it 0x4FF40000..0x4FFA0000 384 KiB RAM handed over whole as __heap_start..__heap_end 0x4FFA0000..0x4FFC0000 128 KiB cache NOT memory: the L2 cache lives here 0x48000000 PSRAM 32 MB fitted, untrained; touching it hangs the core ``` That last RAM line cost a bug worth repeating, because the first version of the probe reported the whole upper 512 KiB as usable. It wrote a pattern to a page and read it straight back, one page at a time — and a store followed immediately by a load of the *same* address returns the stored value whether the backing store is real memory, an address mirror, or merely a dirty cache line. Writing every page before reading any page separates the three, and the top 128 KiB then failed. ESP-IDF's own arithmetic agrees exactly: `SRAM_HIGH_SIZE = 0x80000 - CONFIG_CACHE_L2_CACHE_SIZE`, with the 128 KiB default from `esp_system/port/soc/esp32p4/Kconfig.cache:5,19`. The wrong number had already been committed to the linker script, where it handed 128 KiB of live L2 cache to an allocator. PSRAM stays untrained deliberately. ESP-IDF's ESP32-P4 implementation runs past a thousand lines — MPLL, MSPI clocking, pin drive and DQS, CS timing, mode registers, a connectivity check, and a whole timing-calibration subsystem — and the mask ROM exposes only MMU mapping (`Cache_PSRAM_MMU_Init`, `Cache_PSRAM_MMU_Set`), no device init. The probe reads `0x48000000` on purpose and hangs there, which is why it prints its cursor before every access rather than after. **The RTC watchdog is armed when the bootloader hands over**, and it expects the application to take it over. Nothing here did, so every example in this repo had been resetting on a ten-second cycle, invisibly, for as long as no run lasted eight seconds. `examples/heapcheck.zig`'s 20,000 allocations is the first run that did. `hal.rwdt.disable()` is the fix and `hal.rwdt.armed()` is worth printing at startup. ## What adversarial review found Five reviewers went at this in two rounds; sixteen findings were applied. Two more rounds went at the memory map and the editor port and found the three below first. The ones worth knowing about, because each is a trap the next person will hit too: * **A linker symbol declared as an object gives the optimiser a zero-sized object, and it will drop your stores.** `extern const __heap_start: anyopaque` plus `@intFromPtr`/`@ptrFromInt` looks like the obvious way to reach a region the linker script defines. The pointer it produces carries provenance for zero bytes, so the ordinary (non-volatile) store the allocator makes through it is dead code the backend may remove — and did. The first block header read back as `size=2988759312 next=0x14284684` instead of `{393216, 0xFFFFFFFF}`, the free-list walk followed garbage, and with asserts compiled out in `ReleaseSmall` that is a silent hang with no console output after `MARK HEAP_INIT`. `@extern([*]u8, .{ .name = "__heap_start" })` has no size to lose. Note that `examples/memprobe.zig` could not have caught this: it writes through a `volatile` pointer, which the optimiser must leave alone. The two files disagreed about whether the same address worked, which is what made it findable. * **A one-page memory probe cannot tell RAM from a mirror or a cache line.** See the section above: writing and reading the same address back-to-back succeeds in all three cases, and the pattern being address-derived does not help because the alias is written *and* read through the alias. This one had already shipped a wrong 512 KiB into the linker script. * **Ignoring `POLL.HUP`/`ERR`/`NVAL` is a hot spin, not a no-op.** `tools/console.zig` polls stdin and the port and acted only on `POLL.IN`. Unplug the CH340 mid-session and `revents` carries `HUP|ERR|NVAL` forever: `poll` returns immediately with a non-zero count, neither branch matches, and the loop burns a core with the terminal still in raw mode and `ISIG` off — so Ctrl-C cannot even end it. `std.posix.poll` cannot report it as an error either, because a dead descriptor is delivered in `revents` rather than errno (`std/posix.zig:1007-1017` maps `INVAL` to `unreachable`). * **The ROM's status byte is at `data[len-4]`, not `data[len-2]`.** The four-byte trailer is a ROM-versus-stub difference (esptool `loader.py:653-655`). Reading the wrong byte made *every* ROM error read as success: a rejected `FLASH_DATA` block was never retried and `zig build flash` printed a cheerful success line over a dead image. Only the MD5 pass caught it. Now the offset is right, the length check is mandatory, and `-Doffset=33554432` (past the end of a 16 MB part) fails with `write failed: CommandFailed` instead of claiming victory. * **Congruence modulo the MMU page is necessary but not sufficient.** The first solver satisfied it by shifting a segment forward, which could leave both mapped segments in one *vaddr* page while their data sat in two different *flash* pages. The bootloader writes one MMU entry per vaddr page, so the second mapping replaced the first and the app booted with every constant reading as zero - no error anywhere. `-Ddescriptor=full` triggered it. The solver now anchors every mapped segment to a single `flash_address - load_address`, and `validate` checks the invariant directly. * **The target enables `f` but nothing enabled the FPU.** `mstatus.FS` is Off after reset and ESP-IDF only turns it on lazily from a trap handler, which this image does not have, so the first `f32` multiply in application code was an unhandled illegal instruction. `_start` now sets FS, and the demo prints a float computed with hardware `fmul.s` as proof. * **`-Dmin-rev` was feeding the descriptor's eFuse *block* revision**, an unrelated field, so `-Dmin-rev=150` produced an image the bootloader refuses with `Image requires efuse blk rev >= v0.50`. * **`flash` and `monitor` are unordered top-level steps** and the build runner runs independent steps concurrently, so `zig build flash monitor` could pull the chip out of download mode mid-write. They now share a mutex - and because a mutex cannot express *order* (measured: monitor won 3 times out of 3, printing the old firmware before the new one was written), `zig build run` exists as the ordered version. * **A cache manifest that does not hash the builder is worse than no cache.** The first version hashed only the ELF and the options, so editing `tools/image.zig` produced a cache hit and shipped the previous image - and `--watch` never noticed the edit at all. The manifest now covers the builder sources too. * **`ALIGN(64)` does not reserve a hole.** The image builder needs 8 spare bytes before the second mapped segment for its header; for one rodata length in eight, 64-byte alignment leaves 0 or 4, and the build failed with `MappedSegmentsTooClose`. The generated script now ends `.flash.rodata` with `. = ALIGN(. + 8, 64) - 8;`, verified over a sweep of rodata sizes. * **A constant the compiler folds proves nothing.** The FPU check computed `7 * 1.5 + 0.25` from a literal, so LLVM folded it and the image contained no float instructions at all: the test would have passed on a board whose FPU was still off. It now loads through a volatile pointer, and the image really does contain `fcvt.s.wu`, `fmul.s`, `fadd.s`, `fcvt.wu.s`. ## Three things that bit, and are now encoded in the code 1. **`standardOptimizeOption` hands out Debug builds.** With `preferred_optimize_mode` set it exposes `-Drelease` and still defaults to Debug, which for this target means panic machinery and formatting code inside a 500-byte image. `build.zig` takes `-Doptimize` and defaults to `ReleaseSmall` explicitly. 2. **An `@import`ed descriptor disappears under ReleaseSmall.** A `comptime _ = descriptor;` reference is enough in Debug; in release the constant folds away, `.flash.rodata` vanishes, `KEEP` has nothing to keep, and the image boots with one mapped segment at the wrong offset. The descriptor is now a separate object linked unconditionally. 3. **`tcsetattr` cannot set the baud.** `TCSETS`' struct has no `ispeed`/`ospeed`, so assigning those fields silently does nothing and the port stays at whatever rate it had — a 1.5 KB flash took 230 ms instead of 60. `tools/serial.zig` uses `TCGETS2`/`TCSETS2` with `BOTHER`, and declares `struct termios2` itself because std's version is 60 bytes where the kernel's is 44, which makes std's ioctl numbers wrong. ## What is still ESP-IDF The second-stage bootloader at `0x2000` and the partition table at `0x8000`. Both are ordinary flash contents and neither is rewritten here. The bootloader is what enforces the three rules this toolchain obeys: * exactly two segments in the mapped range (`bootloader_utility.c:842`) * every segment length a multiple of 4 (`esp_image_format.c:857`) * mapped segments congruent modulo the MMU page (`esp_image_format.c:903-909`) `tools/image.zig` re-checks all three after building an image, so a violation is a build error rather than a board that resets in a loop. ## The HAL: registers from ESP-IDF, sequences in Zig A toolchain that can only blink an LED is a demo. The rest of the chip needs a hardware layer, and the ESP32-P4 has a lot of chip: 96 `*_ll.h` headers in ESP-IDF v6.0.2 holding **3,081** inline functions over **5,170** registers and **20,052** fields. Hand-transcribing that is not a plan — the hand-written predecessor of `src/hal/gpio.zig` had a wrong matrix constant with a comment warning about exactly that mistake. So the register layer is not written at all. `build.zig` runs `zig translate-c` over a generated C file that `#include`s every one of ESP-IDF's own `*_reg.h` headers, and the result is imported as a module: ```zig const regs = @import("regs"); // 87,373 constants, straight from IDF's macros ``` That takes 0.26 s and costs about 0.16 s of parse on top of a 1.7 s cold build; Zig only analyses the declarations actually referenced, so the unused 87,000 are free. Every address, shift and mask in this HAL is therefore not a re-derivation of ESP-IDF's number — it *is* ESP-IDF's number, as evaluated by clang. `*_struct.h` is deliberately unused: translate-c demotes each of those register structs to `opaque {}` ("has bitfield"), so the C bitfields buy nothing. `src/mmio.zig` is the ~200 lines that turn flat constants into checked accessors: ```zig const conf0 = mmio.Reg.at(regs.LEDC_CH0_CONF0_REG); const timer_sel = mmio.Field.of(regs.LEDC_TIMER_SEL_CH0_S, regs.LEDC_TIMER_SEL_CH0_V); conf0.modify(.{ timer_sel.is(2) }); // read-modify-write, preserving the rest ``` Fields are built from the `_S`/`_V` pair and never from `_M`, which is not a style rule: **153 `_M` macros are broken C inside ESP-IDF itself** (`INTERRUPT_CORE0_LP_RTC_INT_MAP_M` expands `CORE0_LP_RTC_INT_MAP_V`, dropping the prefix). Nothing in C ever expanded them, so nobody noticed; `translate-c` surfaces them as poisoned declarations. `Field.of` also rejects a pre-shifted mask at comptime, so passing `_M` by hand is a build error rather than a wrong bit position. `modify` is the default and `write` is the exception, because `write` zeroes what it does not name and 46.7% of the registers a low-level driver touches have a field whose reset value is not zero. ### The build refuses to hide what it could not translate `translate-c` exits 0, prints nothing, and still emits `pub const X = @compileError(...)` for every macro it could not handle — invisible until a driver names one. So a build step counts them and fails if the number grows past a recorded 524, whose composition is written down: 333 register addresses whose `DR_REG_*_BASE` **ESP-IDF references and never defines anywhere**, 153 broken `_M` masks, and 38 function-like macros with Zig equivalents. Three of those missing bases were recovered from `esp32p4.peripherals.ld` — which cross-checks against `reg_base.h` on all 65 peripherals both files name, 0 disagreements — putting 176 registers back. The rest stay unreachable on purpose rather than reachable at a guessed address. ### The oracle: differential testing against ESP-IDF, on the die Register numbers being right does not make a *sequence* right. So ESP-IDF's own `*_ll.h` functions are compiled by Zig's clang into the same image as this HAL, and `zig build diff` drives each operation both ways on the chip and compares the register block afterwards: ``` MARK DIFF_CFG gpio_ll_uses_rom_api=0 expect=0 MARK DIFF ok gpio.set_level(1) 400 words identical ... MARK DIFF_TOTAL cases=26 failures=0 ``` It found two real bugs in this HAL on its first run. `GPIO_FUNC0_OEN_SEL` reads backwards from its name — 1 means "use `GPIO_ENABLE_REG`", 0 means "use the peripheral's own output enable" — so `matrixOut` had it inverted and then set the matching `GPIO_ENABLE` bit to compensate. It worked, by the wrong mechanism, leaving the pad latently output-enabled. The first version of the harness also only compared 112 words and so saw the symptom without the cause; the window now reaches the matrix configuration registers at +0x558. Four things the harness has to get right, each measured on this board rather than assumed: * **A snapshot can have side effects.** `UART_FIFO_REG` is at offset 0x000 of every UART block — the first word a "read the whole block" loop touches — and reading it pops the RX FIFO. The header annotates it `RO`. Peripherals declare offsets that must not be read. * **A block cannot be restored by writing its snapshot back.** ~10% of this chip's fields act when written; writing one saved word back to a UART's offset 0 transmits a character. Restore is the peripheral's reset bit, or a deliberate configure function. * **A clock-gated block reads stale data, silently** — the last value latched, not zeros, so two meaningless snapshots can compare equal. The bus clock is checked before every comparison. * **Equal registers do not prove equal sequences.** LEDC commits shadow registers through a self-clearing bit that leaves no trace afterwards. ### What is ported | peripheral | what it covers | differential cases | |---|---|---| | `hal.gpio` | 57 pins across both banks, IO MUX pads, pulls, drive strength, open drain, the GPIO matrix in and out | 46 | | `hal.timg` | TIMG0/1, two timers each: dividers, direction, auto-reload, alarms, the latch-then-read counter, and MWDT behind its write-protect key | 29 | | `hal.uart` | UART0-4: the fractional baud divider, data format, FIFOs, loopback, pin routing, and the `_SYNC`/`REG_UPDATE` commit | 19 | | `hal.intr` | the CLIC (not a PLIC): the 122-source interrupt matrix, per-line enable/trigger/priority, the memory-mapped threshold, a 48-entry vector table | 20 | | `hal.ledc` | Q10.8 timer dividers, channels, duty, idle level, the shadow-register commits, pin routing | 28 | | `hal.i2c` | I2C0/1 master: the full ten-register timing set plus its out-of-block clock divider, a typed command list, FIFOs, transaction status | 43 | | `hal.clkrst` | peripheral clock gates and resets, under an interrupt-masked read-modify-write guard | 8 | | `hal.systimer` | the two 52-bit counters, through their update/valid handshake | — | | `hal.rwdt` | the RTC and super watchdogs, which the bootloader leaves armed | — | **193 cases, 0 failures** on the die. That number was 177 until an adversarial audit of the harness pointed out that about seventy of the cases could not fail for any implementation error, which is worth more than the passing count was. Two structural causes, both now fixed: * **A window that missed the registers under test.** Every pad-configuration function writes the IO MUX at `0x500E1004 + 4*pin`, and GPIO's compared window ended at `0x500E063F` - 0xC00 bytes short. Eight operations across two pins were comparing two identical snapshots of a register file none of them touch. `iomux_suite` is the same cases against the right window. * **Restores written with the code under test.** The harness runs restore, IDF, snapshot, restore, ours, snapshot. If restore calls the HAL, run B starts from whatever IDF just wrote, and a HAL function that does nothing at all compares equal - so each suite was blind to exactly the failure it existed to catch. `clkrst`'s restore called `setClockEnabled`, which is the function whose wrong-register bug started this whole line of work. Every restore is now built from register macros or from ESP-IDF's own LL, never from ours. `src/hal` and `src/mmio.zig` are 4,590 lines; the oracle that checks them is 4,107. `examples/halcheck.zig` exercises the first three on hardware directly. One number out of it is worth keeping: the CPU runs at **89,995 kHz**, measured by counting cycles against the systimer's fixed 16 MHz over 50 ms rather than estimated — two independent clocks, not one clock and an assumption. ### Three traps now encoded in the code **Peripheral clocks are already on.** `esp_system/port/soc/esp32p4/clk.c:200` says so and the reset defaults agree, so the hazard is not gating but atomicity: every gate and reset bit shares a register with unrelated peripherals, and ESP-IDF makes an unguarded call *uncompilable* by referencing an identifier it never defines. `clkrst.maskInterrupts()` is the replacement for the spinlock IDF uses. **Resetting a timer group re-arms flash-boot watchdog protection**, which reboots the board a moment later with nothing on the console to explain it. `resetPeripheral(.timg0)` clears it as part of the reset. **The bootloader leaves the RTC watchdog running**, and expects the application to take it over — which `esp_system` does and a bare image never did. Every demo in this repo had been resetting on a ten-second cycle since the first one, invisibly, because no run had ever lasted eight seconds. It surfaced when the differential grew past 64 cases and stopped fitting inside one watchdog period, where it looked exactly like "the newest suite crashes the board". `hal.rwdt.disable()` is the fix and `hal.rwdt.armed()` is the thing to print at startup. ## Reproducing the report's RF experiment `04-report` section 5 is the only real SDR measurement in that document: a 25 MHz carrier on GPIO20, OOK-keyed with `0x4200`, detected at +12.96 dB and decoded back from raw IQ. `examples/rf.zig` is that emitter driven by this HAL instead of ESP-IDF, and the report's own capture and decode tools are used unmodified. Full write-up and the numbers: `captures/RF-REPRODUCTION.md`. The emitter reproduces exactly. The report derives its carrier from the divider — 1-bit resolution off 80 MHz can only give `80e6 × 256 / (410 × 2)` = 24,975,609 Hz, never a round 25 MHz — and this HAL computes divider 410 and 24,975,610 Hz from IDF's own Q10.8 arithmetic. The source clock was then measured rather than assumed: 80,104,000 Hz, timed against the systimer. The differential harness gained a case that runs the whole bring-up both ways in one image, and the LEDC block comes out **96 words identical**; an edge count on the pad agrees to 2 parts in 112,500 at every rate the instrument can measure. **The detection does not reproduce, and not because of this toolchain.** ESP-IDF's own binary was reflashed and driven by the report's own tool with the report's parameters: | firmware | ON − OFF | |---|---| | report, §5.2 | **+12.96 dB** | | ESP-IDF, re-run today | −0.063 dB | | this toolchain, today | +0.05 dB | The receive chain checks out — an FM carrier sits 31.8 dB above the span median, against the report's 30.9 dB — so the instrument is fine and both firmwares now give the same null. The report says its emission reached the dongle through whatever wire happened to be on JP1 pin 17; that coupling is gone. A measurement whose apparatus is incidental coupling is not reproducible, and this one is not. That is a correction to the report rather than a difference between toolchains. Wi-Fi, BLE and 802.15.4 stay out of scope by construction: the P4 has no radio, and those went over an ESP32-C6 across SDIO under `esp_hosted` plus a prebuilt coprocessor binary. Reproducing them means porting `esp_hosted`, not porting a HAL. ## The vendor ISA extensions, from a stock Zig The P4's cores carry Espressif's `xespv2p1` vector unit ("PIE") and the `xesploop1p0` hardware loops, which upstream LLVM does not know. That is really two gaps: 1. **The optimiser will never choose one.** No cost model, no intrinsics, no autovectorisation. That ceiling does not lift — but Espressif's own fork does not autovectorise into this unit either, which is why `esp-dsp` is hand-written assembly. 2. **The assembler cannot spell one.** Inline assembly does *not* help; it goes through the same integrated assembler: `asm volatile ("esp.vld.128.ip q0, a0, 16")` → `error: :1:2: unrecognized instruction mnemonic`. This is what stops Zig from assembling ESP-IDF's FreeRTOS context switch (`portasm.S`, 34 instructions under `#if SOC_CPU_HAS_PIE`). Gap 2 closes without forking a compiler: `.insn` accepts any word, so generate the words offline and pin the registers the fixed encoding names. `tools/encode.sh` is that oracle — it drives Espressif's GAS at build-authoring time, never at build time, and prints a `.insn` line per mnemonic. All 34 of `portasm.S`'s instructions encode, and the vld/vst/qacc words match what Espressif's objdump reads out of an ESP-IDF build of this board byte for byte. Watch the version: the extension is versioned and the versions disagree. `esp.vld.128.ip q0, a0, 16` is `0x0311009f` under `xespv` (which defaults to 1.0) and `0x0201223b` under `xespv2p1`. The P4 wants the latter — the pair in every IDF image's `.riscv.attributes` here is `xesploop1p0_xespv2p1`. `examples/pie.zig` runs the whole thing on the die: enable the unit (`csrw 0x7f2, 1`), a 16-byte vector load/store round trip, sixteen 8-bit MACs in one `esp.vmulas.s8.qacc`, and a 256-element dot product both ways for a price tag on gap 1: ``` MARK PIE_MAC lanes=16 distinct=0000ffff expect=0000ffff sum_ok=1 MARK PIE_DOT scalar=3045 in 2390 cyc, vector=3045 in 1004 cyc, agree=1 ``` Same answer, 2.4× fewer cycles, naive port — and the compiler would never have written it. What you give up: no operand checking, registers pinned by hand (`"{a0}"`), and the q registers are invisible to LLVM, so keeping one live across two `asm` blocks is only sound while nothing else in the image touches them. ## Known limits * The CPU runs at whatever the bootloader left it at — measured 90 MHz, not 360. Raising it means programming the PLL and the MSPI timings; nothing here does that yet. * No PSRAM init, no cache tuning, no interrupt controller setup, no FreeRTOS. That is the point of the floor, but it means `esp_wifi` and friends are not available: the radio on this board is an ESP32-C6 reached over SDIO by ESP-IDF's `esp_hosted`, which is a much larger dependency. * The flasher speaks the ROM protocol only. No software stub, so no compressed writes; for a 1 KB image that costs nothing, and for a 4 MB one it would cost about a factor of two. * Linux only: `tools/serial.zig` uses `termios2` and Linux ioctl numbers directly.