<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/examples, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-09-16T14:28:40Z</updated>
<entry>
<title>Make the toolchain a package another build can drive, and move the editor's glue to the editor</title>
<updated>2026-09-16T14:28:40Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T16:28:33Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=b42ecaed412be2e30b9e780eb7c9e46e1535f26f'/>
<id>urn:sha1:b42ecaed412be2e30b9e780eb7c9e46e1535f26f</id>
<content type='text'>
</content>
</entry>
<entry>
<title>A test suite that runs on the die, and two build steps for reading an image</title>
<updated>2026-08-26T12:56:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T12:56:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=55743cc5d564f6ef6d6f8a0eb1a614a74b021d3a'/>
<id>urn:sha1:55743cc5d564f6ef6d6f8a0eb1a614a74b021d3a</id>
<content type='text'>
## zig build selftest

Seventeen checks, on the board, chosen by one rule: a check belongs there only if the die
can answer it and a host cannot. Every serious bug this port has produced was invisible to
a host test. `std.mem.eql` compares a byte at a time on this target, so the firmware's
largest read was three times slower than it needed to be and nothing on a laptop could
tell. A lone ESC resolves to the Escape key, which is right when a kernel hands over a
whole sequence and wrong when a 115200 line hands over one byte every 87 us. A full
transmit FIFO stopped anything draining the receiver, and the FIFO depth is a hardware
number.

So: the volatile promise (two reads of a live counter are two reads, which is what `Peek`
rests on), word-wise equality checked against `std.mem.eql` itself at every difference
position and both alignments, the input rescue against a fake port with a FIFO that loses
what arrives into a full one, the allocator on real L2MEM, and the cycle counter against
the systimer - which is clocked from the crystal and therefore cannot be flattered by a
wrong CPU divider.

Anything that is pure logic stays in `zig build test`, which is faster and needs no
hardware. Duplicating those here would make the suite longer and no stronger.

Not a `zig test` binary, deliberately: Zig's runner wants an OS and `std.testing.allocator`
is a debug allocator over the page allocator, which on freestanding is either a compile
error or a lie. The harness is thirty lines.

`selftest` has its own application, image and flash chain so it is one command with no
flags to remember, and it makes the board's verdict the build's - a suite whose result a
human has to read out of a scrolling log is a suite that gets ignored on the first busy
afternoon. Proven both ways: 17/17 with exit 0, and exit 1 naming the check when one is
deliberately inverted.

The clock check earned its place immediately. The first version read `config.cpu_mhz` and
compared the die against it without ever performing the raise, so `-Dcpu-mhz=360` failed
with `khz=90001 want=360000`. The check was right and the expectation was wrong; it now
calls `setCpuFreq` itself, which makes it a test of the raise rather than a tautology.
90001 kHz at 90, 360004 at 360.

## zig build layout, and a loader error that explains itself

Both of these exist because of an hour I spent that they would have saved.

`NotTwoMappedSegments` said the count was wrong and nothing about what the segments were,
which is the only thing that says which section grew, shrank or stopped being emitted. I
diagnosed one by hand with readelf on an artifact that turned out to be a stale install,
then guessing at the linker script. The image step now prints the segment table with the
error - it has to happen there, because a rejected image is never written, so no later step
can show it - and `zig build layout` prints the same table for an ELF plus the loader's
verdict as TEXT, which `size` cannot do because it reads a finished image and the moment you
need the table is when there isn't one.

They immediately paid for themselves. The reason the suite would not build was that it had
no `_start`: `-fentry=_start` found no such symbol, `--gc-sections` discarded every
function as unreachable, and `.flash.text` was empty. `layout` prints `entry 0x0` for
exactly that, in one line. The second failure - a silent board - was a missing app
descriptor, which is the same class of thing and now has a comment where it happened.

## The panic handler already existed, and was printing past the end of its message

`msg` is a Zig slice and `%s` reads until a NUL, so handing `msg.ptr` to the ROM's printf
printed the message and then whatever followed it in memory. String literals get away with
it; std's own panics do not, because they are formatted into a buffer - "index out of
bounds: index 5, len 3" - and carry no terminator. It now goes out through `uart.write`,
which takes a length, with the fault address after it so addr2line can find the line.
</content>
</entry>
<entry>
<title>A measuring instrument, and what it says about where the latency goes</title>
<updated>2026-08-25T21:41:18Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T21:41:18Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=1cef9c2e4bd873ebe13f5df635899231bcc467d2'/>
<id>urn:sha1:1cef9c2e4bd873ebe13f5df635899231bcc467d2</id>
<content type='text'>
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.

## The instrument

`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.

The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.

`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.

`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.

## What it measured

The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.

Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.

And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.

One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.

## Three bugs found while building it

The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.

Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.

`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.

Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.

## The report

`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
</content>
</entry>
<entry>
<title>pardes as P4 firmware: the seam, and a flash-mapping bug in this toolchain</title>
<updated>2026-08-25T17:38:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T16:06:05Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=174991b8f3f8e9c792eede7a52ad7beb10a08b05'/>
<id>urn:sha1:174991b8f3f8e9c792eede7a52ad7beb10a08b05</id>
<content type='text'>
The editor arrives as one freestanding OBJECT exporting a seven-function C ABI
(src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not
as a package dependency. A build.zig.zon path dependency was built first and
reverted: merely DECLARING it nested pardes's ~30-package graph under this one and
broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime
quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed
in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy.

The seam is bytes in and bytes out, which is what a serial line is anyway: the
editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the
clock, and neither names the other's types. It is versioned, because linkers do
not type-check C symbols and a drifted signature would link cleanly and then
corrupt the stack.

THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's
anchor makes two mapped segments share an MMU page safely - as long as rodata does
not END inside the page where text BEGINS. With a 578 KB image it does. A volatile
read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to
"lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared
page read as code, so the first thing the firmware tried to print was machine code
and it died on an instruction access fault. .flash.text is now ALIGN(0x10000),
making the segments page-disjoint. The packing trick this project opened with only
ever mattered when the alternative was 64 KiB of zeros in a 1 KB image.

Two more findings, both recorded in README.md:

  * A linker symbol declared as an anyopaque OBJECT gives the optimiser a
    zero-sized object, so ordinary stores through a pointer derived from its
    address are dead code it may drop - and did, silently. The allocator's first
    block header read back as size=2988759312 next=0x14284684 and the free-list
    walk never terminated. @extern with a many-pointer has no size to lose.
    examples/memprobe.zig could not have caught it: it writes through a volatile
    pointer, which the optimiser must leave alone.

  * The RTC watchdog is armed at handover. Every example here had been resetting on
    a ten-second cycle, invisibly, because no run had ever lasted eight seconds.

State, honestly: the firmware boots, clears .bss, brings up the console, disables
the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB
heap and calls into the editor, which sets up its sink and its environment. It then
faults inside pardes_p4_init on the first allocation. The cause is measured but not
fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000
higher - exactly the vaddr distance between the rodata and text segments - while
pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in
place, deliberately, because the next step needs them.


--- correction, measured after the above was written ---

Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig
was right for a reason I initially got wrong and then measured.

I first read bootloader_utility.c's `#else` branch, which classifies segments by
address window with two independent ifs - and since the P4's DROM and IROM windows
are the identical range (soc.h:146-149), I concluded the last mapped segment wins
both roles and the first is never mapped. That branch does not run on this chip.
The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851),
whose own comment says it: "On chips with shared D/I external vaddr, we don't
divide them into either D or I, as essentially they are the same." It collects
mapped segments POSITIONALLY into rom_addr[2] and ends with

    assert(rom_index == 2);

Shipping a one-segment image proved it, on the board:

    Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2)

So the split stays, image.zig keeps enforcing exactly two - turning that boot-time
abort into a build-time error - and both are now documented with the branch that
actually runs and the assert that actually fires.

What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may
share a 64 KiB MMU page only if they also share a flash page, which the image
builder's anchor guarantees - and that holds right up until an application is large
enough for rodata to END inside the page where text BEGINS. With a 578 KB image it
does. Measured on the die: a volatile read of a string literal at 0x4004a1d1
returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own
.flash.text. Every literal in that shared page read as code, so the first thing the
firmware tried to print was machine code, and it died on an instruction access
fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint.
It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing
trick this project opened with only mattered when the alternative was 64 KiB of
zeros in a 1 KB image.

With that fixed the firmware gets much further: entry, .bss cleared, console up,
watchdog disabled, systimer running, ABI version checked, the 384 KiB heap
initialised, into the editor, its sink and environment ready - and the literal at
0x4004a1d1 now reads back correctly.

Still open, and characterised rather than guessed: pardes_p4_init faults on its
first allocation. The allocator struct crosses the seam intact (its function
pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads
back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it.
Ruled out with measurements: the ELF and the image agree at that address, the flash
is MD5-verified against the image, the wrong bytes are identical across three
resets and two reflashes (so not a stale cache), the corruption is a contiguous run
rather than 64-byte lines, and mmu_hal_map_region's arithmetic
(page_num = ceil(len/page), entry from vaddr) is correct for the segments as now
laid out. The next measurement is the one that settles it: read the MMU entry
registers from the running application and print vaddr -&gt; flash for every page. The
register model in src/soc.zig can do that; the bisect markers are left in place for
it.
</content>
</entry>
<entry>
<title>zig-p4: pure-Zig ESP32-P4 toolchain</title>
<updated>2026-08-25T15:46:51Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T15:40:53Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=f5f8068fac59b4f16046c2022c2fc7c7e447ef4c'/>
<id>urn:sha1:f5f8068fac59b4f16046c2022c2fc7c7e447ef4c</id>
<content type='text'>
build.zig generates the linker script and drives Zig's own LLD; tools/image.zig
turns the ELF into a flashable image and tools/{rom,serial}.zig speak the mask
ROM loader over the UART. No CMake, ninja, idf.py, esptool, or external linker.

src/soc.zig is a comptime register model over ESP-IDF's own *_reg.h headers;
src/hal/ adds peripheral sequences; src/io/ implements std.Io for the chip;
src/oracle/ diffs this HAL against ESP-IDF's on the die.
</content>
</entry>
</feed>
