<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/src/pardes/uart.zig, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-09-16T14:28:40Z</updated>
<entry>
<title>Make the toolchain a package another build can drive, and move the editor's glue to the editor</title>
<updated>2026-09-16T14:28:40Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T16:28:33Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=b42ecaed412be2e30b9e780eb7c9e46e1535f26f'/>
<id>urn:sha1:b42ecaed412be2e30b9e780eb7c9e46e1535f26f</id>
<content type='text'>
</content>
</entry>
<entry>
<title>Drain the receiver while the transmitter is full: keystrokes were being lost</title>
<updated>2026-08-26T04:13:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T04:13:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=e1526395233aad15dc9e2fdb49d5842de3e7e77b'/>
<id>urn:sha1:e1526395233aad15dc9e2fdb49d5842de3e7e77b</id>
<content type='text'>
Reported as "a key is stuck and is only sent when I send a new event". It was neither
stuck nor late - it was gone, and a later frame repainting those cells is what made it
look like it arrived eventually.

The loop is read, apply, render, write, and `uart.write` blocks while the transmit FIFO is
full. That wait is real backpressure and should stay: dropping half an escape sequence
leaves the host terminal in the wrong colour for the rest of the session. But NOTHING
drained the receive FIFO during it, and that FIFO is 128 bytes - 11 ms of wire at 115200.

Measured on the die, typing a burst in one host write and counting what the firmware's
loop actually took off the UART:

    burst    before    after   after + chunked input
      128       128      128       128
      200       197      200       200
      300       257      300       300
      600       478      600       600
     1200         -     1200      1200
     2400         -     2316      2400
     4096         -     3611      4096

Two windows had to close, and the second was only visible once the first was shut.

`input_rescue.pump` drains the receiver on every iteration of the wait for transmitter
room. That is the big one, and it is the whole reason this policy lives in its own file:
`uart.zig` cannot be tested without the chip because every line of it is an MMIO access,
while `pump` takes its port as `anytype` and runs against a fake with a two-byte transmit
FIFO and an eight-byte receive FIFO in `zig build test`. The fake models the receive FIFO
the way the hardware behaves - a byte arriving into a full FIFO is simply gone - so the
test fails by 67 lost bytes with the rescue removed, which is the die's 88-of-200 in
miniature. It also caught a flaw in its own first draft: a fake whose transmit FIFO drains
as fast as it fills never blocks, so `pump` never waits and the test proves nothing.

The second window was APPLYING the input. A keystroke costs 44 us on an empty line and
63 us at 640 characters, so handing the editor a full 128-byte batch is up to 8 ms in
which nothing drains the receiver - against 11 ms of FIFO. The loop now feeds the editor
eight bytes at a time and rescues between chunks. Splitting a burst at an arbitrary byte
is already safe, because `pardes_p4_input` keeps whatever it could not parse; that is how
it survives an escape sequence split across two UART reads. One render still happens per
loop iteration, so this costs no extra wire. Eight rather than thirty-two by measurement:
32 left 2400 and 4096 lossy, 8 does not.

Beyond 4096 bytes in one burst the editor genuinely cannot keep up, and the ring reports
what it abandoned instead of losing it silently - `rxdrop` in the PROF line, alongside a
running count of received bytes. That counter is the other lesson here: the first attempt
at measuring this counted characters on the reconstructed screen, which cannot distinguish
"never arrived" from "arrived but off the edge of the viewport", and it disagreed with the
hardware in both directions.

No cost to latency: round trip median 3687 us over 60 trials against 3687 before, maximum
3866, screen byte-identical to the vaxis reference, host tests green.
</content>
</entry>
<entry>
<title>Flush the caches at startup: the bootloader hands over stale lines</title>
<updated>2026-08-25T17:56:24Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T17:56:24Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=d69785542ced9bd24d210e38788e8eab42d678ad'/>
<id>urn:sha1:d69785542ced9bd24d210e38788e8eab42d678ad</id>
<content type='text'>
The image reads its own .rodata and gets its own .text back, and after ruling out
everything cheaper the answer is the cache.

What was eliminated first, each by measurement rather than argument:

  * The MMU table is CORRECT. Read from the running application through
    SPI_MEM_C_MMU_ITEM_INDEX_REG/CONTENT_REG (hal/esp32p4/mmu_ll.h:311-330),
    entries 0..9 hold 0x1001..0x100a - the valid bit plus physical page N+1 -
    which is exactly what the image builder's single flash-to-vaddr anchor
    requires, and entries 10..11 are unmapped as they should be.
  * The page size is not in question: hardwired to 64 KiB on this chip
    (mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it), which is
    what tools/image.zig already assumed.
  * The flash is correct. The flasher verifies an MD5 of what the ROM stored, and
    app.bin matches the ELF byte for byte at the addresses that misread.
  * Not a write failure and not nondeterminism: identical across three resets and
    two reflashes with the same MD5.
  * Not 64-byte cache-line granularity either: the wrong bytes come in a
    contiguous run of at least 192.

The measurement that settles it: a load at 0x40035a1c returned 93 85 85 0f, and
reading 512 KiB to force capacity eviction made the SAME load return 3c ee 08 40,
which is what the image holds there. So the second-stage bootloader hands over
with cache lines that do not match the mapping it finally installed. It is
perfectly deterministic - the bootloader does the same thing every boot, so it
leaves the same lines - which is precisely why it looked like anything other than
a cache for so long.

The ROM's own Cache_Invalidate_All (0x4fc00404, the same address in both
esp32p4.rom.ld and the eco5 table) would be the right instrument and is NOT used:
called from here it faults inside ROM code with its argument stranded in a2, so it
wants a precondition this image does not know. A capacity flush needs no such
knowledge, costs one pass over 512 KiB of already-mapped flash once at boot, and
is four times the 128 KiB the L2 measured at.

Effect: the firmware now gets through the editor's allocator round-trip and
pardes.allocators.init, which is two steps further than before.

Also here, and correct independently of any of the above: uart.write and
writeByte no longer spin forever on a stalled transmitter. hal/uart.zig:182-186
already made this point about update() - "on a board with no debugger an infinite
spin is indistinguishable from a crash" - and this port proved it by spending an
afternoon reading a stalled console as a hang in whatever code came next. The wait
is bounded and abandoned bytes are counted.

Still open: the console stops immediately after pardes.allocators.init. Bounded
writes did not change it, so it is not the transmit spin; there is no Guru
Meditation, so it is not a trap the ROM can report. The p4 allocator tier's
zero-capacity StackFallbackAllocators are the one unusual thing in that call and
their reasoning against lib/std/heap.zig is written down in src/allocators.zig,
but it has not been tested with a nonzero floor. The bisect markers are left in
place for that.
</content>
</entry>
<entry>
<title>pardes as P4 firmware: the seam, and a flash-mapping bug in this toolchain</title>
<updated>2026-08-25T17:38:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T16:06:05Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=174991b8f3f8e9c792eede7a52ad7beb10a08b05'/>
<id>urn:sha1:174991b8f3f8e9c792eede7a52ad7beb10a08b05</id>
<content type='text'>
The editor arrives as one freestanding OBJECT exporting a seven-function C ABI
(src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not
as a package dependency. A build.zig.zon path dependency was built first and
reverted: merely DECLARING it nested pardes's ~30-package graph under this one and
broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime
quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed
in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy.

The seam is bytes in and bytes out, which is what a serial line is anyway: the
editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the
clock, and neither names the other's types. It is versioned, because linkers do
not type-check C symbols and a drifted signature would link cleanly and then
corrupt the stack.

THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's
anchor makes two mapped segments share an MMU page safely - as long as rodata does
not END inside the page where text BEGINS. With a 578 KB image it does. A volatile
read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to
"lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared
page read as code, so the first thing the firmware tried to print was machine code
and it died on an instruction access fault. .flash.text is now ALIGN(0x10000),
making the segments page-disjoint. The packing trick this project opened with only
ever mattered when the alternative was 64 KiB of zeros in a 1 KB image.

Two more findings, both recorded in README.md:

  * A linker symbol declared as an anyopaque OBJECT gives the optimiser a
    zero-sized object, so ordinary stores through a pointer derived from its
    address are dead code it may drop - and did, silently. The allocator's first
    block header read back as size=2988759312 next=0x14284684 and the free-list
    walk never terminated. @extern with a many-pointer has no size to lose.
    examples/memprobe.zig could not have caught it: it writes through a volatile
    pointer, which the optimiser must leave alone.

  * The RTC watchdog is armed at handover. Every example here had been resetting on
    a ten-second cycle, invisibly, because no run had ever lasted eight seconds.

State, honestly: the firmware boots, clears .bss, brings up the console, disables
the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB
heap and calls into the editor, which sets up its sink and its environment. It then
faults inside pardes_p4_init on the first allocation. The cause is measured but not
fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000
higher - exactly the vaddr distance between the rodata and text segments - while
pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in
place, deliberately, because the next step needs them.


--- correction, measured after the above was written ---

Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig
was right for a reason I initially got wrong and then measured.

I first read bootloader_utility.c's `#else` branch, which classifies segments by
address window with two independent ifs - and since the P4's DROM and IROM windows
are the identical range (soc.h:146-149), I concluded the last mapped segment wins
both roles and the first is never mapped. That branch does not run on this chip.
The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851),
whose own comment says it: "On chips with shared D/I external vaddr, we don't
divide them into either D or I, as essentially they are the same." It collects
mapped segments POSITIONALLY into rom_addr[2] and ends with

    assert(rom_index == 2);

Shipping a one-segment image proved it, on the board:

    Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2)

So the split stays, image.zig keeps enforcing exactly two - turning that boot-time
abort into a build-time error - and both are now documented with the branch that
actually runs and the assert that actually fires.

What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may
share a 64 KiB MMU page only if they also share a flash page, which the image
builder's anchor guarantees - and that holds right up until an application is large
enough for rodata to END inside the page where text BEGINS. With a 578 KB image it
does. Measured on the die: a volatile read of a string literal at 0x4004a1d1
returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own
.flash.text. Every literal in that shared page read as code, so the first thing the
firmware tried to print was machine code, and it died on an instruction access
fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint.
It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing
trick this project opened with only mattered when the alternative was 64 KiB of
zeros in a 1 KB image.

With that fixed the firmware gets much further: entry, .bss cleared, console up,
watchdog disabled, systimer running, ABI version checked, the 384 KiB heap
initialised, into the editor, its sink and environment ready - and the literal at
0x4004a1d1 now reads back correctly.

Still open, and characterised rather than guessed: pardes_p4_init faults on its
first allocation. The allocator struct crosses the seam intact (its function
pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads
back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it.
Ruled out with measurements: the ELF and the image agree at that address, the flash
is MD5-verified against the image, the wrong bytes are identical across three
resets and two reflashes (so not a stale cache), the corruption is a contiguous run
rather than 64-byte lines, and mmu_hal_map_region's arithmetic
(page_num = ceil(len/page), entry from vaddr) is correct for the segments as now
laid out. The next measurement is the one that settles it: read the MMU entry
registers from the running application and print vaddr -&gt; flash for every page. The
register model in src/soc.zig can do that; the bisect markers are left in place for
it.
</content>
</entry>
</feed>
