<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/src, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-09-16T14:28:40Z</updated>
<entry>
<title>Make the toolchain a package another build can drive, and move the editor's glue to the editor</title>
<updated>2026-09-16T14:28:40Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T16:28:33Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=b42ecaed412be2e30b9e780eb7c9e46e1535f26f'/>
<id>urn:sha1:b42ecaed412be2e30b9e780eb7c9e46e1535f26f</id>
<content type='text'>
</content>
</entry>
<entry>
<title>Offer the editor the board's pads, and check one flips on the die</title>
<updated>2026-08-26T15:40:24Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T15:40:24Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=38bb891dd6bd0074894cbfedbf9185e303cc549e'/>
<id>urn:sha1:38bb891dd6bd0074894cbfedbf9185e303cc549e</id>
<content type='text'>
## The board half of Gpio

`pardes_p4_init` gained a `GpioFn` and the ABI version went to 2, which is what turns a mixed
pair of builds into a refusal to boot rather than five arguments read as six.

`gpioToggle` is four lines over `hal.gpio`: configure the pad as a readable output, read the level
it is driving, drive the other one, read it again. It is deliberately here and not in the editor.
A toggle is not a write to GPIO_OUT - `configureOutput` sets the IO MUX function, the GPIO matrix
route, the drive strength, the input buffer and the pulls, then the output enable, indexed by a
per-pin table - and that code is already in this repo, already the call `src/main.zig` blinks with,
and already checked against ESP-IDF's headers by `zig build diff`. The editor object gets a
function pointer instead of a second copy nobody tests.

`getDrivenLevel`, not `getLevel`: the answer is the level the board is driving, which is defined
for every pin. The pad's own level is what the outside world says, and on an unconnected header pin
that is noise. The input buffer is enabled anyway so `Peek` of GPIO_IN_REG can be compared to it.

## A fifth hardware check

`p4-bench --check` runs `Gpio 33` three times and requires 0-&gt;1, 1-&gt;0, 0-&gt;1.

The alternation is the oracle, not either answer. `0-&gt;1` alone is what a firmware printing a
hardcoded string would also say; two runs that disagree can only come from a level that was stored
and read again. Three, so the third rules out an ordering coincidence. GPIO33 because
`src/oracle/ledc_cases.zig` already documents it as a free pin on this board's JP1 header - pin 21
on the diagram the editor now draws. GPIO20 is the blink demo's pin and may have a wire on it.

This is the only check here that crosses the whole seam: editor word, C ABI, HAL, pad, and the
level back out through the message row. Nothing smaller exercises the ABI at all.

## -Dtheme-animation forwarded

Same path as the geometry, for the same reason: baked into the object, wanted from here.

All five checks pass on the die; host tests green; the board is flashed with md5 932898f34ca1dc31.
</content>
</entry>
<entry>
<title>A test suite that runs on the die, and two build steps for reading an image</title>
<updated>2026-08-26T12:56:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T12:56:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=55743cc5d564f6ef6d6f8a0eb1a614a74b021d3a'/>
<id>urn:sha1:55743cc5d564f6ef6d6f8a0eb1a614a74b021d3a</id>
<content type='text'>
## zig build selftest

Seventeen checks, on the board, chosen by one rule: a check belongs there only if the die
can answer it and a host cannot. Every serious bug this port has produced was invisible to
a host test. `std.mem.eql` compares a byte at a time on this target, so the firmware's
largest read was three times slower than it needed to be and nothing on a laptop could
tell. A lone ESC resolves to the Escape key, which is right when a kernel hands over a
whole sequence and wrong when a 115200 line hands over one byte every 87 us. A full
transmit FIFO stopped anything draining the receiver, and the FIFO depth is a hardware
number.

So: the volatile promise (two reads of a live counter are two reads, which is what `Peek`
rests on), word-wise equality checked against `std.mem.eql` itself at every difference
position and both alignments, the input rescue against a fake port with a FIFO that loses
what arrives into a full one, the allocator on real L2MEM, and the cycle counter against
the systimer - which is clocked from the crystal and therefore cannot be flattered by a
wrong CPU divider.

Anything that is pure logic stays in `zig build test`, which is faster and needs no
hardware. Duplicating those here would make the suite longer and no stronger.

Not a `zig test` binary, deliberately: Zig's runner wants an OS and `std.testing.allocator`
is a debug allocator over the page allocator, which on freestanding is either a compile
error or a lie. The harness is thirty lines.

`selftest` has its own application, image and flash chain so it is one command with no
flags to remember, and it makes the board's verdict the build's - a suite whose result a
human has to read out of a scrolling log is a suite that gets ignored on the first busy
afternoon. Proven both ways: 17/17 with exit 0, and exit 1 naming the check when one is
deliberately inverted.

The clock check earned its place immediately. The first version read `config.cpu_mhz` and
compared the die against it without ever performing the raise, so `-Dcpu-mhz=360` failed
with `khz=90001 want=360000`. The check was right and the expectation was wrong; it now
calls `setCpuFreq` itself, which makes it a test of the raise rather than a tautology.
90001 kHz at 90, 360004 at 360.

## zig build layout, and a loader error that explains itself

Both of these exist because of an hour I spent that they would have saved.

`NotTwoMappedSegments` said the count was wrong and nothing about what the segments were,
which is the only thing that says which section grew, shrank or stopped being emitted. I
diagnosed one by hand with readelf on an artifact that turned out to be a stale install,
then guessing at the linker script. The image step now prints the segment table with the
error - it has to happen there, because a rejected image is never written, so no later step
can show it - and `zig build layout` prints the same table for an ELF plus the loader's
verdict as TEXT, which `size` cannot do because it reads a finished image and the moment you
need the table is when there isn't one.

They immediately paid for themselves. The reason the suite would not build was that it had
no `_start`: `-fentry=_start` found no such symbol, `--gc-sections` discarded every
function as unreachable, and `.flash.text` was empty. `layout` prints `entry 0x0` for
exactly that, in one line. The second failure - a silent board - was a missing app
descriptor, which is the same class of thing and now has a comment where it happened.

## The panic handler already existed, and was printing past the end of its message

`msg` is a Zig slice and `%s` reads until a NUL, so handing `msg.ptr` to the ROM's printf
printed the message and then whatever followed it in memory. String literals get away with
it; std's own panics do not, because they are formatted into a buffer - "index out of
bounds: index 5, len 3" - and carry no terminator. It now goes out through `uart.write`,
which takes a length, with the fault address after it so addr2line can find the line.
</content>
</entry>
<entry>
<title>Ask the shell for its own grid, and report the heap on every boot</title>
<updated>2026-08-26T05:04:57Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T05:04:57Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=c64012e35c6370804f8588a62964e358631323b9'/>
<id>urn:sha1:c64012e35c6370804f8588a62964e358631323b9</id>
<content type='text'>
Two small changes that the geometry sweep needed.

The firmware asked for 80x24 and let the shell clamp it, which made this file a second
opinion about the board's geometry - one opinion too many, and wrong the moment the shell
could render more than that. It now asks for 255x255 so the shell's own ceiling is what
governs, and the shell reports what it settled on.

And the heap is printed on every boot rather than only when init fails. A geometry that
fits with 2 KB to spare and one that fits with 80 KB are not the same answer, and from the
host the difference was invisible. That line is what showed the sweep that memory had
stopped being the constraint at all: 336 KB free at every size tried, while '.bss' - the
shadow grid, sized at comptime - is what actually runs out.
</content>
</entry>
<entry>
<title>Drain the receiver while the transmitter is full: keystrokes were being lost</title>
<updated>2026-08-26T04:13:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T04:13:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=e1526395233aad15dc9e2fdb49d5842de3e7e77b'/>
<id>urn:sha1:e1526395233aad15dc9e2fdb49d5842de3e7e77b</id>
<content type='text'>
Reported as "a key is stuck and is only sent when I send a new event". It was neither
stuck nor late - it was gone, and a later frame repainting those cells is what made it
look like it arrived eventually.

The loop is read, apply, render, write, and `uart.write` blocks while the transmit FIFO is
full. That wait is real backpressure and should stay: dropping half an escape sequence
leaves the host terminal in the wrong colour for the rest of the session. But NOTHING
drained the receive FIFO during it, and that FIFO is 128 bytes - 11 ms of wire at 115200.

Measured on the die, typing a burst in one host write and counting what the firmware's
loop actually took off the UART:

    burst    before    after   after + chunked input
      128       128      128       128
      200       197      200       200
      300       257      300       300
      600       478      600       600
     1200         -     1200      1200
     2400         -     2316      2400
     4096         -     3611      4096

Two windows had to close, and the second was only visible once the first was shut.

`input_rescue.pump` drains the receiver on every iteration of the wait for transmitter
room. That is the big one, and it is the whole reason this policy lives in its own file:
`uart.zig` cannot be tested without the chip because every line of it is an MMIO access,
while `pump` takes its port as `anytype` and runs against a fake with a two-byte transmit
FIFO and an eight-byte receive FIFO in `zig build test`. The fake models the receive FIFO
the way the hardware behaves - a byte arriving into a full FIFO is simply gone - so the
test fails by 67 lost bytes with the rescue removed, which is the die's 88-of-200 in
miniature. It also caught a flaw in its own first draft: a fake whose transmit FIFO drains
as fast as it fills never blocks, so `pump` never waits and the test proves nothing.

The second window was APPLYING the input. A keystroke costs 44 us on an empty line and
63 us at 640 characters, so handing the editor a full 128-byte batch is up to 8 ms in
which nothing drains the receiver - against 11 ms of FIFO. The loop now feeds the editor
eight bytes at a time and rescues between chunks. Splitting a burst at an arbitrary byte
is already safe, because `pardes_p4_input` keeps whatever it could not parse; that is how
it survives an escape sequence split across two UART reads. One render still happens per
loop iteration, so this costs no extra wire. Eight rather than thirty-two by measurement:
32 left 2400 and 4096 lossy, 8 does not.

Beyond 4096 bytes in one burst the editor genuinely cannot keep up, and the ring reports
what it abandoned instead of losing it silently - `rxdrop` in the PROF line, alongside a
running count of received bytes. That counter is the other lesson here: the first attempt
at measuring this counted characters on the reconstructed screen, which cannot distinguish
"never arrived" from "arrived but off the edge of the viewport", and it disagreed with the
hardware in both directions.

No cost to latency: round trip median 3687 us over 60 trials against 3687 before, maximum
3866, screen byte-identical to the vaxis reference, host tests green.
</content>
</entry>
<entry>
<title>Run the CPU at 360 MHz: -Dcpu-mhz, and a keystroke lands at 4.37 ms</title>
<updated>2026-08-26T00:11:24Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:11:24Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=aa524eae414f4875794638df0586c011c6b0f2c5'/>
<id>urn:sha1:aa524eae414f4875794638df0586c011c6b0f2c5</id>
<content type='text'>
The board was executing at 90 MHz because the stock second-stage bootloader is built
with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured
here against the systimer, which is XTAL/2.5 and therefore an independent reference:
4,500,367 cycles in 50,004 us = exactly 90 MHz.

The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing
else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage
step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse).
`hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order -
APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own
comment says the other order passes through a state where APB or MEM violates its
timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every
`ets_delay_us` in the image is wrong by exactly the frequency ratio.

Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from
a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the
systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from
the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only
the divisor in experiments/ had to move.

## What it bought

    step                fixed    per char   at 160 chars
    ReleaseSmall      16.99 ms    54.3 us      25.56 ms
    ReleaseFast       14.85 ms    34.7 us      20.30 ms  0.79x
    + ASCII grapheme  14.56 ms    12.0 us      16.46 ms  0.64x
    + ASCII print     14.27 ms     6.9 us      15.36 ms  0.60x
    + shadow grid      8.87 ms     7.3 us      10.02 ms  0.39x
    + byte compare     8.37 ms     7.1 us       9.48 ms  0.37x
    + 360 MHz          4.37 ms     1.9 us       4.67 ms  0.18x

Compute went 6.42 -&gt; 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the
point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte,
so that stage is bounded by L2MEM bandwidth and does not care how fast the core is.
The prediction that this would happen was made before the measurement and held.

## The goal was 4 ms and this is 4.37

Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency
that no firmware change touches (measured independently against the protocol
responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's
Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large
enough to close the gap alone, and it is redundant work - `present` already computes
exactly which cells moved - so emitting ANSI straight from the shadow grid would do it.
I did not, because it is a from-scratch renderer and the honest verification for it
needs more than the harness currently proves.

## Verification, and a bug in my own instrument

Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the
A/B compares screens across clocks as well as across the shadow-grid flag. The first
attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw
SGR parameters applied to each cell, which is history-dependent, and a faster board
splits the same keystrokes across different frames. Decoding SGR into actual state -
resolved foreground, background and attribute set per cell - it is identical: same
characters and same style everywhere, both across clocks and across the flag.

Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked
like it would quantise every sample. It does not - poll(2) returns when data arrives,
not when the timeout expires - and switching to a non-blocking spin moved the measured
round trip by 0 us. The comment now says so, since the next reader will wonder too.

snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is
opt-in because every number in experiments/ up to this commit was taken at 90.
</content>
</entry>
<entry>
<title>Measure inside a frame, and cut a keystroke from 17.0 ms to 8.9 ms</title>
<updated>2026-08-25T23:20:49Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:20:49Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=ce25f7950234f44dbb67af2cfde8da846d01aeeb'/>
<id>urn:sha1:ce25f7950234f44dbb67af2cfde8da846d01aeeb</id>
<content type='text'>
The board half of the run to 4 ms: the instrumentation that found the cost, and the
measurements that judged each change.

`-Dprof` grew two things. It now renders a SECOND time with nothing changed, which
splits a frame's cost cleanly: whatever the second render still costs is the price of
walking and diffing the whole editor state, and the difference between the two is the
price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a
keystroke was work done regardless of what the keystroke did.

It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three
stages in CPU cycles. That is what turned "render is slow" into an address:

    stage                        before      after
    copy Surface -&gt; vaxis       6 750 us   1 450 us
    vaxis diff + emit           2 460 us   2 455 us
    push into the UART                 1 us       1 us
    (pardes's own Surface build) ~2 600 us  ~2 000 us

The copy was 57% of a keystroke and it was in this repo's own `present`, not in
pardes and not in vaxis.

Measured on the die, five document lengths x seven trials per configuration:

    configuration      fixed     per char    at 160 chars
    ReleaseSmall      16.99 ms    54.3 us       25.56 ms
    ReleaseFast       14.85 ms    34.7 us       20.30 ms   0.79x
    + ASCII grapheme  14.56 ms    12.0 us       16.46 ms   0.64x
    + ASCII print     14.27 ms     6.9 us       15.36 ms   0.60x
    + shadow grid      8.87 ms     7.3 us       10.02 ms   0.39x

## Where the remaining 4.9 ms is, and why the goal is not met

Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB
latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a
further 3.4x. The three remaining pieces are known and measured - our walk of the
grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole
Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the
Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke.

Raising the baud does NOT help this number, and that is worth writing down because it
is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost
none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms
and leaves the round trip where it is.

## A bug found on the way

`vx.resize` fails on this board. A runtime geometry change takes its allocation
failure path, restores the previous size and returns: 80 bytes go out where 1,392
should, and the screen keeps its old shape. Reproduced with the shadow grid compiled
out, so it predates it. The board has one geometry per session.

That is also why the shadow grid's correctness test compares two firmwares rather
than forcing a repaint with a resize - the forcing mechanism does not work here. The
reference path and the incremental path were each run against the same 19-step
workload and their reconstructed screens are byte-identical.
</content>
</entry>
<entry>
<title>Bound the edit path's document scans, and find out they were never the problem</title>
<updated>2026-08-25T22:16:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T21:49:57Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=1935944a0352e9d176f0718063327e686ece1948'/>
<id>urn:sha1:1935944a0352e9d176f0718063327e686ece1948</id>
<content type='text'>
The board half: the -Dprof attribution that overturned the conclusion, its data, and
the report correction.

`-Dprof` adds two cycle-counter reads around `pardes_p4_input` and
`pardes_p4_render` and prints both. Off by default: it puts a line on the wire per
frame, which is the very resource being measured, so it answers "where did the 15 ms
go" and not "how fast is it".

It answered. Input is flat at ~220 us regardless of document size - 1.5% of a
keystroke - and the entire ~15 ms floor plus every microsecond of the per-character
slope live in `render`. The edit-path fix that the source reading implied (committed
next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5
conditions x 7 trials, exactly 0% on this board.

`experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's
mechanism claim was wrong, says so, and carries the disproof beside it. The ranked
recommendations are reordered with the renderer at #1.

Also here: `--sweep position` in p4-bench, which holds the document fixed at one
320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28
bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on
this board - the signature of a walk from the start of a line.
</content>
</entry>
<entry>
<title>A measuring instrument, and what it says about where the latency goes</title>
<updated>2026-08-25T21:41:18Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T21:41:18Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=1cef9c2e4bd873ebe13f5df635899231bcc467d2'/>
<id>urn:sha1:1cef9c2e4bd873ebe13f5df635899231bcc467d2</id>
<content type='text'>
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.

## The instrument

`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.

The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.

`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.

`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.

## What it measured

The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.

Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.

And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.

One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.

## Three bugs found while building it

The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.

Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.

`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.

Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.

## The report

`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
</content>
</entry>
<entry>
<title>pardes runs on the board: install a trap vector, clamp the grid</title>
<updated>2026-08-25T18:15:03Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T18:15:03Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=894524a2e10296d6db4e82e0fe98b1802910ab77'/>
<id>urn:sha1:894524a2e10296d6db4e82e0fe98b1802910ab77</id>
<content type='text'>
It works. `zig build interact -Dpardes` flashes the editor, attaches a terminal,
and typing changes the screen.

TWO FIXES, and the first is the one that mattered.

**Install mtvec.** The console went silent immediately after the editor's
allocators.init and nothing could explain it: bounded writes did not change it, no
Guru Meditation was printed, and execution did not return from pardes_p4_init even
when that function was made to return immediately after the marker that DID print.
Setting mtvec to this image's own handler, in DIRECT mode, fixed it - and the
handler never fires, which is the tell. hal/intr.zig:302 names the mechanism: the
CLIC can fetch handlers from MTVT instead of trapping to mtvec, and the bootloader
leaves that vectored mode on with a table this image does not own. `systimer.init`
is enabled a few lines earlier, so its first tick dispatched through a vector table
belonging to nobody. Writing mtvec with the low two bits clear selects direct mode
and the interrupt has somewhere legitimate to go.

That single register write took the port from "faults before the editor starts" to
the whole of pardes_p4_init succeeding, including vaxis's capability handshake going
out over the wire:

    \e[?1049h \e[?1016$p \e[?2027$p \e[?2031$p \e[?2048h \e[6n \e[&gt;q \e[?u
    \e_Gi=1,a=q \e[c

alt screen, in-band resize, cursor report, kitty keyboard, kitty graphics, DA1 -
every one of them answered by the terminal emulator on the far end of the CH340,
which is the whole design.

**Clamp the grid, and make a failed resize atomic.** Pardes.init then returned
OutOfMemory with 9,128 bytes left of 393,216: every cell is paid for four times
(vaxis Screen + InternalScreen, pardes Surface + previous_cells). Measured: 40x12
initialises with room to spare, 80x24 does not. So max_cols/max_rows cap the
geometry and the host's larger terminal simply hosts a corner of itself.

The second half of that is subtler and cost a working editor. `Vaxis.resize` deinits
both screens BEFORE allocating the replacements (Vaxis.zig:194-206), so a failed
resize leaves vaxis with freed screens and renders nothing at all - and the host
bridge injects a size report on attach, so an unclamped 80x24 killed an editor that
had already drawn its interface. A failed resize now restores the previous geometry.

Measured end to end on the die: the first frame is ~1.5 KB of ANSI drawing the acme
tag bars ("New Newcol Joincol Find Grep Help Change" / "Save New Newtty Del
Filter"), and typing two characters produces a 116-byte incremental update that
draws the `^` modified-marker and moves the cursor to row 3 column 3. vaxis's damage
tracking is doing exactly what a 11.9 KB/s link needs.

Image is 598,528 B of the 1,536,000 B factory partition, 39%.

Also retired here: the I1..I10 and B1..B9 bring-up markers, the MMU table dump, the
cache before/after probe, the allocator and vtable pointer dumps, and the RX byte
probe. Each answered its question and each answer now lives in a comment beside the
code it explains. The trap handler stays - it is the diagnostic this port most
needed and did not have.
</content>
</entry>
</feed>
