<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/experiments, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-08-26T05:06:49Z</updated>
<entry>
<title>Report: Experiment 5, the three bugs, and a resize that fixed itself</title>
<updated>2026-08-26T05:06:49Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T05:06:49Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=ddee44e2a70e8b9be7b2778390059de6279dd454'/>
<id>urn:sha1:ddee44e2a70e8b9be7b2778390059de6279dd454</id>
<content type='text'>
Adds the geometry curve - 0.87 us per cell, measured across nine grids from 40x12 to
140x42 - and the two ways it ends: 200x60 refused by the linker because '.bss' cannot hold
the shadow grid, 160x48 linking and then trapping. Neither is a heap problem any more,
which is the finding: the heap has 336 KB spare while '.bss' runs out, because vaxis's two
unread grids stopped being allocated.

Also the three bugs the latency work walked straight past, none of which any number in this
document could see: keystrokes lost while the transmitter was full, every escape sequence
shredded because a lone ESC arrives alone on a 115200 line, and a mouse that was never
enabled. The middle one had been there since the port began and explains the other two -
a click arriving as ten key presses, with the '0' among them moving the cursor to column
zero, which is what made it look like a coordinate bug.

And Table 15 changes from 'not fixed' to 'fixed, and not on purpose'. The vx.resize failure
that three experiments characterised and left alone was an allocation of vaxis's two grids;
the emitter stopped needing them, so the path that failed is gone rather than repaired.
Verified on the die - an in-band report re-lays the screen from 56 columns to 30 and back.
Recorded as a consequence rather than as a fix, because nobody set out to repair it.

Eleven pages.
</content>
</entry>
<entry>
<title>Report: the pad target swept, and the byte-at-a-time read that was hiding in std</title>
<updated>2026-08-26T01:41:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T01:41:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=110eeeb80181b990901c22d28fa4fd41dd10cdd7'/>
<id>urn:sha1:110eeeb80181b990901c22d28fa4fd41dd10cdd7</id>
<content type='text'>
Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target,
and the closing section now reports it the way it should be reported: over 80
phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document
length rather than as one intercept.

Three things this adds that are findings rather than steps.

## The pad target, swept

Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the
cursor-positioning frame 21 -&gt; 49 bytes made it a millisecond faster, padding the
cursor-hiding frame 6 -&gt; 36 made it slower - so the target was swept as the only variable.
0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary
is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either:
it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what
confirms the descriptor is the thing to believe.

## The largest read in the firmware ran a byte at a time

The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame,
comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte,
about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32
at a time: 223 -&gt; 66 us, 0.95 cycles per byte, at every document length.

Recorded alongside it are the two candidates that were measured and REVERTED, which is the
more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in
Surface.set removed 52 million instructions per host run and zero cycles on either host or
board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are
cheap; it was the reads that were slow.

## Where it stopped, honestly

The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty
line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12
viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and
4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are
now in the table with a "cells changed: none" row under them, because a round trip for an
edit that displays nothing is worth reporting as exactly that and not as a failure to hit
a number.

Also corrected: the ranked-recommendations table said the UART0 raise was abandoned
because a higher line rate moves settle and leaves the round trip alone. That reasoning
assumed the first byte reaches the host as soon as it is sent, and the bridge finding says
otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so
the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that
one 115200-baud line is the premise of this port rather than a free variable, and meeting
the target by changing the link would answer a different question.

Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones.
</content>
</entry>
<entry>
<title>Report: the clock, the renderer, and a minimum frame owned by the USB bridge</title>
<updated>2026-08-26T00:40:14Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:39:03Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=6bb2023a419e56a50cd40fb407abd9dda250d088'/>
<id>urn:sha1:6bb2023a419e56a50cd40fb407abd9dda250d088</id>
<content type='text'>
Brings the document up to the end of Experiment 4. Two rows on the progression table -
360 MHz and direct emission - plus the sections behind them, the corrected verifier, and
a summary that no longer says the target was missed. It is met: 3.74 ms from 16.99.

The two findings worth more than the number, and both are written up as findings rather
than as steps:

A 4x clock bought 2.6x. The grid walk reads 27 KB a frame at about six cycles a byte, so
it is bounded by L2MEM bandwidth and does not care how fast the core runs. Predicted
before the measurement.

The direct renderer measured SLOWER at first - less computation, a quarter of the bytes,
worse round trip - because the CH340 forwards a bulk IN packet only when the packet is
full, and a 21-byte frame does not fill 32. It waits about a millisecond for a timer.
So the frame has a minimum size and it belongs to the transport, not the terminal; the
emitter pads to it with repeated cursor positioning. The table of 21/49/81-byte frames is
in the report because the MINIMUM column is the tell: the small frame's floor was already
530 us below vaxis's, exactly the compute saved, and only the median was hostage.

Also corrected: the verifier section. It described hashing SGR parameters per cell, which
is a history rather than a state, and that version reported a difference on the die that
did not exist - a faster board split the same keystrokes across different frames and
reached the same colours by another route. The document now says what it does instead,
and why the from-scratch renderer could not have been verified without the fix.

Table 13, the ranked recommendations from Experiment 3, is kept as the prediction it was
with a note on what has since been applied and that row 3 - raise the line rate - was
abandoned. The threats section no longer states the link floor as a single number, since
the bridge's packet granularity is now part of it and is specific to this bridge.

Nine pages. Every figure still reads the raw per-trial CSVs, including the two new ones.
</content>
</entry>
<entry>
<title>Run the CPU at 360 MHz: -Dcpu-mhz, and a keystroke lands at 4.37 ms</title>
<updated>2026-08-26T00:11:24Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:11:24Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=aa524eae414f4875794638df0586c011c6b0f2c5'/>
<id>urn:sha1:aa524eae414f4875794638df0586c011c6b0f2c5</id>
<content type='text'>
The board was executing at 90 MHz because the stock second-stage bootloader is built
with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured
here against the systimer, which is XTAL/2.5 and therefore an independent reference:
4,500,367 cycles in 50,004 us = exactly 90 MHz.

The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing
else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage
step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse).
`hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order -
APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own
comment says the other order passes through a state where APB or MEM violates its
timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every
`ets_delay_us` in the image is wrong by exactly the frequency ratio.

Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from
a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the
systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from
the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only
the divisor in experiments/ had to move.

## What it bought

    step                fixed    per char   at 160 chars
    ReleaseSmall      16.99 ms    54.3 us      25.56 ms
    ReleaseFast       14.85 ms    34.7 us      20.30 ms  0.79x
    + ASCII grapheme  14.56 ms    12.0 us      16.46 ms  0.64x
    + ASCII print     14.27 ms     6.9 us      15.36 ms  0.60x
    + shadow grid      8.87 ms     7.3 us      10.02 ms  0.39x
    + byte compare     8.37 ms     7.1 us       9.48 ms  0.37x
    + 360 MHz          4.37 ms     1.9 us       4.67 ms  0.18x

Compute went 6.42 -&gt; 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the
point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte,
so that stage is bounded by L2MEM bandwidth and does not care how fast the core is.
The prediction that this would happen was made before the measurement and held.

## The goal was 4 ms and this is 4.37

Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency
that no firmware change touches (measured independently against the protocol
responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's
Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large
enough to close the gap alone, and it is redundant work - `present` already computes
exactly which cells moved - so emitting ANSI straight from the shadow grid would do it.
I did not, because it is a from-scratch renderer and the honest verification for it
needs more than the harness currently proves.

## Verification, and a bug in my own instrument

Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the
A/B compares screens across clocks as well as across the shadow-grid flag. The first
attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw
SGR parameters applied to each cell, which is history-dependent, and a faster board
splits the same keystrokes across different frames. Decoding SGR into actual state -
resolved foreground, background and attribute set per cell - it is identical: same
characters and same style everywhere, both across clocks and across the flag.

Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked
like it would quantise every sample. It does not - poll(2) returns when data arrives,
not when the timeout expires - and switching to a non-blocking spin moved the measured
round trip by 0 us. The comment now says so, since the next reader will wonder too.

snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is
opt-in because every number in experiments/ up to this commit was taken at 90.
</content>
</entry>
<entry>
<title>Report: the optimisation campaign, and the 3.1x that is left</title>
<updated>2026-08-25T23:33:51Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:33:51Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=7664c57d93cd3f3bdb719f7ca915d657668ddfd9'/>
<id>urn:sha1:7664c57d93cd3f3bdb719f7ca915d657668ddfd9</id>
<content type='text'>
Experiment 4 records what each cut was worth, measured on the die at every step, and
the stage table that made the cuts findable at all. Also two corrections to the
document itself:

The data plumbing loaded four of the eight datasets, so the progression table would
have been computed from a subset. It now loads all eight.

A continuation line beginning with `+` is list markup to Typst, so the CSV file names
were rendering as a numbered item at the top of page one. One expression, one line.

The summary no longer ends on the ReleaseFast flag as the best available change; it
ends on the measured 16.99 -&gt; 8.37 ms and on the fact that the target is not met, with
the remaining 6.42 ms of compute broken into the three items it actually consists of.
Saying 'not met' in the summary matters more than the table: the number that was
missed is the one a reader should see first.
</content>
</entry>
<entry>
<title>A keystroke is now 8.37 ms, from 16.99; the remaining 6.4 ms is located</title>
<updated>2026-08-25T23:30:52Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:30:52Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=05610a2a42b3d026473aa84ddfef1592feda5f17'/>
<id>urn:sha1:05610a2a42b3d026473aa84ddfef1592feda5f17</id>
<content type='text'>
Progression on the die, five document lengths x seven trials at each step:

    step                fixed    per char   at 160 chars
    ReleaseSmall      16.99 ms    54.3 us      25.56 ms
    ReleaseFast       14.85 ms    34.7 us      20.30 ms  0.79x
    + ASCII grapheme  14.56 ms    12.0 us      16.46 ms  0.64x
    + ASCII print     14.27 ms     6.9 us      15.36 ms  0.60x
    + shadow grid      8.87 ms     7.3 us      10.02 ms  0.39x
    + byte compare     8.37 ms     7.1 us       9.48 ms  0.37x

Round trip is time to the FIRST response byte: ~1.95 ms of host and USB latency plus
compute. Compute is 6.42 ms and the 4 ms goal needs it under 2.05 ms, so 3.1x remains.

Every microsecond of it is now measured rather than guessed, by stage, on the die:

    our walk of the grid        979 us   this repo's own present()
    vaxis diff + emit          2455 us   vaxis walks all 480 cells, per-cell strings
    pardes Surface rebuild     2000 us   pardes rebuilds every cell every frame
    input parse + edit          250 us

The honest reading is that the two large items are not each other's alternative.
vaxis's diff is redundant work - present() already computes exactly which cells moved,
so emitting ANSI directly from the shadow grid would remove most of that 2.46 ms. But
even a perfect renderer leaves pardes rebuilding a whole Surface per keystroke, and 4 ms
at 40x12 needs that too.

Raising the baud does not move this number. At 115200 an 81-byte reply is 7.0 ms of
wire, but almost none of it lands before the first byte; 921600 takes `settle` from
23 ms to ~16 ms and leaves the round trip where it is. Worth doing, not for this.
</content>
</entry>
<entry>
<title>Measure inside a frame, and cut a keystroke from 17.0 ms to 8.9 ms</title>
<updated>2026-08-25T23:20:49Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:20:49Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=ce25f7950234f44dbb67af2cfde8da846d01aeeb'/>
<id>urn:sha1:ce25f7950234f44dbb67af2cfde8da846d01aeeb</id>
<content type='text'>
The board half of the run to 4 ms: the instrumentation that found the cost, and the
measurements that judged each change.

`-Dprof` grew two things. It now renders a SECOND time with nothing changed, which
splits a frame's cost cleanly: whatever the second render still costs is the price of
walking and diffing the whole editor state, and the difference between the two is the
price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a
keystroke was work done regardless of what the keystroke did.

It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three
stages in CPU cycles. That is what turned "render is slow" into an address:

    stage                        before      after
    copy Surface -&gt; vaxis       6 750 us   1 450 us
    vaxis diff + emit           2 460 us   2 455 us
    push into the UART                 1 us       1 us
    (pardes's own Surface build) ~2 600 us  ~2 000 us

The copy was 57% of a keystroke and it was in this repo's own `present`, not in
pardes and not in vaxis.

Measured on the die, five document lengths x seven trials per configuration:

    configuration      fixed     per char    at 160 chars
    ReleaseSmall      16.99 ms    54.3 us       25.56 ms
    ReleaseFast       14.85 ms    34.7 us       20.30 ms   0.79x
    + ASCII grapheme  14.56 ms    12.0 us       16.46 ms   0.64x
    + ASCII print     14.27 ms     6.9 us       15.36 ms   0.60x
    + shadow grid      8.87 ms     7.3 us       10.02 ms   0.39x

## Where the remaining 4.9 ms is, and why the goal is not met

Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB
latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a
further 3.4x. The three remaining pieces are known and measured - our walk of the
grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole
Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the
Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke.

Raising the baud does NOT help this number, and that is worth writing down because it
is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost
none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms
and leaves the round trip where it is.

## A bug found on the way

`vx.resize` fails on this board. A runtime geometry change takes its allocation
failure path, restores the previous size and returns: 80 bytes go out where 1,392
should, and the screen keeps its old shape. Reproduced with the shadow grid compiled
out, so it predates it. The board has one geometry per session.

That is also why the shadow grid's correctness test compares two firmwares rather
than forcing a repaint with a resize - the forcing mechanism does not work here. The
reference path and the incremental path were each run against the same 19-step
workload and their reconstructed screens are byte-identical.
</content>
</entry>
<entry>
<title>Measure all four configurations; the code fix is worth nothing here, the flag is worth 21%</title>
<updated>2026-08-25T22:29:22Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T22:29:22Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=341a9fe36840ec227284b6fac35e247d5d20ac18'/>
<id>urn:sha1:341a9fe36840ec227284b6fac35e247d5d20ac18</id>
<content type='text'>
The obvious question after the last commit was what it did on the actual target. The
answer is nothing, and the useful part is that the same run says what DOES work.

Four builds on the die, 5 document lengths x 7 trials each:

    configuration            fixed    per char   at 160    vs base
    ReleaseSmall          16.99 ms     54.3 us  25.56 ms     1.00x
    ReleaseSmall+lineSpan 17.10 ms     54.0 us  25.63 ms     1.00x
    ReleaseFast           14.85 ms     34.7 us  20.30 ms     0.79x
    ReleaseFast+lineSpan  14.88 ms     34.2 us  20.25 ms     0.79x

The edit-path change is invisible in BOTH modes, and ReleaseFast+lineSpan is
indistinguishable from ReleaseFast alone. The optimisation mode is the whole of the
difference, which is what Experiment 3 predicted: the edit path is 220 us of a 17 ms
keystroke, so making it cheaper cannot move the total, while the mode makes the
RENDERER faster and the renderer is where the time is.

`experiments/report.typ` gains that table and the reason the four-way comparison
exists rather than four indistinguishable lines on figure 1.

The board is now flashed with the default build, which as of the pin next door means
ReleaseFast: 809,552 B of the 1,536,000 B partition, 53%.
</content>
</entry>
<entry>
<title>Bound the edit path's document scans, and find out they were never the problem</title>
<updated>2026-08-25T22:16:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T21:49:57Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=1935944a0352e9d176f0718063327e686ece1948'/>
<id>urn:sha1:1935944a0352e9d176f0718063327e686ece1948</id>
<content type='text'>
The board half: the -Dprof attribution that overturned the conclusion, its data, and
the report correction.

`-Dprof` adds two cycle-counter reads around `pardes_p4_input` and
`pardes_p4_render` and prints both. Off by default: it puts a line on the wire per
frame, which is the very resource being measured, so it answers "where did the 15 ms
go" and not "how fast is it".

It answered. Input is flat at ~220 us regardless of document size - 1.5% of a
keystroke - and the entire ~15 ms floor plus every microsecond of the per-character
slope live in `render`. The edit-path fix that the source reading implied (committed
next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5
conditions x 7 trials, exactly 0% on this board.

`experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's
mechanism claim was wrong, says so, and carries the disproof beside it. The ranked
recommendations are reordered with the renderer at #1.

Also here: `--sweep position` in p4-bench, which holds the document fixed at one
320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28
bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on
this board - the signature of a walk from the start of a line.
</content>
</entry>
<entry>
<title>A measuring instrument, and what it says about where the latency goes</title>
<updated>2026-08-25T21:41:18Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T21:41:18Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=1cef9c2e4bd873ebe13f5df635899231bcc467d2'/>
<id>urn:sha1:1cef9c2e4bd873ebe13f5df635899231bcc467d2</id>
<content type='text'>
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.

## The instrument

`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.

The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.

`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.

`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.

## What it measured

The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.

Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.

And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.

One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.

## Three bugs found while building it

The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.

Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.

`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.

Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.

## The report

`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
</content>
</entry>
</feed>
