diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 20:20:49 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 20:20:49 -0300 |
| commit | ce25f7950234f44dbb67af2cfde8da846d01aeeb (patch) | |
| tree | 97b06e759fd13680d98987602d4215e06e1f3f37 /examples/halcheck.zig | |
| parent | 341a9fe36840ec227284b6fac35e247d5d20ac18 (diff) | |
| download | esp32p4-ce25f7950234f44dbb67af2cfde8da846d01aeeb.tar.gz esp32p4-ce25f7950234f44dbb67af2cfde8da846d01aeeb.zip | |
Measure inside a frame, and cut a keystroke from 17.0 ms to 8.9 ms
The board half of the run to 4 ms: the instrumentation that found the cost, and the
measurements that judged each change.
`-Dprof` grew two things. It now renders a SECOND time with nothing changed, which
splits a frame's cost cleanly: whatever the second render still costs is the price of
walking and diffing the whole editor state, and the difference between the two is the
price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a
keystroke was work done regardless of what the keystroke did.
It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three
stages in CPU cycles. That is what turned "render is slow" into an address:
stage before after
copy Surface -> vaxis 6 750 us 1 450 us
vaxis diff + emit 2 460 us 2 455 us
push into the UART 1 us 1 us
(pardes's own Surface build) ~2 600 us ~2 000 us
The copy was 57% of a keystroke and it was in this repo's own `present`, not in
pardes and not in vaxis.
Measured on the die, five document lengths x seven trials per configuration:
configuration fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
## Where the remaining 4.9 ms is, and why the goal is not met
Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB
latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a
further 3.4x. The three remaining pieces are known and measured - our walk of the
grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole
Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the
Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke.
Raising the baud does NOT help this number, and that is worth writing down because it
is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost
none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms
and leaves the round trip where it is.
## A bug found on the way
`vx.resize` fails on this board. A runtime geometry change takes its allocation
failure path, restores the previous size and returns: 80 bytes go out where 1,392
should, and the screen keeps its old shape. Reproduced with the shadow grid compiled
out, so it predates it. The board has one geometry per session.
That is also why the shadow grid's correctness test compares two firmwares rather
than forcing a repaint with a resize - the forcing mechanism does not work here. The
reference path and the incremental path were each run against the same 19-step
workload and their reconstructed screens are byte-identical.
Diffstat (limited to 'examples/halcheck.zig')
0 files changed, 0 insertions, 0 deletions
