summaryrefslogtreecommitdiff
path: root/experiments/length-pad16.csv
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-25 20:20:49 -0300
committerGabriel Schneider <[email protected]>2026-08-25 20:20:49 -0300
commitce25f7950234f44dbb67af2cfde8da846d01aeeb (patch)
tree97b06e759fd13680d98987602d4215e06e1f3f37 /experiments/length-pad16.csv
parent341a9fe36840ec227284b6fac35e247d5d20ac18 (diff)
downloadesp32p4-ce25f7950234f44dbb67af2cfde8da846d01aeeb.tar.gz
esp32p4-ce25f7950234f44dbb67af2cfde8da846d01aeeb.zip
Measure inside a frame, and cut a keystroke from 17.0 ms to 8.9 ms
The board half of the run to 4 ms: the instrumentation that found the cost, and the measurements that judged each change. `-Dprof` grew two things. It now renders a SECOND time with nothing changed, which splits a frame's cost cleanly: whatever the second render still costs is the price of walking and diffing the whole editor state, and the difference between the two is the price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a keystroke was work done regardless of what the keystroke did. It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three stages in CPU cycles. That is what turned "render is slow" into an address: stage before after copy Surface -> vaxis 6 750 us 1 450 us vaxis diff + emit 2 460 us 2 455 us push into the UART 1 us 1 us (pardes's own Surface build) ~2 600 us ~2 000 us The copy was 57% of a keystroke and it was in this repo's own `present`, not in pardes and not in vaxis. Measured on the die, five document lengths x seven trials per configuration: configuration fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x ## Where the remaining 4.9 ms is, and why the goal is not met Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a further 3.4x. The three remaining pieces are known and measured - our walk of the grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke. Raising the baud does NOT help this number, and that is worth writing down because it is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms and leaves the round trip where it is. ## A bug found on the way `vx.resize` fails on this board. A runtime geometry change takes its allocation failure path, restores the previous size and returns: 80 bytes go out where 1,392 should, and the screen keeps its old shape. Reproduced with the shadow grid compiled out, so it predates it. The board has one geometry per session. That is also why the shadow grid's correctness test compares two firmwares rather than forcing a repaint with a resize - the forcing mechanism does not work here. The reference path and the incremental path were each run against the same 19-step workload and their reconstructed screens are byte-identical.
Diffstat (limited to 'experiments/length-pad16.csv')
0 files changed, 0 insertions, 0 deletions