<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/experiments/length-RFast-fastcmp.csv, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-08-25T23:30:52Z</updated>
<entry>
<title>A keystroke is now 8.37 ms, from 16.99; the remaining 6.4 ms is located</title>
<updated>2026-08-25T23:30:52Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:30:52Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=05610a2a42b3d026473aa84ddfef1592feda5f17'/>
<id>urn:sha1:05610a2a42b3d026473aa84ddfef1592feda5f17</id>
<content type='text'>
Progression on the die, five document lengths x seven trials at each step:

    step                fixed    per char   at 160 chars
    ReleaseSmall      16.99 ms    54.3 us      25.56 ms
    ReleaseFast       14.85 ms    34.7 us      20.30 ms  0.79x
    + ASCII grapheme  14.56 ms    12.0 us      16.46 ms  0.64x
    + ASCII print     14.27 ms     6.9 us      15.36 ms  0.60x
    + shadow grid      8.87 ms     7.3 us      10.02 ms  0.39x
    + byte compare     8.37 ms     7.1 us       9.48 ms  0.37x

Round trip is time to the FIRST response byte: ~1.95 ms of host and USB latency plus
compute. Compute is 6.42 ms and the 4 ms goal needs it under 2.05 ms, so 3.1x remains.

Every microsecond of it is now measured rather than guessed, by stage, on the die:

    our walk of the grid        979 us   this repo's own present()
    vaxis diff + emit          2455 us   vaxis walks all 480 cells, per-cell strings
    pardes Surface rebuild     2000 us   pardes rebuilds every cell every frame
    input parse + edit          250 us

The honest reading is that the two large items are not each other's alternative.
vaxis's diff is redundant work - present() already computes exactly which cells moved,
so emitting ANSI directly from the shadow grid would remove most of that 2.46 ms. But
even a perfect renderer leaves pardes rebuilding a whole Surface per keystroke, and 4 ms
at 40x12 needs that too.

Raising the baud does not move this number. At 115200 an 81-byte reply is 7.0 ms of
wire, but almost none of it lands before the first byte; 921600 takes `settle` from
23 ms to ~16 ms and leaves the round trip where it is. Worth doing, not for this.
</content>
</entry>
</feed>
