diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 22:41:28 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 22:41:28 -0300 |
| commit | 110eeeb80181b990901c22d28fa4fd41dd10cdd7 (patch) | |
| tree | 4251af13a5fc283565bb379d93074c29c076ff18 /experiments/length-DirectEmit.csv | |
| parent | 6bb2023a419e56a50cd40fb407abd9dda250d088 (diff) | |
| download | esp32p4-110eeeb80181b990901c22d28fa4fd41dd10cdd7.tar.gz esp32p4-110eeeb80181b990901c22d28fa4fd41dd10cdd7.zip | |
Report: the pad target swept, and the byte-at-a-time read that was hiding in std
Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target,
and the closing section now reports it the way it should be reported: over 80
phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document
length rather than as one intercept.
Three things this adds that are findings rather than steps.
## The pad target, swept
Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the
cursor-positioning frame 21 -> 49 bytes made it a millisecond faster, padding the
cursor-hiding frame 6 -> 36 made it slower - so the target was swept as the only variable.
0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary
is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either:
it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what
confirms the descriptor is the thing to believe.
## The largest read in the firmware ran a byte at a time
The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame,
comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte,
about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32
at a time: 223 -> 66 us, 0.95 cycles per byte, at every document length.
Recorded alongside it are the two candidates that were measured and REVERTED, which is the
more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in
Surface.set removed 52 million instructions per host run and zero cycles on either host or
board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are
cheap; it was the reads that were slow.
## Where it stopped, honestly
The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty
line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12
viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and
4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are
now in the table with a "cells changed: none" row under them, because a round trip for an
edit that displays nothing is worth reporting as exactly that and not as a failure to hit
a number.
Also corrected: the ranked-recommendations table said the UART0 raise was abandoned
because a higher line rate moves settle and leaves the round trip alone. That reasoning
assumed the first byte reaches the host as soon as it is sent, and the bridge finding says
otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so
the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that
one 115200-baud line is the premise of this port rather than a free variable, and meeting
the target by changing the link would answer a different question.
Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones.
Diffstat (limited to 'experiments/length-DirectEmit.csv')
0 files changed, 0 insertions, 0 deletions
