From 110eeeb80181b990901c22d28fa4fd41dd10cdd7 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Tue, 25 Aug 2026 22:41:28 -0300 Subject: Report: the pad target swept, and the byte-at-a-time read that was hiding in std Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target, and the closing section now reports it the way it should be reported: over 80 phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document length rather than as one intercept. Three things this adds that are findings rather than steps. ## The pad target, swept Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the cursor-positioning frame 21 -> 49 bytes made it a millisecond faster, padding the cursor-hiding frame 6 -> 36 made it slower - so the target was swept as the only variable. 0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either: it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what confirms the descriptor is the thing to believe. ## The largest read in the firmware ran a byte at a time The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame, comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte, about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32 at a time: 223 -> 66 us, 0.95 cycles per byte, at every document length. Recorded alongside it are the two candidates that were measured and REVERTED, which is the more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in Surface.set removed 52 million instructions per host run and zero cycles on either host or board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are cheap; it was the reads that were slow. ## Where it stopped, honestly The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12 viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and 4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are now in the table with a "cells changed: none" row under them, because a round trip for an edit that displays nothing is worth reporting as exactly that and not as a failure to hit a number. Also corrected: the ranked-recommendations table said the UART0 raise was abandoned because a higher line rate moves settle and leaves the round trip alone. That reasoning assumed the first byte reaches the host as soon as it is sent, and the bridge finding says otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that one 115200-baud line is the premise of this port rather than a free variable, and meeting the target by changing the link would answer a different question. Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones. --- experiments/report.typ | 152 ++++++++++++++++++++++++++++++++++++++----------- 1 file changed, 120 insertions(+), 32 deletions(-) (limited to 'experiments/report.typ') diff --git a/experiments/report.typ b/experiments/report.typ index a15b162..1601cef 100644 --- a/experiments/report.typ +++ b/experiments/report.typ @@ -48,7 +48,8 @@ rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") + rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") + rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") + - rows("length-CPU360-final.csv") + rows("length-DirectEmit.csv") + rows("length-CPU360-final.csv") + rows("length-DirectEmit.csv") + + rows("length-WordCmp.csv") ) #let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv") @@ -107,20 +108,23 @@ implied, and it was only reachable by instrumenting the firmware; the fix that source reading suggested was written, measured, and found to be worth 20% on a 19 MB file and nothing at all on this board. Experiment 3. -Both were then cut. A keystroke is *3.74 ms, from 16.99*, and 4.06 ms at a -160-character line, from 25.56 — the per-character term is 2.0 µs, from 54.3. Three of -the six steps that did it came from *not asking Unicode about ASCII*; the largest single -one came from this repository's own `present` copying all 480 cells into vaxis every -frame whether or not any had changed; and the last two were a clock that was only ever a -divider away, and a renderer that stopped asking vaxis to recompute a diff `present` had -already done. Experiment 4. - -The 4 ms target is met. Two findings on the way there are worth more than the number. -A 4× clock bought 2.6× because the grid walk is bounded by memory bandwidth, not by the -core. And the direct renderer, with less computation and a quarter of the bytes, first -measured *slower* — because the CH340 bridging the board to the host forwards a bulk -packet only when it is full, so a reply too small to fill one waits about a millisecond -for a timer. The frame has a minimum size and it belongs to the transport. +Both were then cut. A keystroke is *3.65 ms, from 16.99*, and the per-character term is +0.93 µs, from 54.3. Three of the seven steps that did it came from *not asking Unicode +about ASCII*; the largest single one came from this repository's own `present` copying all +480 cells into vaxis every frame whether or not any had changed; and the last three were a +clock that was only ever a divider away, a renderer that stopped asking vaxis to recompute +a diff `present` had already done, and a row comparison that turned out to be running one +byte at a time. Experiment 4. + +The 4 ms target is met, over 80 trials, with the slowest of them at 3 912 µs. Three +findings on the way there are worth more than the number. A 4× clock bought 2.6×, because +the grid walk is bounded by memory rather than by the core. The direct renderer, with less +computation and a quarter of the bytes, first measured *slower* — the CH340 bridging the +board to the host forwards a bulk packet only when it is full, so a reply too small to +fill one waits about a millisecond for a timer; the frame has a minimum size and it +belongs to the transport. And the single largest remaining win was not an algorithm at +all: the firmware's biggest read ran at 3.2 cycles per byte because `std.mem.eql` +compares a byte at a time. = The instrument @@ -584,7 +588,7 @@ tree-sitter, which the board never runs. table.hline(stroke: 0.5pt), ..( ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", - "RFast-fastcmp", "CPU360-final", "DirectEmit") + "RFast-fastcmp", "CPU360-final", "DirectEmit", "WordCmp") ).map(b => { let xs = lengths.map(l => l * 1.0) let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) @@ -603,8 +607,9 @@ tree-sitter, which the board never runs. ), caption: [Each row is a separate firmware measured on the die, 5 document lengths × 7--9 trials. The per-character column and the fixed column move for different reasons - and are worth reading separately. The last two rows are the clock raise and the - renderer, and neither is a source optimisation in the sense the four above it are.], + and are worth reading separately. The last three rows are the clock raise, the renderer + and the row comparison, and none of them is a source optimisation in the sense the four + above them are.], ) *Not asking Unicode about ASCII* accounts for the per-character column. Three fast @@ -741,20 +746,98 @@ round trip a staircase in board time — where a real saving can present as a re Sleeping a uniform random 0–2 ms before each keystroke decorrelates them. That was not what was happening here, but it had to be ruled out before the CH340 could be believed. +== The pad target, swept + +Two anecdotes disagreed about whether a bigger frame arrives sooner: padding the +cursor-positioning frame from 21 to 49 bytes made it a millisecond faster, while padding +the cursor-hiding frame from 6 to 36 made it slower. So the pad target was swept as the +only variable — same firmware, same workload, same clock. + +#figure( + table( + columns: (auto, auto, auto, auto, 1fr), + align: (right, right, right, right, left), + stroke: none, + table.hline(), + table.header([target], [frame], [fixed], [worst], []), + table.hline(stroke: 0.5pt), + [0 B], [21 B], [4 743 µs], [4 906 µs], [never fills a packet], + [16 B], [21 B], [4 817 µs], [5 079 µs], [still never fills one], + [*32 B*], [*35 B*], [*3 799 µs*], [*4 335 µs*], [*exactly `wMaxPacketSize`*], + [48 B], [49 B], [3 811 µs], [4 334 µs], [past it, and no better], + [64 B], [70 B], [3 813 µs], [4 317 µs], [past it, and no better], + table.hline(), + ), + caption: [Crossing the packet boundary is worth about 950 µs; going past it buys + nothing at all. 32 is not a fitted constant — it is `wMaxPacketSize` of endpoint 0x82 + as the device reports it in its own descriptor, and the sweep is what confirms the + descriptor is the right thing to believe.], +) + +That also found a hole in the padding: it covered the branch that positions the cursor +and not the branch that hides it, so a frame which only hid the cursor was six bytes and +waited out the timer. Hiding an already-hidden cursor is as idempotent as positioning it +twice. + +== The largest read in the firmware was a byte at a time + +The last win was not a fast path or a clock: it was noticing that the shadow-grid diff — +two 13 KB streams, every frame, comfortably the biggest memory access the firmware makes — +ran at 3.2 cycles per byte, about four times what word-wide loads need. That is the shape +of a byte-at-a-time loop, and `std.mem.eql` was it. + +Comparing a `u32` at a time took the grid walk from *223 µs to 66 µs*, 0.95 cycles per +byte, off every keystroke at every document length. The alignment test has to be made at +runtime because `Cell` is all `u8` fields and so has alignment 1: whether a row starts on +a word boundary is a property of whoever allocated the Surface rather than of the type. +The answer is bit-for-bit identical, which is what matters — the diff still rests on byte +equality implying visual equality, so it can never claim two different cells are the same. + +Two other candidates were measured and *reverted*, which is the more useful half of the +result. Replacing `Surface.fill`'s per-cell writes with a row-at-a-time `@memset` and +giving `Surface.set` a one-byte store instead of a runtime-length `@memcpy` removed 52 +million instructions per host run — 3.9% — and changed the cycle count on the host by +nothing and the board by nothing. Ablating the whole-surface fill priced it at 41 µs, or +1.18 cycles per byte written. Writes on this part are cheap and already pipelined; it was +the READS that were slow, and only the read was worth fixing. + == Where it stopped -A keystroke is *3.74 ms*, from 16.99 — and 4.06 ms at a 160-character line, from 25.56. -The target was 4 ms and it is met, with the phase-randomised instrument agreeing over 60 -trials: median 3 829 µs, minimum 3 722, ninetieth percentile 3 930. The reply is 35 -bytes, down from 81. +A keystroke is *3.65 ms*, from 16.99. The target was 4 ms and it is met with margin: over +80 phase-randomised trials the median is 3 687 µs, the minimum 3 571 and *the maximum +3 912* — every trial under 4 ms. Across document length it holds wherever the editor +actually shows the keystroke: + +#figure( + table( + columns: (auto, auto, auto, auto, auto, auto, auto, auto), + align: (left, right, right, right, right, right, right, right), + stroke: none, + table.hline(), + table.header([chars], [0], [20], [40], [80], [160], [320], [640]), + table.hline(stroke: 0.5pt), + [round trip], [3 602], [3 624], [3 638], [3 790], [3 868], [4 026], [4 192], + [cells changed], [some], [some], [some], [some], [some], [*none*], [*none*], + table.hline(), + ), + caption: [Microseconds, median of nine trials each. At 320 characters and beyond the + line has outgrown a 40×12 viewport, the cursor is off screen and the keystroke changes + no cell at all: the frame is 36 bytes of cursor-hide and padding. Those two columns are + the cost of an edit that displays *nothing*.], +) + +What remains is not something this repository owns. Of the round trip, about 3.0 ms is +host, USB and wire — the floor measured independently against the protocol responder, and +now partly explained by the bridge's packet granularity — and 0.55 ms is board at an empty +line, of which pardes rebuilding all 480 cells of the Surface is 436 µs, the grid walk +66 µs, and parsing and applying the edit 44 µs. -What remains is no longer dominated by anything this repository owns. Of the round trip, -roughly 2.9 ms is host, USB and wire — the floor measured independently against the -protocol responder, and now partly explained by the bridge's packet granularity — and -about 0.84 ms is board: 552 µs of pardes rebuilding every cell of the Surface every -frame, 223 µs walking the grid, 60 µs parsing and editing. The Surface rebuild is the -one architectural item left, and it is worth 552 µs against a 2.9 ms floor, so the next -order of magnitude is not in the firmware at all. +The Surface rebuild is the one architectural item left, and the two columns above say +exactly why it is architectural rather than a fast path: at 640 characters it costs 985 µs +to produce a frame in which nothing changed. Every micro-optimisation attempted against it +here — the grapheme walk, the per-cell writes, the fill — returned between nothing and +21 µs, because the cost is not in any one of those but in doing the whole frame again. +Nothing in this repository can avoid work pardes has already done. #figure( table( @@ -806,9 +889,14 @@ order of magnitude is not in the firmware at all. are derived from measured quantities and cited source. The last row is listed unranked because it is already applied and, on *this* target, buys nothing -- which is precisely why it is worth recording. Kept as the prediction it was: - Experiment 4 has since applied rows 1 and 2 and *abandoned row 3*, because raising - the line rate moves `settle` and leaves the round trip alone --- and the bridge - finding above is the reason row 3 looked attractive in the first place.], + Experiment 4 has since applied rows 1 and 2 and left row 3 unapplied. Row 3 deserves a + correction rather than a dismissal: it was abandoned because a higher line rate moves + `settle` and leaves the round trip alone, and that reasoning assumed the first byte + reaches the host as soon as it is sent. The bridge finding says otherwise --- delivery + waits for 32 bytes, which is 278 µs of wire at 115200 and would be 35 µs at 921600, so + the raise is worth roughly 240 µs of the round trip after all. It stays unapplied + because one 115200-baud line is the premise of this port, not a free variable, and + meeting the target by changing the link would be answering a different question.], ) = Threats to validity -- cgit v1.3