diff options
Diffstat (limited to 'experiments')
| -rw-r--r-- | experiments/length-DirectEmit.csv | 64 | ||||
| -rw-r--r-- | experiments/pg-1.png | bin | 264089 -> 313785 bytes | |||
| -rw-r--r-- | experiments/pg-2.png | bin | 181594 -> 220244 bytes | |||
| -rw-r--r-- | experiments/pg-3.png | bin | 191333 -> 218770 bytes | |||
| -rw-r--r-- | experiments/pg-4.png | bin | 245170 -> 279719 bytes | |||
| -rw-r--r-- | experiments/pg-5.png | bin | 277654 -> 319282 bytes | |||
| -rw-r--r-- | experiments/pg-6.png | bin | 270483 -> 308166 bytes | |||
| -rw-r--r-- | experiments/pg-7.png | bin | 248684 -> 348546 bytes | |||
| -rw-r--r-- | experiments/pg-8.png | bin | 22182 -> 290076 bytes | |||
| -rw-r--r-- | experiments/pg-9.png | bin | 0 -> 66553 bytes | |||
| -rw-r--r-- | experiments/report.typ | 166 |
11 files changed, 194 insertions, 36 deletions
diff --git a/experiments/length-DirectEmit.csv b/experiments/length-DirectEmit.csv new file mode 100644 index 0000000..f4fbd29 --- /dev/null +++ b/experiments/length-DirectEmit.csv @@ -0,0 +1,64 @@ +label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes +DirectEmit,length,0,0,0,insert,0,3987,12964,113 +DirectEmit,length,0,0,0,insert,1,3799,5930,34 +DirectEmit,length,0,0,0,insert,2,3779,5830,35 +DirectEmit,length,0,0,0,insert,3,3746,5886,35 +DirectEmit,length,0,0,0,insert,4,3756,5981,35 +DirectEmit,length,0,0,0,insert,5,3727,6013,35 +DirectEmit,length,0,0,0,insert,6,3738,5915,35 +DirectEmit,length,0,0,0,insert,7,3688,5865,35 +DirectEmit,length,0,0,0,insert,8,3665,5810,35 +DirectEmit,length,0,0,20,insert,0,3800,5952,35 +DirectEmit,length,0,0,20,insert,1,3763,5980,35 +DirectEmit,length,0,0,20,insert,2,3771,6016,35 +DirectEmit,length,0,0,20,insert,3,3815,14251,130 +DirectEmit,length,0,0,20,insert,4,3718,5861,34 +DirectEmit,length,0,0,20,insert,5,3742,6050,35 +DirectEmit,length,0,0,20,insert,6,3758,6007,35 +DirectEmit,length,0,0,20,insert,7,3758,6026,35 +DirectEmit,length,0,0,20,insert,8,3782,6024,35 +DirectEmit,length,0,0,40,insert,0,3802,6077,35 +DirectEmit,length,0,0,40,insert,1,3793,6013,35 +DirectEmit,length,0,0,40,insert,2,3855,6037,35 +DirectEmit,length,0,0,40,insert,3,3831,5958,35 +DirectEmit,length,0,0,40,insert,4,3816,5968,35 +DirectEmit,length,0,0,40,insert,5,3814,5974,35 +DirectEmit,length,0,0,40,insert,6,3924,14310,130 +DirectEmit,length,0,0,40,insert,7,3856,5972,34 +DirectEmit,length,0,0,40,insert,8,3854,6004,35 +DirectEmit,length,0,0,80,insert,0,3970,6116,35 +DirectEmit,length,0,0,80,insert,1,3883,6117,35 +DirectEmit,length,0,0,80,insert,2,3923,6127,35 +DirectEmit,length,0,0,80,insert,3,3928,6089,35 +DirectEmit,length,0,0,80,insert,4,3826,6113,35 +DirectEmit,length,0,0,80,insert,5,3825,6048,35 +DirectEmit,length,0,0,80,insert,6,3867,6074,35 +DirectEmit,length,0,0,80,insert,7,3860,5989,35 +DirectEmit,length,0,0,80,insert,8,3872,6023,35 +DirectEmit,length,0,0,160,insert,0,4063,6130,35 +DirectEmit,length,0,0,160,insert,1,4096,6196,35 +DirectEmit,length,0,0,160,insert,2,4043,6236,35 +DirectEmit,length,0,0,160,insert,3,4081,6154,35 +DirectEmit,length,0,0,160,insert,4,4062,6090,35 +DirectEmit,length,0,0,160,insert,5,4066,6154,35 +DirectEmit,length,0,0,160,insert,6,4029,6203,35 +DirectEmit,length,0,0,160,insert,7,3995,6181,35 +DirectEmit,length,0,0,160,insert,8,4032,6182,35 +DirectEmit,length,0,0,320,insert,0,3921,3921,6 +DirectEmit,length,0,0,320,insert,1,3863,3863,6 +DirectEmit,length,0,0,320,insert,2,3822,3823,6 +DirectEmit,length,0,0,320,insert,3,3949,3949,6 +DirectEmit,length,0,0,320,insert,4,3853,3853,6 +DirectEmit,length,0,0,320,insert,5,3909,3909,6 +DirectEmit,length,0,0,320,insert,6,3796,3797,6 +DirectEmit,length,0,0,320,insert,7,3863,3864,6 +DirectEmit,length,0,0,320,insert,8,3953,3954,6 +DirectEmit,length,0,0,640,insert,0,4169,4169,6 +DirectEmit,length,0,0,640,insert,1,4021,4021,6 +DirectEmit,length,0,0,640,insert,2,4070,4070,6 +DirectEmit,length,0,0,640,insert,3,3996,3996,6 +DirectEmit,length,0,0,640,insert,4,3940,3940,6 +DirectEmit,length,0,0,640,insert,5,4014,4014,6 +DirectEmit,length,0,0,640,insert,6,4029,4029,6 +DirectEmit,length,0,0,640,insert,7,3972,3972,6 +DirectEmit,length,0,0,640,insert,8,3981,3981,6 diff --git a/experiments/pg-1.png b/experiments/pg-1.png Binary files differindex 0f48914..b808a6c 100644 --- a/experiments/pg-1.png +++ b/experiments/pg-1.png diff --git a/experiments/pg-2.png b/experiments/pg-2.png Binary files differindex e7a2e6b..e0ac1f2 100644 --- a/experiments/pg-2.png +++ b/experiments/pg-2.png diff --git a/experiments/pg-3.png b/experiments/pg-3.png Binary files differindex 31f7d9a..9dc00e0 100644 --- a/experiments/pg-3.png +++ b/experiments/pg-3.png diff --git a/experiments/pg-4.png b/experiments/pg-4.png Binary files differindex 450547e..ae3fab6 100644 --- a/experiments/pg-4.png +++ b/experiments/pg-4.png diff --git a/experiments/pg-5.png b/experiments/pg-5.png Binary files differindex 32caf02..9c3e3f2 100644 --- a/experiments/pg-5.png +++ b/experiments/pg-5.png diff --git a/experiments/pg-6.png b/experiments/pg-6.png Binary files differindex 87e34d4..d1264cc 100644 --- a/experiments/pg-6.png +++ b/experiments/pg-6.png diff --git a/experiments/pg-7.png b/experiments/pg-7.png Binary files differindex cb9c41e..0a08184 100644 --- a/experiments/pg-7.png +++ b/experiments/pg-7.png diff --git a/experiments/pg-8.png b/experiments/pg-8.png Binary files differindex 4f875a7..f86fbb9 100644 --- a/experiments/pg-8.png +++ b/experiments/pg-8.png diff --git a/experiments/pg-9.png b/experiments/pg-9.png Binary files differnew file mode 100644 index 0000000..3f8874a --- /dev/null +++ b/experiments/pg-9.png diff --git a/experiments/report.typ b/experiments/report.typ index 2f27684..a15b162 100644 --- a/experiments/report.typ +++ b/experiments/report.typ @@ -47,7 +47,8 @@ rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") + rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") + - rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") + rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") + + rows("length-CPU360-final.csv") + rows("length-DirectEmit.csv") ) #let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv") @@ -106,16 +107,20 @@ implied, and it was only reachable by instrumenting the firmware; the fix that source reading suggested was written, measured, and found to be worth 20% on a 19 MB file and nothing at all on this board. Experiment 3. -Both were then cut. A keystroke is 8.37 ms, from 16.99, and 9.48 ms at a -160-character line, from 25.56 — the per-character term is 7.1 µs, from 54.3. Three of -the four steps that did it came from *not asking Unicode about ASCII*, and the largest -single one came from this repository's own `present` copying all 480 cells into vaxis -every frame whether or not any had changed. Experiment 4. +Both were then cut. A keystroke is *3.74 ms, from 16.99*, and 4.06 ms at a +160-character line, from 25.56 — the per-character term is 2.0 µs, from 54.3. Three of +the six steps that did it came from *not asking Unicode about ASCII*; the largest single +one came from this repository's own `present` copying all 480 cells into vaxis every +frame whether or not any had changed; and the last two were a clock that was only ever a +divider away, and a renderer that stopped asking vaxis to recompute a diff `present` had +already done. Experiment 4. -The target was 4 ms and it is not met: compute is 6.42 ms against a budget of about -2.05, so 3.1× remains. What is different from the start is that every microsecond of it -is located — our walk of the grid, vaxis's own diff, and pardes rebuilding a whole -Surface per keystroke — and that the two largest are not alternatives to each other. +The 4 ms target is met. Two findings on the way there are worth more than the number. +A 4× clock bought 2.6× because the grid walk is bounded by memory bandwidth, not by the +core. And the direct renderer, with less computation and a quarter of the bytes, first +measured *slower* — because the CH340 bridging the board to the host forwards a bulk +packet only when it is full, so a reply too small to fill one waits about a millisecond +for a timer. The frame has a minimum size and it belongs to the transport. = The instrument @@ -578,7 +583,8 @@ tree-sitter, which the board never runs. table.header([step], [fixed cost], [per char], [at 160 chars], [vs start]), table.hline(stroke: 0.5pt), ..( - ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", "RFast-fastcmp") + ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", + "RFast-fastcmp", "CPU360-final", "DirectEmit") ).map(b => { let xs = lengths.map(l => l * 1.0) let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) @@ -596,8 +602,9 @@ tree-sitter, which the board never runs. table.hline(), ), caption: [Each row is a separate firmware measured on the die, 5 document lengths × - 7 trials. The per-character column and the fixed column move for different reasons - and are worth reading separately.], + 7--9 trials. The per-character column and the fixed column move for different reasons + and are worth reading separately. The last two rows are the clock raise and the + renderer, and neither is a source optimisation in the sense the four above it are.], ) *Not asking Unicode about ASCII* accounts for the per-character column. Three fast @@ -643,30 +650,111 @@ the editor's heap it was enough to make `vx.resize` fail. == How a rendering change was made safe to believe A latency benchmark cannot see a corrupted screen, and `src/p4.zig` compiles for this -board alone — no host suite covers it. So the optimisation carries a comptime `shadow_grid` -switch whose `false` arm is the old behaviour, clear and write everything, and the two -firmwares were each run against the same 19-step workload on the die: inserts, deletes, -motions that move the modified-marker, a line outgrowing the viewport, backspaces that -shrink it. The screens were reconstructed from the wire and compared. +board alone — no host suite covers it. So each optimisation carries a comptime switch +whose `false` arm is the previous behaviour — `shadow_grid` for the incremental walk, +`direct_emit` for the renderer — and the firmwares were each run against the same +18-step workload on the die: inserts, deletes, motions that move the modified-marker, a +line outgrowing the viewport, backspaces that shrink it. The screens were reconstructed +from the wire and compared. -The first version of that comparison tracked characters only and would have passed a -colour regression in silence; it now hashes the SGR state of every cell as well. Both -text and per-row style hash are identical between the two paths. On the host side, -where the shared code does live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0 -mismatches, `hxparity` 561 cases with 0 mismatches. +That comparison had to be corrected twice, and the second time it had already produced a +false result. Tracking characters only would have passed a colour regression in silence, +so it began hashing SGR as well — but it hashed the escape *parameters* applied to each +cell, which is a history and not a state. Raising the clock changed which keystrokes +shared a frame, the same final colours were reached by a different route, and the +verifier *reported a difference on the die that did not exist*. It now decodes SGR into +resolved state — foreground, background, attribute set per cell — and two screens compare +equal when every cell holds the same character and the same resolved style, whatever +sequence of escapes the emitter chose to get there. Without that correction the +from-scratch renderer could not have been verified at all: it reaches identical screens +by deliberately different escapes. -== Where it stopped, and what 4 ms still needs +All three pairs are identical under it: shadow grid against full repaint, 90 MHz against +360, and direct emission against vaxis. On the host side, where the shared code does +live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0 mismatches, `hxparity` 561 cases +with 0 mismatches. -A keystroke is 8.37 ms, from 16.99 — and 9.48 ms at a 160-character line, from 25.56. -Compute is 6.42 ms against a budget of 2.05, so 3.1× remains, and all of it is -measured rather than guessed: our walk of the grid (979 µs), vaxis's own diff (2 455 µs), -and pardes rebuilding every cell of the Surface every frame (≈2 000 µs). +== The clock was a divider, not a design -The two large items are not alternatives to each other. vaxis's diff is redundant work -— `present` already computes exactly which cells moved, so emitting ANSI straight from -the shadow grid would remove most of it — but even a perfect renderer leaves pardes -rebuilding a whole frame per keystroke, and 4 ms at 40×12 needs that too. The -remaining path is therefore two changes, not one, and the second is architectural. +The board had been executing at 90 MHz for every measurement above, because the stock +second-stage bootloader is built for it. The CPLL was *already* running at 360 MHz — 90 +is exactly a quarter of it — so reaching 360 is four divider writes and a ROM call, with +no PLL to enable and no voltage step to sequence against. Verified against the systimer, +which is clocked from the crystal and therefore independent: 90 000 kHz before, +359 991 kHz after. Nothing else moves, and that is what makes it safe from a live +console: the UART0 baud clock comes from the crystal, the systimer from the crystal, and +the flash interface from a different PLL entirely. + +Compute fell 6.42 → 2.44 ms: *2.6× for a 4× clock*, and the shortfall is the +interesting part. The grid walk reads 27 KB per frame and takes 226 µs at 360 MHz, about +six cycles a byte, so it is bounded by L2MEM bandwidth rather than by the core. A +prediction made before the measurement, and held. + +== The renderer, and a minimum frame that belongs to the USB bridge + +vaxis's diff was redundant: `present` has already worked out exactly which cells moved, +so vaxis was being told the answer and then computing it again against its own copy. +Emitting the escapes directly — one absolute cursor position per run of changed cells, +absolute SGR rather than a delta, a hand-rolled formatter for the two digits that +sequence needs — took the board's render from 1 362 µs to 779 µs and the reply from 81 +bytes to 21. + +And it measured *slower*: 4.72 ms against 4.37. Fewer bytes, less computation, worse +round trip is the shape of a wrong model, so the search moved off the firmware. + +It is the USB bridge. The board reaches the host through a CH340, a full-speed part +whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet +is *full*. A 21-byte frame never fills one, so it waits in the bridge for an internal +timer to give up on more — worth about a millisecond, a quarter of the entire budget. +Routing through vaxis had only looked competitive because 81-byte frames fill a packet +by accident. + +#figure( + table( + columns: (auto, auto, auto, 1fr), + align: (right, right, right, left), + stroke: none, + table.hline(), + table.header([frame], [min], [median], []), + table.hline(stroke: 0.5pt), + [21 B], [3 843 µs], [4 817 µs], [direct, never fills a packet], + [49 B], [3 719 µs], [3 814 µs], [direct, padded past the boundary], + [81 B], [4 373 µs], [4 475 µs], [through vaxis, fills one by accident], + table.hline(), + ), + caption: [Identical board cost and a byte-identical screen across all three; only the + size of the reply differs. Note the *minimum*: the 21-byte frame's floor is already + 530 µs below vaxis's, which is exactly the computation that was saved. Only the median + was hostage to the bridge's timer.], +) + +So the frame has a minimum size and it is the transport's, not the terminal's. The +emitter pays it, padding with repeated absolute cursor positioning: idempotent, already +the sequence a frame ends on, incapable of altering a cell. This is the bargain an +Ethernet runt frame makes — the medium has a minimum and the sender pays it — and it is +a real trade rather than a free one, since the filler is wire time that delays a later +frame. It applies only while the frame is small, which is when there is wire to spare. + +A second instrument came out of this. The original bench sends keystrokes on a fixed +cadence, which locks the send phase to the host's 1 ms USB frame clock and makes the +round trip a staircase in board time — where a real saving can present as a regression. +Sleeping a uniform random 0–2 ms before each keystroke decorrelates them. That was not +what was happening here, but it had to be ruled out before the CH340 could be believed. + +== Where it stopped + +A keystroke is *3.74 ms*, from 16.99 — and 4.06 ms at a 160-character line, from 25.56. +The target was 4 ms and it is met, with the phase-randomised instrument agreeing over 60 +trials: median 3 829 µs, minimum 3 722, ninetieth percentile 3 930. The reply is 35 +bytes, down from 81. + +What remains is no longer dominated by anything this repository owns. Of the round trip, +roughly 2.9 ms is host, USB and wire — the floor measured independently against the +protocol responder, and now partly explained by the bridge's packet granularity — and +about 0.84 ms is board: 552 µs of pardes rebuilding every cell of the Surface every +frame, 223 µs walking the grid, 60 µs parsing and editing. The Surface rebuild is the +one architectural item left, and it is worth 552 µs against a 2.9 ms floor, so the next +order of magnitude is not in the firmware at all. #figure( table( @@ -717,7 +805,10 @@ remaining path is therefore two changes, not one, and the second is architectura reordered it. Rows 2 and the last row were established by measurement; rows 3--5 are derived from measured quantities and cited source. The last row is listed unranked because it is already applied and, on *this* target, buys nothing -- - which is precisely why it is worth recording.], + which is precisely why it is worth recording. Kept as the prediction it was: + Experiment 4 has since applied rows 1 and 2 and *abandoned row 3*, because raising + the line rate moves `settle` and leaves the round trip alone --- and the bridge + finding above is the reason row 3 looked attractive in the first place.], ) = Threats to validity @@ -729,8 +820,11 @@ is a no-op, so trials after the first measured nothing, and the genuine full-repaint figure quoted throughout (1 392 bytes, 150 ms to settle) comes from a single first resize rather than from a median. -Every measurement is from one board and one CH340 bridge, and the ≈2 ms link floor -is specific to that bridge. Round trip is time to first byte and therefore says +Every measurement is from one board and one CH340 bridge, and the link floor --- ≈2 ms +as measured here, ≈2.9 ms once the bridge's packet granularity is accounted for --- is +specific to that bridge; a device that forwards short packets promptly would show a +different floor and would not reward the frame padding of Experiment 4 at all. +Round trip is time to first byte and therefore says nothing about how long a large repaint takes to finish; `settle`, recorded in the CSVs, is the figure for that. Both builds were measured in a single session each, so slow drift — temperature, or the host's own scheduling — would appear as a |
