diff options
| -rw-r--r-- | experiments/attribution-stages.csv | 6 | ||||
| -rw-r--r-- | experiments/pg-1.png | bin | 0 -> 264089 bytes | |||
| -rw-r--r-- | experiments/pg-2.png | bin | 0 -> 181594 bytes | |||
| -rw-r--r-- | experiments/pg-3.png | bin | 0 -> 191333 bytes | |||
| -rw-r--r-- | experiments/pg-4.png | bin | 0 -> 245170 bytes | |||
| -rw-r--r-- | experiments/pg-5.png | bin | 0 -> 277654 bytes | |||
| -rw-r--r-- | experiments/pg-6.png | bin | 0 -> 270483 bytes | |||
| -rw-r--r-- | experiments/pg-7.png | bin | 0 -> 248684 bytes | |||
| -rw-r--r-- | experiments/pg-8.png | bin | 0 -> 22182 bytes | |||
| -rw-r--r-- | experiments/report.typ | 153 | ||||
| -rw-r--r-- | experiments/stages.csv | 9 |
11 files changed, 163 insertions, 5 deletions
diff --git a/experiments/attribution-stages.csv b/experiments/attribution-stages.csv new file mode 100644 index 0000000..90e62ce --- /dev/null +++ b/experiments/attribution-stages.csv @@ -0,0 +1,6 @@ +label,experiment,chars,input_cy,render_cy,input_us,render_us +ReleaseSmall-prof,attribution,1,19800000,1332360000,220,14804 +ReleaseSmall-prof,attribution,40,17550000,1380690000,195,15341 +ReleaseSmall-prof,attribution,80,19080000,1540260000,212,17114 +ReleaseSmall-prof,attribution,160,20610000,1859130000,229,20657 +ReleaseSmall-prof,attribution,240,22500000,2215350000,250,24615 diff --git a/experiments/pg-1.png b/experiments/pg-1.png Binary files differnew file mode 100644 index 0000000..0f48914 --- /dev/null +++ b/experiments/pg-1.png diff --git a/experiments/pg-2.png b/experiments/pg-2.png Binary files differnew file mode 100644 index 0000000..e7a2e6b --- /dev/null +++ b/experiments/pg-2.png diff --git a/experiments/pg-3.png b/experiments/pg-3.png Binary files differnew file mode 100644 index 0000000..31f7d9a --- /dev/null +++ b/experiments/pg-3.png diff --git a/experiments/pg-4.png b/experiments/pg-4.png Binary files differnew file mode 100644 index 0000000..450547e --- /dev/null +++ b/experiments/pg-4.png diff --git a/experiments/pg-5.png b/experiments/pg-5.png Binary files differnew file mode 100644 index 0000000..32caf02 --- /dev/null +++ b/experiments/pg-5.png diff --git a/experiments/pg-6.png b/experiments/pg-6.png Binary files differnew file mode 100644 index 0000000..87e34d4 --- /dev/null +++ b/experiments/pg-6.png diff --git a/experiments/pg-7.png b/experiments/pg-7.png Binary files differnew file mode 100644 index 0000000..cb9c41e --- /dev/null +++ b/experiments/pg-7.png diff --git a/experiments/pg-8.png b/experiments/pg-8.png Binary files differnew file mode 100644 index 0000000..4f875a7 --- /dev/null +++ b/experiments/pg-8.png diff --git a/experiments/report.typ b/experiments/report.typ index 8134db2..2f27684 100644 --- a/experiments/report.typ +++ b/experiments/report.typ @@ -41,8 +41,14 @@ s.at(int(s.len() / 2)) } -#let length_rows = rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") - + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") +// One expression on one line: a continuation beginning with `+` is list markup to Typst, and it +// rendered the file names as a numbered item on page one. +#let length_rows = ( + rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") + + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") + + rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") + + rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") +) #let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv") // Figure 1 compares the two optimisation modes only; the `lineSpan` variants are the SAME source @@ -100,9 +106,16 @@ implied, and it was only reachable by instrumenting the firmware; the fix that source reading suggested was written, measured, and found to be worth 20% on a 19 MB file and nothing at all on this board. Experiment 3. -One build-flag change — compiling the editor object `ReleaseFast` instead of -`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost, -for 35% more flash. It remains the best ratio measured here. +Both were then cut. A keystroke is 8.37 ms, from 16.99, and 9.48 ms at a +160-character line, from 25.56 — the per-character term is 7.1 µs, from 54.3. Three of +the four steps that did it came from *not asking Unicode about ASCII*, and the largest +single one came from this repository's own `present` copying all 480 cells into vaxis +every frame whether or not any had changed. Experiment 4. + +The target was 4 ms and it is not met: compute is 6.42 ms against a budget of about +2.05, so 3.1× remains. What is different from the start is that every microsecond of it +is located — our walk of the grid, vaxis's own diff, and pardes rebuilding a whole +Surface per keystroke — and that the two largest are not alternatives to each other. = The instrument @@ -544,6 +557,136 @@ of the renderer with a large race surface on a board that has no debugger, and i would not shorten a single keystroke's round trip anyway, because the host still waits for that render. += Experiment 4 --- cutting it, and what each cut was worth + +The target set after Experiment 3 was 4 ms. Round trip is time to the first response +byte, so it is ≈2 ms of host and USB latency plus computation; 4 ms therefore means a +compute budget of about 2 ms, and raising the line rate cannot help — at 115200 an +81-byte reply is 7 ms of wire but almost none of it lands before the first byte. + +Every step below was profiled first, in the board's *exact* configuration: 40×12 and +`-Dtree-sitter=disabled`, because the P4 build has no tree-sitter and a profile that +includes it is a profile of a different program. Half of the first profile was +tree-sitter, which the board never runs. + +#figure( + table( + columns: (auto, auto, auto, auto, auto), + align: (left, right, right, right, right), + stroke: none, + table.hline(), + table.header([step], [fixed cost], [per char], [at 160 chars], [vs start]), + table.hline(stroke: 0.5pt), + ..( + ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", "RFast-fastcmp") + ).map(b => { + let xs = lengths.map(l => l * 1.0) + let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) + let f = fit(xs, ys) + let at160 = ys.last() + let ref160 = med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == 160) + ( + raw(b), + [#calc.round(f.intercept, digits: 2) ms], + [#calc.round(f.slope * 1000, digits: 1) µs], + [#calc.round(at160, digits: 2) ms], + [#calc.round(at160 / ref160, digits: 2)×], + ) + }).flatten(), + table.hline(), + ), + caption: [Each row is a separate firmware measured on the die, 5 document lengths × + 7 trials. The per-character column and the fixed column move for different reasons + and are worth reading separately.], +) + +*Not asking Unicode about ASCII* accounts for the per-character column. Three fast +paths, each guarded so non-ASCII text takes exactly the road it took before: +`modal.graphemeStart` (21.5% of a keystroke, and the reason cost followed the cursor's +column — it walks graphemes from the start of the text with the full UAX #29 state +machine), `Surface.print` (26.2%, which per character took a UTF-8 length, a decode, a +*freshly constructed* grapheme iterator, a validation and a width lookup to conclude +that `y` is one cell), and `graphemeDisplayWidth` (6.9%, all of it asking a Unicode +table about ASCII). The guard is the same in each: an ASCII scalar is its own grapheme +cluster unless it is a CR before an LF, because every rule that could join it — +Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator — is spelled with non-ASCII +scalars. Per character: 54.3 → 6.9 µs. + +*The fixed cost needed the frame broken open.* A second render with nothing changed +cost the same as the first, so the ~15 ms was unconditional; `pardes_p4_frame_prof` +then reported the three stages separately. + +#figure( + table( + columns: (auto, auto, auto), + align: (left, right, right), + stroke: none, + table.hline(), + table.header([stage of one frame], [before], [after]), + table.hline(stroke: 0.5pt), + [copy the Surface into vaxis's grid], [6 750 µs], [*979 µs*], + [vaxis diffs its grid and emits], [2 460 µs], [2 455 µs], + [pardes rebuilds the whole Surface], [≈2 600 µs], [≈2 000 µs], + [push the bytes into the UART], [1 µs], [1 µs], + table.hline(), + ), + caption: [Cycle counts read on the die. The largest item was in *this repository's* + own `present`, not in pardes and not in vaxis: it copied all 480 cells every frame, + whether or not any had changed.], +) + +So `present` now keeps the previous Surface and tells vaxis only what moved, comparing +cells as bytes rather than through `std.meta.eql` on a colour union and eight booleans. +The grid lives in `.bss`, and that is a bug fix rather than a flourish: allocated from +the editor's heap it was enough to make `vx.resize` fail. + +== How a rendering change was made safe to believe + +A latency benchmark cannot see a corrupted screen, and `src/p4.zig` compiles for this +board alone — no host suite covers it. So the optimisation carries a comptime `shadow_grid` +switch whose `false` arm is the old behaviour, clear and write everything, and the two +firmwares were each run against the same 19-step workload on the die: inserts, deletes, +motions that move the modified-marker, a line outgrowing the viewport, backspaces that +shrink it. The screens were reconstructed from the wire and compared. + +The first version of that comparison tracked characters only and would have passed a +colour regression in silence; it now hashes the SGR state of every cell as well. Both +text and per-row style hash are identical between the two paths. On the host side, +where the shared code does live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0 +mismatches, `hxparity` 561 cases with 0 mismatches. + +== Where it stopped, and what 4 ms still needs + +A keystroke is 8.37 ms, from 16.99 — and 9.48 ms at a 160-character line, from 25.56. +Compute is 6.42 ms against a budget of 2.05, so 3.1× remains, and all of it is +measured rather than guessed: our walk of the grid (979 µs), vaxis's own diff (2 455 µs), +and pardes rebuilding every cell of the Surface every frame (≈2 000 µs). + +The two large items are not alternatives to each other. vaxis's diff is redundant work +— `present` already computes exactly which cells moved, so emitting ANSI straight from +the shadow grid would remove most of it — but even a perfect renderer leaves pardes +rebuilding a whole frame per keystroke, and 4 ms at 40×12 needs that too. The +remaining path is therefore two changes, not one, and the second is architectural. + +#figure( + table( + columns: (auto, 1fr), + align: (left, left), + stroke: none, + table.hline(), + table.header([found], [not fixed]), + table.hline(stroke: 0.5pt), + [`vx.resize` fails], [A runtime geometry change takes its allocation-failure path, + restores the previous size and returns: 80 bytes go out where 1 392 should, and + the screen keeps its old shape. Reproduced with the shadow grid compiled out, so + it predates it. The board has one geometry per session — which is also why the + staleness test compares two firmwares rather than forcing a repaint by resizing.], + table.hline(), + ), + caption: [A bug the optimisation work walked into and characterised but did not + repair.], +) + = Ranked by measured benefit #figure( diff --git a/experiments/stages.csv b/experiments/stages.csv new file mode 100644 index 0000000..6c901e7 --- /dev/null +++ b/experiments/stages.csv @@ -0,0 +1,9 @@ +label,stage,us +before,copy Surface into vaxis,6750 +before,vaxis diff and emit,2460 +before,pardes Surface rebuild,2600 +before,push into the UART,1 +after,copy Surface into vaxis,979 +after,vaxis diff and emit,2455 +after,pardes Surface rebuild,2000 +after,push into the UART,1 |
