diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 20:33:51 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 20:33:51 -0300 |
| commit | 7664c57d93cd3f3bdb719f7ca915d657668ddfd9 (patch) | |
| tree | 73b0dca1633126ddf3c2aaf3989f92cfb0813996 | |
| parent | 05610a2a42b3d026473aa84ddfef1592feda5f17 (diff) | |
| download | esp32p4-7664c57d93cd3f3bdb719f7ca915d657668ddfd9.tar.gz esp32p4-7664c57d93cd3f3bdb719f7ca915d657668ddfd9.zip | |
Report: the optimisation campaign, and the 3.1x that is left
Experiment 4 records what each cut was worth, measured on the die at every step, and
the stage table that made the cuts findable at all. Also two corrections to the
document itself:
The data plumbing loaded four of the eight datasets, so the progression table would
have been computed from a subset. It now loads all eight.
A continuation line beginning with `+` is list markup to Typst, so the CSV file names
were rendering as a numbered item at the top of page one. One expression, one line.
The summary no longer ends on the ReleaseFast flag as the best available change; it
ends on the measured 16.99 -> 8.37 ms and on the fact that the target is not met, with
the remaining 6.42 ms of compute broken into the three items it actually consists of.
Saying 'not met' in the summary matters more than the table: the number that was
missed is the one a reader should see first.
| -rw-r--r-- | experiments/attribution-stages.csv | 6 | ||||
| -rw-r--r-- | experiments/pg-1.png | bin | 0 -> 264089 bytes | |||
| -rw-r--r-- | experiments/pg-2.png | bin | 0 -> 181594 bytes | |||
| -rw-r--r-- | experiments/pg-3.png | bin | 0 -> 191333 bytes | |||
| -rw-r--r-- | experiments/pg-4.png | bin | 0 -> 245170 bytes | |||
| -rw-r--r-- | experiments/pg-5.png | bin | 0 -> 277654 bytes | |||
| -rw-r--r-- | experiments/pg-6.png | bin | 0 -> 270483 bytes | |||
| -rw-r--r-- | experiments/pg-7.png | bin | 0 -> 248684 bytes | |||
| -rw-r--r-- | experiments/pg-8.png | bin | 0 -> 22182 bytes | |||
| -rw-r--r-- | experiments/report.typ | 153 | ||||
| -rw-r--r-- | experiments/stages.csv | 9 |
11 files changed, 163 insertions, 5 deletions
diff --git a/experiments/attribution-stages.csv b/experiments/attribution-stages.csv new file mode 100644 index 0000000..90e62ce --- /dev/null +++ b/experiments/attribution-stages.csv @@ -0,0 +1,6 @@ +label,experiment,chars,input_cy,render_cy,input_us,render_us +ReleaseSmall-prof,attribution,1,19800000,1332360000,220,14804 +ReleaseSmall-prof,attribution,40,17550000,1380690000,195,15341 +ReleaseSmall-prof,attribution,80,19080000,1540260000,212,17114 +ReleaseSmall-prof,attribution,160,20610000,1859130000,229,20657 +ReleaseSmall-prof,attribution,240,22500000,2215350000,250,24615 diff --git a/experiments/pg-1.png b/experiments/pg-1.png Binary files differnew file mode 100644 index 0000000..0f48914 --- /dev/null +++ b/experiments/pg-1.png diff --git a/experiments/pg-2.png b/experiments/pg-2.png Binary files differnew file mode 100644 index 0000000..e7a2e6b --- /dev/null +++ b/experiments/pg-2.png diff --git a/experiments/pg-3.png b/experiments/pg-3.png Binary files differnew file mode 100644 index 0000000..31f7d9a --- /dev/null +++ b/experiments/pg-3.png diff --git a/experiments/pg-4.png b/experiments/pg-4.png Binary files differnew file mode 100644 index 0000000..450547e --- /dev/null +++ b/experiments/pg-4.png diff --git a/experiments/pg-5.png b/experiments/pg-5.png Binary files differnew file mode 100644 index 0000000..32caf02 --- /dev/null +++ b/experiments/pg-5.png diff --git a/experiments/pg-6.png b/experiments/pg-6.png Binary files differnew file mode 100644 index 0000000..87e34d4 --- /dev/null +++ b/experiments/pg-6.png diff --git a/experiments/pg-7.png b/experiments/pg-7.png Binary files differnew file mode 100644 index 0000000..cb9c41e --- /dev/null +++ b/experiments/pg-7.png diff --git a/experiments/pg-8.png b/experiments/pg-8.png Binary files differnew file mode 100644 index 0000000..4f875a7 --- /dev/null +++ b/experiments/pg-8.png diff --git a/experiments/report.typ b/experiments/report.typ index 8134db2..2f27684 100644 --- a/experiments/report.typ +++ b/experiments/report.typ @@ -41,8 +41,14 @@ s.at(int(s.len() / 2)) } -#let length_rows = rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") - + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") +// One expression on one line: a continuation beginning with `+` is list markup to Typst, and it +// rendered the file names as a numbered item on page one. +#let length_rows = ( + rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") + + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") + + rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") + + rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") +) #let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv") // Figure 1 compares the two optimisation modes only; the `lineSpan` variants are the SAME source @@ -100,9 +106,16 @@ implied, and it was only reachable by instrumenting the firmware; the fix that source reading suggested was written, measured, and found to be worth 20% on a 19 MB file and nothing at all on this board. Experiment 3. -One build-flag change — compiling the editor object `ReleaseFast` instead of -`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost, -for 35% more flash. It remains the best ratio measured here. +Both were then cut. A keystroke is 8.37 ms, from 16.99, and 9.48 ms at a +160-character line, from 25.56 — the per-character term is 7.1 µs, from 54.3. Three of +the four steps that did it came from *not asking Unicode about ASCII*, and the largest +single one came from this repository's own `present` copying all 480 cells into vaxis +every frame whether or not any had changed. Experiment 4. + +The target was 4 ms and it is not met: compute is 6.42 ms against a budget of about +2.05, so 3.1× remains. What is different from the start is that every microsecond of it +is located — our walk of the grid, vaxis's own diff, and pardes rebuilding a whole +Surface per keystroke — and that the two largest are not alternatives to each other. = The instrument @@ -544,6 +557,136 @@ of the renderer with a large race surface on a board that has no debugger, and i would not shorten a single keystroke's round trip anyway, because the host still waits for that render. += Experiment 4 --- cutting it, and what each cut was worth + +The target set after Experiment 3 was 4 ms. Round trip is time to the first response +byte, so it is ≈2 ms of host and USB latency plus computation; 4 ms therefore means a +compute budget of about 2 ms, and raising the line rate cannot help — at 115200 an +81-byte reply is 7 ms of wire but almost none of it lands before the first byte. + +Every step below was profiled first, in the board's *exact* configuration: 40×12 and +`-Dtree-sitter=disabled`, because the P4 build has no tree-sitter and a profile that +includes it is a profile of a different program. Half of the first profile was +tree-sitter, which the board never runs. + +#figure( + table( + columns: (auto, auto, auto, auto, auto), + align: (left, right, right, right, right), + stroke: none, + table.hline(), + table.header([step], [fixed cost], [per char], [at 160 chars], [vs start]), + table.hline(stroke: 0.5pt), + ..( + ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", "RFast-fastcmp") + ).map(b => { + let xs = lengths.map(l => l * 1.0) + let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) + let f = fit(xs, ys) + let at160 = ys.last() + let ref160 = med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == 160) + ( + raw(b), + [#calc.round(f.intercept, digits: 2) ms], + [#calc.round(f.slope * 1000, digits: 1) µs], + [#calc.round(at160, digits: 2) ms], + [#calc.round(at160 / ref160, digits: 2)×], + ) + }).flatten(), + table.hline(), + ), + caption: [Each row is a separate firmware measured on the die, 5 document lengths × + 7 trials. The per-character column and the fixed column move for different reasons + and are worth reading separately.], +) + +*Not asking Unicode about ASCII* accounts for the per-character column. Three fast +paths, each guarded so non-ASCII text takes exactly the road it took before: +`modal.graphemeStart` (21.5% of a keystroke, and the reason cost followed the cursor's +column — it walks graphemes from the start of the text with the full UAX #29 state +machine), `Surface.print` (26.2%, which per character took a UTF-8 length, a decode, a +*freshly constructed* grapheme iterator, a validation and a width lookup to conclude +that `y` is one cell), and `graphemeDisplayWidth` (6.9%, all of it asking a Unicode +table about ASCII). The guard is the same in each: an ASCII scalar is its own grapheme +cluster unless it is a CR before an LF, because every rule that could join it — +Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator — is spelled with non-ASCII +scalars. Per character: 54.3 → 6.9 µs. + +*The fixed cost needed the frame broken open.* A second render with nothing changed +cost the same as the first, so the ~15 ms was unconditional; `pardes_p4_frame_prof` +then reported the three stages separately. + +#figure( + table( + columns: (auto, auto, auto), + align: (left, right, right), + stroke: none, + table.hline(), + table.header([stage of one frame], [before], [after]), + table.hline(stroke: 0.5pt), + [copy the Surface into vaxis's grid], [6 750 µs], [*979 µs*], + [vaxis diffs its grid and emits], [2 460 µs], [2 455 µs], + [pardes rebuilds the whole Surface], [≈2 600 µs], [≈2 000 µs], + [push the bytes into the UART], [1 µs], [1 µs], + table.hline(), + ), + caption: [Cycle counts read on the die. The largest item was in *this repository's* + own `present`, not in pardes and not in vaxis: it copied all 480 cells every frame, + whether or not any had changed.], +) + +So `present` now keeps the previous Surface and tells vaxis only what moved, comparing +cells as bytes rather than through `std.meta.eql` on a colour union and eight booleans. +The grid lives in `.bss`, and that is a bug fix rather than a flourish: allocated from +the editor's heap it was enough to make `vx.resize` fail. + +== How a rendering change was made safe to believe + +A latency benchmark cannot see a corrupted screen, and `src/p4.zig` compiles for this +board alone — no host suite covers it. So the optimisation carries a comptime `shadow_grid` +switch whose `false` arm is the old behaviour, clear and write everything, and the two +firmwares were each run against the same 19-step workload on the die: inserts, deletes, +motions that move the modified-marker, a line outgrowing the viewport, backspaces that +shrink it. The screens were reconstructed from the wire and compared. + +The first version of that comparison tracked characters only and would have passed a +colour regression in silence; it now hashes the SGR state of every cell as well. Both +text and per-row style hash are identical between the two paths. On the host side, +where the shared code does live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0 +mismatches, `hxparity` 561 cases with 0 mismatches. + +== Where it stopped, and what 4 ms still needs + +A keystroke is 8.37 ms, from 16.99 — and 9.48 ms at a 160-character line, from 25.56. +Compute is 6.42 ms against a budget of 2.05, so 3.1× remains, and all of it is +measured rather than guessed: our walk of the grid (979 µs), vaxis's own diff (2 455 µs), +and pardes rebuilding every cell of the Surface every frame (≈2 000 µs). + +The two large items are not alternatives to each other. vaxis's diff is redundant work +— `present` already computes exactly which cells moved, so emitting ANSI straight from +the shadow grid would remove most of it — but even a perfect renderer leaves pardes +rebuilding a whole frame per keystroke, and 4 ms at 40×12 needs that too. The +remaining path is therefore two changes, not one, and the second is architectural. + +#figure( + table( + columns: (auto, 1fr), + align: (left, left), + stroke: none, + table.hline(), + table.header([found], [not fixed]), + table.hline(stroke: 0.5pt), + [`vx.resize` fails], [A runtime geometry change takes its allocation-failure path, + restores the previous size and returns: 80 bytes go out where 1 392 should, and + the screen keeps its old shape. Reproduced with the shadow grid compiled out, so + it predates it. The board has one geometry per session — which is also why the + staleness test compares two firmwares rather than forcing a repaint by resizing.], + table.hline(), + ), + caption: [A bug the optimisation work walked into and characterised but did not + repair.], +) + = Ranked by measured benefit #figure( diff --git a/experiments/stages.csv b/experiments/stages.csv new file mode 100644 index 0000000..6c901e7 --- /dev/null +++ b/experiments/stages.csv @@ -0,0 +1,9 @@ +label,stage,us +before,copy Surface into vaxis,6750 +before,vaxis diff and emit,2460 +before,pardes Surface rebuild,2600 +before,push into the UART,1 +after,copy Surface into vaxis,979 +after,vaxis diff and emit,2455 +after,pardes Surface rebuild,2000 +after,push into the UART,1 |
