summaryrefslogtreecommitdiff
path: root/experiments
diff options
context:
space:
mode:
Diffstat (limited to 'experiments')
-rw-r--r--experiments/attribution-stages.csv6
-rw-r--r--experiments/pg-1.pngbin0 -> 264089 bytes
-rw-r--r--experiments/pg-2.pngbin0 -> 181594 bytes
-rw-r--r--experiments/pg-3.pngbin0 -> 191333 bytes
-rw-r--r--experiments/pg-4.pngbin0 -> 245170 bytes
-rw-r--r--experiments/pg-5.pngbin0 -> 277654 bytes
-rw-r--r--experiments/pg-6.pngbin0 -> 270483 bytes
-rw-r--r--experiments/pg-7.pngbin0 -> 248684 bytes
-rw-r--r--experiments/pg-8.pngbin0 -> 22182 bytes
-rw-r--r--experiments/report.typ153
-rw-r--r--experiments/stages.csv9
11 files changed, 163 insertions, 5 deletions
diff --git a/experiments/attribution-stages.csv b/experiments/attribution-stages.csv
new file mode 100644
index 0000000..90e62ce
--- /dev/null
+++ b/experiments/attribution-stages.csv
@@ -0,0 +1,6 @@
+label,experiment,chars,input_cy,render_cy,input_us,render_us
+ReleaseSmall-prof,attribution,1,19800000,1332360000,220,14804
+ReleaseSmall-prof,attribution,40,17550000,1380690000,195,15341
+ReleaseSmall-prof,attribution,80,19080000,1540260000,212,17114
+ReleaseSmall-prof,attribution,160,20610000,1859130000,229,20657
+ReleaseSmall-prof,attribution,240,22500000,2215350000,250,24615
diff --git a/experiments/pg-1.png b/experiments/pg-1.png
new file mode 100644
index 0000000..0f48914
--- /dev/null
+++ b/experiments/pg-1.png
Binary files differ
diff --git a/experiments/pg-2.png b/experiments/pg-2.png
new file mode 100644
index 0000000..e7a2e6b
--- /dev/null
+++ b/experiments/pg-2.png
Binary files differ
diff --git a/experiments/pg-3.png b/experiments/pg-3.png
new file mode 100644
index 0000000..31f7d9a
--- /dev/null
+++ b/experiments/pg-3.png
Binary files differ
diff --git a/experiments/pg-4.png b/experiments/pg-4.png
new file mode 100644
index 0000000..450547e
--- /dev/null
+++ b/experiments/pg-4.png
Binary files differ
diff --git a/experiments/pg-5.png b/experiments/pg-5.png
new file mode 100644
index 0000000..32caf02
--- /dev/null
+++ b/experiments/pg-5.png
Binary files differ
diff --git a/experiments/pg-6.png b/experiments/pg-6.png
new file mode 100644
index 0000000..87e34d4
--- /dev/null
+++ b/experiments/pg-6.png
Binary files differ
diff --git a/experiments/pg-7.png b/experiments/pg-7.png
new file mode 100644
index 0000000..cb9c41e
--- /dev/null
+++ b/experiments/pg-7.png
Binary files differ
diff --git a/experiments/pg-8.png b/experiments/pg-8.png
new file mode 100644
index 0000000..4f875a7
--- /dev/null
+++ b/experiments/pg-8.png
Binary files differ
diff --git a/experiments/report.typ b/experiments/report.typ
index 8134db2..2f27684 100644
--- a/experiments/report.typ
+++ b/experiments/report.typ
@@ -41,8 +41,14 @@
s.at(int(s.len() / 2))
}
-#let length_rows = rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv")
- + rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv")
+// One expression on one line: a continuation beginning with `+` is list markup to Typst, and it
+// rendered the file names as a numbered item on page one.
+#let length_rows = (
+ rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") +
+ rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") +
+ rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") +
+ rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv")
+)
#let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv")
// Figure 1 compares the two optimisation modes only; the `lineSpan` variants are the SAME source
@@ -100,9 +106,16 @@ implied, and it was only reachable by instrumenting the firmware; the fix that
source reading suggested was written, measured, and found to be worth 20% on a 19 MB
file and nothing at all on this board. Experiment 3.
-One build-flag change — compiling the editor object `ReleaseFast` instead of
-`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost,
-for 35% more flash. It remains the best ratio measured here.
+Both were then cut. A keystroke is 8.37 ms, from 16.99, and 9.48 ms at a
+160-character line, from 25.56 — the per-character term is 7.1 µs, from 54.3. Three of
+the four steps that did it came from *not asking Unicode about ASCII*, and the largest
+single one came from this repository's own `present` copying all 480 cells into vaxis
+every frame whether or not any had changed. Experiment 4.
+
+The target was 4 ms and it is not met: compute is 6.42 ms against a budget of about
+2.05, so 3.1× remains. What is different from the start is that every microsecond of it
+is located — our walk of the grid, vaxis's own diff, and pardes rebuilding a whole
+Surface per keystroke — and that the two largest are not alternatives to each other.
= The instrument
@@ -544,6 +557,136 @@ of the renderer with a large race surface on a board that has no debugger, and i
would not shorten a single keystroke's round trip anyway, because the host still
waits for that render.
+= Experiment 4 --- cutting it, and what each cut was worth
+
+The target set after Experiment 3 was 4 ms. Round trip is time to the first response
+byte, so it is ≈2 ms of host and USB latency plus computation; 4 ms therefore means a
+compute budget of about 2 ms, and raising the line rate cannot help — at 115200 an
+81-byte reply is 7 ms of wire but almost none of it lands before the first byte.
+
+Every step below was profiled first, in the board's *exact* configuration: 40×12 and
+`-Dtree-sitter=disabled`, because the P4 build has no tree-sitter and a profile that
+includes it is a profile of a different program. Half of the first profile was
+tree-sitter, which the board never runs.
+
+#figure(
+ table(
+ columns: (auto, auto, auto, auto, auto),
+ align: (left, right, right, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([step], [fixed cost], [per char], [at 160 chars], [vs start]),
+ table.hline(stroke: 0.5pt),
+ ..(
+ ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", "RFast-fastcmp")
+ ).map(b => {
+ let xs = lengths.map(l => l * 1.0)
+ let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l))
+ let f = fit(xs, ys)
+ let at160 = ys.last()
+ let ref160 = med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == 160)
+ (
+ raw(b),
+ [#calc.round(f.intercept, digits: 2) ms],
+ [#calc.round(f.slope * 1000, digits: 1) µs],
+ [#calc.round(at160, digits: 2) ms],
+ [#calc.round(at160 / ref160, digits: 2)×],
+ )
+ }).flatten(),
+ table.hline(),
+ ),
+ caption: [Each row is a separate firmware measured on the die, 5 document lengths ×
+ 7 trials. The per-character column and the fixed column move for different reasons
+ and are worth reading separately.],
+)
+
+*Not asking Unicode about ASCII* accounts for the per-character column. Three fast
+paths, each guarded so non-ASCII text takes exactly the road it took before:
+`modal.graphemeStart` (21.5% of a keystroke, and the reason cost followed the cursor's
+column — it walks graphemes from the start of the text with the full UAX #29 state
+machine), `Surface.print` (26.2%, which per character took a UTF-8 length, a decode, a
+*freshly constructed* grapheme iterator, a validation and a width lookup to conclude
+that `y` is one cell), and `graphemeDisplayWidth` (6.9%, all of it asking a Unicode
+table about ASCII). The guard is the same in each: an ASCII scalar is its own grapheme
+cluster unless it is a CR before an LF, because every rule that could join it —
+Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator — is spelled with non-ASCII
+scalars. Per character: 54.3 → 6.9 µs.
+
+*The fixed cost needed the frame broken open.* A second render with nothing changed
+cost the same as the first, so the ~15 ms was unconditional; `pardes_p4_frame_prof`
+then reported the three stages separately.
+
+#figure(
+ table(
+ columns: (auto, auto, auto),
+ align: (left, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([stage of one frame], [before], [after]),
+ table.hline(stroke: 0.5pt),
+ [copy the Surface into vaxis's grid], [6 750 µs], [*979 µs*],
+ [vaxis diffs its grid and emits], [2 460 µs], [2 455 µs],
+ [pardes rebuilds the whole Surface], [≈2 600 µs], [≈2 000 µs],
+ [push the bytes into the UART], [1 µs], [1 µs],
+ table.hline(),
+ ),
+ caption: [Cycle counts read on the die. The largest item was in *this repository's*
+ own `present`, not in pardes and not in vaxis: it copied all 480 cells every frame,
+ whether or not any had changed.],
+)
+
+So `present` now keeps the previous Surface and tells vaxis only what moved, comparing
+cells as bytes rather than through `std.meta.eql` on a colour union and eight booleans.
+The grid lives in `.bss`, and that is a bug fix rather than a flourish: allocated from
+the editor's heap it was enough to make `vx.resize` fail.
+
+== How a rendering change was made safe to believe
+
+A latency benchmark cannot see a corrupted screen, and `src/p4.zig` compiles for this
+board alone — no host suite covers it. So the optimisation carries a comptime `shadow_grid`
+switch whose `false` arm is the old behaviour, clear and write everything, and the two
+firmwares were each run against the same 19-step workload on the die: inserts, deletes,
+motions that move the modified-marker, a line outgrowing the viewport, backspaces that
+shrink it. The screens were reconstructed from the wire and compared.
+
+The first version of that comparison tracked characters only and would have passed a
+colour regression in silence; it now hashes the SGR state of every cell as well. Both
+text and per-row style hash are identical between the two paths. On the host side,
+where the shared code does live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0
+mismatches, `hxparity` 561 cases with 0 mismatches.
+
+== Where it stopped, and what 4 ms still needs
+
+A keystroke is 8.37 ms, from 16.99 — and 9.48 ms at a 160-character line, from 25.56.
+Compute is 6.42 ms against a budget of 2.05, so 3.1× remains, and all of it is
+measured rather than guessed: our walk of the grid (979 µs), vaxis's own diff (2 455 µs),
+and pardes rebuilding every cell of the Surface every frame (≈2 000 µs).
+
+The two large items are not alternatives to each other. vaxis's diff is redundant work
+— `present` already computes exactly which cells moved, so emitting ANSI straight from
+the shadow grid would remove most of it — but even a perfect renderer leaves pardes
+rebuilding a whole frame per keystroke, and 4 ms at 40×12 needs that too. The
+remaining path is therefore two changes, not one, and the second is architectural.
+
+#figure(
+ table(
+ columns: (auto, 1fr),
+ align: (left, left),
+ stroke: none,
+ table.hline(),
+ table.header([found], [not fixed]),
+ table.hline(stroke: 0.5pt),
+ [`vx.resize` fails], [A runtime geometry change takes its allocation-failure path,
+ restores the previous size and returns: 80 bytes go out where 1 392 should, and
+ the screen keeps its old shape. Reproduced with the shadow grid compiled out, so
+ it predates it. The board has one geometry per session — which is also why the
+ staleness test compares two firmwares rather than forcing a repaint by resizing.],
+ table.hline(),
+ ),
+ caption: [A bug the optimisation work walked into and characterised but did not
+ repair.],
+)
+
= Ranked by measured benefit
#figure(
diff --git a/experiments/stages.csv b/experiments/stages.csv
new file mode 100644
index 0000000..6c901e7
--- /dev/null
+++ b/experiments/stages.csv
@@ -0,0 +1,9 @@
+label,stage,us
+before,copy Surface into vaxis,6750
+before,vaxis diff and emit,2460
+before,pardes Surface rebuild,2600
+before,push into the UART,1
+after,copy Surface into vaxis,979
+after,vaxis diff and emit,2455
+after,pardes Surface rebuild,2000
+after,push into the UART,1