summaryrefslogtreecommitdiff
path: root/experiments
diff options
context:
space:
mode:
Diffstat (limited to 'experiments')
-rw-r--r--experiments/length-DirectEmit.csv64
-rw-r--r--experiments/pg-1.pngbin264089 -> 313785 bytes
-rw-r--r--experiments/pg-2.pngbin181594 -> 220244 bytes
-rw-r--r--experiments/pg-3.pngbin191333 -> 218770 bytes
-rw-r--r--experiments/pg-4.pngbin245170 -> 279719 bytes
-rw-r--r--experiments/pg-5.pngbin277654 -> 319282 bytes
-rw-r--r--experiments/pg-6.pngbin270483 -> 308166 bytes
-rw-r--r--experiments/pg-7.pngbin248684 -> 348546 bytes
-rw-r--r--experiments/pg-8.pngbin22182 -> 290076 bytes
-rw-r--r--experiments/pg-9.pngbin0 -> 66553 bytes
-rw-r--r--experiments/report.typ166
11 files changed, 194 insertions, 36 deletions
diff --git a/experiments/length-DirectEmit.csv b/experiments/length-DirectEmit.csv
new file mode 100644
index 0000000..f4fbd29
--- /dev/null
+++ b/experiments/length-DirectEmit.csv
@@ -0,0 +1,64 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+DirectEmit,length,0,0,0,insert,0,3987,12964,113
+DirectEmit,length,0,0,0,insert,1,3799,5930,34
+DirectEmit,length,0,0,0,insert,2,3779,5830,35
+DirectEmit,length,0,0,0,insert,3,3746,5886,35
+DirectEmit,length,0,0,0,insert,4,3756,5981,35
+DirectEmit,length,0,0,0,insert,5,3727,6013,35
+DirectEmit,length,0,0,0,insert,6,3738,5915,35
+DirectEmit,length,0,0,0,insert,7,3688,5865,35
+DirectEmit,length,0,0,0,insert,8,3665,5810,35
+DirectEmit,length,0,0,20,insert,0,3800,5952,35
+DirectEmit,length,0,0,20,insert,1,3763,5980,35
+DirectEmit,length,0,0,20,insert,2,3771,6016,35
+DirectEmit,length,0,0,20,insert,3,3815,14251,130
+DirectEmit,length,0,0,20,insert,4,3718,5861,34
+DirectEmit,length,0,0,20,insert,5,3742,6050,35
+DirectEmit,length,0,0,20,insert,6,3758,6007,35
+DirectEmit,length,0,0,20,insert,7,3758,6026,35
+DirectEmit,length,0,0,20,insert,8,3782,6024,35
+DirectEmit,length,0,0,40,insert,0,3802,6077,35
+DirectEmit,length,0,0,40,insert,1,3793,6013,35
+DirectEmit,length,0,0,40,insert,2,3855,6037,35
+DirectEmit,length,0,0,40,insert,3,3831,5958,35
+DirectEmit,length,0,0,40,insert,4,3816,5968,35
+DirectEmit,length,0,0,40,insert,5,3814,5974,35
+DirectEmit,length,0,0,40,insert,6,3924,14310,130
+DirectEmit,length,0,0,40,insert,7,3856,5972,34
+DirectEmit,length,0,0,40,insert,8,3854,6004,35
+DirectEmit,length,0,0,80,insert,0,3970,6116,35
+DirectEmit,length,0,0,80,insert,1,3883,6117,35
+DirectEmit,length,0,0,80,insert,2,3923,6127,35
+DirectEmit,length,0,0,80,insert,3,3928,6089,35
+DirectEmit,length,0,0,80,insert,4,3826,6113,35
+DirectEmit,length,0,0,80,insert,5,3825,6048,35
+DirectEmit,length,0,0,80,insert,6,3867,6074,35
+DirectEmit,length,0,0,80,insert,7,3860,5989,35
+DirectEmit,length,0,0,80,insert,8,3872,6023,35
+DirectEmit,length,0,0,160,insert,0,4063,6130,35
+DirectEmit,length,0,0,160,insert,1,4096,6196,35
+DirectEmit,length,0,0,160,insert,2,4043,6236,35
+DirectEmit,length,0,0,160,insert,3,4081,6154,35
+DirectEmit,length,0,0,160,insert,4,4062,6090,35
+DirectEmit,length,0,0,160,insert,5,4066,6154,35
+DirectEmit,length,0,0,160,insert,6,4029,6203,35
+DirectEmit,length,0,0,160,insert,7,3995,6181,35
+DirectEmit,length,0,0,160,insert,8,4032,6182,35
+DirectEmit,length,0,0,320,insert,0,3921,3921,6
+DirectEmit,length,0,0,320,insert,1,3863,3863,6
+DirectEmit,length,0,0,320,insert,2,3822,3823,6
+DirectEmit,length,0,0,320,insert,3,3949,3949,6
+DirectEmit,length,0,0,320,insert,4,3853,3853,6
+DirectEmit,length,0,0,320,insert,5,3909,3909,6
+DirectEmit,length,0,0,320,insert,6,3796,3797,6
+DirectEmit,length,0,0,320,insert,7,3863,3864,6
+DirectEmit,length,0,0,320,insert,8,3953,3954,6
+DirectEmit,length,0,0,640,insert,0,4169,4169,6
+DirectEmit,length,0,0,640,insert,1,4021,4021,6
+DirectEmit,length,0,0,640,insert,2,4070,4070,6
+DirectEmit,length,0,0,640,insert,3,3996,3996,6
+DirectEmit,length,0,0,640,insert,4,3940,3940,6
+DirectEmit,length,0,0,640,insert,5,4014,4014,6
+DirectEmit,length,0,0,640,insert,6,4029,4029,6
+DirectEmit,length,0,0,640,insert,7,3972,3972,6
+DirectEmit,length,0,0,640,insert,8,3981,3981,6
diff --git a/experiments/pg-1.png b/experiments/pg-1.png
index 0f48914..b808a6c 100644
--- a/experiments/pg-1.png
+++ b/experiments/pg-1.png
Binary files differ
diff --git a/experiments/pg-2.png b/experiments/pg-2.png
index e7a2e6b..e0ac1f2 100644
--- a/experiments/pg-2.png
+++ b/experiments/pg-2.png
Binary files differ
diff --git a/experiments/pg-3.png b/experiments/pg-3.png
index 31f7d9a..9dc00e0 100644
--- a/experiments/pg-3.png
+++ b/experiments/pg-3.png
Binary files differ
diff --git a/experiments/pg-4.png b/experiments/pg-4.png
index 450547e..ae3fab6 100644
--- a/experiments/pg-4.png
+++ b/experiments/pg-4.png
Binary files differ
diff --git a/experiments/pg-5.png b/experiments/pg-5.png
index 32caf02..9c3e3f2 100644
--- a/experiments/pg-5.png
+++ b/experiments/pg-5.png
Binary files differ
diff --git a/experiments/pg-6.png b/experiments/pg-6.png
index 87e34d4..d1264cc 100644
--- a/experiments/pg-6.png
+++ b/experiments/pg-6.png
Binary files differ
diff --git a/experiments/pg-7.png b/experiments/pg-7.png
index cb9c41e..0a08184 100644
--- a/experiments/pg-7.png
+++ b/experiments/pg-7.png
Binary files differ
diff --git a/experiments/pg-8.png b/experiments/pg-8.png
index 4f875a7..f86fbb9 100644
--- a/experiments/pg-8.png
+++ b/experiments/pg-8.png
Binary files differ
diff --git a/experiments/pg-9.png b/experiments/pg-9.png
new file mode 100644
index 0000000..3f8874a
--- /dev/null
+++ b/experiments/pg-9.png
Binary files differ
diff --git a/experiments/report.typ b/experiments/report.typ
index 2f27684..a15b162 100644
--- a/experiments/report.typ
+++ b/experiments/report.typ
@@ -47,7 +47,8 @@
rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") +
rows("length-ReleaseSmall-lineSpan.csv") + rows("length-ReleaseFast-lineSpan.csv") +
rows("length-RFast-grapheme.csv") + rows("length-RFast-print.csv") +
- rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv")
+ rows("length-RFast-shadow.csv") + rows("length-RFast-fastcmp.csv") +
+ rows("length-CPU360-final.csv") + rows("length-DirectEmit.csv")
)
#let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv")
@@ -106,16 +107,20 @@ implied, and it was only reachable by instrumenting the firmware; the fix that
source reading suggested was written, measured, and found to be worth 20% on a 19 MB
file and nothing at all on this board. Experiment 3.
-Both were then cut. A keystroke is 8.37 ms, from 16.99, and 9.48 ms at a
-160-character line, from 25.56 — the per-character term is 7.1 µs, from 54.3. Three of
-the four steps that did it came from *not asking Unicode about ASCII*, and the largest
-single one came from this repository's own `present` copying all 480 cells into vaxis
-every frame whether or not any had changed. Experiment 4.
+Both were then cut. A keystroke is *3.74 ms, from 16.99*, and 4.06 ms at a
+160-character line, from 25.56 — the per-character term is 2.0 µs, from 54.3. Three of
+the six steps that did it came from *not asking Unicode about ASCII*; the largest single
+one came from this repository's own `present` copying all 480 cells into vaxis every
+frame whether or not any had changed; and the last two were a clock that was only ever a
+divider away, and a renderer that stopped asking vaxis to recompute a diff `present` had
+already done. Experiment 4.
-The target was 4 ms and it is not met: compute is 6.42 ms against a budget of about
-2.05, so 3.1× remains. What is different from the start is that every microsecond of it
-is located — our walk of the grid, vaxis's own diff, and pardes rebuilding a whole
-Surface per keystroke — and that the two largest are not alternatives to each other.
+The 4 ms target is met. Two findings on the way there are worth more than the number.
+A 4× clock bought 2.6× because the grid walk is bounded by memory bandwidth, not by the
+core. And the direct renderer, with less computation and a quarter of the bytes, first
+measured *slower* — because the CH340 bridging the board to the host forwards a bulk
+packet only when it is full, so a reply too small to fill one waits about a millisecond
+for a timer. The frame has a minimum size and it belongs to the transport.
= The instrument
@@ -578,7 +583,8 @@ tree-sitter, which the board never runs.
table.header([step], [fixed cost], [per char], [at 160 chars], [vs start]),
table.hline(stroke: 0.5pt),
..(
- ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow", "RFast-fastcmp")
+ ("ReleaseSmall", "ReleaseFast", "RFast-grapheme", "RFast-print", "RFast-shadow",
+ "RFast-fastcmp", "CPU360-final", "DirectEmit")
).map(b => {
let xs = lengths.map(l => l * 1.0)
let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l))
@@ -596,8 +602,9 @@ tree-sitter, which the board never runs.
table.hline(),
),
caption: [Each row is a separate firmware measured on the die, 5 document lengths ×
- 7 trials. The per-character column and the fixed column move for different reasons
- and are worth reading separately.],
+ 7--9 trials. The per-character column and the fixed column move for different reasons
+ and are worth reading separately. The last two rows are the clock raise and the
+ renderer, and neither is a source optimisation in the sense the four above it are.],
)
*Not asking Unicode about ASCII* accounts for the per-character column. Three fast
@@ -643,30 +650,111 @@ the editor's heap it was enough to make `vx.resize` fail.
== How a rendering change was made safe to believe
A latency benchmark cannot see a corrupted screen, and `src/p4.zig` compiles for this
-board alone — no host suite covers it. So the optimisation carries a comptime `shadow_grid`
-switch whose `false` arm is the old behaviour, clear and write everything, and the two
-firmwares were each run against the same 19-step workload on the die: inserts, deletes,
-motions that move the modified-marker, a line outgrowing the viewport, backspaces that
-shrink it. The screens were reconstructed from the wire and compared.
+board alone — no host suite covers it. So each optimisation carries a comptime switch
+whose `false` arm is the previous behaviour — `shadow_grid` for the incremental walk,
+`direct_emit` for the renderer — and the firmwares were each run against the same
+18-step workload on the die: inserts, deletes, motions that move the modified-marker, a
+line outgrowing the viewport, backspaces that shrink it. The screens were reconstructed
+from the wire and compared.
-The first version of that comparison tracked characters only and would have passed a
-colour regression in silence; it now hashes the SGR state of every cell as well. Both
-text and per-row style hash are identical between the two paths. On the host side,
-where the shared code does live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0
-mismatches, `hxparity` 561 cases with 0 mismatches.
+That comparison had to be corrected twice, and the second time it had already produced a
+false result. Tracking characters only would have passed a colour regression in silence,
+so it began hashing SGR as well — but it hashed the escape *parameters* applied to each
+cell, which is a history and not a state. Raising the clock changed which keystrokes
+shared a frame, the same final colours were reached by a different route, and the
+verifier *reported a difference on the die that did not exist*. It now decodes SGR into
+resolved state — foreground, background, attribute set per cell — and two screens compare
+equal when every cell holds the same character and the same resolved style, whatever
+sequence of escapes the emitter chose to get there. Without that correction the
+from-scratch renderer could not have been verified at all: it reaches identical screens
+by deliberately different escapes.
-== Where it stopped, and what 4 ms still needs
+All three pairs are identical under it: shadow grid against full repaint, 90 MHz against
+360, and direct emission against vaxis. On the host side, where the shared code does
+live: `snap` 95/95 scripts, `hxdiff` 481 cases with 0 mismatches, `hxparity` 561 cases
+with 0 mismatches.
-A keystroke is 8.37 ms, from 16.99 — and 9.48 ms at a 160-character line, from 25.56.
-Compute is 6.42 ms against a budget of 2.05, so 3.1× remains, and all of it is
-measured rather than guessed: our walk of the grid (979 µs), vaxis's own diff (2 455 µs),
-and pardes rebuilding every cell of the Surface every frame (≈2 000 µs).
+== The clock was a divider, not a design
-The two large items are not alternatives to each other. vaxis's diff is redundant work
-— `present` already computes exactly which cells moved, so emitting ANSI straight from
-the shadow grid would remove most of it — but even a perfect renderer leaves pardes
-rebuilding a whole frame per keystroke, and 4 ms at 40×12 needs that too. The
-remaining path is therefore two changes, not one, and the second is architectural.
+The board had been executing at 90 MHz for every measurement above, because the stock
+second-stage bootloader is built for it. The CPLL was *already* running at 360 MHz — 90
+is exactly a quarter of it — so reaching 360 is four divider writes and a ROM call, with
+no PLL to enable and no voltage step to sequence against. Verified against the systimer,
+which is clocked from the crystal and therefore independent: 90 000 kHz before,
+359 991 kHz after. Nothing else moves, and that is what makes it safe from a live
+console: the UART0 baud clock comes from the crystal, the systimer from the crystal, and
+the flash interface from a different PLL entirely.
+
+Compute fell 6.42 → 2.44 ms: *2.6× for a 4× clock*, and the shortfall is the
+interesting part. The grid walk reads 27 KB per frame and takes 226 µs at 360 MHz, about
+six cycles a byte, so it is bounded by L2MEM bandwidth rather than by the core. A
+prediction made before the measurement, and held.
+
+== The renderer, and a minimum frame that belongs to the USB bridge
+
+vaxis's diff was redundant: `present` has already worked out exactly which cells moved,
+so vaxis was being told the answer and then computing it again against its own copy.
+Emitting the escapes directly — one absolute cursor position per run of changed cells,
+absolute SGR rather than a delta, a hand-rolled formatter for the two digits that
+sequence needs — took the board's render from 1 362 µs to 779 µs and the reply from 81
+bytes to 21.
+
+And it measured *slower*: 4.72 ms against 4.37. Fewer bytes, less computation, worse
+round trip is the shape of a wrong model, so the search moved off the firmware.
+
+It is the USB bridge. The board reaches the host through a CH340, a full-speed part
+whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet
+is *full*. A 21-byte frame never fills one, so it waits in the bridge for an internal
+timer to give up on more — worth about a millisecond, a quarter of the entire budget.
+Routing through vaxis had only looked competitive because 81-byte frames fill a packet
+by accident.
+
+#figure(
+ table(
+ columns: (auto, auto, auto, 1fr),
+ align: (right, right, right, left),
+ stroke: none,
+ table.hline(),
+ table.header([frame], [min], [median], []),
+ table.hline(stroke: 0.5pt),
+ [21 B], [3 843 µs], [4 817 µs], [direct, never fills a packet],
+ [49 B], [3 719 µs], [3 814 µs], [direct, padded past the boundary],
+ [81 B], [4 373 µs], [4 475 µs], [through vaxis, fills one by accident],
+ table.hline(),
+ ),
+ caption: [Identical board cost and a byte-identical screen across all three; only the
+ size of the reply differs. Note the *minimum*: the 21-byte frame's floor is already
+ 530 µs below vaxis's, which is exactly the computation that was saved. Only the median
+ was hostage to the bridge's timer.],
+)
+
+So the frame has a minimum size and it is the transport's, not the terminal's. The
+emitter pays it, padding with repeated absolute cursor positioning: idempotent, already
+the sequence a frame ends on, incapable of altering a cell. This is the bargain an
+Ethernet runt frame makes — the medium has a minimum and the sender pays it — and it is
+a real trade rather than a free one, since the filler is wire time that delays a later
+frame. It applies only while the frame is small, which is when there is wire to spare.
+
+A second instrument came out of this. The original bench sends keystrokes on a fixed
+cadence, which locks the send phase to the host's 1 ms USB frame clock and makes the
+round trip a staircase in board time — where a real saving can present as a regression.
+Sleeping a uniform random 0–2 ms before each keystroke decorrelates them. That was not
+what was happening here, but it had to be ruled out before the CH340 could be believed.
+
+== Where it stopped
+
+A keystroke is *3.74 ms*, from 16.99 — and 4.06 ms at a 160-character line, from 25.56.
+The target was 4 ms and it is met, with the phase-randomised instrument agreeing over 60
+trials: median 3 829 µs, minimum 3 722, ninetieth percentile 3 930. The reply is 35
+bytes, down from 81.
+
+What remains is no longer dominated by anything this repository owns. Of the round trip,
+roughly 2.9 ms is host, USB and wire — the floor measured independently against the
+protocol responder, and now partly explained by the bridge's packet granularity — and
+about 0.84 ms is board: 552 µs of pardes rebuilding every cell of the Surface every
+frame, 223 µs walking the grid, 60 µs parsing and editing. The Surface rebuild is the
+one architectural item left, and it is worth 552 µs against a 2.9 ms floor, so the next
+order of magnitude is not in the firmware at all.
#figure(
table(
@@ -717,7 +805,10 @@ remaining path is therefore two changes, not one, and the second is architectura
reordered it. Rows 2 and the last row were established by measurement; rows 3--5
are derived from measured quantities and cited source. The last row is listed
unranked because it is already applied and, on *this* target, buys nothing --
- which is precisely why it is worth recording.],
+ which is precisely why it is worth recording. Kept as the prediction it was:
+ Experiment 4 has since applied rows 1 and 2 and *abandoned row 3*, because raising
+ the line rate moves `settle` and leaves the round trip alone --- and the bridge
+ finding above is the reason row 3 looked attractive in the first place.],
)
= Threats to validity
@@ -729,8 +820,11 @@ is a no-op, so trials after the first measured nothing, and the genuine
full-repaint figure quoted throughout (1 392 bytes, 150 ms to settle) comes from a
single first resize rather than from a median.
-Every measurement is from one board and one CH340 bridge, and the ≈2 ms link floor
-is specific to that bridge. Round trip is time to first byte and therefore says
+Every measurement is from one board and one CH340 bridge, and the link floor --- ≈2 ms
+as measured here, ≈2.9 ms once the bridge's packet granularity is accounted for --- is
+specific to that bridge; a device that forwards short packets promptly would show a
+different floor and would not reward the frame padding of Experiment 4 at all.
+Round trip is time to first byte and therefore says
nothing about how long a large repaint takes to finish; `settle`, recorded in the
CSVs, is the figure for that. Both builds were measured in a single session each,
so slow drift — temperature, or the host's own scheduling — would appear as a