summaryrefslogtreecommitdiff
path: root/experiments
diff options
context:
space:
mode:
Diffstat (limited to 'experiments')
-rw-r--r--experiments/attribution.csv6
-rw-r--r--experiments/length-ReleaseSmall-lineSpan.csv36
-rw-r--r--experiments/report.typ150
3 files changed, 176 insertions, 16 deletions
diff --git a/experiments/attribution.csv b/experiments/attribution.csv
new file mode 100644
index 0000000..90e62ce
--- /dev/null
+++ b/experiments/attribution.csv
@@ -0,0 +1,6 @@
+label,experiment,chars,input_cy,render_cy,input_us,render_us
+ReleaseSmall-prof,attribution,1,19800000,1332360000,220,14804
+ReleaseSmall-prof,attribution,40,17550000,1380690000,195,15341
+ReleaseSmall-prof,attribution,80,19080000,1540260000,212,17114
+ReleaseSmall-prof,attribution,160,20610000,1859130000,229,20657
+ReleaseSmall-prof,attribution,240,22500000,2215350000,250,24615
diff --git a/experiments/length-ReleaseSmall-lineSpan.csv b/experiments/length-ReleaseSmall-lineSpan.csv
new file mode 100644
index 0000000..d63a815
--- /dev/null
+++ b/experiments/length-ReleaseSmall-lineSpan.csv
@@ -0,0 +1,36 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+ReleaseSmall-lineSpan,length,0,0,0,insert,0,17217,28585,141
+ReleaseSmall-lineSpan,length,0,0,0,insert,1,16969,23098,80
+ReleaseSmall-lineSpan,length,0,0,0,insert,2,17008,23148,81
+ReleaseSmall-lineSpan,length,0,0,0,insert,3,17132,23282,81
+ReleaseSmall-lineSpan,length,0,0,0,insert,4,17087,23328,81
+ReleaseSmall-lineSpan,length,0,0,0,insert,5,17182,23308,81
+ReleaseSmall-lineSpan,length,0,0,0,insert,6,17210,23331,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,0,17895,24074,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,1,17884,24065,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,2,17926,24139,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,3,18017,24223,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,4,17978,24163,81
+ReleaseSmall-lineSpan,length,0,0,20,insert,5,18474,31354,158
+ReleaseSmall-lineSpan,length,0,0,20,insert,6,18422,24604,80
+ReleaseSmall-lineSpan,length,0,0,40,insert,0,19185,25443,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,1,19201,25321,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,2,19269,25491,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,3,19285,25415,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,4,19403,25647,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,5,19418,25606,81
+ReleaseSmall-lineSpan,length,0,0,40,insert,6,19360,25597,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,0,21696,27929,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,1,21665,27895,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,2,21611,27873,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,3,21675,27909,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,4,21775,27912,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,5,21867,27997,81
+ReleaseSmall-lineSpan,length,0,0,80,insert,6,21802,28024,81
+ReleaseSmall-lineSpan,length,0,0,160,insert,0,25434,31562,81
+ReleaseSmall-lineSpan,length,0,0,160,insert,1,25476,31624,81
+ReleaseSmall-lineSpan,length,0,0,160,insert,2,25470,31572,81
+ReleaseSmall-lineSpan,length,0,0,160,insert,3,25632,31867,81
+ReleaseSmall-lineSpan,length,0,0,160,insert,4,26049,38935,158
+ReleaseSmall-lineSpan,length,0,0,160,insert,5,26056,32189,80
+ReleaseSmall-lineSpan,length,0,0,160,insert,6,25978,32263,81
diff --git a/experiments/report.typ b/experiments/report.typ
index 1d859b8..8ed2b07 100644
--- a/experiments/report.typ
+++ b/experiments/report.typ
@@ -83,13 +83,21 @@ measured here:
*A cost proportional to the document*, at
#calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 1) µs
- per character already in the line, per keystroke. This is an $O(n)$ edit path,
- and it is what makes the editor feel worse the more you have written.
+ per character already in the line, per keystroke, which is what makes the editor
+ feel worse the more you have written.
]
+*Both live entirely in the renderer.* Timed separately on the die, parsing the
+keystroke and applying the edit takes a flat ≈220 µs regardless of document size —
+1.5% of the total — while `render` carries the whole ≈15 ms floor and every
+microsecond of the slope. That result contradicted the mechanism the source reading
+implied, and it was only reachable by instrumenting the firmware; the fix that
+source reading suggested was written, measured, and found to be worth 20% on a 19 MB
+file and nothing at all on this board. Experiment 3.
+
One build-flag change — compiling the editor object `ReleaseFast` instead of
`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost,
-for 35% more flash. Nothing else measured here comes close to that ratio.
+for 35% more flash. It remains the best ratio measured here.
= The instrument
@@ -315,17 +323,18 @@ measured quantity is unchanged: the round trip of one further inserted character
}).flatten(),
table.hline(),
),
- caption: [Fitted from the medians. The slope is the interesting column: it is a
- per-keystroke re-copy of the whole buffer.],
+ caption: [Fitted from the medians. The slope is the interesting column. Its
+ *mechanism* is settled by Experiment 3, not by this fit.],
)
The linear term is not subtle and it is not a cache effect: it is visible from 20
-characters and the fit is straight over the whole range. The likely mechanism is a
-full-buffer allocate-and-copy for every edit, which is what the source does; at
+characters and the fit is straight over the whole range. At
#calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs
-per character on a ≈90 MHz core, though, it is far more work than a copy alone —
-roughly 5 000 cycles per character of buffer per keystroke — so the copy is
-accompanied by at least one further full pass. *H3 is confirmed.*
+per character on a ≈90 MHz core — roughly 5 000 cycles for every character already
+in the line — it is far too much work to be a `memcpy`, so something is making a
+substantial pass per character. *H3 is confirmed as an observation.* What that pass
+actually is turned out not to be what the source reading suggested, which is
+Experiment 3.
The two series also settle H2. `ReleaseFast` lowers the fixed cost by
#calc.round(
@@ -340,6 +349,109 @@ The two series also settle H2. `ReleaseFast` lowers the fixed cost by
35% more of a 1 536 000-byte partition — affordable, and the only change measured
here that improves both terms at once. *H2 is confirmed.*
+= Experiment 3 --- it is all in the renderer, and the obvious fix was wrong
+
+Everything above measures a keystroke from the host, which cannot see *what* the
+firmware spent the time on. Reading the source suggested an answer: the edit path
+builds each new document with `modal.spliceAlloc`, a fresh allocation and a copy of
+the whole buffer, and `insertAt` called `modal.lineCount` — `std.mem.count` over
+every byte — *twice*, merely to clamp a row. That is three whole-document passes
+before a single character can be inserted, which fits a linear slope exactly.
+
+It is also, on this board, almost entirely irrelevant. Timing the two phases
+separately on the die (`-Dprof`, two reads of the cycle counter around
+`pardes_p4_input` and `pardes_p4_render`) gives:
+
+#figure(
+ {
+ let a = ()
+ for r in csv("attribution.csv") {
+ if r.at(0) == "label" { continue }
+ a.push((chars: int(r.at(2)), input: int(r.at(5)), render: int(r.at(6))))
+ }
+ table(
+ columns: (auto, auto, auto, auto),
+ align: (right, right, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([characters in line], [input: parse + edit], [render], [render share]),
+ table.hline(stroke: 0.5pt),
+ ..a.map(r => (
+ [#r.chars],
+ [#r.input µs],
+ [*#r.render µs*],
+ [#calc.round(r.render / (r.input + r.render) * 100, digits: 1)%],
+ )).flatten(),
+ table.hline(),
+ )
+ },
+ caption: [On-board cycle counts, `ReleaseSmall`. Input is flat; render carries
+ both the fixed cost and the entire slope.],
+)
+
+*Input is flat at ≈220 µs and does not grow with the document at all.* The fixed
+≈15 ms and every microsecond of the per-character slope are inside
+`pardes_p4_render`. The edit path — the allocation, the copy, the double line count
+— is 1.5% of a keystroke and could be made free without anyone noticing.
+
+This was worth proving rather than assuming, because the fix implied by the source
+reading was written and measured. `modal.insertAt` now takes one *bounded* scan
+through a new `modal.lineSpan`, which stops at the row it wants instead of counting
+the whole document, and pays for a full count only on the rare clamping path where
+the cursor is past the end. On the host harness (`zig build perf`, which drives the
+same core over 61 KB to 19 MB fixtures) that is a real win, reproduced over three
+independent runs:
+
+#figure(
+ table(
+ columns: (auto, auto, auto, auto),
+ align: (left, right, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([fixture], [1 000 lines], [50 000 lines], [300 000 lines]),
+ table.hline(stroke: 0.5pt),
+ [`edit-char`, before], [730 µs], [2 996 µs], [15 030 µs],
+ [`edit-char`, after], [721 µs], [2 414 µs], [10 955 µs],
+ [ratio, three runs], [0.98×], [0.79–0.82×], [0.78–0.83×],
+ table.hline(),
+ ),
+ caption: [Host harness, 25 samples per cell. Every untouched operation stayed at
+ 1.00×, which is stronger evidence than any single cell.],
+)
+
+And on the board it changed *nothing*: the slope was
+#calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 1) µs
+per character before and 54.0 µs after, a ratio of 1.00. That is not a
+contradiction, it is the same fact seen twice: the removed passes are $O(#h(0.1em)$document$)$,
+and this board's document is a few hundred *bytes*, so two scans of it cost nothing
+worth measuring. The identical change is worth 20% on a 19 MB file and 0% on a
+240-character one.
+
+The lesson is the one the instrument exists to enforce. A plausible mechanism, read
+off the source and consistent with the shape of the data, was wrong about where the
+time went — and it took a measurement *inside* the firmware to say so. The renderer
+is the target; the next question is what in it is proportional to the line, and the
+position experiment already narrows that: cost follows the cursor's column as well
+as the document's size, which is the signature of a walk from the start of a line.
+
+#figure(
+ table(
+ columns: (auto, auto, auto),
+ align: (left, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([insert position in a fixed 320-character line], [round trip], [bytes emitted]),
+ table.hline(stroke: 0.5pt),
+ [column 320 (end)], [33.9 ms], [28 B],
+ [column 0 (start)], [26.0 ms], [81 B],
+ [column 320 again], [33.7 ms], [28 B],
+ table.hline(),
+ ),
+ caption: [Same document throughout; only the cursor moved. The end of the line
+ costs 7.8 ms more than the start while emitting *a third* as many bytes — output
+ size and latency are not merely uncorrelated here, they are inverted.],
+)
+
= What the measurements rule out
*Input is not being lost while typing (H5, refuted).* At every rate from 6 to 100
@@ -399,21 +511,27 @@ waits for that render.
table.hline(),
table.header([], [change], [effect, measured or derived]),
table.hline(stroke: 0.5pt),
- [1], [Make the edit path stop copying the whole buffer per keystroke.],
- [removes the #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs/char term entirely],
+ [1], [Find what in `render` is proportional to the line, and to the cursor's
+ column. Experiment 3 puts 98.5% of a keystroke there; the position table
+ narrows it to a walk from the start of a line.],
+ [target: the #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs/char slope *and* most of the ≈15 ms floor],
[2], [Build the editor object `ReleaseFast`.],
- [measured: −13% fixed, −36% per character, +35% flash],
+ [measured on the die: −13% fixed, −36% per character, +35% flash],
[3], [Raise UART0 to 921600.],
[derived: repaints #calc.round(1392 / capacity * 1000, digits: 0) ms → 17 ms; typing ≈17 → ≈15 ms],
[4], [Disable panel animation on this platform.],
[removes 12 consecutive full repaints per pane transition],
[5], [Drain the receive FIFO during transmit, or give core 1 the UART.],
[closes the only window in which input is silently lost],
+ [--], [Stop the edit path re-scanning the document (`modal.lineSpan`, done).],
+ [measured: 20% off `edit-char` at 19 MB, 0% on this board],
table.hline(),
),
- caption: [Ordered by benefit per line of code changed. Only rows 2 and the
- measurements underlying row 1 were established on the die; rows 3--5 are derived
- from measured quantities and cited source.],
+ caption: [Ordered by benefit per line of code changed, after Experiment 3
+ reordered it. Rows 2 and the last row were established by measurement; rows 3--5
+ are derived from measured quantities and cited source. The last row is listed
+ unranked because it is already applied and, on *this* target, buys nothing --
+ which is precisely why it is worth recording.],
)
= Threats to validity