summaryrefslogtreecommitdiff
path: root/experiments
Commit message (Collapse)AuthorAge
* Measure inside a frame, and cut a keystroke from 17.0 ms to 8.9 msGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The board half of the run to 4 ms: the instrumentation that found the cost, and the measurements that judged each change. `-Dprof` grew two things. It now renders a SECOND time with nothing changed, which splits a frame's cost cleanly: whatever the second render still costs is the price of walking and diffing the whole editor state, and the difference between the two is the price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a keystroke was work done regardless of what the keystroke did. It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three stages in CPU cycles. That is what turned "render is slow" into an address: stage before after copy Surface -> vaxis 6 750 us 1 450 us vaxis diff + emit 2 460 us 2 455 us push into the UART 1 us 1 us (pardes's own Surface build) ~2 600 us ~2 000 us The copy was 57% of a keystroke and it was in this repo's own `present`, not in pardes and not in vaxis. Measured on the die, five document lengths x seven trials per configuration: configuration fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x ## Where the remaining 4.9 ms is, and why the goal is not met Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a further 3.4x. The three remaining pieces are known and measured - our walk of the grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke. Raising the baud does NOT help this number, and that is worth writing down because it is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms and leaves the round trip where it is. ## A bug found on the way `vx.resize` fails on this board. A runtime geometry change takes its allocation failure path, restores the previous size and returns: 80 bytes go out where 1,392 should, and the screen keeps its old shape. Reproduced with the shadow grid compiled out, so it predates it. The board has one geometry per session. That is also why the shadow grid's correctness test compares two firmwares rather than forcing a repaint with a resize - the forcing mechanism does not work here. The reference path and the incremental path were each run against the same 19-step workload and their reconstructed screens are byte-identical.
* Measure all four configurations; the code fix is worth nothing here, the ↵Gabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | flag is worth 21% The obvious question after the last commit was what it did on the actual target. The answer is nothing, and the useful part is that the same run says what DOES work. Four builds on the die, 5 document lengths x 7 trials each: configuration fixed per char at 160 vs base ReleaseSmall 16.99 ms 54.3 us 25.56 ms 1.00x ReleaseSmall+lineSpan 17.10 ms 54.0 us 25.63 ms 1.00x ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x ReleaseFast+lineSpan 14.88 ms 34.2 us 20.25 ms 0.79x The edit-path change is invisible in BOTH modes, and ReleaseFast+lineSpan is indistinguishable from ReleaseFast alone. The optimisation mode is the whole of the difference, which is what Experiment 3 predicted: the edit path is 220 us of a 17 ms keystroke, so making it cheaper cannot move the total, while the mode makes the RENDERER faster and the renderer is where the time is. `experiments/report.typ` gains that table and the reason the four-way comparison exists rather than four indistinguishable lines on figure 1. The board is now flashed with the default build, which as of the pin next door means ReleaseFast: 809,552 B of the 1,536,000 B partition, 53%.
* Bound the edit path's document scans, and find out they were never the problemGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | The board half: the -Dprof attribution that overturned the conclusion, its data, and the report correction. `-Dprof` adds two cycle-counter reads around `pardes_p4_input` and `pardes_p4_render` and prints both. Off by default: it puts a line on the wire per frame, which is the very resource being measured, so it answers "where did the 15 ms go" and not "how fast is it". It answered. Input is flat at ~220 us regardless of document size - 1.5% of a keystroke - and the entire ~15 ms floor plus every microsecond of the per-character slope live in `render`. The edit-path fix that the source reading implied (committed next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5 conditions x 7 trials, exactly 0% on this board. `experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's mechanism claim was wrong, says so, and carries the disproof beside it. The ranked recommendations are reordered with the renderer at #1. Also here: `--sweep position` in p4-bench, which holds the document fixed at one 320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28 bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on this board - the signature of a walk from the start of a line.
* A measuring instrument, and what it says about where the latency goesGabriel Schneider2026-08-25
"Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result.