summaryrefslogtreecommitdiff
path: root/experiments/report.typ
Commit message (Collapse)AuthorAge
* Report: Experiment 5, the three bugs, and a resize that fixed itselfGabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | Adds the geometry curve - 0.87 us per cell, measured across nine grids from 40x12 to 140x42 - and the two ways it ends: 200x60 refused by the linker because '.bss' cannot hold the shadow grid, 160x48 linking and then trapping. Neither is a heap problem any more, which is the finding: the heap has 336 KB spare while '.bss' runs out, because vaxis's two unread grids stopped being allocated. Also the three bugs the latency work walked straight past, none of which any number in this document could see: keystrokes lost while the transmitter was full, every escape sequence shredded because a lone ESC arrives alone on a 115200 line, and a mouse that was never enabled. The middle one had been there since the port began and explains the other two - a click arriving as ten key presses, with the '0' among them moving the cursor to column zero, which is what made it look like a coordinate bug. And Table 15 changes from 'not fixed' to 'fixed, and not on purpose'. The vx.resize failure that three experiments characterised and left alone was an allocation of vaxis's two grids; the emitter stopped needing them, so the path that failed is gone rather than repaired. Verified on the die - an in-band report re-lays the screen from 56 columns to 30 and back. Recorded as a consequence rather than as a fix, because nobody set out to repair it. Eleven pages.
* Report: the pad target swept, and the byte-at-a-time read that was hiding in stdGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target, and the closing section now reports it the way it should be reported: over 80 phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document length rather than as one intercept. Three things this adds that are findings rather than steps. ## The pad target, swept Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the cursor-positioning frame 21 -> 49 bytes made it a millisecond faster, padding the cursor-hiding frame 6 -> 36 made it slower - so the target was swept as the only variable. 0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either: it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what confirms the descriptor is the thing to believe. ## The largest read in the firmware ran a byte at a time The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame, comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte, about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32 at a time: 223 -> 66 us, 0.95 cycles per byte, at every document length. Recorded alongside it are the two candidates that were measured and REVERTED, which is the more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in Surface.set removed 52 million instructions per host run and zero cycles on either host or board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are cheap; it was the reads that were slow. ## Where it stopped, honestly The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12 viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and 4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are now in the table with a "cells changed: none" row under them, because a round trip for an edit that displays nothing is worth reporting as exactly that and not as a failure to hit a number. Also corrected: the ranked-recommendations table said the UART0 raise was abandoned because a higher line rate moves settle and leaves the round trip alone. That reasoning assumed the first byte reaches the host as soon as it is sent, and the bridge finding says otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that one 115200-baud line is the premise of this port rather than a free variable, and meeting the target by changing the link would answer a different question. Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones.
* Report: the clock, the renderer, and a minimum frame owned by the USB bridgeGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Brings the document up to the end of Experiment 4. Two rows on the progression table - 360 MHz and direct emission - plus the sections behind them, the corrected verifier, and a summary that no longer says the target was missed. It is met: 3.74 ms from 16.99. The two findings worth more than the number, and both are written up as findings rather than as steps: A 4x clock bought 2.6x. The grid walk reads 27 KB a frame at about six cycles a byte, so it is bounded by L2MEM bandwidth and does not care how fast the core runs. Predicted before the measurement. The direct renderer measured SLOWER at first - less computation, a quarter of the bytes, worse round trip - because the CH340 forwards a bulk IN packet only when the packet is full, and a 21-byte frame does not fill 32. It waits about a millisecond for a timer. So the frame has a minimum size and it belongs to the transport, not the terminal; the emitter pads to it with repeated cursor positioning. The table of 21/49/81-byte frames is in the report because the MINIMUM column is the tell: the small frame's floor was already 530 us below vaxis's, exactly the compute saved, and only the median was hostage. Also corrected: the verifier section. It described hashing SGR parameters per cell, which is a history rather than a state, and that version reported a difference on the die that did not exist - a faster board split the same keystrokes across different frames and reached the same colours by another route. The document now says what it does instead, and why the from-scratch renderer could not have been verified without the fix. Table 13, the ranked recommendations from Experiment 3, is kept as the prediction it was with a note on what has since been applied and that row 3 - raise the line rate - was abandoned. The threats section no longer states the link floor as a single number, since the bridge's packet granularity is now part of it and is specific to this bridge. Nine pages. Every figure still reads the raw per-trial CSVs, including the two new ones.
* Report: the optimisation campaign, and the 3.1x that is leftGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | Experiment 4 records what each cut was worth, measured on the die at every step, and the stage table that made the cuts findable at all. Also two corrections to the document itself: The data plumbing loaded four of the eight datasets, so the progression table would have been computed from a subset. It now loads all eight. A continuation line beginning with `+` is list markup to Typst, so the CSV file names were rendering as a numbered item at the top of page one. One expression, one line. The summary no longer ends on the ReleaseFast flag as the best available change; it ends on the measured 16.99 -> 8.37 ms and on the fact that the target is not met, with the remaining 6.42 ms of compute broken into the three items it actually consists of. Saying 'not met' in the summary matters more than the table: the number that was missed is the one a reader should see first.
* Measure all four configurations; the code fix is worth nothing here, the ↵Gabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | flag is worth 21% The obvious question after the last commit was what it did on the actual target. The answer is nothing, and the useful part is that the same run says what DOES work. Four builds on the die, 5 document lengths x 7 trials each: configuration fixed per char at 160 vs base ReleaseSmall 16.99 ms 54.3 us 25.56 ms 1.00x ReleaseSmall+lineSpan 17.10 ms 54.0 us 25.63 ms 1.00x ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x ReleaseFast+lineSpan 14.88 ms 34.2 us 20.25 ms 0.79x The edit-path change is invisible in BOTH modes, and ReleaseFast+lineSpan is indistinguishable from ReleaseFast alone. The optimisation mode is the whole of the difference, which is what Experiment 3 predicted: the edit path is 220 us of a 17 ms keystroke, so making it cheaper cannot move the total, while the mode makes the RENDERER faster and the renderer is where the time is. `experiments/report.typ` gains that table and the reason the four-way comparison exists rather than four indistinguishable lines on figure 1. The board is now flashed with the default build, which as of the pin next door means ReleaseFast: 809,552 B of the 1,536,000 B partition, 53%.
* Bound the edit path's document scans, and find out they were never the problemGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | The board half: the -Dprof attribution that overturned the conclusion, its data, and the report correction. `-Dprof` adds two cycle-counter reads around `pardes_p4_input` and `pardes_p4_render` and prints both. Off by default: it puts a line on the wire per frame, which is the very resource being measured, so it answers "where did the 15 ms go" and not "how fast is it". It answered. Input is flat at ~220 us regardless of document size - 1.5% of a keystroke - and the entire ~15 ms floor plus every microsecond of the per-character slope live in `render`. The edit-path fix that the source reading implied (committed next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5 conditions x 7 trials, exactly 0% on this board. `experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's mechanism claim was wrong, says so, and carries the disproof beside it. The ranked recommendations are reordered with the renderer at #1. Also here: `--sweep position` in p4-bench, which holds the document fixed at one 320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28 bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on this board - the signature of a walk from the start of a line.
* A measuring instrument, and what it says about where the latency goesGabriel Schneider2026-08-25
"Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result.