summaryrefslogtreecommitdiff
path: root/build/snap.zig
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-25 22:10:36 -0300
committerGabriel Schneider <[email protected]>2026-08-25 22:39:09 -0300
commit4ab24352873ed7bf8db93ef6bfec36a34b0357e8 (patch)
tree810024497aee863be550dcad8c221b3c23e7e89d /build/snap.zig
parent97f9ba329081c23d3ac9c9a46338e0178e692742 (diff)
downloadpardes-4ab24352873ed7bf8db93ef6bfec36a34b0357e8.tar.gz
pardes-4ab24352873ed7bf8db93ef6bfec36a34b0357e8.zip
The frame diff was comparing byte at a time; compare words, and pay the bridge on every frame
Two findings, both in the P4 shell's own `present`. ## std.mem.eql was the largest read in the firmware, one byte at a time The shadow-grid diff compares each row against the previous frame: two 13 KB streams, every frame, and by far the biggest memory access the firmware makes. It measured 3.2 cycles per byte, which is about four times what word-wide loads need - the shape of a byte-at-a-time loop, and `std.mem.eql` is what it was. `sameBytes` compares a `u32` at a time and falls back to the byte loop when the spans are not aligned for it. The alignment test has to be a RUNTIME one because `Cell` is all `u8` fields and therefore has alignment 1: whether a row begins on a word boundary is a property of whoever allocated the Surface, not of the type. A row is 40 cells of 26 bytes, divisible by four, so an aligned base makes every row aligned. The answer is bit-for-bit the same - this is still exact byte equality - so it keeps the property the whole diff rests on: byte equality implies visual equality, so the diff can never claim two different cells are the same. Measured on the die: the grid walk 223 -> 66 us, 0.95 cycles per byte. 157 us off every keystroke at every document length, and the single largest win since the clock raise. ## A frame that only hides the cursor still has to fill a USB packet The padding added for the bridge's 32-byte bulk-IN packet covered the branch that positions the cursor and not the branch that hides it. A frame that only hid the cursor was six bytes and waited out the bridge's timer. Hiding an already-hidden cursor is as idempotent as positioning it twice, so it pads the same way. The packet size is no longer inferred from an experiment either: 32 is `wMaxPacketSize` of endpoint 0x82 as the device reports it, and the sweep over pad targets confirms what it implies - 0 and 16 sit at 4.7-5.1 ms, while 32, 48 and 64 all sit at 3.6-3.8 ms. Crossing the boundary is worth about 950 us; going past it buys nothing. ## Result length 0 20 40 80 160 320 640 chars RTT 3602 3624 3638 3790 3868 4026 4192 us Fixed cost 3652 us against a 4 ms target, from 16.99 ms where this started. A phase-randomised instrument agrees over 80 trials: median 3687 us, minimum 3571, maximum 3912 - every trial under 4 ms. The two lengths still above 4 ms are the ones where the line has outgrown the viewport, so the cursor is off screen and the keystroke changes NOTHING: the frame is 36 bytes of cursor-hide and padding, zero cells changed, while pardes still rebuilds all 480 cells of the Surface for 858-985 us. That is the one architectural item left and it is not a micro -optimisation: nothing in this repository can avoid work pardes has already done. Verified: screen byte-identical to the vaxis reference on the 18-step workload, with canonical style decoding rather than escape history. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, both A/B arms build, tty, p4 and gui all build.
Diffstat (limited to 'build/snap.zig')
0 files changed, 0 insertions, 0 deletions