summaryrefslogtreecommitdiff
path: root/src/detached/client.zig
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-25 21:36:37 -0300
committerGabriel Schneider <[email protected]>2026-08-25 21:36:37 -0300
commit81c149b4343c125051c829bac13eea0c38f4262c (patch)
tree787eb694455206ca698473863718e3b9c0933593 /src/detached/client.zig
parentf5d5221d683a97de24fba54597a278decf6b8e98 (diff)
downloadpardes-81c149b4343c125051c829bac13eea0c38f4262c.tar.gz
pardes-81c149b4343c125051c829bac13eea0c38f4262c.zip
Emit the ANSI directly, and pay the USB bridge its minimum frame
`present` already knows exactly which cells moved - that is what the shadow grid is for - and then handed every one of them to vaxis so that vaxis could work it out again against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of it redundant. This emits the escapes itself and skips it. The emitter is small because it is allowed to be: one absolute CUP per run of changed cells rather than per cell, absolute SGR rather than a delta from whatever is currently on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style change and buys the property that no cell can inherit an earlier cell's colour if a frame is cut short. Cursor column tracking gives up after anything that is not a single printable ASCII byte, and at the last column, because deferred wrap makes the answer terminal-dependent and wrong by a whole row. Board cost: `render` 1362 -> 779 us. Bytes per keystroke: 81 -> 21. Image 26.6 KB smaller, since vaxis's renderer is now unreachable. ## And it measured SLOWER 4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of result that means the model is wrong, so I stopped optimising and went looking. It is the USB bridge. The board talks to the host through a CH340, a full-speed part whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal timer gives up waiting for more - about a millisecond, a quarter of the whole budget. Routing through vaxis only looked competitive because its frames are 81 bytes and fill a packet by accident. The evidence, all at identical board cost and with a byte-identical screen: frame min median 21 B 3843 us 4817 us never fills a packet 49 B 3719 us 3814 us padded past the boundary 81 B 4373 us 4475 us vaxis, fills one by accident Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the compute that was saved. Only the median was hostage to the timer. So the frame has a minimum size and it belongs to the transport, not the terminal. Pad to it, with repeated absolute cursor positioning: idempotent, already the sequence the frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has a minimum and the sender pays it - and it is a real trade rather than free, since the filler is wire time that delays a later frame. It only applies when the frame is small, which is when there is wire to spare. ## Result: 3.74 ms, and the goal was 4.00 step fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x + byte compare 8.37 ms 7.1 us 9.48 ms 0.37x + 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x + direct emit 3.74 ms 2.0 us 4.06 ms 0.16x 35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials, median 3829 us, min 3722, p90 3930. That second instrument exists because of this commit. The original bench sends keystrokes on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and makes the round trip a staircase in board time - a real saving can measure as a regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two. It was not what was happening here, but it had to be excluded before the CH340 could be believed, and it is the right default for anything measured across this link. ## Verification `direct_emit = false` routes every cell back through vaxis and is the reference. Both arms, same 18-step workload, same clock: identical characters and identical resolved style in every cell - resolved, not raw SGR, because two emitters reaching the same colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly the change that can be right about latency and wrong about the screen, and until the verifier compared canonical style rather than escape history it could not have told the difference. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, both A/B arms build, tty, p4 and gui all build.
Diffstat (limited to 'src/detached/client.zig')
0 files changed, 0 insertions, 0 deletions