summaryrefslogtreecommitdiff
path: root/src/p4.zig
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-25 21:36:37 -0300
committerGabriel Schneider <[email protected]>2026-08-25 21:36:37 -0300
commit81c149b4343c125051c829bac13eea0c38f4262c (patch)
tree787eb694455206ca698473863718e3b9c0933593 /src/p4.zig
parentf5d5221d683a97de24fba54597a278decf6b8e98 (diff)
downloadpardes-81c149b4343c125051c829bac13eea0c38f4262c.tar.gz
pardes-81c149b4343c125051c829bac13eea0c38f4262c.zip
Emit the ANSI directly, and pay the USB bridge its minimum frame
`present` already knows exactly which cells moved - that is what the shadow grid is for - and then handed every one of them to vaxis so that vaxis could work it out again against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of it redundant. This emits the escapes itself and skips it. The emitter is small because it is allowed to be: one absolute CUP per run of changed cells rather than per cell, absolute SGR rather than a delta from whatever is currently on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style change and buys the property that no cell can inherit an earlier cell's colour if a frame is cut short. Cursor column tracking gives up after anything that is not a single printable ASCII byte, and at the last column, because deferred wrap makes the answer terminal-dependent and wrong by a whole row. Board cost: `render` 1362 -> 779 us. Bytes per keystroke: 81 -> 21. Image 26.6 KB smaller, since vaxis's renderer is now unreachable. ## And it measured SLOWER 4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of result that means the model is wrong, so I stopped optimising and went looking. It is the USB bridge. The board talks to the host through a CH340, a full-speed part whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal timer gives up waiting for more - about a millisecond, a quarter of the whole budget. Routing through vaxis only looked competitive because its frames are 81 bytes and fill a packet by accident. The evidence, all at identical board cost and with a byte-identical screen: frame min median 21 B 3843 us 4817 us never fills a packet 49 B 3719 us 3814 us padded past the boundary 81 B 4373 us 4475 us vaxis, fills one by accident Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the compute that was saved. Only the median was hostage to the timer. So the frame has a minimum size and it belongs to the transport, not the terminal. Pad to it, with repeated absolute cursor positioning: idempotent, already the sequence the frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has a minimum and the sender pays it - and it is a real trade rather than free, since the filler is wire time that delays a later frame. It only applies when the frame is small, which is when there is wire to spare. ## Result: 3.74 ms, and the goal was 4.00 step fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x + byte compare 8.37 ms 7.1 us 9.48 ms 0.37x + 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x + direct emit 3.74 ms 2.0 us 4.06 ms 0.16x 35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials, median 3829 us, min 3722, p90 3930. That second instrument exists because of this commit. The original bench sends keystrokes on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and makes the round trip a staircase in board time - a real saving can measure as a regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two. It was not what was happening here, but it had to be excluded before the CH340 could be believed, and it is the right default for anything measured across this link. ## Verification `direct_emit = false` routes every cell back through vaxis and is the reference. Both arms, same 18-step workload, same clock: identical characters and identical resolved style in every cell - resolved, not raw SGR, because two emitters reaching the same colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly the change that can be right about latency and wrong about the screen, and until the verifier compared canonical style rather than escape history it could not have told the difference. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, both A/B arms build, tty, p4 and gui all build.
Diffstat (limited to 'src/p4.zig')
-rw-r--r--src/p4.zig223
1 files changed, 206 insertions, 17 deletions
diff --git a/src/p4.zig b/src/p4.zig
index 75d660a7..0f23ddd5 100644
--- a/src/p4.zig
+++ b/src/p4.zig
@@ -507,9 +507,10 @@ export fn pardes_p4_quit() callconv(.c) bool {
const pardes_host: pardes.Host.VTable = .{ .push_present = present };
-/// The canonical surface -> vaxis, cell for cell, then one render. Same shape as the tty shell's
-/// (`src/tty/tty.zig:1096`) minus the panel compositor and the kitty image path: neither has a
-/// reason to exist on a board with no pixels.
+/// The canonical surface -> the wire. Same shape as the tty shell's (`src/tty/tty.zig:1096`) minus
+/// the panel compositor and the kitty image path: neither has a reason to exist on a board with no
+/// pixels. Where the tty shell hands every cell to vaxis and lets it diff, this diffs against the
+/// Surface itself and can then emit the ANSI directly - see `direct_emit`.
fn present(_: ?*anyopaque, surface: *const pardes.Surface) void {
const t0 = cycles();
const win = vx.window();
@@ -532,10 +533,16 @@ fn present(_: ?*anyopaque, surface: *const pardes.Surface) void {
// bytes where it had emitted 1,392. The grid is bounded by `max_cols` x `max_rows` at comptime,
// so it belongs in `.bss` where it cannot compete with anything.
const full = !shadow_grid or prev_cols != surface.cols or prev_rows != surface.rows;
+ emit_bytes = 0;
if (full) {
prev_cols = surface.cols;
prev_rows = surface.rows;
- win.clear();
+ if (direct_emit) {
+ // Reset first: a `2J` while a non-default background is active fills the screen with it.
+ emitRaw("\x1b[0m\x1b[2J") catch return;
+ emit_style = .{};
+ emit_col = -1;
+ } else win.clear();
}
const usable = shadow_grid and n <= prev_cells.len;
@@ -565,26 +572,28 @@ fn present(_: ?*anyopaque, surface: *const pardes.Surface) void {
prev_cells[idx] = cell.*;
} else if (cell.default) continue;
- if (cell.default) {
- // Changed TO default. `win.clear()` is what used to blank these, and it is not run
- // on an incremental frame, so say it explicitly.
- win.writeCell(x, y, .{ .char = .{ .grapheme = " " }, .style = .{} });
- continue;
- }
- win.writeCell(x, y, .{
- .char = .{ .grapheme = cell.grapheme() },
- .style = vaxisStyle(cell.style),
- });
+ writeOne(win, x, y, cell, surface.cols) catch return;
}
}
- if (surface.cursor) |cur| {
+ if (direct_emit) {
+ if (surface.cursor) |cur| {
+ cup(cur.y, cur.x) catch return;
+ emitRaw("\x1b[?25h") catch return;
+ emit_col = -1;
+ // Up to the bridge's packet boundary, and no further. See `emit_min_frame`: this is the
+ // one place that knows how many bytes the frame came to, and repeating the sequence the
+ // frame already ended on is the only filler that cannot change what is on the screen.
+ while (emit_bytes < emit_min_frame) cup(cur.y, cur.x) catch return;
+ } else emitRaw("\x1b[?25l") catch return;
+ } else if (surface.cursor) |cur| {
win.showCursor(cur.x, cur.y);
} else win.hideCursor();
const t1 = cycles();
// vaxis diffs against its own shadow grid, so this writes only what changed - which is what
- // makes an editor usable at 11.9 KB/s.
- vx.render(&out) catch return;
+ // makes an editor usable at 11.9 KB/s. With `direct_emit` that diff has already happened, one
+ // stage earlier and against the Surface itself, so there is nothing left here to do.
+ if (!direct_emit) vx.render(&out) catch return;
const t2 = cycles();
out.flush() catch return;
const t3 = cycles();
@@ -594,6 +603,186 @@ fn present(_: ?*anyopaque, surface: *const pardes.Surface) void {
prof_flush_cy = t3 -% t2;
}
+/// One cell to the wire, either through vaxis or straight out.
+inline fn writeOne(win: vaxis.Window, x: u16, y: u16, cell: *const pardes.Cell, cols: u16) !void {
+ if (!direct_emit) {
+ // Changed TO default. `win.clear()` is what used to blank these, and it is not run on an
+ // incremental frame, so say it explicitly.
+ if (cell.default) return win.writeCell(x, y, .{ .char = .{ .grapheme = " " }, .style = .{} });
+ return win.writeCell(x, y, .{
+ .char = .{ .grapheme = cell.grapheme() },
+ .style = vaxisStyle(cell.style),
+ });
+ }
+
+ if (emit_row != y or emit_col != x) {
+ try cup(y, x);
+ emit_row = y;
+ emit_col = @intCast(x);
+ }
+
+ const style: pardes.CellStyle = if (cell.default) .{} else cell.style;
+ if (!std.meta.eql(emit_style, style)) {
+ try emitStyle(style);
+ emit_style = style;
+ }
+
+ try emitRaw(if (cell.default) " " else cell.grapheme());
+
+ // Where the terminal's cursor now is. A single printable ASCII byte advanced it exactly one
+ // column; anything else - a wide glyph, a cluster, the spacer cell pardes writes after a wide
+ // one - is not worth predicting, so give up and let the next cell emit an absolute CUP. The last
+ // column is given up on too, because whether the cursor rests on it or has wrapped past it
+ // depends on the terminal's deferred-wrap behaviour, and the two disagree by a whole row.
+ if (x + 1 < cols and cell.len == 1 and cell.text[0] >= 0x20 and cell.text[0] < 0x7f) {
+ emit_col += 1;
+ } else emit_col = -1;
+}
+
+/// A style as an absolute SGR, always opening with a reset.
+///
+/// Absolute rather than a delta from whatever is currently on, and that is what keeps it short
+/// enough to be worth having: no per-attribute off-codes, no state to keep beyond the last style
+/// emitted, and a frame that gets cut off cannot leave a later cell wearing an earlier one's colour.
+/// It costs a few bytes on a style change, against the ~9 of CUP a changed cell is paying anyway.
+fn emitStyle(s: pardes.CellStyle) !void {
+ try emitRaw("\x1b[0");
+ if (s.bold) try emitRaw(";1");
+ if (s.dim) try emitRaw(";2");
+ if (s.italic) try emitRaw(";3");
+ if (s.blink) try emitRaw(";5");
+ if (s.reverse) try emitRaw(";7");
+ if (s.invisible) try emitRaw(";8");
+ if (s.strikethrough) try emitRaw(";9");
+ try emitRaw(switch (s.ul) {
+ .off => "",
+ .single => ";4",
+ .double => ";4:2",
+ .curly => ";4:3",
+ .dotted => ";4:4",
+ .dashed => ";4:5",
+ });
+ try emitColor(s.fg, 30);
+ try emitColor(s.bg, 40);
+ try emitRaw("m");
+}
+
+/// `base` is 30 for a foreground and 40 for a background, which is the only thing separating the two
+/// in every form SGR has for a colour: 30-37 against 40-47, 90-97 against 100-107, 38 against 48.
+fn emitColor(c: pardes.Color, comptime base: u16) !void {
+ var b: [20]u8 = undefined;
+ var i: usize = 0;
+ switch (c) {
+ // Already said by the reset this SGR opens with.
+ .default => return,
+ .index => |n| {
+ b[i] = ';';
+ i += 1;
+ if (n < 8) {
+ i += dec(b[i..], base + n);
+ } else if (n < 16) {
+ i += dec(b[i..], base + 60 + (n - 8));
+ } else {
+ i += dec(b[i..], base + 8);
+ i += lit(b[i..], ";5;");
+ i += dec(b[i..], n);
+ }
+ },
+ .rgb => |v| {
+ b[i] = ';';
+ i += 1;
+ i += dec(b[i..], base + 8);
+ i += lit(b[i..], ";2;");
+ for (v, 0..) |component, k| {
+ if (k != 0) {
+ b[i] = ';';
+ i += 1;
+ }
+ i += dec(b[i..], component);
+ }
+ },
+ }
+ try emitRaw(b[0..i]);
+}
+
+/// Absolute cursor positioning, hand-rolled rather than through `out.print`.
+///
+/// Not for elegance: this is the single most frequent sequence the emitter produces, at least one per
+/// changed run, and `std.fmt` brings a whole format-string interpreter to write at most two digits.
+/// The grid is bounded by `max_cols` x `max_rows`, so nothing here can exceed three.
+fn cup(row: u16, col: u16) !void {
+ var b: [12]u8 = undefined;
+ var i: usize = lit(&b, "\x1b[");
+ i += dec(b[i..], row + 1);
+ b[i] = ';';
+ i += 1;
+ i += dec(b[i..], col + 1);
+ b[i] = 'H';
+ i += 1;
+ try emitRaw(b[0..i]);
+}
+
+fn dec(buf: []u8, v: u16) usize {
+ if (v >= 10000) return std.fmt.printInt(buf, v, 10, .lower, .{});
+ var digits: [5]u8 = undefined;
+ var n: usize = 0;
+ var rest = v;
+ while (true) {
+ digits[n] = '0' + @as(u8, @intCast(rest % 10));
+ n += 1;
+ rest /= 10;
+ if (rest == 0) break;
+ }
+ for (0..n) |k| buf[k] = digits[n - 1 - k];
+ return n;
+}
+
+inline fn lit(buf: []u8, comptime s: []const u8) usize {
+ @memcpy(buf[0..s.len], s);
+ return s.len;
+}
+
+/// Every direct-emit byte goes through here, because the count is what the padding below needs.
+inline fn emitRaw(bytes: []const u8) !void {
+ emit_bytes += bytes.len;
+ try out.writeAll(bytes);
+}
+
+/// A/B switch for the emitter above, on the same terms as `shadow_grid`: false routes every cell back
+/// through vaxis, which is the reference. vaxis's own diff measured 631 us of a 4.37 ms keystroke and
+/// all of it was redundant - `present` has already worked out which cells moved, so vaxis was being
+/// told the answer and then computing it again from scratch.
+const direct_emit = true;
+
+/// THE FRAME HAS A MINIMUM SIZE, and it is the USB bridge's, not the terminal's.
+///
+/// The board is wired to the host through a CH340, a full-speed device whose bulk IN endpoint takes
+/// 32-byte packets. It forwards a packet when the packet is FULL, and a frame shorter than that sits
+/// in the bridge until an internal timer gives up on more - which is worth about a millisecond, and
+/// a millisecond is a quarter of the entire keystroke budget.
+///
+/// Measured, at the same board cost and with the screen byte-identical: a 21-byte frame round-trips
+/// in 4817 us and the same frame padded to 49 bytes in 3814 us. MORE BYTES, ARRIVING SOONER. It also
+/// explains why routing through vaxis looked competitive - its frames are 81 bytes, so they fill a
+/// packet by accident and never wait.
+///
+/// So pad to the packet boundary. The filler is repeated absolute cursor positioning: idempotent,
+/// already the sequence the emitter ends on, and it cannot alter a cell. This is the same bargain as
+/// an Ethernet runt frame - the medium has a minimum and the sender pays it - and it is a real
+/// trade, not free: the wasted bytes are wire time that delays a LATER frame, so it is only worth it
+/// while the frame is small, which is exactly when it applies.
+const emit_min_frame: usize = 32;
+
+/// Bytes emitted this frame, for `emit_min_frame`.
+var emit_bytes: usize = 0;
+
+/// What the terminal is currently wearing and where its cursor is, so that a run of changed cells in
+/// one row costs one CUP and one SGR rather than one of each per cell. `emit_col` is signed because
+/// -1 means "no longer known" - see `writeOne`.
+var emit_style: pardes.CellStyle = .{};
+var emit_row: u16 = 0;
+var emit_col: i32 = -1;
+
/// A/B switch, kept because this optimisation is exactly the kind that can be right about latency
/// and wrong about the screen. With it false, `present` behaves as it did before the shadow grid -
/// clear and write every cell - which is the reference any measurement of it should be compared