diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 21:10:50 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 21:10:50 -0300 |
| commit | f5d5221d683a97de24fba54597a278decf6b8e98 (patch) | |
| tree | 9bdc6e23b1f9abec88edc21d302a8a0dfba289fd /src/p4.zig | |
| parent | 649fd0e983d6a207932f1ec43f7ded00e26a69c0 (diff) | |
| download | pardes-f5d5221d683a97de24fba54597a278decf6b8e98.tar.gz pardes-f5d5221d683a97de24fba54597a278decf6b8e98.zip | |
Compare a whole row with one memcmp before looking at cells
`Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the
shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell
loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes.
Byte equality implies visual equality, which is what makes the shortcut sound: a row
that compares equal cannot be hiding a changed cell, and a row that differs only in
padding falls through to the per-cell path, which is correct and merely slower.
Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and
the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte,
this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there
is little left in it. It is also why board compute scaled 2.6x rather than 4x when the
core clock went up 4x.
Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and
incremental path, same 18-step workload, same clock - identical characters and
identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0,
unit-test, tty and p4 both build.
Diffstat (limited to 'src/p4.zig')
| -rw-r--r-- | src/p4.zig | 21 |
1 files changed, 17 insertions, 4 deletions
@@ -541,12 +541,25 @@ fn present(_: ?*anyopaque, surface: *const pardes.Surface) void { var y: u16 = 0; while (y < surface.rows) : (y += 1) { + const row0 = @as(usize, y) * @as(usize, surface.cols); + const src = surface.cells[row0..][0..surface.cols]; + + // A ROW AT A TIME FIRST. `Surface.cells` is contiguous and row-major, so a whole row is one + // `memcmp` against the shadow - and on a keystroke eleven of twelve rows are untouched. The + // per-cell loop below is ~40 branchy comparisons where this is one call over 1,120 bytes; + // measured, the walk fell from 246 us to a fraction of it. Byte equality implies visual + // equality (see `sameCell`), so a row that compares equal cannot be hiding a changed cell - + // and a row that differs only in padding falls through to the per-cell path, which is + // correct and merely slower. + if (usable and !full) { + const shadow = prev_cells[row0..][0..surface.cols]; + if (std.mem.eql(u8, std.mem.sliceAsBytes(src), std.mem.sliceAsBytes(shadow))) continue; + } + var x: u16 = 0; while (x < surface.cols) : (x += 1) { - // `at` takes a mutable Surface but only reads; the tty shell does the same const-cast - // for the same reason (src/tty/tty.zig:1105). - const cell = @constCast(surface).at(x, y); - const idx = @as(usize, y) * @as(usize, surface.cols) + @as(usize, x); + const cell = &src[x]; + const idx = row0 + @as(usize, x); if (usable) { if (!full and sameCell(cell, &prev_cells[idx])) continue; prev_cells[idx] = cell.*; |
