<feed xmlns='http://www.w3.org/2005/Atom'>
<title>pardes.git/src/p4.zig, branch main</title>
<subtitle>Pardes</subtitle>
<id>https://git.0x4200.cafe/pardes.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/pardes.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/'/>
<updated>2026-08-27T12:47:39Z</updated>
<entry>
<title>One core behind N frontends, the board's own runner moved in, and every board cap on one screen</title>
<updated>2026-08-27T12:47:39Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T16:27:46Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=11f380f6d7222f2cad93c2cdf13701ea1f903d47'/>
<id>urn:sha1:11f380f6d7222f2cad93c2cdf13701ea1f903d47</id>
<content type='text'>
## The wire is the effect stream, not a new protocol

`pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns
a terminal and a socket and nothing else. N frontends on one core all look at the same screen —
`screen -x`, not N sessions.

The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per
message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the
core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is
a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin
command line, and a command line cannot carry a frame.

ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be
riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit
little-endian fixed width and no message is a blit of a native struct. A protocol that only works
between two builds of the same compiler would have thrown away the one frontend that motivated it.

## The board comes in; its toolchain stays out

`src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over-
serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves
into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so
`zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the
board from this repo's `build.zig`.

The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes
core over a serial line. What stayed is everything a second project would also want — the HAL, the
register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its
own and its `build()` early-returns when it is not the root package, so this costs the package
graph exactly zero packages and the editor's own builds nothing at all.

## limits.zig: nine forgettable places become one budget

Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions —
they are ONE decision, how much memory this build may spend, taken nine times where no reader could
see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against
what it is measured against, derived from two booleans.

The payoff is testability on a machine that is not the board: the caps are ordinary comptime values,
so a host build can be compiled against the board's numbers and the parking, eviction and clamping
paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART.

## A bare `zig build`

`zig build` with no arguments now builds the tty and GUI binaries and installs them into
`~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built
one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and
"use it" two different commands for no reason.
</content>
</entry>
<entry>
<title>A Gpio word that flips one pin, JP1 drawn in ASCII, and these words only on the P4</title>
<updated>2026-08-26T15:40:03Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T15:40:03Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=fbc194068687e49a8490c85c9f1257a2f2bb9079'/>
<id>urn:sha1:fbc194068687e49a8490c85c9f1257a2f2bb9079</id>
<content type='text'>
## Gpio

`Gpio 33` flips one pad and answers on the message row with what it did:

    GPIO 33: 0-&gt;1
    GPIO 33: 1-&gt;0

Bare `Gpio` draws the header instead, because the first question about a header is which pins it
has. The pin number is DECIMAL and it is the only literal in board_memory.zig that is - every
other one is an address, and addresses come off datasheets and linker maps that print hex, which
is why that file made everything hex two commits ago. A GPIO number is not an address, it is part
of a NAME: the schematic says GPIO47, the datasheet's pin table says 47, and `Gpio 20` meaning pin
32 would be a trap laid for the one argument anybody types from memory.

## The toggle is the host's, not the editor's

New `Host.VTable.pull_gpio_toggle`, and a `GpioFn` in the p4 ABI (hence version 2), rather than
board_memory reaching for GPIO_OUT the way `Poke` two functions above it would happily do.

Writing that register is not the job. A pad has to be pointed at the GPIO function in the IO MUX,
routed in the GPIO matrix, given drive strength and an input buffer with its pulls cleared, and
only then driven - four register files behind a per-pin table. That code already exists in
`05-zig-p4/src/hal/gpio.zig`, it is the same `configureOutput` the blink demo has always used, and
its register numbers are checked against ESP-IDF's own headers on the die by `zig build diff`. A
second copy inside the editor object would be a second copy under no test, and getting it wrong on
a pin that boots as something else is how you lose the console you are typing on.

Reported levels are the OUTPUT bits, before and after, because that is what a toggle means: the
level this board is driving. A pad's input buffer on an unconnected header pin reads the air.

## JP1, read off the schematic rather than remembered

The diagram is the vendor's own wiring, from sheet 2 "Expand IO" of
`01-esp32p4-m3/docs/JC-ESP32P4-M3_schematic.pdf` - the only document that carries this mapping. The
specification PDF's "Interface Description" page turned out to be a marketing render, and there is
no board user guide; the chip datasheet has a package pinout, which is not a header.

That sheet is a 872x1168 raster (`pdfimages -list` - the PDF embeds no vectors, so rendering it
larger adds nothing), and at that size the rows around pin 14 are genuinely ambiguous by eye. So
the mapping came from the drawing's geometry instead: thirteen wires leave each side of the symbol,
a net wire runs ~100 px to its label and a power stub ~21 px. Pin 8's wire is 21 px, which is what
identifies it as unconnected rather than as the first of the GPIO4x labels - the reading that had
GPIO47 one row higher and shorted GPIO45 to the ground bracket.

Cross-checked against a second source that has been in the tree all along: `05-zig-p4/build.zig`
documents `-Dled=20` as "JP1 pin 17", and GPIO20 lands on pin 17 here. Both facts are asserted in
the test, so the diagram cannot drift from either.

## Peek, Poke, Hexdump and Gpio are now the P4 build's alone

`board_memory.enabled` was `os.tag == .freestanding and !isWasm()`, on the argument that these
words are a property of having no operating system rather than a product configuration, and that a
predicate spelled out of `builtin` cannot drift the way a hand-maintained enum can.

Tidy, and it answered the wrong question. A word only exists if some shell offers it, and the
shells are the platforms. `Gpio` settles it beyond argument: its whole content is one board's
header, and a second freestanding port would need its own pinout rather than inheriting this one.
"Bare metal" was never the requirement, "this board" was, and the two only looked identical
because there is currently one of them. The old predicate's real work was excluding wasm -
`freestanding` too, where an address is an offset into a linear memory the engine owns - and naming
`p4` excludes it by construction instead of by a term somebody has to keep remembering. The target
is now the witness rather than the gate.

Absent means not compiled: the tty binary contains no `+Gpio`, no `+Hexdump`, no `ES_I2C_SDA` and
no `MisalignedAddress`.

## The boot buffer's lines are checked, not eyeballed

Three times now a line in that tour has been one or two characters too long for a 56-column grid,
and every time it was found by reading the die's screen - the expensive way to measure a string
literal. The text is a named `boot_buffer` with a test over it, six lines came down to fit with
margin, and the tour gained `Gpio`.

Tests: the pinout's width, its thirteen aligned pin rows, GPIO20-on-17 and pin-8-unconnected; the
decimal-versus-hex distinction; every boot-buffer line. Full suite green - unit-test, snap 95/95,
hxdiff 481/0, hxparity 561/0, image-harness, pdf-harness, mupdf-check - and tty, p4, gui,
p4 at 80x24, p4 with the fade forced on. On the die `p4-bench --check` is 5/5, the fifth being a
new one: three `Gpio 33` runs must report 0-&gt;1, 1-&gt;0, 0-&gt;1, because the alternation is the only
oracle a hardcoded string could not fake.
</content>
</entry>
<entry>
<title>Raise the board's grid to 56x14, and make it a build option</title>
<updated>2026-08-26T05:04:47Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T05:04:47Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=fc263fc4b1ee36a3c4bfd7d730cccb57ebe1064a'/>
<id>urn:sha1:fc263fc4b1ee36a3c4bfd7d730cccb57ebe1064a</id>
<content type='text'>
The 40x12 ceiling was never about the screen. It was about memory, and the comment above
`max_cols` said so: "every cell is paid for four times over: vaxis keeps a Screen and an
InternalScreen, pardes keeps its own Surface and previous_cells". Two of those four are now
dead weight - with `direct_emit` the emitter diffs the Surface against its own shadow and
writes the escapes itself, so vaxis's two grids are allocated, never read, and were the
largest single claim on a 384 KiB heap. `init` sizes them to ONE CELL. vaxis still does the
work only it can do: the alternate screen, the capability queries, and parsing everything
that comes back.

That removes the memory ceiling entirely - the heap now reports 336 KB free at every
geometry tried, including ones that used to fail - and leaves latency as the only limit,
which is the honest one: every frame walks the whole grid.

## Measured on the die, 0.87 us per cell

    geometry   cells   round trip
      40x12      480     3,628 us   the old default
      56x14      784     3,930 us   the new one
      56x16      896     3,965 us
      60x18    1,080     4,114 us
      64x20    1,280     4,281 us
      80x24    1,920     4,809 us
     100x30    3,000     5,743 us
     120x36    4,320     6,923 us
     140x42    5,880     8,310 us   the largest that runs
     160x48    7,680     links, then traps
     200x60   12,000     does not link

56x14 is 63% more area and 40% more width than 40x12 and still holds the 4 ms this port was
built to. 56x16 was tried first: 3,965 us on the bench instrument but 4,029 on the
phase-randomised one, which is over, and the two instruments differ by about 50 us
systematically - so the wider grid went and two rows stayed behind. Width is worth more than
height for reading code.

The two failures at the top are worth naming precisely because they are different failures.
200x60 does not link: `.bss will not fit in region l2mem, overflowed by 76036 bytes`, that
`.bss` being the shell's shadow copy of the grid, sized at comptime. 160x48 links and then
TRAPS at boot - the same region pressure arriving at runtime as a collision rather than as a
diagnostic. Neither is a heap problem any more, which is the interesting part: the heap has
336 KB spare while `.bss` runs out.

`-Dp4-cols` / `-Dp4-rows` because none of the above is a constant. 80x24 is one flag away for
anyone who would rather have the classic terminal than the millisecond.

Verified at the new geometry rather than assumed: the A/B against the reference path - vaxis
rendering, full repaint, `shadow_grid` and `direct_emit` both off - is identical in every
cell, characters and resolved style. That matters more here than usual because the emitter's
column arithmetic has a special case at the last column, and 40 was the only width it had
ever been asked about. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, both A/B arms,
tty/p4/gui, and the board's own `p4-bench --check`.
</content>
</entry>
<entry>
<title>Hold a lone ESC: every escape sequence on this wire was being shredded</title>
<updated>2026-08-26T04:42:16Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T04:42:15Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=2d3148247e6555b7b33bd532ae538d7a4358e160'/>
<id>urn:sha1:2d3148247e6555b7b33bd532ae538d7a4358e160</id>
<content type='text'>
The mouse did not work. Chasing that found something much larger: NO escape sequence
worked on this transport, and had not since the port began.

`vaxis.Parser` resolves a buffer containing nothing but 0x1b as the Escape KEY. That is
deliberate and correct for a terminal, where the kernel hands over a whole escape sequence
in a single read, so a solitary ESC really does mean somebody pressed Escape. A 115200
serial line hands over ONE BYTE AT A TIME - 87 us apart, an eternity to a loop running at
360 MHz - so the first byte of every sequence arrived alone and was resolved as Escape,
and the remaining bytes arrived as ordinary keys.

A mouse click therefore came through as TEN key presses: Escape, `[`, `&lt;`, `0`, `;`, `1`,
`8`, `;`, `3`, `M`. The `0` among them is "go to column zero" in normal mode, which is
exactly where the cursor kept landing, and why the first attempt at this looked like a
coordinate bug. Arrow keys, function keys, and the host bridge's in-band resize reports
were all being taken apart the same way.

Longer partial sequences were never affected: the CSI scanner returns `n == 0` for "no
final byte yet" and the shell already keeps those bytes. Only the one-byte case needed an
answer, because it is the only one the parser answers WRONGLY instead of declining. So the
shell holds a buffer that is exactly one ESC and lets `pardes_p4_tick` release it after
10 ms - two orders of magnitude longer than the 87 us until the next byte of a real
sequence, and imperceptible to a person pressing Escape. The same trade every terminal
editor makes, for the same reason.

Finding it took instrumenting the ABI: printing `@tagName` of every event the shell
applied. Ten `key_press` where one `mouse` belonged is not a thing any amount of reading
the coordinate arithmetic would have shown, and I had already read it twice.

## Mouse reporting, and the 1003 that is not requested

With the sequences intact, `apply` already handled `.mouse` - it mirrors the tty shell - so
enabling reporting was the only missing piece. Spelled out here rather than taken from
`vx.setMouseMode`, which asks for `1002;1003;1004;1006`: 1003 is ANY-MOTION tracking, a
report per cell the pointer crosses with no button held. On a 115200 line that is dozens of
15-byte reports for one sweep, arriving as input the editor must parse while it paints, and
arriving whether or not anyone wants it - moving the mouse over the window would starve
typing. 1002 reports presses, releases and motion while a button is held, which is exactly
what a click and a drag-select need.

Verified on the die: a click at column 12 puts the cursor at column 12 and one at column 22
puts it at column 22, a drag paints a selection, and the wheel scrolls. A press alone paints
the new position and then reverts - the caret does not move until the gesture ends - so the
release is what commits it, which cost an hour of believing a working click was broken.

Screen byte-identical to the vaxis reference, round trip median 3682 us against 3682, snap
95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty/p4/gui all build.
</content>
</entry>
<entry>
<title>The frame diff was comparing byte at a time; compare words, and pay the bridge on every frame</title>
<updated>2026-08-26T01:39:09Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T01:10:36Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=4ab24352873ed7bf8db93ef6bfec36a34b0357e8'/>
<id>urn:sha1:4ab24352873ed7bf8db93ef6bfec36a34b0357e8</id>
<content type='text'>
Two findings, both in the P4 shell's own `present`.

## std.mem.eql was the largest read in the firmware, one byte at a time

The shadow-grid diff compares each row against the previous frame: two 13 KB streams,
every frame, and by far the biggest memory access the firmware makes. It measured 3.2
cycles per byte, which is about four times what word-wide loads need - the shape of a
byte-at-a-time loop, and `std.mem.eql` is what it was.

`sameBytes` compares a `u32` at a time and falls back to the byte loop when the spans are
not aligned for it. The alignment test has to be a RUNTIME one because `Cell` is all `u8`
fields and therefore has alignment 1: whether a row begins on a word boundary is a
property of whoever allocated the Surface, not of the type. A row is 40 cells of 26 bytes,
divisible by four, so an aligned base makes every row aligned.

The answer is bit-for-bit the same - this is still exact byte equality - so it keeps the
property the whole diff rests on: byte equality implies visual equality, so the diff can
never claim two different cells are the same.

Measured on the die: the grid walk 223 -&gt; 66 us, 0.95 cycles per byte. 157 us off every
keystroke at every document length, and the single largest win since the clock raise.

## A frame that only hides the cursor still has to fill a USB packet

The padding added for the bridge's 32-byte bulk-IN packet covered the branch that
positions the cursor and not the branch that hides it. A frame that only hid the cursor
was six bytes and waited out the bridge's timer. Hiding an already-hidden cursor is as
idempotent as positioning it twice, so it pads the same way.

The packet size is no longer inferred from an experiment either: 32 is `wMaxPacketSize` of
endpoint 0x82 as the device reports it, and the sweep over pad targets confirms what it
implies - 0 and 16 sit at 4.7-5.1 ms, while 32, 48 and 64 all sit at 3.6-3.8 ms. Crossing
the boundary is worth about 950 us; going past it buys nothing.

## Result

    length   0     20     40     80    160    320    640 chars
    RTT   3602   3624   3638   3790   3868   4026   4192 us

Fixed cost 3652 us against a 4 ms target, from 16.99 ms where this started. A
phase-randomised instrument agrees over 80 trials: median 3687 us, minimum 3571, maximum
3912 - every trial under 4 ms.

The two lengths still above 4 ms are the ones where the line has outgrown the viewport, so
the cursor is off screen and the keystroke changes NOTHING: the frame is 36 bytes of
cursor-hide and padding, zero cells changed, while pardes still rebuilds all 480 cells of
the Surface for 858-985 us. That is the one architectural item left and it is not a micro
-optimisation: nothing in this repository can avoid work pardes has already done.

Verified: screen byte-identical to the vaxis reference on the 18-step workload, with
canonical style decoding rather than escape history. snap 95/95, hxdiff 481/0, hxparity
561/0, unit-test, both A/B arms build, tty, p4 and gui all build.
</content>
</entry>
<entry>
<title>Drop the dead std.fmt fallback in the emitter's integer writer</title>
<updated>2026-08-26T00:42:56Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:42:56Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=209b48a36db527904ddc4b436f79df45cea3ab3e'/>
<id>urn:sha1:209b48a36db527904ddc4b436f79df45cea3ab3e</id>
<content type='text'>
Five digits is every u16, so the `v &gt;= 10000` branch could never be taken and it was
dragging `std.fmt.printInt` into a firmware whose whole reason for hand-rolling this
was to keep the format machinery out of the hottest sequence it emits.

Behaviour is identical, and re-verified rather than assumed: screen byte-identical to
the vaxis reference on the 18-step workload, round trip median 3830 us over 60 trials
(3829 before), snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test.
</content>
</entry>
<entry>
<title>Emit the ANSI directly, and pay the USB bridge its minimum frame</title>
<updated>2026-08-26T00:36:37Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:36:37Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=81c149b4343c125051c829bac13eea0c38f4262c'/>
<id>urn:sha1:81c149b4343c125051c829bac13eea0c38f4262c</id>
<content type='text'>
`present` already knows exactly which cells moved - that is what the shadow grid is
for - and then handed every one of them to vaxis so that vaxis could work it out again
against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of
it redundant. This emits the escapes itself and skips it.

The emitter is small because it is allowed to be: one absolute CUP per run of changed
cells rather than per cell, absolute SGR rather than a delta from whatever is currently
on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it
writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style
change and buys the property that no cell can inherit an earlier cell's colour if a
frame is cut short. Cursor column tracking gives up after anything that is not a single
printable ASCII byte, and at the last column, because deferred wrap makes the answer
terminal-dependent and wrong by a whole row.

Board cost: `render` 1362 -&gt; 779 us. Bytes per keystroke: 81 -&gt; 21. Image 26.6 KB
smaller, since vaxis's renderer is now unreachable.

## And it measured SLOWER

4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of
result that means the model is wrong, so I stopped optimising and went looking.

It is the USB bridge. The board talks to the host through a CH340, a full-speed part
whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet
is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal
timer gives up waiting for more - about a millisecond, a quarter of the whole budget.
Routing through vaxis only looked competitive because its frames are 81 bytes and fill a
packet by accident.

The evidence, all at identical board cost and with a byte-identical screen:

    frame     min      median
     21 B    3843 us   4817 us   never fills a packet
     49 B    3719 us   3814 us   padded past the boundary
     81 B    4373 us   4475 us   vaxis, fills one by accident

Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the
compute that was saved. Only the median was hostage to the timer.

So the frame has a minimum size and it belongs to the transport, not the terminal. Pad
to it, with repeated absolute cursor positioning: idempotent, already the sequence the
frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper
so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has
a minimum and the sender pays it - and it is a real trade rather than free, since the
filler is wire time that delays a later frame. It only applies when the frame is small,
which is when there is wire to spare.

## Result: 3.74 ms, and the goal was 4.00

    step                fixed    per char   at 160 chars
    ReleaseSmall      16.99 ms    54.3 us      25.56 ms
    ReleaseFast       14.85 ms    34.7 us      20.30 ms  0.79x
    + ASCII grapheme  14.56 ms    12.0 us      16.46 ms  0.64x
    + ASCII print     14.27 ms     6.9 us      15.36 ms  0.60x
    + shadow grid      8.87 ms     7.3 us      10.02 ms  0.39x
    + byte compare     8.37 ms     7.1 us       9.48 ms  0.37x
    + 360 MHz          4.37 ms     1.9 us       4.67 ms  0.18x
    + direct emit      3.74 ms     2.0 us       4.06 ms  0.16x

35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials,
median 3829 us, min 3722, p90 3930.

That second instrument exists because of this commit. The original bench sends keystrokes
on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and
makes the round trip a staircase in board time - a real saving can measure as a
regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two.
It was not what was happening here, but it had to be excluded before the CH340 could be
believed, and it is the right default for anything measured across this link.

## Verification

`direct_emit = false` routes every cell back through vaxis and is the reference. Both
arms, same 18-step workload, same clock: identical characters and identical resolved
style in every cell - resolved, not raw SGR, because two emitters reaching the same
colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly
the change that can be right about latency and wrong about the screen, and until the
verifier compared canonical style rather than escape history it could not have told the
difference.

snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
both A/B arms build, tty, p4 and gui all build.
</content>
</entry>
<entry>
<title>Compare a whole row with one memcmp before looking at cells</title>
<updated>2026-08-26T00:10:50Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:10:50Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=f5d5221d683a97de24fba54597a278decf6b8e98'/>
<id>urn:sha1:f5d5221d683a97de24fba54597a278decf6b8e98</id>
<content type='text'>
`Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the
shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell
loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes.

Byte equality implies visual equality, which is what makes the shortcut sound: a row
that compares equal cannot be hiding a changed cell, and a row that differs only in
padding falls through to the per-cell path, which is correct and merely slower.

Measured on the die at 360 MHz: the grid walk 246 -&gt; 226 us. That is a small win and
the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte,
this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there
is little left in it. It is also why board compute scaled 2.6x rather than 4x when the
core clock went up 4x.

Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and
incremental path, same 18-step workload, same clock - identical characters and
identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0,
unit-test, tty and p4 both build.
</content>
</entry>
<entry>
<title>Compare shadow-grid cells as bytes, not through std.meta.eql</title>
<updated>2026-08-25T23:30:52Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:30:52Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=649fd0e983d6a207932f1ec43f7ded00e26a69c0'/>
<id>urn:sha1:649fd0e983d6a207932f1ec43f7ded00e26a69c0</id>
<content type='text'>
`Cell.visuallyEqual` is the semantically exact answer and too slow to ask 480 times a
frame: `std.meta.eql` on a `CellStyle` recurses through a colour union and eight
booleans, and the walk measured 1.45 ms on the die - about 270 cycles to compare a
28-byte struct.

`sameCell` in src/p4.zig does it as bytes. That is safe in the direction that
matters: byte equality IMPLIES visual equality, so it can never claim two different
cells are the same. It can miss an equality - scratch bytes past `len`, or padding -
and the only cost of that is one redundant `writeCell` which vaxis then diffs away.
Defaults are still compared by meaning, because an unpainted cell's text and style are
whatever the previous frame left in them.

Measured: the grid walk 1.45 -&gt; 0.98 ms, a keystroke 8.87 -&gt; 8.37 ms fixed.

Verified the way a rendering change has to be. The A/B harness now hashes the SGR
state of every cell as well as its character, because the first version compared text
only and would have passed a colour regression in silence. Reference path
(`shadow_grid = false`, clear and write everything) and incremental path were each run
against the same 19-step workload on the die and the reconstructed screens are
identical in both text and per-row style hash.

snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
and tty / p4 / gui all build.
</content>
</entry>
<entry>
<title>Make a keystroke 2.6x cheaper by not asking Unicode about ASCII</title>
<updated>2026-08-25T23:20:24Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:20:24Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/pardes.git/commit/?id=767a35d3ae3b87d1bfdd960f2187f0a51ad32333'/>
<id>urn:sha1:767a35d3ae3b87d1bfdd960f2187f0a51ad32333</id>
<content type='text'>
A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in
that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf
-Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it
was Unicode machinery answering questions about the letter `y`.

Four changes, each a fast path guarded so that non-ASCII text takes exactly the road
it took before.

`modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates
graphemes FROM THE START of the text with the full UAX #29 break state machine until
it passes the offset, and the render path calls it once per visible row with a column
offset - so the cost followed the cursor's distance along its line. That is the shape
measured on the die, where inserting at column 320 of a fixed 320-character line cost
7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII
scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule
that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator -
is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII,
and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate.

`Surface.print` then became the largest at 26.2%: per character it took a UTF-8
length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a
width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII
takes none of that now. Same guard, same reason.

`file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth`
about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one
printable cell and `gwidth` stays the authority on them.

`modal.lineSlice` searched for "\n" with the generic substring search where a memchr
does; it is called once per visible row per frame.

Measured at the P4's geometry and configuration, on the host: render 55 -&gt; 12 us,
key-down 483 -&gt; 24 us, key-right 327 -&gt; 13 us, edit-char 205 -&gt; 46 us. On the die,
the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a
keystroke at a 160-character line from 25.56 ms to 15.36 ms.

## The shadow grid, and why it is static

`src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured
6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything
changed: a second render with nothing new cost the same as the first. vaxis diffs its
own grid, but only after being told every cell, and being told is the expensive part.
So `present` now keeps the previous Surface and tells vaxis only what moved.
`Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -&gt; 1.45 ms.

The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is
not a micro-optimisation. The first version allocated it from the editor's heap; on a
board whose 384 KiB is nearly spoken for, that is exactly the kind of change that
works and then breaks something else three steps away.

`shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present`
behaves as it did before - clear and write every cell - which is the reference any
measurement should be compared against, and the way to tell a rendering bug from a
rendering difference. It earned its keep immediately: the two paths were run against
the same 19-step workload on the die - inserts, deletes, motions that move the
modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the
reconstructed screens are byte-identical.

## Verification

`snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0
mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 /
gui all build. The rendering changes are exactly the sort that pass a latency
benchmark while corrupting a screen, so the snapshot parity suite is the one that
matters here and it is unchanged.

`test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the
things that table exists to hold constant, and 40x12 is not a scaled guess at the
board - it is the board. `--only` exists because under `perf record` one 63 ms cell on
the largest fixture swamps every sample from the case being asked about.

## Found, not fixed

`vx.resize` fails on this board: a runtime geometry change hits its allocation
failure path, restores the previous size and returns, so 80 bytes go out where 1,392
should. Verified independent of everything above - it reproduces with `shadow_grid`
false. The board therefore has one geometry for the life of a session, which is why
the staleness test above compares two firmwares rather than resizing one.
</content>
</entry>
</feed>
