summaryrefslogtreecommitdiff
path: root/src/p4.zig
Commit message (Collapse)AuthorAge
* One core behind N frontends, the board's own runner moved in, and every ↵Gabriel Schneider2026-08-27
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | board cap on one screen ## The wire is the effect stream, not a new protocol `pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns a terminal and a socket and nothing else. N frontends on one core all look at the same screen — `screen -x`, not N sessions. The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin command line, and a command line cannot carry a frame. ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit little-endian fixed width and no message is a blit of a native struct. A protocol that only works between two builds of the same compiler would have thrown away the one frontend that motivated it. ## The board comes in; its toolchain stays out `src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over- serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so `zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the board from this repo's `build.zig`. The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes core over a serial line. What stayed is everything a second project would also want — the HAL, the register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its own and its `build()` early-returns when it is not the root package, so this costs the package graph exactly zero packages and the editor's own builds nothing at all. ## limits.zig: nine forgettable places become one budget Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions — they are ONE decision, how much memory this build may spend, taken nine times where no reader could see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against what it is measured against, derived from two booleans. The payoff is testability on a machine that is not the board: the caps are ordinary comptime values, so a host build can be compiled against the board's numbers and the parking, eviction and clamping paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART. ## A bare `zig build` `zig build` with no arguments now builds the tty and GUI binaries and installs them into `~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and "use it" two different commands for no reason.
* A Gpio word that flips one pin, JP1 drawn in ASCII, and these words only on ↵Gabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | the P4 ## Gpio `Gpio 33` flips one pad and answers on the message row with what it did: GPIO 33: 0->1 GPIO 33: 1->0 Bare `Gpio` draws the header instead, because the first question about a header is which pins it has. The pin number is DECIMAL and it is the only literal in board_memory.zig that is - every other one is an address, and addresses come off datasheets and linker maps that print hex, which is why that file made everything hex two commits ago. A GPIO number is not an address, it is part of a NAME: the schematic says GPIO47, the datasheet's pin table says 47, and `Gpio 20` meaning pin 32 would be a trap laid for the one argument anybody types from memory. ## The toggle is the host's, not the editor's New `Host.VTable.pull_gpio_toggle`, and a `GpioFn` in the p4 ABI (hence version 2), rather than board_memory reaching for GPIO_OUT the way `Poke` two functions above it would happily do. Writing that register is not the job. A pad has to be pointed at the GPIO function in the IO MUX, routed in the GPIO matrix, given drive strength and an input buffer with its pulls cleared, and only then driven - four register files behind a per-pin table. That code already exists in `05-zig-p4/src/hal/gpio.zig`, it is the same `configureOutput` the blink demo has always used, and its register numbers are checked against ESP-IDF's own headers on the die by `zig build diff`. A second copy inside the editor object would be a second copy under no test, and getting it wrong on a pin that boots as something else is how you lose the console you are typing on. Reported levels are the OUTPUT bits, before and after, because that is what a toggle means: the level this board is driving. A pad's input buffer on an unconnected header pin reads the air. ## JP1, read off the schematic rather than remembered The diagram is the vendor's own wiring, from sheet 2 "Expand IO" of `01-esp32p4-m3/docs/JC-ESP32P4-M3_schematic.pdf` - the only document that carries this mapping. The specification PDF's "Interface Description" page turned out to be a marketing render, and there is no board user guide; the chip datasheet has a package pinout, which is not a header. That sheet is a 872x1168 raster (`pdfimages -list` - the PDF embeds no vectors, so rendering it larger adds nothing), and at that size the rows around pin 14 are genuinely ambiguous by eye. So the mapping came from the drawing's geometry instead: thirteen wires leave each side of the symbol, a net wire runs ~100 px to its label and a power stub ~21 px. Pin 8's wire is 21 px, which is what identifies it as unconnected rather than as the first of the GPIO4x labels - the reading that had GPIO47 one row higher and shorted GPIO45 to the ground bracket. Cross-checked against a second source that has been in the tree all along: `05-zig-p4/build.zig` documents `-Dled=20` as "JP1 pin 17", and GPIO20 lands on pin 17 here. Both facts are asserted in the test, so the diagram cannot drift from either. ## Peek, Poke, Hexdump and Gpio are now the P4 build's alone `board_memory.enabled` was `os.tag == .freestanding and !isWasm()`, on the argument that these words are a property of having no operating system rather than a product configuration, and that a predicate spelled out of `builtin` cannot drift the way a hand-maintained enum can. Tidy, and it answered the wrong question. A word only exists if some shell offers it, and the shells are the platforms. `Gpio` settles it beyond argument: its whole content is one board's header, and a second freestanding port would need its own pinout rather than inheriting this one. "Bare metal" was never the requirement, "this board" was, and the two only looked identical because there is currently one of them. The old predicate's real work was excluding wasm - `freestanding` too, where an address is an offset into a linear memory the engine owns - and naming `p4` excludes it by construction instead of by a term somebody has to keep remembering. The target is now the witness rather than the gate. Absent means not compiled: the tty binary contains no `+Gpio`, no `+Hexdump`, no `ES_I2C_SDA` and no `MisalignedAddress`. ## The boot buffer's lines are checked, not eyeballed Three times now a line in that tour has been one or two characters too long for a 56-column grid, and every time it was found by reading the die's screen - the expensive way to measure a string literal. The text is a named `boot_buffer` with a test over it, six lines came down to fit with margin, and the tour gained `Gpio`. Tests: the pinout's width, its thirteen aligned pin rows, GPIO20-on-17 and pin-8-unconnected; the decimal-versus-hex distinction; every boot-buffer line. Full suite green - unit-test, snap 95/95, hxdiff 481/0, hxparity 561/0, image-harness, pdf-harness, mupdf-check - and tty, p4, gui, p4 at 80x24, p4 with the fade forced on. On the die `p4-bench --check` is 5/5, the fifth being a new one: three `Gpio 33` runs must report 0->1, 1->0, 0->1, because the alternation is the only oracle a hardcoded string could not fake.
* Raise the board's grid to 56x14, and make it a build optionGabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The 40x12 ceiling was never about the screen. It was about memory, and the comment above `max_cols` said so: "every cell is paid for four times over: vaxis keeps a Screen and an InternalScreen, pardes keeps its own Surface and previous_cells". Two of those four are now dead weight - with `direct_emit` the emitter diffs the Surface against its own shadow and writes the escapes itself, so vaxis's two grids are allocated, never read, and were the largest single claim on a 384 KiB heap. `init` sizes them to ONE CELL. vaxis still does the work only it can do: the alternate screen, the capability queries, and parsing everything that comes back. That removes the memory ceiling entirely - the heap now reports 336 KB free at every geometry tried, including ones that used to fail - and leaves latency as the only limit, which is the honest one: every frame walks the whole grid. ## Measured on the die, 0.87 us per cell geometry cells round trip 40x12 480 3,628 us the old default 56x14 784 3,930 us the new one 56x16 896 3,965 us 60x18 1,080 4,114 us 64x20 1,280 4,281 us 80x24 1,920 4,809 us 100x30 3,000 5,743 us 120x36 4,320 6,923 us 140x42 5,880 8,310 us the largest that runs 160x48 7,680 links, then traps 200x60 12,000 does not link 56x14 is 63% more area and 40% more width than 40x12 and still holds the 4 ms this port was built to. 56x16 was tried first: 3,965 us on the bench instrument but 4,029 on the phase-randomised one, which is over, and the two instruments differ by about 50 us systematically - so the wider grid went and two rows stayed behind. Width is worth more than height for reading code. The two failures at the top are worth naming precisely because they are different failures. 200x60 does not link: `.bss will not fit in region l2mem, overflowed by 76036 bytes`, that `.bss` being the shell's shadow copy of the grid, sized at comptime. 160x48 links and then TRAPS at boot - the same region pressure arriving at runtime as a collision rather than as a diagnostic. Neither is a heap problem any more, which is the interesting part: the heap has 336 KB spare while `.bss` runs out. `-Dp4-cols` / `-Dp4-rows` because none of the above is a constant. 80x24 is one flag away for anyone who would rather have the classic terminal than the millisecond. Verified at the new geometry rather than assumed: the A/B against the reference path - vaxis rendering, full repaint, `shadow_grid` and `direct_emit` both off - is identical in every cell, characters and resolved style. That matters more here than usual because the emitter's column arithmetic has a special case at the last column, and 40 was the only width it had ever been asked about. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, both A/B arms, tty/p4/gui, and the board's own `p4-bench --check`.
* Hold a lone ESC: every escape sequence on this wire was being shreddedGabriel Schneider2026-08-26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The mouse did not work. Chasing that found something much larger: NO escape sequence worked on this transport, and had not since the port began. `vaxis.Parser` resolves a buffer containing nothing but 0x1b as the Escape KEY. That is deliberate and correct for a terminal, where the kernel hands over a whole escape sequence in a single read, so a solitary ESC really does mean somebody pressed Escape. A 115200 serial line hands over ONE BYTE AT A TIME - 87 us apart, an eternity to a loop running at 360 MHz - so the first byte of every sequence arrived alone and was resolved as Escape, and the remaining bytes arrived as ordinary keys. A mouse click therefore came through as TEN key presses: Escape, `[`, `<`, `0`, `;`, `1`, `8`, `;`, `3`, `M`. The `0` among them is "go to column zero" in normal mode, which is exactly where the cursor kept landing, and why the first attempt at this looked like a coordinate bug. Arrow keys, function keys, and the host bridge's in-band resize reports were all being taken apart the same way. Longer partial sequences were never affected: the CSI scanner returns `n == 0` for "no final byte yet" and the shell already keeps those bytes. Only the one-byte case needed an answer, because it is the only one the parser answers WRONGLY instead of declining. So the shell holds a buffer that is exactly one ESC and lets `pardes_p4_tick` release it after 10 ms - two orders of magnitude longer than the 87 us until the next byte of a real sequence, and imperceptible to a person pressing Escape. The same trade every terminal editor makes, for the same reason. Finding it took instrumenting the ABI: printing `@tagName` of every event the shell applied. Ten `key_press` where one `mouse` belonged is not a thing any amount of reading the coordinate arithmetic would have shown, and I had already read it twice. ## Mouse reporting, and the 1003 that is not requested With the sequences intact, `apply` already handled `.mouse` - it mirrors the tty shell - so enabling reporting was the only missing piece. Spelled out here rather than taken from `vx.setMouseMode`, which asks for `1002;1003;1004;1006`: 1003 is ANY-MOTION tracking, a report per cell the pointer crosses with no button held. On a 115200 line that is dozens of 15-byte reports for one sweep, arriving as input the editor must parse while it paints, and arriving whether or not anyone wants it - moving the mouse over the window would starve typing. 1002 reports presses, releases and motion while a button is held, which is exactly what a click and a drag-select need. Verified on the die: a click at column 12 puts the cursor at column 12 and one at column 22 puts it at column 22, a drag paints a selection, and the wheel scrolls. A press alone paints the new position and then reverts - the caret does not move until the gesture ends - so the release is what commits it, which cost an hour of believing a working click was broken. Screen byte-identical to the vaxis reference, round trip median 3682 us against 3682, snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty/p4/gui all build.
* The frame diff was comparing byte at a time; compare words, and pay the ↵Gabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | bridge on every frame Two findings, both in the P4 shell's own `present`. ## std.mem.eql was the largest read in the firmware, one byte at a time The shadow-grid diff compares each row against the previous frame: two 13 KB streams, every frame, and by far the biggest memory access the firmware makes. It measured 3.2 cycles per byte, which is about four times what word-wide loads need - the shape of a byte-at-a-time loop, and `std.mem.eql` is what it was. `sameBytes` compares a `u32` at a time and falls back to the byte loop when the spans are not aligned for it. The alignment test has to be a RUNTIME one because `Cell` is all `u8` fields and therefore has alignment 1: whether a row begins on a word boundary is a property of whoever allocated the Surface, not of the type. A row is 40 cells of 26 bytes, divisible by four, so an aligned base makes every row aligned. The answer is bit-for-bit the same - this is still exact byte equality - so it keeps the property the whole diff rests on: byte equality implies visual equality, so the diff can never claim two different cells are the same. Measured on the die: the grid walk 223 -> 66 us, 0.95 cycles per byte. 157 us off every keystroke at every document length, and the single largest win since the clock raise. ## A frame that only hides the cursor still has to fill a USB packet The padding added for the bridge's 32-byte bulk-IN packet covered the branch that positions the cursor and not the branch that hides it. A frame that only hid the cursor was six bytes and waited out the bridge's timer. Hiding an already-hidden cursor is as idempotent as positioning it twice, so it pads the same way. The packet size is no longer inferred from an experiment either: 32 is `wMaxPacketSize` of endpoint 0x82 as the device reports it, and the sweep over pad targets confirms what it implies - 0 and 16 sit at 4.7-5.1 ms, while 32, 48 and 64 all sit at 3.6-3.8 ms. Crossing the boundary is worth about 950 us; going past it buys nothing. ## Result length 0 20 40 80 160 320 640 chars RTT 3602 3624 3638 3790 3868 4026 4192 us Fixed cost 3652 us against a 4 ms target, from 16.99 ms where this started. A phase-randomised instrument agrees over 80 trials: median 3687 us, minimum 3571, maximum 3912 - every trial under 4 ms. The two lengths still above 4 ms are the ones where the line has outgrown the viewport, so the cursor is off screen and the keystroke changes NOTHING: the frame is 36 bytes of cursor-hide and padding, zero cells changed, while pardes still rebuilds all 480 cells of the Surface for 858-985 us. That is the one architectural item left and it is not a micro -optimisation: nothing in this repository can avoid work pardes has already done. Verified: screen byte-identical to the vaxis reference on the 18-step workload, with canonical style decoding rather than escape history. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, both A/B arms build, tty, p4 and gui all build.
* Drop the dead std.fmt fallback in the emitter's integer writerGabriel Schneider2026-08-25
| | | | | | | | | | Five digits is every u16, so the `v >= 10000` branch could never be taken and it was dragging `std.fmt.printInt` into a firmware whose whole reason for hand-rolling this was to keep the format machinery out of the hottest sequence it emits. Behaviour is identical, and re-verified rather than assumed: screen byte-identical to the vaxis reference on the 18-step workload, round trip median 3830 us over 60 trials (3829 before), snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test.
* Emit the ANSI directly, and pay the USB bridge its minimum frameGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | `present` already knows exactly which cells moved - that is what the shadow grid is for - and then handed every one of them to vaxis so that vaxis could work it out again against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of it redundant. This emits the escapes itself and skips it. The emitter is small because it is allowed to be: one absolute CUP per run of changed cells rather than per cell, absolute SGR rather than a delta from whatever is currently on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style change and buys the property that no cell can inherit an earlier cell's colour if a frame is cut short. Cursor column tracking gives up after anything that is not a single printable ASCII byte, and at the last column, because deferred wrap makes the answer terminal-dependent and wrong by a whole row. Board cost: `render` 1362 -> 779 us. Bytes per keystroke: 81 -> 21. Image 26.6 KB smaller, since vaxis's renderer is now unreachable. ## And it measured SLOWER 4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of result that means the model is wrong, so I stopped optimising and went looking. It is the USB bridge. The board talks to the host through a CH340, a full-speed part whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal timer gives up waiting for more - about a millisecond, a quarter of the whole budget. Routing through vaxis only looked competitive because its frames are 81 bytes and fill a packet by accident. The evidence, all at identical board cost and with a byte-identical screen: frame min median 21 B 3843 us 4817 us never fills a packet 49 B 3719 us 3814 us padded past the boundary 81 B 4373 us 4475 us vaxis, fills one by accident Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the compute that was saved. Only the median was hostage to the timer. So the frame has a minimum size and it belongs to the transport, not the terminal. Pad to it, with repeated absolute cursor positioning: idempotent, already the sequence the frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has a minimum and the sender pays it - and it is a real trade rather than free, since the filler is wire time that delays a later frame. It only applies when the frame is small, which is when there is wire to spare. ## Result: 3.74 ms, and the goal was 4.00 step fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x + byte compare 8.37 ms 7.1 us 9.48 ms 0.37x + 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x + direct emit 3.74 ms 2.0 us 4.06 ms 0.16x 35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials, median 3829 us, min 3722, p90 3930. That second instrument exists because of this commit. The original bench sends keystrokes on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and makes the round trip a staircase in board time - a real saving can measure as a regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two. It was not what was happening here, but it had to be excluded before the CH340 could be believed, and it is the right default for anything measured across this link. ## Verification `direct_emit = false` routes every cell back through vaxis and is the reference. Both arms, same 18-step workload, same clock: identical characters and identical resolved style in every cell - resolved, not raw SGR, because two emitters reaching the same colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly the change that can be right about latency and wrong about the screen, and until the verifier compared canonical style rather than escape history it could not have told the difference. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, both A/B arms build, tty, p4 and gui all build.
* Compare a whole row with one memcmp before looking at cellsGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | `Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes. Byte equality implies visual equality, which is what makes the shortcut sound: a row that compares equal cannot be hiding a changed cell, and a row that differs only in padding falls through to the per-cell path, which is correct and merely slower. Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte, this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there is little left in it. It is also why board compute scaled 2.6x rather than 4x when the core clock went up 4x. Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and incremental path, same 18-step workload, same clock - identical characters and identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty and p4 both build.
* Compare shadow-grid cells as bytes, not through std.meta.eqlGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | `Cell.visuallyEqual` is the semantically exact answer and too slow to ask 480 times a frame: `std.meta.eql` on a `CellStyle` recurses through a colour union and eight booleans, and the walk measured 1.45 ms on the die - about 270 cycles to compare a 28-byte struct. `sameCell` in src/p4.zig does it as bytes. That is safe in the direction that matters: byte equality IMPLIES visual equality, so it can never claim two different cells are the same. It can miss an equality - scratch bytes past `len`, or padding - and the only cost of that is one redundant `writeCell` which vaxis then diffs away. Defaults are still compared by meaning, because an unpainted cell's text and style are whatever the previous frame left in them. Measured: the grid walk 1.45 -> 0.98 ms, a keystroke 8.87 -> 8.37 ms fixed. Verified the way a rendering change has to be. The A/B harness now hashes the SGR state of every cell as well as its character, because the first version compared text only and would have passed a colour regression in silence. Reference path (`shadow_grid = false`, clear and write everything) and incremental path were each run against the same 19-step workload on the die and the reconstructed screens are identical in both text and per-row style hash. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, and tty / p4 / gui all build.
* Make a keystroke 2.6x cheaper by not asking Unicode about ASCIIGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf -Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it was Unicode machinery answering questions about the letter `y`. Four changes, each a fast path guarded so that non-ASCII text takes exactly the road it took before. `modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates graphemes FROM THE START of the text with the full UAX #29 break state machine until it passes the offset, and the render path calls it once per visible row with a column offset - so the cost followed the cursor's distance along its line. That is the shape measured on the die, where inserting at column 320 of a fixed 320-character line cost 7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator - is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII, and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate. `Surface.print` then became the largest at 26.2%: per character it took a UTF-8 length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII takes none of that now. Same guard, same reason. `file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth` about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one printable cell and `gwidth` stays the authority on them. `modal.lineSlice` searched for "\n" with the generic substring search where a memchr does; it is called once per visible row per frame. Measured at the P4's geometry and configuration, on the host: render 55 -> 12 us, key-down 483 -> 24 us, key-right 327 -> 13 us, edit-char 205 -> 46 us. On the die, the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a keystroke at a 160-character line from 25.56 ms to 15.36 ms. ## The shadow grid, and why it is static `src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured 6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything changed: a second render with nothing new cost the same as the first. vaxis diffs its own grid, but only after being told every cell, and being told is the expensive part. So `present` now keeps the previous Surface and tells vaxis only what moved. `Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -> 1.45 ms. The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is not a micro-optimisation. The first version allocated it from the editor's heap; on a board whose 384 KiB is nearly spoken for, that is exactly the kind of change that works and then breaks something else three steps away. `shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present` behaves as it did before - clear and write every cell - which is the reference any measurement should be compared against, and the way to tell a rendering bug from a rendering difference. It earned its keep immediately: the two paths were run against the same 19-step workload on the die - inserts, deletes, motions that move the modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the reconstructed screens are byte-identical. ## Verification `snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0 mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 / gui all build. The rendering changes are exactly the sort that pass a latency benchmark while corrupting a screen, so the snapshot parity suite is the one that matters here and it is unchanged. `test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the things that table exists to hold constant, and 40x12 is not a scaled guess at the board - it is the board. `--only` exists because under `perf record` one 63 ms cell on the largest fixture swamps every sample from the case being asked about. ## Found, not fixed `vx.resize` fails on this board: a runtime geometry change hits its allocation failure path, restores the previous size and returns, so 80 bytes go out where 1,392 should. Verified independent of everything above - it reproduces with `shadow_grid` false. The board therefore has one geometry for the life of a session, which is why the staleness test above compares two firmwares rather than resizing one.
* A fourth platform: pardes as ESP32-P4 firmware, bytes in and bytes outGabriel Schneider2026-08-25
`-Dplatform=p4 -Dtarget=riscv32-freestanding` emits a single freestanding OBJECT exporting a seven-function C ABI, not an executable. The board's toolchain (../05-zig-p4) owns `_start`, the linker script and the UART driver and links this in. The seam is bytes rather than types, so neither side can accidentally depend on the other's internals, and a signature that drifts fails at link time. The serial line is the whole of the I/O. `src/p4.zig` drives vaxis unchanged over it: the renderer is a byte writer and `queryTerminalSend` is a byte writer, so the terminal emulator on the host answers the capability handshake and the firmware sees a real terminal. Measured going out over the wire on attach: alt screen, in-band resize, cursor report, kitty keyboard, kitty graphics, DA1. THREE WORDS EXIST ONLY HERE. `src/board_memory.zig` implements `Peek`, `Poke` and `Hexdump`, gated on `builtin.os.tag == .freestanding and !isWasm()` - derived from the TARGET, because they are a property of running with no OS under you rather than a product option, and because wasm is freestanding too and is exactly what must be excluded: in a browser an address is an offset into the linear memory this editor's own heap lives in. Every access goes through `*allowzero volatile`: a peripheral register is not memory, and address 0 is an ordinary unmapped address on this bus. One 4 KiB cap per command, set by the console rather than the memory - an unbounded dump would wedge the only console the board has for eleven hours. Measured on ESP32-P4 rev v1.3 silicon, driven from a host terminal: Peek 0x501101a4 0x0e63ce71, then 0xaeaa6919 on a second read - the RNG register, so the volatile loads are not folded Poke 0x5011002c 0xdeadbeef LP_STORE0; a later Peek returned 0xdeadbeef Hexdump 0x5011002c 32 16 bytes a row, hex columns and an ASCII gutter Peek 0x50110001 `peek: MisalignedAddress` on the message row That last line is the one that matters. A misaligned 32-bit access traps, and a trap in firmware is a watchdog reset that takes the session with it, so the check that turns it into a message is the reason the file is hand-written rather than a generic reader. BARE METAL BOOTS AN EMPTY OUTPUT BUFFER. Every other boot layout in `init` makes a shell, and on this platform that is not a preference but an impossibility: nothing to fork, no pty to give a terminal pane. Booting one anyway produced precisely what that describes - a pane whose tag ends in `Filter`, no gutter, no buffer, and every keystroke vanishing into the Fallback's silent pty. An output buffer is also what the platform's own words want, since Peek, Poke and Hexdump each fill one. Sized for the board rather than for a desktop: * `allocators.zig` gains a p4 tier that is ALL fallback - every capacity is zero, so each arena spills immediately to the 384 KiB heap the firmware hands over, and no megabyte-shaped static reservation lands in `.bss`. * `source_manifest.zig`'s allowlist is EMPTY on p4. The table is ~0.95 MiB of rodata against a 1.5 MiB flash partition; the firmware's filesystem is the serial host's, through the Host vtable. * The grid is clamped and the clamp is measured, not guessed: every cell is paid for four times (vaxis Screen + InternalScreen, pardes Surface + previous_cells), so 40x12 fits and 80x24 exhausts the heap during `Pardes.init`. * `Vaxis.resize` deinits both screens before allocating replacements, so a failed resize leaves vaxis rendering nothing. The p4 shell keeps the previous geometry on failure instead of leaving a half-applied one. Also here: `output_pane_integration_test.zig` had an exhaustive switch over `Platform` that adding `.p4` left unhandled, which broke `zig build unit-test` outright - the native test binary is the one consumer no platform build compiles. 346 tests pass again.