| Commit message (Collapse) | Author | Age |
| |
|
|
|
|
|
|
|
|
| |
Five digits is every u16, so the `v >= 10000` branch could never be taken and it was
dragging `std.fmt.printInt` into a firmware whose whole reason for hand-rolling this
was to keep the format machinery out of the hottest sequence it emits.
Behaviour is identical, and re-verified rather than assumed: screen byte-identical to
the vaxis reference on the 18-step workload, round trip median 3830 us over 60 trials
(3829 before), snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`present` already knows exactly which cells moved - that is what the shadow grid is
for - and then handed every one of them to vaxis so that vaxis could work it out again
against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of
it redundant. This emits the escapes itself and skips it.
The emitter is small because it is allowed to be: one absolute CUP per run of changed
cells rather than per cell, absolute SGR rather than a delta from whatever is currently
on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it
writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style
change and buys the property that no cell can inherit an earlier cell's colour if a
frame is cut short. Cursor column tracking gives up after anything that is not a single
printable ASCII byte, and at the last column, because deferred wrap makes the answer
terminal-dependent and wrong by a whole row.
Board cost: `render` 1362 -> 779 us. Bytes per keystroke: 81 -> 21. Image 26.6 KB
smaller, since vaxis's renderer is now unreachable.
## And it measured SLOWER
4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of
result that means the model is wrong, so I stopped optimising and went looking.
It is the USB bridge. The board talks to the host through a CH340, a full-speed part
whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet
is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal
timer gives up waiting for more - about a millisecond, a quarter of the whole budget.
Routing through vaxis only looked competitive because its frames are 81 bytes and fill a
packet by accident.
The evidence, all at identical board cost and with a byte-identical screen:
frame min median
21 B 3843 us 4817 us never fills a packet
49 B 3719 us 3814 us padded past the boundary
81 B 4373 us 4475 us vaxis, fills one by accident
Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the
compute that was saved. Only the median was hostage to the timer.
So the frame has a minimum size and it belongs to the transport, not the terminal. Pad
to it, with repeated absolute cursor positioning: idempotent, already the sequence the
frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper
so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has
a minimum and the sender pays it - and it is a real trade rather than free, since the
filler is wire time that delays a later frame. It only applies when the frame is small,
which is when there is wire to spare.
## Result: 3.74 ms, and the goal was 4.00
step fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
+ byte compare 8.37 ms 7.1 us 9.48 ms 0.37x
+ 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x
+ direct emit 3.74 ms 2.0 us 4.06 ms 0.16x
35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials,
median 3829 us, min 3722, p90 3930.
That second instrument exists because of this commit. The original bench sends keystrokes
on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and
makes the round trip a staircase in board time - a real saving can measure as a
regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two.
It was not what was happening here, but it had to be excluded before the CH340 could be
believed, and it is the right default for anything measured across this link.
## Verification
`direct_emit = false` routes every cell back through vaxis and is the reference. Both
arms, same 18-step workload, same clock: identical characters and identical resolved
style in every cell - resolved, not raw SGR, because two emitters reaching the same
colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly
the change that can be right about latency and wrong about the screen, and until the
verifier compared canonical style rather than escape history it could not have told the
difference.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
both A/B arms build, tty, p4 and gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the
shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell
loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes.
Byte equality implies visual equality, which is what makes the shortcut sound: a row
that compares equal cannot be hiding a changed cell, and a row that differs only in
padding falls through to the per-cell path, which is correct and merely slower.
Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and
the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte,
this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there
is little left in it. It is also why board compute scaled 2.6x rather than 4x when the
core clock went up 4x.
Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and
incremental path, same 18-step workload, same clock - identical characters and
identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0,
unit-test, tty and p4 both build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Cell.visuallyEqual` is the semantically exact answer and too slow to ask 480 times a
frame: `std.meta.eql` on a `CellStyle` recurses through a colour union and eight
booleans, and the walk measured 1.45 ms on the die - about 270 cycles to compare a
28-byte struct.
`sameCell` in src/p4.zig does it as bytes. That is safe in the direction that
matters: byte equality IMPLIES visual equality, so it can never claim two different
cells are the same. It can miss an equality - scratch bytes past `len`, or padding -
and the only cost of that is one redundant `writeCell` which vaxis then diffs away.
Defaults are still compared by meaning, because an unpainted cell's text and style are
whatever the previous frame left in them.
Measured: the grid walk 1.45 -> 0.98 ms, a keystroke 8.87 -> 8.37 ms fixed.
Verified the way a rendering change has to be. The A/B harness now hashes the SGR
state of every cell as well as its character, because the first version compared text
only and would have passed a colour regression in silence. Reference path
(`shadow_grid = false`, clear and write everything) and incremental path were each run
against the same 19-step workload on the die and the reconstructed screens are
identical in both text and per-row style hash.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
and tty / p4 / gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in
that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf
-Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it
was Unicode machinery answering questions about the letter `y`.
Four changes, each a fast path guarded so that non-ASCII text takes exactly the road
it took before.
`modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates
graphemes FROM THE START of the text with the full UAX #29 break state machine until
it passes the offset, and the render path calls it once per visible row with a column
offset - so the cost followed the cursor's distance along its line. That is the shape
measured on the die, where inserting at column 320 of a fixed 320-character line cost
7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII
scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule
that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator -
is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII,
and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate.
`Surface.print` then became the largest at 26.2%: per character it took a UTF-8
length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a
width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII
takes none of that now. Same guard, same reason.
`file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth`
about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one
printable cell and `gwidth` stays the authority on them.
`modal.lineSlice` searched for "\n" with the generic substring search where a memchr
does; it is called once per visible row per frame.
Measured at the P4's geometry and configuration, on the host: render 55 -> 12 us,
key-down 483 -> 24 us, key-right 327 -> 13 us, edit-char 205 -> 46 us. On the die,
the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a
keystroke at a 160-character line from 25.56 ms to 15.36 ms.
## The shadow grid, and why it is static
`src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured
6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything
changed: a second render with nothing new cost the same as the first. vaxis diffs its
own grid, but only after being told every cell, and being told is the expensive part.
So `present` now keeps the previous Surface and tells vaxis only what moved.
`Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -> 1.45 ms.
The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is
not a micro-optimisation. The first version allocated it from the editor's heap; on a
board whose 384 KiB is nearly spoken for, that is exactly the kind of change that
works and then breaks something else three steps away.
`shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present`
behaves as it did before - clear and write every cell - which is the reference any
measurement should be compared against, and the way to tell a rendering bug from a
rendering difference. It earned its keep immediately: the two paths were run against
the same 19-step workload on the die - inserts, deletes, motions that move the
modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the
reconstructed screens are byte-identical.
## Verification
`snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0
mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 /
gui all build. The rendering changes are exactly the sort that pass a latency
benchmark while corrupting a screen, so the snapshot parity suite is the one that
matters here and it is unchanged.
`test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the
things that table exists to hold constant, and 40x12 is not a scaled guess at the
board - it is the board. `--only` exists because under `perf record` one 63 ms cell on
the largest fixture swamps every sample from the case being asked about.
## Found, not fixed
`vx.resize` fails on this board: a runtime geometry change hits its allocation
failure path, restores the previous size and returns, so 80 bytes go out where 1,392
should. Verified independent of everything above - it reproduces with `shadow_grid`
false. The board therefore has one geometry for the life of a session, which is why
the staleness test above compares two firmwares rather than resizing one.
|
|
|
`-Dplatform=p4 -Dtarget=riscv32-freestanding` emits a single freestanding OBJECT
exporting a seven-function C ABI, not an executable. The board's toolchain
(../05-zig-p4) owns `_start`, the linker script and the UART driver and links this
in. The seam is bytes rather than types, so neither side can accidentally depend
on the other's internals, and a signature that drifts fails at link time.
The serial line is the whole of the I/O. `src/p4.zig` drives vaxis unchanged over
it: the renderer is a byte writer and `queryTerminalSend` is a byte writer, so the
terminal emulator on the host answers the capability handshake and the firmware
sees a real terminal. Measured going out over the wire on attach: alt screen,
in-band resize, cursor report, kitty keyboard, kitty graphics, DA1.
THREE WORDS EXIST ONLY HERE. `src/board_memory.zig` implements `Peek`, `Poke` and
`Hexdump`, gated on `builtin.os.tag == .freestanding and !isWasm()` - derived from
the TARGET, because they are a property of running with no OS under you rather
than a product option, and because wasm is freestanding too and is exactly what
must be excluded: in a browser an address is an offset into the linear memory this
editor's own heap lives in. Every access goes through `*allowzero volatile`: a
peripheral register is not memory, and address 0 is an ordinary unmapped address
on this bus. One 4 KiB cap per command, set by the console rather than the memory -
an unbounded dump would wedge the only console the board has for eleven hours.
Measured on ESP32-P4 rev v1.3 silicon, driven from a host terminal:
Peek 0x501101a4 0x0e63ce71, then 0xaeaa6919 on a second read - the
RNG register, so the volatile loads are not folded
Poke 0x5011002c 0xdeadbeef LP_STORE0; a later Peek returned 0xdeadbeef
Hexdump 0x5011002c 32 16 bytes a row, hex columns and an ASCII gutter
Peek 0x50110001 `peek: MisalignedAddress` on the message row
That last line is the one that matters. A misaligned 32-bit access traps, and a
trap in firmware is a watchdog reset that takes the session with it, so the check
that turns it into a message is the reason the file is hand-written rather than a
generic reader.
BARE METAL BOOTS AN EMPTY OUTPUT BUFFER. Every other boot layout in `init` makes a
shell, and on this platform that is not a preference but an impossibility: nothing
to fork, no pty to give a terminal pane. Booting one anyway produced precisely what
that describes - a pane whose tag ends in `Filter`, no gutter, no buffer, and every
keystroke vanishing into the Fallback's silent pty. An output buffer is also what
the platform's own words want, since Peek, Poke and Hexdump each fill one.
Sized for the board rather than for a desktop:
* `allocators.zig` gains a p4 tier that is ALL fallback - every capacity is zero,
so each arena spills immediately to the 384 KiB heap the firmware hands over,
and no megabyte-shaped static reservation lands in `.bss`.
* `source_manifest.zig`'s allowlist is EMPTY on p4. The table is ~0.95 MiB of
rodata against a 1.5 MiB flash partition; the firmware's filesystem is the
serial host's, through the Host vtable.
* The grid is clamped and the clamp is measured, not guessed: every cell is paid
for four times (vaxis Screen + InternalScreen, pardes Surface + previous_cells),
so 40x12 fits and 80x24 exhausts the heap during `Pardes.init`.
* `Vaxis.resize` deinits both screens before allocating replacements, so a failed
resize leaves vaxis rendering nothing. The p4 shell keeps the previous geometry
on failure instead of leaving a half-applied one.
Also here: `output_pane_integration_test.zig` had an exhaustive switch over
`Platform` that adding `.p4` left unhandled, which broke `zig build unit-test`
outright - the native test binary is the one consumer no platform build compiles.
346 tests pass again.
|