| Commit message (Collapse) | Author | Age |
| |
|
|
|
|
| |
Consolidate pane, layout, memory and host code. Serve 9P by default over Unix sockets, with runtime mounts and optional TCP/QUIC transports. Remove FUSE and obsolete proof-of-concept examples.
Fix highlighting and terminal-history performance, expand differential and stress-test infrastructure, sort navigation results while preserving the next occurrence, add syntax-colored Braille minimaps, remove SPC-k, and document 9P interaction as a repository skill.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
board cap on one screen
## The wire is the effect stream, not a new protocol
`pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns
a terminal and a socket and nothing else. N frontends on one core all look at the same screen —
`screen -x`, not N sessions.
The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per
message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the
core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is
a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin
command line, and a command line cannot carry a frame.
ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be
riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit
little-endian fixed width and no message is a blit of a native struct. A protocol that only works
between two builds of the same compiler would have thrown away the one frontend that motivated it.
## The board comes in; its toolchain stays out
`src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over-
serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves
into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so
`zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the
board from this repo's `build.zig`.
The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes
core over a serial line. What stayed is everything a second project would also want — the HAL, the
register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its
own and its `build()` early-returns when it is not the root package, so this costs the package
graph exactly zero packages and the editor's own builds nothing at all.
## limits.zig: nine forgettable places become one budget
Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions —
they are ONE decision, how much memory this build may spend, taken nine times where no reader could
see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against
what it is measured against, derived from two booleans.
The payoff is testability on a machine that is not the board: the caps are ordinary comptime values,
so a host build can be compiled against the board's numbers and the parking, eviction and clamping
paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART.
## A bare `zig build`
`zig build` with no arguments now builds the tty and GUI binaries and installs them into
`~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built
one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and
"use it" two different commands for no reason.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the
shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell
loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes.
Byte equality implies visual equality, which is what makes the shortcut sound: a row
that compares equal cannot be hiding a changed cell, and a row that differs only in
padding falls through to the per-cell path, which is correct and merely slower.
Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and
the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte,
this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there
is little left in it. It is also why board compute scaled 2.6x rather than 4x when the
core clock went up 4x.
Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and
incremental path, same 18-step workload, same clock - identical characters and
identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0,
unit-test, tty and p4 both build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in
that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf
-Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it
was Unicode machinery answering questions about the letter `y`.
Four changes, each a fast path guarded so that non-ASCII text takes exactly the road
it took before.
`modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates
graphemes FROM THE START of the text with the full UAX #29 break state machine until
it passes the offset, and the render path calls it once per visible row with a column
offset - so the cost followed the cursor's distance along its line. That is the shape
measured on the die, where inserting at column 320 of a fixed 320-character line cost
7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII
scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule
that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator -
is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII,
and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate.
`Surface.print` then became the largest at 26.2%: per character it took a UTF-8
length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a
width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII
takes none of that now. Same guard, same reason.
`file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth`
about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one
printable cell and `gwidth` stays the authority on them.
`modal.lineSlice` searched for "\n" with the generic substring search where a memchr
does; it is called once per visible row per frame.
Measured at the P4's geometry and configuration, on the host: render 55 -> 12 us,
key-down 483 -> 24 us, key-right 327 -> 13 us, edit-char 205 -> 46 us. On the die,
the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a
keystroke at a 160-character line from 25.56 ms to 15.36 ms.
## The shadow grid, and why it is static
`src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured
6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything
changed: a second render with nothing new cost the same as the first. vaxis diffs its
own grid, but only after being told every cell, and being told is the expensive part.
So `present` now keeps the previous Surface and tells vaxis only what moved.
`Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -> 1.45 ms.
The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is
not a micro-optimisation. The first version allocated it from the editor's heap; on a
board whose 384 KiB is nearly spoken for, that is exactly the kind of change that
works and then breaks something else three steps away.
`shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present`
behaves as it did before - clear and write every cell - which is the reference any
measurement should be compared against, and the way to tell a rendering bug from a
rendering difference. It earned its keep immediately: the two paths were run against
the same 19-step workload on the die - inserts, deletes, motions that move the
modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the
reconstructed screens are byte-identical.
## Verification
`snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0
mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 /
gui all build. The rendering changes are exactly the sort that pass a latency
benchmark while corrupting a screen, so the snapshot parity suite is the one that
matters here and it is unchanged.
`test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the
things that table exists to hold constant, and 40x12 is not a scaled guess at the
board - it is the board. `--only` exists because under `perf record` one 63 ms cell on
the largest fixture swamps every sample from the case being asked about.
## Found, not fixed
`vx.resize` fails on this board: a runtime geometry change hits its allocation
failure path, restores the previous size and returns, so 80 bytes go out where 1,392
should. Verified independent of everything above - it reproduces with `shadow_grid`
false. The board therefore has one geometry for the life of a session, which is why
the staleness test above compares two firmwares rather than resizing one.
|
| |
|
|
| |
+ snapshot refresh
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Every keystroke in a shell pane rebuilt the motion surface from scratch:
shellRows dumped ghostty's WHOLE history+active grid, split it, blanked the
prompt rows and handed back slices into the scratch arena, which the next
update threw away. A pane sitting on a multi-megabyte agent transcript paid an
O(scrollback) dump per press of `j`, and paid it once or twice per key, since
flatSurface then rebuilt the same rows joined by '\n' beside it.
The dump is now memoized against the pane it was built for (term_pane.RowsCache
on Pardes.shell_rows), gpa-owned rather than scratch-arena because the whole
point is to outlive the update that built it. One entry, not a table: the
surface is built for the pane the cursor is in, and a second pane asking would
only double a multi-megabyte buffer for a slot it is about to lose again. A
pane that is not the live one is answered from the arena as before.
The lifetime rule is the part that would have rotted silently, so it is one
rule and it is written down: `rows` is handed out to callers, so everything
that notices the entry has gone bad — output arrived, the grid reflowed, the
pane died, another pane wants the slot — only marks it `stale`, and the
buffers are freed in exactly two places, `sweep` at the TOP of an update
before any handler can be holding them, and `reset` when the editor goes away.
Nothing frees mid-update. dropPane clears the pointer immediately though: a
freed pane's address comes back from the allocator as a different pane, and an
entry still naming it would answer for the wrong grid.
Two things fall out of having the join already:
- flatSurface returns the memo's `text` verbatim when the lines it was handed
are the cached rows untouched, instead of rebuilding the join.
- paneCursorLines returns `rows` directly when there is no edit buffer, where
it used to copy the array one slice at a time to produce exactly what it was
given.
One bug on the way past, in the same function: an EMPTY edit buffer writes one
line but modal.lineCount("") is 0, so `ls` was sized one short of what the
loop writes — the same floor the paste site needs. Killing a whole line
(`A<C-u>`, `d%`) on a buffer covering the last row made that a length of zero.
And test/perf.zig grows the axis that would have caught this: a terminal
scoreboard beside the file one, three scrollback fixtures (64 KiB, 1 MiB,
8 MiB — half the ceiling) against render / output / resize-rows / resize-cols /
key-down / edit-char, sharing the existing text and JSON reports and the
--base comparison. resize-cols and resize-rows are both there because a COLUMN
change reflows every page in the list and a row change does not.
Measured on that table: key-down is 142 / 630 / 636 us across the three
fixtures — flat from 1 MiB to 8 MiB, which is the dump being gone, and render
flat at ~110 us throughout. What remains of key-down's step at 1 MiB is the
linear scan indexOf refuses to index for a terminal; that is now a ponytail
waiver naming its own price (615 us against 140 us) and the threading through
paneOff/panePos/paneLineStart it would cost, to be done the day 0.6 ms shows
up next to something anybody can feel.
|
| | |
|
|
|
zig build perf drives the core directly — event, effects, one frame, no pty —
over four generated fixtures: 1k lines, 50k, 300k, and 400 lines of 8000
columns, because a file that is long and a file that is wide fail differently.
Every sample seeks somewhere else in the file first, since measuring at line 3
of a 300k-line file hides exactly the bug.
perf record said half the run was scanning for newlines from byte 0. So File
carries a line index, built on demand and invalidated in exactly ONE place —
setContent, the funnel every content swap already goes through. That killed the
scrollbar's per-frame line count (12.6% of the whole run by itself), scrollBy,
ensureCursorVisible, lastNavRow, the syntax window bounds and two O(scroll)
walks. normalKey computed max_line as a const at the top: two full passes over
the buffer on every keystroke of every kind, for three g/G branches. It is lazy
now. The modal primitives each walked the text twice for the same line.
And the visible window was re-parsed on every scrolled row — a third of a
megabyte per keypress on the wide fixture. The highlighted range is remembered,
a scroll inside it is free, and only a re-parse that FOLLOWS a scroll takes
slack: doing it unconditionally made typing 2.1x slower, since every character
paid for a band it could never amortise.
One j on a 19 MB file: 37.8ms -> 266us. Render: 4.1ms -> 77us. Open costs 1.25x
more for the one extra pass, which buys 54x on every frame after, and 8 bytes
per line of memory.
Left standing, measured and named: edit-char is 14ms on 19MB because content is
immutable and every keystroke copies the buffer. A third of that is the index
rebuild, which could be a shift if setContent knew the edit offset; the rest
wants a rope. bodyText's double copy and Surface.print's per-cell decode never
rose above 2% of the profile afterwards, so they were left alone.
No golden moved.
|