| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Every keystroke in a shell pane rebuilt the motion surface from scratch:
shellRows dumped ghostty's WHOLE history+active grid, split it, blanked the
prompt rows and handed back slices into the scratch arena, which the next
update threw away. A pane sitting on a multi-megabyte agent transcript paid an
O(scrollback) dump per press of `j`, and paid it once or twice per key, since
flatSurface then rebuilt the same rows joined by '\n' beside it.
The dump is now memoized against the pane it was built for (term_pane.RowsCache
on Pardes.shell_rows), gpa-owned rather than scratch-arena because the whole
point is to outlive the update that built it. One entry, not a table: the
surface is built for the pane the cursor is in, and a second pane asking would
only double a multi-megabyte buffer for a slot it is about to lose again. A
pane that is not the live one is answered from the arena as before.
The lifetime rule is the part that would have rotted silently, so it is one
rule and it is written down: `rows` is handed out to callers, so everything
that notices the entry has gone bad — output arrived, the grid reflowed, the
pane died, another pane wants the slot — only marks it `stale`, and the
buffers are freed in exactly two places, `sweep` at the TOP of an update
before any handler can be holding them, and `reset` when the editor goes away.
Nothing frees mid-update. dropPane clears the pointer immediately though: a
freed pane's address comes back from the allocator as a different pane, and an
entry still naming it would answer for the wrong grid.
Two things fall out of having the join already:
- flatSurface returns the memo's `text` verbatim when the lines it was handed
are the cached rows untouched, instead of rebuilding the join.
- paneCursorLines returns `rows` directly when there is no edit buffer, where
it used to copy the array one slice at a time to produce exactly what it was
given.
One bug on the way past, in the same function: an EMPTY edit buffer writes one
line but modal.lineCount("") is 0, so `ls` was sized one short of what the
loop writes — the same floor the paste site needs. Killing a whole line
(`A<C-u>`, `d%`) on a buffer covering the last row made that a length of zero.
And test/perf.zig grows the axis that would have caught this: a terminal
scoreboard beside the file one, three scrollback fixtures (64 KiB, 1 MiB,
8 MiB — half the ceiling) against render / output / resize-rows / resize-cols /
key-down / edit-char, sharing the existing text and JSON reports and the
--base comparison. resize-cols and resize-rows are both there because a COLUMN
change reflows every page in the list and a row change does not.
Measured on that table: key-down is 142 / 630 / 636 us across the three
fixtures — flat from 1 MiB to 8 MiB, which is the dump being gone, and render
flat at ~110 us throughout. What remains of key-down's step at 1 MiB is the
linear scan indexOf refuses to index for a terminal; that is now a ponytail
waiver naming its own price (615 us against 140 us) and the threading through
paneOff/panePos/paneLineStart it would cost, to be done the day 0.6 ms shows
up next to something anybody can feel.
|
|
|
zig build perf drives the core directly — event, effects, one frame, no pty —
over four generated fixtures: 1k lines, 50k, 300k, and 400 lines of 8000
columns, because a file that is long and a file that is wide fail differently.
Every sample seeks somewhere else in the file first, since measuring at line 3
of a 300k-line file hides exactly the bug.
perf record said half the run was scanning for newlines from byte 0. So File
carries a line index, built on demand and invalidated in exactly ONE place —
setContent, the funnel every content swap already goes through. That killed the
scrollbar's per-frame line count (12.6% of the whole run by itself), scrollBy,
ensureCursorVisible, lastNavRow, the syntax window bounds and two O(scroll)
walks. normalKey computed max_line as a const at the top: two full passes over
the buffer on every keystroke of every kind, for three g/G branches. It is lazy
now. The modal primitives each walked the text twice for the same line.
And the visible window was re-parsed on every scrolled row — a third of a
megabyte per keypress on the wide fixture. The highlighted range is remembered,
a scroll inside it is free, and only a re-parse that FOLLOWS a scroll takes
slack: doing it unconditionally made typing 2.1x slower, since every character
paid for a band it could never amortise.
One j on a 19 MB file: 37.8ms -> 266us. Render: 4.1ms -> 77us. Open costs 1.25x
more for the one extra pass, which buys 54x on every frame after, and 8 bytes
per line of memory.
Left standing, measured and named: edit-char is 14ms on 19MB because content is
immutable and every keystroke copies the buffer. A third of that is the index
rebuild, which could be a shift if setContent knew the edit offset; the rest
wants a rope. bodyText's double copy and Surface.print's per-cell decode never
rose above 2% of the profile afterwards, so they were left alone.
No golden moved.
|