summaryrefslogtreecommitdiff
path: root/test/perf.zig
Commit message (Collapse)AuthorAge
* Zerox opens a second pane on a file pane's buffer: one text, undo and ↵Gabriel Schneider7 hours
| | | | | | | | | | | | | | | | unsaved state, a scroll and a cursor each; an edit in either is spliced into both at File.setContentSpan, O(the edit) more, and what a Save, a get, the disk watch or a rename changes on one is copied to the rest after each step Twins share the content buffer, the History and, while it is one, the line index; the other twin's cursor and scroll move past the splice as acme moves a clone's q0. Del of one twin asks nothing while another holds the text, Exit lists the text once, and a closed twin keeps a copy of its own until it is torn down. Both are in the index, each its serial, one name. An event reader on one hears no edit made through another (first cut). A dump keeps the twins as two panes on the file, not linked again. perf: edit-twin, typing beside a twin of the 50k-line file, gated within 3x of edit-char in the same run; the three baselines re-recorded for it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* A selection deep in a long line costs a render what one at its start does: 5 ↵Gabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | | | MB in, /screen goes from 14.7 s to 5 ms Every render measured a selection's and the cursor's columns from the start of their line (lineDisplayOffset walked the whole prefix twice), so a selection near the end of a 5 MB line stalled each frame, 3 s at 1 MB in and 14.7 s at 5 MB. The cost grew with the offset, with Wrap on or off. Three changes make it independent of where the selection is: - display width adds (a tab is tab_width wherever it stands), so the offset between two columns is the width of the text between them, never the prefix; - the selection's rows are clamped to what is shown: a row wholly before the selection is skipped, a start before the row's first column starts there, and an end past its last column ends there; - a head or cursor whose column lies past its visible row (the last visible row of a line owns every column after its start) is not measured, since it is not drawn. Measured on a Debug build, 5 MB line: offset 10, 1 MB and 5 MB now take 11, 6 and 5 ms per /screen. A new perf gate case, deep-sel, renders with the selection 2 MB into a line: 83 us on Debug. The baselines were re-recorded for the new harness. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* An append to a body through 9P costs its own bytes, not five passes over the ↵Gabriel Schneider27 hours
| | | | | | | | whole body: 60 KB appends go from 35.7 to 2.1 ms each Round 25 measured bulk body writes at about 8 ms a write. There is no frame wait in it: a profile of open, write, clunk in a loop found the flush of each close walking the whole body five times. dotOf, setDot and showOffset turned the cursor between offsets and rows by counting every newline from the top; setContent found the line index's changed span by comparing old and new byte for byte, and hashed the new text to see whether it was back to the saved one. The rows now come from the file's line index by binary search (checked against the counting at every offset), a splice tells setContent the span it changed, and the text is hashed only when its length is the saved text's. Measured over 9P on a Debug build: 60 KB appends 35.7 to 2.1 ms, 8 KB 5.3 to 0.8 ms, 1 KB 1.3 to 0.9 ms. What is left is two copies of the text an edit (the splice's and the undo snapshot's). The perf gate gains body-appends, 64 appends of 8 KB on the 50k-line file, and its three baselines are recorded again. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* Terminal.zig is terminal.zig: a file of functions and no fields takes a ↵Gabriel Schneider27 hours
| | | | | | | | namespace's lowercase name The last of the deferred renames, now that the tty and theme agents have landed: panes.terminal at its importers, the alias lines in Text.zig and File.zig and panes.zig's own uses following. dump.zig's Terminal struct, a dump record, is not this. The served sources list and docs/design.typ name the new file. test/perf.zig's references change, so the three perf baselines take its new harness id with their numbers as recorded. No behaviour changes. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* Mini.zig is mini.zig: a file of functions and no fields takes a namespace's ↵Gabriel Schneider27 hours
| | | | | | | | lowercase name The naming the split agreed on, deferred until now; panes.mini is its name at its importers. File.zig keeps its alias line (const Mini = panes.mini) so the file another agent is working in changes by that line alone. The served sources list and docs/design.typ name the new file. No behaviour changes; test/perf.zig's two references change, and since the perf harness is keyed by its own text, its three baselines are recorded again (under other agents' builds, so a little slower). Co-Authored-By: Claude Opus 5.5 <[email protected]>
* An open's writes in a row to one place go in as one edit: a 10 MB body write ↵Gabriel Schneider27 hours
| | | | | | | | is linear Each body or data write copied and hashed the whole buffer, so a write the mount cut in 8 KB pieces was quadratic: 10 MB took 50 s and a 1 MB insert into 10 MB 11 s, holding the editor's turn. An open's appends to body, or inserts going on at data's address, are now held and put in as one splice (one copy, one undo step, one line-starts pass) before any other request, the close, or the editor's step once the writes pause 20 ms. fs.py-driven: 2 MB 2.19 -> 0.16 s, 10 MB 50.17 -> 0.78 s, 1 MB data into 10 MB 11.49 -> 0.26 s. A body-2m case (2 MB in 256 KB writes on one open) joins the perf gate: 84023 -> 15785 us; baselines re-recorded. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* The theme's chrome is worked out once, not per grapheme; still post passes ↵Gabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | | | | | | let the GUI rest; a perf gate somtrmsz's contrast floors made ChromeTheme.fromTheme run its focus-tint and separator searches (pow calls each), and Output's RowDecoration.styleAt called it for every grapheme of every highlighted row: the 50k-line file's render went from 0.11 ms to 13 ms. Pardes.bodyChrome now keeps the theme's chrome, worked out again only when the theme differs; recolorSyntax asks for it once a decorated row, and a plain row (nearly every row of a file) never asks. ReleaseFast, medium fixture, median us, before -> after (main): render 13150 -> 76 (111), key-down 14658 -> 79 (110), wheel 13589 -> 76 (109), open 14891 -> 2095 (2055), edit-char 16703 -> 3006 (1251; the rest of that gap is editing and tree-sitter, not this). Post.animating asked for frames whenever the window had focus and any pass was ready, so a still pass kept the GUI drawing at the display's rate (Bloom: 49% of a core idle, in a hidden test window). A pass now says whether it moves on its own: the CRT (its hum and dither) and a Shadertoy file whose source reads iTime, iFrame or iDate (shader_build.readsTime; the flag rides the wire's post message); Bloom, Vignette and Grain are still. Bloom idle: 49% -> 1.2%, as with no pass. zig build perf-gate: the 50k-line file's gestures, each's fastest sample within 3x of the recorded baseline's (test/perf-baseline-<platform>- <optimize>.json, recorded from this build), run with every unit-test. On the regressed code it fails at 180x for render.
* Give a pane's body its own Text holding the cursor, selections, mode and undoGabriel Schneider27 hours
| | | | | | | | | | | acme keeps what edits a text in its Text (dat.h:171-190) and the window holds a body and a tag of that type. The cursor, the selections, the modal state and the edit-buffer undo move off Pane into Text.zig, Pane holds them as its body, and the edit and normal-mode operations take the Text they edit. Nothing changes in behaviour; this is the step that lets the tag become a second Text. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* Move looking out of pardes.zig into look.zigGabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | Pure move, no behaviour change (acme keeps this in look.c): expanding the word under a click (ExpandedWord, expandedWord, expandedSel, cursorWordSel), the n/N walk (LookFrom, lookEdge, lookStand, lookWalkPanes, lookPast, wholeRowSpan, lookSpanIn, lookWalk, landLookSpot, noteLookSource, armLookWalk), search results (Search, SearchStart, submitSearch, lookFirstHit, runSearch, searchStep, jumpResult), the look-hover preview (LookHoverWait, LookHoverPreview, FileWordSpan, PdfWordPreview, invalidateLookHover, cancelLookHover, lookHoverPane, noteLookHover, refreshLookHoverFromRaw, advanceLookHover) and lookAt with its targets (focusPaneLine, selectSpan, openPaneTarget, focusPaneByPath, clearNavigationSelection, resolveLookTarget, locationText, canonicalLookLocation, pdfLinkLocation, followPdfLink), with five tests, go verbatim to the end of look.zig after its word and target resolution. The methods become free functions taking `p: *Pardes`; their 143 call sites change from `p.lookAt(..)` to `look.lookAt(p, ..)` (tests reach them as `pardes.look.x`). Inside look.zig the moved code's `look.` prefix drops. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* Move text editing out of pardes.zig into edit.zigGabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | Pure move, no behaviour change (acme keeps this in text.c): which text a pane edits and how it maps to the screen (editText, editTextEol, setEditText, paneCursorLines, paneByteAtDisplay, pinPaneCursor, flatSurface, paneWrapWidth), insert mode (enterInsert, handleInsert, insertKey, insertTab, exitInsert, clampFileCursor), the d/c/y/p edit operations with replace, case, join, indent, comment, number, textobjects and surround, undo and redo, yank/clipboard/paste (setYank, setClipboard, ClipRequest, clipRequest, typeToTty, applyPaste, clipYank), and the pointer selections as text (PointerTextSelection, pointerTextSelection, capturePointerSelection, paneText, pointerSourceLine, selectionText, spanHas, currentSelText), with four tests, go verbatim to edit.zig. The methods become free functions taking `p: *Pardes`; calls change from `p.insertKey(..)` to `edit.insertKey(p, ..)` (pardes.zig, body_layer.zig, selection_pipe.zig, builtins.zig, test/hxdiff.zig, test/perf.zig). In executeNormalAction the `.edit => |edit|` capture becomes `|op|`, since it would now shadow the edit import. test/lspbench.zig's first anchor follows its needle into src/edit.zig. Co-Authored-By: Claude Opus 5.5 <[email protected]>
* Ask from the keyboard which neighbour a closed pane's rows go toGabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | | | | Del takes a side: `Del k` gives the closed pane's rows to the nearest expanded pane above it, `Del j` to the one below, each falling back to the other side when it has none. DelAbove and DelBelow are those two lines, with no path of their own under SPC. A bare Del started from a key -- SPC d, Enter on the tag word, a row run from an output buffer -- on a pane with expanded panes both above and below asks instead of guessing. The question is a prompt like Save's or a search's (Pane.Prompt.del_side), so it is painted on the pane's notice band by the same path, and the next key answers it before any mode sees it: k or Up, j or Down, anything else keeps the pane, as does a click. Only a key press sets Pardes.can_ask, so a click, a 9P ctl or event write, a startup line, a restore and a shell exiting all close the pane at once, the rows going where layout.absorbVWeight has always sent them. A collapsed pane is not asked about (it has only a tag row to give), and collapsed neighbours are passed over (layout.expandedNeighbor, which Collapse now uses too). removePane and absorbVWeight take the recipient; every other caller passes null. Three scripts that closed a middle pane with SPC d answer k, which is where the rows went before, and their goldens are unchanged. delask.snap covers the question, Esc, j, a clicked DelBelow and a clicked Del. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
* Fixes from three adversarial reviews, and a destructive one among themGabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The registry sweep could delete a live socket, anywhere on the filesystem. A reviewer reproduced it: a socket that is bound but has not reached listen(2) answers ECONNREFUSED exactly like a dead one -- that window is every server's startup -- and the sweep then followed the entry's symlink and unlinked whatever absolute path it named. It now follows a target only into the directory our own sockets live in and only to a `pardes-9p-*.sock` name, it re-probes immediately before deleting rather than trusting a probe that is by then several syscalls old, and a readlink that exactly filled its buffer is treated as the truncation it is. The test grew a case for an entry whose target is not ours: the entry goes, the file does not. Ctrl-V in raw tty mode was a black hole when the yank register was empty -- neither typed nor forwarded -- so vim's visual block, readline's quoted-insert and every other program's Ctrl-V simply vanished. With nothing to paste the chord belongs to the program again. The lone-ESC flush added earlier was dead code. vaxis already returns Escape for a one-byte 0x1b (`Parser.parseGround` asserts `input.len == 1`), so the carried byte it waited for can never exist; a reviewer showed a 3 ms gap and a 60 ms gap behaving identically. Removed rather than left to imply a guarantee it never provided. A shell whose editor is gone can start one again. Naming a live but unreachable session made `pardes <file>` exit 1, which let a stale environment variable lock someone out of their own editor; it falls through to an ordinary session, as it did before the variable existed. Also: the macOS ABI check for `pardes_topbar_pane_border_px` had been replaced by a duplicate of the line above it; `--startup` now fails on a leak the way every other measurement in that file does, and stops calling its maximum a p95 below twenty samples; the served README and the skill no longer tell you to write to `data` with `>`, which truncates the whole body before the write lands; `docs/v9fs.md` described the allocate-on-walk design that was rejected; and `test/fs.py` keys nesting off `PARDES_PID`, so its forwarding case stops passing only when the runner happens to be inside a live pardes. fs-test now reaches its one documented pre-existing failure instead of dying early. Suite 778/783 with the two known crashes. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
* Measure startup, and find there is nothing in it worth optimizingGabriel Schneider27 hours
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | A --startup case in the perf harness times Pardes.init and the first frame for each boot layout, with the same --json and --base shape as the others, so these are a baseline rather than a one-off reading. The core boots and paints in 0.22 ms (tty) to 0.30 ms (classic): about one percent of what starting the editor costs, and nothing to win. The other 99% is one thing, found with perf record rather than guessed at. Zig's start.main allocates through the Debug-mode DebugAllocator, which captures a six-frame stack trace per allocation, and the first capture parses and sorts the DWARF unwind tables of an 800 MB binary -- 22% of all samples sit in mem.swap under that pdq sort. It is not the dynamic loader (25 us), not static initializers (the binary has no .init_array), not paging (435 page faults), and not lockStderr. std/start.zig:694 hardcodes DebugAllocator(.{}), so there is no knob short of the build mode. Controls that pin it down: a bare std.process.Init hello-world starts in 3 ms, the ReleaseFast pardes-perf binary in 4 ms, /bin/true in 1 ms, and a Debug pardes in 22 ms. The gui binary's 107 ms first run was cold page cache; warm it matches the tty one. So the conclusion recorded in features.txt is to leave it alone. Twenty-two milliseconds is imperceptible for an editor, and the only lever is switching the default build mode, which would multiply an already ten-minute build to save eighteen milliseconds of startup. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
* Refactor panes and filesystem; replace FUSE with 9PGabriel Schneider2026-09-07
| | | | | | Consolidate pane, layout, memory and host code. Serve 9P by default over Unix sockets, with runtime mounts and optional TCP/QUIC transports. Remove FUSE and obsolete proof-of-concept examples. Fix highlighting and terminal-history performance, expand differential and stress-test infrastructure, sort navigation results while preserving the next occurrence, add syntax-colored Braille minimaps, remove SPC-k, and document 9P interaction as a repository skill.
* One core behind N frontends, the board's own runner moved in, and every ↵Gabriel Schneider2026-08-27
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | board cap on one screen ## The wire is the effect stream, not a new protocol `pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns a terminal and a socket and nothing else. N frontends on one core all look at the same screen — `screen -x`, not N sessions. The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin command line, and a command line cannot carry a frame. ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit little-endian fixed width and no message is a blit of a native struct. A protocol that only works between two builds of the same compiler would have thrown away the one frontend that motivated it. ## The board comes in; its toolchain stays out `src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over- serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so `zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the board from this repo's `build.zig`. The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes core over a serial line. What stayed is everything a second project would also want — the HAL, the register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its own and its `build()` early-returns when it is not the root package, so this costs the package graph exactly zero packages and the editor's own builds nothing at all. ## limits.zig: nine forgettable places become one budget Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions — they are ONE decision, how much memory this build may spend, taken nine times where no reader could see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against what it is measured against, derived from two booleans. The payoff is testability on a machine that is not the board: the caps are ordinary comptime values, so a host build can be compiled against the board's numbers and the parking, eviction and clamping paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART. ## A bare `zig build` `zig build` with no arguments now builds the tty and GUI binaries and installs them into `~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and "use it" two different commands for no reason.
* Compare a whole row with one memcmp before looking at cellsGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | `Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes. Byte equality implies visual equality, which is what makes the shortcut sound: a row that compares equal cannot be hiding a changed cell, and a row that differs only in padding falls through to the per-cell path, which is correct and merely slower. Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte, this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there is little left in it. It is also why board compute scaled 2.6x rather than 4x when the core clock went up 4x. Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and incremental path, same 18-step workload, same clock - identical characters and identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty and p4 both build.
* Make a keystroke 2.6x cheaper by not asking Unicode about ASCIIGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf -Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it was Unicode machinery answering questions about the letter `y`. Four changes, each a fast path guarded so that non-ASCII text takes exactly the road it took before. `modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates graphemes FROM THE START of the text with the full UAX #29 break state machine until it passes the offset, and the render path calls it once per visible row with a column offset - so the cost followed the cursor's distance along its line. That is the shape measured on the die, where inserting at column 320 of a fixed 320-character line cost 7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator - is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII, and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate. `Surface.print` then became the largest at 26.2%: per character it took a UTF-8 length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII takes none of that now. Same guard, same reason. `file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth` about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one printable cell and `gwidth` stays the authority on them. `modal.lineSlice` searched for "\n" with the generic substring search where a memchr does; it is called once per visible row per frame. Measured at the P4's geometry and configuration, on the host: render 55 -> 12 us, key-down 483 -> 24 us, key-right 327 -> 13 us, edit-char 205 -> 46 us. On the die, the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a keystroke at a 160-character line from 25.56 ms to 15.36 ms. ## The shadow grid, and why it is static `src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured 6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything changed: a second render with nothing new cost the same as the first. vaxis diffs its own grid, but only after being told every cell, and being told is the expensive part. So `present` now keeps the previous Surface and tells vaxis only what moved. `Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -> 1.45 ms. The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is not a micro-optimisation. The first version allocated it from the editor's heap; on a board whose 384 KiB is nearly spoken for, that is exactly the kind of change that works and then breaks something else three steps away. `shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present` behaves as it did before - clear and write every cell - which is the reference any measurement should be compared against, and the way to tell a rendering bug from a rendering difference. It earned its keep immediately: the two paths were run against the same 19-step workload on the die - inserts, deletes, motions that move the modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the reconstructed screens are byte-identical. ## Verification `snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0 mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 / gui all build. The rendering changes are exactly the sort that pass a latency benchmark while corrupting a screen, so the snapshot parity suite is the one that matters here and it is unchanged. `test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the things that table exists to hold constant, and 40x12 is not a scaled guess at the board - it is the board. `--only` exists because under `perf record` one 63 ms cell on the largest fixture swamps every sample from the case being asked about. ## Found, not fixed `vx.resize` fails on this board: a runtime geometry change hits its allocation failure path, restores the previous size and returns, so 80 bytes go out where 1,392 should. Verified independent of everything above - it reproduces with `shadow_grid` false. The board therefore has one geometry for the life of a session, which is why the staleness test above compares two firmwares rather than resizing one.
* big slow change: prebuilt shaders (SPIR-V/Metal), core gui reflow, docs, web ↵Gabriel Schneider2026-08-18
| | | | + snapshot refresh
* terminal: memoize the motion surface instead of dumping the scrollback per keyGabriel Schneider2026-08-12
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Every keystroke in a shell pane rebuilt the motion surface from scratch: shellRows dumped ghostty's WHOLE history+active grid, split it, blanked the prompt rows and handed back slices into the scratch arena, which the next update threw away. A pane sitting on a multi-megabyte agent transcript paid an O(scrollback) dump per press of `j`, and paid it once or twice per key, since flatSurface then rebuilt the same rows joined by '\n' beside it. The dump is now memoized against the pane it was built for (term_pane.RowsCache on Pardes.shell_rows), gpa-owned rather than scratch-arena because the whole point is to outlive the update that built it. One entry, not a table: the surface is built for the pane the cursor is in, and a second pane asking would only double a multi-megabyte buffer for a slot it is about to lose again. A pane that is not the live one is answered from the arena as before. The lifetime rule is the part that would have rotted silently, so it is one rule and it is written down: `rows` is handed out to callers, so everything that notices the entry has gone bad — output arrived, the grid reflowed, the pane died, another pane wants the slot — only marks it `stale`, and the buffers are freed in exactly two places, `sweep` at the TOP of an update before any handler can be holding them, and `reset` when the editor goes away. Nothing frees mid-update. dropPane clears the pointer immediately though: a freed pane's address comes back from the allocator as a different pane, and an entry still naming it would answer for the wrong grid. Two things fall out of having the join already: - flatSurface returns the memo's `text` verbatim when the lines it was handed are the cached rows untouched, instead of rebuilding the join. - paneCursorLines returns `rows` directly when there is no edit buffer, where it used to copy the array one slice at a time to produce exactly what it was given. One bug on the way past, in the same function: an EMPTY edit buffer writes one line but modal.lineCount("") is 0, so `ls` was sized one short of what the loop writes — the same floor the paste site needs. Killing a whole line (`A<C-u>`, `d%`) on a buffer covering the last row made that a length of zero. And test/perf.zig grows the axis that would have caught this: a terminal scoreboard beside the file one, three scrollback fixtures (64 KiB, 1 MiB, 8 MiB — half the ceiling) against render / output / resize-rows / resize-cols / key-down / edit-char, sharing the existing text and JSON reports and the --base comparison. resize-cols and resize-rows are both there because a COLUMN change reflows every page in the list and a row change does not. Measured on that table: key-down is 142 / 630 / 636 us across the three fixtures — flat from 1 MiB to 8 MiB, which is the dump being gone, and render flat at ~110 us throughout. What remains of key-down's step at 1 MiB is the linear scan indexOf refuses to index for a terminal; that is now a ponytail waiver naming its own price (615 us against 140 us) and the threading through paneOff/panePos/paneLineStart it would cost, to be done the day 0.6 ms shows up next to something anybody can feel.
* replace ArrayLists with bounded storageGabriel Schneider2026-08-10
|
* motion, scrolling and redraw are flat in file size nowGabriel Schneider2026-08-01
zig build perf drives the core directly — event, effects, one frame, no pty — over four generated fixtures: 1k lines, 50k, 300k, and 400 lines of 8000 columns, because a file that is long and a file that is wide fail differently. Every sample seeks somewhere else in the file first, since measuring at line 3 of a 300k-line file hides exactly the bug. perf record said half the run was scanning for newlines from byte 0. So File carries a line index, built on demand and invalidated in exactly ONE place — setContent, the funnel every content swap already goes through. That killed the scrollbar's per-frame line count (12.6% of the whole run by itself), scrollBy, ensureCursorVisible, lastNavRow, the syntax window bounds and two O(scroll) walks. normalKey computed max_line as a const at the top: two full passes over the buffer on every keystroke of every kind, for three g/G branches. It is lazy now. The modal primitives each walked the text twice for the same line. And the visible window was re-parsed on every scrolled row — a third of a megabyte per keypress on the wide fixture. The highlighted range is remembered, a scroll inside it is free, and only a re-parse that FOLLOWS a scroll takes slack: doing it unconditionally made typing 2.1x slower, since every character paid for a band it could never amortise. One j on a 19 MB file: 37.8ms -> 266us. Render: 4.1ms -> 77us. Open costs 1.25x more for the one extra pass, which buys 54x on every frame after, and 8 bytes per line of memory. Left standing, measured and named: edit-char is 14ms on 19MB because content is immutable and every keystroke copies the buffer. A third of that is the index rebuild, which could be a shift if setContent knew the edit offset; the rest wants a rope. bodyText's double copy and Surface.print's per-cell decode never rose above 2% of the profile afterwards, so they were left alone. No golden moved.