| Commit message (Collapse) | Author | Age |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The measurement said a keystroke costs 54 us per character already in the line. The
source said why: `insertAt` called `lineCount` - `std.mem.count` over every byte -
TWICE merely to clamp a row, and then `spliceAlloc` allocated and copied the whole
document. Three whole-document passes before one character can be inserted, which
fits a linear slope exactly.
So it was fixed. `modal.lineSpan` finds a row's byte span in ONE scan that stops at
that row, and `insertAt` uses it; a full count is paid only on the rare clamping
path where the cursor is past the end. Two tests pin the equivalence, including the
edges that make line counting awkward - an empty document, a trailing newline (its
own empty last line), and a row past the end. The first version of `lineSpan`
disagreed with `lineCount` about an empty document and the test caught it.
On the host harness this is a real win, reproduced over three independent runs at
matched sample counts:
edit-char 1k lines 50k lines 300k lines
before 730 us 2996 us 15030 us
after 721 us 2414 us 10955 us
ratio 0.98x 0.79-0.82x 0.78-0.83x
Every operation I did not touch stayed at 1.00x, which is better evidence than any
single cell.
On the board it changed NOTHING. The slope was 54.3 us/char before and 54.0 after,
a ratio of 1.00 over 5 conditions x 7 trials. Not a contradiction - the same fact
seen twice. The removed passes are O(document), and this board's document is a few
hundred BYTES, so two scans of it cost nothing worth measuring.
## Where the time actually goes
`-Dprof` times the two phases on the die with the cycle counter around
`pardes_p4_input` and `pardes_p4_render`:
chars in line input (parse+edit) render
1 220 us 14804 us
80 212 us 17114 us
240 250 us 24615 us
Input is FLAT at ~220 us - 1.5% of a keystroke - and does not grow with the document
at all. The ~15 ms floor and every microsecond of the slope are inside `render`. The
edit path could be made free and nobody would notice.
`soc.flushFlashCache`'s home in soc.zig is what let the profiling build exist at all
alongside the responder; `-Dprof` defaults off because it puts a line on the wire per
frame, which is the resource being measured.
## The report
`experiments/report.typ` gains Experiment 3 and, more importantly, a correction:
Experiment 2's mechanism claim was wrong and now says so, with the disproof next to
it. The ranked recommendations are reordered - the renderer is now #1 and the change
this commit makes is listed unranked, because on this target it buys nothing, which
is exactly why it is worth recording.
The position table earns its place there too: in a fixed 320-character line, an
insert at column 320 costs 33.9 ms and emits 28 bytes, while one at column 0 costs
26.0 ms and emits 81. Output size and latency are not merely uncorrelated on this
board, they are inverted - which is the signature of a walk from the start of a
line, and the next thing to go looking for.
The lesson is the one the instrument exists to enforce. A plausible mechanism, read
off the source and consistent with the shape of the data, was wrong about where the
time went, and only a measurement inside the firmware could say so.
|
| |
|
|
| |
additions, unit tests
|
| |
|
|
| |
+ snapshot refresh
|
| | |
|
| | |
|
| | |
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
zig build perf drives the core directly — event, effects, one frame, no pty —
over four generated fixtures: 1k lines, 50k, 300k, and 400 lines of 8000
columns, because a file that is long and a file that is wide fail differently.
Every sample seeks somewhere else in the file first, since measuring at line 3
of a 300k-line file hides exactly the bug.
perf record said half the run was scanning for newlines from byte 0. So File
carries a line index, built on demand and invalidated in exactly ONE place —
setContent, the funnel every content swap already goes through. That killed the
scrollbar's per-frame line count (12.6% of the whole run by itself), scrollBy,
ensureCursorVisible, lastNavRow, the syntax window bounds and two O(scroll)
walks. normalKey computed max_line as a const at the top: two full passes over
the buffer on every keystroke of every kind, for three g/G branches. It is lazy
now. The modal primitives each walked the text twice for the same line.
And the visible window was re-parsed on every scrolled row — a third of a
megabyte per keypress on the wide fixture. The highlighted range is remembered,
a scroll inside it is free, and only a re-parse that FOLLOWS a scroll takes
slack: doing it unconditionally made typing 2.1x slower, since every character
paid for a band it could never amortise.
One j on a 19 MB file: 37.8ms -> 266us. Render: 4.1ms -> 77us. Open costs 1.25x
more for the one extra pass, which buys 54x on every frame after, and 8 bytes
per line of memory.
Left standing, measured and named: edit-char is 14ms on 19MB because content is
immutable and every keystroke copies the buffer. A third of that is the index
rebuild, which could be a shift if setContent knew the edit offset; the rest
wants a rope. bodyText's double copy and Surface.print's per-cell decode never
rose above 2% of the profile afterwards, so they were left alone.
No golden moved.
|
| | |
|
|
|
test/ (snapshot parity harness + 18 frozen goldens). One sans-IO core, vaxis tty + SDL3 GPU native + wasm web shells, 18/18 parity with the purged prototype, 7.6k lines vs 12.1k. Fix: gui shell pre-sized the core at init so the greet-releasing resize never fired (blank panes until first interaction); live sessions now init at defaults and get the real grid as a resize event (the shell contract, documented on Options).
|