diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 18:49:57 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 19:16:25 -0300 |
| commit | 1935944a0352e9d176f0718063327e686ece1948 (patch) | |
| tree | 1169c5ac00fdbfa2cc8268620ec03348ea0d4164 /src/pardes/app.zig | |
| parent | 1cef9c2e4bd873ebe13f5df635899231bcc467d2 (diff) | |
| download | esp32p4-1935944a0352e9d176f0718063327e686ece1948.tar.gz esp32p4-1935944a0352e9d176f0718063327e686ece1948.zip | |
Bound the edit path's document scans, and find out they were never the problem
The board half: the -Dprof attribution that overturned the conclusion, its data, and
the report correction.
`-Dprof` adds two cycle-counter reads around `pardes_p4_input` and
`pardes_p4_render` and prints both. Off by default: it puts a line on the wire per
frame, which is the very resource being measured, so it answers "where did the 15 ms
go" and not "how fast is it".
It answered. Input is flat at ~220 us regardless of document size - 1.5% of a
keystroke - and the entire ~15 ms floor plus every microsecond of the per-character
slope live in `render`. The edit-path fix that the source reading implied (committed
next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5
conditions x 7 trials, exactly 0% on this board.
`experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's
mechanism claim was wrong, says so, and carries the disproof beside it. The ranked
recommendations are reordered with the renderer at #1.
Also here: `--sweep position` in p4-bench, which holds the document fixed at one
320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28
bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on
this board - the signature of a walk from the start of a line.
Diffstat (limited to 'src/pardes/app.zig')
| -rw-r--r-- | src/pardes/app.zig | 26 |
1 files changed, 25 insertions, 1 deletions
diff --git a/src/pardes/app.zig b/src/pardes/app.zig index a83485a..23da203 100644 --- a/src/pardes/app.zig +++ b/src/pardes/app.zig @@ -37,6 +37,10 @@ const std = @import("std"); const soc = @import("soc"); const config = @import("config"); + +/// `-Dprof`: time the two phases of a keystroke on the board and print the cycle counts. A +/// diagnostic, not a feature - see the loop. +const prof = config.prof; const hal = @import("hal"); const heapmod = @import("heap"); const uart = @import("uart.zig"); @@ -203,16 +207,36 @@ export fn zig_main() noreturn { var in: [256]u8 = undefined; while (!pardes_p4_quit()) { + // ATTRIBUTION. The host can time a keystroke's round trip but cannot see what the firmware + // spent it on, and the two candidates - parsing and editing, versus rendering - want + // opposite fixes. `soc.cycles()` is the unprivileged cycle counter, so this costs two CSR + // reads per phase and quantises at one cycle, which is four orders of magnitude below the + // milliseconds being attributed. Gated on `prof` so the shipping build carries none of it. const n = uart.read(&in); - if (n > 0) pardes_p4_input(&in, n); + var input_cy: u64 = 0; + if (n > 0) { + const t0 = if (prof) soc.cycles() else 0; + pardes_p4_input(&in, n); + if (prof) input_cy = soc.cycles() - t0; + } pardes_p4_tick(nowMs()); // Only when there is something to show. On a link this slow an unconditional repaint per // iteration would saturate the wire and starve input. if (pardes_p4_wants_frame()) { + const t0 = if (prof) soc.cycles() else 0; const err = pardes_p4_render(); if (err != 0) soc.rom.print("MARK PARDES_RENDER_FAIL rc=%u\r\n", .{err}); + if (prof) { + const render_cy = soc.cycles() - t0; + // Reported in cycles, not microseconds: the divisor is the CPU clock, which this + // firmware does not set and has only ever measured, so converting here would bake a + // guess into the data. `experiments/` divides by the clock it measured. + soc.rom.print("PROF in=%u render=%u\r\n", .{ + @as(u32, @intCast(input_cy)), @as(u32, @intCast(render_cy)), + }); + } } } |
