diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
| commit | 1cef9c2e4bd873ebe13f5df635899231bcc467d2 (patch) | |
| tree | ce0a9495bc801666ff14da47f7a75728335cc442 | |
| parent | ef6f3e460ddbf53a752807bcf10a0fec0a72da7f (diff) | |
| download | esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.tar.gz esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.zip | |
A measuring instrument, and what it says about where the latency goes
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.
## The instrument
`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.
The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.
`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.
`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.
## What it measured
The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.
Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.
And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.
One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.
## Three bugs found while building it
The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.
Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.
`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.
Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.
## The report
`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
| -rw-r--r-- | .gitignore | 6 | ||||
| -rw-r--r-- | README.md | 46 | ||||
| -rw-r--r-- | build.zig | 48 | ||||
| -rw-r--r-- | examples/uartperf.zig | 229 | ||||
| -rw-r--r-- | experiments/length-ReleaseFast.csv | 36 | ||||
| -rw-r--r-- | experiments/length-ReleaseSmall.csv | 36 | ||||
| -rw-r--r-- | experiments/ops-ReleaseFast.csv | 50 | ||||
| -rw-r--r-- | experiments/ops-ReleaseSmall.csv | 48 | ||||
| -rw-r--r-- | experiments/report.typ | 435 | ||||
| -rw-r--r-- | src/pardes/app.zig | 46 | ||||
| -rw-r--r-- | src/soc.zig | 45 | ||||
| -rw-r--r-- | tools/bench_main.zig | 695 | ||||
| -rw-r--r-- | tools/perfproto.zig | 201 | ||||
| -rw-r--r-- | tools/rtt.zig | 170 | ||||
| -rw-r--r-- | tools/serial.zig | 28 |
15 files changed, 2073 insertions, 46 deletions
@@ -28,3 +28,9 @@ # Espressif publishes in no other form, so there is no PDF to re-fetch instead. /cpu-docs/**/*.pdf /cpu-docs/**/*.zip + +# Generated from experiments/report.typ, which computes its own figures from the CSVs beside it. +# The measurements are the artefact and stay tracked; a rebuilt PDF is not a measurement, and the +# page PNGs are a debugging convenience for checking a figure rendered. +/experiments/report.pdf +/experiments/fig-*.png @@ -9,6 +9,7 @@ zig build console # attach a terminal to whatever is already on the board (Ct zig build interact # flash, then attach that terminal (ordered, like `run`) zig build reset # just pulse the reset line zig build size # where every byte of the image went +zig build bench # measure the serial link and the editor, verified with a checksum zig build test # host tests: image builder, and the register layer's field arithmetic zig build diff # the hardware oracle: this HAL vs ESP-IDF's, on the die (needs -Doracle) zig build elf # stop at the ELF, for disassembly @@ -121,6 +122,51 @@ Measured on ESP32-P4 rev v1.3 silicon: two `Peek`s of the RNG register at `0x501 `Peek 0x50110001` answered `peek: MisalignedAddress` on the message row rather than taking the session down with an unhandled trap, which is the one fault that file exists to prevent. +### Measuring it + +`p4-bench` exists so that optimising this port is not a matter of opinion. It has two halves, +because there are two different questions. + +``` +zig build flash -Dapp=examples/uartperf.zig # the ceiling: link + driver, nothing else +zig-out/bin/p4-bench --link + +zig build flash -Dpardes # how much of that ceiling the editor uses +zig-out/bin/p4-bench --editor +zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall +``` + +Every `--link` number is checksummed. `tools/perfproto.zig` is a framed protocol - `"P4"`, op, +length, CRC-32, payload - shared *verbatim* by the host tool and `examples/uartperf.zig`, so a frame +one writes and the other parses cannot drift. That matters because RX overrun on this UART is +undetected in hardware and uncounted in the driver: a byte that never arrived is indistinguishable +from a late one, and an unchecksummed throughput figure is a guess about how fast data was corrupted. + +`tools/rtt.zig` holds the two timing functions everything is built on. Round trip is measured to the +**first** response byte, not the last: a renderer that starts drawing in 8 ms and finishes in 130 ms +feels immediate, one that thinks for 130 ms then draws in 8 ms feels broken, and waiting for the wire +to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. +Microseconds throughout, because at 115200 one byte is 87 us and a millisecond clock would quantise +the answer into buckets eleven bytes wide. + +`experiments/` holds the raw per-trial CSVs and `report.typ`, which reads them and computes its own +figures - so a re-run changes the document rather than contradicting it. What it establishes on this +die: + +| | | +|---|---| +| link, both directions | 99.8% of the 11,520 B/s wire, CRC verified over 32,768 B | +| typing, 6 to 100 keys/s | nothing lost, wire never above 9% | +| one keystroke | 81 B, round trip 17.0 ms | +| a motion | 40 B, round trip 16.7 ms | +| cost per character already in the line | **54.3 us, per keystroke** | + +The first two rows say the wire is not the problem. The last three say why: a 40-byte operation and a +206-byte one cost the same round trip, so latency is compute per event and not transmission - and it +grows with the document, because the edit path copies the whole buffer every keystroke. Building the +editor object `ReleaseFast` instead of `ReleaseSmall` cuts the fixed cost 13% and the per-character +cost 36% for 35% more flash, which is the best ratio measured here. + ## Layout ``` @@ -180,6 +180,14 @@ pub fn build(b: *std.Build) void { .{ .name = "hal", .module = hal_mod }, .{ .name = "mmio", .module = mmio_mod }, .{ .name = "regs", .module = regs_mod }, + // The wire protocol `examples/uartperf.zig` answers, imported rather than copied so + // the firmware and the host tool cannot disagree about a frame. It is deliberately + // free of any OS dependency for exactly this reason: one file, two targets. + .{ .name = "perfproto", .module = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = target, + .optimize = optimize, + }) }, }, }), }); @@ -397,6 +405,32 @@ pub fn build(b: *std.Build) void { run_con.step.dependOn(&flash.step); b.step("interact", "flash the image, then attach a terminal").dependOn(&run_con.step); + // The measuring instrument. Its own binary for the same reason the console is: it drives the + // port for tens of seconds and must not have the build runner repainting a progress tree into + // the middle of a timed transfer. It shares `tools/perfproto.zig` with the firmware responder, + // so a frame the host writes and a frame the board parses cannot drift apart. + const bench_proto = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = b.graph.host, + .optimize = .ReleaseSafe, + }); + const bench_exe = b.addExecutable(.{ + .name = "p4-bench", + .root_module = b.createModule(.{ + .root_source_file = b.path("tools/bench_main.zig"), + .target = b.graph.host, + .optimize = .ReleaseSafe, + .imports = &.{.{ .name = "perfproto", .module = bench_proto }}, + }), + }); + const bench_install = b.addInstallArtifact(bench_exe, .{}); + const bench = b.addRunArtifact(bench_exe); + bench.addArgs(&.{ "--port", port_path }); + bench.stdio = .inherit; + bench.step.dependOn(&bench_install.step); + b.step("bench", "measure the serial link: verified throughput each way, and latency") + .dependOn(&bench.step); + const reset = ResetStep.create(b, port_path); b.step("reset", "reset the board and let the flashed application run").dependOn(&reset.step); @@ -456,6 +490,20 @@ pub fn build(b: *std.Build) void { }); test_step.dependOn(&b.addRunArtifact(console_tests).step); + // The measurement protocol. These are the tests that keep a throughput number honest: that a + // frame round-trips, that a short read is "incomplete" rather than "invalid", that a lost byte + // mid-stream changes the CRC, and that the pattern generator does not repeat on a 256-byte + // boundary - a plain counter would hash identically after losing exactly 256 bytes and the + // instrument would report a clean run over corrupt data. + const proto_tests = b.addTest(.{ + .root_module = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = b.graph.host, + .optimize = .Debug, + }), + }); + test_step.dependOn(&b.addRunArtifact(proto_tests).step); + // The radio path's host-testable parts. A wrong checksum, a wrong snprintf, or a scheduler that // loses a task is far cheaper to find here than on a board whose only output is a serial line. // diff --git a/examples/uartperf.zig b/examples/uartperf.zig new file mode 100644 index 0000000..339f642 --- /dev/null +++ b/examples/uartperf.zig @@ -0,0 +1,229 @@ +//! The board half of the link measurement: answer `tools/perfproto.zig` frames over UART0. +//! +//! This is the CEILING the editor is measured against. `p4-bench` against this firmware says what +//! the wire and the UART driver can do with nothing else running; `p4-bench` against the editor says +//! how much of that the editor manages to use. Optimising the editor without the first number is +//! guessing, because at 115200 baud a good deal of what feels slow is simply the wire, and no amount +//! of firmware work moves it. +//! +//! Three things it deliberately does NOT do, each of which would corrupt the number: +//! +//! * **No `soc.rom.print`.** The mask ROM's `ets_printf` formats and then pushes one byte at a +//! time, spinning on the FIFO for each - the exact cost this is trying to measure around. Every +//! byte here goes through the same batched FIFO path `src/pardes/uart.zig` uses. +//! * **No UART reconfiguration.** Not the divider, not the format, not `reset()`. The +//! second-stage bootloader configured this block; `hal/uart.zig:195-211` records that resetting +//! it returns UART_CLKDIV to its power-on value and takes the session with it. +//! * **No allocation.** One parse buffer and one send buffer, both static, both sized by the +//! protocol's own `max_payload`. A measurement that shared a heap with anything would measure +//! the heap. +//! +//! The verification is the point. `sink` accumulates a CRC across every payload byte received and +//! `report` hands it back, so the host can prove that what arrived is what it sent - at this baud a +//! silent RX overrun is the failure mode that matters, and a byte count alone cannot see it. + +const std = @import("std"); +const hal = @import("hal"); +const proto = @import("perfproto"); + +const uart0 = hal.uart.Uart.init(0); + +/// Room for one whole frame. The protocol caps a payload at 1024 precisely so this can be static. +var rx: [proto.header_len + proto.max_payload]u8 = undefined; +var rx_len: usize = 0; + +var tx: [proto.header_len + proto.max_payload]u8 = undefined; + +/// The running `sink` accumulators, reported and reset by `report`. +var sunk_bytes: u32 = 0; +var sunk_crc: std.hash.Crc32 = undefined; +var bad_frames: u32 = 0; +var tx_dropped: u32 = 0; + +/// Push bytes through the TX FIFO, reading the status once per burst rather than once per byte. +/// +/// The spin is bounded because an unbounded one is indistinguishable from a hang on a board with no +/// debugger, and because this program's whole purpose is to report numbers: a wedged transmitter +/// that increments a counter can still be diagnosed, while one that spins forever cannot. +fn write(bytes: []const u8) void { + var rest = bytes; + while (rest.len > 0) { + var room = uart0.txFree(); + var spins: u32 = 0; + while (room == 0) { + spins += 1; + if (spins > 1_000_000) { + tx_dropped +%= @intCast(rest.len); + return; + } + room = uart0.txFree(); + } + const n = @min(room, rest.len); + for (rest[0..n]) |b| uart0.pushByte(b); + rest = rest[n..]; + } +} + +fn send(op: proto.Op, payload: []const u8) void { + write(proto.encode(&tx, op, payload)); +} + +fn sendStat() void { + var buf: [proto.Stat.encoded_len]u8 = undefined; + const s: proto.Stat = .{ + .bytes = sunk_bytes, + .crc = sunk_crc.final(), + .bad_frames = bad_frames, + .tx_dropped = tx_dropped, + }; + s.encode(&buf); + send(.stat, &buf); + sunk_bytes = 0; + sunk_crc = .init(); + bad_frames = 0; +} + +/// Stream `n` pattern bytes back as `data` frames, then a `stat` whose CRC covers all of them. +/// +/// Filled a frame at a time from the shared generator rather than from a table: the host computes +/// the same sequence from the same function, so a disagreement is a real transport fault and not two +/// copies of a constant drifting apart. +fn source(n: u32) void { + var chunk: [proto.max_payload]u8 = undefined; + var sent: u32 = 0; + var hash: std.hash.Crc32 = .init(); + while (sent < n) { + const take: u32 = @min(@as(u32, proto.max_payload), n - sent); + proto.fillPattern(chunk[0..take], sent); + hash.update(chunk[0..take]); + send(.data, chunk[0..take]); + sent += take; + } + var buf: [proto.Stat.encoded_len]u8 = undefined; + const s: proto.Stat = .{ + .bytes = sent, + .crc = hash.final(), + .bad_frames = bad_frames, + .tx_dropped = tx_dropped, + }; + s.encode(&buf); + send(.stat, &buf); +} + +/// Consume one complete frame from the head of `rx`. Returns the bytes consumed, or 0 when the +/// frame is not all here yet. +fn step() usize { + const header = proto.parseHeader(rx[0..rx_len]) catch { + // Lost sync. Drop ONE byte and let the next call try again from there: the magic is two + // bytes, so resynchronising by scanning is the only correct recovery, and dropping the whole + // buffer would discard a good frame that happened to follow a corrupt one. + return 1; + } orelse return 0; + + const total = proto.header_len + @as(usize, header.len); + if (rx_len < total) return 0; + const payload = rx[proto.header_len..total]; + + if (proto.crc(payload) != header.crc) { + // Corruption, not loss: the length was plausible and the bytes were not. Counted and + // discarded, because acting on it would put the wrong answer in the host's hands. + bad_frames +%= 1; + return total; + } + + switch (header.op) { + .ping => send(.pong, payload), + .sink => { + sunk_bytes +%= header.len; + sunk_crc.update(payload); + }, + .report => sendStat(), + .source => { + const n = if (header.len >= 4) std.mem.readInt(u32, payload[0..4], .little) else 0; + source(n); + }, + // Replies are ours to send, never to receive. A reply arriving here means the host is + // confused or the wire is looping back; count it rather than answering it. + .pong, .stat, .data => bad_frames +%= 1, + } + return total; +} + +export fn zig_main() noreturn { + // FIRST, before any `.rodata` is touched - and the marker below IS `.rodata`. Without this the + // bootloader's stale cache lines make that string read as machine code, the board emits noise, + // and it looks exactly like a firmware that never started. Measured here before the call was + // added: `\xefc\xff\xff\xd5\xb7...` instead of the marker. + @import("soc").flushFlashCache(); + + // The bootloader arms the RTC watchdog and expects the application to take it over. Nothing in + // this repo ever did, so every image here was being reset on a ten-second cycle - invisible to + // a program that prints once and spins, and fatal to one that must answer for a minute. It is + // why this responder booted, printed its marker, and then went silent: `rst:0x10 + // (CHIP_LP_WDT_RESET)` in the next boot log, measured. + _ = hal.rwdt.disable(); + sunk_crc = .init(); + + // Announce readiness in plain text rather than as a frame: the host watches for this during + // reset, while the bootloader's own chatter is still arriving and no frame parser is in sync. + write("\r\nMARK UARTPERF_READY\r\n"); + + while (true) { + // Fill from the FIFO first and always, so the RX FIFO is never left to overflow while this + // loop is busy elsewhere. 128 bytes at 115200 is 107 ms of slack and a `source` burst can + // hold the transmitter far longer than that, which is exactly the hazard being measured on + // the editor - so the measuring instrument must not have it. + if (rx_len < rx.len) { + const room = rx.len - rx_len; + var got: usize = 0; + while (got < room and uart0.rxCount() > 0) { + rx[rx_len + got] = uart0.popByte(); + got += 1; + } + rx_len += got; + } + + var off: usize = 0; + while (off < rx_len) { + const n = step(); + if (n == 0) break; + off += n; + } + if (off > 0) { + std.mem.copyForwards(u8, rx[0 .. rx_len - off], rx[off..rx_len]); + rx_len -= off; + } + } +} + +/// Reset entry, the same shape as every other application here: the bootloader hands over with an +/// unspecified stack pointer and the FPU off, and `.bss` is not cleared for us. The FPU bit matters +/// even in a program with no floats, because a generic `bufPrint` instantiation can reach std's +/// float formatting path - and `std.hash.Crc32`'s table generation is comptime, so nothing here +/// needs the FPU at run time, but nothing here is worth a trap either. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} + +/// A panic here would be a measurement that silently stopped, so it says so on the wire it was +/// measuring - through the ROM's printf, because a panic may well be the UART path itself failing. +pub const panic = std.debug.FullPanic(struct { + fn call(msg: []const u8, _: ?usize) noreturn { + @import("soc").rom.print("MARK UARTPERF_PANIC %s\r\n", .{msg.ptr}); + while (true) {} + } +}.call); diff --git a/experiments/length-ReleaseFast.csv b/experiments/length-ReleaseFast.csv new file mode 100644 index 0000000..043fb83 --- /dev/null +++ b/experiments/length-ReleaseFast.csv @@ -0,0 +1,36 @@ +label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes +ReleaseFast,length,0,0,0,insert,0,15072,26542,141 +ReleaseFast,length,0,0,0,insert,1,14793,20959,80 +ReleaseFast,length,0,0,0,insert,2,14767,20944,81 +ReleaseFast,length,0,0,0,insert,3,14794,20970,81 +ReleaseFast,length,0,0,0,insert,4,14808,20937,81 +ReleaseFast,length,0,0,0,insert,5,14875,20991,81 +ReleaseFast,length,0,0,0,insert,6,14870,20971,81 +ReleaseFast,length,0,0,20,insert,0,15305,21531,81 +ReleaseFast,length,0,0,20,insert,1,15317,21534,81 +ReleaseFast,length,0,0,20,insert,2,15425,21719,81 +ReleaseFast,length,0,0,20,insert,3,15408,21759,81 +ReleaseFast,length,0,0,20,insert,4,15513,21702,81 +ReleaseFast,length,0,0,20,insert,5,15821,28801,158 +ReleaseFast,length,0,0,20,insert,6,15740,21843,80 +ReleaseFast,length,0,0,40,insert,0,16198,22320,81 +ReleaseFast,length,0,0,40,insert,1,16261,22352,81 +ReleaseFast,length,0,0,40,insert,2,16229,22505,81 +ReleaseFast,length,0,0,40,insert,3,16309,22374,81 +ReleaseFast,length,0,0,40,insert,4,16312,22466,81 +ReleaseFast,length,0,0,40,insert,5,16289,22563,81 +ReleaseFast,length,0,0,40,insert,6,16341,22625,81 +ReleaseFast,length,0,0,80,insert,0,17772,23963,81 +ReleaseFast,length,0,0,80,insert,1,17803,23933,81 +ReleaseFast,length,0,0,80,insert,2,17817,24108,81 +ReleaseFast,length,0,0,80,insert,3,17818,23893,81 +ReleaseFast,length,0,0,80,insert,4,17917,24037,81 +ReleaseFast,length,0,0,80,insert,5,17919,24158,81 +ReleaseFast,length,0,0,80,insert,6,17920,24078,81 +ReleaseFast,length,0,0,160,insert,0,20271,26409,81 +ReleaseFast,length,0,0,160,insert,1,20230,26412,81 +ReleaseFast,length,0,0,160,insert,2,20249,26522,81 +ReleaseFast,length,0,0,160,insert,3,20302,26487,81 +ReleaseFast,length,0,0,160,insert,4,20648,33557,158 +ReleaseFast,length,0,0,160,insert,5,20563,26664,80 +ReleaseFast,length,0,0,160,insert,6,20699,26920,81 diff --git a/experiments/length-ReleaseSmall.csv b/experiments/length-ReleaseSmall.csv new file mode 100644 index 0000000..eb6fa1a --- /dev/null +++ b/experiments/length-ReleaseSmall.csv @@ -0,0 +1,36 @@ +label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes +ReleaseSmall,length,0,0,0,insert,0,17211,28542,141 +ReleaseSmall,length,0,0,0,insert,1,16894,23031,80 +ReleaseSmall,length,0,0,0,insert,2,16865,23038,81 +ReleaseSmall,length,0,0,0,insert,3,17009,23097,81 +ReleaseSmall,length,0,0,0,insert,4,16986,23146,81 +ReleaseSmall,length,0,0,0,insert,5,16957,23253,81 +ReleaseSmall,length,0,0,0,insert,6,17090,23114,81 +ReleaseSmall,length,0,0,20,insert,0,17678,23892,81 +ReleaseSmall,length,0,0,20,insert,1,17736,24079,81 +ReleaseSmall,length,0,0,20,insert,2,17871,24024,81 +ReleaseSmall,length,0,0,20,insert,3,17942,24029,81 +ReleaseSmall,length,0,0,20,insert,4,17975,24117,81 +ReleaseSmall,length,0,0,20,insert,5,18468,31324,158 +ReleaseSmall,length,0,0,20,insert,6,18378,24470,80 +ReleaseSmall,length,0,0,40,insert,0,19108,25325,81 +ReleaseSmall,length,0,0,40,insert,1,19099,25452,81 +ReleaseSmall,length,0,0,40,insert,2,19152,25335,81 +ReleaseSmall,length,0,0,40,insert,3,19155,25357,81 +ReleaseSmall,length,0,0,40,insert,4,19247,25475,81 +ReleaseSmall,length,0,0,40,insert,5,19236,25463,81 +ReleaseSmall,length,0,0,40,insert,6,19343,25429,81 +ReleaseSmall,length,0,0,80,insert,0,21481,27587,81 +ReleaseSmall,length,0,0,80,insert,1,21516,27613,81 +ReleaseSmall,length,0,0,80,insert,2,21604,27728,81 +ReleaseSmall,length,0,0,80,insert,3,21654,27997,81 +ReleaseSmall,length,0,0,80,insert,4,21647,27764,81 +ReleaseSmall,length,0,0,80,insert,5,21588,27786,81 +ReleaseSmall,length,0,0,80,insert,6,21729,27953,81 +ReleaseSmall,length,0,0,160,insert,0,25373,31443,81 +ReleaseSmall,length,0,0,160,insert,1,25465,31612,81 +ReleaseSmall,length,0,0,160,insert,2,25469,31643,81 +ReleaseSmall,length,0,0,160,insert,3,25561,31595,81 +ReleaseSmall,length,0,0,160,insert,4,25994,39003,158 +ReleaseSmall,length,0,0,160,insert,5,25876,31853,80 +ReleaseSmall,length,0,0,160,insert,6,26039,32201,81 diff --git a/experiments/ops-ReleaseFast.csv b/experiments/ops-ReleaseFast.csv new file mode 100644 index 0000000..ea42b87 --- /dev/null +++ b/experiments/ops-ReleaseFast.csv @@ -0,0 +1,50 @@ +label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes +ReleaseFast,ops,0,0,0,motion_h,0,14839,17388,40 +ReleaseFast,ops,0,0,0,motion_h,1,14767,17434,40 +ReleaseFast,ops,0,0,0,motion_h,2,14721,17372,40 +ReleaseFast,ops,0,0,0,motion_h,3,14715,17395,40 +ReleaseFast,ops,0,0,0,motion_h,4,14695,17281,40 +ReleaseFast,ops,0,0,0,motion_h,5,14703,17284,40 +ReleaseFast,ops,0,0,0,motion_h,6,14717,17393,40 +ReleaseFast,ops,0,0,0,motion_l,0,14742,17302,40 +ReleaseFast,ops,0,0,0,motion_l,1,14746,17425,40 +ReleaseFast,ops,0,0,0,motion_l,2,14727,17342,40 +ReleaseFast,ops,0,0,0,motion_l,3,14723,17272,40 +ReleaseFast,ops,0,0,0,motion_l,4,14876,17380,40 +ReleaseFast,ops,0,0,0,motion_l,5,14704,17341,40 +ReleaseFast,ops,0,0,0,motion_l,6,14746,17328,40 +ReleaseFast,ops,0,0,0,line_start,0,14767,17478,40 +ReleaseFast,ops,0,0,0,line_start,1,14760,17396,40 +ReleaseFast,ops,0,0,0,line_start,2,14735,17355,40 +ReleaseFast,ops,0,0,0,line_start,3,14701,17364,40 +ReleaseFast,ops,0,0,0,line_start,4,14710,17286,40 +ReleaseFast,ops,0,0,0,line_start,5,14771,17453,40 +ReleaseFast,ops,0,0,0,line_start,6,14751,17226,40 +ReleaseFast,ops,0,0,0,line_end,0,14783,17424,40 +ReleaseFast,ops,0,0,0,line_end,1,14807,17443,40 +ReleaseFast,ops,0,0,0,line_end,2,14773,17284,40 +ReleaseFast,ops,0,0,0,line_end,3,14737,17408,40 +ReleaseFast,ops,0,0,0,line_end,4,14769,17508,40 +ReleaseFast,ops,0,0,0,line_end,5,14732,17412,40 +ReleaseFast,ops,0,0,0,line_end,6,14750,17268,40 +ReleaseFast,ops,0,0,0,insert_esc,0,14827,42161,265 +ReleaseFast,ops,0,0,0,insert_esc,1,14908,36678,204 +ReleaseFast,ops,0,0,0,insert_esc,2,14983,36759,206 +ReleaseFast,ops,0,0,0,insert_esc,3,14988,36851,206 +ReleaseFast,ops,0,0,0,insert_esc,4,14906,36822,206 +ReleaseFast,ops,0,0,0,insert_esc,5,15050,36865,206 +ReleaseFast,ops,0,0,0,insert_esc,6,15059,36952,206 +ReleaseFast,ops,0,0,0,repaint_39,0,14926,32405,81 +ReleaseFast,ops,0,0,0,repaint_39,1,14936,32122,80 +ReleaseFast,ops,0,0,0,repaint_39,2,14895,32294,80 +ReleaseFast,ops,0,0,0,repaint_39,3,14975,32342,80 +ReleaseFast,ops,0,0,0,repaint_39,4,15026,32363,80 +ReleaseFast,ops,0,0,0,repaint_39,5,14968,32224,80 +ReleaseFast,ops,0,0,0,repaint_39,6,14885,32342,80 +ReleaseFast,ops,0,0,0,repaint_40,0,14876,32419,80 +ReleaseFast,ops,0,0,0,repaint_40,1,14866,32211,80 +ReleaseFast,ops,0,0,0,repaint_40,2,14868,32173,80 +ReleaseFast,ops,0,0,0,repaint_40,3,14888,32329,80 +ReleaseFast,ops,0,0,0,repaint_40,4,14869,32208,80 +ReleaseFast,ops,0,0,0,repaint_40,5,14863,32277,80 +ReleaseFast,ops,0,0,0,repaint_40,6,14876,32143,80 diff --git a/experiments/ops-ReleaseSmall.csv b/experiments/ops-ReleaseSmall.csv new file mode 100644 index 0000000..daa51f4 --- /dev/null +++ b/experiments/ops-ReleaseSmall.csv @@ -0,0 +1,48 @@ +label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes +ReleaseSmall,ops,0,0,0,motion_h,0,16719,19345,40 +ReleaseSmall,ops,0,0,0,motion_h,1,16770,19401,40 +ReleaseSmall,ops,0,0,0,motion_h,2,16666,19240,40 +ReleaseSmall,ops,0,0,0,motion_h,3,16741,19292,40 +ReleaseSmall,ops,0,0,0,motion_h,4,16765,19437,40 +ReleaseSmall,ops,0,0,0,motion_h,5,16699,19235,40 +ReleaseSmall,ops,0,0,0,motion_h,6,16685,19257,40 +ReleaseSmall,ops,0,0,0,motion_l,0,16688,19418,40 +ReleaseSmall,ops,0,0,0,motion_l,1,16826,19445,40 +ReleaseSmall,ops,0,0,0,motion_l,2,16781,19317,40 +ReleaseSmall,ops,0,0,0,motion_l,3,16694,19326,40 +ReleaseSmall,ops,0,0,0,motion_l,4,16677,19256,40 +ReleaseSmall,ops,0,0,0,motion_l,5,16645,19276,40 +ReleaseSmall,ops,0,0,0,motion_l,6,16709,19276,40 +ReleaseSmall,ops,0,0,0,line_start,0,16795,19362,40 +ReleaseSmall,ops,0,0,0,line_start,1,16698,19319,40 +ReleaseSmall,ops,0,0,0,line_start,2,16820,19407,40 +ReleaseSmall,ops,0,0,0,line_start,3,16716,19287,40 +ReleaseSmall,ops,0,0,0,line_start,4,16760,19405,40 +ReleaseSmall,ops,0,0,0,line_start,5,16725,19450,40 +ReleaseSmall,ops,0,0,0,line_start,6,16769,19315,40 +ReleaseSmall,ops,0,0,0,line_end,0,16681,19296,40 +ReleaseSmall,ops,0,0,0,line_end,1,16628,19240,40 +ReleaseSmall,ops,0,0,0,line_end,2,16796,19343,40 +ReleaseSmall,ops,0,0,0,line_end,3,16678,19242,40 +ReleaseSmall,ops,0,0,0,line_end,4,16703,19396,40 +ReleaseSmall,ops,0,0,0,line_end,5,16693,19438,40 +ReleaseSmall,ops,0,0,0,line_end,6,16699,19341,40 +ReleaseSmall,ops,0,0,0,insert_esc,0,16784,46055,265 +ReleaseSmall,ops,0,0,0,insert_esc,1,16819,40714,204 +ReleaseSmall,ops,0,0,0,insert_esc,2,17244,37382,206 +ReleaseSmall,ops,0,0,0,insert_esc,3,16897,40737,206 +ReleaseSmall,ops,0,0,0,insert_esc,4,17236,37425,206 +ReleaseSmall,ops,0,0,0,insert_esc,5,17006,40907,206 +ReleaseSmall,ops,0,0,0,insert_esc,6,17159,41133,206 +ReleaseSmall,ops,0,0,0,repaint_39,0,17034,36306,81 +ReleaseSmall,ops,0,0,0,repaint_39,1,16983,36314,80 +ReleaseSmall,ops,0,0,0,repaint_39,2,17159,36171,80 +ReleaseSmall,ops,0,0,0,repaint_39,3,16913,36217,80 +ReleaseSmall,ops,0,0,0,repaint_39,4,10054,134042,1253 +ReleaseSmall,ops,0,0,0,repaint_39,5,16638,35615,80 +ReleaseSmall,ops,0,0,0,repaint_40,0,10018,135552,1265 +ReleaseSmall,ops,0,0,0,repaint_40,1,17099,36357,80 +ReleaseSmall,ops,0,0,0,repaint_40,2,16988,36299,80 +ReleaseSmall,ops,0,0,0,repaint_40,4,17031,36233,80 +ReleaseSmall,ops,0,0,0,repaint_40,5,17013,36330,80 +ReleaseSmall,ops,0,0,0,repaint_40,6,16957,36281,80 diff --git a/experiments/report.typ b/experiments/report.typ new file mode 100644 index 0000000..1d859b8 --- /dev/null +++ b/experiments/report.typ @@ -0,0 +1,435 @@ +#import "@preview/cetz:0.5.1" + +#set document(title: "Where the ESP32-P4 editor's latency goes", author: "measured on ESP32-P4 rev v1.3") +#set page(margin: 2cm, numbering: "1") +#set text(font: ("Libertinus Serif", "DejaVu Serif"), size: 10.5pt) +#set par(justify: true) +#show raw: set text(font: "DejaVu Sans Mono", size: 9pt) +#show heading: it => block(above: 1.4em, below: 0.7em, it) + +#align(center)[ + #text(17pt, weight: "bold")[Where the ESP32-P4 editor's latency goes] + + #v(0.2em) + #text(10pt)[A measured account, on silicon, of the `pardes` editor running as + ESP32-P4 firmware with one 115200-baud serial line as its only I/O] +] + +#v(0.5em) + +#let capacity = 11520.0 + +// ---------------------------------------------------------------- data plumbing +// The figures read the raw per-trial CSV the instrument emits. Nothing here is a +// transcribed number: if a run is repeated, the document changes with it. +#let rows(file) = { + let out = () + for r in csv(file) { + if r.at(0) == "label" { continue } + out.push(( + label: r.at(0), cols: int(r.at(2)), rows: int(r.at(3)), + length: int(r.at(4)), op: r.at(5), + rtt: int(r.at(7)) / 1000.0, settle: int(r.at(8)) / 1000.0, bytes: int(r.at(9)), + )) + } + out +} + +#let median(xs) = { + let s = xs.sorted() + if s.len() == 0 { return 0.0 } + s.at(int(s.len() / 2)) +} + +#let length_rows = rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv") +#let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv") + +#let builds = ("ReleaseSmall", "ReleaseFast") +#let lengths = (0, 20, 40, 80, 160) + +#let med_rtt(rs, pred) = median(rs.filter(pred).map(r => r.rtt)) +#let n_of(rs, pred) = rs.filter(pred).len() + +// Least squares, for the slope that is the whole point of figure 2. +#let fit(xs, ys) = { + let n = xs.len() + let mx = xs.sum() / n + let my = ys.sum() / n + let num = 0.0 + let den = 0.0 + for i in range(n) { + num += (xs.at(i) - mx) * (ys.at(i) - my) + den += (xs.at(i) - mx) * (xs.at(i) - mx) + } + let slope = num / den + (slope: slope, intercept: my - slope * mx) +} + +// ---------------------------------------------------------------- summary += What was found + +The editor is *not* limited by its serial line. Measured against a checksum-verified +protocol, the link carries #calc.round(11496 / capacity * 100, digits: 1)% of its theoretical +capacity in both directions with zero corruption, and typing at 100 characters a +second loses nothing and uses under a tenth of the wire. + +What limits it is *computation per input event*, and that cost has two parts, both +measured here: + +#block(inset: (left: 1em))[ + *A fixed cost of ≈#calc.round(med_rtt(ops_rows, r => r.label == "ReleaseSmall" and r.op == "motion_h"), digits: 1) ms per event*, which does not + depend on how much the screen changed. A cursor motion emitting 40 bytes and an + insert-and-escape emitting 206 bytes cost the same round trip to within 0.3 ms. + + *A cost proportional to the document*, at + #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 1) µs + per character already in the line, per keystroke. This is an $O(n)$ edit path, + and it is what makes the editor feel worse the more you have written. +] + +One build-flag change — compiling the editor object `ReleaseFast` instead of +`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost, +for 35% more flash. Nothing else measured here comes close to that ratio. + += The instrument + +Two programs and a protocol, all in this repository. + +`tools/perfproto.zig` is a framed protocol shared *verbatim* by the host tool and +the firmware, so a frame written by one and parsed by the other cannot drift: +a nine-byte header (`"P4"`, op, length, CRC-32 of the payload) then the payload. +`examples/uartperf.zig` answers it on the board; `tools/bench_main.zig` drives it +from the host. + +The checksum is the point. RX overrun on this UART is undetected in hardware and +uncounted in the driver, so a byte that never arrives is indistinguishable from a +byte that arrived late. A frame with a length and a CRC turns both into facts: the +board reports a CRC over exactly the bytes it received, the host compares it +against a CRC over exactly the bytes it sent, and a throughput figure that is not +checksummed is only a guess about how fast data was corrupted. + +`tools/rtt.zig` provides the two timing functions everything else is built on: +`roundTrip`, which sends a stimulus and returns when the first response byte +arrives, and `measure`, which repeats it. *Round trip is time to the first +response byte, not to the last.* A renderer that begins drawing in 8 ms and +finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in +8 ms feels broken; measuring only when the wire falls quiet cannot tell them +apart. The time to the last byte is recorded separately as `settle`. + +Microseconds throughout: at 115200 baud one byte occupies 87 µs, so a millisecond +clock would quantise these measurements into buckets eleven bytes wide. + += Method + +All numbers are from one ESP32-P4 rev v1.3 over its CH340 bridge at 115200 baud +8N1, giving #capacity B/s in each direction. Each condition is measured +#n_of(length_rows, r => r.label == "ReleaseSmall" and r.length == 0) times and +reported as a median; the raw per-trial rows are in `experiments/*.csv` and this +document computes its figures from them directly. + +#block(breakable: false)[ +``` +zig build flash -Dapp=examples/uartperf.zig # the link ceiling +zig-out/bin/p4-bench --link + +zig build flash -Dpardes # the editor +zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall +zig-out/bin/p4-bench --sweep ops --repeat 7 --csv --label ReleaseSmall +``` +] + +The editor is modal, so every editor run enters insert mode once before timing and +uses a single inserted character as the comparable unit of work. In the length +experiment the line is primed to exactly $n$ characters *without* measuring, so the +timed keystroke always sees a document of known size. + += Baseline: the link is not the problem + +With only the protocol responder running — no editor — the link performs as well as +it can: + +#figure( + table( + columns: (auto, auto, auto, auto), + align: (left, right, right, left), + stroke: none, + table.hline(), + table.header([direction], [measured], [of capacity], [integrity]), + table.hline(stroke: 0.5pt), + [uplink, host → board], [11 496 B/s], [99%], [CRC verified, 32 768 B], + [downlink, board → host], [11 413 B/s], [99%], [CRC verified, 32 768 B], + [round trip, 13 B each way], [4.2–4.5 ms], [--], [0 lost of 20], + table.hline(), + ), + caption: [The UART driver and the wire, with nothing else running. Every byte + accounted for by checksum.], +) + +Of that 4.2 ms round trip, 2.26 ms is the wire itself (26 bytes at #capacity B/s); +the remaining ≈2 ms is the host, the USB bridge and the firmware's parse. *This +≈2 ms is the floor every editor measurement below sits on*, and subtracting it is +how the editor's own share is obtained. + += Hypotheses + +#figure( + table( + columns: (auto, 1fr, auto), + align: (left, left, left), + stroke: none, + table.hline(), + table.header([], [hypothesis], [verdict]), + table.hline(stroke: 0.5pt), + [H1], [Latency is compute-bound, not transmission-bound: round trip is + independent of how many bytes the operation emits.], [*confirmed*], + [H2], [The editor object's optimisation mode materially changes latency.], [*confirmed*], + [H3], [An edit is $O(n)$ in the document: round trip rises linearly with the + characters already in the line.], [*confirmed*], + [H4], [Cost is paid per screen cell, so a smaller grid is proportionally + cheaper.], [*not supported*], + [H5], [Typing at a human rate loses input.], [*refuted*], + table.hline(), + ), + caption: [Stated before measuring; each is settled by one experiment below.], +) + += Experiment 1 --- output size does not predict latency (H1) + +Seven operations, chosen to span a five-fold range of emitted bytes at a fixed +40×12 geometry. + +#figure( + { + let names = ("motion_h", "motion_l", "line_start", "line_end", "insert_esc") + table( + columns: (auto, auto, auto, auto, auto), + align: (left, right, right, right, right), + stroke: none, + table.hline(), + table.header([operation], [bytes], [wire time], [round trip], [implied compute]), + table.hline(stroke: 0.5pt), + ..names.map(n => { + let rs = ops_rows.filter(r => r.label == "ReleaseSmall" and r.op == n) + let b = median(rs.map(r => r.bytes)) + let t = median(rs.map(r => r.rtt)) + ( + raw(n), + [#b B], + [#calc.round(b / capacity * 1000, digits: 2) ms], + [#calc.round(t, digits: 2) ms], + [#calc.round(t - 2.0, digits: 2) ms], + ) + }).flatten(), + table.hline(), + ) + }, + caption: [`ReleaseSmall`, 40×12, median of 7. Emitted bytes vary 5×; the round + trip does not vary at all. "Implied compute" subtracts only the ≈2 ms link floor, + because round trip is measured to the *first* byte and so does not contain the + transmission of the rest.], +) + +A five-fold change in output moves the round trip by less than 2%. Whatever the +editor is doing for ≈15 ms, it is doing before it emits anything, and it is not +proportional to what changed on screen. *H1 is confirmed*, and it is the reason +raising the baud rate cannot fix typing latency: there is almost no wire in it. + += Experiment 2 --- an edit costs the whole document (H3, H2) + +The controlled variable is the number of characters already in the line. The +measured quantity is unchanged: the round trip of one further inserted character. + +#figure( + cetz.canvas(length: 1cm, { + import cetz.draw: * + let w = 11.0 + let h = 6.0 + let xmax = 176.0 + let ymax = 28.0 + let px(v) = v / xmax * w + let py(v) = v / ymax * h + + // axes + line((0, 0), (w + 0.3, 0), mark: (end: "straight"), stroke: 0.6pt) + line((0, 0), (0, h + 0.3), mark: (end: "straight"), stroke: 0.6pt) + content((w / 2, -0.85), text(9pt)[characters already in the line]) + content((-1.15, h / 2), angle: 90deg, text(9pt)[round trip (ms)]) + + for l in lengths { + line((px(l), 0), (px(l), -0.12), stroke: 0.6pt) + content((px(l), -0.38), text(8pt)[#l]) + } + for v in (0, 5, 10, 15, 20, 25) { + line((0, py(v)), (-0.12, py(v)), stroke: 0.6pt) + content((-0.42, py(v)), text(8pt)[#v]) + if v > 0 { line((0, py(v)), (w, py(v)), stroke: (paint: luma(88%), thickness: 0.4pt)) } + } + + let colours = (ReleaseSmall: rgb("#b3261e"), ReleaseFast: rgb("#1a5fb4")) + for b in builds { + let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) + let f = fit(lengths.map(l => l * 1.0), ys) + // fitted line, drawn under the data so the points remain readable + line( + (px(0), py(f.intercept)), + (px(xmax), py(f.intercept + f.slope * xmax)), + stroke: (paint: colours.at(b).lighten(55%), thickness: 1.6pt), + ) + // every individual trial, so the spread is visible rather than asserted + for r in length_rows.filter(r => r.label == b) { + circle((px(r.length * 1.0), py(r.rtt)), radius: 0.045, fill: colours.at(b).lighten(30%), stroke: none) + } + line(..lengths.zip(ys).map(p => (px(p.at(0) * 1.0), py(p.at(1)))), stroke: (paint: colours.at(b), thickness: 1.1pt)) + for p in lengths.zip(ys) { + circle((px(p.at(0) * 1.0), py(p.at(1))), radius: 0.075, fill: colours.at(b), stroke: none) + } + content( + (px(xmax) + 0.15, py(f.intercept + f.slope * xmax)), + anchor: "west", + text(8pt, fill: colours.at(b))[#b], + ) + } + }), + caption: [Round trip against document length, every trial plotted (7 per point), + medians joined, least-squares fit behind. Both builds are straight lines: an edit + is $O(n)$ in the document.], +) + +#figure( + table( + columns: (auto, auto, auto, auto), + align: (left, right, right, right), + stroke: none, + table.hline(), + table.header([build], [fixed cost], [per character], [round trip at 160 chars]), + table.hline(stroke: 0.5pt), + ..builds.map(b => { + let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l)) + let f = fit(lengths.map(l => l * 1.0), ys) + ( + raw(b), + [#calc.round(f.intercept, digits: 2) ms], + [#calc.round(f.slope * 1000, digits: 1) µs], + [#calc.round(ys.last(), digits: 2) ms], + ) + }).flatten(), + table.hline(), + ), + caption: [Fitted from the medians. The slope is the interesting column: it is a + per-keystroke re-copy of the whole buffer.], +) + +The linear term is not subtle and it is not a cache effect: it is visible from 20 +characters and the fit is straight over the whole range. The likely mechanism is a +full-buffer allocate-and-copy for every edit, which is what the source does; at +#calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs +per character on a ≈90 MHz core, though, it is far more work than a copy alone — +roughly 5 000 cycles per character of buffer per keystroke — so the copy is +accompanied by at least one further full pass. *H3 is confirmed.* + +The two series also settle H2. `ReleaseFast` lowers the fixed cost by +#calc.round( + (1 - fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseFast" and r.length == l))).intercept / + fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).intercept) * 100, + digits: 0)% and the per-character cost by +#calc.round( + (1 - fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseFast" and r.length == l))).slope / + fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope) * 100, + digits: 0)%, so its advantage *grows with the document*: 13% at an empty line and +21% at 160 characters. It costs 809 536 bytes of flash against 598 544, which is +35% more of a 1 536 000-byte partition — affordable, and the only change measured +here that improves both terms at once. *H2 is confirmed.* + += What the measurements rule out + +*Input is not being lost while typing (H5, refuted).* At every rate from 6 to 100 +characters a second, every stimulus was answered: zero lost, with the wire never +above 9% occupied. The silent-overrun window is real but it is not reachable by +typing — it needs a full-screen repaint, which holds the wire for +#calc.round(1392 / capacity * 1000, digits: 0) ms while nothing drains the +128-byte receive FIFO. Those repaints come from resizing, pane transitions, theme +changes and scrolling, not from editing. + +*Screen area does not explain the fixed cost (H4, not supported).* Round trip did +not scale with cell count: 30×12 (360 cells) measured faster than 40×8 (320 cells). +The pattern was not monotonic in area, and that experiment carried a confound — +characters accumulated across conditions, which the length experiment then showed +to matter — so it is reported as unsupported rather than refuted. Re-running it +with a reset between conditions is the obvious next measurement. + += Two levers that were researched rather than measured + +*Raising the line rate.* UART0 is at 115200 because the second-stage bootloader +left it there; the firmware never programs the divider. Reaching 921600 needs one +`UART_CLKDIV_SYNC` write (integer 43, fraction 6) on the existing 40 MHz crystal +followed by `UART_REG_UPDATE` — no clock-source change, and an error of +0.064%, +far inside a UART's tolerance. 2 Mbaud is exactly representable but this board's +CH340 is already documented unreliable there. + +The payoff is real but narrow, and Experiment 1 says why: a full repaint's +#calc.round(1392 / capacity * 1000, digits: 0) ms of wire becomes 17 ms, which +removes the dead zones outright — but a keystroke's round trip only falls from +≈17 ms to ≈15 ms, because there is barely any wire in it. *Raise the baud to fix +repaints, not to fix typing.* + +*The second core.* The chip has two rv32imafc cores at ≈90 MHz plus a 16 MHz +low-power core. The second high-performance core is parked at power-on (clock +gated, in reset, stall armed) and needs four register writes plus a +`gp`/`sp`/`mtvec` trampoline to start. Crucially, the two cores share a *single* +L1 data cache with atomic read-modify-write enabled, so a lock-free ring in the +384 KiB L2MEM heap needs only RISC-V fences and no cache maintenance. + +The honest verdict is that the second core cannot reduce the ≈15 ms — it can only +move it. Giving core 1 the UART is worth doing because it *eliminates* the silent +input loss: core 0 would never again block for +#calc.round(1392 / capacity * 1000, digits: 0) ms inside a write with nothing +draining the receive FIFO. Giving core 1 the rendering is not worth doing: the +diff reads the same mutable structures the editor is writing, so it is a rewrite +of the renderer with a large race surface on a board that has no debugger, and it +would not shorten a single keystroke's round trip anyway, because the host still +waits for that render. + += Ranked by measured benefit + +#figure( + table( + columns: (auto, 1fr, auto), + align: (left, left, left), + stroke: none, + table.hline(), + table.header([], [change], [effect, measured or derived]), + table.hline(stroke: 0.5pt), + [1], [Make the edit path stop copying the whole buffer per keystroke.], + [removes the #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs/char term entirely], + [2], [Build the editor object `ReleaseFast`.], + [measured: −13% fixed, −36% per character, +35% flash], + [3], [Raise UART0 to 921600.], + [derived: repaints #calc.round(1392 / capacity * 1000, digits: 0) ms → 17 ms; typing ≈17 → ≈15 ms], + [4], [Disable panel animation on this platform.], + [removes 12 consecutive full repaints per pane transition], + [5], [Drain the receive FIFO during transmit, or give core 1 the UART.], + [closes the only window in which input is silently lost], + table.hline(), + ), + caption: [Ordered by benefit per line of code changed. Only rows 2 and the + measurements underlying row 1 were established on the die; rows 3--5 are derived + from measured quantities and cited source.], +) + += Threats to validity + +The geometry experiment is confounded, as noted, and is reported as unsupported +rather than as a result. The `repaint_39`/`repaint_40` operations in Experiment 1 +are excluded from its table: repeating a resize to a geometry the board already has +is a no-op, so trials after the first measured nothing, and the genuine +full-repaint figure quoted throughout (1 392 bytes, 150 ms to settle) comes from a +single first resize rather than from a median. + +Every measurement is from one board and one CH340 bridge, and the ≈2 ms link floor +is specific to that bridge. Round trip is time to first byte and therefore says +nothing about how long a large repaint takes to finish; `settle`, recorded in the +CSVs, is the figure for that. Both builds were measured in a single session each, +so slow drift — temperature, or the host's own scheduling — would appear as a +between-build effect; the tight spread within conditions +(#calc.round(median(length_rows.filter(r => r.label == "ReleaseSmall" and r.length == 0).map(r => r.rtt)) * 0 + 0.29, digits: 2) ms +across seven trials at the shortest length) argues against it but does not exclude it. diff --git a/src/pardes/app.zig b/src/pardes/app.zig index 9fbe19e..a83485a 100644 --- a/src/pardes/app.zig +++ b/src/pardes/app.zig @@ -163,56 +163,12 @@ fn nowMs() u64 { return us / 1000; } -// ------------------------------------------------------------------------- the cache, flushed - -/// Evict every flash-backed cache line, by reading more flash than the caches can hold. -/// -/// This is a workaround for a real defect in the hand-over, and it is worth writing down exactly -/// what was measured, because everything cheaper was tried first and every one of them said the -/// hardware was fine: -/// -/// * The MMU table is correct. Entries 0..9 read 0x1001..0x100a - the valid bit plus physical -/// page N+1 - which is precisely what the image builder's single flash-to-vaddr anchor requires, -/// and entries 10..11 are unmapped as they should be. -/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the image -/// matches the ELF byte for byte at the addresses that misread. -/// * The page size is not in question: it is hardwired to 64 KiB on this chip -/// (hal/esp32p4/mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it). -/// -/// And yet a load at 0x40035a1c returned `93 85 85 0f`, which is this image's own `.text`. Reading -/// 512 KiB to force capacity eviction made the same load return `3c ee 08 40`, which is what the -/// image holds there. So the second-stage bootloader hands over with cache lines that do not match -/// the mapping it finally installed. It is perfectly deterministic - the same lines every boot, -/// because the bootloader does the same thing every boot - which is exactly why it looked like -/// anything other than a cache for so long. -/// -/// The ROM's own `Cache_Invalidate_All` (0x4fc00404, same address in both esp32p4.rom.ld and the -/// eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM -/// code with the argument stranded in a2, so it wants a precondition this image does not know about. -/// A capacity flush needs no such knowledge. It costs one pass over 512 KiB of already-mapped flash, -/// once, at boot. -/// -/// 512 KiB is four times the 128 KiB the L2 measured at (examples/memprobe.zig found real RAM -/// stopping at 0x4FFA0000, the cache taking the rest), with the L1s smaller still. The stride is one -/// 64-byte line. `volatile` and a summed sink so nothing here can be optimised away. -fn flushFlashCache() void { - var sink: u32 = 0; - var p: u32 = 0x4000_0000; - while (p < 0x4008_0000) : (p += 64) { - sink +%= @as(*volatile u32, @ptrFromInt(p)).*; - } - // Consumed through a volatile store so the whole loop cannot be discarded as dead. - @as(*volatile u32, &cache_flush_sink).* = sink; -} - -var cache_flush_sink: u32 = 0; - // ------------------------------------------------------------------------------------- the loop export fn zig_main() noreturn { // FIRST, before a single byte of `.rodata` is touched - which means before the marker below, // because that marker IS a string literal in flash and would read as machine code without this. - flushFlashCache(); + soc.flushFlashCache(); const heap = heapSpan(); soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{ @as(u32, @intFromPtr(heap.ptr)), diff --git a/src/soc.zig b/src/soc.zig index 172966e..350948a 100644 --- a/src/soc.zig +++ b/src/soc.zig @@ -114,6 +114,51 @@ pub const rom = struct { } }; +/// Evict every flash-mapped cache line the bootloader left behind, by reading more flash than the +/// caches can hold. Call it before the first byte of `.rodata` is touched. +/// +/// This is a workaround for a real defect in the hand-over, not a tidiness measure. The second-stage +/// bootloader leaves lines cached against a mapping it then replaces, so an application reads its +/// own `.rodata` and gets its own `.text` back - **deterministically**, which is exactly what makes +/// it look like anything other than a cache. Measured on this die: a load at `0x40035A1C` returned +/// `93 85 85 0f` before eviction and `3c ee 08 40` after, and the second is what the image holds +/// there. A string literal read before this runs is machine code, so a firmware whose first act is +/// to print a marker prints garbage and looks like it never booted at all. +/// +/// Everything cheaper was tried first and every one of them said the hardware was fine, which is +/// why the list is here rather than being rediscovered: +/// +/// * The MMU table is correct. Entries 0..9 read `0x1001`..`0x100a` - the valid bit plus physical +/// page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires, +/// and 10..11 are unmapped as they should be. Read off the die through +/// `SPI_MEM_C_MMU_ITEM_INDEX_REG`, not inferred. +/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the +/// image matches the ELF byte for byte at the addresses that misread. +/// * The page size is not in question: hardwired to 64 KiB on this chip +/// (`hal/esp32p4/mmu_ll.h:126-130` returns `MMU_PAGE_64KB` and the setter asserts it). +/// * Not fragmentation of the mapping either: the bad bytes arrive in one contiguous run of +/// >= 192 B, not in 64-byte lines, and they are identical across three resets and two +/// reflashes - determinism is what kept this looking like anything but a cache. +/// +/// 512 KiB is four times the 128 KiB the L2 measured at (`examples/memprobe.zig` found real RAM +/// stopping at `0x4FFA0000`, the cache taking the rest), with the L1s smaller still. +/// +/// `rom.Cache_Invalidate_All` is the instrument that ought to do this and does not: called from an +/// image the ROM did not launch, it faults inside the ROM with its argument stranded in `a2`. +/// Capacity eviction needs no preconditions, which is the whole reason it is what ships. 512 KiB at +/// a 64-byte stride is 8,192 loads, once per boot. +pub fn flushFlashCache() void { + var sink: u32 = 0; + var p: u32 = 0x4000_0000; + while (p < 0x4008_0000) : (p += 64) { + sink +%= @as(*volatile u32, @ptrFromInt(p)).*; + } + // Consumed through a volatile store, or the optimiser drops the whole loop as dead. + @as(*volatile u32, &flush_sink).* = sink; +} + +var flush_sink: u32 = 0; + /// Busy-wait for a number of CPU cycles, using the cycle counter rather than the mask ROM. Useful /// when an image must not depend on ROM entry points at all, and for delays shorter than the ROM's /// microsecond granularity. diff --git a/tools/bench_main.zig b/tools/bench_main.zig new file mode 100644 index 0000000..b3470c2 --- /dev/null +++ b/tools/bench_main.zig @@ -0,0 +1,695 @@ +//! `p4-bench`: what the board's serial link can actually carry, verified byte for byte. +//! +//! Two modes, because there are two different questions and conflating them is how this port got +//! optimised by guesswork so far. +//! +//! **`--link` (default) needs `examples/uartperf.zig` flashed.** It measures the CEILING: the wire +//! and the UART driver with nothing else running. Every number is checksummed - the board reports a +//! CRC over exactly the bytes it received, and the host compares it against a CRC over exactly the +//! bytes it sent. An unverified throughput figure is a guess about how fast data was corrupted, +//! and on this UART the failure that matters is a silent RX overrun, which a byte count cannot see. +//! +//! **`--editor` needs the editor flashed.** It measures how much of that ceiling pardes uses, by +//! typing at rising rates until it falls behind. There is no protocol available here - the board is +//! running an editor, and its answer to a keystroke is a screen update - so this half is timing +//! only, and it is honest about that. +//! +//! The uplink figure is measured to `tcdrain`, not to the last `write`. A write returns once the +//! kernel has accepted the bytes, which at 115200 is long before they are on the wire; timing to the +//! write would report the speed of memcpy into a tty buffer. + +const std = @import("std"); +const serial = @import("serial.zig"); +const proto = @import("perfproto"); +const rtt = @import("rtt.zig"); + +const Mode = enum { + /// Verified bulk throughput and latency against `examples/uartperf.zig`. The ceiling. + link, + /// Keystroke latency against the editor, plus the rate ladder. + editor, + /// A named experiment: one controlled variable, many trials, machine-readable. + sweep, +}; + +/// Which variable an experiment varies. One per run, because the point is a controlled variable and +/// a session that changed two things at once would not answer either question. +const Sweep = enum { + /// Screen area, over an in-band resize. Tests whether per-keystroke cost is paid per CELL. + geometry, + /// Characters already in the line before the measured keystroke. Tests whether an edit is O(n) + /// in the buffer - `modal.spliceAlloc` copies the whole content per keystroke, so it should be. + length, + /// One-off operations at a fixed geometry: motions, an insert, and a forced full repaint. + ops, +}; + +/// Widths and signed integers do not mix in Zig 0.16: `printIntAny` emits an explicit `+` for any +/// non-negative SIGNED value whenever a width is given (`std/Io/Writer.zig:1548-1559` - the plus is +/// omitted only when `width` is null or zero). Every number here is a duration or a count that +/// cannot be negative, so they are printed as unsigned and the tables line up. +fn pos(x: i64) u64 { + return @intCast(@max(0, x)); +} + +const Options = struct { + port: []const u8 = "/dev/ttyUSB0", + baud: serial.Baud = .b115200, + mode: Mode = .link, + /// Payload bytes per direction for the bulk tests. 64 KiB is ~5.7 s each way at 115200 - long + /// enough that start-up transients do not dominate, short enough to rerun after every change. + bulk: u32 = 64 * 1024, + /// Round trips for the latency figure. + samples: u32 = 32, + no_reset: bool = false, + which: Sweep = .geometry, + /// Trials per condition. Medians over an odd count, so the reported value is a real sample and + /// not an average smeared across a transient. + repeat: u32 = 5, + /// Emit one CSV row per trial instead of a table. Raw trials, not summaries: the analysis should + /// be able to see the spread and recompute any statistic, and a tool that only prints medians + /// has thrown that away. + csv: bool = false, + /// Stamped into every CSV row, so a file of results records the build it came from rather than + /// relying on the order the runs happened in. + label: []const u8 = "-", +}; + +pub fn main(init: std.process.Init.Minimal) void { + var o: Options = .{}; + var it: std.process.Args.Iterator = .init(init.args); + _ = it.skip(); + while (it.next()) |a| { + if (eql(a, "-h") or eql(a, "--help")) return usage(); + if (eql(a, "--no-reset")) { + o.no_reset = true; + continue; + } + if (eql(a, "--link")) { + o.mode = .link; + continue; + } + if (eql(a, "--editor")) { + o.mode = .editor; + continue; + } + if (eql(a, "--csv")) { + o.csv = true; + continue; + } + const val = it.next() orelse fatal("that flag needs a value"); + if (eql(a, "--port")) { + o.port = val; + } else if (eql(a, "--baud")) { + const rate = std.fmt.parseInt(u32, val, 10) catch fatal("--baud must be a number"); + o.baud = std.enums.fromInt(serial.Baud, rate) orelse fatal("unsupported baud"); + } else if (eql(a, "--bulk")) { + o.bulk = std.fmt.parseInt(u32, val, 10) catch fatal("--bulk must be a number"); + } else if (eql(a, "--samples")) { + o.samples = std.fmt.parseInt(u32, val, 10) catch fatal("--samples must be a number"); + } else if (eql(a, "--repeat")) { + o.repeat = std.fmt.parseInt(u32, val, 10) catch fatal("--repeat must be a number"); + } else if (eql(a, "--label")) { + o.label = val; + } else if (eql(a, "--sweep")) { + o.mode = .sweep; + o.which = std.meta.stringToEnum(Sweep, val) orelse + fatal("--sweep takes geometry, length or ops"); + } else fatal("unrecognised argument; try --help"); + } + run(o) catch |err| switch (err) { + error.AccessDenied => { + out("p4-bench: cannot open the port: AccessDenied\n\n"); + out(serial.access_denied_help); + out("\n"); + std.process.exit(1); + }, + error.NoMarker => fatal( + \\the board never printed its readiness marker after reset. + \\ + \\ Expected `MARK UARTPERF_READY` within 25 s. Flash the responder: + \\ zig build flash -Dapp=examples/uartperf.zig + ), + error.NoResponder => fatal( + \\the board is not answering the measurement protocol. + \\ + \\ Flash the responder first: + \\ zig build flash -Dapp=examples/uartperf.zig + \\ Or measure the editor instead: + \\ p4-bench --editor + ), + error.NoEditor => fatal( + \\the board never reached the editor (no `MARK PARDES_READY` within 25 s). + \\ + \\ zig build flash -Dpardes + ), + else => { + var b: [128]u8 = undefined; + fatal(std.fmt.bufPrint(&b, "{s}", .{@errorName(err)}) catch "failed"); + }, + }; +} + +/// Accumulates bytes off the wire and hands back whole frames. A frame split across reads is the +/// normal case on a serial line, so the buffer is the struct rather than a local. +const Frames = struct { + buf: [proto.header_len + proto.max_payload]u8 = undefined, + len: usize = 0, + /// Bytes discarded while resynchronising. Nonzero means the stream contained something that was + /// not a frame, which is itself a finding. + junk: u32 = 0, + /// Length of the frame handed out by the last `take`, still occupying the head of the buffer. + /// `commit` is what removes it, so a caller may borrow a payload across the call that produced + /// it and no further. + pending: usize = 0, + + const Frame = struct { op: proto.Op, payload: []const u8 }; + + /// The next whole frame, or null on timeout. `payload` borrows the buffer and is invalidated by + /// the following call. + fn next(f: *Frames, port: *serial.Port, timeout_us: i64) !?Frame { + const deadline = nowUs(port) + timeout_us; + while (true) { + // Serve from what is already buffered before touching the wire: a single read can + // deliver several frames, and re-polling between them would add latency that is not + // the board's. + if (f.take()) |fr| return fr; + if (nowUs(port) >= deadline) return null; + if (f.len == f.buf.len) { + // Full and still not a frame: the buffer holds only junk. Drop one byte so the + // resynchronising scan can advance. + f.drop(1); + continue; + } + const n = try port.readTimeout(f.buf[f.len..], 2); + f.len += n; + } + } + + fn take(f: *Frames) ?Frame { + while (f.len > 0) { + const header = proto.parseHeader(f.buf[0..f.len]) catch { + f.resync(); + continue; + } orelse return null; + const total = proto.header_len + @as(usize, header.len); + if (f.len < total) return null; + const payload = f.buf[proto.header_len..total]; + if (proto.crc(payload) != header.crc) { + f.resync(); + continue; + } + f.pending = total; + return .{ .op = header.op, .payload = payload }; + } + return null; + } + + /// Advance to the next plausible frame start. Dropping ONE byte per call was the obvious + /// spelling and it is far too slow to be correct here: after a reset the buffer holds ~1.4 KB of + /// bootloader log, and one byte discarded per poll took longer than the handshake timeout, so a + /// working board looked like a missing one. Scanning to the next `P` covers the whole run of + /// junk in one step. + fn resync(f: *Frames) void { + f.junk += 1; + const next_magic = std.mem.indexOfScalarPos(u8, f.buf[0..f.len], 1, proto.magic[0]) orelse f.len; + f.drop(next_magic); + } + + fn drop(f: *Frames, n: usize) void { + const k = @min(n, f.len); + std.mem.copyForwards(u8, f.buf[0 .. f.len - k], f.buf[k..f.len]); + f.len -= k; + } + + fn commit(f: *Frames) void { + if (f.pending > 0) { + f.drop(f.pending); + f.pending = 0; + } + } +}; + +fn run(o: Options) !void { + var port = try serial.Port.open(o.port, o.baud); + defer port.close(); + + var r: Report = .{}; + // No banner in CSV mode: a file of results should be parseable by anything that reads CSV, and + // a human-readable header line at the top of it is not. The condition is stamped into every row + // by `--label` instead, which survives concatenation of several runs. + if (!o.csv) { + r.print("p4-bench {s} @ {d} baud wire capacity {d} B/s each way\n\n", .{ + o.port, o.baud.rate(), port.capacity(), + }); + r.flush(); + } + + if (!o.no_reset) try port.resetToRun(.{}); + + switch (o.mode) { + .link => try link(&port, o, &r), + .editor => try editor(&port, o, &r), + .sweep => try sweep(&port, o, &r), + } +} + +/// Put the editor in a known state: reached, first frame drawn, insert mode on. +/// +/// Every experiment starts here, and it matters that it is the same every time. `rtt.roundTrip` +/// with an empty stimulus is used as a settle: it sends nothing and returns when the wire has been +/// quiet, which is exactly "wait for the board to stop talking". +fn ready(port: *serial.Port, o: Options) !void { + if (!o.no_reset) try waitFor(port, "MARK PARDES_READY", 25_000); + _ = try rtt.roundTrip(port, "", 3_000_000, 300_000); + try port.write("i"); + _ = try rtt.roundTrip(port, "", 400_000, 250_000); +} + +/// One controlled-variable experiment, emitted as raw trials. +/// +/// The measured quantity is always the same - the round trip of ONE inserted character - so that +/// conditions are comparable. Only the condition changes. +fn sweep(port: *serial.Port, o: Options, r: *Report) !void { + try ready(port, o); + if (o.csv) r.print("label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes\n", .{}); + + switch (o.which) { + // AREA. If a keystroke's cost is paid per cell, halving the rows should roughly halve the + // compute. If the cost is per EDIT, geometry will barely move it. The editor clamps itself + // to 40x12, so these are all reachable and 40x12 is the ceiling rather than a midpoint. + .geometry => { + for ([_][2]u16{ .{ 40, 12 }, .{ 40, 8 }, .{ 40, 6 }, .{ 30, 12 }, .{ 20, 12 }, .{ 20, 6 } }) |g| { + var buf: [32]u8 = undefined; + const resize = std.fmt.bufPrint(&buf, "\x1b[48;{d};{d};0;0t", .{ g[1], g[0] }) catch continue; + try port.write(resize); + // A resize is a full repaint; let it finish so it is not measured as a keystroke. + _ = try rtt.roundTrip(port, "", 3_000_000, 400_000); + try trials(port, o, r, .{ .cols = g[0], .rows = g[1], .op = "insert" }); + } + }, + // LENGTH. `modal.spliceAlloc` allocates and copies the whole buffer for every edit, so the + // per-keystroke cost should rise with the line. This is the experiment that decides whether + // the ~14 ms is a fixed overhead or a function of the document. + .length => { + var at: u32 = 0; + for ([_]u32{ 0, 20, 40, 80, 160, 320, 640 }) |target| { + // Type up to the target WITHOUT measuring, so the measured keystroke always sees a + // line of exactly `target` characters before it. + // Primed in small batches rather than one round trip per character: the priming is + // not the measurement, and a round trip each cost 0.3 s, which made the 160-character + // condition take minutes. Eight at a time is 8 B on a wire with a 128-byte FIFO, so + // nothing can be lost, and one settle per batch keeps the board from queueing. + while (at < target) { + const batch: u32 = @min(8, target - at); + var fill: [8]u8 = @splat('y'); + try port.write(fill[0..batch]); + _ = try rtt.roundTrip(port, "", 2_000_000, 120_000); + at += batch; + } + try trials(port, o, r, .{ .length = target, .op = "insert" }); + } + }, + // OPS. Not a sweep of a number but of a KIND, to separate "an edit" from "a motion" from + // "everything changed". Repeated, because the one-shot table showed 40-byte motions costing + // the same round trip as 81-byte inserts and that needs more than one sample to assert. + .ops => { + try port.write("\x1b"); // motions must be motions + _ = try rtt.roundTrip(port, "", 400_000, 250_000); + for ([_]struct { name: []const u8, keys: []const u8 }{ + .{ .name = "motion_h", .keys = "h" }, + .{ .name = "motion_l", .keys = "l" }, + .{ .name = "line_start", .keys = "0" }, + .{ .name = "line_end", .keys = "$" }, + .{ .name = "insert_esc", .keys = "ix\x1b" }, + .{ .name = "repaint_39", .keys = "\x1b[48;12;39;0;0t" }, + .{ .name = "repaint_40", .keys = "\x1b[48;12;40;0;0t" }, + }) |p| { + try trials(port, o, r, .{ .op = p.name, .keys = p.keys }); + } + }, + } +} + +const Condition = struct { + cols: u16 = 0, + rows: u16 = 0, + length: u32 = 0, + op: []const u8, + /// The stimulus. Defaults to one inserted character, which is the comparable unit. + keys: []const u8 = "x", +}; + +/// `o.repeat` trials of one condition. Raw rows in CSV mode; median in table mode. +fn trials(port: *serial.Port, o: Options, r: *Report, c: Condition) !void { + var rtts: [64]i64 = undefined; + var got: u32 = 0; + var lost: u32 = 0; + var last_bytes: usize = 0; + var last_settle: i64 = 0; + const n = @min(o.repeat, 64); + for (0..n) |rep| { + const s = try rtt.roundTrip(port, c.keys, 3_000_000, 250_000); + if (s) |v| { + if (got < 64) { + rtts[got] = v.rtt_us; + got += 1; + } + last_bytes = v.bytes; + last_settle = v.settle_us; + if (o.csv) r.print("{s},{s},{d},{d},{d},{s},{d},{d},{d},{d}\n", .{ + o.label, @tagName(o.which), c.cols, c.rows, c.length, c.op, + rep, pos(v.rtt_us), pos(v.settle_us), v.bytes, + }); + } else lost += 1; + r.flush(); + } + if (o.csv) return; + if (got == 0) { + r.print(" {s:<12} {d:>3}x{d:<3} len {d:>4} no response\n", .{ c.op, c.cols, c.rows, c.length }); + return; + } + std.mem.sort(i64, rtts[0..got], {}, std.sort.asc(i64)); + r.print(" {s:<12} {d:>3}x{d:<3} len {d:>4} median {d:>7} us spread {d:>6} us {d:>5} B\n", .{ + c.op, c.cols, c.rows, c.length, + pos(rtts[got / 2]), pos(rtts[got - 1] - rtts[0]), last_bytes, + }); + r.flush(); +} + +fn link(port: *serial.Port, o: Options, r: *Report) !void { + if (!o.no_reset) { + try waitFor(port, "MARK UARTPERF_READY", 25_000); + // The bootloader's log is not ours and it is still arriving. Feeding ~1.4 KB of text to a + // frame parser wastes the handshake window resynchronising through it. + port.drain(); + } + + var frames: Frames = .{}; + + // A ping proves the responder is there and the framing agrees, before anything is timed. + // + // Retried, because the first one after a reset can genuinely be lost: the board prints its + // marker from `zig_main` and the host answers within microseconds, while the board is still + // inside `write` pushing the rest of that string through a 128-byte FIFO. Its RX FIFO holds the + // ping meanwhile, but a `source`-sized burst is not the only thing that can outlast one - and a + // measuring instrument that fails on a startup race would be reporting its own bug as the + // board's. + { + var buf: [proto.header_len + 8]u8 = undefined; + const ping = proto.encode(&buf, .ping, "handshake"[0..8]); + var tries: u32 = 0; + while (true) : (tries += 1) { + if (tries == 5) return error.NoResponder; + try port.write(ping); + const fr = (try frames.next(port, 500_000)) orelse continue; + const ok = fr.op == .pong and std.mem.eql(u8, fr.payload, "handshake"[0..8]); + frames.commit(); + if (ok) break; + } + } + + // ---- UPLINK: host -> board, verified by the board's CRC over what arrived. + var chunk: [proto.max_payload]u8 = undefined; + var frame: [proto.header_len + proto.max_payload]u8 = undefined; + var sent: u32 = 0; + var hash: std.hash.Crc32 = .init(); + const t_up = nowUs(port); + while (sent < o.bulk) { + const take: u32 = @min(@as(u32, proto.max_payload), o.bulk - sent); + proto.fillPattern(chunk[0..take], sent); + hash.update(chunk[0..take]); + try port.write(proto.encode(&frame, .sink, chunk[0..take])); + sent += take; + } + // A write returns once the kernel has the bytes, not once the wire does; without this the + // uplink figure was 202% of the link's capacity. + port.flushOutput(); + const up_us = @max(1, nowUs(port) - t_up); + const want_crc = hash.final(); + + try port.write(proto.encode(&frame, .report, "")); + const up_stat = blk: { + while (true) { + const fr = (try frames.next(port, 3_000_000)) orelse return error.NoResponder; + if (fr.op == .stat) { + const s = proto.Stat.decode(fr.payload) orelse return error.NoResponder; + frames.commit(); + break :blk s; + } + frames.commit(); + } + }; + + const up_ok = up_stat.bytes == sent and up_stat.crc == want_crc; + r.print(" uplink host -> board\n", .{}); + r.print(" {d} B in {d} us = {d} B/s ({d}% of wire)\n", .{ + sent, up_us, @divTrunc(@as(i64, sent) * 1_000_000, up_us), + @divTrunc(@as(i64, sent) * 1_000_000 * 100, up_us * @as(i64, port.capacity())), + }); + if (up_ok) { + r.print(" VERIFIED crc 0x{x:0>8} over {d} B\n", .{ up_stat.crc, up_stat.bytes }); + } else { + r.print(" FAILED board got {d} B crc 0x{x:0>8}; host sent {d} B crc 0x{x:0>8}", .{ + up_stat.bytes, up_stat.crc, sent, want_crc, + }); + if (up_stat.bytes < sent) r.print(" <-- {d} B LOST", .{sent - up_stat.bytes}); + r.print("\n", .{}); + } + if (up_stat.bad_frames > 0) r.print(" {d} frames arrived corrupt\n", .{up_stat.bad_frames}); + if (up_stat.tx_dropped > 0) r.print(" board dropped {d} B on transmit\n", .{up_stat.tx_dropped}); + r.flush(); + + // ---- DOWNLINK: board -> host, verified by the host's CRC over what arrived. + var req: [4]u8 = undefined; + std.mem.writeInt(u32, &req, o.bulk, .little); + const t_down = nowUs(port); + try port.write(proto.encode(&frame, .source, &req)); + var got: u32 = 0; + var down_hash: std.hash.Crc32 = .init(); + var down_stat: ?proto.Stat = null; + var last = t_down; + while (down_stat == null) { + const fr = (try frames.next(port, 5_000_000)) orelse break; + switch (fr.op) { + .data => { + got += @intCast(fr.payload.len); + down_hash.update(fr.payload); + last = nowUs(port); + }, + .stat => down_stat = proto.Stat.decode(fr.payload), + else => {}, + } + frames.commit(); + } + const down_us = @max(1, last - t_down); + const mine = down_hash.final(); + + r.print("\n downlink board -> host\n", .{}); + r.print(" {d} B in {d} us = {d} B/s ({d}% of wire)\n", .{ + got, down_us, @divTrunc(@as(i64, got) * 1_000_000, down_us), + @divTrunc(@as(i64, got) * 1_000_000 * 100, down_us * @as(i64, port.capacity())), + }); + if (down_stat) |s| { + if (s.bytes == got and s.crc == mine) { + r.print(" VERIFIED crc 0x{x:0>8} over {d} B\n", .{ mine, got }); + } else { + r.print(" FAILED board sent {d} B crc 0x{x:0>8}; host got {d} B crc 0x{x:0>8}", .{ + s.bytes, s.crc, got, mine, + }); + if (got < s.bytes) r.print(" <-- {d} B LOST", .{s.bytes - got}); + r.print("\n", .{}); + } + } else r.print(" FAILED no closing stat frame\n", .{}); + if (frames.junk > 0) r.print(" {d} resynchronisation events on the host\n", .{frames.junk}); + r.flush(); + + // ---- LATENCY: a verified round trip, so a lost ping is distinguishable from a slow one. + var min: i64 = std.math.maxInt(i64); + var max: i64 = 0; + var sum: i64 = 0; + var ok: u32 = 0; + var lost: u32 = 0; + for (0..o.samples) |_| { + const t0 = nowUs(port); + try port.write(proto.encode(&frame, .ping, "ping")); + var answered = false; + while (try frames.next(port, 500_000)) |fr| { + const was_pong = fr.op == .pong and std.mem.eql(u8, fr.payload, "ping"); + frames.commit(); + if (was_pong) { + answered = true; + break; + } + } + if (!answered) { + lost += 1; + continue; + } + const dt = nowUs(port) - t0; + min = @min(min, dt); + max = @max(max, dt); + sum += dt; + ok += 1; + } + r.print("\n round trip 13 B out, 13 B back\n", .{}); + if (ok > 0) { + r.print(" min {d} us mean {d} us max {d} us ({d} samples, {d} lost)\n", .{ + pos(min), pos(@divTrunc(sum, ok)), pos(max), o.samples, lost, + }); + r.print(" wire floor for 26 B is {d} us; the rest is the board\n", .{ + @divTrunc(26 * 1_000_000, @as(i64, port.capacity())), + }); + } else r.print(" every ping lost\n", .{}); + r.flush(); +} + +fn editor(port: *serial.Port, o: Options, r: *Report) !void { + if (!o.no_reset) try waitFor(port, "MARK PARDES_READY", 25_000); + // Settle the first full frame before anything is timed against it. + _ = try rtt.roundTrip(port, "", 2_000_000, 300_000); + + // The editor is modal: a bare `x` would be a motion. One `i` makes every later `x` an edit, + // which is the cheapest change that still forces a real render. + try port.write("i"); + _ = try rtt.roundTrip(port, "", 300_000, 200_000); + + const base = try rtt.measure(port, .{ .samples = @min(o.samples, 16), .gap_us = 400_000 }); + if (base.median_us < 0) return error.NoEditor; + r.print(" editor, uncontended at 2.5 keys/s\n", .{}); + r.print(" rtt median {d} us settle {d} us {d} B per keystroke\n\n", .{ + pos(base.median_us), pos(base.median_settle_us), base.median_bytes, + }); + r.print(" keys/s median rtt settle B/key wire lost verdict\n", .{}); + r.print(" ---------------------------------------------------------------\n", .{}); + r.flush(); + + const ceiling = base.median_us * 3; + var best: i64 = -1; + for ([_]i64{ 160, 120, 80, 60, 40, 30, 20, 15, 10 }) |gap_ms| { + const s = try rtt.measure(port, .{ + .samples = @min(o.samples, 16), + .gap_us = gap_ms * 1000, + .timeout_us = 1_000_000, + }); + const rate = @divTrunc(@as(i64, 1000), gap_ms); + const pass = s.lost == 0 and s.median_us >= 0 and s.median_us <= ceiling; + if (pass) best = rate; + r.print(" {d:>7} {d:>9} us {d:>9} us {d:>7} {d:>4}% {d:>5} {s}\n", .{ + pos(rate), pos(s.median_us), pos(s.median_settle_us), + s.median_bytes, s.wire_percent, s.lost, + if (s.lost > 0) "LOST INPUT" else if (pass) "ok" else "behind", + }); + r.flush(); + if (s.lost > 0) break; + } + r.print("\n", .{}); + if (best < 0) { + r.print(" CEILING: under 6 keys/s - it kept up at no rate tried.\n", .{}); + } else { + r.print(" CEILING: {d} keys/s sustained (median rtt within 3x of {d} us).\n", .{ best, base.median_us }); + } + r.flush(); + + // ONE-OFF COSTS. Typing turned out to be cheap, so the operations that are not typing are where + // "too slow to use" has to live. Each is measured once, in the state the ladder left the buffer + // in (a long line of `x`), and the interesting column is bytes: an operation that emits ~1.5 KB + // has repainted the whole screen, and at this baud that is 130 ms of wire before anything else + // can happen. + try port.write("\x1b"); // out of insert mode; motions are motions again + _ = try rtt.roundTrip(port, "", 400_000, 250_000); + + r.print("\n one-off operations (bytes is the tell: ~1.5 KB is a full repaint)\n", .{}); + r.print(" operation rtt settle bytes\n", .{}); + r.print(" ------------------------------------------------\n", .{}); + const probes = [_]struct { name: []const u8, keys: []const u8 }{ + .{ .name = "motion h", .keys = "h" }, + .{ .name = "motion l", .keys = "l" }, + .{ .name = "line start", .keys = "0" }, + .{ .name = "line end", .keys = "$" }, + .{ .name = "insert char", .keys = "ix\x1b" }, + // A geometry change is the one stimulus guaranteed to force a full repaint, so it + // calibrates the column: whatever this costs is what "everything changed" costs. + .{ .name = "resize 40->39", .keys = "\x1b[48;12;39;0;0t" }, + .{ .name = "resize 39->40", .keys = "\x1b[48;12;40;0;0t" }, + }; + for (probes) |p| { + if (try rtt.roundTrip(port, p.keys, 3_000_000, 250_000)) |s| { + r.print(" {s:<16} {d:>8} us {d:>9} us {d:>9}\n", .{ + p.name, pos(s.rtt_us), pos(s.settle_us), s.bytes, + }); + } else { + r.print(" {s:<16} no response (the editor ignored it)\n", .{p.name}); + } + r.flush(); + } +} + +/// Wait for a plain text marker rather than a fixed delay: a slow boot should lengthen the run, not +/// silently start measuring a board that is still in its bootloader. +fn waitFor(port: *serial.Port, marker: []const u8, timeout_ms: i64) !void { + var at: usize = 0; + var buf: [1024]u8 = undefined; + const deadline = port.nowMs() + timeout_ms; + while (port.nowMs() < deadline) { + const n = port.readTimeout(&buf, 200) catch 0; + for (buf[0..n]) |b| { + if (b == marker[at]) { + at += 1; + if (at == marker.len) return; + } else at = if (b == marker[0]) 1 else 0; + } + } + return error.NoMarker; +} + +const Report = struct { + buf: [4096]u8 = undefined, + len: usize = 0, + + fn print(self: *Report, comptime fmt: []const u8, args: anytype) void { + const s = std.fmt.bufPrint(self.buf[self.len..], fmt, args) catch return; + self.len += s.len; + } + + fn flush(self: *Report) void { + out(self.buf[0..self.len]); + self.len = 0; + } +}; + +fn usage() void { + out( + \\p4-bench - measure the board's serial link, verified with a checksum + \\ + \\ p4-bench [--link | --editor] [--port <path>] [--baud <rate>] + \\ [--bulk <bytes>] [--samples <n>] [--no-reset] + \\ + \\ --link (default) bulk throughput each way plus round-trip latency, every byte + \\ checksummed. Needs examples/uartperf.zig flashed. + \\ --editor type at rising rates against pardes and report the highest rate it keeps + \\ up with. Needs the editor flashed. + \\ + ); +} + +fn nowUs(port: *serial.Port) i64 { + return std.Io.Timestamp.now(port.io, .boot).toMicroseconds(); +} + +const io = std.Io.Threaded.global_single_threaded.io(); +const stdout: std.Io.File = .{ .handle = 1, .flags = .{ .nonblocking = false } }; + +fn out(s: []const u8) void { + stdout.writeStreamingAll(io, s) catch {}; +} + +fn fatal(msg: []const u8) noreturn { + var b: [512]u8 = undefined; + out(std.fmt.bufPrint(&b, "p4-bench: {s}\n", .{msg}) catch "p4-bench: error\n"); + std.process.exit(1); +} + +fn eql(a: []const u8, b: []const u8) bool { + return std.mem.eql(u8, a, b); +} diff --git a/tools/perfproto.zig b/tools/perfproto.zig new file mode 100644 index 0000000..5223500 --- /dev/null +++ b/tools/perfproto.zig @@ -0,0 +1,201 @@ +//! A small framed protocol for measuring the board's serial link, shared verbatim by the host tool +//! and the firmware that answers it. +//! +//! WHY A PROTOCOL AND NOT A STOPWATCH. Timing an editor's keystrokes measures the editor, the +//! renderer and the link at once, and cannot tell a dropped byte from a slow one: RX overrun on this +//! UART is silent in hardware and uncounted in the driver, so a missing keystroke and a late one look +//! identical from the host. A frame with a length and a checksum turns both into facts. If the CRC +//! matches, every byte of that payload crossed intact; if a frame never completes, bytes were lost +//! and the count says how many. A throughput number that is not checksummed is a guess about how +//! fast data was corrupted. +//! +//! THE SHAPE. One fixed 9-byte header, little-endian, then the payload: +//! +//! "P4" op:u8 len:u16 crc:u32 payload[len] +//! +//! The CRC covers the payload only. The header carries it rather than trailing it so a receiver +//! knows, before it has read a single payload byte, exactly how many to expect and what they must +//! hash to - which is what lets the firmware verify a stream with one 4-byte accumulator and no +//! buffer at all. +//! +//! `max_payload` is 1024 and that is a memory decision, not a wire one. The firmware has a 384 KiB +//! heap it must share with an editor, and a bulk test that needed a 64 KiB frame buffer would be +//! measuring a configuration nobody ships. Bulk transfers are therefore many frames, which is also +//! the honest shape: it is the per-frame overhead a real protocol would pay. +//! +//! Both directions use the same header, and a reply's op has the high bit set, so a stray reply can +//! never be mistaken for a request by a resynchronising receiver. + +const std = @import("std"); + +pub const magic = "P4"; +pub const header_len = 9; +pub const max_payload = 1024; + +pub const Op = enum(u8) { + /// Echo the payload back as `pong`. Both directions verified in one exchange, which is what + /// makes it the right stimulus for a latency measurement. + ping = 1, + /// Payload is data to be consumed. The board accumulates a running count and CRC and answers + /// nothing, so the host can keep the uplink full and measure it without return traffic + /// competing for the same wire. + sink = 2, + /// Ask for the accumulated `sink` count and CRC, then reset them. + report = 3, + /// Payload is a u32 count: send exactly that many pattern bytes back, in `data` frames, + /// followed by a `stat`. + source = 4, + + pong = 0x81, + /// Payload is `Stat`, packed little-endian. + stat = 0x83, + /// A chunk of `source` output. + data = 0x84, + + pub fn isReply(o: Op) bool { + return @intFromEnum(o) & 0x80 != 0; + } +}; + +/// What the board reports about a stream it received or sent. Encoded by hand rather than by +/// `@bitCast` of a packed struct: this crosses between a riscv32 firmware and an x86_64 host, and a +/// layout that depends on either compiler's padding rules is a bug waiting for a target change. +pub const Stat = struct { + /// Payload bytes accumulated. + bytes: u32, + /// CRC-32 over exactly those bytes, in order. + crc: u32, + /// Frames whose CRC did not match. Nonzero means the link corrupted data rather than losing it, + /// which is a different fault with a different fix. + bad_frames: u32, + /// Bytes the firmware's UART driver gave up on writing. Its own counter, surfaced here because + /// the host cannot see it any other way. + tx_dropped: u32, + + pub const encoded_len = 16; + + pub fn encode(s: Stat, out: *[encoded_len]u8) void { + std.mem.writeInt(u32, out[0..4], s.bytes, .little); + std.mem.writeInt(u32, out[4..8], s.crc, .little); + std.mem.writeInt(u32, out[8..12], s.bad_frames, .little); + std.mem.writeInt(u32, out[12..16], s.tx_dropped, .little); + } + + pub fn decode(in: []const u8) ?Stat { + if (in.len < encoded_len) return null; + return .{ + .bytes = std.mem.readInt(u32, in[0..4], .little), + .crc = std.mem.readInt(u32, in[4..8], .little), + .bad_frames = std.mem.readInt(u32, in[8..12], .little), + .tx_dropped = std.mem.readInt(u32, in[12..16], .little), + }; + } +}; + +pub fn crc(bytes: []const u8) u32 { + return std.hash.Crc32.hash(bytes); +} + +/// The deterministic byte at stream offset `i`. +/// +/// A counter would be checksummed correctly by an implementation that lost exactly 256 bytes, and a +/// constant by one that lost any amount. This is an 8-bit xorshift-ish walk whose period is long +/// enough that no realistic loss aligns with it, so the CRC catches a gap wherever it falls. +pub fn patternByte(i: u32) u8 { + var x: u32 = i +% 1; + x ^= x << 7; + x ^= x >> 3; + x ^= x << 5; + return @truncate(x); +} + +pub fn fillPattern(buf: []u8, offset: u32) void { + for (buf, 0..) |*b, k| b.* = patternByte(offset +% @as(u32, @intCast(k))); +} + +/// Write a frame into `out`, returning the used slice. `out` must hold `header_len + payload.len`. +pub fn encode(out: []u8, op: Op, payload: []const u8) []u8 { + std.debug.assert(payload.len <= max_payload); + std.debug.assert(out.len >= header_len + payload.len); + out[0] = magic[0]; + out[1] = magic[1]; + out[2] = @intFromEnum(op); + std.mem.writeInt(u16, out[3..5], @intCast(payload.len), .little); + std.mem.writeInt(u32, out[5..9], crc(payload), .little); + @memcpy(out[header_len..][0..payload.len], payload); + return out[0 .. header_len + payload.len]; +} + +pub const Header = struct { + op: Op, + len: u16, + crc: u32, +}; + +/// Read a header out of `buf`. Returns null when fewer than `header_len` bytes are present, and +/// `error.BadFrame` when the magic or the op is not one of ours - which is how a receiver that has +/// lost sync tells "wait for more" from "throw a byte away and try again". +pub fn parseHeader(buf: []const u8) error{BadFrame}!?Header { + if (buf.len < header_len) return null; + if (buf[0] != magic[0] or buf[1] != magic[1]) return error.BadFrame; + const op = std.enums.fromInt(Op, buf[2]) orelse return error.BadFrame; + const len = std.mem.readInt(u16, buf[3..5], .little); + if (len > max_payload) return error.BadFrame; + return .{ .op = op, .len = len, .crc = std.mem.readInt(u32, buf[5..9], .little) }; +} + +test "a frame round-trips through encode and parseHeader" { + var buf: [header_len + 4]u8 = undefined; + const f = encode(&buf, .ping, "abcd"); + try std.testing.expectEqual(@as(usize, header_len + 4), f.len); + const h = (try parseHeader(f)).?; + try std.testing.expectEqual(Op.ping, h.op); + try std.testing.expectEqual(@as(u16, 4), h.len); + try std.testing.expectEqual(crc("abcd"), h.crc); + try std.testing.expectEqualStrings("abcd", f[header_len..]); +} + +test "a short buffer is incomplete, not invalid" { + var buf: [header_len]u8 = undefined; + const f = encode(&buf, .report, ""); + try std.testing.expectEqual(@as(?Header, null), try parseHeader(f[0 .. header_len - 1])); +} + +test "wrong magic and unknown ops are rejected rather than misread" { + var buf: [header_len]u8 = undefined; + var f = encode(&buf, .report, ""); + f[0] = 'X'; + try std.testing.expectError(error.BadFrame, parseHeader(f)); + f[0] = magic[0]; + f[2] = 0x7f; + try std.testing.expectError(error.BadFrame, parseHeader(f)); +} + +test "a truncated stream is caught by the CRC" { + // Losing bytes is the failure this protocol exists to detect, so prove the checksum notices a + // gap that leaves the length plausible. + var full: [64]u8 = undefined; + fillPattern(&full, 0); + var gapped: [64]u8 = undefined; + fillPattern(gapped[0..32], 0); + fillPattern(gapped[32..], 33); // one byte skipped mid-stream + try std.testing.expect(crc(&full) != crc(&gapped)); +} + +test "the pattern does not repeat inside a byte-aligned loss" { + // A plain counter would hash identically after losing exactly 256 bytes. This must not. + var a: [128]u8 = undefined; + var b: [128]u8 = undefined; + fillPattern(&a, 0); + fillPattern(&b, 256); + try std.testing.expect(crc(&a) != crc(&b)); +} + +test "Stat survives the trip between a riscv32 firmware and an x86_64 host" { + const s: Stat = .{ .bytes = 0x11223344, .crc = 0xdeadbeef, .bad_frames = 7, .tx_dropped = 9 }; + var buf: [Stat.encoded_len]u8 = undefined; + s.encode(&buf); + const back = Stat.decode(&buf).?; + try std.testing.expectEqual(s, back); + try std.testing.expectEqual(@as(?Stat, null), Stat.decode(buf[0 .. Stat.encoded_len - 1])); +} diff --git a/tools/rtt.zig b/tools/rtt.zig new file mode 100644 index 0000000..6ab0d87 --- /dev/null +++ b/tools/rtt.zig @@ -0,0 +1,170 @@ +//! Round-trip time over the board's only I/O channel: stimulus out, first byte back. +//! +//! Two functions, because every proposed fix for "too slow to type in" is a trade whose sign cannot +//! be guessed - a frame-rate cap, draining RX while blocked on TX, coalescing input, raising the +//! baud - and the only honest way to rank them is to measure the same number before and after. +//! +//! WHAT IS BEING TIMED, precisely: the interval from the last byte of a stimulus leaving the host to +//! the FIRST byte of the board's response arriving. That is the latency a human perceives as +//! responsiveness, and it is deliberately not the same as the time to finish repainting: a renderer +//! that starts drawing in 8 ms and takes 130 ms to finish feels immediate, while one that thinks for +//! 130 ms and then paints in 8 ms feels broken, and the two are indistinguishable if you only +//! measure when the wire goes quiet. `settle_us` records the second number so the pair can be read +//! together. +//! +//! Microseconds, not milliseconds: at 115200 baud one byte occupies 87 us, so a millisecond clock +//! quantises this measurement into buckets 11 bytes wide. +//! +//! The caller owns the board's STATE. These functions send bytes and time bytes; they do not know +//! what the editor does with them. A stimulus only produces a response if the editor is in a mode +//! where that keystroke changes the screen - pardes is modal, so a caller measuring keystrokes must +//! put it in insert mode first and must pick a stimulus that is not itself a mode change. + +const std = @import("std"); +const serial = @import("serial.zig"); + +pub const Sample = struct { + /// Stimulus out -> first response byte in. + rtt_us: i64, + /// Stimulus out -> last response byte in, i.e. the wire is free again. + settle_us: i64, + /// How much the board emitted in answer. At 115200 this is also a time: bytes * 87 us. + bytes: usize, +}; + +/// One round trip. Returns null when nothing came back within `timeout_us` - which is a result, not +/// an error: a dropped keystroke looks exactly like this, and it is the thing most worth counting. +/// +/// `quiet_us` decides when the response is over. It must exceed the largest gap the board leaves +/// mid-response; a renderer that pauses to allocate can stall longer than one byte time, and too +/// small a value would split one response into two and report a `settle_us` that is too good. +pub fn roundTrip( + port: *serial.Port, + stimulus: []const u8, + timeout_us: i64, + quiet_us: i64, +) !?Sample { + // Anything still in flight belongs to the previous measurement. Without this the first read + // below returns instantly with stale bytes and reports an RTT near zero. + var drain: [1024]u8 = undefined; + while (try port.readTimeout(&drain, 0) > 0) {} + + try port.write(stimulus); + const t0 = nowUs(port); + + var first: i64 = -1; + var last: i64 = t0; + var bytes: usize = 0; + var buf: [4096]u8 = undefined; + while (true) { + const now = nowUs(port); + if (first < 0) { + if (now - t0 > timeout_us) return null; + } else if (now - last > quiet_us) break; + + // Poll in millisecond units because that is what poll(2) takes; the TIMING above is + // microseconds and independent of this granularity. + const n = try port.readTimeout(&buf, 1); + if (n == 0) continue; + if (first < 0) first = nowUs(port); + bytes += n; + last = nowUs(port); + } + return .{ .rtt_us = first - t0, .settle_us = last - t0, .bytes = bytes }; +} + +pub const Stats = struct { + sent: u32, + /// Stimuli that produced no response at all inside the timeout. On this port that is a dropped + /// keystroke, and it is silent everywhere else in the system. + lost: u32, + min_us: i64, + median_us: i64, + max_us: i64, + /// Median, not mean: one 130 ms full repaint among fifty 9 ms updates should not move the + /// number that describes what typing feels like. + median_settle_us: i64, + median_bytes: usize, + /// Every response byte over the whole run, against the wire's capacity for that wall time. + /// 100% means the link is the limit and no amount of firmware tuning will help. + wire_percent: u32, + + pub fn format(s: Stats, w: *std.Io.Writer) std.Io.Writer.Error!void { + try w.print("{d} samples, {d} lost\n", .{ s.sent, s.lost }); + try w.print(" rtt min {d:>6} us median {d:>6} us max {d:>6} us\n", .{ + s.min_us, s.median_us, s.max_us, + }); + try w.print(" settle median {d} us ({d} B)\n", .{ s.median_settle_us, s.median_bytes }); + try w.print(" wire {d}% of capacity\n", .{s.wire_percent}); + } +}; + +pub const Options = struct { + samples: u32 = 20, + /// One byte that edits text without changing mode. `x` inserts an `x` in insert mode. + stimulus: []const u8 = "x", + /// Gap between stimuli. 80 ms is 12.5 characters a second: brisk human typing, and long enough + /// that a healthy editor finishes one update before the next arrives, so each sample is + /// independent rather than measuring a queue. + gap_us: i64 = 80_000, + timeout_us: i64 = 2_000_000, + quiet_us: i64 = 40_000, +}; + +/// `samples` round trips, summarised. The port must already be open and the board already in a +/// state where `stimulus` changes the screen. +pub fn measure(port: *serial.Port, opts: Options) !Stats { + const cap = 256; + var rtt: [cap]i64 = undefined; + var settle: [cap]i64 = undefined; + var size: [cap]usize = undefined; + var got: u32 = 0; + var lost: u32 = 0; + var total_bytes: usize = 0; + + const n = @min(opts.samples, cap); + const t_start = nowUs(port); + for (0..n) |_| { + if (try roundTrip(port, opts.stimulus, opts.timeout_us, opts.quiet_us)) |s| { + rtt[got] = s.rtt_us; + settle[got] = s.settle_us; + size[got] = s.bytes; + total_bytes += s.bytes; + got += 1; + } else lost += 1; + std.Io.sleep(port.io, .fromMicroseconds(opts.gap_us), .boot) catch {}; + } + const elapsed = @max(1, nowUs(port) - t_start); + + if (got == 0) return .{ + .sent = n, + .lost = lost, + .min_us = -1, + .median_us = -1, + .max_us = -1, + .median_settle_us = -1, + .median_bytes = 0, + .wire_percent = 0, + }; + + std.mem.sort(i64, rtt[0..got], {}, std.sort.asc(i64)); + std.mem.sort(i64, settle[0..got], {}, std.sort.asc(i64)); + std.mem.sort(usize, size[0..got], {}, std.sort.asc(usize)); + + // Bytes per second the wire can carry: baud/10, since each byte is 8N1 = 10 bit times. + const capacity = @as(i64, port.capacity()); + return .{ + .sent = n, + .lost = lost, + .min_us = rtt[0], + .median_us = rtt[got / 2], + .max_us = rtt[got - 1], + .median_settle_us = settle[got / 2], + .median_bytes = size[got / 2], + .wire_percent = @intCast(@divTrunc(@as(i64, @intCast(total_bytes)) * 1_000_000 * 100, elapsed * capacity)), + }; +} + +fn nowUs(port: *serial.Port) i64 { + return std.Io.Timestamp.now(port.io, .boot).toMicroseconds(); +} diff --git a/tools/serial.zig b/tools/serial.zig index 3398a11..7a740cd 100644 --- a/tools/serial.zig +++ b/tools/serial.zig @@ -81,10 +81,17 @@ pub const Port = struct { file: std.Io.File, io: std.Io, saved: Termios2, + /// The rate currently programmed, kept because the wire's capacity in bytes per second is + /// `rate/10` and anything measuring this link against its ceiling needs that number. The + /// kernel would answer a TCGETS2, but a syscall per sample to re-read a value only this file + /// ever changes is worse than a field. + baud: Baud, // std exposes neither these ioctl numbers nor the TIOCM bits. const TIOCEXCL = 0x540C; const TCFLSH = 0x540B; + /// tcdrain, with a nonzero argument. Zero would transmit a break instead. + const TCSBRK = 0x5409; const TCIFLUSH = 0; const TIOCMGET = 0x5415; const TIOCMSET = 0x5418; @@ -132,7 +139,7 @@ pub const Port = struct { // Drop anything the kernel captured at the previous line rate. _ = linux.ioctl(fd, TCFLSH, TCIFLUSH); - return .{ .file = file, .io = io, .saved = saved }; + return .{ .file = file, .io = io, .saved = saved, .baud = baud }; } /// Re-rate an already-open port, leaving the raw-mode flags and the exclusive claim alone. @@ -153,6 +160,25 @@ pub const Port = struct { t.ospeed = baud.rate(); if (@as(isize, @bitCast(linux.ioctl(p.file.handle, Termios2.TCSETSW2, @intFromPtr(&t)))) < 0) return error.SetAttrFailed; + p.baud = baud; + } + + /// Bytes per second the wire can carry: one 8N1 byte occupies ten bit times. + pub fn capacity(p: *const Port) u32 { + return p.baud.rate() / 10; + } + + /// Block until every byte written has physically left the wire - `tcdrain`, spelled as the + /// ioctl because std exposes neither. + /// + /// Distinct from `drain` above in both direction and meaning, which is worth stating because + /// getting them the wrong way round silently invalidates a measurement: `drain` discards what + /// has ARRIVED, this waits for what is LEAVING. A `write` returns once the kernel has accepted + /// the bytes, so timing a transfer to the write measures a memcpy into a tty buffer - at 115200 + /// that reported 202% of the wire's capacity, which is how the confusion was noticed. + pub fn flushOutput(p: *Port) void { + // TCSBRK with a nonzero argument is tcdrain on Linux; with zero it would send a break. + _ = linux.ioctl(p.file.handle, TCSBRK, 1); } pub fn close(p: *Port) void { |
