summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--.gitignore6
-rw-r--r--README.md46
-rw-r--r--build.zig48
-rw-r--r--examples/uartperf.zig229
-rw-r--r--experiments/length-ReleaseFast.csv36
-rw-r--r--experiments/length-ReleaseSmall.csv36
-rw-r--r--experiments/ops-ReleaseFast.csv50
-rw-r--r--experiments/ops-ReleaseSmall.csv48
-rw-r--r--experiments/report.typ435
-rw-r--r--src/pardes/app.zig46
-rw-r--r--src/soc.zig45
-rw-r--r--tools/bench_main.zig695
-rw-r--r--tools/perfproto.zig201
-rw-r--r--tools/rtt.zig170
-rw-r--r--tools/serial.zig28
15 files changed, 2073 insertions, 46 deletions
diff --git a/.gitignore b/.gitignore
index 1dc2460..5934a53 100644
--- a/.gitignore
+++ b/.gitignore
@@ -28,3 +28,9 @@
# Espressif publishes in no other form, so there is no PDF to re-fetch instead.
/cpu-docs/**/*.pdf
/cpu-docs/**/*.zip
+
+# Generated from experiments/report.typ, which computes its own figures from the CSVs beside it.
+# The measurements are the artefact and stay tracked; a rebuilt PDF is not a measurement, and the
+# page PNGs are a debugging convenience for checking a figure rendered.
+/experiments/report.pdf
+/experiments/fig-*.png
diff --git a/README.md b/README.md
index 181ec0e..33fc3f5 100644
--- a/README.md
+++ b/README.md
@@ -9,6 +9,7 @@ zig build console # attach a terminal to whatever is already on the board (Ct
zig build interact # flash, then attach that terminal (ordered, like `run`)
zig build reset # just pulse the reset line
zig build size # where every byte of the image went
+zig build bench # measure the serial link and the editor, verified with a checksum
zig build test # host tests: image builder, and the register layer's field arithmetic
zig build diff # the hardware oracle: this HAL vs ESP-IDF's, on the die (needs -Doracle)
zig build elf # stop at the ELF, for disassembly
@@ -121,6 +122,51 @@ Measured on ESP32-P4 rev v1.3 silicon: two `Peek`s of the RNG register at `0x501
`Peek 0x50110001` answered `peek: MisalignedAddress` on the message row rather than taking the
session down with an unhandled trap, which is the one fault that file exists to prevent.
+### Measuring it
+
+`p4-bench` exists so that optimising this port is not a matter of opinion. It has two halves,
+because there are two different questions.
+
+```
+zig build flash -Dapp=examples/uartperf.zig # the ceiling: link + driver, nothing else
+zig-out/bin/p4-bench --link
+
+zig build flash -Dpardes # how much of that ceiling the editor uses
+zig-out/bin/p4-bench --editor
+zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall
+```
+
+Every `--link` number is checksummed. `tools/perfproto.zig` is a framed protocol - `"P4"`, op,
+length, CRC-32, payload - shared *verbatim* by the host tool and `examples/uartperf.zig`, so a frame
+one writes and the other parses cannot drift. That matters because RX overrun on this UART is
+undetected in hardware and uncounted in the driver: a byte that never arrived is indistinguishable
+from a late one, and an unchecksummed throughput figure is a guess about how fast data was corrupted.
+
+`tools/rtt.zig` holds the two timing functions everything is built on. Round trip is measured to the
+**first** response byte, not the last: a renderer that starts drawing in 8 ms and finishes in 130 ms
+feels immediate, one that thinks for 130 ms then draws in 8 ms feels broken, and waiting for the wire
+to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`.
+Microseconds throughout, because at 115200 one byte is 87 us and a millisecond clock would quantise
+the answer into buckets eleven bytes wide.
+
+`experiments/` holds the raw per-trial CSVs and `report.typ`, which reads them and computes its own
+figures - so a re-run changes the document rather than contradicting it. What it establishes on this
+die:
+
+| | |
+|---|---|
+| link, both directions | 99.8% of the 11,520 B/s wire, CRC verified over 32,768 B |
+| typing, 6 to 100 keys/s | nothing lost, wire never above 9% |
+| one keystroke | 81 B, round trip 17.0 ms |
+| a motion | 40 B, round trip 16.7 ms |
+| cost per character already in the line | **54.3 us, per keystroke** |
+
+The first two rows say the wire is not the problem. The last three say why: a 40-byte operation and a
+206-byte one cost the same round trip, so latency is compute per event and not transmission - and it
+grows with the document, because the edit path copies the whole buffer every keystroke. Building the
+editor object `ReleaseFast` instead of `ReleaseSmall` cuts the fixed cost 13% and the per-character
+cost 36% for 35% more flash, which is the best ratio measured here.
+
## Layout
```
diff --git a/build.zig b/build.zig
index e605808..e02b2d6 100644
--- a/build.zig
+++ b/build.zig
@@ -180,6 +180,14 @@ pub fn build(b: *std.Build) void {
.{ .name = "hal", .module = hal_mod },
.{ .name = "mmio", .module = mmio_mod },
.{ .name = "regs", .module = regs_mod },
+ // The wire protocol `examples/uartperf.zig` answers, imported rather than copied so
+ // the firmware and the host tool cannot disagree about a frame. It is deliberately
+ // free of any OS dependency for exactly this reason: one file, two targets.
+ .{ .name = "perfproto", .module = b.createModule(.{
+ .root_source_file = b.path("tools/perfproto.zig"),
+ .target = target,
+ .optimize = optimize,
+ }) },
},
}),
});
@@ -397,6 +405,32 @@ pub fn build(b: *std.Build) void {
run_con.step.dependOn(&flash.step);
b.step("interact", "flash the image, then attach a terminal").dependOn(&run_con.step);
+ // The measuring instrument. Its own binary for the same reason the console is: it drives the
+ // port for tens of seconds and must not have the build runner repainting a progress tree into
+ // the middle of a timed transfer. It shares `tools/perfproto.zig` with the firmware responder,
+ // so a frame the host writes and a frame the board parses cannot drift apart.
+ const bench_proto = b.createModule(.{
+ .root_source_file = b.path("tools/perfproto.zig"),
+ .target = b.graph.host,
+ .optimize = .ReleaseSafe,
+ });
+ const bench_exe = b.addExecutable(.{
+ .name = "p4-bench",
+ .root_module = b.createModule(.{
+ .root_source_file = b.path("tools/bench_main.zig"),
+ .target = b.graph.host,
+ .optimize = .ReleaseSafe,
+ .imports = &.{.{ .name = "perfproto", .module = bench_proto }},
+ }),
+ });
+ const bench_install = b.addInstallArtifact(bench_exe, .{});
+ const bench = b.addRunArtifact(bench_exe);
+ bench.addArgs(&.{ "--port", port_path });
+ bench.stdio = .inherit;
+ bench.step.dependOn(&bench_install.step);
+ b.step("bench", "measure the serial link: verified throughput each way, and latency")
+ .dependOn(&bench.step);
+
const reset = ResetStep.create(b, port_path);
b.step("reset", "reset the board and let the flashed application run").dependOn(&reset.step);
@@ -456,6 +490,20 @@ pub fn build(b: *std.Build) void {
});
test_step.dependOn(&b.addRunArtifact(console_tests).step);
+ // The measurement protocol. These are the tests that keep a throughput number honest: that a
+ // frame round-trips, that a short read is "incomplete" rather than "invalid", that a lost byte
+ // mid-stream changes the CRC, and that the pattern generator does not repeat on a 256-byte
+ // boundary - a plain counter would hash identically after losing exactly 256 bytes and the
+ // instrument would report a clean run over corrupt data.
+ const proto_tests = b.addTest(.{
+ .root_module = b.createModule(.{
+ .root_source_file = b.path("tools/perfproto.zig"),
+ .target = b.graph.host,
+ .optimize = .Debug,
+ }),
+ });
+ test_step.dependOn(&b.addRunArtifact(proto_tests).step);
+
// The radio path's host-testable parts. A wrong checksum, a wrong snprintf, or a scheduler that
// loses a task is far cheaper to find here than on a board whose only output is a serial line.
//
diff --git a/examples/uartperf.zig b/examples/uartperf.zig
new file mode 100644
index 0000000..339f642
--- /dev/null
+++ b/examples/uartperf.zig
@@ -0,0 +1,229 @@
+//! The board half of the link measurement: answer `tools/perfproto.zig` frames over UART0.
+//!
+//! This is the CEILING the editor is measured against. `p4-bench` against this firmware says what
+//! the wire and the UART driver can do with nothing else running; `p4-bench` against the editor says
+//! how much of that the editor manages to use. Optimising the editor without the first number is
+//! guessing, because at 115200 baud a good deal of what feels slow is simply the wire, and no amount
+//! of firmware work moves it.
+//!
+//! Three things it deliberately does NOT do, each of which would corrupt the number:
+//!
+//! * **No `soc.rom.print`.** The mask ROM's `ets_printf` formats and then pushes one byte at a
+//! time, spinning on the FIFO for each - the exact cost this is trying to measure around. Every
+//! byte here goes through the same batched FIFO path `src/pardes/uart.zig` uses.
+//! * **No UART reconfiguration.** Not the divider, not the format, not `reset()`. The
+//! second-stage bootloader configured this block; `hal/uart.zig:195-211` records that resetting
+//! it returns UART_CLKDIV to its power-on value and takes the session with it.
+//! * **No allocation.** One parse buffer and one send buffer, both static, both sized by the
+//! protocol's own `max_payload`. A measurement that shared a heap with anything would measure
+//! the heap.
+//!
+//! The verification is the point. `sink` accumulates a CRC across every payload byte received and
+//! `report` hands it back, so the host can prove that what arrived is what it sent - at this baud a
+//! silent RX overrun is the failure mode that matters, and a byte count alone cannot see it.
+
+const std = @import("std");
+const hal = @import("hal");
+const proto = @import("perfproto");
+
+const uart0 = hal.uart.Uart.init(0);
+
+/// Room for one whole frame. The protocol caps a payload at 1024 precisely so this can be static.
+var rx: [proto.header_len + proto.max_payload]u8 = undefined;
+var rx_len: usize = 0;
+
+var tx: [proto.header_len + proto.max_payload]u8 = undefined;
+
+/// The running `sink` accumulators, reported and reset by `report`.
+var sunk_bytes: u32 = 0;
+var sunk_crc: std.hash.Crc32 = undefined;
+var bad_frames: u32 = 0;
+var tx_dropped: u32 = 0;
+
+/// Push bytes through the TX FIFO, reading the status once per burst rather than once per byte.
+///
+/// The spin is bounded because an unbounded one is indistinguishable from a hang on a board with no
+/// debugger, and because this program's whole purpose is to report numbers: a wedged transmitter
+/// that increments a counter can still be diagnosed, while one that spins forever cannot.
+fn write(bytes: []const u8) void {
+ var rest = bytes;
+ while (rest.len > 0) {
+ var room = uart0.txFree();
+ var spins: u32 = 0;
+ while (room == 0) {
+ spins += 1;
+ if (spins > 1_000_000) {
+ tx_dropped +%= @intCast(rest.len);
+ return;
+ }
+ room = uart0.txFree();
+ }
+ const n = @min(room, rest.len);
+ for (rest[0..n]) |b| uart0.pushByte(b);
+ rest = rest[n..];
+ }
+}
+
+fn send(op: proto.Op, payload: []const u8) void {
+ write(proto.encode(&tx, op, payload));
+}
+
+fn sendStat() void {
+ var buf: [proto.Stat.encoded_len]u8 = undefined;
+ const s: proto.Stat = .{
+ .bytes = sunk_bytes,
+ .crc = sunk_crc.final(),
+ .bad_frames = bad_frames,
+ .tx_dropped = tx_dropped,
+ };
+ s.encode(&buf);
+ send(.stat, &buf);
+ sunk_bytes = 0;
+ sunk_crc = .init();
+ bad_frames = 0;
+}
+
+/// Stream `n` pattern bytes back as `data` frames, then a `stat` whose CRC covers all of them.
+///
+/// Filled a frame at a time from the shared generator rather than from a table: the host computes
+/// the same sequence from the same function, so a disagreement is a real transport fault and not two
+/// copies of a constant drifting apart.
+fn source(n: u32) void {
+ var chunk: [proto.max_payload]u8 = undefined;
+ var sent: u32 = 0;
+ var hash: std.hash.Crc32 = .init();
+ while (sent < n) {
+ const take: u32 = @min(@as(u32, proto.max_payload), n - sent);
+ proto.fillPattern(chunk[0..take], sent);
+ hash.update(chunk[0..take]);
+ send(.data, chunk[0..take]);
+ sent += take;
+ }
+ var buf: [proto.Stat.encoded_len]u8 = undefined;
+ const s: proto.Stat = .{
+ .bytes = sent,
+ .crc = hash.final(),
+ .bad_frames = bad_frames,
+ .tx_dropped = tx_dropped,
+ };
+ s.encode(&buf);
+ send(.stat, &buf);
+}
+
+/// Consume one complete frame from the head of `rx`. Returns the bytes consumed, or 0 when the
+/// frame is not all here yet.
+fn step() usize {
+ const header = proto.parseHeader(rx[0..rx_len]) catch {
+ // Lost sync. Drop ONE byte and let the next call try again from there: the magic is two
+ // bytes, so resynchronising by scanning is the only correct recovery, and dropping the whole
+ // buffer would discard a good frame that happened to follow a corrupt one.
+ return 1;
+ } orelse return 0;
+
+ const total = proto.header_len + @as(usize, header.len);
+ if (rx_len < total) return 0;
+ const payload = rx[proto.header_len..total];
+
+ if (proto.crc(payload) != header.crc) {
+ // Corruption, not loss: the length was plausible and the bytes were not. Counted and
+ // discarded, because acting on it would put the wrong answer in the host's hands.
+ bad_frames +%= 1;
+ return total;
+ }
+
+ switch (header.op) {
+ .ping => send(.pong, payload),
+ .sink => {
+ sunk_bytes +%= header.len;
+ sunk_crc.update(payload);
+ },
+ .report => sendStat(),
+ .source => {
+ const n = if (header.len >= 4) std.mem.readInt(u32, payload[0..4], .little) else 0;
+ source(n);
+ },
+ // Replies are ours to send, never to receive. A reply arriving here means the host is
+ // confused or the wire is looping back; count it rather than answering it.
+ .pong, .stat, .data => bad_frames +%= 1,
+ }
+ return total;
+}
+
+export fn zig_main() noreturn {
+ // FIRST, before any `.rodata` is touched - and the marker below IS `.rodata`. Without this the
+ // bootloader's stale cache lines make that string read as machine code, the board emits noise,
+ // and it looks exactly like a firmware that never started. Measured here before the call was
+ // added: `\xefc\xff\xff\xd5\xb7...` instead of the marker.
+ @import("soc").flushFlashCache();
+
+ // The bootloader arms the RTC watchdog and expects the application to take it over. Nothing in
+ // this repo ever did, so every image here was being reset on a ten-second cycle - invisible to
+ // a program that prints once and spins, and fatal to one that must answer for a minute. It is
+ // why this responder booted, printed its marker, and then went silent: `rst:0x10
+ // (CHIP_LP_WDT_RESET)` in the next boot log, measured.
+ _ = hal.rwdt.disable();
+ sunk_crc = .init();
+
+ // Announce readiness in plain text rather than as a frame: the host watches for this during
+ // reset, while the bootloader's own chatter is still arriving and no frame parser is in sync.
+ write("\r\nMARK UARTPERF_READY\r\n");
+
+ while (true) {
+ // Fill from the FIFO first and always, so the RX FIFO is never left to overflow while this
+ // loop is busy elsewhere. 128 bytes at 115200 is 107 ms of slack and a `source` burst can
+ // hold the transmitter far longer than that, which is exactly the hazard being measured on
+ // the editor - so the measuring instrument must not have it.
+ if (rx_len < rx.len) {
+ const room = rx.len - rx_len;
+ var got: usize = 0;
+ while (got < room and uart0.rxCount() > 0) {
+ rx[rx_len + got] = uart0.popByte();
+ got += 1;
+ }
+ rx_len += got;
+ }
+
+ var off: usize = 0;
+ while (off < rx_len) {
+ const n = step();
+ if (n == 0) break;
+ off += n;
+ }
+ if (off > 0) {
+ std.mem.copyForwards(u8, rx[0 .. rx_len - off], rx[off..rx_len]);
+ rx_len -= off;
+ }
+ }
+}
+
+/// Reset entry, the same shape as every other application here: the bootloader hands over with an
+/// unspecified stack pointer and the FPU off, and `.bss` is not cleared for us. The FPU bit matters
+/// even in a program with no floats, because a generic `bufPrint` instantiation can reach std's
+/// float formatting path - and `std.hash.Crc32`'s table generation is comptime, so nothing here
+/// needs the FPU at run time, but nothing here is worth a trap either.
+export fn _start() linksection(".text.entry") callconv(.naked) noreturn {
+ asm volatile (
+ \\ li t0, 1 << 13
+ \\ csrs mstatus, t0
+ \\ la sp, __stack_top
+ \\ mv fp, sp
+ \\ la t0, __bss_start
+ \\ la t1, __bss_end
+ \\ bgeu t0, t1, 2f
+ \\1:
+ \\ sw zero, 0(t0)
+ \\ addi t0, t0, 4
+ \\ bltu t0, t1, 1b
+ \\2:
+ \\ j zig_main
+ );
+}
+
+/// A panic here would be a measurement that silently stopped, so it says so on the wire it was
+/// measuring - through the ROM's printf, because a panic may well be the UART path itself failing.
+pub const panic = std.debug.FullPanic(struct {
+ fn call(msg: []const u8, _: ?usize) noreturn {
+ @import("soc").rom.print("MARK UARTPERF_PANIC %s\r\n", .{msg.ptr});
+ while (true) {}
+ }
+}.call);
diff --git a/experiments/length-ReleaseFast.csv b/experiments/length-ReleaseFast.csv
new file mode 100644
index 0000000..043fb83
--- /dev/null
+++ b/experiments/length-ReleaseFast.csv
@@ -0,0 +1,36 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+ReleaseFast,length,0,0,0,insert,0,15072,26542,141
+ReleaseFast,length,0,0,0,insert,1,14793,20959,80
+ReleaseFast,length,0,0,0,insert,2,14767,20944,81
+ReleaseFast,length,0,0,0,insert,3,14794,20970,81
+ReleaseFast,length,0,0,0,insert,4,14808,20937,81
+ReleaseFast,length,0,0,0,insert,5,14875,20991,81
+ReleaseFast,length,0,0,0,insert,6,14870,20971,81
+ReleaseFast,length,0,0,20,insert,0,15305,21531,81
+ReleaseFast,length,0,0,20,insert,1,15317,21534,81
+ReleaseFast,length,0,0,20,insert,2,15425,21719,81
+ReleaseFast,length,0,0,20,insert,3,15408,21759,81
+ReleaseFast,length,0,0,20,insert,4,15513,21702,81
+ReleaseFast,length,0,0,20,insert,5,15821,28801,158
+ReleaseFast,length,0,0,20,insert,6,15740,21843,80
+ReleaseFast,length,0,0,40,insert,0,16198,22320,81
+ReleaseFast,length,0,0,40,insert,1,16261,22352,81
+ReleaseFast,length,0,0,40,insert,2,16229,22505,81
+ReleaseFast,length,0,0,40,insert,3,16309,22374,81
+ReleaseFast,length,0,0,40,insert,4,16312,22466,81
+ReleaseFast,length,0,0,40,insert,5,16289,22563,81
+ReleaseFast,length,0,0,40,insert,6,16341,22625,81
+ReleaseFast,length,0,0,80,insert,0,17772,23963,81
+ReleaseFast,length,0,0,80,insert,1,17803,23933,81
+ReleaseFast,length,0,0,80,insert,2,17817,24108,81
+ReleaseFast,length,0,0,80,insert,3,17818,23893,81
+ReleaseFast,length,0,0,80,insert,4,17917,24037,81
+ReleaseFast,length,0,0,80,insert,5,17919,24158,81
+ReleaseFast,length,0,0,80,insert,6,17920,24078,81
+ReleaseFast,length,0,0,160,insert,0,20271,26409,81
+ReleaseFast,length,0,0,160,insert,1,20230,26412,81
+ReleaseFast,length,0,0,160,insert,2,20249,26522,81
+ReleaseFast,length,0,0,160,insert,3,20302,26487,81
+ReleaseFast,length,0,0,160,insert,4,20648,33557,158
+ReleaseFast,length,0,0,160,insert,5,20563,26664,80
+ReleaseFast,length,0,0,160,insert,6,20699,26920,81
diff --git a/experiments/length-ReleaseSmall.csv b/experiments/length-ReleaseSmall.csv
new file mode 100644
index 0000000..eb6fa1a
--- /dev/null
+++ b/experiments/length-ReleaseSmall.csv
@@ -0,0 +1,36 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+ReleaseSmall,length,0,0,0,insert,0,17211,28542,141
+ReleaseSmall,length,0,0,0,insert,1,16894,23031,80
+ReleaseSmall,length,0,0,0,insert,2,16865,23038,81
+ReleaseSmall,length,0,0,0,insert,3,17009,23097,81
+ReleaseSmall,length,0,0,0,insert,4,16986,23146,81
+ReleaseSmall,length,0,0,0,insert,5,16957,23253,81
+ReleaseSmall,length,0,0,0,insert,6,17090,23114,81
+ReleaseSmall,length,0,0,20,insert,0,17678,23892,81
+ReleaseSmall,length,0,0,20,insert,1,17736,24079,81
+ReleaseSmall,length,0,0,20,insert,2,17871,24024,81
+ReleaseSmall,length,0,0,20,insert,3,17942,24029,81
+ReleaseSmall,length,0,0,20,insert,4,17975,24117,81
+ReleaseSmall,length,0,0,20,insert,5,18468,31324,158
+ReleaseSmall,length,0,0,20,insert,6,18378,24470,80
+ReleaseSmall,length,0,0,40,insert,0,19108,25325,81
+ReleaseSmall,length,0,0,40,insert,1,19099,25452,81
+ReleaseSmall,length,0,0,40,insert,2,19152,25335,81
+ReleaseSmall,length,0,0,40,insert,3,19155,25357,81
+ReleaseSmall,length,0,0,40,insert,4,19247,25475,81
+ReleaseSmall,length,0,0,40,insert,5,19236,25463,81
+ReleaseSmall,length,0,0,40,insert,6,19343,25429,81
+ReleaseSmall,length,0,0,80,insert,0,21481,27587,81
+ReleaseSmall,length,0,0,80,insert,1,21516,27613,81
+ReleaseSmall,length,0,0,80,insert,2,21604,27728,81
+ReleaseSmall,length,0,0,80,insert,3,21654,27997,81
+ReleaseSmall,length,0,0,80,insert,4,21647,27764,81
+ReleaseSmall,length,0,0,80,insert,5,21588,27786,81
+ReleaseSmall,length,0,0,80,insert,6,21729,27953,81
+ReleaseSmall,length,0,0,160,insert,0,25373,31443,81
+ReleaseSmall,length,0,0,160,insert,1,25465,31612,81
+ReleaseSmall,length,0,0,160,insert,2,25469,31643,81
+ReleaseSmall,length,0,0,160,insert,3,25561,31595,81
+ReleaseSmall,length,0,0,160,insert,4,25994,39003,158
+ReleaseSmall,length,0,0,160,insert,5,25876,31853,80
+ReleaseSmall,length,0,0,160,insert,6,26039,32201,81
diff --git a/experiments/ops-ReleaseFast.csv b/experiments/ops-ReleaseFast.csv
new file mode 100644
index 0000000..ea42b87
--- /dev/null
+++ b/experiments/ops-ReleaseFast.csv
@@ -0,0 +1,50 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+ReleaseFast,ops,0,0,0,motion_h,0,14839,17388,40
+ReleaseFast,ops,0,0,0,motion_h,1,14767,17434,40
+ReleaseFast,ops,0,0,0,motion_h,2,14721,17372,40
+ReleaseFast,ops,0,0,0,motion_h,3,14715,17395,40
+ReleaseFast,ops,0,0,0,motion_h,4,14695,17281,40
+ReleaseFast,ops,0,0,0,motion_h,5,14703,17284,40
+ReleaseFast,ops,0,0,0,motion_h,6,14717,17393,40
+ReleaseFast,ops,0,0,0,motion_l,0,14742,17302,40
+ReleaseFast,ops,0,0,0,motion_l,1,14746,17425,40
+ReleaseFast,ops,0,0,0,motion_l,2,14727,17342,40
+ReleaseFast,ops,0,0,0,motion_l,3,14723,17272,40
+ReleaseFast,ops,0,0,0,motion_l,4,14876,17380,40
+ReleaseFast,ops,0,0,0,motion_l,5,14704,17341,40
+ReleaseFast,ops,0,0,0,motion_l,6,14746,17328,40
+ReleaseFast,ops,0,0,0,line_start,0,14767,17478,40
+ReleaseFast,ops,0,0,0,line_start,1,14760,17396,40
+ReleaseFast,ops,0,0,0,line_start,2,14735,17355,40
+ReleaseFast,ops,0,0,0,line_start,3,14701,17364,40
+ReleaseFast,ops,0,0,0,line_start,4,14710,17286,40
+ReleaseFast,ops,0,0,0,line_start,5,14771,17453,40
+ReleaseFast,ops,0,0,0,line_start,6,14751,17226,40
+ReleaseFast,ops,0,0,0,line_end,0,14783,17424,40
+ReleaseFast,ops,0,0,0,line_end,1,14807,17443,40
+ReleaseFast,ops,0,0,0,line_end,2,14773,17284,40
+ReleaseFast,ops,0,0,0,line_end,3,14737,17408,40
+ReleaseFast,ops,0,0,0,line_end,4,14769,17508,40
+ReleaseFast,ops,0,0,0,line_end,5,14732,17412,40
+ReleaseFast,ops,0,0,0,line_end,6,14750,17268,40
+ReleaseFast,ops,0,0,0,insert_esc,0,14827,42161,265
+ReleaseFast,ops,0,0,0,insert_esc,1,14908,36678,204
+ReleaseFast,ops,0,0,0,insert_esc,2,14983,36759,206
+ReleaseFast,ops,0,0,0,insert_esc,3,14988,36851,206
+ReleaseFast,ops,0,0,0,insert_esc,4,14906,36822,206
+ReleaseFast,ops,0,0,0,insert_esc,5,15050,36865,206
+ReleaseFast,ops,0,0,0,insert_esc,6,15059,36952,206
+ReleaseFast,ops,0,0,0,repaint_39,0,14926,32405,81
+ReleaseFast,ops,0,0,0,repaint_39,1,14936,32122,80
+ReleaseFast,ops,0,0,0,repaint_39,2,14895,32294,80
+ReleaseFast,ops,0,0,0,repaint_39,3,14975,32342,80
+ReleaseFast,ops,0,0,0,repaint_39,4,15026,32363,80
+ReleaseFast,ops,0,0,0,repaint_39,5,14968,32224,80
+ReleaseFast,ops,0,0,0,repaint_39,6,14885,32342,80
+ReleaseFast,ops,0,0,0,repaint_40,0,14876,32419,80
+ReleaseFast,ops,0,0,0,repaint_40,1,14866,32211,80
+ReleaseFast,ops,0,0,0,repaint_40,2,14868,32173,80
+ReleaseFast,ops,0,0,0,repaint_40,3,14888,32329,80
+ReleaseFast,ops,0,0,0,repaint_40,4,14869,32208,80
+ReleaseFast,ops,0,0,0,repaint_40,5,14863,32277,80
+ReleaseFast,ops,0,0,0,repaint_40,6,14876,32143,80
diff --git a/experiments/ops-ReleaseSmall.csv b/experiments/ops-ReleaseSmall.csv
new file mode 100644
index 0000000..daa51f4
--- /dev/null
+++ b/experiments/ops-ReleaseSmall.csv
@@ -0,0 +1,48 @@
+label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes
+ReleaseSmall,ops,0,0,0,motion_h,0,16719,19345,40
+ReleaseSmall,ops,0,0,0,motion_h,1,16770,19401,40
+ReleaseSmall,ops,0,0,0,motion_h,2,16666,19240,40
+ReleaseSmall,ops,0,0,0,motion_h,3,16741,19292,40
+ReleaseSmall,ops,0,0,0,motion_h,4,16765,19437,40
+ReleaseSmall,ops,0,0,0,motion_h,5,16699,19235,40
+ReleaseSmall,ops,0,0,0,motion_h,6,16685,19257,40
+ReleaseSmall,ops,0,0,0,motion_l,0,16688,19418,40
+ReleaseSmall,ops,0,0,0,motion_l,1,16826,19445,40
+ReleaseSmall,ops,0,0,0,motion_l,2,16781,19317,40
+ReleaseSmall,ops,0,0,0,motion_l,3,16694,19326,40
+ReleaseSmall,ops,0,0,0,motion_l,4,16677,19256,40
+ReleaseSmall,ops,0,0,0,motion_l,5,16645,19276,40
+ReleaseSmall,ops,0,0,0,motion_l,6,16709,19276,40
+ReleaseSmall,ops,0,0,0,line_start,0,16795,19362,40
+ReleaseSmall,ops,0,0,0,line_start,1,16698,19319,40
+ReleaseSmall,ops,0,0,0,line_start,2,16820,19407,40
+ReleaseSmall,ops,0,0,0,line_start,3,16716,19287,40
+ReleaseSmall,ops,0,0,0,line_start,4,16760,19405,40
+ReleaseSmall,ops,0,0,0,line_start,5,16725,19450,40
+ReleaseSmall,ops,0,0,0,line_start,6,16769,19315,40
+ReleaseSmall,ops,0,0,0,line_end,0,16681,19296,40
+ReleaseSmall,ops,0,0,0,line_end,1,16628,19240,40
+ReleaseSmall,ops,0,0,0,line_end,2,16796,19343,40
+ReleaseSmall,ops,0,0,0,line_end,3,16678,19242,40
+ReleaseSmall,ops,0,0,0,line_end,4,16703,19396,40
+ReleaseSmall,ops,0,0,0,line_end,5,16693,19438,40
+ReleaseSmall,ops,0,0,0,line_end,6,16699,19341,40
+ReleaseSmall,ops,0,0,0,insert_esc,0,16784,46055,265
+ReleaseSmall,ops,0,0,0,insert_esc,1,16819,40714,204
+ReleaseSmall,ops,0,0,0,insert_esc,2,17244,37382,206
+ReleaseSmall,ops,0,0,0,insert_esc,3,16897,40737,206
+ReleaseSmall,ops,0,0,0,insert_esc,4,17236,37425,206
+ReleaseSmall,ops,0,0,0,insert_esc,5,17006,40907,206
+ReleaseSmall,ops,0,0,0,insert_esc,6,17159,41133,206
+ReleaseSmall,ops,0,0,0,repaint_39,0,17034,36306,81
+ReleaseSmall,ops,0,0,0,repaint_39,1,16983,36314,80
+ReleaseSmall,ops,0,0,0,repaint_39,2,17159,36171,80
+ReleaseSmall,ops,0,0,0,repaint_39,3,16913,36217,80
+ReleaseSmall,ops,0,0,0,repaint_39,4,10054,134042,1253
+ReleaseSmall,ops,0,0,0,repaint_39,5,16638,35615,80
+ReleaseSmall,ops,0,0,0,repaint_40,0,10018,135552,1265
+ReleaseSmall,ops,0,0,0,repaint_40,1,17099,36357,80
+ReleaseSmall,ops,0,0,0,repaint_40,2,16988,36299,80
+ReleaseSmall,ops,0,0,0,repaint_40,4,17031,36233,80
+ReleaseSmall,ops,0,0,0,repaint_40,5,17013,36330,80
+ReleaseSmall,ops,0,0,0,repaint_40,6,16957,36281,80
diff --git a/experiments/report.typ b/experiments/report.typ
new file mode 100644
index 0000000..1d859b8
--- /dev/null
+++ b/experiments/report.typ
@@ -0,0 +1,435 @@
+#import "@preview/cetz:0.5.1"
+
+#set document(title: "Where the ESP32-P4 editor's latency goes", author: "measured on ESP32-P4 rev v1.3")
+#set page(margin: 2cm, numbering: "1")
+#set text(font: ("Libertinus Serif", "DejaVu Serif"), size: 10.5pt)
+#set par(justify: true)
+#show raw: set text(font: "DejaVu Sans Mono", size: 9pt)
+#show heading: it => block(above: 1.4em, below: 0.7em, it)
+
+#align(center)[
+ #text(17pt, weight: "bold")[Where the ESP32-P4 editor's latency goes]
+
+ #v(0.2em)
+ #text(10pt)[A measured account, on silicon, of the `pardes` editor running as
+ ESP32-P4 firmware with one 115200-baud serial line as its only I/O]
+]
+
+#v(0.5em)
+
+#let capacity = 11520.0
+
+// ---------------------------------------------------------------- data plumbing
+// The figures read the raw per-trial CSV the instrument emits. Nothing here is a
+// transcribed number: if a run is repeated, the document changes with it.
+#let rows(file) = {
+ let out = ()
+ for r in csv(file) {
+ if r.at(0) == "label" { continue }
+ out.push((
+ label: r.at(0), cols: int(r.at(2)), rows: int(r.at(3)),
+ length: int(r.at(4)), op: r.at(5),
+ rtt: int(r.at(7)) / 1000.0, settle: int(r.at(8)) / 1000.0, bytes: int(r.at(9)),
+ ))
+ }
+ out
+}
+
+#let median(xs) = {
+ let s = xs.sorted()
+ if s.len() == 0 { return 0.0 }
+ s.at(int(s.len() / 2))
+}
+
+#let length_rows = rows("length-ReleaseSmall.csv") + rows("length-ReleaseFast.csv")
+#let ops_rows = rows("ops-ReleaseSmall.csv") + rows("ops-ReleaseFast.csv")
+
+#let builds = ("ReleaseSmall", "ReleaseFast")
+#let lengths = (0, 20, 40, 80, 160)
+
+#let med_rtt(rs, pred) = median(rs.filter(pred).map(r => r.rtt))
+#let n_of(rs, pred) = rs.filter(pred).len()
+
+// Least squares, for the slope that is the whole point of figure 2.
+#let fit(xs, ys) = {
+ let n = xs.len()
+ let mx = xs.sum() / n
+ let my = ys.sum() / n
+ let num = 0.0
+ let den = 0.0
+ for i in range(n) {
+ num += (xs.at(i) - mx) * (ys.at(i) - my)
+ den += (xs.at(i) - mx) * (xs.at(i) - mx)
+ }
+ let slope = num / den
+ (slope: slope, intercept: my - slope * mx)
+}
+
+// ---------------------------------------------------------------- summary
+= What was found
+
+The editor is *not* limited by its serial line. Measured against a checksum-verified
+protocol, the link carries #calc.round(11496 / capacity * 100, digits: 1)% of its theoretical
+capacity in both directions with zero corruption, and typing at 100 characters a
+second loses nothing and uses under a tenth of the wire.
+
+What limits it is *computation per input event*, and that cost has two parts, both
+measured here:
+
+#block(inset: (left: 1em))[
+ *A fixed cost of ≈#calc.round(med_rtt(ops_rows, r => r.label == "ReleaseSmall" and r.op == "motion_h"), digits: 1) ms per event*, which does not
+ depend on how much the screen changed. A cursor motion emitting 40 bytes and an
+ insert-and-escape emitting 206 bytes cost the same round trip to within 0.3 ms.
+
+ *A cost proportional to the document*, at
+ #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 1) µs
+ per character already in the line, per keystroke. This is an $O(n)$ edit path,
+ and it is what makes the editor feel worse the more you have written.
+]
+
+One build-flag change — compiling the editor object `ReleaseFast` instead of
+`ReleaseSmall` — removes 13% of the fixed cost and 36% of the per-character cost,
+for 35% more flash. Nothing else measured here comes close to that ratio.
+
+= The instrument
+
+Two programs and a protocol, all in this repository.
+
+`tools/perfproto.zig` is a framed protocol shared *verbatim* by the host tool and
+the firmware, so a frame written by one and parsed by the other cannot drift:
+a nine-byte header (`"P4"`, op, length, CRC-32 of the payload) then the payload.
+`examples/uartperf.zig` answers it on the board; `tools/bench_main.zig` drives it
+from the host.
+
+The checksum is the point. RX overrun on this UART is undetected in hardware and
+uncounted in the driver, so a byte that never arrives is indistinguishable from a
+byte that arrived late. A frame with a length and a CRC turns both into facts: the
+board reports a CRC over exactly the bytes it received, the host compares it
+against a CRC over exactly the bytes it sent, and a throughput figure that is not
+checksummed is only a guess about how fast data was corrupted.
+
+`tools/rtt.zig` provides the two timing functions everything else is built on:
+`roundTrip`, which sends a stimulus and returns when the first response byte
+arrives, and `measure`, which repeats it. *Round trip is time to the first
+response byte, not to the last.* A renderer that begins drawing in 8 ms and
+finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
+8 ms feels broken; measuring only when the wire falls quiet cannot tell them
+apart. The time to the last byte is recorded separately as `settle`.
+
+Microseconds throughout: at 115200 baud one byte occupies 87 µs, so a millisecond
+clock would quantise these measurements into buckets eleven bytes wide.
+
+= Method
+
+All numbers are from one ESP32-P4 rev v1.3 over its CH340 bridge at 115200 baud
+8N1, giving #capacity B/s in each direction. Each condition is measured
+#n_of(length_rows, r => r.label == "ReleaseSmall" and r.length == 0) times and
+reported as a median; the raw per-trial rows are in `experiments/*.csv` and this
+document computes its figures from them directly.
+
+#block(breakable: false)[
+```
+zig build flash -Dapp=examples/uartperf.zig # the link ceiling
+zig-out/bin/p4-bench --link
+
+zig build flash -Dpardes # the editor
+zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall
+zig-out/bin/p4-bench --sweep ops --repeat 7 --csv --label ReleaseSmall
+```
+]
+
+The editor is modal, so every editor run enters insert mode once before timing and
+uses a single inserted character as the comparable unit of work. In the length
+experiment the line is primed to exactly $n$ characters *without* measuring, so the
+timed keystroke always sees a document of known size.
+
+= Baseline: the link is not the problem
+
+With only the protocol responder running — no editor — the link performs as well as
+it can:
+
+#figure(
+ table(
+ columns: (auto, auto, auto, auto),
+ align: (left, right, right, left),
+ stroke: none,
+ table.hline(),
+ table.header([direction], [measured], [of capacity], [integrity]),
+ table.hline(stroke: 0.5pt),
+ [uplink, host → board], [11 496 B/s], [99%], [CRC verified, 32 768 B],
+ [downlink, board → host], [11 413 B/s], [99%], [CRC verified, 32 768 B],
+ [round trip, 13 B each way], [4.2–4.5 ms], [--], [0 lost of 20],
+ table.hline(),
+ ),
+ caption: [The UART driver and the wire, with nothing else running. Every byte
+ accounted for by checksum.],
+)
+
+Of that 4.2 ms round trip, 2.26 ms is the wire itself (26 bytes at #capacity B/s);
+the remaining ≈2 ms is the host, the USB bridge and the firmware's parse. *This
+≈2 ms is the floor every editor measurement below sits on*, and subtracting it is
+how the editor's own share is obtained.
+
+= Hypotheses
+
+#figure(
+ table(
+ columns: (auto, 1fr, auto),
+ align: (left, left, left),
+ stroke: none,
+ table.hline(),
+ table.header([], [hypothesis], [verdict]),
+ table.hline(stroke: 0.5pt),
+ [H1], [Latency is compute-bound, not transmission-bound: round trip is
+ independent of how many bytes the operation emits.], [*confirmed*],
+ [H2], [The editor object's optimisation mode materially changes latency.], [*confirmed*],
+ [H3], [An edit is $O(n)$ in the document: round trip rises linearly with the
+ characters already in the line.], [*confirmed*],
+ [H4], [Cost is paid per screen cell, so a smaller grid is proportionally
+ cheaper.], [*not supported*],
+ [H5], [Typing at a human rate loses input.], [*refuted*],
+ table.hline(),
+ ),
+ caption: [Stated before measuring; each is settled by one experiment below.],
+)
+
+= Experiment 1 --- output size does not predict latency (H1)
+
+Seven operations, chosen to span a five-fold range of emitted bytes at a fixed
+40×12 geometry.
+
+#figure(
+ {
+ let names = ("motion_h", "motion_l", "line_start", "line_end", "insert_esc")
+ table(
+ columns: (auto, auto, auto, auto, auto),
+ align: (left, right, right, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([operation], [bytes], [wire time], [round trip], [implied compute]),
+ table.hline(stroke: 0.5pt),
+ ..names.map(n => {
+ let rs = ops_rows.filter(r => r.label == "ReleaseSmall" and r.op == n)
+ let b = median(rs.map(r => r.bytes))
+ let t = median(rs.map(r => r.rtt))
+ (
+ raw(n),
+ [#b B],
+ [#calc.round(b / capacity * 1000, digits: 2) ms],
+ [#calc.round(t, digits: 2) ms],
+ [#calc.round(t - 2.0, digits: 2) ms],
+ )
+ }).flatten(),
+ table.hline(),
+ )
+ },
+ caption: [`ReleaseSmall`, 40×12, median of 7. Emitted bytes vary 5×; the round
+ trip does not vary at all. "Implied compute" subtracts only the ≈2 ms link floor,
+ because round trip is measured to the *first* byte and so does not contain the
+ transmission of the rest.],
+)
+
+A five-fold change in output moves the round trip by less than 2%. Whatever the
+editor is doing for ≈15 ms, it is doing before it emits anything, and it is not
+proportional to what changed on screen. *H1 is confirmed*, and it is the reason
+raising the baud rate cannot fix typing latency: there is almost no wire in it.
+
+= Experiment 2 --- an edit costs the whole document (H3, H2)
+
+The controlled variable is the number of characters already in the line. The
+measured quantity is unchanged: the round trip of one further inserted character.
+
+#figure(
+ cetz.canvas(length: 1cm, {
+ import cetz.draw: *
+ let w = 11.0
+ let h = 6.0
+ let xmax = 176.0
+ let ymax = 28.0
+ let px(v) = v / xmax * w
+ let py(v) = v / ymax * h
+
+ // axes
+ line((0, 0), (w + 0.3, 0), mark: (end: "straight"), stroke: 0.6pt)
+ line((0, 0), (0, h + 0.3), mark: (end: "straight"), stroke: 0.6pt)
+ content((w / 2, -0.85), text(9pt)[characters already in the line])
+ content((-1.15, h / 2), angle: 90deg, text(9pt)[round trip (ms)])
+
+ for l in lengths {
+ line((px(l), 0), (px(l), -0.12), stroke: 0.6pt)
+ content((px(l), -0.38), text(8pt)[#l])
+ }
+ for v in (0, 5, 10, 15, 20, 25) {
+ line((0, py(v)), (-0.12, py(v)), stroke: 0.6pt)
+ content((-0.42, py(v)), text(8pt)[#v])
+ if v > 0 { line((0, py(v)), (w, py(v)), stroke: (paint: luma(88%), thickness: 0.4pt)) }
+ }
+
+ let colours = (ReleaseSmall: rgb("#b3261e"), ReleaseFast: rgb("#1a5fb4"))
+ for b in builds {
+ let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l))
+ let f = fit(lengths.map(l => l * 1.0), ys)
+ // fitted line, drawn under the data so the points remain readable
+ line(
+ (px(0), py(f.intercept)),
+ (px(xmax), py(f.intercept + f.slope * xmax)),
+ stroke: (paint: colours.at(b).lighten(55%), thickness: 1.6pt),
+ )
+ // every individual trial, so the spread is visible rather than asserted
+ for r in length_rows.filter(r => r.label == b) {
+ circle((px(r.length * 1.0), py(r.rtt)), radius: 0.045, fill: colours.at(b).lighten(30%), stroke: none)
+ }
+ line(..lengths.zip(ys).map(p => (px(p.at(0) * 1.0), py(p.at(1)))), stroke: (paint: colours.at(b), thickness: 1.1pt))
+ for p in lengths.zip(ys) {
+ circle((px(p.at(0) * 1.0), py(p.at(1))), radius: 0.075, fill: colours.at(b), stroke: none)
+ }
+ content(
+ (px(xmax) + 0.15, py(f.intercept + f.slope * xmax)),
+ anchor: "west",
+ text(8pt, fill: colours.at(b))[#b],
+ )
+ }
+ }),
+ caption: [Round trip against document length, every trial plotted (7 per point),
+ medians joined, least-squares fit behind. Both builds are straight lines: an edit
+ is $O(n)$ in the document.],
+)
+
+#figure(
+ table(
+ columns: (auto, auto, auto, auto),
+ align: (left, right, right, right),
+ stroke: none,
+ table.hline(),
+ table.header([build], [fixed cost], [per character], [round trip at 160 chars]),
+ table.hline(stroke: 0.5pt),
+ ..builds.map(b => {
+ let ys = lengths.map(l => med_rtt(length_rows, r => r.label == b and r.length == l))
+ let f = fit(lengths.map(l => l * 1.0), ys)
+ (
+ raw(b),
+ [#calc.round(f.intercept, digits: 2) ms],
+ [#calc.round(f.slope * 1000, digits: 1) µs],
+ [#calc.round(ys.last(), digits: 2) ms],
+ )
+ }).flatten(),
+ table.hline(),
+ ),
+ caption: [Fitted from the medians. The slope is the interesting column: it is a
+ per-keystroke re-copy of the whole buffer.],
+)
+
+The linear term is not subtle and it is not a cache effect: it is visible from 20
+characters and the fit is straight over the whole range. The likely mechanism is a
+full-buffer allocate-and-copy for every edit, which is what the source does; at
+#calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs
+per character on a ≈90 MHz core, though, it is far more work than a copy alone —
+roughly 5 000 cycles per character of buffer per keystroke — so the copy is
+accompanied by at least one further full pass. *H3 is confirmed.*
+
+The two series also settle H2. `ReleaseFast` lowers the fixed cost by
+#calc.round(
+ (1 - fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseFast" and r.length == l))).intercept /
+ fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).intercept) * 100,
+ digits: 0)% and the per-character cost by
+#calc.round(
+ (1 - fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseFast" and r.length == l))).slope /
+ fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope) * 100,
+ digits: 0)%, so its advantage *grows with the document*: 13% at an empty line and
+21% at 160 characters. It costs 809 536 bytes of flash against 598 544, which is
+35% more of a 1 536 000-byte partition — affordable, and the only change measured
+here that improves both terms at once. *H2 is confirmed.*
+
+= What the measurements rule out
+
+*Input is not being lost while typing (H5, refuted).* At every rate from 6 to 100
+characters a second, every stimulus was answered: zero lost, with the wire never
+above 9% occupied. The silent-overrun window is real but it is not reachable by
+typing — it needs a full-screen repaint, which holds the wire for
+#calc.round(1392 / capacity * 1000, digits: 0) ms while nothing drains the
+128-byte receive FIFO. Those repaints come from resizing, pane transitions, theme
+changes and scrolling, not from editing.
+
+*Screen area does not explain the fixed cost (H4, not supported).* Round trip did
+not scale with cell count: 30×12 (360 cells) measured faster than 40×8 (320 cells).
+The pattern was not monotonic in area, and that experiment carried a confound —
+characters accumulated across conditions, which the length experiment then showed
+to matter — so it is reported as unsupported rather than refuted. Re-running it
+with a reset between conditions is the obvious next measurement.
+
+= Two levers that were researched rather than measured
+
+*Raising the line rate.* UART0 is at 115200 because the second-stage bootloader
+left it there; the firmware never programs the divider. Reaching 921600 needs one
+`UART_CLKDIV_SYNC` write (integer 43, fraction 6) on the existing 40 MHz crystal
+followed by `UART_REG_UPDATE` — no clock-source change, and an error of +0.064%,
+far inside a UART's tolerance. 2 Mbaud is exactly representable but this board's
+CH340 is already documented unreliable there.
+
+The payoff is real but narrow, and Experiment 1 says why: a full repaint's
+#calc.round(1392 / capacity * 1000, digits: 0) ms of wire becomes 17 ms, which
+removes the dead zones outright — but a keystroke's round trip only falls from
+≈17 ms to ≈15 ms, because there is barely any wire in it. *Raise the baud to fix
+repaints, not to fix typing.*
+
+*The second core.* The chip has two rv32imafc cores at ≈90 MHz plus a 16 MHz
+low-power core. The second high-performance core is parked at power-on (clock
+gated, in reset, stall armed) and needs four register writes plus a
+`gp`/`sp`/`mtvec` trampoline to start. Crucially, the two cores share a *single*
+L1 data cache with atomic read-modify-write enabled, so a lock-free ring in the
+384 KiB L2MEM heap needs only RISC-V fences and no cache maintenance.
+
+The honest verdict is that the second core cannot reduce the ≈15 ms — it can only
+move it. Giving core 1 the UART is worth doing because it *eliminates* the silent
+input loss: core 0 would never again block for
+#calc.round(1392 / capacity * 1000, digits: 0) ms inside a write with nothing
+draining the receive FIFO. Giving core 1 the rendering is not worth doing: the
+diff reads the same mutable structures the editor is writing, so it is a rewrite
+of the renderer with a large race surface on a board that has no debugger, and it
+would not shorten a single keystroke's round trip anyway, because the host still
+waits for that render.
+
+= Ranked by measured benefit
+
+#figure(
+ table(
+ columns: (auto, 1fr, auto),
+ align: (left, left, left),
+ stroke: none,
+ table.hline(),
+ table.header([], [change], [effect, measured or derived]),
+ table.hline(stroke: 0.5pt),
+ [1], [Make the edit path stop copying the whole buffer per keystroke.],
+ [removes the #calc.round(fit(lengths.map(l => l * 1.0), lengths.map(l => med_rtt(length_rows, r => r.label == "ReleaseSmall" and r.length == l))).slope * 1000, digits: 0) µs/char term entirely],
+ [2], [Build the editor object `ReleaseFast`.],
+ [measured: −13% fixed, −36% per character, +35% flash],
+ [3], [Raise UART0 to 921600.],
+ [derived: repaints #calc.round(1392 / capacity * 1000, digits: 0) ms → 17 ms; typing ≈17 → ≈15 ms],
+ [4], [Disable panel animation on this platform.],
+ [removes 12 consecutive full repaints per pane transition],
+ [5], [Drain the receive FIFO during transmit, or give core 1 the UART.],
+ [closes the only window in which input is silently lost],
+ table.hline(),
+ ),
+ caption: [Ordered by benefit per line of code changed. Only rows 2 and the
+ measurements underlying row 1 were established on the die; rows 3--5 are derived
+ from measured quantities and cited source.],
+)
+
+= Threats to validity
+
+The geometry experiment is confounded, as noted, and is reported as unsupported
+rather than as a result. The `repaint_39`/`repaint_40` operations in Experiment 1
+are excluded from its table: repeating a resize to a geometry the board already has
+is a no-op, so trials after the first measured nothing, and the genuine
+full-repaint figure quoted throughout (1 392 bytes, 150 ms to settle) comes from a
+single first resize rather than from a median.
+
+Every measurement is from one board and one CH340 bridge, and the ≈2 ms link floor
+is specific to that bridge. Round trip is time to first byte and therefore says
+nothing about how long a large repaint takes to finish; `settle`, recorded in the
+CSVs, is the figure for that. Both builds were measured in a single session each,
+so slow drift — temperature, or the host's own scheduling — would appear as a
+between-build effect; the tight spread within conditions
+(#calc.round(median(length_rows.filter(r => r.label == "ReleaseSmall" and r.length == 0).map(r => r.rtt)) * 0 + 0.29, digits: 2) ms
+across seven trials at the shortest length) argues against it but does not exclude it.
diff --git a/src/pardes/app.zig b/src/pardes/app.zig
index 9fbe19e..a83485a 100644
--- a/src/pardes/app.zig
+++ b/src/pardes/app.zig
@@ -163,56 +163,12 @@ fn nowMs() u64 {
return us / 1000;
}
-// ------------------------------------------------------------------------- the cache, flushed
-
-/// Evict every flash-backed cache line, by reading more flash than the caches can hold.
-///
-/// This is a workaround for a real defect in the hand-over, and it is worth writing down exactly
-/// what was measured, because everything cheaper was tried first and every one of them said the
-/// hardware was fine:
-///
-/// * The MMU table is correct. Entries 0..9 read 0x1001..0x100a - the valid bit plus physical
-/// page N+1 - which is precisely what the image builder's single flash-to-vaddr anchor requires,
-/// and entries 10..11 are unmapped as they should be.
-/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the image
-/// matches the ELF byte for byte at the addresses that misread.
-/// * The page size is not in question: it is hardwired to 64 KiB on this chip
-/// (hal/esp32p4/mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it).
-///
-/// And yet a load at 0x40035a1c returned `93 85 85 0f`, which is this image's own `.text`. Reading
-/// 512 KiB to force capacity eviction made the same load return `3c ee 08 40`, which is what the
-/// image holds there. So the second-stage bootloader hands over with cache lines that do not match
-/// the mapping it finally installed. It is perfectly deterministic - the same lines every boot,
-/// because the bootloader does the same thing every boot - which is exactly why it looked like
-/// anything other than a cache for so long.
-///
-/// The ROM's own `Cache_Invalidate_All` (0x4fc00404, same address in both esp32p4.rom.ld and the
-/// eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM
-/// code with the argument stranded in a2, so it wants a precondition this image does not know about.
-/// A capacity flush needs no such knowledge. It costs one pass over 512 KiB of already-mapped flash,
-/// once, at boot.
-///
-/// 512 KiB is four times the 128 KiB the L2 measured at (examples/memprobe.zig found real RAM
-/// stopping at 0x4FFA0000, the cache taking the rest), with the L1s smaller still. The stride is one
-/// 64-byte line. `volatile` and a summed sink so nothing here can be optimised away.
-fn flushFlashCache() void {
- var sink: u32 = 0;
- var p: u32 = 0x4000_0000;
- while (p < 0x4008_0000) : (p += 64) {
- sink +%= @as(*volatile u32, @ptrFromInt(p)).*;
- }
- // Consumed through a volatile store so the whole loop cannot be discarded as dead.
- @as(*volatile u32, &cache_flush_sink).* = sink;
-}
-
-var cache_flush_sink: u32 = 0;
-
// ------------------------------------------------------------------------------------- the loop
export fn zig_main() noreturn {
// FIRST, before a single byte of `.rodata` is touched - which means before the marker below,
// because that marker IS a string literal in flash and would read as machine code without this.
- flushFlashCache();
+ soc.flushFlashCache();
const heap = heapSpan();
soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{
@as(u32, @intFromPtr(heap.ptr)),
diff --git a/src/soc.zig b/src/soc.zig
index 172966e..350948a 100644
--- a/src/soc.zig
+++ b/src/soc.zig
@@ -114,6 +114,51 @@ pub const rom = struct {
}
};
+/// Evict every flash-mapped cache line the bootloader left behind, by reading more flash than the
+/// caches can hold. Call it before the first byte of `.rodata` is touched.
+///
+/// This is a workaround for a real defect in the hand-over, not a tidiness measure. The second-stage
+/// bootloader leaves lines cached against a mapping it then replaces, so an application reads its
+/// own `.rodata` and gets its own `.text` back - **deterministically**, which is exactly what makes
+/// it look like anything other than a cache. Measured on this die: a load at `0x40035A1C` returned
+/// `93 85 85 0f` before eviction and `3c ee 08 40` after, and the second is what the image holds
+/// there. A string literal read before this runs is machine code, so a firmware whose first act is
+/// to print a marker prints garbage and looks like it never booted at all.
+///
+/// Everything cheaper was tried first and every one of them said the hardware was fine, which is
+/// why the list is here rather than being rediscovered:
+///
+/// * The MMU table is correct. Entries 0..9 read `0x1001`..`0x100a` - the valid bit plus physical
+/// page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires,
+/// and 10..11 are unmapped as they should be. Read off the die through
+/// `SPI_MEM_C_MMU_ITEM_INDEX_REG`, not inferred.
+/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the
+/// image matches the ELF byte for byte at the addresses that misread.
+/// * The page size is not in question: hardwired to 64 KiB on this chip
+/// (`hal/esp32p4/mmu_ll.h:126-130` returns `MMU_PAGE_64KB` and the setter asserts it).
+/// * Not fragmentation of the mapping either: the bad bytes arrive in one contiguous run of
+/// >= 192 B, not in 64-byte lines, and they are identical across three resets and two
+/// reflashes - determinism is what kept this looking like anything but a cache.
+///
+/// 512 KiB is four times the 128 KiB the L2 measured at (`examples/memprobe.zig` found real RAM
+/// stopping at `0x4FFA0000`, the cache taking the rest), with the L1s smaller still.
+///
+/// `rom.Cache_Invalidate_All` is the instrument that ought to do this and does not: called from an
+/// image the ROM did not launch, it faults inside the ROM with its argument stranded in `a2`.
+/// Capacity eviction needs no preconditions, which is the whole reason it is what ships. 512 KiB at
+/// a 64-byte stride is 8,192 loads, once per boot.
+pub fn flushFlashCache() void {
+ var sink: u32 = 0;
+ var p: u32 = 0x4000_0000;
+ while (p < 0x4008_0000) : (p += 64) {
+ sink +%= @as(*volatile u32, @ptrFromInt(p)).*;
+ }
+ // Consumed through a volatile store, or the optimiser drops the whole loop as dead.
+ @as(*volatile u32, &flush_sink).* = sink;
+}
+
+var flush_sink: u32 = 0;
+
/// Busy-wait for a number of CPU cycles, using the cycle counter rather than the mask ROM. Useful
/// when an image must not depend on ROM entry points at all, and for delays shorter than the ROM's
/// microsecond granularity.
diff --git a/tools/bench_main.zig b/tools/bench_main.zig
new file mode 100644
index 0000000..b3470c2
--- /dev/null
+++ b/tools/bench_main.zig
@@ -0,0 +1,695 @@
+//! `p4-bench`: what the board's serial link can actually carry, verified byte for byte.
+//!
+//! Two modes, because there are two different questions and conflating them is how this port got
+//! optimised by guesswork so far.
+//!
+//! **`--link` (default) needs `examples/uartperf.zig` flashed.** It measures the CEILING: the wire
+//! and the UART driver with nothing else running. Every number is checksummed - the board reports a
+//! CRC over exactly the bytes it received, and the host compares it against a CRC over exactly the
+//! bytes it sent. An unverified throughput figure is a guess about how fast data was corrupted,
+//! and on this UART the failure that matters is a silent RX overrun, which a byte count cannot see.
+//!
+//! **`--editor` needs the editor flashed.** It measures how much of that ceiling pardes uses, by
+//! typing at rising rates until it falls behind. There is no protocol available here - the board is
+//! running an editor, and its answer to a keystroke is a screen update - so this half is timing
+//! only, and it is honest about that.
+//!
+//! The uplink figure is measured to `tcdrain`, not to the last `write`. A write returns once the
+//! kernel has accepted the bytes, which at 115200 is long before they are on the wire; timing to the
+//! write would report the speed of memcpy into a tty buffer.
+
+const std = @import("std");
+const serial = @import("serial.zig");
+const proto = @import("perfproto");
+const rtt = @import("rtt.zig");
+
+const Mode = enum {
+ /// Verified bulk throughput and latency against `examples/uartperf.zig`. The ceiling.
+ link,
+ /// Keystroke latency against the editor, plus the rate ladder.
+ editor,
+ /// A named experiment: one controlled variable, many trials, machine-readable.
+ sweep,
+};
+
+/// Which variable an experiment varies. One per run, because the point is a controlled variable and
+/// a session that changed two things at once would not answer either question.
+const Sweep = enum {
+ /// Screen area, over an in-band resize. Tests whether per-keystroke cost is paid per CELL.
+ geometry,
+ /// Characters already in the line before the measured keystroke. Tests whether an edit is O(n)
+ /// in the buffer - `modal.spliceAlloc` copies the whole content per keystroke, so it should be.
+ length,
+ /// One-off operations at a fixed geometry: motions, an insert, and a forced full repaint.
+ ops,
+};
+
+/// Widths and signed integers do not mix in Zig 0.16: `printIntAny` emits an explicit `+` for any
+/// non-negative SIGNED value whenever a width is given (`std/Io/Writer.zig:1548-1559` - the plus is
+/// omitted only when `width` is null or zero). Every number here is a duration or a count that
+/// cannot be negative, so they are printed as unsigned and the tables line up.
+fn pos(x: i64) u64 {
+ return @intCast(@max(0, x));
+}
+
+const Options = struct {
+ port: []const u8 = "/dev/ttyUSB0",
+ baud: serial.Baud = .b115200,
+ mode: Mode = .link,
+ /// Payload bytes per direction for the bulk tests. 64 KiB is ~5.7 s each way at 115200 - long
+ /// enough that start-up transients do not dominate, short enough to rerun after every change.
+ bulk: u32 = 64 * 1024,
+ /// Round trips for the latency figure.
+ samples: u32 = 32,
+ no_reset: bool = false,
+ which: Sweep = .geometry,
+ /// Trials per condition. Medians over an odd count, so the reported value is a real sample and
+ /// not an average smeared across a transient.
+ repeat: u32 = 5,
+ /// Emit one CSV row per trial instead of a table. Raw trials, not summaries: the analysis should
+ /// be able to see the spread and recompute any statistic, and a tool that only prints medians
+ /// has thrown that away.
+ csv: bool = false,
+ /// Stamped into every CSV row, so a file of results records the build it came from rather than
+ /// relying on the order the runs happened in.
+ label: []const u8 = "-",
+};
+
+pub fn main(init: std.process.Init.Minimal) void {
+ var o: Options = .{};
+ var it: std.process.Args.Iterator = .init(init.args);
+ _ = it.skip();
+ while (it.next()) |a| {
+ if (eql(a, "-h") or eql(a, "--help")) return usage();
+ if (eql(a, "--no-reset")) {
+ o.no_reset = true;
+ continue;
+ }
+ if (eql(a, "--link")) {
+ o.mode = .link;
+ continue;
+ }
+ if (eql(a, "--editor")) {
+ o.mode = .editor;
+ continue;
+ }
+ if (eql(a, "--csv")) {
+ o.csv = true;
+ continue;
+ }
+ const val = it.next() orelse fatal("that flag needs a value");
+ if (eql(a, "--port")) {
+ o.port = val;
+ } else if (eql(a, "--baud")) {
+ const rate = std.fmt.parseInt(u32, val, 10) catch fatal("--baud must be a number");
+ o.baud = std.enums.fromInt(serial.Baud, rate) orelse fatal("unsupported baud");
+ } else if (eql(a, "--bulk")) {
+ o.bulk = std.fmt.parseInt(u32, val, 10) catch fatal("--bulk must be a number");
+ } else if (eql(a, "--samples")) {
+ o.samples = std.fmt.parseInt(u32, val, 10) catch fatal("--samples must be a number");
+ } else if (eql(a, "--repeat")) {
+ o.repeat = std.fmt.parseInt(u32, val, 10) catch fatal("--repeat must be a number");
+ } else if (eql(a, "--label")) {
+ o.label = val;
+ } else if (eql(a, "--sweep")) {
+ o.mode = .sweep;
+ o.which = std.meta.stringToEnum(Sweep, val) orelse
+ fatal("--sweep takes geometry, length or ops");
+ } else fatal("unrecognised argument; try --help");
+ }
+ run(o) catch |err| switch (err) {
+ error.AccessDenied => {
+ out("p4-bench: cannot open the port: AccessDenied\n\n");
+ out(serial.access_denied_help);
+ out("\n");
+ std.process.exit(1);
+ },
+ error.NoMarker => fatal(
+ \\the board never printed its readiness marker after reset.
+ \\
+ \\ Expected `MARK UARTPERF_READY` within 25 s. Flash the responder:
+ \\ zig build flash -Dapp=examples/uartperf.zig
+ ),
+ error.NoResponder => fatal(
+ \\the board is not answering the measurement protocol.
+ \\
+ \\ Flash the responder first:
+ \\ zig build flash -Dapp=examples/uartperf.zig
+ \\ Or measure the editor instead:
+ \\ p4-bench --editor
+ ),
+ error.NoEditor => fatal(
+ \\the board never reached the editor (no `MARK PARDES_READY` within 25 s).
+ \\
+ \\ zig build flash -Dpardes
+ ),
+ else => {
+ var b: [128]u8 = undefined;
+ fatal(std.fmt.bufPrint(&b, "{s}", .{@errorName(err)}) catch "failed");
+ },
+ };
+}
+
+/// Accumulates bytes off the wire and hands back whole frames. A frame split across reads is the
+/// normal case on a serial line, so the buffer is the struct rather than a local.
+const Frames = struct {
+ buf: [proto.header_len + proto.max_payload]u8 = undefined,
+ len: usize = 0,
+ /// Bytes discarded while resynchronising. Nonzero means the stream contained something that was
+ /// not a frame, which is itself a finding.
+ junk: u32 = 0,
+ /// Length of the frame handed out by the last `take`, still occupying the head of the buffer.
+ /// `commit` is what removes it, so a caller may borrow a payload across the call that produced
+ /// it and no further.
+ pending: usize = 0,
+
+ const Frame = struct { op: proto.Op, payload: []const u8 };
+
+ /// The next whole frame, or null on timeout. `payload` borrows the buffer and is invalidated by
+ /// the following call.
+ fn next(f: *Frames, port: *serial.Port, timeout_us: i64) !?Frame {
+ const deadline = nowUs(port) + timeout_us;
+ while (true) {
+ // Serve from what is already buffered before touching the wire: a single read can
+ // deliver several frames, and re-polling between them would add latency that is not
+ // the board's.
+ if (f.take()) |fr| return fr;
+ if (nowUs(port) >= deadline) return null;
+ if (f.len == f.buf.len) {
+ // Full and still not a frame: the buffer holds only junk. Drop one byte so the
+ // resynchronising scan can advance.
+ f.drop(1);
+ continue;
+ }
+ const n = try port.readTimeout(f.buf[f.len..], 2);
+ f.len += n;
+ }
+ }
+
+ fn take(f: *Frames) ?Frame {
+ while (f.len > 0) {
+ const header = proto.parseHeader(f.buf[0..f.len]) catch {
+ f.resync();
+ continue;
+ } orelse return null;
+ const total = proto.header_len + @as(usize, header.len);
+ if (f.len < total) return null;
+ const payload = f.buf[proto.header_len..total];
+ if (proto.crc(payload) != header.crc) {
+ f.resync();
+ continue;
+ }
+ f.pending = total;
+ return .{ .op = header.op, .payload = payload };
+ }
+ return null;
+ }
+
+ /// Advance to the next plausible frame start. Dropping ONE byte per call was the obvious
+ /// spelling and it is far too slow to be correct here: after a reset the buffer holds ~1.4 KB of
+ /// bootloader log, and one byte discarded per poll took longer than the handshake timeout, so a
+ /// working board looked like a missing one. Scanning to the next `P` covers the whole run of
+ /// junk in one step.
+ fn resync(f: *Frames) void {
+ f.junk += 1;
+ const next_magic = std.mem.indexOfScalarPos(u8, f.buf[0..f.len], 1, proto.magic[0]) orelse f.len;
+ f.drop(next_magic);
+ }
+
+ fn drop(f: *Frames, n: usize) void {
+ const k = @min(n, f.len);
+ std.mem.copyForwards(u8, f.buf[0 .. f.len - k], f.buf[k..f.len]);
+ f.len -= k;
+ }
+
+ fn commit(f: *Frames) void {
+ if (f.pending > 0) {
+ f.drop(f.pending);
+ f.pending = 0;
+ }
+ }
+};
+
+fn run(o: Options) !void {
+ var port = try serial.Port.open(o.port, o.baud);
+ defer port.close();
+
+ var r: Report = .{};
+ // No banner in CSV mode: a file of results should be parseable by anything that reads CSV, and
+ // a human-readable header line at the top of it is not. The condition is stamped into every row
+ // by `--label` instead, which survives concatenation of several runs.
+ if (!o.csv) {
+ r.print("p4-bench {s} @ {d} baud wire capacity {d} B/s each way\n\n", .{
+ o.port, o.baud.rate(), port.capacity(),
+ });
+ r.flush();
+ }
+
+ if (!o.no_reset) try port.resetToRun(.{});
+
+ switch (o.mode) {
+ .link => try link(&port, o, &r),
+ .editor => try editor(&port, o, &r),
+ .sweep => try sweep(&port, o, &r),
+ }
+}
+
+/// Put the editor in a known state: reached, first frame drawn, insert mode on.
+///
+/// Every experiment starts here, and it matters that it is the same every time. `rtt.roundTrip`
+/// with an empty stimulus is used as a settle: it sends nothing and returns when the wire has been
+/// quiet, which is exactly "wait for the board to stop talking".
+fn ready(port: *serial.Port, o: Options) !void {
+ if (!o.no_reset) try waitFor(port, "MARK PARDES_READY", 25_000);
+ _ = try rtt.roundTrip(port, "", 3_000_000, 300_000);
+ try port.write("i");
+ _ = try rtt.roundTrip(port, "", 400_000, 250_000);
+}
+
+/// One controlled-variable experiment, emitted as raw trials.
+///
+/// The measured quantity is always the same - the round trip of ONE inserted character - so that
+/// conditions are comparable. Only the condition changes.
+fn sweep(port: *serial.Port, o: Options, r: *Report) !void {
+ try ready(port, o);
+ if (o.csv) r.print("label,experiment,cols,rows,length,op,rep,rtt_us,settle_us,bytes\n", .{});
+
+ switch (o.which) {
+ // AREA. If a keystroke's cost is paid per cell, halving the rows should roughly halve the
+ // compute. If the cost is per EDIT, geometry will barely move it. The editor clamps itself
+ // to 40x12, so these are all reachable and 40x12 is the ceiling rather than a midpoint.
+ .geometry => {
+ for ([_][2]u16{ .{ 40, 12 }, .{ 40, 8 }, .{ 40, 6 }, .{ 30, 12 }, .{ 20, 12 }, .{ 20, 6 } }) |g| {
+ var buf: [32]u8 = undefined;
+ const resize = std.fmt.bufPrint(&buf, "\x1b[48;{d};{d};0;0t", .{ g[1], g[0] }) catch continue;
+ try port.write(resize);
+ // A resize is a full repaint; let it finish so it is not measured as a keystroke.
+ _ = try rtt.roundTrip(port, "", 3_000_000, 400_000);
+ try trials(port, o, r, .{ .cols = g[0], .rows = g[1], .op = "insert" });
+ }
+ },
+ // LENGTH. `modal.spliceAlloc` allocates and copies the whole buffer for every edit, so the
+ // per-keystroke cost should rise with the line. This is the experiment that decides whether
+ // the ~14 ms is a fixed overhead or a function of the document.
+ .length => {
+ var at: u32 = 0;
+ for ([_]u32{ 0, 20, 40, 80, 160, 320, 640 }) |target| {
+ // Type up to the target WITHOUT measuring, so the measured keystroke always sees a
+ // line of exactly `target` characters before it.
+ // Primed in small batches rather than one round trip per character: the priming is
+ // not the measurement, and a round trip each cost 0.3 s, which made the 160-character
+ // condition take minutes. Eight at a time is 8 B on a wire with a 128-byte FIFO, so
+ // nothing can be lost, and one settle per batch keeps the board from queueing.
+ while (at < target) {
+ const batch: u32 = @min(8, target - at);
+ var fill: [8]u8 = @splat('y');
+ try port.write(fill[0..batch]);
+ _ = try rtt.roundTrip(port, "", 2_000_000, 120_000);
+ at += batch;
+ }
+ try trials(port, o, r, .{ .length = target, .op = "insert" });
+ }
+ },
+ // OPS. Not a sweep of a number but of a KIND, to separate "an edit" from "a motion" from
+ // "everything changed". Repeated, because the one-shot table showed 40-byte motions costing
+ // the same round trip as 81-byte inserts and that needs more than one sample to assert.
+ .ops => {
+ try port.write("\x1b"); // motions must be motions
+ _ = try rtt.roundTrip(port, "", 400_000, 250_000);
+ for ([_]struct { name: []const u8, keys: []const u8 }{
+ .{ .name = "motion_h", .keys = "h" },
+ .{ .name = "motion_l", .keys = "l" },
+ .{ .name = "line_start", .keys = "0" },
+ .{ .name = "line_end", .keys = "$" },
+ .{ .name = "insert_esc", .keys = "ix\x1b" },
+ .{ .name = "repaint_39", .keys = "\x1b[48;12;39;0;0t" },
+ .{ .name = "repaint_40", .keys = "\x1b[48;12;40;0;0t" },
+ }) |p| {
+ try trials(port, o, r, .{ .op = p.name, .keys = p.keys });
+ }
+ },
+ }
+}
+
+const Condition = struct {
+ cols: u16 = 0,
+ rows: u16 = 0,
+ length: u32 = 0,
+ op: []const u8,
+ /// The stimulus. Defaults to one inserted character, which is the comparable unit.
+ keys: []const u8 = "x",
+};
+
+/// `o.repeat` trials of one condition. Raw rows in CSV mode; median in table mode.
+fn trials(port: *serial.Port, o: Options, r: *Report, c: Condition) !void {
+ var rtts: [64]i64 = undefined;
+ var got: u32 = 0;
+ var lost: u32 = 0;
+ var last_bytes: usize = 0;
+ var last_settle: i64 = 0;
+ const n = @min(o.repeat, 64);
+ for (0..n) |rep| {
+ const s = try rtt.roundTrip(port, c.keys, 3_000_000, 250_000);
+ if (s) |v| {
+ if (got < 64) {
+ rtts[got] = v.rtt_us;
+ got += 1;
+ }
+ last_bytes = v.bytes;
+ last_settle = v.settle_us;
+ if (o.csv) r.print("{s},{s},{d},{d},{d},{s},{d},{d},{d},{d}\n", .{
+ o.label, @tagName(o.which), c.cols, c.rows, c.length, c.op,
+ rep, pos(v.rtt_us), pos(v.settle_us), v.bytes,
+ });
+ } else lost += 1;
+ r.flush();
+ }
+ if (o.csv) return;
+ if (got == 0) {
+ r.print(" {s:<12} {d:>3}x{d:<3} len {d:>4} no response\n", .{ c.op, c.cols, c.rows, c.length });
+ return;
+ }
+ std.mem.sort(i64, rtts[0..got], {}, std.sort.asc(i64));
+ r.print(" {s:<12} {d:>3}x{d:<3} len {d:>4} median {d:>7} us spread {d:>6} us {d:>5} B\n", .{
+ c.op, c.cols, c.rows, c.length,
+ pos(rtts[got / 2]), pos(rtts[got - 1] - rtts[0]), last_bytes,
+ });
+ r.flush();
+}
+
+fn link(port: *serial.Port, o: Options, r: *Report) !void {
+ if (!o.no_reset) {
+ try waitFor(port, "MARK UARTPERF_READY", 25_000);
+ // The bootloader's log is not ours and it is still arriving. Feeding ~1.4 KB of text to a
+ // frame parser wastes the handshake window resynchronising through it.
+ port.drain();
+ }
+
+ var frames: Frames = .{};
+
+ // A ping proves the responder is there and the framing agrees, before anything is timed.
+ //
+ // Retried, because the first one after a reset can genuinely be lost: the board prints its
+ // marker from `zig_main` and the host answers within microseconds, while the board is still
+ // inside `write` pushing the rest of that string through a 128-byte FIFO. Its RX FIFO holds the
+ // ping meanwhile, but a `source`-sized burst is not the only thing that can outlast one - and a
+ // measuring instrument that fails on a startup race would be reporting its own bug as the
+ // board's.
+ {
+ var buf: [proto.header_len + 8]u8 = undefined;
+ const ping = proto.encode(&buf, .ping, "handshake"[0..8]);
+ var tries: u32 = 0;
+ while (true) : (tries += 1) {
+ if (tries == 5) return error.NoResponder;
+ try port.write(ping);
+ const fr = (try frames.next(port, 500_000)) orelse continue;
+ const ok = fr.op == .pong and std.mem.eql(u8, fr.payload, "handshake"[0..8]);
+ frames.commit();
+ if (ok) break;
+ }
+ }
+
+ // ---- UPLINK: host -> board, verified by the board's CRC over what arrived.
+ var chunk: [proto.max_payload]u8 = undefined;
+ var frame: [proto.header_len + proto.max_payload]u8 = undefined;
+ var sent: u32 = 0;
+ var hash: std.hash.Crc32 = .init();
+ const t_up = nowUs(port);
+ while (sent < o.bulk) {
+ const take: u32 = @min(@as(u32, proto.max_payload), o.bulk - sent);
+ proto.fillPattern(chunk[0..take], sent);
+ hash.update(chunk[0..take]);
+ try port.write(proto.encode(&frame, .sink, chunk[0..take]));
+ sent += take;
+ }
+ // A write returns once the kernel has the bytes, not once the wire does; without this the
+ // uplink figure was 202% of the link's capacity.
+ port.flushOutput();
+ const up_us = @max(1, nowUs(port) - t_up);
+ const want_crc = hash.final();
+
+ try port.write(proto.encode(&frame, .report, ""));
+ const up_stat = blk: {
+ while (true) {
+ const fr = (try frames.next(port, 3_000_000)) orelse return error.NoResponder;
+ if (fr.op == .stat) {
+ const s = proto.Stat.decode(fr.payload) orelse return error.NoResponder;
+ frames.commit();
+ break :blk s;
+ }
+ frames.commit();
+ }
+ };
+
+ const up_ok = up_stat.bytes == sent and up_stat.crc == want_crc;
+ r.print(" uplink host -> board\n", .{});
+ r.print(" {d} B in {d} us = {d} B/s ({d}% of wire)\n", .{
+ sent, up_us, @divTrunc(@as(i64, sent) * 1_000_000, up_us),
+ @divTrunc(@as(i64, sent) * 1_000_000 * 100, up_us * @as(i64, port.capacity())),
+ });
+ if (up_ok) {
+ r.print(" VERIFIED crc 0x{x:0>8} over {d} B\n", .{ up_stat.crc, up_stat.bytes });
+ } else {
+ r.print(" FAILED board got {d} B crc 0x{x:0>8}; host sent {d} B crc 0x{x:0>8}", .{
+ up_stat.bytes, up_stat.crc, sent, want_crc,
+ });
+ if (up_stat.bytes < sent) r.print(" <-- {d} B LOST", .{sent - up_stat.bytes});
+ r.print("\n", .{});
+ }
+ if (up_stat.bad_frames > 0) r.print(" {d} frames arrived corrupt\n", .{up_stat.bad_frames});
+ if (up_stat.tx_dropped > 0) r.print(" board dropped {d} B on transmit\n", .{up_stat.tx_dropped});
+ r.flush();
+
+ // ---- DOWNLINK: board -> host, verified by the host's CRC over what arrived.
+ var req: [4]u8 = undefined;
+ std.mem.writeInt(u32, &req, o.bulk, .little);
+ const t_down = nowUs(port);
+ try port.write(proto.encode(&frame, .source, &req));
+ var got: u32 = 0;
+ var down_hash: std.hash.Crc32 = .init();
+ var down_stat: ?proto.Stat = null;
+ var last = t_down;
+ while (down_stat == null) {
+ const fr = (try frames.next(port, 5_000_000)) orelse break;
+ switch (fr.op) {
+ .data => {
+ got += @intCast(fr.payload.len);
+ down_hash.update(fr.payload);
+ last = nowUs(port);
+ },
+ .stat => down_stat = proto.Stat.decode(fr.payload),
+ else => {},
+ }
+ frames.commit();
+ }
+ const down_us = @max(1, last - t_down);
+ const mine = down_hash.final();
+
+ r.print("\n downlink board -> host\n", .{});
+ r.print(" {d} B in {d} us = {d} B/s ({d}% of wire)\n", .{
+ got, down_us, @divTrunc(@as(i64, got) * 1_000_000, down_us),
+ @divTrunc(@as(i64, got) * 1_000_000 * 100, down_us * @as(i64, port.capacity())),
+ });
+ if (down_stat) |s| {
+ if (s.bytes == got and s.crc == mine) {
+ r.print(" VERIFIED crc 0x{x:0>8} over {d} B\n", .{ mine, got });
+ } else {
+ r.print(" FAILED board sent {d} B crc 0x{x:0>8}; host got {d} B crc 0x{x:0>8}", .{
+ s.bytes, s.crc, got, mine,
+ });
+ if (got < s.bytes) r.print(" <-- {d} B LOST", .{s.bytes - got});
+ r.print("\n", .{});
+ }
+ } else r.print(" FAILED no closing stat frame\n", .{});
+ if (frames.junk > 0) r.print(" {d} resynchronisation events on the host\n", .{frames.junk});
+ r.flush();
+
+ // ---- LATENCY: a verified round trip, so a lost ping is distinguishable from a slow one.
+ var min: i64 = std.math.maxInt(i64);
+ var max: i64 = 0;
+ var sum: i64 = 0;
+ var ok: u32 = 0;
+ var lost: u32 = 0;
+ for (0..o.samples) |_| {
+ const t0 = nowUs(port);
+ try port.write(proto.encode(&frame, .ping, "ping"));
+ var answered = false;
+ while (try frames.next(port, 500_000)) |fr| {
+ const was_pong = fr.op == .pong and std.mem.eql(u8, fr.payload, "ping");
+ frames.commit();
+ if (was_pong) {
+ answered = true;
+ break;
+ }
+ }
+ if (!answered) {
+ lost += 1;
+ continue;
+ }
+ const dt = nowUs(port) - t0;
+ min = @min(min, dt);
+ max = @max(max, dt);
+ sum += dt;
+ ok += 1;
+ }
+ r.print("\n round trip 13 B out, 13 B back\n", .{});
+ if (ok > 0) {
+ r.print(" min {d} us mean {d} us max {d} us ({d} samples, {d} lost)\n", .{
+ pos(min), pos(@divTrunc(sum, ok)), pos(max), o.samples, lost,
+ });
+ r.print(" wire floor for 26 B is {d} us; the rest is the board\n", .{
+ @divTrunc(26 * 1_000_000, @as(i64, port.capacity())),
+ });
+ } else r.print(" every ping lost\n", .{});
+ r.flush();
+}
+
+fn editor(port: *serial.Port, o: Options, r: *Report) !void {
+ if (!o.no_reset) try waitFor(port, "MARK PARDES_READY", 25_000);
+ // Settle the first full frame before anything is timed against it.
+ _ = try rtt.roundTrip(port, "", 2_000_000, 300_000);
+
+ // The editor is modal: a bare `x` would be a motion. One `i` makes every later `x` an edit,
+ // which is the cheapest change that still forces a real render.
+ try port.write("i");
+ _ = try rtt.roundTrip(port, "", 300_000, 200_000);
+
+ const base = try rtt.measure(port, .{ .samples = @min(o.samples, 16), .gap_us = 400_000 });
+ if (base.median_us < 0) return error.NoEditor;
+ r.print(" editor, uncontended at 2.5 keys/s\n", .{});
+ r.print(" rtt median {d} us settle {d} us {d} B per keystroke\n\n", .{
+ pos(base.median_us), pos(base.median_settle_us), base.median_bytes,
+ });
+ r.print(" keys/s median rtt settle B/key wire lost verdict\n", .{});
+ r.print(" ---------------------------------------------------------------\n", .{});
+ r.flush();
+
+ const ceiling = base.median_us * 3;
+ var best: i64 = -1;
+ for ([_]i64{ 160, 120, 80, 60, 40, 30, 20, 15, 10 }) |gap_ms| {
+ const s = try rtt.measure(port, .{
+ .samples = @min(o.samples, 16),
+ .gap_us = gap_ms * 1000,
+ .timeout_us = 1_000_000,
+ });
+ const rate = @divTrunc(@as(i64, 1000), gap_ms);
+ const pass = s.lost == 0 and s.median_us >= 0 and s.median_us <= ceiling;
+ if (pass) best = rate;
+ r.print(" {d:>7} {d:>9} us {d:>9} us {d:>7} {d:>4}% {d:>5} {s}\n", .{
+ pos(rate), pos(s.median_us), pos(s.median_settle_us),
+ s.median_bytes, s.wire_percent, s.lost,
+ if (s.lost > 0) "LOST INPUT" else if (pass) "ok" else "behind",
+ });
+ r.flush();
+ if (s.lost > 0) break;
+ }
+ r.print("\n", .{});
+ if (best < 0) {
+ r.print(" CEILING: under 6 keys/s - it kept up at no rate tried.\n", .{});
+ } else {
+ r.print(" CEILING: {d} keys/s sustained (median rtt within 3x of {d} us).\n", .{ best, base.median_us });
+ }
+ r.flush();
+
+ // ONE-OFF COSTS. Typing turned out to be cheap, so the operations that are not typing are where
+ // "too slow to use" has to live. Each is measured once, in the state the ladder left the buffer
+ // in (a long line of `x`), and the interesting column is bytes: an operation that emits ~1.5 KB
+ // has repainted the whole screen, and at this baud that is 130 ms of wire before anything else
+ // can happen.
+ try port.write("\x1b"); // out of insert mode; motions are motions again
+ _ = try rtt.roundTrip(port, "", 400_000, 250_000);
+
+ r.print("\n one-off operations (bytes is the tell: ~1.5 KB is a full repaint)\n", .{});
+ r.print(" operation rtt settle bytes\n", .{});
+ r.print(" ------------------------------------------------\n", .{});
+ const probes = [_]struct { name: []const u8, keys: []const u8 }{
+ .{ .name = "motion h", .keys = "h" },
+ .{ .name = "motion l", .keys = "l" },
+ .{ .name = "line start", .keys = "0" },
+ .{ .name = "line end", .keys = "$" },
+ .{ .name = "insert char", .keys = "ix\x1b" },
+ // A geometry change is the one stimulus guaranteed to force a full repaint, so it
+ // calibrates the column: whatever this costs is what "everything changed" costs.
+ .{ .name = "resize 40->39", .keys = "\x1b[48;12;39;0;0t" },
+ .{ .name = "resize 39->40", .keys = "\x1b[48;12;40;0;0t" },
+ };
+ for (probes) |p| {
+ if (try rtt.roundTrip(port, p.keys, 3_000_000, 250_000)) |s| {
+ r.print(" {s:<16} {d:>8} us {d:>9} us {d:>9}\n", .{
+ p.name, pos(s.rtt_us), pos(s.settle_us), s.bytes,
+ });
+ } else {
+ r.print(" {s:<16} no response (the editor ignored it)\n", .{p.name});
+ }
+ r.flush();
+ }
+}
+
+/// Wait for a plain text marker rather than a fixed delay: a slow boot should lengthen the run, not
+/// silently start measuring a board that is still in its bootloader.
+fn waitFor(port: *serial.Port, marker: []const u8, timeout_ms: i64) !void {
+ var at: usize = 0;
+ var buf: [1024]u8 = undefined;
+ const deadline = port.nowMs() + timeout_ms;
+ while (port.nowMs() < deadline) {
+ const n = port.readTimeout(&buf, 200) catch 0;
+ for (buf[0..n]) |b| {
+ if (b == marker[at]) {
+ at += 1;
+ if (at == marker.len) return;
+ } else at = if (b == marker[0]) 1 else 0;
+ }
+ }
+ return error.NoMarker;
+}
+
+const Report = struct {
+ buf: [4096]u8 = undefined,
+ len: usize = 0,
+
+ fn print(self: *Report, comptime fmt: []const u8, args: anytype) void {
+ const s = std.fmt.bufPrint(self.buf[self.len..], fmt, args) catch return;
+ self.len += s.len;
+ }
+
+ fn flush(self: *Report) void {
+ out(self.buf[0..self.len]);
+ self.len = 0;
+ }
+};
+
+fn usage() void {
+ out(
+ \\p4-bench - measure the board's serial link, verified with a checksum
+ \\
+ \\ p4-bench [--link | --editor] [--port <path>] [--baud <rate>]
+ \\ [--bulk <bytes>] [--samples <n>] [--no-reset]
+ \\
+ \\ --link (default) bulk throughput each way plus round-trip latency, every byte
+ \\ checksummed. Needs examples/uartperf.zig flashed.
+ \\ --editor type at rising rates against pardes and report the highest rate it keeps
+ \\ up with. Needs the editor flashed.
+ \\
+ );
+}
+
+fn nowUs(port: *serial.Port) i64 {
+ return std.Io.Timestamp.now(port.io, .boot).toMicroseconds();
+}
+
+const io = std.Io.Threaded.global_single_threaded.io();
+const stdout: std.Io.File = .{ .handle = 1, .flags = .{ .nonblocking = false } };
+
+fn out(s: []const u8) void {
+ stdout.writeStreamingAll(io, s) catch {};
+}
+
+fn fatal(msg: []const u8) noreturn {
+ var b: [512]u8 = undefined;
+ out(std.fmt.bufPrint(&b, "p4-bench: {s}\n", .{msg}) catch "p4-bench: error\n");
+ std.process.exit(1);
+}
+
+fn eql(a: []const u8, b: []const u8) bool {
+ return std.mem.eql(u8, a, b);
+}
diff --git a/tools/perfproto.zig b/tools/perfproto.zig
new file mode 100644
index 0000000..5223500
--- /dev/null
+++ b/tools/perfproto.zig
@@ -0,0 +1,201 @@
+//! A small framed protocol for measuring the board's serial link, shared verbatim by the host tool
+//! and the firmware that answers it.
+//!
+//! WHY A PROTOCOL AND NOT A STOPWATCH. Timing an editor's keystrokes measures the editor, the
+//! renderer and the link at once, and cannot tell a dropped byte from a slow one: RX overrun on this
+//! UART is silent in hardware and uncounted in the driver, so a missing keystroke and a late one look
+//! identical from the host. A frame with a length and a checksum turns both into facts. If the CRC
+//! matches, every byte of that payload crossed intact; if a frame never completes, bytes were lost
+//! and the count says how many. A throughput number that is not checksummed is a guess about how
+//! fast data was corrupted.
+//!
+//! THE SHAPE. One fixed 9-byte header, little-endian, then the payload:
+//!
+//! "P4" op:u8 len:u16 crc:u32 payload[len]
+//!
+//! The CRC covers the payload only. The header carries it rather than trailing it so a receiver
+//! knows, before it has read a single payload byte, exactly how many to expect and what they must
+//! hash to - which is what lets the firmware verify a stream with one 4-byte accumulator and no
+//! buffer at all.
+//!
+//! `max_payload` is 1024 and that is a memory decision, not a wire one. The firmware has a 384 KiB
+//! heap it must share with an editor, and a bulk test that needed a 64 KiB frame buffer would be
+//! measuring a configuration nobody ships. Bulk transfers are therefore many frames, which is also
+//! the honest shape: it is the per-frame overhead a real protocol would pay.
+//!
+//! Both directions use the same header, and a reply's op has the high bit set, so a stray reply can
+//! never be mistaken for a request by a resynchronising receiver.
+
+const std = @import("std");
+
+pub const magic = "P4";
+pub const header_len = 9;
+pub const max_payload = 1024;
+
+pub const Op = enum(u8) {
+ /// Echo the payload back as `pong`. Both directions verified in one exchange, which is what
+ /// makes it the right stimulus for a latency measurement.
+ ping = 1,
+ /// Payload is data to be consumed. The board accumulates a running count and CRC and answers
+ /// nothing, so the host can keep the uplink full and measure it without return traffic
+ /// competing for the same wire.
+ sink = 2,
+ /// Ask for the accumulated `sink` count and CRC, then reset them.
+ report = 3,
+ /// Payload is a u32 count: send exactly that many pattern bytes back, in `data` frames,
+ /// followed by a `stat`.
+ source = 4,
+
+ pong = 0x81,
+ /// Payload is `Stat`, packed little-endian.
+ stat = 0x83,
+ /// A chunk of `source` output.
+ data = 0x84,
+
+ pub fn isReply(o: Op) bool {
+ return @intFromEnum(o) & 0x80 != 0;
+ }
+};
+
+/// What the board reports about a stream it received or sent. Encoded by hand rather than by
+/// `@bitCast` of a packed struct: this crosses between a riscv32 firmware and an x86_64 host, and a
+/// layout that depends on either compiler's padding rules is a bug waiting for a target change.
+pub const Stat = struct {
+ /// Payload bytes accumulated.
+ bytes: u32,
+ /// CRC-32 over exactly those bytes, in order.
+ crc: u32,
+ /// Frames whose CRC did not match. Nonzero means the link corrupted data rather than losing it,
+ /// which is a different fault with a different fix.
+ bad_frames: u32,
+ /// Bytes the firmware's UART driver gave up on writing. Its own counter, surfaced here because
+ /// the host cannot see it any other way.
+ tx_dropped: u32,
+
+ pub const encoded_len = 16;
+
+ pub fn encode(s: Stat, out: *[encoded_len]u8) void {
+ std.mem.writeInt(u32, out[0..4], s.bytes, .little);
+ std.mem.writeInt(u32, out[4..8], s.crc, .little);
+ std.mem.writeInt(u32, out[8..12], s.bad_frames, .little);
+ std.mem.writeInt(u32, out[12..16], s.tx_dropped, .little);
+ }
+
+ pub fn decode(in: []const u8) ?Stat {
+ if (in.len < encoded_len) return null;
+ return .{
+ .bytes = std.mem.readInt(u32, in[0..4], .little),
+ .crc = std.mem.readInt(u32, in[4..8], .little),
+ .bad_frames = std.mem.readInt(u32, in[8..12], .little),
+ .tx_dropped = std.mem.readInt(u32, in[12..16], .little),
+ };
+ }
+};
+
+pub fn crc(bytes: []const u8) u32 {
+ return std.hash.Crc32.hash(bytes);
+}
+
+/// The deterministic byte at stream offset `i`.
+///
+/// A counter would be checksummed correctly by an implementation that lost exactly 256 bytes, and a
+/// constant by one that lost any amount. This is an 8-bit xorshift-ish walk whose period is long
+/// enough that no realistic loss aligns with it, so the CRC catches a gap wherever it falls.
+pub fn patternByte(i: u32) u8 {
+ var x: u32 = i +% 1;
+ x ^= x << 7;
+ x ^= x >> 3;
+ x ^= x << 5;
+ return @truncate(x);
+}
+
+pub fn fillPattern(buf: []u8, offset: u32) void {
+ for (buf, 0..) |*b, k| b.* = patternByte(offset +% @as(u32, @intCast(k)));
+}
+
+/// Write a frame into `out`, returning the used slice. `out` must hold `header_len + payload.len`.
+pub fn encode(out: []u8, op: Op, payload: []const u8) []u8 {
+ std.debug.assert(payload.len <= max_payload);
+ std.debug.assert(out.len >= header_len + payload.len);
+ out[0] = magic[0];
+ out[1] = magic[1];
+ out[2] = @intFromEnum(op);
+ std.mem.writeInt(u16, out[3..5], @intCast(payload.len), .little);
+ std.mem.writeInt(u32, out[5..9], crc(payload), .little);
+ @memcpy(out[header_len..][0..payload.len], payload);
+ return out[0 .. header_len + payload.len];
+}
+
+pub const Header = struct {
+ op: Op,
+ len: u16,
+ crc: u32,
+};
+
+/// Read a header out of `buf`. Returns null when fewer than `header_len` bytes are present, and
+/// `error.BadFrame` when the magic or the op is not one of ours - which is how a receiver that has
+/// lost sync tells "wait for more" from "throw a byte away and try again".
+pub fn parseHeader(buf: []const u8) error{BadFrame}!?Header {
+ if (buf.len < header_len) return null;
+ if (buf[0] != magic[0] or buf[1] != magic[1]) return error.BadFrame;
+ const op = std.enums.fromInt(Op, buf[2]) orelse return error.BadFrame;
+ const len = std.mem.readInt(u16, buf[3..5], .little);
+ if (len > max_payload) return error.BadFrame;
+ return .{ .op = op, .len = len, .crc = std.mem.readInt(u32, buf[5..9], .little) };
+}
+
+test "a frame round-trips through encode and parseHeader" {
+ var buf: [header_len + 4]u8 = undefined;
+ const f = encode(&buf, .ping, "abcd");
+ try std.testing.expectEqual(@as(usize, header_len + 4), f.len);
+ const h = (try parseHeader(f)).?;
+ try std.testing.expectEqual(Op.ping, h.op);
+ try std.testing.expectEqual(@as(u16, 4), h.len);
+ try std.testing.expectEqual(crc("abcd"), h.crc);
+ try std.testing.expectEqualStrings("abcd", f[header_len..]);
+}
+
+test "a short buffer is incomplete, not invalid" {
+ var buf: [header_len]u8 = undefined;
+ const f = encode(&buf, .report, "");
+ try std.testing.expectEqual(@as(?Header, null), try parseHeader(f[0 .. header_len - 1]));
+}
+
+test "wrong magic and unknown ops are rejected rather than misread" {
+ var buf: [header_len]u8 = undefined;
+ var f = encode(&buf, .report, "");
+ f[0] = 'X';
+ try std.testing.expectError(error.BadFrame, parseHeader(f));
+ f[0] = magic[0];
+ f[2] = 0x7f;
+ try std.testing.expectError(error.BadFrame, parseHeader(f));
+}
+
+test "a truncated stream is caught by the CRC" {
+ // Losing bytes is the failure this protocol exists to detect, so prove the checksum notices a
+ // gap that leaves the length plausible.
+ var full: [64]u8 = undefined;
+ fillPattern(&full, 0);
+ var gapped: [64]u8 = undefined;
+ fillPattern(gapped[0..32], 0);
+ fillPattern(gapped[32..], 33); // one byte skipped mid-stream
+ try std.testing.expect(crc(&full) != crc(&gapped));
+}
+
+test "the pattern does not repeat inside a byte-aligned loss" {
+ // A plain counter would hash identically after losing exactly 256 bytes. This must not.
+ var a: [128]u8 = undefined;
+ var b: [128]u8 = undefined;
+ fillPattern(&a, 0);
+ fillPattern(&b, 256);
+ try std.testing.expect(crc(&a) != crc(&b));
+}
+
+test "Stat survives the trip between a riscv32 firmware and an x86_64 host" {
+ const s: Stat = .{ .bytes = 0x11223344, .crc = 0xdeadbeef, .bad_frames = 7, .tx_dropped = 9 };
+ var buf: [Stat.encoded_len]u8 = undefined;
+ s.encode(&buf);
+ const back = Stat.decode(&buf).?;
+ try std.testing.expectEqual(s, back);
+ try std.testing.expectEqual(@as(?Stat, null), Stat.decode(buf[0 .. Stat.encoded_len - 1]));
+}
diff --git a/tools/rtt.zig b/tools/rtt.zig
new file mode 100644
index 0000000..6ab0d87
--- /dev/null
+++ b/tools/rtt.zig
@@ -0,0 +1,170 @@
+//! Round-trip time over the board's only I/O channel: stimulus out, first byte back.
+//!
+//! Two functions, because every proposed fix for "too slow to type in" is a trade whose sign cannot
+//! be guessed - a frame-rate cap, draining RX while blocked on TX, coalescing input, raising the
+//! baud - and the only honest way to rank them is to measure the same number before and after.
+//!
+//! WHAT IS BEING TIMED, precisely: the interval from the last byte of a stimulus leaving the host to
+//! the FIRST byte of the board's response arriving. That is the latency a human perceives as
+//! responsiveness, and it is deliberately not the same as the time to finish repainting: a renderer
+//! that starts drawing in 8 ms and takes 130 ms to finish feels immediate, while one that thinks for
+//! 130 ms and then paints in 8 ms feels broken, and the two are indistinguishable if you only
+//! measure when the wire goes quiet. `settle_us` records the second number so the pair can be read
+//! together.
+//!
+//! Microseconds, not milliseconds: at 115200 baud one byte occupies 87 us, so a millisecond clock
+//! quantises this measurement into buckets 11 bytes wide.
+//!
+//! The caller owns the board's STATE. These functions send bytes and time bytes; they do not know
+//! what the editor does with them. A stimulus only produces a response if the editor is in a mode
+//! where that keystroke changes the screen - pardes is modal, so a caller measuring keystrokes must
+//! put it in insert mode first and must pick a stimulus that is not itself a mode change.
+
+const std = @import("std");
+const serial = @import("serial.zig");
+
+pub const Sample = struct {
+ /// Stimulus out -> first response byte in.
+ rtt_us: i64,
+ /// Stimulus out -> last response byte in, i.e. the wire is free again.
+ settle_us: i64,
+ /// How much the board emitted in answer. At 115200 this is also a time: bytes * 87 us.
+ bytes: usize,
+};
+
+/// One round trip. Returns null when nothing came back within `timeout_us` - which is a result, not
+/// an error: a dropped keystroke looks exactly like this, and it is the thing most worth counting.
+///
+/// `quiet_us` decides when the response is over. It must exceed the largest gap the board leaves
+/// mid-response; a renderer that pauses to allocate can stall longer than one byte time, and too
+/// small a value would split one response into two and report a `settle_us` that is too good.
+pub fn roundTrip(
+ port: *serial.Port,
+ stimulus: []const u8,
+ timeout_us: i64,
+ quiet_us: i64,
+) !?Sample {
+ // Anything still in flight belongs to the previous measurement. Without this the first read
+ // below returns instantly with stale bytes and reports an RTT near zero.
+ var drain: [1024]u8 = undefined;
+ while (try port.readTimeout(&drain, 0) > 0) {}
+
+ try port.write(stimulus);
+ const t0 = nowUs(port);
+
+ var first: i64 = -1;
+ var last: i64 = t0;
+ var bytes: usize = 0;
+ var buf: [4096]u8 = undefined;
+ while (true) {
+ const now = nowUs(port);
+ if (first < 0) {
+ if (now - t0 > timeout_us) return null;
+ } else if (now - last > quiet_us) break;
+
+ // Poll in millisecond units because that is what poll(2) takes; the TIMING above is
+ // microseconds and independent of this granularity.
+ const n = try port.readTimeout(&buf, 1);
+ if (n == 0) continue;
+ if (first < 0) first = nowUs(port);
+ bytes += n;
+ last = nowUs(port);
+ }
+ return .{ .rtt_us = first - t0, .settle_us = last - t0, .bytes = bytes };
+}
+
+pub const Stats = struct {
+ sent: u32,
+ /// Stimuli that produced no response at all inside the timeout. On this port that is a dropped
+ /// keystroke, and it is silent everywhere else in the system.
+ lost: u32,
+ min_us: i64,
+ median_us: i64,
+ max_us: i64,
+ /// Median, not mean: one 130 ms full repaint among fifty 9 ms updates should not move the
+ /// number that describes what typing feels like.
+ median_settle_us: i64,
+ median_bytes: usize,
+ /// Every response byte over the whole run, against the wire's capacity for that wall time.
+ /// 100% means the link is the limit and no amount of firmware tuning will help.
+ wire_percent: u32,
+
+ pub fn format(s: Stats, w: *std.Io.Writer) std.Io.Writer.Error!void {
+ try w.print("{d} samples, {d} lost\n", .{ s.sent, s.lost });
+ try w.print(" rtt min {d:>6} us median {d:>6} us max {d:>6} us\n", .{
+ s.min_us, s.median_us, s.max_us,
+ });
+ try w.print(" settle median {d} us ({d} B)\n", .{ s.median_settle_us, s.median_bytes });
+ try w.print(" wire {d}% of capacity\n", .{s.wire_percent});
+ }
+};
+
+pub const Options = struct {
+ samples: u32 = 20,
+ /// One byte that edits text without changing mode. `x` inserts an `x` in insert mode.
+ stimulus: []const u8 = "x",
+ /// Gap between stimuli. 80 ms is 12.5 characters a second: brisk human typing, and long enough
+ /// that a healthy editor finishes one update before the next arrives, so each sample is
+ /// independent rather than measuring a queue.
+ gap_us: i64 = 80_000,
+ timeout_us: i64 = 2_000_000,
+ quiet_us: i64 = 40_000,
+};
+
+/// `samples` round trips, summarised. The port must already be open and the board already in a
+/// state where `stimulus` changes the screen.
+pub fn measure(port: *serial.Port, opts: Options) !Stats {
+ const cap = 256;
+ var rtt: [cap]i64 = undefined;
+ var settle: [cap]i64 = undefined;
+ var size: [cap]usize = undefined;
+ var got: u32 = 0;
+ var lost: u32 = 0;
+ var total_bytes: usize = 0;
+
+ const n = @min(opts.samples, cap);
+ const t_start = nowUs(port);
+ for (0..n) |_| {
+ if (try roundTrip(port, opts.stimulus, opts.timeout_us, opts.quiet_us)) |s| {
+ rtt[got] = s.rtt_us;
+ settle[got] = s.settle_us;
+ size[got] = s.bytes;
+ total_bytes += s.bytes;
+ got += 1;
+ } else lost += 1;
+ std.Io.sleep(port.io, .fromMicroseconds(opts.gap_us), .boot) catch {};
+ }
+ const elapsed = @max(1, nowUs(port) - t_start);
+
+ if (got == 0) return .{
+ .sent = n,
+ .lost = lost,
+ .min_us = -1,
+ .median_us = -1,
+ .max_us = -1,
+ .median_settle_us = -1,
+ .median_bytes = 0,
+ .wire_percent = 0,
+ };
+
+ std.mem.sort(i64, rtt[0..got], {}, std.sort.asc(i64));
+ std.mem.sort(i64, settle[0..got], {}, std.sort.asc(i64));
+ std.mem.sort(usize, size[0..got], {}, std.sort.asc(usize));
+
+ // Bytes per second the wire can carry: baud/10, since each byte is 8N1 = 10 bit times.
+ const capacity = @as(i64, port.capacity());
+ return .{
+ .sent = n,
+ .lost = lost,
+ .min_us = rtt[0],
+ .median_us = rtt[got / 2],
+ .max_us = rtt[got - 1],
+ .median_settle_us = settle[got / 2],
+ .median_bytes = size[got / 2],
+ .wire_percent = @intCast(@divTrunc(@as(i64, @intCast(total_bytes)) * 1_000_000 * 100, elapsed * capacity)),
+ };
+}
+
+fn nowUs(port: *serial.Port) i64 {
+ return std.Io.Timestamp.now(port.io, .boot).toMicroseconds();
+}
diff --git a/tools/serial.zig b/tools/serial.zig
index 3398a11..7a740cd 100644
--- a/tools/serial.zig
+++ b/tools/serial.zig
@@ -81,10 +81,17 @@ pub const Port = struct {
file: std.Io.File,
io: std.Io,
saved: Termios2,
+ /// The rate currently programmed, kept because the wire's capacity in bytes per second is
+ /// `rate/10` and anything measuring this link against its ceiling needs that number. The
+ /// kernel would answer a TCGETS2, but a syscall per sample to re-read a value only this file
+ /// ever changes is worse than a field.
+ baud: Baud,
// std exposes neither these ioctl numbers nor the TIOCM bits.
const TIOCEXCL = 0x540C;
const TCFLSH = 0x540B;
+ /// tcdrain, with a nonzero argument. Zero would transmit a break instead.
+ const TCSBRK = 0x5409;
const TCIFLUSH = 0;
const TIOCMGET = 0x5415;
const TIOCMSET = 0x5418;
@@ -132,7 +139,7 @@ pub const Port = struct {
// Drop anything the kernel captured at the previous line rate.
_ = linux.ioctl(fd, TCFLSH, TCIFLUSH);
- return .{ .file = file, .io = io, .saved = saved };
+ return .{ .file = file, .io = io, .saved = saved, .baud = baud };
}
/// Re-rate an already-open port, leaving the raw-mode flags and the exclusive claim alone.
@@ -153,6 +160,25 @@ pub const Port = struct {
t.ospeed = baud.rate();
if (@as(isize, @bitCast(linux.ioctl(p.file.handle, Termios2.TCSETSW2, @intFromPtr(&t)))) < 0)
return error.SetAttrFailed;
+ p.baud = baud;
+ }
+
+ /// Bytes per second the wire can carry: one 8N1 byte occupies ten bit times.
+ pub fn capacity(p: *const Port) u32 {
+ return p.baud.rate() / 10;
+ }
+
+ /// Block until every byte written has physically left the wire - `tcdrain`, spelled as the
+ /// ioctl because std exposes neither.
+ ///
+ /// Distinct from `drain` above in both direction and meaning, which is worth stating because
+ /// getting them the wrong way round silently invalidates a measurement: `drain` discards what
+ /// has ARRIVED, this waits for what is LEAVING. A `write` returns once the kernel has accepted
+ /// the bytes, so timing a transfer to the write measures a memcpy into a tty buffer - at 115200
+ /// that reported 202% of the wire's capacity, which is how the confusion was noticed.
+ pub fn flushOutput(p: *Port) void {
+ // TCSBRK with a nonzero argument is tcdrain on Linux; with zero it would send a break.
+ _ = linux.ioctl(p.file.handle, TCSBRK, 1);
}
pub fn close(p: *Port) void {