From 1cef9c2e4bd873ebe13f5df635899231bcc467d2 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Tue, 25 Aug 2026 18:41:18 -0300 Subject: A measuring instrument, and what it says about where the latency goes "Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result. --- build.zig | 48 ++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 48 insertions(+) (limited to 'build.zig') diff --git a/build.zig b/build.zig index e605808..e02b2d6 100644 --- a/build.zig +++ b/build.zig @@ -180,6 +180,14 @@ pub fn build(b: *std.Build) void { .{ .name = "hal", .module = hal_mod }, .{ .name = "mmio", .module = mmio_mod }, .{ .name = "regs", .module = regs_mod }, + // The wire protocol `examples/uartperf.zig` answers, imported rather than copied so + // the firmware and the host tool cannot disagree about a frame. It is deliberately + // free of any OS dependency for exactly this reason: one file, two targets. + .{ .name = "perfproto", .module = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = target, + .optimize = optimize, + }) }, }, }), }); @@ -397,6 +405,32 @@ pub fn build(b: *std.Build) void { run_con.step.dependOn(&flash.step); b.step("interact", "flash the image, then attach a terminal").dependOn(&run_con.step); + // The measuring instrument. Its own binary for the same reason the console is: it drives the + // port for tens of seconds and must not have the build runner repainting a progress tree into + // the middle of a timed transfer. It shares `tools/perfproto.zig` with the firmware responder, + // so a frame the host writes and a frame the board parses cannot drift apart. + const bench_proto = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = b.graph.host, + .optimize = .ReleaseSafe, + }); + const bench_exe = b.addExecutable(.{ + .name = "p4-bench", + .root_module = b.createModule(.{ + .root_source_file = b.path("tools/bench_main.zig"), + .target = b.graph.host, + .optimize = .ReleaseSafe, + .imports = &.{.{ .name = "perfproto", .module = bench_proto }}, + }), + }); + const bench_install = b.addInstallArtifact(bench_exe, .{}); + const bench = b.addRunArtifact(bench_exe); + bench.addArgs(&.{ "--port", port_path }); + bench.stdio = .inherit; + bench.step.dependOn(&bench_install.step); + b.step("bench", "measure the serial link: verified throughput each way, and latency") + .dependOn(&bench.step); + const reset = ResetStep.create(b, port_path); b.step("reset", "reset the board and let the flashed application run").dependOn(&reset.step); @@ -456,6 +490,20 @@ pub fn build(b: *std.Build) void { }); test_step.dependOn(&b.addRunArtifact(console_tests).step); + // The measurement protocol. These are the tests that keep a throughput number honest: that a + // frame round-trips, that a short read is "incomplete" rather than "invalid", that a lost byte + // mid-stream changes the CRC, and that the pattern generator does not repeat on a 256-byte + // boundary - a plain counter would hash identically after losing exactly 256 bytes and the + // instrument would report a clean run over corrupt data. + const proto_tests = b.addTest(.{ + .root_module = b.createModule(.{ + .root_source_file = b.path("tools/perfproto.zig"), + .target = b.graph.host, + .optimize = .Debug, + }), + }); + test_step.dependOn(&b.addRunArtifact(proto_tests).step); + // The radio path's host-testable parts. A wrong checksum, a wrong snprintf, or a scheduler that // loses a task is far cheaper to find here than on a board whose only output is a serial line. // -- cgit v1.3