diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
| commit | 1cef9c2e4bd873ebe13f5df635899231bcc467d2 (patch) | |
| tree | ce0a9495bc801666ff14da47f7a75728335cc442 /tools/perfproto.zig | |
| parent | ef6f3e460ddbf53a752807bcf10a0fec0a72da7f (diff) | |
| download | esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.tar.gz esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.zip | |
A measuring instrument, and what it says about where the latency goes
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.
## The instrument
`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.
The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.
`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.
`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.
## What it measured
The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.
Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.
And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.
One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.
## Three bugs found while building it
The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.
Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.
`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.
Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.
## The report
`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
Diffstat (limited to 'tools/perfproto.zig')
| -rw-r--r-- | tools/perfproto.zig | 201 |
1 files changed, 201 insertions, 0 deletions
diff --git a/tools/perfproto.zig b/tools/perfproto.zig new file mode 100644 index 0000000..5223500 --- /dev/null +++ b/tools/perfproto.zig @@ -0,0 +1,201 @@ +//! A small framed protocol for measuring the board's serial link, shared verbatim by the host tool +//! and the firmware that answers it. +//! +//! WHY A PROTOCOL AND NOT A STOPWATCH. Timing an editor's keystrokes measures the editor, the +//! renderer and the link at once, and cannot tell a dropped byte from a slow one: RX overrun on this +//! UART is silent in hardware and uncounted in the driver, so a missing keystroke and a late one look +//! identical from the host. A frame with a length and a checksum turns both into facts. If the CRC +//! matches, every byte of that payload crossed intact; if a frame never completes, bytes were lost +//! and the count says how many. A throughput number that is not checksummed is a guess about how +//! fast data was corrupted. +//! +//! THE SHAPE. One fixed 9-byte header, little-endian, then the payload: +//! +//! "P4" op:u8 len:u16 crc:u32 payload[len] +//! +//! The CRC covers the payload only. The header carries it rather than trailing it so a receiver +//! knows, before it has read a single payload byte, exactly how many to expect and what they must +//! hash to - which is what lets the firmware verify a stream with one 4-byte accumulator and no +//! buffer at all. +//! +//! `max_payload` is 1024 and that is a memory decision, not a wire one. The firmware has a 384 KiB +//! heap it must share with an editor, and a bulk test that needed a 64 KiB frame buffer would be +//! measuring a configuration nobody ships. Bulk transfers are therefore many frames, which is also +//! the honest shape: it is the per-frame overhead a real protocol would pay. +//! +//! Both directions use the same header, and a reply's op has the high bit set, so a stray reply can +//! never be mistaken for a request by a resynchronising receiver. + +const std = @import("std"); + +pub const magic = "P4"; +pub const header_len = 9; +pub const max_payload = 1024; + +pub const Op = enum(u8) { + /// Echo the payload back as `pong`. Both directions verified in one exchange, which is what + /// makes it the right stimulus for a latency measurement. + ping = 1, + /// Payload is data to be consumed. The board accumulates a running count and CRC and answers + /// nothing, so the host can keep the uplink full and measure it without return traffic + /// competing for the same wire. + sink = 2, + /// Ask for the accumulated `sink` count and CRC, then reset them. + report = 3, + /// Payload is a u32 count: send exactly that many pattern bytes back, in `data` frames, + /// followed by a `stat`. + source = 4, + + pong = 0x81, + /// Payload is `Stat`, packed little-endian. + stat = 0x83, + /// A chunk of `source` output. + data = 0x84, + + pub fn isReply(o: Op) bool { + return @intFromEnum(o) & 0x80 != 0; + } +}; + +/// What the board reports about a stream it received or sent. Encoded by hand rather than by +/// `@bitCast` of a packed struct: this crosses between a riscv32 firmware and an x86_64 host, and a +/// layout that depends on either compiler's padding rules is a bug waiting for a target change. +pub const Stat = struct { + /// Payload bytes accumulated. + bytes: u32, + /// CRC-32 over exactly those bytes, in order. + crc: u32, + /// Frames whose CRC did not match. Nonzero means the link corrupted data rather than losing it, + /// which is a different fault with a different fix. + bad_frames: u32, + /// Bytes the firmware's UART driver gave up on writing. Its own counter, surfaced here because + /// the host cannot see it any other way. + tx_dropped: u32, + + pub const encoded_len = 16; + + pub fn encode(s: Stat, out: *[encoded_len]u8) void { + std.mem.writeInt(u32, out[0..4], s.bytes, .little); + std.mem.writeInt(u32, out[4..8], s.crc, .little); + std.mem.writeInt(u32, out[8..12], s.bad_frames, .little); + std.mem.writeInt(u32, out[12..16], s.tx_dropped, .little); + } + + pub fn decode(in: []const u8) ?Stat { + if (in.len < encoded_len) return null; + return .{ + .bytes = std.mem.readInt(u32, in[0..4], .little), + .crc = std.mem.readInt(u32, in[4..8], .little), + .bad_frames = std.mem.readInt(u32, in[8..12], .little), + .tx_dropped = std.mem.readInt(u32, in[12..16], .little), + }; + } +}; + +pub fn crc(bytes: []const u8) u32 { + return std.hash.Crc32.hash(bytes); +} + +/// The deterministic byte at stream offset `i`. +/// +/// A counter would be checksummed correctly by an implementation that lost exactly 256 bytes, and a +/// constant by one that lost any amount. This is an 8-bit xorshift-ish walk whose period is long +/// enough that no realistic loss aligns with it, so the CRC catches a gap wherever it falls. +pub fn patternByte(i: u32) u8 { + var x: u32 = i +% 1; + x ^= x << 7; + x ^= x >> 3; + x ^= x << 5; + return @truncate(x); +} + +pub fn fillPattern(buf: []u8, offset: u32) void { + for (buf, 0..) |*b, k| b.* = patternByte(offset +% @as(u32, @intCast(k))); +} + +/// Write a frame into `out`, returning the used slice. `out` must hold `header_len + payload.len`. +pub fn encode(out: []u8, op: Op, payload: []const u8) []u8 { + std.debug.assert(payload.len <= max_payload); + std.debug.assert(out.len >= header_len + payload.len); + out[0] = magic[0]; + out[1] = magic[1]; + out[2] = @intFromEnum(op); + std.mem.writeInt(u16, out[3..5], @intCast(payload.len), .little); + std.mem.writeInt(u32, out[5..9], crc(payload), .little); + @memcpy(out[header_len..][0..payload.len], payload); + return out[0 .. header_len + payload.len]; +} + +pub const Header = struct { + op: Op, + len: u16, + crc: u32, +}; + +/// Read a header out of `buf`. Returns null when fewer than `header_len` bytes are present, and +/// `error.BadFrame` when the magic or the op is not one of ours - which is how a receiver that has +/// lost sync tells "wait for more" from "throw a byte away and try again". +pub fn parseHeader(buf: []const u8) error{BadFrame}!?Header { + if (buf.len < header_len) return null; + if (buf[0] != magic[0] or buf[1] != magic[1]) return error.BadFrame; + const op = std.enums.fromInt(Op, buf[2]) orelse return error.BadFrame; + const len = std.mem.readInt(u16, buf[3..5], .little); + if (len > max_payload) return error.BadFrame; + return .{ .op = op, .len = len, .crc = std.mem.readInt(u32, buf[5..9], .little) }; +} + +test "a frame round-trips through encode and parseHeader" { + var buf: [header_len + 4]u8 = undefined; + const f = encode(&buf, .ping, "abcd"); + try std.testing.expectEqual(@as(usize, header_len + 4), f.len); + const h = (try parseHeader(f)).?; + try std.testing.expectEqual(Op.ping, h.op); + try std.testing.expectEqual(@as(u16, 4), h.len); + try std.testing.expectEqual(crc("abcd"), h.crc); + try std.testing.expectEqualStrings("abcd", f[header_len..]); +} + +test "a short buffer is incomplete, not invalid" { + var buf: [header_len]u8 = undefined; + const f = encode(&buf, .report, ""); + try std.testing.expectEqual(@as(?Header, null), try parseHeader(f[0 .. header_len - 1])); +} + +test "wrong magic and unknown ops are rejected rather than misread" { + var buf: [header_len]u8 = undefined; + var f = encode(&buf, .report, ""); + f[0] = 'X'; + try std.testing.expectError(error.BadFrame, parseHeader(f)); + f[0] = magic[0]; + f[2] = 0x7f; + try std.testing.expectError(error.BadFrame, parseHeader(f)); +} + +test "a truncated stream is caught by the CRC" { + // Losing bytes is the failure this protocol exists to detect, so prove the checksum notices a + // gap that leaves the length plausible. + var full: [64]u8 = undefined; + fillPattern(&full, 0); + var gapped: [64]u8 = undefined; + fillPattern(gapped[0..32], 0); + fillPattern(gapped[32..], 33); // one byte skipped mid-stream + try std.testing.expect(crc(&full) != crc(&gapped)); +} + +test "the pattern does not repeat inside a byte-aligned loss" { + // A plain counter would hash identically after losing exactly 256 bytes. This must not. + var a: [128]u8 = undefined; + var b: [128]u8 = undefined; + fillPattern(&a, 0); + fillPattern(&b, 256); + try std.testing.expect(crc(&a) != crc(&b)); +} + +test "Stat survives the trip between a riscv32 firmware and an x86_64 host" { + const s: Stat = .{ .bytes = 0x11223344, .crc = 0xdeadbeef, .bad_frames = 7, .tx_dropped = 9 }; + var buf: [Stat.encoded_len]u8 = undefined; + s.encode(&buf); + const back = Stat.decode(&buf).?; + try std.testing.expectEqual(s, back); + try std.testing.expectEqual(@as(?Stat, null), Stat.decode(buf[0 .. Stat.encoded_len - 1])); +} |
