From 1cef9c2e4bd873ebe13f5df635899231bcc467d2 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Tue, 25 Aug 2026 18:41:18 -0300 Subject: A measuring instrument, and what it says about where the latency goes "Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result. --- src/pardes/app.zig | 46 +--------------------------------------------- src/soc.zig | 45 +++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 46 insertions(+), 45 deletions(-) (limited to 'src') diff --git a/src/pardes/app.zig b/src/pardes/app.zig index 9fbe19e..a83485a 100644 --- a/src/pardes/app.zig +++ b/src/pardes/app.zig @@ -163,56 +163,12 @@ fn nowMs() u64 { return us / 1000; } -// ------------------------------------------------------------------------- the cache, flushed - -/// Evict every flash-backed cache line, by reading more flash than the caches can hold. -/// -/// This is a workaround for a real defect in the hand-over, and it is worth writing down exactly -/// what was measured, because everything cheaper was tried first and every one of them said the -/// hardware was fine: -/// -/// * The MMU table is correct. Entries 0..9 read 0x1001..0x100a - the valid bit plus physical -/// page N+1 - which is precisely what the image builder's single flash-to-vaddr anchor requires, -/// and entries 10..11 are unmapped as they should be. -/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the image -/// matches the ELF byte for byte at the addresses that misread. -/// * The page size is not in question: it is hardwired to 64 KiB on this chip -/// (hal/esp32p4/mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it). -/// -/// And yet a load at 0x40035a1c returned `93 85 85 0f`, which is this image's own `.text`. Reading -/// 512 KiB to force capacity eviction made the same load return `3c ee 08 40`, which is what the -/// image holds there. So the second-stage bootloader hands over with cache lines that do not match -/// the mapping it finally installed. It is perfectly deterministic - the same lines every boot, -/// because the bootloader does the same thing every boot - which is exactly why it looked like -/// anything other than a cache for so long. -/// -/// The ROM's own `Cache_Invalidate_All` (0x4fc00404, same address in both esp32p4.rom.ld and the -/// eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM -/// code with the argument stranded in a2, so it wants a precondition this image does not know about. -/// A capacity flush needs no such knowledge. It costs one pass over 512 KiB of already-mapped flash, -/// once, at boot. -/// -/// 512 KiB is four times the 128 KiB the L2 measured at (examples/memprobe.zig found real RAM -/// stopping at 0x4FFA0000, the cache taking the rest), with the L1s smaller still. The stride is one -/// 64-byte line. `volatile` and a summed sink so nothing here can be optimised away. -fn flushFlashCache() void { - var sink: u32 = 0; - var p: u32 = 0x4000_0000; - while (p < 0x4008_0000) : (p += 64) { - sink +%= @as(*volatile u32, @ptrFromInt(p)).*; - } - // Consumed through a volatile store so the whole loop cannot be discarded as dead. - @as(*volatile u32, &cache_flush_sink).* = sink; -} - -var cache_flush_sink: u32 = 0; - // ------------------------------------------------------------------------------------- the loop export fn zig_main() noreturn { // FIRST, before a single byte of `.rodata` is touched - which means before the marker below, // because that marker IS a string literal in flash and would read as machine code without this. - flushFlashCache(); + soc.flushFlashCache(); const heap = heapSpan(); soc.rom.print("\r\nMARK B3 rom.print heap 0x%08x..0x%08x %u KiB\r\n", .{ @as(u32, @intFromPtr(heap.ptr)), diff --git a/src/soc.zig b/src/soc.zig index 172966e..350948a 100644 --- a/src/soc.zig +++ b/src/soc.zig @@ -114,6 +114,51 @@ pub const rom = struct { } }; +/// Evict every flash-mapped cache line the bootloader left behind, by reading more flash than the +/// caches can hold. Call it before the first byte of `.rodata` is touched. +/// +/// This is a workaround for a real defect in the hand-over, not a tidiness measure. The second-stage +/// bootloader leaves lines cached against a mapping it then replaces, so an application reads its +/// own `.rodata` and gets its own `.text` back - **deterministically**, which is exactly what makes +/// it look like anything other than a cache. Measured on this die: a load at `0x40035A1C` returned +/// `93 85 85 0f` before eviction and `3c ee 08 40` after, and the second is what the image holds +/// there. A string literal read before this runs is machine code, so a firmware whose first act is +/// to print a marker prints garbage and looks like it never booted at all. +/// +/// Everything cheaper was tried first and every one of them said the hardware was fine, which is +/// why the list is here rather than being rediscovered: +/// +/// * The MMU table is correct. Entries 0..9 read `0x1001`..`0x100a` - the valid bit plus physical +/// page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires, +/// and 10..11 are unmapped as they should be. Read off the die through +/// `SPI_MEM_C_MMU_ITEM_INDEX_REG`, not inferred. +/// * The flash is correct. `zig build flash` verifies an MD5 of what the ROM stored, and the +/// image matches the ELF byte for byte at the addresses that misread. +/// * The page size is not in question: hardwired to 64 KiB on this chip +/// (`hal/esp32p4/mmu_ll.h:126-130` returns `MMU_PAGE_64KB` and the setter asserts it). +/// * Not fragmentation of the mapping either: the bad bytes arrive in one contiguous run of +/// >= 192 B, not in 64-byte lines, and they are identical across three resets and two +/// reflashes - determinism is what kept this looking like anything but a cache. +/// +/// 512 KiB is four times the 128 KiB the L2 measured at (`examples/memprobe.zig` found real RAM +/// stopping at `0x4FFA0000`, the cache taking the rest), with the L1s smaller still. +/// +/// `rom.Cache_Invalidate_All` is the instrument that ought to do this and does not: called from an +/// image the ROM did not launch, it faults inside the ROM with its argument stranded in `a2`. +/// Capacity eviction needs no preconditions, which is the whole reason it is what ships. 512 KiB at +/// a 64-byte stride is 8,192 loads, once per boot. +pub fn flushFlashCache() void { + var sink: u32 = 0; + var p: u32 = 0x4000_0000; + while (p < 0x4008_0000) : (p += 64) { + sink +%= @as(*volatile u32, @ptrFromInt(p)).*; + } + // Consumed through a volatile store, or the optimiser drops the whole loop as dead. + @as(*volatile u32, &flush_sink).* = sink; +} + +var flush_sink: u32 = 0; + /// Busy-wait for a number of CPU cycles, using the cycle counter rather than the mask ROM. Useful /// when an image must not depend on ROM entry points at all, and for delays shorter than the ROM's /// microsecond granularity. -- cgit v1.3