diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 18:41:18 -0300 |
| commit | 1cef9c2e4bd873ebe13f5df635899231bcc467d2 (patch) | |
| tree | ce0a9495bc801666ff14da47f7a75728335cc442 /README.md | |
| parent | ef6f3e460ddbf53a752807bcf10a0fec0a72da7f (diff) | |
| download | esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.tar.gz esp32p4-1cef9c2e4bd873ebe13f5df635899231bcc467d2.zip | |
A measuring instrument, and what it says about where the latency goes
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.
## The instrument
`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.
The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.
`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.
`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.
## What it measured
The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.
Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.
And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.
One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.
## Three bugs found while building it
The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.
Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.
`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.
Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.
## The report
`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
Diffstat (limited to 'README.md')
| -rw-r--r-- | README.md | 46 |
1 files changed, 46 insertions, 0 deletions
@@ -9,6 +9,7 @@ zig build console # attach a terminal to whatever is already on the board (Ct zig build interact # flash, then attach that terminal (ordered, like `run`) zig build reset # just pulse the reset line zig build size # where every byte of the image went +zig build bench # measure the serial link and the editor, verified with a checksum zig build test # host tests: image builder, and the register layer's field arithmetic zig build diff # the hardware oracle: this HAL vs ESP-IDF's, on the die (needs -Doracle) zig build elf # stop at the ELF, for disassembly @@ -121,6 +122,51 @@ Measured on ESP32-P4 rev v1.3 silicon: two `Peek`s of the RNG register at `0x501 `Peek 0x50110001` answered `peek: MisalignedAddress` on the message row rather than taking the session down with an unhandled trap, which is the one fault that file exists to prevent. +### Measuring it + +`p4-bench` exists so that optimising this port is not a matter of opinion. It has two halves, +because there are two different questions. + +``` +zig build flash -Dapp=examples/uartperf.zig # the ceiling: link + driver, nothing else +zig-out/bin/p4-bench --link + +zig build flash -Dpardes # how much of that ceiling the editor uses +zig-out/bin/p4-bench --editor +zig-out/bin/p4-bench --sweep length --repeat 7 --csv --label ReleaseSmall +``` + +Every `--link` number is checksummed. `tools/perfproto.zig` is a framed protocol - `"P4"`, op, +length, CRC-32, payload - shared *verbatim* by the host tool and `examples/uartperf.zig`, so a frame +one writes and the other parses cannot drift. That matters because RX overrun on this UART is +undetected in hardware and uncounted in the driver: a byte that never arrived is indistinguishable +from a late one, and an unchecksummed throughput figure is a guess about how fast data was corrupted. + +`tools/rtt.zig` holds the two timing functions everything is built on. Round trip is measured to the +**first** response byte, not the last: a renderer that starts drawing in 8 ms and finishes in 130 ms +feels immediate, one that thinks for 130 ms then draws in 8 ms feels broken, and waiting for the wire +to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. +Microseconds throughout, because at 115200 one byte is 87 us and a millisecond clock would quantise +the answer into buckets eleven bytes wide. + +`experiments/` holds the raw per-trial CSVs and `report.typ`, which reads them and computes its own +figures - so a re-run changes the document rather than contradicting it. What it establishes on this +die: + +| | | +|---|---| +| link, both directions | 99.8% of the 11,520 B/s wire, CRC verified over 32,768 B | +| typing, 6 to 100 keys/s | nothing lost, wire never above 9% | +| one keystroke | 81 B, round trip 17.0 ms | +| a motion | 40 B, round trip 16.7 ms | +| cost per character already in the line | **54.3 us, per keystroke** | + +The first two rows say the wire is not the problem. The last three say why: a 40-byte operation and a +206-byte one cost the same round trip, so latency is compute per event and not transmission - and it +grows with the document, because the edit path copies the whole buffer every keystroke. Building the +editor object `ReleaseFast` instead of `ReleaseSmall` cuts the fixed cost 13% and the per-character +cost 36% for 35% more flash, which is the best ratio measured here. + ## Layout ``` |
