summaryrefslogtreecommitdiff
path: root/src/soc.zig
Commit message (Collapse)AuthorAge
* Run the CPU at 360 MHz: -Dcpu-mhz, and a keystroke lands at 4.37 msGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The board was executing at 90 MHz because the stock second-stage bootloader is built with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured here against the systimer, which is XTAL/2.5 and therefore an independent reference: 4,500,367 cycles in 50,004 us = exactly 90 MHz. The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse). `hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order - APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own comment says the other order passes through a state where APB or MEM violates its timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every `ets_delay_us` in the image is wrong by exactly the frequency ratio. Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only the divisor in experiments/ had to move. ## What it bought step fixed per char at 160 chars ReleaseSmall 16.99 ms 54.3 us 25.56 ms ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x + ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x + ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x + shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x + byte compare 8.37 ms 7.1 us 9.48 ms 0.37x + 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x Compute went 6.42 -> 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte, so that stage is bounded by L2MEM bandwidth and does not care how fast the core is. The prediction that this would happen was made before the measurement and held. ## The goal was 4 ms and this is 4.37 Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency that no firmware change touches (measured independently against the protocol responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large enough to close the gap alone, and it is redundant work - `present` already computes exactly which cells moved - so emitting ANSI straight from the shadow grid would do it. I did not, because it is a from-scratch renderer and the honest verification for it needs more than the harness currently proves. ## Verification, and a bug in my own instrument Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the A/B compares screens across clocks as well as across the shadow-grid flag. The first attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw SGR parameters applied to each cell, which is history-dependent, and a faster board splits the same keystrokes across different frames. Decoding SGR into actual state - resolved foreground, background and attribute set per cell - it is identical: same characters and same style everywhere, both across clocks and across the flag. Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked like it would quantise every sample. It does not - poll(2) returns when data arrives, not when the timeout expires - and switching to a non-blocking spin moved the measured round trip by 0 us. The comment now says so, since the next reader will wonder too. snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test, zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is opt-in because every number in experiments/ up to this commit was taken at 90.
* A measuring instrument, and what it says about where the latency goesGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | "Too slow for interactive use" is a real complaint and not a number. This adds the number, and the number says the wire is innocent. ## The instrument `tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so a frame one writes and the other parses cannot drift. It is imported as a module by both, not copied. The checksum is the whole point. RX overrun on this UART is undetected in hardware and uncounted in the driver, so a byte that never arrived is indistinguishable from a late one; a throughput figure that is not checksummed is a guess about how fast data was corrupted. `sink` accumulates a CRC over every payload byte the board received and `report` hands it back, so the host can prove that what arrived is what it sent. `tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in 8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time to the last byte is recorded separately as `settle`. Microseconds, because at 115200 one byte is 87 us and a millisecond clock quantises the answer into buckets eleven bytes wide. `tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how much of it the editor uses, `--sweep` for one controlled variable at a time with `--csv` raw per-trial output. ## What it measured The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of capacity in both directions, CRC verified over 32,768 B each way, zero corruption. Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so H5 - "typing loses input" - is refuted. Latency is compute per input event, not transmission. A 40-byte motion and a 206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a five-fold range of output. That is why raising the baud cannot fix typing: there is almost no wire in it. And an edit costs the whole document. Round trip against characters already in the line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far more than a copy alone, so the full-buffer copy the source does is accompanied by at least one more full pass. One controlled intervention: building the editor object ReleaseFast instead of ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document. Nothing else measured comes close to that ratio. ## Three bugs found while building it The responder printed garbage and looked dead: it read `.rodata` before evicting the bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to `soc.zig` with its measured evidence, since every application that touches `.rodata` after hand-over needs it and exactly one file knew that. Then it booted, printed its marker and went silent after ten seconds: `rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the application to take it over. Only the editor ever did. `serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported 202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named so the two cannot be confused again. Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every number in the first tables. ## The report `experiments/report.typ` reads the raw CSVs and computes its own figures, so a re-run changes the document instead of contradicting it. It states five hypotheses, settles each against one experiment, and is explicit about the one that failed: the geometry sweep is confounded, because characters accumulated across conditions and the length experiment then proved that matters. It is reported as unsupported rather than dressed up as a result.
* Flush the caches at startup: the bootloader hands over stale linesGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The image reads its own .rodata and gets its own .text back, and after ruling out everything cheaper the answer is the cache. What was eliminated first, each by measurement rather than argument: * The MMU table is CORRECT. Read from the running application through SPI_MEM_C_MMU_ITEM_INDEX_REG/CONTENT_REG (hal/esp32p4/mmu_ll.h:311-330), entries 0..9 hold 0x1001..0x100a - the valid bit plus physical page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires, and entries 10..11 are unmapped as they should be. * The page size is not in question: hardwired to 64 KiB on this chip (mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it), which is what tools/image.zig already assumed. * The flash is correct. The flasher verifies an MD5 of what the ROM stored, and app.bin matches the ELF byte for byte at the addresses that misread. * Not a write failure and not nondeterminism: identical across three resets and two reflashes with the same MD5. * Not 64-byte cache-line granularity either: the wrong bytes come in a contiguous run of at least 192. The measurement that settles it: a load at 0x40035a1c returned 93 85 85 0f, and reading 512 KiB to force capacity eviction made the SAME load return 3c ee 08 40, which is what the image holds there. So the second-stage bootloader hands over with cache lines that do not match the mapping it finally installed. It is perfectly deterministic - the bootloader does the same thing every boot, so it leaves the same lines - which is precisely why it looked like anything other than a cache for so long. The ROM's own Cache_Invalidate_All (0x4fc00404, the same address in both esp32p4.rom.ld and the eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM code with its argument stranded in a2, so it wants a precondition this image does not know. A capacity flush needs no such knowledge, costs one pass over 512 KiB of already-mapped flash once at boot, and is four times the 128 KiB the L2 measured at. Effect: the firmware now gets through the editor's allocator round-trip and pardes.allocators.init, which is two steps further than before. Also here, and correct independently of any of the above: uart.write and writeByte no longer spin forever on a stalled transmitter. hal/uart.zig:182-186 already made this point about update() - "on a board with no debugger an infinite spin is indistinguishable from a crash" - and this port proved it by spending an afternoon reading a stalled console as a hang in whatever code came next. The wait is bounded and abandoned bytes are counted. Still open: the console stops immediately after pardes.allocators.init. Bounded writes did not change it, so it is not the transmit spin; there is no Guru Meditation, so it is not a trap the ROM can report. The p4 allocator tier's zero-capacity StackFallbackAllocators are the one unusual thing in that call and their reasoning against lib/std/heap.zig is written down in src/allocators.zig, but it has not been tested with a nonzero floor. The bisect markers are left in place for that.
* zig-p4: pure-Zig ESP32-P4 toolchainGabriel Schneider2026-08-25
build.zig generates the linker script and drives Zig's own LLD; tools/image.zig turns the ELF into a flashable image and tools/{rom,serial}.zig speak the mask ROM loader over the UART. No CMake, ninja, idf.py, esptool, or external linker. src/soc.zig is a comptime register model over ESP-IDF's own *_reg.h headers; src/hal/ adds peripheral sequences; src/io/ implements std.Io for the chip; src/oracle/ diffs this HAL against ESP-IDF's on the die.