diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 21:11:24 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 21:11:24 -0300 |
| commit | aa524eae414f4875794638df0586c011c6b0f2c5 (patch) | |
| tree | 2a1a7e1794195f7864b69b1ae838b67c9f33b98e /build.zig | |
| parent | 7664c57d93cd3f3bdb719f7ca915d657668ddfd9 (diff) | |
| download | esp32p4-aa524eae414f4875794638df0586c011c6b0f2c5.tar.gz esp32p4-aa524eae414f4875794638df0586c011c6b0f2c5.zip | |
Run the CPU at 360 MHz: -Dcpu-mhz, and a keystroke lands at 4.37 ms
The board was executing at 90 MHz because the stock second-stage bootloader is built
with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured
here against the systimer, which is XTAL/2.5 and therefore an independent reference:
4,500,367 cycles in 50,004 us = exactly 90 MHz.
The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing
else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage
step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse).
`hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order -
APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own
comment says the other order passes through a state where APB or MEM violates its
timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every
`ets_delay_us` in the image is wrong by exactly the frequency ratio.
Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from
a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the
systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from
the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only
the divisor in experiments/ had to move.
## What it bought
step fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
+ byte compare 8.37 ms 7.1 us 9.48 ms 0.37x
+ 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x
Compute went 6.42 -> 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the
point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte,
so that stage is bounded by L2MEM bandwidth and does not care how fast the core is.
The prediction that this would happen was made before the measurement and held.
## The goal was 4 ms and this is 4.37
Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency
that no firmware change touches (measured independently against the protocol
responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's
Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large
enough to close the gap alone, and it is redundant work - `present` already computes
exactly which cells moved - so emitting ANSI straight from the shadow grid would do it.
I did not, because it is a from-scratch renderer and the honest verification for it
needs more than the harness currently proves.
## Verification, and a bug in my own instrument
Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the
A/B compares screens across clocks as well as across the shadow-grid flag. The first
attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw
SGR parameters applied to each cell, which is history-dependent, and a faster board
splits the same keystrokes across different frames. Decoding SGR into actual state -
resolved foreground, background and attribute set per cell - it is identical: same
characters and same style everywhere, both across clocks and across the flag.
Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked
like it would quantise every sample. It does not - poll(2) returns when data arrives,
not when the timeout expires - and switching to a non-blocking spin moved the measured
round trip by 0 us. The comment now says so, since the next reader will wonder too.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is
opt-in because every number in experiments/ up to this commit was taken at 90.
Diffstat (limited to 'build.zig')
| -rw-r--r-- | build.zig | 11 |
1 files changed, 11 insertions, 0 deletions
@@ -108,6 +108,14 @@ pub fn build(b: *std.Build) void { // cycle counts. Off by default because it puts a line on the wire per frame, which is the very // resource being measured - it answers "where did the 34 ms go", not "how fast is it". options.addOption(bool, "prof", b.option(bool, "prof", "print per-phase cycle counts (pardes)") orelse false); + // The CPU clock, in MHz. The bootloader leaves 90; the CPLL is already at 360, so 180 and 360 + // are a divider change away and nothing else (see hal/clkrst.zig:setCpuFreq). Opt-in rather + // than default because it is the one setting here that changes how every other timing in the + // image behaves, and because 90 is what every measurement in experiments/ was taken against. + const cpu_mhz = b.option(u16, "cpu-mhz", "HP CPU clock: 90 (bootloader default), 180 or 360") orelse 90; + if (cpu_mhz != 90 and cpu_mhz != 180 and cpu_mhz != 360) + std.debug.panic("-Dcpu-mhz must be 90, 180 or 360; the P4's CPLL divides 360 by 4, 2 or 1", .{}); + options.addOption(u16, "cpu_mhz", cpu_mhz); options.addOption([]const u8, "wifi_ssid", wifi_ssid); options.addOption([]const u8, "wifi_psk", if (psk_file) |path| blk: { const raw = std.Io.Dir.cwd().readFileAlloc(b.graph.io, path, b.allocator, .limited(256)) catch @@ -1216,6 +1224,9 @@ fn linkerScript(b: *std.Build, stack_size: u32, peripherals_ld: ?[]const u8) []c \\/* ESP32-P4 mask ROM entry points, the only "library" this image links against */ \\ets_printf = 0x4fc00024; \\ets_delay_us = 0x4fc0003c; + \\/* Recalibrate the ROM's cycles-per-microsecond after a CPU clock change; without it every + \\ ets_delay_us is wrong by exactly the frequency ratio (esp32p4.rom.ld:32). */ + \\ets_update_cpu_frequency = 0x4fc00044; \\/* Invalidate the caches. The bootloader leaves lines that do not match the mapping it \\ finally installs, so a large image reads its own .rodata and gets its own .text back. */ \\Cache_Invalidate_All = 0x4fc00404; |
