diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-25 13:06:05 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-25 14:38:25 -0300 |
| commit | 174991b8f3f8e9c792eede7a52ad7beb10a08b05 (patch) | |
| tree | 3dcc42604752251227e233294d956e18cb264ad8 /examples | |
| parent | f5f8068fac59b4f16046c2022c2fc7c7e447ef4c (diff) | |
| download | esp32p4-174991b8f3f8e9c792eede7a52ad7beb10a08b05.tar.gz esp32p4-174991b8f3f8e9c792eede7a52ad7beb10a08b05.zip | |
pardes as P4 firmware: the seam, and a flash-mapping bug in this toolchain
The editor arrives as one freestanding OBJECT exporting a seven-function C ABI
(src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not
as a package dependency. A build.zig.zon path dependency was built first and
reverted: merely DECLARING it nested pardes's ~30-package graph under this one and
broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime
quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed
in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy.
The seam is bytes in and bytes out, which is what a serial line is anyway: the
editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the
clock, and neither names the other's types. It is versioned, because linkers do
not type-check C symbols and a drifted signature would link cleanly and then
corrupt the stack.
THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's
anchor makes two mapped segments share an MMU page safely - as long as rodata does
not END inside the page where text BEGINS. With a 578 KB image it does. A volatile
read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to
"lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared
page read as code, so the first thing the firmware tried to print was machine code
and it died on an instruction access fault. .flash.text is now ALIGN(0x10000),
making the segments page-disjoint. The packing trick this project opened with only
ever mattered when the alternative was 64 KiB of zeros in a 1 KB image.
Two more findings, both recorded in README.md:
* A linker symbol declared as an anyopaque OBJECT gives the optimiser a
zero-sized object, so ordinary stores through a pointer derived from its
address are dead code it may drop - and did, silently. The allocator's first
block header read back as size=2988759312 next=0x14284684 and the free-list
walk never terminated. @extern with a many-pointer has no size to lose.
examples/memprobe.zig could not have caught it: it writes through a volatile
pointer, which the optimiser must leave alone.
* The RTC watchdog is armed at handover. Every example here had been resetting on
a ten-second cycle, invisibly, because no run had ever lasted eight seconds.
State, honestly: the firmware boots, clears .bss, brings up the console, disables
the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB
heap and calls into the editor, which sets up its sink and its environment. It then
faults inside pardes_p4_init on the first allocation. The cause is measured but not
fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000
higher - exactly the vaddr distance between the rodata and text segments - while
pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in
place, deliberately, because the next step needs them.
--- correction, measured after the above was written ---
Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig
was right for a reason I initially got wrong and then measured.
I first read bootloader_utility.c's `#else` branch, which classifies segments by
address window with two independent ifs - and since the P4's DROM and IROM windows
are the identical range (soc.h:146-149), I concluded the last mapped segment wins
both roles and the first is never mapped. That branch does not run on this chip.
The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851),
whose own comment says it: "On chips with shared D/I external vaddr, we don't
divide them into either D or I, as essentially they are the same." It collects
mapped segments POSITIONALLY into rom_addr[2] and ends with
assert(rom_index == 2);
Shipping a one-segment image proved it, on the board:
Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2)
So the split stays, image.zig keeps enforcing exactly two - turning that boot-time
abort into a build-time error - and both are now documented with the branch that
actually runs and the assert that actually fires.
What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may
share a 64 KiB MMU page only if they also share a flash page, which the image
builder's anchor guarantees - and that holds right up until an application is large
enough for rodata to END inside the page where text BEGINS. With a 578 KB image it
does. Measured on the die: a volatile read of a string literal at 0x4004a1d1
returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own
.flash.text. Every literal in that shared page read as code, so the first thing the
firmware tried to print was machine code, and it died on an instruction access
fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint.
It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing
trick this project opened with only mattered when the alternative was 64 KiB of
zeros in a 1 KB image.
With that fixed the firmware gets much further: entry, .bss cleared, console up,
watchdog disabled, systimer running, ABI version checked, the 384 KiB heap
initialised, into the editor, its sink and environment ready - and the literal at
0x4004a1d1 now reads back correctly.
Still open, and characterised rather than guessed: pardes_p4_init faults on its
first allocation. The allocator struct crosses the seam intact (its function
pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads
back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it.
Ruled out with measurements: the ELF and the image agree at that address, the flash
is MD5-verified against the image, the wrong bytes are identical across three
resets and two reflashes (so not a stale cache), the corruption is a contiguous run
rather than 64-byte lines, and mmu_hal_map_region's arithmetic
(page_num = ceil(len/page), entry from vaddr) is correct for the segments as now
laid out. The next measurement is the one that settles it: read the MMU entry
registers from the running application and print vaddr -> flash for every page. The
register model in src/soc.zig can do that; the bisect markers are left in place for
it.
Diffstat (limited to 'examples')
| -rw-r--r-- | examples/echo.zig | 146 | ||||
| -rw-r--r-- | examples/heapcheck.zig | 240 | ||||
| -rw-r--r-- | examples/memprobe.zig | 215 | ||||
| -rw-r--r-- | examples/minimal.zig | 2 |
4 files changed, 602 insertions, 1 deletions
diff --git a/examples/echo.zig b/examples/echo.zig new file mode 100644 index 0000000..6f17c98 --- /dev/null +++ b/examples/echo.zig @@ -0,0 +1,146 @@ +//! UART0 as a duplex byte pipe, which is the one thing this toolchain had never done. +//! +//! Everything else here talks to the host through `soc.rom.print`, a mask-ROM `ets_printf`. That is +//! one-way and it is slow: the ROM formats, then pushes a byte at a time and spins on the FIFO. An +//! editor rendering a screen needs the other direction and needs the fast path, so this example +//! exists to prove three things on the die before anything larger depends on them: +//! +//! 1. **RX works at all.** `hal/uart.zig` has had `rxCount`/`popByte` since the differential +//! suite needed them, but that suite runs on UART1 in internal loopback - no byte has ever +//! arrived from the outside world on UART0. +//! 2. **UART0 can be driven without reconfiguring it.** The header of `hal/uart.zig` is blunt +//! about the hazard: resetting UART0 clears UART_CLKDIV, the console turns to garbage +//! mid-sentence and the board dies on a watchdog reset. So this touches no configuration +//! register - the second-stage bootloader already set the divider, the format and the pad +//! routing, and this code only reads and writes the FIFO. +//! 3. **What the wire rate actually is.** Printed by asking the hardware +//! (`Uart.baudrate`), not by assuming the 115200 the host tooling opens with. +//! +//! Protocol, so the host side has something unambiguous to assert on: +//! +//! any byte -> echoed back verbatim +//! CR (0x0d) -> echoed as CRLF, so a human sees lines +//! '!' -> also emit `bulk_len` bytes of a counted pattern and report the cycles it took +//! Ctrl-D -> print the byte/frame counters +//! +//! The echo is verbatim rather than uppercased or otherwise transformed on purpose: a transform +//! that happens to be idempotent hides a duplicated byte, and a duplicated byte is exactly the +//! failure a FIFO-polling loop produces when `rxCount` is misread. + +const std = @import("std"); +const soc = @import("soc"); +const hal = @import("hal"); + +/// UART0. Instance 0 because that is the pad pair the CH340 is wired to and the one the ROM +/// configured; nothing here may reset it. +const con = hal.uart.Uart.init(0); + +/// The bulk burst `!` emits. 4 KiB is ~35 ms of wire time at 115200 and ~4.5 ms at 921600, so the +/// difference between the two is obvious to the naked eye on the host. +const bulk_len = 4096; + +/// The clock the UART's baud generator is dividing. XTAL is the reset default and what the ROM +/// leaves selected; `Uart.clockSource` is read below rather than assumed, so a bootloader that +/// switched to PLL_F80M shows up as a wrong rate instead of a silent 2x error. +fn sourceHz() u32 { + return con.clockSource().nominalHz(); +} + +/// Block until the TX FIFO has room, then push. The spin is bounded by the wire: at 115200 a full +/// 128-byte FIFO drains in 11 ms, and there is nothing else for this core to do. +/// +/// `txFree` and not "is the FIFO empty": pushing whenever there is a single free slot keeps the +/// transmitter fed, which is what makes this ~10x the throughput of the ROM's per-byte printf. +fn put(byte: u8) void { + while (con.txFree() == 0) {} + con.pushByte(byte); +} + +fn puts(bytes: []const u8) void { + for (bytes) |b| put(b); +} + +export fn zig_main() noreturn { + // Deliberately the ROM path for the banner: if the direct-FIFO writes below are wrong, the + // banner still arrives and says so. Mixing the two is safe because both end up in the same + // FIFO and this is the only writer. + soc.rom.print("\r\nMARK ECHO_BOOT uart0 duplex echo\r\n", .{}); + soc.rom.print("MARK ECHO_BAUD hw=%u src=%u Hz\r\n", .{ con.baudrate(sourceHz()), sourceHz() }); + + // Not `resetRxFifo`: that is a CONF0_SYNC read-modify-write plus two commits on the console + // UART, and this file's whole premise is that UART0's configuration is untouchable. Draining by + // popping has the same effect on the FIFO and touches only offset 0x000. + var dropped: u32 = 0; + while (con.rxCount() > 0) : (dropped += 1) _ = con.popByte(); + soc.rom.print("MARK ECHO_DRAIN dropped=%u stale bytes\r\n", .{dropped}); + puts("MARK ECHO_FIFO direct-fifo tx works\r\n"); + puts("type; '!' bulk, ctrl-D stats\r\n"); + + var rx_total: u32 = 0; + var bursts: u32 = 0; + while (true) { + if (con.rxCount() == 0) continue; + const byte = con.popByte(); + rx_total += 1; + + switch (byte) { + '\r' => puts("\r\n"), + 0x04 => { + var buf: [96]u8 = undefined; + const line = std.fmt.bufPrint( + &buf, + "\r\nMARK ECHO_STATS rx={d} bursts={d} baud={d}\r\n", + .{ rx_total, bursts, con.baudrate(sourceHz()) }, + ) catch "\r\nMARK ECHO_STATS fmt failed\r\n"; + puts(line); + }, + '!' => { + bursts += 1; + put(byte); + const t0 = soc.cycles(); + // A counted pattern, not a constant: a run of identical bytes cannot reveal a + // dropped or reordered one, and the host asserts on the sequence. + var i: u32 = 0; + while (i < bulk_len) : (i += 1) put('0' + @as(u8, @intCast(i % 10))); + const cycles = soc.cycles() - t0; + var buf: [96]u8 = undefined; + const line = std.fmt.bufPrint( + &buf, + "\r\nMARK ECHO_BULK {d} bytes in {d} cycles\r\n", + .{ bulk_len, cycles }, + ) catch "\r\nMARK ECHO_BULK fmt failed\r\n"; + puts(line); + }, + else => put(byte), + } + } +} + +/// Reset entry. Identical in shape to `src/main.zig`'s and for the same reasons - the bootloader +/// hands over with an unspecified stack pointer and the FPU off - but `std.fmt.bufPrint` above is +/// the reason the FPU bit matters here too: its float formatting path is reachable from a generic +/// `bufPrint` instantiation even when no argument is a float. +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} + +pub const panic = std.debug.FullPanic(struct { + fn call(msg: []const u8, _: ?usize) noreturn { + soc.rom.print("MARK ECHO_PANIC %s\r\n", .{msg.ptr}); + while (true) {} + } +}.call); diff --git a/examples/heapcheck.zig b/examples/heapcheck.zig new file mode 100644 index 0000000..9d61f71 --- /dev/null +++ b/examples/heapcheck.zig @@ -0,0 +1,240 @@ +//! Does `src/net/heap.zig` survive an editor's allocation pattern in 512 KiB? +//! +//! `Heap` was written for ESP-Hosted and sized by it: its own header says the workload it is +//! "sized for is the one the parent measured - `mempool.c` recycling fixed-size buffers - which is +//! exactly the pattern that keeps a coalescing free list short and its first fit O(1) in practice". +//! An editor is not that workload. pardes allocates panes, text buffers, cell grids and per-frame +//! scratch in a dozen different sizes and frees them in an order nobody chose. Two properties that +//! were free under fixed-size recycling stop being free: +//! +//! * **First fit fragments.** Varied sizes leave gaps too small for the next request, and the +//! symptom is not a clean failure - it is `largest_free` collapsing while `free` stays healthy, +//! so the heap reports plenty of room and cannot satisfy a grid reallocation. +//! * **First fit is O(n) in the free list.** One long-lived allocation in the middle of the arena +//! splits it permanently, and every subsequent walk pays for it. +//! +//! So this measures both, on the die, before the editor depends on it. Each phase prints a line +//! whether it passed or not: a number is the deliverable, not a verdict. +//! +//! The interesting column is `largest/free`. At 1.00 the free space is one block and the heap is +//! pristine; as it falls, that fraction is the largest single allocation still possible. An editor +//! that cannot get one contiguous cell grid is dead regardless of how many bytes are notionally +//! free. + +const std = @import("std"); +const soc = @import("soc"); +const hal = @import("hal"); +const heapmod = @import("heap"); + +/// The upper RAM chunk, from the linker script - the same span the firmware gets. +/// +/// Reached with `@extern`, NOT with `extern const __heap_start: anyopaque` plus +/// `@intFromPtr`/`@ptrFromInt`. That spelling cost real debugging on this board. Declaring a linker +/// symbol as an `anyopaque` OBJECT gives the optimiser a zero-sized object to reason about, so a +/// pointer derived from its address carries provenance for zero bytes - and the ordinary +/// (non-volatile) store `Heap.init` makes through it was simply dropped. The symptom was the block +/// header reading back as `size=2988759312 next=0x14284684` instead of `{393216, 0xFFFFFFFF}`, after +/// which the first free-list walk followed garbage and never terminated. With asserts compiled out +/// in ReleaseSmall that is a silent hang, and `examples/memprobe.zig` could not see it because it +/// writes through a `volatile` pointer, which the optimiser must not touch. +/// +/// `@extern` with a `[*]u8` result has no size to lose. +const heap_start = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_start" }); +const heap_end = @extern([*]align(heapmod.Heap.granule) u8, .{ .name = "__heap_end" }); + +var gpa_heap: heapmod.Heap = undefined; + +fn span() []align(heapmod.Heap.granule) u8 { + return heap_start[0 .. @intFromPtr(heap_end) - @intFromPtr(heap_start)]; +} + +fn report(tag: [*:0]const u8) void { + const s = gpa_heap.stats(); + // Split across two calls, and no `%%`: the mask ROM's printf is size-optimised and this file's + // first version passed it six varargs plus a literal `%%`, which hung. Four is known to work + // (examples/memprobe.zig's range lines), so this stays inside what has been demonstrated. + soc.rom.print("MARK HEAP %s total=%u free=%u blocks=%u\r\n", .{ + tag, s.total, s.free, s.free_blocks, + }); + // `largest` is the number that matters under fragmentation: it is the largest single allocation + // still possible, whatever `free` claims. + soc.rom.print("MARK HEAP %s largest=%u\r\n", .{ tag, s.largest_free }); +} + +/// A deterministic LCG, so a bad run is reproducible. Numerical Recipes' constants. +var rng_state: u32 = 0x1234_5678; +fn rand() u32 { + rng_state = rng_state *% 1664525 +% 1013904223; + return rng_state; +} + +/// How many live pointers the phases below track. 512 slots at an average of ~512 B is ~256 KiB, +/// half the arena, which is enough to fragment it without trivially exhausting it. +const slots = 512; +var live: [slots][]u8 = undefined; +var live_len: usize = 0; + +export fn zig_main() noreturn { + const s = span(); + // The bootloader leaves the RTC watchdog running and expects the application to take it over. + // A bare image never did, so every demo in this repo has been resetting on a ten-second cycle + // (README.md:289-293) - invisible until a run lasted longer than eight seconds. The churn phase + // below does 20,000 allocations, so this is the first example here that would have hit it: the + // symptom was HEAP_BOOT printed twice and nothing after. + const was_armed = hal.rwdt.disable(); + soc.rom.print("MARK HEAP_RWDT was_armed=%u now_armed=%u\r\n", .{ + @as(u32, @intFromBool(was_armed)), @as(u32, @intFromBool(hal.rwdt.armed())), + }); + soc.rom.print("\r\nMARK HEAP_BOOT 0x%08x..0x%08x %u KiB\r\n", .{ + @as(u32, @intFromPtr(s.ptr)), + @as(u32, @intFromPtr(s.ptr)) + @as(u32, @intCast(s.len)), + @as(u32, @intCast(s.len / 1024)), + }); + gpa_heap = heapmod.Heap.init(s); + soc.rom.print("MARK HEAP_INIT done\r\n", .{}); + // Read the block header straight back through a volatile pointer. `stats()` walks the free list + // starting here, so if these two words are not {len, 0xFFFFFFFF} the walk follows garbage and + // never terminates - and with asserts compiled out in ReleaseSmall that is a silent hang. + { + const hdr: *volatile [2]u32 = @ptrFromInt(0x4FF4_0000); + soc.rom.print("MARK HEAP_HDR size=%u next=0x%08x\r\n", .{ hdr[0], hdr[1] }); + } + const a = gpa_heap.allocator(); + soc.rom.print("MARK HEAP_VTABLE done\r\n", .{}); + const probe_stats = gpa_heap.stats(); + soc.rom.print("MARK HEAP_STATS total=%u\r\n", .{probe_stats.total}); + report("fresh"); + + // ---- phase 1: how much of the arena is actually reachable, and what does the header cost? + // Allocate one block at a time until refusal, to separate "512 KiB of RAM" from "512 KiB of + // usable allocations". Each block costs an 8-byte header, so the answer is strictly less. + { + var n: u32 = 0; + var total: usize = 0; + while (n < slots) { + const blk = a.alloc(u8, 512) catch break; + live[n] = blk; + total += blk.len; + n += 1; + } + live_len = n; + soc.rom.print("MARK HEAP_FILL %u blocks of 512 B = %u B payload\r\n", .{ n, @as(u32, @intCast(total)) }); + report("filled"); + for (live[0..live_len]) |blk| a.free(blk); + live_len = 0; + // Coalescing is the whole design: after freeing everything the arena must be ONE block + // again. If it is not, `insert`'s merge is wrong and every later number is meaningless. + report("emptied"); + } + + // ---- phase 2: the fragmentation case. Fill with alternating sizes, free every other block. + // This is the adversarial pattern for first fit: the holes are all the smaller size, and the + // next larger request has to walk past every one of them. + { + var n: usize = 0; + while (n < slots) : (n += 1) { + const size: usize = if (n % 2 == 0) 192 else 320; + live[n] = a.alloc(u8, size) catch break; + } + live_len = n; + var i: usize = 0; + while (i < live_len) : (i += 2) a.free(live[i]); + report("holed"); + + // Now ask for something that fits in no single hole and see what the heap does. + if (a.alloc(u8, 4096)) |big| { + soc.rom.print("MARK HEAP_BIG 4096 B satisfied after holing\r\n", .{}); + a.free(big); + } else |_| { + soc.rom.print("MARK HEAP_BIG 4096 B REFUSED after holing\r\n", .{}); + } + i = 1; + while (i < live_len) : (i += 2) a.free(live[i]); + live_len = 0; + report("unholed"); + } + + // ---- phase 3: the O(n) first-fit walk, measured rather than argued. + // Build a long free list, then time one allocation that has to traverse it. `soc.cycles()` is + // the cycle counter, so this is in real CPU cycles. + { + var n: usize = 0; + while (n < slots) : (n += 1) live[n] = a.alloc(u8, 256) catch break; + live_len = n; + // Free every other block to make `free_blocks` large, then measure a request that no hole + // can satisfy, which is the worst case: the full walk. + var i: usize = 0; + while (i < live_len) : (i += 2) a.free(live[i]); + const before = gpa_heap.stats().free_blocks; + + const t0 = soc.cycles(); + const probe = a.alloc(u8, 1024) catch null; + const cycles = soc.cycles() - t0; + soc.rom.print("MARK HEAP_WALK %u free blocks, alloc took %u cycles\r\n", .{ before, @as(u32, @intCast(cycles)) }); + if (probe) |p| a.free(p); + + i = 1; + while (i < live_len) : (i += 2) a.free(live[i]); + live_len = 0; + report("after-walk"); + } + + // ---- phase 4: a long random churn, which is the closest thing here to a real session. + // Random sizes, random free order, and a running count of refusals. A heap that fragments + // itself to death shows up as refusals climbing while `free` stays large. + { + var refusals: u32 = 0; + var ops: u32 = 0; + live_len = 0; + while (ops < 20000) : (ops += 1) { + const keep = live_len < slots and (live_len < 64 or rand() % 100 < 55); + if (keep) { + // 24 B to ~6 KiB: the spread pardes actually shows, from a small string to a + // reallocated row buffer. + const size = 24 + (rand() % 6000); + if (a.alloc(u8, size)) |blk| { + live[live_len] = blk; + live_len += 1; + } else |_| refusals += 1; + } else if (live_len > 0) { + const victim = rand() % @as(u32, @intCast(live_len)); + a.free(live[victim]); + live[victim] = live[live_len - 1]; + live_len -= 1; + } + } + soc.rom.print("MARK HEAP_CHURN %u ops, %u live, %u refusals\r\n", .{ ops, @as(u32, @intCast(live_len)), refusals }); + report("churned"); + for (live[0..live_len]) |blk| a.free(blk); + live_len = 0; + report("drained"); + } + + soc.rom.print("MARK HEAP_DONE\r\n", .{}); + while (true) {} +} + +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} + +pub const panic = std.debug.FullPanic(struct { + fn call(msg: []const u8, _: ?usize) noreturn { + soc.rom.print("MARK HEAP_PANIC %s\r\n", .{msg.ptr}); + while (true) {} + } +}.call); diff --git a/examples/memprobe.zig b/examples/memprobe.zig new file mode 100644 index 0000000..1ce8a0b --- /dev/null +++ b/examples/memprobe.zig @@ -0,0 +1,215 @@ +//! What RAM does this board actually have, and where? +//! +//! The linker script maps one 128 KiB window at 0x4FF00000 and has never needed more. Hosting a +//! real application needs an answer with more than one digit in it, and the answer cannot be read +//! off ESP-IDF's linker fragments, because the fragment that matters +//! (`esp_system/ld/esp32p4/memory.ld.in:18-33`) is parameterised on two things this image does not +//! have: +//! +//! * `CONFIG_ESP32P4_SELECTS_REV_LESS_V3` - true for this rev v1.3 die, which selects a SPLIT +//! layout: a low region 0x4FF00000..0x4FF2BBD0 and a high region from 0x4FF40000, with the +//! mask ROM's own .data/.bss in between at 0x4FF3FBA4..0x4FF40000. +//! * `CONFIG_CACHE_L2_CACHE_SIZE` - the L2 cache is carved out of the SAME 768 KiB array, from +//! the TOP, so `SRAM_HIGH_SIZE = 0x80000 - cache_size`. Its Kconfig default is 128 KiB, but the +//! help text says "to be set on application startup" - the APPLICATION sets it, and this +//! application does not. So the live size is whatever the ROM and the second-stage bootloader +//! left behind, which is exactly the sort of thing that has to be measured. +//! +//! So: probe. For every 4 KiB page in the array, save the first word, write a value derived from +//! the page's own address, read it back, and restore. An address-derived pattern is the point - a +//! constant cannot distinguish real memory from an alias, and aliasing is the specific failure mode +//! of poking at a region the cache controller owns. A page that reads back what it was given is +//! RAM; anything else is reported with what it actually returned. +//! +//! Two pages are never touched: the one holding this image's own .data/.bss/stack, and the mask +//! ROM's reserved window - `soc.rom.print` is the only way this program can report anything, and +//! corrupting the ROM's statics would take the console down with it. +//! +//! A page that is neither RAM nor decoded may raise a bus fault, and this image has no trap +//! handler, so a fault is a silent hang. That is why the scan prints its cursor as it goes: if this +//! stops, the last address printed is the one that killed it, which is itself the result. + +const std = @import("std"); +const soc = @import("soc"); + +/// The whole L2MEM array, per `soc/esp32p4/include/soc/soc.h:161-164` +/// (SOC_DRAM_LOW 0x4ff00000, SOC_DRAM_HIGH 0x4ffc0000). +const l2mem_low: u32 = 0x4FF0_0000; +const l2mem_high: u32 = 0x4FFC_0000; + +/// The mask ROM's .data/.bss, from `bootloader.memory.ld.in:13-16`. Not reclaimable while anything +/// still calls into the ROM, and `soc.rom.print` does. +const rom_data_low: u32 = 0x4FF3_FBA4; +const rom_data_high: u32 = 0x4FF4_0000; + +/// Where PSRAM appears once a driver has trained it (`soc.h:151-153`). Nothing here trains it, so +/// this is expected to fail; it is probed anyway because the cost is four instructions and the +/// alternative is assuming. +const psram_base: u32 = 0x4800_0000; + +const page: u32 = 0x1000; + +/// This image's own footprint, from the linker script's symbols. `.data` starts at the region base +/// and `__stack_top` is the last thing in it, so [l2mem_low, __stack_top) is off limits. +extern const __stack_top: anyopaque; + +fn selfEnd() u32 { + return @intFromPtr(&__stack_top); +} + +/// The heap span the linker script hands over (`build.zig`'s MEMORY block defines both from +/// `l2high`). Referenced here so the report states the same numbers the linker will give the real +/// application, rather than a second copy of them written down in Zig. +extern const __heap_start: anyopaque; +extern const __heap_end: anyopaque; + +/// The value page `addr` must return if it is real, distinct memory. +inline fn pattern(addr: u32) u32 { + // Not `addr` itself: an address bus stuck high would return something that looks plausible. + // XOR with a constant that has bits set where an address never does. + return addr ^ 0xA5A5_0F0F; +} + +/// One saved word per page, so the array can be written whole and read back whole. +const max_pages = (l2mem_high - l2mem_low) / page; +var saved: [max_pages]u32 = @splat(0); + +/// Is this page one the program refuses to touch? +fn skipped(addr: u32) bool { + return (addr < selfEnd()) or (addr + page > rom_data_low and addr < rom_data_high); +} + +/// WHY THIS IS TWO PASSES, and not a save/write/read/restore per page. +/// +/// The per-page version is what this file did first, and it cannot distinguish the three things it +/// most needs to: a store immediately followed by a load of the SAME address returns the stored +/// value under real distinct SRAM, under an address mirror, and under a dirty line in any cache +/// covering L2MEM. The pattern being address-derived does not help, because the alias is written and +/// read through the alias. Cache residency is not excluded by scan length either: 192 pages touch +/// one cache line each, ~12 KiB in total, which fits in any plausible L1 and so is never evicted. +/// +/// Writing every page before reading any page fixes both. If 0x4FF80000 mirrors 0x4FF00000, the +/// later write lands on the earlier page and the read pass sees the WRONG pattern at one of them. +/// The distance between the two passes is 192 pages of traffic, which no L1 holds. +/// +/// This matters more than a tidier loop: `__heap_end` hands the upper span straight to an allocator, +/// so a mirror reported as RAM is silent heap corruption. +fn writePass() void { + var addr = l2mem_low; + while (addr < l2mem_high) : (addr += page) { + if (skipped(addr)) continue; + const p: *volatile u32 = @ptrFromInt(addr); + saved[(addr - l2mem_low) / page] = p.*; + p.* = pattern(addr); + } +} + +/// Read every page back, then put the original word back. Restoring in the same pass is safe: the +/// comparison for this page is already done, and a mirror has by now already been detected at +/// whichever of the two aliases was read second. +fn readPass(addr: u32) Kind { + if (skipped(addr)) return .skipped; + const p: *volatile u32 = @ptrFromInt(addr); + const got = p.*; + p.* = saved[(addr - l2mem_low) / page]; + return if (got == pattern(addr)) .ram else .dead; +} + +/// What a page turned out to be. Three outcomes, not two: a page this program refuses to write is +/// neither RAM nor dead, and folding "skipped" into "dead" is what made the first run of this +/// report `dead 0x00000000..0x4ff02000`, a range that does not exist. +const Kind = enum { + ram, + dead, + skipped, + + fn label(k: Kind) [*:0]const u8 { + return switch (k) { + .ram => "ram", + .dead => "dead", + .skipped => "skipped", + }; + } +}; + +/// Report a maximal run of pages that all behaved the same way. +fn flush(kind: Kind, start: u32, end: u32) void { + if (end <= start) return; + soc.rom.print("MARK MEM_RANGE %s 0x%08x..0x%08x %u KiB\r\n", .{ + kind.label(), start, end, (end - start) / 1024, + }); +} + +export fn zig_main() noreturn { + soc.rom.print("\r\nMARK MEM_BOOT probing L2MEM 0x%08x..0x%08x\r\n", .{ l2mem_low, l2mem_high }); + soc.rom.print("MARK MEM_SELF image occupies 0x%08x..0x%08x\r\n", .{ l2mem_low, selfEnd() }); + soc.rom.print("MARK MEM_ROMRSV rom .data 0x%08x..0x%08x (never written)\r\n", .{ rom_data_low, rom_data_high }); + soc.rom.print("MARK MEM_HEAP linker gives 0x%08x..0x%08x %u KiB\r\n", .{ + @as(u32, @intFromPtr(&__heap_start)), + @as(u32, @intFromPtr(&__heap_end)), + (@as(u32, @intFromPtr(&__heap_end)) - @as(u32, @intFromPtr(&__heap_start))) / 1024, + }); + + // Write every page first, read every page second. See `writePass` for why one pass cannot + // answer this question at all. + soc.rom.print("MARK MEM_PASS write\r\n", .{}); + writePass(); + soc.rom.print("MARK MEM_PASS read\r\n", .{}); + + // Runs are coalesced so the output is a map rather than 192 lines. Every page belongs to + // exactly one run, and every run is printed, so the ranges tile the array with no gaps - which + // is the property that makes the report checkable. + var run: Kind = .skipped; + var run_start: u32 = l2mem_low; + var addr: u32 = l2mem_low; + while (addr < l2mem_high) : (addr += page) { + const kind = readPass(addr); + if (kind != run) { + flush(run, run_start, addr); + run = kind; + run_start = addr; + } + } + flush(run, run_start, l2mem_high); + + // PSRAM, untrained. The question here is only "does the bus answer at all", not "is it + // distinct", so a single write-read-restore is the right shape - and it is expected to fault. + // The line is printed BEFORE the access so a hang is unambiguous. + soc.rom.print("MARK MEM_PSRAM probing 0x%08x (untrained, may hang)\r\n", .{psram_base}); + const pp: *volatile u32 = @ptrFromInt(psram_base); + const ps_saved = pp.*; + pp.* = pattern(psram_base); + const ps = pp.*; + pp.* = ps_saved; + soc.rom.print("MARK MEM_PSRAM read 0x%08x expect 0x%08x %s\r\n", .{ + ps, pattern(psram_base), if (ps == pattern(psram_base)) "answers".ptr else "absent".ptr, + }); + + soc.rom.print("MARK MEM_DONE\r\n", .{}); + while (true) {} +} + +export fn _start() linksection(".text.entry") callconv(.naked) noreturn { + asm volatile ( + \\ li t0, 1 << 13 + \\ csrs mstatus, t0 + \\ la sp, __stack_top + \\ mv fp, sp + \\ la t0, __bss_start + \\ la t1, __bss_end + \\ bgeu t0, t1, 2f + \\1: + \\ sw zero, 0(t0) + \\ addi t0, t0, 4 + \\ bltu t0, t1, 1b + \\2: + \\ j zig_main + ); +} + +pub const panic = std.debug.FullPanic(struct { + fn call(msg: []const u8, _: ?usize) noreturn { + soc.rom.print("MARK MEM_PANIC %s\r\n", .{msg.ptr}); + while (true) {} + } +}.call); diff --git a/examples/minimal.zig b/examples/minimal.zig index a70f0ef..bf93de8 100644 --- a/examples/minimal.zig +++ b/examples/minimal.zig @@ -15,7 +15,7 @@ pub const panic = std.debug.FullPanic(struct { const led: u6 = @intCast(config.led_pin); export fn zig_main() noreturn { - soc.gpio.configureOutput(led); + soc.gpio.configureOutput(led, .{}); while (true) { soc.gpio.setHigh(led); soc.rom.ets_delay_us(100_000); |
