summaryrefslogtreecommitdiff
path: root/src/pardes/uart.zig
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-08-25 14:56:24 -0300
committerGabriel Schneider <[email protected]>2026-08-25 14:56:24 -0300
commitd69785542ced9bd24d210e38788e8eab42d678ad (patch)
treea56fcdfbb20fdee03bf5d7b366620485f50870a2 /src/pardes/uart.zig
parent174991b8f3f8e9c792eede7a52ad7beb10a08b05 (diff)
downloadesp32p4-d69785542ced9bd24d210e38788e8eab42d678ad.tar.gz
esp32p4-d69785542ced9bd24d210e38788e8eab42d678ad.zip
Flush the caches at startup: the bootloader hands over stale lines
The image reads its own .rodata and gets its own .text back, and after ruling out everything cheaper the answer is the cache. What was eliminated first, each by measurement rather than argument: * The MMU table is CORRECT. Read from the running application through SPI_MEM_C_MMU_ITEM_INDEX_REG/CONTENT_REG (hal/esp32p4/mmu_ll.h:311-330), entries 0..9 hold 0x1001..0x100a - the valid bit plus physical page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires, and entries 10..11 are unmapped as they should be. * The page size is not in question: hardwired to 64 KiB on this chip (mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it), which is what tools/image.zig already assumed. * The flash is correct. The flasher verifies an MD5 of what the ROM stored, and app.bin matches the ELF byte for byte at the addresses that misread. * Not a write failure and not nondeterminism: identical across three resets and two reflashes with the same MD5. * Not 64-byte cache-line granularity either: the wrong bytes come in a contiguous run of at least 192. The measurement that settles it: a load at 0x40035a1c returned 93 85 85 0f, and reading 512 KiB to force capacity eviction made the SAME load return 3c ee 08 40, which is what the image holds there. So the second-stage bootloader hands over with cache lines that do not match the mapping it finally installed. It is perfectly deterministic - the bootloader does the same thing every boot, so it leaves the same lines - which is precisely why it looked like anything other than a cache for so long. The ROM's own Cache_Invalidate_All (0x4fc00404, the same address in both esp32p4.rom.ld and the eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM code with its argument stranded in a2, so it wants a precondition this image does not know. A capacity flush needs no such knowledge, costs one pass over 512 KiB of already-mapped flash once at boot, and is four times the 128 KiB the L2 measured at. Effect: the firmware now gets through the editor's allocator round-trip and pardes.allocators.init, which is two steps further than before. Also here, and correct independently of any of the above: uart.write and writeByte no longer spin forever on a stalled transmitter. hal/uart.zig:182-186 already made this point about update() - "on a board with no debugger an infinite spin is indistinguishable from a crash" - and this port proved it by spending an afternoon reading a stalled console as a hang in whatever code came next. The wait is bounded and abandoned bytes are counted. Still open: the console stops immediately after pardes.allocators.init. Bounded writes did not change it, so it is not the transmit spin; there is no Guru Meditation, so it is not a trap the ROM can report. The p4 allocator tier's zero-capacity StackFallbackAllocators are the one unusual thing in that call and their reasoning against lib/std/heap.zig is written down in src/allocators.zig, but it has not been tested with a nonzero floor. The bisect markers are left in place for that.
Diffstat (limited to 'src/pardes/uart.zig')
-rw-r--r--src/pardes/uart.zig47
1 files changed, 38 insertions, 9 deletions
diff --git a/src/pardes/uart.zig b/src/pardes/uart.zig
index ce386fe..d98ac62 100644
--- a/src/pardes/uart.zig
+++ b/src/pardes/uart.zig
@@ -32,26 +32,55 @@ const uart0 = hal.uart.Uart.init(0);
/// Push `bytes` into the TX FIFO, blocking while it is full.
///
-/// The spin is bounded by the wire and there is nothing else for this core to do: a full 128-byte
-/// FIFO drains in 11 ms at 115200. It is also the only backpressure in the system - dropping
-/// instead would truncate an escape sequence, and a half-written SGR leaves the host terminal in
-/// the wrong colour for the rest of the session.
+/// The spin is normally bounded by the wire - a full 128-byte FIFO drains in 11 ms at 115200 - and
+/// dropping instead of waiting would truncate an escape sequence, leaving the host terminal in the
+/// wrong colour for the rest of the session. So the wait is real backpressure.
+///
+/// But it is BOUNDED, for the reason `hal/uart.zig:182-186` gives about `update()`: a UART whose
+/// core clock has been gated never makes progress, and "on a board with no debugger an infinite
+/// spin is indistinguishable from a crash". That is not hypothetical here - it is how this port
+/// spent an afternoon: output stopped mid-boot with no panic and no watchdog (the RTC watchdog
+/// having been correctly disabled), which looked like a hang in whatever code came next rather than
+/// a stalled transmitter. A bounded wait turns that into visibly dropped output plus a counter,
+/// which is a diagnosis instead of a mystery.
+///
+/// The limit is per burst, not per call, and generous: 1,000,000 status reads is far longer than
+/// any legitimate drain and still a fraction of a second.
pub fn write(bytes: []const u8) void {
var rest = bytes;
while (rest.len > 0) {
// One status read per burst, not per byte.
var room = uart0.txFree();
- while (room == 0) room = uart0.txFree();
+ var spins: u32 = 0;
+ while (room == 0) {
+ spins += 1;
+ if (spins > 1_000_000) {
+ dropped +%= @intCast(rest.len);
+ return;
+ }
+ room = uart0.txFree();
+ }
const n = @min(room, rest.len);
for (rest[0..n]) |b| uart0.pushByte(b);
rest = rest[n..];
}
}
+/// Bytes abandoned because the transmitter stopped making progress. Nonzero means the console is
+/// lying about what happened, so it is worth printing.
+pub var dropped: u32 = 0;
+
/// One byte, for callers that must not touch `.rodata` to say anything - which during bring-up is
/// the difference between a diagnostic and a second copy of the bug being diagnosed.
pub fn writeByte(b: u8) void {
- while (uart0.txFree() == 0) {}
+ var spins: u32 = 0;
+ while (uart0.txFree() == 0) {
+ spins += 1;
+ if (spins > 1_000_000) {
+ dropped +%= 1;
+ return;
+ }
+ }
uart0.pushByte(b);
}
@@ -102,9 +131,9 @@ pub fn read(buf: []u8) usize {
/// Pops rather than calling `resetRxFifo`, which is a CONF0_SYNC read-modify-write plus two commits
/// on the console UART - see this file's header.
pub fn drainInput() u32 {
- var dropped: u32 = 0;
- while (uart0.rxCount() > 0) : (dropped += 1) _ = uart0.popByte();
- return dropped;
+ var discarded: u32 = 0;
+ while (uart0.rxCount() > 0) : (discarded += 1) _ = uart0.popByte();
+ return discarded;
}
/// The rate the hardware is actually producing, by reading its dividers back. Reported rather than