summaryrefslogtreecommitdiff
path: root/src
Commit message (Collapse)AuthorAge
* pardes runs on the board: install a trap vector, clamp the gridGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | It works. `zig build interact -Dpardes` flashes the editor, attaches a terminal, and typing changes the screen. TWO FIXES, and the first is the one that mattered. **Install mtvec.** The console went silent immediately after the editor's allocators.init and nothing could explain it: bounded writes did not change it, no Guru Meditation was printed, and execution did not return from pardes_p4_init even when that function was made to return immediately after the marker that DID print. Setting mtvec to this image's own handler, in DIRECT mode, fixed it - and the handler never fires, which is the tell. hal/intr.zig:302 names the mechanism: the CLIC can fetch handlers from MTVT instead of trapping to mtvec, and the bootloader leaves that vectored mode on with a table this image does not own. `systimer.init` is enabled a few lines earlier, so its first tick dispatched through a vector table belonging to nobody. Writing mtvec with the low two bits clear selects direct mode and the interrupt has somewhere legitimate to go. That single register write took the port from "faults before the editor starts" to the whole of pardes_p4_init succeeding, including vaxis's capability handshake going out over the wire: \e[?1049h \e[?1016$p \e[?2027$p \e[?2031$p \e[?2048h \e[6n \e[>q \e[?u \e_Gi=1,a=q \e[c alt screen, in-band resize, cursor report, kitty keyboard, kitty graphics, DA1 - every one of them answered by the terminal emulator on the far end of the CH340, which is the whole design. **Clamp the grid, and make a failed resize atomic.** Pardes.init then returned OutOfMemory with 9,128 bytes left of 393,216: every cell is paid for four times (vaxis Screen + InternalScreen, pardes Surface + previous_cells). Measured: 40x12 initialises with room to spare, 80x24 does not. So max_cols/max_rows cap the geometry and the host's larger terminal simply hosts a corner of itself. The second half of that is subtler and cost a working editor. `Vaxis.resize` deinits both screens BEFORE allocating the replacements (Vaxis.zig:194-206), so a failed resize leaves vaxis with freed screens and renders nothing at all - and the host bridge injects a size report on attach, so an unclamped 80x24 killed an editor that had already drawn its interface. A failed resize now restores the previous geometry. Measured end to end on the die: the first frame is ~1.5 KB of ANSI drawing the acme tag bars ("New Newcol Joincol Find Grep Help Change" / "Save New Newtty Del Filter"), and typing two characters produces a 116-byte incremental update that draws the `^` modified-marker and moves the cursor to row 3 column 3. vaxis's damage tracking is doing exactly what a 11.9 KB/s link needs. Image is 598,528 B of the 1,536,000 B factory partition, 39%. Also retired here: the I1..I10 and B1..B9 bring-up markers, the MMU table dump, the cache before/after probe, the allocator and vtable pointer dumps, and the RX byte probe. Each answered its question and each answer now lives in a comment beside the code it explains. The trap handler stays - it is the diagnostic this port most needed and did not have.
* Bound the transmit spin; retire the probes that have answeredGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Housekeeping on the instrumentation, plus one fix that stands on its own. uart.write and uart.writeByte no longer spin forever waiting for TX FIFO space. hal/uart.zig:182-186 already made this argument about update() - "on a board with no debugger an infinite spin is indistinguishable from a crash" - and this port demonstrated it: the console going quiet mid-boot read as a hang in whatever code came next, for hours, when a stalled transmitter would have looked identical. The wait is bounded per burst and abandoned bytes are counted in `uart.dropped`, so a lying console is at least a countable one. The bound is deliberately generous: 1,000,000 status reads against an 11 ms drain at 115200. Removed, because each has answered its question and the answers are recorded in comments where they matter: * the MMU table dump - the table is CORRECT, entries 0..9 holding 0x1001..0x100a, exactly the valid bit plus physical page N+1 that the image builder's anchor requires. That is now stated in flushFlashCache's doc comment rather than re-measured every boot. * the before/after eviction read - it established that the same load returns 93 85 85 0f before a capacity flush and 3c ee 08 40 after, which is what the image holds there. The flush itself stays; the proof of why it is needed is in the comment. * the allocator and vtable pointer dumps in src/p4.zig - they showed the struct crosses the seam intact, with its function pointers landing in .flash.text. * the bisect early return - it showed that execution does not come back from pardes_p4_init at all. What is left in place, on purpose: the I1..I10 markers inside pardes_p4_init and the B1..B9 markers in the firmware. The port does not work yet and they are how the next person finds out where it stops. Where it stops: the console goes silent immediately after the editor's pardes.allocators.init and never resumes. It is not the transmit spin - lowering the bound to 20,000 and watching for 30 seconds changed nothing - and it is not a trap the ROM can report, because no Guru Meditation is printed. Execution does not return from pardes_p4_init even when that function is made to return immediately after the marker that does print. So the CPU is lost inside a call whose only work is handing a string to a function pointer, which points at the firmware's own writeOut. That needs an instrument this setup does not have: JTAG, or a GPIO-based tracer that does not depend on the UART at all. Everything cheaper has been tried and is recorded above.
* Flush the caches at startup: the bootloader hands over stale linesGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The image reads its own .rodata and gets its own .text back, and after ruling out everything cheaper the answer is the cache. What was eliminated first, each by measurement rather than argument: * The MMU table is CORRECT. Read from the running application through SPI_MEM_C_MMU_ITEM_INDEX_REG/CONTENT_REG (hal/esp32p4/mmu_ll.h:311-330), entries 0..9 hold 0x1001..0x100a - the valid bit plus physical page N+1 - which is exactly what the image builder's single flash-to-vaddr anchor requires, and entries 10..11 are unmapped as they should be. * The page size is not in question: hardwired to 64 KiB on this chip (mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it), which is what tools/image.zig already assumed. * The flash is correct. The flasher verifies an MD5 of what the ROM stored, and app.bin matches the ELF byte for byte at the addresses that misread. * Not a write failure and not nondeterminism: identical across three resets and two reflashes with the same MD5. * Not 64-byte cache-line granularity either: the wrong bytes come in a contiguous run of at least 192. The measurement that settles it: a load at 0x40035a1c returned 93 85 85 0f, and reading 512 KiB to force capacity eviction made the SAME load return 3c ee 08 40, which is what the image holds there. So the second-stage bootloader hands over with cache lines that do not match the mapping it finally installed. It is perfectly deterministic - the bootloader does the same thing every boot, so it leaves the same lines - which is precisely why it looked like anything other than a cache for so long. The ROM's own Cache_Invalidate_All (0x4fc00404, the same address in both esp32p4.rom.ld and the eco5 table) would be the right instrument and is NOT used: called from here it faults inside ROM code with its argument stranded in a2, so it wants a precondition this image does not know. A capacity flush needs no such knowledge, costs one pass over 512 KiB of already-mapped flash once at boot, and is four times the 128 KiB the L2 measured at. Effect: the firmware now gets through the editor's allocator round-trip and pardes.allocators.init, which is two steps further than before. Also here, and correct independently of any of the above: uart.write and writeByte no longer spin forever on a stalled transmitter. hal/uart.zig:182-186 already made this point about update() - "on a board with no debugger an infinite spin is indistinguishable from a crash" - and this port proved it by spending an afternoon reading a stalled console as a hang in whatever code came next. The wait is bounded and abandoned bytes are counted. Still open: the console stops immediately after pardes.allocators.init. Bounded writes did not change it, so it is not the transmit spin; there is no Guru Meditation, so it is not a trap the ROM can report. The p4 allocator tier's zero-capacity StackFallbackAllocators are the one unusual thing in that call and their reasoning against lib/std/heap.zig is written down in src/allocators.zig, but it has not been tested with a nonzero floor. The bisect markers are left in place for that.
* pardes as P4 firmware: the seam, and a flash-mapping bug in this toolchainGabriel Schneider2026-08-25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The editor arrives as one freestanding OBJECT exporting a seven-function C ABI (src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not as a package dependency. A build.zig.zon path dependency was built first and reverted: merely DECLARING it nested pardes's ~30-package graph under this one and broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy. The seam is bytes in and bytes out, which is what a serial line is anyway: the editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the clock, and neither names the other's types. It is versioned, because linkers do not type-check C symbols and a drifted signature would link cleanly and then corrupt the stack. THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's anchor makes two mapped segments share an MMU page safely - as long as rodata does not END inside the page where text BEGINS. With a 578 KB image it does. A volatile read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared page read as code, so the first thing the firmware tried to print was machine code and it died on an instruction access fault. .flash.text is now ALIGN(0x10000), making the segments page-disjoint. The packing trick this project opened with only ever mattered when the alternative was 64 KiB of zeros in a 1 KB image. Two more findings, both recorded in README.md: * A linker symbol declared as an anyopaque OBJECT gives the optimiser a zero-sized object, so ordinary stores through a pointer derived from its address are dead code it may drop - and did, silently. The allocator's first block header read back as size=2988759312 next=0x14284684 and the free-list walk never terminated. @extern with a many-pointer has no size to lose. examples/memprobe.zig could not have caught it: it writes through a volatile pointer, which the optimiser must leave alone. * The RTC watchdog is armed at handover. Every example here had been resetting on a ten-second cycle, invisibly, because no run had ever lasted eight seconds. State, honestly: the firmware boots, clears .bss, brings up the console, disables the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB heap and calls into the editor, which sets up its sink and its environment. It then faults inside pardes_p4_init on the first allocation. The cause is measured but not fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000 higher - exactly the vaddr distance between the rodata and text segments - while pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in place, deliberately, because the next step needs them. --- correction, measured after the above was written --- Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig was right for a reason I initially got wrong and then measured. I first read bootloader_utility.c's `#else` branch, which classifies segments by address window with two independent ifs - and since the P4's DROM and IROM windows are the identical range (soc.h:146-149), I concluded the last mapped segment wins both roles and the first is never mapped. That branch does not run on this chip. The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851), whose own comment says it: "On chips with shared D/I external vaddr, we don't divide them into either D or I, as essentially they are the same." It collects mapped segments POSITIONALLY into rom_addr[2] and ends with assert(rom_index == 2); Shipping a one-segment image proved it, on the board: Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2) So the split stays, image.zig keeps enforcing exactly two - turning that boot-time abort into a build-time error - and both are now documented with the branch that actually runs and the assert that actually fires. What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may share a 64 KiB MMU page only if they also share a flash page, which the image builder's anchor guarantees - and that holds right up until an application is large enough for rodata to END inside the page where text BEGINS. With a 578 KB image it does. Measured on the die: a volatile read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own .flash.text. Every literal in that shared page read as code, so the first thing the firmware tried to print was machine code, and it died on an instruction access fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint. It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing trick this project opened with only mattered when the alternative was 64 KiB of zeros in a 1 KB image. With that fixed the firmware gets much further: entry, .bss cleared, console up, watchdog disabled, systimer running, ABI version checked, the 384 KiB heap initialised, into the editor, its sink and environment ready - and the literal at 0x4004a1d1 now reads back correctly. Still open, and characterised rather than guessed: pardes_p4_init faults on its first allocation. The allocator struct crosses the seam intact (its function pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it. Ruled out with measurements: the ELF and the image agree at that address, the flash is MD5-verified against the image, the wrong bytes are identical across three resets and two reflashes (so not a stale cache), the corruption is a contiguous run rather than 64-byte lines, and mmu_hal_map_region's arithmetic (page_num = ceil(len/page), entry from vaddr) is correct for the segments as now laid out. The next measurement is the one that settles it: read the MMU entry registers from the running application and print vaddr -> flash for every page. The register model in src/soc.zig can do that; the bisect markers are left in place for it.
* zig-p4: pure-Zig ESP32-P4 toolchainGabriel Schneider2026-08-25
build.zig generates the linker script and drives Zig's own LLD; tools/image.zig turns the ELF into a flashable image and tools/{rom,serial}.zig speak the mask ROM loader over the UART. No CMake, ninja, idf.py, esptool, or external linker. src/soc.zig is a comptime register model over ESP-IDF's own *_reg.h headers; src/hal/ adds peripheral sequences; src/io/ implements std.Io for the chip; src/oracle/ diffs this HAL against ESP-IDF's on the die.