| Commit message (Collapse) | Author | Age |
| |
|
|
| |
glue to the editor
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
## The board half of Gpio
`pardes_p4_init` gained a `GpioFn` and the ABI version went to 2, which is what turns a mixed
pair of builds into a refusal to boot rather than five arguments read as six.
`gpioToggle` is four lines over `hal.gpio`: configure the pad as a readable output, read the level
it is driving, drive the other one, read it again. It is deliberately here and not in the editor.
A toggle is not a write to GPIO_OUT - `configureOutput` sets the IO MUX function, the GPIO matrix
route, the drive strength, the input buffer and the pulls, then the output enable, indexed by a
per-pin table - and that code is already in this repo, already the call `src/main.zig` blinks with,
and already checked against ESP-IDF's headers by `zig build diff`. The editor object gets a
function pointer instead of a second copy nobody tests.
`getDrivenLevel`, not `getLevel`: the answer is the level the board is driving, which is defined
for every pin. The pad's own level is what the outside world says, and on an unconnected header pin
that is noise. The input buffer is enabled anyway so `Peek` of GPIO_IN_REG can be compared to it.
## A fifth hardware check
`p4-bench --check` runs `Gpio 33` three times and requires 0->1, 1->0, 0->1.
The alternation is the oracle, not either answer. `0->1` alone is what a firmware printing a
hardcoded string would also say; two runs that disagree can only come from a level that was stored
and read again. Three, so the third rules out an ordering coincidence. GPIO33 because
`src/oracle/ledc_cases.zig` already documents it as a free pin on this board's JP1 header - pin 21
on the diagram the editor now draws. GPIO20 is the blink demo's pin and may have a wire on it.
This is the only check here that crosses the whole seam: editor word, C ABI, HAL, pad, and the
level back out through the message row. Nothing smaller exercises the ABI at all.
## -Dtheme-animation forwarded
Same path as the geometry, for the same reason: baked into the object, wanted from here.
All five checks pass on the die; host tests green; the board is flashed with md5 932898f34ca1dc31.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The firmware build already rebuilds pardes's object when the grid changes, for the reason
that changing it in one directory and linking in another silently links the previous answer.
The chrome fade is the same shape of setting - baked into the object, wanted from here - so
it travels the same path: -Dtheme-animation is forwarded to the nested build, and like the
geometry it is skipped entirely when -Dpardes-obj names an object explicitly.
Measured either way on the die: the fade costs 1,097 ms of wire per theme change against 215
without it, and compiling it out takes 2,336 bytes off the image. Both numbers are in the
comment above the option in ../02-pardes-code/build.zig.
Board reflashed with the default (md5 018a9554e4a8ae51); p4-bench --check 4/4; host tests
green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
## -Dcols / -Drows
The geometry is baked into pardes's object, and that object is built by the other
repository - so changing the grid was two commands in two directories, and the second one
silently linked whatever the first had left behind. That is the shape of mistake that ends
with a firmware whose grid is not the grid you think you flashed.
`-Dcols` and `-Drows` now drive the editor's own build before linking it. A nested
`zig build`, not a package dependency: the seam between these repositories stays a file,
for the reasons the comment above it has always given. This only reaches across to ask for
the file to be made a particular way, and it stands aside entirely when `-Dpardes-obj`
names an object explicitly, because then the caller has already said which one they want.
Verified by the segment table: `-Dcols=80 -Drows=24` takes the loaded segment from 20,408
to 49,944 bytes, which is the shadow grid going from 784 cells to 1,920.
## Poke and Peek, checked against each other
Two more checks in `p4-bench --check`, and they are checked against each other because
that is the only way to check either one without a second debugger: a Peek alone cannot
tell a correct read from a stuck one, and a Poke alone cannot tell a write from a no-op.
Both oracles were wrong before they were right, and both failures are worth recording.
The Escape check asked whether an `x` came back on the wire. That was true until the boot
buffer gained a line beginning "x selects a line", at which point a passing check began
failing for a reason that had nothing to do with Escape. It now reads the CURSOR COLUMN
after `0`, which the document cannot spell: 8 in normal mode, three further right if the
`0` had been inserted instead.
The Poke and Peek checks searched the raw byte stream for `deadbeef` and could not find it -
because this renderer emits only the cells that CHANGED, jumping between runs with absolute
cursor positioning, so a word on screen is frequently not a word on the wire. `stripAnsi`
removes the escapes first. The Poke assertion then had to be loosened from one phrase to
three facts, because the message row is 49 columns here and a wrapped line puts a row
boundary inside whichever phrase straddles it.
Four checks, all passing on the die:
lone Escape still leaves insert mode ok (cursor col 8)
a click split byte-by-byte lands at 18 ok (cursor 3,18)
Poke writes a word and reads it back ok
and a later Peek still finds it ok
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
## zig build selftest
Seventeen checks, on the board, chosen by one rule: a check belongs there only if the die
can answer it and a host cannot. Every serious bug this port has produced was invisible to
a host test. `std.mem.eql` compares a byte at a time on this target, so the firmware's
largest read was three times slower than it needed to be and nothing on a laptop could
tell. A lone ESC resolves to the Escape key, which is right when a kernel hands over a
whole sequence and wrong when a 115200 line hands over one byte every 87 us. A full
transmit FIFO stopped anything draining the receiver, and the FIFO depth is a hardware
number.
So: the volatile promise (two reads of a live counter are two reads, which is what `Peek`
rests on), word-wise equality checked against `std.mem.eql` itself at every difference
position and both alignments, the input rescue against a fake port with a FIFO that loses
what arrives into a full one, the allocator on real L2MEM, and the cycle counter against
the systimer - which is clocked from the crystal and therefore cannot be flattered by a
wrong CPU divider.
Anything that is pure logic stays in `zig build test`, which is faster and needs no
hardware. Duplicating those here would make the suite longer and no stronger.
Not a `zig test` binary, deliberately: Zig's runner wants an OS and `std.testing.allocator`
is a debug allocator over the page allocator, which on freestanding is either a compile
error or a lie. The harness is thirty lines.
`selftest` has its own application, image and flash chain so it is one command with no
flags to remember, and it makes the board's verdict the build's - a suite whose result a
human has to read out of a scrolling log is a suite that gets ignored on the first busy
afternoon. Proven both ways: 17/17 with exit 0, and exit 1 naming the check when one is
deliberately inverted.
The clock check earned its place immediately. The first version read `config.cpu_mhz` and
compared the die against it without ever performing the raise, so `-Dcpu-mhz=360` failed
with `khz=90001 want=360000`. The check was right and the expectation was wrong; it now
calls `setCpuFreq` itself, which makes it a test of the raise rather than a tautology.
90001 kHz at 90, 360004 at 360.
## zig build layout, and a loader error that explains itself
Both of these exist because of an hour I spent that they would have saved.
`NotTwoMappedSegments` said the count was wrong and nothing about what the segments were,
which is the only thing that says which section grew, shrank or stopped being emitted. I
diagnosed one by hand with readelf on an artifact that turned out to be a stale install,
then guessing at the linker script. The image step now prints the segment table with the
error - it has to happen there, because a rejected image is never written, so no later step
can show it - and `zig build layout` prints the same table for an ELF plus the loader's
verdict as TEXT, which `size` cannot do because it reads a finished image and the moment you
need the table is when there isn't one.
They immediately paid for themselves. The reason the suite would not build was that it had
no `_start`: `-fentry=_start` found no such symbol, `--gc-sections` discarded every
function as unreachable, and `.flash.text` was empty. `layout` prints `entry 0x0` for
exactly that, in one line. The second failure - a silent board - was a missing app
descriptor, which is the same class of thing and now has a comment where it happened.
## The panic handler already existed, and was printing past the end of its message
`msg` is a Zig slice and `%s` reads until a NUL, so handing `msg.ptr` to the ROM's printf
printed the message and then whatever followed it in memory. String literals get away with
it; std's own panics do not, because they are formatted into a buffer - "index out of
bounds: index 5, len 3" - and carry no terminator. It now goes out through `uart.write`,
which takes a length, with the fault address after it so addr2line can find the line.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
could not see it
`zig build interact` did not compile. The mouse coalescer dated its held report with
`std.time.milliTimestamp`, and this Zig's `std.time` has constants and `epoch` and no clock
at all: the repo's own idiom is `std.Io.Timestamp`, already wrapped as `serial.Port.nowMs`,
which is what the three call sites now use.
The reason it survived every check I ran is worth writing down, because the same hole will
swallow the next one. `zig build test` compiles `tools/console.zig` as a test module, and
Zig only analyses what is REACHABLE: no test calls `attach`, so the loop containing the bad
call was never looked at. Seven passing MouseFilter tests and a green `zig build test` said
nothing whatsoever about whether the file compiles as part of an executable. Nor does plain
`zig build` - p4-console is not in the default install step - so the only thing that would
have caught it is building or running the named step, which is also the documented way to
use this repo: `zig build interact -Dpardes`.
Verified the way it should have been the first time: `zig build interact -Dpardes
-Dcpu-mhz=360` with piped stdin flashes, attaches, takes the keystrokes and detaches at
EOF. Also compiled every step's artifacts - elf, console, bench - and re-ran the matrix
that matters for the editor side: default, Debug, ReleaseFast, ReleaseSafe, ReleaseSmall,
and -Dplatform=p4 at Debug and ReleaseSmall, plus -Dp4-cols/-Dp4-rows at 40x12 and 80x24.
Note for anyone reading a failing `zig build console` or `bench`: those RUN their tool, so
they fail with DeviceBusy when something else is attached to the port. That is not a build
error, and it is how this one hid in plain sight in the middle of the output.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Adds the geometry curve - 0.87 us per cell, measured across nine grids from 40x12 to
140x42 - and the two ways it ends: 200x60 refused by the linker because '.bss' cannot hold
the shadow grid, 160x48 linking and then trapping. Neither is a heap problem any more,
which is the finding: the heap has 336 KB spare while '.bss' runs out, because vaxis's two
unread grids stopped being allocated.
Also the three bugs the latency work walked straight past, none of which any number in this
document could see: keystrokes lost while the transmitter was full, every escape sequence
shredded because a lone ESC arrives alone on a 115200 line, and a mouse that was never
enabled. The middle one had been there since the port began and explains the other two -
a click arriving as ten key presses, with the '0' among them moving the cursor to column
zero, which is what made it look like a coordinate bug.
And Table 15 changes from 'not fixed' to 'fixed, and not on purpose'. The vx.resize failure
that three experiments characterised and left alone was an allocation of vaxis's two grids;
the emitter stopped needing them, so the path that failed is gone rather than repaired.
Verified on the die - an in-band report re-lays the screen from 56 columns to 30 and back.
Recorded as a consequence rather than as a fix, because nobody set out to repair it.
Eleven pages.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Two small changes that the geometry sweep needed.
The firmware asked for 80x24 and let the shell clamp it, which made this file a second
opinion about the board's geometry - one opinion too many, and wrong the moment the shell
could render more than that. It now asks for 255x255 so the shell's own ceiling is what
governs, and the shell reports what it settled on.
And the heap is printed on every boot rather than only when init fails. A geometry that
fits with 2 KB to spare and one that fits with 80 KB are not the same answer, and from the
host the difference was invisible. That line is what showed the sweep that memory had
stopped being the constraint at all: 336 KB free at every size tried, while '.bss' - the
shadow grid, sized at comptime - is what actually runs out.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
hardware can
## The console thins mouse reports
The board asks for DEC 1002, so the terminal reports presses, releases and motion while a
button is held. A press is one report; a DRAG is one report per cell crossed, each about a
dozen bytes. A hand crosses forty cells in a tenth of a second, which is ~500 bytes, which
is 43 ms of a 115200 line - and every one of those bytes is input the board must parse
while it is trying to paint the result of the previous one. Unthinned, a drag makes the
editor unusable for as long as the drag lasts and for a while after.
`MouseFilter` holds the newest motion report and drops the ones it supersedes, on a 40 ms
window: ~25 reports a second, about 2.6% of the line. Newest-wins is right for motion and
only for motion - where the pointer PASSED THROUGH is not information the editor can use,
since a selection is defined by where the drag began and where it is now, so an intermediate
report already stale by the time it reaches the wire is pure cost. Presses, releases and
wheel notches are never held: each one means something different and dropping one loses a
click.
Ordering is the part worth testing. A held report is released before any non-motion byte
that follows it, so a release cannot overtake the motion it ends, and a drag that stops
moving still delivers its final position when the window expires. The seven tests cover
those, a report split across a read boundary, a non-mouse escape sequence passing through
untouched, and - the one that would hurt most - a lone ESC not being swallowed, because
that is how you leave insert mode.
## Two hardware checks: p4-bench --check
Both are regressions that no host test can see and no latency number can show.
A lone Escape still leaves insert mode. The shell now holds a solitary ESC for 10 ms
because on this wire the first byte of every sequence arrives alone; if that hold ever
stops expiring, Escape stops working and the editor is unusable.
A click split byte-by-byte lands at the column clicked. This is the bug that made the mouse
look unimplemented, and it only appears when the bytes arrive separately - which the wire
does anyway, 87 us apart. The check sends them as ten separate writes with no gap, because
a gap longer than the hold would expire it and the check would be exercising nothing.
Two things this file deliberately does NOT check, both because a check that cannot fail
honestly is worse than no check. Input loss during transmit belongs to the deterministic
host test in `src/pardes/input_rescue.zig`, which loses 67 bytes with the fix removed and
needs no board. And every hardware oracle for it that was tried here was worse: the cursor
stops being reported past 160 characters because the wrapped line outgrows the viewport, and
a screen reconstruction cannot be rebuilt mid-session because the board only sends what
changed. Both false starts are recorded in the file so the next person does not repeat them.
The check also found its own bugs before it found any of the firmware's: 12 ms gaps between
the click's bytes expired the very hold it meant to test, `$` produces no frame when the
cursor is already at the end of the line, and reading the cursor after a press alone
measures the revert rather than the click.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Reported as "a key is stuck and is only sent when I send a new event". It was neither
stuck nor late - it was gone, and a later frame repainting those cells is what made it
look like it arrived eventually.
The loop is read, apply, render, write, and `uart.write` blocks while the transmit FIFO is
full. That wait is real backpressure and should stay: dropping half an escape sequence
leaves the host terminal in the wrong colour for the rest of the session. But NOTHING
drained the receive FIFO during it, and that FIFO is 128 bytes - 11 ms of wire at 115200.
Measured on the die, typing a burst in one host write and counting what the firmware's
loop actually took off the UART:
burst before after after + chunked input
128 128 128 128
200 197 200 200
300 257 300 300
600 478 600 600
1200 - 1200 1200
2400 - 2316 2400
4096 - 3611 4096
Two windows had to close, and the second was only visible once the first was shut.
`input_rescue.pump` drains the receiver on every iteration of the wait for transmitter
room. That is the big one, and it is the whole reason this policy lives in its own file:
`uart.zig` cannot be tested without the chip because every line of it is an MMIO access,
while `pump` takes its port as `anytype` and runs against a fake with a two-byte transmit
FIFO and an eight-byte receive FIFO in `zig build test`. The fake models the receive FIFO
the way the hardware behaves - a byte arriving into a full FIFO is simply gone - so the
test fails by 67 lost bytes with the rescue removed, which is the die's 88-of-200 in
miniature. It also caught a flaw in its own first draft: a fake whose transmit FIFO drains
as fast as it fills never blocks, so `pump` never waits and the test proves nothing.
The second window was APPLYING the input. A keystroke costs 44 us on an empty line and
63 us at 640 characters, so handing the editor a full 128-byte batch is up to 8 ms in
which nothing drains the receiver - against 11 ms of FIFO. The loop now feeds the editor
eight bytes at a time and rescues between chunks. Splitting a burst at an arbitrary byte
is already safe, because `pardes_p4_input` keeps whatever it could not parse; that is how
it survives an escape sequence split across two UART reads. One render still happens per
loop iteration, so this costs no extra wire. Eight rather than thirty-two by measurement:
32 left 2400 and 4096 lossy, 8 does not.
Beyond 4096 bytes in one burst the editor genuinely cannot keep up, and the ring reports
what it abandoned instead of losing it silently - `rxdrop` in the PROF line, alongside a
running count of received bytes. That counter is the other lesson here: the first attempt
at measuring this counted characters on the reconstructed screen, which cannot distinguish
"never arrived" from "arrived but off the edge of the viewport", and it disagreed with the
hardware in both directions.
No cost to latency: round trip median 3687 us over 60 trials against 3687 before, maximum
3866, screen byte-identical to the vaxis reference, host tests green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target,
and the closing section now reports it the way it should be reported: over 80
phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document
length rather than as one intercept.
Three things this adds that are findings rather than steps.
## The pad target, swept
Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the
cursor-positioning frame 21 -> 49 bytes made it a millisecond faster, padding the
cursor-hiding frame 6 -> 36 made it slower - so the target was swept as the only variable.
0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary
is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either:
it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what
confirms the descriptor is the thing to believe.
## The largest read in the firmware ran a byte at a time
The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame,
comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte,
about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32
at a time: 223 -> 66 us, 0.95 cycles per byte, at every document length.
Recorded alongside it are the two candidates that were measured and REVERTED, which is the
more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in
Surface.set removed 52 million instructions per host run and zero cycles on either host or
board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are
cheap; it was the reads that were slow.
## Where it stopped, honestly
The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty
line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12
viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and
4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are
now in the table with a "cells changed: none" row under them, because a round trip for an
edit that displays nothing is worth reporting as exactly that and not as a failure to hit
a number.
Also corrected: the ranked-recommendations table said the UART0 raise was abandoned
because a higher line rate moves settle and leaves the round trip alone. That reasoning
assumed the first byte reaches the host as soon as it is sent, and the bridge finding says
otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so
the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that
one 115200-baud line is the premise of this port rather than a free variable, and meeting
the target by changing the link would answer a different question.
Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Brings the document up to the end of Experiment 4. Two rows on the progression table -
360 MHz and direct emission - plus the sections behind them, the corrected verifier, and
a summary that no longer says the target was missed. It is met: 3.74 ms from 16.99.
The two findings worth more than the number, and both are written up as findings rather
than as steps:
A 4x clock bought 2.6x. The grid walk reads 27 KB a frame at about six cycles a byte, so
it is bounded by L2MEM bandwidth and does not care how fast the core runs. Predicted
before the measurement.
The direct renderer measured SLOWER at first - less computation, a quarter of the bytes,
worse round trip - because the CH340 forwards a bulk IN packet only when the packet is
full, and a 21-byte frame does not fill 32. It waits about a millisecond for a timer.
So the frame has a minimum size and it belongs to the transport, not the terminal; the
emitter pads to it with repeated cursor positioning. The table of 21/49/81-byte frames is
in the report because the MINIMUM column is the tell: the small frame's floor was already
530 us below vaxis's, exactly the compute saved, and only the median was hostage.
Also corrected: the verifier section. It described hashing SGR parameters per cell, which
is a history rather than a state, and that version reported a difference on the die that
did not exist - a faster board split the same keystrokes across different frames and
reached the same colours by another route. The document now says what it does instead,
and why the from-scratch renderer could not have been verified without the fix.
Table 13, the ranked recommendations from Experiment 3, is kept as the prediction it was
with a note on what has since been applied and that row 3 - raise the line rate - was
abandoned. The threats section no longer states the link floor as a single number, since
the bridge's packet granularity is now part of it and is specific to this bridge.
Nine pages. Every figure still reads the raw per-trial CSVs, including the two new ones.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The board was executing at 90 MHz because the stock second-stage bootloader is built
with CONFIG_BOOTLOADER_CPU_CLK_FREQ_MHZ=90 (bootloader_clock_init.c:27-37). Measured
here against the systimer, which is XTAL/2.5 and therefore an independent reference:
4,500,367 cycles in 50,004 us = exactly 90 MHz.
The CPLL is ALREADY at 360 MHz - 90 is 360/4 - so this is a divider change and nothing
else. No PLL to enable, no lock to wait for, and the P4 has no per-frequency voltage
step to order it against (rtc_clk_init.c:58-80 sets HP_ACTIVE DBIAS once from efuse).
`hal/clkrst.zig:setCpuFreq` writes the four dividers in ESP-IDF's upscale order -
APB, SYS, MEM, then CPU, with a bus-update handshake after each - because IDF's own
comment says the other order passes through a state where APB or MEM violates its
timing. Then it calls the ROM's `ets_update_cpu_frequency`, without which every
`ets_delay_us` in the image is wrong by exactly the frequency ratio.
Measured after: 359,991 kHz. Nothing else moved, which is the reason this is safe from
a running console: UART0's baud clock comes from XTAL (hal/uart.zig:116-139), the
systimer from XTAL/2.5, and the flash interface from SPLL 480 MHz - none of them from
the CPU. The `cycle` CSR simply counts faster, and the board never converts it, so only
the divisor in experiments/ had to move.
## What it bought
step fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
+ byte compare 8.37 ms 7.1 us 9.48 ms 0.37x
+ 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x
Compute went 6.42 -> 2.44 ms: 2.6x for a 4x clock, not 4x, and the shortfall is the
point. At 360 MHz the grid walk reads 27 KB per frame in 226 us, about 6 cycles a byte,
so that stage is bounded by L2MEM bandwidth and does not care how fast the core is.
The prediction that this would happen was made before the measurement and held.
## The goal was 4 ms and this is 4.37
Short by 372 us, and the remaining budget is known: ~2.0 ms of host and USB latency
that no firmware change touches (measured independently against the protocol
responder), plus 2.4 ms of board compute of which vaxis's own diff is 631 us, pardes's
Surface rebuild ~505 us and our grid walk 226 us. vaxis's diff is the only item large
enough to close the gap alone, and it is redundant work - `present` already computes
exactly which cells moved - so emitting ANSI straight from the shadow grid would do it.
I did not, because it is a from-scratch renderer and the honest verification for it
needs more than the harness currently proves.
## Verification, and a bug in my own instrument
Raising a core clock 4x is exactly the change that corrupts a screen quietly, so the
A/B compares screens across clocks as well as across the shadow-grid flag. The first
attempt REPORTED A DIFFERENCE at 360 MHz, and it was the verifier: it hashed the raw
SGR parameters applied to each cell, which is history-dependent, and a faster board
splits the same keystrokes across different frames. Decoding SGR into actual state -
resolved foreground, background and attribute set per cell - it is identical: same
characters and same style everywhere, both across clocks and across the flag.
Also checked and found innocent: `rtt.zig` polled with a 1 ms timeout, which looked
like it would quantise every sample. It does not - poll(2) returns when data arrives,
not when the timeout expires - and switching to a non-blocking spin moved the measured
round trip by 0 us. The comment now says so, since the next reader will wonder too.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
zig-p4 host tests, tty and p4 both build. 90 MHz remains the default; -Dcpu-mhz=360 is
opt-in because every number in experiments/ up to this commit was taken at 90.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Experiment 4 records what each cut was worth, measured on the die at every step, and
the stage table that made the cuts findable at all. Also two corrections to the
document itself:
The data plumbing loaded four of the eight datasets, so the progression table would
have been computed from a subset. It now loads all eight.
A continuation line beginning with `+` is list markup to Typst, so the CSV file names
were rendering as a numbered item at the top of page one. One expression, one line.
The summary no longer ends on the ReleaseFast flag as the best available change; it
ends on the measured 16.99 -> 8.37 ms and on the fact that the target is not met, with
the remaining 6.42 ms of compute broken into the three items it actually consists of.
Saying 'not met' in the summary matters more than the table: the number that was
missed is the one a reader should see first.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Progression on the die, five document lengths x seven trials at each step:
step fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
+ byte compare 8.37 ms 7.1 us 9.48 ms 0.37x
Round trip is time to the FIRST response byte: ~1.95 ms of host and USB latency plus
compute. Compute is 6.42 ms and the 4 ms goal needs it under 2.05 ms, so 3.1x remains.
Every microsecond of it is now measured rather than guessed, by stage, on the die:
our walk of the grid 979 us this repo's own present()
vaxis diff + emit 2455 us vaxis walks all 480 cells, per-cell strings
pardes Surface rebuild 2000 us pardes rebuilds every cell every frame
input parse + edit 250 us
The honest reading is that the two large items are not each other's alternative.
vaxis's diff is redundant work - present() already computes exactly which cells moved,
so emitting ANSI directly from the shadow grid would remove most of that 2.46 ms. But
even a perfect renderer leaves pardes rebuilding a whole Surface per keystroke, and 4 ms
at 40x12 needs that too.
Raising the baud does not move this number. At 115200 an 81-byte reply is 7.0 ms of
wire, but almost none of it lands before the first byte; 921600 takes `settle` from
23 ms to ~16 ms and leaves the round trip where it is. Worth doing, not for this.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The board half of the run to 4 ms: the instrumentation that found the cost, and the
measurements that judged each change.
`-Dprof` grew two things. It now renders a SECOND time with nothing changed, which
splits a frame's cost cleanly: whatever the second render still costs is the price of
walking and diffing the whole editor state, and the difference between the two is the
price of the change itself. On the die those measured 10.9 ms and 0.1 ms - so 99% of a
keystroke was work done regardless of what the keystroke did.
It also reads `pardes_p4_frame_prof`, a new export that reports the last frame's three
stages in CPU cycles. That is what turned "render is slow" into an address:
stage before after
copy Surface -> vaxis 6 750 us 1 450 us
vaxis diff + emit 2 460 us 2 455 us
push into the UART 1 us 1 us
(pardes's own Surface build) ~2 600 us ~2 000 us
The copy was 57% of a keystroke and it was in this repo's own `present`, not in
pardes and not in vaxis.
Measured on the die, five document lengths x seven trials per configuration:
configuration fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
## Where the remaining 4.9 ms is, and why the goal is not met
Round trip is time to the FIRST response byte, so it is ~1.95 ms of host and USB
latency plus compute. Compute is now ~6.9 ms and 4 ms needs it under 2.05 ms: a
further 3.4x. The three remaining pieces are known and measured - our walk of the
grid (1.45 ms), vaxis's own diff and emit (2.46 ms), and pardes rebuilding the whole
Surface (~2.0 ms) - and the honest reading is that even a perfect renderer leaves the
Surface rebuild, so 4 ms needs pardes to stop rebuilding a whole frame per keystroke.
Raising the baud does NOT help this number, and that is worth writing down because it
is the obvious next idea: at 115200 an 81-byte reply is 7.0 ms of wire, but almost
none of it lands before the first byte. 921600 takes `settle` from 23 ms to ~16 ms
and leaves the round trip where it is.
## A bug found on the way
`vx.resize` fails on this board. A runtime geometry change takes its allocation
failure path, restores the previous size and returns: 80 bytes go out where 1,392
should, and the screen keeps its old shape. Reproduced with the shadow grid compiled
out, so it predates it. The board has one geometry per session.
That is also why the shadow grid's correctness test compares two firmwares rather
than forcing a repaint with a resize - the forcing mechanism does not work here. The
reference path and the incremental path were each run against the same 19-step
workload and their reconstructed screens are byte-identical.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
flag is worth 21%
The obvious question after the last commit was what it did on the actual target. The
answer is nothing, and the useful part is that the same run says what DOES work.
Four builds on the die, 5 document lengths x 7 trials each:
configuration fixed per char at 160 vs base
ReleaseSmall 16.99 ms 54.3 us 25.56 ms 1.00x
ReleaseSmall+lineSpan 17.10 ms 54.0 us 25.63 ms 1.00x
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
ReleaseFast+lineSpan 14.88 ms 34.2 us 20.25 ms 0.79x
The edit-path change is invisible in BOTH modes, and ReleaseFast+lineSpan is
indistinguishable from ReleaseFast alone. The optimisation mode is the whole of the
difference, which is what Experiment 3 predicted: the edit path is 220 us of a 17 ms
keystroke, so making it cheaper cannot move the total, while the mode makes the
RENDERER faster and the renderer is where the time is.
`experiments/report.typ` gains that table and the reason the four-way comparison
exists rather than four indistinguishable lines on figure 1.
The board is now flashed with the default build, which as of the pin next door means
ReleaseFast: 809,552 B of the 1,536,000 B partition, 53%.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The board half: the -Dprof attribution that overturned the conclusion, its data, and
the report correction.
`-Dprof` adds two cycle-counter reads around `pardes_p4_input` and
`pardes_p4_render` and prints both. Off by default: it puts a line on the wire per
frame, which is the very resource being measured, so it answers "where did the 15 ms
go" and not "how fast is it".
It answered. Input is flat at ~220 us regardless of document size - 1.5% of a
keystroke - and the entire ~15 ms floor plus every microsecond of the per-character
slope live in `render`. The edit-path fix that the source reading implied (committed
next door in 02-pardes-code) is worth 20% on a 19 MB file and, measured here over 5
conditions x 7 trials, exactly 0% on this board.
`experiments/report.typ` gains Experiment 3 and a correction: Experiment 2's
mechanism claim was wrong, says so, and carries the disproof beside it. The ranked
recommendations are reordered with the renderer at #1.
Also here: `--sweep position` in p4-bench, which holds the document fixed at one
320-character line and moves only the cursor. Column 320 costs 33.9 ms and emits 28
bytes; column 0 costs 26.0 ms and emits 81. Latency and output size are inverted on
this board - the signature of a walk from the start of a line.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
"Too slow for interactive use" is a real complaint and not a number. This adds the
number, and the number says the wire is innocent.
## The instrument
`tools/perfproto.zig` is a small framed protocol - "P4", op, length, CRC-32 of the
payload, payload - shared VERBATIM by the host tool and `examples/uartperf.zig`, so
a frame one writes and the other parses cannot drift. It is imported as a module by
both, not copied.
The checksum is the whole point. RX overrun on this UART is undetected in hardware
and uncounted in the driver, so a byte that never arrived is indistinguishable from
a late one; a throughput figure that is not checksummed is a guess about how fast
data was corrupted. `sink` accumulates a CRC over every payload byte the board
received and `report` hands it back, so the host can prove that what arrived is
what it sent.
`tools/rtt.zig` is the two timing functions: `roundTrip` and `measure`. Round trip
is to the FIRST response byte, deliberately. A renderer that starts drawing in 8 ms
and finishes in 130 ms feels immediate; one that thinks for 130 ms and then draws in
8 ms feels broken; waiting for the wire to fall quiet cannot tell them apart. Time
to the last byte is recorded separately as `settle`. Microseconds, because at 115200
one byte is 87 us and a millisecond clock quantises the answer into buckets eleven
bytes wide.
`tools/bench_main.zig` is `p4-bench`: `--link` for the ceiling, `--editor` for how
much of it the editor uses, `--sweep` for one controlled variable at a time with
`--csv` raw per-trial output.
## What it measured
The link is essentially perfect: 11,496 B/s up and 11,413 B/s down, 99.8% of
capacity in both directions, CRC verified over 32,768 B each way, zero corruption.
Typing at 6 to 100 keys/s loses nothing and never uses more than 9% of the wire, so
H5 - "typing loses input" - is refuted.
Latency is compute per input event, not transmission. A 40-byte motion and a
206-byte insert-and-escape cost the SAME round trip to within 0.3 ms, across a
five-fold range of output. That is why raising the baud cannot fix typing: there is
almost no wire in it.
And an edit costs the whole document. Round trip against characters already in the
line is a straight line at 54.3 us per character per keystroke - 17.0 ms at an empty
line, 25.6 ms at 160. On a ~90 MHz core that is ~5,000 cycles per character, far
more than a copy alone, so the full-buffer copy the source does is accompanied by at
least one more full pass.
One controlled intervention: building the editor object ReleaseFast instead of
ReleaseSmall cuts the fixed cost 13% and the per-character cost 36%, for 35% more
flash (809,536 B of a 1,536,000 B partition). Its advantage grows with the document.
Nothing else measured comes close to that ratio.
## Three bugs found while building it
The responder printed garbage and looked dead: it read `.rodata` before evicting the
bootloader's stale cache lines. `flushFlashCache` moved from `src/pardes/app.zig` to
`soc.zig` with its measured evidence, since every application that touches `.rodata`
after hand-over needs it and exactly one file knew that.
Then it booted, printed its marker and went silent after ten seconds:
`rst:0x10 (CHIP_LP_WDT_RESET)`. The bootloader arms the RTC watchdog and expects the
application to take it over. Only the editor ever did.
`serial.Port.drain()` drains INPUT, not output - so timing a transfer to it reported
202% of the wire's capacity and ate the reply. Added `flushOutput` (tcdrain), named
so the two cannot be confused again.
Also: Zig 0.16 emits an explicit `+` for a non-negative SIGNED integer whenever a
width is given (std/Io/Writer.zig:1548-1559), which put a `+` in front of every
number in the first tables.
## The report
`experiments/report.typ` reads the raw CSVs and computes its own figures, so a
re-run changes the document instead of contradicting it. It states five hypotheses,
settles each against one experiment, and is explicit about the one that failed: the
geometry sweep is confounded, because characters accumulated across conditions and
the length experiment then proved that matters. It is reported as unsupported rather
than dressed up as a result.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Peek`, `Poke` and `Hexdump` are the reason this port exists and nothing here said
how to reach them. Two things are not guessable.
The boot layout: the editor now comes up as ONE EMPTY OUTPUT BUFFER and no shell
pane, because a shell is not a layout preference on bare metal but an
impossibility - nothing to fork, no pty. That is a pardes-side change; what belongs
here is that it is what `zig build interact -Dpardes` puts on the screen.
The chord: there is no mouse on a serial line, so execution is
`i <cmd> ESC` then `x` then TAB - write the command into the buffer, select the
line, execute the selection (pardes.zig:7669). `o` opens a fresh line for the
next one. Nothing about that is discoverable from the screen, and every other
spelling I tried failed in a way that looked like the port was broken: `:` opens
the tag command line in NORMAL mode, so typing there moves the cursor instead of
inserting, and the tag's editable region is seeded with the builtin words, so text
typed at its start merges into `Save` and the chord executes the wrong word.
Every number in the new section is off the die, including the one that matters:
`Peek 0x50110001` answers `peek: MisalignedAddress` on the message row instead of
trapping, and a trap here is a watchdog reset that takes the session with it.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`zig build interact` shredded the editor's screen with fragments of
`[11/13] steps +- console`. The cause is structural, not cosmetic: an interactive
step lives for as long as the human does, and `std.Progress` redraws the build
runner's step tree on stderr every 80 ms for all of it.
std already solves this, and the solution is a lock rather than a flag.
`Step.Run` with `stdio == .inherit` holds `io.lockStderr()` for the entire
lifetime of the child (std/Build/Step/Run.zig:1588-1592) - the same lock
`Progress` must take to draw. A hand-rolled step gets none of that unless it
takes the lock itself, and `ConsoleStep` did not. So the console is now a real
host program, `tools/console_main.zig`, and `console`/`interact` are `Run` steps
on it. `--color off` is not needed and is no longer suggested anywhere.
Measured under a real pty (a pipe hides the bug, because progress only draws to
a terminal - which is why every earlier end-to-end test here looked clean):
zero bytes of progress output across an 11-second session, 3 frames, the editor's
`^` modified-marker landing at row 2 after two keystrokes.
The second reason for a program is that "just connect" should not imply a build.
Once the firmware is in flash the board runs it across resets, so the common case
during use is to open the port and nothing else:
zig-out/bin/p4-console # the board is already programmed
zig-out/bin/p4-console --no-reset # ...and leave a live session running
`--no-reset` is the interesting one and it is proven on the die: attaching to the
running editor produced 309 bytes with no ESP-ROM banner and zero `boot:` lines,
then a `^` frame in response to typing. The session survived a detach and
reattach with no repaint, which is exactly what a 11.9 KB/s link wants.
`zig build console` is 4 steps and builds no image, no app and no object. It
installs `p4-console` itself rather than going through `installArtifact`, so
`zig build` alone still lands exactly one file in `zig-out` - verified from an
empty tree.
Also here:
* `tools/console.zig`'s `ModeWatch` tests were reachable by nothing. `zig build
test` runs them now; the matcher has to resynchronise when the byte that broke
a match is the next match's ESC, which is worth a regression test.
* The port-permission advice existed twice, in `failPort` and in the new program.
It now lives once, in `tools/serial.zig` beside the code that opens ports, and
says nothing about the path so each caller can name its own on the first line.
|
| |
|
|
|
|
|
|
|
|
| |
AccessDenied on /dev/ttyUSB0 is the first thing anyone hits, and "cannot open
/dev/ttyUSB0: AccessDenied" does not help. The steps that open the port now share
one failure path that names the group, and names the part that actually confuses
people: adding yourself with usermod -aG does NOT affect a shell that is already
running, because credentials are captured at login. So the same error greets
someone who has already fixed it. The message offers newgrp for this shell and sg
for one command.
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
`cpu-docs/` holds the documents behind the reverse-engineering in this repo, and
`cpu-docs/manifests/` records each source URL with a sha256, a byte count and a
page count. The manifests ARE the archive as far as this history is concerned:
121 MB of vendor PDFs are ignored, exactly as `/zig-pkg/` is, because a re-fetch
is one command away and the manifest makes a drifted document detectable.
The small text sources stay tracked instead of ignored, because Espressif
publishes them in no other form - the esptool serial protocol, firmware image
format and boot-mode selection pages, and the P4's custom PIE/SIMD instruction
reference. There is no PDF to re-fetch in their place.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
It works. `zig build interact -Dpardes` flashes the editor, attaches a terminal,
and typing changes the screen.
TWO FIXES, and the first is the one that mattered.
**Install mtvec.** The console went silent immediately after the editor's
allocators.init and nothing could explain it: bounded writes did not change it, no
Guru Meditation was printed, and execution did not return from pardes_p4_init even
when that function was made to return immediately after the marker that DID print.
Setting mtvec to this image's own handler, in DIRECT mode, fixed it - and the
handler never fires, which is the tell. hal/intr.zig:302 names the mechanism: the
CLIC can fetch handlers from MTVT instead of trapping to mtvec, and the bootloader
leaves that vectored mode on with a table this image does not own. `systimer.init`
is enabled a few lines earlier, so its first tick dispatched through a vector table
belonging to nobody. Writing mtvec with the low two bits clear selects direct mode
and the interrupt has somewhere legitimate to go.
That single register write took the port from "faults before the editor starts" to
the whole of pardes_p4_init succeeding, including vaxis's capability handshake going
out over the wire:
\e[?1049h \e[?1016$p \e[?2027$p \e[?2031$p \e[?2048h \e[6n \e[>q \e[?u
\e_Gi=1,a=q \e[c
alt screen, in-band resize, cursor report, kitty keyboard, kitty graphics, DA1 -
every one of them answered by the terminal emulator on the far end of the CH340,
which is the whole design.
**Clamp the grid, and make a failed resize atomic.** Pardes.init then returned
OutOfMemory with 9,128 bytes left of 393,216: every cell is paid for four times
(vaxis Screen + InternalScreen, pardes Surface + previous_cells). Measured: 40x12
initialises with room to spare, 80x24 does not. So max_cols/max_rows cap the
geometry and the host's larger terminal simply hosts a corner of itself.
The second half of that is subtler and cost a working editor. `Vaxis.resize` deinits
both screens BEFORE allocating the replacements (Vaxis.zig:194-206), so a failed
resize leaves vaxis with freed screens and renders nothing at all - and the host
bridge injects a size report on attach, so an unclamped 80x24 killed an editor that
had already drawn its interface. A failed resize now restores the previous geometry.
Measured end to end on the die: the first frame is ~1.5 KB of ANSI drawing the acme
tag bars ("New Newcol Joincol Find Grep Help Change" / "Save New Newtty Del
Filter"), and typing two characters produces a 116-byte incremental update that
draws the `^` modified-marker and moves the cursor to row 3 column 3. vaxis's damage
tracking is doing exactly what a 11.9 KB/s link needs.
Image is 598,528 B of the 1,536,000 B factory partition, 39%.
Also retired here: the I1..I10 and B1..B9 bring-up markers, the MMU table dump, the
cache before/after probe, the allocator and vtable pointer dumps, and the RX byte
probe. Each answered its question and each answer now lives in a comment beside the
code it explains. The trap handler stays - it is the diagnostic this port most
needed and did not have.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Housekeeping on the instrumentation, plus one fix that stands on its own.
uart.write and uart.writeByte no longer spin forever waiting for TX FIFO space.
hal/uart.zig:182-186 already made this argument about update() - "on a board with
no debugger an infinite spin is indistinguishable from a crash" - and this port
demonstrated it: the console going quiet mid-boot read as a hang in whatever code
came next, for hours, when a stalled transmitter would have looked identical. The
wait is bounded per burst and abandoned bytes are counted in `uart.dropped`, so a
lying console is at least a countable one. The bound is deliberately generous:
1,000,000 status reads against an 11 ms drain at 115200.
Removed, because each has answered its question and the answers are recorded in
comments where they matter:
* the MMU table dump - the table is CORRECT, entries 0..9 holding 0x1001..0x100a,
exactly the valid bit plus physical page N+1 that the image builder's anchor
requires. That is now stated in flushFlashCache's doc comment rather than
re-measured every boot.
* the before/after eviction read - it established that the same load returns
93 85 85 0f before a capacity flush and 3c ee 08 40 after, which is what the
image holds there. The flush itself stays; the proof of why it is needed is in
the comment.
* the allocator and vtable pointer dumps in src/p4.zig - they showed the struct
crosses the seam intact, with its function pointers landing in .flash.text.
* the bisect early return - it showed that execution does not come back from
pardes_p4_init at all.
What is left in place, on purpose: the I1..I10 markers inside pardes_p4_init and
the B1..B9 markers in the firmware. The port does not work yet and they are how
the next person finds out where it stops.
Where it stops: the console goes silent immediately after the editor's
pardes.allocators.init and never resumes. It is not the transmit spin - lowering
the bound to 20,000 and watching for 30 seconds changed nothing - and it is not a
trap the ROM can report, because no Guru Meditation is printed. Execution does not
return from pardes_p4_init even when that function is made to return immediately
after the marker that does print. So the CPU is lost inside a call whose only
work is handing a string to a function pointer, which points at the firmware's own
writeOut. That needs an instrument this setup does not have: JTAG, or a GPIO-based
tracer that does not depend on the UART at all. Everything cheaper has been tried
and is recorded above.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The image reads its own .rodata and gets its own .text back, and after ruling out
everything cheaper the answer is the cache.
What was eliminated first, each by measurement rather than argument:
* The MMU table is CORRECT. Read from the running application through
SPI_MEM_C_MMU_ITEM_INDEX_REG/CONTENT_REG (hal/esp32p4/mmu_ll.h:311-330),
entries 0..9 hold 0x1001..0x100a - the valid bit plus physical page N+1 -
which is exactly what the image builder's single flash-to-vaddr anchor
requires, and entries 10..11 are unmapped as they should be.
* The page size is not in question: hardwired to 64 KiB on this chip
(mmu_ll.h:126-130 returns MMU_PAGE_64KB and the setter asserts it), which is
what tools/image.zig already assumed.
* The flash is correct. The flasher verifies an MD5 of what the ROM stored, and
app.bin matches the ELF byte for byte at the addresses that misread.
* Not a write failure and not nondeterminism: identical across three resets and
two reflashes with the same MD5.
* Not 64-byte cache-line granularity either: the wrong bytes come in a
contiguous run of at least 192.
The measurement that settles it: a load at 0x40035a1c returned 93 85 85 0f, and
reading 512 KiB to force capacity eviction made the SAME load return 3c ee 08 40,
which is what the image holds there. So the second-stage bootloader hands over
with cache lines that do not match the mapping it finally installed. It is
perfectly deterministic - the bootloader does the same thing every boot, so it
leaves the same lines - which is precisely why it looked like anything other than
a cache for so long.
The ROM's own Cache_Invalidate_All (0x4fc00404, the same address in both
esp32p4.rom.ld and the eco5 table) would be the right instrument and is NOT used:
called from here it faults inside ROM code with its argument stranded in a2, so it
wants a precondition this image does not know. A capacity flush needs no such
knowledge, costs one pass over 512 KiB of already-mapped flash once at boot, and
is four times the 128 KiB the L2 measured at.
Effect: the firmware now gets through the editor's allocator round-trip and
pardes.allocators.init, which is two steps further than before.
Also here, and correct independently of any of the above: uart.write and
writeByte no longer spin forever on a stalled transmitter. hal/uart.zig:182-186
already made this point about update() - "on a board with no debugger an infinite
spin is indistinguishable from a crash" - and this port proved it by spending an
afternoon reading a stalled console as a hang in whatever code came next. The wait
is bounded and abandoned bytes are counted.
Still open: the console stops immediately after pardes.allocators.init. Bounded
writes did not change it, so it is not the transmit spin; there is no Guru
Meditation, so it is not a trap the ROM can report. The p4 allocator tier's
zero-capacity StackFallbackAllocators are the one unusual thing in that call and
their reasoning against lib/std/heap.zig is written down in src/allocators.zig,
but it has not been tested with a nonzero floor. The bisect markers are left in
place for that.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The editor arrives as one freestanding OBJECT exporting a seven-function C ABI
(src/pardes/app.zig declares it, ../02-pardes-code/src/p4.zig implements it), not
as a package dependency. A build.zig.zon path dependency was built first and
reverted: merely DECLARING it nested pardes's ~30-package graph under this one and
broke every build here - std/Build.zig:2091 exceeded its 1000-branch comptime
quota via ghostty's lazyImport, seven cached tree_sitter versions use APIs removed
in 0.16, and the fetch wrote 2.6 GB across 42,736 files into this working copy.
The seam is bytes in and bytes out, which is what a serial line is anyway: the
editor owns vaxis and the ANSI encoding, this side owns the UART, the heap and the
clock, and neither names the other's types. It is versioned, because linkers do
not type-check C symbols and a drifted signature would link cleanly and then
corrupt the stack.
THE BUG WORTH THE COMMIT. .flash.text was ALIGN(64), and the image builder's
anchor makes two mapped segments share an MMU page safely - as long as rodata does
not END inside the page where text BEGINS. With a 578 KB image it does. A volatile
read of a string literal at 0x4004a1d1 returned 37 09 fa 4f, which disassembles to
"lui s2, 0x4ffa0": this image's own .flash.text. Every literal in that last shared
page read as code, so the first thing the firmware tried to print was machine code
and it died on an instruction access fault. .flash.text is now ALIGN(0x10000),
making the segments page-disjoint. The packing trick this project opened with only
ever mattered when the alternative was 64 KiB of zeros in a 1 KB image.
Two more findings, both recorded in README.md:
* A linker symbol declared as an anyopaque OBJECT gives the optimiser a
zero-sized object, so ordinary stores through a pointer derived from its
address are dead code it may drop - and did, silently. The allocator's first
block header read back as size=2988759312 next=0x14284684 and the free-list
walk never terminated. @extern with a many-pointer has no size to lose.
examples/memprobe.zig could not have caught it: it writes through a volatile
pointer, which the optimiser must leave alone.
* The RTC watchdog is armed at handover. Every example here had been resetting on
a ten-second cycle, invisibly, because no run had ever lasted eight seconds.
State, honestly: the firmware boots, clears .bss, brings up the console, disables
the watchdog, starts the systimer, checks the ABI version, initialises the 384 KiB
heap and calls into the editor, which sets up its sink and its environment. It then
faults inside pardes_p4_init on the first allocation. The cause is measured but not
fixed: a load from .flash.rodata page 3 returns the contents of the page 0x50000
higher - exactly the vaddr distance between the rodata and text segments - while
pages 0, 2 and 4 read correctly. The bisect markers that localised it are still in
place, deliberately, because the next step needs them.
--- correction, measured after the above was written ---
Two mapped segments is NOT a choice, and the earlier comment in tools/image.zig
was right for a reason I initially got wrong and then measured.
I first read bootloader_utility.c's `#else` branch, which classifies segments by
address window with two independent ifs - and since the P4's DROM and IROM windows
are the identical range (soc.h:146-149), I concluded the last mapped segment wins
both roles and the first is never mapped. That branch does not run on this chip.
The P4 takes the SOC_MMU_DI_VADDR_SHARED branch (bootloader_utility.c:805-851),
whose own comment says it: "On chips with shared D/I external vaddr, we don't
divide them into either D or I, as essentially they are the same." It collects
mapped segments POSITIONALLY into rom_addr[2] and ends with
assert(rom_index == 2);
Shipping a one-segment image proved it, on the board:
Assert failed in unpack_load_app, bootloader_utility.c:842 (rom_index == 2)
So the split stays, image.zig keeps enforcing exactly two - turning that boot-time
abort into a build-time error - and both are now documented with the branch that
actually runs and the assert that actually fires.
What DOES change is alignment. .flash.text was ALIGN(64). Two mapped segments may
share a 64 KiB MMU page only if they also share a flash page, which the image
builder's anchor guarantees - and that holds right up until an application is large
enough for rodata to END inside the page where text BEGINS. With a 578 KB image it
does. Measured on the die: a volatile read of a string literal at 0x4004a1d1
returned 37 09 fa 4f, which disassembles to "lui s2, 0x4ffa0" - this image's own
.flash.text. Every literal in that shared page read as code, so the first thing the
firmware tried to print was machine code, and it died on an instruction access
fault. .flash.text is now ALIGN(0x10000), which makes the segments page-disjoint.
It costs up to 64 KiB of image padding against a 1.5 MiB partition; the packing
trick this project opened with only mattered when the alternative was 64 KiB of
zeros in a 1 KB image.
With that fixed the firmware gets much further: entry, .bss cleared, console up,
watchdog disabled, systimer running, ABI version checked, the 384 KiB heap
initialised, into the editor, its sink and environment ready - and the literal at
0x4004a1d1 now reads back correctly.
Still open, and characterised rather than guessed: pardes_p4_init faults on its
first allocation. The allocator struct crosses the seam intact (its function
pointers land in .flash.text), but the std.mem.Allocator vtable at 0x40035a1c reads
back as instruction bytes, and the dispatch at .flash.text+0xade2 jumps through it.
Ruled out with measurements: the ELF and the image agree at that address, the flash
is MD5-verified against the image, the wrong bytes are identical across three
resets and two reflashes (so not a stale cache), the corruption is a contiguous run
rather than 64-byte lines, and mmu_hal_map_region's arithmetic
(page_num = ceil(len/page), entry from vaddr) is correct for the segments as now
laid out. The next measurement is the one that settles it: read the MMU entry
registers from the running application and print vaddr -> flash for every page. The
register model in src/soc.zig can do that; the bisect markers are left in place for
it.
|
|
|
build.zig generates the linker script and drives Zig's own LLD; tools/image.zig
turns the ELF into a flashable image and tools/{rom,serial}.zig speak the mask
ROM loader over the UART. No CMake, ninja, idf.py, esptool, or external linker.
src/soc.zig is a comptime register model over ESP-IDF's own *_reg.h headers;
src/hal/ adds peripheral sequences; src/io/ implements std.Io for the chip;
src/oracle/ diffs this HAL against ESP-IDF's on the die.
|