<feed xmlns='http://www.w3.org/2005/Atom'>
<title>esp32p4.git/experiments/pg-6.png, branch main</title>
<subtitle>ESP32-P4</subtitle>
<id>https://git.0x4200.cafe/esp32p4.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/esp32p4.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/'/>
<updated>2026-08-26T01:41:28Z</updated>
<entry>
<title>Report: the pad target swept, and the byte-at-a-time read that was hiding in std</title>
<updated>2026-08-26T01:41:28Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T01:41:28Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=110eeeb80181b990901c22d28fa4fd41dd10cdd7'/>
<id>urn:sha1:110eeeb80181b990901c22d28fa4fd41dd10cdd7</id>
<content type='text'>
Brings the document to the end of the work. A keystroke is 3.65 ms against a 4 ms target,
and the closing section now reports it the way it should be reported: over 80
phase-randomised trials with the SLOWEST at 3,912 us, and broken out across document
length rather than as one intercept.

Three things this adds that are findings rather than steps.

## The pad target, swept

Two anecdotes disagreed about whether a bigger frame arrives sooner - padding the
cursor-positioning frame 21 -&gt; 49 bytes made it a millisecond faster, padding the
cursor-hiding frame 6 -&gt; 36 made it slower - so the target was swept as the only variable.
0 and 16 sit at 4.7-5.1 ms, 32, 48 and 64 all sit at 3.6-3.8. Crossing the packet boundary
is worth ~950 us and going past it buys nothing. 32 is no longer a fitted constant either:
it is wMaxPacketSize of endpoint 0x82 as the device reports it, and the sweep is what
confirms the descriptor is the thing to believe.

## The largest read in the firmware ran a byte at a time

The last win was not an algorithm. The shadow-grid diff - two 13 KB streams every frame,
comfortably the biggest memory access the firmware makes - ran at 3.2 cycles per byte,
about four times what word-wide loads need. `std.mem.eql` was the reason. Comparing a u32
at a time: 223 -&gt; 66 us, 0.95 cycles per byte, at every document length.

Recorded alongside it are the two candidates that were measured and REVERTED, which is the
more useful half: a row-at-a-time memset in Surface.fill plus a one-byte store in
Surface.set removed 52 million instructions per host run and zero cycles on either host or
board, and ablating the whole-surface fill priced it at 41 us. Writes on this part are
cheap; it was the reads that were slow.

## Where it stopped, honestly

The target holds wherever the editor actually SHOWS the keystroke - 3,602 us at an empty
line through 3,868 at 160 characters. At 320 and beyond the line has outgrown a 40x12
viewport, the cursor is off screen, and the keystroke changes no cell at all: 4,026 and
4,192 us to produce a frame of 36 bytes in which nothing changed. Those two columns are
now in the table with a "cells changed: none" row under them, because a round trip for an
edit that displays nothing is worth reporting as exactly that and not as a failure to hit
a number.

Also corrected: the ranked-recommendations table said the UART0 raise was abandoned
because a higher line rate moves settle and leaves the round trip alone. That reasoning
assumed the first byte reaches the host as soon as it is sent, and the bridge finding says
otherwise - delivery waits for 32 bytes, 278 us of wire at 115200 against 35 at 921600, so
the raise is worth ~240 us of round trip after all. It stays unapplied on the grounds that
one 115200-baud line is the premise of this port rather than a free variable, and meeting
the target by changing the link would answer a different question.

Ten pages. Every figure still computes from the raw per-trial CSVs, including the new ones.
</content>
</entry>
<entry>
<title>Report: the clock, the renderer, and a minimum frame owned by the USB bridge</title>
<updated>2026-08-26T00:40:14Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-26T00:39:03Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=6bb2023a419e56a50cd40fb407abd9dda250d088'/>
<id>urn:sha1:6bb2023a419e56a50cd40fb407abd9dda250d088</id>
<content type='text'>
Brings the document up to the end of Experiment 4. Two rows on the progression table -
360 MHz and direct emission - plus the sections behind them, the corrected verifier, and
a summary that no longer says the target was missed. It is met: 3.74 ms from 16.99.

The two findings worth more than the number, and both are written up as findings rather
than as steps:

A 4x clock bought 2.6x. The grid walk reads 27 KB a frame at about six cycles a byte, so
it is bounded by L2MEM bandwidth and does not care how fast the core runs. Predicted
before the measurement.

The direct renderer measured SLOWER at first - less computation, a quarter of the bytes,
worse round trip - because the CH340 forwards a bulk IN packet only when the packet is
full, and a 21-byte frame does not fill 32. It waits about a millisecond for a timer.
So the frame has a minimum size and it belongs to the transport, not the terminal; the
emitter pads to it with repeated cursor positioning. The table of 21/49/81-byte frames is
in the report because the MINIMUM column is the tell: the small frame's floor was already
530 us below vaxis's, exactly the compute saved, and only the median was hostage.

Also corrected: the verifier section. It described hashing SGR parameters per cell, which
is a history rather than a state, and that version reported a difference on the die that
did not exist - a faster board split the same keystrokes across different frames and
reached the same colours by another route. The document now says what it does instead,
and why the from-scratch renderer could not have been verified without the fix.

Table 13, the ranked recommendations from Experiment 3, is kept as the prediction it was
with a note on what has since been applied and that row 3 - raise the line rate - was
abandoned. The threats section no longer states the link floor as a single number, since
the bridge's packet granularity is now part of it and is specific to this bridge.

Nine pages. Every figure still reads the raw per-trial CSVs, including the two new ones.
</content>
</entry>
<entry>
<title>Report: the optimisation campaign, and the 3.1x that is left</title>
<updated>2026-08-25T23:33:51Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-08-25T23:33:51Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/esp32p4.git/commit/?id=7664c57d93cd3f3bdb719f7ca915d657668ddfd9'/>
<id>urn:sha1:7664c57d93cd3f3bdb719f7ca915d657668ddfd9</id>
<content type='text'>
Experiment 4 records what each cut was worth, measured on the die at every step, and
the stage table that made the cuts findable at all. Also two corrections to the
document itself:

The data plumbing loaded four of the eight datasets, so the progression table would
have been computed from a subset. It now loads all eight.

A continuation line beginning with `+` is list markup to Typst, so the CSV file names
were rendering as a numbered item at the top of page one. One expression, one line.

The summary no longer ends on the ReleaseFast flag as the best available change; it
ends on the measured 16.99 -&gt; 8.37 ms and on the fact that the target is not met, with
the remaining 6.42 ms of compute broken into the three items it actually consists of.
Saying 'not met' in the summary matters more than the table: the number that was
missed is the one a reader should see first.
</content>
</entry>
</feed>
