| Commit message (Collapse) | Author | Age |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
with no thread
Step 1 of the 9P chain (docs/9p.typ 12.1, docs/registry.typ 9P-14).
A detached session was the one configuration no script could drive. The core,
the panes and the undo history outlive every frontend that attaches -- and the
filesystem that would let a program read or change any of it was never mounted,
because `push_fs_reply` was one of the host methods this process left null.
Nothing prevented it; the call was simply not there.
It costs less here than in the desktop shells. They start a thread that blocks
on poll() and pokes a loop it does not otherwise share (`fs_service.wake`);
this process already runs ONE poll over its listener, its frontends, its pane
shells and inotify, so /dev/fuse is one more descriptor in the same syscall and
there is no thread at all. `Source.fuse`'s arm does nothing on purpose: being
in the set is the whole point, because the wake must end the sleep so that
`pollFrame` -- which runs after `pull_wait_input` returns, where re-entering
the core is legal -- reaches the drain.
`main.zig` refused `--detach --fs` outright, with a comment saying that
serving it would mean mounting FUSE in the detached core and that this was a
feature rather than a fix. It was right, and this is the feature. `--attach`
is still refused: a frontend has no core to serve.
Verified against the project's own clients: examples/acmefs/pardesctl panes,
new, send, body and del all drive a daemon, and the pane shells it forks now
inherit PARDES_FS/PARDES_PANE like every other host's.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Esc stops recentring
## A terminal row's ANSI colours survive being edited
The loudest colour bug this editor had: one keystroke anywhere in a coloured shell row turned EVERY
column of it grey. `EditAnchors` anchored a buffer line only when it was BYTE-IDENTICAL to the shell
row it stood over, so a single differing byte dropped the whole row's colour projection. Worst shape
is invisible: append past the pane's right edge, where the text is clipped, and the row looks the
same and only its colour goes.
Anchoring is byte-level now. An edit leaves the row's own bytes at both ends, and being the same
bytes they keep the same colours; only what was typed has no cell under it, so only that takes none.
Live, on real `fastfetch`: a 32-column blue run split into 6 + 26 around one typed character.
Three defects underneath it, all found by machinery rather than by reading:
* A JOIN removes a buffer line while the buffer's covered span grows, so `lines == covered` and both
aligned guesses — Nth line over the Nth covered row, and the same counted from the bottom —
resolved to the SAME wrong row. Every untouched row below a join went plain. Anchoring is now a
streaming monotone matching: one shell-row cursor that only ever moves forward, advanced once per
buffer line, linear in the buffer where the version before it was quadratic.
* An EMPTY line is not evidence. Splitting a row makes one, it equals every blank row in the span,
and left free to look ahead it claimed the blank row below the last output and took every coloured
row in between out of reach of the lines that owned them.
* Reflow under a scrolled viewport. `PageList.getTopLeft(.viewport)` returns the viewport pin
verbatim, x and all, while `PageList.pin` forces x to 0 — so after a reflow remapped a tracked pin
into the middle of a row, the text pass dumped row 0 from that column while the colour pass paired
the fragment with the row's FIRST cells. Row 0 wore its left half's colours until the pane snapped
back to live output. `bodyText` dumps from column zero now, which is also what ghostty's own
renderer draws.
Also here: DECSCNM (reverse video) was silently dropped whenever `tty_filter` was off, because the
raw path resolved a `.none` colour by role and never consulted the mode.
The test that found the first two is the one worth keeping: random editing against an ABSOLUTE
oracle — every row's own text names the colour it must have — because the differential oracle it
replaced was blind by construction. It skipped the edited row, which is the row the user is
complaining about.
## Esc returns to a pane without moving its view
Esc in body normal mode runs `Last`, "the pane you were in before this one", and that went through
`focusPaneLine`, which recentred a file on the target line unconditionally. So returning to a buffer
repainted the whole screen to show a line that was already on it.
`focusPaneLine` takes a landing now: `.center` for the three callers going somewhere you have not
been (a look target, a path a pane already holds, `@pN:LINE:COL`), `.keep` for Esc. `.keep` leaves
the view alone and lets `ensureCursorVisible` — which already existed and already scrolls by the
minimum into the `scroll_off` band — be the only thing that may move anything.
Not `line = 0`, which `focusPaneLine` already understands as "focus and touch nothing": a background
pane's view can move while you are away, because the wheel scrolls the pane under the POINTER and a
resize reveals no cursor, so the recorded cursor plus a minimal nudge is what actually gets you back.
Ctrl-o and Ctrl-i keep centring, and the asymmetry is structural rather than arbitrary: `Last` only
ever CROSSES panes, so the pane it lands on already holds the view you left it with, while `jumpBy`
can land in the SAME pane, where a long in-file jump would arrive on the very top or bottom row with
`scroll_off` lines of context on one side. Helix splits the same pair the same way — its jumplist
centres, its buffer switch does not.
One deliberate consequence: under `.keep` a PDF's page is not restored AT ALL, because a page reveal
IS that pane's view and a reveal of the page you are already on still snaps `document_scroll_y` to
that page's start, discarding where you had read to. When something moved the pane while you were
away — the wheel again — Esc leaves it where the wheel left it, and Ctrl-o is how you reach the
recorded page.
## host_io.zig: the machine-local half of a host, once
`host.zig` is the seam. The part of the answer that is identical on every host with an operating
system under it — fork a pane's shell, put bytes on a disk — was written FOUR times: in tty.zig,
gui.zig, macos.zig and detached/server.zig. What those copies had in common says what they were for:
all four were missing FD_CLOEXEC on the pty master, so in every shell pardes has shipped, a program
in one pane could read another pane's terminal.
One copy now, and the wire got smaller for it: `ServerMsg.spawn` is gone. A frontend never asked the
server to fork anything — the server has an operating system under it and forks through `host_io`
like every other host — and `decodeClient` lost the scratch buffer that message needed.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
board cap on one screen
## The wire is the effect stream, not a new protocol
`pardes --detach` leaves a core running with no terminal; `pardes --attach` is a frontend that owns
a terminal and a socket and nothing else. N frontends on one core all look at the same screen —
`screen -x`, not N sessions.
The codec (`src/detached/wire.zig`) carries exactly one `Event` or one `Host.VTable` call per
message. That is not a coincidence and it is why there is no third vocabulary to keep in step: the
core's IO seam was already a struct of function pointers with plain-data arguments, so a socket is
a legal implementation of it. `nested.zig`'s socket could not be reused — it carries a builtin
command line, and a command line cannot carry a frame.
ARCHITECTURE-NEUTRAL on purpose, not as decoration. The frontend on the far end may be
riscv32-freestanding on the ESP32-P4 while the core is x86_64 Linux, so every field is an explicit
little-endian fixed width and no message is a blit of a native struct. A protocol that only works
between two builds of the same compiler would have thrown away the one frontend that motivated it.
## The board comes in; its toolchain stays out
`src/p4.zig` becomes `src/esp32p4.zig`, and the pardes half of `../05-zig-p4` — the vaxis-over-
serial runner, the UART editor terminal, the keystroke rescue ring, the on-die test suite — moves
into `src/esp32p4/`. `build.zig.zon` gains `.zig_p4 = .{ .path = "../05-zig-p4" }`, so
`zig build -Dplatform=esp32p4 -Desp32p4-firmware` builds, flashes, monitors and self-tests the
board from this repo's `build.zig`.
The DIVISION is the point. What moved is what only pardes wants: the runner that drives a pardes
core over a serial line. What stayed is everything a second project would also want — the HAL, the
register/radio/oracle layers, the linker script, `_start`. `zig_p4` declares no dependencies of its
own and its `build()` early-returns when it is not the root package, so this costs the package
graph exactly zero packages and the editor's own builds nothing at all.
## limits.zig: nine forgettable places become one budget
Nine `platform == .esp32p4` capacity tests lived in nine files. They were never nine decisions —
they are ONE decision, how much memory this build may spend, taken nine times where no reader could
see the total. `src/limits.zig` puts the whole budget on one screen with every cap named against
what it is measured against, derived from two booleans.
The payoff is testability on a machine that is not the board: the caps are ordinary comptime values,
so a host build can be compiled against the board's numbers and the parking, eviction and clamping
paths a 240 KiB core takes get exercised by the normal test suite instead of only over a UART.
## A bare `zig build`
`zig build` with no arguments now builds the tty and GUI binaries and installs them into
`~/.local/bin`, and says so once on stdout with the flag that overrides it. The old default built
one binary into `zig-out` — a path nothing on a `PATH` ever looks at, which made "build it" and
"use it" two different commands for no reason.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
the P4
## Gpio
`Gpio 33` flips one pad and answers on the message row with what it did:
GPIO 33: 0->1
GPIO 33: 1->0
Bare `Gpio` draws the header instead, because the first question about a header is which pins it
has. The pin number is DECIMAL and it is the only literal in board_memory.zig that is - every
other one is an address, and addresses come off datasheets and linker maps that print hex, which
is why that file made everything hex two commits ago. A GPIO number is not an address, it is part
of a NAME: the schematic says GPIO47, the datasheet's pin table says 47, and `Gpio 20` meaning pin
32 would be a trap laid for the one argument anybody types from memory.
## The toggle is the host's, not the editor's
New `Host.VTable.pull_gpio_toggle`, and a `GpioFn` in the p4 ABI (hence version 2), rather than
board_memory reaching for GPIO_OUT the way `Poke` two functions above it would happily do.
Writing that register is not the job. A pad has to be pointed at the GPIO function in the IO MUX,
routed in the GPIO matrix, given drive strength and an input buffer with its pulls cleared, and
only then driven - four register files behind a per-pin table. That code already exists in
`05-zig-p4/src/hal/gpio.zig`, it is the same `configureOutput` the blink demo has always used, and
its register numbers are checked against ESP-IDF's own headers on the die by `zig build diff`. A
second copy inside the editor object would be a second copy under no test, and getting it wrong on
a pin that boots as something else is how you lose the console you are typing on.
Reported levels are the OUTPUT bits, before and after, because that is what a toggle means: the
level this board is driving. A pad's input buffer on an unconnected header pin reads the air.
## JP1, read off the schematic rather than remembered
The diagram is the vendor's own wiring, from sheet 2 "Expand IO" of
`01-esp32p4-m3/docs/JC-ESP32P4-M3_schematic.pdf` - the only document that carries this mapping. The
specification PDF's "Interface Description" page turned out to be a marketing render, and there is
no board user guide; the chip datasheet has a package pinout, which is not a header.
That sheet is a 872x1168 raster (`pdfimages -list` - the PDF embeds no vectors, so rendering it
larger adds nothing), and at that size the rows around pin 14 are genuinely ambiguous by eye. So
the mapping came from the drawing's geometry instead: thirteen wires leave each side of the symbol,
a net wire runs ~100 px to its label and a power stub ~21 px. Pin 8's wire is 21 px, which is what
identifies it as unconnected rather than as the first of the GPIO4x labels - the reading that had
GPIO47 one row higher and shorted GPIO45 to the ground bracket.
Cross-checked against a second source that has been in the tree all along: `05-zig-p4/build.zig`
documents `-Dled=20` as "JP1 pin 17", and GPIO20 lands on pin 17 here. Both facts are asserted in
the test, so the diagram cannot drift from either.
## Peek, Poke, Hexdump and Gpio are now the P4 build's alone
`board_memory.enabled` was `os.tag == .freestanding and !isWasm()`, on the argument that these
words are a property of having no operating system rather than a product configuration, and that a
predicate spelled out of `builtin` cannot drift the way a hand-maintained enum can.
Tidy, and it answered the wrong question. A word only exists if some shell offers it, and the
shells are the platforms. `Gpio` settles it beyond argument: its whole content is one board's
header, and a second freestanding port would need its own pinout rather than inheriting this one.
"Bare metal" was never the requirement, "this board" was, and the two only looked identical
because there is currently one of them. The old predicate's real work was excluding wasm -
`freestanding` too, where an address is an offset into a linear memory the engine owns - and naming
`p4` excludes it by construction instead of by a term somebody has to keep remembering. The target
is now the witness rather than the gate.
Absent means not compiled: the tty binary contains no `+Gpio`, no `+Hexdump`, no `ES_I2C_SDA` and
no `MisalignedAddress`.
## The boot buffer's lines are checked, not eyeballed
Three times now a line in that tour has been one or two characters too long for a 56-column grid,
and every time it was found by reading the die's screen - the expensive way to measure a string
literal. The text is a named `boot_buffer` with a test over it, six lines came down to fit with
margin, and the tour gained `Gpio`.
Tests: the pinout's width, its thirteen aligned pin rows, GPIO20-on-17 and pin-8-unconnected; the
decimal-versus-hex distinction; every boot-buffer line. Full suite green - unit-test, snap 95/95,
hxdiff 481/0, hxparity 561/0, image-harness, pdf-harness, mupdf-check - and tty, p4, gui,
p4 at 80x24, p4 with the fade forced on. On the die `p4-bench --check` is 5/5, the fifth being a
new one: three `Gpio 33` runs must report 0->1, 1->0, 0->1, because the alternation is the only
oracle a hardcoded string could not fake.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A theme change moves the anchored chrome palette - taglines, boxes, line numbers, scroll
bars - from the old colors to the new ones over ten display frames. On a screen that
repaints in microseconds that is a short legible transition, and it is why the code exists:
a palette that teleports reads as a glitch.
On a 115200 serial line it is not a fade. Each of the ten steps recolors every anchored
cell, so the diff finds the whole chrome dirty and spends a frame's worth of wire on it, ten
times over, with nothing else on screen to look at. Measured on the die, one `NextColor`:
fade on 12,593 bytes 1,097 ms of saturated wire
fade off 2,425 bytes 215 ms
A second of the editor talking to itself about a color, on the one transport where a second
is noticeable, for a gradient nobody can watch arrive at 11.5 KB/s.
## Comptime, so the code is not there
`ChromeAnimation` now selects between `animation.Transition` and a new `animation.Immediate`
- the same interface with the animation taken out, a value that is only ever what it was
last set to. That is what makes `ChromeTheme.interpolate` unreachable, and unreachable is
what makes it absent: the flashed image drops 2,336 bytes, and the object 13,180.
A bool tested at runtime would have kept every one of those bytes and still paid the
branch. It also would have needed a second meaning bolted onto `animate_theme_changes`,
whose job is the startup window and nothing else; that field is untouched here.
The option is `-Dtheme-animation`, defaulting to off for `p4` and on everywhere else, and it
is an option rather than a platform test because "is a frame expensive" is a property of the
transport: a P4 driven over something faster than a UART would want the fade back, and
`-Dtheme-animation=true` gives it to them.
## What was checked
`Immediate` is new code with one contract worth pinning, and it is the one a caller could
get wrong: it must arrive at the SAME palette a completed fade arrives at. An endpoint that
differed by a rounding step would make the option a change of colors rather than a change of
how long they take. Tested against a fully advanced `Transition` in `animation.zig`.
Full suite: unit-test, snap 95/95, hxdiff 481/0, hxparity 561/0, image-harness, pdf-harness,
mupdf-check. Builds: tty, p4, gui, and tty/gui with the fade forced off. On the die the
canonical verifier reports the screen IDENTICAL across both arms - the workload contains no
theme change, so this is the check that ordinary rendering was not perturbed - and
`p4-bench --check` stays 4/4.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Two columns back, and a better reason than the two columns.
Every number these words read is hex - there is no other kind, and they refuse a decimal
one - so a prefix on the output restates what the whole file already says. Dropping it buys
something worth more than the width: an address in a dump can be typed straight back into a
Peek without editing it, because bare hex is exactly what the parser now wants. Output that
is valid input beats output that is decorated.
No platform question to answer either: `enabled` is freestanding-and-not-wasm, so these
three words exist only on bare metal. There is no host format to stay consistent with.
On the die, 44 columns of a 48-column body:
40000020 32 54 cd ab 00 00 00 00 |2T......|
40000030 30 2e 31 00 00 00 00 00 |0.1.....|
5011002c: wrote deadbeef, reads deadbeef
5011002c: deadbeef
501101a4: fc48777d
501101a4: 4b4ae238
The last two are the same command twice - LP_SYSTEM_REG_RNG_DATA, which is what makes it
the honest demonstration that a register is not memory.
unit-test, and `p4-bench --check` 4/4 on the board.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
## Eight bytes a row on the P4
`hexdump -C`'s sixteen needs 79 columns: ten for the address, forty-eight of hex, a gap,
and eighteen of ASCII gutter. The board drives 56 columns of which seven go to the line
numbers, so every row wrapped onto a second display line and the columns stopped lining
up - which is the entire value of the layout. Eight fits in 46 and keeps every property
that matters, including a gap at the halfway mark, because the eye counts in fours and
eights rather than in sixteens.
Verified on the die:
0x40000020 32 54 cd ab 00 00 00 00 |2T......|
0x40000030 30 2e 31 00 00 00 00 00 |0.1.....|
That is the app descriptor: 0xABCD5432 and the version string, read out of flash by a
command typed with no 0x on either argument.
## The boot buffer is shorter, and its addresses are named
The first draft opened with four lines of prose explaining that there is no operating
system. True, unhelpful, and it cost a third of a fourteen-row window before the first
command. One header line earns its place; the rest of the screen is addresses.
The two LP registers at the end are now named, because they are named in ESP-IDF's own
headers and the names are the interesting part: 0x5011002c is LP_SYSTEM_REG_LP_STORE0, a
general-purpose retention register that holds what you put in it, and 0x501101a4 is
LP_SYSTEM_REG_RNG_DATA, the hardware random generator. Between them they demonstrate the
whole point of a volatile read - one address gives back what was written, the other never
gives the same answer twice:
Poke 5011002c deadbeef -> 0x5011002c: wrote 0xdeadbeef, reads 0xdeadbeef
Peek 5011002c -> 0x5011002c: 0xdeadbeef
Peek 501101a4 -> 0x501101a4: 0x0b099791
Peek 501101a4 -> 0x501101a4: 0xfc97f3b7
All four run on the die, all with bare hex. Peek and Poke had not been tested there before
this - only Hexdump had, which I had let stand as though it covered all three.
snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty/p4/gui.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
the bus
## Hex, always
Base-0 parsing accepted `0x4ff40000` and `1341390848` and refused a bare `4ff40000`, on
the grounds that guessing between hex and decimal would let one typo address somewhere
else entirely. The reasoning was sound and the conclusion was still wrong: the ambiguity
it guarded against is not a real one. Every address anybody has ever typed at these three
words is hex - it came off a datasheet, a linker map, or a previous dump's own output, all
of which print hex - so the base was never in doubt, and demanding `0x` on every one of
them was a toll on the common case to protect a case that does not arise.
The COUNTS go with them, and that is the part worth saying out loud rather than leaving as
a surprise: `Hexdump 4ff40000 100` shows 0x100 bytes, which is 256, not one hundred. One
rule for every literal beats two rules that each fit their own argument better, because
the second kind has to be remembered at the moment you are concentrating on something
else. What these words PRINT is hex too now, clamp notes included, so a number can go back
in where it came out.
## And the board boots into somewhere worth looking
The empty output buffer was honest and useless. The three words that make this port
interesting all take an address, and a board's address space is precisely the thing you
cannot guess - so the boot buffer is now a tour of it: the image's own rodata and code in
flash, the firmware's data and the editor's heap in L2MEM, the mask ROM, UART0, the
systimer, GPIO_OUT and an IO_MUX pad, and one harmless Poke.
Every address comes from this repository rather than from memory, which is what makes them
worth trusting: the flash and RAM figures are the linker script's own ORIGINs in
`05-zig-p4/build.zig`, and the peripheral bases are the `DR_REG_*` values `05-zig-p4/src/hal`
uses. Each command sits alone on its line because an argument list ends at the last
argument - a trailing comment would be `ExtraArgument` - so the notes go above the lines
they describe. Lines are kept inside 48 columns because the first draft wrapped every one
of them at the 56-column grid, which reads like a bug.
Verified on the die: the buffer renders one line per line, and putting the cursor on
`Hexdump 40000020 60`, selecting with `x` and pressing Tab opens a dump whose first bytes
are `32 54 cd ab` - 0xABCD5432, the ESP app-descriptor magic - with the version string
right behind it. Bare hex, no prefix, reading real flash.
snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty/p4/gui.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The 40x12 ceiling was never about the screen. It was about memory, and the comment above
`max_cols` said so: "every cell is paid for four times over: vaxis keeps a Screen and an
InternalScreen, pardes keeps its own Surface and previous_cells". Two of those four are now
dead weight - with `direct_emit` the emitter diffs the Surface against its own shadow and
writes the escapes itself, so vaxis's two grids are allocated, never read, and were the
largest single claim on a 384 KiB heap. `init` sizes them to ONE CELL. vaxis still does the
work only it can do: the alternate screen, the capability queries, and parsing everything
that comes back.
That removes the memory ceiling entirely - the heap now reports 336 KB free at every
geometry tried, including ones that used to fail - and leaves latency as the only limit,
which is the honest one: every frame walks the whole grid.
## Measured on the die, 0.87 us per cell
geometry cells round trip
40x12 480 3,628 us the old default
56x14 784 3,930 us the new one
56x16 896 3,965 us
60x18 1,080 4,114 us
64x20 1,280 4,281 us
80x24 1,920 4,809 us
100x30 3,000 5,743 us
120x36 4,320 6,923 us
140x42 5,880 8,310 us the largest that runs
160x48 7,680 links, then traps
200x60 12,000 does not link
56x14 is 63% more area and 40% more width than 40x12 and still holds the 4 ms this port was
built to. 56x16 was tried first: 3,965 us on the bench instrument but 4,029 on the
phase-randomised one, which is over, and the two instruments differ by about 50 us
systematically - so the wider grid went and two rows stayed behind. Width is worth more than
height for reading code.
The two failures at the top are worth naming precisely because they are different failures.
200x60 does not link: `.bss will not fit in region l2mem, overflowed by 76036 bytes`, that
`.bss` being the shell's shadow copy of the grid, sized at comptime. 160x48 links and then
TRAPS at boot - the same region pressure arriving at runtime as a collision rather than as a
diagnostic. Neither is a heap problem any more, which is the interesting part: the heap has
336 KB spare while `.bss` runs out.
`-Dp4-cols` / `-Dp4-rows` because none of the above is a constant. 80x24 is one flag away for
anyone who would rather have the classic terminal than the millisecond.
Verified at the new geometry rather than assumed: the A/B against the reference path - vaxis
rendering, full repaint, `shadow_grid` and `direct_emit` both off - is identical in every
cell, characters and resolved style. That matters more here than usual because the emitter's
column arithmetic has a special case at the last column, and 40 was the only width it had
ever been asked about. snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test, both A/B arms,
tty/p4/gui, and the board's own `p4-bench --check`.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The mouse did not work. Chasing that found something much larger: NO escape sequence
worked on this transport, and had not since the port began.
`vaxis.Parser` resolves a buffer containing nothing but 0x1b as the Escape KEY. That is
deliberate and correct for a terminal, where the kernel hands over a whole escape sequence
in a single read, so a solitary ESC really does mean somebody pressed Escape. A 115200
serial line hands over ONE BYTE AT A TIME - 87 us apart, an eternity to a loop running at
360 MHz - so the first byte of every sequence arrived alone and was resolved as Escape,
and the remaining bytes arrived as ordinary keys.
A mouse click therefore came through as TEN key presses: Escape, `[`, `<`, `0`, `;`, `1`,
`8`, `;`, `3`, `M`. The `0` among them is "go to column zero" in normal mode, which is
exactly where the cursor kept landing, and why the first attempt at this looked like a
coordinate bug. Arrow keys, function keys, and the host bridge's in-band resize reports
were all being taken apart the same way.
Longer partial sequences were never affected: the CSI scanner returns `n == 0` for "no
final byte yet" and the shell already keeps those bytes. Only the one-byte case needed an
answer, because it is the only one the parser answers WRONGLY instead of declining. So the
shell holds a buffer that is exactly one ESC and lets `pardes_p4_tick` release it after
10 ms - two orders of magnitude longer than the 87 us until the next byte of a real
sequence, and imperceptible to a person pressing Escape. The same trade every terminal
editor makes, for the same reason.
Finding it took instrumenting the ABI: printing `@tagName` of every event the shell
applied. Ten `key_press` where one `mouse` belonged is not a thing any amount of reading
the coordinate arithmetic would have shown, and I had already read it twice.
## Mouse reporting, and the 1003 that is not requested
With the sequences intact, `apply` already handled `.mouse` - it mirrors the tty shell - so
enabling reporting was the only missing piece. Spelled out here rather than taken from
`vx.setMouseMode`, which asks for `1002;1003;1004;1006`: 1003 is ANY-MOTION tracking, a
report per cell the pointer crosses with no button held. On a 115200 line that is dozens of
15-byte reports for one sweep, arriving as input the editor must parse while it paints, and
arriving whether or not anyone wants it - moving the mouse over the window would starve
typing. 1002 reports presses, releases and motion while a button is held, which is exactly
what a click and a drag-select need.
Verified on the die: a click at column 12 puts the cursor at column 12 and one at column 22
puts it at column 22, a drag paints a selection, and the wheel scrolls. A press alone paints
the new position and then reverts - the caret does not move until the gesture ends - so the
release is what commits it, which cost an hour of believing a working click was broken.
Screen byte-identical to the vaxis reference, round trip median 3682 us against 3682, snap
95/95, hxdiff 481/0, hxparity 561/0, unit-test, tty/p4/gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
bridge on every frame
Two findings, both in the P4 shell's own `present`.
## std.mem.eql was the largest read in the firmware, one byte at a time
The shadow-grid diff compares each row against the previous frame: two 13 KB streams,
every frame, and by far the biggest memory access the firmware makes. It measured 3.2
cycles per byte, which is about four times what word-wide loads need - the shape of a
byte-at-a-time loop, and `std.mem.eql` is what it was.
`sameBytes` compares a `u32` at a time and falls back to the byte loop when the spans are
not aligned for it. The alignment test has to be a RUNTIME one because `Cell` is all `u8`
fields and therefore has alignment 1: whether a row begins on a word boundary is a
property of whoever allocated the Surface, not of the type. A row is 40 cells of 26 bytes,
divisible by four, so an aligned base makes every row aligned.
The answer is bit-for-bit the same - this is still exact byte equality - so it keeps the
property the whole diff rests on: byte equality implies visual equality, so the diff can
never claim two different cells are the same.
Measured on the die: the grid walk 223 -> 66 us, 0.95 cycles per byte. 157 us off every
keystroke at every document length, and the single largest win since the clock raise.
## A frame that only hides the cursor still has to fill a USB packet
The padding added for the bridge's 32-byte bulk-IN packet covered the branch that
positions the cursor and not the branch that hides it. A frame that only hid the cursor
was six bytes and waited out the bridge's timer. Hiding an already-hidden cursor is as
idempotent as positioning it twice, so it pads the same way.
The packet size is no longer inferred from an experiment either: 32 is `wMaxPacketSize` of
endpoint 0x82 as the device reports it, and the sweep over pad targets confirms what it
implies - 0 and 16 sit at 4.7-5.1 ms, while 32, 48 and 64 all sit at 3.6-3.8 ms. Crossing
the boundary is worth about 950 us; going past it buys nothing.
## Result
length 0 20 40 80 160 320 640 chars
RTT 3602 3624 3638 3790 3868 4026 4192 us
Fixed cost 3652 us against a 4 ms target, from 16.99 ms where this started. A
phase-randomised instrument agrees over 80 trials: median 3687 us, minimum 3571, maximum
3912 - every trial under 4 ms.
The two lengths still above 4 ms are the ones where the line has outgrown the viewport, so
the cursor is off screen and the keystroke changes NOTHING: the frame is 36 bytes of
cursor-hide and padding, zero cells changed, while pardes still rebuilds all 480 cells of
the Surface for 858-985 us. That is the one architectural item left and it is not a micro
-optimisation: nothing in this repository can avoid work pardes has already done.
Verified: screen byte-identical to the vaxis reference on the 18-step workload, with
canonical style decoding rather than escape history. snap 95/95, hxdiff 481/0, hxparity
561/0, unit-test, both A/B arms build, tty, p4 and gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`fitEnd` decides where a wrapped row breaks, and it asked two function calls per
character to learn what arithmetic knows. `modal.nextGrapheme` and
`graphemeDisplayWidth` each already answer ASCII in constant time - that was earlier
work - but they answer once per character, and a 640-column line asks 640 times.
A printable ASCII byte whose successor is also ASCII is a complete grapheme cluster one
column wide. That is the same guard, for the same reason, as the three fast paths already
in `Surface.print`, `modal.nextGrapheme` and `graphemeDisplayWidth`: every rule that
could join an ASCII base into a longer cluster - Extend, ZWJ, SpacingMark, Prepend,
Regional_Indicator - is spelled with non-ASCII scalars. Tabs and the C0 controls are
excluded by the range test and keep the general path, as does anything wide.
Measured on the die: 21 us of a 640-character keystroke, 5 us at 160. Small, and reported
as small - the interesting part is that it is small, because it says the per-character
grapheme walk was NOT where a long line's cost lives.
The test pins the fast path to the general walk it replaces rather than to transcribed
expectations: same inputs through both routes, every start offset, every width from zero
to past the end, over strings chosen to land the boundary inside a combining sequence, a
wide glyph, a regional-indicator pair, a tab and a CR. A break that moved by one column
would move text on screen, so this is the invariant worth holding.
|
| |
|
|
|
|
|
|
|
|
| |
Five digits is every u16, so the `v >= 10000` branch could never be taken and it was
dragging `std.fmt.printInt` into a firmware whose whole reason for hand-rolling this
was to keep the format machinery out of the hottest sequence it emits.
Behaviour is identical, and re-verified rather than assumed: screen byte-identical to
the vaxis reference on the 18-step workload, round trip median 3830 us over 60 trials
(3829 before), snap 95/95, hxdiff 481/0, hxparity 561/0, unit-test.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`present` already knows exactly which cells moved - that is what the shadow grid is
for - and then handed every one of them to vaxis so that vaxis could work it out again
against its own copy. That second diff measured 631 us of a 4.37 ms keystroke, all of
it redundant. This emits the escapes itself and skips it.
The emitter is small because it is allowed to be: one absolute CUP per run of changed
cells rather than per cell, absolute SGR rather than a delta from whatever is currently
on, and a hand-rolled two-digit formatter instead of `std.fmt` for the sequence it
writes most. Absolute SGR is the interesting choice - it costs a few bytes on a style
change and buys the property that no cell can inherit an earlier cell's colour if a
frame is cut short. Cursor column tracking gives up after anything that is not a single
printable ASCII byte, and at the last column, because deferred wrap makes the answer
terminal-dependent and wrong by a whole row.
Board cost: `render` 1362 -> 779 us. Bytes per keystroke: 81 -> 21. Image 26.6 KB
smaller, since vaxis's renderer is now unreachable.
## And it measured SLOWER
4.72 ms against 4.37. Fewer bytes, less compute, worse round trip - which is the sort of
result that means the model is wrong, so I stopped optimising and went looking.
It is the USB bridge. The board talks to the host through a CH340, a full-speed part
whose bulk IN endpoint carries 32-byte packets, and it forwards a packet when the packet
is FULL. A 21-byte frame does not fill one, so it sits in the bridge until an internal
timer gives up waiting for more - about a millisecond, a quarter of the whole budget.
Routing through vaxis only looked competitive because its frames are 81 bytes and fill a
packet by accident.
The evidence, all at identical board cost and with a byte-identical screen:
frame min median
21 B 3843 us 4817 us never fills a packet
49 B 3719 us 3814 us padded past the boundary
81 B 4373 us 4475 us vaxis, fills one by accident
Note the minimum: the 21-byte frame's floor is already 530 us below vaxis's, exactly the
compute that was saved. Only the median was hostage to the timer.
So the frame has a minimum size and it belongs to the transport, not the terminal. Pad
to it, with repeated absolute cursor positioning: idempotent, already the sequence the
frame ends on, cannot alter a cell. Every emitted byte goes through one counting helper
so the epilogue knows how much is owed. This is an Ethernet runt frame - the medium has
a minimum and the sender pays it - and it is a real trade rather than free, since the
filler is wire time that delays a later frame. It only applies when the frame is small,
which is when there is wire to spare.
## Result: 3.74 ms, and the goal was 4.00
step fixed per char at 160 chars
ReleaseSmall 16.99 ms 54.3 us 25.56 ms
ReleaseFast 14.85 ms 34.7 us 20.30 ms 0.79x
+ ASCII grapheme 14.56 ms 12.0 us 16.46 ms 0.64x
+ ASCII print 14.27 ms 6.9 us 15.36 ms 0.60x
+ shadow grid 8.87 ms 7.3 us 10.02 ms 0.39x
+ byte compare 8.37 ms 7.1 us 9.48 ms 0.37x
+ 360 MHz 4.37 ms 1.9 us 4.67 ms 0.18x
+ direct emit 3.74 ms 2.0 us 4.06 ms 0.16x
35 bytes per keystroke, down from 81. A phase-randomised instrument agrees: 60 trials,
median 3829 us, min 3722, p90 3930.
That second instrument exists because of this commit. The original bench sends keystrokes
on a fixed cadence, which locks the send phase to the host's 1 ms USB frame clock and
makes the round trip a staircase in board time - a real saving can measure as a
regression. Sleeping a uniform random 0-2 ms before each keystroke decorrelates the two.
It was not what was happening here, but it had to be excluded before the CH340 could be
believed, and it is the right default for anything measured across this link.
## Verification
`direct_emit = false` routes every cell back through vaxis and is the reference. Both
arms, same 18-step workload, same clock: identical characters and identical resolved
style in every cell - resolved, not raw SGR, because two emitters reaching the same
colour by different escapes are the same screen. A from-scratch ANSI emitter is exactly
the change that can be right about latency and wrong about the screen, and until the
verifier compared canonical style rather than escape history it could not have told the
difference.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
both A/B arms build, tty, p4 and gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Surface.cells` is contiguous and row-major, so a row is a single `memcmp` against the
shadow grid - and on a keystroke eleven of twelve rows are untouched. The per-cell
loop was ~40 branchy comparisons per row where this is one call over 1,120 bytes.
Byte equality implies visual equality, which is what makes the shortcut sound: a row
that compares equal cannot be hiding a changed cell, and a row that differs only in
padding falls through to the per-cell path, which is correct and merely slower.
Measured on the die at 360 MHz: the grid walk 246 -> 226 us. That is a small win and
the reason is worth recording - at 27 KB read per frame and about 6 cycles per byte,
this stage is now bounded by L2MEM bandwidth rather than by comparison work, so there
is little left in it. It is also why board compute scaled 2.6x rather than 4x when the
core clock went up 4x.
Verified with a canonical-style A/B: reference path (`shadow_grid = false`) and
incremental path, same 18-step workload, same clock - identical characters and
identical resolved style in every cell. snap 95/95, hxdiff 481/0, hxparity 561/0,
unit-test, tty and p4 both build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`Cell.visuallyEqual` is the semantically exact answer and too slow to ask 480 times a
frame: `std.meta.eql` on a `CellStyle` recurses through a colour union and eight
booleans, and the walk measured 1.45 ms on the die - about 270 cycles to compare a
28-byte struct.
`sameCell` in src/p4.zig does it as bytes. That is safe in the direction that
matters: byte equality IMPLIES visual equality, so it can never claim two different
cells are the same. It can miss an equality - scratch bytes past `len`, or padding -
and the only cost of that is one redundant `writeCell` which vaxis then diffs away.
Defaults are still compared by meaning, because an unpainted cell's text and style are
whatever the previous frame left in them.
Measured: the grid walk 1.45 -> 0.98 ms, a keystroke 8.87 -> 8.37 ms fixed.
Verified the way a rendering change has to be. The A/B harness now hashes the SGR
state of every cell as well as its character, because the first version compared text
only and would have passed a colour regression in silence. Reference path
(`shadow_grid = false`, clear and write everything) and incremental path were each run
against the same 19-step workload on the die and the reconstructed screens are
identical in both text and per-row style hash.
snap 95/95, hxdiff 481 cases 0 mismatches, hxparity 561 cases 0 mismatches, unit-test,
and tty / p4 / gui all build.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A keystroke on the ESP32-P4 cost 17.0 ms and the goal is 4. Profiling the core in
that board's exact configuration - 40x12, tree-sitter disabled, via `zig build perf
-Dtree-sitter=disabled -- --cols 40 --rows 12 --only small` - named the cost, and it
was Unicode machinery answering questions about the letter `y`.
Four changes, each a fast path guarded so that non-ASCII text takes exactly the road
it took before.
`modal.graphemeStart` was 21.5% of a keystroke, the single largest item. It iterates
graphemes FROM THE START of the text with the full UAX #29 break state machine until
it passes the offset, and the render path calls it once per visible row with a column
offset - so the cost followed the cursor's distance along its line. That is the shape
measured on the die, where inserting at column 320 of a fixed 320-character line cost
7.8 ms more than inserting at column 0 of the same line. In UAX #29 every ASCII
scalar is its own cluster with ONE exception, GB3 (CR joined to LF); every other rule
that could extend a cluster - Extend, ZWJ, SpacingMark, Prepend, Regional_Indicator -
is spelled with non-ASCII scalars. So an ASCII byte whose predecessor is also ASCII,
and not that CR-LF pair, IS a boundary. O(1), and sound rather than approximate.
`Surface.print` then became the largest at 26.2%: per character it took a UTF-8
length, a decode, a FRESHLY CONSTRUCTED grapheme iterator, a slice validation and a
width lookup, to conclude that `y` is one cell. Printable ASCII followed by ASCII
takes none of that now. Same guard, same reason.
`file_pane.graphemeDisplayWidth` was 6.9%, essentially all of it asking `gwidth`
about ASCII. Bounded to 0x20..0x7e on purpose: DEL and the C0 controls are not one
printable cell and `gwidth` stays the authority on them.
`modal.lineSlice` searched for "\n" with the generic substring search where a memchr
does; it is called once per visible row per frame.
Measured at the P4's geometry and configuration, on the host: render 55 -> 12 us,
key-down 483 -> 24 us, key-right 327 -> 13 us, edit-char 205 -> 46 us. On the die,
the per-character cost of a keystroke fell from 54.3 to 6.9 us - 7.9x - and a
keystroke at a 160-character line from 25.56 ms to 15.36 ms.
## The shadow grid, and why it is static
`src/p4.zig`'s `present` copied all 480 cells into vaxis every frame, which measured
6.75 ms on the die - 57% of a keystroke - and was paid whether or not anything
changed: a second render with nothing new cost the same as the first. vaxis diffs its
own grid, but only after being told every cell, and being told is the expensive part.
So `present` now keeps the previous Surface and tells vaxis only what moved.
`Cell.visuallyEqual` is the right comparison and already existed. Copy: 6.75 -> 1.45 ms.
The grid lives in `.bss`, sized by `max_cols` x `max_rows` at comptime, and that is
not a micro-optimisation. The first version allocated it from the editor's heap; on a
board whose 384 KiB is nearly spoken for, that is exactly the kind of change that
works and then breaks something else three steps away.
`shadow_grid` is a comptime A/B switch, kept deliberately. With it false, `present`
behaves as it did before - clear and write every cell - which is the reference any
measurement should be compared against, and the way to tell a rendering bug from a
rendering difference. It earned its keep immediately: the two paths were run against
the same 19-step workload on the die - inserts, deletes, motions that move the
modified-marker, a line outgrowing the viewport, backspaces that shrink it - and the
reconstructed screens are byte-identical.
## Verification
`snap` 95/95 scripts, `hxdiff` 481 cases 0 mismatches, `hxparity` 561 cases 0
mismatches, `unit-test`, `image-harness`, `pdf-harness`, `mupdf-check`, and tty / p4 /
gui all build. The rendering changes are exactly the sort that pass a latency
benchmark while corrupting a screen, so the snapshot parity suite is the one that
matters here and it is unchanged.
`test/perf.zig` gains `--cols`/`--rows`/`--only`. The screen's shape is one of the
things that table exists to hold constant, and 40x12 is not a scaled guess at the
board - it is the board. `--only` exists because under `perf record` one 63 ms cell on
the largest fixture swamps every sample from the case being asked about.
## Found, not fixed
`vx.resize` fails on this board: a runtime geometry change hits its allocation
failure path, restores the previous size and returns, so 80 bytes go out where 1,392
should. Verified independent of everything above - it reproduces with `shadow_grid`
false. The board therefore has one geometry for the life of a session, which is why
the staleness test above compares two firmwares rather than resizing one.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The measurement said a keystroke costs 54 us per character already in the line. The
source said why: `insertAt` called `lineCount` - `std.mem.count` over every byte -
TWICE merely to clamp a row, and then `spliceAlloc` allocated and copied the whole
document. Three whole-document passes before one character can be inserted, which
fits a linear slope exactly.
So it was fixed. `modal.lineSpan` finds a row's byte span in ONE scan that stops at
that row, and `insertAt` uses it; a full count is paid only on the rare clamping
path where the cursor is past the end. Two tests pin the equivalence, including the
edges that make line counting awkward - an empty document, a trailing newline (its
own empty last line), and a row past the end. The first version of `lineSpan`
disagreed with `lineCount` about an empty document and the test caught it.
On the host harness this is a real win, reproduced over three independent runs at
matched sample counts:
edit-char 1k lines 50k lines 300k lines
before 730 us 2996 us 15030 us
after 721 us 2414 us 10955 us
ratio 0.98x 0.79-0.82x 0.78-0.83x
Every operation I did not touch stayed at 1.00x, which is better evidence than any
single cell.
On the board it changed NOTHING. The slope was 54.3 us/char before and 54.0 after,
a ratio of 1.00 over 5 conditions x 7 trials. Not a contradiction - the same fact
seen twice. The removed passes are O(document), and this board's document is a few
hundred BYTES, so two scans of it cost nothing worth measuring.
## Where the time actually goes
`-Dprof` times the two phases on the die with the cycle counter around
`pardes_p4_input` and `pardes_p4_render`:
chars in line input (parse+edit) render
1 220 us 14804 us
80 212 us 17114 us
240 250 us 24615 us
Input is FLAT at ~220 us - 1.5% of a keystroke - and does not grow with the document
at all. The ~15 ms floor and every microsecond of the slope are inside `render`. The
edit path could be made free and nobody would notice.
`soc.flushFlashCache`'s home in soc.zig is what let the profiling build exist at all
alongside the responder; `-Dprof` defaults off because it puts a line on the wire per
frame, which is the resource being measured.
## The report
`experiments/report.typ` gains Experiment 3 and, more importantly, a correction:
Experiment 2's mechanism claim was wrong and now says so, with the disproof next to
it. The ranked recommendations are reordered - the renderer is now #1 and the change
this commit makes is listed unranked, because on this target it buys nothing, which
is exactly why it is worth recording.
The position table earns its place there too: in a fixed 320-character line, an
insert at column 320 costs 33.9 ms and emits 28 bytes, while one at column 0 costs
26.0 ms and emits 81. Output size and latency are not merely uncorrelated on this
board, they are inverted - which is the signature of a walk from the start of a
line, and the next thing to go looking for.
The lesson is the one the instrument exists to enforce. A plausible mechanism, read
off the source and consistent with the shape of the data, was wrong about where the
time went, and only a measurement inside the firmware could say so.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
`-Dplatform=p4 -Dtarget=riscv32-freestanding` emits a single freestanding OBJECT
exporting a seven-function C ABI, not an executable. The board's toolchain
(../05-zig-p4) owns `_start`, the linker script and the UART driver and links this
in. The seam is bytes rather than types, so neither side can accidentally depend
on the other's internals, and a signature that drifts fails at link time.
The serial line is the whole of the I/O. `src/p4.zig` drives vaxis unchanged over
it: the renderer is a byte writer and `queryTerminalSend` is a byte writer, so the
terminal emulator on the host answers the capability handshake and the firmware
sees a real terminal. Measured going out over the wire on attach: alt screen,
in-band resize, cursor report, kitty keyboard, kitty graphics, DA1.
THREE WORDS EXIST ONLY HERE. `src/board_memory.zig` implements `Peek`, `Poke` and
`Hexdump`, gated on `builtin.os.tag == .freestanding and !isWasm()` - derived from
the TARGET, because they are a property of running with no OS under you rather
than a product option, and because wasm is freestanding too and is exactly what
must be excluded: in a browser an address is an offset into the linear memory this
editor's own heap lives in. Every access goes through `*allowzero volatile`: a
peripheral register is not memory, and address 0 is an ordinary unmapped address
on this bus. One 4 KiB cap per command, set by the console rather than the memory -
an unbounded dump would wedge the only console the board has for eleven hours.
Measured on ESP32-P4 rev v1.3 silicon, driven from a host terminal:
Peek 0x501101a4 0x0e63ce71, then 0xaeaa6919 on a second read - the
RNG register, so the volatile loads are not folded
Poke 0x5011002c 0xdeadbeef LP_STORE0; a later Peek returned 0xdeadbeef
Hexdump 0x5011002c 32 16 bytes a row, hex columns and an ASCII gutter
Peek 0x50110001 `peek: MisalignedAddress` on the message row
That last line is the one that matters. A misaligned 32-bit access traps, and a
trap in firmware is a watchdog reset that takes the session with it, so the check
that turns it into a message is the reason the file is hand-written rather than a
generic reader.
BARE METAL BOOTS AN EMPTY OUTPUT BUFFER. Every other boot layout in `init` makes a
shell, and on this platform that is not a preference but an impossibility: nothing
to fork, no pty to give a terminal pane. Booting one anyway produced precisely what
that describes - a pane whose tag ends in `Filter`, no gutter, no buffer, and every
keystroke vanishing into the Fallback's silent pty. An output buffer is also what
the platform's own words want, since Peek, Poke and Hexdump each fill one.
Sized for the board rather than for a desktop:
* `allocators.zig` gains a p4 tier that is ALL fallback - every capacity is zero,
so each arena spills immediately to the 384 KiB heap the firmware hands over,
and no megabyte-shaped static reservation lands in `.bss`.
* `source_manifest.zig`'s allowlist is EMPTY on p4. The table is ~0.95 MiB of
rodata against a 1.5 MiB flash partition; the firmware's filesystem is the
serial host's, through the Host vtable.
* The grid is clamped and the clamp is measured, not guessed: every cell is paid
for four times (vaxis Screen + InternalScreen, pardes Surface + previous_cells),
so 40x12 fits and 80x24 exhausts the heap during `Pardes.init`.
* `Vaxis.resize` deinits both screens before allocating replacements, so a failed
resize leaves vaxis rendering nothing. The p4 shell keeps the previous geometry
on failure instead of leaving a half-applied one.
Also here: `output_pane_integration_test.zig` had an exhaustive switch over
`Platform` that adding `.p4` left unhandled, which broke `zig build unit-test`
outright - the native test binary is the one consumer no platform build compiles.
346 tests pass again.
|
| |
|
|
| |
stops rewriting the suite
|
| | |
|
| |
|
|
| |
takes a path argument
|
| |
|
|
| |
walk wrapped rows
|
| |
|
|
| |
optional methods
|
| |
|
|
| |
backends agree
|
| | |
|
| | |
|
| | |
|
| |
|
|
| |
docs
|
| |
|
|
| |
additions, unit tests
|
| |
|
|
| |
over grid cells
|
| |
|
|
| |
+ snapshot refresh
|
| | |
|
| |
|
|
| |
snapshots
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
glslc is the one build input that wants a tool a stock machine does not have,
and it is also the input that changes least often: eight GLSL files that have
outlived several rewrites of everything around them. Asking every machine that
wants to run the SDL shell for shaderc is the wrong trade.
The SPIR-V is now COMMITTED, under shaders/prebuilt/, and -Dprebuilt-shaders
embeds that copy instead of shelling out. The default stays the honest one --
compile the shaders that are actually in the tree -- because the flag trades a
dependency for a freshness problem: with it on, the .glsl sources are not build
inputs at all, so editing one changes nothing.
`zig build shaders` is the other half, and it is deliberately independent of
-Dplatform: it recompiles every shader and writes the result back into the
tracked directory, so whoever changes a shader refreshes the cache on a machine
that has the compiler and commits the diff. `jj diff shaders/prebuilt` after it
is the freshness check -- empty means the cache was already current.
The shader list is also spelled once now (gui_shaders): the eight embeds, the
eight glslc runs and the refresh step all read it, so adding a shader is a name
there plus the @embedFile in gui.zig, not three edits in two places.
Verified: -Dplatform=gui -Dprebuilt-shaders builds with glslc absent from PATH,
and image-harness passes on that binary -- real SDL GPU pipelines built from the
committed SPIR-V, 512 source pixels read back. The default gui build still runs
the eight glslc steps; tty runs none. The committed bytes are identical to a
fresh glslc run, and `zig build shaders` is idempotent.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The documentation had gone stale in the ordinary way -- claims that were true
when they were written and that nothing since had been obliged to re-read.
Some of them were load-bearing.
THE TUTOR. It still said there is no multi-cursor, that NextColor cycles three
themes, and that its practice blocks "are also run as unit tests (generated
from this file by tutor_gen)" -- a tool that appears nowhere in the tree, and
nothing anywhere parses a `# keys:` block. Left alone, that claim is what
makes the next wrong block survive.
Three of those blocks WERE wrong, and all three for one reason: since the
helix motion model landed, w/e/f/t SELECT the range they cross, so `i` after
one inserts at the SELECTION'S START. `w i Z esc` on "foo bar" gives
"Zfoo bar", not the "foo Zbar" the file promised. They were written against a
vim reading of the same keys. Every block in the file has now been run through
`zig build hxdiff` against the real core and matches byte for byte, and the
trap itself is written down in 3.3 rather than left to be rediscovered.
The tutor gains a PART 4 for everything added since it was written -- PDF
panes, the in-process ZLS backend, themes and fonts, the startup file -- and
PART 3 gains counts (and which keys ignore one), f/F/t/T, the whole g table
(bare `G` is a no-op; `ge` is the START of the last line), multiple cursors
and the s/S regex pair, `m`, `]`/`[`, `|`, insert mode, and all fifty leader
paths.
THE REST. design.typ's line table claimed 7,626 lines against a real 38,048,
and its rows did not sum to its own total; its Event/Effect boundary contract
-- the part a shell author writes against -- named four variants that do not
exist and omitted fourteen that do. lsp.md's probe count. config.md's
theme-name rules, which as written could not reach a zed theme at all.
helix-keys.md's Skipped section, holding five families that have since landed.
macos.md's menu bar, undocumented, along with sixteen other claims. web.md on
what the browser build can actually do.
SOURCE COMMENTS that had rotted alongside them: `tag_normal` is a space, not
the `•` its own comment describes; Wrap is ON by default, not off; a FontSel
row is SELECTED by n and RUN by Tab, not run by n; the SPC paths in lsp.zig
lost their `l` group prefix when the language group moved; and the
differential suites are 481 and 561 cases, not 360 and 440.
TWO THINGS FOUND BY DOCUMENTING THEM, both left standing and written down
rather than papered over. Typing `[^\n]` at an s/S prompt panics: the live
preview compiles every prefix, and `[^\` indexes an empty slice in mvzr's
parseCharSet. Both the tutor and a waiver recommended that pattern as the
workaround for `.` matching a newline; they now say what it costs and what
would make it sayable. And `Exec` is a builtin, so an `Exec` line in the
startup config types that command into a shell before the first frame -- the
tutor said nothing in that file is ever sent to one.
Nine adversarial reviews over two rounds, each with the hxdiff harness to
execute what it doubted. The second round exists because the first round's
fixes needed checking too, and it caught three regressions of my own -- one of
them a probe count I had "corrected" away from the truth.
Verified: unit-test, snap 87/87, hxdiff 481/0, hxparity 561/0, mupdf-check.
docs/design.pdf regenerated. The tutor's first seventeen lines are byte-
identical, which is what tutor.golden pins.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Three changes that all turned out to be the same shape -- a feature that
worked in one direction, or for one pane kind, and quietly did not in the
others.
CLIPBOARD. Every register write emitted set_clipboard, so deleting one
character threw away whatever the desktop was holding; multi-cursor yank took
the join's early return and emitted nothing at all, so the same key reached
the clipboard on one cursor and not on two. Nothing could READ the clipboard:
the SDL shell had no SDL_GetClipboardText anywhere in it, and the tty shell
never asked for OSC 52, so `p` from another application was dead in both.
Now it is helix's split. y/d/c/p/P/R and the acme chords are the DEFAULT
REGISTER and nothing else; the system clipboard is five words on helix's own
letters -- SPC y, SPC Y, SPC p, SPC P, SPC R -- spelled as builtins so they
land in Help and are executable like every other verb. The one exception is
the tag `y` chord, which still mirrors out because a tag is always insert, so
SPC cannot be pressed there, and copying the path out is the whole point of
the chord.
Reading is a new read_clipboard effect answered by an ordinary Event.paste, so
the round trip is honest about being one: SDL and NSPasteboard answer inside
the same drain, the browser answers a promise, and a terminal answers over
OSC 52 or -- far more often -- refuses. A refused read is a paste that does
not happen, and the request dies at the next keystroke rather than landing
minutes late in whatever pane is focused by then.
The tty shell also enables BRACKETED PASTE now and coalesces
paste_start..paste_end into one event. Before this a paste arrived as a flood
of individual key presses: plausible in insert mode, and in normal mode every
pasted character ran as a command.
n/N. They stepped the armed results buffer and immediately Looked each row, so
you could not walk past a hit without opening it. They are a MOTION now:
select the next look-able text, open nothing, and let Enter decide. What they
step is the largest whitespace-delimited run look.resolve can act on
(look.lookableSpan, wrapper punctuation peeled), over a RING of panes -- every
pane that has performed a look, most recent first, then the output buffers
that have not, newest first, and only if both are empty the pane in front of
you. N is the exact inverse of n, computed rather than remembered: both
directions ask the same question about the same spans and compare against the
column the walk parks on, so x presses one way and x back land exactly where
you started, pane boundaries and the ring's seam included.
A ring rather than a list with two ends because a shell's cursor sits at the
prompt, below everything it has printed, so a walk that could not come round
would have nowhere to go on the very first press -- which is the case n/N were
written for.
One motion everywhere, no pane-kind or buffer-kind special case. The only
thing a buffer may change is the GRAIN of what a step selects, and it does it
with one flag rather than a branch: output_pane.Traits.commands (renamed from
`executes`, which named one reader's behaviour rather than the fact) makes a
row select WHOLE, because a ThemeSel line is a word to run and has no path
inside it to pick out. `]d`/`[d` are not n/N -- they are helix's diagnostic
motions, their job is to ARRIVE, and they still reach searchStep.
THE TTY PROMPT. Leaving raw tty blanked the prompt row, and the command you
had typed at that prompt shares the row, so it went too -- a shell out of tty
read as output only. OSC 133 marks the row CELL by cell, so the two are
separable: config.tty_blank = .prompt cuts the prompt's own columns and leaves
the command, left-hugged at column 0 in line with the output under it rather
than in a bay of blanks. .prompt_and_input is the old behaviour, kept.
Because the row is now something you can put a cursor in, enterTty adds the
hidden prompt width back before asking ghostty to walk the shell's own cursor
to it -- the modal column on a cut row is short by exactly that much.
Verified: unit-test 186/186 (nine new), snap 87/87 (new ttyprompt.snap),
hxdiff 481 and hxparity 561 with 0 mismatches, tty and gui both build. And
against the real binaries rather than the harness: in a pty, SPC y emits OSC
52 carrying exactly the selection while plain y emits nothing, SPC p issues
the read and pastes the reply, and a bracketed paste of "dd..." inserts text
instead of deleting two lines. In a real SDL window, SPC y then SPC p round
trips through the system clipboard while the default register holds different
text. Setting tty_blank back to .prompt_and_input reproduces all 86 old
goldens byte for byte.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
max_fonts was 512 and this desktop has 1071 monospace faces installed. The
walk stopped at the cap, and because the sort runs AFTER the cut, the picker
did not look truncated -- it ran A to z with four hundred faces missing out of
the middle of it, which is a far worse way to be wrong than a short list.
The cap is now 4096 and the array comes from the arena rather than the stack:
4096 * {name, path} is 128 KiB, which is a fine thing to hand an arena that
resets at the end of the keystroke and not a thing to put on a call stack. It
was only ever a MEMORY bound anyway -- the work is bounded by max_steps, since
a face has to be walked past before it can be found -- and the comment now
says so instead of implying the number was about how long a list can be read.
Costs nothing measurable: the walk is what takes the time, not the four sfnt
reads per file. 512 faces warm was 31ms, 1071 is 36ms. (The 11s I first
measured was a cold page cache reading every font file on the disk once.)
Two guards, because the reason this went unnoticed is more interesting than
the off-by-a-cap:
- src/fonts.zig is imported behind `platform == .gui or .macos`, so on the tty
build nothing analyses it and zig collected no tests from it. It HAD tests;
they never ran. It now has its own libc-linked module in unit-test, which is
the hazard build.zig already writes down next to shell_bin.zig.
- a canary test asserting installed.len < max_fonts. Reaching the cap means
the list handed to the picker is a lie, and it should fail loudly rather
than quietly serve half a machine. Verified it fails at 512 and passes at
4096.
Verified in a real SDL window: SPC t f then a dump, 1092 rows in +Fonts.
unit-test 186/186 (5 of them newly reachable).
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Every keystroke in a shell pane rebuilt the motion surface from scratch:
shellRows dumped ghostty's WHOLE history+active grid, split it, blanked the
prompt rows and handed back slices into the scratch arena, which the next
update threw away. A pane sitting on a multi-megabyte agent transcript paid an
O(scrollback) dump per press of `j`, and paid it once or twice per key, since
flatSurface then rebuilt the same rows joined by '\n' beside it.
The dump is now memoized against the pane it was built for (term_pane.RowsCache
on Pardes.shell_rows), gpa-owned rather than scratch-arena because the whole
point is to outlive the update that built it. One entry, not a table: the
surface is built for the pane the cursor is in, and a second pane asking would
only double a multi-megabyte buffer for a slot it is about to lose again. A
pane that is not the live one is answered from the arena as before.
The lifetime rule is the part that would have rotted silently, so it is one
rule and it is written down: `rows` is handed out to callers, so everything
that notices the entry has gone bad — output arrived, the grid reflowed, the
pane died, another pane wants the slot — only marks it `stale`, and the
buffers are freed in exactly two places, `sweep` at the TOP of an update
before any handler can be holding them, and `reset` when the editor goes away.
Nothing frees mid-update. dropPane clears the pointer immediately though: a
freed pane's address comes back from the allocator as a different pane, and an
entry still naming it would answer for the wrong grid.
Two things fall out of having the join already:
- flatSurface returns the memo's `text` verbatim when the lines it was handed
are the cached rows untouched, instead of rebuilding the join.
- paneCursorLines returns `rows` directly when there is no edit buffer, where
it used to copy the array one slice at a time to produce exactly what it was
given.
One bug on the way past, in the same function: an EMPTY edit buffer writes one
line but modal.lineCount("") is 0, so `ls` was sized one short of what the
loop writes — the same floor the paste site needs. Killing a whole line
(`A<C-u>`, `d%`) on a buffer covering the last row made that a length of zero.
And test/perf.zig grows the axis that would have caught this: a terminal
scoreboard beside the file one, three scrollback fixtures (64 KiB, 1 MiB,
8 MiB — half the ceiling) against render / output / resize-rows / resize-cols /
key-down / edit-char, sharing the existing text and JSON reports and the
--base comparison. resize-cols and resize-rows are both there because a COLUMN
change reflows every page in the list and a row change does not.
Measured on that table: key-down is 142 / 630 / 636 us across the three
fixtures — flat from 1 MiB to 8 MiB, which is the dump being gone, and render
flat at ~110 us throughout. What remains of key-down's step at 1 MiB is the
linear scan indexOf refuses to index for a terminal; that is now a ponytail
waiver naming its own price (615 us against 140 us) and the threading through
paneOff/panePos/paneLineStart it would cost, to be done the day 0.6 ms shows
up next to something anybody can feel.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The SDL shell painted 18,18,18 behind the grid and inside every cell the core
left at its default background, whatever the theme was wearing. Under a light
theme that is a black line along the bottom and right edges of the window --
the strip left over when the window is not a whole number of cells -- measured
two pixels tall with acme at 1728x2102.
The ground is now core.theme().bg, read once a frame, the way PardesView reads
pardes_theme_bg: a Theme command takes hold without a relaunch, and it is the
theme's OWN background rather than the animated chrome colour, because
document backgrounds switch the instant the theme does.
The other half of the macOS fix does not port. A theme that declares no
background (the curated dark, every vendored *_transparent) goes see-through
over an NSVisualEffectView there; here it keeps the terminal-native dark,
because SDL's GPU API refuses to claim a SDL_WINDOW_TRANSPARENT window at all
-- 'The GPU API doesn't support transparent windows', SDL_gpu.c, since D3D12
has no transparent swapchain. Tried it: the shell fails at
ClaimWindowForGPUDevice and exits. bg_default now says so where the next
person will look for it.
Verified on a real window under niri: with Theme acme the bottom strip is
#ffffea where it was #121212.
|
| |\
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
theming, and mupdf -Djpx
Three commits off 38e9919 (macos-app@upstream) merged into main's ghostty bump.
No textual conflicts, and two things the merge needed:
- nested.zig asked libc for fstatat. Darwin has it; on linux std.c declares it
`void` (glibc hides it behind a versioned symbol std cannot name), so the tty
build stopped at 'type void not a function'. statNoFollow keeps fstatat on
darwin and asks statx on linux for the same three fields, which is what this
file did before the branch generalized it to both platforms.
- .DS_Store rode along with a797a1a. Deleted, and .gitignore now says so.
linux: snap 86/86, unit-test, image-harness and mupdf-check green. nested.zig
also type-checks for aarch64-macos.
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
The AppKit shell now draws what the core renders, follows the theme without a
relaunch, and builds into something you can hand to someone.
- Pixel attachments. Surface.images was dropped on the floor here, so a PDF
pane showed nothing at all: native_images is now set, pardes_image_s carries
the geometry the core already clipped, and PardesView keeps one CGImage per
(serial, page, revision) so scrolling costs a draw and not a decode. Image
panes get real pixels instead of the petscii fallback.
- Themes take hold live. pardes_tick never advanced the chrome animation, so
every tagline kept the previous theme's colours until the next launch and
the 16 ms re-pump spun for the rest of the session. pardes_theme_bg retires
the hand-agreed #121212 and drives the window background and the titlebar
appearance; a theme with no background of its own now gets a transparent
window over an NSVisualEffectView.
- The cell snaps to whole DEVICE pixels rather than whole points. Monaco
advances 8.4014pt at 14, so ceiling to 9 spaced every column 7.1% wider than
the face was drawn for.
- The dial is one notch per 10 degrees instead of 20, and a release keeps
turning in proportion to how hard it was thrown -- ramping up from zero at
the floor, so a slow twist coasts not a little but not at all.
- A file dropped on the grid is a click plus Look, so it opens beside the pane
it was dropped on. No drop concept was added to the core.
- The titlebar follows the focused pane: proxy icon, filename, and the dirty
dot. File.saved_revision is the watermark that last one needed.
- Config (SPC f c) prints the resolved startup config path.
- build.zig assembles, signs and packages the bundle itself; build-app.sh is
gone. -Dmacos-identity= takes a Developer ID, macos-dmg makes the image, and
the icon is Glenda.
|
| | | |
|
| |/
|
|
|
|
|
| |
- build.zig.zon: bump ghostty ba38b493 -> 82e53e3f (translate-c backport, fixes 404 on cold cache), update hash
- src/pardes.zig: adapt Terminal.init/resize to new std.Io signatures; fix display-column conversion for tabs and clicks past EOL (fileRawDisplayCol/fileRawAtDisplay)
- test/e2e_harness.zig: adapt to new ghostty_vt signatures
- snap suite 86/86 green
|
| | |
|
| | |
|
| | |
|
| | |
|
| | |
|
| | |
|