diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-27 15:32:02 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-27 16:42:08 -0300 |
| commit | def843b2f59b867ee9b1d501f559f59fb335d4cc (patch) | |
| tree | d7c1650c045653ebc93a77d7a90985e5e31725c2 /docs/registry.typ | |
| parent | f5927a033f0c83753b5cc004e514568eec8c24f8 (diff) | |
| download | pardes-def843b2f59b867ee9b1d501f559f59fb335d4cc.tar.gz pardes-def843b2f59b867ee9b1d501f559f59fb335d4cc.zip | |
9p: serve the acme tree over 9P2000 on a unix socket, beside the mount
Step 4 of the 9P chain (docs/9p.typ 12.4, docs/registry.typ 9P-15/16/17/4/5).
src/9p.zig is a base 9P2000 codec and a SANS-IO server: it never touches a
descriptor, takes no allocator, starts no thread, and builds for
wasm32-freestanding and riscv32-freestanding. That is what lets the same code
serve a unix socket here and a UART on the board later.
Server(comptime fs: type) duck-typed on fs.Req/fs.Reply/fs.Reply.Attr, so it
never imports acmefs and acmefs never learns 9P
init{ in, out, root } the caller owns the buffers; msize is derived
retry/next/reply the three fs_service.Transport ops, by name
push/output/wrote/hangup bytes in, bytes out, partial writes supported
next() is a PUMP, not one-message-one-request: a 3-element Twalk is three
lookups, Topen|OTRUNC is a setattr then an open, Tversion is none at all.
Decisions that were open and are now taken, each recorded in the file:
* qid.version is ALWAYS 0, which makes Linux set P9L_DIRECT and skip its
cache -- the 9P equivalent of the FOPEN_DIRECT_IO fuse.zig relies on.
* Every Rread is clamped to the client's count. An over-long one is a hard
-EIO in Linux, not a truncation.
* Rerror carries Linux's exact strerror text (registry 9P-4 option A), so a
mount recovers the errno instead of ESERVERFAULT. Asserted as literals,
because a typo there is 'Unknown error 526' on every mount.
* `.` and `..` are resolved BY THE SERVER. Under FUSE the kernel does it
and acmefs says so; 9P has no kernel, and forwarding `..` as a lookup
would break every client that normalises a path.
* Topen checks the perm bits itself. Under FUSE the kernel enforced them;
over 9P nobody is above the server, and `errors` would have been readable.
* Tcreate and Tremove are Rerror: `new/` creates a pane on WALK, so the
capability exists and is not spelled Tcreate.
THE INTEGRATION BUG, which was not in the protocol: the daemon's push_fs_reply
sent every reply to the FUSE mount, whose park table has no 9P tag, so it
dropped it -- Tversion worked (no core involved) and Tattach hung forever. That
is exactly the 'no routing origin for the 9P descriptor' cell in the layering
table of docs/9p.typ. Session.fs_origin now carries the transport that asked.
Proved with plan9port against a live daemon serving BOTH transports at once:
9p ls / and /1, read index/ctl/tag, write /1/body, stat, a walk through
/1/../index, pane creation through `new/body`, and the two refusals arriving as
strings -- 'permission denied' and 'No such file or directory' -- confirmed on
the raw wire as Rerror text rather than numbers. A write over 9P reads back
through FUSE and a write through FUSE reads back over 9P.
msize 8192, 34,072 bytes per connection (Server 9,488 + in 8,192 + out 16,384,
out being two msize so that every reply is infallible), four connections.
zig build unit-test: 468 tests before, 503 after.
Diffstat (limited to 'docs/registry.typ')
| -rw-r--r-- | docs/registry.typ | 86 |
1 files changed, 86 insertions, 0 deletions
diff --git a/docs/registry.typ b/docs/registry.typ index 597e930f..3db77c65 100644 --- a/docs/registry.typ +++ b/docs/registry.typ @@ -298,6 +298,17 @@ rejected whole. entry boundary is silently accepted and a changing directory tears. The 40-line budget above buys correctness `ad` does not have. ] + + #note("built", "2026-08-27")[ + Built, and the carry-over entry turned out to be unnecessary. u9fs needs one + because `readdir(3)` has already consumed the entry it could not fit; + `acmefs` re-stages the whole listing from an index on every call and says + why (`src/acmefs.zig:1013-1016`), so an entry that does not fit is simply + not counted and the next read asks for it by index. The server keeps a + per-fid PAIR — a byte cursor for the client's rule and an entry index for + the core's — advanced together. That removes ≈300 bytes per fid and a class + of staleness bug. The ≈40-line budget held. + ] ] #entry("9P-6", "Is `addr` per-fid or per-window?", state: "open", tags: ("semantics", "divergence"))[ @@ -470,6 +481,25 @@ rejected whole. ] #q[Should the parked second RISC-V core own the 9P server? `report.typ:552-563` says it needs four register writes plus a trampoline, shares one L1 D-cache so a lock-free ring needs only fences, and that giving core 1 the UART "eliminates the silent input loss". That is the one arrangement where the board serves 9P *and* keeps the editor. Uncosted.] + + #note("built", "2026-08-27")[ + *The RAM figure above is wrong and the built one is 21,776 B, not 8,832.* + Measured from the real structs: `Fid` is 64 B × 32 = 2,048, `Slot` is 208 B + × 32 = 6,656, and the whole `Server` is 9,488 B before buffers; a 4,096 + msize adds `in` 4,096 + `out` 8,192. + + Three reasons, all of them things the estimate did not know. `out` is TWO + msize — one message being written, one being built — which is what makes + every reply infallible and removes "can I write yet" from the whole file. A + fid entry is 64 B and not 16, because `Rstat` carries a NAME that a node id + does not, plus the open handle and the two-coordinate cursor. And the + estimate did not cost the park table at all, which is 6,656 B of the total. + + It still fits with room: 6.5% of the board's ≈336 KB free heap, and about + 11 KB in total at the 512-byte msize Plan 9 accepts. A test bounds `Fid` and + `Slot` so that a change to either shows up as a diff in the board's budget + rather than as a surprise on the die. + ] ] #entry("BOARD-1", "Raise UART0 to 921600 before quoting any board latency", state: "decided", tags: ("board", "cheap", "prerequisite"))[ @@ -734,6 +764,20 @@ Small, certain, and independent of every argument above. (`22041ab`); flush tracking (`7170e6a`, `5a6b5bf`). #q[Most of these are `Tcreate`/`Tremove` bugs, and a synthetic tree could simply answer `Rerror` to both — the draft's §9 already argues our trees are invented and have no cases we did not choose. Does `new/` need real `Tcreate`, or is walk-to-create enough as it is under FUSE?] + + #note("built", "2026-08-27")[ + *Answered, and most of the list is moot.* `new/` creates a pane on WALK + (`src/acmefs.zig:901-919`), so the capability exists and is not spelled + `Tcreate`. The server answers `Rerror` to both `Tcreate` and `Tremove`, + which takes the create-permission-masking, create-`..`, remove-on-close and + write-then-remove bugs off the table entirely — four of `ad`'s twelve. + `Tremove` still clunks the fid first, because remove(5) says the fid is + invalid even if the remove fails. + + Of the rest, the partial-walk rule and the flush tracking are implemented + and tested by name; the root qid's name is `"/"`; the fid cache is cleared + on clunk and the release the core is owed is issued there. + ] ] #entry("FIX-2", "The board's free-heap comment is two refactors out of date", state: "decided", tags: ("comment", "board", "one-line"))[ @@ -984,3 +1028,45 @@ Three of those four are achievable. One is not, and `9P-20` says which. "NINETEEN each". The 20 figure was wrong by one and is superseded here. ] ] + +#entry("9P-24", "A dead descriptor left in a poll set is a whole core, forever", state: "landed", tags: ("bug", "lesson"))[ + Found by an adversarial pass over step 1 after it was committed and after the + happy path had been demonstrated with the project's own example clients. It is + recorded because the shape recurs, not because the fix was hard. + + Step 1 put `/dev/fuse` in the daemon's `poll(2)` set whenever the mount + existed, and gave `Source.fuse` an empty arm on the grounds that being in the + set was the whole point. + + #ev("linux/fs/fuse/dev.c")[`fuse_dev_poll` answers `EPOLLERR` once the connection is gone — and POSIX reports `POLLERR` whatever the `events` mask asked for, so an empty arm cannot decline it.] + #ev("src/fuse.zig")[`pollLoop`, the mount's own poll thread, has carried `if (revents & (ERR|HUP|NVAL) != 0) return;` all along. That is why the tty and SDL shells never showed this and the daemon did: they never poll the descriptor themselves.] + + So an external `fusermount3 -u`, a sysfs abort, or systemd taking + `/run/user/$UID` away at final logout — *exactly the moment a detached session + is supposed to keep running* — made `poll(2)` return instantly and forever. + Measured on the committed change: 0 CPU ticks over 10 s idle, then 1000 ticks + over the next 10 s. After the fix, 0 ticks over 8 s in the same scenario, with + the process alive and in state `S`. + + #verdict[ + Gate the insertion on `!f.dead` and let the arm consume `POLLERR`/`POLLHUP`/ + `POLLNVAL` by marking the mount dead. Both, not either: the gate is what + ends the spin, and the arm is what saves the one spinning round before + `Fs.next`'s first failed read would have set the flag anyway. + + The general rule, which is the reason this entry exists: *a descriptor + added to a shared poll set needs an error arm even when it needs no data + arm.* An empty arm is a decision about `POLLIN` only, and the kernel does + not ask permission before reporting `POLLERR`. + ] + + #note("review", "2026-08-27")[ + Worth noticing how it was caught. The feature was demonstrated working — + `pardesctl panes`, `new`, `send`, `body`, `del` against a live daemon — and + the bug was nowhere near the happy path. It took an adversary told to + *assume the happy path works and look elsewhere*, who then went and measured + `/proc/<pid>/stat` before and after an unmount. Four of that pass's other + seven findings were false comments rather than false code, which in this + codebase is the same severity: the comments are how the next change is made. + ] +] |
