diff options
| author | Gabriel Schneider <[email protected]> | 2026-08-27 15:32:02 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-08-27 16:42:08 -0300 |
| commit | def843b2f59b867ee9b1d501f559f59fb335d4cc (patch) | |
| tree | d7c1650c045653ebc93a77d7a90985e5e31725c2 /src/detached/server.zig | |
| parent | f5927a033f0c83753b5cc004e514568eec8c24f8 (diff) | |
| download | pardes-def843b2f59b867ee9b1d501f559f59fb335d4cc.tar.gz pardes-def843b2f59b867ee9b1d501f559f59fb335d4cc.zip | |
9p: serve the acme tree over 9P2000 on a unix socket, beside the mount
Step 4 of the 9P chain (docs/9p.typ 12.4, docs/registry.typ 9P-15/16/17/4/5).
src/9p.zig is a base 9P2000 codec and a SANS-IO server: it never touches a
descriptor, takes no allocator, starts no thread, and builds for
wasm32-freestanding and riscv32-freestanding. That is what lets the same code
serve a unix socket here and a UART on the board later.
Server(comptime fs: type) duck-typed on fs.Req/fs.Reply/fs.Reply.Attr, so it
never imports acmefs and acmefs never learns 9P
init{ in, out, root } the caller owns the buffers; msize is derived
retry/next/reply the three fs_service.Transport ops, by name
push/output/wrote/hangup bytes in, bytes out, partial writes supported
next() is a PUMP, not one-message-one-request: a 3-element Twalk is three
lookups, Topen|OTRUNC is a setattr then an open, Tversion is none at all.
Decisions that were open and are now taken, each recorded in the file:
* qid.version is ALWAYS 0, which makes Linux set P9L_DIRECT and skip its
cache -- the 9P equivalent of the FOPEN_DIRECT_IO fuse.zig relies on.
* Every Rread is clamped to the client's count. An over-long one is a hard
-EIO in Linux, not a truncation.
* Rerror carries Linux's exact strerror text (registry 9P-4 option A), so a
mount recovers the errno instead of ESERVERFAULT. Asserted as literals,
because a typo there is 'Unknown error 526' on every mount.
* `.` and `..` are resolved BY THE SERVER. Under FUSE the kernel does it
and acmefs says so; 9P has no kernel, and forwarding `..` as a lookup
would break every client that normalises a path.
* Topen checks the perm bits itself. Under FUSE the kernel enforced them;
over 9P nobody is above the server, and `errors` would have been readable.
* Tcreate and Tremove are Rerror: `new/` creates a pane on WALK, so the
capability exists and is not spelled Tcreate.
THE INTEGRATION BUG, which was not in the protocol: the daemon's push_fs_reply
sent every reply to the FUSE mount, whose park table has no 9P tag, so it
dropped it -- Tversion worked (no core involved) and Tattach hung forever. That
is exactly the 'no routing origin for the 9P descriptor' cell in the layering
table of docs/9p.typ. Session.fs_origin now carries the transport that asked.
Proved with plan9port against a live daemon serving BOTH transports at once:
9p ls / and /1, read index/ctl/tag, write /1/body, stat, a walk through
/1/../index, pane creation through `new/body`, and the two refusals arriving as
strings -- 'permission denied' and 'No such file or directory' -- confirmed on
the raw wire as Rerror text rather than numbers. A write over 9P reads back
through FUSE and a write through FUSE reads back over 9P.
msize 8192, 34,072 bytes per connection (Server 9,488 + in 8,192 + out 16,384,
out being two msize so that every reply is infallible), four connections.
zig build unit-test: 468 tests before, 503 after.
Diffstat (limited to 'src/detached/server.zig')
| -rw-r--r-- | src/detached/server.zig | 155 |
1 files changed, 148 insertions, 7 deletions
diff --git a/src/detached/server.zig b/src/detached/server.zig index 8e6b92d4..2d7a5238 100644 --- a/src/detached/server.zig +++ b/src/detached/server.zig @@ -174,6 +174,10 @@ const file_watch = @import("../file_watch.zig"); /// drain, at the same point in the frame that tty.zig drains. const fuse = @import("../fuse.zig"); const fs_service = @import("../fs_service.zig"); +/// ...and the SECOND transport onto that same tree, on a unix socket of its +/// own. `Source.ninep_listener` and `Source.ninep` are its arms; `pollFrame` +/// is its drain, beside the mount's, at the same point in the frame. +const fs9_service = @import("../fs9_service.zig"); /// Host-lifetime storage for the OSC 133 rc files a forked shell sources, held /// by `Session` because a Session is exactly one host's lifetime. @@ -250,9 +254,11 @@ const read_chunk = 16 * 1024; const pty_chunk = 64 * 1024; /// Descriptors in the ONE poll this process runs: the listener, every frontend, -/// every pane's pty master, the inotify descriptor behind every watch, and -/// `/dev/fuse`. 51 on a full house, and one syscall covers all of them. -const poll_slots = 1 + max_clients + pardes.MAX_PANES + 2; +/// every pane's pty master, the inotify descriptor behind every watch, +/// `/dev/fuse`, the 9P listener, and every 9P connection. That is +/// 1 + 32 + 16 + 1 + 1 + 1 + 4 = 56 on a full house, and one syscall covers +/// all of them. +const poll_slots = 1 + max_clients + pardes.MAX_PANES + 2 + 1 + fs9_service.max_conns; /// How many times `reloadWatched` will honour `file_watch.reloadChanged`'s /// request for another pass within one round. See `reloadWatched`. @@ -411,6 +417,16 @@ const Source = union(enum) { /// one, and the `fd < 0` guard cannot see it because the fd is valid, /// merely different. Same hazard as the client slots, same reason. fuse, + /// The 9P listener, and one arm per connection on it. Same shape and same + /// hazard as `fuse` above, and stated separately only because the hazard + /// is WORSE here: a 9P `Twrite` to `ctl` reaches `core.perform` exactly as + /// a FUSE write does, so it can `push_spawn` and close a pane's master + /// under a later `.pty` entry in this same pass. So `dispatch` moves BYTES + /// — accept, read into `push`, drain `output` — and never serves a + /// request; `pollFrame` does the serving, after the snapshot is done with. + ninep_listener, + /// A connection's index in `Listener.conns`. + ninep: u8, }; pub const Session = struct { @@ -463,6 +479,31 @@ pub const Session = struct { /// Same role as `check_files`: nothing else will wake us, because no /// acknowledgement has gone back, so the next round must not sleep. fs_pending: bool = false, + /// The same tree on a unix socket, or null when `--fs9` was not asked for + /// or the bind failed. Owned here, like `fs`, so that `deinit` closes the + /// socket and unlinks its path on every way out. + ninep: ?*fs9_service.Listener = null, + /// A 9P drain stopped with work still owed. `fs_pending`'s twin, and it + /// needs its own field rather than sharing that one: they are cleared by + /// different drains, and or'ing them into one flag would make a busy 9P + /// script keep the FUSE drain's `pending` set forever. + ninep_pending: bool = false, + /// WHICH TRANSPORT THE REQUEST BEING SERVED CAME FROM, for the whole of + /// one `fs_service.drain` and never outside one. + /// + /// A filesystem reply reaches its transport through `push_fs_reply`, which + /// is a HOST method: the core answers a `fs_req` and does not know, and + /// must not know, that this process has two filesystems on one tree. This + /// field is the routing origin docs/9p.typ's layering table says a second + /// listener costs, and it is one pointer set by the drain that already + /// knows the answer. + /// + /// Null outside a drain, and `fsReply` then falls back to the mount. That + /// fallback is for a reply the core produced with no request outstanding, + /// which `Fs.reply` drops on a slot lookup; it is NOT a routing guess, and + /// a 9P reply cannot reach it — every 9P request is answered inside the + /// `step` that made it, which is inside the drain that set this. + fs_origin: ?fs_service.Transport = null, /// Which directory mark belongs to which pane, and the generation the core /// has already accepted from each. file_watch.zig owns the shape and the /// transaction; this host owns only the descriptor and the wake. @@ -513,6 +554,15 @@ pub const Session = struct { f.deinit(); s.fs = null; } + // ...and the 9P socket, for the mount's reason above: a script blocked + // on `event` over 9P is woken by the EOF its own connection closing + // produces, while its shell is still alive to run its exit path. The + // orphaned fids are not pumped — the core they would report releases to + // is going with them. + if (s.ninep) |l| { + l.deinit(s.gpa); + s.ninep = null; + } // Tell everyone the session is over before the socket disappears, so a // frontend exits on a `quit` rather than on a read error whose meaning // it has to guess. Best effort by construction: these descriptors are @@ -960,9 +1010,17 @@ pub const Session = struct { /// transport holding it. `bytes` was resolved by `pardes.fsPayload` inside /// `perform` and is borrowed only for this call, so a body read is a window /// onto the pane's live text and copies nothing. `.again` needs no case: - /// `Fs.reply` reads the status and re-parks the request itself. + /// both transports read the status and re-park the request themselves. + /// + /// WHICH transport is `fs_origin`, set by the drain that asked. This used + /// to be `s.fs` unconditionally, which was right while a mount was the only + /// answer there was and became a silent misroute the moment `--fs9` gave + /// the session a second one: `Fs.reply` looks the tag up in ITS park table, + /// finds nothing, and returns — so a 9P `Tattach` was answered into the + /// void and its client waited forever. Measured against plan9port's `9p`. fn fsReply(ctx: ?*anyopaque, reply: *const pardes.acmefs.Reply, bytes: []const u8) void { const s = of(ctx); + if (s.fs_origin) |t| return t.reply(reply, bytes); if (s.fs) |f| f.reply(reply, bytes); } @@ -985,7 +1043,29 @@ pub const Session = struct { // in the surface this frame composes, not the next one. The flag is // read by `waitInput`, which must not sleep while the kernel still has // requests we have not acknowledged. - if (s.fs) |f| s.fs_pending = fs_service.drain(f.transport(), s.core).pending; + if (s.fs) |f| { + s.fs_origin = f.transport(); + s.fs_pending = fs_service.drain(s.fs_origin.?, s.core).pending; + s.fs_origin = null; + } + // ...and the 9P connections, in the same breath and for the same + // reason. One connection is one `Transport`, so this is + // `fs_service.drain` per connection per frame, with `fs_origin` naming + // the one being served so its replies come back to it (see `fsReply`). + // HERE rather than in `dispatch` for `Source.ninep`'s reason: a + // `Twrite` to `ctl` can `push_spawn`, and `dispatch` is mid-iteration + // over a descriptor snapshot when it runs. + if (s.ninep) |l| { + var pending = false; + for (0..fs9_service.max_conns) |i| { + const t = l.transport(@intCast(i)) orelse continue; + s.fs_origin = t; + const d = fs_service.drain(t, s.core); + s.fs_origin = null; + if (l.settle(@intCast(i), d)) pending = true; + } + s.ninep_pending = pending; + } for (&s.ptys, 0..) |*pt, pane| { if (pt.fd < 0) continue; var lbuf: [1024]u8 = undefined; @@ -1345,6 +1425,28 @@ pub const Session = struct { src[n] = .fuse; n += 1; }; + // ...and the 9P socket, when there is one: the listener, plus one + // descriptor per live connection. POLLOUT only while a connection owes + // bytes, which is the same rule and the same reason as a client's and + // a pty's — asking for it unconditionally makes every idle socket a + // ready descriptor and turns the poll into a spin. + if (s.ninep) |l| { + if (l.fd >= 0) { + fds[n] = .{ .fd = l.fd, .events = poll_in, .revents = 0 }; + src[n] = .ninep_listener; + n += 1; + } + for (0..fs9_service.max_conns) |i| { + if (!l.live(@intCast(i))) continue; + fds[n] = .{ + .fd = l.conns[i].fd, + .events = if (l.owes(@intCast(i))) poll_in | poll_out else poll_in, + .revents = 0, + }; + src[n] = .{ .ninep = @intCast(i) }; + n += 1; + } + } // A session with no listener, no clients, no shells and no watches has // no event source at all. Returning immediately would spin the outer // `while (!core.quit)` at full speed, so sleep the interval the core @@ -1355,7 +1457,13 @@ pub const Session = struct { // `reconcile` posts a resize only when a client is ATTACHED, which means // its socket is in the set; and `fs_pending` is only ever set by a drain, // which runs only when `fs` is live, which puts `/dev/fuse` in the set. - // So `n == 0` implies `!regridded` and `!fs_pending` too. + // So `n == 0` implies `!regridded` and `!fs_pending` too. `--fs9` is + // the ONE exception, and it is why `ninep_pending` is named on the + // `timeout = 0` line below rather than here: a hung-up 9P connection + // still owing the core its orphaned fids has no descriptor at all, so + // it can be the only work left with `n == 0`. It is also finite — one + // release per fid — so the 16 ms nap that path takes costs it a couple + // of frames and never a stall. if (n == 0) return nap(if (timeout_ms == 0) 16 else timeout_ms); // Zero is the core's word for "sleep until something happens" (see // pardes.zig `pump`: it passes a frame interval only while an animation @@ -1373,7 +1481,7 @@ pub const Session = struct { // during `perform`, outside this function) and a regrid this round has // already performed. Both are consumed before this function returns, so // the round must not sleep before reaching them. - if (s.check_files or regridded or s.fs_pending) timeout = 0; + if (s.check_files or regridded or s.fs_pending or s.ninep_pending) timeout = 0; const ready = libc.poll(&fds, @intCast(n), timeout); // A timeout is an ordinary frame boundary and EINTR is a signal we do not // handle here. Neither skips anything below any more: what used to be an @@ -1455,6 +1563,26 @@ pub const Session = struct { .fuse => if (pfd.revents & (poll_hup | poll_err | poll_nval) != 0) { if (s.fs) |f| f.dead = true; }, + .ninep_listener => if (pfd.revents != 0) { + if (s.ninep) |l| l.accept(); + }, + // BYTES ONLY. The requests those bytes decode into are served by + // `pollFrame`, after this snapshot is done with — see + // `Source.ninep`. POLLOUT before POLLIN so a reply the last frame + // could not finish writing goes before more work arrives, and + // POLLIN before the hangup because a script that wrote a `Tclunk` + // and closed has bytes still worth reading; the `live` guard is the + // client slots' `c.fd < 0`, for its reason. + .ninep => |i| if (s.ninep) |l| { + if (!l.live(i)) continue; + if (pfd.revents & poll_out != 0) l.flush(i); + if (!l.live(i)) continue; + if (pfd.revents & poll_in != 0) { + l.fill(i); + } else if (pfd.revents & (poll_hup | poll_err | poll_nval) != 0) { + l.drop(i); + } + }, }; } @@ -1898,6 +2026,19 @@ pub fn run(init: std.process.Init, opts: pardes.Options, name: []const u8) !void // process polls `/dev/fuse` itself, in the same syscall as everything else. // See `Source.fuse`. session.fs = fs_service.start(gpa, core); + // ...and the same tree on a unix socket, independently: `--fs` and `--fs9` + // are two transports and neither is the other's prerequisite, so a daemon + // may serve one, both or neither. Null on every failure, for + // `fs_service.start`'s reason — a transport that will not bind must cost + // the operator their scripting, never their session — and NOT a + // `return error` the way the frontend socket above is, because a session + // whose frontend socket did not bind is one nobody can ever find, while + // this one is merely one nobody can script over 9P. + // + // A bare `--fs9` is named by the SESSION rather than by the pid: the + // operator typed that name to find the daemon again, and having to look up + // a pid to reach its filesystem would undo it. `--fs9=<name>` wins. + if (opts.fs9) |named| session.ninep = fs9_service.open(gpa, named, name); const h = session.host(); core.host = h; |
