diff options
Diffstat (limited to 'src')
| -rw-r--r-- | src/CHANGELOG.md | 14 | ||||
| -rw-r--r-- | src/acmefs.zig | 2982 | ||||
| -rw-r--r-- | src/config.zig | 4 | ||||
| -rw-r--r-- | src/file_pane.zig | 26 | ||||
| -rw-r--r-- | src/fs_service.zig | 270 | ||||
| -rw-r--r-- | src/fuse.zig | 2709 | ||||
| -rw-r--r-- | src/gui/gui.zig | 96 | ||||
| -rw-r--r-- | src/host.zig | 8 | ||||
| -rw-r--r-- | src/main.zig | 45 | ||||
| -rw-r--r-- | src/output_pane.zig | 11 | ||||
| -rw-r--r-- | src/pardes.zig | 225 | ||||
| -rw-r--r-- | src/tty/tty.zig | 99 | ||||
| -rw-r--r-- | src/tutor.txt | 38 |
13 files changed, 6508 insertions, 19 deletions
diff --git a/src/CHANGELOG.md b/src/CHANGELOG.md index 3b5056b5..11df45ab 100644 --- a/src/CHANGELOG.md +++ b/src/CHANGELOG.md @@ -2,6 +2,20 @@ ## 0.0.1 +- `pardes --fs` serves acme's control filesystem over Linux FUSE: a directory + per pane holding `addr`, `body`, `ctl`, `data`, `errors`, `event`, `rdsel`, + `tag`, `wrsel`, `xdata`, plus `index`, `cons` and `new/` at the root. A + program that opens those files is an editor extension — no plugin API, no + interpreter. `--fs=<dir>` names the mount point instead of deriving one under + `$XDG_RUNTIME_DIR`; every pane shell gets `PARDES_FS` and `PARDES_PANE`. + See docs/acme-fs.md and examples/acmefs/. +- While a script holds a pane's `event` file open, buttons 2 and 3 in that pane + are REPORTED to it instead of performed, so the words in that pane's tag are + the script's commands (acme's extension model). Keyboard Enter/Tab still + execute, so a scripted pane stays editable. +- Writing an event record back (`origin type q0 q1`) makes pardes perform the + Look or Exec it names — the same door acme has always had, and the reason + the mount is 0700. - Version tracking: the compiled version is embedded from build.zig.zon and shown by Changelog. - New builtins: Changelog, Newtty, and Joincol. - Markdown tree-sitter highlighting, including nested fenced code blocks. diff --git a/src/acmefs.zig b/src/acmefs.zig new file mode 100644 index 00000000..db610f67 --- /dev/null +++ b/src/acmefs.zig @@ -0,0 +1,2982 @@ +//! ACME'S CONTROL FILESYSTEM, as a pure transaction over the core. +//! +//! plan9's acme serves `/mnt/acme`: a directory per window holding `addr`, +//! `body`, `ctl`, `data`, `event`, `tag`..., and a program that opens those +//! files IS an editor extension — no plugin API, no embedded interpreter, no +//! rebuild. `pardes --fs` serves the same tree over Linux FUSE (src/fuse.zig), +//! and this file is the whole of what the files MEAN. Read it beside acme's +//! `fsys.c` (the tree) and `xfid.c` (the handlers). +//! +//! THE SHAPE, and why it is this shape. A filesystem is a request/response +//! protocol driven by other processes, i.e. exactly the kind of concurrency +//! the core does not have and must not grow. acme answers it with a thread per +//! in-flight request (`xfidallocthread`, a `Channel` per `Xfid`, a `QLock` per +//! window); pardes cannot and should not, so: +//! +//! * This module is a PURE MAIN-THREAD TRANSACTION: `handle(p, req) Reply`. +//! No thread, no waiting, no callback, no allocation on the hot path. It +//! is freestanding-safe (no libc, no OS) and unit-testable with no FUSE +//! anywhere near it — the tests below post requests and read replies. +//! * Requests arrive as an ordinary `Event.fs_req` and answers leave as an +//! ordinary `Effect.fs_reply`, so the transport is the queue every other +//! host<->core message already uses. A backend with no threads at all +//! (the browser, a test) is not a special case: it either never sends a +//! request, or sends one from its own frame loop. +//! * BLOCKING — acme's `event` file, whose read waits for the user to do +//! something (acme parks the `Xfid` in `w->eventx` and a later `winevent` +//! sends it a message) — is `Status.again` here: "nothing consumed, ask me +//! again". The waiting lives in the host, which is where the kernel's +//! request already is. The core keeps no waiter list and no wakeups. +//! +//! DIVERGENCE FROM ACME, deliberate: acme counts RUNES, pardes counts BYTES +//! (clamped to grapheme boundaries). Every offset in this filesystem — `addr`, +//! `data`, the event records' q0/q1, `index`'s lengths — is a byte offset, +//! because pardes is byte-addressed end to end (selections, look spots, LSP +//! offsets) and a second coordinate system would mean an O(n) conversion at +//! every boundary and a lossy `addr=dot`. acme pays that cost the other way +//! round: it keeps the document as `Rune*` and converts on every utf read +//! (`xfidutfread`, which carries a "BUG: stupid code: scan from beginning" +//! comment for its cache miss). Identical for ASCII, which is what scripts +//! compute with. +const std = @import("std"); +const mvzr = @import("mvzr"); +const pardes = @import("pardes.zig"); +const config = @import("config.zig"); +const modal = @import("modal.zig"); +const file_pane = @import("file_pane.zig"); +const output_pane = @import("output_pane.zig"); +/// only for `look.readFile`, which is what a `get` verb IS — the same +/// synchronous path-backed read `file_pane.open` does, and the one place this +/// module touches a disk. +const look = @import("look.zig"); +const Pardes = pardes.Pardes; +const Pane = pardes.Pane; +const MAX_PANES = pardes.MAX_PANES; + +// ============================================================================ +// THE ABI — what a transport hands in and gets back. +// ============================================================================ + +/// The filesystem operations the core answers. Protocol-neutral on purpose: +/// FUSE opcodes, 9P messages and a unit test all reduce to these. +pub const Op = enum(u8) { + /// resolve `data` (a name) inside the directory `node` + lookup, + getattr, + /// only `truncate` is honoured; a filesystem of live editor state has no + /// mode, owner or timestamps to set + setattr, + open, + read, + write, + release, + readdir, + statfs, +}; + +pub const Status = enum(u8) { + ok, + /// NO DATA YET, nothing consumed: the transport must hold this request and + /// re-submit it unchanged on a later frame. The one blocking primitive, + /// and the reason the core needs no waiters (see the header). + again, + err, +}; + +/// One operation. `data` is BORROWED for the length of the single +/// `update(.{ .fs_req = ... })` call that carries it — the same rule as +/// `.pty_read`'s bytes — so this never goes through `postEvent`. +pub const Req = struct { + /// opaque echo token; the transport's request id (FUSE `unique`) + tag: u64, + op: Op, + node: u64, + /// from `.open`, on read/write/release + handle: u32 = 0, + /// read/write: byte offset. readdir: how many entries to skip. + off: u64 = 0, + /// read/readdir: bytes wanted. Writes carry their length in `data`. + size: u32 = 0, + /// lookup: the name. write: the bytes. + data: []const u8 = &.{}, + /// setattr: a size was set (only 0 means anything here) + truncate: bool = false, +}; + +pub const Reply = struct { + tag: u64, + status: Status = .ok, + /// positive errno when `status == .err` + errno: u16 = 0, + /// lookup/getattr/setattr answer this; `open` leaves it zeroed + attr: Attr = .{}, + /// `open` answers this; every later read/write/release repeats it + handle: u32 = 0, + payload: Payload = .none, + /// write: how many of the offered bytes were taken. A short count is a + /// real answer (`data` refusing a partial grapheme), not an error. + written: u32 = 0, + + pub const Attr = struct { + node: u64 = 0, + dir: bool = false, + size: u64 = 0, + /// permission bits only; the transport adds the format bits + mode: u16 = 0o600, + }; + + /// WHERE THE ANSWER'S BYTES ARE. Resolved by `pardes.fsPayload` inside the + /// effect drain — a borrow window identical to `.save_text`'s — so reading + /// a megabyte of body copies nothing. + pub const Payload = union(enum) { + none, + /// `State.out[0..len]`: formatted answers (ctl, addr, index, dirents, + /// event records). Valid until the next `handle` call. + staged: u32, + /// a slice of a live pane's text. `serial` rejects a reused slot + /// exactly like `.save_text` does. + region: struct { pane: u8, serial: u32, off: u32, len: u32 }, + }; + + pub fn fail(tag: u64, e: u16) Reply { + return .{ .tag = tag, .status = .err, .errno = e }; + } +}; + +/// The errno values this filesystem returns, standing in for acme's error +/// strings (`fsys.c`/`xfid.c`: Eperm, Ebadctl, Ebadaddr, Ebadevent, Edel...). +/// A filesystem has one channel for "no": the number. +pub const E = struct { + pub const PERM: u16 = 1; + pub const NOENT: u16 = 2; + pub const IO: u16 = 5; + pub const NOMEM: u16 = 12; + pub const NOTDIR: u16 = 20; + pub const INVAL: u16 = 22; + pub const NFILE: u16 = 23; + /// the one FILE here with a fixed capacity: a pane's editable tag tail is + /// a bounded one-line buffer, so a write with no room left is full rather + /// than refused (`writeTag`) + pub const NOSPC: u16 = 28; + pub const NOSYS: u16 = 38; +}; + +// ============================================================================ +// THE TREE — nodes, names, and the packing that makes both cheap. +// ============================================================================ + +/// One file inside a pane's directory: acme's `dirtabw` minus the plan9 +/// compatibility stubs (`editout` needs acme's Edit language; `draw`, +/// `consctl` and `label` are rio artefacts acme keeps for other programs' +/// sake), plus nothing. +pub const PaneFile = enum(u4) { + dir = 0, + addr, + body, + ctl, + data, + errors, + event, + tag, + xdata, + rdsel, + wrsel, + + /// Every name IS the variant's name; only the directory itself is spelled + /// differently, because `.` is not an identifier. + pub fn name(f: PaneFile) []const u8 { + return if (f == .dir) "." else @tagName(f); + } + + /// acme's dirtabw modes: 0400 read, 0200 write, 0600 both. + pub fn mode(f: PaneFile) u16 { + return switch (f) { + .dir => 0o500, + .errors, .wrsel => 0o200, + .rdsel => 0o400, + else => 0o600, + }; + } +}; + +/// The files at the root, and the root itself. `new` is a directory whose +/// every lookup CREATES a pane (acme(4): "Accessing any file in new creates a +/// new window"), which is how a script opens one without a keystroke. +pub const TopFile = enum(u4) { + root = 1, + index = 2, + cons = 3, + new = 4, + + pub fn name(f: TopFile) []const u8 { + return if (f == .root) "." else @tagName(f); + } + + pub fn mode(f: TopFile) u16 { + return switch (f) { + .root, .new => 0o500, + .index => 0o400, + .cons => 0o200, + }; + } + + pub fn dir(f: TopFile) bool { + return f == .root or f == .new; + } +}; + +/// A NODE ID, which is one integer to the kernel and two fields to us: the +/// file within a pane's directory, and the pane's SERIAL — never reused, so a +/// node id can never come to mean a different pane. acme does the same packing +/// with `QID(w->id, f)` / `WIN(q)` / `FILE(q)` macros over an int; a packed +/// struct is the same bits with the shifts and masks checked by the compiler, +/// and `serial == 0` (no pane) is what keeps the top-level ids 1..5 out of the +/// way with no separate range check. +pub const Node = packed struct(u64) { + file: u4 = 0, + serial: u60 = 0, + + pub fn of(serial: u32, file: PaneFile) u64 { + std.debug.assert(serial != 0); + return @bitCast(Node{ .file = @intFromEnum(file), .serial = serial }); + } + + /// What this id points at, or null when it names neither a top-level file + /// nor a possible pane file. Validity is decided HERE so no handler has to. + pub fn target(node: u64) ?Target { + const n: Node = @bitCast(node); + if (n.serial == 0) { + return .{ .top = std.enums.fromInt(TopFile, n.file) orelse return null }; + } + return .{ .pane = .{ + .serial = std.math.cast(u32, n.serial) orelse return null, + .file = std.enums.fromInt(PaneFile, n.file) orelse return null, + } }; + } +}; + +/// What a node id points at. +pub const Target = union(enum) { + top: TopFile, + pane: struct { serial: u32, file: PaneFile }, +}; + +// ============================================================================ +// STATE — everything the filesystem remembers between requests. +// ============================================================================ + +/// What a formatted answer starts out able to hold before it grows: `index` +/// over every pane, a directory listing, one event record. The buffer is kept +/// between requests and cleared, not freed, so the steady state allocates +/// nothing and nothing is capped by a number picked here. +pub const out_reserve = 4 * 1024; + +/// Records a reader has not taken yet, per pane. Beyond this the oldest are +/// dropped: an editor must not stall or grow without bound because a script +/// stopped reading, and a reader that fell this far behind has already lost +/// the thread — it can re-read `body` and resynchronise. (acme grows +/// `w->events` with `realloc` and has no bound at all.) +pub const queue_cap = 64 * 1024; + +/// A byte queue of formatted event records, each framed by its length so a +/// record whose TEXT contains newlines still comes out whole. +pub const Queue = struct { + buf: std.ArrayList(u8) = .empty, + /// How much of `buf` has been consumed. Popping moves this instead of + /// sliding the remainder down: a full queue holds thousands of ~20-byte + /// records, and a memmove per pop made draining one quadratic. The space + /// is reclaimed when the head passes half the buffer, so the amortised + /// cost of a pop is a pointer bump. + head: usize = 0, + + pub fn deinit(q: *Queue, gpa: std.mem.Allocator) void { + q.buf.deinit(gpa); + q.head = 0; + } + + pub fn push(q: *Queue, gpa: std.mem.Allocator, record: []const u8) void { + if (record.len > std.math.maxInt(u32)) return; + while (q.buf.items.len - q.head + record.len + 4 > queue_cap) { + if (q.peek() == null) return; + q.pop(); + } + q.compact(); + var head: [4]u8 = undefined; + std.mem.writeInt(u32, &head, @intCast(record.len), .little); + q.buf.appendSlice(gpa, &head) catch return; + q.buf.appendSlice(gpa, record) catch { + q.buf.shrinkRetainingCapacity(q.buf.items.len - 4); + return; + }; + } + + /// The oldest record, or null when empty. Does not consume. + pub fn peek(q: *const Queue) ?[]const u8 { + const rest = q.buf.items[@min(q.head, q.buf.items.len)..]; + if (rest.len < 4) return null; + const len = std.mem.readInt(u32, rest[0..4], .little); + if (rest.len < 4 + len) return null; + return rest[4 .. 4 + len]; + } + + pub fn pop(q: *Queue) void { + const record = q.peek() orelse return; + q.head += 4 + record.len; + if (q.head == q.buf.items.len) { + q.buf.clearRetainingCapacity(); + q.head = 0; + } + } + + fn compact(q: *Queue) void { + if (q.head == 0 or q.head * 2 < q.buf.items.len) return; + const rest = q.buf.items.len - q.head; + std.mem.copyForwards(u8, q.buf.items[0..rest], q.buf.items[q.head..]); + q.buf.shrinkRetainingCapacity(rest); + q.head = 0; + } + + pub fn empty(q: *const Queue) bool { + return q.peek() == null; + } +}; + +/// Per-pane filesystem state, indexed by pane SLOT (not serial): it dies with +/// the pane, and a reused slot must start clean. +pub const PaneFs = struct { + /// acme's `w->addr`: where `data`/`xdata` read and write. Byte offsets. + addr: Range = .{}, + /// `limit=addr`: the range regex searches are confined to, or none. + limit: ?Range = null, + /// how many opens of this pane's `event` file are live. Non-zero means the + /// pane is SCRIPT-DRIVEN: its Look and Exec are reported, not performed. + readers: u16 = 0, + events: Queue = .{}, + /// `nomark`: writes stop pushing an undo point each, so a script's batch + /// of edits is one Undo (acme: `w->nomark`). + nomark: bool = false, + /// `noscroll`: a body write does not drag the view to the new text. + noscroll: bool = false, + /// The tag as it was at the end of the last update, so tag edits can be + /// reported without a hook in every tag mutation. Only kept while somebody + /// is listening. + tag_snap: std.ArrayList(u8) = .empty, + + pub const Range = struct { q0: u32 = 0, q1: u32 = 0 }; + + fn deinit(pf: *PaneFs, gpa: std.mem.Allocator) void { + pf.events.deinit(gpa); + pf.tag_snap.deinit(gpa); + pf.* = .{}; + } +}; + +/// The core's filesystem state. Lives on `Pardes`; zero-initialised, so a core +/// that never serves a filesystem pays one branch per frame and no memory +/// beyond this struct. +pub const State = struct { + /// Formatted answers, valid until the next `handle` call (Payload.staged). + /// Kept and cleared rather than freed: after the first few requests the + /// capacity is there and staging an answer allocates nothing. + out: std.ArrayList(u8) = .empty, + panes: [MAX_PANES]PaneFs = @splat(.{}), + /// How many `event` files are open anywhere. The one gate every recording + /// hook in the core is behind: nobody listening, nothing recorded, no diff + /// computed, no bytes copied. + listeners: u16 = 0, + /// Which input the core is handling, as acme's origin character: `K` + /// keyboard, `M` mouse, `E` a write to body/tag through this filesystem, + /// `F` an action through one of its other files. Set once per update. + origin: u8 = 'K', + + pub fn deinit(st: *State, gpa: std.mem.Allocator) void { + for (&st.panes) |*pf| pf.deinit(gpa); + st.out.deinit(gpa); + } + + /// Start a fresh answer. The previous one's bytes are dead the moment the + /// next request arrives, which is exactly the borrow window `fsPayload` + /// documents. + pub fn stage(st: *State, gpa: std.mem.Allocator) *std.ArrayList(u8) { + st.out.clearRetainingCapacity(); + st.out.ensureTotalCapacity(gpa, out_reserve) catch {}; + return &st.out; + } + + /// A pane died: drop its filesystem state, and with it any listener count + /// it held, so a script killed with its pane cannot leave the editor + /// suppressing button actions forever. + pub fn forget(st: *State, gpa: std.mem.Allocator, id: usize) void { + if (id >= MAX_PANES) return; + st.listeners -= @min(st.listeners, st.panes[id].readers); + st.panes[id].deinit(gpa); + } + + /// Is anybody reading this pane's events? The suppression rule and every + /// recording hook ask this. + pub fn scripted(st: *const State, id: usize) bool { + return id < MAX_PANES and st.panes[id].readers != 0; + } +}; + +// ============================================================================ +// EVENT RECORDS — what the core reports, in acme's wire format. +// ============================================================================ + +/// acme's record, byte for byte: origin char, type char, then four +/// blank-separated decimals (q0, q1, flag, text length), a blank, the text, +/// and a newline — `winevent`'s `"%c%d %d %d %d %.*S\n"` with the owner char +/// pushed in front (`wind.c`). Text of 256 bytes or more is elided (the +/// reader fetches it from `data`), which is also what bounds this buffer. +pub const max_record_text = 256; + +/// The action characters. Lower case is the tag, upper case the body, which is +/// how a reader tells them apart with no extra field. +pub const Action = enum(u8) { + body_delete = 'D', + tag_delete = 'd', + body_insert = 'I', + tag_insert = 'i', + body_look = 'L', + tag_look = 'l', + body_exec = 'X', + tag_exec = 'x', + + /// The enum IS the character, the way acme's `winevent` takes a `char` and + /// prints `%c` (`wind.c`) — so these three are expressions rather than the + /// three parallel switches that spelled the same alphabet out again. + pub fn char(a: Action) u8 { + return @intFromEnum(a); + } + + pub fn fromChar(c: u8) ?Action { + return std.enums.fromInt(Action, c); + } + + /// Lower case is the tag, upper case the body: acme's whole encoding of + /// WHICH TEXT a record is about, with no extra field. + pub fn onTag(a: Action) bool { + return @intFromEnum(a) >= 'a'; + } +}; + +/// Flag bits, acme(4). Look and exec are different vocabularies at the same +/// bit positions, so they get separate names rather than one enum. +pub const flag_builtin: u32 = 1; +pub const flag_expansion: u32 = 2; +pub const flag_filename: u32 = 4; +pub const flag_chorded: u32 = 8; + +/// Format one record into `buf` and return the bytes. +pub fn formatRecord( + buf: []u8, + origin: u8, + action: Action, + q0: u32, + q1: u32, + flag: u32, + text: []const u8, +) []const u8 { + const sent = if (text.len >= max_record_text) text[0..0] else text; + return std.fmt.bufPrint(buf, "{c}{c}{d} {d} {d} {d} {s}\n", .{ + origin, + action.char(), + q0, + q1, + flag, + sent.len, + sent, + }) catch buf[0..0]; +} + +/// THE SPAN TWO VERSIONS OF A TEXT DIFFER IN: everything outside their common +/// prefix and common suffix. +/// +/// Chunked through `std.mem.eql`, which lowers to vectorised compares. That is +/// not premature: this runs on EVERY edit of a scripted pane, over the whole +/// buffer, and the byte-at-a-time loop it replaces cost 2.4x per keystroke on +/// a 40 KB body (`zig build fs-bench`, the two `keystroke` rows). +pub const Span = struct { at: u32, removed: u32, inserted: u32 }; + +pub fn diffSpan(old: []const u8, new: []const u8) Span { + const both = @min(old.len, new.len); + const stride = 64; + var head: usize = 0; + while (head + stride <= both and + std.mem.eql(u8, old[head..][0..stride], new[head..][0..stride])) head += stride; + while (head < both and old[head] == new[head]) head += 1; + var tail: usize = 0; + const rest = both - head; + while (tail + stride <= rest and std.mem.eql( + u8, + old[old.len - tail - stride ..][0..stride], + new[new.len - tail - stride ..][0..stride], + )) tail += stride; + while (tail < rest and old[old.len - 1 - tail] == new[new.len - 1 - tail]) tail += 1; + return .{ + .at = @intCast(head), + .removed = @intCast(old.len - tail - head), + .inserted = @intCast(new.len - tail - head), + }; +} + +/// Report a whole-text replacement the way acme reports an edit: the deletion +/// first and then the insertion, because that is the order `textdelete` and +/// `textinsert` would have run in. pardes replaces whole buffers, so the pair +/// is recovered here — one implementation, one place that knows the order, and +/// the only cost paid by an unscripted editor is the `scripted` check. +pub fn noteReplace(p: *Pardes, id: usize, on_tag: bool, old: []const u8, new: []const u8) void { + if (!p.fs.scripted(id)) return; + const span = diffSpan(old, new); + if (span.removed == 0 and span.inserted == 0) return; + if (span.removed > 0) _ = noteAction( + p, + id, + if (on_tag) .tag_delete else .body_delete, + span.at, + span.at + span.removed, + 0, + "", + ); + if (span.inserted > 0) _ = noteAction( + p, + id, + if (on_tag) .tag_insert else .body_insert, + span.at, + span.at + span.inserted, + 0, + new[span.at..][0..span.inserted], + ); +} + +/// Record a Look or an Exec, and say whether THE CORE MUST NOT PERFORM IT. +/// +/// That inversion is acme's whole extension model: while a script holds a +/// pane's `event` file open, buttons 2 and 3 in that pane belong to the script +/// — the words in its tag are its commands, not pardes's. A script that dies +/// closes the file and the pane goes back to being an editor. +pub fn noteAction( + p: *Pardes, + id: usize, + action: Action, + q0: u32, + q1: u32, + flag: u32, + text: []const u8, +) bool { + if (!p.fs.scripted(id)) return false; + var buf: [max_record_text + 64]u8 = undefined; + const record = formatRecord(&buf, p.fs.origin, action, q0, q1, flag, text); + p.fs.panes[id].events.push(p.gpa, record); + return true; +} + +// ============================================================================ +// THE TRANSACTION. +// ============================================================================ + +/// Answer one filesystem request against the live editor. The only entry +/// point: `Event.fs_req` lands here and the `Reply` leaves as +/// `Effect.fs_reply`. +pub fn handle(p: *Pardes, req: Req) Reply { + const target = Node.target(req.node) orelse return Reply.fail(req.tag, E.NOENT); + // THE ORIGIN CHARACTER for everything this request goes on to cause — + // including the TAG DIFF the core takes at the end of the update, after + // this function has returned. acme sets `w->owner` in `winlock` and calls + // `winsettag` before `winunlock`, so the tag change a body write provokes + // (the dirty marker appearing) is attributed to that write and not to + // whatever touched the editor last. Set once, here, for the same reason. + // + // Nothing restores it: `Pardes.update` sets the origin afresh on every + // keystroke and every mouse event, which is what owns it the rest of the + // time. A read cannot cause a record, so only the mutating ops set it. + if (req.op == .write or req.op == .setattr) p.fs.origin = switch (target) { + // acme's `xfidwrite` opens with exactly this: `c = 'F'; if(qid==QWtag + // || qid==QWbody) c = 'E';` — `E` is "writes to the body or tag file", + // `F` is "actions through the window's other files" (acme(4)). + .pane => |t| @as(u8, if (t.file == .body or t.file == .tag) 'E' else 'F'), + .top => 'F', + }; + return switch (req.op) { + .lookup => lookup(p, req, target), + .getattr => switch (attrOf(p, target)) { + .ok => |a| .{ .tag = req.tag, .attr = a }, + .missing => Reply.fail(req.tag, E.NOENT), + }, + .setattr => setattr(p, req, target), + .open => open(p, req, target), + .release => release(p, req), + .readdir => readdir(p, req, target), + .read => read(p, req, target), + .write => write(p, req, target), + // A synthetic filesystem has no blocks. Answering successfully with + // zeros keeps `df` and anything that stats the mount working. + .statfs => .{ .tag = req.tag }, + }; +} + +const AttrResult = union(enum) { ok: Reply.Attr, missing }; + +fn attrOf(p: *Pardes, target: Target) AttrResult { + switch (target) { + .top => |f| return .{ .ok = .{ + .node = @intFromEnum(f), + .dir = f.dir(), + .mode = f.mode(), + .size = topSize(p, f), + } }, + .pane => |t| { + const id = p.paneBySerial(t.serial) orelse return .missing; + return .{ .ok = .{ + .node = Node.of(t.serial, t.file), + .dir = t.file == .dir, + .mode = t.file.mode(), + .size = paneFileSize(p, id, t.file), + } }; + }, + } +} + +/// A size for `stat`. Exact where it is cheap and honest (`body`, `tag`), zero +/// where the file is a stream whose length is not a property (`event`, `log`); +/// FUSE serves these with direct IO, so a zero-length file still reads. +fn topSize(p: *Pardes, f: TopFile) u64 { + return switch (f) { + .root, .new, .cons => 0, + .index => indexLen(p), + }; +} + +fn paneFileSize(p: *Pardes, id: usize, f: PaneFile) u64 { + const pane = p.panes[id] orelse return 0; + return switch (f) { + .body, .data, .xdata => bodyLen(p, pane), + .tag => tagLen(p, pane), + .dir, .addr, .ctl, .errors, .event, .rdsel, .wrsel => 0, + }; +} + +// --------------------------------------------------------------------------- +// PER-FILE SEMANTICS. Everything above is the frame: the ABI, the tree, the +// state, the records. Everything below is what acme's xfid.c does. +// --------------------------------------------------------------------------- + +// =========================================================================== +// THE TWO TEXTS A PANE HAS. Every handler below asks these, so "what is this +// pane's body" has one answer here and not eleven answers scattered about. +// =========================================================================== + +/// A pane's BODY, BORROWED. A file pane — which includes every output buffer +/// — lends its content, and that is the whole reason a `body` read costs +/// nothing (`Payload.region`). A terminal has no such buffer: its body is the +/// emulator's scrollback, which has to be RENDERED before it is bytes, so it +/// is not lendable and `readBody` produces one instead. Empty here therefore +/// means "nothing to lend", which for a terminal is not "empty document". +fn bodyOf(pane: *const Pane) []const u8 { + if (pane.file) |*f| return f.content; + return ""; +} + +/// ...and the writable side of the same question. Null is "this pane has no +/// document", which is every terminal and the answer to every write that +/// would need one. +fn fileOf(pane: *Pane) ?*file_pane.State { + return if (pane.file) |*f| f else null; +} + +/// The pane's TAG exactly as it is drawn: the live read-only prefix (the path, +/// the dirty marker, the pane's builtin words, the alignment gap) then the +/// editable tail. +/// +/// Scratch-owned — and `Pardes.update` resets that arena before the transport +/// ever reads a payload, so every tag answer is COPIED into `State.out`. +/// `body` is the only text lent out, because it is the only one that is a +/// buffer rather than a rendering. +fn tagOf(p: *Pardes, pane: *Pane) []const u8 { + return p.tagText(p.scratch.allocator(), pane) catch ""; +} + +/// The directory a pane belongs to: acme's "the directory currently named in +/// the tag", which is where this pane's `+Errors` goes. +fn dirOf(pane: *Pane) []const u8 { + if (pane.file) |*f| return std.fs.path.dirname(f.path) orelse "/"; + const cwd = pane.cwdSlice(); + return if (cwd.len > 0) cwd else "/"; +} + +/// acme's `w->dirty`: the body differs from what is on disk. A terminal and an +/// output buffer have nothing on disk, so they are never dirty — the same +/// `saves` trait the tag's `*` marker already asks. +fn dirtyOf(pane: *const Pane) bool { + const f = if (pane.file) |*x| x else return false; + if (!output_pane.fileTraits(f.output).saves) return false; + return f.revision != f.saved_revision; +} + +/// A byte offset as a `Range` field. A pane holding four gigabytes of text is +/// not something this editor does; saturating is honest where a silent wrap +/// would hand a script an address pointing at the wrong end of the file. +fn clip(n: usize) u32 { + return std.math.cast(u32, n) orelse std.math.maxInt(u32); +} + +fn cellOf(row: i32, col: i32) modal.Cursor { + return .{ .row = @intCast(@max(0, row)), .col = @intCast(@max(0, col)) }; +} + +fn firstLine(s: []const u8) []const u8 { + return s[0 .. std.mem.indexOfScalar(u8, s, '\n') orelse s.len]; +} + +/// acme's DOT — the user's selection — as a byte range over the body. +/// +/// pardes keeps the selection as two (row, col) cells with a HELIX block +/// cursor, i.e. the head cell is INSIDE the range; acme's dot is gap to gap. +/// This is the one place that conversion lives and `setDot` is its inverse, so +/// `addr=dot` followed by `dot=addr` is the identity rather than a range that +/// creeps by one grapheme each round trip. +fn dotOf(pane: *Pane) PaneFs.Range { + const text = bodyOf(pane); + const head = modal.hxOff(text, cellOf(pane.cur_row, pane.cur_col)); + if (!pane.vsel.active) return .{ .q0 = clip(head), .q1 = clip(head) }; + const anchor = modal.hxOff(text, cellOf(pane.vsel.row, pane.vsel.col)); + var hi = @max(head, anchor); + if (hi < text.len) hi = modal.nextGrapheme(text, hi); + return .{ .q0 = clip(@min(head, anchor)), .q1 = clip(hi) }; +} + +/// acme's `textsetselect`. The head lands ON the last grapheme of the range, +/// never one past it, because that is where every pardes motion leaves it and +/// a cursor sitting one cell right of its own selection is a selection the +/// acme chords will not act on. +fn setDot(pane: *Pane, r: PaneFs.Range) void { + const text = bodyOf(pane); + const q0 = @min(@as(usize, r.q0), text.len); + const q1 = @max(q0, @min(@as(usize, r.q1), text.len)); + const a = modal.hxPos(text, q0); + pane.vsel = .{ .active = q1 > q0, .row = @intCast(a.row), .col = @intCast(a.col), .explicit = true }; + const h = modal.hxPos(text, if (q1 > q0) modal.prevGrapheme(text, q1) else q0); + pane.cur_row = @intCast(h.row); + pane.cur_col = @intCast(h.col); + pane.cur_pinned = true; + pane.sticky_col = -1; + pane.msel.active = false; + pane.ensureCursorVisible(); +} + +/// acme's `textshow`: put a spot on screen. Suppressed by `noscroll`. +fn showOffset(pane: *Pane, off: usize) void { + const text = bodyOf(pane); + const c = modal.hxPos(text, @min(off, text.len)); + pane.cur_row = @intCast(c.row); + pane.cur_col = @intCast(c.col); + pane.cur_pinned = true; + pane.sticky_col = -1; + pane.ensureCursorVisible(); +} + +/// acme's `clampaddr`. Its `Range` is signed and it clamps both ends; ours is +/// unsigned, so only the top can be wrong — and it can, the moment a pane's +/// body shrinks under a stored address. +fn clampAddr(pf: *PaneFs, len: usize) void { + const n = clip(len); + pf.addr.q0 = @min(pf.addr.q0, n); + pf.addr.q1 = @min(pf.addr.q1, n); + if (pf.limit) |*l| { + l.q0 = @min(l.q0, n); + l.q1 = @min(l.q1, n); + } +} + +/// Move a range across an edit at `at` that replaced `removed` bytes with +/// `inserted` — acme's `if(tq0 >= q0) tq0 += nr;`, applied to both ends, which +/// is what keeps a script rewriting text under your cursor from dragging the +/// cursor onto a different word. +fn shiftBy(r: PaneFs.Range, at: u32, removed: u32, inserted: u32) PaneFs.Range { + return .{ .q0 = shiftOne(r.q0, at, removed, inserted), .q1 = shiftOne(r.q1, at, removed, inserted) }; +} + +fn shiftOne(v: u32, at: u32, removed: u32, inserted: u32) u32 { + if (v <= at) return v; + if (v <= at +| removed) return at +| inserted; + return v - removed +| inserted; +} + +/// How many of these bytes end on a character boundary. +/// +/// acme buffers a partial rune on the Fid (`fullrunewrite` plus `f->rpart`) +/// and stitches it onto the next write. A SHORT COUNT is the POSIX spelling of +/// the same promise — the writer's libc retries with the tail — and it needs +/// no per-handle state at all. Never zero for a +/// non-empty write: a writer handed 0 retries the same bytes forever. +fn wholeUtf8(data: []const u8) usize { + var i = data.len; + var back: usize = 0; + while (i > 0 and back < 4) : (back += 1) { + i -= 1; + const c = data[i]; + if (c < 0x80) return data.len; // an ASCII tail is always complete + if (c & 0xC0 == 0xC0) { // a lead byte: is its sequence all here? + const need = std.unicode.utf8ByteSequenceLength(c) catch return data.len; + if (i + need <= data.len or i == 0) return data.len; + return i; + } + } + // four trailing continuation bytes and no lead: not UTF-8 at all. acme's + // `cvttorunes` substitutes for bad bytes rather than refusing them, and so + // does storing them verbatim. + return data.len; +} + +// =========================================================================== +// LOOKUP — acme's `fsyswalk`, minus 9P's fid bookkeeping. +// =========================================================================== + +/// The files inside a pane's directory, by name. `.dir` is the directory +/// itself and is never a name to resolve. +/// A name inside a pane's directory. `.` and `..` are the kernel's business, +/// never ours, and the directory variant is not nameable — so a hit on the +/// variant names is the whole lookup. +fn paneFileNamed(name: []const u8) ?PaneFile { + const f = std.meta.stringToEnum(PaneFile, name) orelse return null; + return if (f == .dir) null else f; +} + +fn topFileNamed(name: []const u8) ?TopFile { + const f = std.meta.stringToEnum(TopFile, name) orelse return null; + return if (f == .root) null else f; +} + +/// acme: "is it a numeric name? yes: it's a directory". A pane's directory is +/// named by its SERIAL, which is never reused, so a stale path can go stale +/// but can never come to mean a different pane. +fn serialNamed(name: []const u8) ?u32 { + if (name.len == 0 or name.len > 10) return null; + for (name) |c| if (c < '0' or c > '9') return null; + return std.fmt.parseInt(u32, name, 10) catch null; +} + +/// The smallest live serial greater than `after`, so a caller can walk every +/// pane in ascending serial without sorting anything. O(panes) per step over +/// at most sixteen slots, and no allocation — the alternative was a scratch +/// array in a function that must not allocate. +fn nextSerialAfter(p: *Pardes, after: u32) ?u32 { + var best: ?u32 = null; + for (p.panes) |slot| { + const pane = slot orelse continue; + if (pane.serial <= after) continue; + if (best == null or pane.serial < best.?) best = pane.serial; + } + return best; +} + +/// Create a pane the way the `New` builtin does — an empty scratch below the +/// active one, in its column — and answer its serial. +/// +/// acme has `newwindowthread` sitting on a channel for exactly this, and its +/// windows go wherever `rowadd` puts them. Going through `newScratchBelow` +/// means a pane a script opened is in every respect a pane you opened: same +/// tag, same builtins, same undo, same Del. +fn newPane(p: *Pardes) ?u32 { + const slot = p.freeSlot() orelse return null; + p.newScratchBelow(p.active); + const pane = p.panes[slot] orelse return null; + return pane.serial; +} + +fn lookup(p: *Pardes, req: Req, target: Target) Reply { + const name = req.data; + if (name.len == 0 or std.mem.indexOfScalar(u8, name, '/') != null) return Reply.fail(req.tag, E.NOENT); + const node: u64 = switch (target) { + .top => |f| switch (f) { + .root => root: { + if (topFileNamed(name)) |t| break :root @intFromEnum(t); + const serial = serialNamed(name) orelse return Reply.fail(req.tag, E.NOENT); + _ = p.paneBySerial(serial) orelse return Reply.fail(req.tag, E.NOENT); + break :root Node.of(serial, .dir); + }, + .new => new: { + // acme(4): "Accessing any file in new creates a new window." + // + // acme creates it one component EARLIER — `fsyswalk` sends on + // `cnewwindow` the moment it walks the name `new` itself. That + // cannot work over FUSE: the kernel CACHES the dentry for + // `new`, so a lookup there would fire once per mount and never + // again. Creating at the CHILD keeps the promise the man page + // makes (`echo hi > $PARDES_FS/new/body` opens a pane holding + // `hi`) under a protocol that caches. + // + // The name is checked BEFORE the pane is made, so a stat of + // `new/nosuchfile` leaves no litter. acme's walk creates the + // window first and then fails the second component, which + // leaves an empty window behind for every typo. + const want = paneFileNamed(name) orelse return Reply.fail(req.tag, E.NOENT); + const serial = newPane(p) orelse return Reply.fail(req.tag, E.NFILE); + break :new Node.of(serial, want); + }, + else => return Reply.fail(req.tag, E.NOTDIR), + }, + .pane => |t| pane: { + if (t.file != .dir) return Reply.fail(req.tag, E.NOTDIR); + _ = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + const f = paneFileNamed(name) orelse return Reply.fail(req.tag, E.NOENT); + break :pane Node.of(t.serial, f); + }, + }; + // A lookup answers with the TARGET's attributes, which is exactly what a + // getattr of that node would say — one spelling, so the two can never + // disagree about a size or a mode. + return switch (attrOf(p, Node.target(node) orelse return Reply.fail(req.tag, E.NOENT))) { + .ok => |a| .{ .tag = req.tag, .attr = a }, + .missing => Reply.fail(req.tag, E.NOENT), + }; +} + +// =========================================================================== +// READDIR +// =========================================================================== + +/// One directory entry in the transport-neutral staging format `src/fuse.zig` +/// decodes: node id, kind, name length, name — packed, little-endian, no +/// padding. A readdir answer is that record repeated. +/// +/// `node` travels because it becomes the `d_ino` a `getdents64` reports, and a +/// `d_ino` that disagrees with the later `st_ino` is a filesystem that lies to +/// `find -inum`. +fn stageDirent(out: *std.ArrayList(u8), gpa: std.mem.Allocator, node: u64, dir: bool, name: []const u8) void { + if (name.len == 0 or name.len > 255) return; + var head: [10]u8 = undefined; + std.mem.writeInt(u64, head[0..8], node, .little); + head[8] = @intFromBool(dir); + head[9] = @intCast(name.len); + out.appendSlice(gpa, &head) catch return; + out.appendSlice(gpa, name) catch return; +} + +/// The pane files, for a pane directory and for `new/`. `serial == 0` is +/// `new/`: there is no pane yet — the LOOKUP is what creates one — so there is +/// no id to report, and the transport substitutes one. +fn stagePaneFiles(p: *Pardes, out: *std.ArrayList(u8), serial: u32, skip: *u64) void { + inline for (comptime std.enums.values(PaneFile)) |f| { + if (f != .dir) { + if (skip.* > 0) skip.* -= 1 else stageDirent(out, p.gpa, Node.of(serial, f), false, f.name()); + } + } +} + +fn readdir(p: *Pardes, req: Req, target: Target) Reply { + const out = p.fs.stage(p.gpa); + var skip = req.off; + switch (target) { + .top => |f| switch (f) { + .root => { + inline for (.{ TopFile.index, TopFile.cons, TopFile.new }) |t| { + if (skip > 0) skip -= 1 else stageDirent(out, p.gpa, @intFromEnum(t), t.dir(), t.name()); + } + // Ascending serial: serials are never reused, so this order is + // stable across a create and a delete — which is what a script + // that walks the tree twice and diffs the two walks needs. + // acme lists windows in SCREEN order (column by column), which + // changes when you drag a window and says nothing a script can + // rely on. + var last: u32 = 0; + while (nextSerialAfter(p, last)) |s| { + last = s; + if (skip > 0) { + skip -= 1; + continue; + } + var buf: [16]u8 = undefined; + const name = std.fmt.bufPrint(&buf, "{d}", .{s}) catch continue; + stageDirent(out, p.gpa, Node.of(s, .dir), true, name); + } + }, + // `new/` ENUMERATES NOTHING, and that is a guarantee rather than a + // shrug: the names it could list are exactly the names whose LOOKUP + // creates a pane, and every tool that lists a directory then stats + // what it found — `ls -l`, `ls --color`, `find`, a shell completing + // `$PARDES_FS/new/` — would make one pane per name. acme never + // lists it either. Naming a file here is what creates one; see + // `lookup`. + .new => {}, + else => return Reply.fail(req.tag, E.NOTDIR), + }, + .pane => |t| { + if (t.file != .dir) return Reply.fail(req.tag, E.NOTDIR); + _ = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + stagePaneFiles(p, out, t.serial, &skip); + }, + } + // Zero bytes is END OF DIRECTORY, never an error: the transport stops + // asking, and re-staging from scratch on every call is what makes a + // partially consumed answer safe to ask for again at a higher cookie. + return .{ .tag = req.tag, .payload = .{ .staged = @intCast(out.items.len) } }; +} + +// =========================================================================== +// OPEN / RELEASE / SETATTR +// =========================================================================== + +/// Open carries no per-open state, because there is none to carry: `addr` and +/// `limit` belong to the pane (as they do in acme, where they are Window +/// fields), and every read brings its own offset. What an open DOES do is +/// arm the two things acme arms on open, and count event readers. +/// +/// So there is no fid table. acme needs one because 9P walks to a fid and +/// every later message names only that fid; FUSE puts the nodeid on every +/// request, RELEASE included, so the handle is decoration. It is answered +/// non-zero only because the transport spells "no handle" as zero. +fn open(p: *Pardes, req: Req, target: Target) Reply { + switch (target) { + .top => {}, + .pane => |t| { + const id = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + const pf = &p.fs.panes[id]; + switch (t.file) { + // acme(4): "When the ctl file is first opened, regular + // expression context searches in addr addresses examine the + // whole file"; `limit=addr` narrows them again. + .ctl => pf.limit = null, + // acme resets both on the FIRST open (`w->nopen[QWaddr]++ == + // 0`) and keeps a per-file open count to know. There is none + // here: `addr` is one piece of per-pane state that a second + // opener would be sharing anyway, so the honest reading of + // "first" is "whenever somebody opens it" — and a script's + // first act on `addr` is always to write one. + .addr => { + pf.addr = .{}; + pf.limit = null; + }, + // THE SUPPRESSION GATE. While this is non-zero the pane is + // script-driven: its Look and Exec are reported, not + // performed (`noteAction`). Counted per OPEN, not per pane, so + // two readers means the second one closing leaves the first + // still in charge. + .event => { + pf.readers +|= 1; + p.fs.listeners +|= 1; + }, + else => {}, + } + }, + } + return .{ .tag = req.tag, .handle = 1 }; +} + +fn release(p: *Pardes, req: Req) Reply { + const target = Node.target(req.node) orelse return .{ .tag = req.tag }; + switch (target) { + .top => {}, + .pane => |t| { + if (t.file != .event) return .{ .tag = req.tag }; + // The pane may have DIED while this was open. `State.forget` has + // then already taken its whole reader count out of `listeners` + // (the core calls it from `deinitPane`), so a serial that no + // longer resolves must not be decremented a second time — that + // underflow is exactly what would leave the editor suppressing + // button actions forever with no script left to interpret them. + const id = p.paneBySerial(t.serial) orelse return .{ .tag = req.tag }; + const pf = &p.fs.panes[id]; + if (pf.readers == 0) return .{ .tag = req.tag }; + pf.readers -= 1; + p.fs.listeners -|= 1; + // The LAST reader leaving takes the tag snapshot with it. It is + // only ever compared against while somebody is listening, so + // keeping it would let the tag drift unobserved and then hand the + // NEXT reader a `d`/`i` pair for a change it never saw. + if (pf.readers == 0) pf.tag_snap.clearAndFree(p.gpa); + }, + } + return .{ .tag = req.tag }; +} + +fn setattr(p: *Pardes, req: Req, target: Target) Reply { + // acme has NO equivalent: 9P has no truncate-on-open, so nothing in + // `xfid.c` answers a Twstat carrying a length. Linux does — `> body` is + // O_TRUNC — and refusing it would make the shell's most natural way to + // REPLACE a pane's text (rather than append to it) fail with EPERM on the + // redirect, before a single byte was written. So exactly one field is + // honoured, only the value zero means anything, and everything else a + // `stat` structure can carry (mode, owner, times) is silently accepted and + // ignored the way a filesystem of live editor state has to. + if (req.truncate) switch (target) { + .pane => |t| switch (t.file) { + .body, .data, .xdata => { + const id = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + const pane = p.panes[id].?; + if (fileOf(pane) != null) { + _ = spliceBody(p, id, pane, 0, bodyOf(pane).len, "") orelse + return Reply.fail(req.tag, E.NOMEM); + p.fs.panes[id].addr = .{}; + setDot(pane, .{}); + } + }, + else => {}, + }, + else => {}, + }; + return switch (attrOf(p, target)) { + .ok => |a| .{ .tag = req.tag, .attr = a }, + .missing => Reply.fail(req.tag, E.NOENT), + }; +} + +// =========================================================================== +// READ +// =========================================================================== + +/// Answer with a WINDOW onto what was just staged. `Payload.staged` is a +/// LENGTH from the start of the buffer, so a read at an offset slides the +/// bytes down rather than growing the payload union with a second field +/// nothing else would ever use. +fn staged(p: *Pardes, req: Req) Reply { + const out = &p.fs.out; + const off = @min(req.off, out.items.len); + const n = @min(out.items.len - off, req.size); + if (off > 0) std.mem.copyForwards(u8, out.items[0..n], out.items[off..][0..n]); + out.shrinkRetainingCapacity(n); + return .{ .tag = req.tag, .payload = .{ .staged = @intCast(n) } }; +} + +fn read(p: *Pardes, req: Req, target: Target) Reply { + switch (target) { + .top => |f| return switch (f) { + .index => readIndex(p, req), + // acme's dirtab: `cons` is 0200 and a directory is not read(2)able. + .cons, .root, .new => Reply.fail(req.tag, E.PERM), + }, + .pane => |t| { + const id = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + const pane = p.panes[id].?; + const pf = &p.fs.panes[id]; + return switch (t.file) { + .addr => readAddr(p, req, pf, pane), + .body => readBody(p, req, id, pane), + .ctl => readCtl(p, req, pane), + .data => readData(req, id, pane, pf, false), + .xdata => readData(req, id, pane, pf, true), + .tag => readTag(p, req, pane), + .event => readQueue(p, req, &pf.events), + .rdsel => readRdsel(req, id, pane), + .dir, .errors, .wrsel => Reply.fail(req.tag, E.PERM), + }; + }, + } +} + +/// acme's `Ctlsize`: five `%11d ` fields = 60 bytes, before the tag. +const ctl_fields = 5 * 12; + +/// acme's `winctlprint(w, buf, 0)` — the five numbers `index` and `ctl` share. +/// +/// COST: acme reads the tag's length off `w->tag.file->nc` for free, because +/// acme's tag IS a buffer. pardes's is COMPUTED every time it is asked for +/// (path, dirty marker, builtins, and the alignment gap, which is measured +/// against every other pane in the same layout column), so these five numbers +/// cost one tag render — a couple of microseconds and a few bumps of the +/// per-update scratch arena, which `Pardes.update` resets. That is the price +/// of the second field being the number a `tag` read will actually hand back; +/// a cheaper approximation that disagreed with `read tag` would be worse than +/// slow, it would be wrong. +fn stageCtlNumbers(p: *Pardes, out: *std.ArrayList(u8), pane: *Pane) void { + out.print(p.gpa, "{d:>11} {d:>11} {d:>11} {d:>11} {d:>11} ", .{ + pane.serial, + tagOf(p, pane).len, + bodyOf(pane).len, + // acme's `isdir` marks a window holding a DIRECTORY LISTING. pardes + // never opens one — a Look at a directory spawns a shell there + // (look.zig) — so this is structurally zero, not unimplemented. + @as(u32, 0), + @intFromBool(dirtyOf(pane)), + }) catch {}; +} + +/// acme's `xfidindexread`: one line per pane, the five numbers then the tag up +/// to its first newline. Seekable, so a script can pread the middle of it — +/// "at character position 5×12 starts the name of the window" (acme(4)). +fn readIndex(p: *Pardes, req: Req) Reply { + const out = p.fs.stage(p.gpa); + var last: u32 = 0; + while (nextSerialAfter(p, last)) |s| { + last = s; + const pane = p.panes[p.paneBySerial(s).?].?; + stageCtlNumbers(p, out, pane); + out.appendSlice(p.gpa, firstLine(tagOf(p, pane))) catch {}; + out.append(p.gpa, '\n') catch {}; + } + return staged(p, req); +} + +/// acme: `sprint(buf, "%11d %11d ", w->addr.q0, w->addr.q1)`. acme's numbers +/// are RUNE offsets; these are bytes (see the header). "Thus a regular +/// expression may be evaluated by writing it to addr and reading it back." +fn readAddr(p: *Pardes, req: Req, pf: *PaneFs, pane: *Pane) Reply { + clampAddr(pf, bodyOf(pane).len); + const out = p.fs.stage(p.gpa); + out.print(p.gpa, "{d:>11} {d:>11} ", .{ pf.addr.q0, pf.addr.q1 }) catch {}; + return staged(p, req); +} + +fn readBody(p: *Pardes, req: Req, id: usize, pane: *Pane) Reply { + if (pane.file != null) { + // ZERO COPY: `.region` is resolved by `fsPayload` during the effect + // drain, so reading a megabyte of body moves no bytes in here at all. + // This is the whole reason `Payload` is a union and not a slice. + const text = bodyOf(pane); + const off = @min(req.off, text.len); + const n = @min(text.len - off, req.size); + return .{ .tag = req.tag, .payload = .{ .region = .{ + .pane = @intCast(id), + .serial = pane.serial, + .off = clip(off), + .len = clip(n), + } } }; + } + // A TERMINAL has no such buffer. acme's body is always a `Text`; pardes's + // is a terminal emulator, and its "body" is the scrollback — which only + // becomes bytes when somebody renders the pages into lines. So it is + // produced, staged, and paid for per read. `win`'s transcript, read side. + const text = pane.vt.screens.active.dumpStringAlloc(p.gpa, .{ .screen = .{} }) catch + return Reply.fail(req.tag, E.NOMEM); + defer p.gpa.free(text); + const out = p.fs.stage(p.gpa); + out.appendSlice(p.gpa, text) catch return Reply.fail(req.tag, E.NOMEM); + return staged(p, req); +} + +/// The face the shell was last asked to wear. acme owns its fonts and prints +/// the real one; the core only knows what it REQUESTED — on a tty the font +/// belongs to the terminal emulator and in the browser to the page — so it +/// prints that, or `default`, which is the same word the Debug overlay shows +/// for the same reason. +fn fontName(p: *Pardes) []const u8 { + const name = p.settings.font.effective_name.get(); + return if (name.len == 0) "default" else name; +} + +/// plan9's `%q` (`quotestrfmt`): a string with nothing special in it prints +/// bare, anything else is wrapped in single quotes with internal quotes +/// doubled. Load-bearing rather than decoration — a script splits the ctl line +/// into shell words, and a font name with a space in it is one word. +fn stageQuoted(out: *std.ArrayList(u8), gpa: std.mem.Allocator, s: []const u8) void { + const plain = s.len > 0 and for (s) |c| { + if (c <= ' ' or c == '\'') break false; + } else true; + if (plain) { + out.appendSlice(gpa, s) catch {}; + return; + } + out.append(gpa, '\'') catch {}; + for (s) |c| { + if (c == '\'') out.append(gpa, '\'') catch {}; + out.append(gpa, c) catch {}; + } + out.append(gpa, '\'') catch {}; +} + +/// acme's `winctlprint(w, buf, 1)`: index's five numbers plus three more. +fn readCtl(p: *Pardes, req: Req, pane: *Pane) Reply { + const out = p.fs.stage(p.gpa); + stageCtlNumbers(p, out, pane); + // acme prints `Dx(w->body.r)` — the body's width in PIXELS — and + // `w->body.maxtab`, a tab's width in pixels too. pardes is a CELL GRID: + // on a tty there is no pixel width to report at all, and on the two pixel + // shells the number a script actually wants is still how many characters + // fit. So both are CELLS. A script that would have divided by the font + // width to get columns gets columns without dividing. + out.print(p.gpa, "{d:>11} ", .{pane.cols}) catch {}; + stageQuoted(out, p.gpa, fontName(p)); + out.print(p.gpa, " {d:>11} ", .{config.tab_width}) catch {}; + return staged(p, req); +} + +fn readTag(p: *Pardes, req: Req, pane: *Pane) Reply { + const out = p.fs.stage(p.gpa); + out.appendSlice(p.gpa, tagOf(p, pane)) catch {}; + return staged(p, req); +} + +/// acme's `xfidruneread`: hand back whole characters from the START of `addr` +/// and move `addr` to the null string just after them; `xdata` additionally +/// stops at the END of `addr` (acme passes `w->addr.q1` where `data` passes +/// `nc`). The file offset is ignored — `addr` is the position. +/// +/// "Whole characters" is acme's partial-rune rule; here it is a GRAPHEME +/// boundary, which is strictly stronger and is what every other offset in +/// pardes already respects. A read too small for the next grapheme returns +/// zero bytes rather than half of one — acme's `if(m == 0) break`. +fn readData(req: Req, id: usize, pane: *Pane, pf: *PaneFs, stop_at_end: bool) Reply { + const text = bodyOf(pane); + clampAddr(pf, text.len); + const q0: usize = pf.addr.q0; + // acme carries a "BUG: what should happen if q1 > q0?" here and answers by + // reading nothing. An inverted address is a legal thing to have written + // (`address()` never normalises), so the empty read is the answer. + const hi: usize = if (stop_at_end) @max(q0, @as(usize, pf.addr.q1)) else text.len; + var end = @min(hi, q0 +| req.size); + end = @max(q0, modal.graphemeStart(text, end)); + // `data` collapses the address onto the point it read up to; `xdata` moves + // only q0 and KEEPS q1, because q1 is the stop address the man page + // promises ("reads stop at the end address") and the next chunked read has + // to be able to continue from where this one stopped. acme spells the same + // difference at xfid.c:331-341: QWdata assigns both, QWxdata only q0. + pf.addr.q0 = clip(end); + if (!stop_at_end) pf.addr.q1 = clip(end); + if (pane.file == null) return .{ .tag = req.tag }; + return .{ .tag = req.tag, .payload = .{ .region = .{ + .pane = @intCast(id), + .serial = pane.serial, + .off = clip(q0), + .len = clip(end - q0), + } } }; +} + +/// acme copies the selection into a TEMP FILE at open, with a comment +/// apologising for it, so a `|sort` cannot see the text change underneath. +/// There is no such window here: the whole request is one main-thread +/// transaction, nothing can run between the open and the read, and the bytes +/// go out of the pane unmoved. +fn readRdsel(req: Req, id: usize, pane: *Pane) Reply { + if (pane.file == null) return .{ .tag = req.tag }; + const text = bodyOf(pane); + const d = dotOf(pane); + const lo = @min(@as(usize, d.q0), text.len); + const hi = @max(lo, @min(@as(usize, d.q1), text.len)); + const off = @min(req.off, hi - lo); + const n = @min(hi - lo - off, req.size); + return .{ .tag = req.tag, .payload = .{ .region = .{ + .pane = @intCast(id), + .serial = pane.serial, + .off = clip(lo + off), + .len = clip(n), + } } }; +} + +/// ONE RECORD PER READ, and `Status.again` when there is none. +/// +/// This is the whole of what acme's blocking `event` read becomes. acme parks +/// the `Xfid` in `w->eventx` and `winevent` sends it a message to wake it up; +/// the waiting lives in a thread per in-flight request, and `xfidflush` exists +/// to cancel one. Here nothing is consumed and nothing is remembered: the +/// transport still holds the kernel's request and asks again. No waiter list, +/// no wakeup, no flush bookkeeping, and no loop anywhere in the core. +fn readQueue(p: *Pardes, req: Req, q: *Queue) Reply { + const record = q.peek() orelse return .{ .tag = req.tag, .status = .again }; + // acme hands back as much of its event buffer as the count allows and + // keeps the rest, which can split a record down the middle; a reader is + // simply expected never to ask for less than one. Refusing is the honest + // version of that contract — half a record is unparseable and silently + // desynchronises the reader for the rest of the session. + if (req.size < record.len) return Reply.fail(req.tag, E.INVAL); + const out = p.fs.stage(p.gpa); + out.appendSlice(p.gpa, record) catch return Reply.fail(req.tag, E.NOMEM); + q.pop(); + return .{ .tag = req.tag, .payload = .{ .staged = @intCast(out.items.len) } }; +} + +// =========================================================================== +// WRITE +// =========================================================================== + +fn write(p: *Pardes, req: Req, target: Target) Reply { + switch (target) { + .top => |f| return switch (f) { + // acme(4): text written to `cons` appears in `dir/+Errors`, where + // `dir` is the directory the command ran in — acme knows which + // from the mount the writer inherited (`x->f->mntdir`, one per + // `win`). A FUSE mount is ONE directory for the whole editor, so + // the writing process is anonymous and the only defensible owner + // is the pane the user is in. A script that wants a specific + // pane's errors writes `<id>/errors`, which is unambiguous. + .cons => if (appendErrors(p, p.active, req.data)) |took| + .{ .tag = req.tag, .written = @intCast(took) } + else + Reply.fail(req.tag, E.IO), + else => Reply.fail(req.tag, E.PERM), + }, + .pane => |t| { + const id = p.paneBySerial(t.serial) orelse return Reply.fail(req.tag, E.NOENT); + const pane = p.panes[id].?; + return switch (t.file) { + .addr => writeAddr(p, req, id, pane), + .body => writeBody(p, req, id, pane), + .ctl => writeCtl(p, req, t.serial), + // acme's `data` and `xdata` differ only in what a READ stops + // at; the writes are the same code path there and here. + .data, .xdata => writeData(p, req, id, pane), + .tag => writeTag(p, req, pane), + .event => writeEvent(p, req, id), + .wrsel => writeWrsel(p, req, id, pane), + .errors => if (appendErrors(p, id, req.data)) |took| + .{ .tag = req.tag, .written = @intCast(took) } + else + Reply.fail(req.tag, E.IO), + .dir, .rdsel => Reply.fail(req.tag, E.PERM), + }; + }, + } +} + +/// THE ONE BODY SPLICE every writing file goes through: replace `[q0, q1)` +/// with `bytes`, via `file_pane.setContent` — which is where the core diffs +/// out the insert/delete event records, so a script's edit is reported exactly +/// once and in exactly the same shape as a keystroke's. One swap per write for +/// the same reason: two swaps would be two `D`/`I` pairs for one write. +/// +/// The origin character the records carry is `handle`'s, set once per request +/// (acme's winlock owner), so nothing here has to know which file it is +/// serving. +fn spliceBody(p: *Pardes, id: usize, pane: *Pane, q0: usize, q1: usize, bytes: []const u8) ?usize { + const f = fileOf(pane) orelse return null; + const take = if (bytes.len == 0) 0 else wholeUtf8(bytes); + const lo = @min(q0, f.content.len); + const hi = @max(lo, @min(q1, f.content.len)); + const new = p.gpa.alloc(u8, f.content.len - (hi - lo) + take) catch return null; + @memcpy(new[0..lo], f.content[0..lo]); + @memcpy(new[lo..][0..take], bytes[0..take]); + @memcpy(new[lo + take ..], f.content[hi..]); + // acme: `if(w->nomark == FALSE){ seq++; filemark(t->file); }` — `nomark` + // is how a script makes a batch of edits one Undo. + // + // COST, and the reason `nomark` matters more here than it does in acme: + // acme's `filemark` is a sequence number on a log-structured, disk-backed + // Buffer, so it is O(1). pardes's undo is a SNAPSHOT of the whole body + // (`file_pane.pushUndo` compares and then duplicates it), so a script that + // appends a line at a time to a megabyte body pays a megabyte per line and + // keeps 256 of them. That is exactly the same cost one KEYSTROKE pays on + // the same body — this is not a filesystem tax, it is the core's edit + // model — but a script can do it ten thousand times a second where a + // typist cannot. `nomark` is the documented remedy and the reason acme + // gave scripts the verb. + if (!p.fs.panes[id].nomark) file_pane.pushUndo(p, pane); + file_pane.setContent(p, f, new); + return take; +} + +/// acme(4): "Text written to body is always appended; the file offset is +/// ignored." So `req.off` is deliberately never read here. +fn writeBody(p: *Pardes, req: Req, id: usize, pane: *Pane) Reply { + if (req.data.len == 0) return .{ .tag = req.tag, .written = 0 }; + // A TERMINAL's body is not a document, it is a program's transcript — and + // the only way to put text into a transcript is to TYPE it. So a body + // write to a terminal pane is a pty write: `echo ls > $PARDES_FS/3/body` + // runs ls in pane 3's shell. That is `win`'s semantics in acme (the shell + // reads what you write to its window's body), reached through the effect + // the core already has instead of through a pipe. + // + // Nothing is RECORDED for it: the insert/delete diff lives in + // `file_pane.setContent`, and a terminal has no `file` to swap. The + // program's output comes back as ordinary `.output` bytes. + if (pane.file == null) { + const take = wholeUtf8(req.data); + p.emitWrite(id, req.data[0..take]); + return .{ .tag = req.tag, .written = @intCast(take) }; + } + const at = bodyOf(pane).len; + const take = spliceBody(p, id, pane, at, at, req.data) orelse + return Reply.fail(req.tag, E.NOMEM); + if (!p.fs.panes[id].noscroll) showOffset(pane, at + take); + return .{ .tag = req.tag, .written = @intCast(take) }; +} + +/// acme's tag is one `Text` and a write appends to all of it. pardes's tag is +/// PREFIX ++ TAIL: the prefix is chrome the core recomputes every frame (the +/// path, the dirty marker, the builtin words, the alignment gap), so bytes +/// appended to it would be gone by the next render. A tag write therefore +/// appends to the TAIL — which is the part that is a buffer, and the part a +/// script means when it writes ` Undo` into a tag. +/// +/// The tail is a fixed one-line buffer (`Pane.tag_tail`), so a write that does +/// not fit is short, and one with no room at all is ENOSPC rather than a zero +/// count the writer would retry forever. +fn writeTag(p: *Pardes, req: Req, pane: *Pane) Reply { + if (req.data.len == 0) return .{ .tag = req.tag, .written = 0 }; + // the laid-out default tail becomes real bytes on first touch, exactly as + // it does when you click into the tag + p.seedTail(pane); + const room = pane.tag_tail.len - pane.tag_tail_len; + if (room == 0) return Reply.fail(req.tag, E.NOSPC); + const take = wholeUtf8(req.data[0..@min(req.data.len, room)]); + @memcpy(pane.tag_tail[pane.tag_tail_len..][0..take], req.data[0..take]); + pane.tag_tail_len += take; + pane.tag_init = true; + return .{ .tag = req.tag, .written = @intCast(take) }; +} + +/// acme(4): text written to `data` "replaces the characters addressed by the +/// addr file and sets the address to the null string at the end of the written +/// text". The file offset is ignored. +fn writeData(p: *Pardes, req: Req, id: usize, pane: *Pane) Reply { + if (fileOf(pane) == null) return Reply.fail(req.tag, E.INVAL); + const pf = &p.fs.panes[id]; + clampAddr(pf, bodyOf(pane).len); + const q0: usize = pf.addr.q0; + const q1: usize = @max(q0, @as(usize, pf.addr.q1)); + const before = dotOf(pane); + // acme's winlock(w, 'F'): everything but body and tag is "an action + // through the window's other files". + const take = spliceBody(p, id, pane, q0, q1, req.data) orelse + return Reply.fail(req.tag, E.NOMEM); + setDot(pane, shiftBy(before, clip(q0), clip(q1 - q0), clip(take))); + pf.addr = .{ .q0 = clip(q0 + take), .q1 = clip(q0 + take) }; + if (!pf.noscroll) showOffset(pane, q0 + take); + return .{ .tag = req.tag, .written = @intCast(take) }; +} + +/// acme's `wrsel` cuts the selection when the file is OPENED and inserts each +/// write at a running point after it (`w->wrselrange`). Same result, no +/// open-time mutation: each write REPLACES the selection, and because the +/// selection is left collapsed just after the inserted text, a second write +/// appends to the first exactly as `wrselrange` does. The only difference is +/// what an open and close with NO write does — acme has already emptied the +/// selection by then, this leaves the pane untouched. A filesystem that edits +/// your document when you `stat` it is a filesystem you cannot explore. +/// +/// acme also forces `nomark` for the file's lifetime so the whole stream is +/// one Undo. That needs open-time state we do not keep; a script that wants it +/// writes `nomark` to `ctl`, which is the same button with a name on it. +fn writeWrsel(p: *Pardes, req: Req, id: usize, pane: *Pane) Reply { + if (fileOf(pane) == null) return Reply.fail(req.tag, E.INVAL); + const d = dotOf(pane); + const q0: usize = d.q0; + const q1: usize = @max(q0, @as(usize, d.q1)); + const take = spliceBody(p, id, pane, q0, q1, req.data) orelse + return Reply.fail(req.tag, E.NOMEM); + setDot(pane, .{ .q0 = clip(q0 + take), .q1 = clip(q0 + take) }); + return .{ .tag = req.tag, .written = @intCast(take) }; +} + +/// acme's `xfidwrite` QWaddr. Two failures, and acme has two error strings for +/// them: `Ebadaddr` (the parser stopped before the end of the expression) and +/// `Eaddr` (it parsed but did not evaluate — out of range, or no match). A +/// filesystem has one channel for "no", so both are EINVAL. +fn writeAddr(p: *Pardes, req: Req, id: usize, pane: *Pane) Reply { + const pf = &p.fs.panes[id]; + const text = bodyOf(pane); + clampAddr(pf, text.len); + // acme's parser stops at a newline of its own accord (`\n` reaches the + // `default:` arm), which is what lets `echo '/foo/' > addr` work from a + // shell. Trimming says the same thing without threading it through every + // arm of the state machine. + const expr = std.mem.trimEnd(u8, req.data, "\n"); + var a: Addr = .{ .text = text, .lim = pf.limit, .expr = expr }; + const r = a.address(pf.addr) orelse return Reply.fail(req.tag, E.INVAL); + if (a.i < expr.len) return Reply.fail(req.tag, E.INVAL); + pf.addr = r; + return .{ .tag = req.tag, .written = @intCast(req.data.len) }; +} + +// =========================================================================== +// THE ADDRESS LANGUAGE — acme's addr.c, byte-addressed. +// =========================================================================== + +/// mvzr PANICS on a pattern that ends inside an escape: `parseCharSet` slices +/// `in[i+1..]` and `valueFor` indexes `[0]` of it, so a trailing backslash is +/// an out-of-bounds read rather than a compile failure (pardes.zig's +/// `applySelRegex` carries the same warning about the prefix `[^\`). A live +/// typist can only reach that by accident; a SCRIPT's regex is untrusted +/// input, so it is screened here before the engine ever sees it. +fn safePattern(pat: []const u8) bool { + var i: usize = 0; + while (i < pat.len) : (i += 1) { + if (pat[i] != '\\') continue; + if (i + 1 >= pat.len) return false; + i += 1; + } + return true; +} + +/// THE ADDRESS PARSER, in acme's shape: one left-to-right pass with three +/// pieces of state — a running range, a DIRECTION (`+`/`-`/none) and a SIZE +/// (line or character) — recursing once per `,` or `;`. +/// +/// What is gone is the C. acme reads the expression through a `getc` callback +/// over a `Rune*` so one parser can serve both the filesystem and the Edit +/// language; it reports failure through two out-parameters (`evalp` for "did +/// not evaluate", `qp` for "stopped here") because it cannot return three +/// things; and it grows the regex pattern with `runerealloc` one rune at a +/// time. Here the expression is a slice, the cursor is a field, a pattern is a +/// subslice of the expression, and "did not evaluate" is `null`. +const Addr = struct { + text: []const u8, + /// `limit=addr`: regex context searches are confined to this. acme applies + /// it FORWARDS only, and so does this. + lim: ?PaneFs.Range, + expr: []const u8, + i: usize = 0, + /// One frame per `,` or `;`. acme recurses without a bound, which is fine + /// when the expression came from a person typing into a tag and is a + /// STACK OVERFLOW when it came from a script: `,,,,,...` a hundred + /// thousand deep is one write(2). A compound address deeper than this is + /// not an address anybody meant. + depth: u8 = 0, + + const max_depth = 32; + const Size = enum { char, line }; + + /// acme's `address()`. `ar` is what `.` means — and `xfidwrite` passes + /// `w->addr`, NOT the user's selection, so `.` is the CURRENT ADDRESS and + /// `addr=dot` is the only door the selection comes in by. (acme(4) + /// describes the language as "the format understood by button 3", where + /// `.` is dot; the code is the authority and this follows the code.) + fn address(a: *Addr, ar_in: PaneFs.Range) ?PaneFs.Range { + const start = a.i; + var ar = ar_in; + var r = ar_in; + var dir: u8 = 0; + var size: Size = .line; + var c: u8 = 0; + while (a.i < a.expr.len) { + const prevc = c; + c = a.expr[a.i]; + a.i += 1; + switch (c) { + ',', ';' => { + // `;` differs from `,` in one way: it makes the RIGHT side + // relative to the left one. + if (c == ';') ar = r; + if (prevc == 0) r.q0 = 0; // lhs defaults to 0 + if (a.i >= a.expr.len) { + r.q1 = clip(a.text.len); // rhs defaults to $ + } else { + if (a.depth >= max_depth) return null; + a.depth += 1; + const nr = a.address(ar) orelse return null; + a.depth -= 1; + r.q1 = nr.q1; + } + return r; + }, + '+', '-' => { + // a pending `+`/`-` with no count of its own means one + // line, unless what follows is itself an operand + if (prevc == '+' or prevc == '-') { + const nc = if (a.i < a.expr.len) a.expr[a.i] else 0; + if (nc != '#' and nc != '/' and nc != '?') + r = a.number(r, 1, prevc, .line) orelse return null; + } + dir = c; + }, + '.', '$' => { + // both are only meaningful as the FIRST character of a + // (sub)expression; anywhere else they end the parse + if (a.i != start + 1) { + a.i -= 1; + return r; + } + r = if (c == '.') ar else .{ .q0 = clip(a.text.len), .q1 = clip(a.text.len) }; + dir = if (a.i < a.expr.len) '+' else 0; + }, + '#', '0'...'9' => { + var digit = c; + if (c == '#') { + if (a.i >= a.expr.len or a.expr[a.i] < '0' or a.expr[a.i] > '9') { + a.i -= 1; + return r; + } + digit = a.expr[a.i]; + a.i += 1; + size = .char; + } + var n: u64 = digit - '0'; + while (a.i < a.expr.len) : (a.i += 1) { + const d = a.expr[a.i]; + if (d < '0' or d > '9') break; + n = @min(n * 10 + (d - '0'), std.math.maxInt(u32)); + } + r = a.number(r, @intCast(n), dir, size) orelse return null; + dir = 0; + size = .line; + }, + '/', '?' => { + const back = c == '?'; + r = a.regexp(r, a.pattern(c), back) orelse return null; + dir = 0; + size = .line; + }, + else => { + a.i -= 1; + return r; + }, + } + } + // a trailing `+` or `-` with nothing after it: one line that way + if (dir != 0) r = a.number(r, 1, dir, .line) orelse return null; + return r; + } + + /// The pattern between the delimiters, with the backslash of an escape + /// KEPT (it belongs to the regex engine, not to this parser). + /// + /// DIVERGENCE: acme closes both `/re/` and `?re?` on a `/` — its scanner + /// has no `case '?'` at all, so `?foo?` yields the pattern `foo?`, which + /// as a regex means `fo` plus an optional `o`. That is a bug you can only + /// find by reading addr.c. Here the OPENING delimiter closes. + fn pattern(a: *Addr, delim: u8) []const u8 { + const s = a.i; + while (a.i < a.expr.len) { + const c = a.expr[a.i]; + if (c == '\n') break; + a.i += 1; + if (c == '\\') { + if (a.i < a.expr.len) a.i += 1; + continue; + } + if (c == delim) return a.expr[s .. a.i - 1]; + } + return a.expr[s..a.i]; + } + + /// acme's `number()`, byte for byte — including its two oddities: a `-` + /// count from offset 0 wraps to the END of the file, and `:1-1` is legal + /// (it means `#0`) while `:1-2` is an error. + fn number(a: *Addr, r_in: PaneFs.Range, n: u32, dir: u8, size: Size) ?PaneFs.Range { + var r = r_in; + if (size == .char) { + var off: i64 = n; + if (dir == '+') { + off = @as(i64, r.q1) + n; + } else if (dir == '-') { + if (r.q0 == 0 and n > 0) r.q0 = clip(a.text.len); + off = @as(i64, r.q0) - n; + } + if (off < 0 or off > @as(i64, @intCast(a.text.len))) return null; + // BYTES, and a byte offset can land inside a grapheme where acme's + // rune offset never could. Clamped to the boundary at or before + // it, which is the rule every other offset in pardes follows. + const g = clip(modal.graphemeStart(a.text, @intCast(off))); + return .{ .q0 = g, .q1 = g }; + } + var line: i64 = n; + var q0: usize = r.q0; + var q1: usize = r.q1; + switch (dir) { + '-' => { + if (q0 < a.text.len) while (q0 > 0 and a.text[q0 - 1] != '\n') { + q0 -= 1; + }; + q1 = q0; + while (line > 0 and q0 > 0) { + if (a.text[q0 - 1] == '\n') { + line -= 1; + q1 = q0; + } + q0 -= 1; + } + if (line > 1) return null; + while (q0 > 0 and a.text[q0 - 1] != '\n') q0 -= 1; + return .{ .q0 = clip(q0), .q1 = clip(q1) }; + }, + '+' => { + if (q1 > 0) while (q1 < a.text.len and a.text[q1 - 1] != '\n') { + q1 += 1; + }; + q0 = q1; + }, + else => { + q0 = 0; + q1 = 0; + }, + } + while (line > 0 and q1 < a.text.len) { + const ch = a.text[q1]; + q1 += 1; + if (ch == '\n' or q1 == a.text.len) { + line -= 1; + if (line > 0) q0 = q1; + } + } + if (line > 0) return null; + return .{ .q0 = clip(q0), .q1 = clip(q1) }; + } + + /// acme's `regexp()`. Forward runs from the END of the running range to + /// the limit (`limit=addr`, else the end of the file); backward runs from + /// its START back to the beginning. + /// + /// The engine is mvzr, the one `%s` and the selection previews already + /// use. Two of its properties come along and cannot be fixed here: `^` and + /// `$` assert against the SLICE being searched rather than against a line, + /// and `.` matches a newline like any other byte. Both are already waived + /// in pardes.zig; an address that needs a line anchor matches `\n`. + fn regexp(a: *Addr, r: PaneFs.Range, pat: []const u8, back: bool) ?PaneFs.Range { + // acme reuses the LAST compiled expression for an empty pattern + // (`rxnull`). There is no such global here — one more piece of hidden + // state for a script to guess wrong about — so `//` is not an address. + if (pat.len == 0 or !safePattern(pat)) return null; + const re = mvzr.compile(pat) orelse return null; + if (back) { + const hi = @min(@as(usize, r.q0), a.text.len); + var best: ?mvzr.Match = null; + var at: usize = 0; + while (at < hi) { + const m = re.matchPos(at, a.text[0..hi]) orelse break; + best = m; + at = if (m.end > m.start) m.end else m.end + 1; + } + const m = best orelse return null; + return .{ .q0 = clip(m.start), .q1 = clip(m.end) }; + } + const hi = if (a.lim) |l| @min(@as(usize, l.q1), a.text.len) else a.text.len; + const from = @min(@as(usize, r.q1), hi); + const m = re.match(a.text[from..hi]) orelse return null; + return .{ .q0 = clip(from + m.start), .q1 = clip(from + m.end) }; + } +}; + +// =========================================================================== +// CTL VERBS — acme's xfidctlwrite. +// =========================================================================== + +/// The verbs that mean something here. acme matches PREFIXES with `strncmp` +/// and advances by the matched length, which is why its arms have to be +/// ordered `delete` before `del`, `nomark` before `mark`, `noscroll` before +/// `scroll` — get that ordering wrong and a verb is silently truncated into a +/// different one. Splitting on the newline the man page already requires and +/// matching WHOLE tokens makes that class of bug unrepresentable. +const Verb = enum { + @"addr=dot", + clean, + cleartag, + del, + delete, + dirty, + @"dot=addr", + get, + @"limit=addr", + mark, + nomark, + noscroll, + put, + scroll, + show, +}; + +/// ...and the ones acme has that pardes REFUSES. Loudly, because a silently +/// accepted no-op is the worse failure: the script believes it holds the lock. +/// +/// menu / nomenu — acme maintains `Undo Redo Put` in the LEFT HALF of the +/// tag and these switch that off. pardes's tag prefix is computed chrome +/// (the path, the dirty marker, the pane's own builtins) with no halves +/// and no writable menu region, so there is nothing to switch. +/// dump / dumpdir — acme's dump file stores a COMMAND that recreates a +/// window. pardes's dump (src/dump.zig) stores the window's TEXT, so a +/// recreation command has nowhere to be kept and nothing to run it. +/// font — the face belongs to the SHELL, not the core: on a tty it is the +/// terminal emulator's and in the browser it is the page's. The `Font` +/// builtin only ASKS; a ctl verb that looked like it set one would be a +/// lie on three of the four platforms. +/// lock / unlock — acme's exclusive-use lock is a `QLock` held against a 9P +/// fid. There is no fid here and the core is single-threaded, so a lock +/// would promise a mutual exclusion nothing can violate and nothing +/// provides. +const refused_verbs = [_][]const u8{ "dump", "dumpdir", "font", "lock", "menu", "nomenu", "unlock" }; + +fn verbIs(line: []const u8, word: []const u8) bool { + if (!std.mem.startsWith(u8, line, word)) return false; + return line.len == word.len or line[word.len] == ' '; +} + +/// acme's ctl write is NOT atomic: it applies verbs until one fails, then +/// answers `Ebadctl` with a count of the bytes it got through, so +/// `dirty\nbogus\n` leaves the window dirty and the write "fails". A short +/// count on a Linux write is not read as "the rest failed" by anybody, so the +/// only honest translation is all-or-nothing: validate every verb first, then +/// apply. `ctlVerb` answers the same yes/no in both passes. +fn writeCtl(p: *Pardes, req: Req, serial: u32) Reply { + for ([2]bool{ false, true }) |apply| { + // `del`'s guard is the one predicate that reads state EARLIER VERBS IN + // THE SAME WRITE change, so the validation pass has to model it or the + // two passes disagree: `clean\ndel` (acme's own idiom, and what + // examples/acmefs/life.py sends on the way out) would fail validation + // while `dirty\ndel` would pass it and then fail half-applied. + var dirty = if (p.paneBySerial(serial)) |id| dirtyOf(p.panes[id].?) else false; + var it = std.mem.splitScalar(u8, req.data, '\n'); + while (it.next()) |raw| { + const line = std.mem.trim(u8, raw, " \t\r"); + if (line.len == 0) continue; + // `del` and `delete` remove the pane, and the verbs after them in + // the same write have nothing left to act on. + const live = p.paneBySerial(serial) orelse if (apply) break else return Reply.fail(req.tag, E.NOENT); + if (!ctlVerb(p, live, line, apply, &dirty)) return Reply.fail(req.tag, E.INVAL); + } + } + return .{ .tag = req.tag, .written = @intCast(req.data.len) }; +} + +/// One verb. `apply` false is the validation pass and must change nothing but +/// `dirty`, which both passes advance identically so that `del`'s guard sees +/// the same answer in each. +fn ctlVerb(p: *Pardes, id: usize, line: []const u8, apply: bool, dirty: *bool) bool { + const pane = p.panes[id] orelse return false; + const pf = &p.fs.panes[id]; + + // The one verb with an argument. acme rejects a name containing any + // character `<= ' '` and an empty one; so does this. + if (verbIs(line, "name")) { + if (line.len <= 5) return false; + const name = std.mem.trim(u8, line[5..], " \t"); + if (name.len == 0) return false; + for (name) |c| if (c <= ' ') return false; + if (!apply) return true; + const f = fileOf(pane) orelse return true; // a terminal has no name to set + const copy = p.gpa.dupe(u8, name) catch return true; + p.gpa.free(f.path); + f.path = copy; + return true; + } + for (refused_verbs) |w| if (verbIs(line, w)) return false; + + const v = std.meta.stringToEnum(Verb, line) orelse return false; + // acme: `del` is "delete, but check dirty", `delete` is "delete for sure". + // pardes's `Del` builtin is unconditional (the guard there is the `*` you + // can see in the tag), so `del` gets acme's guard here and `delete` does + // not — which is the whole difference between the two words. + if (v == .del and dirty.*) return false; + switch (v) { + .dirty => dirty.* = true, + .clean, .get, .put => dirty.* = false, + else => {}, + } + if (!apply) return true; + + switch (v) { + .@"addr=dot" => pf.addr = dotOf(pane), + .@"dot=addr" => { + clampAddr(pf, bodyOf(pane).len); + setDot(pane, pf.addr); + }, + .@"limit=addr" => { + clampAddr(pf, bodyOf(pane).len); + pf.limit = pf.addr; + }, + // acme marks the window clean by resetting the file's sequence number; + // pardes's equivalent is "the revision on screen IS the saved one". + .clean => if (fileOf(pane)) |f| { + f.saved_revision = f.revision; + }, + .dirty => if (fileOf(pane)) |f| { + f.saved_revision = f.revision -% 1; + }, + // acme: "wipe tag right of bar". pardes's bar is the boundary between + // the computed prefix and the editable tail, so this empties the tail + // — and leaves it SEEDED, or the next render would put the default + // builtins straight back. + .cleartag => { + pane.tag_tail_len = 0; + pane.tag_init = true; + }, + .del, .delete => _ = p.executeBuiltinLine(id, "Del"), + .put => _ = p.executeBuiltinLine(id, "Save"), + // acme's `get`: "Equivalent to the Get interactive command with no + // arguments". pardes has no such builtin, so this is what Get would + // be — the same synchronous read `file_pane.open` does, through the + // same content swap, with an undo point in front of it so a script + // cannot discard your edits irrecoverably. + .get => if (fileOf(pane)) |f| { + if (output_pane.fileTraits(f.output).saves) { + if (look.readFile(p.gpa, f.path)) |bytes| { + file_pane.pushUndo(p, pane); + file_pane.setContent(p, f, bytes); + f.saved_revision = f.revision; + } else |_| {} + } + }, + // acme's `mark` both cancels `nomark` AND pushes a mark, so the edits + // made while nomark was on stay one Undo and the next one starts fresh. + .mark => { + pf.nomark = false; + file_pane.pushUndo(p, pane); + }, + .nomark => pf.nomark = true, + .noscroll => pf.noscroll = true, + .scroll => pf.noscroll = false, + .show => showOffset(pane, dotOf(pane).q0), + } + return true; +} + +// =========================================================================== +// EVENT WRITE-BACK — acme's xfideventwrite. +// =========================================================================== + +const EventRecord = struct { action: Action, q0: u32, q1: u32 }; + +/// `{origin}{type}{q0} {q1}\n`, acme's `xfideventwrite` parse: two characters, +/// two blank-separated decimals, a newline. Everything a full record carries +/// after that — the flag, the count, the text — is omitted on the way back in, +/// which is what acme(4) means by "with the flag, count, and text omitted". +/// +/// acme walks this with `strtoul`, pointer arithmetic and `goto Rescue`; here +/// the failure is `null` and the position stays in the struct, so the caller +/// can tell "ran out cleanly" from "stopped on garbage" by looking at `i`. +const EventReader = struct { + data: []const u8, + i: usize = 0, + + fn next(er: *EventReader) ?EventRecord { + if (er.i >= er.data.len) return null; + var i = er.i; + if (i + 2 > er.data.len) return null; + // acme stores the first character as `w->owner` (with a + // `/* disgusting */` beside it) so later records inherit whatever the + // writer claimed. Read and dropped here — see `writeEvent`. + i += 1; + const action = Action.fromChar(er.data[i]) orelse return null; + i += 1; + const q0 = scanNumber(er.data, &i) orelse return null; + const q1 = scanNumber(er.data, &i) orelse return null; + while (i < er.data.len and er.data[i] == ' ') i += 1; + if (i >= er.data.len or er.data[i] != '\n') return null; + er.i = i + 1; + return .{ .action = action, .q0 = q0, .q1 = q1 }; + } +}; + +fn scanNumber(data: []const u8, i: *usize) ?u32 { + while (i.* < data.len and data[i.*] == ' ') i.* += 1; + const s = i.*; + var n: u64 = 0; + while (i.* < data.len and data[i.*] >= '0' and data[i.*] <= '9') : (i.* += 1) + n = @min(n * 10 + (data[i.*] - '0'), std.math.maxInt(u32)); + if (i.* == s) return null; + return @intCast(n); +} + +/// Writing a record back PERFORMS the action it names, "exactly as it would +/// have been if the event file had not been open" (acme(4)). This is the +/// documented remote-control door and the point of the whole suppression rule: +/// a script reads an `X` record, decides the text is not one of its own tag +/// commands, and hands it back for pardes to run. +/// +/// It is also, deliberately, arbitrary code execution — an `X` record is an +/// Exec — which is why the mount is 0700 under the user's runtime directory. +/// +/// NOTHING in the write applies unless all of it parses: acme validates each +/// record just before executing it and leaves the earlier ones done, which +/// makes a malformed batch half-applied and unrepeatable. +fn writeEvent(p: *Pardes, req: Req, id: usize) Reply { + const pane0 = p.panes[id] orelse return Reply.fail(req.tag, E.NOENT); + const serial = pane0.serial; + { + const body = bodyOf(pane0); + const tag = tagOf(p, pane0); + var check: EventReader = .{ .data = req.data }; + while (check.next()) |r| { + switch (r.action) { + // acme accepts only `xXlL` on the way back in. A `D` or an `I` + // is a REPORT, not a request; writing one back would mean + // "pretend the user typed this", which nothing implements and + // acme's switch rejects with `Ebadevent`. + .body_look, .tag_look, .body_exec, .tag_exec => {}, + else => return Reply.fail(req.tag, E.INVAL), + } + // lower case is the tag, upper case the body — how a reader tells + // the two texts apart with no extra field + const n = if (r.action.onTag()) tag.len else body.len; + if (r.q0 > r.q1 or r.q1 > n) return Reply.fail(req.tag, E.INVAL); + } + if (check.i != req.data.len) return Reply.fail(req.tag, E.INVAL); + } + // acme(4): `F` is "actions through the window's other files", which is + // exactly what this is, and `handle` has already set it. acme takes the + // origin from the RECORD instead (`w->owner = *p++`, with a + // `/* disgusting */` beside it), so a writer can attribute its own action + // to the keyboard; the character is parsed here and dropped, because a + // record saying where it came from is worth nothing if the sender picks. + var run: EventReader = .{ .data = req.data }; + while (run.next()) |r| { + // an earlier action in this same write may have deleted the pane + const now = p.paneBySerial(serial) orelse break; + const pane = p.panes[now].?; + const whole = if (r.action.onTag()) tagOf(p, pane) else bodyOf(pane); + const lo = @min(@as(usize, r.q0), whole.len); + const hi = @max(lo, @min(@as(usize, r.q1), whole.len)); + // the action can replace the very text it is reading from + const text = p.scratch.allocator().dupe(u8, whole[lo..hi]) catch continue; + switch (r.action) { + .body_exec, .tag_exec => _ = p.execute(now, text), + .body_look, .tag_look => p.lookAt(now, text), + else => unreachable, + } + } + return .{ .tag = req.tag, .written = @intCast(req.data.len) }; +} + +// =========================================================================== +// +Errors — acme's `errorwin`. +// =========================================================================== + +/// acme(4): writing to `errors` "appends to the body of the dir/+Errors +/// window, where dir is the directory currently named in the tag. The window +/// is created if necessary, but not until text is actually written." +/// +/// One buffer per DIRECTORY, not per pane — which is why a search for an +/// existing one matches on the dirname and not on the writer. Answers HOW +/// MANY BYTES WERE TAKEN, or null for failure. +/// +/// The count matters because the append goes through `spliceBody`, which stops +/// at a whole-character boundary: the kernel splits a large `write(2)` at +/// `max_write` wherever it lands, so a multi-byte character straddling that +/// boundary must be reported short and retried by the writer's libc, exactly +/// as `body` and `data` do. Acknowledging the whole buffer would drop it. +fn appendErrors(p: *Pardes, id: usize, text: []const u8) ?usize { + if (text.len == 0) return 0; + const pane = p.panes[id] orelse return null; + const dir = dirOf(pane); + for (p.panes, 0..) |slot, i| { + const q = slot orelse continue; + const qf = fileOf(q) orelse continue; + const o = qf.output orelse continue; + if (std.meta.activeTag(o.from) != .errors) continue; + if (!std.mem.eql(u8, std.fs.path.dirname(qf.path) orelse "", dir)) continue; + return spliceBody(p, i, q, qf.content.len, qf.content.len, text); + } + const free = p.freeSlot() orelse return null; + const content = p.gpa.dupe(u8, text) catch return null; + const np = output_pane.open(p, free, dir, .errors, "", content) catch { + p.gpa.free(content); + return null; + }; + p.placeDoc(id, free, np); + return text.len; +} + +// =========================================================================== +// SIZES — what `stat` reports. +// =========================================================================== + +/// The whole `index`, measured. acme's `xfidindexread` walks every window to +/// size its buffer too; the tag of each is FORMATTED to be measured, into the +/// per-update scratch arena that is reset anyway, so this is bump allocation +/// rather than sixteen allocations a frame. +fn indexLen(p: *Pardes) u64 { + var n: u64 = 0; + for (p.panes) |slot| { + const pane = slot orelse continue; + n += ctl_fields + firstLine(tagOf(p, pane)).len + 1; + } + return n; +} + +/// A terminal's body has no length that is cheap AND honest — measuring it +/// means rendering the whole scrollback — so it reports zero and is served +/// with direct IO, exactly like `event` and `log`. +fn bodyLen(p: *Pardes, pane: *Pane) u64 { + _ = p; + return bodyOf(pane).len; +} + +fn tagLen(p: *Pardes, pane: *Pane) u64 { + return tagOf(p, pane).len; +} + +// =========================================================================== +// TESTS. +// +// The whole point of the split: every one of these drives `handle()` through +// the ordinary event queue with NO FUSE, NO mount, NO thread and no /dev/fuse +// anywhere. A filesystem whose semantics are a pure function of the core is a +// filesystem you can unit-test at the speed of a function call, and one whose +// blocking is a return value is one you can test without a scheduler. +// =========================================================================== + +const testing = std.testing; + +/// What a transport sees: the reply, and the bytes `fsPayload` resolved for it +/// inside the drain's borrow window. The two effects a filesystem operation +/// can additionally cause are captured too, because for `put` and for a write +/// to a terminal's body THE EFFECT IS THE ANSWER. +const Answer = struct { + reply: Reply = .{ .tag = 0, .status = .err, .errno = E.IO }, + bytes: []const u8 = "", + saved: bool = false, + pty_buf: [256]u8 = undefined, + pty_len: usize = 0, + + fn pty(a: *const Answer) []const u8 { + return a.pty_buf[0..a.pty_len]; + } + + fn errno(a: Answer) u16 { + return if (a.reply.status == .err) a.reply.errno else 0; + } +}; + +/// One request in, one answer out. Effects are DRAINED but not performed: a +/// `put` must be observable as a `.save_file` without a test writing to the +/// real filesystem. +fn call(p: *Pardes, req: Req) Answer { + p.update(.{ .fs_req = req }); + var ans: Answer = .{}; + while (p.nextEffect()) |e| switch (e) { + .fs_reply => |r| { + ans.reply = r; + ans.bytes = p.fsPayload(r); + }, + .save_file, .save_text => ans.saved = true, + .write => |w| { + const b = w.bytes.slice(); + const n = @min(b.len, ans.pty_buf.len - ans.pty_len); + @memcpy(ans.pty_buf[ans.pty_len..][0..n], b[0..n]); + ans.pty_len += n; + }, + else => {}, + }; + return ans; +} + +fn rd(p: *Pardes, node: u64, off: u64, size: u32) Answer { + return call(p, .{ .tag = 1, .op = .read, .node = node, .off = off, .size = size }); +} + +fn wr(p: *Pardes, node: u64, data: []const u8) Answer { + return call(p, .{ .tag = 2, .op = .write, .node = node, .data = data }); +} + +fn rdir(p: *Pardes, node: u64, skip: u64) Answer { + return call(p, .{ .tag = 4, .op = .readdir, .node = node, .off = skip, .size = 4096 }); +} + +fn look_up(p: *Pardes, dir: u64, name: []const u8) Answer { + return call(p, .{ .tag = 3, .op = .lookup, .node = dir, .data = name }); +} + +/// A core with one FILE pane holding `text`, which is what most of acme's +/// window files are about. Slot 0, and its serial is the directory name. +fn withFile(gpa: std.mem.Allocator, text: []const u8) !*Pardes { + const p = try Pardes.init(gpa, .{ .tty_only = true, .cols = 80, .rows = 24 }); + errdefer p.deinit(); + while (p.nextEffect()) |_| {} + _ = try p.hxOpenFileContent(text); + while (p.nextEffect()) |_| {} + return p; +} + +fn serialOf(p: *Pardes) u32 { + return p.panes[0].?.serial; +} + +const Dirent = struct { node: u64, dir: bool, name: []const u8 }; + +/// Decode the readdir staging format `src/fuse.zig` agreed to. +fn dirents(bytes: []const u8, out: []Dirent) []Dirent { + var n: usize = 0; + var i: usize = 0; + while (i + 10 <= bytes.len and n < out.len) { + const node = std.mem.readInt(u64, bytes[i..][0..8], .little); + const kind = bytes[i + 8]; + const len = bytes[i + 9]; + i += 10; + if (i + len > bytes.len) break; + out[n] = .{ .node = node, .dir = kind == 1, .name = bytes[i .. i + len] }; + i += len; + n += 1; + } + return out[0..n]; +} + +fn nameAt(list: []const Dirent, want: []const u8) ?Dirent { + for (list) |d| if (std.mem.eql(u8, d.name, want)) return d; + return null; +} + +test "readdir lists the root, a pane directory, and new/ without creating anything" { + const gpa = testing.allocator; + const p = try withFile(gpa, "hello\n"); + defer p.deinit(); + const serial = serialOf(p); + var buf: [32]Dirent = undefined; + + const root = rdir(p, @intFromEnum(TopFile.root), 0); + try testing.expectEqual(Status.ok, root.reply.status); + const top = dirents(root.bytes, &buf); + try testing.expectEqual(@as(usize, 4), top.len); + try testing.expectEqualStrings("index", top[0].name); + try testing.expectEqualStrings("cons", top[1].name); + try testing.expectEqualStrings("new", top[2].name); + try testing.expect(top[2].dir and !top[0].dir); + var idbuf: [16]u8 = undefined; + try testing.expectEqualStrings(try std.fmt.bufPrint(&idbuf, "{d}", .{serial}), top[3].name); + try testing.expect(top[3].dir); + // the id a readdir reports is the id a getattr will report + try testing.expectEqual(Node.of(serial, .dir), top[3].node); + + // `off` skips entries, and past the end is EOF, not an error + const rest = rdir(p, @intFromEnum(TopFile.root), 3); + try testing.expectEqual(@as(usize, 1), dirents(rest.bytes, &buf).len); + const eof = rdir(p, @intFromEnum(TopFile.root), 99); + try testing.expectEqual(Status.ok, eof.reply.status); + try testing.expectEqual(@as(usize, 0), eof.bytes.len); + + const dir = rdir(p, Node.of(serial, .dir), 0); + const files = dirents(dir.bytes, &buf); + try testing.expectEqual(@as(usize, 10), files.len); // dirtabw minus "." + try testing.expect(nameAt(files, "addr") != null); + try testing.expect(nameAt(files, "xdata") != null); + try testing.expect(nameAt(files, ".") == null); + try testing.expectEqual(Node.of(serial, .body), nameAt(files, "body").?.node); + + // acme(4) says accessing a file in `new` creates a window, so LISTING it + // must enumerate nothing at all: every name it could report is a name + // whose lookup creates a pane, and `ls -l` stats what a listing reported. + const before = p.next_serial; + const new = rdir(p, @intFromEnum(TopFile.new), 0); + try testing.expectEqual(Status.ok, new.reply.status); + try testing.expectEqual(@as(usize, 0), new.bytes.len); + try testing.expectEqual(before, p.next_serial); + + // a file is not a directory + try testing.expectEqual(E.NOTDIR, rdir(p, Node.of(serial, .body), 0).errno()); +} + +test "lookup resolves top files, pane serials and pane files" { + const gpa = testing.allocator; + const p = try withFile(gpa, "hello\n"); + defer p.deinit(); + const serial = serialOf(p); + const root = @intFromEnum(TopFile.root); + + try testing.expectEqual(@as(u64, @intFromEnum(TopFile.index)), look_up(p, root, "index").reply.attr.node); + try testing.expect(look_up(p, root, "new").reply.attr.dir); + try testing.expectEqual(E.NOENT, look_up(p, root, "nosuchthing").errno()); + + var idbuf: [16]u8 = undefined; + const dir = look_up(p, root, try std.fmt.bufPrint(&idbuf, "{d}", .{serial})); + try testing.expectEqual(Node.of(serial, .dir), dir.reply.attr.node); + try testing.expect(dir.reply.attr.dir); + // a serial that is not a live pane, and a serial that never existed + try testing.expectEqual(E.NOENT, look_up(p, root, "99999").errno()); + + const body = look_up(p, Node.of(serial, .dir), "body"); + try testing.expectEqual(Node.of(serial, .body), body.reply.attr.node); + // a lookup answers exactly what a getattr of the same node would + const stat = call(p, .{ .tag = 4, .op = .getattr, .node = Node.of(serial, .body) }); + try testing.expectEqual(body.reply.attr.size, stat.reply.attr.size); + try testing.expectEqual(@as(u64, "hello\n".len), stat.reply.attr.size); + try testing.expectEqual(E.NOENT, look_up(p, Node.of(serial, .dir), "editout").errno()); + try testing.expectEqual(E.NOTDIR, look_up(p, Node.of(serial, .body), "x").errno()); +} + +test "a lookup inside new/ creates a pane and resolves that pane's file" { + const gpa = testing.allocator; + const p = try withFile(gpa, "first\n"); + defer p.deinit(); + const before = serialOf(p); + + // a name that is not a pane file creates nothing + try testing.expectEqual(E.NOENT, look_up(p, @intFromEnum(TopFile.new), "bogus").errno()); + try testing.expectEqual(before, p.next_serial); + + const a = look_up(p, @intFromEnum(TopFile.new), "body"); + try testing.expectEqual(Status.ok, a.reply.status); + const made: Node = @bitCast(a.reply.attr.node); + try testing.expect(made.serial != before); + try testing.expectEqual(@intFromEnum(PaneFile.body), made.file); + + // ...and it is a real pane: `echo hi > new/body` leaves a pane holding hi + _ = wr(p, a.reply.attr.node, "hi"); + const id = p.paneBySerial(@intCast(made.serial)).?; + try testing.expectEqualStrings("hi", p.panes[id].?.file.?.content); +} + +test "index prints winctlprint's five fields then the tag" { + const gpa = testing.allocator; + const p = try withFile(gpa, "hello\nthere\n"); + defer p.deinit(); + const pane = p.panes[0].?; + + const a = rd(p, @intFromEnum(TopFile.index), 0, 4096); + try testing.expectEqual(Status.ok, a.reply.status); + var got: [512]u8 = undefined; + @memcpy(got[0..a.bytes.len], a.bytes); + const line = got[0..a.bytes.len]; + + const tag = tagOf(p, pane); + var want: std.ArrayList(u8) = .empty; + defer want.deinit(gpa); + try want.print(gpa, "{d:>11} {d:>11} {d:>11} {d:>11} {d:>11} {s}\n", .{ + pane.serial, tag.len, @as(usize, "hello\nthere\n".len), 0, 0, firstLine(tag), + }); + try testing.expectEqualStrings(want.items, line); + // acme(4): "at character position 5x12 starts the name of the window" + try testing.expectEqual(@as(usize, 60), std.mem.indexOf(u8, line, firstLine(tag)).?); + + // seekable: a script may pread the middle of it + const mid = rd(p, @intFromEnum(TopFile.index), 60, 5); + try testing.expectEqualStrings(firstLine(tag)[0..5], mid.bytes); + + // ...and a dirty pane says so in the fifth field + pane.file.?.saved_revision = pane.file.?.revision -% 1; + const dirty = rd(p, @intFromEnum(TopFile.index), 48, 12); + try testing.expectEqualStrings(" 1 ", dirty.bytes); +} + +test "ctl read is index's five fields plus width in cells, font and tab width" { + const gpa = testing.allocator; + const p = try withFile(gpa, "x\n"); + defer p.deinit(); + const pane = p.panes[0].?; + + const a = rd(p, Node.of(pane.serial, .ctl), 0, 4096); + try testing.expectEqual(Status.ok, a.reply.status); + var want: std.ArrayList(u8) = .empty; + defer want.deinit(gpa); + try want.print(gpa, "{d:>11} {d:>11} {d:>11} {d:>11} {d:>11} {d:>11} {s} {d:>11} ", .{ + pane.serial, tagOf(p, pane).len, @as(usize, 2), 0, 0, pane.cols, "default", config.tab_width, + }); + try testing.expectEqualStrings(want.items, a.bytes); + + // plan9 %q: a name with a space in it becomes one shell word + var quoted: std.ArrayList(u8) = .empty; + defer quoted.deinit(gpa); + stageQuoted("ed, gpa, "DejaVu Sans Mono"); + try testing.expectEqualStrings("'DejaVu Sans Mono'", quoted.items); + quoted.clearRetainingCapacity(); + stageQuoted("ed, gpa, "it's"); + try testing.expectEqualStrings("'it''s'", quoted.items); +} + +test "body reads at any offset and writes append" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + const body = Node.of(serial, .body); + + try testing.expectEqualStrings("one\ntwo\n", rd(p, body, 0, 100).bytes); + try testing.expectEqualStrings("two\n", rd(p, body, 4, 100).bytes); + try testing.expectEqualStrings("wo", rd(p, body, 5, 2).bytes); + try testing.expectEqualStrings("", rd(p, body, 999, 2).bytes); + // zero copy: the answer points INTO the pane, it is not a staged copy + try testing.expect(rd(p, body, 0, 100).bytes.ptr == p.panes[0].?.file.?.content.ptr); + + // acme(4): "Text written to body is always appended; the file offset is + // ignored" — so a write at offset 0 still lands at the end. + const w = call(p, .{ .tag = 5, .op = .write, .node = body, .off = 0, .data = "three\n" }); + try testing.expectEqual(@as(u32, 6), w.reply.written); + try testing.expectEqualStrings("one\ntwo\nthree\n", p.panes[0].?.file.?.content); + + // a write cut mid-character is SHORT, never split + const short = wr(p, body, "a\xC3"); + try testing.expectEqual(@as(u32, 1), short.reply.written); + try testing.expectEqualStrings("one\ntwo\nthree\na", p.panes[0].?.file.?.content); +} + +test "a body write to a terminal pane types at its shell" { + const gpa = testing.allocator; + const p = try Pardes.init(gpa, .{ .tty_only = true, .cols = 40, .rows = 10 }); + defer p.deinit(); + while (p.nextEffect()) |_| {} + const pane = p.panes[0].?; + try testing.expect(pane.isTerminal()); + + // `win`'s transcript semantics: the only way into a program's transcript + // is to type at it, so a body write becomes a pty write. + const a = wr(p, Node.of(pane.serial, .body), "ls -l\r"); + try testing.expectEqual(@as(u32, 6), a.reply.written); + try testing.expectEqualStrings("ls -l\r", a.pty()); + + // and a body READ renders the scrollback rather than lending a buffer + const r = rd(p, Node.of(pane.serial, .body), 0, 64); + try testing.expectEqual(Status.ok, r.reply.status); +} + +test "tag reads the whole tag and writes append to the editable tail" { + const gpa = testing.allocator; + const p = try withFile(gpa, "x\n"); + defer p.deinit(); + const pane = p.panes[0].?; + const node = Node.of(pane.serial, .tag); + + const whole = rd(p, node, 0, 4096); + try testing.expect(std.mem.startsWith(u8, whole.bytes, "/hxcase.txt")); + try testing.expect(std.mem.indexOf(u8, whole.bytes, "Del") != null); + + const before = rd(p, node, 0, 4096).bytes.len; + const w = wr(p, node, " Mine"); + try testing.expectEqual(@as(u32, 5), w.reply.written); + try testing.expect(std.mem.endsWith(u8, pane.tag_tail[0..pane.tag_tail_len], " Mine")); + const after = rd(p, node, 0, 4096); + try testing.expectEqual(before + 5, after.bytes.len); + try testing.expect(std.mem.endsWith(u8, after.bytes, " Mine")); + + // the tail is one bounded line; with no room left the file is FULL + pane.tag_tail_len = pane.tag_tail.len; + try testing.expectEqual(E.NOSPC, wr(p, node, "x").errno()); +} + +test "the address language, form by form" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\nthree\n"); // 14 bytes, three lines + defer p.deinit(); + const serial = serialOf(p); + const addr = Node.of(serial, .addr); + + const Case = struct { expr: []const u8, q0: u32, q1: u32 }; + for ([_]Case{ + .{ .expr = "#0", .q0 = 0, .q1 = 0 }, + .{ .expr = "#5", .q0 = 5, .q1 = 5 }, + .{ .expr = "0", .q0 = 0, .q1 = 0 }, + .{ .expr = "1", .q0 = 0, .q1 = 4 }, + .{ .expr = "2", .q0 = 4, .q1 = 8 }, + .{ .expr = "$", .q0 = 14, .q1 = 14 }, + .{ .expr = ",", .q0 = 0, .q1 = 14 }, + .{ .expr = "1,2", .q0 = 0, .q1 = 8 }, + .{ .expr = "#1,#4", .q0 = 1, .q1 = 4 }, + .{ .expr = "2+1", .q0 = 8, .q1 = 14 }, + .{ .expr = "$-1", .q0 = 8, .q1 = 14 }, + .{ .expr = "/two/", .q0 = 4, .q1 = 7 }, + .{ .expr = "/t.o/", .q0 = 4, .q1 = 7 }, + // a trailing newline is what a shell redirect leaves behind + .{ .expr = "1\n", .q0 = 0, .q1 = 4 }, + }) |c| { + // every case starts from a known address, so `.` and `+`/`-` are + // measured against the same place each time + _ = wr(p, addr, "#0"); + const w = wr(p, addr, c.expr); + try testing.expectEqual(Status.ok, w.reply.status); + const got = rd(p, addr, 0, 64); + var want: [32]u8 = undefined; + try testing.expectEqualStrings( + try std.fmt.bufPrint(&want, "{d:>11} {d:>11} ", .{ c.q0, c.q1 }), + got.bytes, + ); + } + + // `.` is the CURRENT ADDRESS (acme passes w->addr as `ar`), not the + // selection: set it, then ask for it back. + _ = wr(p, addr, "1"); + _ = wr(p, addr, "."); + try testing.expectEqual(@as(u32, 0), p.fs.panes[0].addr.q0); + try testing.expectEqual(@as(u32, 4), p.fs.panes[0].addr.q1); + + // `?re?` searches BACKWARD from the start of the running range and takes + // the LAST match before it — acme's `rxbexecute`. + _ = wr(p, addr, "$"); + _ = wr(p, addr, "?o?"); + try testing.expectEqual(@as(u32, 6), p.fs.panes[0].addr.q0); // the `o` in "two" + try testing.expectEqual(@as(u32, 7), p.fs.panes[0].addr.q1); + + // limit=addr confines a forward search + _ = wr(p, addr, "1"); + _ = wr(p, Node.of(serial, .ctl), "limit=addr\n"); + _ = wr(p, addr, "#0"); + try testing.expectEqual(E.INVAL, wr(p, addr, "/three/").errno()); + _ = wr(p, Node.of(serial, .ctl), "clean\n"); // any ctl write; limit stays + // ...and opening ctl clears it again (acme(4)) + _ = call(p, .{ .tag = 6, .op = .open, .node = Node.of(serial, .ctl) }); + try testing.expect(p.fs.panes[0].limit == null); + _ = wr(p, addr, "#0"); + try testing.expectEqual(Status.ok, wr(p, addr, "/three/").reply.status); + + // refusals + for ([_][]const u8{ "zzz", "#", "//", "/nomatch/", "1 2", "99", "/a\\" }) |bad| { + _ = wr(p, addr, "#0"); + try testing.expectEqual(E.INVAL, wr(p, addr, bad).errno()); + } + + // acme recurses once per `,` with no bound at all, which a script turns + // into a stack overflow with one write(2). Refused, not crashed. + const nested = "," ** 4096; + _ = wr(p, addr, "#0"); + try testing.expectEqual(E.INVAL, wr(p, addr, nested).errno()); +} + +test "data and xdata read from addr, move it, and write through it" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + const addr = Node.of(serial, .addr); + const data = Node.of(serial, .data); + const xdata = Node.of(serial, .xdata); + + _ = wr(p, addr, "#0"); + try testing.expectEqualStrings("one", rd(p, data, 0, 3).bytes); + // ...and the address is now the null string after what was returned + try testing.expectEqual(@as(u32, 3), p.fs.panes[0].addr.q0); + try testing.expectEqual(@as(u32, 3), p.fs.panes[0].addr.q1); + + // xdata stops at the END of the address where data would run on + _ = wr(p, addr, "1"); + try testing.expectEqualStrings("one\n", rd(p, xdata, 0, 100).bytes); + _ = wr(p, addr, "1"); + try testing.expectEqualStrings("one\ntwo\n", rd(p, data, 0, 100).bytes); + + // a write REPLACES the addressed text and leaves the address after it + _ = wr(p, addr, "1"); + const w = wr(p, data, "ONE\n"); + try testing.expectEqual(@as(u32, 4), w.reply.written); + try testing.expectEqualStrings("ONE\ntwo\n", p.panes[0].?.file.?.content); + try testing.expectEqual(@as(u32, 4), p.fs.panes[0].addr.q0); +} + +test "data never splits a grapheme, in either direction" { + const gpa = testing.allocator; + const p = try withFile(gpa, "\u{00e9}x\n"); // é is two bytes + defer p.deinit(); + const serial = serialOf(p); + _ = wr(p, Node.of(serial, .addr), "#0"); + // one byte is not enough for the first character: acme's `if(m == 0) break` + try testing.expectEqualStrings("", rd(p, Node.of(serial, .data), 0, 1).bytes); + _ = wr(p, Node.of(serial, .addr), "#0"); + try testing.expectEqualStrings("\u{00e9}", rd(p, Node.of(serial, .data), 0, 2).bytes); + + // and a write ending mid-character is short rather than corrupting + _ = wr(p, Node.of(serial, .addr), "#0"); + try testing.expectEqual(@as(u32, 1), wr(p, Node.of(serial, .data), "a\xC3").reply.written); +} + +test "rdsel reads the selection and wrsel replaces it" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + const ctl = Node.of(serial, .ctl); + + _ = wr(p, Node.of(serial, .addr), "#0,#3"); + try testing.expectEqual(Status.ok, wr(p, ctl, "dot=addr\n").reply.status); + try testing.expectEqualStrings("one", rd(p, Node.of(serial, .rdsel), 0, 100).bytes); + + // ...and the round trip back out is the identity, not a range that creeps + _ = wr(p, ctl, "addr=dot\n"); + try testing.expectEqual(@as(u32, 0), p.fs.panes[0].addr.q0); + try testing.expectEqual(@as(u32, 3), p.fs.panes[0].addr.q1); + + try testing.expectEqual(Status.ok, wr(p, Node.of(serial, .wrsel), "ONE").reply.status); + try testing.expectEqualStrings("ONE\ntwo\n", p.panes[0].?.file.?.content); + // a second write appends after the first, acme's `wrselrange` + _ = wr(p, Node.of(serial, .wrsel), "!"); + try testing.expectEqualStrings("ONE!\ntwo\n", p.panes[0].?.file.?.content); +} + +test "every ctl verb, and every refusal" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + const ctl = Node.of(serial, .ctl); + const pane = p.panes[0].?; + const pf = &p.fs.panes[0]; + + // several verbs in one write, which is what the man page promises + try testing.expectEqual(Status.ok, wr(p, ctl, "nomark\nnoscroll\ndirty\n").reply.status); + try testing.expect(pf.nomark and pf.noscroll and dirtyOf(pane)); + try testing.expectEqual(Status.ok, wr(p, ctl, "mark\nscroll\nclean\n").reply.status); + try testing.expect(!pf.nomark and !pf.noscroll and !dirtyOf(pane)); + + _ = wr(p, ctl, "cleartag\n"); + try testing.expectEqual(@as(usize, 0), pane.tag_tail_len); + + _ = wr(p, Node.of(serial, .addr), "2"); + _ = wr(p, ctl, "limit=addr\n"); + try testing.expectEqual(@as(u32, 4), pf.limit.?.q0); + _ = wr(p, ctl, "dot=addr\nshow\n"); + try testing.expectEqual(@as(i32, 1), pane.cur_row); + + try testing.expectEqual(Status.ok, wr(p, ctl, "name /tmp/renamed.txt\n").reply.status); + try testing.expectEqualStrings("/tmp/renamed.txt", pane.file.?.path); + // acme rejects a name with any character <= ' ' in it + try testing.expectEqual(E.INVAL, wr(p, ctl, "name two words\n").errno()); + try testing.expectEqual(E.INVAL, wr(p, ctl, "name\n").errno()); + try testing.expectEqualStrings("/tmp/renamed.txt", pane.file.?.path); + + // `put` is acme's Put, which is pardes's Save + try testing.expect(wr(p, ctl, "put\n").saved); + + // REFUSED, each for a reason that is not "unimplemented" — see + // `refused_verbs`. Silently accepting these is the worse failure. + for ([_][]const u8{ + "menu", "nomenu", "dump echo hi", "dumpdir /tmp", "font Go Mono", "lock", "unlock", "bogus", "DEL", + }) |bad| try testing.expectEqual(E.INVAL, wr(p, ctl, bad).errno()); + + // ATOMIC, which acme is not: an unknown verb aborts the WHOLE write. + try testing.expect(!dirtyOf(pane)); + try testing.expectEqual(E.INVAL, wr(p, ctl, "dirty\nbogus\n").errno()); + try testing.expect(!dirtyOf(pane)); +} + +test "ctl get reloads the pane from disk and del honours a dirty body" { + const gpa = testing.allocator; + var tmp = testing.tmpDir(.{}); + defer tmp.cleanup(); + try tmp.dir.writeFile(testing.io, .{ .sub_path = "note.txt", .data = "from disk\n" }); + var path_buf: [256]u8 = undefined; + const path = try std.fmt.bufPrint(&path_buf, ".zig-cache/tmp/{s}/note.txt", .{tmp.sub_path}); + + const p = try withFile(gpa, "in memory\n"); + defer p.deinit(); + const serial = serialOf(p); + const ctl = Node.of(serial, .ctl); + const pane = p.panes[0].?; + + var name: [std.fs.max_path_bytes + 8]u8 = undefined; + _ = wr(p, ctl, try std.fmt.bufPrint(&name, "name {s}\n", .{path})); + try testing.expectEqual(Status.ok, wr(p, ctl, "get\n").reply.status); + try testing.expectEqualStrings("from disk\n", pane.file.?.content); + // Get leaves the pane clean and the previous text one Undo away + try testing.expect(!dirtyOf(pane)); + try testing.expect(pane.file.?.undo_len > 0); + + // acme: `del` is "delete, but check dirty"; `delete` is "delete for sure" + _ = wr(p, ctl, "dirty\n"); + try testing.expectEqual(E.INVAL, wr(p, ctl, "del\n").errno()); + try testing.expect(p.paneBySerial(serial) != null); + // ...and a second pane so the last one closing does not quit the editor + _ = look_up(p, @intFromEnum(TopFile.new), "body"); + try testing.expectEqual(Status.ok, wr(p, ctl, "delete\n").reply.status); + try testing.expect(p.paneBySerial(serial) == null); +} + +test "errors and cons append to one +Errors buffer per directory" { + const gpa = testing.allocator; + const p = try withFile(gpa, "x\n"); + defer p.deinit(); + const serial = serialOf(p); + + const live = for (p.panes) |slot| { + if (slot) |q| if (q.file) |f| if (f.output) |o| if (std.meta.activeTag(o.from) == .errors) break q; + } else null; + try testing.expect(live == null); // "not until text is actually written" + + try testing.expectEqual(Status.ok, wr(p, Node.of(serial, .errors), "boom\n").reply.status); + _ = wr(p, @intFromEnum(TopFile.cons), "again\n"); + + var found: usize = 0; + for (p.panes) |slot| { + const q = slot orelse continue; + const f = q.file orelse continue; + const o = f.output orelse continue; + if (std.meta.activeTag(o.from) != .errors) continue; + found += 1; + try testing.expectEqualStrings("boom\nagain\n", f.content); + try testing.expectEqualStrings("/+Errors", f.path); + } + try testing.expectEqual(@as(usize, 1), found); +} + +test "setattr truncation empties the body and answers fresh attributes" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + + const a = call(p, .{ .tag = 7, .op = .setattr, .node = Node.of(serial, .body), .truncate = true }); + try testing.expectEqual(Status.ok, a.reply.status); + try testing.expectEqual(@as(u64, 0), a.reply.attr.size); + try testing.expectEqualStrings("", p.panes[0].?.file.?.content); + + // `> body` then a write is the shell's way of REPLACING a pane's text + _ = wr(p, Node.of(serial, .body), "new text\n"); + try testing.expectEqualStrings("new text\n", p.panes[0].?.file.?.content); + + // a setattr that sets no size changes nothing + const noop = call(p, .{ .tag = 8, .op = .setattr, .node = Node.of(serial, .body) }); + try testing.expectEqual(@as(u64, 9), noop.reply.attr.size); +} + +test "event records are acme's bytes, one per read, and .again when empty" { + const gpa = testing.allocator; + const p = try withFile(gpa, "Msg fs-ran\n"); + defer p.deinit(); + const serial = serialOf(p); + const event = Node.of(serial, .event); + + // nothing is recorded while nobody is listening + _ = noteAction(p, 0, .body_exec, 1, 4, flag_builtin, "sg "); + try testing.expect(p.fs.panes[0].events.empty()); + + const h = call(p, .{ .tag = 10, .op = .open, .node = event }); + try testing.expect(h.reply.handle != 0); + try testing.expectEqual(@as(u16, 1), p.fs.listeners); + + // an empty queue is `.again`: nothing consumed, ask me later. NEVER an + // error, and never a loop. + try testing.expectEqual(Status.again, rd(p, event, 0, 4096).reply.status); + + p.fs.origin = 'M'; + _ = noteAction(p, 0, .body_exec, 1, 4, flag_builtin, "ell"); + _ = noteAction(p, 0, .body_delete, 0, 3, 0, ""); + // `%c%c%d %d %d %d %s\n`, wind.c's winevent with the owner char in front + try testing.expectEqualStrings("MX1 4 1 3 ell\n", rd(p, event, 0, 4096).bytes); + try testing.expectEqualStrings("MD0 3 0 0 \n", rd(p, event, 0, 4096).bytes); + try testing.expectEqual(Status.again, rd(p, event, 0, 4096).reply.status); + + // one record per read: a read too small to hold one is refused rather + // than answered with half a record the reader cannot resynchronise from + _ = noteAction(p, 0, .body_look, 0, 3, flag_filename, "one"); + try testing.expectEqual(E.INVAL, rd(p, event, 0, 4).errno()); + try testing.expectEqualStrings("ML0 3 4 3 one\n", rd(p, event, 0, 4096).bytes); + + // text of 256 bytes or more is elided; the reader fetches it from `data` + const big = "z" ** max_record_text; + _ = noteAction(p, 0, .body_exec, 0, max_record_text, 0, big); + try testing.expectEqualStrings("MX0 256 0 0 \n", rd(p, event, 0, 4096).bytes); + + _ = call(p, .{ .tag = 11, .op = .release, .node = event, .handle = h.reply.handle }); + try testing.expectEqual(@as(u16, 0), p.fs.listeners); +} + +/// Drain a queue into `store` and return the records. Reading is destructive +/// and a `.staged` answer is only valid until the next request, so each record +/// is copied out as it arrives. +fn drainEvents(p: *Pardes, node: u64, store: []u8, out: [][]const u8) [][]const u8 { + var used: usize = 0; + var n: usize = 0; + while (n < out.len) { + const a = rd(p, node, 0, 4096); + if (a.reply.status != .ok) break; + @memcpy(store[used..][0..a.bytes.len], a.bytes); + out[n] = store[used..][0..a.bytes.len]; + used += a.bytes.len; + n += 1; + } + return out[0..n]; +} + +test "a write through the filesystem is reported once, attributed to the file it came through" { + const gpa = testing.allocator; + const p = try withFile(gpa, "one\ntwo\n"); + defer p.deinit(); + const serial = serialOf(p); + const event = Node.of(serial, .event); + _ = call(p, .{ .tag = 40, .op = .open, .node = event }); + var store: [4096]u8 = undefined; + var slots: [16][]const u8 = undefined; + _ = drainEvents(p, event, &store, &slots); + + // A body write is acme's `E`: "writes to the body or tag file". ONE pair + // per write, because the diff lives in `file_pane.setContent` and a write + // is one content swap — that is the contract with the core's hook, and + // emitting records from the handler as well is what it forbids. + _ = wr(p, Node.of(serial, .body), "three\n"); + const body_recs = drainEvents(p, event, &store, &slots); + try testing.expect(body_recs.len >= 1); + // ...and the record's TEXT here contains a newline of its own, which is + // exactly why `Queue` frames records by length instead of by line + try testing.expectEqualStrings("EI8 14 0 6 three\n\n", body_recs[0]); + // the write also made the pane dirty, so its TAG changed — and acme + // attributes that to the write too (`winsettag` runs inside the same + // `winlock(w, 'E')`), which is why the origin is set for the whole + // request and not just for the mutation. + for (body_recs[1..]) |r| { + try testing.expectEqual(@as(u8, 'E'), r[0]); + try testing.expect(Action.fromChar(r[1]).?.onTag()); + } + + // A `data` write is acme's `F`: "actions through the window's other + // files" — and a replacement is a delete then an insert, acme's order, + // with no text on the delete. + _ = wr(p, Node.of(serial, .addr), "1"); + _ = wr(p, Node.of(serial, .data), "ONE\n"); + const data_recs = drainEvents(p, event, &store, &slots); + try testing.expectEqual(@as(usize, 2), data_recs.len); + try testing.expectEqualStrings("FD0 3 0 0 \n", data_recs[0]); + try testing.expectEqualStrings("FI0 3 0 3 ONE\n", data_recs[1]); + try testing.expectEqual(Status.again, rd(p, event, 0, 4096).reply.status); +} + +test "two event readers each count once, and the second closing leaves the first" { + const gpa = testing.allocator; + const p = try withFile(gpa, "x\n"); + defer p.deinit(); + const event = Node.of(serialOf(p), .event); + + _ = call(p, .{ .tag = 12, .op = .open, .node = event }); + _ = call(p, .{ .tag = 13, .op = .open, .node = event }); + try testing.expectEqual(@as(u16, 2), p.fs.panes[0].readers); + try testing.expectEqual(@as(u16, 2), p.fs.listeners); + + // A release names the NODE, not a handle: FUSE carries the nodeid on every + // request, so there is no fid table to look one up in, and one release + // answers for one open. + _ = call(p, .{ .tag = 14, .op = .release, .node = event }); + try testing.expectEqual(@as(u16, 1), p.fs.panes[0].readers); + try testing.expect(p.fs.scripted(0)); // the pane is STILL script-driven + + // a release with nothing left to release changes nothing and is not an + // error, and neither is one naming a node that never counted + _ = call(p, .{ .tag = 15, .op = .release, .node = Node.of(serialOf(p), .body) }); + try testing.expectEqual(@as(u16, 1), p.fs.listeners); + + _ = call(p, .{ .tag = 16, .op = .release, .node = event }); + try testing.expectEqual(@as(u16, 0), p.fs.listeners); + try testing.expect(!p.fs.scripted(0)); + _ = call(p, .{ .tag = 17, .op = .release, .node = event }); + try testing.expectEqual(@as(u16, 0), p.fs.listeners); +} + +test "a pane deleted while its event file is open leaves no suppression behind" { + const gpa = testing.allocator; + const p = try withFile(gpa, "x\n"); + defer p.deinit(); + const serial = serialOf(p); + const event = Node.of(serial, .event); + // a second pane, so deleting the first does not quit the editor + _ = look_up(p, @intFromEnum(TopFile.new), "body"); + + const a = call(p, .{ .tag = 18, .op = .open, .node = event }); + const b = call(p, .{ .tag = 19, .op = .open, .node = event }); + try testing.expectEqual(@as(u16, 2), p.fs.listeners); + + _ = wr(p, Node.of(serial, .ctl), "delete\n"); + try testing.expect(p.paneBySerial(serial) == null); + // the core's `State.forget` took BOTH readers out with the pane + try testing.expectEqual(@as(u16, 0), p.fs.listeners); + + // ...and the two late releases must not underflow it back to 65535, which + // would suppress every button action in the editor forever + _ = call(p, .{ .tag = 20, .op = .release, .node = event, .handle = a.reply.handle }); + _ = call(p, .{ .tag = 21, .op = .release, .node = event, .handle = b.reply.handle }); + try testing.expectEqual(@as(u16, 0), p.fs.listeners); + + // every operation on the dead pane is ENOENT — acme's Edel + try testing.expectEqual(E.NOENT, rd(p, event, 0, 64).errno()); + try testing.expectEqual(E.NOENT, rd(p, Node.of(serial, .body), 0, 64).errno()); + try testing.expectEqual(E.NOENT, wr(p, Node.of(serial, .ctl), "clean\n").errno()); + try testing.expectEqual(E.NOENT, call(p, .{ .tag = 22, .op = .open, .node = event }).errno()); +} + +test "writing an event record back performs the action it names" { + const gpa = testing.allocator; + const p = try withFile(gpa, "Msg fs-ran\n"); + defer p.deinit(); + const serial = serialOf(p); + const event = Node.of(serial, .event); + const pane = p.panes[0].?; + + // an `X` record over the body text `Msg fs-ran` is an Exec of it + const w = wr(p, event, "FX0 10\n"); + try testing.expectEqual(Status.ok, w.reply.status); + try testing.expectEqualStrings("fs-ran", pane.msg[0..pane.msg_len]); + + // several records in one write + pane.msg_len = 0; + try testing.expectEqual(Status.ok, wr(p, event, "FX0 10\nFX0 10\n").reply.status); + try testing.expectEqualStrings("fs-ran", pane.msg[0..pane.msg_len]); + + // ...and nothing applies when any of it is malformed: acme's Ebadevent + pane.msg_len = 0; + for ([_][]const u8{ + "FX0 10\nFQ0 1\n", // unknown type character + "FX0 999\n", // out of range + "FX0 10", // no newline + "FX5 1\n", // q0 > q1 + "FD0 3\n", // a report, not a request + "F\n", + }) |bad| { + try testing.expectEqual(E.INVAL, wr(p, event, bad).errno()); + try testing.expectEqual(@as(usize, 0), pane.msg_len); + } + + // The action is attributed to the FILESYSTEM (`F`), never to whatever the + // writer put in the record's origin character — acme copies that byte + // into `w->owner` and lets a script claim its Exec came from the + // keyboard. + _ = call(p, .{ .tag = 23, .op = .open, .node = event }); + p.fs.origin = 'K'; + _ = wr(p, event, "KX0 10\n"); + try testing.expectEqual(@as(u8, 'F'), p.fs.origin); +} + diff --git a/src/config.zig b/src/config.zig index 3e3f2cb0..79fd7fc9 100644 --- a/src/config.zig +++ b/src/config.zig @@ -790,6 +790,10 @@ pub const changelog_buffer = "+Changelog"; /// The empty buffer New and Newcol open: no file behind it yet, so Save asks /// for a path (prefilled with the inherited directory). pub const scratch_buffer = "+New"; +/// acme's own `+Errors`, and the one output buffer no keystroke opens: a +/// script writes it, through a pane's `errors` file or the top-level `cons` +/// (src/acmefs.zig). +pub const errors_buffer = "+Errors"; // ============================================================================ // PART 3 — THE HELIX KEYMAP. READ THIS BEFORE RETARGETING ANYTHING BELOW. diff --git a/src/file_pane.zig b/src/file_pane.zig index 80c31edb..be3ff201 100644 --- a/src/file_pane.zig +++ b/src/file_pane.zig @@ -388,9 +388,33 @@ pub fn deinit(p: *Pardes, pane: *Pane, file: *State) void { for (file.redo[0..file.redo_len]) |snap| p.gpa.free(snap.content); } +/// TELL A SCRIPT WHAT CHANGED, when one is listening. +/// +/// acme reports edits from the two places that make them — `textinsert` and +/// `textdelete`, which already know their range — so a replacement arrives as +/// a `D` record and then an `I`. pardes has no such pair: every edit lands +/// here as a whole new buffer, so the range is recovered by DIFFING, and +/// `acmefs.noteReplace` owns both the diff and the D-then-I order. +/// +/// The cost is two vectorised scans of the content, and it is paid only while +/// a script holds an `event` file open (`p.fs.listeners`); the editor nobody +/// is scripting does one branch. The pane lookup is a walk of at most +/// MAX_PANES slots comparing the FILE pointer — a file pane's state is stored +/// inline in its pane, so that identifies the pane exactly. +fn reportEdit(p: *Pardes, f: *State, new: []const u8) void { + if (p.fs.listeners == 0) return; + const id = for (p.panes, 0..) |slot, i| { + const pane = slot orelse continue; + if (pane.file) |*state| if (state == f) break i; + } else return; + pardes.acmefs.noteReplace(p, id, false, f.content, new); +} + /// The ONE content swap. Everything that edits a file pane lands here, which -/// is what lets the line index above have a single invalidation point. +/// is what lets the line index above have a single invalidation point — and +/// is why one diff HERE is every body edit a script can be told about. pub fn setContent(p: *Pardes, f: *State, new: []u8) void { + reportEdit(p, f, new); p.gpa.free(f.content); f.content = new; f.revision +%= 1; diff --git a/src/fs_service.zig b/src/fs_service.zig new file mode 100644 index 00000000..6fdb9991 --- /dev/null +++ b/src/fs_service.zig @@ -0,0 +1,270 @@ +//! `pardes --fs`: what a native HOST has to decide to serve acme's control +//! filesystem. `acmefs.zig` owns the semantics and `fuse.zig` owns the kernel; +//! what is left, and lives here, is three decisions — WHERE to mount (derive a +//! per-session point, or take the one the user named), WHEN to drain (one +//! frame's batch, in the order fuse.zig's two queues require), and WHAT A PANE +//! SHELL IS TOLD about it (`PARDES_FS`/`PARDES_PANE`, exported before the +//! fork). +//! +//! It exists because tty.zig and gui.zig would otherwise each carry the same +//! forty lines through two different loops; the only thing that genuinely +//! differs between them is how a background thread wakes the loop, and that is +//! a function pointer. A session without `--fs` allocates nothing here, starts +//! no thread, and costs one null check per frame. +const std = @import("std"); +const libc = std.c; +const pardes = @import("pardes.zig"); +const fuse = @import("fuse.zig"); + +/// Diagnostics land on a pane's message row, not on stderr: in the tty shell +/// stderr IS the screen (see main.zig's logFn, which drops every scope for +/// exactly that reason). The log line is the `PARDES_LOG=1` copy, where the +/// mount point and the errno name are worth having. +const log = std.log.scoped(.fs); + +// std.c has getenv but neither setter, same as nested.zig. +extern "c" fn setenv(name: [*:0]const u8, value: [*:0]const u8, overwrite: c_int) c_int; +extern "c" fn unsetenv(name: [*:0]const u8) c_int; + +/// How many kernel requests one frame will answer before handing the loop back +/// to the renderer. A `find $PARDES_FS` or a script in a `while true` loop can +/// produce them faster than a frame takes, and an uncapped drain would render +/// only when the script paused. Hitting the cap is not a stall: `drain` says so +/// and the caller wakes its own loop, so the batch continues on the next pass +/// with one frame drawn in between. +const max_batch = 64; + +/// Where per-session mounts live: `$XDG_RUNTIME_DIR/pardes` else +/// `~/.local/state/pardes`, and `<that>/<pid>` is this session's mount point. +/// +/// NOT `nested.socketDir`, though it answers a related question. That one +/// returns `$XDG_RUNTIME_DIR` itself, because a socket is a FILE whose name +/// (`pardes-<pid>.sock`) already namespaces it. A mount point is a DIRECTORY +/// per pid, and `fuse.sweepStale` unmounts and removes every `<digits>` entry +/// it finds — so it needs a parent that contains nothing but our mounts, which +/// under `$XDG_RUNTIME_DIR` means one more level. The HOME fallback already has +/// that level, which is why the two strings coincide there and only there. +/// The two strings are parameters rather than `getenv` calls so the tests below +/// need not mutate the process environment. That is not fastidiousness: a test +/// binary shares one environ, and unsetting HOME here once took down an +/// unrelated subprocess test three files away. +fn parentFrom(buf: *[std.fs.max_path_bytes:0]u8, xdg: ?[]const u8, home: ?[]const u8) ?[:0]const u8 { + if (xdg) |x| return std.fmt.bufPrintSentinel(buf, "{s}/pardes", .{x}, 0) catch null; + const h = home orelse return null; + return std.fmt.bufPrintSentinel(buf, "{s}/.local/state/pardes", .{h}, 0) catch null; +} + +fn envSlice(name: [*:0]const u8) ?[]const u8 { + return if (libc.getenv(name)) |v| std.mem.span(v) else null; +} + +fn parentDir(buf: *[std.fs.max_path_bytes:0]u8) ?[:0]const u8 { + return parentFrom(buf, envSlice("XDG_RUNTIME_DIR"), envSlice("HOME")); +} + +/// The mount point itself, from `Options.fs`: EMPTY means a bare `--fs`, so +/// derive `<parent>/<pid>`, and anything else is the `--fs=<dir>` the user +/// named, which wins verbatim — scripts and the snapshot harness need a name +/// they can predict. +/// +/// A named point must be absolute for the reason fuse.zig gives: the path is +/// handed to a setuid helper that resolves it against its OWN cwd, so a +/// relative one names somewhere else. Passing it through unresolved rather than +/// rooting it here keeps that one rule in one place; `Fs.mount` returns +/// `error.MountPathNotAbsolute`. +fn mountPoint(buf: *[std.fs.max_path_bytes:0]u8, named: []const u8, parent: ?[]const u8) ?[:0]const u8 { + if (named.len != 0) return std.fmt.bufPrintSentinel(buf, "{s}", .{named}, 0) catch null; + const dir = parent orelse return null; + // unsigned: {d} prints a leading '+' for a positive SIGNED int + return std.fmt.bufPrintSentinel(buf, "{s}/{d}", .{ dir, @as(u32, @intCast(libc.getpid())) }, 0) catch null; +} + +/// Sweep, derive, mount. Null when the session did not ask for a filesystem — +/// and also when it asked and the mount failed, which is deliberately the same +/// answer: a missing `fuse3`, a `user_allow_other`-less config or a kernel +/// without FUSE must cost the user their scripting, never their session. The +/// failure is reported once, on pane 0's message row, and everything else runs. +/// +/// Call after the core exists and before the first frame: the mount is live the +/// moment it returns, so a script racing startup finds a filesystem whose panes +/// are already there. +pub fn start(gpa: std.mem.Allocator, core: *pardes.Pardes) ?*fuse.Fs { + const named = core.opts.fs orelse return null; + var parent_buf: [std.fs.max_path_bytes:0]u8 = undefined; + const parent = parentDir(&parent_buf); + var buf: [std.fs.max_path_bytes:0]u8 = undefined; + const point = mountPoint(&buf, named, parent) orelse { + core.reportError(0, "fs mount", error.NoRuntimeDirectory); + return null; + }; + // Both of these are about a point we DERIVED. A `--fs=<dir>` the user named + // is not a directory we are entitled to unmount other things out of, its + // siblings are not ours to guess about, and it is not ours to remove on the + // way out either — `owns_dir` is what keeps `Fs.deinit` from rmdir'ing a + // directory the user made. + const derived = named.len == 0; + if (derived) if (parent) |dir| fuse.sweepStale(dir); + const fs = fuse.Fs.mount(gpa, .{ .mount = point, .owns_dir = derived }) catch |err| { + log.warn("--fs: cannot mount at {s}: {t}", .{ point, err }); + core.reportError(0, "fs mount", err); + return null; + }; + log.info("--fs: serving {s}", .{point}); + return fs; +} + +/// Start the one background thread, if there is a filesystem to start it for. +/// It waits for POLLIN on `/dev/fuse` and calls `wake(ctx)` — nothing else; it +/// never touches the core, the descriptor's data, or a request. Both hosts pass +/// a one-line callback that posts their own wake event, which is the ONLY thing +/// that differs between them here. +/// +/// A thread that will not spawn is not a filesystem that will not work: the +/// frame poll drains the same requests either way, so the loss is wake latency +/// (a script waits for the next event to arrive from anywhere) and the session +/// is not worth failing over it. That is also the documented no-parallelism +/// backend: skip this call entirely and everything still works. +pub fn wake(fs: ?*fuse.Fs, ctx: ?*anyopaque, callback: *const fn (?*anyopaque) void) void { + const f = fs orelse return; + f.wakeThread(ctx, callback) catch |err| + log.warn("--fs: no poll thread ({t}); draining once per frame instead", .{err}); +} + +/// What one frame's worth of filesystem work amounted to. Two separate facts, +/// because the two hosts need different ones: an interactive loop asks whether +/// to re-arm itself, while the headless grid harness asks whether anything +/// happened at all — its contract is one frame per event, and a request that +/// changed a pane IS an event. +pub const Drained = struct { + /// Requests answered, parked retries included. + count: usize = 0, + /// The cap stopped the batch with requests still waiting in the kernel. + pending: bool = false, +}; + +/// One frame's worth of filesystem work. +/// +/// The two loops are both to null and in this order, which is fuse.zig's +/// contract rather than a preference: +/// +/// - `retry()`'s null ENDS AND RESETS the round, so a caller that took one +/// parked request per frame would leave the second-oldest blocked reader +/// waiting 32 frames. The round is bounded by the park table, so it needs +/// no cap of its own. +/// - `next()`'s null is what acknowledges the drain to the poll thread. That +/// handshake is what stops a level-triggered `poll()` from spinning a core, +/// which is why `pending` has to keep the loop hot: no ack has been sent, +/// so nothing else will wake us. +pub fn drain(fs: *fuse.Fs, core: *pardes.Pardes) Drained { + var d: Drained = .{}; + while (fs.retry()) |req| { + step(fs, core, req); + d.count += 1; + } + while (d.count < max_batch) { + const req = fs.next() orelse return d; + step(fs, core, req); + d.count += 1; + } + d.pending = true; + return d; +} + +/// One request, one answer, and nothing in between: `req.data` borrows storage +/// the next `next()` overwrites, and the `.fs_reply` this emits is drained +/// before the loop can move on — so the borrow window is a single step, exactly +/// as the design contract requires. The reply normally reaches `Fs.reply` +/// through the host's `push_fs_reply`, because the payload bytes are resolved +/// by `pardes.fsPayload` inside `perform` and are only valid there. +/// +/// The exception is the `if` at the end. The core's effect ring is bounded and +/// `emit` DROPS on overflow, which for every other effect costs a repaint and +/// for this one costs a foreign process: an unanswered FUSE request leaves its +/// writer in uninterruptible sleep and its park slot used forever, and 32 of +/// those make the whole mount answer EAGAIN. One `ctl` write reaches the cap +/// (`put` emits a `.save_file` per line). So this loop, which is the only place +/// that knows a request is outstanding, watches the effects it performs for the +/// answer and invents an EIO when none came. +fn step(fs: *fuse.Fs, core: *pardes.Pardes, req: pardes.acmefs.Req) void { + core.update(.{ .fs_req = req }); + var answered = false; + while (core.nextEffect()) |e| { + if (e == .fs_reply and e.fs_reply.tag == req.tag) answered = true; + core.perform(e); + } + if (!answered) { + const eio = pardes.acmefs.Reply.fail(req.tag, pardes.acmefs.E.IO); + fs.reply(&eio, ""); + } +} + +/// What a pane shell is told about the filesystem: `PARDES_FS` is the mount and +/// `PARDES_PANE` is this pane's serial, so a script run inside a pane addresses +/// its own window with no arguments. That pair is acme's `winid` (exec.c), and +/// the serial rather than the slot index because slots are reused and serials +/// never are — `$PARDES_FS/$PARDES_PANE/body` must not start naming somebody +/// else's pane after a close. +/// +/// Exported in the PARENT, immediately before the fork, and this is the one +/// place pardes cannot copy acme. acme calls `putenv` in the child, which is +/// safe there because `rfork(RFENVG)` has just given that child a private +/// environment group. A Linux fork has no such thing, and `setenv` between fork +/// and exec can deadlock on an allocator lock some other thread held at fork +/// time — the same rule that already forces `shell_bin.resolve` above the fork +/// in both hosts. The cost is that pardes's own environ carries the +/// last-spawned pane's number; nothing in pardes reads it, and a subprocess +/// that inherits it was spawned on behalf of a pane anyway. +/// +/// With no filesystem the pair is REMOVED rather than left alone. A pardes +/// started inside a pardes that does serve one inherits both variables from its +/// parent's pane shell, and a session with no mount of its own must not hand +/// its panes an address that resolves to a window in someone else's session. +pub fn exportPaneEnv(fs: ?*const fuse.Fs, serial: u32) void { + const f = fs orelse { + _ = unsetenv("PARDES_FS"); + _ = unsetenv("PARDES_PANE"); + return; + }; + _ = setenv("PARDES_FS", f.path.ptr, 1); + var buf: [16:0]u8 = undefined; + const id = std.fmt.bufPrintSentinel(&buf, "{d}", .{serial}, 0) catch return; + _ = setenv("PARDES_PANE", id.ptr, 1); +} + +const testing = std.testing; + +test "the mount point is one level below a per-user parent, named by our pid" { + // $XDG_RUNTIME_DIR is shared with every other program in the session, so + // the mounts need a `pardes/` of their own under it — the level + // nested.socketDir does not have, and the reason this is not that function. + // Asserted rather than merely described, because `fuse.sweepStale` unmounts + // and removes every `<digits>` entry in whatever directory it is handed. + var parent: [std.fs.max_path_bytes:0]u8 = undefined; + const dir = parentFrom(&parent, "/run/user/1000", "/home/tester").?; + try testing.expectEqualStrings("/run/user/1000/pardes", dir); + var buf: [std.fs.max_path_bytes:0]u8 = undefined; + var expect: [std.fs.max_path_bytes]u8 = undefined; + try testing.expectEqualStrings( + try std.fmt.bufPrint(&expect, "{s}/{d}", .{ dir, @as(u32, @intCast(libc.getpid())) }), + mountPoint(&buf, "", dir).?, + ); +} + +test "no XDG_RUNTIME_DIR falls back to the home state directory, which has the level already" { + var parent: [std.fs.max_path_bytes:0]u8 = undefined; + try testing.expectEqualStrings( + "/home/tester/.local/state/pardes", + parentFrom(&parent, null, "/home/tester").?, + ); +} + +test "a session with no filesystem removes an inherited address rather than passing it on" { + // PARDES_FS/PARDES_PANE are ours alone, and this leaves them the way an + // --fs-less session leaves them: absent. Nothing else in the test binary + // reads either name, which is why this is the one env-touching test here. + _ = setenv("PARDES_FS", "/run/user/1000/pardes/999", 1); + _ = setenv("PARDES_PANE", "7", 1); + exportPaneEnv(null, 3); + try testing.expect(libc.getenv("PARDES_FS") == null); + try testing.expect(libc.getenv("PARDES_PANE") == null); +} diff --git a/src/fuse.zig b/src/fuse.zig new file mode 100644 index 00000000..311d887b --- /dev/null +++ b/src/fuse.zig @@ -0,0 +1,2709 @@ +//! The `/dev/fuse` transport for pardes's acme control filesystem: wire codec, +//! mount and unmount through `fusermount3`, one `poll()` thread, and the park +//! table that turns acme's blocking `event` read into "ask me again later". +//! +//! Raw protocol, no libfuse. libfuse is a thread pool, a request dispatcher and +//! a session lifetime — three things pardes already has and would have to fight. +//! What is left once those are removed is a struct layout and a read/write loop, +//! which is this file. It links nothing; the only external program it runs is +//! the setuid `fusermount3` helper, because an unprivileged process cannot +//! `mount(2)` in the initial user namespace and that helper exists precisely to +//! hand back a `/dev/fuse` descriptor for a mount it made on our behalf. +//! +//! The whole file is one side of a strict division of labour: +//! +//! - `acmefs.zig` owns the semantics and knows nothing about FUSE. It speaks +//! `Req`/`Reply` and never blocks. +//! - this file owns the kernel's opinions and knows nothing about panes. It +//! answers, in place, every request the core has no business seeing (INIT, +//! FORGET, INTERRUPT, DESTROY and the whole ENOSYS family), and translates +//! the eleven that remain. +//! - the host loop (tty/gui) owns the ordering: `retry()` to null, `next()` +//! to null, one `update()` per request, effects drained in between. +//! +//! THREADING. The main thread owns the descriptor for read and for write. The +//! poll thread never touches its data, never sees a `Req`, and never calls into +//! the core; it waits for POLLIN, calls the host's wake callback, and then +//! blocks until the main thread has drained. That last handshake is not +//! decoration: `poll()` is level triggered, so a poller that re-polls +//! immediately would spin a core at 100% for as long as one unanswered request +//! sits in the kernel queue. A host with no threads at all skips `wakeThread` +//! and drains from its frame poll; it loses wake latency and nothing else. +//! +//! BLOCKING. A FUSE server blocks a reader by simply not answering, and that is +//! the one and only way (the kernel gives no meaning to an EAGAIN reply). So +//! `Status.again` means "held": the request moves into the park table with its +//! bytes copied out of the read buffer, and `retry()` offers it back once per +//! frame until the core has something to say. Two obligations come with that: +//! +//! 1. a SIGKILLed reader whose request is never answered ends in +//! *uninterruptible* sleep (`fuse_dev`'s final `wait_event` is not +//! killable), so it survives its own kill until we reply. FUSE_INTERRUPT +//! is the escape hatch and is honoured below. +//! 2. teardown must answer everything still parked, and must abort the +//! connection by closing the descriptor before unmounting, or a reader +//! that raced the shutdown is stuck in D state with nobody left to wake +//! it. +//! +//! Linux only, guarded the way `file_watch.zig` guards inotify: every entry +//! point returns the inert answer off Linux, so a macOS or web build compiles +//! and mounts nothing. Only `mount()` can create an `Fs`, so off Linux no other +//! function in this file is ever reached. +//! +//! Verified against `/usr/include/linux/fuse.h` (7.45) and `fs/fuse/{dev,inode, +//! file,dir,readdir}.c`; the comptime size assertions below turn a header drift +//! into a compile error rather than a wedged mount nobody can unmount. +const std = @import("std"); +const builtin = @import("builtin"); +const libc = std.c; +const linux = std.os.linux; +const acmefs = @import("acmefs.zig"); + +/// Everything below the mount is Linux kernel ABI. Off Linux the module still +/// compiles (it is imported by the shared native shell) and does nothing. +const supported = builtin.os.tag == .linux; + +// --------------------------------------------------------------------------- +// wire protocol +// --------------------------------------------------------------------------- + +/// The protocol version this server speaks. A mismatch in the *major* aborts +/// the connection outright (`fuse_init_finish`: `arg->major != +/// FUSE_KERNEL_VERSION` -> `ok = false` -> the mount is dead on arrival), so +/// there is nothing to negotiate there. +const kernel_version: u32 = 7; + +/// The highest minor these structs were checked against (see the module +/// header). The INIT reply carries `@min(kernel_minor, what the kernel +/// offered)`: `fuse_init_finish` stores our number as `fc->minor`, and the +/// kernel then sizes the replies it reads back from us by it (the +/// `FUSE_COMPAT_*_SIZE` family in `fs/fuse/`), so echoing a *newer* kernel's +/// minor promises reply fields these structs do not have. Capping costs +/// nothing: with `flags = 0` no feature depends on the number. +const kernel_minor: u32 = 45; + +/// `fuse_dev_do_read` refuses to hand over a request when the server's read +/// buffer is smaller than this, and answers the *client* EIO instead: every +/// syscall through the mount fails and nothing says why. +const min_read_buffer: usize = 8192; + +/// `FUSE_REC_ALIGN`. A dirent record that is not a multiple of 8 desynchronises +/// the kernel's parse of the rest of the reply, so one bad name turns the whole +/// directory into garbage rather than into an error. +const rec_align: usize = 8; + +/// `FUSE_NAME_OFFSET` — the fixed part of a `fuse_dirent`, before the name. +const dirent_name_offset: usize = @sizeOf(fuse_dirent); + +fn recAlign(n: usize) usize { + return (n + rec_align - 1) & ~(rec_align - 1); +} + +/// The subset of `enum fuse_opcode` this server can receive. Non-exhaustive on +/// purpose: a newer kernel adds opcodes, and `@enumFromInt` of an unlisted +/// value into an exhaustive enum is undefined behaviour — the one bug in a +/// protocol decoder that cannot be diagnosed from the outside. +const Opcode = enum(u32) { + lookup = 1, + forget = 2, + getattr = 3, + setattr = 4, + readlink = 5, + symlink = 6, + mknod = 8, + mkdir = 9, + unlink = 10, + rmdir = 11, + rename = 12, + link = 13, + open = 14, + read = 15, + write = 16, + statfs = 17, + release = 18, + fsync = 20, + setxattr = 21, + getxattr = 22, + listxattr = 23, + removexattr = 24, + flush = 25, + init = 26, + opendir = 27, + readdir = 28, + releasedir = 29, + fsyncdir = 30, + getlk = 31, + setlk = 32, + setlkw = 33, + access = 34, + create = 35, + interrupt = 36, + bmap = 37, + destroy = 38, + ioctl = 39, + poll = 40, + notify_reply = 41, + batch_forget = 42, + fallocate = 43, + readdirplus = 44, + rename2 = 45, + lseek = 46, + copy_file_range = 47, + setupmapping = 48, + removemapping = 49, + syncfs = 50, + tmpfile = 51, + statx = 52, + copy_file_range_64 = 53, + _, +}; + +/// `FATTR_SIZE`. The only setattr bit this filesystem reads: without +/// `FUSE_ATOMIC_O_TRUNC` (which `flags = 0` deliberately does not negotiate) +/// the kernel strips `O_TRUNC` from the OPEN and issues a separate +/// `SETATTR(size = 0)`, so this bit *is* how `> file` reaches the core. +const FATTR_SIZE: u32 = 1 << 3; + +/// `FUSE_GETATTR_FH` — says the `fh` field of `fuse_getattr_in` is meaningful. +/// Reading `fh` without checking it hands the core a stale handle from an +/// unrelated open. +const FUSE_GETATTR_FH: u32 = 1 << 0; + +/// `FOPEN_DIRECT_IO`. Without it the kernel serves reads out of the page cache +/// and coalesces them, which for this filesystem is wrong in both directions: +/// a second `cat` of `index` would return the first one's bytes, and a blocking +/// `event` read would never reach us at all. +const FOPEN_DIRECT_IO: u32 = 1 << 0; + +const fuse_in_header = extern struct { + len: u32, + opcode: u32, + unique: u64, + nodeid: u64, + uid: u32, + gid: u32, + pid: u32, + total_extlen: u16, + padding: u16, +}; + +const fuse_out_header = extern struct { + len: u32, + @"error": i32, + unique: u64, +}; + +const fuse_init_in = extern struct { + major: u32, + minor: u32, + max_readahead: u32, + flags: u32, + flags2: u32, + unused: [11]u32, +}; + +const fuse_init_out = extern struct { + major: u32, + minor: u32, + max_readahead: u32, + flags: u32, + max_background: u16, + congestion_threshold: u16, + max_write: u32, + time_gran: u32, + max_pages: u16, + map_alignment: u16, + flags2: u32, + max_stack_depth: u32, + request_timeout: u16, + unused: [11]u16, +}; + +const fuse_attr = extern struct { + ino: u64, + size: u64, + blocks: u64, + atime: u64, + mtime: u64, + ctime: u64, + atimensec: u32, + mtimensec: u32, + ctimensec: u32, + mode: u32, + nlink: u32, + uid: u32, + gid: u32, + rdev: u32, + blksize: u32, + flags: u32, +}; + +const fuse_entry_out = extern struct { + nodeid: u64, + generation: u64, + entry_valid: u64, + attr_valid: u64, + entry_valid_nsec: u32, + attr_valid_nsec: u32, + attr: fuse_attr, +}; + +const fuse_attr_out = extern struct { + attr_valid: u64, + attr_valid_nsec: u32, + dummy: u32, + attr: fuse_attr, +}; + +const fuse_getattr_in = extern struct { + getattr_flags: u32, + dummy: u32, + fh: u64, +}; + +const fuse_setattr_in = extern struct { + valid: u32, + padding: u32, + fh: u64, + size: u64, + lock_owner: u64, + atime: u64, + mtime: u64, + ctime: u64, + atimensec: u32, + mtimensec: u32, + ctimensec: u32, + mode: u32, + unused4: u32, + uid: u32, + gid: u32, + unused5: u32, +}; + +const fuse_open_in = extern struct { + flags: u32, + open_flags: u32, +}; + +const fuse_open_out = extern struct { + fh: u64, + open_flags: u32, + backing_id: i32, +}; + +const fuse_read_in = extern struct { + fh: u64, + offset: u64, + size: u32, + read_flags: u32, + lock_owner: u64, + flags: u32, + padding: u32, +}; + +const fuse_write_in = extern struct { + fh: u64, + offset: u64, + size: u32, + write_flags: u32, + lock_owner: u64, + flags: u32, + padding: u32, +}; + +const fuse_write_out = extern struct { + size: u32, + padding: u32, +}; + +const fuse_release_in = extern struct { + fh: u64, + flags: u32, + release_flags: u32, + lock_owner: u64, +}; + +const fuse_flush_in = extern struct { + fh: u64, + unused: u32, + padding: u32, + lock_owner: u64, +}; + +const fuse_forget_in = extern struct { + nlookup: u64, +}; + +const fuse_batch_forget_in = extern struct { + count: u32, + dummy: u32, +}; + +const fuse_interrupt_in = extern struct { + unique: u64, +}; + +const fuse_kstatfs = extern struct { + blocks: u64, + bfree: u64, + bavail: u64, + files: u64, + ffree: u64, + bsize: u32, + namelen: u32, + frsize: u32, + padding: u32, + spare: [6]u32, +}; + +const fuse_statfs_out = extern struct { + st: fuse_kstatfs, +}; + +/// The `name` array is flexible in C and therefore absent here; this struct IS +/// `FUSE_NAME_OFFSET`, and `dirent_name_offset` is taken from its size so the +/// encoder and the kernel cannot disagree about where a name starts. +const fuse_dirent = extern struct { + ino: u64, + off: u64, + namelen: u32, + type: u32, +}; + +/// `DT_*` from `linux/dirent.h`, as `fuse_dirent.type` wants them. +const DT_DIR: u32 = 4; +const DT_REG: u32 = 8; + +/// `S_IFMT` bits. `Reply.Attr.mode` carries permissions only, so the format +/// nibble is ours to add; a `fuse_attr.mode` with no format bits is a file of +/// no type and `stat(2)` through the mount returns something no tool expects. +const S_IFDIR: u32 = 0o040000; +const S_IFREG: u32 = 0o100000; + +// A drifted header is a mount that hangs with no diagnostic, so every struct +// on the wire asserts its size here. These numbers are `sizeof` from +// /usr/include/linux/fuse.h at FUSE_KERNEL_MINOR_VERSION 45; they are frozen +// ABI and are not allowed to change under us silently. +comptime { + std.debug.assert(@sizeOf(fuse_in_header) == 40); + std.debug.assert(@sizeOf(fuse_out_header) == 16); + std.debug.assert(@sizeOf(fuse_init_in) == 64); + std.debug.assert(@sizeOf(fuse_init_out) == 64); + std.debug.assert(@sizeOf(fuse_attr) == 88); + std.debug.assert(@sizeOf(fuse_entry_out) == 128); + std.debug.assert(@sizeOf(fuse_attr_out) == 104); + std.debug.assert(@sizeOf(fuse_getattr_in) == 16); + std.debug.assert(@sizeOf(fuse_setattr_in) == 88); + std.debug.assert(@sizeOf(fuse_open_in) == 8); + std.debug.assert(@sizeOf(fuse_open_out) == 16); + std.debug.assert(@sizeOf(fuse_read_in) == 40); + std.debug.assert(@sizeOf(fuse_write_in) == 40); + std.debug.assert(@sizeOf(fuse_write_out) == 8); + std.debug.assert(@sizeOf(fuse_release_in) == 24); + std.debug.assert(@sizeOf(fuse_flush_in) == 24); + std.debug.assert(@sizeOf(fuse_forget_in) == 8); + std.debug.assert(@sizeOf(fuse_batch_forget_in) == 8); + std.debug.assert(@sizeOf(fuse_interrupt_in) == 8); + std.debug.assert(@sizeOf(fuse_kstatfs) == 80); + std.debug.assert(@sizeOf(fuse_statfs_out) == 80); + std.debug.assert(@sizeOf(fuse_dirent) == 24); + // The one field offset the codec depends on beyond struct sizes: the body + // of every request starts here, and 40 is a multiple of 8, which is what + // lets the parse point a struct at the read buffer instead of copying. + std.debug.assert(@sizeOf(fuse_in_header) % rec_align == 0); +} + +// --------------------------------------------------------------------------- +// the neutral readdir staging format +// --------------------------------------------------------------------------- + +/// How `acmefs` hands a directory listing to this file. The core is protocol +/// neutral by design, so it must not stage `fuse_dirent`s: those carry an +/// alignment rule, a cookie rule and a `DT_*` table that are the kernel's +/// business, not the editor's. It stages this instead, packed and repeated, +/// little endian, into `State.out`: +/// +/// node: u64 the acmefs node id of the entry, never 0 (see below) +/// kind: u8 0 = regular file, 1 = directory +/// namelen: u8 1..255, never 0 +/// name: [namelen]u8 +/// +/// `node` travels so that the `d_ino` a `getdents64` sees is the same number a +/// later `stat` reports. Synthesising one here instead would make `find -inum` +/// and every hardlink-detecting tool lie about this filesystem. +/// +/// `node` is never 0. It used to be, for the entries under `new/`: those name +/// panes that do not exist, because acme creates the pane when the name is +/// LOOKED UP. `new/` now stages nothing at all — every name in it is a +/// *creating* lookup, so any tool that stats what a readdir reported (`ls -l`, +/// `find`, tab completion) would make one pane per entry — which is why there +/// is no longer a sentinel `d_ino` for an unresolved name on the wire. +/// +/// The core stages entries starting at index `req.off` (the cookie the kernel +/// echoed back) in a stable order. This encoder assigns cookie `off = req.off + +/// n + 1` to the nth entry it emits, and may emit only a *prefix* of what was +/// staged when the kernel's requested `size` runs out — the remainder comes +/// back as another readdir at the higher cookie, so staging has to be +/// idempotent per cookie rather than a stream. Zero staged bytes means EOF; it +/// is not an error, and the kernel stops asking. +/// +/// No `.` or `..`: the kernel synthesises neither and needs neither, and a +/// filesystem that emits them has to answer `LOOKUP("..")` too. +pub const dirent_stage_prefix = 10; + +/// Encode staged entries into kernel `fuse_dirent` records. Returns the bytes +/// written to `out`. Pure: this is where the alignment and cookie rules live, +/// and it is tested directly. +fn encodeDirents(out: []u8, staged: []const u8, cookie: u64) usize { + var in: usize = 0; + var w: usize = 0; + var n: u64 = 0; + while (in + dirent_stage_prefix <= staged.len) { + const node = std.mem.readInt(u64, staged[in..][0..8], .little); + const kind = staged[in + 8]; + const namelen: usize = staged[in + 9]; + // A zero name length would make the record self-referential (the + // kernel would parse the padding as the next entry), and a truncated + // record means the core staged something we cannot read. Stop rather + // than guess: a short reply is a legal readdir, a malformed one is not. + if (namelen == 0 or in + dirent_stage_prefix + namelen > staged.len) break; + const name = staged[in + dirent_stage_prefix ..][0..namelen]; + const record = recAlign(dirent_name_offset + namelen); + if (w + record > out.len) break; + + // Written field by field rather than through a struct pointer: `out` + // is a caller's slice of unknown alignment, and one @alignCast that is + // wrong here is a misaligned store into a kernel-bound buffer. + std.mem.writeInt(u64, out[w..][0..8], node, .little); + std.mem.writeInt(u64, out[w + 8 ..][0..8], cookie + n + 1, .little); + std.mem.writeInt(u32, out[w + 16 ..][0..4], @intCast(namelen), .little); + std.mem.writeInt(u32, out[w + 20 ..][0..4], if (kind == 1) DT_DIR else DT_REG, .little); + @memcpy(out[w + dirent_name_offset ..][0..namelen], name); + // The kernel never shows the padding to anyone, but zeroing it keeps + // the wire deterministic, which is what the encoder test asserts on. + @memset(out[w + dirent_name_offset + namelen ..][0 .. record - dirent_name_offset - namelen], 0); + + in += dirent_stage_prefix + namelen; + w += record; + n += 1; + } + return w; +} + +// --------------------------------------------------------------------------- +// fusermount3 +// --------------------------------------------------------------------------- + +/// The environment variable `fusermount3` reads to find the socket it must send +/// the `/dev/fuse` descriptor back over. Spelled with the leading underscore in +/// libfuse (`FUSE_COMMFD_ENV`); it is a private contract between the two +/// programs, not a user knob. +const commfd_env = "_FUSE_COMMFD"; + +/// Where the helper might be. Arch puts it in /usr/bin with /usr/sbin a symlink +/// to it, Debian derivatives use /usr/bin, and a machine with only libfuse2 +/// installed spells it without the 3 — that binary speaks the same +/// socketpair/SCM_RIGHTS protocol, so it is a real fallback and not a guess. +/// Searched by absolute path rather than through PATH because the thing being +/// executed is setuid root: PATH is attacker-influenced input. +const fusermount_paths = [_][:0]const u8{ + "/usr/bin/fusermount3", + "/usr/sbin/fusermount3", + "/bin/fusermount3", + "/sbin/fusermount3", + "/usr/local/bin/fusermount3", + "/usr/bin/fusermount", + "/usr/sbin/fusermount", + "/bin/fusermount", +}; + +/// The `-o` string. Every option here is a deliberate refusal: +/// +/// - `fsname`/`subtype` are cosmetic but load bearing: they are what `mount`, +/// `df` and `/proc/self/mountinfo` show, and an unnamed fuse mount in a bug +/// report is indistinguishable from anyone else's. +/// - `nosuid,nodev` are what fusermount3 forces anyway; naming them keeps the +/// intent in the source rather than in someone else's default. +/// - NOT `allow_other`: it needs `user_allow_other` in /etc/fuse.conf, which +/// is commented out on a stock Arch install, and asking for it makes +/// fusermount3 fail the whole mount instead of ignoring the option. It +/// would also be wrong — this filesystem executes text on write. +/// - NOT `default_permissions`: with it the kernel enforces the mode bits we +/// report, which sounds like a free wall but moves access control from the +/// core (which knows that `cons` is write-only) into a mode field, so a +/// wrong nibble in a table becomes an EACCES nobody can explain. Same +/// reason INIT negotiates no flags: fewer kernel behaviours to honour. +fn mountOpts(buf: *[128:0]u8) [:0]const u8 { + return std.fmt.bufPrintSentinel(buf, "fsname=pardes,subtype=pardes,nosuid,nodev", .{}, 0) catch unreachable; +} + +/// `_FUSE_COMMFD=<n>`, the child's end of the socketpair by number. libfuse +/// passes the descriptor this way rather than on the command line because +/// fusermount3 is setuid: its argv is world readable through /proc, its +/// environment is not. +fn commfdEnv(buf: *[32:0]u8, fd: c_int) [:0]const u8 { + return std.fmt.bufPrintSentinel(buf, commfd_env ++ "={d}", .{fd}, 0) catch unreachable; +} + +/// `fusermount3 -o <opts> -- <mountpoint>`. The `--` is not optional: a +/// mountpoint that begins with a dash would otherwise be parsed as a flag by a +/// setuid program. +fn mountArgv( + argv: *[6:null]?[*:0]const u8, + prog: [*:0]const u8, + opts: [*:0]const u8, + mountpoint: [*:0]const u8, +) void { + argv.* = .{ prog, "-o", opts, "--", mountpoint, null }; +} + +/// `fusermount3 -u -q -z -- <mountpoint>`. Lazy (`-z`) because the mount may +/// still have an open descriptor on it — a pane shell that inherited a cwd +/// inside the mount, say — and a non-lazy unmount would fail with EBUSY and +/// leave the mount behind for good. Quiet (`-q`) because the common case at +/// exit is a mount the kernel already tore down, and its complaint would be the +/// last thing on the user's terminal. +fn unmountArgv(argv: *[7:null]?[*:0]const u8, prog: [*:0]const u8, mountpoint: [*:0]const u8) void { + argv.* = .{ prog, "-u", "-q", "-z", "--", mountpoint, null }; +} + +/// CMSG_ALIGN/CMSG_LEN/CMSG_SPACE. Only ever evaluated on the Linux path, +/// where the alignment is `sizeof(size_t)`; other platforms align control +/// messages to 4 and would need their own numbers. +fn cmsgAlign(n: usize) usize { + const a: usize = @alignOf(usize); + return (n + a - 1) & ~(a - 1); +} +fn cmsgLen(n: usize) usize { + return cmsgAlign(@sizeOf(libc.cmsghdr)) + n; +} +fn cmsgSpace(n: usize) usize { + return cmsgAlign(@sizeOf(libc.cmsghdr)) + cmsgAlign(n); +} + +/// Build the child's environment: ours, plus `_FUSE_COMMFD`, minus any +/// `_FUSE_COMMFD` we inherited. The subtraction matters — `getenv` returns the +/// *first* match, so an inherited stale entry (pardes launched from inside +/// something that mounts) would win over the one we just appended and +/// fusermount3 would send the descriptor to a closed socket. +fn buildEnv(gpa: std.mem.Allocator, commfd: [:0]const u8) ![]?[*:0]const u8 { + var count: usize = 0; + while (libc.environ[count] != null) count += 1; + const env = try gpa.alloc(?[*:0]const u8, count + 2); + var n: usize = 0; + for (0..count) |i| { + const entry = libc.environ[i].?; + if (std.mem.startsWith(u8, std.mem.span(entry), commfd_env ++ "=")) continue; + env[n] = entry; + n += 1; + } + env[n] = commfd.ptr; + env[n + 1] = null; + return env[0 .. n + 2]; +} + +/// Resolve the helper once, by absolute path. Doing it in the parent rather +/// than by chaining execve attempts in the child keeps `argv[0]` honest (it is +/// what `ps` and fusermount3's own diagnostics print) and turns "fuse3 is not +/// installed" into its own error instead of an exit status. +fn findFusermount() ?[:0]const u8 { + for (fusermount_paths) |candidate| { + if (libc.access(candidate.ptr, libc.X_OK) == 0) return candidate; + } + return null; +} + +/// fork + execve the helper and wait for it. Not `std.process.Child`: that has +/// no way to hand a child an arbitrary descriptor, and the entire protocol here +/// is "the child writes to descriptor N". Everything the child does before +/// execve is async-signal-safe (close, execve, _exit) because the parent may +/// well be multithreaded by the time this runs. +fn spawnHelper( + prog: [*:0]const u8, + argv: [*:null]const ?[*:0]const u8, + envp: [*:null]const ?[*:0]const u8, + close_in_child: c_int, +) !u8 { + const pid = libc.fork(); + if (pid < 0) return error.ForkFailed; + if (pid == 0) { + // The parent's end of the socketpair. Left open, the parent's recvmsg + // could never see EOF when the helper dies without sending anything, + // and a refused mount would hang instead of failing. + if (close_in_child >= 0) _ = libc.close(close_in_child); + _ = libc.execve(prog, argv, envp); + // 127 is the shell's convention for "not found". Reachable only when + // the binary vanished between the access(2) above and now. + libc._exit(127); + } + var status: c_int = 0; + while (true) { + const got = libc.waitpid(pid, &status, 0); + if (got == pid) break; + if (got < 0 and libc.errno(got) == .INTR) continue; + // Reaped by somebody else's SIGCHLD handler: the status is gone, and + // the descriptor either arrived or it did not. Claim success and let + // the recvmsg be the judge. + return 0; + } + // WIFEXITED/WEXITSTATUS spelled out: std has no portable macro, and a + // helper killed by a signal is not a helper that refused the mount. + if (status & 0x7f != 0) return error.FusermountKilled; + return @intCast((status >> 8) & 0xff); +} + +/// Receive the `/dev/fuse` descriptor. fusermount3 sends it as an SCM_RIGHTS +/// control message alongside exactly one byte of ordinary data, and the byte is +/// not padding: a control message with no data attached may be dropped, so both +/// sides are required to send at least one. +/// +/// `MSG_CMSG_CLOEXEC` is the important flag. Every pane shell is forked from +/// this process and inherits open descriptors; a bash holding a copy of this +/// one keeps the FUSE connection alive after pardes exits, and the mount stays +/// up, unkillable, answering nothing, until that shell dies. +fn receiveFd(sock: c_int) !c_int { + var byte: [1]u8 = undefined; + var iov = [1]std.posix.iovec{.{ .base = &byte, .len = 1 }}; + var control: [cmsgSpace(@sizeOf(c_int))]u8 align(@alignOf(libc.cmsghdr)) = undefined; + while (true) { + var msg: libc.msghdr = .{ + .name = null, + .namelen = 0, + .iov = &iov, + .iovlen = 1, + .control = &control, + .controllen = @intCast(control.len), + .flags = 0, + }; + const n = libc.recvmsg(sock, &msg, linux.MSG.CMSG_CLOEXEC); + if (n < 0) { + if (libc.errno(n) == .INTR) continue; + return error.CommSocketFailed; + } + // EOF: the helper exited without sending anything, which is what a + // refused mount looks like from here. + if (n == 0) return error.FusermountRefused; + if (@as(usize, @intCast(msg.controllen)) < cmsgLen(@sizeOf(c_int))) return error.NoDescriptor; + const cmsg: *const libc.cmsghdr = @ptrCast(&control); + if (cmsg.level != libc.SOL.SOCKET or cmsg.type != libc.SCM.RIGHTS) return error.NoDescriptor; + if (@as(usize, @intCast(cmsg.len)) < cmsgLen(@sizeOf(c_int))) return error.NoDescriptor; + var fd: c_int = -1; + @memcpy( + std.mem.asBytes(&fd), + control[cmsgAlign(@sizeOf(libc.cmsghdr))..][0..@sizeOf(c_int)], + ); + if (fd < 0) return error.NoDescriptor; + return fd; + } +} + +/// `mkdir -p` for the mount point, 0700. The leaf is this process's own pid +/// directory and the parent is `.../pardes`, which on a fresh machine does not +/// exist; without the -p the whole feature would switch itself off in silence +/// on exactly the machines that never used it before. Same shape as +/// `nested.zig`'s ensureSocketDir, and 0700 for the same reason: what lives +/// under here takes commands. +fn ensureDir(path: [:0]const u8) void { + var partial: [4096:0]u8 = undefined; + if (path.len >= partial.len) return; + @memcpy(partial[0 .. path.len + 1], path[0 .. path.len + 1]); + for (1..path.len) |i| { + if (path[i] != '/') continue; + partial[i] = 0; + _ = libc.mkdir(partial[0..i :0], 0o700); + partial[i] = '/'; + } + _ = libc.mkdir(path, 0o700); +} + +/// Unmount and remove `<dir>/<pid>` for every pid that is gone. A pardes killed +/// with SIGKILL runs no defer, so its mount outlives it as an ENOTCONN stump +/// that `ls` reports as a permission error and that nothing else will ever +/// clean up — the snapshot suite alone would leave one per aborted run. +/// Bounded: one readdir of a directory only we write to, one kill(0) each. +/// Mirrors nested.zig's socket sweep deliberately, including the ESRCH rule: +/// 0 means alive, EPERM means alive and someone else's, only ESRCH is a corpse. +pub fn sweepStale(dir: []const u8) void { + if (comptime !supported) return; + var dir_buf: [4096:0]u8 = undefined; + const dir_z = std.fmt.bufPrintSentinel(&dir_buf, "{s}", .{dir}, 0) catch return; + const d = libc.opendir(dir_z) orelse return; + defer _ = libc.closedir(d); + const me = libc.getpid(); + while (libc.readdir(d)) |ent| { + const name = std.mem.sliceTo(&ent.name, 0); + // Strictly digits: parseInt would accept `+7` and `-7`, and this + // function unmounts and removes whatever it answers about. + if (name.len == 0) continue; + for (name) |ch| if (!std.ascii.isDigit(ch)) break; + if (std.mem.indexOfNone(u8, name, "0123456789") != null) continue; + const pid = std.fmt.parseInt(libc.pid_t, name, 10) catch continue; + if (pid == me) continue; + const rc = libc.kill(pid, @enumFromInt(0)); + if (rc == 0 or libc.errno(rc) != .SRCH) continue; + var path_buf: [4096:0]u8 = undefined; + const path = std.fmt.bufPrintSentinel(&path_buf, "{s}/{s}", .{ dir, name }, 0) catch continue; + // Always ours to remove: the name is a pid under a directory only + // pardes writes to, and taking the stump away is the point of a sweep. + unmountPath(path, true); + } +} + +/// Run the helper's unmount and, when the directory is ours, take it away. +/// Best effort in both halves: an already-unmounted point makes fusermount3 +/// complain (which -q swallows) and a non-empty one makes rmdir fail, and +/// neither is worth a diagnostic at exit. +/// +/// `remove_dir` is not a convenience. The *unmount* is always right — the mount +/// is ours whoever made the directory — but the *rmdir* is only right for a +/// point pardes derived itself (`<parent>/<pid>`, which `ensureDir` created). +/// A `--fs=<dir>` the user named is theirs, and removing it is the same +/// overreach `sweepStale` is already refused under an explicit `--fs` for. +fn unmountPath(path: [:0]const u8, remove_dir: bool) void { + if (findFusermount()) |prog| { + var argv: [7:null]?[*:0]const u8 = undefined; + unmountArgv(&argv, prog.ptr, path.ptr); + // A minimal environment: the helper wants nothing of ours, and the one + // variable that WOULD change its behaviour is the comm descriptor it + // must not find here. + const envp = [_:null]?[*:0]const u8{null}; + _ = spawnHelper(prog.ptr, &argv, &envp, -1) catch {}; + } + if (remove_dir) _ = libc.rmdir(path); +} + +// --------------------------------------------------------------------------- +// the park table +// --------------------------------------------------------------------------- + +/// How many kernel requests may be outstanding at once. Every slot is either in +/// flight (handed to the core, not yet answered) or parked (the core said +/// `.again`). In-flight slots are transient — the host answers each request +/// inside the same drain step — so in practice this counts BLOCKED READERS: one +/// slot per process sitting on `event` or `log`. A session with 32 of those has +/// 32 scripts watching it. +/// +/// Overflow is a refusal, not a queue: `take` answers EAGAIN and the descriptor +/// keeps being read. See its comment for why the tempting alternative (stop +/// reading and let the kernel hold the surplus) is a deadlock. +const max_slots = 32; + +/// Bytes of request payload a slot can own. A parked request's `data` cannot go +/// on borrowing the read buffer (the next `next()` overwrites it), so it is +/// copied in at parse time when it fits. This covers every payload that can +/// realistically block: a LOOKUP name is at most 255 bytes and a ctl verb line +/// or an event write-back is a few dozen. A WRITE larger than this is left +/// borrowed and answered EAGAIN if the core ever tries to park it — a write is +/// a transaction in this design and is not supposed to block, and growing this +/// table by 64 KiB a slot to make an impossible case zero-copy is the wrong +/// trade. +const park_data_max = 512; + +const Slot = struct { + used: bool = false, + /// The core answered `.again`; `retry()` will offer it back. + parked: bool = false, + /// Already offered in this retry round. Reset when a round finds nothing, + /// which is what gives every parked request exactly one attempt per frame + /// instead of letting the oldest one starve the rest. + retried: bool = false, + /// `req.data` points into `data` below rather than into the read buffer. + copied: bool = false, + /// Arrival order, so retries are FIFO: the reader that blocked first is + /// offered first. + seq: u64 = 0, + op: Opcode = @enumFromInt(0), + req: acmefs.Req = undefined, + data: [park_data_max]u8 = undefined, +}; + +// --------------------------------------------------------------------------- +// Fs +// --------------------------------------------------------------------------- + +pub const Fs = struct { + pub const Options = struct { + /// Absolute path of the mount point. Absolute because it is handed to a + /// setuid program that resolves it against its own cwd, and because the + /// unmount at exit must name the same place after any chdir. + mount: []const u8, + /// The largest WRITE payload the kernel may send in one request, and + /// therefore the size of the read buffer. 64 KiB matches what a `cp` + /// into `body` will use; smaller only splits the same bytes into more + /// round trips. + max_write: u32 = 64 * 1024, + /// Whether pardes made this directory and may therefore remove it at + /// exit. True for the derived `<parent>/<pid>`, false for a + /// `--fs=<dir>` the user named. See `unmountPath`. + owns_dir: bool = false, + }; + + gpa: std.mem.Allocator, + /// The `/dev/fuse` descriptor. -1 once torn down; every entry point checks + /// it, so a double deinit and a post-unmount drain are both no-ops. + fd: c_int = -1, + /// Set when the connection is gone (ENODEV/ECONNABORTED, or DESTROY). + /// `next()` stops reading; replies are still written because a slot may be + /// mid-flight and the write simply fails. + dead: bool = false, + path: [:0]u8, + /// Mirrors `Options.owns_dir`; gates the rmdir in `deinit`. + owns_dir: bool = false, + /// One request per read(2), so this is sized for the largest request that + /// exists: header + fuse_write_in + max_write. Below FUSE_MIN_READ_BUFFER + /// the kernel refuses to hand over requests at all and answers the client + /// EIO. 8-aligned so the parse can point structs at it. + buf: []align(8) u8, + /// Encoded `fuse_dirent`s. Separate from `buf` because a readdir reply is + /// built while its request is still being read from `buf`. + dirents: [8192]u8 align(8) = undefined, + + uid: u32, + gid: u32, + max_write: u32, + /// The minor the kernel offered, echoed back at INIT. Kept for the record: + /// it is the one number in this file that a future feature would consult. + minor: u32 = 0, + + slots: [max_slots]Slot = @splat(.{}), + seq: u64 = 0, + thread: ?std.Thread = null, + /// main -> poller, an `eventfd(2)`. The main thread adds 1 per completed + /// drain and the poller's blocking read takes the whole counter in one go, + /// which is the "collapse the acknowledgements that piled up while we were + /// not waiting" behaviour a pipe needed three functions and a nonblocking + /// toggle to fake. Not a condition variable, because the poller is blocked + /// in `poll()` most of the time and an fd is the only thing that both + /// `poll()` and a blocking read can wait on — which is what lets shutdown + /// break it out of either state. + /// + /// The counter cannot say "stop": a stop and a drain acknowledgement that + /// race are summed into one indistinguishable number. `stopping` is the + /// sticky half of the signal, and is re-read after every wake; the eventfd + /// only ever means "look again". The store/write and read/load pair is a + /// release/acquire edge over the eventfd's own wait-queue lock, so a poller + /// that observes the increment observes the flag with it. + ctl: c_int = -1, + stopping: std.atomic.Value(bool) = .init(false), + wake_ctx: ?*anyopaque = null, + wake_fn: ?*const fn (?*anyopaque) void = null, + + /// Mount, hand out the descriptor, and complete the INIT handshake. On + /// return the filesystem is live: the kernel will start sending lookups the + /// moment anything touches the directory. + pub fn mount(gpa: std.mem.Allocator, opts: Options) !*Fs { + if (comptime !supported) return error.Unsupported; + if (opts.mount.len == 0 or opts.mount[0] != '/') return error.MountPathNotAbsolute; + + const path = try gpa.dupeZ(u8, opts.mount); + errdefer gpa.free(path); + ensureDir(path); + + const buf_len = @max( + min_read_buffer, + @sizeOf(fuse_in_header) + @sizeOf(fuse_write_in) + @as(usize, opts.max_write), + ); + const buf = try gpa.alignedAlloc(u8, .@"8", buf_len); + errdefer gpa.free(buf); + + const fd = try mountFusermount(gpa, path); + errdefer _ = libc.close(fd); + + const fs = try gpa.create(Fs); + errdefer gpa.destroy(fs); + fs.* = .{ + .gpa = gpa, + .fd = fd, + .path = path, + .buf = buf, + .uid = libc.getuid(), + .gid = libc.getgid(), + .max_write = opts.max_write, + .owns_dir = opts.owns_dir, + }; + // Still blocking here on purpose: INIT is already queued (fusermount3 + // completed mount(2) before it sent us the descriptor), and a + // non-blocking read would make the handshake a spin loop. + try fs.handshake(); + try fs.setNonblocking(); + return fs; + } + + /// socketpair, fork the setuid helper, take the descriptor it sends back. + fn mountFusermount(gpa: std.mem.Allocator, path: [:0]const u8) !c_int { + const prog = findFusermount() orelse return error.FusermountMissing; + var sv: [2]c_int = undefined; + if (libc.socketpair(libc.AF.UNIX, libc.SOCK.STREAM, 0, &sv) != 0) return error.SocketPairFailed; + // Both ends close-on-exec first, then the child's end is un-marked just + // before the fork. The window in between is what any *other* thread's + // fork would inherit, and pane shells are forked with forkpty and + // inherit everything open. + setCloexec(sv[0]); + setCloexec(sv[1]); + errdefer _ = libc.close(sv[0]); + + var opts_buf: [128:0]u8 = undefined; + var commfd_buf: [32:0]u8 = undefined; + const opts = mountOpts(&opts_buf); + const commfd = commfdEnv(&commfd_buf, sv[1]); + + const envp = try buildEnv(gpa, commfd); + defer gpa.free(envp); + var argv: [6:null]?[*:0]const u8 = undefined; + mountArgv(&argv, prog.ptr, opts.ptr, path.ptr); + + clearCloexec(sv[1]); + const code = spawnHelper(prog.ptr, &argv, @ptrCast(envp.ptr), sv[0]) catch |err| { + _ = libc.close(sv[1]); + return err; + }; + // Ours to close either way: the child has its own copy, and while we + // hold one the recvmsg below can never see EOF when the helper dies. + _ = libc.close(sv[1]); + if (code == 127) return error.FusermountMissing; + + const fd = try receiveFd(sv[0]); + if (code != 0) { + _ = libc.close(fd); + return error.FusermountFailed; + } + _ = libc.close(sv[0]); + return fd; + } + + /// Read the kernel's INIT and answer it. Negotiating nothing is the design: + /// every flag is a kernel behaviour we would then have to honour forever, + /// and this filesystem wants none of them — no readdirplus (whose ENOSYS + /// has no fallback and would fail every getdents), no atomic O_TRUNC (so + /// `> file` arrives as a plain SETATTR the core already handles), no POSIX + /// or BSD locks (flags = 0 makes the kernel set `no_lock`/`no_flock` and + /// answer them itself). + fn handshake(fs: *Fs) !void { + const n = readFull(fs.fd, fs.buf); + if (n < @sizeOf(fuse_in_header) + @sizeOf(fuse_init_in)) return error.InitFailed; + const h: *const fuse_in_header = @ptrCast(fs.buf.ptr); + if (@as(Opcode, @enumFromInt(h.opcode)) != .init) return error.InitFailed; + const in: *const fuse_init_in = @ptrCast(@as([*]align(8) u8, @alignCast(fs.buf.ptr + @sizeOf(fuse_in_header)))); + // A major mismatch is fatal and there is nothing to negotiate: the + // kernel aborts the connection, and answering anyway just delays the + // failure to the first syscall through the mount. + if (in.major != kernel_version) return error.InitVersion; + fs.minor = in.minor; + + const out: fuse_init_out = .{ + .major = kernel_version, + // Capped, not echoed: see `kernel_minor`. + .minor = @min(in.minor, kernel_minor), + // Zero, not "some readahead": with FOPEN_DIRECT_IO there is no page + // cache to read ahead into, and a nonzero value here only invites + // the kernel to ask for bytes nobody wanted. + .max_readahead = 0, + .flags = 0, + // Left at zero so the kernel keeps its own defaults; a nonzero + // max_background is the one that silently caps concurrency. + .max_background = 0, + .congestion_threshold = 0, + .max_write = fs.max_write, + // 1 ns. Timestamps on this filesystem are all zero anyway, but a + // time_gran of 0 is not a legal granularity. + .time_gran = 1, + .max_pages = 0, + .map_alignment = 0, + .flags2 = 0, + .max_stack_depth = 0, + // 0 = no server timeout. A timeout would let the kernel abort the + // connection while a legitimately parked `event` read waits. + .request_timeout = 0, + .unused = @splat(0), + }; + fs.answer(h.unique, std.mem.asBytes(&out), &.{}); + return; + } + + fn setNonblocking(fs: *Fs) !void { + const flags = libc.fcntl(fs.fd, libc.F.GETFL, @as(c_int, 0)); + if (flags < 0) return error.FcntlFailed; + var o: libc.O = @bitCast(@as(u32, @bitCast(flags))); + o.NONBLOCK = true; + if (libc.fcntl(fs.fd, libc.F.SETFL, @as(c_int, @bitCast(@as(u32, @bitCast(o))))) < 0) + return error.FcntlFailed; + } + + /// Answer everything still held, abort the connection, unmount, remove the + /// directory. The order is not interchangeable: + /// + /// 1. reply -ENODEV to every slot, so a reader blocked on `event` gets an + /// error rather than being left in uninterruptible sleep. + /// 2. close the descriptor, which aborts the connection — the backstop + /// for anything that raced step 1, since the kernel then fails every + /// pending request itself. + /// 3. only then unmount, because a mount whose server is gone is exactly + /// what `fusermount3 -u -z` is for. + /// 4. remove the directory, but only when pardes made it: the derived + /// `<parent>/<pid>` is ours, a `--fs=<dir>` the user named is not. + pub fn deinit(fs: *Fs) void { + const gpa = fs.gpa; + fs.stopThread(); + if (fs.fd >= 0) { + for (&fs.slots) |*s| { + if (!s.used) continue; + fs.answerErr(s.req.tag, .NODEV); + s.* = .{}; + } + _ = libc.close(fs.fd); + fs.fd = -1; + } + if (comptime supported) unmountPath(fs.path, fs.owns_dir); + gpa.free(fs.path); + gpa.free(fs.buf); + gpa.destroy(fs); + } + + // -- request pump ------------------------------------------------------- + + /// Parse the next pending kernel request, or null when the descriptor is + /// drained. Call in a loop until null; the loop is the batch, and one wake + /// serves all of it. + /// + /// The returned `Req.data` borrows storage owned by this `Fs` and is valid + /// until the next `next()` call. The core copies whatever it keeps — the + /// same rule as `.pty_read`. + /// + /// Requests the core has no business seeing are answered here and the loop + /// continues, so a caller never observes them. + /// + /// Running this to null is also what acknowledges the batch to the poll + /// thread, so a host that stops early keeps the poller waiting and loses + /// wake latency until the next frame. It is not a correctness bug — the + /// remaining requests simply wait in the kernel — but the loop is the + /// contract. + pub fn next(fs: *Fs) ?acmefs.Req { + if (comptime !supported) return null; + // Only the EAGAIN arm below releases the poller, and deliberately so. + // Every other null return from here implies `dead`, which is write-once + // and means reads on the descriptor are failing: posting would send the + // poller back into `poll()` on a still-open fd that reports POLLIN + // forever, wake the host, drain to this same null, and spin two threads + // at 100%. Parking the poller in `consume()` is the right resting state + // for a connection that can never produce work again; `stopThread` + // releases it. `fd < 0` is unreachable here, since only `deinit` sets it + // and it joins the poller first. + if (fs.fd < 0 or fs.dead) return null; + while (true) { + const n = libc.read(fs.fd, fs.buf.ptr, fs.buf.len); + if (n < 0) switch (libc.errno(n)) { + .INTR => continue, + .AGAIN => { + // Drained: release the poller (see `post`). + fs.post(); + return null; + }, + // The request was interrupted or aborted between being queued + // and being read; there is nothing to answer. + .NOENT => continue, + // ENODEV (connection aborted, or we were unmounted from under + // ourselves) and ECONNABORTED are terminal. Anything else here + // is not a thing /dev/fuse does, and treating the unknown as + // terminal beats a loop that reads -1 forever. + else => { + fs.dead = true; + return null; + }, + }; + if (n == 0) { + fs.dead = true; + return null; + } + const total: usize = @intCast(n); + // Cannot happen (the kernel writes whole requests) but the parse + // below indexes on it. + if (total < @sizeOf(fuse_in_header)) continue; + if (fs.dispatch(total)) |req| return req; + } + } + + /// Offer parked requests back, one per call. Call in a loop until null, + /// once per frame, before `next()`: the null both ends the round and resets + /// it, so every parked request gets exactly one attempt per frame and a + /// permanently blocked reader cannot starve the others. + pub fn retry(fs: *Fs) ?acmefs.Req { + if (comptime !supported) return null; + if (fs.fd < 0) return null; + var best: ?usize = null; + for (&fs.slots, 0..) |*s, i| { + if (!s.used or !s.parked or s.retried) continue; + if (best == null or s.seq < fs.slots[best.?].seq) best = i; + } + const i = best orelse { + for (&fs.slots) |*s| s.retried = false; + return null; + }; + fs.slots[i].retried = true; + // In flight again: `reply()` re-parks it if the core still has nothing. + fs.slots[i].parked = false; + return fs.slots[i].req; + } + + /// Write the core's answer, or park the request when it said `.again`. + /// Called from the `.fs_reply` effect; `bytes` is the payload resolved by + /// `pardes.fsPayload` and is borrowed only for the duration of this call. + pub fn reply(fs: *Fs, r: *const acmefs.Reply, bytes: []const u8) void { + if (comptime !supported) return; + const i = fs.findSlot(r.tag) orelse return; // interrupted, or torn down + const s = &fs.slots[i]; + + if (r.status == .again) { + // The one case a park is refused: a payload too large to have been + // copied at parse time still borrows the read buffer, so parking it + // would park a dangling slice. EAGAIN is honest — the writer can + // retry — and by construction unreachable, since the core answers + // writes as transactions and only reads ever block. + if (!s.copied and s.req.data.len != 0) { + fs.answerErr(s.req.tag, .AGAIN); + fs.release(i); + return; + } + s.parked = true; + return; + } + + if (r.status == .err) { + fs.answerErr(s.req.tag, @enumFromInt(if (r.errno == 0) @intFromEnum(libc.E.IO) else r.errno)); + fs.release(i); + return; + } + + switch (s.req.op) { + .lookup => { + const out: fuse_entry_out = .{ + .nodeid = r.attr.node, + // Node ids are never reused in this filesystem (pane + // serials are monotonic), which is exactly the condition + // for a constant generation to be safe. + .generation = 0, + // No caching, at all. Every file here changes under the + // reader's feet, and a cached negative lookup would make + // `new/<name>` (which CREATES a pane) work exactly once. + .entry_valid = 0, + .attr_valid = 0, + .entry_valid_nsec = 0, + .attr_valid_nsec = 0, + .attr = fs.attr(r.attr, r.attr.node), + }; + fs.answer(s.req.tag, std.mem.asBytes(&out), &.{}); + }, + .getattr, .setattr => { + const out: fuse_attr_out = .{ + .attr_valid = 0, + .attr_valid_nsec = 0, + .dummy = 0, + .attr = fs.attr(r.attr, s.req.node), + }; + fs.answer(s.req.tag, std.mem.asBytes(&out), &.{}); + }, + .open => { + const out: fuse_open_out = .{ + .fh = r.handle, + // Direct IO for files; nothing for directories, where the + // flag has no meaning and FOPEN_CACHE_DIR (which we do not + // set) is the caching knob. An uncached directory is the + // point: `new/` and the pane list change constantly. + .open_flags = if (s.op == .opendir) 0 else FOPEN_DIRECT_IO, + .backing_id = 0, + }; + fs.answer(s.req.tag, std.mem.asBytes(&out), &.{}); + }, + .read => { + // Never more than was asked for: a read reply longer than + // `size` is a protocol error the kernel answers with EIO. + const len = @min(bytes.len, s.req.size); + fs.answer(s.req.tag, &.{}, bytes[0..len]); + }, + .readdir => { + const room = @min(@as(usize, s.req.size), fs.dirents.len); + const len = encodeDirents(fs.dirents[0..room], bytes, s.req.off); + fs.answer(s.req.tag, &.{}, fs.dirents[0..len]); + }, + .write => { + // The core's own count, not the request size: `data` refusing a + // partial grapheme is a real short write, and claiming the + // whole request would tell the writer its trailing bytes + // landed when they did not. Clamped anyway, because a count + // larger than what was offered makes the kernel advance a file + // offset past bytes that never existed. + const out: fuse_write_out = .{ + .size = @min(r.written, s.req.size), + .padding = 0, + }; + fs.answer(s.req.tag, std.mem.asBytes(&out), &.{}); + }, + .release => fs.answer(s.req.tag, &.{}, &.{}), + .statfs => { + // Synthetic numbers, but not arbitrary ones: `namelen` is what + // pathconf(_PC_NAME_MAX) returns and a zero there makes some + // tools refuse to create any name at all, and `bsize` is what + // `stat` reports as the IO block size. + const out: fuse_statfs_out = .{ .st = .{ + .blocks = 0, + .bfree = 0, + .bavail = 0, + .files = 0, + .ffree = 0, + .bsize = 4096, + .namelen = 255, + .frsize = 4096, + .padding = 0, + .spare = @splat(0), + } }; + fs.answer(s.req.tag, std.mem.asBytes(&out), &.{}); + }, + } + fs.release(i); + } + + /// Translate one request. Null means it was answered here. + fn dispatch(fs: *Fs, total: usize) ?acmefs.Req { + const h: *const fuse_in_header = @ptrCast(fs.buf.ptr); + // Bounded by the header's own length, not just by what the read + // returned. They agree on /dev/fuse, and taking the smaller of the two + // is what keeps a WRITE from claiming payload it did not bring even if + // some future kernel ever pads a request. + const end = @min(total, @max(@as(usize, h.len), @sizeOf(fuse_in_header))); + const body: []align(8) const u8 = @alignCast(fs.buf[@sizeOf(fuse_in_header)..end]); + const op: Opcode = @enumFromInt(h.opcode); + switch (op) { + // Already answered in the handshake. A second INIT cannot happen; + // answering it again is cheaper than a special case that could. + .init => { + fs.answerErr(h.unique, .INVAL); + return null; + }, + // NEVER replied to. The kernel does not track these as pending + // requests, so a reply carries a `unique` it will not recognise — + // -ENOENT at best, and at worst a reply matched against a *live* + // request that happens to share the number. Ignoring the refcount + // itself is fine: this filesystem's node table is bounded by the + // pane count, so nothing grows. + .forget, .batch_forget => return null, + // Answer the ORIGINAL with EINTR and drop it. This is the only + // thing standing between a SIGKILLed reader of `event` and + // permanent uninterruptible sleep: after the fatal signal the + // kernel's last wait is not killable, so the process survives its + // own kill until this reply lands. No reply to the interrupt + // itself — its unique is `original | 1` and the kernel keeps no + // pending entry for it, while answering -ENOSYS would switch + // interrupts off for the whole connection and take the escape + // hatch away. + .interrupt => { + if (body.len >= @sizeOf(fuse_interrupt_in)) { + const in: *const fuse_interrupt_in = @ptrCast(body.ptr); + if (fs.findSlot(in.unique)) |i| { + fs.answerErr(fs.slots[i].req.tag, .INTR); + fs.release(i); + } + } + return null; + }, + // A missing reply here hangs `umount` outright. + .destroy => { + fs.answer(h.unique, &.{}, &.{}); + fs.dead = true; + return null; + }, + // -ENOSYS rather than an empty reply: the kernel sets `no_flush` + // and stops sending them, so this costs one round trip for the + // whole connection instead of one per close(2). Nothing here has + // buffered state for a flush to commit. + .flush => { + fs.answerErr(h.unique, .NOSYS); + return null; + }, + .lookup => { + // The name is the whole body, NUL terminated. An empty name is + // not a lookup of anything. + const name = std.mem.sliceTo(body, 0); + if (name.len == 0) { + fs.answerErr(h.unique, .INVAL); + return null; + } + return fs.take(op, .{ + .tag = h.unique, + .op = .lookup, + .node = h.nodeid, + .data = name, + }); + }, + .getattr => { + const in = fs.arg(fuse_getattr_in, body) orelse return null; + return fs.take(op, .{ + .tag = h.unique, + .op = .getattr, + .node = h.nodeid, + // `fh` is only meaningful with the flag; reading it blind + // hands the core a handle from an unrelated open. + .handle = if (in.getattr_flags & FUSE_GETATTR_FH != 0) @truncate(in.fh) else 0, + }); + }, + .setattr => { + const in = fs.arg(fuse_setattr_in, body) orelse return null; + return fs.take(op, .{ + .tag = h.unique, + .op = .setattr, + .node = h.nodeid, + .handle = @truncate(in.fh), + // The `> file` path, and the only setattr this filesystem + // has an opinion about. A truncate to a nonzero length is + // not expressible in the core's ABI and is reported as no + // truncate at all: the reply still carries the current + // attributes, so ftruncate(fd, n) succeeds and changes + // nothing, which is what every synthetic file here wants. + .truncate = in.valid & FATTR_SIZE != 0 and in.size == 0, + }); + }, + .open, .opendir => { + // The flags are read only to reject a short body: this + // filesystem's permission model is the mode bits each synthetic + // file reports from GETATTR, which the kernel enforces itself, + // so the access mode has nothing left to say here. + if (fs.arg(fuse_open_in, body) == null) return null; + return fs.take(op, .{ + .tag = h.unique, + .op = .open, + .node = h.nodeid, + }); + }, + .read, .readdir => { + const in = fs.arg(fuse_read_in, body) orelse return null; + return fs.take(op, .{ + .tag = h.unique, + .op = if (op == .readdir) .readdir else .read, + .node = h.nodeid, + .handle = @truncate(in.fh), + .off = in.offset, + .size = in.size, + }); + }, + .write => { + const in = fs.arg(fuse_write_in, body) orelse return null; + const payload = body[@sizeOf(fuse_write_in)..]; + // Trust the header's length over the struct's: a `size` larger + // than what arrived would read past the request. + const len = @min(@as(usize, in.size), payload.len); + return fs.take(op, .{ + .tag = h.unique, + .op = .write, + .node = h.nodeid, + .handle = @truncate(in.fh), + .off = in.offset, + .size = @intCast(len), + .data = payload[0..len], + }); + }, + .release, .releasedir => { + const in = fs.arg(fuse_release_in, body) orelse return null; + return fs.take(op, .{ + .tag = h.unique, + .op = .release, + .node = h.nodeid, + .handle = @truncate(in.fh), + }); + }, + .statfs => return fs.take(op, .{ + .tag = h.unique, + .op = .statfs, + .node = h.nodeid, + }), + // Everything else. -ENOSYS is not a shrug: for most of these the + // kernel caches the answer and stops asking (`no_access`, + // `no_getxattr`, `no_statx`, `no_poll`, `no_lseek`, `no_create`), + // so one refusal switches the whole feature off for the connection. + // The mutations (mkdir, unlink, rename, link, symlink) are refused + // because this tree is generated: its shape follows the pane list + // and there is nothing for a user to create or remove in it. + // READDIRPLUS is not in this list by accident — it is unreachable, + // because INIT never sets FUSE_DO_READDIRPLUS, and it has to stay + // that way: its -ENOSYS has NO fallback in the kernel and would + // fail every getdents through the mount. + else => { + fs.answerErr(h.unique, .NOSYS); + return null; + }, + } + } + + /// Point a request struct at the read buffer. Null (and an EINVAL reply) + /// when the kernel sent less than the struct, which cannot happen but would + /// otherwise be a read past the buffer. + fn arg(fs: *Fs, comptime T: type, body: []align(8) const u8) ?*const T { + if (body.len < @sizeOf(T)) { + const h: *const fuse_in_header = @ptrCast(fs.buf.ptr); + fs.answerErr(h.unique, .INVAL); + return null; + } + return @ptrCast(body.ptr); + } + + /// Move a parsed request into a slot and hand it to the caller. Small + /// payloads are copied in here so that a later park has stable bytes; a + /// large one stays borrowed (see `park_data_max`). + /// + /// Null (and an EAGAIN reply) when the table is full. That is the whole + /// reason `next` reads unconditionally instead of gating on a free slot: + /// gating looks like polite backpressure and is a deadlock. With 32 readers + /// blocked on `event`, refusing to read the descriptor means the INTERRUPT + /// that would free a slot is never read either, so a SIGKILLed reader stays + /// in uninterruptible sleep forever and every unrelated `ls` of the mount + /// hangs behind it. Reading and answering EAGAIN keeps FORGET, INTERRUPT, + /// DESTROY and the ENOSYS family flowing — none of which need a slot — and + /// turns "too many blocked readers" into one failed syscall the caller can + /// see and retry. + fn take(fs: *Fs, op: Opcode, req: acmefs.Req) ?acmefs.Req { + const i = fs.freeSlot() orelse { + fs.answerErr(req.tag, .AGAIN); + return null; + }; + const s = &fs.slots[i]; + s.* = .{ + .used = true, + .seq = fs.seq, + .op = op, + .req = req, + }; + fs.seq += 1; + if (req.data.len != 0 and req.data.len <= park_data_max) { + @memcpy(s.data[0..req.data.len], req.data); + s.copied = true; + s.req.data = s.data[0..req.data.len]; + } + return s.req; + } + + fn freeSlot(fs: *Fs) ?usize { + for (&fs.slots, 0..) |*s, i| if (!s.used) return i; + return null; + } + + fn findSlot(fs: *Fs, tag: u64) ?usize { + for (&fs.slots, 0..) |*s, i| if (s.used and s.req.tag == tag) return i; + return null; + } + + fn release(fs: *Fs, i: usize) void { + fs.slots[i] = .{}; + } + + /// `Reply.Attr` -> `fuse_attr`. `node` is the fallback inode for replies + /// that do not name one (a getattr answers about a node the request already + /// identified); a zero `st_ino` is a value no filesystem is allowed to + /// report and some tools treat it as a deleted entry. + fn attr(fs: *const Fs, a: acmefs.Reply.Attr, node: u64) fuse_attr { + const ino = if (a.node != 0) a.node else node; + return .{ + .ino = ino, + .size = a.size, + // 512-byte units, as `stat` wants them. Rounded up so a nonempty + // file never reports zero blocks, which `du` reads as a hole. + .blocks = (a.size + 511) / 512, + .atime = 0, + .mtime = 0, + .ctime = 0, + .atimensec = 0, + .mtimensec = 0, + .ctimensec = 0, + .mode = (if (a.dir) S_IFDIR else S_IFREG) | @as(u32, a.mode), + // 2 for a directory (itself and `.`) is what every tool expects; + // `find` in particular uses it to decide whether to recurse. + .nlink = if (a.dir) 2 else 1, + // The mounting user owns everything: without `allow_other` nobody + // else can reach the mount at all, and reporting some other owner + // would only make `ls -l` lie. + .uid = fs.uid, + .gid = fs.gid, + .rdev = 0, + .blksize = 4096, + .flags = 0, + }; + } + + // -- reply framing ------------------------------------------------------ + + /// One `writev` per reply: header, then the op's fixed out struct, then the + /// payload. Split into iovecs rather than assembled in a buffer so that a + /// megabyte read out of a pane's text is written straight from the core's + /// bytes — the whole point of `Reply.Payload.region`. + fn answer(fs: *Fs, unique: u64, fixed: []const u8, payload: []const u8) void { + var header: fuse_out_header = .{ + .len = @intCast(@sizeOf(fuse_out_header) + fixed.len + payload.len), + .@"error" = 0, + .unique = unique, + }; + var iov: [3]std.posix.iovec_const = undefined; + var n: usize = 1; + iov[0] = .{ .base = std.mem.asBytes(&header).ptr, .len = @sizeOf(fuse_out_header) }; + if (fixed.len != 0) { + iov[n] = .{ .base = fixed.ptr, .len = fixed.len }; + n += 1; + } + if (payload.len != 0) { + iov[n] = .{ .base = payload.ptr, .len = payload.len }; + n += 1; + } + fs.writeReply(iov[0..n], header.len); + } + + /// An error reply is header-only: the kernel checks `nbytes == + /// sizeof(oh)` when `error != 0` and answers -EINVAL otherwise, which + /// leaves the original request pending forever. + fn answerErr(fs: *Fs, unique: u64, e: libc.E) void { + var header: fuse_out_header = .{ + .len = @sizeOf(fuse_out_header), + .@"error" = -@as(i32, @intFromEnum(e)), + .unique = unique, + }; + const iov = [1]std.posix.iovec_const{ + .{ .base = std.mem.asBytes(&header).ptr, .len = @sizeOf(fuse_out_header) }, + }; + fs.writeReply(&iov, header.len); + } + + fn writeReply(fs: *Fs, iov: []const std.posix.iovec_const, expect: u32) void { + if (fs.fd < 0) return; + while (true) { + const n = libc.writev(fs.fd, iov.ptr, @intCast(iov.len)); + if (n < 0) switch (libc.errno(n)) { + .INTR => continue, + // /dev/fuse writes never block, so this is not the usual + // EAGAIN; retrying is the only thing that can make progress and + // it cannot loop forever because the kernel is not waiting on + // us. + .AGAIN => continue, + // The request is no longer pending: it was interrupted or the + // connection was aborted between the read and this write. + // Dropping it is correct — there is nothing left to answer. + .NOENT => return, + else => { + fs.dead = true; + return; + }, + }; + // A short write to /dev/fuse is not a thing (the kernel takes the + // whole reply or none of it), so this can only mean the reply was + // malformed and the request is still pending. Nothing useful is + // left to do about it here, and pretending otherwise would hide it. + std.debug.assert(@as(u32, @intCast(n)) == expect); + return; + } + } + + // -- poll thread -------------------------------------------------------- + + /// Start the one background thread: it waits for POLLIN and calls `wake`. + /// It never touches the descriptor's data, never sees a request and never + /// calls the core; the host's `wake` is expected to do nothing but post an + /// event on the loop, exactly like the inotify thread's. + /// + /// Optional by design. A host with no threads simply does not call this and + /// drains from its frame poll instead; it loses wake latency and nothing + /// else, which is what makes the no-parallelism backend work unchanged. + pub fn wakeThread(fs: *Fs, ctx: ?*anyopaque, wake: *const fn (?*anyopaque) void) !void { + if (comptime !supported) return; + if (fs.thread != null) return; + // Blocking on purpose: `consume` is a blocking read on this descriptor. + // The write side cannot block anyway — an eventfd write only waits for + // a counter one short of `maxInt(u64)` to be drained, which is not + // reachable at one increment per drain. + const efd = libc.eventfd(0, linux.EFD.CLOEXEC); + if (efd < 0) return error.EventFdFailed; + fs.ctl = efd; + fs.wake_ctx = ctx; + fs.wake_fn = wake; + fs.thread = std.Thread.spawn(.{}, pollLoop, .{fs}) catch |err| { + _ = libc.close(efd); + fs.ctl = -1; + return err; + }; + } + + fn stopThread(fs: *Fs) void { + if (comptime !supported) return; + // `ctl` and `thread` are set and cleared together, so there is no + // descriptor to close on the path where no poller was ever started. + const t = fs.thread orelse return; + // The flag before the wake, never after: a poller that reads the + // increment must not then find `stopping` false and go back to sleep on + // a counter nobody will raise again. With this order every state the + // poller can be in ends in an exit — the loop condition, the `poll()` + // (the eventfd becomes readable) and the blocking wait for a drain + // acknowledgement (the read returns) all re-read the flag. + fs.stopping.store(true, .release); + fs.post(); + t.join(); + fs.thread = null; + _ = libc.close(fs.ctl); + fs.ctl = -1; + } + + /// Raise the counter by one: "the descriptor has been drained, you may poll + /// again", or during teardown "look at `stopping`". Without the drain half + /// of that handshake the poller re-polls a level-triggered descriptor that + /// is still readable and spins a core until the main thread catches up; + /// with it, one wake serves one batch. + fn post(fs: *Fs) void { + if (fs.ctl < 0) return; + const one: u64 = 1; + _ = libc.write(fs.ctl, std.mem.asBytes(&one), @sizeOf(u64)); + } + + fn pollLoop(fs: *Fs) void { + if (comptime !supported) return; + while (!fs.stopping.load(.acquire)) { + var fds = [2]libc.pollfd{ + .{ .fd = fs.fd, .events = libc.POLL.IN, .revents = 0 }, + .{ .fd = fs.ctl, .events = libc.POLL.IN, .revents = 0 }, + }; + const rc = libc.poll(&fds, 2, -1); + if (rc < 0) { + if (libc.errno(rc) == .INTR) continue; + return; + } + // Shutdown, or an acknowledgement for a drain that happened without + // us. Take the whole counter and re-poll either way: a leftover + // count would make the wait below return instantly and turn the + // next wake into a spin. + if (fds[1].revents != 0 and fs.consume()) return; + if (fds[0].revents & (libc.POLL.ERR | libc.POLL.HUP | libc.POLL.NVAL) != 0) return; + if (fds[0].revents & libc.POLL.IN == 0) continue; + + (fs.wake_fn.?)(fs.wake_ctx); + // Wait for the main thread to finish the batch. This is the whole + // anti-spin mechanism; see `post`. + if (fs.consume()) return; + } + } + + /// Block until the counter is nonzero, then take all of it. True when the + /// poller must exit, which is `stopping` and nothing else: the count itself + /// carries no meaning beyond "look again". + /// + /// The read blocks, including on the branch that reached here from a + /// `poll()` that only *said* the descriptor was readable. That is safe + /// because `stopThread` closes `ctl` after `join()` and never before: an + /// eventfd raises neither POLLERR nor POLLHUP, so the one revents value + /// that would be readable-but-not-readable is POLLNVAL, and a closed + /// descriptor is the only thing that produces it. + fn consume(fs: *Fs) bool { + var v: u64 = undefined; + while (true) { + const n = libc.read(fs.ctl, std.mem.asBytes(&v), @sizeOf(u64)); + // A short read and an EOF do not exist on an eventfd: the read + // returns 8 or -1. So anything but EINTR means this descriptor is + // not the one we opened, and exiting beats spinning on it. + if (n < 0) { + if (libc.errno(n) == .INTR) continue; + return true; + } + return fs.stopping.load(.acquire); + } + } +}; + +// --------------------------------------------------------------------------- +// descriptor flags +// --------------------------------------------------------------------------- + +fn setCloexec(fd: c_int) void { + const FD_CLOEXEC: c_int = 1; + _ = libc.fcntl(fd, libc.F.SETFD, FD_CLOEXEC); +} + +/// The child of the mount fork must KEEP this descriptor across execve — it is +/// the whole channel the setuid helper answers on. +fn clearCloexec(fd: c_int) void { + _ = libc.fcntl(fd, libc.F.SETFD, @as(c_int, 0)); +} + +fn setNonblock(fd: c_int) void { + const flags = libc.fcntl(fd, libc.F.GETFL, @as(c_int, 0)); + if (flags < 0) return; + var o: libc.O = @bitCast(@as(u32, @bitCast(flags))); + o.NONBLOCK = true; + _ = libc.fcntl(fd, libc.F.SETFL, @as(c_int, @bitCast(@as(u32, @bitCast(o))))); +} + +/// One blocking read, EINTR-safe. Used only for the INIT handshake, where the +/// descriptor is still blocking; every later read goes through `next()`. +fn readFull(fd: c_int, buf: []u8) usize { + while (true) { + const n = libc.read(fd, buf.ptr, buf.len); + if (n < 0) { + if (libc.errno(n) == .INTR) continue; + return 0; + } + return @intCast(n); + } +} + +// --------------------------------------------------------------------------- +// tests +// --------------------------------------------------------------------------- +// +// No test here mounts anything: a real mount needs the setuid helper, a +// writable runtime directory and a kernel that will let go of it again, which +// is a snapshot test's job and not a unit test's. What is testable without a +// mount is everything that has ever actually been wrong in a FUSE server — +// struct sizes, dirent alignment, cookies, the INIT reply, the park table, and +// the argv handed to a setuid program. Those are what follows, driven through a +// socketpair standing in for /dev/fuse. + +const testing = std.testing; + +/// Build an `Fs` with no mount, wired to `fd`. The socketpair replaces +/// /dev/fuse for the codec tests: the kernel's side of the conversation is +/// written by hand and the reply is read back and compared byte for byte. +fn testFs(gpa: std.mem.Allocator, fd: c_int) !*Fs { + const fs = try gpa.create(Fs); + fs.* = .{ + .gpa = gpa, + .fd = fd, + .path = try gpa.dupeZ(u8, "/nonexistent"), + .buf = try gpa.alignedAlloc(u8, .@"8", min_read_buffer), + .uid = 1000, + .gid = 1000, + .max_write = 4096, + }; + return fs; +} + +fn testFsFree(fs: *Fs) void { + const gpa = fs.gpa; + gpa.free(fs.path); + gpa.free(fs.buf); + gpa.destroy(fs); +} + +/// Frame a request the way the kernel does and push it at the server. +fn pushRequest(fd: c_int, unique: u64, op: Opcode, nodeid: u64, body: []const u8) !void { + var buf: [4096]u8 align(8) = undefined; + const h: fuse_in_header = .{ + .len = @intCast(@sizeOf(fuse_in_header) + body.len), + .opcode = @intFromEnum(op), + .unique = unique, + .nodeid = nodeid, + .uid = 1000, + .gid = 1000, + .pid = 1, + .total_extlen = 0, + .padding = 0, + }; + @memcpy(buf[0..@sizeOf(fuse_in_header)], std.mem.asBytes(&h)); + @memcpy(buf[@sizeOf(fuse_in_header)..][0..body.len], body); + const total = @sizeOf(fuse_in_header) + body.len; + try testing.expectEqual(@as(isize, @intCast(total)), libc.write(fd, &buf, total)); +} + +/// Read one reply back off the socketpair. +fn readReply(fd: c_int, buf: []u8) ![]u8 { + const n = libc.read(fd, buf.ptr, buf.len); + try testing.expect(n >= @sizeOf(fuse_out_header)); + return buf[0..@intCast(n)]; +} + +fn outHeader(bytes: []const u8) fuse_out_header { + var h: fuse_out_header = undefined; + @memcpy(std.mem.asBytes(&h), bytes[0..@sizeOf(fuse_out_header)]); + return h; +} + +/// A socketpair standing in for /dev/fuse. SEQPACKET, not STREAM, and that is +/// the whole point: the kernel's character device hands over exactly one +/// request per read(2) and takes exactly one reply per write(2), and a stream +/// socket would coalesce three requests into one read and let a codec that +/// ignores `fuse_in_header.len` pass anyway. +/// +/// Both ends non-blocking. The server's end so `next()` meets EAGAIN where it +/// would on the real descriptor; the kernel's end so a test can assert that +/// NOTHING was written — which is what "a held request has no reply" and "a +/// FORGET is never answered" mean, and a blocking read would simply hang there +/// instead of failing. +fn testPair() ![2]c_int { + var sv: [2]c_int = undefined; + if (libc.socketpair(libc.AF.UNIX, libc.SOCK.SEQPACKET, 0, &sv) != 0) return error.SocketPairFailed; + setNonblock(sv[0]); + setNonblock(sv[1]); + return sv; +} + +test "lookup round trip: parse borrows the name, reply frames an entry" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + try pushRequest(sv[1], 100, .lookup, 1, "index\x00"); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.lookup, req.op); + try testing.expectEqual(@as(u64, 100), req.tag); + try testing.expectEqual(@as(u64, 1), req.node); + try testing.expectEqualStrings("index", req.data); + // Drained, and nothing else was invented. + try testing.expect(fs.next() == null); + + fs.reply(&.{ + .tag = 100, + .attr = .{ .node = 7, .size = 42, .mode = 0o444 }, + }, &.{}); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + const h = outHeader(got); + try testing.expectEqual(@as(u32, @sizeOf(fuse_out_header) + @sizeOf(fuse_entry_out)), h.len); + try testing.expectEqual(@as(u32, @intCast(got.len)), h.len); + try testing.expectEqual(@as(i32, 0), h.@"error"); + try testing.expectEqual(@as(u64, 100), h.unique); + + var entry: fuse_entry_out = undefined; + @memcpy(std.mem.asBytes(&entry), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_entry_out)]); + try testing.expectEqual(@as(u64, 7), entry.nodeid); + // Caching off in both directions, or `new/<name>` creates a pane once and + // then serves the cached negative lookup forever. + try testing.expectEqual(@as(u64, 0), entry.entry_valid); + try testing.expectEqual(@as(u64, 0), entry.attr_valid); + try testing.expectEqual(@as(u64, 7), entry.attr.ino); + try testing.expectEqual(@as(u64, 42), entry.attr.size); + try testing.expectEqual(S_IFREG | @as(u32, 0o444), entry.attr.mode); + try testing.expectEqual(@as(u32, 1), entry.attr.nlink); + try testing.expectEqual(@as(u32, 1000), entry.attr.uid); + // The slot went back. + try testing.expect(fs.freeSlot() != null); + try testing.expectEqual(@as(?usize, null), fs.findSlot(100)); +} + +test "read reply is capped at the requested size and written as one frame" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const in: fuse_read_in = .{ + .fh = 3, + .offset = 8, + .size = 4, + .read_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + try pushRequest(sv[1], 200, .read, 5, std.mem.asBytes(&in)); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.read, req.op); + try testing.expectEqual(@as(u32, 3), req.handle); + try testing.expectEqual(@as(u64, 8), req.off); + try testing.expectEqual(@as(u32, 4), req.size); + + // The core offers more than was asked for; a reply longer than `size` is + // answered EIO by the kernel, so it has to be clamped here. + fs.reply(&.{ .tag = 200, .payload = .{ .staged = 9 } }, "abcdefghi"); + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header) + 4), got.len); + try testing.expectEqual(@as(u32, @intCast(got.len)), outHeader(got).len); + try testing.expectEqualStrings("abcd", got[@sizeOf(fuse_out_header)..]); +} + +test "an error reply is header only" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + try pushRequest(sv[1], 300, .lookup, 1, "nope\x00"); + _ = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = 300, .status = .err, .errno = @intFromEnum(libc.E.NOENT) }, &.{}); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + // len MUST be exactly the header when error is set; anything else makes the + // kernel answer -EINVAL and leaves the request pending forever. + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header)), got.len); + const h = outHeader(got); + try testing.expectEqual(@as(u32, @sizeOf(fuse_out_header)), h.len); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.NOENT)), h.@"error"); +} + +test "opcodes the core never sees are answered here" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + var buf: [512]u8 = undefined; + + // FORGET and BATCH_FORGET get NO reply, ever: the kernel keeps no pending + // entry for them, so a reply would carry a unique it does not recognise. + const forget: fuse_forget_in = .{ .nlookup = 1 }; + try pushRequest(sv[1], 400, .forget, 7, std.mem.asBytes(&forget)); + const batch: fuse_batch_forget_in = .{ .count = 0, .dummy = 0 }; + try pushRequest(sv[1], 402, .batch_forget, 0, std.mem.asBytes(&batch)); + // ...and a mutation is refused, which is the first thing that produces a + // reply, proving nothing was written for the two above. + try pushRequest(sv[1], 404, .mkdir, 1, "x\x00"); + try testing.expect(fs.next() == null); + + const got = try readReply(sv[1], &buf); + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header)), got.len); + const h = outHeader(got); + try testing.expectEqual(@as(u64, 404), h.unique); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.NOSYS)), h.@"error"); + + // DESTROY must be answered or umount hangs. + try pushRequest(sv[1], 406, .destroy, 0, &.{}); + try testing.expect(fs.next() == null); + const destroyed = try readReply(sv[1], &buf); + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header)), destroyed.len); + try testing.expectEqual(@as(i32, 0), outHeader(destroyed).@"error"); + try testing.expectEqual(@as(u64, 406), outHeader(destroyed).unique); + + // FLUSH is refused so the kernel stops sending one per close(2). + fs.dead = false; + const flush: fuse_flush_in = .{ .fh = 1, .unused = 0, .padding = 0, .lock_owner = 0 }; + try pushRequest(sv[1], 408, .flush, 1, std.mem.asBytes(&flush)); + try testing.expect(fs.next() == null); + const flushed = try readReply(sv[1], &buf); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.NOSYS)), outHeader(flushed).@"error"); +} + +test "setattr size=0 is the truncate the kernel sends instead of O_TRUNC" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + var in: fuse_setattr_in = std.mem.zeroes(fuse_setattr_in); + in.valid = FATTR_SIZE; + in.size = 0; + try pushRequest(sv[1], 500, .setattr, 9, std.mem.asBytes(&in)); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.setattr, req.op); + try testing.expect(req.truncate); + + // A nonzero size is not a truncate this ABI can express, and must not be + // reported as one: the core would clear a pane on `ftruncate(fd, 10)`. + in.size = 10; + try pushRequest(sv[1], 502, .setattr, 9, std.mem.asBytes(&in)); + fs.reply(&.{ .tag = 500, .attr = .{ .node = 9 } }, &.{}); + const req2 = fs.next() orelse return error.NoRequest; + try testing.expect(!req2.truncate); + + fs.reply(&.{ .tag = 502, .attr = .{ .node = 9, .size = 3, .dir = true } }, &.{}); + var buf: [512]u8 = undefined; + _ = try readReply(sv[1], &buf); // the first reply + const got = try readReply(sv[1], &buf); + var out: fuse_attr_out = undefined; + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_attr_out)]); + try testing.expectEqual(S_IFDIR | @as(u32, 0o600), out.attr.mode); + try testing.expectEqual(@as(u32, 2), out.attr.nlink); + try testing.expectEqual(@as(u64, 0), out.attr_valid); +} + +test "open reports direct io for files and nothing for directories" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + var buf: [512]u8 = undefined; + + // O_WRONLY + const wr: fuse_open_in = .{ .flags = 1, .open_flags = 0 }; + try pushRequest(sv[1], 600, .open, 4, std.mem.asBytes(&wr)); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.open, req.op); + fs.reply(&.{ .tag = 600, .handle = 11 }, &.{}); + var got = try readReply(sv[1], &buf); + var open_out: fuse_open_out = undefined; + @memcpy(std.mem.asBytes(&open_out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_open_out)]); + try testing.expectEqual(@as(u64, 11), open_out.fh); + try testing.expectEqual(FOPEN_DIRECT_IO, open_out.open_flags); + + // O_RDONLY on a directory + const rd: fuse_open_in = .{ .flags = 0, .open_flags = 0 }; + try pushRequest(sv[1], 602, .opendir, 1, std.mem.asBytes(&rd)); + const dir_req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.open, dir_req.op); + fs.reply(&.{ .tag = 602, .handle = 12 }, &.{}); + got = try readReply(sv[1], &buf); + @memcpy(std.mem.asBytes(&open_out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_open_out)]); + // No FOPEN_CACHE_DIR either: the pane list changes between two `ls`. + try testing.expectEqual(@as(u32, 0), open_out.open_flags); +} + +test "write borrows the payload and reports the core's own count" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + var body: [@sizeOf(fuse_write_in) + 5]u8 = undefined; + const in: fuse_write_in = .{ + .fh = 2, + .offset = 0, + .size = 5, + .write_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + @memcpy(body[0..@sizeOf(fuse_write_in)], std.mem.asBytes(&in)); + @memcpy(body[@sizeOf(fuse_write_in)..], "hello"); + try pushRequest(sv[1], 700, .write, 6, &body); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.write, req.op); + try testing.expectEqualStrings("hello", req.data); + + fs.reply(&.{ .tag = 700, .written = 5 }, &.{}); + var buf: [512]u8 = undefined; + var got = try readReply(sv[1], &buf); + var out: fuse_write_out = undefined; + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_write_out)]); + try testing.expectEqual(@as(u32, 5), out.size); + + // A short count is a real answer — `data` refusing a partial grapheme — + // and must reach write(2) as a short write rather than as a full one. + try pushRequest(sv[1], 704, .write, 6, &body); + _ = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = 704, .written = 3 }, &.{}); + got = try readReply(sv[1], &buf); + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_write_out)]); + try testing.expectEqual(@as(u32, 3), out.size); + + // A count larger than what was offered would advance the file offset past + // bytes that never existed. + try pushRequest(sv[1], 706, .write, 6, &body); + _ = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = 706, .written = 99 }, &.{}); + got = try readReply(sv[1], &buf); + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_write_out)]); + try testing.expectEqual(@as(u32, 5), out.size); +} + +test "a write whose size lies about the payload is clamped to what arrived" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + var body: [@sizeOf(fuse_write_in) + 2]u8 = undefined; + var in: fuse_write_in = std.mem.zeroes(fuse_write_in); + in.size = 4096; // more than the two bytes that follow + @memcpy(body[0..@sizeOf(fuse_write_in)], std.mem.asBytes(&in)); + @memcpy(body[@sizeOf(fuse_write_in)..], "hi"); + try pushRequest(sv[1], 702, .write, 6, &body); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqualStrings("hi", req.data); + try testing.expectEqual(@as(u32, 2), req.size); +} + +test "dirent encoding: 8-byte records, cookies from the request offset" { + var staged: [64]u8 = undefined; + var w: usize = 0; + // node=2 kind=file name="addr" + std.mem.writeInt(u64, staged[w..][0..8], 2, .little); + staged[w + 8] = 0; + staged[w + 9] = 4; + @memcpy(staged[w + 10 ..][0..4], "addr"); + w += 14; + // node=3 kind=dir name="new" + std.mem.writeInt(u64, staged[w..][0..8], 3, .little); + staged[w + 8] = 1; + staged[w + 9] = 3; + @memcpy(staged[w + 10 ..][0..3], "new"); + w += 13; + + var out: [128]u8 = undefined; + const n = encodeDirents(&out, staged[0..w], 5); + // 24 + 4 -> 32; 24 + 3 -> 32. A record that is not a multiple of 8 + // desynchronises the kernel's parse of everything after it. + try testing.expectEqual(@as(usize, 64), n); + try testing.expectEqual(@as(u64, 0), n % rec_align); + + try testing.expectEqual(@as(u64, 2), std.mem.readInt(u64, out[0..8], .little)); + // Cookies continue from the request's offset: the kernel sends the last + // `off` it saw as the next request's offset, so restarting at 1 would loop + // the directory forever. + try testing.expectEqual(@as(u64, 6), std.mem.readInt(u64, out[8..16], .little)); + try testing.expectEqual(@as(u32, 4), std.mem.readInt(u32, out[16..20], .little)); + try testing.expectEqual(DT_REG, std.mem.readInt(u32, out[20..24], .little)); + try testing.expectEqualStrings("addr", out[24..28]); + // Padding zeroed, so the wire is deterministic. + try testing.expectEqualSlices(u8, &.{ 0, 0, 0, 0 }, out[28..32]); + + try testing.expectEqual(@as(u64, 3), std.mem.readInt(u64, out[32..40], .little)); + try testing.expectEqual(@as(u64, 7), std.mem.readInt(u64, out[40..48], .little)); + try testing.expectEqual(DT_DIR, std.mem.readInt(u32, out[52..56], .little)); + try testing.expectEqualStrings("new", out[56..59]); +} + +test "dirent encoding stops cleanly when the reply buffer or the staging runs out" { + var staged: [64]u8 = undefined; + std.mem.writeInt(u64, staged[0..8], 9, .little); + staged[8] = 0; + staged[9] = 4; + @memcpy(staged[10..14], "body"); + std.mem.writeInt(u64, staged[14..22], 10, .little); + staged[22] = 0; + staged[23] = 4; + @memcpy(staged[24..28], "ctl!"); + + // Room for one record only: the second comes back at the higher cookie. + var out: [40]u8 = undefined; + try testing.expectEqual(@as(usize, 32), encodeDirents(&out, staged[0..28], 0)); + + // A truncated staging record is dropped rather than guessed at. + try testing.expectEqual(@as(usize, 32), encodeDirents(&out, staged[0..26], 0)); + // Zero staged bytes is EOF, not an error. + try testing.expectEqual(@as(usize, 0), encodeDirents(&out, &.{}, 4)); + // A zero name length would make the kernel parse the padding as an entry. + var bad: [10]u8 = @splat(0); + try testing.expectEqual(@as(usize, 0), encodeDirents(&out, &bad, 0)); +} + +test "readdir reply carries encoded dirents built from the staged names" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const in: fuse_read_in = .{ + .fh = 1, + .offset = 0, + .size = 4096, + .read_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + try pushRequest(sv[1], 800, .readdir, 1, std.mem.asBytes(&in)); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.readdir, req.op); + + var staged: [16]u8 = undefined; + std.mem.writeInt(u64, staged[0..8], 4, .little); + staged[8] = 1; + staged[9] = 5; + @memcpy(staged[10..15], "panes"); + fs.reply(&.{ .tag = 800, .payload = .{ .staged = 15 } }, staged[0..15]); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header) + 32), got.len); + const rec = got[@sizeOf(fuse_out_header)..]; + try testing.expectEqual(@as(u64, 4), std.mem.readInt(u64, rec[0..8], .little)); + try testing.expectEqual(@as(u64, 1), std.mem.readInt(u64, rec[8..16], .little)); + try testing.expectEqual(DT_DIR, std.mem.readInt(u32, rec[20..24], .little)); + try testing.expectEqualStrings("panes", rec[24..29]); +} + +test "INIT reply negotiates nothing and caps the minor at ours" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + fs.max_write = 64 * 1024; + + const in: fuse_init_in = .{ + .major = 7, + .minor = 45, + .max_readahead = 131072, + // Everything the kernel is willing to do. The point of the test is that + // none of it comes back. + .flags = 0xffff_ffff, + .flags2 = 0xffff_ffff, + .unused = @splat(0), + }; + try pushRequest(sv[1], 1, .init, 0, std.mem.asBytes(&in)); + try fs.handshake(); + try testing.expectEqual(@as(u32, 45), fs.minor); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header) + @sizeOf(fuse_init_out)), got.len); + try testing.expectEqual(@as(u64, 1), outHeader(got).unique); + var out: fuse_init_out = undefined; + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_init_out)]); + try testing.expectEqual(@as(u32, 7), out.major); + try testing.expectEqual(@as(u32, 45), out.minor); + // The one assertion this test exists for. Every bit here is a kernel + // behaviour we would owe forever: readdirplus whose ENOSYS has no fallback, + // atomic O_TRUNC that would bypass the SETATTR the core handles, locks. + try testing.expectEqual(@as(u32, 0), out.flags); + try testing.expectEqual(@as(u32, 0), out.flags2); + try testing.expectEqual(@as(u32, 0), out.max_readahead); + try testing.expectEqual(@as(u32, 64 * 1024), out.max_write); + // A time granularity of zero is not a legal value. + try testing.expectEqual(@as(u32, 1), out.time_gran); + try testing.expectEqual(@as(u16, 0), out.request_timeout); + + // A newer kernel's minor is CAPPED, not echoed. `fc->minor` is our own + // declared level and it is what sizes the replies the kernel reads back + // from us, so claiming 7.99 on these structs promises fields they do not + // have. This assertion is the one the old `@min(in.minor, in.minor)` could + // not make. + const newer: fuse_init_in = .{ + .major = 7, + .minor = kernel_minor + 54, + .max_readahead = 0, + .flags = 0, + .flags2 = 0, + .unused = @splat(0), + }; + try pushRequest(sv[1], 2, .init, 0, std.mem.asBytes(&newer)); + try fs.handshake(); + const capped = try readReply(sv[1], &buf); + @memcpy(std.mem.asBytes(&out), capped[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_init_out)]); + try testing.expectEqual(kernel_minor, out.minor); + + // A foreign major is fatal, and answering it anyway only moves the failure + // to the first syscall through the mount. + const bad: fuse_init_in = .{ + .major = 8, + .minor = 0, + .max_readahead = 0, + .flags = 0, + .flags2 = 0, + .unused = @splat(0), + }; + try pushRequest(sv[1], 3, .init, 0, std.mem.asBytes(&bad)); + try testing.expectError(error.InitVersion, fs.handshake()); +} + +test "park table: again holds the request, retry offers it back once per round" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const in: fuse_read_in = .{ + .fh = 1, + .offset = 0, + .size = 64, + .read_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + // Two blocked readers of `event`, in arrival order. + try pushRequest(sv[1], 900, .read, 20, std.mem.asBytes(&in)); + try pushRequest(sv[1], 902, .read, 21, std.mem.asBytes(&in)); + const a = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = a.tag, .status = .again }, &.{}); + const b = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = b.tag, .status = .again }, &.{}); + try testing.expect(fs.next() == null); + // Nothing was written: a held request has no reply, which is the only way + // FUSE expresses blocking. + var buf: [512]u8 = undefined; + try testing.expect(libc.read(sv[1], &buf, buf.len) < 0); + + // One round offers each parked request exactly once, oldest first, and then + // ends. Without the per-round flag the oldest would be offered forever and + // the second reader would never be looked at again. + const r1 = fs.retry() orelse return error.NoRetry; + try testing.expectEqual(@as(u64, 900), r1.tag); + fs.reply(&.{ .tag = r1.tag, .status = .again }, &.{}); + const r2 = fs.retry() orelse return error.NoRetry; + try testing.expectEqual(@as(u64, 902), r2.tag); + fs.reply(&.{ .tag = r2.tag, .status = .again }, &.{}); + try testing.expect(fs.retry() == null); + + // ...and the next round starts over. + const r3 = fs.retry() orelse return error.NoRetry; + try testing.expectEqual(@as(u64, 900), r3.tag); + fs.reply(&.{ .tag = r3.tag, .payload = .{ .staged = 3 } }, "ev\n"); + const got = try readReply(sv[1], &buf); + try testing.expectEqualStrings("ev\n", got[@sizeOf(fuse_out_header)..]); + // The answered one is gone; the other is still held. + try testing.expectEqual(@as(?usize, null), fs.findSlot(900)); + try testing.expect(fs.findSlot(902) != null); +} + +test "park table: interrupt answers the original with EINTR and drops it" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const in: fuse_read_in = .{ + .fh = 1, + .offset = 0, + .size = 64, + .read_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + try pushRequest(sv[1], 1000, .read, 20, std.mem.asBytes(&in)); + try pushRequest(sv[1], 1002, .read, 21, std.mem.asBytes(&in)); + const a = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = a.tag, .status = .again }, &.{}); + const b = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = b.tag, .status = .again }, &.{}); + + // The kernel's interrupt names the ORIGINAL unique in its body; its own + // unique is `original | 1`, which is why it must not be echoed. + const intr: fuse_interrupt_in = .{ .unique = 1002 }; + try pushRequest(sv[1], 1002 | 1, .interrupt, 0, std.mem.asBytes(&intr)); + try testing.expect(fs.next() == null); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + // Exactly one reply, to the interrupted request, not to the interrupt. + // Getting this wrong leaves a SIGKILLed reader in uninterruptible sleep. + try testing.expectEqual(@as(usize, @sizeOf(fuse_out_header)), got.len); + const h = outHeader(got); + try testing.expectEqual(@as(u64, 1002), h.unique); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.INTR)), h.@"error"); + try testing.expectEqual(@as(?usize, null), fs.findSlot(1002)); + try testing.expect(fs.findSlot(1000) != null); + + // An interrupt for something we do not hold is ignored, not answered. + const stale: fuse_interrupt_in = .{ .unique = 4242 }; + try pushRequest(sv[1], 4243, .interrupt, 0, std.mem.asBytes(&stale)); + try testing.expect(fs.next() == null); + try testing.expect(libc.read(sv[1], &buf, buf.len) < 0); +} + +test "park table: a full table answers EAGAIN and keeps the descriptor flowing" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const in: fuse_read_in = .{ + .fh = 1, + .offset = 0, + .size = 8, + .read_flags = 0, + .lock_owner = 0, + .flags = 0, + .padding = 0, + }; + for (0..max_slots) |i| { + try pushRequest(sv[1], 2000 + i * 2, .read, 30, std.mem.asBytes(&in)); + const req = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = req.tag, .status = .again }, &.{}); + } + var buf: [512]u8 = undefined; + + // One more than the table holds. It is READ and refused, not left queued. + // Gating the read on a free slot is a deadlock dressed as backpressure: + // the INTERRUPT that frees a slot would never be read either, so a + // SIGKILLed reader would stay in uninterruptible sleep and every unrelated + // `ls` of the mount would hang behind the 32 blocked ones. Measured: that + // wedges a real mount. + try pushRequest(sv[1], 9998, .read, 30, std.mem.asBytes(&in)); + try testing.expect(fs.next() == null); + const refused = try readReply(sv[1], &buf); + try testing.expectEqual(@as(u64, 9998), outHeader(refused).unique); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.AGAIN)), outHeader(refused).@"error"); + + // And the requests that need no slot keep being answered with the table + // still full — DESTROY above all, since a missing reply to it hangs umount. + try pushRequest(sv[1], 9990, .access, 1, &.{}); + try testing.expect(fs.next() == null); + const nosys = try readReply(sv[1], &buf); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.NOSYS)), outHeader(nosys).@"error"); + + // An interrupt still lands, which is what lets a full table recover at all. + const intr: fuse_interrupt_in = .{ .unique = 2000 }; + try pushRequest(sv[1], 2001, .interrupt, 0, std.mem.asBytes(&intr)); + try testing.expect(fs.next() == null); + const killed = try readReply(sv[1], &buf); + try testing.expectEqual(@as(u64, 2000), outHeader(killed).unique); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.INTR)), outHeader(killed).@"error"); + + // ...and the freed slot takes the next request. + try pushRequest(sv[1], 9996, .read, 30, std.mem.asBytes(&in)); + const late = fs.next() orelse return error.NoRequest; + try testing.expectEqual(@as(u64, 9996), late.tag); +} + +test "park table: a payload too large to copy is refused rather than dangled" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const payload_len = park_data_max + 1; + var body: [@sizeOf(fuse_write_in) + payload_len]u8 = undefined; + var in: fuse_write_in = std.mem.zeroes(fuse_write_in); + in.size = payload_len; + @memcpy(body[0..@sizeOf(fuse_write_in)], std.mem.asBytes(&in)); + @memset(body[@sizeOf(fuse_write_in)..], 'z'); + try pushRequest(sv[1], 3000, .write, 6, &body); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(@as(usize, payload_len), req.data.len); + + // Parking this would park a slice of the read buffer, which the next + // `next()` overwrites. EAGAIN is the honest answer. + fs.reply(&.{ .tag = 3000, .status = .again }, &.{}); + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + try testing.expectEqual(-@as(i32, @intFromEnum(libc.E.AGAIN)), outHeader(got).@"error"); + try testing.expectEqual(@as(?usize, null), fs.findSlot(3000)); + + // A payload that fits IS copied, so parking it is safe even after the read + // buffer has been reused. + var small: [@sizeOf(fuse_write_in) + 4]u8 = undefined; + in.size = 4; + @memcpy(small[0..@sizeOf(fuse_write_in)], std.mem.asBytes(&in)); + @memcpy(small[@sizeOf(fuse_write_in)..], "keep"); + try pushRequest(sv[1], 3002, .write, 6, &small); + const kept = fs.next() orelse return error.NoRequest; + fs.reply(&.{ .tag = kept.tag, .status = .again }, &.{}); + // Something else lands in the read buffer... + try pushRequest(sv[1], 3004, .statfs, 1, &.{}); + _ = fs.next() orelse return error.NoRequest; + // ...and the parked bytes survived it. + const again = fs.retry() orelse return error.NoRetry; + try testing.expectEqualStrings("keep", again.data); +} + +test "a reply for a tag we no longer hold is dropped, not written" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + // The interrupt path already answered and freed this one; a second reply + // would carry a unique the kernel does not recognise, and could in + // principle be matched against a live request that reused the number. + fs.reply(&.{ .tag = 12345 }, &.{}); + var buf: [512]u8 = undefined; + try testing.expect(libc.read(sv[1], &buf, buf.len) < 0); +} + +test "statfs reports a usable namelen" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + try pushRequest(sv[1], 1100, .statfs, 1, &.{}); + const req = fs.next() orelse return error.NoRequest; + try testing.expectEqual(acmefs.Op.statfs, req.op); + fs.reply(&.{ .tag = 1100 }, &.{}); + + var buf: [512]u8 = undefined; + const got = try readReply(sv[1], &buf); + var out: fuse_statfs_out = undefined; + @memcpy(std.mem.asBytes(&out), got[@sizeOf(fuse_out_header)..][0..@sizeOf(fuse_statfs_out)]); + // Zero here makes pathconf(_PC_NAME_MAX) return 0 and some tools then + // refuse to create any name at all. + try testing.expectEqual(@as(u32, 255), out.st.namelen); + try testing.expectEqual(@as(u32, 4096), out.st.bsize); +} + +test "the fusermount command line and environment" { + var opts_buf: [128:0]u8 = undefined; + const opts = mountOpts(&opts_buf); + try testing.expectEqualStrings("fsname=pardes,subtype=pardes,nosuid,nodev", opts); + // allow_other needs user_allow_other in /etc/fuse.conf, which is commented + // out on a stock install, and asking for it FAILS the whole mount rather + // than being ignored. default_permissions would move access control out of + // the core and into a mode nibble. + try testing.expect(std.mem.indexOf(u8, opts, "allow_other") == null); + try testing.expect(std.mem.indexOf(u8, opts, "default_permissions") == null); + + var env_buf: [32:0]u8 = undefined; + try testing.expectEqualStrings("_FUSE_COMMFD=7", commfdEnv(&env_buf, 7)); + + var argv: [6:null]?[*:0]const u8 = undefined; + mountArgv(&argv, "/usr/bin/fusermount3", opts.ptr, "/run/user/1000/pardes/42"); + try testing.expectEqualStrings("/usr/bin/fusermount3", std.mem.span(argv[0].?)); + try testing.expectEqualStrings("-o", std.mem.span(argv[1].?)); + try testing.expectEqualStrings("fsname=pardes,subtype=pardes,nosuid,nodev", std.mem.span(argv[2].?)); + // Without the `--` a mountpoint beginning with a dash is parsed as a flag + // by a setuid program. + try testing.expectEqualStrings("--", std.mem.span(argv[3].?)); + try testing.expectEqualStrings("/run/user/1000/pardes/42", std.mem.span(argv[4].?)); + try testing.expectEqual(@as(?[*:0]const u8, null), argv[5]); + + var uargv: [7:null]?[*:0]const u8 = undefined; + unmountArgv(&uargv, "/usr/bin/fusermount3", "/run/user/1000/pardes/42"); + try testing.expectEqualStrings("-u", std.mem.span(uargv[1].?)); + try testing.expectEqualStrings("-q", std.mem.span(uargv[2].?)); + // Lazy, or a pane shell with a cwd inside the mount makes the unmount fail + // with EBUSY and the mount outlives the editor. + try testing.expectEqualStrings("-z", std.mem.span(uargv[3].?)); + try testing.expectEqualStrings("--", std.mem.span(uargv[4].?)); + try testing.expectEqual(@as(?[*:0]const u8, null), uargv[6]); +} + +test "the child environment drops an inherited comm descriptor" { + if (comptime !supported) return; + const gpa = testing.allocator; + var buf: [32:0]u8 = undefined; + const commfd = commfdEnv(&buf, 5); + const env = try buildEnv(gpa, commfd); + defer gpa.free(env); + + // Exactly one _FUSE_COMMFD, and it is ours: getenv returns the FIRST match, + // so an inherited stale entry would win and fusermount3 would send the + // descriptor to a closed socket. + var seen: usize = 0; + var i: usize = 0; + while (env[i]) |entry| : (i += 1) { + if (std.mem.startsWith(u8, std.mem.span(entry), commfd_env ++ "=")) { + seen += 1; + try testing.expectEqualStrings("_FUSE_COMMFD=5", std.mem.span(entry)); + } + } + try testing.expectEqual(@as(usize, 1), seen); + try testing.expectEqual(@as(?[*:0]const u8, null), env[env.len - 1]); +} + +test "poll thread: one wake per drained batch, and stop joins from either state" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + var wakes: std.atomic.Value(u32) = .init(0); + const Sink = struct { + fn wake(ctx: ?*anyopaque) void { + const c: *std.atomic.Value(u32) = @ptrCast(@alignCast(ctx.?)); + _ = c.fetchAdd(1, .release); + } + }; + try fs.wakeThread(&wakes, Sink.wake); + + // One pending request, one wake. A FORGET is answered inside `next()` and + // never surfaces, so draining to null is the whole batch — and it is that + // null which raises the eventfd and lets the poller poll again. + const forget: fuse_forget_in = .{ .nlookup = 1 }; + try pushRequest(sv[1], 7000, .forget, 2, std.mem.asBytes(&forget)); + while (wakes.load(.acquire) == 0) std.Thread.yield() catch {}; + try testing.expectEqual(@as(?acmefs.Req, null), fs.next()); + + // The poller is now in one of the two states a stop has to break: still in + // the blocking wait, or back in `poll()` because the drain above beat the + // stop there. Which one is a race, deliberately unresolved — the assertion + // is that either joins, and a hang here is this test's only failure mode. + fs.stopThread(); + try testing.expect(fs.thread == null); + try testing.expectEqual(@as(c_int, -1), fs.ctl); +} + +test "poll thread: stop breaks a poller that never saw a request" { + if (comptime !supported) return; + const gpa = testing.allocator; + const sv = try testPair(); + defer { + _ = libc.close(sv[0]); + _ = libc.close(sv[1]); + } + const fs = try testFs(gpa, sv[0]); + defer testFsFree(fs); + + const Sink = struct { + fn wake(_: ?*anyopaque) void { + unreachable; // nothing is ever pending on this descriptor + } + }; + try fs.wakeThread(null, Sink.wake); + // Covers the two states with no acknowledgement in them at all: blocked in + // `poll()` with an idle descriptor, and not yet past the loop condition. + fs.stopThread(); + try testing.expect(fs.thread == null); +} + +test "mount refuses a relative point" { + if (comptime !supported) return; + try testing.expectError( + error.MountPathNotAbsolute, + Fs.mount(testing.allocator, .{ .mount = "relative/dir" }), + ); +} diff --git a/src/gui/gui.zig b/src/gui/gui.zig index e080dd9a..092ac2e4 100644 --- a/src/gui/gui.zig +++ b/src/gui/gui.zig @@ -26,6 +26,8 @@ const selection_pipe = @import("../selection_pipe.zig"); const shell_bin = @import("../shell_bin.zig"); const nested = @import("../nested.zig"); +const fuse = @import("../fuse.zig"); +const fs_service = @import("../fs_service.zig"); pub const c = @cImport({ @cDefine("SDL_DISABLE_OLD_NAMES", "1"); @@ -851,6 +853,12 @@ const Msg = union(enum) { /// a pardes launched inside this one sent us a builtin command line (see /// lookThread); gpa-owned, like `output` bytes command: []u8, + /// `--fs`: the /dev/fuse descriptor has requests on it. Carries nothing — + /// the drain lives in pollFrame, and this only ends a blocking + /// SDL_WaitEventTimeout. Posted by the poll thread and, when a batch hits + /// its cap, by pollFrame itself. Lossy under backpressure on purpose: a + /// full queue already holds something that will wake the loop. + fs_ready, fn deinit(m: Msg, gpa: std.mem.Allocator, lsp_allocator: std.mem.Allocator) void { switch (m) { @@ -861,7 +869,7 @@ const Msg = union(enum) { response.deinit(gpa); }, .command => |line| gpa.free(line), - .eof, .files_changed => {}, + .eof, .files_changed, .fs_ready => {}, } } }; @@ -1903,6 +1911,18 @@ fn runNative(init: std.process.Init, opts_in: pardes.Options) !void { const sock_fd: c_int = if (opts.nested) -1 else nested.listen(); defer nested.unlisten(sock_fd); + // `--fs`: mounted before the initial spawns (they are the shells that need + // PARDES_FS) and before any thread of ours exists (the mount forks the + // setuid fusermount3 helper). Null covers both "no --fs" and "--fs but the + // mount failed"; the second is reported on a message row inside `start` and + // the session runs on without a filesystem. Teardown answers everything + // held, aborts the connection, unmounts and removes `<parent>/<pid>`; the + // parent stays, like nested.zig's socket directory. + var fs = fs_service.start(gpa, core); + // Covers the error paths only: the ordinary exit unmounts at the END OF + // THE LOOP instead, see there. + defer if (fs) |f| f.deinit(); + var shell: Shell = .{ .core = core, .gui = &g, @@ -1916,6 +1936,7 @@ fn runNative(init: std.process.Init, opts_in: pardes.Options) !void { .pipe_tasks = &pipe_tasks, .inotify_fd = inotify_fd, .watches = &watches, + .fs = fs, .test_mode = test_mode, }; const host = shell.host(); @@ -1935,6 +1956,10 @@ fn runNative(init: std.process.Init, opts_in: pardes.Options) !void { // ...and the nested-instance listener, detached like every other blocking // worker here if (sock_fd >= 0) if (std.Thread.spawn(.{}, lookThread, .{ gpa, sock_fd, &queue })) |th| th.detach() else |_| {}; + // ...and the /dev/fuse poller, which is the same kind of thread again — + // except joined by `Fs.deinit` rather than detached, because fuse.zig gives + // it a control pipe that CAN wake it out of poll(). + fs_service.wake(fs, &queue, wakeFs); _ = c.SDL_StartTextInput(window); @@ -1987,6 +2012,16 @@ fn runNative(init: std.process.Init, opts_in: pardes.Options) !void { syncTaglineFont(&g, core); } } + + // THE FILESYSTEM GOES FIRST, ahead of every deferred teardown below: a + // session that has decided to exit must not spend its teardown holding a + // mount nobody is serving, so a client blocked on `<id>/event` when the + // last pane is deleted through `ctl` gets ENOTCONN at once. + if (fs) |f| { + f.deinit(); + fs = null; + shell.fs = null; + } } /// PARDES_TEST_GRID=1: headless. No SDL at all — stdin escape sequences in, @@ -2059,6 +2094,11 @@ fn runGrid(init: std.process.Init, opts_in: pardes.Options) !void { // scripted input event, and a reload that arrives on its own clock would // put a frame in the stream nothing asked for. -1 makes watchPane a no-op. var watches: file_watch.Table = @splat(null); + // The filesystem IS served here, unlike the file watcher above: `--fs=<dir>` + // names a predictable mount point precisely so a snapshot can drive this + // mode through it. No poll thread though — see gridPollFrame. + const fs = fs_service.start(gpa, core); + defer if (fs) |f| f.deinit(); var shell: Shell = .{ .core = core, .io = io, @@ -2071,6 +2111,7 @@ fn runGrid(init: std.process.Init, opts_in: pardes.Options) !void { .pipe_tasks = &pipe_tasks, .inotify_fd = -1, .watches = &watches, + .fs = fs, }; const host = shell.host(); // The core owns the loop here too, but not the scripted stdin: EOF ends @@ -2997,6 +3038,10 @@ const Shell = struct { gui: ?*Gui = null, io: std.Io, gpa: std.mem.Allocator, + /// acme's control filesystem for this session, or null when `--fs` was not + /// given (or its mount failed, or this is the headless grid harness, which + /// serves nothing). Owned by `run`. + fs: ?*fuse.Fs = null, lsp_allocator: std.mem.Allocator, prompt_rcs: *const shell_bin.PromptRcs, ptys: *[pardes.MAX_PANES]?Pty, @@ -3048,6 +3093,7 @@ const Shell = struct { .push_open_link = openLink, .pull_lsp = lsp, .pull_pipe = pipe, + .push_fs_reply = fsReply, }; /// The headless grid harness. It reads its scripted stdin itself, because @@ -3110,12 +3156,34 @@ const Shell = struct { // `git checkout`) collapses into ONE pass below, so it cannot queue // a reload — or an undo entry — per write. .files_changed => check_files = true, + // A wake and nothing more; the requests behind it are drained in + // pollFrame, which is where a whole batch can be answered against + // one render instead of one render per request. + .fs_ready => {}, }; if (check_files and file_watch.reloadChanged(s.core, s.io, s.gpa, s.watches)) s.queue.push(.files_changed); } }; +/// The core's answer to one filesystem request, handed straight back to the +/// transport holding it. `bytes` was resolved by `pardes.fsPayload` inside +/// `perform` and is borrowed only for this call, so a megabyte body read copies +/// nothing. `.again` needs no case here: `Fs.reply` reads the status and +/// re-parks the request itself. +fn fsReply(ctx: ?*anyopaque, reply: *const pardes.acmefs.Reply, bytes: []const u8) void { + const s = shellOf(ctx); + if (s.fs) |f| f.reply(reply, bytes); +} + +/// The /dev/fuse poller's wake. `Queue.push` is the thread-safe door and +/// already raises the SDL user event that ends a blocking WaitEventTimeout, so +/// this is the whole callback — the same shape as watchThread's. +fn wakeFs(ctx: ?*anyopaque) void { + const q: *Queue = @ptrCast(@alignCast(ctx.?)); + q.push(.fs_ready); +} + fn shellOf(ctx: ?*anyopaque) *Shell { return @ptrCast(@alignCast(ctx.?)); } @@ -3159,6 +3227,14 @@ fn waitInput(ctx: ?*anyopaque, timeout_ms: u32) void { fn pollFrame(ctx: ?*anyopaque) void { const s = shellOf(ctx); const core = s.core; + // acme's filesystem: one batch per frame, answered before anything else in + // the pass, so an edit a script just made through `body` is in the surface + // this frame composes. Ahead of the `s.gui orelse return` below because it + // has nothing to do with pixels. Hitting the cap means no ack reached the + // poll thread, so nothing else will wake us — re-arm the loop ourselves; + // `Queue.push` is lossy for this variant, which is correct, because a queue + // too full to take a wake is already holding one. + if (s.fs) |f| if (fs_service.drain(f, core).pending) s.queue.push(.fs_ready); const g = s.gui orelse return; // TaglineSize is pure renderer state: update the smaller face and its // visual band immediately, without changing the body metrics or grid. @@ -3257,6 +3333,16 @@ fn postPresent(ctx: ?*anyopaque) void { /// animation step, and the text of the grid itself. fn gridPollFrame(ctx: ?*anyopaque) void { const s = shellOf(ctx); + // The filesystem, drained on the pass rather than woken by a thread: this + // mode's contract is one frame per scripted event, and a poller posting on + // its own clock would put frames in the stream nothing asked for. fuse.zig + // supports exactly this — skip `wakeThread` and drain from the frame poll — + // and here it is not a degradation but the point. A request that changed + // something IS an event, so the pass renders: that is what lets a snapshot + // `wait` for text a script wrote through the mount. + if (s.fs) |f| { + if (fs_service.drain(f, s.core).count != 0) s.saw_event = true; + } pollCwds(s.core, s.ptys); // The harness polls stdin at the same 16 ms cadence as native SDL, so a // pass IS a frame interval and the tick is due here rather than behind a @@ -3307,7 +3393,7 @@ fn spawnPane(ctx: ?*anyopaque, pane: u8, cwd: []const u8) void { cwd_buf[cwd.len] = 0; cwd_z = @ptrCast(&cwd_buf); } - const pt = forkShell(s.core, pane, s.prompt_rcs, s.core.shellBin(), cwd_z, s.core.screen_h, s.core.screen_w); + const pt = forkShell(s.core, pane, s.prompt_rcs, s.core.shellBin(), cwd_z, s.core.screen_h, s.core.screen_w, s.fs); s.ptys[pane] = pt; // report the pane's starting directory back to the core (tags); the slot // needs no occupancy reset, nothing about it is remembered @@ -3448,12 +3534,16 @@ fn pipe(ctx: ?*anyopaque, id: u32) void { if (s.threads_ok) spawnPipe(s.core, s.io, s.gpa, s.queue, s.pipe_tasks, id); } -fn forkShell(core: *pardes.Pardes, pane: usize, prompt_rcs: *const shell_bin.PromptRcs, bin: []const u8, cwd: ?[*:0]const u8, rows: u16, cols: u16) Pty { +fn forkShell(core: *pardes.Pardes, pane: usize, prompt_rcs: *const shell_bin.PromptRcs, bin: []const u8, cwd: ?[*:0]const u8, rows: u16, cols: u16, fs: ?*const fuse.Fs) Pty { var master: c_int = undefined; // resolved BEFORE the fork, into this frame, which the child inherits: // nothing between fork and exec may allocate, and a PATH search would var path_buf: [std.fs.max_path_bytes]u8 = undefined; const spawn = shell_bin.resolve(bin, &path_buf, prompt_rcs); + // ...and so is the pane's own address on the control filesystem: acme puts + // `winid` in the child, which is safe there only because rfork(RFENVG) has + // just given it a private environment group. See fs_service.exportPaneEnv. + fs_service.exportPaneEnv(fs, if (core.panes[pane]) |pn| pn.serial else 0); const ws = posix.winsize{ .row = rows, .col = cols, .xpixel = 0, .ypixel = 0 }; const pid = forkpty(&master, null, null, &ws); if (pid == 0) { diff --git a/src/host.zig b/src/host.zig index d3e93d36..e44318da 100644 --- a/src/host.zig +++ b/src/host.zig @@ -107,6 +107,14 @@ pub const Host = struct { /// completed and the core would apply it to whatever holds that id now. pull_lsp: ?*const fn (ctx: ?*anyopaque, req: LspRequest) void = null, pull_pipe: ?*const fn (ctx: ?*anyopaque, id: u32) void = null, + /// Hand one filesystem answer back to whoever asked for it (a FUSE + /// `write(2)` to /dev/fuse). `bytes` is the payload the core resolved + /// for this reply and is borrowed for the length of this call — it may + /// point straight into a pane's text, so a host that needs it later + /// copies it. A push and not a pull: the answer is already computed, + /// and a second host serving the same mount is not a thing that + /// happens (the transport that asked is the one holding the request). + push_fs_reply: ?*const fn (ctx: ?*anyopaque, reply: *const pardes.acmefs.Reply, bytes: []const u8) void = null, }; }; diff --git a/src/main.zig b/src/main.zig index be188200..69cc1715 100644 --- a/src/main.zig +++ b/src/main.zig @@ -80,6 +80,11 @@ const help_text = \\ Without it, a pardes started inside a pardes \\ hands its FILE argument to the outer one. This \\ session will not serve its own children either. + \\ --fs serve acme's control filesystem for this session + \\ under $XDG_RUNTIME_DIR/pardes/<pid>, and export + \\ PARDES_FS and PARDES_PANE into every pane shell + \\ --fs=<dir> ...at <dir> instead. Must be absolute; pardes + \\ unmounts it on exit but leaves the directory \\ -h, --help show this help and exit \\ ; @@ -120,11 +125,17 @@ fn nativeMain(init: std.process.Init) !void { var opts: pardes.Options = .{}; const arena = init.arena.allocator(); const args = try init.minimal.args.toSlice(arena); - // bare `pardes` boots straight into tty mode — and `pardes --nested` still - // counts as bare, because --nested says something about the session's - // relationship to its parent and nothing about its layout - if (args.len == 1 or (args.len == 2 and std.mem.eql(u8, args[1], "--nested"))) - opts.tty_only = true; + // bare `pardes` boots straight into tty mode — and so does a `pardes` + // carrying nothing but flags that say something about the SESSION rather + // than about its layout: --nested is about this session's relationship to + // its parent, --fs is about who may script it, and neither says anything + // about what should be on screen. Anything else (a FILE, -n, --tty) is + // layout, and answers this question itself further down. + opts.tty_only = for (args[1..]) |a| { + if (!std.mem.eql(u8, a, "--nested") and + !std.mem.eql(u8, a, "--fs") and + !std.mem.startsWith(u8, a, "--fs=")) break false; + } else true; // Kept RAW until every flag is parsed: classifying it means chdir'ing into // a directory and recording nothing, and the nested client below still // needs the word itself to resolve. @@ -148,6 +159,16 @@ fn nativeMain(init: std.process.Init) !void { i += 1; if (i >= args.len) return error.BadArgs; opts.load_path = args[i]; + } else if (std.mem.eql(u8, a, "--fs")) { + opts.fs = ""; + } else if (std.mem.startsWith(u8, a, "--fs=")) { + // `--fs=<dir>` and NEVER `--fs <dir>`, which is the one place this + // parser cannot follow --tty-toggle: --tty-toggle's argument is + // mandatory, so consuming the next word is unambiguous. `--fs` is + // useful bare, so a two-word form would make `pardes --fs README` + // mount at ./README and open no file — the flag would silently eat + // the FILE argument. One spelling, and it carries its own value. + opts.fs = a["--fs=".len..]; } else if (std.mem.eql(u8, a, "--nested")) { opts.nested = true; } else if (std.mem.eql(u8, a, "-h") or std.mem.eql(u8, a, "--help")) { @@ -241,6 +262,20 @@ fn parseCtrlKey(raw: []const u8) ?u21 { test { _ = @import("user_config.zig"); _ = @import("allocators.zig"); + _ = @import("fs_service.zig"); + // acme's control filesystem, both halves, and NOT their own b.addTest + // modules in build.zig the way temp_file/nested/fonts are: both reach + // src/pardes.zig (acmefs takes a *Pardes, fuse.zig speaks its Req/Reply), + // so a standalone module would have to re-wire ghostty-vt, tree-sitter, the + // themes and every option the core imports. This module already has them. + // + // fuse.zig genuinely needs its name here (nothing the core analyses reaches + // it — only the shells import it). acmefs.zig does not today, because + // pardes.zig re-exports it unconditionally; it is named anyway, because the + // day that re-export grows a comptime gate is the day 23 tests disappear in + // silence. That is the fonts.zig story above, told once already. + _ = @import("acmefs.zig"); + _ = @import("fuse.zig"); if (comptime pardes.platform == .tty) { _ = @import("tty/tty.zig"); // tty.zig calls the compositor only from its runtime loop, so merely diff --git a/src/output_pane.zig b/src/output_pane.zig index a7763bd3..9021e479 100644 --- a/src/output_pane.zig +++ b/src/output_pane.zig @@ -48,6 +48,11 @@ pub const Origin = union(enum) { query: lsp.Kind, /// the bare `/` — the pane's own text, searched in core search, + /// acme's `+Errors`: whatever a script wrote to a pane's `errors` file or + /// to the top-level `cons`. Not a command at all — the third vocabulary + /// is "somebody else's output", and it has no word to click because the + /// writer is a process, not a keystroke. + errors, }; /// The exact command identity. A search prompt is bounded by the same one-line @@ -92,6 +97,10 @@ pub fn traits(o: Origin) Traits { return switch (o) { // rows are `location text`, so n/N walk them .search => .{ .name = config.search_buffer, .steps = true }, + // A transcript, not a list: rows are whatever a program printed, so + // n/N walks its words like any prose buffer, and there is nothing to + // Save — acme's +Errors is not a file either. + .errors => .{ .name = config.errors_buffer, .doc = true }, .cmd => |b| builtins.registry.outputTraits(b) orelse unreachable, .query => |k| switch (k) { .hover => .{ .name = config.hover_buffer }, @@ -158,6 +167,7 @@ pub fn word(o: Origin) []const u8 { .cmd => |b| @tagName(b), .query => |k| @tagName(k), .search => "/", + .errors => config.errors_buffer, }; } @@ -166,6 +176,7 @@ pub fn word(o: Origin) []const u8 { pub fn fromWord(w: []const u8) ?Origin { if (w.len == 0) return null; if (std.mem.eql(u8, w, "/")) return .search; + if (std.mem.eql(u8, w, config.errors_buffer)) return .errors; if (std.meta.stringToEnum(Builtin, w)) |b| if (builtins.registry.outputTraits(b) != null) return .{ .cmd = b }; if (std.meta.stringToEnum(lsp.Kind, w)) |k| return .{ .query = k }; diff --git a/src/pardes.zig b/src/pardes.zig index b4a0b974..c5eda6aa 100644 --- a/src/pardes.zig +++ b/src/pardes.zig @@ -39,6 +39,10 @@ const output_pane = @import("output_pane.zig"); const builtins = @import("builtins.zig"); const runtime_cfg = @import("runtime_config.zig"); const selection_pipe = @import("selection_pipe.zig"); +/// acme's control filesystem, as a pure transaction over this core: the FILES +/// a script opens (`body`, `ctl`, `event`, ...) and what they mean. The +/// transport that carries requests in is a host's business (src/fuse.zig). +pub const acmefs = @import("acmefs.zig"); pub const config = @import("config.zig"); pub const pdf_enabled = pdf_pane.enabled; pub const pdf = pdf_pane.pdf; @@ -3428,6 +3432,14 @@ pub const Event = union(enum) { /// this must not clamp onto and preview the final grid cell. pointer_leave, tick, + /// ONE FILESYSTEM REQUEST from a process that opened a file under the + /// acme-style control mount (src/acmefs.zig, served by src/fuse.zig). + /// `data` is borrowed for this call exactly like `output` bytes, which is + /// why this — like them — never goes through `postEvent`. The answer + /// leaves as an `Effect.fs_reply` in the same update, so the transport + /// that asked is the one that writes it back: no thread, no waiting and + /// no filesystem knowledge anywhere in here. + fs_req: acmefs.Req, }; /// IO the core wants done. Payloads are inline (fixed buffers): effects are @@ -3483,6 +3495,14 @@ pub const Effect = union(enum) { /// Write the build-time theme ring below the per-user config directory. /// The pane receives the completion/error message from the native host. dump_themes: struct { pane: u8 }, + /// The answer to an `Event.fs_req`. The bytes are NOT in here: `payload` + /// says where they live (a staging buffer in the core, or a range of a + /// pane's live text) and `fsPayload` resolves it during the drain, so a + /// megabyte read costs one `writev` and no copy. `.again` means the core + /// has nothing yet and the transport must ask again later — acme's + /// blocking `event` read, with the waiting left where the kernel's + /// request already is. + fs_reply: acmefs.Reply, quit, fn Buf(comptime n: usize) type { @@ -5188,6 +5208,18 @@ pub const Options = struct { /// deterministic; when present, each line is dispatched as a builtin /// before init returns and therefore before any frontend can render. startup_config: ?[]const u8 = null, + /// SERVE ACME'S CONTROL FILESYSTEM for this session (`--fs`), and where. + /// `null` is off; `""` means "derive the mount point" (a per-session + /// directory under `$XDG_RUNTIME_DIR`); anything else is the directory + /// `--fs=<dir>` named, which scripts and the snapshot harness need because + /// they have to predict it. + /// + /// One field rather than a flag plus a path: two of those encode a state + /// ("no filesystem, mounted here") that means nothing. The core never + /// mounts anything — a mount is a host's business, and one host (the + /// browser) has no filesystem at all — but the option rides here because + /// argv already reaches the shells this way, like `nested`. + fs: ?[]const u8 = null, /// Native launcher's resolved per-user `pardes` directory. Relative /// ThemeFile operands and DumpThemes are rooted here. Null for web and /// direct core callers which did not opt into per-user configuration. @@ -5512,6 +5544,11 @@ pub const Pardes = struct { pipe_seq: u32 = 0, pipe_wait: ?PendingPipe = null, + /// acme's control filesystem, when a host serves one (`pardes --fs`). + /// Zero-initialised and inert: a core nobody scripts pays for one branch + /// per edit and nothing else. See src/acmefs.zig. + fs: acmefs.State = .{}, + /// Pending effects, drained by the shell after each update. The bounded /// ring preserves byte order; once full, later effects are refused so no /// already-queued write can be reordered or silently evicted. @@ -5652,6 +5689,7 @@ pub const Pardes = struct { if (p.custom_theme) |theme_value| std.zon.parse.free(gpa, theme_value); if (p.chord_arg) |a| gpa.free(a); if (p.pipe_wait) |*wait| wait.deinit(gpa); + p.fs.deinit(gpa); p.shell_rows.reset(gpa); p.scratch.deinit(); p.frame_arena.deinit(); @@ -5693,8 +5731,13 @@ pub const Pardes = struct { // watch keyed to this slot and drop any hover it owns. The heap // teardown is deferred (reapPanes) so pointers to it survive the frame. const watched = (if (pane.file) |f| f.output == null else false) or hasPdf(pane); - if (watched) for (p.panes, 0..) |slot, id| { - if (slot == pane) p.emit(.{ .watch = .{ .pane = @intCast(id), .on = false } }); + for (p.panes, 0..) |slot, id| if (slot == pane) { + if (watched) p.emit(.{ .watch = .{ .pane = @intCast(id), .on = false } }); + // A pane's filesystem state dies WITH the pane, here, while the + // slot still names it: the alternative is a script that held its + // `event` file open leaving the editor suppressing button actions + // for whatever pane lands in this slot next. + p.fs.forget(p.gpa, id); }; if (p.lookHoverPane()) |h| if (h < p.panes.len and p.panes[h] == pane) p.cancelLookHover(); p.pane_alloc.doom(pane); @@ -6200,7 +6243,7 @@ pub const Pardes = struct { pub fn postEvent(p: *Pardes, ev: Event) void { switch (ev) { .key => |k| std.debug.assert(k.text.len == 0), - .output, .paste, .lsp_resp, .pipe_resp, .file_changed, .command => unreachable, + .output, .paste, .lsp_resp, .pipe_resp, .file_changed, .command, .fs_req => unreachable, else => {}, } if (p.in_len == p.in_q.len) return; @@ -6238,6 +6281,27 @@ pub const Pardes = struct { return pane.pdfPath(); } + /// THE BYTES BEHIND AN `.fs_reply`, resolved in the drain. A filesystem + /// read answers with either something the handler formatted (staged in the + /// core, valid until the next request) or a window onto a pane's live text, + /// which is handed over WITHOUT A COPY — the same trick, and the same + /// serial check, `.save_text` uses to write a megabyte it never duplicated. + /// A slot reused between the answer and this call resolves to nothing + /// rather than to another pane's text. + pub fn fsPayload(p: *const Pardes, r: acmefs.Reply) []const u8 { + return switch (r.payload) { + .none => &.{}, + .staged => |n| p.fs.out.items[0..@min(n, p.fs.out.items.len)], + .region => |g| region: { + const pane = p.panes[g.pane] orelse break :region &.{}; + if (pane.serial != g.serial) break :region &.{}; + const text = if (pane.file) |*f| f.content else break :region &.{}; + const lo = @min(g.off, text.len); + break :region text[lo..@min(lo + g.len, text.len)]; + }, + }; + } + /// Perform one effect through the host, falling back per METHOD (not per /// host) to the in-process implementation. This is the switch that used to /// be copied into all four shells. @@ -6306,6 +6370,11 @@ pub const Pardes = struct { }, .theme_file => |t| if (v.push_watch_theme) |f| f(p.host.ctx, t.generation, t.on), .dump_themes => |d| if (v.push_dump_themes) |f| f(p.host.ctx, d.pane), + // The bytes are read off the core HERE, in the drain, exactly as + // save_file reads a file pane: the reply named where they live and + // this is the borrow window. A host with no filesystem serving + // cannot have asked, so a null method is not a dropped answer. + .fs_reply => |r| if (v.push_fs_reply) |f| f(p.host.ctx, &r, p.fsPayload(r)), // the loop's own condition; a host tears down after its own loop .quit => p.quit = true, } @@ -6399,6 +6468,15 @@ pub const Pardes = struct { }, else => {}, } + // Which input this update IS, in acme's origin alphabet, so every + // event record the handlers below produce is attributed without any + // of them being told: `K` for the keyboard, `M` for the mouse. A + // filesystem write says `E`/`F` for itself (see acmefs). + if (p.fs.listeners != 0) p.fs.origin = switch (ev) { + .key => 'K', + .mouse => 'M', + else => p.fs.origin, + }; switch (ev) { .resize => |sz| { p.snap_panel_layout_once = true; @@ -6463,6 +6541,10 @@ pub const Pardes = struct { }, .paste => |bytes| p.applyPaste(bytes), .command => |line| _ = p.executeBuiltinLine(p.active, line), + // One filesystem request in, one answer out, in this update. The + // whole of the concurrency is that the transport asked from the + // loop thread; see acmefs.zig's header. + .fs_req => |r| p.emit(.{ .fs_reply = acmefs.handle(p, r) }), .pinch => |scale| p.ov_pinch_scale = scale, .touch_scroll => |delta| p.ov_touch_scroll_delta = delta, .pointer_leave => p.pointer_inside = false, @@ -6474,6 +6556,34 @@ pub const Pardes = struct { } _ = p.scratch.reset(.retain_capacity); p.sync(); + p.fsReport(); + } + + /// TAG EDITS, which no single call site owns: a tag is assembled from a + /// live prefix and an editable tail by half a dozen paths (typing, a prompt + /// arming, a Save clearing the dirty marker, a shell reporting a new cwd), + /// so it is diffed at the END of an update, where it is finally settled. + /// acme can hook `textinsert` on the tag itself because its tag IS a text + /// buffer; pardes's is a rendering, so the diff is the honest equivalent. + /// + /// Costs nothing when nobody is listening: one branch, and the snapshots + /// are only allocated for panes a script has opened. + fn fsReport(p: *Pardes) void { + if (p.fs.listeners == 0) return; + for (p.panes, 0..) |slot, id| { + const pane = slot orelse continue; + if (!p.fs.scripted(id)) continue; + const tag = p.tagText(p.scratch.allocator(), pane) catch continue; + const snap = &p.fs.panes[id].tag_snap; + if (std.mem.eql(u8, snap.items, tag)) continue; + // First sight of a tag is not an edit: the script just opened the + // file and can read `tag` for itself. `release` drops the snapshot + // with the last reader, so this stays true across re-opens. + if (snap.capacity != 0 or snap.items.len != 0) + acmefs.noteReplace(p, id, true, snap.items, tag); + snap.clearRetainingCapacity(); + snap.appendSlice(p.gpa, tag) catch {}; + } } /// THE DEFAULT REGISTER, and nothing else. helix: an ordinary `y`/`d`/`c` @@ -6629,7 +6739,7 @@ pub const Pardes = struct { /// retained through the frame and a cwd can be rewritten under us by the /// next shell report, so this owns its bytes rather than lending the /// pane's. - fn tagPrefix(p: *Pardes, pane: *Pane) ![]u8 { + pub fn tagPrefix(p: *Pardes, pane: *Pane) ![]u8 { const arena = p.scratch.allocator(); if (comptime pdf_enabled) if (pane.pdf) |pv| return std.fmt.allocPrint( arena, @@ -6750,7 +6860,7 @@ pub const Pardes = struct { /// the tag exactly as it is rendered: prefix ++ gap ++ tail. THE text /// tag_col and tag_anchor index, so the renderer, the mouse, the motions /// and the chord all read the same bytes at the same columns. - fn tagText(p: *Pardes, arena: std.mem.Allocator, pane: *Pane) ![]u8 { + pub fn tagText(p: *Pardes, arena: std.mem.Allocator, pane: *Pane) ![]u8 { const prefix = try p.tagPrefix(pane); const tail = curTail(pane); const gap = p.tagGap(pane, file_pane.displayWidth(prefix) + file_pane.displayWidth(tail)); @@ -6764,7 +6874,7 @@ pub const Pardes = struct { /// Take the laid-out tail into the pane's own buffer, once, on first touch. /// The gap comes along as ordinary characters — that is what hands the /// padding to you to edit, and what freezes it against reflow from here on. - fn seedTail(p: *Pardes, pane: *Pane) void { + pub fn seedTail(p: *Pardes, pane: *Pane) void { if (pane.tag_init) return; const tail = curTail(pane); const prefix = (p.tagPrefix(pane) catch return); @@ -12240,16 +12350,109 @@ pub const Pardes = struct { } } + /// THE ACME INVERSION: while a script holds this pane's `event` file open, + /// buttons 2 and 3 in it belong to the script. The words in its tag are + /// ITS commands — `Step`, `Run`, `Clear` — and pardes has never heard of + /// them, so it reports the click and performs nothing. Returns true when + /// the caller must not run the builtin. + /// + /// Keyboard Enter/Tab are deliberately NOT suppressed, unlike acme (which + /// has no keyboard equivalent to suppress): a scripted pane stays + /// editable, and a script that dies mid-run cannot leave you unable to + /// execute anything in it. + fn reportGesture( + p: *Pardes, + id: usize, + cmd: Builtin, + text: []const u8, + on_tag: bool, + operand: PointerOperand, + chorded: bool, + ) bool { + if (!p.fs.scripted(id)) return false; + const is_look = cmd == config.look_cmd; + const action: acmefs.Action = if (is_look) + (if (on_tag) .tag_look else .body_look) + else + (if (on_tag) .tag_exec else .body_exec); + const named = std.meta.stringToEnum(Builtin, commandText(text)) != null; + const range = p.gestureRange(id, text, on_tag, operand); + // acme(4)'s two flag vocabularies at the same bit positions. For an + // exec, bit 1 is "this is a builtin"; for a look, it is "pardes can + // act on this without loading a file", which is the same fact plus a + // word that is not a path. Bit 2 is acme's "the text indicated is a + // null string that has a non-null expansion", which is exactly what an + // empty range plus text means here. + var flag: u32 = if (named) acmefs.flag_builtin else 0; + if (range.q0 == range.q1 and text.len > 0) flag |= acmefs.flag_expansion; + if (is_look) { + if (!named and std.mem.indexOfAny(u8, text, "/.:") != null) flag |= acmefs.flag_filename; + } else if (chorded) flag |= acmefs.flag_chorded; + return acmefs.noteAction(p, id, action, range.q0, range.q1, flag, text); + } + + /// THE BYTE RANGE A GESTURE NAMES, in the coordinates `addr` and `data` + /// speak — and it must name the TEXT REPORTED WITH IT, because a script + /// that does not recognise a record writes it back and pardes then + /// re-derives the text from these two numbers. A range that started at the + /// pointer instead of at the operand would execute `l D` for a click on + /// the `l` of `Del`. + /// + /// So the start comes from whatever the operand actually resolved: an + /// expanded word's own column, or the leading corner of the selection it + /// reused. When neither is available — a modal `v`/`x` selection, a + /// terminal's projected screen, a PDF — the answer is acme's null range at + /// the click, which its flag bit 2 already has a meaning for: the text + /// travels, the range does not claim to be it. + fn gestureRange(p: *Pardes, id: usize, text: []const u8, on_tag: bool, operand: PointerOperand) acmefs.PaneFs.Range { + const pane = p.panes[id] orelse return .{}; + if (on_tag) { + const tag = p.tagText(p.scratch.allocator(), pane) catch return .{}; + const sel = operand.expanded orelse operand.preview orelse return .{}; + const lead = @min(sel.c0, sel.c1); + const at = file_pane.rawAtDisplay(tag, @intCast(@max(0, lead))); + const q0: u32 = @intCast(@min(at, tag.len)); + return .{ .q0 = q0, .q1 = @intCast(@min(q0 + text.len, tag.len)) }; + } + const f = if (pane.file) |*file| file else return .{}; + const start: ?modal.Cursor = if (operand.file_word) |w| + .{ .row = @intCast(@max(0, w.row)), .col = @intCast(@max(0, w.lo)) } + else if (operand.expanded orelse operand.preview) |sel| lead: { + // A selection's leading corner in reading order, converted the way + // `pointerOperand` converts the click itself. + const top = @min(sel.r0, sel.r1) - @as(i32, BOX_H); + if (top < 0) break :lead null; + const w = pane.wrapAt(top); + const col = p.paneByteAtDisplay(pane, w.line, w.at, @min(sel.c0, sel.c1) - config.PREFIX_W); + break :lead .{ + .row = @intCast(@max(0, w.line)), + .col = @intCast(@max(0, col)), + }; + } else null; + const cursor = start orelse return .{}; + const q0: u32 = @intCast(@min(modal.hxOff(f.content, cursor), f.content.len)); + return .{ .q0 = q0, .q1 = @intCast(@min(q0 + text.len, f.content.len)) }; + } + fn dispatchPointerBuiltin( p: *Pardes, id: usize, cmd: Builtin, text: ?[]const u8, + gesture: ?struct { on_tag: bool, operand: PointerOperand }, ) void { const arg = p.chord_arg; p.chord_arg = null; defer if (arg) |a| p.gpa.free(a); - const operand = text orelse return; + const operand = text orelse { + // Nothing expanded — a click on `#`, `*`, `|`, a blank cell. A + // WATCHED pane still owns it: the script decides what a character + // pardes has no word for means, and life.py's grid is made of + // exactly those characters. + if (gesture) |g| _ = p.reportGesture(id, cmd, "", g.on_tag, g.operand, arg != null); + return; + }; + if (gesture) |g| if (p.reportGesture(id, cmd, operand, g.on_tag, g.operand, arg != null)) return; p.runBuiltin(cmd, id, "", p.withArg(operand, arg)); } @@ -12307,10 +12510,13 @@ pub const Pardes = struct { s.chorded, ); defer release.deinit(p.pdf_gpa); + // A native PDF page has no byte offsets of its own to + // report, so a scripted pane cannot intercept this one. if (release.action) |action| p.dispatchPointerBuiltin( s.id, if (action == .look) config.look_cmd else config.exec_cmd, release.text, + null, ); return; } @@ -12383,7 +12589,10 @@ pub const Pardes = struct { } const txt = operand.text; const cmd = if (s.button == config.look_button) config.look_cmd else config.exec_cmd; - p.dispatchPointerBuiltin(s.id, cmd, txt); + p.dispatchPointerBuiltin(s.id, cmd, txt, .{ + .on_tag = clk.r0 < BOX_H, + .operand = operand, + }); } }, .tag => {}, // dragUpdate already left the tag cursor + selection set diff --git a/src/tty/tty.zig b/src/tty/tty.zig index 404fe582..c5abf0fe 100644 --- a/src/tty/tty.zig +++ b/src/tty/tty.zig @@ -17,6 +17,8 @@ const file_watch = @import("../file_watch.zig"); const user_config = @import("../user_config.zig"); const selection_pipe = @import("../selection_pipe.zig"); const nested = @import("../nested.zig"); +const fuse = @import("../fuse.zig"); +const fs_service = @import("../fs_service.zig"); const panel_compositor = @import("panel_compositor.zig"); const host_api = @import("../host.zig"); @@ -61,6 +63,14 @@ pub const Command = struct { /// a pardes launched inside this one sent us a builtin command line /// (see lookServer); gpa-owned, like pty_read command: []u8, + /// `--fs`: the /dev/fuse descriptor has requests on it. Carries + /// nothing and is applied as a no-op — its whole job is to end the + /// blocking `nextEvent`, because the drain itself lives in pollFrame + /// beside the file-watch reload. Same shape and same reason as + /// `files_changed` above, and posted from two places: the poll thread + /// when the kernel makes the descriptor readable, and pollFrame itself + /// when a batch hit its cap with requests still pending. + fs_ready, } = .nop; }; const Loop = vaxis.Loop(@TypeOf(Command.value)); @@ -521,6 +531,24 @@ pub fn run(init: std.process.Init, opts: pardes.Options) !void { defer paste_buf.deinit(); var loop: Loop = .init(io, &tty, &vx); + // `--fs`: mount before the initial spawns, because those shells are the + // ones that need PARDES_FS in their environment, and before the first + // frame, because a script racing startup must find panes that are already + // there. Also before any thread of ours exists — the mount forks the + // setuid fusermount3 helper, and forking from a multithreaded process is + // the hazard this whole region is ordered around. Null covers both "no + // --fs" and "--fs but the mount failed"; the second is reported on a + // message row inside `start` and the session runs on regardless. + // + // The teardown answers every held request, aborts the connection, + // unmounts and removes `<parent>/<pid>`. The PARENT (`.../pardes`) stays, + // like nested.zig's socket directory: another session may be living in it, + // and rmdir of a shared directory is not ours to attempt. + var fs = fs_service.start(gpa, core); + // Covers the error paths only: the ordinary exit unmounts at the END OF + // THE LOOP instead, see there. + defer if (fs) |f| f.deinit(); + var sh: Shell = .{ .io = io, .gpa = gpa, @@ -538,6 +566,7 @@ pub fn run(init: std.process.Init, opts: pardes.Options) !void { // mark the file a positional path argument opened. -1 off linux: // watchPane goes quiet and the core simply never gets a file_changed. .inotify_fd = if (builtin.os.tag == .linux) libc.inotify_init1(linux.IN.CLOEXEC) else -1, + .fs = fs, }; defer { // reap the reader tasks (cancel interrupts a blocked read) before @@ -604,6 +633,11 @@ pub fn run(init: std.process.Init, opts: pardes.Options) !void { // blocking accept(2) never returns either, so an io.concurrent task would // hang the teardown that joins it. if (sock_fd >= 0) (try std.Thread.spawn(.{}, lookServer, .{ gpa, sock_fd, &loop })).detach(); + // ...and the /dev/fuse poller, which is the same kind of thread again: it + // waits for POLLIN and posts, never touching the core or the descriptor's + // data. Joined by `Fs.deinit` rather than detached, because unlike accept4 + // it CAN be woken — fuse.zig gives it a control pipe for exactly that. + fs_service.wake(fs, &loop, wakeFs); // Capability handshake — SEND the probes, do not wait on them. This was // queryTerminal(2ms), which blocks on a futex until DA1 comes back. The // number has to beat one terminal round trip: a local terminal answers in @@ -705,6 +739,18 @@ pub fn run(init: std.process.Init, opts: pardes.Options) !void { } } } + // THE FILESYSTEM GOES FIRST, ahead of every deferred teardown below. + // `loop.stop()` joins a reader parked in `read(2)` on the tty, so it does + // not return until the next keystroke — and a session that has decided to + // exit must not spend that wait holding a mount nobody is serving. A + // client blocked on `<id>/event` when the last pane is deleted through + // `ctl` then gets ENOTCONN at once instead of hanging until somebody + // touches the keyboard. + if (fs) |f| { + f.deinit(); + fs = null; + sh.fs = null; + } } /// Everything the terminal shell owns and the core cannot: the ptys, the @@ -727,6 +773,10 @@ const Shell = struct { /// it, and the markers are the only thing that says where it ends. paste_buf: *std.Io.Writer.Allocating, inotify_fd: c_int, + /// acme's control filesystem for this session, or null when `--fs` was not + /// given (or its mount failed). Owned by `run`, which mounts it before the + /// first fork and tears it down on every path out. + fs: ?*fuse.Fs = null, ptys: [pardes.MAX_PANES]?Pty = @splat(null), /// per-slot spawn generation: a reused pane id ignores the old shell's /// late pty_eof (which would otherwise close the NEW pty on that slot) @@ -790,6 +840,7 @@ const Shell = struct { .push_open_link = openLink, .pull_lsp = lsp, .pull_pipe = pipe, + .push_fs_reply = fsReply, }; // ---- input ------------------------------------------------------------ @@ -960,6 +1011,12 @@ const Shell = struct { // `git checkout`) collapses into ONE pass below, so it cannot // queue a reload — or an undo entry — per write. .files_changed => s.check_files = true, + // A wake and nothing more. The requests behind it are drained in + // pollFrame, where the file-watch reload also happens: both want to + // run once per frame with the whole batch already in, not once per + // event. So this arm has nothing to do — which is the point, since + // its only job was ending the blocking wait above. + .fs_ready => {}, .lsp_done => |d| { core.update(.{ .lsp_resp = .{ .id = d.id, .rows = d.rows } }); s.lsp_gpa.free(d.rows); @@ -995,6 +1052,18 @@ const Shell = struct { fn pollFrame(ctx: ?*anyopaque) void { const s = of(ctx); + // acme's filesystem, answered here for the same reason the file-watch + // reload is (see reloadWatched): one batch per frame, not one frame per + // request. First in the pass, so an edit a script just made through + // `body` is in the surface this frame composes rather than the next. + if (s.fs) |f| if (fs_service.drain(f, s.core).pending) { + // The batch hit its cap with requests still pending, and no ack has + // gone to the poll thread — so nothing else will wake us. Re-arm + // the loop ourselves. tryPostEvent, not postEvent: this runs on the + // only thread that drains the queue, so blocking on a full one + // would deadlock, and a full queue already holds a wake. + _ = s.loop.tryPostEvent(.fs_ready) catch {}; + }; // live cwd for tags/look: cheap per-pane lookup, per frame. Whether the // pane's tty still belongs to the prompt we forked is NOT polled here — // it is a walk through /proc and nothing draws it, so the core pulls it @@ -1156,7 +1225,7 @@ const Shell = struct { cwd_buf[cwd.len] = 0; cwd_z = @ptrCast(&cwd_buf); } - const child = forkShell(s.core, pane, s.prompt_rcs, s.core.shellBin(), cwd_z, s.core.screen_h, s.core.screen_w); + const child = forkShell(s.core, pane, s.prompt_rcs, s.core.shellBin(), cwd_z, s.core.screen_h, s.core.screen_w, s.fs); s.ptys[pane] = .{ .file = child.file, .pid = child.pid, .reader = .{ .any_future = null, .result = {} } }; // report the pane's starting directory back to the core (tags). The // slot needs no occupancy reset: nothing is remembered, and the next @@ -1246,6 +1315,17 @@ const Shell = struct { s.loop.postEvent(.files_changed) catch {}; } + /// The core's answer to one filesystem request, handed straight back to the + /// transport that is holding it. `bytes` was resolved by `pardes.fsPayload` + /// inside `perform` and is borrowed only for this call — a body read is a + /// window onto the pane's live text, so there is nothing to copy and + /// nothing to free. `.again` needs no special case here: `Fs.reply` reads + /// the status and re-parks the request itself. + fn fsReply(ctx: ?*anyopaque, reply: *const pardes.acmefs.Reply, bytes: []const u8) void { + const s = of(ctx); + if (s.fs) |f| f.reply(reply, bytes); + } + fn dumpThemes(ctx: ?*anyopaque, pane: u8) void { const s = of(ctx); const config_dir = s.core.opts.config_dir orelse return; @@ -1454,12 +1534,27 @@ fn winchWatch(loop: *Loop, vx: *vaxis.Vaxis, tty: *vaxis.Tty) void { } } -fn forkShell(core: *pardes.Pardes, pane: usize, prompt_rcs: *const shell_bin.PromptRcs, bin: []const u8, cwd: ?[*:0]const u8, rows: u16, cols: u16) struct { file: std.Io.File, pid: posix.pid_t } { +/// The /dev/fuse poller's wake, and deliberately nothing else — the thread that +/// calls this has no business in the core, so all it does is end the blocking +/// `nextEvent`. tryPostEvent rather than postEvent for the same reason readPty's +/// final post uses it: this can fire after the loop has already been left, and a +/// blocking push into a full queue nobody is draining would never return. +fn wakeFs(ctx: ?*anyopaque) void { + const loop: *Loop = @ptrCast(@alignCast(ctx.?)); + _ = loop.tryPostEvent(.fs_ready) catch {}; +} + +fn forkShell(core: *pardes.Pardes, pane: usize, prompt_rcs: *const shell_bin.PromptRcs, bin: []const u8, cwd: ?[*:0]const u8, rows: u16, cols: u16, fs: ?*const fuse.Fs) struct { file: std.Io.File, pid: posix.pid_t } { var master: c_int = undefined; // resolved BEFORE the fork, into this frame, which the child inherits: // nothing between fork and exec may allocate, and a PATH search would var path_buf: [std.fs.max_path_bytes]u8 = undefined; const spawn = shell_bin.resolve(bin, &path_buf, prompt_rcs); + // ...and so is the pane's own address on the control filesystem, for a + // second reason on top of that one: acme puts `winid` in the child, which + // is safe there only because rfork(RFENVG) has just given it a private + // environment group. See fs_service.exportPaneEnv. + fs_service.exportPaneEnv(fs, if (core.panes[pane]) |pn| pn.serial else 0); const ws = posix.winsize{ .row = rows, .col = cols, .xpixel = 0, .ypixel = 0 }; const pid = forkpty(&master, null, null, &ws); if (pid == 0) { diff --git a/src/tutor.txt b/src/tutor.txt index 7dfac0bf..f75734e4 100644 --- a/src/tutor.txt +++ b/src/tutor.txt @@ -1179,6 +1179,40 @@ abc special case.) ================================================================= += PART 5 — THE CONTROL FILESYSTEM (pardes --fs) = +================================================================= + + Started as `pardes --fs`, this session serves plan9 acme's control + files over FUSE: a directory per pane, holding `body`, `tag`, `ctl`, + `addr`, `data`, `event` and the rest, plus `index`, `cons` and `new/` + at the top. Every pane shell gets `$PARDES_FS` (where it is mounted) + and `$PARDES_PANE` (which pane it is), so a script run in a pane can + talk to its own window with no arguments: + + echo hello >> $PARDES_FS/$PARDES_PANE/body append to this pane + cat $PARDES_FS/index one line per pane + echo hi > $PARDES_FS/new/body open a pane holding "hi" + + That IS the extension mechanism: no plugin API, no interpreter, no + rebuild — a program that opens files. Two things about it are worth + knowing before you use it. + + While a program holds a pane's `event` file open, MIDDLE AND RIGHT + CLICKS IN THAT PANE BELONG TO IT: pardes reports them and performs + nothing, so the words in that pane's tag can be the program's own + commands (examples/acmefs/life.py puts Step/Run/Clear there). The + keyboard is not suppressed, so the pane stays editable either way, and + when the program exits the clicks come back. + + And writing an event record back — `origin type q0 q1` — makes pardes + perform the Look or Exec it names. That is arbitrary command execution + by design, the same door acme has always had, which is why the mount + lives under $XDG_RUNTIME_DIR at 0700 and why `--fs` is opt-in. + + examples/README.md is the client's guide; docs/acme-fs.md compares this + implementation with acme's, file by file. + +================================================================= = SUMMARY = ================================================================= @@ -1198,6 +1232,10 @@ abc landing its cursor where you navigated (prompt start if the input is empty, mid-text otherwise) Filter in its tag toggles a pane-local, theme-keyed palette + --fs: acme's control files over FUSE: $PARDES_FS/<id>/{body,tag, + ctl,addr,data,event,...} and $PARDES_FS/{index,cons,new} + a script holding a pane's `event` open owns that pane's + middle and right clicks; writing a record back runs it KEYS (helix): h j k l w b e 0 $ ^ gg ge f F t T v x X d c y p i a I A o O u undo U redo Enter=look Tab=execute w b e f F t T SELECT what they cross, so `i` after one |
