diff options
Diffstat (limited to 'src/detached/server.zig')
| -rw-r--r-- | src/detached/server.zig | 1169 |
1 files changed, 929 insertions, 240 deletions
diff --git a/src/detached/server.zig b/src/detached/server.zig index 6d13a4ae..93c964b4 100644 --- a/src/detached/server.zig +++ b/src/detached/server.zig @@ -1,53 +1,96 @@ //! THE DETACHED CORE: one `Pardes` instance in a process with no terminal, //! serving N frontends over one unix socket. //! -//! THIS SIDE OWNS THE CORE. `Session` is a `host.Host` implementation whose -//! methods encode wire messages instead of doing IO, and whose -//! `pull_wait_input` is a `poll(2)` over the listener and every attached -//! frontend. The frontends own terminals and nothing else (client.zig). So the -//! `Pardes` is here, `update` is called from here, and the same screen is on -//! every attached frontend at once — `screen -x`, not N sessions. +//! THIS SIDE OWNS THE CORE, AND EVERYTHING UNDER IT. `Session` is a `host.Host` +//! implementation that performs the machine-local half of a host itself — it +//! forks the pane shells, writes the files, watches the paths — and whose +//! `pull_wait_input` is one `poll(2)` over the listener, every attached +//! frontend, every pane's pty master and the inotify descriptor. A frontend owns +//! a screen and a keyboard and nothing else (client.zig). So the `Pardes` is +//! here, `update` is called from here, the shells are forked from here, and the +//! same screen is on every attached frontend at once — `screen -x`, not N +//! sessions. //! -//! WHAT THIS SIDE SERVES ITSELF. Every method this vtable leaves null falls -//! through to the core's own `host.Fallback`: the embedded source filesystem, -//! the in-process clipboard, silent ptys. host.zig says in as many words that a -//! zero-method host is a complete pardes, and that is exactly what a session -//! with nothing attached is. Everything a real frontend can do BETTER — fork a -//! shell on a real tty, put bytes on a real disk, reach a real desktop -//! clipboard — is asked of a frontend, and the routing table below says which. +//! WHAT THIS SIDE SERVES ITSELF, WHICH IS NOW ALL OF IT. A unix socket means +//! the core and its frontends are on the SAME machine, so there is no question +//! of whose process table, whose disk or whose inotify descriptor a call is +//! about — and given that, the process that must hold them is the long-lived +//! one. A shell forked by a frontend dies with that frontend, and a session +//! whose whole promise is outliving the frontend attached to it cannot keep its +//! panes that way. So this file forks the pane shells (`host_io.forkShell`), +//! writes the files (`host_io.writeFileBytes`), marks the directories +//! (file_watch.zig) and drains the pty masters in its own `poll(2)`. THE PANE +//! SHELLS OUTLIVE EVERY FRONTEND: attach, detach, kill the terminal, attach +//! from another one, and the build that was running in pane 3 is still running +//! and has been scrolling into the core the whole time. +//! +//! Only what this vtable leaves null falls through to the core's own +//! `host.Fallback` — and host.zig says in as many words that a zero-method host +//! is a complete pardes. What a frontend can still do BETTER is exactly what +//! needs the human's own display, and nothing else: put a yank on the clipboard +//! in front of them, take a paste off it, open a link in their browser. Three +//! messages, which is why the routing table below is as short as it is. //! //! ROUTING, and it is not "push means broadcast". A push reaches every HOST //! (host.zig's rule, which `Fanout.isPull` enforces); this is ONE host that //! happens to be backed by several frontends, and how it spreads a call inside -//! itself is its own business. Three rules, one per kind of side effect: +//! itself is its own business. Two rules over four messages: //! * BROADCAST — the frame, and `set_clipboard`. Every screen must show the //! same thing, and a yank in a shared session is a session-wide fact that //! every attached desktop is entitled to. -//! * PRIMARY ONLY — `spawn`, `pty_write`, `pty_resize`, `write_file`, -//! `write_dump`, `watch_file`, `watch_theme`, `dump_themes`. Each of these -//! has ONE real resource behind it, and doing it twice is not doing it -//! twice as well: two frontends forking a shell for pane 3 gives the pane -//! two shells, and two frontends writing one path race each other. Primary -//! is the lowest attached slot, i.e. the oldest surviving attachment — a -//! rule that is stable while frontends come and go and needs no election. -//! A pane's shell therefore lives in the frontend that forked it: when that -//! frontend leaves, its panes stop producing output and the session's text, -//! files and layout carry on. That is a real limit and it is stated here -//! rather than papered over, because migrating a live pty between processes -//! is a different feature. -//! * ORIGIN, ELSE PRIMARY — `read_clipboard` (the one `pull_` on the wire) -//! and `open_link`. Both answer a thing a HUMAN just did, and the answer -//! belongs on that human's machine: the paste must come from the keyboard -//! that asked for it, and a link must open in front of the person who -//! clicked it. `origin` is the frontend whose event was applied most -//! recently. Effects drain after a whole batch of events (pardes.zig +//! * ORIGIN, ELSE PRIMARY — `read_clipboard` (the one `pull_` on the wire), +//! `open_link`, and `detach`. Each answers a thing a HUMAN just did, and the +//! answer belongs to that human: the paste must come from the keyboard that +//! asked for it, a link must open in front of the person who clicked it, and +//! a `Detach` typed in one frontend must send THAT frontend away and leave +//! the others painting. `origin` is the frontend whose event was applied +//! most recently. Effects drain after a whole batch of events (pardes.zig //! `pump`), so in the rare case where two frontends type in the same -//! millisecond the second one wins; the fallback to primary covers an +//! millisecond the second one wins; the fallback to `primary` — the lowest +//! attached slot, i.e. the oldest surviving attachment, a rule that is +//! stable while frontends come and go and needs no election — covers an //! effect that no input caused at all. //! -//! FAIRNESS, and why no client can stall the core or another client: -//! * every descriptor is non-blocking, and there is no thread per client. One -//! `poll(2)` per pump covers the listener and all `max_clients` frontends. +//! `detach` is the odd one and is worth naming as such: it is not an EFFECT the +//! session performs on the world, it is SESSION CONTROL — one frontend asking to +//! stop being a frontend. That is why wire.zig gives it 0x05, in the +//! 0x01..0x0f session range beside `quit`, rather than a number in the 0x10.. +//! range where every tag is one `push_` method that reaches a disk, a clipboard +//! or a browser. And it is why this side does nothing but send it: see `detach`. +//! There is no third rule, and the class of message it used to serve is gone: +//! `spawn`, `pty_write`, `pty_resize`, `write_file`, `write_dump`, `watch_file`, +//! `watch_theme` and `dump_themes` were routed to ONE frontend precisely +//! because each has one real resource behind it, and every one of them is now +//! performed HERE, once, by the process that owns the resource. Two frontends +//! can no longer fork two shells for pane 3 or race each other writing one +//! path, because neither of them writes anything. +//! +//! FAIRNESS, and why no client — and no shell — can stall the core or another +//! client. The property the bullets below add up to is worth stating as one +//! sentence, because it is what a detached session is FOR: there is no path on +//! which this process blocks indefinitely. Every descriptor it holds is +//! non-blocking, the single `poll(2)` is the only place it sleeps, and every +//! queue that could grow without bound has a ceiling with a stated answer for +//! reaching it. A daemon nobody is looking at cannot be made to stop looking +//! after the shells nobody else is keeping. +//! * ONE `poll(2)` per pump covers the listener, all `max_clients` frontends, +//! all `pardes.MAX_PANES` pty masters and the inotify descriptor: +//! `poll_slots` descriptors, one syscall, no thread per client and none per +//! pty. Putting the shells in the poll set the clients were already in is +//! what lets a daemon own sixteen of them and stay single-threaded. +//! * one read per pty per round, which is `receive`'s rule for clients +//! applied to shells: a `yes` in pane 1 gets one turn and the loop moves on +//! to the other panes, the frontends and the frame. +//! * EVERY descriptor is non-blocking, sockets and pty masters alike, and a +//! pane owes its bytes the same way a client does. A blocking write to a +//! master was the one hole this file's own comment used to argue was safe — +//! "the peer on a pty is a shell this process forked, not a stranger who can +//! stop reading on purpose" — and that was wrong, because the peer is +//! whatever program the human ran in that pane. `sleep 3600` plus a paste +//! larger than the pty's input buffer parked the WHOLE daemon inside +//! `write(2)`: no frame to any frontend, fifteen other masters unread, no +//! `accept`, no `expire`, no inotify drain. So a pane has an out-queue and a +//! POLLOUT, on the descriptor that was already in the set. See `ptyWrite`. //! * FRAMES ARE NOT QUEUED. A client with bytes still owed to the kernel is //! SKIPPED for this frame and its mirror is left alone, so the next frame //! it does get is a diff against what it actually has. A slow frontend @@ -96,11 +139,48 @@ //! collects, so the client checks the directory and the socket before it //! connects, exactly as this side checks them before it binds. const std = @import("std"); +const builtin = @import("builtin"); const libc = std.c; +const linux = std.os.linux; +const posix = std.posix; const pardes = @import("../pardes.zig"); const host_api = @import("../host.zig"); const wire = @import("wire.zig"); +/// The machine-local half of a host — fork a shell onto a pty, put bytes on a +/// disk — shared verbatim with the tty shell, and the sharing is the point: +/// `spawn` and `writeFile` below are the same two operations tty.zig performs, +/// and having them in one file is what keeps a daemon's pane and a terminal's +/// pane the same pane. See host_io.zig's header for why the daemon is the side +/// that performs them. +const host_io = @import("../host_io.zig"); + +/// ...and the inotify half, likewise shared: `applyEffect` is the mark-then- +/// reconcile transaction the tty and sdl shells run, and this session runs the +/// identical one. All that differs is who waits on the descriptor — a thread +/// there, `waitInput`'s poll set here. +const file_watch = @import("../file_watch.zig"); + +/// Host-lifetime storage for the OSC 133 rc files a forked shell sources, held +/// by `Session` because a Session is exactly one host's lifetime. +const shell_bin = @import("../shell_bin.zig"); + +/// `shellCwd` for a pane's shell, `ttyTaken` for a pane the core is about to +/// type a command line into, `readFile` for `run`'s `--load`. +const look = @import("../look.zig"); + +/// The "saved <path>" / "dumped themes <path>" message row, stamped the way +/// every other host stamps it — one clock format across every frontend. +const message = @import("../message.zig"); + +/// `Dump themes` writes the reference set out as .zon, into the core's own +/// `opts.config_dir`. +const user_config = @import("../user_config.zig"); + +// TIOCSWINSZ: absent from std.c.T on darwin — _IOW('t', 103, winsize). The +// same constant the tty, gui and macos shells spell, for the same reason. +const TIOCSWINSZ: c_int = @bitCast(@as(u32, if (@hasDecl(posix.T, "IOCSWINSZ")) posix.T.IOCSWINSZ else 0x80087467)); + /// Diagnostics for whoever is running the daemon. Every one of these is a /// `debug`, and the level is not a judgement about how bad the thing is: /// main.zig's logFn drops this scope entirely unless PARDES_LOG is set, so what @@ -136,11 +216,11 @@ const sun_path_len = nested.sun_path_len; pub const max_clients = 32; /// Bytes of un-drained CONTROL messages a client may owe before it is closed. -/// Frames are not in here (see the module header), so this is a backlog of -/// ACTIONS — spawns, clipboard mirrors, file writes — and a frontend that has -/// not taken 1 MiB of those has stopped reading its socket. Checked before an -/// append rather than after, so one oversized message is never the thing that -/// trips it. +/// Frames are not in here (see the module header), so this bounds a backlog of +/// the three things that are still on the wire — a welcome, a clipboard mirror, +/// a link to open — and a frontend that has not taken 1 MiB of those has +/// stopped reading its socket. Checked before an append rather than after, so +/// one oversized message is never the thing that trips it. const out_backlog = 1 << 20; /// One read per client per poll round (see `receive`). 16 KiB is two orders of @@ -148,16 +228,73 @@ const out_backlog = 1 << 20; /// 4 MiB paste arrives across several rounds, which is the point. const read_chunk = 16 * 1024; +/// Bytes taken off one pane's pty per poll round. 64 KiB is what every other +/// host's pty reader uses (`readPty` in tty.zig, gui.zig and macos.zig), and it +/// sits on `readPty`'s own frame rather than the loop's. Nothing is copied out +/// of it: `Event.output` borrows the buffer for one `update` call, so a daemon +/// serving a shell that is printing a build log asks the allocator for nothing. +const pty_chunk = 64 * 1024; + +/// Descriptors in the ONE poll this process runs: the listener, every frontend, +/// every pane's pty master, and the inotify descriptor behind every watch. 50 +/// on a full house, and one syscall covers all of them. +const poll_slots = 1 + max_clients + pardes.MAX_PANES + 1; + +/// How many times `reloadWatched` will honour `file_watch.reloadChanged`'s +/// request for another pass within one round. See `reloadWatched`. +const reload_retries = 4; + /// Bytes of client traffic — every in-queue and out-queue together — this /// session may hold before it starts closing the peers holding it. /// `out_backlog` bounds ONE slot and this bounds the table, which is not the -/// same ceiling: 32 clients each a byte under their own cap is 32 MiB of a -/// daemon nobody is looking at. 4 MiB is one whole paste in flight plus every -/// frame queue a real session builds, and past it the fattest peer is the peer -/// that stopped reading. The mirrors are NOT in this number: a mirror is this -/// session's own bookkeeping for a client it chose to serve, not something a -/// peer can grow. -const session_backlog = 4 << 20; +/// same ceiling: a client's `out` tops out at `out_backlog` plus the one +/// oversized message allowed through whole, so `out_backlog` alone permits +/// 32 * (1 + 1.6) MiB, about 83 MiB of a daemon nobody is looking at. +/// +/// DERIVED, and the derivation IS the fix. This was the literal `4 << 20`, +/// which was by coincidence the exact value of tty.zig's `max_paste_bytes` — +/// and `in` grows to hold one WHOLE message, so a frontend assembling the very +/// paste wire.zig names as one of the two messages that set `max_payload` +/// crossed the table's ceiling while still receiving it. The session then +/// closed the only frontend it had, mid-paste, with `.backlog`, which is the +/// diagnostic for a peer that STOPPED reading. The documented maximum paste +/// could not complete. Two whole `max_payload`s is the smallest number that is +/// headroom rather than another coincidence: one peer may legitimately be +/// assembling a message of the largest size `framed` will accept while the rest +/// of the table holds frames, and past 32 MiB the fattest peer is the peer that +/// stopped draining. Neither the mirrors nor the pane queues are in this +/// number: a mirror is this session's own bookkeeping for a client it chose to +/// serve, and a pane is bounded per pane by `pty_backlog` because it is not a +/// peer and cannot be closed to reclaim anything. +const session_backlog = 2 * @as(usize, wire.max_payload); + +/// Bytes of un-drained INPUT one pane's shell may owe before more is refused. +/// +/// A pane is not a client, so the answer cannot be `out_backlog`'s: a client +/// that stops draining is closed, and the thing at the other end of a pty is a +/// program the human is running. This refuses the write and says so on the +/// pane's message row instead, which is the only honest answer left — dropping +/// input silently loses half a command line, and killing a shell to reclaim a +/// megabyte destroys work. +/// +/// Checked BEFORE the append, exactly as `queue` checks `out_backlog`, and that +/// is what makes 1 MiB enough: any single write lands whole, so a maximum paste +/// into an empty queue is never truncated. What gets refused is MORE input typed +/// at a program that has stopped reading its input at all — `sleep 3600`, a +/// stopped job, anything blocked on its own output. +const pty_backlog = 1 << 20; + +/// How long a connection has to say `hello`, and the ONE number both ends of +/// this transport time the handshake against. `pub` because a frontend that +/// waited longer than the session is willing to hold its slot would report a +/// timeout for a slot that had already been taken back, and a frontend that +/// waited less would give up on a session that was still going to answer — two +/// halves of one deadline, and two literals is how they drift apart. +/// +/// The `Session` field it initialises is a field and not this constant for +/// exactly one reason: the test for expiry would otherwise have to sleep five +/// seconds. See `Session.greet_deadline_ms`. +pub const greet_deadline_default_ms: u32 = 5_000; /// What a DRAINED client is allowed to keep. `in` grows to hold one whole /// message, so a single 4 MiB paste otherwise leaves 4 MiB resident in that @@ -206,36 +343,52 @@ const Client = struct { accepted_ms: i64 = 0, }; -/// A pane's shell: which frontend was asked to fork it, and where. +/// A pane's shell: forked by THIS process, drained by its poll set, reaped by +/// it. /// -/// WHY THE SESSION REMEMBERS THIS. A spawn is the one primary-only call that -/// has to survive having no frontend to serve it. Every boot layout creates its -/// panes before the socket exists, so a `--detach` performs its startup spawns -/// with nobody attached — and dropping them meant a session that opened with -/// panes whose shells had never been forked, forever, in silence. So a spawn -/// with no primary is OWED, and asked of whoever attaches next. -/// -/// It is also what makes `pty_write` reach the right process. A pane's pty -/// lives in the frontend that forked it, which is not always the primary: A -/// attaches and forks the shells, B attaches, A leaves — the panes are re-owed -/// to B, and then C attaching into A's freed slot becomes primary while the -/// ptys are in B. Routing a pane's bytes by its OWNER rather than by the -/// primary is the difference between typing into a shell and typing into -/// nothing. -const Shell = struct { - /// The slot that was asked to fork this pane's shell, or null when nobody - /// has been. - owner: ?u8 = null, - /// A spawn owed to whoever attaches next: either it was never asked, or the - /// frontend holding it left and took the pty with it. - owed: bool = false, - /// Copied, because `push_spawn`'s `cwd` borrows the core's memory for the - /// length of that one call and this outlives it by definition. - cwd: std.ArrayListUnmanaged(u8) = .empty, +/// There is no owner here and nothing is owed, and the absence is the whole +/// change. This struct used to record which frontend had been asked to fork a +/// pane and re-ask the next arrival when that frontend left, because the pty +/// lived in the frontend that forked it; and a `--detach`, whose panes always +/// exist before its socket does, had nobody to ask at all and had to remember +/// the request instead. Both were one problem, and forking here dissolves both: +/// a startup layout's shells are forked during `run`'s pre-loop drain with +/// nobody attached, and they are still those same shells when the tenth +/// frontend attaches an hour later. +const Pty = struct { + /// The pty master, non-negative exactly when this pane has a live shell. + /// While it is here it is in the poll set (`waitInput`), NON-BLOCKING like + /// every other descriptor this file holds — `spawn` flips it, because + /// `forkpty` hands it back blocking and tty.zig's streaming reader wants it + /// that way. + fd: c_int = -1, + /// Kept past the fork for `look.shellCwd` and `look.ttyTaken`, both of which + /// ask /proc about this pid rather than about the descriptor. + pid: posix.pid_t = 0, + /// Bytes owed to this shell's stdin, drained by POLLOUT and bounded by + /// `pty_backlog`. The same shape as `Client.out`, for the same reason: the + /// thing on the far side may not be reading, and this process must not wait + /// to find out. See `ptyWrite`. + out: std.ArrayListUnmanaged(u8) = .empty, +}; + +/// What one descriptor in `waitInput`'s poll set is. A tagged union rather than +/// the bare slot index this loop used to carry alongside its `pollfd`s, because +/// the set now holds four different kinds of thing and a `u8` cannot say which. +const Source = union(enum) { + listener, + client: u8, + pty: u8, + inotify, }; pub const Session = struct { gpa: std.mem.Allocator, + /// The Io every filesystem read this host performs goes through: + /// file_watch.zig's reload of a changed pane, and `user_config.dumpThemes`. + /// Required and not optional — a session that owns the disk work cannot be + /// handed a null disk. + io: std.Io, core: *pardes.Pardes, /// -1 when nothing is bound: an unsupported platform, or a bind that /// failed. A session with no listener is a session nobody can attach to, @@ -258,9 +411,33 @@ pub const Session = struct { /// needed rather than sized from `wire.max_payload`, which would be 16 MiB /// of resident memory for a session whose frames are six kilobytes. scratch: std.ArrayListUnmanaged(u8) = .empty, - /// Where each pane's shell lives, and which spawns are still owed. See - /// `Shell`. - shells: [pardes.MAX_PANES]Shell = @splat(.{}), + /// Each pane's shell. Forked here, drained by the poll set, reaped by + /// `harvest`. See `Pty`. + ptys: [pardes.MAX_PANES]Pty = @splat(.{}), + /// The OSC 133 rc files a forked shell sources, staged once for the life of + /// this host exactly as tty.zig stages them for the life of a terminal: + /// `shell_bin.resolve` hands a child pointers into these buffers and the + /// child holds them until it execs, so they must not live in a stack frame. + /// The default is the empty one, which `resolve` reads as "this shell gets + /// no prompt marks"; `run` supplies a staged one. + prompt_rcs: shell_bin.PromptRcs = .{}, + /// The one inotify descriptor behind every watch this session holds, and the + /// last member of the poll set. Opened lazily — see `inotify`. + inotify_fd: c_int = -1, + /// Which directory mark belongs to which pane, and the generation the core + /// has already accepted from each. file_watch.zig owns the shape and the + /// transaction; this host owns only the descriptor and the wake. + watches: file_watch.Table = @splat(null), + /// A reconcile pass is due: a watched directory had an edge, or the last + /// pass asked for another one. Consumed at the end of `waitInput`, which is + /// where a core change still makes the current frame — `Pardes.pump` renders + /// after `pull_wait_input` returns. + check_files: bool = false, + /// False during `run`'s pre-loop effect drain. Read in exactly one place: + /// `Pardes.loadThemeFile` animates an interactive theme change and must not + /// animate a startup one, and a `ThemeFile` in a boot layout is a startup + /// one. tty.zig spells the same distinction `threads_ok`. + in_loop: bool = false, /// How long a connection may stay silent before the session takes its slot /// back. `Client.open` writes its `hello` in the same call that connects, /// so a peer that has said nothing for five seconds is not a frontend that @@ -268,8 +445,10 @@ pub const Session = struct { /// real frontend out with a `refuse .full`. /// /// A field rather than a constant for exactly one reason: the test for that - /// would otherwise have to sleep five seconds. Nothing else changes it. - greet_deadline_ms: u32 = 5_000, + /// would otherwise have to sleep five seconds. Nothing else changes it, and + /// the number itself is `greet_deadline_default_ms`, which the frontend half + /// of this transport reads too. + greet_deadline_ms: u32 = greet_deadline_default_ms, /// Monotonic milliseconds until which the LISTENER is left out of the poll /// set, because an `accept` failed for a reason that persists. See `accept`. accept_paused_ms: i64 = 0, @@ -286,7 +465,19 @@ pub const Session = struct { for (&s.clients) |*c| if (c.fd >= 0) s.close(c, .quitting); s.unlisten(); s.scratch.deinit(s.gpa); - for (&s.shells) |*sh| sh.cwd.deinit(s.gpa); + // The pane shells go with the SESSION and not with a frontend, which is + // this file's whole change. Closing a master is what hangs its shell up; + // `harvest` collects whatever has already exited, and the process is + // about to leave, so anything slower than that is the kernel's job. + for (0..s.ptys.len) |pane| s.closePty(@intCast(pane)); + s.harvest(); + // Every mark dies with the descriptor, so there is nothing to unmark. + if (s.inotify_fd >= 0) { + _ = libc.close(s.inotify_fd); + s.inotify_fd = -1; + } + // Unlinks the two rc files staged for this host's shells. + s.prompt_rcs.deinit(); } /// Bind and listen. False when there is no socket, and a session without @@ -362,17 +553,31 @@ pub const Session = struct { return @ptrCast(@alignCast(ctx.?)); } - /// Thirteen methods, and the seven that are missing are missing on purpose - /// — see wire.zig's header for each one's reason. `push_poll_frame` and - /// `push_post_present` carry no information a frame does not; the four - /// synchronous or dispatched pulls and `push_fs_reply` belong to whoever - /// owns the core, which is this process. + /// Sixteen methods, and NOT the fullest host in the tree — that claim stood + /// here, was believed, and was copied into docs/detached.md before an audit + /// counted the others. The tty and SDL shells fill NINETEEN each (everything + /// but `pull_gpio_toggle` and `push_detach`) and macOS fourteen, so this host + /// is the only one that implements `push_detach` and otherwise the least + /// complete of the three desktop hosts. What is true is narrower and is the + /// point anyway: it performs every MACHINE-LOCAL effect there is, and the + /// five of host.zig's twenty-one it leaves null are null because there is + /// nothing here for them to do. Three of those five are real losses a person + /// can notice — no `pull_lsp` and no `pull_pipe`, because both want the + /// worker pool this deliberately single-threaded loop does not have, and no + /// `push_fs_reply`, because this process mounted no /dev/fuse. The other two + /// are not losses at all: `push_post_present` marks the moment a frame + /// reached a screen and this process has no screen, and `pull_gpio_toggle` + /// wants pads. + /// + /// `push_detach` is the one entry here that is not an effect. See `detach`. const vtable: host_api.Host.VTable = .{ .pull_wait_input = waitInput, .push_present = present, + .push_poll_frame = pollFrame, .push_spawn = spawn, .push_pty_write = ptyWrite, .push_pty_resize = ptyResize, + .pull_tty_taken = ttyTaken, .push_write_file = writeFile, .push_write_dump = writeDump, .push_watch_file = watchFile, @@ -381,6 +586,7 @@ pub const Session = struct { .push_set_clipboard = setClipboard, .pull_read_clipboard = readClipboard, .push_open_link = openLink, + .push_detach = detach, }; // ---- routing ---------------------------------------------------------- @@ -402,16 +608,6 @@ pub const Session = struct { return s.primary(); } - /// The frontend holding pane `pane`'s pty, which is NOT the primary in - /// general — see `Shell`. Null when no frontend holds it, which is a pane - /// with no child: the core's own answer to that is silence, and so is this. - fn holder(s: *Session, pane: u8) ?*Client { - if (pane >= s.shells.len) return null; - const owner = s.shells[pane].owner orelse return null; - const c = &s.clients[owner]; - return if (c.attached) c else null; - } - /// Which slot this client is. From the pointer because every caller here /// holds a `*Client` and not its index. fn slotOf(s: *Session, c: *const Client) u8 { @@ -422,84 +618,202 @@ pub const Session = struct { for (&s.clients) |*c| if (c.attached) s.send(c, msg); } - // ---- the host methods ------------------------------------------------- + // ---- the host methods: pseudo-terminals ------------------------------- fn spawn(ctx: ?*anyopaque, pane: u8, cwd: []const u8) void { const s = of(ctx); - if (pane >= s.shells.len) return; // the core indexes its own panes - const sh = &s.shells[pane]; - // Kept whether or not there is somebody to ask, because the pane now - // exists either way and the cwd is the only thing that cannot be - // reconstructed later. - sh.cwd.clearRetainingCapacity(); - sh.cwd.appendSlice(s.gpa, cwd) catch {}; - if (s.primary()) |c| { - sh.owner = s.slotOf(c); - sh.owed = false; - return s.send(c, .{ .spawn = .{ .pane = pane, .cwd = cwd } }); + if (pane >= s.ptys.len) return; // the core indexes its own panes + // The in-process host's reaping rule and its reason, verbatim from + // tty.zig `spawn`: the core reuses pane ids and there is no close + // effect, so a deleted pane's shell lives in its slot until a respawn + // lands here. + s.closePty(pane); + var cwd_buf: [256:0]u8 = undefined; + var cwd_z: ?[*:0]const u8 = null; + if (cwd.len > 0 and cwd.len < cwd_buf.len) { + @memcpy(cwd_buf[0..cwd.len], cwd); + cwd_buf[cwd.len] = 0; + cwd_z = @ptrCast(&cwd_buf); } - // Nobody can fork a shell right now — a startup layout, or every - // frontend gone. NOT dropped: `reconcile` asks the next arrival. - sh.owner = null; - sh.owed = true; + const child = host_io.forkShell( + s.core, + pane, + &s.prompt_rcs, + s.core.shellBin(), + cwd_z, + // The SESSION grid — which `reconcile` already made the smallest + // common one across everyone attached, and which survives every + // frontend leaving, so a shell forked into an empty session is + // still sized like the pane the core reflowed. + s.core.screen_h, + s.core.screen_w, + null, // no `--fs` mount in a daemon: nothing here answers /dev/fuse + ); + // A `forkpty` that failed left `master` holding a number this process + // does not own. The shells get away with not checking because they hand + // the descriptor to a reader task that simply ends; this one would go + // into `poll(2)`, come back POLLNVAL, and be closed out from under + // whoever really owns it. + if (child.pid < 0) return; + s.ptys[pane] = .{ .fd = child.file.handle, .pid = child.pid }; + // ...and the master joins the rule every other descriptor in this file + // obeys. `forkpty` hands it back BLOCKING, and host_io.zig leaves it that + // way because tty.zig streams it from a thread that wants a blocking + // read; a poll loop wants the opposite, and one blocking `write(2)` here + // is the whole session parked. Only this side of the pty is affected — + // the master and the slave are separate open file descriptions, so the + // shell's own stdin stays exactly as `forkpty` made it. + setNonblock(child.file.handle); + // The pane's starting directory, for the tags. `pollFrame` keeps it + // current after a `cd`; this is the one before the first frame. + var lbuf: [1024]u8 = undefined; + if (look.shellCwd(child.pid, &lbuf)) |wd| s.core.setCwd(pane, wd); } + /// Keystrokes and pastes into the shell, QUEUED and never blocked on. + /// + /// This was one blocking `host_io.writeFd`, and the comment defending it + /// argued that "the peer is a shell this process forked rather than a + /// stranger who can stop reading on purpose". The peer is whatever program + /// the human ran in the pane: `sleep 3600`, a job stopped with ^Z, anything + /// blocked writing its own output. Any of those plus a paste larger than the + /// pty's input buffer — four kilobytes, and a frontend is entitled to send a + /// four-MEGABYTE paste — put this single-threaded process to sleep inside + /// `write(2)` with the whole session behind it: no frame to any frontend, + /// fifteen other masters unread, no `accept`, no `expire`, no inotify drain. + /// + /// So a pane owes bytes the way a client does, and the answer is the shape + /// this file already had for exactly this problem. What differs is what + /// happens when the queue will not drain: a client that stops reading is + /// CLOSED, and a pane cannot be, because closing it kills a program the + /// human is running. See `pty_backlog` — the write is refused and said out + /// loud on the pane's own message row. fn ptyWrite(ctx: ?*anyopaque, pane: u8, bytes: []const u8) void { const s = of(ctx); - // To the frontend that forked this pane's shell, not to the primary: - // the pty is in that process and nowhere else. A pane whose holder is - // gone is silent, which is what the core does with a null method. - if (s.holder(pane)) |c| s.send(c, .{ .pty_write = .{ .pane = pane, .bytes = bytes } }); + if (pane >= s.ptys.len) return; + const pt = &s.ptys[pane]; + // A pane with no shell swallows what is typed at it, which is exactly + // what the core does with a null method. + if (pt.fd < 0) return; + // BEFORE the append, which is `queue`'s rule and gives `queue`'s + // guarantee: one write always lands whole, so the biggest paste anyone + // can send is never truncated on arrival, and what is refused is the + // NEXT one typed at a program that has read nothing. + if (pt.out.items.len > pty_backlog) { + var mbuf: [256]u8 = undefined; + const text = std.fmt.bufPrint( + &mbuf, + "input refused: pane not reading ({d} bytes queued)", + .{pt.out.items.len}, + ) catch "input refused: pane not reading"; + return s.core.setMessage(pane, text); + } + pt.out.appendSlice(s.gpa, bytes) catch { + // Out of memory for a keystroke. The shell is fine and the session + // is fine; this one write is not, and saying so is all there is. + return s.core.setMessage(pane, "input refused: out of memory"); + }; + // Try immediately. On an idle pty this empties the queue in one write and + // the descriptor never asks for a POLLOUT at all, which keeps a session + // of keystrokes exactly as cheap as it was. + s.flushPty(pane); } fn ptyResize(ctx: ?*anyopaque, pane: u8, cols: u16, rows: u16) void { const s = of(ctx); - if (s.holder(pane)) |c| s.send(c, .{ .pty_resize = .{ .pane = pane, .cols = cols, .rows = rows } }); + if (pane >= s.ptys.len) return; + const fd = s.ptys[pane].fd; + if (fd < 0) return; + const ws: posix.winsize = .{ .row = rows, .col = cols, .xpixel = 0, .ypixel = 0 }; + _ = posix.system.ioctl(fd, TIOCSWINSZ, @intFromPtr(&ws)); + } + + /// Is this pane's tty still the prompt we forked, or has a program taken it? + /// + /// Answerable at all only because the pty is HERE. While a pane's shell + /// lived in a frontend this method had to stay null, and a null one means + /// the core types every `Exec` at the shell — into vim, into a pager, into + /// an agent waiting on stdin. Lazy by construction (host.zig): it runs where + /// the core is about to type a command line, so the /proc walk costs an + /// ordinary frame nothing. + fn ttyTaken(ctx: ?*anyopaque, pane: u8) bool { + const s = of(ctx); + if (pane >= s.ptys.len) return false; + const pt = s.ptys[pane]; + if (pt.fd < 0) return false; + return look.ttyTaken(pt.pid, pt.fd); } + // ---- the host methods: the filesystem --------------------------------- + fn writeFile(ctx: ?*anyopaque, pane: u8, path: []const u8, bytes: []const u8) void { const s = of(ctx); - if (s.primary()) |c| return s.send(c, .{ .write_file = .{ .pane = pane, .path = path, .bytes = bytes } }); - // Nobody attached, and a Put must not evaporate. This method being - // non-null means the core did NOT reach for its own filesystem, so the - // obligation a null method would have discharged is discharged here by - // hand — the same shape `readClipboard` below has, and host.zig's rule - // that a zero-method host is a complete pardes. - s.core.fallback.writeFile(path, bytes); + if (!host_io.writeFileBytes(path, bytes)) return; + // Our own write is about to come back as an inotify edge: restamp from + // the bytes we just put there so the reconcile reads as "no change". + // Only when this IS the pane's watched file — a `Save <elsewhere>` must + // not silence a real change to the file the pane has open. Six lines + // shared with tty.zig `writeFile` over the same `file_watch.Table`, + // which is what makes a save in a detached pane behave like a save in a + // terminal one. + if (s.core.panes[pane]) |pn| if (pn.file) |f| if (std.mem.eql(u8, f.path, path)) { + if (s.watches[pane]) |*w| if (w.serial == pn.serial) switch (w.generation) { + .text => w.generation = .{ .text = std.hash.Wyhash.hash(0, bytes) }, + .pdf => {}, + }; + }; + // ...and say so on the pane's message row. AFTER the write, not beside + // it: the early return above is a save that did not happen and must not + // be reported as one. + var mbuf: [256]u8 = undefined; + s.core.setMessage(pane, message.stamp(&mbuf, "saved", path)); } fn writeDump(ctx: ?*anyopaque, bytes: []const u8) void { const s = of(ctx); - if (s.primary()) |c| return s.send(c, .{ .write_dump = bytes }); - // ...and the same for a Dump, including the part that makes the bytes - // reachable again: a real host reports where it landed, which is what - // puts `Restore <path>` in the topbar (pardes.zig `write_dump`). - s.core.fallback.writeFile(pardes.fallback_dump_path, bytes); - s.core.setLastDump(pardes.fallback_dump_path); + var pbuf: [1024:0]u8 = undefined; + const path = pardes.dump.outPath(&pbuf) orelse return; + if (!host_io.writeFileBytes(path, bytes)) return; + // Where it landed, which is what puts `Restore <path>` in the topbar + // (pardes.zig `write_dump`). A dump of a detached session now lands in + // the same directory a terminal session's does, rather than in whatever + // directory the frontend that happened to be primary was started from. + s.core.setLastDump(path); } - fn watchFile(ctx: ?*anyopaque, pane: u8, path: []const u8, on: bool) void { + fn watchFile(ctx: ?*anyopaque, pane: u8, _: []const u8, on: bool) void { const s = of(ctx); - if (s.primary()) |c| return s.send(c, .{ .watch_file = .{ .pane = pane, .path = path, .on = on } }); - // No frontend to watch a path, so the core's own record of what was - // asked is the whole of what a watch means here — exactly what a null - // method leaves behind. - if (pane < s.core.fallback.watched.len) s.core.fallback.watched[pane] = on; + // The path argument is unused because `applyEffect` takes it off the + // core's own pane, together with the serial and the generation that make + // the reconcile safe. That is the one thing a frontend could not do — it + // had no core — and it is why the frontend's copy of this method needed a + // second table of pathnames to go with the watch table. + if (file_watch.applyEffect(s.core, s.io, s.gpa, s.inotify(), &s.watches, pane, on)) + s.check_files = true; } fn watchTheme(ctx: ?*anyopaque, generation: u32, on: bool) void { const s = of(ctx); - // Dropped with nobody attached, and that is the whole of it: the core's - // own `theme_file` effect does nothing for a null method either, so - // there is no obligation left over. Same for `dump_themes` below. - if (s.primary()) |c| s.send(c, .{ .watch_theme = .{ .generation = generation, .on = on } }); + if (file_watch.applyThemeEffect(s.core, s.gpa, s.inotify(), &s.watches, generation, on, s.in_loop)) + s.check_files = true; } fn dumpThemes(ctx: ?*anyopaque, pane: u8) void { const s = of(ctx); - if (s.primary()) |c| s.send(c, .{ .dump_themes = .{ .pane = pane } }); + // A session started without one has nowhere to put them; the core's + // options are the only place that answer lives. + const config_dir = s.core.opts.config_dir orelse return; + const out_dir = user_config.dumpThemes(s.io, s.gpa, config_dir, pardes.themes) catch |err| { + s.core.reportError(pane, "dump themes", err); + return; + }; + defer s.gpa.free(out_dir); + var mbuf: [256]u8 = undefined; + s.core.setMessage(pane, message.stamp(&mbuf, "dumped themes", out_dir)); } + // ---- the host methods: the desktop ------------------------------------ + fn setClipboard(ctx: ?*anyopaque, text: []const u8) void { const s = of(ctx); // Mirrored into the core's own clipboard ALWAYS, not only when nobody @@ -533,6 +847,254 @@ pub const Session = struct { s.core.fallback.setLink(url); } + // ---- the host methods: session control -------------------------------- + + /// `Detach` in an attached frontend: that frontend leaves, the session and + /// every other frontend carry on. tmux's `detach-client`. + /// + /// A `send` and NOTHING ELSE, and each of the three things it does not do is + /// deliberate. It does not quit — the whole point is that the session + /// survives, and a detach that took the daemon with it would be `quit` under + /// another name. It does not touch the core — no pane closes, no shell dies, + /// no frame changes; the grid is retaken by `reconcile` from the frontends + /// that remain, on the ordinary path, because a frontend leaving is already a + /// case this file handles. And it does not close the connection: the frontend + /// closes its own socket when it reads the message, and the peer-hangup path + /// then frees the slot exactly as it does for a frontend somebody killed. + /// Closing from this side would race the frontend's own teardown for no gain. + /// + /// It is therefore the one vtable entry here that is not an effect on the + /// world but SESSION CONTROL — one frontend asking to stop being a frontend + /// — which is why wire.zig numbers it 0x05, in the session range beside + /// `quit`, rather than in 0x10.. where every tag reaches a disk, a clipboard + /// or a browser. The module header's routing table says the same. + /// + /// ORIGIN, ELSE PRIMARY, for `read_clipboard`'s and `open_link`'s reason: it + /// answers something one particular human just typed, so it has to reach that + /// human's screen and not somebody else's — sending a detach to the wrong + /// frontend takes away a session from a person who did not ask. Nobody + /// attached at all is a no-op, and correctly so: there is no frontend to + /// detach, and the core has nothing to record about one. + fn detach(ctx: ?*anyopaque) void { + const s = of(ctx); + if (s.origins()) |c| s.send(c, .detach); + } + + // ---- pane shells ------------------------------------------------------ + + /// Each pane's live cwd, for the tags. One readlink of /proc per pane that + /// has a shell, per frame, which is what tty.zig's `pollFrame` costs — and + /// why the far more expensive question, whether a program has taken the + /// pane's tty, is a pull asked at the `Exec` that cares (`ttyTaken`) + /// instead of polled here. + /// + /// This could not exist before. A detached session's shells lived in a + /// frontend, and a frontend has no core to report a cwd TO, so a `cd` in a + /// detached pane never reached its tag no matter how many frontends were + /// watching. The pids are here now, so it does. + fn pollFrame(ctx: ?*anyopaque) void { + const s = of(ctx); + for (&s.ptys, 0..) |*pt, pane| { + if (pt.fd < 0) continue; + var lbuf: [1024]u8 = undefined; + if (look.shellCwd(pt.pid, &lbuf)) |cwd| s.core.setCwd(pane, cwd); + } + } + + /// Drop a pane's shell: out of the poll set, out of the process. Closing the + /// master is what hangs the shell up — which is true only because + /// `host_io.forkShell` puts FD_CLOEXEC on it, so no LATER pane's shell is + /// still holding a copy open. The pid is left to `harvest`, because a + /// `waitpid` here would return 0 for a shell that has not noticed the hangup + /// yet and that answer is worth nothing. + /// + /// The core is NOT told. Its two callers are `spawn` — a respawn, where the + /// core is the thing that asked — and `deinit`, where there is no core left + /// to tell. The path that does tell it is `paneEof`. + fn closePty(s: *Session, pane: u8) void { + const pt = &s.ptys[pane]; + if (pt.fd < 0) return; + _ = libc.close(pt.fd); + // Before the reset, or the queue's allocation goes with the slot: what + // is in it is input a program that is not reading never took, and there + // is nobody left to hand it to. + pt.out.deinit(s.gpa); + pt.* = .{}; + } + + /// Push what the kernel will take of what this pane owes its shell, and + /// leave the rest for a POLLOUT. `flush`'s body, on a pty instead of a + /// socket, down to the `retire` that hands a drained megabyte back. + /// + /// The one difference is what an error means. A failed write to a SOCKET + /// closes a client; a failed write to a master means the slave side is gone, + /// which is the same event as a read of 0. It is NOT the same moment, + /// though, and that is why the error arm reads the pane before it ends it: + /// linux's `n_tty_write` returns EIO the instant the slave has no open + /// descriptors left, while `n_tty_read` on that same master still hands back + /// what the shell wrote before it went — so the write fails while the last + /// line is still retrievable, and ending the pane first would throw it away. + /// That is the very thing the dispatch's `.pty` branch protects against when + /// it takes POLLIN before POLLHUP, and it has to hold here too, because two + /// paths reach this arm with no read of their own in between: the dispatch + /// runs POLLOUT before POLLIN, and `ptyWrite` calls this during `perform`, + /// after this round's `readPty` has already been and gone. + fn flushPty(s: *Session, pane: u8) void { + const pt = &s.ptys[pane]; + var off: usize = 0; + while (off < pt.out.items.len) { + const n = libc.write(pt.fd, pt.out.items.ptr + off, pt.out.items.len - off); + if (n < 0) switch (libc.errno(n)) { + .INTR => continue, + // The pty's input buffer is full: the rest waits for POLLOUT, + // and this is the case the whole change exists for. + .AGAIN => break, + else => { + // `readPty` either takes that last chunk or reaches the end + // itself and has already ended the pane; the guard is what + // stops the second `paneEof` from being a double-end. + s.readPty(pane); + if (s.ptys[pane].fd >= 0) s.paneEof(pane); + return; + }, + }; + // No progress and no error. host_io.zig's `writeFd` says why this is + // a `break` and never a retry: looping on a zero-byte write is a + // spin, and a spin in here is the whole session at 100% of a core + // with no syscall for a signal to interrupt. + if (n == 0) break; + off += @intCast(n); + } + if (off == 0) return; + if (off == pt.out.items.len) { + pt.out.clearRetainingCapacity(); + return retire(s.gpa, &pt.out); + } + std.mem.copyForwards(u8, pt.out.items, pt.out.items[off..]); + pt.out.items.len -= off; + } + + /// One read per readable pty per round — `receive`'s rule for clients, + /// applied to shells: a `yes` in pane 1 gets one turn and the loop moves on + /// to the other panes, the frontends and the frame. + /// + /// Nothing is copied. `Event.output` borrows the buffer for the length of + /// one `update` call, which is the same borrow window every other host gives + /// a pty chunk — tty.zig frees its duplicate the line after the update — + /// except that this one never allocated a duplicate to free. A daemon + /// serving sixteen shells printing build logs asks the allocator for + /// nothing. + fn readPty(s: *Session, pane: u8) void { + var buf: [pty_chunk]u8 = undefined; + const got = libc.read(s.ptys[pane].fd, &buf, buf.len); + if (got == 0) return s.paneEof(pane); + if (got < 0) return switch (libc.errno(got)) { + // A master that said POLLIN and then had nothing is not an error; + // the next round asks again. + .INTR, .AGAIN => {}, + // EIO is how linux reports the slave side going away, which is the + // ordinary end of a shell rather than a fault. + else => s.paneEof(pane), + }; + s.core.update(.{ .output = .{ .pane = pane, .bytes = buf[0..@intCast(got)] } }); + } + + /// The shell in `pane` is gone. The descriptor leaves the poll set BEFORE + /// the core is told, because an `eof` is what makes the core offer a respawn + /// and a respawn into a slot still holding the old fd would leak it. + fn paneEof(s: *Session, pane: u8) void { + s.closePty(pane); + s.harvest(); + s.core.update(.{ .eof = .{ .pane = pane } }); + } + + /// Collect every child that has exited. + /// + /// `waitpid(-1)` and not a pid list, because the only children this process + /// forks are pane shells (`host_io.forkShell`) — so "any exited child" and + /// "an exited pane shell" are the same set — and because the pids a list + /// would hold are exactly the ones it cannot help with: a respawn closes a + /// master, and the shell that gets the hangup exits some milliseconds later + /// with its slot already reused by a different shell. + /// + /// Nothing here waits, so a session whose shells are all running pays one + /// syscall that returns 0. Called once per poll round and again wherever a + /// shell is dropped, which is what keeps a daemon that runs for a week and + /// spawns a thousand shells free of zombies — the one bookkeeping cost a + /// long-lived process pays that a frontend, which exits, never did. + fn harvest(_: *Session) void { + while (true) { + // 0: there are children and none has exited. -1: no children at all. + if (libc.waitpid(-1, null, libc.W.NOHANG) <= 0) return; + } + } + + // ---- watched files ---------------------------------------------------- + + /// The one inotify descriptor behind every watch, opened on first use. + /// + /// Lazy for two reasons pointing the same way: a `Session` is built as a + /// struct literal (client.zig's test harness is one) and so has no init hook + /// to open it in, and a session whose panes are all shells never watches a + /// path and has no use for one. -1 on anything but linux and on a failed + /// `inotify_init1`, which file_watch.zig reads as "mark nothing" — the core + /// then keeps its own record of what was asked and simply never gets a + /// reload, which is what a host with no watcher has always done. + /// + /// NONBLOCK because this descriptor is drained from `poll`, not from a + /// thread parked in `read` (tty.zig `watchFiles`): `drainInotify` must be + /// able to stop. + fn inotify(s: *Session) c_int { + if (s.inotify_fd >= 0) return s.inotify_fd; + // The whole body is inside the comptime branch so that neither + // `inotify_init1` nor `linux.IN` is even analysed on a platform that has + // no inotify — the same shape file_watch.zig's `watchPath` uses. + if (comptime builtin.os.tag == .linux) { + s.inotify_fd = libc.inotify_init1(linux.IN.CLOEXEC | linux.IN.NONBLOCK); + } + return s.inotify_fd; + } + + /// A directory this session marked had an edge. The CONTENTS are discarded + /// on purpose, exactly as tty.zig's watcher thread discards them: a record + /// names a mark and a filename, and reconciling every mark against the + /// generation the core accepted is both cheaper and safer than deciding from + /// the record which pane it meant. + /// + /// DRAINED TO EMPTY, in a loop, and one read was a real cost rather than the + /// coalescing this comment used to claim. The descriptor is level-triggered, + /// so a queue left partly full makes `poll` return ready again immediately — + /// and each of those rounds is a whole `pump`: `reloadChanged` over all 17 + /// slots, every watched text pane re-read from disk and re-hashed, a render, + /// a present. A `git checkout` can queue the kernel's whole 16384 events; at + /// roughly 128 records per 4 KiB that was ~128 spin rounds and some two + /// thousand whole-file reads for one command, at 100% of a core, while every + /// frontend got a frame per round it could not use. The fd is IN_NONBLOCK + /// (`inotify`), so the loop ends on EAGAIN. + fn drainInotify(s: *Session) void { + var buf: [4096]u8 = undefined; + while (libc.read(s.inotify_fd, &buf, buf.len) > 0) s.check_files = true; + } + + /// Reconcile every marked pane and the theme file. Called at the END of + /// `waitInput`, which is what puts a reload in THIS frame: `Pardes.pump` + /// renders after `pull_wait_input` returns. + /// + /// The loop is `reloadChanged`'s contract. It asks for another pass when a + /// PDF's pathname changed between the stat before MuPDF reopened it and the + /// stat after — a save that landed mid-reconcile, where committing either + /// identity would lose a generation. tty.zig posts that request back into + /// its event queue; this loop has no queue, so it is retried here and + /// BOUNDED, because a file being rewritten in a loop must not hold the core. + /// What is left over is picked up by the next directory edge. + fn reloadWatched(s: *Session) void { + if (!s.check_files) return; + s.check_files = false; + for (0..reload_retries) |_| { + if (!file_watch.reloadChanged(s.core, s.io, s.gpa, &s.watches)) return; + } + } + // ---- the frame -------------------------------------------------------- fn present(ctx: ?*anyopaque, surface: *const pardes.Surface) void { @@ -590,19 +1152,37 @@ pub const Session = struct { // ---- the loop --------------------------------------------------------- /// The only place this process sleeps, which is what `pull_wait_input`'s - /// comment in host.zig requires of whoever serves it. One `poll(2)` covers - /// the listener and every attached frontend; there is no thread per client - /// and nothing here blocks on a single peer. + /// comment in host.zig requires of whoever serves it, and the only place it + /// waits on ANYTHING: one `poll(2)` over the listener, every attached + /// frontend, every pane's pty master and the inotify descriptor. No thread + /// per client, no thread per shell, no watcher thread, and nothing here + /// blocks on a single peer. + /// + /// That the shells are in this set and not on threads of their own is what + /// lets a daemon own sixteen of them and stay a single-threaded state + /// machine — and it costs the shells nothing, because a pty master is + /// pollable and a pane's output has nowhere to go but the core this loop is + /// driving anyway. fn waitInput(ctx: ?*anyopaque, timeout_ms: u32) void { const s = of(ctx); // Push what the kernel will take before sleeping: a client that becomes // writable while we are inside poll(2) would otherwise be a frame late, // and a frame late is a frame skipped (see `present`). for (&s.clients) |*c| if (c.fd >= 0) s.flush(c); + // The session grid, retaken BEFORE the sleep as well as after it. A + // client can leave OUTSIDE this function — a `set_clipboard` broadcast + // whose write failed during `perform` closes it — and the minimum across + // attached frontends would then stay sized for a frontend that is gone + // until some descriptor happened to become readable, which on an idle + // session is never. That is what this call buys, and `regridded` is what + // it costs: a round that has just told the core to reflow must not then + // sleep on it, because the frame carrying that reflow is the one + // `Pardes.pump` composes the moment this returns. + const regridded = s.reconcile(); const now = monotonicMs(); - var fds: [max_clients + 1]libc.pollfd = undefined; - var slots: [max_clients + 1]u8 = undefined; + var fds: [poll_slots]libc.pollfd = undefined; + var src: [poll_slots]Source = undefined; var n: usize = 0; // The listener is left OUT of the set while accepting is paused, which // is how an EMFILE is waited out without the core sleeping (see @@ -610,7 +1190,7 @@ pub const Session = struct { const watching_listener = s.listener >= 0 and now >= s.accept_paused_ms; if (watching_listener) { fds[n] = .{ .fd = s.listener, .events = poll_in, .revents = 0 }; - slots[n] = 0; + src[n] = .listener; n += 1; } for (&s.clients, 0..) |*c, i| { @@ -620,13 +1200,43 @@ pub const Session = struct { .events = if (c.out.items.len != 0) poll_in | poll_out else poll_in, .revents = 0, }; - slots[n] = @intCast(i); + src[n] = .{ .client = @intCast(i) }; + n += 1; + } + // The pane shells, and note what is NOT conditional on a frontend: a + // session with nobody attached still polls these, still reads them and + // still feeds the core. That is the difference between a detach that + // pauses your build and a detach that does not. + for (&s.ptys, 0..) |*pt, pane| { + if (pt.fd < 0) continue; + fds[n] = .{ + .fd = pt.fd, + // POLLOUT only while this pane owes its shell bytes, which is + // the same rule and the same reason as a client's: asking for it + // unconditionally makes every idle pty a ready descriptor and + // turns the poll into a spin. + .events = if (pt.out.items.len != 0) poll_in | poll_out else poll_in, + .revents = 0, + }; + src[n] = .{ .pty = @intCast(pane) }; + n += 1; + } + // Opened only once something asked to be watched, so an unwatched + // session simply has one fewer descriptor here (see `inotify`). + if (s.inotify_fd >= 0) { + fds[n] = .{ .fd = s.inotify_fd, .events = poll_in, .revents = 0 }; + src[n] = .inotify; n += 1; } - // A detached session with no listener and no clients has no event - // source at all. Returning immediately would spin the outer + // A session with no listener, no clients, no shells and no watches has + // no event source at all. Returning immediately would spin the outer // `while (!core.quit)` at full speed, so sleep the interval the core // offered and, when it offered none, a frame's worth. + // + // Nothing is owed on this path. `check_files` is only ever set by a + // watch, and a watch means the inotify descriptor is in the set; and + // `reconcile` posts a resize only when a client is ATTACHED, which means + // its socket is in the set — so `n == 0` implies `!regridded` too. if (n == 0) return nap(if (timeout_ms == 0) 16 else timeout_ms); // Zero is the core's word for "sleep until something happens" (see // pardes.zig `pump`: it passes a frame interval only while an animation @@ -639,37 +1249,87 @@ pub const Session = struct { // idle, IS the denial `greet_deadline_ms` exists to answer — so the // wait is clamped to whichever is due first. if (s.nextWake(now)) |due| timeout = if (timeout < 0) due else @min(timeout, due); + // ...and two things are due on nothing at all rather than on a + // descriptor: a reconcile pass a watch effect asked for (`watchFile` ran + // during `perform`, outside this function) and a regrid this round has + // already performed. Both are consumed before this function returns, so + // the round must not sleep before reaching them. + if (s.check_files or regridded) timeout = 0; const ready = libc.poll(&fds, @intCast(n), timeout); - // Expired unconditionally: a slot held by silence comes back on a - // timeout exactly as it does on a wakeup, and a poll that returned - // nothing is the ordinary way this deadline is reached. + // A timeout is an ordinary frame boundary and EINTR is a signal we do not + // handle here. Neither skips anything below any more: what used to be an + // early `return` here is why a client closed without any descriptor being + // readable — which is every `expire` — left the session grid sized for a + // frontend that had gone, until the next readable event, on an idle + // session possibly hours later. + if (ready > 0) s.dispatch(fds[0..n], src[0..n]); + // AFTER the dispatch, and that ordering is itself a fix. `expire` frees a + // client slot and `accept` — which runs INSIDE the dispatch — fills the + // lowest free one, so an expire that ran first could hand a slot to a new + // connection within this same round and the dispatch would then apply the + // OLD connection's `revents` to the new descriptor: a POLLHUP from the + // peer that left, closing the peer that just arrived. The dispatch's + // `c.fd < 0` guard cannot see that, because the fd is perfectly valid — + // it is simply a different fd. Expiring after means a freed slot is + // refilled no earlier than the next round, which builds a fresh `fds` for + // it. It fixes a smaller thing for free, too: a connection whose `hello` + // arrived in THIS round is attached before its deadline is judged, + // instead of being taken back with its handshake still unread. s.expire(monotonicMs()); - // A timeout is an ordinary frame boundary and EINTR is a signal we do - // not handle here; both simply come back next pump. - if (ready <= 0) return; + // Unconditional, and not only where a shell is noticed to have died: a + // shell whose master `spawn` closed on a respawn exits after that close, + // with no descriptor left for anyone to see it on. See `harvest`. + s.harvest(); + // Both before this function returns, so a file that changed on disk and a + // frontend that left during this round are in the frame `Pardes.pump` + // composes next rather than the one after it. + s.reloadWatched(); + _ = s.reconcile(); + } - var k: usize = 0; - if (watching_listener) { - if (fds[0].revents != 0) s.accept(); - k = 1; - } - while (k < n) : (k += 1) { - const c = &s.clients[slots[k]]; - // A slot closed earlier in this same pass (its peer hung up, a - // decode failed, its handshake expired) must not be touched - // through a stale revents. - if (c.fd < 0) continue; - if (fds[k].revents & poll_out != 0) s.flush(c); - if (c.fd < 0) continue; - if (fds[k].revents & poll_in != 0) { - s.receive(c, slots[k]); - } else if (fds[k].revents & (poll_hup | poll_err | poll_nval) != 0) { - // POLLIN wins when both are set: a peer that wrote and then - // closed has bytes still worth reading. - s.close(c, .peer); - } - } - s.reconcile(); + /// One pass over the descriptors `poll` reported ready. Split out of + /// `waitInput` for one reason: everything that must happen AFTER it — + /// `expire`, `harvest`, `reloadWatched`, `reconcile` — is then stated once, + /// in one order, where no early return can skip it. An early return past + /// that list is exactly what findings 5 and 6 were. + fn dispatch(s: *Session, fds: []const libc.pollfd, src: []const Source) void { + for (fds, src) |pfd, source| switch (source) { + .listener => if (pfd.revents != 0) s.accept(), + .client => |i| { + const c = &s.clients[i]; + // A slot closed earlier in this same pass (its peer hung up, a + // decode failed) must not be touched through a stale revents. + if (c.fd < 0) continue; + if (pfd.revents & poll_out != 0) s.flush(c); + if (c.fd < 0) continue; + if (pfd.revents & poll_in != 0) { + s.receive(c, i); + } else if (pfd.revents & (poll_hup | poll_err | poll_nval) != 0) { + // POLLIN wins when both are set: a peer that wrote and then + // closed has bytes still worth reading. + s.close(c, .peer); + } + }, + .pty => |pane| { + if (s.ptys[pane].fd < 0) continue; + // What this pane still owes its shell, which is the whole of + // finding 1's drain: `ptyWrite` queued it and stopped at EAGAIN + // rather than sleeping, and this is where the rest goes. + if (pfd.revents & poll_out != 0) s.flushPty(pane); + // `flushPty` ends the pane when the slave side has gone. + if (s.ptys[pane].fd < 0) continue; + // The same precedence as a client's, and it matters more here: a + // shell that printed its last line and exited reports + // POLLIN|POLLHUP together, and taking the hangup first would + // throw that line away. `readPty` reaches the EOF by reading 0. + if (pfd.revents & poll_in != 0) { + s.readPty(pane); + } else if (pfd.revents & (poll_hup | poll_err | poll_nval) != 0) { + s.paneEof(pane); + } + }, + .inotify => if (pfd.revents & poll_in != 0) s.drainInotify(), + }; } /// Milliseconds until the next deadline that is kept by the CLOCK rather @@ -802,16 +1462,30 @@ pub const Session = struct { } fn apply(s: *Session, c: *Client, slot: u8, tag: u8, payload: []const u8) wire.Error!void { - var scratch: wire.Scratch = .{}; - switch (try wire.decodeClient(tag, payload, &scratch)) { + // `Hello.version` BEFORE the payload is decoded, which is the whole + // point of wire.zig putting it first at a fixed offset: a mismatch has to + // stay diagnosable when the rest of the layout is the part that changed. + // Checking it inside the `.hello` arm defeated exactly that guarantee — + // `decodeClient` refuses a cols/rows this build does not like and refuses + // trailing bytes, so a v2 hello with one extra field came back as + // `.protocol` and the `refuse .version` the frontend needs to say + // something useful was never sent. `wire.helloVersion` reads the one + // field without decoding the rest, and lives in the file that owns the + // layout. + if (tag == @intFromEnum(wire.ClientTag.hello)) { + // A second hello on one connection is not a resize; it is a peer + // that is not speaking this protocol. Judged here rather than in the + // arm below so that a repeat hello is a protocol error whatever + // version it claims. + if (c.attached) return error.BadValue; + const claimed = try wire.helloVersion(payload); + if (claimed != wire.version) { + log.debug("frontend speaks protocol {d}, this session speaks {d}", .{ claimed, wire.version }); + return s.refuse(c, .version); + } + } + switch (try wire.decodeClient(tag, payload)) { .hello => |h| { - // A second hello on one connection is not a resize; it is a - // peer that is not speaking this protocol. - if (c.attached) return error.BadValue; - if (h.version != wire.version) { - log.debug("frontend speaks protocol {d}, this session speaks {d}", .{ h.version, wire.version }); - return s.refuse(c, .version); - } if (s.core.quit) return s.refuse(c, .quitting); c.cols = h.cols; c.rows = h.rows; @@ -847,7 +1521,12 @@ pub const Session = struct { /// Settle the session grid and greet whoever arrived, once per poll round /// rather than once per message: three frontends attaching in the same /// round are one resize, not three reflows of every pane. - fn reconcile(s: *Session) void { + /// + /// True when the CORE was told to reflow, which is the one thing a caller + /// has to react to: the frame carrying that reflow is the next one + /// `Pardes.pump` composes, so a `waitInput` that hears true must not go to + /// sleep before returning. See its `regridded`. + fn reconcile(s: *Session) bool { var cols: u16 = 0; var rows: u16 = 0; for (&s.clients) |*c| { @@ -858,6 +1537,7 @@ pub const Session = struct { // Nobody attached: keep the grid we had. A detached session is not a // session of no size, it is one nobody is looking at, and reflowing // every pane to nothing for zero readers is work with no reader. + var regridded = false; if (cols != 0 and (cols != s.cols or rows != s.rows)) { s.cols = cols; s.rows = rows; @@ -867,33 +1547,14 @@ pub const Session = struct { // safe too. for (&s.clients) |*c| c.need_full = true; s.core.update(.{ .resize = .{ .cols = cols, .rows = rows } }); + regridded = true; } for (&s.clients, 0..) |*c, i| { if (!c.greet) continue; c.greet = false; s.send(c, .{ .welcome = .{ .slot = @intCast(i), .cols = s.cols, .rows = s.rows } }); } - s.flushOwed(); - } - - /// Hand every owed spawn to the frontend that can serve it. Runs at the end - /// of a poll round, so a frontend that has just been greeted is asked for - /// its panes' shells in the same round it arrived — and a session that was - /// started with panes and no frontend (which is every `--detach`) is a - /// session whose panes get their shells from the first attach rather than - /// never. See `Shell`. - fn flushOwed(s: *Session) void { - const c = s.primary() orelse return; - const slot = s.slotOf(c); - for (&s.shells, 0..) |*sh, pane| { - if (!sh.owed) continue; - sh.owed = false; - sh.owner = slot; - s.send(c, .{ .spawn = .{ .pane = @intCast(pane), .cwd = sh.cwd.items } }); - // The send closed it, and `close` put its panes back on the owed - // list; the ones this loop has not reached are still owed anyway. - if (c.fd < 0) return; - } + return regridded; } // ---- bytes ------------------------------------------------------------ @@ -902,11 +1563,12 @@ pub const Session = struct { const want = wire.serverBound(msg); s.scratch.ensureTotalCapacity(s.gpa, want) catch return s.close(c, .oom); const bytes = wire.encodeServer(s.scratch.allocatedSlice()[0..want], msg) catch |err| { - // The only reachable case is a payload past `max_payload`: a save - // of a pane holding more text than this protocol carries. The - // session keeps it (the core's own filesystem already has it) and - // the frontend's copy does not happen — said out loud rather than - // silently. + // The only reachable case is a payload past `max_payload`, and with + // every effect that carried a whole file gone from this protocol the + // only payload that can still get there is a yank of more than + // 16 MiB. The session keeps it — `setClipboard` put it in the core's + // own clipboard before this was ever queued — and the frontends' + // desktop clipboards do not get it, out loud rather than silently. log.debug("message {t} not encodable: {t}", .{ msg, err }); return; }; @@ -918,7 +1580,8 @@ pub const Session = struct { // what this refuses is a client that has stopped draining. if (c.out.items.len > out_backlog) return s.close(c, .backlog); // ...and the table as a whole, which `out_backlog` does not bound: 32 - // slots one byte under it each is 32 MiB. See `session_backlog`. + // slots one byte under it each, plus a frame apiece. See + // `session_backlog`. s.account(); if (c.fd < 0) return; // the fattest peer was this one c.out.appendSlice(s.gpa, bytes) catch return s.close(c, .oom); @@ -933,13 +1596,23 @@ pub const Session = struct { /// peer holding the most of it is the peer that stopped reading. The next /// append asks again, so a second offender is closed a message later rather /// than in a loop that could empty the table on one bad frame. + /// + /// `items.len` and NOT `capacity`, which was half of the bug in + /// `session_backlog`'s history. An ArrayList grows geometrically, so a + /// client's `in.capacity` crossed a 4 MiB ceiling while it was still + /// assembling a paste of roughly 2.8 MiB — the peer was punished for the + /// allocator's rounding rather than for anything it held. What this is + /// asking is "how much is a peer making this session hold RIGHT NOW", and + /// that is `items.len`; capacity above it is transient by construction, + /// because `retire` hands back anything over `idle_retain` the moment a + /// buffer empties. fn account(s: *Session) void { var total: usize = 0; var worst: ?*Client = null; var worst_bytes: usize = 0; for (&s.clients) |*c| { if (c.fd < 0) continue; - const held = c.in.capacity + c.out.capacity; + const held = c.in.items.len + c.out.items.len; total += held; if (held > worst_bytes) { worst_bytes = held; @@ -998,8 +1671,9 @@ pub const Session = struct { } /// Free one slot. A frontend dying takes NOTHING with it: not the core, not - /// the listener, not another frontend's frames. Its buffers go back and the - /// slot is reusable on the next connect. + /// the listener, not another frontend's frames, and — since this file + /// forks — not its panes' shells either. Its buffers go back and the slot is + /// reusable on the next connect. fn close(s: *Session, c: *Client, why: Closed) void { if (c.fd < 0) return; log.debug("frontend detached: {t}", .{why}); @@ -1013,19 +1687,12 @@ pub const Session = struct { if (s.origin) |i| if (i == gone) { s.origin = null; }; - // ...and its panes' shells died with the process that forked them. They - // go back on the owed list, so the frontend that replaces this one is - // asked to fork them again in the directory they were forked in: the - // alternative — which is what this did — is a pane that looks alive, - // produces nothing, and swallows everything typed into it. Migrating a - // live pty between processes is the other answer and is a different - // feature; a fresh shell is the one this transport can keep. - for (&s.shells) |*sh| { - const owner = sh.owner orelse continue; - if (owner != gone) continue; - sh.owner = null; - sh.owed = true; - } + // ...and that is the whole of it. A frontend used to take its panes' + // shells with it and leave them owed to whoever attached next, because + // the ptys were in its process; a pane that survived a detach looked + // alive, produced nothing and swallowed everything typed into it. The + // shells are here now, so a frontend leaving is a screen going away and + // nothing else. c.* = .{}; } }; @@ -1038,12 +1705,14 @@ pub const Session = struct { /// core's own `pump`, exactly as the tty and gui shells run it — this frontend /// simply has no window of its own. /// -/// The pre-loop effect drain is here for the same reason tty.zig has one: the -/// startup spawns are already queued, and they have to be PERFORMED before the -/// loop rather than left in the queue. They reach no frontend — there is none -/// yet — and are remembered instead, then asked of the first attach; `Shell` -/// says why that is the only shape that works for a session whose panes exist -/// before its socket does. +/// The pre-loop effect drain is here for the same reason tty.zig has one, and +/// it is no longer half a promise: the startup spawns are already queued and are +/// PERFORMED here, on this process's own process table. So a session binds its +/// socket with every pane's shell already forked and already in the poll set, +/// and the first frontend to attach — whether that is a second later or the +/// next morning — is sent a frame of shells that have been printing into the +/// core since before it existed. Nothing is remembered for a later frontend, +/// because nothing is owed to one. pub fn run(init: std.process.Init, opts: pardes.Options, name: []const u8) !void { const gpa = init.gpa; const allocs = pardes.allocators.init(gpa); @@ -1066,13 +1735,23 @@ pub fn run(init: std.process.Init, opts: pardes.Options, name: []const u8) !void } const core = if (options.load_path) |lp| blk: { - const bytes = try @import("../look.zig").readFile(gpa, lp); + const bytes = try look.readFile(gpa, lp); defer gpa.free(bytes); break :blk try pardes.Pardes.initFromDump(allocs.pardes, options, bytes); } else try pardes.Pardes.init(allocs.pardes, options); defer core.deinit(); - var session: Session = .{ .gpa = gpa, .core = core, .cols = options.cols, .rows = options.rows }; + var session: Session = .{ + .gpa = gpa, + .io = init.io, + .core = core, + .cols = options.cols, + .rows = options.rows, + // Staged before the first fork and owned by the Session for exactly as + // long as it can fork: `shell_bin.resolve` hands a child pointers into + // these buffers, and the child holds them until it execs. + .prompt_rcs = .init(), + }; defer session.deinit(); if (!session.listen(name)) { // Loud, and on stderr rather than through the log: a `--detach` whose @@ -1085,6 +1764,9 @@ pub fn run(init: std.process.Init, opts: pardes.Options, name: []const u8) !void const h = session.host(); core.host = h; while (core.nextEffect()) |effect| core.perform(effect); + // Past the startup drain: a `ThemeFile` reload from here on is a human's + // and animates. See `in_loop`. + session.in_loop = true; while (!core.quit) try core.pump(h); } @@ -1137,7 +1819,14 @@ fn retire(gpa: std.mem.Allocator, list: *std.ArrayListUnmanaged(u8)) void { /// Zero on failure, and every caller treats zero as "no clock" and enforces no /// deadline at all — a session that cannot read a clock keeps every slot rather /// than dropping every slot. -fn monotonicMs() i64 { +/// +/// `pub` for the same reason `setNonblock`, `nosignal` and the `poll_*` +/// constants are: this file owns the transport's conventions and BOTH ends of +/// it, and the clock a handshake is timed against is one of them. client.zig +/// times its wait for a `welcome` on this and against +/// `greet_deadline_default_ms`, so the two ends cannot disagree about how long +/// the handshake is allowed to take. +pub fn monotonicMs() i64 { var ts: libc.timespec = undefined; if (libc.clock_gettime(.MONOTONIC, &ts) != 0) return 0; return @as(i64, ts.sec) * std.time.ms_per_s + @divTrunc(ts.nsec, std.time.ns_per_ms); @@ -1254,8 +1943,8 @@ fn alive(path: [:0]const u8) bool { } /// Unlink the sockets of detached sessions that are gone — our own litter, -/// which `--attach`'s "the one session there is" would otherwise count as a -/// session (tty.zig `sessionName`). `alive` is the whole of the judgement. +/// which the bare `Attach`'s "whichever session is there" would otherwise count +/// as a session (client.zig `resolve`). `alive` is the whole of the judgement. /// /// Bounded: one readdir of a directory only we write to, one connect each. fn sweep(dir: [:0]const u8) void { |
