# Language intelligence in pardes Three things landed together, and only the first two are permanent: 1. **An async execution model.** The core stays a state machine; slow work goes to a worker and comes back as an event. 2. **A helix-exact keymap** for every LSP command. 3. **A seam** (`src/lsp/lsp.zig`) with exactly one function behind it, so competing backends can be swapped, measured, and thrown away. ## The async model There was none before this: every effect the core emitted was fire-and-forget (`spawn`, `write`, `save_file`) or instantaneous. A language query is the first thing pardes asks for that *answers later*, so it needed a request/response shape — and got the smallest one that works. ``` core shell worker | Effect .lsp{id,kind, | | | pane,offset,arg} | | |-------------------------->| | | | snapshot path + content | | |--------------------------->| | | | lsp.query(...) | | Event .lsp_resp{id,rows}| |<--------------------------|<---------------------------| | lspResponse -> jump, or open a results buffer | ``` The shell already ran this exact pattern for pty readers, so the async part is about thirty lines per shell: `tty.zig` uses `io.concurrent` + the vaxis loop queue, `gui.zig` uses a detached thread + the mutex queue it already had. The web shell has no threads and no-ops the effect. Three rules make it safe: - **The worker never touches the core.** Path, source, arg and root are copied into an `LspJob` before it starts (`tty.zig`). The user keeps typing while a query is in flight; a borrowed slice would be a use-after-free the length of one keystroke. - **One query in flight, identified by a monotonic id.** A second press bumps the id, which makes the older answer stale — `lspResponse` drops any id it is not waiting for. This is also what makes a closed pane safe. - **No rows is a legal answer.** A backend that cannot answer appends nothing, which is indistinguishable from a language server still starting up, and the core does nothing. There is no error path to render. ## Results are `+Search` rows Every backend renders into one format: ``` /abs/path/to/file.zig:LINE:COL text /abs/path/to/file.zig:LINE:COL-ENDCOL text ``` 1-based line and column, absolute path. The second form carries the answer's RANGE where the protocol gave one on a single line (a token, a symbol's name), and a look on it SELECTS that span rather than parking at its first cell — so `gd` lands on the whole name and `gr` steps references with each one highlighted. This is the format `look.zig` already resolves, `runSearch` already produces and `n`/`N` already step — so: - **one row from a goto** → jump straight there (`lookAt`) - **several rows** → an output buffer, which `n`/`N` walk which means helix's multi-result picker required **no picker code at all**. The `+Search` buffer *is* the picker. Kinds whose answer is prose rather than locations (`hover`, `code_action`, `format`, `rename`) open `+Hover`/`+Lsp` instead and do not arm the stepper — `n` over a documentation blurb would step to nowhere. `completion` is the kind this shape changes the most. Every other editor answers a dot with a popup of NAMES to insert; a seam that returns locations cannot insert anything, so this one answers with the candidates' **declarations** — one `path:LINE:COL-ENDCOL` row each, in the same `+Search` buffer, steppable with `n`. That is a different and arguably better answer to "what goes here": you read the definitions rather than a list of words. It is the one location kind that does NOT jump on a single row, because with one candidate you still want to see the list rather than be teleported into it. A results buffer is REFILLED rather than reopened when the same kind is asked again — the rule `runSearch` always had, and which the language path was missing. It survived being missing while every query was a deliberate press (`gr` twice left two identical lists and you closed one); Tab after a dot is an ordinary typing keystroke, and measured, twenty of them stacked **fifteen** byte-identical `+Search` panes, crushed the file to one visible line, and then ran `freeSlot` out so the key was silently eaten for the rest of the session. Unlike a search the ARGUMENT is not part of the identity: a language query is asked about a different symbol every time with the same (usually empty) arg, so the kind is the unit. Making it work needed one trick. A completion is asked for exactly when the line is half-typed, and a half-typed line does not parse: `switch (e) { . }` loses the whole switch to the parser's error recovery, taking with it every ancestor an expected-type resolution needs. ZLS answers this with a private token scanner welded to its `*Server`. `lsp_zls.completionSource` instead makes the tree PARSE — it splices a placeholder in after the dot, in the six spellings a half-typed line can need (an identifier, and the same again closing a prong, a statement, a paren or a brace), and keeps the one that both makes the dot reachable in the tree and leaves the fewest parse errors. Everything after that is ZLS's ordinary public resolution over an ordinary tree. **A Tab the backend cannot answer still indents.** The keystroke has already diverted by the time "no rows" comes back, so `lspResponse` performs the indent the Tab prong skipped — on the condition that the cursor has not moved since, so nobody who kept typing gets four spaces landing behind their hands. Without that, a dot in a comment, in a string, or on a line nothing can be made of ate the keystroke outright. With several cursors Tab never diverts at all: a language query is a per-keystroke action inside a per-selection replay, so asking would stop the replay dead and collapse the multicursor. ### What it costs, and what it cannot do Per press, measured by `zig build lspbench` on this repo: | | ReleaseFast | Debug (what `zig build` installs) | |---|---|---| | a switch arm in `src/pardes.zig` (12.8k lines) | 8.8 ms | 87 ms | | `std.` — 91 candidates, each alias-resolved into the stdlib | 26 ms | 204 ms | It is a worker thread, so the editor does not block; but the second press of Tab joins the first query on the UI thread (`old.cancel(io)` in the shell) and that wait is real. Pre-existing and shared by every LSP kind — not this feature's to fix, but it is what a fast double-Tab feels like. Known limitations, in the order you will meet them: - **`@This()` anywhere in a container makes the whole container unresolvable**, so `var list: std.ArrayList(u8) = .` — the most common decl literal in this codebase — answers nothing. This is not the completion filter: `hover` and a plain field access on the same struct return nothing either. It is the case a user hits first, and it is upstream of everything here. - **Only the break AT THE CURSOR is repaired.** Zig's error recovery runs forward, so an unrepaired break earlier in the file swallows the declaration the cursor is in and the answer is empty. While typing you normally have one broken spot, which is the case this works for. - **A dependency module** (`@import("vaxis")`) cannot be typed at all, for the same reason `gd` on `vaxis.init` finds nothing. - **`error.`** is not handled — the position context is `.error_access`, which no branch claims. ## The keymap is helix's, exactly Verified against `helix-term/src/keymap/default.rs`, not from memory. | keys | command | notes | |---|---|---| | `gd` | definition | jumps on a single result, lists on several | | `gD` | declaration | | | `gy` | type definition | | | `gi` | implementation | | | `gr` | references | | | `SPC l k` | hover | opens `+Hover` | | `SPC l r` | rename | tag input, like Find/Grep | | `SPC l a` | code action | | | `SPC l h` | select references | | | `SPC l s` / `SPC l S` | document / workspace symbols | `S` takes a query | | `SPC l d` / `SPC l D` | document / workspace diagnostics | | | `]d` / `[d` | next / prev diagnostic | steps the list, asks for one if absent | | `]D` / `[D` | last / first diagnostic | | | `=` | format | | | `Ctrl`+left-click | definition | the mouse spelling of `gd` | | `Tab` in INSERT mode, right after a `.` | completion | what could go here, and where each of those is defined | Tab is the one key here that is not helix's and not a goto. helix's `Tab` completes; pardes's shows you the CANDIDATES' DECLARATIONS in a `+Search` buffer and inserts nothing, because that is what a seam returning locations can honestly do — see below. It only diverts where an answer is possible: on a terminal, in an output buffer, or in a file the backend does not speak (`lsp.speaks`, which the core asks and the backend answers), Tab indents exactly as it always did. A Tab that silently does nothing would be worse than not having the feature. Nothing about the mode changes either — the pane is still in insert, so typing goes on and walking the answer with `n` means pressing `Esc` first, the same as for every other results buffer. Ctrl-click rides the ordinary left-click drag rather than firing on the press: a click does not place the modal cursor until RELEASE, so a query asked at press time would answer about wherever the cursor previously sat. A ctrl-DRAG still selects, and still asks about where it started. **The gotos are helix's exactly; the leader commands are helix's letters under an `l` prefix.** `g` and `[`/`]` had no conflicts — `gd/gD/gy/gi/gr` and `]d/[d` were all free, so they stay where helix puts them. The bare `` letters were NOT free, and an earlier pass took them anyway, displacing `SPC d` (Del), `SPC k` (Kill), `SPC s?` (Dump/Restore) and `SPC h t` (Tutor). That is the wrong trade: those are pardes's most-pressed keys and predate the language work, whereas an LSP command is something you reach for deliberately and can afford one keystroke more. So every one of them keeps helix's own letter and gains the prefix — `k` becomes `SPC l k`, `d` becomes `SPC l d` — and nothing pardes had moved at all. Two more live in the same group because they belong to it, not to helix: | keys | builtin | what it shows | |---|---|---| | `SPC l i` | `Lspinfo` | which ZLS, which stdlib (and whether it opens), what the backend answers and refuses, plus the last 24 queries with timings, row counts and **the errors `query` swallowed** | | `SPC l w` | `Lspwhy` | why the definition query at the cursor answers what it does | These exist because of the seam's own contract: a backend never fails loudly, which is right for an editor — a thrown analyser must not take the process with it — but it makes a broken backend and a correct one that found nothing look identical. `Lspwhy` narrates the REAL resolution path (the position context is the analyser's own answer, threaded out through a trace) rather than re-deriving it beside the code, because a debug view that reimplements the logic is one that can disagree with it. `Lspinfo` answers from ANY pane, including one with no file, since it is about the backend rather than a document — which matters precisely when the pane you are in is the problem. `K` is **not** hover — in helix it is `keep_selections`. It was checked; do not "fix" it. ## Writing a backend `src/lsp/lsp.zig` is the seam. An implementation supplies three things and touches nothing else: ```zig pub fn query(gpa, arena, req: Req, out: *std.Io.Writer) void pub fn speaks(path: []const u8) bool pub const supports: std.EnumSet(Kind) pub const backend_name = "..." ``` `supports` says what a backend can do; `speaks` says what it can do it TO, and exists for the one key that must not be eaten when the answer is no — see the Tab note above. `query` runs on a worker thread with no access to the core — everything it may read is in `req` (`path`, `source` (NUL-terminated), `offset`, `arg`, `root`). `out` is a plain `std.Io.Writer`: the shell owns the buffer behind it (an `Io.Writer.Allocating`), so a backend never allocates the result, never frees it, and cannot get the allocator wrong. `arena` is freed wholesale on return; `gpa` is for a backend's own scratch. Use `lsp.row()` to emit a location and `lsp.lineCol()` to convert an offset, so every backend's rows are byte-identical in shape. ## How the implementations are judged `zig build lspbench` — same harness, same corpus (pardes's own `src/`), same 22 probes, every backend. - **Feature completeness.** Which kinds return rows, and whether the rows contain what they should. The harness trusts *results*, not the `supports` flag: a kind claimed but returning nothing is reported as `CLAIMED-EMPTY`, and a kind that answers without claiming is `unclaimed-works`. Correctness is a substring the rows must contain, so returning a confident wrong location scores worse than returning nothing. - **Latency.** `cold` (first query, index construction included) and `warm` (median of 20). They differ by orders of magnitude for an indexing backend and both matter: cold is what the first keypress costs, warm is what every one after it costs. - **Memory.** Peak RSS delta (`VmHWM`), so a backend that frees its index before returning still pays for having built it. - **Lines of code.** Not measured by the harness — it is `jj diff --stat` against the base commit. Less is better, and vendoring a library is not free but is charged in build time and dependency surface rather than in lines we maintain. Run `zig build lspbench -- --json` for machine-readable output.