1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
|
# Language intelligence in pardes
Three things landed together, and only the first two are permanent:
1. **An async execution model.** The core stays a state machine; slow work goes
to a worker and comes back as an event.
2. **A helix-exact keymap** for every LSP command.
3. **A seam** (`src/lsp/lsp.zig`) with exactly one function behind it, so competing
backends can be swapped, measured, and thrown away.
## The async model
There was none before this: every effect the core emitted was fire-and-forget
(`spawn`, `write`, `save_file`) or instantaneous. A language query is the first
thing pardes asks for that *answers later*, so it needed a request/response
shape — and got the smallest one that works.
```
core shell worker
| Effect .lsp{id,kind, | |
| pane,offset,arg} | |
|-------------------------->| |
| | snapshot path + content |
| |--------------------------->|
| | | lsp.query(...)
| | Event .lsp_resp{id,rows} |
|<--------------------------|<---------------------------|
| lspResponse -> atomic edit, jump, or results buffer |
```
The shell already ran this exact pattern for pty readers, so the async part is
about thirty lines per shell: `tty.zig` uses `io.concurrent` + the vaxis loop
queue, `gui.zig` uses a detached thread + the mutex queue it already had. The
web shell compiles in no backend at all (`zls_backend` is off for wasm), which
makes `lsp.supports` empty, which makes `lspRequest` return before it emits —
so on the web the effect is never even raised. `web.zig` carries a prong for it
and exports `pardes_lsp_response` anyway; both are unreachable in the shipped
configuration and wait for a host that links a backend.
Three rules make it safe:
- **The worker never touches the core.** Path, source, arg and root are copied
into an `LspJob` before it starts (`tty.zig`). The user keeps typing while a
query is in flight; a borrowed slice would be a use-after-free the length of
one keystroke.
- **One query in flight, identified by a monotonic id.** A second press bumps
the id, which makes the older answer stale. The pending request also records
the pane serial, so a closed-and-reused slot cannot accept its response.
- **Mutating answers are revision-checked.** Rename records the file revision
sent to the worker and applies nothing if the user edited before it answered.
- **No rows is a legal answer.** A backend that cannot answer appends nothing,
which is indistinguishable from a language server still starting up, and the
core does nothing. There is no error path to render.
## Results are `+Search` rows
Every backend renders into one format:
```
sub/file.zig:LINE:COL text under the asking window's dir
sub/file.zig:LINE:COL-ENDCOL text
/abs/path/elsewhere.zig:LINE:COL text anywhere else
```
1-based line and column. The second form carries the answer's
RANGE where the protocol gave one on a single line (a token, a symbol's name),
and a look on it SELECTS that span rather than parking at its first cell — so
`gd` lands on the whole name, and a `gr` reference stepped to with `n`/`N` and
opened with Enter arrives with the reference itself selected. This is the
format `look.zig` already resolves, `runSearch` already produces and `n`/`N`
already walk — so:
- **one row from a goto** → jump straight there (`lookAt`)
- **several rows** → an output buffer, which `n`/`N` walk: a step SELECTS a row
and Enter opens it
which means helix's multi-result picker required **no picker code at all**. The
`+Search` buffer *is* the picker. Non-location answers (`hover`, `code_action`,
`format`) open `+Hover`/`+Lsp` instead, which are prose and not places — `n`/`N`
find nothing look-able in a documentation blurb and walk straight past it to the
next pane on the ring.
Rename is deliberately the one exception to rows as presentation. The backend
emits `@edit START END` records through `lsp.edit()`, using half-open byte
offsets into the exact `Req.source` snapshot it resolved. The core validates
that every range is ordered, non-overlapping and in bounds, checks that the
pane serial and file revision still match, then substitutes the requested name
across all ranges with one allocation and one undo transaction. A malformed,
stale, or empty response changes nothing. The current ZLS backend resolves and
renames references in the **current file only**; it does not claim a workspace
rename.
**A path UNDER `Req.root` is written relative to it; everything else keeps its
full absolute path** (`lsp.rel`). `Req.root` is the directory of the file the
query was asked about — the window that generated the buffer — and a `gr` over
one file was otherwise the same forty-character prefix repeated down the whole
pane, with the part you came to read pushed off the right edge. The short form
resolves because `Req.root` is also the directory the results buffer is *named*
in (`output_pane.open`), and a look resolves a relative word against the
directory of the pane it was clicked in — which is that buffer.
Under, never "shorter": a hit outside the tree is **not** walked up to with
`../`. An absolute path resolves from anywhere and says where it is; a `../..`
chain says neither, and stops being true the moment the row is read anywhere but
beside its own buffer. `test/snapshots/lsprelpath.snap` pins both directions end
to end — the enum one level up (absolute row) and one level down (`inner/tint.zig`,
a stripped path that still has a separator in it), each with the `n` step that
selects it and the Enter that opens it, plus a right click.
This is the rule `look.grep` already follows for its own rows (`look.zig`, the
`shown` computation), written a second time; the two are now the same function
and want to become one.
`completion` is the kind this shape changes the most. Every other editor answers
a dot with a popup of NAMES to insert; a seam that returns locations cannot
insert anything, so this one answers with the candidates' **declarations** —
one row each, in the same `+Search` buffer, steppable with `n`:
```
path:LINE:COL-ENDCOL name the candidate's declaration line
a.zig:2:5-13 verdigris verdigris,
```
The name comes first because that is the thing you would type — the row answers
"what goes here" before "where does it come from", which is the order the
question was asked in; every other kind here answers a WHERE, and for this one
the location is the evidence rather than the answer. It is not padded into a
column: the location in front of it is already ragged, so there is nothing to
align to. That is a different and arguably better answer to "what goes here":
you get the word AND you can read the definition rather than a list of words. It
is the one location kind that does NOT jump on a single row, because with one
candidate you still want to see the list rather than be teleported into it.
A results buffer is REFILLED rather than reopened when the same kind is asked
again — the rule `runSearch` always had, and which the language path was
missing. It survived being missing while every query was a deliberate press
(`gr` twice left two identical lists and you closed one); Tab after a dot is an
ordinary typing keystroke, and measured, twenty of them stacked **fifteen**
byte-identical `+Search` panes, crushed the file to one visible line, and then
ran `freeSlot` out so the key was silently eaten for the rest of the session.
Unlike a search the ARGUMENT is not part of the identity: a language query is
asked about a different symbol every time with the same (usually empty) arg, so
the kind is the unit.
Making it work needed one trick. A completion is asked for exactly when the
line is half-typed, and a half-typed line does not parse: `switch (e) { . }`
loses the whole switch to the parser's error recovery, taking with it every
ancestor an expected-type resolution needs. ZLS answers this with a private
token scanner welded to its `*Server`. `lsp_zls.completionSource` instead makes
the tree PARSE — it splices a placeholder in after the dot, in the six
spellings a half-typed line can need — three shapes, each with and without a
closer still hanging: a switch prong (`_p => {},` / `_p => {}, }`), an
unterminated statement (`_p;` / `_p)`), and a bare identifier (`_p` / `_p }`)
— and keeps the one that both makes
the dot reachable in the tree and leaves the fewest parse errors. Everything
after that is ZLS's ordinary public resolution over an ordinary tree.
**A Tab the backend cannot answer still indents.** The keystroke has already
diverted by the time "no rows" comes back, so `lspResponse` performs the indent
the Tab prong skipped — on the condition that the cursor has not moved since,
so nobody who kept typing gets four spaces landing behind their hands. Without
that, a dot in a comment, in a string, or on a line nothing can be made of ate
the keystroke outright. With several cursors Tab never diverts at all: a
language query is a per-keystroke action inside a per-selection replay, so
asking would stop the replay dead and collapse the multicursor.
### What it costs, and what it cannot do
Per press, measured by `zig build lspbench` on this repo:
| | ReleaseFast | Debug (what `zig build` installs) |
|---|---|---|
| a switch arm in `src/pardes.zig` (14.6k lines) | 8.8 ms | 87 ms |
| `std.` — 91 candidates, each alias-resolved into the stdlib | 26 ms | 204 ms |
It is a worker thread, so the editor does not block; but the second press of
Tab joins the first query on the UI thread (`old.cancel(io)` in the shell) and
that wait is real. Pre-existing and shared by every LSP kind — not this
feature's to fix, but it is what a fast double-Tab feels like.
Known limitations, in the order you will meet them:
- **`@This()` anywhere in a container makes the whole container unresolvable**,
so `var list: std.ArrayList(u8) = .` — the most common decl literal in this
codebase — answers nothing. This is not the completion filter: `hover` and a
plain field access on the same struct return nothing either. It is the case a
user hits first, and it is upstream of everything here.
- **Only the break AT THE CURSOR is repaired.** Zig's error recovery runs
forward, so an unrepaired break earlier in the file swallows the declaration
the cursor is in and the answer is empty. While typing you normally have one
broken spot, which is the case this works for.
- **A dependency module** (`@import("vaxis")`) cannot be typed at all, for the
same reason `gd` on `vaxis.init` finds nothing.
- **`error.`** is not handled — the position context is `.error_access`, which
no branch claims.
## The keymap is helix's, exactly
Verified against `helix-term/src/keymap/default.rs`, not from memory.
| keys | command | notes |
|---|---|---|
| `gd` | definition | jumps on a single result, lists on several |
| `gD` | declaration | |
| `gy` | type definition | |
| `gi` | implementation | |
| `gr` | references | |
| `SPC l k` | hover | opens `+Hover` |
| `SPC l r` | rename | tag input; applies current-file references in one undo step |
| `SPC l a` | code action | |
| `SPC l h` | select references | |
| `SPC l s` / `SPC l S` | document / workspace symbols | `S` takes a query |
| `SPC l d` / `SPC l D` | document / workspace diagnostics | |
| `]d` / `[d` | next / prev diagnostic | steps the list, asks for one if absent |
| `]D` / `[D` | last / first diagnostic | |
| `=` | format | |
| `Ctrl`+left-click | definition | the mouse spelling of `gd` |
| `Tab` in INSERT mode, right after a `.` | completion | what could go here, and where each of those is defined |
Tab is the one key here that is not helix's and not a goto. helix's `Tab`
completes; pardes's shows you the CANDIDATES' DECLARATIONS in a `+Search`
buffer and inserts nothing, because that is what a seam returning locations can
honestly do — see below. It only diverts where an answer is possible: on a
terminal, in an output buffer, or in a file the backend does not speak
(`lsp.speaks`, which the core asks and the backend answers), Tab indents
exactly as it always did. A Tab that silently does nothing would be worse than
not having the feature. Nothing about the mode changes either — the pane is
still in insert, so typing goes on and walking the answer with `n` means
pressing `Esc` first, the same as for every other results buffer.
Ctrl-click rides the ordinary left-click drag rather than firing on the press:
a click does not place the modal cursor until RELEASE, so a query asked at
press time would answer about wherever the cursor previously sat. A ctrl-DRAG
still selects, and still asks about where it started.
**The gotos are helix's exactly; the leader commands are helix's letters under
an `l` prefix.** `g` and `[`/`]` had no conflicts — `gd/gD/gy/gi/gr` and
`]d/[d` were all free, so they stay where helix puts them. The bare `<space>`
letters were NOT free, and an earlier pass took them anyway, displacing `SPC d`
(Del), `SPC k` (Kill), `SPC s?` (Dump/Restore) and `SPC h t` (Tutor). That is
the wrong trade: those are pardes's most-pressed keys and predate the language
work, whereas an LSP command is something you reach for deliberately and can
afford one keystroke more. So every one of them keeps helix's own letter and
gains the prefix — `<space>k` becomes `SPC l k`, `<space>d` becomes `SPC l d`
— and nothing pardes had moved at all.
Two more live in the same group because they belong to it, not to helix:
| keys | builtin | what it shows |
|---|---|---|
| `SPC l i` | `Lspinfo` | which ZLS, which stdlib (and whether it opens), what the backend answers and refuses, plus the last 24 queries with timings, row counts and **the errors `query` swallowed** |
| `SPC l w` | `Lspwhy` | why the definition query at the cursor answers what it does |
These exist because of the seam's own contract: a backend never fails loudly,
which is right for an editor — a thrown analyser must not take the process with
it — but it makes a broken backend and a correct one that found nothing look
identical. `Lspwhy` narrates the REAL resolution path (the position context is
the analyser's own answer, threaded out through a trace) rather than
re-deriving it beside the code, because a debug view that reimplements the
logic is one that can disagree with it. `Lspinfo` answers from ANY pane,
including one with no file, since it is about the backend rather than a
document — which matters precisely when the pane you are in is the problem.
`K` is **not** hover — in helix it is `keep_selections`. It was checked; do not
"fix" it.
## Writing a backend
`src/lsp/lsp.zig` is the seam. An implementation supplies three things and touches
nothing else (`backend_name` below is the SEAM's, not yours — it is a literal in
`lsp.zig` naming whichever backend was compiled in):
```zig
pub fn query(gpa, arena, req: Req, out: *std.Io.Writer) void
pub fn speaks(path: []const u8) bool
pub const supports: std.EnumSet(Kind)
pub const backend_name = "..."
```
`supports` says what a backend can do; `speaks` says what it can do it TO, and
exists for the one key that must not be eaten when the answer is no — see the
Tab note above.
`query` runs on a worker thread with no access to the core — everything it may
read is in `req` (`path`, `source` (NUL-terminated), `offset`, `arg`, `root`).
`out` is a plain `std.Io.Writer`: the shell owns the buffer behind it (an
`Io.Writer.Allocating`), so a backend never allocates the result, never frees
it, and cannot get the allocator wrong. `arena` is freed wholesale on return;
`gpa` is for a backend's own scratch. Use `lsp.row()` to emit a location,
`lsp.spanRow()` for one that carries the RANGE it matched (the
`path:LINE:COL-ENDCOL` form above), `lsp.rel()` to spell a path against
`req.root`, `lsp.lineCol()` to convert an offset, and `lsp.edit()` for each
half-open range of a rename response. Location
rows are byte-identical across backends; rename ranges are consumed by the core
and never rendered. `rel` allocates nothing — it returns a slice of what you
hand it.
## How the implementations are judged
`zig build lspbench` — same harness, same corpus, same 22 probes, every backend.
The corpus is pardes's own `src/`, plus `test/lspfixture/`: five of the probes
point at fixtures rather than at real source, because on clean, already
formatted code the correct answer to `diagnostics` and `format` is nothing, and
that is indistinguishable from a backend that has neither. `broken.zig` carries
an unused local and a misformatted fn; `dotcomplete.zig` and `dothalf.zig`
carry the two shapes of half-typed dot. The 22 probes cover 15 of the 17
`lsp.Kind`s
— `definition` three times, `document_symbols` twice, `completion` five times,
and the two introspection kinds (`status`, `explain`) not at all.
- **Feature completeness.** Which kinds return rows, and whether the rows
contain what they should. The harness trusts *results*, not the `supports`
flag: a kind claimed but returning nothing is reported as `CLAIMED-EMPTY`,
and a kind that answers without claiming is `unclaimed-works`. Correctness
is a substring the rows must contain, so returning a confident wrong location
scores worse than returning nothing.
- **Latency.** `cold` (first query, index construction included) and `warm`
(median of 20). They differ by orders of magnitude for an indexing backend
and both matter: cold is what the first keypress costs, warm is what every
one after it costs.
- **Memory.** Peak RSS delta (`VmHWM`), so a backend that frees its index
before returning still pays for having built it.
- **Lines of code.** Not measured by the harness — it is `jj diff --stat`
against the base commit. Less is better, and vendoring a library is not free
but is charged in build time and dependency surface rather than in lines we
maintain.
The user-visible rename contract is also pinned through the actual TTY,
leader prompt, worker and ZLS backend by `test/snapshots/lsp-rename.snap`: both
resolved occurrences change, a shadowed local does not, and undo/redo treats
the response as one transaction.
Run `zig build lspbench -- --json` for machine-readable output.
|