summaryrefslogtreecommitdiff
path: root/docs/lsp-evaluation.md
blob: 6aa5840977d6c71c992347ecf57963f7082fdbf7 (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
# Three language backends, measured

Same core, same seam (`src/lsp/lsp.zig`), same harness (`zig build lspbench`)
over the same corpus — 17 probes at the time, 22 today. Three jj workspaces,
three independent implementations, one function each. C landed; see
**Decision** below.

> **This is a dated report, not a description of the tree.** Every MEASUREMENT
> below is the one taken at decision time and is deliberately left alone,
> because renumbering a comparison nobody can re-run turns a record into a
> guess. What has moved since, re-checked 2026-08-26: `lspbench` now has **22**
> probes rather than 17 (the dot-completion work added five),
> `src/lsp/lsp_zls.zig` has grown from 540 lines to **1 432**,
> `test/snapshots/` holds **95** scripts rather than 56 (`.snap` + `.golden`
> pairs), and the differential suites are `hxdiff` **481** (`cases.jsonl`) and
> `hxparity` **561** (`cases.jsonl` + the 80 editing extras in
> `parity.jsonl`). The one number that IS kept current is the grammar count in
> the Recommendation, because that is an argument about today rather than a
> measurement of then. For what the harness does TODAY, read `docs/typ/building.typ` and the guide's Keys section.

| | **A · stdlib** | **B · client** | **C · in-process** |
|---|---|---|---|
| change | `nvqpknzm` | `stuurqqt` | **landed** |
| what it is | `std.zig.Ast` + `AstGen`, hand-rolled scope walk | `zls` as a child process, JSON-RPC over a socketpair | ZLS linked as a Zig module, analyser called directly |
| **probes correct** | 12 / 17 | **17 / 17** | **17 / 17** |
| **false claims** | 0 | 0 | 0 |
| **implementation LOC** | 622 (+307 test) | 721 | **540** (+29 build) |
| new dependency | **none** | a `zls` binary at runtime | ZLS source (fetched into `zig-pkg/`) |
| binary size vs base | +6 MiB | +14 MiB | +35 MiB |
| memory | 3.7 MiB | 1.1 MiB **+ 37 MiB in the child** | 5.5 MiB |
| processes | **1** | 2 | **1** |

### Latency (warm, µs unless noted)

| operation | A · stdlib | B · client | C · in-process |
|---|---|---|---|
| goto, same file | **22** | 59 | 127 |
| goto, into `std` | 1 433 | **154** | 2 873 |
| goto, 5 000-line file | 2 699 | 182 | 4 181 |
| hover | **24** | 49 | 125 |
| document symbols | **28** | 196 | 89 |
| references | 6 624 | **65** | 168 |
| workspace symbols | 15 042 | **985** | 6 746 |
| diagnostics | 37 | **9** | 78 |
| **first query of the session** | 32 µs | **33 ms** | 4.4 ms |

## What the numbers say

**A is fastest at the thing you do most and slowest at everything that needs an
index.** A same-file `gd` in 22 µs is below the frame budget by three orders of
magnitude, and cold *is* warm because there is no index to build — but
references (6.6 ms) and workspace symbols (15 ms) are linear walks that will
grow with the tree. It answers 12 of 17 kinds and is honest about the other
five: field access, method calls and generics need a type resolver, and it is
not one. `gy`/`gi`/`=`/`SPC a`/`SPC r` stay inert rather than guessing.

**B is the only one with a real index, and the only one that generalizes.**
References at 65 µs and workspace symbols at 985 µs are 100× and 15× the others
because zls maintains state across queries and pardes does not. It pays 33 ms
once per session for fork+exec+handshake, then 59–182 µs forever. Nothing in
`lsp_client.zig` knows Zig except a binary name and a `languageId` string —
pointing it at rust-analyzer or gopls is a table edit. Its costs are a second
process, 37 MiB of RSS that this harness does not charge it, and an external
binary that must exist and match the toolchain.

**C is the smallest and the most complete.** 540 lines buy all 17 probes with no
subprocess, no JSON, no handshake and no version skew — the analyser is compiled
in. This is the "gut ZLS and call it directly" idea, and it works. Its latency is
middling and flat for the same reason A's is: it builds a `DocumentStore`,
answers, and throws it away, so there is no cache to be warm. A cross-query cache
is the obvious next step and is a real design (a global, a mutex, an invalidation
story), not a line of code. It costs +35 MiB of binary and couples pardes to a
ZLS checkout.

## Verification notes

Every number above was re-measured independently, not taken from the
implementers' reports. Three corrections came out of that:

1. **The harness was wrong, and B caught it.** The diagnostics and format probes
   originally pointed at clean, already-formatted source, where the correct
   answer is nothing — indistinguishable from a backend that has no diagnostics.
   That rewarded A and C for emitting a filler "no diagnostics" row and scored B
   with 3 false claims for answering honestly. `test/lspfixture/broken.zig` now
   gives those probes real work; B scores 17/17 on the corrected harness.
2. **A's unit tests never ran.** 23 tests compiled but were not collected —
   `zig build unit-test` roots at `pardes.zig`, and Zig only collects tests from
   files it actually analyzes, so the lazy `pub const lsp = @import(...)` reached
   nothing. Once wired, 16 of 23 crashed on an invalid free in a test helper that
   passed one arena as both the `gpa` and `arena` parameters. Fixed; 71/71 pass.
3. **C fuzzed itself and found two real crashes** (a decl's name token indexes
   its own file, not the requesting one; and it is not necessarily an
   `.identifier` on a half-typed line). Both fixed before reporting.

All three keep the 56 existing snapshot scripts green, add a 57th driving the
real binary through the helix keymap, and pass `hxdiff` (360) and `hxparity`
(440) with no new waivers.

## Decision

**C landed.** It is what `src/lsp/lsp_zls.zig` is, and ZLS is now a fetched
dependency (`build.zig.zon`, pinned to commit `3e0d0820` on the 0.16.x branch)
so it sits in `zig-pkg/` like everything else and the build is reproducible
from the .zon alone — no path dependency on a local checkout. The same commit
is spelled a second time in `build.zig` as `const zls_version =
"0.16.1-dev+3e0d0820"`, for ZLS's own `-Dversion-string`; the semver half
exists nowhere in the manifest, so the duplication cannot be removed, but a
`comptime` block beside the constant `@compileError`s unless the two name the
same commit. See `docs/typ/building.typ`.

A and B were not deleted, only un-worktree'd. They remain whole commits:

- **A · stdlib** — change `nvqpknzm`
- **B · client** — change `stuurqqt`

Both are one `jj new <change>` away if the multi-language argument below wins
later, or if the ZLS coupling ever needs backing out.

> **Addendum, 2026-09-01.** The multi-language argument won. The tree now has
> a REAL protocol client beside the in-process analyser —
> `src/lsp/lsp_client.zig`, routed per file by the seam — but it is a new
> implementation, not `stuurqqt` resurrected: B pumped the socket only while
> a query waited, which cannot carry an indexing server's `$/progress` or
> unsolicited diagnostics, so the client has a reader thread per server and a
> status sink narrating server state onto the transient message row. B's
> transport bones (socketpair, deadline-bounded writes, MSG_NOSIGNAL,
> believe-the-server encoding) survive in it. See "The protocol client" in
> docs/typ/building.typ.

## Recommendation (as written before the decision)

**C (in-process), with B as the answer to a question pardes has not asked yet.**

On the stated criteria — fewest lines, feature completeness, performance — C
wins: it is the smallest implementation, it answers everything, it needs no
second process, and its weakest number is a cache that does not exist rather
than a limit that cannot be lifted.

The one thing that should override that: **pardes ships tree-sitter grammars for
28 languages** — `src/grammar_manifest.zig`'s `all` has 29 entries, but
`markdown_inline` carries `.exts = &.{}` on purpose: no filename ever selects
it, and the block grammar re-enters it by name for its `inline` nodes. So 29
grammars, 28 buffers a language server could ever be asked about. C is
Zig-only forever; B is the only implementation that will
ever answer `gd` in a Rust or Go buffer. If language intelligence is meant to
follow the syntax highlighting, B is the strategic choice and its 721 lines are
the cheapest multi-language client anyone will write.

A is not the answer to "add LSP support", but it is a genuinely good answer to a
different question — 622 lines and zero dependencies for instant same-file
navigation and real semantic diagnostics. It is the one to keep if the ZLS
coupling in B and C ever becomes a problem.