From 841038666a45b40107ee26d1f8e54ed19d0891a0 Mon Sep 17 00:00:00 2001 From: Gabriel Schneider Date: Wed, 29 Jul 2026 00:13:08 -0300 Subject: lspbench: probe diagnostics and format against a BROKEN fixture Plus docs/lsp-evaluation.md: the three backends measured head to head. --- docs/lsp-evaluation.md | 103 +++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 103 insertions(+) create mode 100644 docs/lsp-evaluation.md (limited to 'docs') diff --git a/docs/lsp-evaluation.md b/docs/lsp-evaluation.md new file mode 100644 index 00000000..ae822213 --- /dev/null +++ b/docs/lsp-evaluation.md @@ -0,0 +1,103 @@ +# Three language backends, measured + +Same core, same seam (`src/lsp.zig`), same 17-probe harness (`zig build lspbench`) +over the same corpus. Three jj workspaces, three independent implementations, +one function each. + +| | **A · stdlib** | **B · client** | **C · in-process** | +|---|---|---|---| +| workspace | `pardes-ws-ast` | `pardes-ws-proto` | `pardes-ws-inproc` | +| what it is | `std.zig.Ast` + `AstGen`, hand-rolled scope walk | `zls` as a child process, JSON-RPC over a socketpair | ZLS linked as a Zig module, analyser called directly | +| **probes correct** | 12 / 17 | **17 / 17** | **17 / 17** | +| **false claims** | 0 | 0 | 0 | +| **implementation LOC** | 622 (+307 test) | 721 | **540** (+29 build) | +| new dependency | **none** | a `zls` binary at runtime | ZLS source (path dep) | +| binary size vs base | +6 MiB | +14 MiB | +35 MiB | +| memory | 3.7 MiB | 1.1 MiB **+ 37 MiB in the child** | 5.5 MiB | +| processes | **1** | 2 | **1** | + +### Latency (warm, µs unless noted) + +| operation | A · stdlib | B · client | C · in-process | +|---|---|---|---| +| goto, same file | **22** | 59 | 127 | +| goto, into `std` | 1 433 | **154** | 2 873 | +| goto, 5 000-line file | 2 699 | 182 | 4 181 | +| hover | **24** | 49 | 125 | +| document symbols | **28** | 196 | 89 | +| references | 6 624 | **65** | 168 | +| workspace symbols | 15 042 | **985** | 6 746 | +| diagnostics | 37 | **9** | 78 | +| **first query of the session** | 32 µs | **33 ms** | 4.4 ms | + +## What the numbers say + +**A is fastest at the thing you do most and slowest at everything that needs an +index.** A same-file `gd` in 22 µs is below the frame budget by three orders of +magnitude, and cold *is* warm because there is no index to build — but +references (6.6 ms) and workspace symbols (15 ms) are linear walks that will +grow with the tree. It answers 12 of 17 kinds and is honest about the other +five: field access, method calls and generics need a type resolver, and it is +not one. `gy`/`gi`/`=`/`SPC a`/`SPC r` stay inert rather than guessing. + +**B is the only one with a real index, and the only one that generalizes.** +References at 65 µs and workspace symbols at 985 µs are 100× and 15× the others +because zls maintains state across queries and pardes does not. It pays 33 ms +once per session for fork+exec+handshake, then 59–182 µs forever. Nothing in +`lsp_client.zig` knows Zig except a binary name and a `languageId` string — +pointing it at rust-analyzer or gopls is a table edit. Its costs are a second +process, 37 MiB of RSS that this harness does not charge it, and an external +binary that must exist and match the toolchain. + +**C is the smallest and the most complete.** 540 lines buy all 17 probes with no +subprocess, no JSON, no handshake and no version skew — the analyser is compiled +in. This is the "gut ZLS and call it directly" idea, and it works. Its latency is +middling and flat for the same reason A's is: it builds a `DocumentStore`, +answers, and throws it away, so there is no cache to be warm. A cross-query cache +is the obvious next step and is a real design (a global, a mutex, an invalidation +story), not a line of code. It costs +35 MiB of binary and couples pardes to a +ZLS checkout. + +## Verification notes + +Every number above was re-measured independently, not taken from the +implementers' reports. Three corrections came out of that: + +1. **The harness was wrong, and B caught it.** The diagnostics and format probes + originally pointed at clean, already-formatted source, where the correct + answer is nothing — indistinguishable from a backend that has no diagnostics. + That rewarded A and C for emitting a filler "no diagnostics" row and scored B + with 3 false claims for answering honestly. `test/lspfixture/broken.zig` now + gives those probes real work; B scores 17/17 on the corrected harness. +2. **A's unit tests never ran.** 23 tests compiled but were not collected — + `zig build unit-test` roots at `pardes.zig`, and Zig only collects tests from + files it actually analyzes, so the lazy `pub const lsp = @import(...)` reached + nothing. Once wired, 16 of 23 crashed on an invalid free in a test helper that + passed one arena as both the `gpa` and `arena` parameters. Fixed; 71/71 pass. +3. **C fuzzed itself and found two real crashes** (a decl's name token indexes + its own file, not the requesting one; and it is not necessarily an + `.identifier` on a half-typed line). Both fixed before reporting. + +All three keep the 56 existing snapshot scripts green, add a 57th driving the +real binary through the helix keymap, and pass `hxdiff` (360) and `hxparity` +(440) with no new waivers. + +## Recommendation + +**C (in-process), with B as the answer to a question pardes has not asked yet.** + +On the stated criteria — fewest lines, feature completeness, performance — C +wins: it is the smallest implementation, it answers everything, it needs no +second process, and its weakest number is a cache that does not exist rather +than a limit that cannot be lifted. + +The one thing that should override that: **pardes ships tree-sitter grammars for +25 languages.** C is Zig-only forever; B is the only implementation that will +ever answer `gd` in a Rust or Go buffer. If language intelligence is meant to +follow the syntax highlighting, B is the strategic choice and its 721 lines are +the cheapest multi-language client anyone will write. + +A is not the answer to "add LSP support", but it is a genuinely good answer to a +different question — 622 lines and zero dependencies for instant same-file +navigation and real semantic diagnostics. It is the one to keep if the ZLS +coupling in B and C ever becomes a problem. -- cgit v1.3