1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
|
# API and ownership
The public entry point is `src/root.zig`. `Msg` covers every base request and
response, excluding the reserved, nonexistent Terror. Decode returns borrowed
strings/data. Encode checks complete output size before writing; input payloads
may already occupy their exact final position in the output frame, but arbitrary
input/output overlap is not supported. Stat uses the protocol's inner length;
Rstat and Twstat add and validate the outer length separately.
`Client` supports 16 ordinary outstanding requests and one additional flush.
Replies can arrive out of order. Tags remain reserved while flushes are pending,
even when the original request has already completed. Multiple flushes and
flushing a flush are accounted for. A completed result's slices remain valid
until the next `take()` or `hangup()`. `push()` only appends and never compacts.
Submit copies request payloads into the output buffer. An iounit returned by an
open/create is per fid; the caller must apply it when selecting atomic I/O sizes.
`maxRead()` and `maxWrite()` describe frame limits, not that per-fid guarantee.
Renegotiation requires an empty client request window.
`Server` supports 64 ordinary outstanding requests and one additional flush.
Receive errors terminate the connection. A full request table produces NoTags;
output backpressure instead returns null without consuming the next request.
`receive()` yields at most one frame. `release()` ends its borrow and advances
input. A backend retaining any request data must copy it before release.
A successful reply copies output and frees the request tag. An output-capacity
error leaves that tag pending so the backend can drain and retry.
The server validates transaction structure, not filesystem policy. Backends
own fid maps, authentication state, file handles, access modes, complete directory
records, atomic wstat, and cancellations. They must retire canceled work before
reusing its tag; a tag alone cannot distinguish a stale backend completion from a
new request. A version event requires backend cancellation and fid cleanup before
`negotiate()`. A partially sent response drains before a new version is accepted.
An Rflush promises no later reply to oldtag. Replying Rerror to a flush is refused.
Caller-provided input/output buffers must be disjoint and remain stable for the
session. No protocol API takes an allocator. Server and client request tables are
fixed-size. Copies and work per frame are bounded by buffer size and fixed table
capacities; transports and application backends can have their own allocations.
# Transports
`transport.readFrame` and `writeFrame` adapt `std.Io.Reader` and `std.Io.Writer`.
The writer remains buffered until its owner flushes it. Reader errors leave the
stream unsuitable for further frame processing; close the connection.
`transport.connect` and `listen` use `std.Io.net` for TCP and Unix streams. The
POSIX `connectFd`, `listenFd`, `acceptFd`, `read`, `write`, and `wait` APIs fit
caller-owned poll loops. They configure nonblocking descriptors and close-on-exec;
read/write return null for retry, and read returns zero for EOF. `wait` takes an
absolute monotonic deadline. The application closes descriptors, limits its
connections, selects addresses and deadlines, and manages Unix path permissions
and removal. No mounting or namespace policy is performed.
`http.WebSocket` adapts established HTTP streams without owning their sockets
or buffers. `http.Client` negotiates HTTP/1.1 upgrades using `std.http.Client`,
including its certificate-verified HTTPS support. `http.Bridge` relays between
an accepted WebSocket and arbitrary standard readers/writers, including UART
adapters. The runnable application's web UI, device leases, deadlines, and
origin policy remain in `web/`. See [HTTP ownership and usage](http.md).
`Quic(OpenSSL, alpn)` is an optional OpenSSL 3.6+ adapter with a single ordered
bidirectional stream per connection. It preserves the previous Pardes transport:
ephemeral self-signed certificates, no peer authentication. Applications needing
identity verification must provide that policy before using it across a trust
boundary. OpenSSL allocations and handshake costs are outside the allocation-free
protocol core. Accepted connections must close before their shared listener.
# File server engine
`fs.Server(Backend, Options)` is a file server built on `Server`: it owns the
fid table, permission checks, the directory-read cursor, flush and hangup, and
asks a backend only for filesystem operations. The backend contract is
`fs.Req` (`Op`: lookup, getattr, setattr, open, read, write, release, readdir;
each names a node, and open/read/write/readdir/release carry the handle open
returned) answered by `fs.Reply` (`Status`, an errno from `fs.E`, `fs.Attr`,
the open handle, the written count) with a read's or readdir's bytes passed
beside the reply. `fs.ReplyWith(Payload)` is the same reply carrying an
application-defined locator for those bytes, which the engine ignores; a
backend declares `Req` and `Reply` as those types. A readdir answers records
of `node:u64le dir:u8 len:u8 name`, which the engine encodes into whole Stat
entries directly in the output frame.
Every 9P request is a job: attach, walk, open, read, readdir, write, clunk,
remove, stat and wstat become one or more backend requests issued in order
(a walk asks one lookup per element; an open with OTRUNC asks setattr then
open; a clunk of an open fid asks release). `next()` yields the next backend
request and null while one is outstanding, the output is full, or no whole
frame has arrived; the backend answers by the request's tag through
`reply()`, at once or later. Tags are unique per connection and never zero.
A stale tag (flushed, hung up) is ignored; the backend must retire canceled
work before answering.
A read, readdir or write may answer `Status.again`: the job parks in a slot,
the engine moves on, and `retry()` re-issues each parked request once per
round, oldest first, until it completes; a parked write keeps a copy of its
data up to `park_data_max` bytes. Any other operation answering `again`, or
a park with no free slot, fails with EAGAIN. Tflush answers a parked
request's tag with EINTR before its Rflush and drops the slot; a flush of an
unknown tag is just Rflush. `hangup()` marks every open fid orphaned and
`next()` then yields the release each one owes before anything else, so a
dropped connection still pays the backend; a Tversion does the same before
negotiating.
Every table is sized at comptime by `fs.Options`: `fid_capacity` (256),
`slot_capacity` (32), `park_data_max` (128), `name_capacity` (28, at most
255; zero keeps no names, a stat's name is then the one its getattr
answers and the backend judges name lengths), `username_capacity` (28) and
`fid_index` (off: an open-addressing index from fid number to slot, two
bytes per bucket, for tables of thousands of fids; `InitOptions.seed`
salts it); the defaults are the editor's, except the name capacity, which
is the board's. `msize_min` (217) is the smallest negotiable msize, one
full Rwalk. Auth and attaching a named tree are refused, and so are
create, remove and any wstat other than a zero-length truncation unless
the backend declares the feature: a backend may carry `pub const features:
fs.Features` naming `create` (Tcreate becomes an `open` with `create`,
`data` the name, `perm` and `omode`, answered with the new node's `attr`
and handle), `remove` (Tremove, and ORCLOSE on clunk or hangup, become a
`release` with `remove`; its error is the Rremove's), `wstat` (a Twstat
becomes a `setattr` whose `set` names the changed fields: `data` the new
name, `perm`, `mtime`, `length`; the fields the engine owns must be "don't
care") and `references` (every lookup result, "." and a fid clone
included, is a reference the backend gets exactly one `release` for, with
`opened` telling whether an open handle goes with it). An `Attr` may add
a qid `version`, `atime`, a qid `path` distinct from the node, and the
append-only and exclusive-use bits; a reply's `ename` replaces
`errString(errno)` as the Rerror text. Rerror strings the engine emits on
its own are the ones Linux v9fs maps back to errnos (`fs.errString`). The
engine allocates nothing and makes no OS calls; a session is `push()`,
`retry()`, `next()`, `reply()`, `output()`, `wrote()`; `fidCount()` counts
the fids held.
# Io runner
There are two ways to run the engine. A freestanding target, or an
application with its own event loop, drives it directly: `push()` bytes in,
answer every `retry()` then every `next()` request through `reply()`, drain
`output()` and report `wrote()`, on whatever thread and schedule it likes
(the board firmware and the 9proc probe do this). A hosted target may
instead hand the engine to `serve.Runner(Backend, Options, Limits)`, an
`std.Io` adapter beside `http`: it owns listeners (`listen()` takes a
`transport.Address`, Unix or TCP, up to `Limits.listeners`), a table of
`Limits.connections` slots with their buffers (`Limits.msize` sizes them; a
connection past the table is accepted and closed), and two tasks per
connection in one `Io.Group`, one reading frames with `transport.readFrame`
and pushing them in, one stepping the engine and writing its output with
the stream's writer. The engine and the backend contract are untouched by
it, and it makes the engine no less Io-free.
The backend is reached through a `Handler`: `serve(ctx, conn, req)` runs on
the connection's task with the engine unlocked and answers with
`Conn.reply()`, at once or later from any task or thread, or takes the
engine over between `Conn.lock()` and `Conn.unlock()`, drives `next()`,
`retry()` and `reply()` itself and calls `Conn.flush()` so the connection
task sends what that produced (the editor answers on its own thread this
way). `opened` and `closed` bracket a slot's life; `Conn.user` is the
application's.
Wake-up is explicit. Each connection has a flag word and an `Io.Event`; the
reader raises "input", `reply()` and `flush()` raise "output", `Conn.wake()`
(or the runner's `wakeAll()`) raises "retry", and the connection task steps
the engine on every signal, retrying parked requests only for "retry". That
keeps a backend that answers `Status.again` from being polled by its own
replies: it is asked again when it says it has news. Tflush goes through the
engine as always. A connection ends on EOF, a read or write error, a
protocol error (`protocol.dead`), `Conn.close()`, the optional greet
timeout (`InitOptions.greet_timeout_ms`: no Tversion in time) or `stop()`;
the runner then calls `hangup()`, pays the backend every release `next()`
still yields through `serve`, closes the socket and frees the slot.
`stop()` cancels the group (`Io.Group.cancel` interrupts the blocking
accepts and reads), joins every task and closes the listeners; a `serve`
that waits on another thread must stop waiting once `stop()` has begun.
Unix paths bound by `listen()` are the application's: the runner neither
unlinks, chmods nor removes them. `listenPosted(env, name, backlog)` is the
exception with a rule of its own: it posts through the registry (below) —
one posted name per runner, a second `listenPosted` is `error.AlreadyPosted`
(its socket would be orphaned in the registry, nothing left to unpost it) —
and `stop()` unposts, but only while the entry is still the runner's own
(`posted_ino`, below): a name re-posted by another server after this
runner's socket file was lost survives the stop.
# Post registry
`post` is the `/srv` translation: a server posts itself under a name in one
per-user registry directory and clients list the names and dial them, the
shape of Plan 9's `devsrv.c` and plan9port's `post9pservice`. The registry
is `$XDG_RUNTIME_DIR/9p/` (per-user tmpfs, the right lifetime); a name's
socket lives at `$XDG_RUNTIME_DIR/9p/<name>`. There is no union daemon and
no per-server directory: mounting and namespace policy are the client's
(9ns `--mntgen` lists the registry for its synthetic root and dials on
walk).
The library never reads the process environment: the environment comes in
as the raw block `main` receives (`post.Env`, scanned by `post.getenv`,
the same threading 9ns uses). `registryDir`/`registryPath` build paths
into caller buffers, zero-terminated, and refuse when `XDG_RUNTIME_DIR` is
unset — there is no `/tmp` fallback. A name is legal when it would be a
legal file name for the engine (`legalName` in fs.zig: non-empty, not "."
or "..", no '/' or NUL) *and* fits the Unix socket path budget
(`transport.sun_path_len`); path traversal through a posted name is the
attack the caps exist for, and `max_name_len` derives from that budget
against a conforming 32-byte `/run/user/<uid>` prefix, with the real
prefix checked again per call.
`post.post(io, env, name, backlog, path_buf)` posts: the registry directory
is created 0o750 (already present is fine) and the claim runs. The claiming
socket is bound and **listening at a private temp path first** —
`$XDG_RUNTIME_DIR/.post.sock.<pid>.<serial>`, beside the registry, never
inside it, so no listing ever sees it — and only then is the name taken,
entirely by atomic `renameat2` calls under an advisory `flock` on
`$XDG_RUNTIME_DIR/.post.lock` (the kernel drops the lock if the poster
dies; the dotfile persists, empty): a free name is claimed with
`RENAME_NOREPLACE`, which is the arbiter between two posts racing on one
name — exactly one ends up bound, the loser re-runs into
`error.AlreadyPosted`; a stale entry (a socket that refuses a connect) is
grabbed with `RENAME_EXCHANGE` against a private dummy file
(`.post.tmp.*`, also beside the registry) and verified by inode *and* a
fresh probe before the dead socket is removed, so a live server that came
up in between is swapped back untouched. **`post` never unlinks the
registry path**: what a claim displaces is parked under private temp names
and deleted only when provably its own dummy or the verified dead entry —
anything a foreign hand put in their place is left alone, intact. An
entry that is not a socket is `error.NotSocket` and is never touched at
all, learned the hard way in zmx. The lone exception is the legacy
fallback on filesystems without the `renameat2` flags: there the stale
entry is removed by an inode-checked in-place unlink (the replace window
that costs is confined to such filesystems).
The claimed entry's inode comes back in `Posted{server, path, inode}`, and
`post.unpost(io, path, inode)` unlinks only that entry: a path whose
content has been replaced (the socket file lost, the name re-posted by
another server) is left strictly alone, so one server's late stop can
never unpost another's live name. Idempotent, errors swallowed. On the
server side, `serve.Runner.listenPosted` posts and `stop()` unposts.
Probing is a nonblocking raw-syscall connect (`post.probe`:
`.none`/`.stale`/`.live`; uncertainty counts live), because `std.Io`'s
Unix connect does not promise ECONNREFUSED. The kernel answers a connect
to a non-socket entry with the same ECONNREFUSED as to a dead server's
socket, so `.stale` covers both — `post` tells them apart by stat before
anything is removed; callers that must know use `Io.Dir.statFile`.
`post.posted(io, env, out)` lists the raw entries into a caller buffer as
`len:u8 name` staged records (the engine's readdir staging pattern; no
allocation, no connection made — a missing registry lists as empty) and
returns a `Names` iterator; staleness is the caller's concern, and a
buffer too small for the listing is `error.NoSpace`, never a silent drop.
`post.dial(io, env, name)` connects and answers `error.NotPosted` (no
entry) and `error.Stale` (entry present, connection refused) distinctly.
`post.Watch` is a Linux inotify watcher on the registry directory for
hosts that cache its listing (`init`, `add`, `next`, `deinit`; hosted
Linux only — everything else in the module is platform-neutral, and
`next` yields `added` and `removed` per name and, with an empty name, the
two signals a caching consumer must not miss: `overflow` — the kernel
dropped events, rescan — and `gone` — the watch itself ended (the
directory was removed); re-`add` and rescan.
# Multiplexer (9web)
`9web` (`web/`) is a gateway and a 9P multiplexer. It keeps **one** upstream
connection — a single `cloud9.Client`, all sixteen tags — and fans it out to
any number of downstream sessions: the browser WASM client over a WebSocket at
`/_cloud9/9p`, plain 9P clients over an optional `--serve` TCP/Unix listener,
and an HTTP view of the tree under `/fs/`. `web/mux.zig` owns the shared
upstream; `web/httpfs.zig` is the HTTP view; `web/probe.zig` embeds 9proc under
`--probe`.
The mux is a **frame-level remux**, not an `fs.Server`/`serve.Runner` backend.
Each downstream request is forwarded upstream on a remapped tag with its fids
remapped into the shared upstream fid space, and each upstream reply is routed
back by tag. `Tversion` is answered locally (the upstream is negotiated once at
connect); `Tattach` is forwarded, so each downstream gets its own upstream tree
root; `Tflush` is forwarded as `Tflush`. Downstream fid spaces are isolated by
construction — two downstreams never share an upstream fid. The choice is
deliberate: the file-server engine re-decomposes each request into filesystem
operations and re-encodes directories with synthetic stats and an entry-index
cursor, which would lose the upstream's real directory stats, iounit and qids
and turn the readdir byte offset into an index mapping. Forwarding frames
unchanged is exactly the "forward each request upstream" contract, the shape
plan9port's 9pserve has, and it keeps the upstream's bytes intact. `serve.Runner`
remains the right tool for a *server* whose backend is a real filesystem; a
transparent *proxy* is not that.
The shared upstream runs one reader task and any number of downstream forwarder
tasks under one mutex; the tag, fid and pending tables are comptime-sized and
nothing allocates per request. Sockets are only written under the mutex with the
client's own output buffer (disjoint from the stream writer's), and replies are
encoded into per-tag buffers and sent to downstreams outside the lock so a slow
downstream cannot stall the client. The WebSocket and `--serve` downstreams
drive the mux the same way — each reads whole frames (WebSocket binary messages
or `transport.readFrame`) and calls `forward`; the runner's socket-only listener
is not used for them because WebSocket framing and the HTTP `/fs` view do not fit
it. When all sixteen upstream tags are in flight, a further downstream request
waits for one to free (Tflush keeps its reserved seventeenth tag). On an upstream
I/O failure the reader fails every outstanding request, reconnects, bumps a
generation, and answers any downstream request that still names a pre-reconnect
fid with EIO until it re-attaches.
The HTTP `/fs` view uses the same shared upstream at the fid level through a
blocking `rpc`: it allocates upstream fids from the mux pool and walks, stats,
reads and writes directly, so `GET`/`PUT`/`DELETE`, directory JSON, `Range` and
`?follow=1` server-sent events all ride the one upstream connection. A parked
`follow` read is cancelled with an upstream `Tflush` when its client goes away.
# Related programs
Programs built on the library ship from this repository as `cloud9/<name>/`,
sibling directories of `src/` with `9`-prefixed names (`web/` builds `9web`;
`9proc/` and `9ns/` build `9proc-demo` and `9ns`), never inside it: each has
its own `src/`, `test/`, `docs/`, README and a `build.zig` fragment
(`pub fn add(b, ctx)`) that the root `build.zig` imports, passes the resolved
target, optimize mode and the `cloud9` module to, and enables with a
`-D<name>` toggle. Fragments register namespaced steps (`<name>`,
`<name>-test`, ...), never call `standardTargetOptions`, and use root-relative
`b.path("<name>/...")`. The library keeps its contract: mounting, namespaces,
threads, allocation and process policy stay in the program. `9proc` is
also exported as a module next to `cloud9` for dependents. New related
programs follow this layout.
# Pardes integration
Pardes consumes the sibling package through build.zig.zon. Its control tree
(`src/ninep/`) is an `fs.Server` backend using the contract types above; its
`src/9p.zig` names the editor's and the board's `fs.Options`.
`src/9p_io.zig` retains mounting, discovery, the editor's event-loop scheduling,
connection limits, and error presentation; its Unix and TCP listeners run on
`serve.Runner` with a handler that hands each request to the editor's thread,
while QUIC keeps its own poll loop. Protocol bytes, session validation,
the file-server engine, and TCP/Unix/QUIC transport implementation come from
cloud9. Pardes selects the existing `pardes-9p` ALPN for compatibility.
The separate `05-zig-p4` firmware build adds cloud9 to the GPIO application's
module map. UART hardware access remains in the board firmware. No hardware was
flashed by this change.
|