<feed xmlns='http://www.w3.org/2005/Atom'>
<title>cloud9.git, branch main</title>
<subtitle>9p for zig</subtitle>
<id>https://git.0x4200.cafe/cloud9.git/atom?h=main</id>
<link rel='self' href='https://git.0x4200.cafe/cloud9.git/atom?h=main'/>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/'/>
<updated>2026-09-25T20:52:33Z</updated>
<entry>
<title>9agents: /active/&lt;h&gt;/&lt;pid&gt;/chat, the conversation as one file per message</title>
<updated>2026-09-25T20:52:33Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-25T19:40:49Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=1c1b192d4ef59199a7196229d56901b1f6512678'/>
<id>urn:sha1:1c1b192d4ef59199a7196229d56901b1f6512678</id>
<content type='text'>
Each message of a claude or codex agent's transcript is a file named
&lt;id&gt;-&lt;kind&gt;: the id is its position in the transcript from 0, as 8
digits so the names sort; the kind is user, assistant, thinking, system
or agent, and a tool call is named after its tool (00000002-bash) and
its result after the call (00000003-bash-result). A file is the message
as text (a tool call: its name, then its input); its mtime is when it
was said.

chat.zig indexes a transcript incrementally, keyed by (dev, ino) and
checked by birth time and a hash of its first and last indexed bytes, so
a file rewritten in place is indexed again. It holds offsets, never
bytes: every read reads the file, and a read resumes where the last one
stopped so a long message is not decoded from its start each time. The
indexes and the runner live in lazily backed mappings (NORESERVE,
NOHUGEPAGE), and the fid table is 32768 so a kernel mount can hold every
message of a long chat.

A 9agents started inside a user namespace (from a 9ns-wrapped terminal)
cannot read /proc/&lt;pid&gt;/fd of processes outside it, so the fd route
never resolved codex there; /active now also finds a codex session
through the thread writer locks it holds (/proc/locks), via = lock.

Fixes to /active found on the way: a fifo in place of a transcript or
record no longer blocks the daemon; a transcript is opened from its
pinned root a component at a time, so a symlinked directory or a
sessionId with .. cannot reach outside it; slot generations are 24 bits,
so a stale node id cannot come to name another agent; listing a harness
rescans; an agent that had not resolved yet is asked again.

Co-Authored-By: Claude Opus 5.5 (1M context) &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>fs, serve: set an engine up in place, so a large fid table costs only the fids used</title>
<updated>2026-09-25T20:52:33Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-25T20:10:23Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=73602127d15d10a1932b6fe916bd18a608054980'/>
<id>urn:sha1:73602127d15d10a1932b6fe916bd18a608054980</id>
<content type='text'>
A kernel mount (9ns) holds a fid for every inode the kernel caches, so a
server that wants `ls -l` of a directory of thousands of files to work
needs a fid table of tens of thousands. Server.initIn sets an engine up
in place and, with fid_index, never writes a fid slot at or past
high_water; the walks over the table (references, reset, orphan) stop
there. Runner.init sets each connection's small fields one by one and
leaves the engine to start(), so connection slots nobody uses are never
written.

Co-Authored-By: Claude Opus 5.5 (1M context) &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>Rename 9harness to 9agents; README, file-backed qids, worded errors</title>
<updated>2026-09-25T20:52:33Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-22T13:30:02Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=df20863879fe2d83077534f4726a985ffc239def'/>
<id>urn:sha1:df20863879fe2d83077534f4726a985ffc239def</id>
<content type='text'>
- 9harness/ becomes 9agents/ (build option -D9agents, package paths).
- 9agents serves /README, reports qid paths from the file's (dev, ino)
  and qid versions that move with the file, and names its refusals.
</content>
</entry>
<entry>
<title>fs: a parked job whose fid was clunked is answered, not asked again</title>
<updated>2026-09-22T14:50:09Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-22T14:50:09Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=eb1a104f385f375a74319695e1a59f6f82e6384e'/>
<id>urn:sha1:eb1a104f385f375a74319695e1a59f6f82e6384e</id>
<content type='text'>
A client that gives up on a parked open sends Tclunk for its fid without
a Tflush; the fid goes, the parked job stays. On retry the job went back
to the backend, which did the work -- opened a handle, made an object --
and the reply then found no fid and dropped it, handle and all. Now the
retry looks for the fid first and answers "fid unknown" without asking.

Found by an adversarial review of pardes's use of the engine.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>9ns --mntgen: a server that never answers stalls only its own name</title>
<updated>2026-09-22T14:39:25Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-22T14:18:05Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=b7fc01550c7bde290cf14276d94193b5b4031dc8'/>
<id>urn:sha1:b7fc01550c7bde290cf14276d94193b5b4031dc8</id>
<content type='text'>
Opening a fish (self-wrapped in `9ns --mntgen`) and running an agent in it
would sometimes freeze the whole session: no input reached it and nothing
under /mnt/9p answered, until the shell was killed from outside. The cause
was one posted server that accepted a connection and then never spoke 9P —
pardes, answering its 9P from the same loop that was walking its own mount,
was the one on this machine, but any wedged or half-dead server does it.

Three things conspired, and each is fixed on its own:

* The dispatcher dialed. A LOOKUP of an undialed name ran connect, Tversion,
  Tattach and Tstat on the one thread that reads /dev/fuse, so while that
  server kept quiet no request for any name was read, and no FUSE_INTERRUPT
  either. Now the dispatcher makes a Mount without touching the network and
  queues the walk to the mount's worker, which dials while serving it. The
  dial is the request in flight, so an interrupt of the walk abandons it at
  once (`Session.abort_on_cancel`: nothing to flush before a session
  exists) and the walk answers EINTR; a failed dial leaves the mount
  undialed for the next walk to retry; a full listen backlog (the server
  stopped accepting) is retried for 5s and then EIO. Every later LOOKUP of
  the name goes through the same queue and is answered from the remembered
  root attr, so the dispatcher never holds a session at all.

* Once the dispatcher had read a request the process behind it was
  unkillable (FUSE waits out a request userspace has taken), and an
  INTERRUPT for a request still sitting in a mount's queue was dropped. The
  dispatcher now takes a queued request out and answers EINTR itself, and
  forwards only in-flight ones to the worker; queue and in-flight unique
  are read under the mount's mutex, where the worker moves a request from
  one to the other. A Tflush the server never answers is given 3s
  (`Session.flush_grace_ms`) and then the session is declared wedged: the
  request answers EINTR, the mount dies, the next walk makes a new one.

* The kernel serialized the directory. Without FUSE_PARALLEL_DIROPS in the
  INIT reply every LOOKUP and READDIR in a directory takes its inode lock,
  so one parked walk held up every other name under /mnt/9p however free
  the dispatcher was (`cat` sat in fuse_lock_inode). The flag is now
  negotiated when the kernel offers it.

What remains is the kernel's own serialization of lookups of one *name*: a
second walker into the parked name waits for the first walk to end, and
only then proceeds (and can be interrupted in its turn).

An adversarial review of the above found three more things, fixed here:
the single-connection bridge's one-slot stash stopped polling the FUSE fd
while a second request was parked, so an INTERRUPT could not arrive (and
parallel dirops make a second request routine) — the stash is now a queue
of copies and the fd is always watched; a dead or wedged mount kept its
socket open until exit, where a late-answering single-threaded server
could block on it — the session is closed when the mount dies; and
teardown after DESTROY or ENODEV (the child still alive, so stop_fd says
nothing) could join a worker parked on a mute server forever — the
sockets are shut down before the join. The flush grace is a deadline now,
not a timer restarted on every wakeup. A black-box run against the binary
(hostile servers: mute, garbage, close-after-accept, full backlog, 100
mute names, interrupt storms, 300 deaths of one server) found that a dead
mount kept its socket, its interrupt pipe and a megabyte of buffers until
exit — three descriptors per death — so `retire` now frees all of it and
keeps only the slot; descriptors, threads and RSS stay flat across 400
deaths. The 4096-slot cap per process remains and is documented.

Reproduced with a socket that accepts and never writes, posted beside
9agents in a scratch registry: before, `cat /mnt/9p/agents/pid` parked
behind `stat /mnt/9p/hang` and SIGINT did nothing; after, it answers at
once, the parked walker dies of its signal within milliseconds, and a
server that answers the handshake but ignores reads and Tflush releases
its reader after the grace. mntgen.sh and adv_bridge_interrupt.sh now
check exactly that; nine.zig gains unit tests for the grace and the
abort. Also in this change: the uncommitted ESTALE-on-death and
FUSE_NOTIFY_INVAL_ENTRY work from the working copy, which the dead-mount
path here builds on.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>fs: let open, truncate, clunk and remove park on again, not only reads and writes</title>
<updated>2026-09-22T14:17:13Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-22T14:17:13Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=1f3aff78702b65c328384bc5b422c448751e809b'/>
<id>urn:sha1:1f3aff78702b65c328384bc5b422c448751e809b</id>
<content type='text'>
A backend that is not ready for a request answers `again` and the engine
parks the request, so the connection goes on answering everything else and
a retry asks the parked request later. That only worked for a read, readdir
or write; every other op answering `again` failed with EAGAIN on the spot.

pardes needs the rest. Its 9P is answered on the connection's task while
the editor's own thread may be out in a syscall in the middle of a step, and
in that window a request that would change a pane -- an open of /pane/new, a
truncating open or wstat, a remove -- has to wait, not fail, and has to wait
without holding the connection up, because the editor's own request may be
the next frame on that very connection (it is, when the syscall goes through
a mount of the editor's own tree). Parking is exactly that.

A parked open, wstat, clunk or remove goes back into the job slot on retry
and its reply then goes through `jobReply` like any other, which is why the
slot keeps only what an open needs (`omode`, `step`); a wstat parks only as
the truncation to zero a Linux client sends for O_TRUNC, and a clunk parks
before it lets its fid go, since a release that was never paid still owns
its handle. A walk, attach, stat, create or renaming wstat keeps names in the
input frame that parking lets go of, so those still answer EAGAIN. A parked
job retried and parked again sits the round out like a retried read does --
without that the first version of this looped forever in `retry`.

Co-Authored-By: Claude Fable 5.1 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>Package 9harness: the published tarball was missing it</title>
<updated>2026-09-21T20:15:19Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T20:15:19Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=9c4d668c926e2d9eae6ef87cd9b391d8185605f9'/>
<id>urn:sha1:9c4d668c926e2d9eae6ef87cd9b391d8185605f9</id>
<content type='text'>
build.zig @imports each program's build fragment unconditionally, so a
program absent from .paths builds fine from a checkout and fails for
every consumer with

  zig-pkg/cloud9-.../9harness/build.zig: unable to load 'build.zig': FileNotFound

9ns and 9proc were listed; 9harness was not, from the change that added
it. A pardes build against cloud9 main found it — the first consumer to
try. Nothing in the suites covers "the package a dependent fetches
actually builds", which is why it went unnoticed.
</content>
</entry>
<entry>
<title>9harness: drop the zmxify script; the mount is the replacement</title>
<updated>2026-09-21T19:56:35Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T19:56:35Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=534c084fb7ab8e3776b25017c901899964490a5d'/>
<id>urn:sha1:534c084fb7ab8e3776b25017c901899964490a5d</id>
<content type='text'>
The script listed /active, ran fzf over it and wrote the chosen name
into an agent's zmx. Everything in it but the picker, a free-name loop
and the attach was `ls` and `cat` against the tree — 130 lines
re-deriving what four lines of shell already read:

  ls /mnt/9p/harness/active/*/*
  cat /mnt/9p/harness/active/claude/345104/session
  echo mine &gt; /mnt/9p/harness/active/claude/345104/zmx
  zmx attach mine

zmxify was 307 lines of rc because there was no filesystem to ask. The
answer to that is the filesystem, not a smaller script on top of it,
which would only become a second and worse interface beside the real
one. DESIGN.md keeps what a caller now arranges itself (a free name,
the attach) and why the old script's own-ancestry rule is gone rather
than moved.
</content>
</entry>
<entry>
<title>9harness: /active, and zmxify as a write to a file</title>
<updated>2026-09-21T19:49:20Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T19:49:20Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=dddd556accea6b6ea7802cd3f622f8b3cf8eb43f'/>
<id>urn:sha1:dddd556accea6b6ea7802cd3f622f8b3cf8eb43f</id>
<content type='text'>
The mirror answers what files exist. /active answers what is running:
one directory per live agent, normalized across harnesses, fields as
small text files, synthesized per request.

  /active/claude/345104/{pid,cwd,session,via,name,status,title,
                         model,started,zmx,transcript,agents/}

This is /proc's shape, and deliberately: a directory per object named
by pid under a directory per harness, rather than a compound
`claude-345104` that would make you parse a name to recover a field
that is already the directory above it. There is no `updated` file —
that is the mtime of `transcript`, which stat already carries.

Each harness is asked in its own terms, and the route is reported in
`via` so a wrong guess is visible rather than silent. Claude Code
publishes sessions/&lt;pid&gt;.json itself, with procStart as a pid-reuse
guard, so nothing there is guessed. omp and dsh are found by the
transcript they hold open, omp falling back to the store named after
its cwd. codex's rollout file carries the session id in its *name*, so
its sqlite is never opened. hermes is the one gap and needs none: its
sessions live only in sqlite, and the only hermes processes that run
are the gateway and the dashboard, which are not sessions.

Liveness is /proc/&lt;pid&gt; plus a matching start time: a pid alone is not
an identity. The daemon never lists its own ancestry, so it cannot show
or act on the tree serving the request.

The write path, and why it is a file and not a ctl: writing a zmx
session name into an agent's `zmx` moves it there. The file means which
zmx session this agent lives in, and writing makes that true. A ctl
taking verbs is the ordinary Plan 9 spelling, and an executable script
served in the tree is the spelling zmx's own `attach` uses, but a
script that shells out to a local binary lies over a remote mount — it
would run against a session that is not on the client's machine. A
write is served where the authority is.

Every refusal comes before anything is destroyed: the name must be
zmx's label charset, unused by a live session, and the agent's session
must have resolved, because nothing is killed that has nowhere to come
back to. The command is fixed per harness and no client byte reaches
exec. It is off unless --allow-move: this is the one place the tree is
not read-only, and anything that can mount it could otherwise kill an
agent.

9harness/zmxify replaces the 307-line rc script. It parses no /proc,
opens no fd table and queries no database; it lists /active, offers the
rows to fzf and writes the chosen name. It no longer excludes the
caller's own session, which the old one had to: that script did the
killing itself, so killing its own parent lost the session it was
rescuing. The daemon completes the kill and the re-exec whether or not
the client is still connected — verified by hanging up immediately
after sending the write — so zmxifying the terminal you are sitting in
now works, which is the common case.

--proc DIR is the fixture seam: the scan, the liveness guard, the
exclusions and the ancestry rule are unit-tested against a fake process
tree, never the live one.

Suites: 87/87 root, 48/48 9ns, 26/26 9harness (+5 for the view),
60/60 9proc, 144/144 programs-test, 51+88 9ns integration, 44/0
9harness end-to-end (+16, including a real move against a fake harness
and a fake zmx), 29/0 9proc debug, 213/0 9ns adversarial, freestanding
green.
</content>
</entry>
<entry>
<title>9harness: the harness fs daemon</title>
<updated>2026-09-21T18:20:40Z</updated>
<author>
<name>Gabriel Schneider</name>
<email>gbrls@0x4200.cafe</email>
</author>
<published>2026-09-21T17:23:27Z</published>
<link rel='alternate' type='text/html' href='https://git.0x4200.cafe/cloud9.git/commit/?id=b4db588dd5b92d647b661c2dc17b40925af92348'/>
<id>urn:sha1:b4db588dd5b92d647b661c2dc17b40925af92348</id>
<content type='text'>
The last item of the 9P plan, replacing zmxify's introspection half: a
read-only, fresh-from-disk 9P view of every agent harness's state on
this machine, posted as `harness` like any other service, so a shell
inside a 9ns --mntgen mount reads it at /mnt/9p/harness with no setup.

  /pid /uptime   /claude/{projects,history,skills}
  /codex/{sessions,session-index,history}
  /omp /hermes /dsh     the mirrors
  /skills/{claude,codex,omp}

Nothing is cached: a lookup, getattr or readdir walks the real
filesystem, so a transcript grows as its harness writes it and a new
session appears as soon as its file lands. Writes answer EPERM, and no
name that looks like a credential, key, token or auth store is ever
answered at any depth.

Three findings from the adversarial pass, each with its regression:

- The read path composed &lt;base&gt;/&lt;rel&gt; and opened it in one call, which
  follows symlinks. A name swapped for a link between the walk and the
  read served bytes from outside every pinned root (proved against
  /etc/passwd). Every stat, read and readdir now resolves through
  openIn, which walks from the base one component at a time with
  O_NOFOLLOW, and O_PATH for the intermediates, so no component can
  redirect the walk. O_PATH also keeps a fifo in a root from parking the
  daemon in open(); a read refuses anything but a regular file.
- Joining a child onto an empty relative path returned an uncopied
  scratch slice, so every file at the top of a mirror root (/hermes/x,
  /dsh/x) listed but read back uninitialized stack bytes.
- A directory past the comptime caps was served short, and a short
  listing cannot be told from a small directory. The caps answer NFILE
  now. Staging also stops at the first record that does not fit instead
  of packing a shorter one behind it, which dropped that entry from the
  listing across the read boundary.

Suites: 13/13 unit (fake HOME, never the live roots), 36/0 end-to-end
including the mntgen money shot and the live ~/.claude/.credentials.json
proved unreachable, 131/131 programs-test.
</content>
</entry>
</feed>
