summaryrefslogtreecommitdiff
path: root/9ns/README.md
blob: 7d6babeb178694f0af5c50eacc47741b695fbe28 (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
# 9ns

Mount a 9P2000 file tree into a fresh mount namespace and run a program in it,
as a plain user, without touching the host's mount table.

```sh
9ns --unix /run/user/1000/acme -- fish        # a shell that sees the tree at /mnt/9p/acme
9ns --tcp 127.0.0.1:564 -- claude              # an agent that sees it at /mnt/9p/tcp-127.0.0.1-564
9ns --spawn '9proc-demo --stdio' -- bash       # start the server yourself; /mnt/9p/9proc-demo
9ns --unix /tmp/9debug.sock --name dbg -- bash # pick the name: /mnt/9p/dbg
```

Inside, the tree is ordinary files: `ls`, `cat`, `echo x > ctl`, editors,
`find`, `rsync`, whatever. Every mount lives under `/mnt/9p/<name>`, where
the name comes from `--name` or, by default, from the transport (the socket's
basename, `tcp-IP-PORT`, the spawned command's basename, `fdN`); `--mount
PATH` puts it anywhere else. `$NINE_MOUNT` tells programs where it is. When
the program exits, 9ns exits with its status and the namespace, mount and
connection disappear. Nesting works: `9ns --name a -- 9ns --name b -- fish`
gives a shell that sees both `/mnt/9p/a` and `/mnt/9p/b`.

## How it works

The kernel's own `9p` filesystem cannot be mounted inside an unprivileged user
namespace, so 9ns is a small FUSE server that speaks 9P2000 to the real
server. There is no libfuse and no libc: `src/fuse.zig` implements the subset
of the kernel FUSE protocol needed, straight from `linux/fuse.h`.

```
 program (fish/bash/claude)      9ns (parent)                9P server
 in a new user+mount namespace     │
   /mnt/9p/<name> ─FUSE─▶ kernel ─▶│ fuse.zig ─▶ bridge.zig ─▶ nine.zig ──▶ unix / tcp / socketpair
                                   │ (framing)   (translation) (cloud9 Client)
```

1. The parent connects to the 9P server (version + attach) so failures are
   reported before anything is forked.
2. The child does `unshare(CLONE_NEWUSER|CLONE_NEWNS)`, maps its own uid/gid,
   makes every mount private, opens `/dev/fuse` (it must be opened inside the
   new user namespace), mounts it on the mountpoint and passes the descriptor
   back to the parent over `SCM_RIGHTS`, then execs the program.
3. The parent serves FUSE requests by translating them into 9P transactions
   (`Twalk`, `Topen`, `Tread`, `Twrite`, `Tcreate`, `Tremove`, `Tstat`,
   `Twstat`) until the child exits or the namespace disappears.

Files are opened with `FOPEN_DIRECT_IO`, so synthetic files that report length
0 (the 9P convention for control files) still read correctly, and `O_TRUNC`
travels inside the 9P open mode (`OTRUNC`) rather than as a separate
truncate. Repeated lookups of the same qid map to the same inode.

The mountpoint `/mnt/9p/<name>` normally does not exist and cannot be
created by a plain user. 9ns then walks down the path to the deepest existing
directory (`/mnt` on a host without `/mnt/9p`, `/mnt/9p` if the host has one),
mounts a tmpfs over it *inside the namespace only*, bind-mounts every existing
entry of that directory back into the tmpfs so nothing is hidden, and creates
the missing components inside. Inside a 9ns namespace `/mnt/9p` is already
that writable tmpfs directory, so a nested 9ns just adds its own name next to
the outer mount (the same name mounts over it). Pass `--mount DIR` to use any
other directory.

## Building and testing

9ns lives in the [cloud9](../) repository as `cloud9/9ns/`, beside
the 9P2000 protocol library it is built on, and is wired into cloud9's
`build.zig` through the fragment `9ns/build.zig`. Everything is run from
the cloud9 root with Zig 0.16:

```sh
zig build                  # zig-out/bin/{9ns,9proc-demo,9web,cloud9-probe}
zig build 9ns              # build and install only zig-out/bin/9ns
zig build 9ns-test         # unit tests (protocol structs, session, bridge, namespace helpers)
zig build 9ns-itest        # integration tests: real namespaces, real FUSE,
                           # 9proc-demo over unix/tcp/socketpair, and plan9port's
                           # ramfs when /usr/lib/plan9/bin/ramfs is installed
zig build 9ns-adv          # adversarial suites: hostile 9P servers, FUSE semantics,
                           # namespace/signal edge cases, stress (several minutes)
zig build -Doptimize=ReleaseSafe
zig build -D9ns=false      # leave 9ns out (the default on non-Linux targets)
```

The integration suites mount the `9proc-demo` server (`../9proc`),
so they need `-D9proc=true` (the default on Linux), unprivileged user
namespaces (`kernel.unprivileged_userns_clone=1` on distributions that have
the knob), `/dev/fuse` and Python 3; they skip themselves otherwise.
`zig build programs-test` and `programs-itest` run the unit and integration
steps of every program in the repository.

## Usage

```
9ns [options] -- PROGRAM [ARGS...]

Transport (exactly one):
  --unix PATH          Unix stream socket
  --tcp IP:PORT        TCP (IPv4/IPv6 literal)
  --fd N               an already-connected inherited descriptor
  --spawn CMD          run CMD via /bin/sh -c with a socketpair on its stdin/stdout

Options:
  --name NAME          mount name: the tree appears at /mnt/9p/NAME (one path
                       component; default derived from the transport, see below)
  --mount PATH         mountpoint inside the new namespace (overrides --name)
  --uname NAME         9P user name (default $USER)
  --aname NAME         9P tree to attach (default "")
  --msize BYTES        maximum 9P message size to request (default 131072, max 16 MiB)
  --cache SECONDS      attr/entry cache validity, fractional allowed (default 1)
  --no-direct-io       let the kernel cache file pages (trusts stat length)
  --debug              trace FUSE and 9P operations on stderr
  --help, --version
```

PROGRAM defaults to `$SHELL`. Exit status is the program's (`128+signal` if it
was killed); 125 means 9ns itself failed (usage, connect, namespace,
mount); 126/127 are exec failures as usual.

The default name comes from the transport: `--unix PATH` → the basename of
PATH without a trailing `.sock`, `.9p` or `.socket` (`/tmp/9debug.sock` →
`9debug`); `--tcp IP:PORT` → `tcp-IP-PORT` with every `:` turned into `-`
(`[::1]:564` → `tcp---1-564`); `--spawn CMD` → the basename of CMD's first
word (`/x/9proc-demo --stdio` → `9proc-demo`); `--fd N` → `fdN`. A name must
be a single path component (not empty, no `/`, not `.` or `..`); an invalid
`--name` is a usage error, an unusable derived name falls back to `9p`. With
`--mount PATH` the name is simply PATH's last component. The program gets
`NINE_MOUNT=<mountpoint>` (replacing any inherited value, so a nested 9ns
overwrites it).

## 9proc-demo: a demo 9P server

`9proc-demo` is a single-binary 9P2000 server whose file tree is the binary
itself: build-time facts, `comptime` reflection and live runtime state. It is
the demo of the [9proc library](../9proc) (`../9proc/demo/main.zig`),
built and installed by `zig build 9proc`.

```
/README
/build/{zig_version,target,optimize,time,change}   captured by build.zig (jj change id, UTC time)
/comptime/types/<T>/{name,size,align,fields}        @sizeOf/@alignOf/@typeInfo, generated at comptime
/comptime/decls                                     pub declarations of the server module
/runtime/{pid,ppid,uptime,argv,cwd,env,clients}
/runtime/fn/<name>                                  reading calls a Zig function (hostname, now, random, uname, fib30);
                                                    the directory is generated from @typeInfo of the Fns struct
/runtime/ctl                                        write "fib N" | "add A B" | "echo TEXT" | "sleep-ms N", read the result
/scratch/                                           in-memory read/write tree
```

`/runtime/env` exposes the server's whole environment, so serve 9proc-demo on
a Unix socket or loopback only.

```sh
zig-out/bin/9proc-demo --unix /tmp/intro.sock &
zig-out/bin/9ns --unix /tmp/intro.sock -- sh -c '
  cat $NINE_MOUNT/build/zig_version; echo
  cat $NINE_MOUNT/comptime/types/Qid/fields
  echo "fib 20" > $NINE_MOUNT/runtime/ctl; cat $NINE_MOUNT/runtime/ctl'
```

The same command with `claude -p "explore /mnt/9p/intro ..."` as the program
gives an agent a live, file-shaped view into a running process; that is the
intended use.

## Limitations

* One 9P request is in flight at a time, so a server read that blocks (event
  files, consoles) stalls the mount while it is outstanding. The blocked
  request can be interrupted, though: killing or Ctrl-C-ing the reader makes
  the kernel send `FUSE_INTERRUPT`, which 9ns forwards as 9P `Tflush`;
  servers that honour it release the reader with `EINTR` at once, servers
  that ignore it keep the mount stalled until they answer.
* Base 9P2000 only: no symlinks, ownership, xattrs or locks. Every file is
  reported as owned by the invoking user. Cross-directory rename is `EXDEV`.
* No PID namespace and no `/proc` remount. `--tcp` takes IP literals only
  (no libc, no resolver).
* Linux only.

## Relation to cloud9

9ns consumes cloud9 as the module `cloud9` and keeps all mounting,
namespace and process policy on its side, which is what cloud9's design asks
of applications. It ships from the cloud9 repository as the sibling directory
`9ns/` (sources in `src/`, suites in `test/`, this README and
`docs/DESIGN.md`) with a build fragment that the root `build.zig` enables with
`-D9ns` on Linux targets; nothing in the code depends on that layout.
`docs/DESIGN.md` has the full module contracts.