diff options
| author | Gabriel Schneider <[email protected]> | 2026-09-21 14:13:43 -0300 |
|---|---|---|
| committer | Gabriel Schneider <[email protected]> | 2026-09-21 14:13:43 -0300 |
| commit | 3a23f6a29e47ace901bd4d82b9db4055fcc12bb9 (patch) | |
| tree | b82d6e7c3ebe108434ce00ca75db59cf037917e0 /9ns/test | |
| parent | f1b53c1533539aecbf16ad19fd9156deae091f92 (diff) | |
| download | cloud9-3a23f6a29e47ace901bd4d82b9db4055fcc12bb9.tar.gz cloud9-3a23f6a29e47ace901bd4d82b9db4055fcc12bb9.zip | |
post registry + 9ns --mntgen: the /srv translation
cloud9.post: servers post their socket under a name in
$XDG_RUNTIME_DIR/9p (post/unpost, posted, dial, Watch) and
serve.Runner.listenPosted posts a server by name, unposting on stop.
Names are budget-checked against the 108-byte socket path; a claim
binds+listens at a private temp path and takes the name with atomic
renames under flock (RENAME_NOREPLACE for free names, RENAME_EXCHANGE
grab-verify-commit for stale ones): the registry path is never unlinked
by a claim, live names refuse with AlreadyPosted, foreign files with
NotSocket, and unpost removes only the caller's inode-matched entry.
Watch surfaces inotify overflow and a replaced registry dir.
9ns --mntgen [--mount DIR] -- PROGRAM: one FUSE mount at /mnt/9p whose
synthetic root lists the posted registry (no connection made); a walk
into an unmounted name dials it and runs the existing bridge dispatch
in a per-server worker thread, routed by mount index in the node id's
top bits (ordinals never reused, cap 4096); a dead server answers EIO
on its subtree and is re-dialed on the next walk. The dial watches
stop_fd through Tversion (connectWatched). All existing 9ns forms are
unchanged.
9proc's unix listener no longer blind-unlinks its path: a foreign
non-socket is refused (Occupied), a live server is refused
(AlreadyListening), only a refused socket is cleared, and stop()
unlinks only the listener's own inode-matched socket.
Hardened by adversarial review (GLM 5.3 x2 + DeepSeek V4.1 Flash, all
high-thinking): double-bind races on one name (0 in 180k rounds),
foreign-file TOCTOU deletions (0 in 4M flips), a 255-byte-name listing
panic, inotify queue overflow silently dropped, listenPosted silently
overwriting, dial-time Tversion hangs wedging the dispatcher, --debug
silently ignored in mntgen, and xattr/statx probes answering EPERM on
the synthetic root (broke `ls -l /mnt/9p`).
Tests: root 80/80, 9ns 47/47, 9proc 60/60, integration 88/88 +
mntgen 37/37, adversarial 213/0, freestanding riscv32 gate green.
Diffstat (limited to '9ns/test')
| -rwxr-xr-x | 9ns/test/mntgen.sh | 222 |
1 files changed, 222 insertions, 0 deletions
diff --git a/9ns/test/mntgen.sh b/9ns/test/mntgen.sh new file mode 100755 index 0000000..ea58e79 --- /dev/null +++ b/9ns/test/mntgen.sh @@ -0,0 +1,222 @@ +#!/usr/bin/env bash +# Integration tests for 9ns --mntgen: one FUSE mount whose synthetic root +# lists the posted-9P registry ($XDG_RUNTIME_DIR/9p), servers dialed lazily +# on the first walk into their name, one worker thread per server. +# Usage: bash 9ns/test/mntgen.sh <9ns> <9proc-demo> (zig build 9ns-itest) +# Exit 0 on success (or when the machine cannot run the tests), 1 on failure. +set -u + +NS=$(realpath "${1:?path to 9ns}") +PROC=$(realpath "${2:?path to 9proc-demo}") +TMP=$(mktemp -d "${TMPDIR:-/tmp}/9ns-mntgen.XXXXXX") +PIDS=() +FAILED=0 +PASSED=0 +M=/mnt/9p +RAMFS=/usr/lib/plan9/bin/ramfs +cleanup() { + # ramfs delegates to 9pserve, and tests may spawn servers inside the + # namespace: catch anything holding a path under $TMP. + for p in "${PIDS[@]:-}"; do [ -n "$p" ] && kill "$p" 2>/dev/null; done + pkill -f "$TMP" 2>/dev/null + rm -rf "$TMP" +} +trap cleanup EXIT + +if ! unshare -Urm true 2>/dev/null; then + echo "SKIP: unprivileged user namespaces unavailable"; exit 0 +fi +if [ ! -c /dev/fuse ]; then + echo "SKIP: /dev/fuse missing"; exit 0 +fi + +pass() { PASSED=$((PASSED + 1)); echo "ok - $1"; } +fail() { FAILED=$((FAILED + 1)); echo "FAIL - $1"; shift; [ $# -gt 0 ] && printf ' %s\n' "$@"; } +expect_eq() { # name expected actual + if [ "$2" = "$3" ]; then pass "$1"; else fail "$1" "expected: $(printf %q "$2")" "actual: $(printf %q "$3")"; fi +} +expect_contains() { # name needle haystack + case "$3" in *"$2"*) pass "$1" ;; *) fail "$1" "missing: $(printf %q "$2")" "in: $(printf %q "$3")" ;; esac +} + +wait_socket() { # path + for _ in $(seq 1 100); do [ -S "$1" ] && return 0; sleep 0.05; done + return 1 +} + +# A scratch registry: the /srv translation is per-user tmpfs keyed by +# XDG_RUNTIME_DIR; tests must never touch the real /run/user/<uid>/9p. +# mktemp paths are short enough for the 108-byte socket name budget. +export XDG_RUNTIME_DIR="$TMP" +mkdir -p "$TMP/9p" +REG="$TMP/9p" + +# run_in "<shell script>" — inside a namespace with the mntgen mount on $M. +run_in() { timeout 60 "$NS" --mntgen --mount "$M" -- sh -c "$1" 2>"$TMP/stderr"; } + +# ============================================================================ +echo "# two servers posted under two names" +# 9proc-demo posts its socket directly in the registry directory: a socket +# at $XDG_RUNTIME_DIR/9p/<name> IS a posted name (the /srv model). +"$PROC" --unix "$REG/alpha" & +ALPHA=$! +PIDS+=($ALPHA) +wait_socket "$REG/alpha" || { echo "9proc-demo did not post $REG/alpha"; exit 1; } +# plan9port ramfs: an independent 9P2000 implementation; NAMESPACE points its +# post9pservice at the registry (-S names the service; the default would +# post under "ramfs"). +HAVE_RAMFS=no +if [ -x "$RAMFS" ]; then + NAMESPACE="$REG" "$RAMFS" -S beta & + PIDS+=($!) + wait_socket "$REG/beta" && HAVE_RAMFS=yes +fi + +echo "# synthetic root: posted names, no connection made for listing" +OUT=$(run_in "ls $M") +expect_contains "root lists alpha" "alpha" "$OUT" +[ "$HAVE_RAMFS" = yes ] && expect_contains "root lists beta" "beta" "$OUT" +# A plain file in the registry is listed but can never be dialed: this also +# proves listing connects to nothing (there is nothing to connect to). +touch "$REG/junk" +expect_contains "root lists a non-socket entry" "junk" "$(run_in "ls $M")" + +echo "# lazy dial and per-server routing in one program run" +OUT=$(run_in "ls $M; ls $M/alpha; cat $M/alpha/build/zig_version") +expect_eq "alpha subtree served after lazy dial" "$(zig version)" "$(run_in "cat $M/alpha/build/zig_version")" +expect_contains "root and subtree in one ls run" "build" "$OUT" +expect_contains "root and subtree in one ls run (zig version)" "$(zig version)" "$OUT" +if [ "$HAVE_RAMFS" = yes ]; then + expect_eq "ramfs file round trip through its own mount" "hello-from-beta" \ + "$(run_in "echo hello-from-beta > $M/beta/f && cat $M/beta/f")" + # Both servers in one run: the dispatcher serves them through two workers. + expect_eq "two subtrees in one run" "alpha ok beta ok" \ + "$(run_in "[ -f $M/alpha/build/zig_version ] && echo alpha ok; [ -f $M/beta/f ] && echo beta ok" | tr '\n' ' ' | sed 's/ $//')" + # Parallel readers on both mounts at once. + expect_eq "parallel reads on both mounts" "ok" \ + "$(run_in 'for i in 1 2 3 4 5 6 7 8; do cat '"$M"'/alpha/build/zig_version >/dev/null & cat '"$M"'/beta/f >/dev/null & done; wait; echo ok')" +fi + +echo "# a post after the mount is visible (snapshot per opendir)" +"$PROC" --unix "$REG/gamma" & +PIDS+=($!) +wait_socket "$REG/gamma" +expect_contains "gamma appears without remounting" "gamma" "$(run_in "ls $M")" +expect_eq "gamma dials on walk" "$(zig version)" "$(run_in "cat $M/gamma/build/zig_version")" + +echo "# walking a non-socket entry fails; the entry is never removed" +expect_eq "walk into junk yields EIO, exit status" "1" "$(run_in "cat $M/junk 2>/dev/null; echo \$?")" +expect_contains "junk still listed" "junk" "$(run_in "ls $M")" +[ -S "$REG/junk" ] && fail "junk was turned into a socket" || pass "junk untouched" + +echo "# server death: EIO on the subtree, name still listed, re-dial on re-post" +# SIGKILL, not SIGTERM: a graceful 9proc-demo unposts (the /srv model — a +# clean exit removes the name), while the acceptance case is a server that +# dies without cleaning up: the stale socket stays and must be listed, +# yield EIO on walks, and be replaced by the next server that posts. +# (the kill itself happens inside the child, at a deterministic point) +# One 9ns process stays mounted through the whole cycle. The child shares +# the host PID namespace, so the program itself kills the server and +# re-posts a fresh one under the same name at deterministic points. +DEATH_OUT=$(run_in ' + cat '"$M"'/alpha/build/zig_version >/dev/null && echo dial=ok + kill -9 '"$ALPHA"' 2>/dev/null + sleep 0.4 + MSG=$(cat '"$M"'/alpha/build/zig_version 2>&1 >/dev/null); echo "dead_status=$? dead_msg=$MSG" + ls '"$M"' | grep -q "^alpha$" && echo listed=yes + '"$PROC"' --unix '"$REG"'/alpha >/dev/null 2>&1 & + NEW=$! + for i in $(seq 1 50); do cat '"$M"'/alpha/build/zig_version >/dev/null 2>&1 && break; sleep 0.1; done + cat '"$M"'/alpha/build/zig_version >/dev/null && echo redial=ok + kill -9 "$NEW" 2>/dev/null; wait "$NEW" 2>/dev/null +') +expect_contains "mount dialed the server first" "dial=ok" "$DEATH_OUT" +expect_contains "dead server yields EIO, not a hang" "dead_status=1" "$DEATH_OUT" +expect_contains "dead server errno is EIO" "Input/output error" "$DEATH_OUT" +expect_contains "dead name still listed (stale socket, listing dials nothing)" "listed=yes" "$DEATH_OUT" +expect_contains "re-dial on next walk after re-post" "redial=ok" "$DEATH_OUT" +# The re-posted server was killed without cleanup inside the child, so a +# fresh process walks a stale entry: EIO, again. +expect_eq "stale name answers EIO in a fresh process" "1" "$(run_in "cat $M/alpha/build/zig_version 2>/dev/null; echo \$?")" +# Leave a live alpha for the remaining sections. +"$PROC" --unix "$REG/alpha" & +ALPHA=$! +PIDS+=($ALPHA) +wait_socket "$REG/alpha" +expect_eq "re-dial picks up the fresh post" "$(zig version)" "$(run_in "cat $M/alpha/build/zig_version")" + +echo "# mount defaults and environment" +expect_eq "default mountpoint is /mnt/9p" "$M" "$(timeout 60 "$NS" --mntgen -- sh -c 'echo $NINE_MOUNT')" +expect_contains "default mount is served" "alpha" "$(timeout 60 "$NS" --mntgen -- sh -c 'ls $NINE_MOUNT')" +expect_eq "mount is fuse" "yes" "$(run_in "grep -q \"^9ns $M fuse\" /proc/mounts && echo yes")" + +echo "# usage errors" +expect_eq "--mntgen with --unix is a usage error" "125" "$(timeout 10 "$NS" --mntgen --unix /tmp/x -- true 2>/dev/null; echo $?)" +expect_eq "--mntgen with --tcp is a usage error" "125" "$(timeout 10 "$NS" --mntgen --tcp 127.0.0.1:564 -- true 2>/dev/null; echo $?)" +expect_eq "--mntgen with --fd is a usage error" "125" "$(timeout 10 "$NS" --mntgen --fd 3 -- true 2>/dev/null; echo $?)" +expect_eq "--mntgen with --spawn is a usage error" "125" "$(timeout 10 "$NS" --spawn "true" --mntgen -- true 2>/dev/null; echo $?)" +expect_eq "--mntgen with --name is a usage error" "125" "$(timeout 10 "$NS" --mntgen --name foo -- true 2>/dev/null; echo $?)" +expect_eq "--mntgen=x is a usage error" "125" "$(timeout 10 "$NS" --mntgen=x -- true 2>/dev/null; echo $?)" +expect_eq "missing XDG_RUNTIME_DIR is fatal" "125" "$(env -u XDG_RUNTIME_DIR timeout 10 "$NS" --mntgen -- true 2>/dev/null; echo $?)" +expect_eq "exit status propagates" "7" "$(run_in 'exit 7'; echo $?)" + +echo "# xattr probes on the synthetic root read as unsupported, not EPERM" +# llistxattr/lgetxattr (ACL and capability probes from ls -l and stat) used to +# get EPERM from the root's read-only fallback, and coreutils blamed the +# mount root itself: "ls: /mnt/9p: Operation not permitted". They must read +# ENOSYS (like every other unimplemented op), which the kernel turns into a +# silent "no xattrs here". +XATTR_OUT=$(run_in "stat $M > /dev/null 2>&1; echo stat_rc=\$?; ls -ld $M > /dev/null 2>&1; echo lsld_rc=\$?") +expect_contains "stat of the mount root is clean" "stat_rc=0" "$XATTR_OUT" +expect_contains "ls -ld of the mount root is clean" "lsld_rc=0" "$XATTR_OUT" +# junk (a plain registry file) makes full ls -l fail with EIO on that entry +# (expected); what must never appear is EPERM blamed on the root. +BAD=$(run_in "ls -l $M 2>&1 | grep -i 'not permitted' | head -1") +expect_eq "no EPERM leak from ls -l" "" "$BAD" + +echo "# --debug and --no-direct-io reach the mntgen dispatcher" +# runMntgen used to drop both flags from the options it handed serveMntgen; +# --debug produced no trace at all. +DEBUG_ERR=$(timeout 60 "$NS" --mntgen --debug -- sh -c "ls $M/alpha >/dev/null" 2>&1 >/dev/null) +expect_contains "--debug traces the dispatcher" "lookup 'alpha'" "$DEBUG_ERR" + +echo "# a dial parked on a mute server unwedges when the program exits" +# A server that accepts the connection but never answers Tversion parks the +# dispatcher in the dial (pinned: no dial timeout, no concurrent dial). The +# dial must watch stop_fd through the whole handshake: when the program's +# main flow exits while a background walk is parked there, 9ns must follow it +# out instead of wedging forever (pre-fix it survived SIGTERM). +if command -v python3 > /dev/null; then + python3 - "$REG/mute" <<'PYEOF' & +import socket, sys, os +path = sys.argv[1] +try: + os.unlink(path) +except FileNotFoundError: + pass +s = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM) +s.bind(path) +s.listen(8) +while True: + conn, _ = s.accept() # accept, then never say a word +PYEOF + MUTEPID=$! + PIDS+=($MUTEPID) + sleep 0.3 + HANG_START=$(date +%s) + timeout 20 "$NS" --mntgen -- bash -c "(stat $M/mute >/dev/null 2>&1) & sleep 1" >/dev/null 2>&1 + HANG_RC=$? + HANG_SECONDS=$(( $(date +%s) - HANG_START )) + if [ "$HANG_RC" -eq 124 ] || [ "$HANG_SECONDS" -ge 15 ]; then + fail "9ns unwedges after the program exits a parked dial" "rc=$HANG_RC after ${HANG_SECONDS}s (wedged)" + else + pass "9ns unwedges after the program exits a parked dial (rc=$HANG_RC after ${HANG_SECONDS}s)" + fi + rm -f "$REG/mute" +else + echo "SKIP: mute-server dial-hang check needs python3" +fi + +echo +echo "passed=$PASSED failed=$FAILED" +[ "$FAILED" -eq 0 ] |
