summaryrefslogtreecommitdiff
path: root/src/CHANGELOG.md
diff options
context:
space:
mode:
authorGabriel Schneider <[email protected]>2026-09-03 13:29:50 -0300
committerGabriel Schneider <[email protected]>2026-09-03 13:29:50 -0300
commite6c9f1726cca910c464fc807aefebc404c8e7a12 (patch)
tree0bbcb79ae26a006afa40a4a8b26215a81610fb49 /src/CHANGELOG.md
parent7c6165f184e45af30ac697f8e8a584529ac4a287 (diff)
downloadpardes-e6c9f1726cca910c464fc807aefebc404c8e7a12.tar.gz
pardes-e6c9f1726cca910c464fc807aefebc404c8e7a12.zip
crash: a panic record must not be able to hang the process it is recording
Adversarial re-review, confirmed against std's source and then measured. `captureCurrentStackTrace` is not the safe half of `writeCurrentStackTrace`. `StackIterator.init` picks the `.di` strategy whenever `SelfInfo` can unwind, `stratOk` accepts `.di` regardless of `allow_unsafe_unwind`, and `.di` takes `SelfInfo`'s rwlock EXCLUSIVELY on its first call — the only kind of call a panic record makes — across `dl_iterate_phdr`, a DWARF CFI machine and an allocation. A panic in there (a smashed stack is a leading reason to be in a panic handler at all) leaves the lock held, because the unlock is a `defer` in a frame that never returns, and `defaultPanic` then waits on it for the life of the process. A crash becomes a hang, which is worse than what this file was added to improve on. The frames stay on stderr, where defaultPanic prints them under the staging that makes them safe; the record keeps what can be gathered without asking the process any questions. ONE record per process, never released. With the guard released on the way out, one panic wrote two records: the real message, then "reached unreachable code" under it. That second panic is this handler's own `vaxis.recover()` running a second time — it closes the vaxis tty and never clears the global saying there is one, so the double close is `recoverableOsBugDetected` and an `unreachable` in a Debug build. Guarded now in both the panic and the segfault handler; that half is a fix older than the crash file. `clock_gettime`'s return is checked, unlike dump.zig's, because a failure here leaves `ts` undefined and an undefined large-positive `sec` walks `calculateYearDay`'s u16 year past 65535 and overflow-panics inside the panic handler. Debug fills it with 0xaa and lands in 1970, which is why it reads as harmless. Verified end to end with a temporary probe: one panic, one record. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016Q4RATpafkwahrovHQLKRf
Diffstat (limited to 'src/CHANGELOG.md')
-rw-r--r--src/CHANGELOG.md24
1 files changed, 16 insertions, 8 deletions
diff --git a/src/CHANGELOG.md b/src/CHANGELOG.md
index 4cef906d..bb7e9ad5 100644
--- a/src/CHANGELOG.md
+++ b/src/CHANGELOG.md
@@ -15,14 +15,22 @@
fact. `src/crash.zig` runs inside the panic handler, so it takes no lock of
this program's, allocates nothing of its own, and every failure is swallowed:
a crash file that could not be written must not become the crash, and stderr
- still gets its own copy either way. The file gets RETURN ADDRESSES rather
- than the symbolised trace, and that is measured rather than chosen —
- `writeCurrentStackTrace` called from a panic handler *before* `defaultPanic`
- wedges the process at 0% CPU: symbolising reads DWARF, that read can panic,
- and the staging which turns a nested panic into "aborting due to recursive
- panic" is `defaultPanic`'s own and private. Walking frames is safe, so the
- addresses go in the file and `addr2line -e` against the build named on the
- line above them finishes the job. The AppKit shell got a panic handler of its
+ still gets its own copy either way. The file carries NO STACK TRACE, and that
+ is measured rather than chosen: `writeCurrentStackTrace` called from a panic
+ handler *before* `defaultPanic` wedges the process at 0% CPU, and
+ `captureCurrentStackTrace` — which looks like the safe half of it — takes
+ `SelfInfo`'s rwlock exclusively on its first call, so a panic inside the walk
+ leaves that lock held and `defaultPanic` waits on it for the life of the
+ process. A crash that becomes a hang is worse than the crash. What makes
+ `defaultPanic` itself survive that is its private `panic_stage`, reachable
+ from nowhere outside `std.debug`, so the frames stay on stderr where they
+ already work. ONE record per process, because the nested panic std's trace
+ printer raises comes back through the handler: the first version of this
+ wrote the real message and then "reached unreachable code" underneath it,
+ which is the panic handler's own second `vaxis.recover()` double-closing the
+ tty — `recover()` never cleared the global saying there was one. That call is
+ guarded now too, in both the panic and the segfault handler, which is a fix
+ older than this file. The AppKit shell got a panic handler of its
own in the process: the macOS build roots at `macos.zig`, so the one in
`main.zig` had never run there — in the shell with the least useful stderr of
the four.