@khorsheed/dsh-ankh-guard 0.3.2 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +9 -0
- package/README.en.md +10 -5
- package/README.i18n.yaml +2 -2
- package/README.md +10 -5
- package/lib/cli.js +136 -47
- package/lib/preflight-runner.js +37 -3
- package/lib/types/cli.d.ts +13 -0
- package/lib/types/cli.js +121 -21
- package/lib/types/preflight-runner.d.ts +27 -0
- package/lib/types/preflight-runner.js +43 -2
- package/lib/types/transition.d.ts +50 -10
- package/lib/types/transition.js +71 -22
- package/package.json +10 -10
- package/scripts/dsh-watchdog.sh +108 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,14 @@
|
|
|
1
1
|
# 变更记录
|
|
2
2
|
|
|
3
|
+
## 0.4.1(2026-09-30)
|
|
4
|
+
|
|
5
|
+
修复 2026-09-29/30 的 0.2.0-rc.2 切换事故暴露的三个缺口:
|
|
6
|
+
|
|
7
|
+
- **guardInvocation() 改用绝对 execPath**:build/source 两种形态均用绝对 `process.execPath`——裸 `node` 在 launcher 的 PATH 下找不到
|
|
8
|
+
- **awaiting-user 收据上的裸 supervise 改为泊驻**:占用 pidfile、不拉起进程、消费 abort/restore 控制标记——不再直接退出把事务楔死在没有活体消费者的状态;`supervise --cutover-id` 释放泊驻并恢复;拒绝消息打印确切可执行的命令
|
|
9
|
+
- **失败 boot 即使日志为空也镜像 attempt 日志**(「produced no output」本身就是信号);`reconfigure`/`schedule-exit`/`supervise` 新增 `--boot-timeout-ms`(穿到 `WD_BOOT_TIMEOUT`)
|
|
10
|
+
- 加宽 `@deepseek-ai/dsh-*` peer 区间以覆盖宿主 0.2.0
|
|
11
|
+
|
|
3
12
|
## 0.3.2(2026-09-27)
|
|
4
13
|
|
|
5
14
|
适配宿主 0.1.7-rc.2 线:verifiedHost 前移至 0.1.7-rc.2(3080 生产实证线随宿主基线切到 rc.2);rc.1→rc.2 对本包无破坏性变更(逐类清点见 [Agent Note](../../.agents/notes/implemented/architecture/2026-09-27-host-017-rc2-breaking-changes.md)),全量构建+测试双绿。
|
package/README.en.md
CHANGED
|
@@ -6,7 +6,7 @@ Let an agent change its own code and restart its own service — without taking
|
|
|
6
6
|
|
|
7
7
|
When the agent wants to restart after editing code, this plugin asks one question first: did the build and tests pass? Yes, go ahead. No, blocked — so broken code can't take the service, and the conversation running inside it, down with it.
|
|
8
8
|
|
|
9
|
-
<img src="https://raw.githubusercontent.com/Khorsheed/dsh-
|
|
9
|
+
<img src="https://raw.githubusercontent.com/Khorsheed/dsh-basic/main/docs/screenshots/ankh-guard.JPG" width="640" alt="a guarded restart: the agent announces its verification plan beforehand, and the canary reactivates the session afterwards to keep verifying">
|
|
10
10
|
|
|
11
11
|
## How it works
|
|
12
12
|
|
|
@@ -79,7 +79,7 @@ Full commands: `verify`, `record`, `status`, `clear`, `checkpoint`, `reset`, `ca
|
|
|
79
79
|
|
|
80
80
|
### preflight: the composition gate
|
|
81
81
|
|
|
82
|
-
`preflight` deep-dry-runs the exact composition a restart would boot: it composes the profile's full patch stack through the same path as the real launcher (bundle layers, user layers, overlays), boots the **whole plugin tree** in a subprocess through the same engine — every plugin's apply runs, because apply is activation — with an overlay pinning the webserver port to 0 (OS-assigned, so it never collides with the live instance), checks every registered client bundle artifact exists, and disposes (registrations are effects, so dispose rolls the dry-run back). Exit codes are the contract:
|
|
82
|
+
`preflight` deep-dry-runs the exact composition a restart would boot: it composes the profile's full patch stack through the same path as the real launcher (bundle layers, user layers, overlays), boots the **whole plugin tree** in a subprocess through the same engine — every plugin's apply runs, because apply is activation — with an overlay pinning the webserver port to 0 (OS-assigned, so it never collides with the live instance), checks every registered client bundle artifact exists, and reads back the agent preset registry's `broken` diagnostic (preset rows mount on the registry's standing scopes, not the profile root, so **a clean boot does not imply a usable preset**: a broken preset shows 加载失败 in the picker and its sessions fail resume with `never started` — 3080 hit exactly that on 2026-09-28 with the dry-run green all the way), and disposes (registrations are effects, so dispose rolls the dry-run back). Exit codes are the contract:
|
|
83
83
|
|
|
84
84
|
- `0` — the composition boots clean.
|
|
85
85
|
- `1` — a composition verdict: the tree a restart would boot is broken; the output names the failing layer.
|
|
@@ -109,9 +109,9 @@ dsh-ankh-guard supervise --port 3080 --start "CMD" --state-dir "$DSH_HOME/state"
|
|
|
109
109
|
|
|
110
110
|
`supervise` also needs the dsh home the supervised instance boots with (the watchdog exports it as the instance's `DSH_HOME`): `--home DIR` wins, else `$DSH_HOME`; with neither set it refuses loudly — a home guessed from `--state-dir` would silently boot the instance on the wrong profiles/credentials. First-time persistence also requires an explicit `--harness-root` or `DSH_HARNESS`; it never guesses the host root from the credential repo.
|
|
111
111
|
|
|
112
|
-
It spawns `scripts/dsh-watchdog.sh` (ships with the package) detached with `--wait-owner`: the watchdog idles while the current instance runs, takes over the port when the instance exits (intentional restart or crash), respawns it, runs the guard canary on intentional restarts (a `restart-requested.json` marker), and clears the marker on pass. Two consecutive boot failures roll the checkout back to the last known-good revision — the healthy-boot stamp (`last-good-boot.json`, written every time the instance comes up, so it names the last revision that genuinely ran in this deployment), else the guard checkpoint, else the credential's HEAD — but only when the boot failure's error subject is a path inside the repository. When the subject lives outside the checkout (a broken profile overlay or an installed plugin), a checkout reset cannot help, so the watchdog instead restores the last healthy **profile composition**: the snapshot of the profile's composition inputs (`last-good-composition/`, taken at every healthy boot) replaces the live bundles layer and manifest, unmounting the newest plugin change, with the failing inputs preserved under `composition-backup-*` and the recovered report naming exactly what was unmounted. The same exemption logic covers a start command that does not bind the supervised port: when the boot window times out while the instance is listening elsewhere — or fails with `EADDRINUSE` naming a port this watchdog does not own — the watchdog names the bound port and skips both rollbacks, because resetting files cannot change a command-line argument. `EADDRINUSE` on the supervised port keeps its free-and-retry escape hatch, now bounded at five attempts. Every reset (watchdog, CLI, or service) first creates `guard-backup-*` branch anchors for the discarded HEAD and for uncommitted tracked changes, so recovery never depends on the reflog. Four failures serve a crash page on the port with a retry button (SIGUSR1 to the watchdog). A `watchdog-stop` marker exits the watchdog for good. The instance itself can adopt supervision before a self-restart — the user never starts the watchdog by hand.
|
|
112
|
+
It spawns `scripts/dsh-watchdog.sh` (ships with the package) detached with `--wait-owner`: the watchdog idles while the current instance runs, takes over the port when the instance exits (intentional restart or crash), respawns it, runs the guard canary on intentional restarts (a `restart-requested.json` marker), and clears the marker on pass. Two consecutive boot failures roll the checkout back to the last known-good revision — the healthy-boot stamp (`last-good-boot.json`, written every time the instance comes up, so it names the last revision that genuinely ran in this deployment), else the guard checkpoint, else the credential's HEAD — but only when the boot failure's error subject is a path inside the repository. When the subject lives outside the checkout (a broken profile overlay or an installed plugin), a checkout reset cannot help, so the watchdog instead restores the last healthy **profile composition**: the snapshot of the profile's composition inputs (`last-good-composition/`, taken at every healthy boot) replaces the live bundles layer and manifest, unmounting the newest plugin change, with the failing inputs preserved under `composition-backup-*` and the recovered report naming exactly what was unmounted. The same exemption logic covers a start command that does not bind the supervised port: when the boot window times out while the instance is listening elsewhere — or fails with `EADDRINUSE` naming a port this watchdog does not own — the watchdog names the bound port and skips both rollbacks, because resetting files cannot change a command-line argument. `EADDRINUSE` on the supervised port keeps its free-and-retry escape hatch, now bounded at five attempts. A failed boot always mirrors its captured attempt log into the main watchdog log — with an explicit `attempt N produced no output` line when the instance never wrote a line at all — so a dead-on-arrival boot is visible where an operator looks first. Every reset (watchdog, CLI, or service) first creates `guard-backup-*` branch anchors for the discarded HEAD and for uncommitted tracked changes, so recovery never depends on the reflog. Four failures serve a crash page on the port with a retry button (SIGUSR1 to the watchdog). A `watchdog-stop` marker exits the watchdog for good. The instance itself can adopt supervision before a self-restart — the user never starts the watchdog by hand.
|
|
113
113
|
|
|
114
|
-
When a watchdog is already supervising, the restart trigger is `schedule-exit`: it gets the port, credential repo, host root, and profile from the durable active launch spec, rejects conflicting explicit flags, and verifies the supervisor's complete command before writing the restart marker and spawning a detached process explicitly labelled `exit-agent pid`. It prefers a credential within the ten-minute freshness window. Once that expires, only an exact `provenDeployment` may take the pure-restart fast path; the selected evidence SHA is copied into the short-lived restart marker and the new watchdog rechecks that same SHA and the live fingerprint during canary, closing the check-to-stop evidence replacement window. A successful canary then promotes or retains the proof. For an agent Bash/tool call, set that call's `timeoutMs` to 180000; this is tool metadata, not a CLI flag, and keeps the caller alive longer than the 120-second preflight gate. The managed shell cannot reap the exit agent, so the scheduled kill lands after the scheduling turn ends. The watchdog respawns, runs the canary, and the new instance reports via `last-restart.json`; watchdog lifecycle lines are timestamped. With no live watchdog, `schedule-exit` hard-refuses: establish supervision first or use `restart`, which owns the complete single-shot loop. The latter lacks durable supervisor/launch ownership and therefore still requires a fresh credential.
|
|
114
|
+
When a watchdog is already supervising, the restart trigger is `schedule-exit`: it gets the port, credential repo, host root, and profile from the durable active launch spec, rejects conflicting explicit flags, and verifies the supervisor's complete command before writing the restart marker and spawning a detached process explicitly labelled `exit-agent pid`. It prefers a credential within the ten-minute freshness window. Once that expires, only an exact `provenDeployment` may take the pure-restart fast path; the selected evidence SHA is copied into the short-lived restart marker and the new watchdog rechecks that same SHA and the live fingerprint during canary, closing the check-to-stop evidence replacement window. A successful canary then promotes or retains the proof. For an agent Bash/tool call, set that call's `timeoutMs` to 180000; this is tool metadata, not a CLI flag, and keeps the caller alive longer than the 120-second preflight gate. The managed shell cannot reap the exit agent, so the scheduled kill lands after the scheduling turn ends. The watchdog respawns, runs the canary, and the new instance reports via `last-restart.json`; watchdog lifecycle lines are timestamped. With no live watchdog, `schedule-exit` hard-refuses: establish supervision first or use `restart`, which owns the complete single-shot loop. The latter lacks durable supervisor/launch ownership and therefore still requires a fresh credential. `--boot-timeout-ms MS` rides the restart marker as a one-boot readiness budget: the respawn's boot window honors it (rounded up to whole seconds) while the marker is pending, then the watchdog falls back to its own `WD_BOOT_TIMEOUT` (default 60 s).
|
|
115
115
|
|
|
116
116
|
### reconfigure: transactional launch changes
|
|
117
117
|
|
|
@@ -128,18 +128,23 @@ dsh-ankh-guard reconfigure \
|
|
|
128
128
|
--transition-file "<optional state-transition plan.json>" \
|
|
129
129
|
--on-failure restore-previous \
|
|
130
130
|
--browser-handoff required \
|
|
131
|
+
--boot-timeout-ms 90000 \
|
|
131
132
|
--state-dir "$DSH_HOME/state"
|
|
132
133
|
```
|
|
133
134
|
|
|
135
|
+
`reconfigure` and `supervise` accept `--boot-timeout-ms MS` to hand the spawned watchdog its per-boot readiness budget (`WD_BOOT_TIMEOUT`, whole seconds; omit it on a slow host and a healthy-but-slow first boot reads as a failure); on `reconfigure` the value rides the supervisor driver's argv to the successor watchdog.
|
|
136
|
+
|
|
134
137
|
The recovery choice is mandatory and therefore approved before the old host stops: `restore-previous` restores the entire previous launch specification, while `wait-for-user` parks for intervention without resetting any repository. The full previous/target pair separately persists command, home, credential repo, harness root, profile, and port. Each new spec also binds the composition preflight's `source|built` surface, runner executable/path/content SHA, the actual `@deepseek-ai/dsh/package.json` install anchor, and the target command SHA. A built successor resolves modules from its npm toolchain; the guard never chooses checkout source merely because its own process happens to use tsx. `reconfigure` additionally requires the caller to supply a one-shot `--candidate-probe-command` committed with the target command SHA. It runs that probe against an isolated home before composition preflight runs on the bound execution surface. The guard binds and executes both command digests; it cannot prove that an arbitrary successful shell command was automatically derived from the target argv. The bundled Skill constructs a DSH `--dump-config` probe from the same executable and launcher argv; integrations outside that caller trust boundary must establish the same semantic provenance themselves. Either failure occurs before previous stops. The mode-0600 `launch-spec.json` contains the commands and selected side; the receipt contains only bindings, hashes, and PASS outcomes, never probe/start commands or bearer URLs.
|
|
135
138
|
|
|
136
139
|
The atomic selected-side rename is the configuration commit point. If no complete durable previous spec exists, initialize it first with the real current values and explicit preflight surface/install anchor through `configure-launch`: legacy `instance-launch.json` cannot supply the missing roles, and a target `--repo` is never backfilled into previous. A replacement watchdog then atomically claims `watchdog.pid` while the old host is still serving. During preparation the guard captures PID/start identities for the old supervisor, its direct child, and its listener. The successor waits for the exact supervisor identity to yield with a 15-second default bound (configurable through `--supervisor-yield-timeout-ms`) and consumes abort/restore throughout that wait. On timeout it may retire only the frozen, revalidated old supervisor tree. It then stops only the proven child/listener identities and confirms port release; it never chooses or kills an arbitrary process found by port. PID reuse, a hung watchdog, a short-lived `reconfigure` caller, an outer launchd/systemd waiter, and nested shells therefore cannot blur ownership or suspend a cutover forever.
|
|
137
140
|
|
|
138
141
|
For a protected target, the watchdog accepts a launch URL only from the final process's output, only for the exact supervised loopback authority, and without depending on a parameter name. It proves 303 cookie exchange and authenticated root 200 with a temporary jar, then spends a three-second default stability window proving that the child is alive, its sole listener remains in that child tree, PID/start identities do not change, and retry is zero. The browser half uses held long polls on a same-origin plugin route rather than a permanent 500 ms loop. The proven previous listener stores only hashes of per-tab capabilities and puts every responsive registered tab into a waiting state. Only after ownership-stable service readiness and canary success does the final listener tell each tab to reload when its cookie is still accepted, or return that final process's same-origin one-time URL for an in-memory `location.replace()` after 401. The authenticated page acknowledges and returns to the original safe pathname with query and fragment discarded. One real ACK gates terminal ready; other registered, unacknowledged tabs remain eligible after state compaction. If no original tab registered or none acknowledges in time, the watchdog requests one system-open fallback and still waits for its authenticated page acknowledgement; opener exit 0 alone never completes handoff. Server readiness/canary and browser handoff are recorded separately, and all required evidence must complete before session wake-up. A naked 401, an old listener's 200, a target's transient 200, or a subsequent target exit can never become ready or hand a rejected target URL to the browser. No raw browser capability or bearer URL enters a state file, durable log, or receipt. `launch-cutover.json` records redacted configuration summaries, process identities, authentication, handoff, stability, and per-role failures. Target readiness/canary remains under `targetValidation`; previous recovery readiness and a passing, failing, or explicitly skipped recovery canary (when only a target-scoped credential exists) live under `recovery.validation`. A restored receipt can no longer carry an unqualified target canary failure beside previous readiness. `launch-status` prints the receipt without exposing either command. During a transaction, `abort-cutover --state-dir "$DSH_HOME/state"` applies the pre-approved recovery policy, while `restore-previous --state-dir "$DSH_HOME/state"` explicitly authorizes stopping the proven target and restoring the complete previous spec. The two actions use separate atomic markers and restore always wins on read, so even concurrent sessions cannot let a later ordinary abort downgrade an explicit restore.
|
|
139
142
|
|
|
143
|
+
**Mid-cutover recovery when the supervisor chain dies.** A receipt parked in `awaiting-user` no longer wedges the transaction. A bare `supervise` — the command launchd/systemd runs — now spawns the watchdog in a parked hold: it claims supervision, launches nothing (the rejected side never boots), and keeps consuming the operator control markers, so `abort-cutover --state-dir "$DSH_HOME/state"` (applies the pre-approved policy) and `restore-previous --state-dir "$DSH_HOME/state"` work immediately instead of refusing with "no live watchdog". The hold announces its own exits once: `touch "$DSH_HOME/state/watchdog-stop"` ends the hold without settling, and `supervise --cutover-id <id> --state-dir "$DSH_HOME/state"` (the id is in `launch-status` and the receipt) resumes the selected side — the resume flips the receipt out of `awaiting-user`, the hold releases its claim in response, and the resuming chain waits for that release before taking over.
|
|
144
|
+
|
|
140
145
|
When a candidate cannot read a reconstructible projection or cache left by the previous host, `--transition-file` can submit a reviewed schema-v1 quarantine plan. A plan accepts only non-overlapping, symlink-free paths below `home` that exclude guard state, with an explicit `quarantine` operation; it contains no host-version or filename knowledge. Example: `{"schemaVersion":1,"home":"/absolute/dsh-home","operations":[{"kind":"quarantine","path":"storages/<reconstructible-cache>","expect":"present"}]}`. Each `expect` is `present` or `absent`; isolated preflight and live apply must observe that same state or refuse before previous stops or target starts. The plan must cover both old paths that need to leave before target starts and new output paths that must leave before previous can recover after a target failure; list the latter explicitly with `expect: "absent"` even when they do not exist at preparation. Do not use this mechanism for authoritative logs, credentials, or irreplaceable data. Formats that require content transformation need a separate reversible migration tool and review.
|
|
141
146
|
|
|
142
|
-
The guard first copies the live home's
|
|
147
|
+
The guard first copies the live home's boot inputs with copy-on-write preference — the allowlist is `profiles/`, the home-level `cordis.patch.yml`, `settings.yaml`, `.credentials.yaml`, and `.anonymous-user-id`; plugin data directories (sessions, state, local-agent sub-homes, tarballs, scratch, …) are excluded by default, so a newly installed plugin's data never silently grows the snapshot, and the list only moves when the host's boot starts reading a new home input (a miss fails the preflight loudly instead of slowing the copy down). It then rebuilds pnpm/Cordis links as relative links whose targets are snapshot-owned copies; legitimate internal dependency cycles remain intact. External targets enter a hash-named, deduplicated materialization area inside the snapshot. A linked `node_modules` target anchors at the package's own parent `node_modules` — the tightest scope that preserves Node's ancestor lookup — rather than the whole store root. A post-copy `realpath` audit requires every writable link target to remain under the snapshot root. Runtime entries without copyable semantics (sockets, FIFOs, and links to them) are skipped and counted; dangling or unresolvable links, any other special files, writable escapes, and read/copy failures make `reconfigure` fail closed before it runs the candidate, creates a cutover, or stops previous. It then applies the same quarantine and runs target composition preflight in that safe copy. If the copy cannot be prepared or the target does not boot, previous keeps serving and the live home stays unchanged. Only after the successor owns supervision and has stopped and revalidated the previous process tree does it apply the hash-bound durable plan with same-filesystem renames. If the target is rejected, the watchdog first stops its proven process, retains replacements it created at transitioned paths under `launch-transitions/<cutover>/rejected-target/`, restores the exact previous bytes and records the result, and only then permits previous to start. Any unproven step parks at `awaiting-user` instead of exposing previous to mixed state. After target success, the quarantined previous content remains in the cutover directory for operator disposition; it is never deleted automatically.
|
|
143
148
|
|
|
144
149
|
**The restart report reaches the model by itself — and waits for its owner.** After a scheduled restart (a pending `last-restart.json` record), the plugin queues the report as the next turn via `agent.followup` — the official wake-the-agent seam the schedule system uses for reminders — so the agent reports the restart result without any user message. Session restore after a restart is lazy (an agent is created only when the UI or an RPC touches the session), so the report goes ONLY to the session that scheduled the exit (`schedule-exit` records `$DSH_SESSION_ID` as the initiator), whenever it resumes — no other session is ever woken for reporting, and the record stays pending until its owner resumes or the next restart replaces it (new `exitAt`). A record without an initiator is claimed by the first root agent created. Only root agents, once (acknowledged on delivery). Config `reportRestartContext`: `followup` (default, autonomous), `step` (ride the first step of whatever turn comes next), or `off`.
|
|
145
150
|
|
package/README.i18n.yaml
CHANGED
|
@@ -2,5 +2,5 @@
|
|
|
2
2
|
# last confirmed-consistent state. Both languages carry equal authority; after
|
|
3
3
|
# editing either side, bring the other along and re-record with:
|
|
4
4
|
# pnpm run verify-translation-pairing --write packages/ankh-guard/README.en.md
|
|
5
|
-
packages/ankh-guard/README.en.md:
|
|
6
|
-
packages/ankh-guard/README.md:
|
|
5
|
+
packages/ankh-guard/README.en.md: fb474172d57154e58e44661bbefb6824c78425e0
|
|
6
|
+
packages/ankh-guard/README.md: e58a1358dcfae6d426b3ced160120de361709e24
|
package/README.md
CHANGED
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
|
|
7
7
|
agent 改完代码想重启的时候,这个插件会先问一句:这次改动,构建和测试都过了吗?过了才放行,没过就拦下来——免得改坏的代码把整个服务、连同正在进行的对话一起带走。
|
|
8
8
|
|
|
9
|
-
<img src="https://raw.githubusercontent.com/Khorsheed/dsh-
|
|
9
|
+
<img src="https://raw.githubusercontent.com/Khorsheed/dsh-basic/main/docs/screenshots/ankh-guard.JPG" width="640" alt="一次受守护的重启:重启前告知验证项,重启后金丝雀自动激活会话并注入上下文继续验证">
|
|
10
10
|
|
|
11
11
|
## 工作原理
|
|
12
12
|
|
|
@@ -79,7 +79,7 @@ dsh-ankh-guard reconfigure --start "NEW CMD" --repo "<credential repo>" \
|
|
|
79
79
|
|
|
80
80
|
### preflight: the composition gate
|
|
81
81
|
|
|
82
|
-
`preflight` 对重启将要 boot 的组合做完全一致的深度干跑:走与真实 launcher 相同的路径组装 profile 的全部 patch 层(bundle 层、用户层、overlay),在子进程里用同一个引擎 boot **整棵插件树**——每个插件的 apply 都真实执行,因为 apply 即激活——同时用 overlay 把 webserver 端口钉到 0(操作系统分配,绝不与在跑实例抢端口),检查每个已注册 client bundle
|
|
82
|
+
`preflight` 对重启将要 boot 的组合做完全一致的深度干跑:走与真实 launcher 相同的路径组装 profile 的全部 patch 层(bundle 层、用户层、overlay),在子进程里用同一个引擎 boot **整棵插件树**——每个插件的 apply 都真实执行,因为 apply 即激活——同时用 overlay 把 webserver 端口钉到 0(操作系统分配,绝不与在跑实例抢端口),检查每个已注册 client bundle 产物存在,并回读 agent preset 注册表的 `broken` 诊断(preset 行挂在注册表的 standing scope 上而不是 profile 根,**boot 干净不代表 preset 可用**:坏 preset 的选择器卡片显示「加载失败」、其会话 resume 报 `never started`——3080 在 2026-09-28 撞过一次,干跑一路全绿),然后 dispose(注册即 effect,dispose 即回滚这次干跑)。退出码即契约:
|
|
83
83
|
|
|
84
84
|
- `0`——组合干净通过。
|
|
85
85
|
- `1`——组合结论:重启将要 boot 的树是坏的;输出会指明坏在哪一层。
|
|
@@ -109,9 +109,9 @@ dsh-ankh-guard supervise --port 3080 --start "CMD" --state-dir "$DSH_HOME/state"
|
|
|
109
109
|
|
|
110
110
|
`supervise` 还需要被监管实例启动时使用的 dsh home(watchdog 会把它 export 为实例的 `DSH_HOME`):`--home DIR` 优先,否则取 `$DSH_HOME`;两者都没有时响亮拒绝——从 `--state-dir` 猜出来的 home 会让实例静默读错 profile/凭据目录。首次持久化还必须显式提供 `--harness-root` 或 `DSH_HARNESS`;它不会把 credential repo 猜成宿主根。
|
|
111
111
|
|
|
112
|
-
它以 `--wait-owner` 模式 detached 拉起随包发布的 `scripts/dsh-watchdog.sh`:watchdog 在当前实例运行期间待机,实例退出(有意重启或崩溃)后接管端口、重新拉起,有意重启时跑 guard canary(读 `restart-requested.json` 标记),通过后清除标记。连续 2 次起不来→回滚到最后已知可用版本:健康启动戳(`last-good-boot.json`,每次实例成功启动时重写,指向本部署里最近一次真正跑起来的版本)优先,其次是 guard checkpoint,最后是凭证 HEAD;但仅当启动失败的错误主体路径在仓库内。主体在仓库之外时(坏掉的 profile overlay 或已装插件),回滚检出修不好,watchdog 改为恢复上次健康的 **profile 组合**:健康启动时快照的组合输入(`last-good-composition/`)覆盖回 live 的 bundles 层与清单,最新插件变更被卸载,故障输入保留在 `composition-backup-*`,恢复报告会点名被卸载的内容。启动命令没有绑到被监督端口时同样豁免:启动窗口超时而实例正监听在别处、或以点名了本 watchdog 并不拥有的端口的 `EADDRINUSE` 失败时,watchdog 会点名实际绑定的端口并跳过两种回滚——重置文件改不了命令行参数。发生在被监督端口上的 `EADDRINUSE`
|
|
112
|
+
它以 `--wait-owner` 模式 detached 拉起随包发布的 `scripts/dsh-watchdog.sh`:watchdog 在当前实例运行期间待机,实例退出(有意重启或崩溃)后接管端口、重新拉起,有意重启时跑 guard canary(读 `restart-requested.json` 标记),通过后清除标记。连续 2 次起不来→回滚到最后已知可用版本:健康启动戳(`last-good-boot.json`,每次实例成功启动时重写,指向本部署里最近一次真正跑起来的版本)优先,其次是 guard checkpoint,最后是凭证 HEAD;但仅当启动失败的错误主体路径在仓库内。主体在仓库之外时(坏掉的 profile overlay 或已装插件),回滚检出修不好,watchdog 改为恢复上次健康的 **profile 组合**:健康启动时快照的组合输入(`last-good-composition/`)覆盖回 live 的 bundles 层与清单,最新插件变更被卸载,故障输入保留在 `composition-backup-*`,恢复报告会点名被卸载的内容。启动命令没有绑到被监督端口时同样豁免:启动窗口超时而实例正监听在别处、或以点名了本 watchdog 并不拥有的端口的 `EADDRINUSE` 失败时,watchdog 会点名实际绑定的端口并跳过两种回滚——重置文件改不了命令行参数。发生在被监督端口上的 `EADDRINUSE` 保留原本的释放并重试逃生口,现在以五次为上限。每次启动失败都会把该次 attempt 捕获的输出镜像进 watchdog 主日志——实例一行都没写时则显式记录 `attempt N produced no output`——让「boot 即死」在运维第一眼就能看到。任何路径的 reset(watchdog、CLI、service)都会先为被丢弃的 HEAD 和未提交改动创建 `guard-backup-*` 分支锚点,恢复不依赖 reflog。4 次失败→在端口上提供带重试按钮的崩溃页(SIGUSR1 通知 watchdog)。`watchdog-stop` 标记让 watchdog 彻底退出。实例可以在自我重启前自行采用监督——用户永远不需要手动启动 watchdog。
|
|
113
113
|
|
|
114
|
-
已有 watchdog 监督时,重启触发用 `schedule-exit`:它从耐久 active launch spec 取得端口、凭证仓库、宿主根与 profile,拒绝任何冲突的显式参数,并核对 supervisor 写下的完整命令后才写 restart 标记、spawn detached 退出代理(输出明确标为 `exit-agent pid`)。它优先使用 10 分钟内的新鲜 credential;credential 过期时,只允许精确匹配的 `provenDeployment` 走纯重启快速路径,并把所选证据 SHA 写入短寿命 restart marker。新 watchdog 在 canary 时再次核对同一 SHA 与现场指纹,防止检查后、停止前的证据替换。通过后才晋升或保留部署证明。从 agent 的 Bash/tool 调用时应把该调用的 `timeoutMs` 设为 180000;这不是 CLI 参数,而是保证调用方不会先于 120 秒 preflight 闸门退出的等待契约。托管 shell 的进程组回收不到退出代理,所以计划中的 kill 会在调度回合结束后真实落地。watchdog 重新拉起、跑 canary,新实例经 `last-restart.json` 回报;watchdog 生命周期日志均带时间戳。没有存活 watchdog 时 `schedule-exit` 硬拒绝,只能先建立监督或使用拥有单次完整循环的 `restart`;后者没有相同的 durable supervisor/launch ownership,因此仍要求新鲜 credential
|
|
114
|
+
已有 watchdog 监督时,重启触发用 `schedule-exit`:它从耐久 active launch spec 取得端口、凭证仓库、宿主根与 profile,拒绝任何冲突的显式参数,并核对 supervisor 写下的完整命令后才写 restart 标记、spawn detached 退出代理(输出明确标为 `exit-agent pid`)。它优先使用 10 分钟内的新鲜 credential;credential 过期时,只允许精确匹配的 `provenDeployment` 走纯重启快速路径,并把所选证据 SHA 写入短寿命 restart marker。新 watchdog 在 canary 时再次核对同一 SHA 与现场指纹,防止检查后、停止前的证据替换。通过后才晋升或保留部署证明。从 agent 的 Bash/tool 调用时应把该调用的 `timeoutMs` 设为 180000;这不是 CLI 参数,而是保证调用方不会先于 120 秒 preflight 闸门退出的等待契约。托管 shell 的进程组回收不到退出代理,所以计划中的 kill 会在调度回合结束后真实落地。watchdog 重新拉起、跑 canary,新实例经 `last-restart.json` 回报;watchdog 生命周期日志均带时间戳。没有存活 watchdog 时 `schedule-exit` 硬拒绝,只能先建立监督或使用拥有单次完整循环的 `restart`;后者没有相同的 durable supervisor/launch ownership,因此仍要求新鲜 credential。`--boot-timeout-ms MS` 随 restart marker 记录一次性就绪预算:marker 待消费期间,重新拉起的启动窗口按它计算(向上取整到整秒),之后 watchdog 回落到自己的 `WD_BOOT_TIMEOUT`(默认 60 秒)。
|
|
115
115
|
|
|
116
116
|
### reconfigure:启动配置事务切换
|
|
117
117
|
|
|
@@ -128,18 +128,23 @@ dsh-ankh-guard reconfigure \
|
|
|
128
128
|
--transition-file "<可选的状态迁移计划.json>" \
|
|
129
129
|
--on-failure restore-previous \
|
|
130
130
|
--browser-handoff required \
|
|
131
|
+
--boot-timeout-ms 90000 \
|
|
131
132
|
--state-dir "$DSH_HOME/state"
|
|
132
133
|
```
|
|
133
134
|
|
|
135
|
+
`reconfigure` 与 `supervise` 接受 `--boot-timeout-ms MS`,把每次启动的就绪预算交给拉起的 watchdog(`WD_BOOT_TIMEOUT`,整秒;慢宿主上省略它,一个健康但缓慢的首次 boot 会被误判为失败);`reconfigure` 下该值随 supervisor 驱动器的 argv 传给继任 watchdog。
|
|
136
|
+
|
|
134
137
|
恢复选择是必填项,因而会在旧宿主停止前获批:`restore-previous` 恢复上一份**完整**启动配置;`wait-for-user` 不重置任何仓库,原地停留等用户处理。完整 previous/target 对分别保存 command、home、credential repo、harness root、profile 与 port。每份新配置还固化 composition preflight 的 `source|built` 面、runner 可执行方式/路径/内容 SHA、实际 `@deepseek-ai/dsh/package.json` 安装锚点和 target command SHA;built successor 从其 npm toolchain 解析模块,绝不因 guard 自己恰由 tsx 启动就误选 checkout source。`reconfigure` 另要求调用方提供一条与 target command SHA 同时提交的一次性 `--candidate-probe-command`,先在隔离 home 执行,再跑同一执行面上的 composition preflight;Guard 只绑定并执行两个 command 摘要,不能证明任意成功的 shell command 是从 target argv 自动派生的。随包 Skill 会用相同 executable 与 launcher argv 构造 DSH `--dump-config` probe;这条调用方信任边界之外的其他集成也必须自行保证语义同源。任一失败都发生在 previous 停止之前。当前选中侧保存在 mode-0600 的 `launch-spec.json`;回执只写各摘要和 PASS 结果,不写 probe/start 命令或 bearer URL。
|
|
135
138
|
|
|
136
139
|
原子切换选中侧是配置提交点。缺少完整耐久 previous 时必须先用当前真实值运行 `configure-launch`,并显式给出它的 preflight surface/install anchor;旧 `instance-launch.json` 不足以推断,目标 `--repo` 也绝不会倒填 previous。随后替代 watchdog 会在旧宿主仍对外服务时原子取得 `watchdog.pid`。准备事务时 guard 同时固化旧 supervisor、直接 child 与 listener 的 PID/启动 identity;successor 按 identity 有界等待旧 supervisor 正常让权(默认 15 秒,可用 `--supervisor-yield-timeout-ms` 调整),等待期间持续消费 abort/restore。超时只会在先冻结并复核旧 supervisor identity 后终止其精确进程树。successor 随后只停止已证明的旧 child/listener 并确认端口释放,绝不凭端口反查后杀任意监听者。因此 PID 复用、卡死 watchdog、短命 `reconfigure` 调用者、外层 launchd/systemd 等待者或嵌套 shell 都不会模糊所有权或无限悬挂切换。
|
|
137
140
|
|
|
138
141
|
目标受保护时,watchdog 只接受最终进程输出、且 authority 与被监督 loopback 完全一致的启动 URL,不依赖任何查询参数名。它用临时 jar 证明 303 Cookie 交换和认证后根路径 200,再在默认 3 秒稳定窗口内持续证明 child 存活、唯一 listener 属于该 child 树、PID/启动 identity 不变且 retry 为 0。浏览器端改用插件同源路由上的持有式长轮询,不再永久每 500ms 请求:已证明的 previous listener 只落盘每标签页随机 capability 的哈希,并让所有仍响应的已登记标签页进入等待。只有在 ownership-stable 服务就绪和 canary 成功后,final listener 才在 Cookie 有效时通知各页刷新,或在收到 401 时把该最终进程的同源一次性 URL 返回内存,由页面执行 `location.replace()`;已认证页面回传 ACK 后,会丢弃 query/fragment 并回到原来的安全 pathname。一个真实 ACK 解除 terminal ready 门禁,其他尚未 ACK 的已登记页面在状态压缩后仍可恢复。没有原标签页登记或都未在超时内确认时,watchdog 才请求一次 system open 兜底,并继续等待新页面回传已认证 ACK;opener 的 exit 0 本身永远不算交接成功。服务端 readiness/canary 与浏览器交接分别记录,所需证据全部完成后才释放会话唤醒。裸 401、旧 listener 的 200、target 的短暂 200 或随后退出都不会成为 ready,也不会把已拒绝 target 的 URL 交给浏览器;原始 capability 和 bearer URL 均不进入状态文件、耐久日志或回执。`launch-cutover.json` 记录脱敏配置摘要、新旧 supervisor/child/listener identity、旧 supervisor 让权、认证就绪、浏览器 ACK 渠道、稳定性证明与分角色失败计数;target 的 readiness/canary 保留在 `targetValidation`,previous 的恢复 readiness 以及可用、失败或因只有 target-scoped credential 而明确跳过的恢复 canary 单独写在 `recovery.validation`,不会再出现 previous 已恢复却挂着一个无主语的 target canary fail。`launch-status` 输出该回执且不暴露两边命令。事务进行中可运行 `abort-cutover --state-dir "$DSH_HOME/state"` 执行事前批准的恢复策略;`restore-previous --state-dir "$DSH_HOME/state"` 显式授权停止已证明的 target 并恢复完整 previous spec。两个动作分别使用原子 marker,读取时 restore 永远优先,因此并发会话的晚到 abort 也不能降级 restore。
|
|
139
142
|
|
|
143
|
+
**监督链中途死亡时的 cutover 恢复。** 回执停在 `awaiting-user` 不再楔死事务。裸 `supervise`(launchd/systemd 跑的那条命令)现在以停泊形态拉起 watchdog:它取得监督所有权、什么都不启动(被拒的一侧绝不 boot),并持续消费运维控制 marker,因此 `abort-cutover --state-dir "$DSH_HOME/state"`(执行事前批准的策略)与 `restore-previous --state-dir "$DSH_HOME/state"` 立即可用,不再以「没有存活 watchdog」拒绝。停泊态会一次性说明自己的退出与恢复路径:`touch "$DSH_HOME/state/watchdog-stop"` 不做终态结算就结束停泊;`supervise --cutover-id <id> --state-dir "$DSH_HOME/state"`(id 见 `launch-status` 与回执)恢复选中侧——resume 把回执翻离 `awaiting-user`,停泊态随即释放所有权,恢复链先等到这次释放再接管。
|
|
144
|
+
|
|
140
145
|
candidate 无法读取旧宿主留下的可重建投影或缓存时,`--transition-file` 可以提交一份经过评审的 schema-v1 隔离计划。计划只接受 `home` 下互不重叠、没有符号链接且不包含 guard state 的相对路径,以及显式的 `quarantine` 操作;它不内置任何宿主版本或文件名知识。示例:`{"schemaVersion":1,"home":"/absolute/dsh-home","operations":[{"kind":"quarantine","path":"storages/<可重建缓存>","expect":"present"}]}`。每项 `expect` 必须是 `present` 或 `absent`,副本 preflight 与 live apply 都必须观察到相同状态,否则在停 previous 前或启动 target 前拒绝。计划既要覆盖 target 启动前必须移开的旧路径,也要覆盖 target 失败后 previous 启动前必须清走的新输出路径;后者即使准备时不存在也必须用 `expect: "absent"` 显式列出。权威日志、凭据或不可重建数据不得借此移出;需要内容转换的格式应使用独立、可逆且另行评审的迁移工具。
|
|
141
146
|
|
|
142
|
-
guard 先以 copy-on-write 优先方式复制 live home
|
|
147
|
+
guard 先以 copy-on-write 优先方式复制 live home 的启动输入——白名单为 `profiles/`、home 级 `cordis.patch.yml`、`settings.yaml`、`.credentials.yaml`、`.anonymous-user-id`;插件数据目录(sessions、state、local-agent 子 home、tarballs、scratch 等)默认不进复制,新插件的数据目录不会悄悄扩大快照,名单只在宿主启动读取范围变化时才需要更新(漏项会以预检 FAIL 显形,而不是慢慢变大)。再把 pnpm/Cordis 链接重建为只指向 snapshot 内副本的相对链接;合法的内部依赖循环保留。外部目标进入 snapshot 自己的哈希命名去重物化区,`node_modules` 目标锚定在包自身的父 `node_modules`(保留 Node 祖先查找语义的最小范围),而不是整个 store 根。复制后逐链接 `realpath` 审计,任何可写目标都必须仍在 snapshot 根内;无可复制语义的运行时条目(socket、FIFO 及指向它们的链接)跳过并计数;悬空或不可解析链接、其余特殊文件、可写逃逸及读取/复制失败都会让 `reconfigure` 在运行 candidate、创建 cutover 或停止 previous 前 fail closed。在安全副本中执行相同隔离后才运行 target composition preflight;副本无法准备或 target 无法 boot 时,previous 继续运行且 live home 不变。successor 取得监督所有权、停止并复核 previous 进程树以后,才按照哈希绑定的耐久计划用同文件系统 rename 隔离原路径。target 被拒绝时,watchdog 必须先停止其已证明的进程,再把它在同路径产生的替代内容保留到 `launch-transitions/<cutover>/rejected-target/`,恢复 previous 原字节并写入回执,最后才允许 previous 启动;任一步无法证明完成都会停在 `awaiting-user`,不会让旧宿主读取混合状态。target 成功后,旧内容仍保存在 cutover 目录,等待 operator 后续处置,不会自动删除。
|
|
143
148
|
|
|
144
149
|
**重启报告自动到达模型——并只等它的主人。** 计划重启后(存在未确认的 `last-restart.json` 记录),插件通过 `agent.followup` 把报告排入下一回合,agent 无需任何用户消息即可回报重启结果。重启后的会话恢复是 lazy 的(只有 UI 或 RPC 碰到某个会话,它的 agent 才会被创建),所以完整报告只发给发起重启的会话(`schedule-exit` 把 `$DSH_SESSION_ID` 记为 initiator),等它何时恢复何时送达——其他会话永远不会为了报告被唤醒;记录保持未确认,直到发起会话恢复或下一次重启替换它(新 `exitAt`)。没有 initiator 的记录由首个创建的根 agent 领走。仅根 agent、仅一次(送达即确认)。配置 `reportRestartContext`:`followup`(默认,自主)、`step`(骑在下一次回合的第一步上)、或 `off`。
|
|
145
150
|
|
package/lib/cli.js
CHANGED
|
@@ -805,12 +805,26 @@ function snapshotCopyError(source, error) {
|
|
|
805
805
|
return /* @__PURE__ */ new Error(`preflight snapshot could not safely copy ${source}: ${String(error)}`);
|
|
806
806
|
}
|
|
807
807
|
/**
|
|
808
|
-
* Top-level home entries
|
|
809
|
-
*
|
|
810
|
-
*
|
|
811
|
-
*
|
|
808
|
+
* Top-level home entries the preflight snapshot copies — the ALLOWLIST of
|
|
809
|
+
* inputs the launcher's boot actually reads: the profile trees, the home-level
|
|
810
|
+
* patch layer, settings, and the credential/identity stores. Everything else
|
|
811
|
+
* (plugin data: sessions, state, local-agent sub-homes, tarballs, scratch, …)
|
|
812
|
+
* is excluded BY DEFAULT, so a newly installed plugin's data directory can
|
|
813
|
+
* never silently join the copy — this list moves only when the HOST's boot
|
|
814
|
+
* starts reading a new home input, and a miss fails the dry-run loudly with
|
|
815
|
+
* the missing path rather than degrading into a slow copy. The denylist this
|
|
816
|
+
* replaced failed twice the other way: a 24 GB scratch tree expired the
|
|
817
|
+
* credential mid-cutover (canary failed, restored), and on 2026-09-27 the
|
|
818
|
+
* local-agent sub-home's absolute links dragged the host checkout's entire
|
|
819
|
+
* node_modules into a ~4 GB / 848 s prepare.
|
|
812
820
|
*/
|
|
813
|
-
const
|
|
821
|
+
const SNAPSHOT_INCLUDED_TOP_LEVEL = [
|
|
822
|
+
"profiles",
|
|
823
|
+
"settings.yaml",
|
|
824
|
+
"cordis.patch.yml",
|
|
825
|
+
".credentials.yaml",
|
|
826
|
+
".anonymous-user-id"
|
|
827
|
+
];
|
|
814
828
|
function canonicalSnapshotSource(source) {
|
|
815
829
|
try {
|
|
816
830
|
return realpathSync(source);
|
|
@@ -859,7 +873,7 @@ function copySnapshotNode(source, destination, context) {
|
|
|
859
873
|
});
|
|
860
874
|
try {
|
|
861
875
|
for (const name of readdirSync(canonical)) {
|
|
862
|
-
if (canonical === context.rootCanonical &&
|
|
876
|
+
if (canonical === context.rootCanonical && !context.includeTopLevel.has(name)) continue;
|
|
863
877
|
copySnapshotNode(join(canonical, name), join(destination, name), context);
|
|
864
878
|
}
|
|
865
879
|
} catch (error) {
|
|
@@ -872,18 +886,32 @@ function copySnapshotNode(source, destination, context) {
|
|
|
872
886
|
copyFileSync(canonical, destination, constants.COPYFILE_FICLONE);
|
|
873
887
|
chmodSync(destination, linkMetadata.mode & 4095);
|
|
874
888
|
utimesSync(destination, linkMetadata.atime, linkMetadata.mtime);
|
|
889
|
+
context.copiedFiles++;
|
|
890
|
+
context.copiedBytes += linkMetadata.size;
|
|
891
|
+
context.onProgress?.({
|
|
892
|
+
files: context.copiedFiles,
|
|
893
|
+
bytes: context.copiedBytes,
|
|
894
|
+
skippedRuntimeEntries: context.skippedRuntimeEntries
|
|
895
|
+
});
|
|
875
896
|
} catch (error) {
|
|
876
897
|
throw snapshotCopyError(source, error);
|
|
877
898
|
}
|
|
878
899
|
}
|
|
879
900
|
/**
|
|
880
901
|
* Preserve Node's ancestor node_modules lookup for an external package while
|
|
881
|
-
* avoiding one copy per package link.
|
|
882
|
-
*
|
|
902
|
+
* avoiding one copy per package link. The anchor is the package's OWN parent
|
|
903
|
+
* node_modules — the LAST node_modules segment in the resolved target: pnpm
|
|
904
|
+
* store links resolve to …/.pnpm/<name>@<version>/node_modules/<name>, where
|
|
905
|
+
* that parent already holds the package's dependency siblings, so the lookup
|
|
906
|
+
* survives at the tightest scope. Anchoring the FIRST segment instead dragged
|
|
907
|
+
* the whole multi-GB store root into the snapshot (observed 2026-09-27: 32
|
|
908
|
+
* profile links pulled in the host checkout's entire root node_modules).
|
|
909
|
+
* Other external targets are materialized individually and still deduplicated
|
|
910
|
+
* by canonical path.
|
|
883
911
|
*/
|
|
884
912
|
function externalMaterializationAnchor(target) {
|
|
885
913
|
const parsed = resolve(target).split(sep);
|
|
886
|
-
const nodeModulesIndex = parsed.
|
|
914
|
+
const nodeModulesIndex = parsed.lastIndexOf("node_modules");
|
|
887
915
|
if (nodeModulesIndex >= 0) return {
|
|
888
916
|
source: parsed.slice(0, nodeModulesIndex + 1).join(sep) || sep,
|
|
889
917
|
destination: "node_modules"
|
|
@@ -958,17 +986,21 @@ function finalizeSnapshotDirectories(context) {
|
|
|
958
986
|
}
|
|
959
987
|
}
|
|
960
988
|
/**
|
|
961
|
-
* Clone a live home while retaining a contained package-link
|
|
962
|
-
*
|
|
963
|
-
*
|
|
964
|
-
*
|
|
965
|
-
*
|
|
966
|
-
*
|
|
967
|
-
*
|
|
989
|
+
* Clone a live home's BOOT INPUTS while retaining a contained package-link
|
|
990
|
+
* graph. Only the allowlisted top-level entries are copied (see
|
|
991
|
+
* {@link SNAPSHOT_INCLUDED_TOP_LEVEL}) — plugin data directories are excluded
|
|
992
|
+
* by default, so the copy's size is bounded by what the composition's boot
|
|
993
|
+
* reads, not by whatever the home happens to hold. Internal links are rebuilt
|
|
994
|
+
* against copied nodes; external targets are deduplicated in a snapshot-owned
|
|
995
|
+
* materialization area. No retained link resolves outside the snapshot root,
|
|
996
|
+
* so writes through pnpm/Cordis links cannot reach live bytes. Runtime entries
|
|
997
|
+
* without copyable content (sockets, FIFOs — and links to them) are skipped
|
|
998
|
+
* and counted, never copied. Device nodes still fail closed.
|
|
968
999
|
* @param sourceHome - Live dsh home to read.
|
|
969
|
-
* @
|
|
1000
|
+
* @param options - Include-list override and progress callback.
|
|
1001
|
+
* @returns Isolated home, an idempotent cleanup callback, and copy statistics.
|
|
970
1002
|
*/
|
|
971
|
-
function createPreflightSnapshot(sourceHome) {
|
|
1003
|
+
function createPreflightSnapshot(sourceHome, options = {}) {
|
|
972
1004
|
const root = mkdtempSync(join(tmpdir(), "ankh-transition-preflight-"));
|
|
973
1005
|
const home = join(root, "home");
|
|
974
1006
|
try {
|
|
@@ -977,10 +1009,14 @@ function createPreflightSnapshot(sourceHome) {
|
|
|
977
1009
|
const context = {
|
|
978
1010
|
externalRoot: join(root, "materialized"),
|
|
979
1011
|
rootCanonical: source,
|
|
1012
|
+
includeTopLevel: new Set(options.includeTopLevel ?? SNAPSHOT_INCLUDED_TOP_LEVEL),
|
|
980
1013
|
destinations: /* @__PURE__ */ new Map(),
|
|
981
1014
|
pendingLinks: [],
|
|
982
1015
|
directories: [],
|
|
983
|
-
skippedRuntimeEntries: 0
|
|
1016
|
+
skippedRuntimeEntries: 0,
|
|
1017
|
+
copiedFiles: 0,
|
|
1018
|
+
copiedBytes: 0,
|
|
1019
|
+
...options.onProgress === void 0 ? {} : { onProgress: options.onProgress }
|
|
984
1020
|
};
|
|
985
1021
|
copySnapshotNode(source, home, context);
|
|
986
1022
|
resolveSnapshotLinks(context);
|
|
@@ -990,6 +1026,8 @@ function createPreflightSnapshot(sourceHome) {
|
|
|
990
1026
|
home,
|
|
991
1027
|
root,
|
|
992
1028
|
skippedRuntimeEntries: context.skippedRuntimeEntries,
|
|
1029
|
+
copiedFiles: context.copiedFiles,
|
|
1030
|
+
copiedBytes: context.copiedBytes,
|
|
993
1031
|
cleanup: () => {
|
|
994
1032
|
rmSync(root, {
|
|
995
1033
|
recursive: true,
|
|
@@ -1005,8 +1043,12 @@ function createPreflightSnapshot(sourceHome) {
|
|
|
1005
1043
|
throw error;
|
|
1006
1044
|
}
|
|
1007
1045
|
}
|
|
1008
|
-
function createTransitionPreflightSnapshot(plan) {
|
|
1009
|
-
const
|
|
1046
|
+
function createTransitionPreflightSnapshot(plan, options = {}) {
|
|
1047
|
+
const operationRoots = plan.operations.map((operation) => operation.path.split(sep)[0]);
|
|
1048
|
+
const snapshot = createPreflightSnapshot(plan.home, {
|
|
1049
|
+
...options,
|
|
1050
|
+
includeTopLevel: [.../* @__PURE__ */ new Set([...options.includeTopLevel ?? SNAPSHOT_INCLUDED_TOP_LEVEL, ...operationRoots])]
|
|
1051
|
+
});
|
|
1010
1052
|
try {
|
|
1011
1053
|
const stateDir = join(snapshot.root, "guard-state");
|
|
1012
1054
|
mkdirSync(stateDir, {
|
|
@@ -1017,10 +1059,7 @@ function createTransitionPreflightSnapshot(plan) {
|
|
|
1017
1059
|
...plan,
|
|
1018
1060
|
home: snapshot.home
|
|
1019
1061
|
}, snapshot.home, stateDir, "preflight"), snapshot.home, stateDir, "preflight");
|
|
1020
|
-
return
|
|
1021
|
-
home: snapshot.home,
|
|
1022
|
-
cleanup: snapshot.cleanup
|
|
1023
|
-
};
|
|
1062
|
+
return snapshot;
|
|
1024
1063
|
} catch (error) {
|
|
1025
1064
|
snapshot.cleanup();
|
|
1026
1065
|
throw error;
|
|
@@ -1234,12 +1273,14 @@ commands:
|
|
|
1234
1273
|
[--home DIR] [--repo DIR] [--harness-root DIR] [--profile NAME] [--browser-handoff required|off]
|
|
1235
1274
|
--preflight-surface source|built [--preflight-runner FILE] --preflight-install-anchor FILE
|
|
1236
1275
|
--candidate-probe-command "CMD"
|
|
1237
|
-
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
1276
|
+
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
1277
|
+
[--boot-timeout-ms MS] [--state-dir DIR]
|
|
1238
1278
|
restart --port N --start "CMD" [--pid PID] [--timeout-ms MS] [--delay-ms MS] [--stop-timeout-ms MS] [--rollback]
|
|
1239
1279
|
[--profile NAME] [--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR] [--max-age MIN]
|
|
1240
1280
|
schedule-exit [--port N] --delay-ms MS [--initiator ID] [--log FILE] [--profile NAME]
|
|
1241
|
-
[--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
1281
|
+
[--harness-root DIR] [--preflight-timeout-ms MS] [--boot-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
1242
1282
|
supervise --port N --start "CMD" [--foreground] [--log FILE] [--state-dir DIR] [--repo DIR] [--harness-root DIR] [--home DIR]
|
|
1283
|
+
[--cutover-id ID] [--boot-timeout-ms MS]
|
|
1243
1284
|
flags:
|
|
1244
1285
|
--state-dir DIR state directory (default: $DSH_HOME/state, else <cwd>/.dsh-guard-state)
|
|
1245
1286
|
--repo DIR repository the credential binds to (default: cwd)
|
|
@@ -1266,6 +1307,16 @@ flags:
|
|
|
1266
1307
|
(agent-driven graceful self-restart: schedule, complete, then restart);
|
|
1267
1308
|
schedule-exit: delay before the detached exit agent kills the host;
|
|
1268
1309
|
reconfigure: grace after successor supervisor claim before old-child stop
|
|
1310
|
+
--boot-timeout-ms MS reconfigure/schedule-exit/supervise: readiness budget for one
|
|
1311
|
+
boot (the watchdog's WD_BOOT_TIMEOUT, whole seconds, default 60).
|
|
1312
|
+
reconfigure/supervise hand it to the spawned watchdog's environment;
|
|
1313
|
+
schedule-exit records it in the restart marker for the respawn's boot
|
|
1314
|
+
window only (the running watchdog then falls back to its own budget)
|
|
1315
|
+
--cutover-id ID supervise: resume the durable launch-cutover transaction after its
|
|
1316
|
+
supervisor chain died (the id is in launch-status / the receipt);
|
|
1317
|
+
without it a bare supervise on an awaiting-user receipt only holds
|
|
1318
|
+
the claim and consumes operator control markers, never booting the
|
|
1319
|
+
rejected side
|
|
1269
1320
|
--log FILE supervise (detached only — with --foreground the external supervisor's
|
|
1270
1321
|
redirection owns the log) / schedule-exit: log file (default: <state-dir>/*.log)
|
|
1271
1322
|
--home DIR supervise: the dsh home the supervised instance boots with (profiles,
|
|
@@ -1322,6 +1373,7 @@ function parse(argv) {
|
|
|
1322
1373
|
delayMs: void 0,
|
|
1323
1374
|
stopTimeoutMs: void 0,
|
|
1324
1375
|
supervisorYieldTimeoutMs: void 0,
|
|
1376
|
+
bootTimeoutMs: void 0,
|
|
1325
1377
|
log: void 0,
|
|
1326
1378
|
foreground: false,
|
|
1327
1379
|
rollback: false,
|
|
@@ -1455,6 +1507,14 @@ function parse(argv) {
|
|
|
1455
1507
|
i++;
|
|
1456
1508
|
break;
|
|
1457
1509
|
}
|
|
1510
|
+
case "--boot-timeout-ms": {
|
|
1511
|
+
const raw = flagValue(arg, true);
|
|
1512
|
+
const n = Number(raw);
|
|
1513
|
+
if (raw === void 0 || !Number.isInteger(n) || n < 1e3) throw new Error("--boot-timeout-ms must be an integer >= 1000");
|
|
1514
|
+
options.bootTimeoutMs = n;
|
|
1515
|
+
i++;
|
|
1516
|
+
break;
|
|
1517
|
+
}
|
|
1458
1518
|
case "--foreground":
|
|
1459
1519
|
options.foreground = true;
|
|
1460
1520
|
break;
|
|
@@ -1644,9 +1704,14 @@ function verifyRepoCredential(stateDir, repoDir, maxAgeMinutes) {
|
|
|
1644
1704
|
return verifyCredential(loadState(stateDir), currentHead(repoDir), Date.now(), maxAgeMinutes, isWorkingTreeClean(repoDir));
|
|
1645
1705
|
}
|
|
1646
1706
|
/**
|
|
1647
|
-
* How the spawned watchdog should invoke the guard CLI
|
|
1648
|
-
*
|
|
1649
|
-
*
|
|
1707
|
+
* How the spawned watchdog should invoke the guard CLI. The executable is
|
|
1708
|
+
* always the ABSOLUTE process.execPath, never a bare `node`: the watchdog runs
|
|
1709
|
+
* under whatever PATH spawned it, and a launcher chain (launchd/systemd) has a
|
|
1710
|
+
* minimal PATH without homebrew — a bare `node` there makes every guard
|
|
1711
|
+
* invocation the watchdog issues (canary, record-proven-deployment, cutover
|
|
1712
|
+
* events) fail with command-not-found (the 2026-09-30 launchd incident). The
|
|
1713
|
+
* source form additionally needs tsx with an absolute path (the watchdog runs
|
|
1714
|
+
* with a deployment cwd that resolves no node_modules).
|
|
1650
1715
|
* @returns the command prefix (verb args are appended by the watchdog).
|
|
1651
1716
|
*/
|
|
1652
1717
|
function guardInvocation() {
|
|
@@ -1654,10 +1719,10 @@ function guardInvocation() {
|
|
|
1654
1719
|
if (cliPath.includes(`${sep}src${sep}`)) {
|
|
1655
1720
|
const nodeModules = resolve(dirname(cliPath), "../../../node_modules");
|
|
1656
1721
|
const tsx = join(nodeModules, "tsx", "dist", "esm", "index.mjs");
|
|
1657
|
-
if (existsSync(tsx)) return
|
|
1658
|
-
return
|
|
1722
|
+
if (existsSync(tsx)) return `${process.execPath} --import ${tsx} ${cliPath}`;
|
|
1723
|
+
return `${process.execPath} ${cliPath}`;
|
|
1659
1724
|
}
|
|
1660
|
-
return
|
|
1725
|
+
return `${process.execPath} ${cliPath}`;
|
|
1661
1726
|
}
|
|
1662
1727
|
/**
|
|
1663
1728
|
* argv (after process.execPath) that runs the exit agent, with the same
|
|
@@ -2513,7 +2578,7 @@ async function runCli(argv, io) {
|
|
|
2513
2578
|
}
|
|
2514
2579
|
const watchdogPid = liveWatchdogPid(stateDir);
|
|
2515
2580
|
if (watchdogPid === null) {
|
|
2516
|
-
io.stderr(`${command} refused: no live watchdog can consume the durable control request\n`);
|
|
2581
|
+
io.stderr(`${command} refused: no live watchdog can consume the durable control request. Start the consumer first: \`supervise --state-dir ${stateDir}\` holds the awaiting-user cutover ${transaction.receipt.id} without launching anything, or \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\` resumes the selected side directly\n`);
|
|
2517
2582
|
return 1;
|
|
2518
2583
|
}
|
|
2519
2584
|
const requested = command === "restore-previous" ? "restore-previous" : "abort";
|
|
@@ -2785,12 +2850,20 @@ async function runCli(argv, io) {
|
|
|
2785
2850
|
const snapshotStartedAt = Date.now();
|
|
2786
2851
|
let snapshot;
|
|
2787
2852
|
try {
|
|
2788
|
-
|
|
2853
|
+
let lastProgressAt = 0;
|
|
2854
|
+
const onProgress = (progress) => {
|
|
2855
|
+
const now = Date.now();
|
|
2856
|
+
if (now - lastProgressAt < 2e3) return;
|
|
2857
|
+
lastProgressAt = now;
|
|
2858
|
+
io.stdout(`preflight snapshot: ${progress.files} files / ${Math.round(progress.bytes / 1024 / 1024)} MB copied…\n`);
|
|
2859
|
+
};
|
|
2860
|
+
snapshot = transitionPlan === void 0 ? createPreflightSnapshot(target.home, { onProgress }) : createTransitionPreflightSnapshot(transitionPlan, { onProgress });
|
|
2789
2861
|
} catch (error) {
|
|
2790
2862
|
return refuse("preflight-snapshot", `reconfigure refused: could not prepare an isolated${transitionPlan === void 0 ? "" : " transitioned"} home: ${String(error)}\n`);
|
|
2791
2863
|
}
|
|
2792
2864
|
const snapshotMs = Date.now() - snapshotStartedAt;
|
|
2793
|
-
|
|
2865
|
+
io.stdout(`isolated home snapshot ready: ${snapshot.copiedFiles} files / ${Math.round(snapshot.copiedBytes / 1024 / 1024)} MB in ${Math.round(snapshotMs / 1e3)}s\n`);
|
|
2866
|
+
if (snapshotMs > options.maxAgeMinutes * 6e4 / 2) io.stdout(`note: the isolated-home snapshot took ${Math.round(snapshotMs / 1e3)}s — over half the ${options.maxAgeMinutes}min credential window; re-record the credential immediately before reconfigure\n`);
|
|
2794
2867
|
try {
|
|
2795
2868
|
const timeout = options.preflightTimeoutMs ?? DEFAULT_PREFLIGHT_TIMEOUT_MS;
|
|
2796
2869
|
if (!await candidateProbeGate(target, timeout, io, snapshot.home)) return refuseQuiet("preflight", "reconfigure refused: the candidate probe failed (see stderr)");
|
|
@@ -2845,6 +2918,7 @@ async function runCli(argv, io) {
|
|
|
2845
2918
|
String(options.delayMs ?? 5e3),
|
|
2846
2919
|
"--supervisor-yield-timeout-ms",
|
|
2847
2920
|
String(options.supervisorYieldTimeoutMs ?? 15e3),
|
|
2921
|
+
...options.bootTimeoutMs !== void 0 ? ["--boot-timeout-ms", String(options.bootTimeoutMs)] : [],
|
|
2848
2922
|
...initiator !== void 0 ? ["--initiator", initiator] : []
|
|
2849
2923
|
];
|
|
2850
2924
|
const cutoverDriverEnv = { ...process.env };
|
|
@@ -3067,10 +3141,9 @@ async function runCli(argv, io) {
|
|
|
3067
3141
|
io.stderr(`supervise refused: cutover ${options.cutoverId} is not the selected launch transaction\n`);
|
|
3068
3142
|
return 1;
|
|
3069
3143
|
}
|
|
3070
|
-
|
|
3071
|
-
|
|
3072
|
-
|
|
3073
|
-
}
|
|
3144
|
+
let parkedCutover = false;
|
|
3145
|
+
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user") parkedCutover = true;
|
|
3146
|
+
const resumeFromAwaitingUser = options.cutoverId !== void 0 && transaction?.receipt.phase === "awaiting-user";
|
|
3074
3147
|
if (options.foreground !== true && !sandboxGate("supervise", options, io)) return 2;
|
|
3075
3148
|
if (options.cutoverId !== void 0) try {
|
|
3076
3149
|
const driverIdentity = processIdentity(process.pid);
|
|
@@ -3105,7 +3178,22 @@ async function runCli(argv, io) {
|
|
|
3105
3178
|
return 1;
|
|
3106
3179
|
}
|
|
3107
3180
|
if (durable === null) writeStableLaunchSpec(stateDir, spec);
|
|
3108
|
-
if (
|
|
3181
|
+
if (resumeFromAwaitingUser) {
|
|
3182
|
+
const existingIdentity = processIdentity(existingPid);
|
|
3183
|
+
if (existingIdentity === null) {
|
|
3184
|
+
io.stderr(`supervise refused: could not capture watchdog ${existingPid} start identity before waiting\n`);
|
|
3185
|
+
return 1;
|
|
3186
|
+
}
|
|
3187
|
+
io.stdout(`watchdog ${existing} holds the parked cutover — waiting for it to release, then resuming\n`);
|
|
3188
|
+
const releaseDeadline = Date.now() + 15e3;
|
|
3189
|
+
while (processIdentityMatches(existingIdentity) && Date.now() < releaseDeadline) await sleep(250);
|
|
3190
|
+
if (processIdentityMatches(existingIdentity)) {
|
|
3191
|
+
io.stderr(`supervise refused: watchdog ${existingPid} did not release the parked cutover ${options.cutoverId ?? ""} within 15000 ms — it is not the awaiting-user hold; settle the transaction with \`abort-cutover --state-dir ${stateDir}\` or \`restore-previous --state-dir ${stateDir}\`, or stop that watchdog and retry\n`);
|
|
3192
|
+
return 1;
|
|
3193
|
+
}
|
|
3194
|
+
waitedForWatchdog = true;
|
|
3195
|
+
io.stdout(`watchdog ${existing} released the parked cutover — resuming\n`);
|
|
3196
|
+
} else if (options.foreground) {
|
|
3109
3197
|
io.stdout(`watchdog ${existing} already supervises the port — waiting for it to exit, then taking over (foreground)\n`);
|
|
3110
3198
|
const existingIdentity = processIdentity(existingPid);
|
|
3111
3199
|
if (existingIdentity === null) {
|
|
@@ -3132,10 +3220,7 @@ async function runCli(argv, io) {
|
|
|
3132
3220
|
io.stderr(`supervise refused after wait: cutover ${options.cutoverId} is no longer the selected launch transaction\n`);
|
|
3133
3221
|
return 1;
|
|
3134
3222
|
}
|
|
3135
|
-
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user")
|
|
3136
|
-
io.stderr(`supervise: cutover ${transaction.receipt.id} settled awaiting-user while this supervisor waited; refusing to restart the rejected target (receipt ${stateFile(stateDir, "launchCutover")})\n`);
|
|
3137
|
-
return 0;
|
|
3138
|
-
}
|
|
3223
|
+
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user") parkedCutover = true;
|
|
3139
3224
|
io.stdout("launch state refreshed after wait — using the durable selected specification\n");
|
|
3140
3225
|
}
|
|
3141
3226
|
const previousOwnership = transaction?.receipt.ownership?.previous;
|
|
@@ -3184,6 +3269,7 @@ async function runCli(argv, io) {
|
|
|
3184
3269
|
WD_PROFILE: spec.profile,
|
|
3185
3270
|
WD_WAIT_OWNER: options.takeoverFrom !== void 0 || !options.foreground ? "1" : "0",
|
|
3186
3271
|
WD_GUARD: guardInvocation(),
|
|
3272
|
+
...options.bootTimeoutMs !== void 0 ? { WD_BOOT_TIMEOUT: String(Math.ceil(options.bootTimeoutMs / 1e3)) } : {},
|
|
3187
3273
|
...options.takeoverFrom !== void 0 ? {
|
|
3188
3274
|
WD_TAKEOVER_FROM: String(options.takeoverFrom),
|
|
3189
3275
|
WD_TAKEOVER_FROM_START: previousSupervisorStart ?? ""
|
|
@@ -3204,9 +3290,11 @@ async function runCli(argv, io) {
|
|
|
3204
3290
|
WD_PREVIOUS_CHILD_START: previousOwnership.childStartToken,
|
|
3205
3291
|
WD_PREVIOUS_LISTENER_PID: String(previousOwnership.listenerPid),
|
|
3206
3292
|
WD_PREVIOUS_LISTENER_START: previousOwnership.listenerStartToken,
|
|
3293
|
+
...parkedCutover ? { WD_CUTOVER_PARKED: "1" } : {},
|
|
3207
3294
|
...transaction.state.transition === void 0 ? {} : { WD_TRANSITION_PLAN_SHA256: transaction.state.transition.planSha256 }
|
|
3208
3295
|
} : {}
|
|
3209
3296
|
};
|
|
3297
|
+
if (parkedCutover && transaction !== null) io.stdout(`supervise: cutover ${transaction.receipt.id} is waiting for user action — holding the supervision claim WITHOUT launching the rejected ${transaction.state.selected} side (receipt ${stateFile(stateDir, "launchCutover")}). The operator verbs now have a live consumer: \`abort-cutover --state-dir ${stateDir}\` applies the pre-approved ${transaction.receipt.recovery.policy} policy, \`restore-previous --state-dir ${stateDir}\` restores the complete previous spec. To end the hold without settling: write the stop marker (\`touch ${stateFile(stateDir, "watchdogStop")}\`). To resume the rejected side: \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\`\n`);
|
|
3210
3298
|
if (options.foreground) {
|
|
3211
3299
|
const child = spawn("bash", [watchdog, "--supervise"], {
|
|
3212
3300
|
stdio: "inherit",
|
|
@@ -3256,7 +3344,7 @@ async function runCli(argv, io) {
|
|
|
3256
3344
|
const claimDeadline = Date.now() + 5e3;
|
|
3257
3345
|
while (Date.now() < claimDeadline && spawnError === void 0) {
|
|
3258
3346
|
if (liveWatchdogPid(stateDir) === spawnedPid) {
|
|
3259
|
-
io.stdout(`watchdog spawned and ready (pid ${spawnedPid}) — supervises :${spec.port}, log ${logPath}\n`);
|
|
3347
|
+
io.stdout(parkedCutover && transaction !== null ? `watchdog spawned and parked on cutover ${transaction.receipt.id} (pid ${spawnedPid}) — nothing launched; consuming abort-cutover/restore-previous markers; log ${logPath}\n` : `watchdog spawned and ready (pid ${spawnedPid}) — supervises :${spec.port}, log ${logPath}\n`);
|
|
3260
3348
|
return 0;
|
|
3261
3349
|
}
|
|
3262
3350
|
try {
|
|
@@ -3327,6 +3415,7 @@ async function runCli(argv, io) {
|
|
|
3327
3415
|
writeFileSync(stateFile(stateDir, "restartRequested"), `${JSON.stringify({
|
|
3328
3416
|
reason: "scheduled self-restart",
|
|
3329
3417
|
requestedAt: Date.now(),
|
|
3418
|
+
...options.bootTimeoutMs === void 0 ? {} : { bootTimeoutMs: options.bootTimeoutMs },
|
|
3330
3419
|
...gate.authorization === void 0 ? {} : { authorization: gate.authorization },
|
|
3331
3420
|
...initiator !== void 0 ? { initiator } : {}
|
|
3332
3421
|
})}\n`);
|
|
@@ -3380,4 +3469,4 @@ if (isDirectInvocation(import.meta.url)) {
|
|
|
3380
3469
|
});
|
|
3381
3470
|
}
|
|
3382
3471
|
//#endregion
|
|
3383
|
-
export { envInternals, parse, preflightInternals, resolveHarnessRoot, resolvePreflightBin, resolveRunnerCommand, resolveWdHome, runCli, runPreflightCheck };
|
|
3472
|
+
export { envInternals, guardInvocation, parse, preflightInternals, resolveHarnessRoot, resolvePreflightBin, resolveRunnerCommand, resolveWdHome, runCli, runPreflightCheck };
|