@khorsheed/dsh-ankh-guard 0.4.0 → 0.4.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.en.md +7 -2
- package/README.i18n.yaml +2 -2
- package/README.md +7 -2
- package/lib/cli.js +62 -20
- package/lib/client.js +20 -0
- package/lib/types/cli.d.ts +13 -0
- package/lib/types/cli.js +102 -16
- package/lib/types/client/index.js +27 -0
- package/package.json +9 -9
- package/scripts/dsh-watchdog.sh +108 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# 变更记录
|
|
2
2
|
|
|
3
|
+
## 0.4.2(2026-09-30)
|
|
4
|
+
|
|
5
|
+
- **修复重启浮层卡死**:标签页带着 pending handoff 时,若 cutover 已不经由本标签页收口(操作员 restore、handoff off、监督链楔死),服务端对该 cutoverId 永远回 409,「正在重启」浮层会无限重试不下台。现在每 8 次失败 ACK 花一次 poll——`idle` 即证明没有在途事务,丢弃 pending 并收浮层。sessionStorage 按标签页存活,所以旧实例上已卡住的标签页不受影响此修复——关掉重开即可(3080 实测案例)
|
|
6
|
+
|
|
7
|
+
## 0.4.1(2026-09-30)
|
|
8
|
+
|
|
9
|
+
修复 2026-09-29/30 的 0.2.0-rc.2 切换事故暴露的三个缺口:
|
|
10
|
+
|
|
11
|
+
- **guardInvocation() 改用绝对 execPath**:build/source 两种形态均用绝对 `process.execPath`——裸 `node` 在 launcher 的 PATH 下找不到
|
|
12
|
+
- **awaiting-user 收据上的裸 supervise 改为泊驻**:占用 pidfile、不拉起进程、消费 abort/restore 控制标记——不再直接退出把事务楔死在没有活体消费者的状态;`supervise --cutover-id` 释放泊驻并恢复;拒绝消息打印确切可执行的命令
|
|
13
|
+
- **失败 boot 即使日志为空也镜像 attempt 日志**(「produced no output」本身就是信号);`reconfigure`/`schedule-exit`/`supervise` 新增 `--boot-timeout-ms`(穿到 `WD_BOOT_TIMEOUT`)
|
|
14
|
+
- 加宽 `@deepseek-ai/dsh-*` peer 区间以覆盖宿主 0.2.0
|
|
15
|
+
|
|
3
16
|
## 0.3.2(2026-09-27)
|
|
4
17
|
|
|
5
18
|
适配宿主 0.1.7-rc.2 线:verifiedHost 前移至 0.1.7-rc.2(3080 生产实证线随宿主基线切到 rc.2);rc.1→rc.2 对本包无破坏性变更(逐类清点见 [Agent Note](../../.agents/notes/implemented/architecture/2026-09-27-host-017-rc2-breaking-changes.md)),全量构建+测试双绿。
|
package/README.en.md
CHANGED
|
@@ -109,9 +109,9 @@ dsh-ankh-guard supervise --port 3080 --start "CMD" --state-dir "$DSH_HOME/state"
|
|
|
109
109
|
|
|
110
110
|
`supervise` also needs the dsh home the supervised instance boots with (the watchdog exports it as the instance's `DSH_HOME`): `--home DIR` wins, else `$DSH_HOME`; with neither set it refuses loudly — a home guessed from `--state-dir` would silently boot the instance on the wrong profiles/credentials. First-time persistence also requires an explicit `--harness-root` or `DSH_HARNESS`; it never guesses the host root from the credential repo.
|
|
111
111
|
|
|
112
|
-
It spawns `scripts/dsh-watchdog.sh` (ships with the package) detached with `--wait-owner`: the watchdog idles while the current instance runs, takes over the port when the instance exits (intentional restart or crash), respawns it, runs the guard canary on intentional restarts (a `restart-requested.json` marker), and clears the marker on pass. Two consecutive boot failures roll the checkout back to the last known-good revision — the healthy-boot stamp (`last-good-boot.json`, written every time the instance comes up, so it names the last revision that genuinely ran in this deployment), else the guard checkpoint, else the credential's HEAD — but only when the boot failure's error subject is a path inside the repository. When the subject lives outside the checkout (a broken profile overlay or an installed plugin), a checkout reset cannot help, so the watchdog instead restores the last healthy **profile composition**: the snapshot of the profile's composition inputs (`last-good-composition/`, taken at every healthy boot) replaces the live bundles layer and manifest, unmounting the newest plugin change, with the failing inputs preserved under `composition-backup-*` and the recovered report naming exactly what was unmounted. The same exemption logic covers a start command that does not bind the supervised port: when the boot window times out while the instance is listening elsewhere — or fails with `EADDRINUSE` naming a port this watchdog does not own — the watchdog names the bound port and skips both rollbacks, because resetting files cannot change a command-line argument. `EADDRINUSE` on the supervised port keeps its free-and-retry escape hatch, now bounded at five attempts. Every reset (watchdog, CLI, or service) first creates `guard-backup-*` branch anchors for the discarded HEAD and for uncommitted tracked changes, so recovery never depends on the reflog. Four failures serve a crash page on the port with a retry button (SIGUSR1 to the watchdog). A `watchdog-stop` marker exits the watchdog for good. The instance itself can adopt supervision before a self-restart — the user never starts the watchdog by hand.
|
|
112
|
+
It spawns `scripts/dsh-watchdog.sh` (ships with the package) detached with `--wait-owner`: the watchdog idles while the current instance runs, takes over the port when the instance exits (intentional restart or crash), respawns it, runs the guard canary on intentional restarts (a `restart-requested.json` marker), and clears the marker on pass. Two consecutive boot failures roll the checkout back to the last known-good revision — the healthy-boot stamp (`last-good-boot.json`, written every time the instance comes up, so it names the last revision that genuinely ran in this deployment), else the guard checkpoint, else the credential's HEAD — but only when the boot failure's error subject is a path inside the repository. When the subject lives outside the checkout (a broken profile overlay or an installed plugin), a checkout reset cannot help, so the watchdog instead restores the last healthy **profile composition**: the snapshot of the profile's composition inputs (`last-good-composition/`, taken at every healthy boot) replaces the live bundles layer and manifest, unmounting the newest plugin change, with the failing inputs preserved under `composition-backup-*` and the recovered report naming exactly what was unmounted. The same exemption logic covers a start command that does not bind the supervised port: when the boot window times out while the instance is listening elsewhere — or fails with `EADDRINUSE` naming a port this watchdog does not own — the watchdog names the bound port and skips both rollbacks, because resetting files cannot change a command-line argument. `EADDRINUSE` on the supervised port keeps its free-and-retry escape hatch, now bounded at five attempts. A failed boot always mirrors its captured attempt log into the main watchdog log — with an explicit `attempt N produced no output` line when the instance never wrote a line at all — so a dead-on-arrival boot is visible where an operator looks first. Every reset (watchdog, CLI, or service) first creates `guard-backup-*` branch anchors for the discarded HEAD and for uncommitted tracked changes, so recovery never depends on the reflog. Four failures serve a crash page on the port with a retry button (SIGUSR1 to the watchdog). A `watchdog-stop` marker exits the watchdog for good. The instance itself can adopt supervision before a self-restart — the user never starts the watchdog by hand.
|
|
113
113
|
|
|
114
|
-
When a watchdog is already supervising, the restart trigger is `schedule-exit`: it gets the port, credential repo, host root, and profile from the durable active launch spec, rejects conflicting explicit flags, and verifies the supervisor's complete command before writing the restart marker and spawning a detached process explicitly labelled `exit-agent pid`. It prefers a credential within the ten-minute freshness window. Once that expires, only an exact `provenDeployment` may take the pure-restart fast path; the selected evidence SHA is copied into the short-lived restart marker and the new watchdog rechecks that same SHA and the live fingerprint during canary, closing the check-to-stop evidence replacement window. A successful canary then promotes or retains the proof. For an agent Bash/tool call, set that call's `timeoutMs` to 180000; this is tool metadata, not a CLI flag, and keeps the caller alive longer than the 120-second preflight gate. The managed shell cannot reap the exit agent, so the scheduled kill lands after the scheduling turn ends. The watchdog respawns, runs the canary, and the new instance reports via `last-restart.json`; watchdog lifecycle lines are timestamped. With no live watchdog, `schedule-exit` hard-refuses: establish supervision first or use `restart`, which owns the complete single-shot loop. The latter lacks durable supervisor/launch ownership and therefore still requires a fresh credential.
|
|
114
|
+
When a watchdog is already supervising, the restart trigger is `schedule-exit`: it gets the port, credential repo, host root, and profile from the durable active launch spec, rejects conflicting explicit flags, and verifies the supervisor's complete command before writing the restart marker and spawning a detached process explicitly labelled `exit-agent pid`. It prefers a credential within the ten-minute freshness window. Once that expires, only an exact `provenDeployment` may take the pure-restart fast path; the selected evidence SHA is copied into the short-lived restart marker and the new watchdog rechecks that same SHA and the live fingerprint during canary, closing the check-to-stop evidence replacement window. A successful canary then promotes or retains the proof. For an agent Bash/tool call, set that call's `timeoutMs` to 180000; this is tool metadata, not a CLI flag, and keeps the caller alive longer than the 120-second preflight gate. The managed shell cannot reap the exit agent, so the scheduled kill lands after the scheduling turn ends. The watchdog respawns, runs the canary, and the new instance reports via `last-restart.json`; watchdog lifecycle lines are timestamped. With no live watchdog, `schedule-exit` hard-refuses: establish supervision first or use `restart`, which owns the complete single-shot loop. The latter lacks durable supervisor/launch ownership and therefore still requires a fresh credential. `--boot-timeout-ms MS` rides the restart marker as a one-boot readiness budget: the respawn's boot window honors it (rounded up to whole seconds) while the marker is pending, then the watchdog falls back to its own `WD_BOOT_TIMEOUT` (default 60 s).
|
|
115
115
|
|
|
116
116
|
### reconfigure: transactional launch changes
|
|
117
117
|
|
|
@@ -128,15 +128,20 @@ dsh-ankh-guard reconfigure \
|
|
|
128
128
|
--transition-file "<optional state-transition plan.json>" \
|
|
129
129
|
--on-failure restore-previous \
|
|
130
130
|
--browser-handoff required \
|
|
131
|
+
--boot-timeout-ms 90000 \
|
|
131
132
|
--state-dir "$DSH_HOME/state"
|
|
132
133
|
```
|
|
133
134
|
|
|
135
|
+
`reconfigure` and `supervise` accept `--boot-timeout-ms MS` to hand the spawned watchdog its per-boot readiness budget (`WD_BOOT_TIMEOUT`, whole seconds; omit it on a slow host and a healthy-but-slow first boot reads as a failure); on `reconfigure` the value rides the supervisor driver's argv to the successor watchdog.
|
|
136
|
+
|
|
134
137
|
The recovery choice is mandatory and therefore approved before the old host stops: `restore-previous` restores the entire previous launch specification, while `wait-for-user` parks for intervention without resetting any repository. The full previous/target pair separately persists command, home, credential repo, harness root, profile, and port. Each new spec also binds the composition preflight's `source|built` surface, runner executable/path/content SHA, the actual `@deepseek-ai/dsh/package.json` install anchor, and the target command SHA. A built successor resolves modules from its npm toolchain; the guard never chooses checkout source merely because its own process happens to use tsx. `reconfigure` additionally requires the caller to supply a one-shot `--candidate-probe-command` committed with the target command SHA. It runs that probe against an isolated home before composition preflight runs on the bound execution surface. The guard binds and executes both command digests; it cannot prove that an arbitrary successful shell command was automatically derived from the target argv. The bundled Skill constructs a DSH `--dump-config` probe from the same executable and launcher argv; integrations outside that caller trust boundary must establish the same semantic provenance themselves. Either failure occurs before previous stops. The mode-0600 `launch-spec.json` contains the commands and selected side; the receipt contains only bindings, hashes, and PASS outcomes, never probe/start commands or bearer URLs.
|
|
135
138
|
|
|
136
139
|
The atomic selected-side rename is the configuration commit point. If no complete durable previous spec exists, initialize it first with the real current values and explicit preflight surface/install anchor through `configure-launch`: legacy `instance-launch.json` cannot supply the missing roles, and a target `--repo` is never backfilled into previous. A replacement watchdog then atomically claims `watchdog.pid` while the old host is still serving. During preparation the guard captures PID/start identities for the old supervisor, its direct child, and its listener. The successor waits for the exact supervisor identity to yield with a 15-second default bound (configurable through `--supervisor-yield-timeout-ms`) and consumes abort/restore throughout that wait. On timeout it may retire only the frozen, revalidated old supervisor tree. It then stops only the proven child/listener identities and confirms port release; it never chooses or kills an arbitrary process found by port. PID reuse, a hung watchdog, a short-lived `reconfigure` caller, an outer launchd/systemd waiter, and nested shells therefore cannot blur ownership or suspend a cutover forever.
|
|
137
140
|
|
|
138
141
|
For a protected target, the watchdog accepts a launch URL only from the final process's output, only for the exact supervised loopback authority, and without depending on a parameter name. It proves 303 cookie exchange and authenticated root 200 with a temporary jar, then spends a three-second default stability window proving that the child is alive, its sole listener remains in that child tree, PID/start identities do not change, and retry is zero. The browser half uses held long polls on a same-origin plugin route rather than a permanent 500 ms loop. The proven previous listener stores only hashes of per-tab capabilities and puts every responsive registered tab into a waiting state. Only after ownership-stable service readiness and canary success does the final listener tell each tab to reload when its cookie is still accepted, or return that final process's same-origin one-time URL for an in-memory `location.replace()` after 401. The authenticated page acknowledges and returns to the original safe pathname with query and fragment discarded. One real ACK gates terminal ready; other registered, unacknowledged tabs remain eligible after state compaction. If no original tab registered or none acknowledges in time, the watchdog requests one system-open fallback and still waits for its authenticated page acknowledgement; opener exit 0 alone never completes handoff. Server readiness/canary and browser handoff are recorded separately, and all required evidence must complete before session wake-up. A naked 401, an old listener's 200, a target's transient 200, or a subsequent target exit can never become ready or hand a rejected target URL to the browser. No raw browser capability or bearer URL enters a state file, durable log, or receipt. `launch-cutover.json` records redacted configuration summaries, process identities, authentication, handoff, stability, and per-role failures. Target readiness/canary remains under `targetValidation`; previous recovery readiness and a passing, failing, or explicitly skipped recovery canary (when only a target-scoped credential exists) live under `recovery.validation`. A restored receipt can no longer carry an unqualified target canary failure beside previous readiness. `launch-status` prints the receipt without exposing either command. During a transaction, `abort-cutover --state-dir "$DSH_HOME/state"` applies the pre-approved recovery policy, while `restore-previous --state-dir "$DSH_HOME/state"` explicitly authorizes stopping the proven target and restoring the complete previous spec. The two actions use separate atomic markers and restore always wins on read, so even concurrent sessions cannot let a later ordinary abort downgrade an explicit restore.
|
|
139
142
|
|
|
143
|
+
**Mid-cutover recovery when the supervisor chain dies.** A receipt parked in `awaiting-user` no longer wedges the transaction. A bare `supervise` — the command launchd/systemd runs — now spawns the watchdog in a parked hold: it claims supervision, launches nothing (the rejected side never boots), and keeps consuming the operator control markers, so `abort-cutover --state-dir "$DSH_HOME/state"` (applies the pre-approved policy) and `restore-previous --state-dir "$DSH_HOME/state"` work immediately instead of refusing with "no live watchdog". The hold announces its own exits once: `touch "$DSH_HOME/state/watchdog-stop"` ends the hold without settling, and `supervise --cutover-id <id> --state-dir "$DSH_HOME/state"` (the id is in `launch-status` and the receipt) resumes the selected side — the resume flips the receipt out of `awaiting-user`, the hold releases its claim in response, and the resuming chain waits for that release before taking over.
|
|
144
|
+
|
|
140
145
|
When a candidate cannot read a reconstructible projection or cache left by the previous host, `--transition-file` can submit a reviewed schema-v1 quarantine plan. A plan accepts only non-overlapping, symlink-free paths below `home` that exclude guard state, with an explicit `quarantine` operation; it contains no host-version or filename knowledge. Example: `{"schemaVersion":1,"home":"/absolute/dsh-home","operations":[{"kind":"quarantine","path":"storages/<reconstructible-cache>","expect":"present"}]}`. Each `expect` is `present` or `absent`; isolated preflight and live apply must observe that same state or refuse before previous stops or target starts. The plan must cover both old paths that need to leave before target starts and new output paths that must leave before previous can recover after a target failure; list the latter explicitly with `expect: "absent"` even when they do not exist at preparation. Do not use this mechanism for authoritative logs, credentials, or irreplaceable data. Formats that require content transformation need a separate reversible migration tool and review.
|
|
141
146
|
|
|
142
147
|
The guard first copies the live home's boot inputs with copy-on-write preference — the allowlist is `profiles/`, the home-level `cordis.patch.yml`, `settings.yaml`, `.credentials.yaml`, and `.anonymous-user-id`; plugin data directories (sessions, state, local-agent sub-homes, tarballs, scratch, …) are excluded by default, so a newly installed plugin's data never silently grows the snapshot, and the list only moves when the host's boot starts reading a new home input (a miss fails the preflight loudly instead of slowing the copy down). It then rebuilds pnpm/Cordis links as relative links whose targets are snapshot-owned copies; legitimate internal dependency cycles remain intact. External targets enter a hash-named, deduplicated materialization area inside the snapshot. A linked `node_modules` target anchors at the package's own parent `node_modules` — the tightest scope that preserves Node's ancestor lookup — rather than the whole store root. A post-copy `realpath` audit requires every writable link target to remain under the snapshot root. Runtime entries without copyable semantics (sockets, FIFOs, and links to them) are skipped and counted; dangling or unresolvable links, any other special files, writable escapes, and read/copy failures make `reconfigure` fail closed before it runs the candidate, creates a cutover, or stops previous. It then applies the same quarantine and runs target composition preflight in that safe copy. If the copy cannot be prepared or the target does not boot, previous keeps serving and the live home stays unchanged. Only after the successor owns supervision and has stopped and revalidated the previous process tree does it apply the hash-bound durable plan with same-filesystem renames. If the target is rejected, the watchdog first stops its proven process, retains replacements it created at transitioned paths under `launch-transitions/<cutover>/rejected-target/`, restores the exact previous bytes and records the result, and only then permits previous to start. Any unproven step parks at `awaiting-user` instead of exposing previous to mixed state. After target success, the quarantined previous content remains in the cutover directory for operator disposition; it is never deleted automatically.
|
package/README.i18n.yaml
CHANGED
|
@@ -2,5 +2,5 @@
|
|
|
2
2
|
# last confirmed-consistent state. Both languages carry equal authority; after
|
|
3
3
|
# editing either side, bring the other along and re-record with:
|
|
4
4
|
# pnpm run verify-translation-pairing --write packages/ankh-guard/README.en.md
|
|
5
|
-
packages/ankh-guard/README.en.md:
|
|
6
|
-
packages/ankh-guard/README.md:
|
|
5
|
+
packages/ankh-guard/README.en.md: fb474172d57154e58e44661bbefb6824c78425e0
|
|
6
|
+
packages/ankh-guard/README.md: e58a1358dcfae6d426b3ced160120de361709e24
|
package/README.md
CHANGED
|
@@ -109,9 +109,9 @@ dsh-ankh-guard supervise --port 3080 --start "CMD" --state-dir "$DSH_HOME/state"
|
|
|
109
109
|
|
|
110
110
|
`supervise` 还需要被监管实例启动时使用的 dsh home(watchdog 会把它 export 为实例的 `DSH_HOME`):`--home DIR` 优先,否则取 `$DSH_HOME`;两者都没有时响亮拒绝——从 `--state-dir` 猜出来的 home 会让实例静默读错 profile/凭据目录。首次持久化还必须显式提供 `--harness-root` 或 `DSH_HARNESS`;它不会把 credential repo 猜成宿主根。
|
|
111
111
|
|
|
112
|
-
它以 `--wait-owner` 模式 detached 拉起随包发布的 `scripts/dsh-watchdog.sh`:watchdog 在当前实例运行期间待机,实例退出(有意重启或崩溃)后接管端口、重新拉起,有意重启时跑 guard canary(读 `restart-requested.json` 标记),通过后清除标记。连续 2 次起不来→回滚到最后已知可用版本:健康启动戳(`last-good-boot.json`,每次实例成功启动时重写,指向本部署里最近一次真正跑起来的版本)优先,其次是 guard checkpoint,最后是凭证 HEAD;但仅当启动失败的错误主体路径在仓库内。主体在仓库之外时(坏掉的 profile overlay 或已装插件),回滚检出修不好,watchdog 改为恢复上次健康的 **profile 组合**:健康启动时快照的组合输入(`last-good-composition/`)覆盖回 live 的 bundles 层与清单,最新插件变更被卸载,故障输入保留在 `composition-backup-*`,恢复报告会点名被卸载的内容。启动命令没有绑到被监督端口时同样豁免:启动窗口超时而实例正监听在别处、或以点名了本 watchdog 并不拥有的端口的 `EADDRINUSE` 失败时,watchdog 会点名实际绑定的端口并跳过两种回滚——重置文件改不了命令行参数。发生在被监督端口上的 `EADDRINUSE`
|
|
112
|
+
它以 `--wait-owner` 模式 detached 拉起随包发布的 `scripts/dsh-watchdog.sh`:watchdog 在当前实例运行期间待机,实例退出(有意重启或崩溃)后接管端口、重新拉起,有意重启时跑 guard canary(读 `restart-requested.json` 标记),通过后清除标记。连续 2 次起不来→回滚到最后已知可用版本:健康启动戳(`last-good-boot.json`,每次实例成功启动时重写,指向本部署里最近一次真正跑起来的版本)优先,其次是 guard checkpoint,最后是凭证 HEAD;但仅当启动失败的错误主体路径在仓库内。主体在仓库之外时(坏掉的 profile overlay 或已装插件),回滚检出修不好,watchdog 改为恢复上次健康的 **profile 组合**:健康启动时快照的组合输入(`last-good-composition/`)覆盖回 live 的 bundles 层与清单,最新插件变更被卸载,故障输入保留在 `composition-backup-*`,恢复报告会点名被卸载的内容。启动命令没有绑到被监督端口时同样豁免:启动窗口超时而实例正监听在别处、或以点名了本 watchdog 并不拥有的端口的 `EADDRINUSE` 失败时,watchdog 会点名实际绑定的端口并跳过两种回滚——重置文件改不了命令行参数。发生在被监督端口上的 `EADDRINUSE` 保留原本的释放并重试逃生口,现在以五次为上限。每次启动失败都会把该次 attempt 捕获的输出镜像进 watchdog 主日志——实例一行都没写时则显式记录 `attempt N produced no output`——让「boot 即死」在运维第一眼就能看到。任何路径的 reset(watchdog、CLI、service)都会先为被丢弃的 HEAD 和未提交改动创建 `guard-backup-*` 分支锚点,恢复不依赖 reflog。4 次失败→在端口上提供带重试按钮的崩溃页(SIGUSR1 通知 watchdog)。`watchdog-stop` 标记让 watchdog 彻底退出。实例可以在自我重启前自行采用监督——用户永远不需要手动启动 watchdog。
|
|
113
113
|
|
|
114
|
-
已有 watchdog 监督时,重启触发用 `schedule-exit`:它从耐久 active launch spec 取得端口、凭证仓库、宿主根与 profile,拒绝任何冲突的显式参数,并核对 supervisor 写下的完整命令后才写 restart 标记、spawn detached 退出代理(输出明确标为 `exit-agent pid`)。它优先使用 10 分钟内的新鲜 credential;credential 过期时,只允许精确匹配的 `provenDeployment` 走纯重启快速路径,并把所选证据 SHA 写入短寿命 restart marker。新 watchdog 在 canary 时再次核对同一 SHA 与现场指纹,防止检查后、停止前的证据替换。通过后才晋升或保留部署证明。从 agent 的 Bash/tool 调用时应把该调用的 `timeoutMs` 设为 180000;这不是 CLI 参数,而是保证调用方不会先于 120 秒 preflight 闸门退出的等待契约。托管 shell 的进程组回收不到退出代理,所以计划中的 kill 会在调度回合结束后真实落地。watchdog 重新拉起、跑 canary,新实例经 `last-restart.json` 回报;watchdog 生命周期日志均带时间戳。没有存活 watchdog 时 `schedule-exit` 硬拒绝,只能先建立监督或使用拥有单次完整循环的 `restart`;后者没有相同的 durable supervisor/launch ownership,因此仍要求新鲜 credential
|
|
114
|
+
已有 watchdog 监督时,重启触发用 `schedule-exit`:它从耐久 active launch spec 取得端口、凭证仓库、宿主根与 profile,拒绝任何冲突的显式参数,并核对 supervisor 写下的完整命令后才写 restart 标记、spawn detached 退出代理(输出明确标为 `exit-agent pid`)。它优先使用 10 分钟内的新鲜 credential;credential 过期时,只允许精确匹配的 `provenDeployment` 走纯重启快速路径,并把所选证据 SHA 写入短寿命 restart marker。新 watchdog 在 canary 时再次核对同一 SHA 与现场指纹,防止检查后、停止前的证据替换。通过后才晋升或保留部署证明。从 agent 的 Bash/tool 调用时应把该调用的 `timeoutMs` 设为 180000;这不是 CLI 参数,而是保证调用方不会先于 120 秒 preflight 闸门退出的等待契约。托管 shell 的进程组回收不到退出代理,所以计划中的 kill 会在调度回合结束后真实落地。watchdog 重新拉起、跑 canary,新实例经 `last-restart.json` 回报;watchdog 生命周期日志均带时间戳。没有存活 watchdog 时 `schedule-exit` 硬拒绝,只能先建立监督或使用拥有单次完整循环的 `restart`;后者没有相同的 durable supervisor/launch ownership,因此仍要求新鲜 credential。`--boot-timeout-ms MS` 随 restart marker 记录一次性就绪预算:marker 待消费期间,重新拉起的启动窗口按它计算(向上取整到整秒),之后 watchdog 回落到自己的 `WD_BOOT_TIMEOUT`(默认 60 秒)。
|
|
115
115
|
|
|
116
116
|
### reconfigure:启动配置事务切换
|
|
117
117
|
|
|
@@ -128,15 +128,20 @@ dsh-ankh-guard reconfigure \
|
|
|
128
128
|
--transition-file "<可选的状态迁移计划.json>" \
|
|
129
129
|
--on-failure restore-previous \
|
|
130
130
|
--browser-handoff required \
|
|
131
|
+
--boot-timeout-ms 90000 \
|
|
131
132
|
--state-dir "$DSH_HOME/state"
|
|
132
133
|
```
|
|
133
134
|
|
|
135
|
+
`reconfigure` 与 `supervise` 接受 `--boot-timeout-ms MS`,把每次启动的就绪预算交给拉起的 watchdog(`WD_BOOT_TIMEOUT`,整秒;慢宿主上省略它,一个健康但缓慢的首次 boot 会被误判为失败);`reconfigure` 下该值随 supervisor 驱动器的 argv 传给继任 watchdog。
|
|
136
|
+
|
|
134
137
|
恢复选择是必填项,因而会在旧宿主停止前获批:`restore-previous` 恢复上一份**完整**启动配置;`wait-for-user` 不重置任何仓库,原地停留等用户处理。完整 previous/target 对分别保存 command、home、credential repo、harness root、profile 与 port。每份新配置还固化 composition preflight 的 `source|built` 面、runner 可执行方式/路径/内容 SHA、实际 `@deepseek-ai/dsh/package.json` 安装锚点和 target command SHA;built successor 从其 npm toolchain 解析模块,绝不因 guard 自己恰由 tsx 启动就误选 checkout source。`reconfigure` 另要求调用方提供一条与 target command SHA 同时提交的一次性 `--candidate-probe-command`,先在隔离 home 执行,再跑同一执行面上的 composition preflight;Guard 只绑定并执行两个 command 摘要,不能证明任意成功的 shell command 是从 target argv 自动派生的。随包 Skill 会用相同 executable 与 launcher argv 构造 DSH `--dump-config` probe;这条调用方信任边界之外的其他集成也必须自行保证语义同源。任一失败都发生在 previous 停止之前。当前选中侧保存在 mode-0600 的 `launch-spec.json`;回执只写各摘要和 PASS 结果,不写 probe/start 命令或 bearer URL。
|
|
135
138
|
|
|
136
139
|
原子切换选中侧是配置提交点。缺少完整耐久 previous 时必须先用当前真实值运行 `configure-launch`,并显式给出它的 preflight surface/install anchor;旧 `instance-launch.json` 不足以推断,目标 `--repo` 也绝不会倒填 previous。随后替代 watchdog 会在旧宿主仍对外服务时原子取得 `watchdog.pid`。准备事务时 guard 同时固化旧 supervisor、直接 child 与 listener 的 PID/启动 identity;successor 按 identity 有界等待旧 supervisor 正常让权(默认 15 秒,可用 `--supervisor-yield-timeout-ms` 调整),等待期间持续消费 abort/restore。超时只会在先冻结并复核旧 supervisor identity 后终止其精确进程树。successor 随后只停止已证明的旧 child/listener 并确认端口释放,绝不凭端口反查后杀任意监听者。因此 PID 复用、卡死 watchdog、短命 `reconfigure` 调用者、外层 launchd/systemd 等待者或嵌套 shell 都不会模糊所有权或无限悬挂切换。
|
|
137
140
|
|
|
138
141
|
目标受保护时,watchdog 只接受最终进程输出、且 authority 与被监督 loopback 完全一致的启动 URL,不依赖任何查询参数名。它用临时 jar 证明 303 Cookie 交换和认证后根路径 200,再在默认 3 秒稳定窗口内持续证明 child 存活、唯一 listener 属于该 child 树、PID/启动 identity 不变且 retry 为 0。浏览器端改用插件同源路由上的持有式长轮询,不再永久每 500ms 请求:已证明的 previous listener 只落盘每标签页随机 capability 的哈希,并让所有仍响应的已登记标签页进入等待。只有在 ownership-stable 服务就绪和 canary 成功后,final listener 才在 Cookie 有效时通知各页刷新,或在收到 401 时把该最终进程的同源一次性 URL 返回内存,由页面执行 `location.replace()`;已认证页面回传 ACK 后,会丢弃 query/fragment 并回到原来的安全 pathname。一个真实 ACK 解除 terminal ready 门禁,其他尚未 ACK 的已登记页面在状态压缩后仍可恢复。没有原标签页登记或都未在超时内确认时,watchdog 才请求一次 system open 兜底,并继续等待新页面回传已认证 ACK;opener 的 exit 0 本身永远不算交接成功。服务端 readiness/canary 与浏览器交接分别记录,所需证据全部完成后才释放会话唤醒。裸 401、旧 listener 的 200、target 的短暂 200 或随后退出都不会成为 ready,也不会把已拒绝 target 的 URL 交给浏览器;原始 capability 和 bearer URL 均不进入状态文件、耐久日志或回执。`launch-cutover.json` 记录脱敏配置摘要、新旧 supervisor/child/listener identity、旧 supervisor 让权、认证就绪、浏览器 ACK 渠道、稳定性证明与分角色失败计数;target 的 readiness/canary 保留在 `targetValidation`,previous 的恢复 readiness 以及可用、失败或因只有 target-scoped credential 而明确跳过的恢复 canary 单独写在 `recovery.validation`,不会再出现 previous 已恢复却挂着一个无主语的 target canary fail。`launch-status` 输出该回执且不暴露两边命令。事务进行中可运行 `abort-cutover --state-dir "$DSH_HOME/state"` 执行事前批准的恢复策略;`restore-previous --state-dir "$DSH_HOME/state"` 显式授权停止已证明的 target 并恢复完整 previous spec。两个动作分别使用原子 marker,读取时 restore 永远优先,因此并发会话的晚到 abort 也不能降级 restore。
|
|
139
142
|
|
|
143
|
+
**监督链中途死亡时的 cutover 恢复。** 回执停在 `awaiting-user` 不再楔死事务。裸 `supervise`(launchd/systemd 跑的那条命令)现在以停泊形态拉起 watchdog:它取得监督所有权、什么都不启动(被拒的一侧绝不 boot),并持续消费运维控制 marker,因此 `abort-cutover --state-dir "$DSH_HOME/state"`(执行事前批准的策略)与 `restore-previous --state-dir "$DSH_HOME/state"` 立即可用,不再以「没有存活 watchdog」拒绝。停泊态会一次性说明自己的退出与恢复路径:`touch "$DSH_HOME/state/watchdog-stop"` 不做终态结算就结束停泊;`supervise --cutover-id <id> --state-dir "$DSH_HOME/state"`(id 见 `launch-status` 与回执)恢复选中侧——resume 把回执翻离 `awaiting-user`,停泊态随即释放所有权,恢复链先等到这次释放再接管。
|
|
144
|
+
|
|
140
145
|
candidate 无法读取旧宿主留下的可重建投影或缓存时,`--transition-file` 可以提交一份经过评审的 schema-v1 隔离计划。计划只接受 `home` 下互不重叠、没有符号链接且不包含 guard state 的相对路径,以及显式的 `quarantine` 操作;它不内置任何宿主版本或文件名知识。示例:`{"schemaVersion":1,"home":"/absolute/dsh-home","operations":[{"kind":"quarantine","path":"storages/<可重建缓存>","expect":"present"}]}`。每项 `expect` 必须是 `present` 或 `absent`,副本 preflight 与 live apply 都必须观察到相同状态,否则在停 previous 前或启动 target 前拒绝。计划既要覆盖 target 启动前必须移开的旧路径,也要覆盖 target 失败后 previous 启动前必须清走的新输出路径;后者即使准备时不存在也必须用 `expect: "absent"` 显式列出。权威日志、凭据或不可重建数据不得借此移出;需要内容转换的格式应使用独立、可逆且另行评审的迁移工具。
|
|
141
146
|
|
|
142
147
|
guard 先以 copy-on-write 优先方式复制 live home 的启动输入——白名单为 `profiles/`、home 级 `cordis.patch.yml`、`settings.yaml`、`.credentials.yaml`、`.anonymous-user-id`;插件数据目录(sessions、state、local-agent 子 home、tarballs、scratch 等)默认不进复制,新插件的数据目录不会悄悄扩大快照,名单只在宿主启动读取范围变化时才需要更新(漏项会以预检 FAIL 显形,而不是慢慢变大)。再把 pnpm/Cordis 链接重建为只指向 snapshot 内副本的相对链接;合法的内部依赖循环保留。外部目标进入 snapshot 自己的哈希命名去重物化区,`node_modules` 目标锚定在包自身的父 `node_modules`(保留 Node 祖先查找语义的最小范围),而不是整个 store 根。复制后逐链接 `realpath` 审计,任何可写目标都必须仍在 snapshot 根内;无可复制语义的运行时条目(socket、FIFO 及指向它们的链接)跳过并计数;悬空或不可解析链接、其余特殊文件、可写逃逸及读取/复制失败都会让 `reconfigure` 在运行 candidate、创建 cutover 或停止 previous 前 fail closed。在安全副本中执行相同隔离后才运行 target composition preflight;副本无法准备或 target 无法 boot 时,previous 继续运行且 live home 不变。successor 取得监督所有权、停止并复核 previous 进程树以后,才按照哈希绑定的耐久计划用同文件系统 rename 隔离原路径。target 被拒绝时,watchdog 必须先停止其已证明的进程,再把它在同路径产生的替代内容保留到 `launch-transitions/<cutover>/rejected-target/`,恢复 previous 原字节并写入回执,最后才允许 previous 启动;任一步无法证明完成都会停在 `awaiting-user`,不会让旧宿主读取混合状态。target 成功后,旧内容仍保存在 cutover 目录,等待 operator 后续处置,不会自动删除。
|
package/lib/cli.js
CHANGED
|
@@ -1273,12 +1273,14 @@ commands:
|
|
|
1273
1273
|
[--home DIR] [--repo DIR] [--harness-root DIR] [--profile NAME] [--browser-handoff required|off]
|
|
1274
1274
|
--preflight-surface source|built [--preflight-runner FILE] --preflight-install-anchor FILE
|
|
1275
1275
|
--candidate-probe-command "CMD"
|
|
1276
|
-
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
1276
|
+
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
1277
|
+
[--boot-timeout-ms MS] [--state-dir DIR]
|
|
1277
1278
|
restart --port N --start "CMD" [--pid PID] [--timeout-ms MS] [--delay-ms MS] [--stop-timeout-ms MS] [--rollback]
|
|
1278
1279
|
[--profile NAME] [--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR] [--max-age MIN]
|
|
1279
1280
|
schedule-exit [--port N] --delay-ms MS [--initiator ID] [--log FILE] [--profile NAME]
|
|
1280
|
-
[--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
1281
|
+
[--harness-root DIR] [--preflight-timeout-ms MS] [--boot-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
1281
1282
|
supervise --port N --start "CMD" [--foreground] [--log FILE] [--state-dir DIR] [--repo DIR] [--harness-root DIR] [--home DIR]
|
|
1283
|
+
[--cutover-id ID] [--boot-timeout-ms MS]
|
|
1282
1284
|
flags:
|
|
1283
1285
|
--state-dir DIR state directory (default: $DSH_HOME/state, else <cwd>/.dsh-guard-state)
|
|
1284
1286
|
--repo DIR repository the credential binds to (default: cwd)
|
|
@@ -1305,6 +1307,16 @@ flags:
|
|
|
1305
1307
|
(agent-driven graceful self-restart: schedule, complete, then restart);
|
|
1306
1308
|
schedule-exit: delay before the detached exit agent kills the host;
|
|
1307
1309
|
reconfigure: grace after successor supervisor claim before old-child stop
|
|
1310
|
+
--boot-timeout-ms MS reconfigure/schedule-exit/supervise: readiness budget for one
|
|
1311
|
+
boot (the watchdog's WD_BOOT_TIMEOUT, whole seconds, default 60).
|
|
1312
|
+
reconfigure/supervise hand it to the spawned watchdog's environment;
|
|
1313
|
+
schedule-exit records it in the restart marker for the respawn's boot
|
|
1314
|
+
window only (the running watchdog then falls back to its own budget)
|
|
1315
|
+
--cutover-id ID supervise: resume the durable launch-cutover transaction after its
|
|
1316
|
+
supervisor chain died (the id is in launch-status / the receipt);
|
|
1317
|
+
without it a bare supervise on an awaiting-user receipt only holds
|
|
1318
|
+
the claim and consumes operator control markers, never booting the
|
|
1319
|
+
rejected side
|
|
1308
1320
|
--log FILE supervise (detached only — with --foreground the external supervisor's
|
|
1309
1321
|
redirection owns the log) / schedule-exit: log file (default: <state-dir>/*.log)
|
|
1310
1322
|
--home DIR supervise: the dsh home the supervised instance boots with (profiles,
|
|
@@ -1361,6 +1373,7 @@ function parse(argv) {
|
|
|
1361
1373
|
delayMs: void 0,
|
|
1362
1374
|
stopTimeoutMs: void 0,
|
|
1363
1375
|
supervisorYieldTimeoutMs: void 0,
|
|
1376
|
+
bootTimeoutMs: void 0,
|
|
1364
1377
|
log: void 0,
|
|
1365
1378
|
foreground: false,
|
|
1366
1379
|
rollback: false,
|
|
@@ -1494,6 +1507,14 @@ function parse(argv) {
|
|
|
1494
1507
|
i++;
|
|
1495
1508
|
break;
|
|
1496
1509
|
}
|
|
1510
|
+
case "--boot-timeout-ms": {
|
|
1511
|
+
const raw = flagValue(arg, true);
|
|
1512
|
+
const n = Number(raw);
|
|
1513
|
+
if (raw === void 0 || !Number.isInteger(n) || n < 1e3) throw new Error("--boot-timeout-ms must be an integer >= 1000");
|
|
1514
|
+
options.bootTimeoutMs = n;
|
|
1515
|
+
i++;
|
|
1516
|
+
break;
|
|
1517
|
+
}
|
|
1497
1518
|
case "--foreground":
|
|
1498
1519
|
options.foreground = true;
|
|
1499
1520
|
break;
|
|
@@ -1683,9 +1704,14 @@ function verifyRepoCredential(stateDir, repoDir, maxAgeMinutes) {
|
|
|
1683
1704
|
return verifyCredential(loadState(stateDir), currentHead(repoDir), Date.now(), maxAgeMinutes, isWorkingTreeClean(repoDir));
|
|
1684
1705
|
}
|
|
1685
1706
|
/**
|
|
1686
|
-
* How the spawned watchdog should invoke the guard CLI
|
|
1687
|
-
*
|
|
1688
|
-
*
|
|
1707
|
+
* How the spawned watchdog should invoke the guard CLI. The executable is
|
|
1708
|
+
* always the ABSOLUTE process.execPath, never a bare `node`: the watchdog runs
|
|
1709
|
+
* under whatever PATH spawned it, and a launcher chain (launchd/systemd) has a
|
|
1710
|
+
* minimal PATH without homebrew — a bare `node` there makes every guard
|
|
1711
|
+
* invocation the watchdog issues (canary, record-proven-deployment, cutover
|
|
1712
|
+
* events) fail with command-not-found (the 2026-09-30 launchd incident). The
|
|
1713
|
+
* source form additionally needs tsx with an absolute path (the watchdog runs
|
|
1714
|
+
* with a deployment cwd that resolves no node_modules).
|
|
1689
1715
|
* @returns the command prefix (verb args are appended by the watchdog).
|
|
1690
1716
|
*/
|
|
1691
1717
|
function guardInvocation() {
|
|
@@ -1693,10 +1719,10 @@ function guardInvocation() {
|
|
|
1693
1719
|
if (cliPath.includes(`${sep}src${sep}`)) {
|
|
1694
1720
|
const nodeModules = resolve(dirname(cliPath), "../../../node_modules");
|
|
1695
1721
|
const tsx = join(nodeModules, "tsx", "dist", "esm", "index.mjs");
|
|
1696
|
-
if (existsSync(tsx)) return
|
|
1697
|
-
return
|
|
1722
|
+
if (existsSync(tsx)) return `${process.execPath} --import ${tsx} ${cliPath}`;
|
|
1723
|
+
return `${process.execPath} ${cliPath}`;
|
|
1698
1724
|
}
|
|
1699
|
-
return
|
|
1725
|
+
return `${process.execPath} ${cliPath}`;
|
|
1700
1726
|
}
|
|
1701
1727
|
/**
|
|
1702
1728
|
* argv (after process.execPath) that runs the exit agent, with the same
|
|
@@ -2552,7 +2578,7 @@ async function runCli(argv, io) {
|
|
|
2552
2578
|
}
|
|
2553
2579
|
const watchdogPid = liveWatchdogPid(stateDir);
|
|
2554
2580
|
if (watchdogPid === null) {
|
|
2555
|
-
io.stderr(`${command} refused: no live watchdog can consume the durable control request\n`);
|
|
2581
|
+
io.stderr(`${command} refused: no live watchdog can consume the durable control request. Start the consumer first: \`supervise --state-dir ${stateDir}\` holds the awaiting-user cutover ${transaction.receipt.id} without launching anything, or \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\` resumes the selected side directly\n`);
|
|
2556
2582
|
return 1;
|
|
2557
2583
|
}
|
|
2558
2584
|
const requested = command === "restore-previous" ? "restore-previous" : "abort";
|
|
@@ -2892,6 +2918,7 @@ async function runCli(argv, io) {
|
|
|
2892
2918
|
String(options.delayMs ?? 5e3),
|
|
2893
2919
|
"--supervisor-yield-timeout-ms",
|
|
2894
2920
|
String(options.supervisorYieldTimeoutMs ?? 15e3),
|
|
2921
|
+
...options.bootTimeoutMs !== void 0 ? ["--boot-timeout-ms", String(options.bootTimeoutMs)] : [],
|
|
2895
2922
|
...initiator !== void 0 ? ["--initiator", initiator] : []
|
|
2896
2923
|
];
|
|
2897
2924
|
const cutoverDriverEnv = { ...process.env };
|
|
@@ -3114,10 +3141,9 @@ async function runCli(argv, io) {
|
|
|
3114
3141
|
io.stderr(`supervise refused: cutover ${options.cutoverId} is not the selected launch transaction\n`);
|
|
3115
3142
|
return 1;
|
|
3116
3143
|
}
|
|
3117
|
-
|
|
3118
|
-
|
|
3119
|
-
|
|
3120
|
-
}
|
|
3144
|
+
let parkedCutover = false;
|
|
3145
|
+
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user") parkedCutover = true;
|
|
3146
|
+
const resumeFromAwaitingUser = options.cutoverId !== void 0 && transaction?.receipt.phase === "awaiting-user";
|
|
3121
3147
|
if (options.foreground !== true && !sandboxGate("supervise", options, io)) return 2;
|
|
3122
3148
|
if (options.cutoverId !== void 0) try {
|
|
3123
3149
|
const driverIdentity = processIdentity(process.pid);
|
|
@@ -3152,7 +3178,22 @@ async function runCli(argv, io) {
|
|
|
3152
3178
|
return 1;
|
|
3153
3179
|
}
|
|
3154
3180
|
if (durable === null) writeStableLaunchSpec(stateDir, spec);
|
|
3155
|
-
if (
|
|
3181
|
+
if (resumeFromAwaitingUser) {
|
|
3182
|
+
const existingIdentity = processIdentity(existingPid);
|
|
3183
|
+
if (existingIdentity === null) {
|
|
3184
|
+
io.stderr(`supervise refused: could not capture watchdog ${existingPid} start identity before waiting\n`);
|
|
3185
|
+
return 1;
|
|
3186
|
+
}
|
|
3187
|
+
io.stdout(`watchdog ${existing} holds the parked cutover — waiting for it to release, then resuming\n`);
|
|
3188
|
+
const releaseDeadline = Date.now() + 15e3;
|
|
3189
|
+
while (processIdentityMatches(existingIdentity) && Date.now() < releaseDeadline) await sleep(250);
|
|
3190
|
+
if (processIdentityMatches(existingIdentity)) {
|
|
3191
|
+
io.stderr(`supervise refused: watchdog ${existingPid} did not release the parked cutover ${options.cutoverId ?? ""} within 15000 ms — it is not the awaiting-user hold; settle the transaction with \`abort-cutover --state-dir ${stateDir}\` or \`restore-previous --state-dir ${stateDir}\`, or stop that watchdog and retry\n`);
|
|
3192
|
+
return 1;
|
|
3193
|
+
}
|
|
3194
|
+
waitedForWatchdog = true;
|
|
3195
|
+
io.stdout(`watchdog ${existing} released the parked cutover — resuming\n`);
|
|
3196
|
+
} else if (options.foreground) {
|
|
3156
3197
|
io.stdout(`watchdog ${existing} already supervises the port — waiting for it to exit, then taking over (foreground)\n`);
|
|
3157
3198
|
const existingIdentity = processIdentity(existingPid);
|
|
3158
3199
|
if (existingIdentity === null) {
|
|
@@ -3179,10 +3220,7 @@ async function runCli(argv, io) {
|
|
|
3179
3220
|
io.stderr(`supervise refused after wait: cutover ${options.cutoverId} is no longer the selected launch transaction\n`);
|
|
3180
3221
|
return 1;
|
|
3181
3222
|
}
|
|
3182
|
-
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user")
|
|
3183
|
-
io.stderr(`supervise: cutover ${transaction.receipt.id} settled awaiting-user while this supervisor waited; refusing to restart the rejected target (receipt ${stateFile(stateDir, "launchCutover")})\n`);
|
|
3184
|
-
return 0;
|
|
3185
|
-
}
|
|
3223
|
+
if (options.cutoverId === void 0 && transaction?.receipt.phase === "awaiting-user") parkedCutover = true;
|
|
3186
3224
|
io.stdout("launch state refreshed after wait — using the durable selected specification\n");
|
|
3187
3225
|
}
|
|
3188
3226
|
const previousOwnership = transaction?.receipt.ownership?.previous;
|
|
@@ -3231,6 +3269,7 @@ async function runCli(argv, io) {
|
|
|
3231
3269
|
WD_PROFILE: spec.profile,
|
|
3232
3270
|
WD_WAIT_OWNER: options.takeoverFrom !== void 0 || !options.foreground ? "1" : "0",
|
|
3233
3271
|
WD_GUARD: guardInvocation(),
|
|
3272
|
+
...options.bootTimeoutMs !== void 0 ? { WD_BOOT_TIMEOUT: String(Math.ceil(options.bootTimeoutMs / 1e3)) } : {},
|
|
3234
3273
|
...options.takeoverFrom !== void 0 ? {
|
|
3235
3274
|
WD_TAKEOVER_FROM: String(options.takeoverFrom),
|
|
3236
3275
|
WD_TAKEOVER_FROM_START: previousSupervisorStart ?? ""
|
|
@@ -3251,9 +3290,11 @@ async function runCli(argv, io) {
|
|
|
3251
3290
|
WD_PREVIOUS_CHILD_START: previousOwnership.childStartToken,
|
|
3252
3291
|
WD_PREVIOUS_LISTENER_PID: String(previousOwnership.listenerPid),
|
|
3253
3292
|
WD_PREVIOUS_LISTENER_START: previousOwnership.listenerStartToken,
|
|
3293
|
+
...parkedCutover ? { WD_CUTOVER_PARKED: "1" } : {},
|
|
3254
3294
|
...transaction.state.transition === void 0 ? {} : { WD_TRANSITION_PLAN_SHA256: transaction.state.transition.planSha256 }
|
|
3255
3295
|
} : {}
|
|
3256
3296
|
};
|
|
3297
|
+
if (parkedCutover && transaction !== null) io.stdout(`supervise: cutover ${transaction.receipt.id} is waiting for user action — holding the supervision claim WITHOUT launching the rejected ${transaction.state.selected} side (receipt ${stateFile(stateDir, "launchCutover")}). The operator verbs now have a live consumer: \`abort-cutover --state-dir ${stateDir}\` applies the pre-approved ${transaction.receipt.recovery.policy} policy, \`restore-previous --state-dir ${stateDir}\` restores the complete previous spec. To end the hold without settling: write the stop marker (\`touch ${stateFile(stateDir, "watchdogStop")}\`). To resume the rejected side: \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\`\n`);
|
|
3257
3298
|
if (options.foreground) {
|
|
3258
3299
|
const child = spawn("bash", [watchdog, "--supervise"], {
|
|
3259
3300
|
stdio: "inherit",
|
|
@@ -3303,7 +3344,7 @@ async function runCli(argv, io) {
|
|
|
3303
3344
|
const claimDeadline = Date.now() + 5e3;
|
|
3304
3345
|
while (Date.now() < claimDeadline && spawnError === void 0) {
|
|
3305
3346
|
if (liveWatchdogPid(stateDir) === spawnedPid) {
|
|
3306
|
-
io.stdout(`watchdog spawned and ready (pid ${spawnedPid}) — supervises :${spec.port}, log ${logPath}\n`);
|
|
3347
|
+
io.stdout(parkedCutover && transaction !== null ? `watchdog spawned and parked on cutover ${transaction.receipt.id} (pid ${spawnedPid}) — nothing launched; consuming abort-cutover/restore-previous markers; log ${logPath}\n` : `watchdog spawned and ready (pid ${spawnedPid}) — supervises :${spec.port}, log ${logPath}\n`);
|
|
3307
3348
|
return 0;
|
|
3308
3349
|
}
|
|
3309
3350
|
try {
|
|
@@ -3374,6 +3415,7 @@ async function runCli(argv, io) {
|
|
|
3374
3415
|
writeFileSync(stateFile(stateDir, "restartRequested"), `${JSON.stringify({
|
|
3375
3416
|
reason: "scheduled self-restart",
|
|
3376
3417
|
requestedAt: Date.now(),
|
|
3418
|
+
...options.bootTimeoutMs === void 0 ? {} : { bootTimeoutMs: options.bootTimeoutMs },
|
|
3377
3419
|
...gate.authorization === void 0 ? {} : { authorization: gate.authorization },
|
|
3378
3420
|
...initiator !== void 0 ? { initiator } : {}
|
|
3379
3421
|
})}\n`);
|
|
@@ -3427,4 +3469,4 @@ if (isDirectInvocation(import.meta.url)) {
|
|
|
3427
3469
|
});
|
|
3428
3470
|
}
|
|
3429
3471
|
//#endregion
|
|
3430
|
-
export { envInternals, parse, preflightInternals, resolveHarnessRoot, resolvePreflightBin, resolveRunnerCommand, resolveWdHome, runCli, runPreflightCheck };
|
|
3472
|
+
export { envInternals, guardInvocation, parse, preflightInternals, resolveHarnessRoot, resolvePreflightBin, resolveRunnerCommand, resolveWdHome, runCli, runPreflightCheck };
|
package/lib/client.js
CHANGED
|
@@ -159,6 +159,8 @@ window.__ModuleLoader__.load({
|
|
|
159
159
|
let timer;
|
|
160
160
|
let pending = consumeFallbackFragment() ?? readPending();
|
|
161
161
|
let cap;
|
|
162
|
+
/** Consecutive failed ACKs for the current pending handoff. */
|
|
163
|
+
let ackFailures = 0;
|
|
162
164
|
let errorRetryMs = ERROR_RETRY_MIN_MS;
|
|
163
165
|
let firstFailureAt;
|
|
164
166
|
let request;
|
|
@@ -226,6 +228,22 @@ window.__ModuleLoader__.load({
|
|
|
226
228
|
location.replace(returnPath);
|
|
227
229
|
return;
|
|
228
230
|
}
|
|
231
|
+
} else {
|
|
232
|
+
ackFailures += 1;
|
|
233
|
+
if (ackFailures >= 8) {
|
|
234
|
+
ackFailures = 0;
|
|
235
|
+
const probe = await post({
|
|
236
|
+
version: 1,
|
|
237
|
+
operation: "poll",
|
|
238
|
+
capability: cap ??= capability()
|
|
239
|
+
}, current.signal);
|
|
240
|
+
if (stale()) return;
|
|
241
|
+
if (probe.ok && (await probe.json()).state === "idle") {
|
|
242
|
+
sessionStorage.removeItem(PENDING_KEY);
|
|
243
|
+
sessionStorage.removeItem(CAPABILITY_KEY);
|
|
244
|
+
pending = null;
|
|
245
|
+
}
|
|
246
|
+
}
|
|
229
247
|
}
|
|
230
248
|
errorRetryMs = ERROR_RETRY_MIN_MS;
|
|
231
249
|
firstFailureAt = void 0;
|
|
@@ -278,6 +296,7 @@ window.__ModuleLoader__.load({
|
|
|
278
296
|
};
|
|
279
297
|
pending = nextPending;
|
|
280
298
|
writePending(nextPending);
|
|
299
|
+
ackFailures = 0;
|
|
281
300
|
waitingOverlay();
|
|
282
301
|
location.reload();
|
|
283
302
|
return;
|
|
@@ -299,6 +318,7 @@ window.__ModuleLoader__.load({
|
|
|
299
318
|
};
|
|
300
319
|
pending = nextPending;
|
|
301
320
|
writePending(nextPending);
|
|
321
|
+
ackFailures = 0;
|
|
302
322
|
waitingOverlay();
|
|
303
323
|
location.replace(launch.href);
|
|
304
324
|
return;
|
package/lib/types/cli.d.ts
CHANGED
|
@@ -19,6 +19,7 @@ interface CliOptions {
|
|
|
19
19
|
delayMs: number | undefined;
|
|
20
20
|
stopTimeoutMs: number | undefined;
|
|
21
21
|
supervisorYieldTimeoutMs: number | undefined;
|
|
22
|
+
bootTimeoutMs: number | undefined;
|
|
22
23
|
log: string | undefined;
|
|
23
24
|
foreground: boolean;
|
|
24
25
|
rollback: boolean;
|
|
@@ -61,6 +62,18 @@ export declare function parse(argv: readonly string[]): {
|
|
|
61
62
|
positionals: readonly string[];
|
|
62
63
|
options: CliOptions;
|
|
63
64
|
};
|
|
65
|
+
/**
|
|
66
|
+
* How the spawned watchdog should invoke the guard CLI. The executable is
|
|
67
|
+
* always the ABSOLUTE process.execPath, never a bare `node`: the watchdog runs
|
|
68
|
+
* under whatever PATH spawned it, and a launcher chain (launchd/systemd) has a
|
|
69
|
+
* minimal PATH without homebrew — a bare `node` there makes every guard
|
|
70
|
+
* invocation the watchdog issues (canary, record-proven-deployment, cutover
|
|
71
|
+
* events) fail with command-not-found (the 2026-09-30 launchd incident). The
|
|
72
|
+
* source form additionally needs tsx with an absolute path (the watchdog runs
|
|
73
|
+
* with a deployment cwd that resolves no node_modules).
|
|
74
|
+
* @returns the command prefix (verb args are appended by the watchdog).
|
|
75
|
+
*/
|
|
76
|
+
export declare function guardInvocation(): string;
|
|
64
77
|
/**
|
|
65
78
|
* How the guard invokes the dsh app's `preflight` mode: the source form runs
|
|
66
79
|
* `node --import <repo>/node_modules/tsx/dist/esm/index.mjs <repo>/apps/cli/src/bin.ts`,
|
package/lib/types/cli.js
CHANGED
|
@@ -220,12 +220,14 @@ commands:
|
|
|
220
220
|
[--home DIR] [--repo DIR] [--harness-root DIR] [--profile NAME] [--browser-handoff required|off]
|
|
221
221
|
--preflight-surface source|built [--preflight-runner FILE] --preflight-install-anchor FILE
|
|
222
222
|
--candidate-probe-command "CMD"
|
|
223
|
-
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
223
|
+
[--transition-file FILE] [--delay-ms MS] [--supervisor-yield-timeout-ms MS] [--preflight-timeout-ms MS]
|
|
224
|
+
[--boot-timeout-ms MS] [--state-dir DIR]
|
|
224
225
|
restart --port N --start "CMD" [--pid PID] [--timeout-ms MS] [--delay-ms MS] [--stop-timeout-ms MS] [--rollback]
|
|
225
226
|
[--profile NAME] [--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR] [--max-age MIN]
|
|
226
227
|
schedule-exit [--port N] --delay-ms MS [--initiator ID] [--log FILE] [--profile NAME]
|
|
227
|
-
[--harness-root DIR] [--preflight-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
228
|
+
[--harness-root DIR] [--preflight-timeout-ms MS] [--boot-timeout-ms MS] [--state-dir DIR] [--repo DIR]
|
|
228
229
|
supervise --port N --start "CMD" [--foreground] [--log FILE] [--state-dir DIR] [--repo DIR] [--harness-root DIR] [--home DIR]
|
|
230
|
+
[--cutover-id ID] [--boot-timeout-ms MS]
|
|
229
231
|
flags:
|
|
230
232
|
--state-dir DIR state directory (default: $DSH_HOME/state, else <cwd>/.dsh-guard-state)
|
|
231
233
|
--repo DIR repository the credential binds to (default: cwd)
|
|
@@ -252,6 +254,16 @@ flags:
|
|
|
252
254
|
(agent-driven graceful self-restart: schedule, complete, then restart);
|
|
253
255
|
schedule-exit: delay before the detached exit agent kills the host;
|
|
254
256
|
reconfigure: grace after successor supervisor claim before old-child stop
|
|
257
|
+
--boot-timeout-ms MS reconfigure/schedule-exit/supervise: readiness budget for one
|
|
258
|
+
boot (the watchdog's WD_BOOT_TIMEOUT, whole seconds, default 60).
|
|
259
|
+
reconfigure/supervise hand it to the spawned watchdog's environment;
|
|
260
|
+
schedule-exit records it in the restart marker for the respawn's boot
|
|
261
|
+
window only (the running watchdog then falls back to its own budget)
|
|
262
|
+
--cutover-id ID supervise: resume the durable launch-cutover transaction after its
|
|
263
|
+
supervisor chain died (the id is in launch-status / the receipt);
|
|
264
|
+
without it a bare supervise on an awaiting-user receipt only holds
|
|
265
|
+
the claim and consumes operator control markers, never booting the
|
|
266
|
+
rejected side
|
|
255
267
|
--log FILE supervise (detached only — with --foreground the external supervisor's
|
|
256
268
|
redirection owns the log) / schedule-exit: log file (default: <state-dir>/*.log)
|
|
257
269
|
--home DIR supervise: the dsh home the supervised instance boots with (profiles,
|
|
@@ -293,6 +305,7 @@ export function parse(argv) {
|
|
|
293
305
|
const options = {
|
|
294
306
|
stateDir: '', repoDir: '', harnessRoot: '', home: '', maxAgeMinutes: 10, port: undefined, command: undefined, run: false, runArgv: undefined, message: undefined, detail: undefined,
|
|
295
307
|
start: undefined, pid: undefined, timeoutMs: undefined, delayMs: undefined, stopTimeoutMs: undefined, supervisorYieldTimeoutMs: undefined,
|
|
308
|
+
bootTimeoutMs: undefined,
|
|
296
309
|
log: undefined,
|
|
297
310
|
foreground: false, rollback: false, force: false, sync: false, initiator: undefined, profile: undefined, preflightTimeoutMs: undefined,
|
|
298
311
|
preflightSurface: undefined, preflightRunner: undefined, preflightInstallAnchor: undefined, candidateProbeCommand: undefined,
|
|
@@ -418,6 +431,17 @@ export function parse(argv) {
|
|
|
418
431
|
i++;
|
|
419
432
|
break;
|
|
420
433
|
}
|
|
434
|
+
case '--boot-timeout-ms': {
|
|
435
|
+
const raw = flagValue(arg, true);
|
|
436
|
+
const n = Number(raw);
|
|
437
|
+
// The wrapper's budget is whole seconds; sub-second values are
|
|
438
|
+
// rejected rather than silently rounded away.
|
|
439
|
+
if (raw === undefined || !Number.isInteger(n) || n < 1000)
|
|
440
|
+
throw new Error('--boot-timeout-ms must be an integer >= 1000');
|
|
441
|
+
options.bootTimeoutMs = n;
|
|
442
|
+
i++;
|
|
443
|
+
break;
|
|
444
|
+
}
|
|
421
445
|
case '--foreground':
|
|
422
446
|
options.foreground = true;
|
|
423
447
|
break;
|
|
@@ -586,12 +610,17 @@ function verifyRepoCredential(stateDir, repoDir, maxAgeMinutes) {
|
|
|
586
610
|
return verifyCredential(loadState(stateDir), currentHead(repoDir), Date.now(), maxAgeMinutes, isWorkingTreeClean(repoDir));
|
|
587
611
|
}
|
|
588
612
|
/**
|
|
589
|
-
* How the spawned watchdog should invoke the guard CLI
|
|
590
|
-
*
|
|
591
|
-
*
|
|
613
|
+
* How the spawned watchdog should invoke the guard CLI. The executable is
|
|
614
|
+
* always the ABSOLUTE process.execPath, never a bare `node`: the watchdog runs
|
|
615
|
+
* under whatever PATH spawned it, and a launcher chain (launchd/systemd) has a
|
|
616
|
+
* minimal PATH without homebrew — a bare `node` there makes every guard
|
|
617
|
+
* invocation the watchdog issues (canary, record-proven-deployment, cutover
|
|
618
|
+
* events) fail with command-not-found (the 2026-09-30 launchd incident). The
|
|
619
|
+
* source form additionally needs tsx with an absolute path (the watchdog runs
|
|
620
|
+
* with a deployment cwd that resolves no node_modules).
|
|
592
621
|
* @returns the command prefix (verb args are appended by the watchdog).
|
|
593
622
|
*/
|
|
594
|
-
function guardInvocation() {
|
|
623
|
+
export function guardInvocation() {
|
|
595
624
|
const cliPath = fileURLToPath(import.meta.url);
|
|
596
625
|
if (cliPath.includes(`${sep}src${sep}`)) {
|
|
597
626
|
// Source form: locate the tsx loader relative to this file's own
|
|
@@ -600,10 +629,10 @@ function guardInvocation() {
|
|
|
600
629
|
const nodeModules = resolve(dirname(cliPath), '../../../node_modules');
|
|
601
630
|
const tsx = join(nodeModules, 'tsx', 'dist', 'esm', 'index.mjs');
|
|
602
631
|
if (existsSync(tsx))
|
|
603
|
-
return
|
|
604
|
-
return
|
|
632
|
+
return `${process.execPath} --import ${tsx} ${cliPath}`;
|
|
633
|
+
return `${process.execPath} ${cliPath}`;
|
|
605
634
|
}
|
|
606
|
-
return
|
|
635
|
+
return `${process.execPath} ${cliPath}`;
|
|
607
636
|
}
|
|
608
637
|
/**
|
|
609
638
|
* argv (after process.execPath) that runs the exit agent, with the same
|
|
@@ -1540,7 +1569,7 @@ export async function runCli(argv, io) {
|
|
|
1540
1569
|
}
|
|
1541
1570
|
const watchdogPid = liveWatchdogPid(stateDir);
|
|
1542
1571
|
if (watchdogPid === null) {
|
|
1543
|
-
io.stderr(`${command} refused: no live watchdog can consume the durable control request\n`);
|
|
1572
|
+
io.stderr(`${command} refused: no live watchdog can consume the durable control request. Start the consumer first: \`supervise --state-dir ${stateDir}\` holds the awaiting-user cutover ${transaction.receipt.id} without launching anything, or \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\` resumes the selected side directly\n`);
|
|
1544
1573
|
return 1;
|
|
1545
1574
|
}
|
|
1546
1575
|
const requested = command === 'restore-previous' ? 'restore-previous' : 'abort';
|
|
@@ -1970,6 +1999,9 @@ export async function runCli(argv, io) {
|
|
|
1970
1999
|
'--takeover-from', String(previousSupervisorPid), '--cutover-id', cutoverId,
|
|
1971
2000
|
'--delay-ms', String(options.delayMs ?? 5000),
|
|
1972
2001
|
'--supervisor-yield-timeout-ms', String(options.supervisorYieldTimeoutMs ?? 15_000),
|
|
2002
|
+
// The successor watchdog's readiness budget rides the driver argv;
|
|
2003
|
+
// the durable spec stays free of per-restart tuning.
|
|
2004
|
+
...(options.bootTimeoutMs !== undefined ? ['--boot-timeout-ms', String(options.bootTimeoutMs)] : []),
|
|
1973
2005
|
...(initiator !== undefined ? ['--initiator', initiator] : []),
|
|
1974
2006
|
];
|
|
1975
2007
|
const cutoverDriverEnv = { ...process.env };
|
|
@@ -2232,10 +2264,21 @@ export async function runCli(argv, io) {
|
|
|
2232
2264
|
io.stderr(`supervise refused: cutover ${options.cutoverId} is not the selected launch transaction\n`);
|
|
2233
2265
|
return 1;
|
|
2234
2266
|
}
|
|
2267
|
+
// A receipt parked in awaiting-user must never auto-boot the rejected
|
|
2268
|
+
// side — but exiting here leaves NO live consumer for the operator
|
|
2269
|
+
// control markers, and abort-cutover/restore-previous then refuse with
|
|
2270
|
+
// "no live watchdog" (the 2026-09-30 mid-cutover wedge). Hold instead:
|
|
2271
|
+
// spawn the watchdog in its parked mode, which claims the pidfile and
|
|
2272
|
+
// consumes those markers without launching anything.
|
|
2273
|
+
let parkedCutover = false;
|
|
2235
2274
|
if (options.cutoverId === undefined && transaction?.receipt.phase === 'awaiting-user') {
|
|
2236
|
-
|
|
2237
|
-
return 0;
|
|
2275
|
+
parkedCutover = true;
|
|
2238
2276
|
}
|
|
2277
|
+
// An explicit --cutover-id resume against a live PARKED holder: the
|
|
2278
|
+
// driver-started event below flips the receipt out of awaiting-user,
|
|
2279
|
+
// which is the hold's release signal. Remember the entry phase so the
|
|
2280
|
+
// pidfile branch below knows the live owner is expected to let go.
|
|
2281
|
+
const resumeFromAwaitingUser = options.cutoverId !== undefined && transaction?.receipt.phase === 'awaiting-user';
|
|
2239
2282
|
// A detached watchdog spawned from a sandboxed turn is reaped with it —
|
|
2240
2283
|
// refuse before claiming anything. Foreground mode is driven by the
|
|
2241
2284
|
// external supervisor (launchd/systemd) and stays exempt.
|
|
@@ -2291,7 +2334,30 @@ export async function runCli(argv, io) {
|
|
|
2291
2334
|
}
|
|
2292
2335
|
if (durable === null)
|
|
2293
2336
|
writeStableLaunchSpec(stateDir, spec);
|
|
2294
|
-
if (
|
|
2337
|
+
if (resumeFromAwaitingUser) {
|
|
2338
|
+
// The live owner is the awaiting-user hold (or a wrapper still
|
|
2339
|
+
// parked on its crash page). The hold releases its claim as
|
|
2340
|
+
// soon as the driver-started event above flips the receipt out
|
|
2341
|
+
// of awaiting-user — wait for that release, bounded: an owner
|
|
2342
|
+
// that keeps the claim is not the hold, and waiting behind it
|
|
2343
|
+
// forever would wedge the explicit resume it never sees.
|
|
2344
|
+
const existingIdentity = processIdentity(existingPid);
|
|
2345
|
+
if (existingIdentity === null) {
|
|
2346
|
+
io.stderr(`supervise refused: could not capture watchdog ${existingPid} start identity before waiting\n`);
|
|
2347
|
+
return 1;
|
|
2348
|
+
}
|
|
2349
|
+
io.stdout(`watchdog ${existing} holds the parked cutover — waiting for it to release, then resuming\n`);
|
|
2350
|
+
const releaseDeadline = Date.now() + 15_000;
|
|
2351
|
+
while (processIdentityMatches(existingIdentity) && Date.now() < releaseDeadline)
|
|
2352
|
+
await sleep(250);
|
|
2353
|
+
if (processIdentityMatches(existingIdentity)) {
|
|
2354
|
+
io.stderr(`supervise refused: watchdog ${existingPid} did not release the parked cutover ${options.cutoverId ?? ''} within 15000 ms — it is not the awaiting-user hold; settle the transaction with \`abort-cutover --state-dir ${stateDir}\` or \`restore-previous --state-dir ${stateDir}\`, or stop that watchdog and retry\n`);
|
|
2355
|
+
return 1;
|
|
2356
|
+
}
|
|
2357
|
+
waitedForWatchdog = true;
|
|
2358
|
+
io.stdout(`watchdog ${existing} released the parked cutover — resuming\n`);
|
|
2359
|
+
}
|
|
2360
|
+
else if (options.foreground) {
|
|
2295
2361
|
// Foreground = an external supervisor (launchd KeepAlive) runs
|
|
2296
2362
|
// THIS process. Exiting 0 here would read as an intentional stop
|
|
2297
2363
|
// under `KeepAlive SuccessfulExit: false`, so the job would go
|
|
@@ -2334,8 +2400,10 @@ export async function runCli(argv, io) {
|
|
|
2334
2400
|
return 1;
|
|
2335
2401
|
}
|
|
2336
2402
|
if (options.cutoverId === undefined && transaction?.receipt.phase === 'awaiting-user') {
|
|
2337
|
-
|
|
2338
|
-
|
|
2403
|
+
// Settled into awaiting-user while this supervisor waited behind the
|
|
2404
|
+
// old owner: park exactly like the entry check above instead of
|
|
2405
|
+
// booting the rejected target — or exiting and leaving no consumer.
|
|
2406
|
+
parkedCutover = true;
|
|
2339
2407
|
}
|
|
2340
2408
|
io.stdout('launch state refreshed after wait — using the durable selected specification\n');
|
|
2341
2409
|
}
|
|
@@ -2412,6 +2480,11 @@ export async function runCli(argv, io) {
|
|
|
2412
2480
|
// adoption; the detached form waits for the current owner to exit.
|
|
2413
2481
|
WD_WAIT_OWNER: options.takeoverFrom !== undefined || !options.foreground ? '1' : '0',
|
|
2414
2482
|
WD_GUARD: guardInvocation(),
|
|
2483
|
+
// One boot's readiness budget, in the wrapper's whole-seconds unit.
|
|
2484
|
+
// Ceiling: never round a requested budget DOWN into a tighter window.
|
|
2485
|
+
...(options.bootTimeoutMs !== undefined
|
|
2486
|
+
? { WD_BOOT_TIMEOUT: String(Math.ceil(options.bootTimeoutMs / 1000)) }
|
|
2487
|
+
: {}),
|
|
2415
2488
|
...(options.takeoverFrom !== undefined ? {
|
|
2416
2489
|
WD_TAKEOVER_FROM: String(options.takeoverFrom),
|
|
2417
2490
|
WD_TAKEOVER_FROM_START: previousSupervisorStart ?? '',
|
|
@@ -2432,11 +2505,18 @@ export async function runCli(argv, io) {
|
|
|
2432
2505
|
WD_PREVIOUS_CHILD_START: previousOwnership.childStartToken,
|
|
2433
2506
|
WD_PREVIOUS_LISTENER_PID: String(previousOwnership.listenerPid),
|
|
2434
2507
|
WD_PREVIOUS_LISTENER_START: previousOwnership.listenerStartToken,
|
|
2508
|
+
// Parked hold: claim supervision and consume operator control
|
|
2509
|
+
// markers, but never launch — the receipt's selected side was
|
|
2510
|
+
// explicitly rejected and waits for a user decision.
|
|
2511
|
+
...(parkedCutover ? { WD_CUTOVER_PARKED: '1' } : {}),
|
|
2435
2512
|
...(transaction.state.transition === undefined ? {} : {
|
|
2436
2513
|
WD_TRANSITION_PLAN_SHA256: transaction.state.transition.planSha256,
|
|
2437
2514
|
}),
|
|
2438
2515
|
} : {}),
|
|
2439
2516
|
};
|
|
2517
|
+
if (parkedCutover && transaction !== null) {
|
|
2518
|
+
io.stdout(`supervise: cutover ${transaction.receipt.id} is waiting for user action — holding the supervision claim WITHOUT launching the rejected ${transaction.state.selected} side (receipt ${stateFile(stateDir, 'launchCutover')}). The operator verbs now have a live consumer: \`abort-cutover --state-dir ${stateDir}\` applies the pre-approved ${transaction.receipt.recovery.policy} policy, \`restore-previous --state-dir ${stateDir}\` restores the complete previous spec. To end the hold without settling: write the stop marker (\`touch ${stateFile(stateDir, 'watchdogStop')}\`). To resume the rejected side: \`supervise --cutover-id ${transaction.receipt.id} --state-dir ${stateDir}\`\n`);
|
|
2519
|
+
}
|
|
2440
2520
|
if (options.foreground) {
|
|
2441
2521
|
// Run the watchdog inline: the CLI process stays alive as the
|
|
2442
2522
|
// watchdog's parent, so an external supervisor (launchd KeepAlive)
|
|
@@ -2474,7 +2554,9 @@ export async function runCli(argv, io) {
|
|
|
2474
2554
|
const claimDeadline = Date.now() + 5_000;
|
|
2475
2555
|
while (Date.now() < claimDeadline && spawnError === undefined) {
|
|
2476
2556
|
if (liveWatchdogPid(stateDir) === spawnedPid) {
|
|
2477
|
-
io.stdout(
|
|
2557
|
+
io.stdout(parkedCutover && transaction !== null
|
|
2558
|
+
? `watchdog spawned and parked on cutover ${transaction.receipt.id} (pid ${spawnedPid}) — nothing launched; consuming abort-cutover/restore-previous markers; log ${logPath}\n`
|
|
2559
|
+
: `watchdog spawned and ready (pid ${spawnedPid}) — supervises :${spec.port}, log ${logPath}\n`);
|
|
2478
2560
|
return 0;
|
|
2479
2561
|
}
|
|
2480
2562
|
try {
|
|
@@ -2584,6 +2666,10 @@ export async function runCli(argv, io) {
|
|
|
2584
2666
|
writeFileSync(stateFile(stateDir, 'restartRequested'), `${JSON.stringify({
|
|
2585
2667
|
reason: 'scheduled self-restart',
|
|
2586
2668
|
requestedAt: Date.now(),
|
|
2669
|
+
// A one-boot readiness budget for the respawn: the live watchdog
|
|
2670
|
+
// applies it over its own WD_BOOT_TIMEOUT for boots attempted
|
|
2671
|
+
// while this marker is pending, then falls back to its default.
|
|
2672
|
+
...(options.bootTimeoutMs === undefined ? {} : { bootTimeoutMs: options.bootTimeoutMs }),
|
|
2587
2673
|
...(gate.authorization === undefined ? {} : { authorization: gate.authorization }),
|
|
2588
2674
|
...(initiator !== undefined ? { initiator } : {}),
|
|
2589
2675
|
})}\n`);
|
|
@@ -155,6 +155,8 @@ export function apply(_ctx) {
|
|
|
155
155
|
let timer;
|
|
156
156
|
let pending = consumeFallbackFragment() ?? readPending();
|
|
157
157
|
let cap;
|
|
158
|
+
/** Consecutive failed ACKs for the current pending handoff. */
|
|
159
|
+
let ackFailures = 0;
|
|
158
160
|
let errorRetryMs = ERROR_RETRY_MIN_MS;
|
|
159
161
|
let firstFailureAt;
|
|
160
162
|
let request;
|
|
@@ -231,6 +233,29 @@ export function apply(_ctx) {
|
|
|
231
233
|
return;
|
|
232
234
|
}
|
|
233
235
|
}
|
|
236
|
+
else {
|
|
237
|
+
// A failed ACK can be permanent: the cutover settled without this tab
|
|
238
|
+
// (operator restore, handoff off, a wedged supervisor) and the server
|
|
239
|
+
// keeps answering 409 'waiting'/'stale' for a dead cutoverId — the
|
|
240
|
+
// veil would otherwise stay up forever (observed 2026-09-30). Every
|
|
241
|
+
// few failures, spend one poll: 'idle' proves nothing is in flight,
|
|
242
|
+
// so the pending handoff is dead — drop it and let the ordinary
|
|
243
|
+
// poll loop remove the veil on the next tick.
|
|
244
|
+
ackFailures += 1;
|
|
245
|
+
if (ackFailures >= 8) {
|
|
246
|
+
ackFailures = 0;
|
|
247
|
+
const probe = await post({
|
|
248
|
+
version: 1, operation: 'poll', capability: cap ??= capability(),
|
|
249
|
+
}, current.signal);
|
|
250
|
+
if (stale())
|
|
251
|
+
return;
|
|
252
|
+
if (probe.ok && (await probe.json()).state === 'idle') {
|
|
253
|
+
sessionStorage.removeItem(PENDING_KEY);
|
|
254
|
+
sessionStorage.removeItem(CAPABILITY_KEY);
|
|
255
|
+
pending = null;
|
|
256
|
+
}
|
|
257
|
+
}
|
|
258
|
+
}
|
|
234
259
|
errorRetryMs = ERROR_RETRY_MIN_MS;
|
|
235
260
|
firstFailureAt = undefined;
|
|
236
261
|
schedule(ACTIVE_RETRY_MS);
|
|
@@ -287,6 +312,7 @@ export function apply(_ctx) {
|
|
|
287
312
|
};
|
|
288
313
|
pending = nextPending;
|
|
289
314
|
writePending(nextPending);
|
|
315
|
+
ackFailures = 0;
|
|
290
316
|
waitingOverlay();
|
|
291
317
|
location.reload();
|
|
292
318
|
return;
|
|
@@ -307,6 +333,7 @@ export function apply(_ctx) {
|
|
|
307
333
|
};
|
|
308
334
|
pending = nextPending;
|
|
309
335
|
writePending(nextPending);
|
|
336
|
+
ackFailures = 0;
|
|
310
337
|
waitingOverlay();
|
|
311
338
|
// The bearer remains a local variable for the shortest possible time.
|
|
312
339
|
location.replace(launch.href);
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@khorsheed/dsh-ankh-guard",
|
|
3
3
|
"description": "Hard gate for self-modification restarts: a green-build credential bound to the git HEAD, checked before any restart of the running instance",
|
|
4
|
-
"version": "0.4.
|
|
4
|
+
"version": "0.4.2",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "lib/index.js",
|
|
7
7
|
"types": "lib/types/index.d.ts",
|
|
@@ -44,14 +44,14 @@
|
|
|
44
44
|
"license": "MIT",
|
|
45
45
|
"peerDependencies": {
|
|
46
46
|
"@deepseek-ai/cordis": "^4.0.1",
|
|
47
|
-
"@deepseek-ai/dsh-agent": "^0.1.0-rc.6",
|
|
48
|
-
"@deepseek-ai/dsh-agent-preset-registry": "^0.1.0-rc.6",
|
|
49
|
-
"@deepseek-ai/dsh-agent-presets": "^0.1.0-rc.6",
|
|
50
|
-
"@deepseek-ai/dsh-client-connection": "^0.1.0-rc.6",
|
|
51
|
-
"@deepseek-ai/dsh-host-webserver": "^0.1.0-rc.6",
|
|
52
|
-
"@deepseek-ai/dsh-invariants": "^0.1.0-rc.6",
|
|
53
|
-
"@deepseek-ai/dsh-llm": "^0.1.0-rc.6",
|
|
54
|
-
"@deepseek-ai/dsh-session-persistence": "^0.1.0-rc.6",
|
|
47
|
+
"@deepseek-ai/dsh-agent": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
48
|
+
"@deepseek-ai/dsh-agent-preset-registry": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
49
|
+
"@deepseek-ai/dsh-agent-presets": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
50
|
+
"@deepseek-ai/dsh-client-connection": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
51
|
+
"@deepseek-ai/dsh-host-webserver": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
52
|
+
"@deepseek-ai/dsh-invariants": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
53
|
+
"@deepseek-ai/dsh-llm": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
54
|
+
"@deepseek-ai/dsh-session-persistence": "^0.1.0-rc.6 || ^0.2.0-rc.1",
|
|
55
55
|
"@deepseek-ai/schemastery": "^3.18.1"
|
|
56
56
|
},
|
|
57
57
|
"peerDependenciesMeta": {
|
package/scripts/dsh-watchdog.sh
CHANGED
|
@@ -48,6 +48,11 @@
|
|
|
48
48
|
# WD_CUTOVER_ID=ID durable launch-cutover receipt transaction
|
|
49
49
|
# WD_CUTOVER_POLICY= restore-previous or wait-for-user (approved pre-stop)
|
|
50
50
|
# WD_CUTOVER_DELAY_SECONDS=N grace after supervisor claim before old-child stop
|
|
51
|
+
# WD_CUTOVER_PARKED=1 hold the supervision claim while the receipt sits in
|
|
52
|
+
# awaiting-user: never launch the rejected side, consume
|
|
53
|
+
# abort/restore control markers, exit on watchdog-stop,
|
|
54
|
+
# and release the claim when a --cutover-id resume flips
|
|
55
|
+
# the receipt out of awaiting-user
|
|
51
56
|
# WD_PREVIOUS_CHILD_*=PID/start token authoritative old supervisor child root
|
|
52
57
|
# WD_PREVIOUS_LISTENER_*=PID/start token listener inside that old child tree
|
|
53
58
|
# WD_READY_STABILITY_SECONDS=N unchanged child/listener proof window (default 3)
|
|
@@ -61,7 +66,9 @@
|
|
|
61
66
|
# WD_TEST_BREAK=1 launch a command that always fails (give-up testing)
|
|
62
67
|
#
|
|
63
68
|
# Markers under the state directory (written by the app or the agent):
|
|
64
|
-
# restart-requested.json -> intentional restart: respawn + canary + clear
|
|
69
|
+
# restart-requested.json -> intentional restart: respawn + canary + clear;
|
|
70
|
+
# an optional integer bootTimeoutMs field overrides
|
|
71
|
+
# WD_BOOT_TIMEOUT for boots attempted while it is pending
|
|
65
72
|
# watchdog-stop -> exit the watchdog without respawn
|
|
66
73
|
set -u
|
|
67
74
|
|
|
@@ -219,6 +226,19 @@ cutover_control_action() {
|
|
|
219
226
|
' "$CONTROL_RESTORE_FILE" "$CONTROL_ABORT_FILE" "$CONTROL_FILE" "$CUTOVER_ID" 2>/dev/null
|
|
220
227
|
}
|
|
221
228
|
|
|
229
|
+
# The durable receipt's current phase, empty when unreadable. The parked hold
|
|
230
|
+
# treats an unreadable receipt as still parked (hold, never flap): only a
|
|
231
|
+
# cleanly read non-awaiting-user phase releases the claim.
|
|
232
|
+
cutover_receipt_phase() {
|
|
233
|
+
node -e '
|
|
234
|
+
const fs = require("fs")
|
|
235
|
+
try {
|
|
236
|
+
const value = JSON.parse(fs.readFileSync(process.argv[1], "utf8"))
|
|
237
|
+
if (typeof value?.phase === "string") process.stdout.write(value.phase)
|
|
238
|
+
} catch {}
|
|
239
|
+
' "$STATE_DIR/launch-cutover.json" 2>/dev/null
|
|
240
|
+
}
|
|
241
|
+
|
|
222
242
|
# The host owns the shape of its per-process launch URL. Discovery is generic:
|
|
223
243
|
# the first HTTP URL printed by THIS attempt whose authority is exactly the
|
|
224
244
|
# supervised loopback authority and whose query is non-empty. No parameter
|
|
@@ -1239,6 +1259,57 @@ if [ -n "${ANKH_GUARD_TEST_RUN_DIR:-}" ]; then wd_sleep 0.01; fi
|
|
|
1239
1259
|
test_event_self "${ANKH_GUARD_TEST_PROCESS_ROLE:-watchdog}" keepalive-first-tick
|
|
1240
1260
|
test_event_self "${ANKH_GUARD_TEST_PROCESS_ROLE:-watchdog}" ready
|
|
1241
1261
|
|
|
1262
|
+
# Parked awaiting-user hold (WD_CUTOVER_PARKED=1, spawned by a bare `supervise`
|
|
1263
|
+
# that found the receipt waiting for a user decision): the selected side was
|
|
1264
|
+
# explicitly rejected, so NOTHING is launched here. The hold keeps exactly one
|
|
1265
|
+
# live consumer for the operator control verbs — without it, a supervisor chain
|
|
1266
|
+
# that died mid-cutover wedges the transaction: abort-cutover/restore-previous
|
|
1267
|
+
# refuse with "no live watchdog", and a bare supervise used to exit 0.
|
|
1268
|
+
if [ "${WD_CUTOVER_PARKED:-0}" = "1" ]; then
|
|
1269
|
+
wd_log "launch cutover $CUTOVER_ID is awaiting user action — holding the supervision claim WITHOUT launching the rejected $CUTOVER_ROLE side; settle: abort-cutover / restore-previous; exit: the watchdog-stop marker; resume the rejected side: supervise --cutover-id $CUTOVER_ID"
|
|
1270
|
+
while true; do
|
|
1271
|
+
# Same ownership discipline as the main loop: self-heal a deleted claim,
|
|
1272
|
+
# yield to a different live owner.
|
|
1273
|
+
if [ "$SUPERVISE" = "1" ]; then
|
|
1274
|
+
if [ ! -f "$PIDFILE" ]; then (set -C; echo $$ > "$PIDFILE") 2>/dev/null || true; fi
|
|
1275
|
+
pidowner=$(cat "$PIDFILE" 2>/dev/null)
|
|
1276
|
+
if [ -n "$pidowner" ] && [ "$pidowner" != "$$" ] && kill -0 "$pidowner" 2>/dev/null; then
|
|
1277
|
+
wd_log "pidfile now owned by live pid $pidowner — yielding the parked hold"
|
|
1278
|
+
yielded=1
|
|
1279
|
+
exit 75
|
|
1280
|
+
fi
|
|
1281
|
+
fi
|
|
1282
|
+
# Deliberate stop: exit without respawn or settlement, exactly like the
|
|
1283
|
+
# main loop's stop-marker path.
|
|
1284
|
+
if [ -f "$STOP_MARKER" ]; then
|
|
1285
|
+
wd_log "stop marker present — exiting the parked hold"
|
|
1286
|
+
rm -f "$STOP_MARKER" "$PIDFILE"
|
|
1287
|
+
exit 0
|
|
1288
|
+
fi
|
|
1289
|
+
# Operator control before the release check: a restore decision consumes
|
|
1290
|
+
# the marker and breaks into the ordinary cutover resume path below, which
|
|
1291
|
+
# boots the previous complete launch specification.
|
|
1292
|
+
if handle_cutover_control; then
|
|
1293
|
+
if [ "$control_result" = "restore" ]; then
|
|
1294
|
+
wd_log "operator control selected the previous complete launch specification"
|
|
1295
|
+
break
|
|
1296
|
+
fi
|
|
1297
|
+
wd_log "operator control consumed; cutover $CUTOVER_ID remains parked awaiting user action"
|
|
1298
|
+
continue
|
|
1299
|
+
fi
|
|
1300
|
+
# A `supervise --cutover-id` resume records driver-started, flipping the
|
|
1301
|
+
# receipt out of awaiting-user — this hold's release signal. Exit non-zero
|
|
1302
|
+
# like a yield so an external supervisor (launchd/systemd) restarts its
|
|
1303
|
+
# stable launcher, which then waits behind the resuming chain.
|
|
1304
|
+
phase=$(cutover_receipt_phase)
|
|
1305
|
+
if [ -n "$phase" ] && [ "$phase" != "awaiting-user" ]; then
|
|
1306
|
+
wd_log "cutover $CUTOVER_ID left awaiting-user (phase $phase) — releasing the claim for the resuming supervisor"
|
|
1307
|
+
exit 75
|
|
1308
|
+
fi
|
|
1309
|
+
wd_sleep 1
|
|
1310
|
+
done
|
|
1311
|
+
fi
|
|
1312
|
+
|
|
1242
1313
|
write_cutover_restart_marker() {
|
|
1243
1314
|
node -e '
|
|
1244
1315
|
const fs = require("fs")
|
|
@@ -1252,6 +1323,21 @@ write_cutover_restart_marker() {
|
|
|
1252
1323
|
' "$RESTART_MARKER" "$CUTOVER_ID" "${WD_INITIATOR:-}"
|
|
1253
1324
|
}
|
|
1254
1325
|
|
|
1326
|
+
# A scheduled exit may carry a one-boot readiness budget (schedule-exit
|
|
1327
|
+
# --boot-timeout-ms), recorded in the restart marker as whole milliseconds.
|
|
1328
|
+
# Validate before trusting: an unreadable or out-of-shape marker changes
|
|
1329
|
+
# nothing — the watchdog's own WD_BOOT_TIMEOUT stays in effect.
|
|
1330
|
+
restart_marker_boot_timeout() {
|
|
1331
|
+
node -e '
|
|
1332
|
+
const fs = require("fs")
|
|
1333
|
+
try {
|
|
1334
|
+
const value = JSON.parse(fs.readFileSync(process.argv[1], "utf8"))
|
|
1335
|
+
const ms = value?.bootTimeoutMs
|
|
1336
|
+
if (Number.isInteger(ms) && ms >= 1000) process.stdout.write(String(Math.ceil(ms / 1000)))
|
|
1337
|
+
} catch {}
|
|
1338
|
+
' "$RESTART_MARKER" 2>/dev/null
|
|
1339
|
+
}
|
|
1340
|
+
|
|
1255
1341
|
consume_takeover_control() {
|
|
1256
1342
|
if handle_cutover_control; then
|
|
1257
1343
|
if [ "$control_result" = "wait" ]; then
|
|
@@ -1460,9 +1546,19 @@ while true; do
|
|
|
1460
1546
|
current_listener_start=''
|
|
1461
1547
|
# Boot window: transport-up is not enough. A protected root can answer 401;
|
|
1462
1548
|
# ready_probe completes the process's announced launch-URL cookie exchange.
|
|
1549
|
+
# The budget is the watchdog's own WD_BOOT_TIMEOUT unless the pending restart
|
|
1550
|
+
# marker carries a one-boot override (schedule-exit --boot-timeout-ms).
|
|
1551
|
+
attempt_boot_timeout=$BOOT_TIMEOUT
|
|
1552
|
+
if [ -f "$RESTART_MARKER" ]; then
|
|
1553
|
+
marker_boot_timeout=$(restart_marker_boot_timeout)
|
|
1554
|
+
case "$marker_boot_timeout" in
|
|
1555
|
+
''|*[!0-9]*) ;;
|
|
1556
|
+
*) attempt_boot_timeout=$marker_boot_timeout ;;
|
|
1557
|
+
esac
|
|
1558
|
+
fi
|
|
1463
1559
|
up=0
|
|
1464
|
-
readiness_failure_detail="readiness not proven within ${
|
|
1465
|
-
boot_limit=$(( $(date +%s) +
|
|
1560
|
+
readiness_failure_detail="readiness not proven within ${attempt_boot_timeout}s"
|
|
1561
|
+
boot_limit=$(( $(date +%s) + attempt_boot_timeout ))
|
|
1466
1562
|
while [ "$(date +%s)" -lt "$boot_limit" ]; do
|
|
1467
1563
|
if [ -n "$(cutover_control_action)" ]; then
|
|
1468
1564
|
readiness_failure_detail="operator control interrupted readiness"
|
|
@@ -1499,8 +1595,15 @@ while true; do
|
|
|
1499
1595
|
redact_launch_urls_in_output
|
|
1500
1596
|
# Mirror the captured output into the watchdog log: with plain redirection
|
|
1501
1597
|
# (see the launch site) the attempt log is the only place the failure was
|
|
1502
|
-
# written, and the watchdog log is where an operator looks first.
|
|
1503
|
-
|
|
1598
|
+
# written, and the watchdog log is where an operator looks first. An EMPTY
|
|
1599
|
+
# attempt log is itself the signal — "the instance never wrote a line" is
|
|
1600
|
+
# what distinguished a dead-on-arrival boot from a noisy one in the
|
|
1601
|
+
# 2026-09-30 launchd incident, and mirroring nothing hid it for an hour.
|
|
1602
|
+
if [ -s "$ATTEMPT_LOG" ]; then
|
|
1603
|
+
sed 's/^/[instance] /' "$ATTEMPT_LOG" 2>/dev/null
|
|
1604
|
+
else
|
|
1605
|
+
wd_log "attempt $current_attempt produced no output — the instance never wrote a line (captured log $ATTEMPT_LOG is empty)"
|
|
1606
|
+
fi
|
|
1504
1607
|
|
|
1505
1608
|
if [ -n "$CUTOVER_ID" ] && [ -n "$(cutover_control_action)" ]; then
|
|
1506
1609
|
failures=$((failures + 1))
|