@sema-agent/client-core 0.82.1 → 0.82.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -49,6 +49,22 @@
49
49
  > 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
50
50
  > commit 漏转在发布前就红,不再靠人记。
51
51
 
52
+ ## 0.82.3(2026-09-24)
53
+
54
+ ### Added
55
+ - **自愈提示可以带「保住这段对话、把卡找回来」的出路**(CC-177):卡重开不了的两句(计划审阅卡 / 审批与提问卡)此前只给两条路 —— `/clear`(开新会话,这段对话就丢了)和手工向引擎发决断请求;而真能救回的往往是退出后续开本会话。`ActiveRunSelfHealCopy` 新增 `keepSessionWayOut`:宿主按此刻走不走得通,返回一个动作半句(小写起头、不带句末标点,例如退出后按会话 id 续开),走不通返回 `null`。包把它排在「放弃」那一路之前,两条路都给;句子说的是「续开后卡若还在等就会回来」,不承诺一定回来。以下情形不给、也不调用宿主函数:注入件、「决断还在路上」、宿主用 `rowFor` 整行覆写了这一结局,以及这道门的待决项已被证实不在引擎上 —— 这一形的结局新增机读位 `pendingRowGone: true`(续开也找不回)。这个证明只在呈卡**之前**算数:清理卡在屏上期间 run 可能转去等另一道门,所以呈卡之后的收口不带这一位,除非之后复读确认 run 已经结束。宿主函数抛错或返回空串按走不通。另新增 `engineDecidePath: false`,关掉交互面各行「直接在引擎上发决断」那半句;取消入口不受影响,headless 面也不受影响。两位都不给时,措辞与 0.82.2 逐字节相同。接入文档 **§95 S-1**。
56
+
57
+ ## 0.82.2(2026-09-24)
58
+
59
+ ### Added
60
+ - **云控制面有效配置回体的降级告警、预算与能力上限**(CC-164):`projectEffectiveBody` 的结果新增 `effectiveWarnings` / `budget` / `runtimeCaps` 三个读数,给人看的句子统一出自新增的 `cloudEffectiveNotices(projection)`。告警按开集读,认不出的种类也出一句提示;一行告警坏了只算坏一行,不连累别的行。预算与能力上限分四态:配了 / 这个主体没解析出(`null`)/ 预览别人时被抹掉(`null` 且带 `previewOf`)/ 判不出。只有「这个主体没解析出」那一态才说没有预算;被抹与判不出都不会被读成 0 或「没有预算」;读不懂的预算字段逐个点名为「没显示」,「没配 / 被抹 / 判不出」三类句子里不含数字。接入文档 **§94 S-1**。
61
+ - **计划复核结局带机读位**(CC-160 ①):投给宿主队列口的结局项新增 `_sema_planReviewOutcome: { taskId, dispatchNo, decision, effect }`,宿主不必再从结局文字里找任务号(仍锁 / 推进 / 没送出 / 拒绝生效这几种的文字本来就不带任务号)。`dispatchNo` 是本包铸的投递序号,只在同一份包实例内单调,进程重启从 1 起,引擎不认识它;被在飞闸拒掉的那次不铸号,也不投结局。结局文字逐字不变;`enqueuePlanReviewOutcome` 多了一个可选第二参,不给则队列项上没有这一键。接入文档 **§94 S-3**。
62
+ - **审批决断操作的起止回调**(CC-160 ③):`HitlBridge` 构造可以带第三参 `{ onDecideOperation }`,`AskGateWireDeps` / `FsApprovalWireDeps` 有同名可选位,透传到两条停泊决断腿。一次决断(含瞬断重试)在第一发请求之前同步报 `start`;最后一发结算、不会再有下一发时恰好报一次 `end`(带 `ok` / `attempts` / `retryExhausted`);重试之间不报。回调抛错或返回被拒的 promise 都被吞掉,不改变这次决断。包内对同一次决断的再投递(老 server 不认会话级放行时回退纯 approve、归因键不在签名集里时去键重发)算同一个操作,只报一对 `start` / `end`;宿主自己要把多发并成一次时,用 `bridge.openDecideOperation()` 开句柄放进 `decideTool` 的 `opts.operation`;句柄只并同一会话、同一 `boundCallId` 的发,拿去发另一道门的那一发单独成一个操作;关句柄时还有一发在飞,`end` 等它结算后才报。计划审阅、续跑与计划复核投递不经这个回调。宿主不必再照着包的重发判据自写一份「这一发还会不会再发」。接入文档 **§94 S-4**。
63
+
64
+ ### Fixed
65
+ - **本机引擎的模型目录补上默认模型与档位组**(CC-165):写给本机引擎的 models 文档此前丢掉 `default` / `tierGroups` / `activeTierGroup`,面板上选的默认模型与档位组在本机不生效。现在宿主在 `cloudModelsToDoc` / `projectEffectiveBody` 注入本机条目判定 `entryAccepted`(用宿主手上 settings-schema 的 `ModelEntry` 判定实现)后,三键都带上;不注入则三键照旧不投 —— 本机引擎会先剔掉格式不合的条目再判引用,本包判不出哪些条目会被剔,投了就可能让整份目录作废。注入时,判定口不收的条目本身不写进文档 —— 本机引擎启动时会逐条剔掉它,运行中刷新时却会因它拒收整份配置;一条都没过判定口时目录原样保留并出一句提示(写成空目录,刷新时会被当成合法配置收下、改用环境目录,上一份就丢了)。指向被拒条目的引用一并剔掉(记 `reason:'rejected'`);档位词只带本机格式认得的那几个,表外词剔掉并记 `reason:'unsupported'`。本机引擎会整个拒收的目录引用(指向不在目录里的模型或档位组、坏组名、重名组)逐条剔掉,并记进新字段 `modelRefIssues` —— 留着其中任何一条,本机引擎都会弃用整份云端目录:启动时改用自己的环境目录,运行中刷新时拒收这份新配置、保留上一份。修前就会触发的「@ 名单里有被授权排除的模型」一并修正;@ 名单整张都不在目录里时原样保留并给一句提示(剔空会被读成「谁都能 @」)。接入文档 **§94 S-2**。
66
+ - **云端改了模型配置,本机引擎的刷新不再因条目上的来源注记整份被拒**:registry 的个人层会给每一条模型条目打上来源注记(`origin`,覆盖团队条目时另带 `overridesTeam`),本机引擎的模型条目格式不认这两个键,刷新时会因此拒收整份配置、继续用上一份。现在写 models 文档前把这两个注记键去掉,条目其余内容逐字不变。接入文档 **§94 S-2**。
67
+
52
68
  ## 0.82.1(2026-09-24)
53
69
 
54
70
  ### Added
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.82.1
38
+ **Version:** 0.82.3
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -348,9 +348,9 @@ guard still cross-checks the table by name).
348
348
  | `scripts/run-background-view-test.mjs` | `createBackgroundView` lifecycle: polling/notify pairing, per-source degrade (`501 → not-configured` vs `unavailable`), the capabilities `scheduler` probe, and dispose really aborting the in-flight fleet snapshot (pure projection lives in the pure suite's W-A segment) |
349
349
  | `scripts/run-fleet-view-keys-test.mjs` | The fleet projection views, **both directions**: `FleetTaskView`/`FleetWorkflowView` ⇄ their key lists (compile-pinned) ⇄ what a maximal/minimal row really projects, plus a wire-key coverage ledger (every `FleetTaskRow` key is either projected or carries a written reason why not) and a drift ledger against the shell's render contract. A one-directional assignability check is blind to optional keys — which is how `startedAt` was silently dropped |
350
350
  | `scripts/run-usage-verbatim-channel-test.mjs` | The two complementary usage disciplines (core 3.0.0 metering semantics): the CC `ModelUsage` mirror stays pure (five pinned keys, `totalInputTokens` has no seat), while the sema-owned channel forwards the engine `turn_end.usage` object **verbatim** (six keys, incl. `totalInputTokens`) via `last_turn_usage.engineUsage` / `handle.latestEngineUsage` — honest absence on pre-3.0.0 engines, no fabricated zeros |
351
- | `scripts/run-plan-review-decide-verify-test.mjs` | `decidePlanReview`'s post-decide honesty ([2315]/[2316], engine RB-471 family): a 2xx from the decide endpoint is **not** a terminal — the wire re-pulls the task status and words the outcome by the real shape (still-locked / legal new gate / genuinely left park / unverified), never claiming success it hasn't earned; when the engine answers that the session's stored resume context cannot be read, the outcome names the unreadable row and says the decision was not applied. Driven against a real fake-engine HTTP server through the shipped dist |
351
+ | `scripts/run-plan-review-decide-verify-test.mjs` | `decidePlanReview`'s post-decide honesty ([2315]/[2316], engine RB-471 family): a 2xx from the decide endpoint is **not** a terminal — the wire re-pulls the task status and words the outcome by the real shape (still-locked / legal new gate / genuinely left park / unverified), never claiming success it hasn't earned; when the engine answers that the session's stored resume context cannot be read, the outcome names the unreadable row and says the decision was not applied. Driven against a real fake-engine HTTP server through the shipped dist. The outcome queue item also carries a machine-readable `_sema_planReviewOutcome` (task id, a package-minted dispatch number, decision, effect) so a host can tell which in-flight decision an outcome belongs to without searching the prose; the prose is byte-identical, the number is minted only for a decision the in-flight latch admits, and a caller that passes no metadata gets no key. |
352
352
  | `scripts/run-shell-gate-durable-allow-test.mjs` | #110: the durable approval leg for **shell** gates. The tool_end HOLD/REJECT predicate must cover Bash the same way park detection already does (otherwise the park poison frame `Operation aborted` hits the transcript, `endedCalls` swallows the real replayed result, and the user who pressed Yes watches a command that really ran be reported as aborted); a replayed, already-decided park must resume reading the stream instead of being reported as a failed turn; `lastEventId` must track numeric `seq` too. Mutation-proven: each of the three fixes reverted turns the gate red |
353
- | `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings |
353
+ | `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings. A final section pins the decide-operation observer: one `start` in the same tick as the first request and exactly one `end` after the last attempt has settled, across success, retried timeouts, exhausted transient failures, semantic refusal, binding mismatch and both kinds of caller abort, with nothing between retries; an observer that throws or rejects — even when logging that fault fails — never changes what is sent or returned, and the stream-level dependency reaches both durable park legs. A package-internal re-delivery of the same decision (plain approve after an older server rejects the session-scope flag, or a re-send without the attribution key) is reported as one operation with a single start and end. An operation handle only groups sends for the same session and bound call — a send for another gate through the same handle is its own operation — and closing a handle while a send is still in flight defers the end until that send settles. |
354
354
  | `scripts/run-park-hop-progress-test.mjs` | L-80: the park re-attach loop budgets **stalled** rounds, not parks. A turn where the model keeps hitting gates and every one of them is really decided (a card was answered, the engine really moved on) must never be cut off by the hop budget — the budget counts consecutive rounds that produced no progress, and "the engine revived and immediately parked again on the same coordinates" is not progress. The three non-progress arms (already-resolved, decide-transport-exhausted, and a re-scan that was adopted but led nowhere) share one same-cause limit instead of one arm having a limit and the others having none, and every non-progress re-attach is announced once through the host callback rather than only to the debug log. When the limit is spent the resolver reads the approval queue once more and puts whatever is decidable in front of the user before it gives up; only when there is genuinely nothing to show does it fail soft, and the terminal message then carries the real cause and a real way out instead of a sentence about a budget. On the self-heal side, a reopen verdict that reports `decidedWithoutCard` — the chain settled the gate by rule, so there was no card to present — is progress, not a reopen failure, and the user is not told their message was NOT sent. Negative control: a genuinely empty queue with a run that never moves still fails soft |
355
355
  | `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
356
356
  | `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
@@ -370,11 +370,11 @@ guard still cross-checks the table by name).
370
370
  | `scripts/run-type-superset-ledger-test.mjs` | The type/wire **superset ledger** (`docs/type-superset.json`): positions this package adds on top of a CC-shaped contract, each carrying the evidence for what CC's own type surface does or does not have there. Completeness is deliberately uneven and the ledger says so. The `_sema_*` private-key class is checked in **both** directions (a key in the source that never entered the ledger is red, naming key and file; a ledger row whose key left the source is red) — but only for keys written as literals, which is the convention the ledger mandates. A key assembled by string arithmetic is beyond what any static rule can enumerate, so the guard fails closed on every shape it *can* decide (a bare `_sema_` prefix is red wherever it appears, save one pinned guard site) and leaves the rest as a convention violation for review to catch, rather than claiming a completeness it does not have. The two hand-surveyed classes are only checked for coordinate and evidence integrity, never discovered. Both directions read the source through the **TypeScript AST**, not a text scan, and they read two different sets out of it. A *key site* is an identifier, or a string whose whole value is the key — so `'_sema_decision-v2'` is carried whole rather than truncated at the first non-identifier character into some *other* key that happens to be registered. A *mention* is the key appearing inside a longer string, which is prose, not usage. The staleness direction counts key sites only: a comment or a doc sentence left behind after the last real mint site is deleted must not keep the row alive (mutation-proven — with both the comment and the prose string untouched, removing the one real site turns the guard red). And because a prefix can be concatenated or interpolated into a key no static set will ever see, the bare `_sema_` literal is refused outright rather than traced: every occurrence is red except the single inline `startsWith` guard the sanitizer needs, because the set of expressions a bare prefix can travel through on its way to a concatenation is open-ended and enumerating it is always one form behind. Every row's `host` must still resolve, with the key being a real **member of that declaration** rather than a string occurring somewhere in the same file — `governanceForced`/`delegation` each live on two different shapes in one file, and a member commented out is a member deleted, which a text-shaped check happily reads as still present. And the direction worth the most: each machine-form `ccAbsenceEvidence` is re-derived from the row's own `key` — the ledger's recorded string must match that derivation verbatim, since a row quietly witnessing `\bnever_present\b` is green forever while watching nothing (mutation-proven: the same edit passes the unbound form and is caught by the bound one) — and the check runs against the names the installed `@sema-agent/agent-types` `.d.ts` set actually declares, parsed with the TypeScript AST rather than grepped, so a name CC merely mentions in a comment cannot force the row into the manual escape hatch and thereby retire the very witness that was supposed to fire the day CC declares that name for real. That escape hatch is gated by an allowlist living **in the guard**, not the ledger, so claiming it costs a reviewed diff. Missing material never reads as a pass, and the verdict splits by *why* it is missing: no TypeScript parser skips the suite before it starts; a missing `agent-types` still runs and prints the first three directions, then exits **1** when `package.json` declares the mirror but it is not installed — a broken install must not retire the repository's only "the day CC declares this name" alarm, and reporting it as a skip would leave "never evaluated" and "evaluated, no drift" indistinguishable to the runner — and exits 3 only when nothing declares the mirror at all, which is the one case where the direction genuinely does not apply. Either way a run that evaluated no witness is never counted as one that did. When the mirror *is* present its **installed version** is witnessed too (the two declared floors must agree with each other and the installed copy must meet them), since four preflight probes are satisfied by an arbitrarily stale mirror — they prove the extractor speaks, not that it is current. Every direction carries a positive control — known-present CC symbols, a comment-only sample proving the extractor distinguishes declaration from mention, and synthetic corpora fed through the **same** discriminator function the real verdict uses, so a verdict quietly rewritten to return nothing takes its own control down with it |
371
371
  | `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" |
372
372
  | `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
373
- | `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited |
373
+ | `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited. The two "card could not be reopened" rows can carry a host-declared way to keep the conversation, which says the card comes back on resume only if it is still waiting: it is placed before the route that abandons it, never offered for an injected submission, while a decision is still on its way, or once the pending approval has been proven gone (the outcome then carries a flag saying so; the proof only counts before the cleanup card is shown, so a fallback after the card carries no flag unless a fresh read finds the run finished, and a recheck that finds the approval back clears it), the host function is not even called in those cases, and it is treated as unavailable when it throws or returns an empty value; a host can also switch off the engine decide route on the interactive rows while the cancel route stays, and with neither given all four rows are pinned byte-for-byte to the text the previous release produced. |
374
374
  | `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `isApiErrorMessage` class flag so brief-mode visibility filtering does not silently drop it. The second half is who the rejected submission belonged to: the self-heal copy told every caller "Your message was NOT sent … send it again", which is three separate untruths for a system injection (a plan-review outcome, a cron wake-up, a task notification) — not the user's message, and not re-sendable, since a host queue marks those non-editable and non-recallable. The injected form says so instead, and the one sentence that promises re-delivery is pinned to the single disposition that earns it: `selfHealSubmissionDisposition` is the same function the host consults before putting the item back on its queue, so the promise and the behaviour cannot drift apart, and the arms where no card could be surfaced state plainly that nothing was delivered and nothing will retry. Since 0.72.6 the same gate pins the **follow intent** after a steer (): a message handed to a live run only pays off if someone tails that run's own event stream, so `steerFollowIntent` decides from the delivery word whether to tail now, after the pending decision, or only after a wake — and the "watch that run" sentence turns into a factual "sema is following that run" **only** when the host declares it attached that tail, so a shell that did not wire it can never claim it did |
375
375
  | `scripts/run-additive-key-passthrough-test.mjs` | The one disease shape behind two legs: a **closed whitelist / flattening arm** dropping a fact that is already on the wire, while both sides of the seam look correct. (1) The `task_progress` projection carries a registered **key ledger** — a frame populated with every key the service really projects is pushed through the shipped `eventToSdkMessage`, and the set of wire keys that survive must equal the registered pass-through list **name for name in both directions**, so quietly forwarding one more key is as red as quietly dropping one. `model` (the child run's model id, minted by core as `prepared.model.id` and projected by the server since 7.52.1) is the key this batch adds, with the same conditional the server itself applies: a non-empty string or no key at all — an empty string is neither a model id nor "unknown". The ledger is also checked against the fenced list in `docs/INTEGRATION-CLIENTS.md` §3d, so a doc that still says seven keys while the code forwards eight is red rather than merely stale. (2) The decide-failure arms carry the server's S-02 `currentPending` pointer key from a 409 `approval_stale` refusal onto the outcome the host reads. The reader is structural rather than `instanceof`, because the client is host-injected and the class identity is not this package's to assume; a half triple never mints (half a pointer cannot relocate anything), an empty string is not presence, and `checkpointToken` never transits. Both the allow and the deny leg are driven end to end through the real durable approval path — as is the accept-session leg, where a refusal carrying the pointer key must now re-raise instead of silently re-sending the human's answer for the **old** card as a plain approve (one decide call, pointer preserved), while a legacy 400 still falls back exactly as before — and all three flattening points must call the one shared reader — the same-shape residue check that makes "fixed one arm and left the twin" red instead of invisible. (3) The same disease growing on the REQUEST side: the `.mcp.json` → server-spec projection rebuilds each server key by key, and the settings schema deliberately leaves some keys parse-transparent — whatever JSON the file carries reaches the engine untouched, because validating them where the whole domain parses all-or-nothing would let one bad declaration take every server down silently. The whitelist had no row for the newest of them, so an operator's per-tool declarations — the ones the write fence reads — were stripped at the package boundary while both sides looked correct. The criterion is not "is that key handled" but the transparent-key table read out of the INSTALLED schema at runtime, reconciled name-for-name against this leg's ledger, so the day upstream adds a third one this turns red and forces an explicit decision. Behaviour is pinned on both transports, by object identity rather than deep equality (a rebuild would be a second judge), and malformed values must transit UNCHANGED rather than be refused here — the engine refuses them loudly and names the server, whereas a package-side judge can only swallow a declared protection quietly. Absence still mints no key, unknown keys still never reach the wire (the fix is the dropped key, not the gate), and the one transparent key this leg deliberately does not forward is a ledger entry with its own exit condition: it belongs to the deployment plane, and the day the request-plane type declares it the entry's premise is gone and the gate says so |
376
376
  | `scripts/run-esc-halt-plan-test.mjs` | The Esc stop decision every client shares: fire the **turn-level** halt first, and escalate to a **run-level** cancel in exactly two cases — the engine itself answered with a 409 from the closed code set (it is saying "there is no in-flight turn here; use cancel for a run-level stop"), or that shot came back with no verdict at all *and* the shell can independently prove a permission card was on screen. Everything else does not escalate. The asymmetry is the whole point and every negative control guards the same direction — deciding *not* to escalate costs the user one more choice on a busy-session card (recoverable), deciding to escalate wrongly tears down a run that was alive and takes every in-flight tool with it (not). So: the closed code set is a **frozen** value, not a `ReadonlySet` — type-level immutability does not stop a consumer's `.add()`, and the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift; the escalation gate is the **conjunction** of that closed set and the 409 status, since honouring the code alone lets a 500 that merely quotes it drive a destructive call; `interrupt.not_held` and `steering.not_running` are deliberately outside the set (the first means *this replica* has no live face — the run may be perfectly alive on another); an unreadable code falls to the no-escalation side; a `parked` flag never overrides a verdict the engine did give, and only strict `true` counts when it did not. The first shot is unconditional by construction — it does not consult `parked`, because the 409 it earns is exactly the verdict the gate wants — and the verdict itself is a closed machine-readable reason word, not display copy. A third escalating case was added once tearing the stream stopped reaping the run: with detach armed, a shot that never lands leaves the run going all the way to the end of the turn, so the Esc the user pressed has no effect at all and nothing on screen says so — the old behaviour had a silent backstop (tearing the stream ended the run) and that backstop is gone. The new fact is held to the same three disciplines as `parked`: it is read only where the engine gave no verdict, it is judged **after** `parked` so an existing host's reason word does not change under it, and only strict `true` counts. Absence is proven to be a no-op rather than asserted — the guard carries its own reference implementation of the previous version's table, runs the full grid through both, requires zero divergence when the new field is omitted, and first shows the comparison really does report a difference on the one cell where the two versions are meant to differ |
377
- | `scripts/run-peer-frame-projection-test.mjs` | The three engine-injected lanes design/385 puts on the **one** `task_notification` carrier, which are not the same kind of thing at all: a delegated child's uplink (`agentMessage`), another session's message drained from this session's own box (`crossSessionMessage`), and a receipt about one of *this* session's own outbound messages (`crossSessionNotice`). The engine renders none of them inside a `<task-notification>` shell, so a client that projects them as the generic completion card shows "background task finished" while the model read a colleague's sentence — two faces describing different events. The discriminator is pinned to the **typed carrier being present**, never to the `summary` text: those carriers can only be minted by the engine's injection legs (the external `notify()` input is a strict subset of the payload and can wear none of them), while `summary` is filled by every notification there is — so anchoring on text would let any background task impersonate a colleague's message by writing `<agent-message from="…">` into its own summary, and a positive control asserts exactly that payload still projects as the generic card. Fail-closed has two tiers rather than one: a broken **required** field (empty `from`, a non-string `body`, a notice `kind` outside the closed set) returns absence so the caller falls back to the generic card — an honest downgrade where the user still sees the notification — while a broken **optional** field drops only itself, because losing an attribution note and losing a colleague's whole message are not the same magnitude. The provenance side record is **required and must agree on four points** (`kind` matches the lane; `from`/`taskId`/`seq` are present and equal the carrier/payload — each equality is anchored on a core mint site and pinned by the cli wire-anchor A-K24), so a carrier signed with a trusted name but a disagreeing provenance falls back to the generic card; peer bodies pass the same authority-envelope neutralization core applies (`<task-notification>` etc. are defused) so a colleague's text can never seed the resume dedup ledger. Lane precedence copies the engine renderer's own order, because the model already read the frame in that order and a client ordering of its own would put a card on screen that disagrees with the frame the model saw. Rendering and parsing of the transcript line live in the same module and are round-tripped in both directions, including a body carrying a forged closing tag (a parser fooled there hands half a message to the next row) and a quote inside the sender label (which must not forge a second attribute); the notice lane is deliberately kept **out** of the parser, since recognising it would mean anchoring the `[Cross-session …]` prefix and a user typing that same line would be rendered as engine speech. Hostile carriers are read as own **data** descriptors only and accessors are never invoked at all — `catch` catches throwing, not never returning — proven by a counting getter that must stay at zero calls, alongside a revoked proxy and a prototype-only carrier; and four legacy payload shapes assert the no-carrier path is byte-identical to before, which is the executable form of "zero difference for an older host" |
377
+ | `scripts/run-peer-frame-projection-test.mjs` | The three engine-injected lanes design/385 puts on the **one** `task_notification` carrier, which are not the same kind of thing at all: a delegated child's uplink (`agentMessage`), another session's message drained from this session's own box (`crossSessionMessage`), and a receipt about one of *this* session's own outbound messages (`crossSessionNotice`). The engine renders none of them inside a `<task-notification>` shell, so a client that projects them as the generic completion card shows "background task finished" while the model read a colleague's sentence — two faces describing different events. The discriminator is pinned to the **typed carrier being present**, never to the `summary` text: those carriers can only be minted by the engine's injection legs (the external `notify()` input is a strict subset of the payload and can wear none of them), while `summary` is filled by every notification there is — so anchoring on text would let any background task impersonate a colleague's message by writing `<agent-message from="…">` into its own summary, and a positive control asserts exactly that payload still projects as the generic card. Fail-closed has two tiers rather than one: a broken **required** field (empty `from`, a non-string `body`, a notice `kind` outside the closed set) returns absence so the caller falls back to the generic card — an honest downgrade where the user still sees the notification — while a broken **optional** field drops only itself, because losing an attribution note and losing a colleague's whole message are not the same magnitude. The provenance side record is **required and must agree on four points** (`kind` matches the lane; `from`/`taskId`/`seq` are present and equal the carrier/payload — each equality is anchored on a core mint site and pinned by the cli wire-anchor A-K24), so a carrier signed with a trusted name but a disagreeing provenance falls back to the generic card; peer bodies pass the same authority-envelope neutralization core applies (`<task-notification>` etc. are defused) so a colleague's text can never seed the resume dedup ledger. Lane precedence copies the engine renderer's own order, because the model already read the frame in that order and a client ordering of its own would put a card on screen that disagrees with the frame the model saw. Rendering and parsing of the transcript line live in the same module and are round-tripped in both directions, including a body carrying a forged closing tag (a parser fooled there hands half a message to the next row) and a quote inside the sender label (which must not forge a second attribute); the notice lane is deliberately kept **out** of the parser, since recognising it would mean anchoring the `[Cross-session …]` prefix and a user typing that same line would be rendered as engine speech. Hostile carriers are read as own **data** descriptors only and accessors are never invoked at all — `catch` catches throwing, not never returning — proven by a counting getter that must stay at zero calls, alongside a revoked proxy and a prototype-only carrier; and four legacy payload shapes assert the no-carrier path is byte-identical to before, which is the executable form of "zero difference for an older host". A re-supplied cross-session message — same task id, status and sequence as the first delivery, handed to the model again after compaction — is not rendered a second time, because the first delivery is still on the user's screen; a record with the next sequence number still renders |
378
378
  | `scripts/run-wiring-manifest-projection-test.mjs` | The two end-user facts carried on the engine's `wiring_manifest` frame (`modelGate`: which tools this run's model gate removed and the verbatim restore hint; `autoMode`: whether auto mode is actually armed and the engine's own reason word). Projection: both sections ride as `_sema_`-prefixed superset keys, verbatim, and no SDK-named key is minted; a frame where neither section is well-formed projects to `none/not_in_slice` (no empty arm); `modelGate` needs all three keys and treats `removed: []` as a bad value rather than a reading; `autoMode` needs a boolean plus a non-empty reason that agrees with it, and the reason word is never mapped onto the capabilities vocabulary; the frame is flat (a nested `manifest:{}` wrapper is not a supply); `eventId` rides like every other arm. Adapter: exactly one chrome event on the main lane, a sub-flow frame (any `parentToolCallId`, `null` included) yields nothing, and an absent `eventId` leaves the key absent. Added at receiving time because the shell-side gate could not see this package's behaviour: two mutations (empty `removed` accepted, sub-flow gate removed) had passed the package suite untouched 0.71.0 adds sections F–I: the fourth/fifth/sixth manifest sections (`tools` via the roster reader, `hooks[]` rows dropped one by one when malformed, `lsp` absent unless `mounted` is a boolean), the `tool_roster_delta` arm (narrowed `delta`, `malformed` when `fromDigest`/`roster` cannot be read, host applies it against its own digest), the `context_usage` arm (finite-gated scalars plus `sections[]` rows dropped one by one), and the `WiringManifestMcpEntryView` rename with `MAX_AGENT_SKILLS` gone from the surface |
379
379
  | `scripts/run-submit-wiring-manifest-test.mjs` | The non-streaming submit receipt can carry the run's opening wiring manifest (`TaskResult.wiringManifest`, additive on newer servers). `readSubmitWiringManifest` answers one of three: the key is absent on the receipt itself (older server, or a deployment whose engine never produced that frame) — not the same as unreadable; the key is present but cannot be read (not an object, or none of the nine sections survive); or a manifest view. The view is the same shape the streaming lane's chrome event carries (minus its two envelope keys) and is assembled by the same code path, so both lanes agree byte for byte on the same object. Liveness fields ride through untouched — this reader never mints a liveness verdict — and an operator-shaped receipt with extra governance sections reads to the same view as a tenant-shaped one. A zero-tool roster is a real reading, not an absence. |
380
380
  | `scripts/run-rule-offers-reader-test.mjs` | The narrowing reader behind the "don't ask again" options, now a public entry point rather than a card-port-only one. Hosts that render the frame themselves (a browser has no three-way terminal card) previously had to rebuild this reader on their side, and what it carries is a **redemption-safety** judgement, not a convenience: the batch arm is redeemed by **index**, so a reader that compacts the array after dropping a malformed entry makes the k-th option a person clicked and the k-th rule the server writes two different rules. So: a bad entry is dropped **on its own** (one bad option must not make a real one disappear) while every surviving entry keeps its **original wire index** — pinned from both ends, with the bad entries leading and trailing. A batch's *members* are the opposite: any malformed member drops the whole batch, because a conjunctive batch is one "yes" to all of them and a batch missing a member is a different grant; its honest-remainder count is a reading, not decoration, so a non-integer or negative value drops the batch rather than rendering a fabricated zero. An empty array, a non-array, an over-cap array and an all-bad array all read as **absence** rather than an empty list, because an empty list renders as "there is an option lane with nothing in it". The two wire generations are ordered by a rule, not a preference: the newer key wins outright, a newer key that is **present but unreadable** does not fall back to the retired key (borrowing the older material would pass someone else's options off as this request's), and a `null` newer key reads as absence so a relaying layer that serialises "missing" as null cannot delete the whole lane on older engines. The public entry is finally reconciled against **both** card-port legs on the same material, byte for byte, so the exported reader and the one the card sees can never become two. Two upstream vocabularies used to be **hand-copied** here, and both had fallen behind: a match word outside the copied pair dropped an otherwise valid option outright, and a batch carrying a directory-read member — a member kind the copy did not know — dropped the whole batch. Both tables now come from one place upstream and are re-exported verbatim, pinned in both directions: every word in the table must be accepted (a narrower copy reds on the words it never learned) and a word constructed to be outside it must still be refused (a reader widened to "any string" reds too), with the retired-key normalising leg sharing the same narrowing so the fix cannot land on one leg only. A member whose kind is genuinely unknown still drops **the whole batch and only that batch** — never one member, because a conjunctive batch one member short renders "yes to N" as "yes to N−1", and never the card, because the honest single beside it is intact — while a member from before the discriminant existed normalises to the historical kind rather than being refused. The additive per-segment reasons ride through verbatim, drop only the row that is malformed, and stay **absent rather than empty** when nothing survives, since an empty list would read as "confirmed nothing uncovered" while the count remains the only source of truth |
@@ -419,6 +419,7 @@ guard still cross-checks the table by name).
419
419
  | `scripts/run-registry-quota-usage-test.mjs` | The projection of the cloud control plane's quota reading, and the three different things a missing number can mean there. This response says `null` in two places and means something different each time: no token quota is configured for this principal on this instance, and this window has no cap at all. Both are **facts the server is asserting**, not gaps in the reading — while a key that is absent or carries the wrong type is a genuine gap. All three have to survive to the screen separately, because folding them is how a user ends up staring at a confident `0`: an uncapped window rendered as if nothing were left, or a deployment that simply never configured quotas rendered as if the quota were exhausted. The complaint that started this was the opposite direction — a centrally configured quota that the command line could not see at all — so the reading also refuses to let an unreadable response masquerade as "no quota configured". Two fields deliberately do not share one signal: whether a window is exhausted and whether there is a recovery time, since the recovery time is only ever populated in the exhausted case and reading its absence as "not exhausted" would answer a question the response never answered. Counts that the server always provides are narrowed no further than the mint: a used counter has no uncapped state, so a null there is unreadable rather than zero. The wording helper carries the only human-facing phrasing, and the sentences for "no cap" and "unknown" are checked to contain no digits at all |
420
420
  | `scripts/run-file-history-capture-capability-test.mjs` | The engine's file-history-capture self-description (`capabilities.fileHistoryCapture`), read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `off`), words are taken as an open set so a newer mode is not mistaken for a malformed answer, `fileHistoryCaptureMode` recognises only `off` and `on-always`, and the wording for `off` speaks about capture only — whether code can be rewound is left to the rewind readings. |
421
421
  | `scripts/run-model-identity-resolvability-test.mjs` | The model-identity judgement a client makes before letting anyone in: can the engine it is about to use start with a model name? Each end reports what it read from each place that can feed a model name to a local engine (complete, partial — a gateway address or a credential but no model name —, absent, or unreadable), or, for an engine that runs elsewhere, whether that engine has been seen answering; `modelIdentityResolvability` answers resolvable, unresolvable or unknown. Having part of an upstream configuration is not having enough of one, so partial lanes never add up to resolvable; a lane that was not reported or could not be read makes the answer unknown rather than unresolvable; an engine that runs elsewhere is never judged unresolvable and local lanes are never consulted for it (it does not start without a model name, so seeing it answer is enough to call it resolvable). `modelSetupDecision` combines that answer with whether this end can configure a model at all: setup is offered only for unresolvable on an end that can configure one, an end that cannot says so and points at whoever runs the engine, and unknown never opens setup. The detail and notice sentences are checked to be pairwise distinct, unknown sentences neither claim a model is configured nor that it is not, and the module is checked to import no platform I/O |
422
+ | `scripts/run-cloud-effective-projection-test.mjs` | The cloud control plane's effective-configuration response beyond its four configuration domains, and what a locally started engine does with the models document derived from it. Three top-level keys are read with the meaning their producer gives them: `warnings` (degradation warnings from the build that produced the served view — an empty list is a clean build, an absent key is an older server that cannot tell), and `budget` / `runtimeCaps`, where `null` means two different things: nothing resolves for this principal when you view yourself, and values withheld when the response previews another principal. A missing key or a wrong type is a third state, unknown, and none of the three is ever folded into a zero, a `false` or "no budget". A malformed warning row, budget field or cap costs only itself, and a known budget field of the wrong type is named as not shown rather than silently read as "no limit on this axis"; warning kinds are an open set, so a kind this client does not recognise still produces a warning line. The budget and cap readers are reconciled against the installed settings schema. The single wording source puts degradation warnings first and keeps every "not set / withheld / unknown" sentence free of digits, while a zero the server really sent is shown as a zero. On the models side, when the host injects an entry check the models document carries the default model, the tier groups and the active tier group (without the check none of the three is written), and every catalog reference the local engine's schema would reject — a default, role, @-mention entry, tier binding or active group that names something outside the catalog served to this principal — is dropped and recorded, because one dangling reference makes the local engine discard the whole models domain and fall back to its environment catalog; this is proven by reading the produced document with the installed file store. An @-mention allowlist that would be pruned to empty is kept as sent, since an empty list means "everything may be mentioned"; that case is recorded, produces its own warning that a locally started engine will reject the cloud model settings and use its environment catalog instead, and the gate reads the document with the installed file store to confirm exactly that outcome, so the sentence turns red the day the local reader becomes lenient. A per-model budget in which no field could be read is never described as having no limits. Registry annotation keys on model entries (`origin`, `overridesTeam`) are removed before the models document is written: the local engine's schema does not accept them, and a configuration refresh would otherwise be rejected as a whole. A model entry the host-injected entry check rejects is left out of the document and references to it are dropped: the local engine drops such an entry at startup, but a refresh rejects the whole configuration over it, so the gate requires a clean read of the produced document; when no entry passes the check, the catalog is kept as sent and gets its own warning, which the gate proves by reading the document back. The check receives a copy, so it cannot alter what is written. Tier words outside the local schema's closed set are dropped and recorded as unsupported, and the package's tier word list is reconciled against the installed schema in both directions; an active tier group is judged against group names, never model names. |
422
423
 
423
424
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
424
425
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
@@ -80,6 +80,7 @@ export type SelfHealOutcome = {
80
80
  taskId: string;
81
81
  decidePath: string | null;
82
82
  decisionInFlight?: true;
83
+ pendingRowGone?: true;
83
84
  } | {
84
85
  kind: 'not-parked';
85
86
  taskId: string;
@@ -101,6 +102,7 @@ export type SelfHealOutcome = {
101
102
  taskId: string;
102
103
  decidePath: string | null;
103
104
  decisionInFlight?: true;
105
+ pendingRowGone?: true;
104
106
  } | {
105
107
  kind: 'ask-decided-without-card';
106
108
  taskId: string;
@@ -200,6 +202,8 @@ export declare function attemptActiveRunSelfHeal(signal: ActiveRunBusySignal, ru
200
202
  export interface ActiveRunSelfHealCopy {
201
203
  wayOut?: string;
202
204
  freshSession?: string;
205
+ keepSessionWayOut?: () => string | null;
206
+ engineDecidePath?: boolean;
203
207
  rowFor?: (outcome: SelfHealOutcome, signal: ActiveRunBusySignal | null | undefined) => string | undefined;
204
208
  headlessRowFor?: (signal: ActiveRunBusySignal) => string | undefined;
205
209
  }
@@ -310,6 +310,7 @@ async function planReviewArm(taskId, signal, deps) {
310
310
  taskId,
311
311
  decidePath: signal.pendingGate?.decidePath ?? null,
312
312
  ...(reopenRefusedForDecisionInFlight(verdict) ? { decisionInFlight: true } : {}),
313
+ ...(verdict.pendingRowGone === true ? { pendingRowGone: true } : {}),
313
314
  };
314
315
  }
315
316
  async function askParkArm(taskId, signal, runs, deps) {
@@ -343,6 +344,7 @@ async function askParkArm(taskId, signal, runs, deps) {
343
344
  taskId,
344
345
  decidePath: signal.pendingGate?.decidePath ?? null,
345
346
  ...(inFlight ? { decisionInFlight: true } : {}),
347
+ ...(verdict.pendingRowGone === true ? { pendingRowGone: true } : {}),
346
348
  };
347
349
  }
348
350
  async function probeRunStatusOnce(taskId, durable, durableGet, deps) {
@@ -366,6 +368,12 @@ async function staleParkArm(taskId, busy, runs, deps) {
366
368
  kind: 'ask-reopen-failed',
367
369
  taskId,
368
370
  decidePath: busy.pendingGate?.decidePath ?? null,
371
+ pendingRowGone: true,
372
+ };
373
+ const reopenFailedAfterCard = {
374
+ kind: 'ask-reopen-failed',
375
+ taskId,
376
+ decidePath: busy.pendingGate?.decidePath ?? null,
369
377
  };
370
378
  if (runs === undefined || typeof runs.get !== 'function')
371
379
  return reopenFailed;
@@ -406,24 +414,25 @@ async function staleParkArm(taskId, busy, runs, deps) {
406
414
  choice = null;
407
415
  }
408
416
  if (choice === null)
409
- return reopenFailed;
417
+ return reopenFailedAfterCard;
410
418
  if (choice === 'wait') {
411
419
  rememberDeclinedOffer(declineKey);
412
420
  return { kind: 'stale-park-wait', taskId, status: fresh };
413
421
  }
414
422
  runningChoiceDeclined.delete(declineKey);
415
423
  if (callerAborted())
416
- return reopenFailed;
424
+ return reopenFailedAfterCard;
417
425
  const afterCard = await probeRunStatusOnce(taskId, runs, durableGet, deps);
418
426
  if (afterCard.ghost)
419
427
  return { kind: 'ask-run-not-found', taskId };
420
428
  if (callerAborted())
421
- return reopenFailed;
422
- if (afterCard.status === null || !ASK_PARK_STATES.includes(afterCard.status))
423
- return reopenFailed;
429
+ return reopenFailedAfterCard;
430
+ if (afterCard.status === null || !ASK_PARK_STATES.includes(afterCard.status)) {
431
+ return afterCard.status !== null && !CLAIM_HELD_STATES.includes(afterCard.status) ? reopenFailed : reopenFailedAfterCard;
432
+ }
424
433
  const stillGone = await readOwnedPendingCount(recheckOwnedPending, deps);
425
434
  if (stillGone !== 0)
426
- return reopenFailed;
435
+ return reopenFailedAfterCard;
427
436
  const verdict = await cancelAndConfirmRelease(taskId, runs, durableGet, deps);
428
437
  if (verdict.outcome === 'released')
429
438
  return { kind: 'stale-park-cancelled', taskId };
@@ -605,12 +614,19 @@ export function activeRunSelfHealRow(outcome, signal, copy, origin, follow) {
605
614
  const following = follow?.following === true && steerFollowIntent(outcome)?.kind === 'tail';
606
615
  if (origin === 'injected')
607
616
  return injectedSubmissionRow(outcome, following);
608
- const base = activeRunSelfHealBaseRow(outcome, signal, copy?.wayOut ?? DEFAULT_WAY_OUT, following);
617
+ const engineDecide = copy?.engineDecidePath !== false;
618
+ const keep = (outcome.kind === 'plan-review-reopen-failed' || outcome.kind === 'ask-reopen-failed') &&
619
+ outcome.decisionInFlight !== true &&
620
+ outcome.pendingRowGone !== true
621
+ ? keepSessionWayOutOf(copy)
622
+ : null;
623
+ const base = activeRunSelfHealBaseRow(outcome, signal, copy?.wayOut ?? DEFAULT_WAY_OUT, following, keep, engineDecide);
624
+ const viaEngine = engineDecide ? decidePathClause(signal) : '';
609
625
  switch (outcome.kind) {
610
626
  case 'not-parked':
611
- return base + lastActivityClause(signal) + decidePathClause(signal) + governanceOriginClause(signal);
627
+ return base + lastActivityClause(signal) + viaEngine + governanceOriginClause(signal);
612
628
  case 'state-unknown':
613
- return base + decidePathClause(signal) + governanceOriginClause(signal);
629
+ return base + viaEngine + governanceOriginClause(signal);
614
630
  default:
615
631
  return base + governanceOriginClause(signal);
616
632
  }
@@ -651,7 +667,19 @@ function decisionInFlightRow(parkedOn) {
651
667
  `sema did not show the card again — answering it a second time would send a second decision. It did NOT ` +
652
668
  `cancel the run. Your message was NOT sent; send it again in a moment, once the engine has applied that decision.`);
653
669
  }
654
- function activeRunSelfHealBaseRow(outcome, signal, wayOut, following = false) {
670
+ function keepSessionWayOutOf(copy) {
671
+ if (typeof copy?.keepSessionWayOut !== 'function')
672
+ return null;
673
+ try {
674
+ const v = copy.keepSessionWayOut();
675
+ const t = typeof v === 'string' ? v.trim().replace(/[.;。]+$/, '').trim() : '';
676
+ return t !== '' ? t : null;
677
+ }
678
+ catch {
679
+ return null;
680
+ }
681
+ }
682
+ function activeRunSelfHealBaseRow(outcome, signal, wayOut, following = false, keepSessionWayOut = null, engineDecide = true) {
655
683
  switch (outcome.kind) {
656
684
  case 'decision-pending':
657
685
  return (`This session is held by an earlier turn${outcome.taskId ? ` (run ${outcome.taskId})` : ''} that ` +
@@ -690,22 +718,26 @@ function activeRunSelfHealBaseRow(outcome, signal, wayOut, following = false) {
690
718
  case 'ask-reopen-failed': {
691
719
  if (outcome.decisionInFlight === true)
692
720
  return decisionInFlightRow(`waiting for your decision (run ${outcome.taskId})`);
693
- const viaEngine = outcome.decidePath
721
+ const viaEngine = engineDecide && outcome.decidePath
694
722
  ? ` You can also decide it on the engine directly: POST ${outcome.decidePath}.`
695
723
  : '';
696
724
  return (`The previous turn is parked waiting for your decision (run ${outcome.taskId}) and sema could ` +
697
725
  `not reopen that card here. It did NOT cancel the run — that would have decided it for you. ` +
698
- `Your message was NOT sent; ${wayOut} if you no longer want it.${viaEngine}`);
726
+ (keepSessionWayOut !== null
727
+ ? `Your message was NOT sent. To keep this conversation, ${keepSessionWayOut}: the card comes back when the session resumes, if it is still waiting. Or ${wayOut} if you no longer want it.${viaEngine}`
728
+ : `Your message was NOT sent; ${wayOut} if you no longer want it.${viaEngine}`));
699
729
  }
700
730
  case 'plan-review-reopen-failed': {
701
731
  if (outcome.decisionInFlight === true)
702
732
  return decisionInFlightRow(`on a plan review (run ${outcome.taskId})`);
703
- const viaEngine = outcome.decidePath
733
+ const viaEngine = engineDecide && outcome.decidePath
704
734
  ? ` You can also decide it on the engine directly: POST ${outcome.decidePath}.`
705
735
  : '';
706
736
  return (`The previous turn is parked waiting for a plan review (run ${outcome.taskId}) and sema could not ` +
707
737
  `reopen that approval card. It did NOT cancel the run — that would have discarded the plan for ` +
708
- `you. Your message was NOT sent; ${wayOut} if you no longer want that plan.${viaEngine}`);
738
+ (keepSessionWayOut !== null
739
+ ? `you. Your message was NOT sent. To keep this conversation, ${keepSessionWayOut}: the card comes back when the session resumes, if it is still waiting. Or ${wayOut} if you no longer want that plan.${viaEngine}`
740
+ : `you. Your message was NOT sent; ${wayOut} if you no longer want that plan.${viaEngine}`));
709
741
  }
710
742
  case 'running-steered': {
711
743
  const handle = `run ${outcome.taskId}`;
@@ -16,7 +16,17 @@ export interface CloudMcpProjection {
16
16
  }
17
17
  export declare function cloudMcpToSpecs(servers: unknown, env: EnvLike): CloudMcpProjection;
18
18
  export declare function cloudSkillsToSpecs(manifestSkills: unknown, resolveBody: (contentHash: string) => string | null): SkillSpec[];
19
- export declare function cloudModelsToDoc(models: unknown): Rec | null;
19
+ export interface CloudModelRefIssue {
20
+ readonly key: 'models' | 'default' | 'roles' | 'atModelAllowlist' | 'tierGroups' | 'activeTierGroup';
21
+ readonly at?: string;
22
+ readonly target?: string;
23
+ readonly reason: 'dangling' | 'rejected' | 'unreadable' | 'unsupported' | 'duplicate';
24
+ readonly outcome: 'dropped' | 'kept';
25
+ }
26
+ export interface CloudModelsDocOptions {
27
+ readonly entryAccepted?: (entry: object) => boolean;
28
+ }
29
+ export declare function cloudModelsToDoc(models: unknown, opts?: CloudModelsDocOptions): Rec | null;
20
30
  export interface CloudDroppedModel {
21
31
  id?: string;
22
32
  name?: string;
@@ -35,6 +45,62 @@ export interface CloudExecutionProjection {
35
45
  }
36
46
  export declare function cloudExecutionToPolicy(execution: unknown): CloudExecutionProjection;
37
47
  export declare function cloudRosterNames(rostersValue: unknown): string[];
48
+ export interface CloudEffectiveWarning {
49
+ readonly kind: string;
50
+ readonly domain?: string;
51
+ readonly droppedNames?: readonly string[];
52
+ readonly keys?: readonly string[];
53
+ readonly path?: string;
54
+ readonly file?: string;
55
+ readonly why?: string;
56
+ readonly truncatedFrom?: number;
57
+ }
58
+ export type CloudEffectiveWarningsReading = {
59
+ readonly kind: 'reported';
60
+ readonly warnings: readonly CloudEffectiveWarning[];
61
+ readonly unreadable: number;
62
+ } | {
63
+ readonly kind: 'unknown';
64
+ };
65
+ export interface CloudBudgetModelLimits {
66
+ readonly maxBudgetUsd?: number;
67
+ readonly tpmLimit?: number;
68
+ readonly rpmLimit?: number;
69
+ }
70
+ export interface CloudBudgetLimits {
71
+ readonly mode?: string;
72
+ readonly maxBudgetUsd?: number;
73
+ readonly budgetDuration?: string;
74
+ readonly tpmLimit?: number;
75
+ readonly rpmLimit?: number;
76
+ readonly maxParallelRequests?: number;
77
+ readonly perModel?: Readonly<Record<string, CloudBudgetModelLimits>>;
78
+ readonly maxRunCostUsd?: number;
79
+ readonly maxIterations?: number;
80
+ readonly enforcement?: string;
81
+ }
82
+ export type CloudBudgetReading = {
83
+ readonly kind: 'configured';
84
+ readonly limits: CloudBudgetLimits;
85
+ readonly uninterpretedKeys: readonly string[];
86
+ } | {
87
+ readonly kind: 'not-configured';
88
+ } | {
89
+ readonly kind: 'redacted';
90
+ } | {
91
+ readonly kind: 'unknown';
92
+ };
93
+ export type CloudRuntimeCapsReading = {
94
+ readonly kind: 'configured';
95
+ readonly caps: Readonly<Record<string, boolean>>;
96
+ readonly uninterpretedKeys: readonly string[];
97
+ } | {
98
+ readonly kind: 'not-configured';
99
+ } | {
100
+ readonly kind: 'redacted';
101
+ } | {
102
+ readonly kind: 'unknown';
103
+ };
38
104
  export interface CloudProjection {
39
105
  mcp: McpServerSpec[];
40
106
  skills: SkillSpec[];
@@ -47,5 +113,14 @@ export interface CloudProjection {
47
113
  droppedMcpServers: CloudMcpDroppedServer[];
48
114
  droppedModels: CloudDroppedModel[];
49
115
  execution: CloudExecutionProjection;
116
+ effectiveWarnings: CloudEffectiveWarningsReading;
117
+ budget: CloudBudgetReading;
118
+ runtimeCaps: CloudRuntimeCapsReading;
119
+ modelRefIssues: CloudModelRefIssue[];
120
+ }
121
+ export declare function projectEffectiveBody(body: unknown, skillBodies: Record<string, string>, env: EnvLike, rostersValue?: unknown, opts?: CloudModelsDocOptions): CloudProjection;
122
+ export interface CloudEffectiveNotice {
123
+ readonly level: 'warn' | 'info';
124
+ readonly text: string;
50
125
  }
51
- export declare function projectEffectiveBody(body: unknown, skillBodies: Record<string, string>, env: EnvLike, rostersValue?: unknown): CloudProjection;
126
+ export declare function cloudEffectiveNotices(p: Pick<CloudProjection, 'effectiveWarnings' | 'budget' | 'runtimeCaps' | 'modelRefIssues'>): CloudEffectiveNotice[];