@sema-agent/client-core 0.79.1 → 0.80.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -49,6 +49,36 @@
49
49
  > 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
50
50
  > commit 漏转在发布前就红,不再靠人记。
51
51
 
52
+ ## 0.80.1(2026-09-22)
53
+
54
+ ### Changed
55
+ - **去键重发判据的措辞按「铸」计**(CC-143):0.80.0 接入文档 §87 把去键重发写成「带键的体恰一次」,而每次投递内部的瞬态传输重试会把**同一只已铸的体**再投一次 —— 首发(带键)投递上断网时按字面会判红。权威读法自本版成文:**带键的体被铸恰一次(POST 几次随传输层)/ 去键的体被铸恰一次 / 绝无第三种体**;实现一个字节未改,只改判据句。§87 已发段不回改,读法以本段与 §88 为准。
56
+ - **模型预设表 `deepseek` 项改上游闭集真名**(CC-144):`deepseek-v4-flash` 是别名(打它回包 `model: deepseek-flash`),预设项 id 改为 `deepseek-flash`,其余字段不变;端若钉了别名字面请换真名。
57
+ - **接入文档段落自带射程限定词**(CC-143):自 §88 起,每条可验判据句旁边写明它在什么条件下不成立(时间射程:按哪一版的读法;车道射程:只覆盖哪条腿),形式为「(本节按 x.y.z 的读法;若与发车帖不一致以发车帖为准)」或「(某某部署形;另一条腿另有值)」;已发段不回改。
58
+
59
+ ## 0.80.0(2026-09-22)
60
+
61
+ > 主题:**策略拒的归因上 wire + 删除规则后果句归包 + `/mcp` 详情卡两腿分说归包**(🔴 **minor**:peer sdk 地板 `>=11.2.1` → **`>=11.3.0`**;根公面运行期导出 1184 → **1195**;型面 additive)。一次策略自动拒从此在 wire 上说得出自己是机器拒的,读回时归进 CC 的规则桶;三端各自铸过的两处壳侧措辞与判据(删除规则确认卡的后果句、`/mcp` 详情卡的两腿分说)收进本包,并同批改掉壳侧带过来的八处说满的话。
62
+
63
+ ### 🔴 BREAKING(地板 + 一处宣告过的删除)
64
+
65
+ - **终帧旧拼法 `structuredOutput` 删除**(0.79.1 CHANGELOG 与接入文档 §86 宣告的过渡窗只有 0.79.1 一版):成功终帧只铸 CC 形 `structured_output`,值 / 引用 / 缺席语义与 0.79.1 逐字不变;按旧键读的端会读到 `undefined`,而 `undefined` 在这条语义上表示「引擎没产出」—— 换钉前先改读法。
66
+ - **peer `@sema-agent/sdk` 地板 `>=11.2.1` → `>=11.3.0`**:11.3.0 在耐久 `decide` 与活卡 `respond` 两条回决口体上声明 `settledBy?: "policy"`、在 `Settlement.kind` 上声明第十三词 `policy_refused`、把三错误码写进 spec;本包对它们编译,老 sdk 不再通过类型检查。地板由门见证:门从装着的声明文件里逐条读出那三处具名位(不是信版本号)。开发钉 core `~7.26.3`(结算词表对账基准换 core 真字节,sdk 降为第二见证)。
67
+
68
+ ### Added
69
+
70
+ - **一次拒绝可以说自己是机器拒的,不是人拒的**(CC-133):宿主把一次 deny 归因给这台部署自己的策略(权限规则 / 无交互姿态 / 根本没有审批口的车道)时,这个归因随决断一起上 wire,耐久腿与活卡腿同律;引擎据它把拒绝措辞成策略的,而不是在一条从没问过人的 run 上告诉模型「用户不想继续」。归因给人的、以及没有归因的,**什么都不发**——wire 上没有第二个词来说「人拒的」,缺席就是那句话。allow 族永不带它:那一形在裁决之前被拒,而放行的副作用已经落地。两处读数各在第一个 `await` 前快照一次,发送 / 回交 / 失败分类三处同用。
71
+ - **同一个区分也读得回来**(CC-133):结算词说是策略答的 ⇒ 转录上两只载体都归进策略桶(`permission-rule`);说是人拒的 ⇒ 人的桶(`user-rejected`);其余十一词整键缺席不猜;引擎写下的「哪一层拒的」原样不动。宿主自己报过归因的以宿主为准 —— 一条重放帧不许改宿主说过的话(同一次调用跨帧合流也按这一律:后到的重放词不推翻宿主报过的归因,先到的引擎词被后到的宿主归因替换;只有同权威来源的异值才判成矛盾一起缺席)。新导出:`SETTLEMENT_KIND_WORDS`(十三词镜像,编译期双向钉)/ `isPolicyRefusedGate` / `ccToolDenialKindForToolEnd` / `ccToolDenialKindForSettledBy`。
72
+ - **三条与归因键有关的拒绝有了读口与人话**(CC-133;`denyAttributionRefusalFromError` / `denyAttributionRefusalContent`):签决断体但不签这一键的部署、决断与归因互相矛盾的体、引擎不收的结算事实。三条都发生在裁决之前,每句都明说「没落定、卡还在等」并各指一个不同的下一步 —— 没有一句是「原样再发」。第三句刻意不断言「那个词拼错了」(引擎在同一格上有几种不相关的因由),只点名它不收的那一格。两个码在别处另有含义:一个只在调用方声明真发过这一键时认领,另一个只在引擎点名两个席位之一时认领;`field` 从 `APIError.extra` 读。
73
+ - **签决断体的部署不再因为这一键丢掉那次 deny**(CC-133):那种部署不签归因键,又没有能力位可探,首发可能恰因此被拒。同一决断随即**去掉那一键再发恰一次**(其余逐字节相同),deny 照常落定;结局仍报宿主说的归因,另带 present-iff 的 `denyAttributionDropped: true` 说明那个词没到引擎。只对那一个码、只在本次真带过这一键、调用方未中止时;再失败按既有分类交回,不做第二次去键重发(每次投递内部的瞬态传输重试照旧,与首发一视同仁)。
74
+ - **删除权限规则的后果句归包**(CC-132):三端「删除这条规则」确认卡此前各自铸一份「这会怎样」的措辞,新导出 `RuleRemovalBehavior` / `removalConsequenceLineForBehavior` / `ruleRemovalBehaviorOf` 把三态(allow / deny / ask)各一句 + 缺席一句(不回落 allow)收进本包唯一措辞铸点;deny 那一句点名 widens(删一条 deny 规则的真实后果是放宽不是收紧)。读口按运行时闭集收,不信静态型,零属性访问。四句文案与端侧现有实现逐字节相同。
75
+ - **`/mcp` 详情卡两腿分说归包**(CC-134):新导出 `mcpEngineLegLivenessOf`(一台 server 在引擎腿上的活性**六档**:没有读得完的名册 / 引擎申报零台 / 同名多行 / 送来一份读不懂的 / 行上没说活性 / 一次观察带开集原词)、`mcpDetailLegNote`(Status 底下那一句范围说明的唯一铸点,**九句**:活性五句 + 名册四句,活性观测优先于名册,不渲活性原词)、`isMcpLivenessState`(活性三词闭集的唯一判据上公面)。三条不许折的分野:没有一张读得完的名册 ≠ 一张空名册;读不出 ≠ 缺席(活性格与名册数各自如此);歧义只许答歧义。只认自有键、每格恰读一次、永不抛,谎报长度的名册不扫。壳侧带过来的三句说满的话同批改口:引擎看过但说不清 / 表外词 / 送来过读不懂的记录各自成句,「这一会话里引擎什么都没说过」改成只说本端手上有什么。
76
+
77
+ ### Fixed
78
+
79
+ - **拒绝归因台账的溢出闩不再连坐矛盾检查**(0.79.0 CC-125 存量):闩落下之后同一 id 再出现相反的结算词,此前不再出表、终帧保留过时值;现在矛盾出表挪出闩(它只往外删,撑不破上限),往里加那一半照旧受闩管。
80
+ - **`ApprovalCardDenyDecision.settledBy` 的公开类型注释**此前仍说「本位不上 wire」;0.80.0 起它同时上 wire 与决定转录那一戳,注释改口(它随声明文件出包给三端读)。
81
+
52
82
  ## 0.79.1(2026-09-21)
53
83
 
54
84
  > 主题:**终帧结构化产出改用 CC 拼法 + 写审批退化卡拿回路径 + 云控制面子路径入口**(patch;peer 地板不动 `>=11.2.1`;根公面运行期导出 1184 不变)。三件:终帧结构化产出同帧铸 CC 形 `structured_output`(旧拼法 `structuredOutput` 过渡一版,0.80.0 删);写审批 `argsOmitted` 退化卡在没有受护写入的部署上重新解得出路径;云控制面十值二型经本包**第二个入口** `@sema-agent/client-core/registry` 原样转口(根入口零变,不用它的端零升级动作)。
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.79.1
38
+ **Version:** 0.80.1
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -67,7 +67,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
67
67
  against — the tables live upstream precisely so this package does not keep a second copy that can
68
68
  fall behind. The browser bundle really bundles the SDK through (the portability guard would
69
69
  exit 3 rather than quietly mark it external).
70
- - The declared floor is `>=11.2.1` (raised from `>=11.0.1` in 0.79.0: `capabilities.mcpProbe`, the MCP probe face (`mcpCapabilities` / `probeMcp` with `McpProbeFace`) and the write receipt's third liveness arm (`stillLive: "unknown"`) are typed from 11.2.x on, 11.2.1 adds `TaskRequest.approverPosture` and the `mandated` approval-frame key to the types and the runtime key anchor, the package now compiles against 11.2.1, and no consumer ships 11.0.x / 11.1.x any more, so the older floor lost its witness; before that raised from `>=9.8.1` in 0.78.0: `DeniedBy` carries its tenth word `read_boundary`, `rules.write` answers a `stillLive`-discriminated body, `RemovalLiveness` / `RuleWriteRequest` / `RuleWriteResult` / `RuleWriteBehavior` are exported from the SDK root and `TaskRequest.excludeAllTools` is typed from 11.x on, the package now compiles against 11.0.1, and no consumer ships 9.8.x any more, so the older floor lost its witness; before that raised from `>=9.7.1` in 0.75.0: `Capabilities.deviceExecutor.management` is typed from 9.8.x on, the package now compiles against 9.8.1, and no consumer ships 9.7.x any more, so the older floor lost its witness; before that raised from `>=9.6.0` in 0.74.0: `Capabilities.approvalsStreamLive` / `.executionLane`, `LivePendingRow.frame`, the `live_*` approval-stream events and `gates[].toolCallId` are typed from 9.7.x on, and no consumer ships 9.6.0 any more, so the older floor lost its witness; before that raised from `>=9.4.0` in 0.71.0: the `tool_disclosure` / `tool_progress` frames and `ToolApprovalFrame.readRootCandidate` are typed there; earlier: raised from `>=8.8.0` in 0.69.0: the `reasoning_end` frame and `McpStatusPanel.lastLegMcp` are typed from 9.4.0 on), and it is *witnessed*: the guard checks that an actually
70
+ - The declared floor is `>=11.3.0` (raised from `>=11.2.1` in 0.80.0: a deny decision may now name its settler — `settledBy: "policy"` is typed on both the durable decide body and the live respond body from 11.3.0 on, and the gate record's settlement vocabulary carries its thirteenth word `policy_refused`, which this package reads to place a refusal in the policy bucket rather than the person's; the package now compiles against 11.3.0 and no consumer ships 11.2.x any more, so the older floor lost its witness; before that raised from `>=11.0.1` in 0.79.0: `capabilities.mcpProbe`, the MCP probe face (`mcpCapabilities` / `probeMcp` with `McpProbeFace`) and the write receipt's third liveness arm (`stillLive: "unknown"`) are typed from 11.2.x on, 11.2.1 adds `TaskRequest.approverPosture` and the `mandated` approval-frame key to the types and the runtime key anchor, the package now compiles against 11.2.1, and no consumer ships 11.0.x / 11.1.x any more, so the older floor lost its witness; before that raised from `>=9.8.1` in 0.78.0: `DeniedBy` carries its tenth word `read_boundary`, `rules.write` answers a `stillLive`-discriminated body, `RemovalLiveness` / `RuleWriteRequest` / `RuleWriteResult` / `RuleWriteBehavior` are exported from the SDK root and `TaskRequest.excludeAllTools` is typed from 11.x on, the package now compiles against 11.0.1, and no consumer ships 9.8.x any more, so the older floor lost its witness; before that raised from `>=9.7.1` in 0.75.0: `Capabilities.deviceExecutor.management` is typed from 9.8.x on, the package now compiles against 9.8.1, and no consumer ships 9.7.x any more, so the older floor lost its witness; before that raised from `>=9.6.0` in 0.74.0: `Capabilities.approvalsStreamLive` / `.executionLane`, `LivePendingRow.frame`, the `live_*` approval-stream events and `gates[].toolCallId` are typed from 9.7.x on, and no consumer ships 9.6.0 any more, so the older floor lost its witness; before that raised from `>=9.4.0` in 0.71.0: the `tool_disclosure` / `tool_progress` frames and `ToolApprovalFrame.readRootCandidate` are typed there; earlier: raised from `>=8.8.0` in 0.69.0: the `reasoning_end` frame and `McpStatusPanel.lastLegMcp` are typed from 9.4.0 on), and it is *witnessed*: the guard checks that an actually
71
71
  installed SDK at that line still exports every value-level symbol this package imports and still
72
72
  declares `TaskStats.costMicroUsd` (the key `costOrNull` reads). A floor nobody ever ran is a
73
73
  promise, not a contract.
@@ -296,21 +296,21 @@ guard still cross-checks the table by name).
296
296
  | `scripts/run-segment-authority-single-source-test.mjs` | The authoritative-segment replacement verdict, single-sourced. `text_end.content` and the `text_delta` stream stopped being byte-identical the day the engine started redacting the former through the same filter as the result, so every consumer now has to decide six ways what to do with the segment it has half-emitted — and until this release that decision existed **twice**: once here for the transcript lane, once in the shell for the print lane, hot-fixed a version apart. The verdict is now one pure function both lanes call, and the guard pins it on the quantity that actually decides the outcome: whether the authoritative text still *starts with* the bytes that already left, not whether a flush has happened — the latter is a precondition, and anchoring on it withholds a perfectly ordinary answer. Each of the six forms is checked with its counter-case, the prefix length is pinned to UTF-16 code units against a non-ASCII sample whose UTF-8 byte count differs (slicing by bytes leaves the very thing being redacted on screen), and the withheld-segment ledger is compared by normalised equality rather than substring, because a short redaction marker quoted in an unrelated later answer would otherwise suppress that answer entirely. The same file pins the session-level memory-capture declaration to one mint point — the wire value is a single-member closed set, and a consumer that spells it wrong gets a loud refusal rather than a silently dropped privacy request — and pins the SDK URL/health transit to be the **same function reference**, since wrapping it would discard the one guarantee the transit exists for. A last section strips comments with the TypeScript parser and asserts the second expression has not grown back |
297
297
  | `scripts/run-print-bash-iserror-test.mjs` | The print lane's Bash `is_error` authority (structured over regex). A second section pins where the denial classification word lands on this lane: on the message envelope, never inside the tool-result block, because that block is forwarded verbatim to the provider on compaction and a self-minted key there is the shape of an old, real defect. A word outside the upstream table — or an empty string, a non-string, or nothing at all — mints no key rather than a guess, and the word never moves the error flag, because attribution does not decide anything |
298
298
  | `scripts/run-bash-benign-exit-interpretation-test.mjs` | Benign non-zero Bash exits (`returnCodeInterpretation`) stay non-errors across all three derivation arms, and the annotation transits to the card |
299
- | `scripts/run-sdk-floor-test.mjs` | The SDK version floor — and, more to the point, that the *installed* type declarations still carry the keys this package reads |
299
+ | `scripts/run-sdk-floor-test.mjs` | The SDK version floor — and, more to the point, that the *installed* type declarations still carry the keys this package reads — including, from 0.80.0, the three declarations that justify the floor itself: the key naming who settled a refusal on both decision legs, and the thirteenth word in the settlement vocabulary. They are found through the syntax tree rather than by searching text, because this guard's own comment stripper blanks string contents and would have made that check permanently, silently green. |
300
300
  | `scripts/run-engine-caps-ledger-test.mjs` | A per-key disposition ledger for `GET /v1/capabilities`. The SDK's `Capabilities` grew from 74 keys to 93 in one release and nothing on the board could see it: this package consumes that table through four synchronous readers, and *nineteen new positions arriving while the package does not move* is exactly the disease shape this repo keeps logging on other axes — the fact is already on the wire, the package boundary is the cell that swallows it, and no client can read it however they write their side. So the ledger is reconciled **element-wise against the SDK interface in both directions**: a key the SDK added with no ledger row is red (someone must classify it), and a row for a key the SDK removed is red too (a registration that no longer does anything). Each row then has to survive its own claim — a `read` row names the source file, and the **code** there (comments stripped) must really mention the key, because prose asserting an alignment is the classic way these guards go hollow; a `not_read` row must have **zero** read sites in the tree, so wiring one up while the ledger still says the package ignores it is red rather than invisible. The census behind those two directions recognises five call shapes, each of which really occurs here — a reader whose base argument carries its own parentheses, a direct `caps.<key>`, a narrowing cast, an own-property read helper, and a `*_CAP` constant — and proves it on fabricated samples first, since a census that recognises one shape reports "nothing here" for the other four. What the guard deliberately does **not** judge is whether a position *ought* to be read: that is a design call, and the ledger only pins that every capability was looked at once by a person and that what they wrote down does not contradict the code |
301
301
  | `scripts/run-sql-engine-capability-test.mjs` | The SQL-posture read face and the four-state capability reader underneath it. One capability cell here carries **four different things**, and each one points an operator somewhere else: nothing has been observed yet in this process (a one-shot doctor run is always in that state), the response arrived but carries no such key (an older engine), the engine explicitly answered `null` — *this deployment has no SQL backend*, which is a **positive fact** rather than an absence — and a full reading. Fold any two together and the screen states something flatly, confidently, and wrongly, so every positive control here is paired with a control pointing the opposite way, and the four sentences the doctor row can print are checked to be pairwise distinct and non-implying. The reading itself is narrowed no tighter than the mint: `txnMode: null` is a **legal value** — two of the three engines always report it that way, and the upstream type note names reading it as "optimistic" as the error — so treating it as malformed would throw away the entire reading for ordinary deployments, which is the same disease this repo logged when a consumer's domain was narrower than the producer's. A response that cannot be parsed **clears** the cell rather than leaving the previous engine's answer in place, and a separate invalidation port exists for the case the generation latch cannot catch — a same-port respawn whose new probe never succeeded, where the stale reading would otherwise be answered as current fact. Untrusted values (the isolation string is read back from a database server variable) are sanitised and bounded before display, and the bound is applied **before** escaping so a visible escape never gets cut in half. Finally the export names are themselves a guard: the shell still carries a copy that is meant to go red on the package's same-named export and be swapped out, so renaming anything here would silently disarm that lock |
302
302
  | `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
303
- | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there |
303
+ | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. |
304
304
  | `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
305
305
  | `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails |
306
306
  | `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever |
307
307
  | `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
308
- | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer |
308
+ | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. |
309
309
  | `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
310
310
  | `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
311
- | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key pair on the success result: the CC-spelled `structured_output` is the authoritative home for the value the wire calls `structuredOutput`, and the camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rides alongside it for one release with the **same value and the same reference**, so a consumer reading either name gets the same object. The wire position is read exactly once, because two reads let a value-changing accessor mint the two names as two different objects; absence is absence on both names; a wire key that is present but `undefined` mints neither, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` still mint both, and so do shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope mints neither, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
311
+ | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key on the success result: the CC-spelled `structured_output` is the only home for the value the wire calls `structuredOutput`. The camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rode alongside it for exactly one release (0.79.1) and is **absent from 0.80.0 on**, pinned both by own-key and by `in`, so a consumer still reading the old name sees `undefined` rather than a stale copy. The wire position is read exactly once, so a value-changing accessor is only ever asked for its first answer; absence stays absence; a wire key that is present but `undefined` mints nothing, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` are still minted, and so are shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope carries no such key, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
312
312
  | `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the sibling client trees are on disk the consumption evidence is checked by name — a `contract` row's claimed consumers must equal the real set, and a `retire` name must not be imported by any client. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. |
313
- | `scripts/run-wire-refusal-copy-test.mjs` | Two wire refusals read the same way on every client: a cancel's 409 carries one of two codes with opposite dispositions (`conflict.approval_settled` — someone else already decided, go read the result; `conflict.run_not_running` — nothing changed, send the cancel again), an unrecognised or codeless 409 is reported as such rather than guessed, and the submit-side 429 `usage.window_exhausted` is read as a waitable refusal whose wait is stated only when the engine supplied one. `ControlRouter.cancel` raises a distinct safety code for the retry-directly case. |
313
+ | `scripts/run-wire-refusal-copy-test.mjs` | Two wire refusals read the same way on every client: a cancel's 409 carries one of two codes with opposite dispositions (`conflict.approval_settled` — someone else already decided, go read the result; `conflict.run_not_running` — nothing changed, send the cancel again), an unrecognised or codeless 409 is reported as such rather than guessed, and the submit-side 429 `usage.window_exhausted` is read as a waitable refusal whose wait is stated only when the engine supplied one. `ControlRouter.cancel` raises a distinct safety code for the retry-directly case. A third family covers the refusals that a deny's *settler note* can draw: a deployment that signs the decisions it accepts but does not sign that note, a body whose decision and note contradict each other, and a word the engine does not recognise. All three are refused **before** the approval is judged, so each sentence states plainly that nothing was decided and the approval is still waiting, and each names a different next step — none of them “send it again unchanged”, which would simply be refused again. Two of the three codes already mean something else in this package: one is shared with the plan-review leg, so the reader refuses to claim it unless the caller states that the note really was sent, and the other is split by the field the engine names, because without that field the same code means a review outcome was rejected for its content. The third sentence deliberately does not say the word was misspelled: upstream mints that same field for at least four different reasons, so the only thing it proves is that the engine would not take the settlement details and named which part — which is what the sentence says, and where it points. The three sentences are pinned verbatim rather than by keyword, because a keyword check passes a sentence that tells the reader to send the same body again unchanged, which is the one next step that is certainly wrong. The field itself is read from the bag the SDK keeps additional response keys in, not off the top of the error, since only the hand-built shapes a test would write carry it there. Both decision legs hand the reading back on their outcome, and a leg that never sent the note claims nothing. |
314
314
  | `scripts/run-tool-disclosure-progress-projection-test.mjs` | The two wire arms sdk 9.6.0 adds — `tool_disclosure` (name-only tool census: open-set `policy`, `thresholdPercent` absent ≠ default, `deferred`/`activated` full snapshots) and `tool_progress` (one frame, two beats: Bash ticks carry an output tail with `totalLines`/`totalBytes` that come and go together; other tools carry only `elapsedSeconds`) — project to neutral internal arms plus chrome arms. Required keys missing ⇒ `malformed`; bad optional keys drop only themselves; the sub-flow three-key gate keeps child frames off the leader lane; both arms are `required: false` in the arm table with duties stated (the output tail is untrusted raw and must never be fed back to the model). |
315
315
  | `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes **0.69.0:** `fetchMcpPanel` fetches the panel through the SDK client's own `sessions.mcp` call (same transport and auth as every other read) and projects it; transport failure, an unreadable body and an empty session id all come back as `undefined`, never as a fabricated empty panel 0.71.0 adds section K: `mcpEngineLegPresence(view)` — the engine-side MCP presence tri-state read only off the panel view (`unknown` when the view could not be read, never rendered as "no MCP configured") |
316
316
  | `scripts/run-absence-fold-census-test.mjs` | A package-wide census of the "absence folded into a positive outcome" defect shape, so that fixing the six sites this release does not merely move the shape somewhere else. The defect is defined by position, not syntax: a fallback position (the unconditional tail return, the `default:` arm, the literal minted when there is nothing to pass on, the value returned from an error path) may only say `unknown` or stay absent, never a positive word. Detection walks the syntax tree of every source file, so comments, strings and multi-line spellings cannot hide or fake a hit, and covers five forms: the right arm of `??` / `\|\|`, the else arm of a ternary, the first return of an explicit `default:`, a `catch` block or `.catch(() => …)` arrow returning a healthy value, and a function whose last statement returns a positive word after other returns. Every remaining hit must be registered with a written reason, an unregistered hit fails the gate naming the file and line, the registered count must equal the real count so a cleared site cannot leave a spare allowance behind, and the gate proves its own teeth behind a fence (a failed self-proof refuses to report any count): each form injected into an in-memory copy must add exactly one hit, two correct spellings are pinned as non-hits, and samples inside comments or strings do not count. It also pins the headline site: the fleet panel projection no longer mints an `end` with `isError: false` on absence |
@@ -323,7 +323,7 @@ guard still cross-checks the table by name).
323
323
  | `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
324
324
  | `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string \| null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
325
325
  | `scripts/run-approvals-feed-unknown-test.mjs` | The approvals feed tells three states apart: **N items waiting**, **nothing waiting**, and **this fetch did not come back, so we do not know**. Every way a fetch can fail (the call throwing or rejecting, a body that is not an object, a `livePending` section that is not an array — including the `null` seen in the field, a `pending` that is not an array or holds a malformed row) publishes `{kind:'unknown', why, at, mode}` on the subscription — never an empty snapshot and never silence. Real snapshots carry `kind:'snapshot'`; `snapshot()` still answers only with the last real one (a fact about the past) while `reading()` answers whether it is current (`unobserved` / `present` / `unknown`). Recovery always publishes a real snapshot again, even when the contents are byte-identical to before the failure. An unknown reading is never counted as zero: the awaiting-decision counts read `null`, the view is empty, and the tracker reports no removals, so cards on screen are not retracted for a failed fetch. Retry, backoff and circuit-breaking are unchanged. |
326
- | `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. |
326
+ | `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. From 0.80.0 that last pin reads the syntax tree instead of scanning lines: the same return had been rewritten across several lines with conditional spreads, a shape a line-wise search misses entirely, which would have quietly turned the pin into a check of nothing. |
327
327
  | `scripts/run-panel-cycle-identity-test.mjs` | The **cycle identity** on agent-panel events and the fleet ledger's **departure read-out**: a background agent may be revived under the same id, so `fleet-row` and `end` events now carry the wire's own `cycleSeq` / `startedAt` when present (absent means the row has no notion of generations, never "generation one"), `isStaleEngineAgentPanelEnd` is the single rule for ignoring a late `end` from a previous cycle (only when both sides carry a comparable identity; absence never drops a real terminal), a changed `cycleSeq` is a new cycle for usage stickiness and buffer coalescing, the notification lane carries `cycleSeq` only when the wire really sent `seq`, and `task_remove` frames reach the host through `onTaskRemoved` with `removeReason` / `cycleSeq` verbatim, a stale previous-generation removal leaving the newer row in place. A terminal row held back because the consumer has no such row yet is also released by the keys it carries itself (its transcript id, or the delegating call id of a subagent already on screen under its wire id), since the key tables are only written once a row has actually been published — the release still goes through the one funnel, so the subagent stays one row; a running frame arriving after such a held terminal row is a stale snapshot when both sides carry a comparable generation and it matches (no event, the held row keeps its final usage), a revival when the frame is provably newer (the held row is dropped), and is treated as a revival when neither side can be compared. A subagent lifecycle event carries `agentType` only from an honest source — the fleet row's own agent type, recorded before it is folded into the row label — and omits the key when there is none, never substituting the display name |
328
328
  | `scripts/run-workflow-size-warning-test.mjs` | The **workflow size warning** verdict shared by every host footer / panel: a three-state result (`warn` / `ok` / `unknown`) read off the optional fleet view keys, where an unknown size is never reported as a normal one (absent `totalCount` / `tokens` without positive over-cap evidence is `unknown`, naming the missing keys), positive evidence on either axis wins regardless of absent keys, the per-agent denominator uses the engine's started count only when it is not below done+failed (a smaller value is a stale reading), otherwise falls back to the done+failed lower bound only when both keys are present — and a lower-bound denominator only yields an upper bound of the projection, which can prove *within cap* (`ok`, flagged) but never *over cap* (`unknown`, with the upper bound exposed) — and the prior is used only when the engine itself reports zero started agents; cap precedence env > explicit guideline > default, prototype keys never act as a guideline, the env reader is pure and does not fall through to the second name on a bad first value; caps, guideline table, env names and the three copy variants are single-sourced |
329
329
  | `scripts/run-approvals-stream-live-capability-test.mjs` | The engine's live-approval-push self-description (`capabilities.approvalsStreamLive`, engine ≥7.87.1), read the same four-state way as its four sibling capability readers: an absent key is reported as not reported (never folded into `false`), the value must be a strict boolean, and the one decision the feed consumer needs — whether it must keep pulling suspended asks itself — is answered by `livePendingNeedsReconcile`, which only says no when the engine explicitly says it pushes. |
@@ -391,7 +391,7 @@ guard still cross-checks the table by name).
391
391
  | `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
392
392
  | `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
393
393
  | `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
394
- | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false |
394
+ | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. |
395
395
  | `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
396
396
  | `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
397
397
  | `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
@@ -412,6 +412,8 @@ guard still cross-checks the table by name).
412
412
  | `scripts/run-mcp-probe-face-test.mjs` | The engine's **run-free MCP status face** (server ≥7.93.0, sdk 11.2.0): the `capabilities.mcpProbe` bit read the same four-state way as its eleven sibling capability readers, and two call ports on top of it — read the deployment's own declared servers, or probe a caller-supplied list. Presence of the bit is carried by the engine version, so an **absent key means an older engine** (that route answers a coded 404) and is read as *not reported*, never as *no*; an explicit `false` is the engine's own no and is reported without inventing a reason (whether caller-supplied declarations are accepted is a different bit's question); a non-boolean is malformed and is dropped rather than folded into a no. The availability verdict answers only *should this call go on the wire*: an explicit no means zero requests, and both kinds of *cannot tell* are sent anyway, so an older engine answers with its own coded refusal instead of being silently skipped. One failure judge serves both ports and asks **provenance before status**: a 4xx, 501 or 503 that carries no machine code proves nothing about who answered and is reported as *no verdict*; twelve coded refusals each get their own arm (three identity codes, one of which sits outside the `auth.` family so a prefix fallback would miss it; three different treatments ride the same 400), and a coded answer this version does not recognise lands in *cannot tell*, never in *this engine has no such face*. Retry-after seconds ride 429 and 503 and are absent rather than 0 when the engine gave none. The 200 body is narrowed through the **same per-row narrower** as the streaming `wiring_manifest.mcp[]` leg and the session panel's replay leg, liveness cell included; on this face rows are paired with the submitted declarations **by index**, so a dropped row makes the whole answer unreadable rather than a half table, a row count that differs from the submitted count is reported as misaligned, and an honest empty `servers: []` is kept apart from *non-empty but nothing readable*. `probedAt` and `ttlSec` pass through untouched — this package mints no freshness verdict — the declaration list is handed over as-is with its length snapshotted once, a list that serialises itself differently from what was counted is refused locally with zero requests, and no port ever retries a dial. |
413
413
  | `scripts/run-sdk-wire-transit-test.mjs` | The package's pass-through of a few SDK names (`sdkWireTransit`), pinned: every value re-export is the **same reference** as the SDK's own (a client class the host recognises with `instanceof`, the two approval-frame predicates, the three session-bundle calls and the SDK error class), not a look-alike wrapper — wrapping would discard the one anti-drift guarantee a pass-through has; every type re-export is present by name in the emitted declaration file; the gate's list and the source file's export lists are compared in both directions so a name added to one without the other turns red; and a name the SDK does not export must fail the same test, so the gate is not vacuously green. |
414
414
  | `scripts/run-sdk-registry-transit-test.mjs` | The package's cloud control-plane surface (`sdkRegistryTransit`), which lives behind its own `./registry` subpath entry point rather than on the root barrel, and this guard holds both halves of that decision. Upstream publishes the same surface behind a subpath of its own, because the subject differs: the engine-wire surface speaks for one engine's service credential, this one for a person's rotating token, and their refresh and error semantics were deliberately never merged. Keeping it behind a second entry point means a client that never touches the control plane neither resolves nor type-checks it. Every value re-export is therefore read from the file the subpath entry actually resolves to, and must be the **same reference** as upstream's own (the control-plane client class, the three config reads, the health probe, the feedback call, the auth-path constant, the content-address helper, and the two typed error classes a host recognises with `instanceof`); every type re-export is compared with the emitted declaration file in both directions; the gate's list and the source file's export lists are likewise compared both ways; and the root entry is checked to carry none of these names, with the root barrel's own source checked to reference the entry file nowhere — a name leaking onto the root would put the cost of this surface back on clients that never asked for it. The subpath is then verified end to end: the installed SDK must really publish its own `./registry` entry and declare every transited name inside it, and this package's own `exports` must point that subpath at exactly the files the gate just judged. Portability is two checks rather than one, done with a parser rather than a text search: the upstream subpath's emitted JavaScript, walked recursively, must contain neither a `node:` specifier (static, side-effect, dynamic, `require` and re-export forms all exercised) nor a Node **global** — because the same package's third entry point is a Node-only surface that imports nothing at all and reaches for the `Buffer` global, so a specifier check alone would pass it as isomorphic, while a byte-level search of it reports two `node:` hits that live entirely inside a documentation example. A text scanner that merely strips comments first gets both directions wrong on ordinary JavaScript — a regular-expression literal containing a slash pair swallows the rest of its line, and the word in a string reads as a reference — so both scanners run off the syntax tree and are checked against fixtures for each failure direction as well as against that real material — including a dynamic import written with a template literal, which a check that accepts only quoted strings misses entirely, and a dynamic import whose target cannot be determined statically, which is refused rather than read as no edge at all. Zero-processing is likewise enforced with a syntax-tree allowlist rather than a keyword search: every top-level statement must be a named re-export carrying that one specifier, so an import followed by an in-place edit of the upstream prototype is refused with a file and line — that shape leaves the name lists untouched and even keeps the same-reference check green, since both sides are then the one object that was damaged. The last section measures, rather than merely notes, one declaration-level gap: two of the upstream subpath's declaration files reference a package the SDK lists only among its own dev dependencies. Moving such a check to a scratch directory is not isolation, because package resolution walks the ancestor directories, so the gate builds a sandbox served by a restricted compiler host and proves the isolation both ways — a decoy copy of the missing package placed one level above the sandbox must silence the errors for an unrestricted host and must not silence them for the restricted one. Then it installs this package into that same sandbox as a real consumer would, and pins the two readings that justify the entry-point split: a consumer that imports only from the root sees no unresolved-module errors at all, while a consumer that imports the subpath sees exactly the two, reported honestly rather than swallowed by this layer. The day upstream ships those declarations, that section turns red and the note comes out with it. The separation itself rests on the root closure being computed correctly, so the portability guard that computes it was extended in the same change: a template-literal dynamic import is followed like any other edge, and an edge whose target cannot be resolved statically is refused on every one of the four graphs — without that, a single line in a third file already reachable from the root would put this surface back into the root runtime while every guard stayed green. |
415
+ | `scripts/run-rule-removal-consequence-test.mjs` | The one sentence a rule-removal confirmation surface shows for what removing the rule will actually do: one line per behaviour the rule could have been enforcing, plus a neutral line for when that behaviour cannot be read back, so a caller that hits an out-of-set or missing value never falls back to a specific claim it cannot support. The guard compares the actual output against frozen text rather than merely checking that some string came back, so a dropped word or a swapped clause is caught the day it lands, and the four sentences are pinned pairwise distinct. The line for a rule that was denying something is pinned to say the removal widens what can run rather than echoing the wording used for a rule that asks again — the two are opposite directions, and sharing a sentence between them would tell the person confirming the removal the opposite of what is about to happen. The neutral line is checked from the other side for the same reason: it must not contain a word that belongs to only one of the three behaviours, because that would answer on behalf of a state the caller was unable to determine. The lookup that turns a raw stored or transmitted value into one of the three behaviours reads by strict equality only, proven with a fully trapped proxy and a counting getter to show it never touches a property on whatever it is handed, so a value with a legitimate-looking word sitting on its prototype chain is rejected exactly like any other out-of-set value rather than being unwrapped. A closing self-check mutates one character out of each frozen sentence and asserts the exact-match comparison actually fails on it, so the guard cannot pass by checking only that a string of some kind came back |
416
+ | `scripts/run-mcp-engine-leg-test.mjs` | The two legs behind one row on the MCP detail card. In a two-process setup the servers are hosted by the **engine**, while the Status cell on the card reports **this client's own** connection to them, and the two are independent truths: this client failing to connect does not mean the server's tools are unavailable, and the engine leg on the same screen may be saying it completed an exchange moments ago. Two mints share the work. The first reads one server's liveness off the engine leg as six readings, and its load-bearing distinctions are three. **No roster that could be read end to end** (`null`, `undefined`, anything that is not an array, a length that is not a non-negative integer, a length beyond the scan bound, or rows that could not be read at all with nothing matched) is **not** the same as an **empty** roster, which is the engine leg's positive statement that it declared no servers; the same narrower that feeds this reader answers `undefined` when a non-empty section yields no readable row, so *I could not read it* never turns into *I know it is zero*, and a roster that was not read to the end never produces *this server is not on it*. **Could not be read is not the same as absent**: absence means only that the row carries no liveness key of its own, while a record that is present but unintelligible — a cell that is not an object, a word that is not a non-empty string, or a read that fails outright — is reported as unreadable and is decided before the observation beside it, since a record that may well have said the opposite is no evidence of reachability. **An ambiguous name is answered as ambiguous**: when a roster that was read end to end matches a name more than once, the reader reports the match count instead of picking one, because the upstream contract for the sibling MCP face states that names are not guaranteed unique, and picking optimistically would contradict the worst-fact-first verdict shown on the same screen. When a row's identity could not be read at all the reader makes **no claim about that name whatsoever**, not even a count, since the row it could not read may well be a second one carrying the same name — a row that is plainly not a row, such as a hole in a sparse array, is a different matter and leaves the roster complete. The observed word is passed through **verbatim as an open set**, so a fourth word one day arrives at consumers untouched. The second mint is the one sentence that goes under Status, and it appears **only** when this client's own connection really did fail — the other Status values already tell the truth, and a sentence on top of them would only muddy the verbatim health vocabulary. **The liveness observation outranks the tool roster, and having heard from the engine includes hearing that it could not tell, and hearing something unreadable**: a word meaning *looked and could not tell*, a word this version does not recognise, and a record that arrived but could not be read each get their own sentence rather than falling back to the roster, because falling back would say this client holds nothing at all while the diagnostics page shows that very record. A value that is not a reading at all is a different case and does fall back, since nothing then establishes that the engine sent anything. Only when liveness genuinely cannot speak does the roster get a turn, and the roster itself has four answers — tools were listed, nothing was said, the engine reported zero, and the count could not be read — because **unreadable, absent and zero are three different facts**. No sentence claims that nothing at all has been seen about the server, because on two of the readings that reach the roster the engine has plainly listed it; the sentence for a silent roster states only what this client holds. The nine sentences are pairwise distinct, each one names the engine leg, and **none of them renders the liveness word itself**. Membership of the word list is decided by the one shared predicate rather than a second copy, and the branch over the known words is pinned so that a new word upstream fails the build instead of silently taking the *not recognised* sentence. Neither mint ever throws, whatever a host hands it: every value is taken through one reader that accepts **own properties only** — an inherited key is not a wire fact, and one on a shared prototype could otherwise manufacture an engine observation or suppress a real one — reads each key exactly once, and keeps a read that fails apart from a value that is absent. The roster is walked by index rather than through the array's own `find`, the array test is guarded because it can throw on its own, and a length that overstates itself would otherwise spin forever |
415
417
 
416
418
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
417
419
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
@@ -1,8 +1,7 @@
1
1
  import { wiringManifestViewOf } from './wiringManifestView.js';
2
2
  import { stamp, snapshotSegmentIdentity, } from '../types.js';
3
3
  import { turnUsageToModelUsage } from './turnUsageToModelUsage.js';
4
- import { gateOutcomeOf } from '../../gateOutcome.js';
5
- import { isCcToolDenialKind } from '../../gateVocabulary.js';
4
+ import { ccToolDenialKindForToolEnd, gateOutcomeOf } from '../../gateOutcome.js';
6
5
  import { projectToolRoster, projectToolRosterDelta } from '../../toolRoster.js';
7
6
  import { readMcpLiveness } from '../../mcpLiveness.js';
8
7
  import { SEGMENT_END_IDENTITY_KEYS } from '../types.js';
@@ -171,8 +170,7 @@ export function eventToSdkMessage(ev, ctx) {
171
170
  const toolEndDelivered = ev.delivered;
172
171
  const toolEndGatedCallId = ev.gatedCallId;
173
172
  const collateralAbort = ev._sema_collateral_abort;
174
- const denialKindRaw = ev._sema_denial_kind;
175
- const denialKind = isCcToolDenialKind(denialKindRaw) ? denialKindRaw : undefined;
173
+ const denialKind = ccToolDenialKindForToolEnd(ev);
176
174
  return projected(stamp(ctx, armBody({
177
175
  type: 'tool_end_result',
178
176
  toolCallId: ev.toolCallId,
@@ -195,7 +195,7 @@ export function readRunCostFacts(stats, observed) {
195
195
  }
196
196
  function structuredOutputParts(r) {
197
197
  const so = r.structuredOutput;
198
- return so !== undefined ? { structured_output: so, structuredOutput: so } : {};
198
+ return so !== undefined ? { structured_output: so } : {};
199
199
  }
200
200
  function effectiveFactParts(rec) {
201
201
  const reasoning = readEffectiveReasoning(rec.effectiveReasoning);
@@ -5,7 +5,7 @@ import { readRunCostFacts, terminalToSdkResult, TOOL_INPUT_JOIN_MAX_CALLS } from
5
5
  import { coerceOutput, publishSubagentContentEvent } from '../subagentContentStore.js';
6
6
  import { ACTIVE_RUN_BUSY_ERROR_CODE, OUTPUT_INVALID, isLimitsExceededCode } from '../engineErrorCodes.js';
7
7
  import { isReviewPark, readRunTerminal, runTerminalCode } from '../runTerminal.js';
8
- import { gateDeniedBy, gateOutcomeOf } from '../gateOutcome.js';
8
+ import { ccToolDenialKindForToolEnd, gateDeniedBy, gateOutcomeOf } from '../gateOutcome.js';
9
9
  import { isCcToolDenialKind, isGateDeniedByWord } from '../gateVocabulary.js';
10
10
  const mainLane = () => ({ lane: 'main' });
11
11
  function emitChromeFireAndForget(ctx, event) {
@@ -182,6 +182,7 @@ async function* runStreamInner(events, ctx, handle = {}) {
182
182
  let toolInputOverflowed = false;
183
183
  const gateDeniedByByCallId = new Map();
184
184
  const denialKindByCallId = new Map();
185
+ const denialKindSourceByCallId = new Map();
185
186
  const denialAmbiguous = new Set();
186
187
  let denialJoinOverflowed = false;
187
188
  let successfulToolEndObserved = false;
@@ -230,8 +231,12 @@ async function* runStreamInner(events, ctx, handle = {}) {
230
231
  if (typeof endCallId === 'string' && endCallId.length > 0) {
231
232
  const deniedByRaw = gateDeniedBy(gateOutcomeOf(ev));
232
233
  const deniedBy = isGateDeniedByWord(deniedByRaw) ? deniedByRaw : undefined;
233
- const kindRaw = ev._sema_denial_kind;
234
- const kind = isCcToolDenialKind(kindRaw) ? kindRaw : undefined;
234
+ const kind = ccToolDenialKindForToolEnd(ev);
235
+ const kindSource = kind === undefined
236
+ ? undefined
237
+ : isCcToolDenialKind(ev._sema_denial_kind)
238
+ ? 'local'
239
+ : 'wire';
235
240
  if (deniedBy !== undefined || kind !== undefined) {
236
241
  const isNew = (deniedBy !== undefined && !gateDeniedByByCallId.has(endCallId)) ||
237
242
  (kind !== undefined && !denialKindByCallId.has(endCallId));
@@ -240,21 +245,33 @@ async function* runStreamInner(events, ctx, handle = {}) {
240
245
  denialKindByCallId.size >= TOOL_INPUT_JOIN_MAX_CALLS)) {
241
246
  denialJoinOverflowed = true;
242
247
  }
243
- if (!denialJoinOverflowed && !denialAmbiguous.has(endCallId)) {
248
+ if (!denialAmbiguous.has(endCallId)) {
244
249
  const seenDeniedBy = gateDeniedByByCallId.get(endCallId);
245
250
  const seenKind = denialKindByCallId.get(endCallId);
251
+ const seenSource = denialKindSourceByCallId.get(endCallId);
252
+ const kindDiffers = kind !== undefined && seenKind !== undefined && seenKind !== kind;
246
253
  const conflict = (deniedBy !== undefined && seenDeniedBy !== undefined && seenDeniedBy !== deniedBy) ||
247
- (kind !== undefined && seenKind !== undefined && seenKind !== kind);
254
+ (kindDiffers && seenSource === kindSource);
255
+ const localSupersedesWire = kindDiffers && seenSource === 'wire' && kindSource === 'local';
248
256
  if (conflict) {
249
257
  gateDeniedByByCallId.delete(endCallId);
250
258
  denialKindByCallId.delete(endCallId);
259
+ denialKindSourceByCallId.delete(endCallId);
251
260
  denialAmbiguous.add(endCallId);
252
261
  }
253
262
  else {
254
- if (deniedBy !== undefined && seenDeniedBy === undefined)
255
- gateDeniedByByCallId.set(endCallId, deniedBy);
256
- if (kind !== undefined && seenKind === undefined)
263
+ if (localSupersedesWire || (kind !== undefined && seenKind === kind && seenSource === 'wire' && kindSource === 'local')) {
257
264
  denialKindByCallId.set(endCallId, kind);
265
+ denialKindSourceByCallId.set(endCallId, 'local');
266
+ }
267
+ if (!denialJoinOverflowed) {
268
+ if (deniedBy !== undefined && seenDeniedBy === undefined)
269
+ gateDeniedByByCallId.set(endCallId, deniedBy);
270
+ if (kind !== undefined && seenKind === undefined) {
271
+ denialKindByCallId.set(endCallId, kind);
272
+ denialKindSourceByCallId.set(endCallId, kindSource);
273
+ }
274
+ }
258
275
  }
259
276
  }
260
277
  }
@@ -1,4 +1,5 @@
1
1
  import type { DeniedBy, Settlement } from '@sema-agent/sdk';
2
+ import { type CcToolDenialKind } from './gateVocabulary.js';
2
3
  export type GateDispositionView = {
3
4
  kind: 'allowed';
4
5
  classifier?: GateClassifierRoundView;
@@ -23,6 +24,7 @@ export interface SettlementWhoView {
23
24
  approver?: string;
24
25
  window?: string;
25
26
  }
27
+ export declare const SETTLEMENT_KIND_WORDS: readonly ["human_allowed", "human_refused", "approval_window_expired", "denial_limit_window_expired", "park_sla_expired", "no_approver", "approver_unavailable", "approver_error", "approver_contract", "presentation_failed", "blanket_allow_refused", "policy_refused", "task_aborted"];
26
28
  export interface SettlementView {
27
29
  kind: Settlement['kind'] | (string & {});
28
30
  who: SettlementWhoView;
@@ -43,3 +45,5 @@ export declare function isApprovalWindowExpiredGate(g: GateOutcomeView | undefin
43
45
  export declare function isDenialLimitAutoDeniedGate(g: GateOutcomeView | undefined): boolean;
44
46
  export declare function isParkSlaExpiredGate(g: GateOutcomeView | undefined): boolean;
45
47
  export declare function isHumanSettledGate(g: GateOutcomeView | undefined): boolean;
48
+ export declare function isPolicyRefusedGate(g: GateOutcomeView | undefined): boolean;
49
+ export declare function ccToolDenialKindForToolEnd(frame: unknown): CcToolDenialKind | undefined;
@@ -1,7 +1,23 @@
1
+ import { ccToolDenialKindForSettledBy, isCcToolDenialKind } from './gateVocabulary.js';
1
2
  import { GATE_PARKED_ERROR_CODE } from './engineErrorCodes.js';
2
3
  function str(v) {
3
4
  return typeof v === 'string' && v.length > 0 ? v : undefined;
4
5
  }
6
+ export const SETTLEMENT_KIND_WORDS = Object.freeze([
7
+ 'human_allowed',
8
+ 'human_refused',
9
+ 'approval_window_expired',
10
+ 'denial_limit_window_expired',
11
+ 'park_sla_expired',
12
+ 'no_approver',
13
+ 'approver_unavailable',
14
+ 'approver_error',
15
+ 'approver_contract',
16
+ 'presentation_failed',
17
+ 'blanket_allow_refused',
18
+ 'policy_refused',
19
+ 'task_aborted',
20
+ ]);
5
21
  function readWho(v) {
6
22
  if (typeof v !== 'object' || v === null || Array.isArray(v))
7
23
  return undefined;
@@ -121,5 +137,25 @@ export function isHumanSettledGate(g) {
121
137
  const k = g?.settlement?.kind;
122
138
  return k === 'human_allowed' || k === 'human_refused';
123
139
  }
140
+ export function isPolicyRefusedGate(g) {
141
+ return g?.settlement?.kind === 'policy_refused';
142
+ }
143
+ export function ccToolDenialKindForToolEnd(frame) {
144
+ if (typeof frame !== 'object' || frame === null || Array.isArray(frame))
145
+ return undefined;
146
+ const stamped = frame._sema_denial_kind;
147
+ if (isCcToolDenialKind(stamped))
148
+ return stamped;
149
+ const kind = gateOutcomeOf(frame)?.settlement?.kind;
150
+ if (kind === 'human_refused')
151
+ return ccToolDenialKindForSettledBy('human');
152
+ if (kind === 'policy_refused')
153
+ return ccToolDenialKindForSettledBy('policy');
154
+ return undefined;
155
+ }
124
156
  const _gateOutcomeShapePin = (w) => w;
125
157
  void _gateOutcomeShapePin;
158
+ const _settlementWordMirrorPin = SETTLEMENT_KIND_WORDS;
159
+ void _settlementWordMirrorPin;
160
+ const _settlementWordsCoveredPin = true;
161
+ void _settlementWordsCoveredPin;
@@ -3,6 +3,7 @@ export declare const CC_TOOL_DENIAL_KINDS: readonly ["user-rejected", "permissio
3
3
  export type CcToolDenialKind = (typeof CC_TOOL_DENIAL_KINDS)[number];
4
4
  export declare function isCcToolDenialKind(v: unknown): v is CcToolDenialKind;
5
5
  export declare function isCcToolDenialKindADenial(v: unknown): boolean;
6
+ export declare function ccToolDenialKindForSettledBy(settledBy: unknown): CcToolDenialKind | undefined;
6
7
  export declare const GATE_DENIED_BY_WORDS: readonly DeniedBy[];
7
8
  export declare function isGateDeniedByWord(v: unknown): v is DeniedBy;
8
9
  export declare function gateDeniedByDetail(deniedBy: unknown): string;
@@ -13,6 +13,13 @@ export function isCcToolDenialKind(v) {
13
13
  export function isCcToolDenialKindADenial(v) {
14
14
  return isCcToolDenialKind(v) && v !== 'interrupted' && v !== 'cancelled';
15
15
  }
16
+ export function ccToolDenialKindForSettledBy(settledBy) {
17
+ if (settledBy === 'human')
18
+ return 'user-rejected';
19
+ if (settledBy === 'policy')
20
+ return 'permission-rule';
21
+ return undefined;
22
+ }
16
23
  export const GATE_DENIED_BY_WORDS = Object.freeze([
17
24
  'policy',
18
25
  'hook',
@@ -5,6 +5,7 @@ import { GATE_PARKED_ERROR_CODE } from '../engineErrorCodes.js';
5
5
  import { readRunTerminal, runTerminalCode, runTerminalGateKind, runTerminalGateToolName } from '../runTerminal.js';
6
6
  import { surfaceForCurrentSession } from './hitlHostSurface.js';
7
7
  import { SEMA_DENIAL_KIND_KEY } from './gateLedger.js';
8
+ import { ccToolDenialKindForSettledBy } from '../gateVocabulary.js';
8
9
  export const HITL_REJECT_MESSAGE = "The user doesn't want to proceed with this tool use. The tool use was rejected (eg. if it was a file edit, the new_string was NOT written to the file). STOP what you are doing and wait for the user to tell you how to proceed.";
9
10
  export const HITL_POLICY_DENY_MESSAGE = 'Permission for this tool use was denied: it requires interactive approval, and permission prompts are ' +
10
11
  'not available in this session. The action was NOT performed. Do not claim it succeeded, and do not retry ' +
@@ -18,11 +19,7 @@ function denyOutputForRender(attribution) {
18
19
  : HITL_POLICY_DENY_MESSAGE;
19
20
  }
20
21
  function denialKindForTranscript(attribution) {
21
- if (attribution?.settledBy === 'human')
22
- return 'user-rejected';
23
- if (attribution?.settledBy === 'policy')
24
- return 'permission-rule';
25
- return undefined;
22
+ return ccToolDenialKindForSettledBy(attribution?.settledBy);
26
23
  }
27
24
  function denyStampedEnd(ev, attribution) {
28
25
  const kind = denialKindForTranscript(attribution);
@@ -47,6 +47,11 @@ export type PlanReviewModeAfter = (typeof PLAN_REVIEW_MODE_AFTER_WORDS)[number];
47
47
  export declare function isPlanReviewModeAfter(v: string | undefined): v is PlanReviewModeAfter;
48
48
  export type HitlSafetyCode = 'binding_mismatch' | 'no_pending' | 'wrong_gate' | 'bad_plan_edit' | 'bad_plan_mode' | 'empty_answer';
49
49
  export type HitlFailureStage = 'fetch' | 'input' | 'card' | 'decide' | 'orchestration';
50
+ export interface HitlDecideDenyOutcome {
51
+ decision: 'deny';
52
+ reason?: string | undefined;
53
+ settledBy?: 'policy' | undefined;
54
+ }
50
55
  export declare class HitlSafetyError extends Error {
51
56
  readonly code: HitlSafetyCode;
52
57
  constructor(message: string, code: HitlSafetyCode);
@@ -90,10 +95,7 @@ export declare class HitlBridge {
90
95
  decision: 'approve';
91
96
  updatedInput?: unknown;
92
97
  remember?: 'session' | undefined;
93
- } | {
94
- decision: 'deny';
95
- reason?: string | undefined;
96
- }, toolUseID?: string, opts?: {
98
+ } | HitlDecideDenyOutcome, toolUseID?: string, opts?: {
97
99
  signal?: AbortSignal;
98
100
  }, preResolvedPending?: PendingCheckpoint): Promise<unknown>;
99
101
  answerQuestion(answers: AskAnswer[], toolUseID?: string, opts?: {
@@ -195,6 +195,7 @@ export class HitlBridge {
195
195
  decision: 'deny',
196
196
  ...this.bindingOf(pending),
197
197
  ...(outcome.reason !== undefined ? { reason: outcome.reason } : {}),
198
+ ...(outcome.settledBy === 'policy' ? { settledBy: 'policy' } : {}),
198
199
  };
199
200
  return this.decideRaw(pending.sessionId, decision, opts);
200
201
  }