@sema-agent/client-core 0.75.1 → 0.76.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/CHANGELOG.md +53 -0
  2. package/README.md +11 -1
  3. package/dist/adapt/arms.js +137 -35
  4. package/dist/adapt/ids.d.ts +29 -2
  5. package/dist/adapt/ids.js +29 -2
  6. package/dist/adapt/panelTasks.d.ts +4 -3
  7. package/dist/adapt/panelTasks.js +29 -11
  8. package/dist/adapt/textStream.js +5 -5
  9. package/dist/adapt/toolCards.js +2 -2
  10. package/dist/adapt/turnFlags.js +5 -5
  11. package/dist/adapt.d.ts +1 -1
  12. package/dist/adapt.js +8 -8
  13. package/dist/adapter/downstream/eventToSdkMessage.d.ts +57 -0
  14. package/dist/adapter/downstream/eventToSdkMessage.js +125 -17
  15. package/dist/adapter/runStream.js +14 -5
  16. package/dist/adapter/types.d.ts +12 -0
  17. package/dist/adapter/types.js +30 -0
  18. package/dist/agentsWireCaps.d.ts +15 -2
  19. package/dist/approvalsStreamLiveCapability.js +15 -35
  20. package/dist/deviceExecutorManagementCapability.js +15 -35
  21. package/dist/engineAgentPanelStore.js +13 -1
  22. package/dist/engineCapReader.d.ts +77 -0
  23. package/dist/engineCapReader.js +87 -0
  24. package/dist/engineErrorCodes.d.ts +4 -0
  25. package/dist/engineErrorCodes.js +14 -0
  26. package/dist/executionLaneCapability.js +15 -38
  27. package/dist/fleet/fleetRowAgentType.js +1 -1
  28. package/dist/hitl/approvalOutcomeNote.d.ts +0 -10
  29. package/dist/hitl/approvalOutcomeNote.js +35 -8
  30. package/dist/hitl/approvalResolution.d.ts +148 -0
  31. package/dist/hitl/approvalResolution.js +199 -0
  32. package/dist/hitl/approvalsFeed.d.ts +100 -2
  33. package/dist/hitl/approvalsFeed.js +234 -18
  34. package/dist/hitl/livePendingAsk.d.ts +26 -6
  35. package/dist/hitl/livePendingAsk.js +52 -12
  36. package/dist/hitl/persistedRulesWire.d.ts +20 -0
  37. package/dist/hitl/persistedRulesWire.js +23 -0
  38. package/dist/index.d.ts +8 -0
  39. package/dist/index.js +25 -1
  40. package/dist/mcpLiveness.d.ts +179 -0
  41. package/dist/mcpLiveness.js +218 -0
  42. package/dist/mcpPanel.d.ts +17 -0
  43. package/dist/mcpPanel.js +21 -0
  44. package/dist/mcpProbeCapability.d.ts +84 -0
  45. package/dist/mcpProbeCapability.js +124 -0
  46. package/dist/memoryComplianceCapability.d.ts +73 -0
  47. package/dist/memoryComplianceCapability.js +109 -0
  48. package/dist/memoryEntriesWire.d.ts +302 -0
  49. package/dist/memoryEntriesWire.js +595 -0
  50. package/dist/memoryOriginCapability.d.ts +68 -0
  51. package/dist/memoryOriginCapability.js +102 -0
  52. package/dist/memorySpecWire.d.ts +175 -0
  53. package/dist/memorySpecWire.js +320 -0
  54. package/dist/peerLaneCapability.d.ts +62 -0
  55. package/dist/peerLaneCapability.js +100 -0
  56. package/dist/permissionRulesWriteCapability.d.ts +64 -0
  57. package/dist/permissionRulesWriteCapability.js +98 -0
  58. package/dist/seam.d.ts +53 -5
  59. package/dist/seam.js +10 -1
  60. package/dist/selfOrchestrationDenial.js +7 -2
  61. package/dist/sqlEngineCapability.js +15 -35
  62. package/dist/webSearchBackendCapability.js +15 -35
  63. package/dist/writeProtectionCapability.js +15 -35
  64. package/docs/INTEGRATION-CLIENTS.md +148 -15
  65. package/package.json +1 -1
package/CHANGELOG.md CHANGED
@@ -49,6 +49,59 @@
49
49
  > 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
50
50
  > commit 漏转在发布前就红,不再靠人记。
51
51
 
52
+ ## 0.76.1(2026-09-20)
53
+
54
+ > 主题:**能力位读器工厂 + 五只新读口 + 一处别名收口**(patch;零型面 BREAKING;公面值导出 1081 → **1120**)—— 读器工厂(CC-75)· MCP 活性观察读口(CC-36)· `capabilities.peerLane` / `capabilities.permissionRulesWrite` 四态读口(CC-102)· 记忆治理面两只四态读口 + 三只条目面回体读口(CC-97 a)· 主车道证明逐次新建(CC-99)。🔴 **已从本版撤出**:`capabilities.mcpProbe` 读口(CC-74)—— 上游三仓源码里尚无这一格与它的 501 码,上游那张票仍是设计稿;公面一旦发出就是承诺,无一字可验的键名不上公面,候上游真码落地后按字节重做。
55
+
56
+ ### Fixed
57
+
58
+ - 🔴 **身份四键的一次性快照只修了投影口一侧(外部验收方对 0.76.0 的 F1 / F2)**:0.76.0 说「修在共享助手上 ⇒ 全部内部臂连带受益」—— 不成立。① `compaction_outcome` 的 adapt 侧臂两版逐字节相同,仍先判子流再取 `parentToolCallId`(读两次),而投影口已换成每键一读 ⇒ 值随读变的帧让「投影口 → 内部臂」与「直喂 adapt()」两条路**分叉**(直喂丢事件、经投影出 subagent chrome)—— 恰是 0.76.0 自己那格「两入口同答」要禁的形,还落在发车帖第一个点名的臂上;② `text_end` / `reasoning_end` 走的投影函数在同一个表达式里把每只身份键读两次,「每键恰读一次」在这两条臂上两版都不成立(方向 fail-closed:合法子代帧被判 malformed)。现在:快照助手 `snapshotSegmentIdentity` 单源(投影口的 `identitySnapshot` 委托它,adapt 侧 compaction 臂与 `prompt_assembled` 臂改读它;两侧各恰读一次、读抛整帧拒),`text_end` / `reasoning_end` 改走与其余七条内部臂同一只 `identityOrMalformed`;adapt 侧另两处「typeof 再取值」的双读(工具结果收卡的 `eventId` / `parentToolCallId`、`task_progress` 的 `parentToolCallId`)改单读 —— 🔴 **这两处也是修,不是纯重构**:对稳定帧行为不变,对值随读变的帧按设计改变(修前判形看第一份、取值看第二份 ⇒ 收卡丢父卡号、tick 落到主车道;修后一份到底),各给一格判据(compaction 门 A11 / A12)。🔴 **射程如实**:这两处是**存量**(2026-08-02 适配内核搬入那批入库,最早含它的发布版 0.67.0;在已发布的 0.75.1 / 0.76.0 制品上用同一夹具实跑读数一致),不是 0.76.0 引入;只对进程内**非幂等帧**(值随读变的 accessor)可见 —— 经 sdk 事件流进来的帧是其 SSE 读器 `JSON.parse` 的产物(普通数据属性,两次读同一个值),不会触发;不经序列化、把活对象直接喂进 `adapt()` 的路(宿主自建管线 / 重放腿 / 引擎与壳同进程直传)才受影响。F1(compaction 两入口分叉)则是 0.76.0 新引入。判据:compaction 门 A8–A10(两入口同答 ∧ 直喂 / 投影各恰读一次 ∧ 读抛不逸出)、文本段落门 Te0a / Te0b(两臂投影恰读一次且身份原样带出);修前实跑各红。🔴 **0.76.0 已发布且含 F1 / F2**:排了 0.76.0 的下游建议直接提 0.76.1。
59
+ - **0.76.0 发车文案勘误(已发布段冻结,勘误只在这里)**:① 「两条行为面改口**只对直接调读器的端**可见」**被证伪** —— 审批 feed 自己就调 `readLivePendingRows`,只订阅 feed 的端对 `livePending:[{}]` / `pending:[{}]` / `pending:[7]` 三形在 0.76.0 也从「真快照 pending:1」变成 `unknown{unreadable_payload}`;修法方向不改,射程措辞改成「feed 订阅者同样可见」。② CC-77 ①「行为一字未动 … 逐键对拍等价」按字面不成立:12 形里 2 形跨版不等(均 `settled` true → false,fail-open 收紧;包内构造不出,只有宿主直调公面才够得着)。③ 发车帖「24 个 `_*ForTest`」按字面只匹配 22 只(另 2 只是 `_reset*` 无后缀);本版 §0a 改按真匹配数写。④ 「新门四只」实列五只。⑤ 流程:0.76.0 权威全门的第五轮来自与旧轮**无重叠**的一跑(前四轮里有一轮是重叠窗内旧轮负控套篡改共享文件报的假红,当时未写进帖)。
60
+
61
+ ### Changed
62
+
63
+ - **六只能力位读器的四口收成一只工厂(CC-75)**:`sql` / `writeProtection` / `webSearch.backend` / `executionLane` / `approvalsStreamLive` / `deviceExecutor.management` 此前各自持有**逐字相同**的四份代码(caps 探测链上的 tee / per-baseUrl 读口 / 换代失效口 / 测试钩),现在共用包内叶 `createEngineCapReader`(不出公面);每只只剩**投影函数**(键路径 / 在场判据 / 值形校验)与诊断措辞表两件。**等价性三证**:六只读器的 `.d.ts` 与 0.76.0 发布制品**按字节一致**;六只各自的既有门一字未改、格数不变;行为逐条留在原位。🔴 刻意**不做**声明式键路径表 —— 投影函数本身就是那一份判据,压进通用描述表 = 给同一份判据立第二个判官。未观测读数改成**每次新铸**(此前六只各自新建、收编后若共享一只常量会被就地改坏,而失效口清不回它)。工厂门新增两格:**失效口的爆破半径恰是被点名那一格**(宿主同时连两台引擎时,失效掉 A 那台不许把 B 那台当代的读数清成「未观测」—— 那是凭空造出的缺席);**名单等值**(`src/` 里每一个调工厂的文件都必须在门的读器表里有一行,新读器逃不出这道门)。
64
+ - **主车道证明改逐次新建(CC-99)**:`laneProof` 是公开输出面上的一位,此前约五十个主车道发点交出去的是**同一只**模块级常量 `{lane:'main'}` —— 跨适配器实例、跨流、跨 turn 全是它;修前实测两个互不相干的消费者各取一次未登记行的证明 `===` 为真,给第一只补上卡号之后第二只当场读出那张卡,开流序幕四条主车道事件也一起变。🔴 **判据如实**:按语法树穷举,本包内对车道证明**零就地写**,三个消费端的用法全是只读 ⇒ 本条修的是**别名风险**(拿证明做身份去重 / WeakMap 键 / 视图层身份记忆的消费者,在两条无关行共享一只对象时会把它们认成同一条,一个字不写也照错),**不是**一条已经在错归因的行。修法 = 每次新建(`mainLane()` 出厂口),落选的「冻结」只挡写不解别名、且只冻主车道半边会让同一写操作随行是否绑卡漂成间歇抛。同批:终态那两处(关卡 / 终态 tick)此前取一次证明给 stop 与 end 两条事件**共用** ⇒ 改成「归属判一次、证明铸多只」—— 刻意不做「每条事件各调一次」,生成器两次 yield 之间挂起,中间绑卡会把同一行分流到两个归属。
65
+ - **`failloud` 豁免上限收到当日实测(42 → 36)**:那 6 格松量是历次抬数后 catch 又被删掉留下的,留着等于这道棘轮在 6 格之内不判任何新增静默吞。
66
+
67
+ ### Added
68
+
69
+ - **引擎托管 MCP 服务器的活性观察读口(CC-36)**:`wiring_manifest.mcp[].liveness` 答的是**「这台服务器还够得着吗」**,与同一行上的申报位 `status`(拨号那一刻的判决,引擎有意冻结)、re-dial 回执的 `outcome`(一次动作的判词)**各说各的**;`status:"failed"` 与 `liveness.state:"reachable"` 同行并存是真行(答了握手、答的是协议错)。读法是**四种读数**:三个词(`reachable` / `unreachable` / `unknown`)+ **缺席 ⇒ `indeterminate`** —— 缺席在 wire 上四源同形(这条腿没走到那台 / 声明从未被拨过 / 老引擎或老服务端 / 上游投影判形没过),唯一诚实的含义是「没有可用的观察记录」,**不许**据它反推拨没拨过号,不许折成 `unknown` / 健康 / 关着。畸形与缺席**分开**:行上立 never-false 的 `livenessUnreadable`,腿级判词至多退到 `unknown`、**绝不**答 `reachable`;只由读不出的格决定的 `unknown` 另说一句(「记录在场而本端读不出」,主语不是「回来的东西答不了」—— 上游哪天加第四个词,老客户端落的正是这一档)。腿级判词按**最坏事实优先**合成,判词后面的观察时刻**与判词同源**(说「至少有一台够不着」时它是最新那次够不着的时刻,不是这条腿上任何一次观察的最大值)。新导出 `MCP_LIVENESS_STATES` / `readMcpLiveness` / `mcpLivenessRollupOf` / `mcpEngineLegHealthOf` / `mcpEngineLegHealthDetail`;行视图 `WiringManifestMcpEntryView` 加两个可选位(纯 additive;活体腿与面板回放仍是同一只读器)。⚠️ 服务端自 **7.91.2** 起把这一格过境,更老的服务端上恒缺席 = `indeterminate`(不是故障);`/mcp` 面板 `servers[]` 那一面上游投影白名单不带这一格,本包不铸那个位。
70
+ - **两枚具名能力位键的四态读口(CC-102;server ≥7.91.1;sdk 9.8.1 无型 ⇒ 结构读)**:`capabilities.peerLane`(跨会话车道;**不是 HTTP 面**,v1 零新端点,消费方式是渲染 / 预期)与 `capabilities.permissionRulesWrite`(收紧方向的单步写口对这个调用方够不够得着)。四态 `unobserved` / `not_reported` / `absent` / `present`,三种「没有」不许互折。🔴 `permissionRulesWrite` 的读法按上游成文:**键缺席 = 这份二进制比单步写口老 ⇒ 藏入口;键在场 ⇒ 按本键自己的值判**(值与 `permissionRules` 是服务端同一个表达式),**与撤销面那一位读法不同、不许混** ⇒ `src/hitl/persistedRulesWire.ts` 新增便利口 `persistedRulesWriteAvailable(baseUrl)`,只由本键的四态派生,**不合取** `permissionRules`、不再读一遍 caps(反钉:`permissionRules:true` ∧ 写键缺席 ⇒ `false`;写键 `true` ∧ `permissionRules` 缺席 ⇒ `true`);显式 `undefined` 不回落装好的默认读锚(与两只兄弟 gate 同律)。`peerLane` 键缺席按上游成文语义是「老 worker,本部署没有跨会话车道」⇒ `peerLaneAvailable` 答 `no`(在场性即版本信号),只有本进程从没观测过才答 `unknown`。新导出 13 件 + 测试钩 2 件。
71
+ - **记忆治理面两只能力位四态读口 + 三只条目面回体读口(CC-97 a;server ≥7.46.0 起铸两位)**:`capabilities.memoryCompliance`(出处问询 / 抹除两口)与 `capabilities.memoryOrigin`(外源审计 / 清标三口)各一只四态读口(工厂薄包装),🔴 **刻意不合并成一只**:上游把它们做成两个产品面并明说不合并(今天恒同值只是巧合),本包各持一张 per-baseUrl 表;两口 / 三口可用性判词各自归包,记忆治理面缺席时的 501 码 `MEMORY_ENGINE_REQUIRED_CODE` 只铸一次(上游源码里真有这一码)。三只回体读口一个动词都不调(动词半场候抬 sdk 地板):**条目导出行逐行窄读**(空数组 = 真 0 / 非空全坏 = `malformed` / 半坏 = `present` + `dropped`,读不懂的行与没有行永不共用一句话);🔴 **「空答不等于本店干净」现在是代码** —— 这条读口在服务端**缺省把带外源标的条目整条扣留**(消费方以查询参数声明认识 origin 键族后才放行),而装着的 sdk 那只动词不带该参数 ⇒ 经它拿到的回体里带标条目恒被扣留、回体形状与「没有带标条目」完全一样;所以「这个 scope 有没有外源条目」的判词口 `scopeExternalOriginVerdict(reading, asked)` 第二个形参**必填**,没声明 ⇒ 空答一律 `unknown(fenced)`,肯定的否定话最远只到调用方点名的那一个 scope 且带着 scope 串,联合里没有「整店」臂;服务端明说扣留了 n 条时那个计数本身就是「有带标条目」的肯定证据。**抹除证明读口**:「这一次一条都没抹掉」/「200 空体」/「回体读不懂」判成**三回事**;版本信封 `v > 1` 一律拒不重解释;`historyUnknown` 原样三态过境(上游今天只铸 `true` 或整键省掉)且**多出第四态** `unreadable`(键在场读不成,不折回缺席);「判决那一刻没有这条记录」**不**授权任何界面说「它从来不存在」(上游明说这份证明不区分从没存在过 / 自然删掉 / 别的请求删的);重发姿势三态(证据腿幂等 / 降级腿由人决定 / 读不出不当幂等)。**外源清标回执读口**:三把承重键 + 被清掉的标原样带出。🔴 **每一段数组只读一遍**(长度恰读一次、每下标恰读一次,不走调用方迭代器;自报行数大到读不下 ⇒ 如实报读不出,绝不静默截断)—— 异源对抗复审实撞:走的时候报一个长度、走完再报另一个的回体,修前能让一条带标的行被静默丢掉而丢行计数仍为零,一路判成「这个 scope 干净」。「键在场但值读不成」永不折回「键缺席」(历史位 / 扣留计数 / 「这是全量吗」三格各自处置,完备性轴 fail-closed)。新导出 21 件 + 测试钩 2 件。
72
+
73
+ ### Gates
74
+
75
+ - 接入文档冻结账门补**「全量 = A + B = C」算式判据**(§0a 同一格手抄着三个数,此前门只盯行首两个;实翻:改了行首而算式仍是旧三数,门照样全绿)。
76
+ - 新门四只:`run-engine-cap-reader-factory-test.mjs`(表驱动等价门,读器表长即真源)/ `run-mcp-liveness-test.mjs` / `run-peer-lane-rules-write-capability-test.mjs` / `run-lane-proof-identity-test.mjs`(含语法树 + 类型检查器的结构反钉)。
77
+
78
+ ## 0.76.0(2026-09-20)
79
+
80
+ > 主题:**六票一批** —— `prompt_assembled` 投影臂(CC-91)· 审批 feed 的「不知道」三态(CC-98,**型面 BREAKING**)· `ApprovalResolution` 单源判别联合(CC-77 ①)· 记忆 spec 读口与未列键判官(CC-96)· 终态词表补上游钉(CC-78 ①)· `tasks_expand` 退役(CC-92,**型面 BREAKING**)。外加一条**跨全部内部臂**的修复(身份四键一次性快照)。**minor**:两处型面 BREAKING + 两条只对「直接调读器的端」可见的行为面改口;公面值导出 1065 → **1081**(+16)。接入面 §76。
81
+
82
+ ### Fixed
83
+
84
+ - 🔴 **身份四键被反复求值,值随读变的帧能让两个入口分叉(跨全部内部臂)**:同一帧经「投影口 → 内部臂 → 适配层」与「宿主自建管线 / 重放腿直喂 `adapt()`」两条路进来时,身份键(`eventId` + 三只子流键)此前在形门、子流判定、取父键三处**各读一次**;遇到 accessor 帧(两次读给不同值)会让两条路把同一帧判成不同 lane —— **子代的组成信息被洗成主会话的**;读身份抛错还会从两个入口逸出。现在两侧都对四键取**一次性快照**:每键恰读一次、全程用同一份、读取抛错整帧拒。修在**共享助手**上,所有内部臂(`compaction_outcome` / `text_end` / `reasoning_end` / `tool_disclosure` / `tool_progress` / `context_usage` …)连带受益。
85
+ - **审批 feed 把「取不到」渲成「没有」(cli L-426 归包)**:`list()` 抛 / `livePending` 读不懂时此前**零发布**,而 `snapshot()` 恒答上一张、计数照旧答具体数 ⇒ 消费端把「这次没取到」读成「一条都没有」,计数冻在旧值;同形第二处是收到空回体直接**撤卡**。现在订阅口发**判别联合** `ApprovalsFeedEmission`(真快照 / `kind:'unknown'{why,at,mode}`),unknown 时计数三格答 `null`、视图空、**零 delta 不撤卡**;恢复后必发真快照;退避与封顶语义逐字未变。
86
+ - **终态词表五张只有字面量快照、没有上游钉(RH-4 第一步)**:「这条 run / 任务结了没」散在 5 个文件的 8 张表里,此前多数只有 `表.join(',') === 'a,b,c'` 那种快照断言 —— 只证明「表今天长这样」,不证明「长这样是对的」,上游加词时原样绿。现在逐表补**真的双向咬**(exact / superset / subset+词数 canary / 边界钉),读的是**源码文本**经 TypeScript 语法树取值,不读编译产物。
87
+
88
+ ### Added
89
+
90
+ - **内部臂 + chrome 臂 `prompt_assembled`(CC-91)**:引擎的 prompt 组装清单(`blocks` / `sections` / `tools` / `totalChars`)此前在投影口 `not_in_slice`、三端零消费。🔴 **本臂只投 `chars`,一个 token 键都不铸** —— 亲核三面真字节:整条 wire 上**不存在 token 真值**,连 `context_usage.sections[].tokens` 都是引擎拿同一份 `chars` 估的(`ceil(chars / 系数)`)。`sections[].id` 是引擎**开集**(`core/<slot>.<name>`),与端的展示分类名对不上,映射归端。
91
+ - **`ApprovalResolution` 单源判别联合(CC-77 ①,additive)**:`decided` / `not_sent{cause}` / `unsettled{cause}` 三臂 + 两张冻结 cause 词表;`'unresolved'` 的三义就此分开(撤卡 / 编辑被拒 / respond 失败各落一处)。既有导出、型、行为一字未动,便签口 `approvalOutcomeNoteOf` 改由本联合派生(逐键对拍等价)。
92
+ - **记忆 spec 读口与未列键判官(CC-96)**:`agents[].memory` 的四键闭白名单三态读口(`scopes` 缺席 ≠ 空集;`writeScope: null` = 本 run 只读 ≠ 缺席);判官 `unknownMemorySpecKeys` 三态(`[]` = 确认零未列键 / 非空 = 看到了这些 / `undefined` = 读不出),读到退役的单数 `scope` 会说出真后果「**整只 agent** 被 400,不是丢这一个键」。🔴 在场判据按**序列化字节**判(按属性上下文取快照),不按进程内对象的 `hasOwn` —— 不可枚举键与 `toJSON` 会让两者分叉,那会让读口一边说「只读、零违规」一边真发退役键。写面 `TaskAgentWireMemory.writeScope` 同批放宽成 `string | null`(读得出却铸不出是半条腿)。
93
+
94
+ ### Changed(🔴 BREAKING)
95
+
96
+ - **审批 feed 订阅口换判别联合**(CC-98):端的回调入参从 `ApprovalsFeedSnapshot` 变 `ApprovalsFeedEmission`,**编译期强制表态**(不留兼容层、零别名、零双读)。
97
+ - **chrome 臂 `tasks_expand` 退役**(CC-92,clay 裁定 C-R76 ②):包内零铸点、消费端空桩,双侧皆死。`ChromeEvent` 联合缩窄 + 臂表行删除,靠 `Record<ChromeArmKind, …>` 的编译期耦合咬住;穷举 switch 与字面量比较的消费者都会拿到编译信号(TS2678 / TS2367)。
98
+ - **两条只对「直接调读器的端」可见**(CC-98):`readLivePendingRows` 对「非空数组却一行都读不出」从 `{present, rows:[]}` 改判 `{malformed}`(半坏仍丢行 + present + dropped,空数组仍是真 0);durable `pending` 段开始窄读(`[{}]` / `[7]` 从「发一张 durable:1 的快照」变「发一张不知道」)。
99
+
100
+ ### Guards
101
+
102
+ - 新门四只:`run-prompt-assembled-projection-test.mjs`(48 断言)· `run-approvals-feed-unknown-test.mjs`(103 格)· `run-approval-resolution-test.mjs`(45 格)· `run-memory-spec-wire-test.mjs`(68 格)· `run-terminal-table-provenance-test.mjs`(51 格,走语法树不走正则)。变异合计 **70 余枚**逐格见红。异源对抗复审:六辆车合计 **14 轮**,`[high]`×6 + `[medium]`×14 全采修;其中三条由复审**纠正了车自己修反的第一版**。
103
+ - 棘轮:公面 1065 → 1081(+16)· typeshape unknown 出境 370 → 373(负控锚同批)· portability index 闭包 180 → 182 · singleton 清单 405 → 408(high 119 不动)· failloud 豁免 41 → 42(带账:身份快照的 catch 是 fail-closed,读身份抛错 ⇒ 整帧拒)。
104
+
52
105
  ## 0.75.1(2026-09-20)
53
106
 
54
107
  > 主题:**面板反序缓行的残洞**(外部复验读数:短命子代的最终用量在一种到达序下永远交不出)+ **缓行遇同代 running 帧的三向定形** + **`spawnName` 透传与 `agentType` 改读诚实来源**(#969 提货,CC-95)。**patch**:型面 additive(tick 臂 +1 可选位),公面值导出零增;行为面三笔(见下)。接入面 §75。
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.75.1
38
+ **Version:** 0.76.1
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -317,8 +317,12 @@ public-surface guard checks that last one).
317
317
  | `scripts/run-run-cancel-context-test.mjs` | The run record's `cancelContext` side-note (engine ≥7.87.3) read structurally, and the cause of a `turn_aborted{engine_error}` classified from machine-readable evidence only: `cancelled` (code `cancelled`, with the cancel-time context when present) / `engine_error` (any other failure code, passed through verbatim) / `run_still_live` (the record is not terminal — a dropped stream is a client-side fact, not the run's cause) / `unknown` (never guessed). An absent `cancelContext` reads as *not reported*, never as "not cancelled"; `elapsedMs` is never folded to 0. |
318
318
  | `scripts/run-suspended-reopen-projection-test.mjs` | The durable `suspended` event's `reopened` key read as three distinct states — `reopened` (with the engine's code, verbatim), `not_reopened` (an explicit `null`), `unstated` (key absent or unreadable) — and carried on the HITL bridge's active gate (`currentGateReopen()`), re-read on every `suspended` and cleared with the gate. |
319
319
  | `scripts/run-panel-identity-normalization-test.mjs` | One background subagent has two ids on the wire — the fleet row id tail and the `task_progress` task id (its transcript id). Every panel event goes through one funnel that rewrites the `tick` / `end` task id onto the fleet row's id once a `fleet-row` has registered the key (`transcriptId` first, `parentToolCallId` as the fallback), carrying the original as `wireTaskId` and marking `taskIdOrigin`; an unbound tick whose row has not arrived yet waits one beat (bounded) and is released verbatim on the next tick / `end`, when the buffer is full, or after `MAX_HELD_WIRE_TICK_BEATS` other fleet-row / end / sweep events (a `sweep` itself leaves it alone: there is no row to settle yet); a normalized `end` that carries no cycle identity borrows the registering row's, so a late close of a revived task is recognized as stale; the key table is an LRU (a task that keeps ticking is never evicted by newer registrations); the residency mark migrates with the id and both keys are cleared on settle — except that a stale (previous-cycle) terminal never clears the revived row's mark — so the notification lane can clear it. Once a tick has been delivered verbatim under its UUID, that UUID is the subagent's key: later fleet rows and fleet-side ends are rewritten onto it (the tail kept in `wireTaskId`), so a consumer sees one row in every arrival order; a late tick from a previous cycle is dropped rather than folded into the revived row. |
320
+ | `scripts/run-prompt-assembled-projection-test.mjs` | The `prompt_assembled` frame (one prepare's prompt-assembly manifest) projected to an internal arm and then to the additive `prompt_assembled` chrome event — the per-section / per-block **character** counts, the mounted tool names and `totalChars`, each key present only when the engine really sent it (the frame's `constitution` is deliberately not carried: no consumer asks for it today, and every published key is a contract to keep). The manifest carries **no token counts** anywhere upstream, so this projection mints none: a token figure derived from characters would be an invented number, and the engine's own estimate lives on `context_usage.sections[].tokens` (same id wordlist, joinable). Bad rows are dropped one by one, and a face that loses every row reads as an absent key rather than an empty array — so an absent face means only "this event carries no readable view of it" (an absent upstream key, an empty array and a fully filtered list all land on the same shape) and is never reported as a diagnosis about the engine. `blocks[].id` and `sections[].id` are two different wordlists with a many-to-one relation, and the token join against `context_usage.sections[].tokens` only holds when both sides carry a section view. A frame with no readable composition key at all is malformed, ids and slots are read as an open set, one chrome event per frame with zero transcript rows, several prepares per task are all handed over (de-duplication — "take the last one" — is the host's move), and the lane is told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). Both entry points obey the same rule: the adapt layer rebuilds every row too, so a host pipeline (or a replayed transcript) that feeds the raw frame straight into `adapt()` cannot smuggle extra keys (`tokens`, digests, aliases), a negative `chars` or a `null` row into the chrome payload, an empty array does not count as a composition face, the identity keys are snapshotted once on both paths (read exactly once each, a throwing accessor rejects the whole frame — reading one twice is what lets an accessor frame land on a different lane on each path), and the two paths are compared verbatim so the two readers cannot drift. |
320
321
  | `scripts/run-compaction-outcome-projection-test.mjs` | The `compaction_outcome` frame (a compaction that did **not** end as compacted: mooted by the task ending, failed, …) projected to an internal arm and then to the additive `compaction_outcome` chrome event — `outcome` required and verbatim (open set), `trigger` / `reason` present only when the engine sent a non-empty string, malformed frames dropped, zero transcript rows, the lane told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). |
321
322
  | `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
323
+ | `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string | null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
324
+ | `scripts/run-approvals-feed-unknown-test.mjs` | The approvals feed tells three states apart: **N items waiting**, **nothing waiting**, and **this fetch did not come back, so we do not know**. Every way a fetch can fail (the call throwing or rejecting, a body that is not an object, a `livePending` section that is not an array — including the `null` seen in the field, a `pending` that is not an array or holds a malformed row) publishes `{kind:'unknown', why, at, mode}` on the subscription — never an empty snapshot and never silence. Real snapshots carry `kind:'snapshot'`; `snapshot()` still answers only with the last real one (a fact about the past) while `reading()` answers whether it is current (`unobserved` / `present` / `unknown`). Recovery always publishes a real snapshot again, even when the contents are byte-identical to before the failure. An unknown reading is never counted as zero: the awaiting-decision counts read `null`, the view is empty, and the tracker reports no removals, so cards on screen are not retracted for a failed fetch. Retry, backoff and circuit-breaking are unchanged. |
325
+ | `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. |
322
326
  | `scripts/run-panel-cycle-identity-test.mjs` | The **cycle identity** on agent-panel events and the fleet ledger's **departure read-out**: a background agent may be revived under the same id, so `fleet-row` and `end` events now carry the wire's own `cycleSeq` / `startedAt` when present (absent means the row has no notion of generations, never "generation one"), `isStaleEngineAgentPanelEnd` is the single rule for ignoring a late `end` from a previous cycle (only when both sides carry a comparable identity; absence never drops a real terminal), a changed `cycleSeq` is a new cycle for usage stickiness and buffer coalescing, the notification lane carries `cycleSeq` only when the wire really sent `seq`, and `task_remove` frames reach the host through `onTaskRemoved` with `removeReason` / `cycleSeq` verbatim, a stale previous-generation removal leaving the newer row in place. A terminal row held back because the consumer has no such row yet is also released by the keys it carries itself (its transcript id, or the delegating call id of a subagent already on screen under its wire id), since the key tables are only written once a row has actually been published — the release still goes through the one funnel, so the subagent stays one row; a running frame arriving after such a held terminal row is a stale snapshot when both sides carry a comparable generation and it matches (no event, the held row keeps its final usage), a revival when the frame is provably newer (the held row is dropped), and is treated as a revival when neither side can be compared. A subagent lifecycle event carries `agentType` only from an honest source — the fleet row's own agent type, recorded before it is folded into the row label — and omits the key when there is none, never substituting the display name |
323
327
  | `scripts/run-workflow-size-warning-test.mjs` | The **workflow size warning** verdict shared by every host footer / panel: a three-state result (`warn` / `ok` / `unknown`) read off the optional fleet view keys, where an unknown size is never reported as a normal one (absent `totalCount` / `tokens` without positive over-cap evidence is `unknown`, naming the missing keys), positive evidence on either axis wins regardless of absent keys, the per-agent denominator uses the engine's started count only when it is not below done+failed (a smaller value is a stale reading), otherwise falls back to the done+failed lower bound only when both keys are present — and a lower-bound denominator only yields an upper bound of the projection, which can prove *within cap* (`ok`, flagged) but never *over cap* (`unknown`, with the upper bound exposed) — and the prior is used only when the engine itself reports zero started agents; cap precedence env > explicit guideline > default, prototype keys never act as a guideline, the env reader is pure and does not fall through to the second name on a bad first value; caps, guideline table, env names and the three copy variants are single-sourced |
324
328
  | `scripts/run-approvals-stream-live-capability-test.mjs` | The engine's live-approval-push self-description (`capabilities.approvalsStreamLive`, engine ≥7.87.1), read the same four-state way as its four sibling capability readers: an absent key is reported as not reported (never folded into `false`), the value must be a strict boolean, and the one decision the feed consumer needs — whether it must keep pulling suspended asks itself — is answered by `livePendingNeedsReconcile`, which only says no when the engine explicitly says it pushes. |
@@ -377,6 +381,7 @@ public-surface guard checks that last one).
377
381
  | `scripts/run-approval-frame-chrome-arms-test.mjs` | The two in-stream approval frames finally reaching every host through the shared pipeline instead of one shell's private branch — the shape of a layering defect: hosts that only consume the package could not rebuild their pending cards after a reconnect, and did not clear a card the engine had withdrawn. The payload is deliberately carried as the **envelope** the upstream types declare rather than the first-version card: the stream parser applies no predicate, so narrowing here would let a legitimately newer frame pass as the older shape and invite consumers to read keys a newer card never promised. The guard therefore pins that every open key survives untouched, that an unknown version still passes through, and that narrowing is left to the host's own predicates — with the fallback being a generic card and a person, **never** an automatic denial. A frame whose version cannot be read at all is reported as malformed rather than dropped in silence, because both frames carry user-visible decisions and state changes. Both arms are registered as **required** host duties, and their duty text names the load-bearing rules a host would otherwise have to rediscover: which predicate to narrow with, that the reconnect preamble — not a replayed historical frame — is the authority on which cards exist, and that a withdrawal frame can be lost entirely. Unlike the sibling arms, these carry **no** sub-stream cutoff: an approval raised under a delegated call still has to reach a person, and filtering it by ownership is the host's job, not a reason to discard it. Finally the upstream bytes that justify the envelope discipline are checked to still be there, since the whole design rests on them |
378
382
  | `scripts/run-terminal-status-vocabulary-test.mjs` | One place that decides whether a run has **ended** and whether it ended badly — written because that judgement had already been hand-copied three times, so the day the engine added a word for *the agent itself reported it cannot continue*, every copy missed it and a panel settled a self-reported failure as a success. The distinction the table exists for is pinned from both sides: that word belongs in it, while the two words meaning *waiting for a person to decide* deliberately do **not** — reading those as endings would bury a run that is actively waiting on the reader. A word this client does not know answers *no*, and the guard states plainly that *no* is not evidence of success: proving success means reading the positive side, so negating this predicate is the very mistake that caused two earlier incidents. The fleet lane gets the same treatment from the other direction: a workflow parked on a durable approval used to fall through to *running*, leaving the person with no hint that a card was waiting, and it now lands on the same rendered word the task lane already used — same fact, same word, checked end to end on a real row. Why the word was added directly rather than carried as a private superset key is checked mechanically against the upstream declaration being open, so the day it closes this reds and the decision gets revisited. The residue sweep is the point: the source tree must contain **no** further inlined copy of the judgement, each of the three former sites is checked to really read the single predicate, and the one reviewed exemption carries its reason **and** a liveness assertion, so an exemption whose justification expires cannot quietly keep standing |
379
383
  | `scripts/run-terminal-word-source-test.mjs` | Two tables of ending words, kept apart by **who owns them** — because they used to be one. The engine's own closed set of reasons a run ended, and the server's set of row states a run can finish in, overlap in three words but not in all of them: one word for *something outside stopped it* exists only on the server side, and one for *it paused and can be resumed* exists only on the engine side and means very nearly the opposite of an ending. Merged into a single list, those two sources became indistinguishable, so a new word on either side looked the same as a new word on the other, and the safest-looking move — folding the unknown word into a known one — is the exact mistake that has caused incidents here before. The engine-owned table is checked as a **copy, not an opinion**: it is reconciled word-for-word and in order against the installed engine package, read from both its declaration and its runtime bytes with the two required to agree, so the day upstream adds a fifth reason this reds before anything ships. The two dividing words are each pinned from both sides, including against the upstream declaration directly rather than only against this package's own list. Why the table is copied rather than re-exported is itself an assertion with an expiry: the day upstream publishes the set as a value, this guard reds and the decision gets revisited. The renamed tables leave **no alias** behind, since an alias would let a reader keep consuming the merged list and the split would have bought nothing |
384
+ | `scripts/run-terminal-table-provenance-test.mjs` | Several tables answering *has this ended*, which until now only asserted their own current wording rather than that the wording was right — a snapshot equality passes forever even the day upstream adds a word this package never learns about. Each is reconciled against a named upstream source instead, one comparator shared across all of them rather than one copy per table: a notification's terminal words are the engine's own closed set minus its one live word; a sub-agent tick's terminal words are the engine's own inline status literal minus *running*; a run row's terminal words must **cover every** engine reason a run can end — missing one is the exact failure mode that once let a client retry a connection until its budget ran out while the ending sat unread in the row the whole time — plus one explicitly named legacy word the engine's current declaration no longer carries; a fleet row's terminal words are pinned to **exact equality** with the transport's own status set minus its known non-terminal words (an adversarial pass found the earlier one-directional form let a real terminal word be quietly deleted from this side and still pass), and separately pinned against that set's current member count so the day it changes a person has to look. A sixth table has no clean upstream owner at all — a fact this guard states as a finding, not hides: every status union collected elsewhere in this suite is checked to barely overlap this table's words, with the overlap threshold itself a living assertion that reds the day something upstream finally does match closely enough to replace the guess, and the table is separately checked against every non-terminal word gathered — including one meaning *durably paused*, sourced from the engine's own outcome vocabulary rather than any of the other five unions, after the same adversarial pass found a caller-documented non-terminal word this boundary had missed. A seventh pair, found by the same pass sweeping the whole tree for the same shape of hand-copied table, answers a related but distinct question — whether a session's claim on a run has been released or is still held — and is pinned as two complementary halves of one upstream set: released-minus-one-named-legacy-word and held must partition the transport's status set exactly, so a real state going missing from either side is caught the same way a fabricated one would be. Every extraction in this guard parses real syntax rather than pattern-matching quoted text, so a comment mentioning a word never counts as that word being present, and single- and double-quoted members are read identically |
380
385
  | `scripts/run-workflow-park-truth-projection-test.mjs` | The read face for *which approvals a workflow run left parked* — and the credential that must never ride along with it. Upstream strips the redemption token from that response, and this package's reader is built so the token **cannot** come back: each row is assembled field by field from the three identity keys, never copied wholesale, so an extra key appearing upstream is structurally unable to reach anything this package hands a UI. The guard proves that rather than asserting it — a poisoned row carrying a secret is read, and the secret is searched for across the **entire** serialized result, with the same search proven to find it in the input so a blind search cannot pass; renaming the credential key does not help it through, because the rule is *only these three*, not a blocklist; and the reader's own source is checked to contain no object spread, since one such line would quietly void all of it. The other half is an absence distinction with opposite consequences: a record with **no** parks field at all was written by an older engine and proves nothing about whether approvals are waiting, while an empty list is a positive statement that none are — collapsing those two would let a run whose parked approvals cannot be proven be resumed anyway, so they are kept literally distinguishable, and a payload whose rows are all unreadable answers *unknown* rather than *none*. The four refusal codes for this family are checked code by code against the engine's real bytes, never matched by name prefix, and the older umbrella code they were split out of is asserted to still be **alive** — treating the whole code as retired would make a family of real refusals vanish silently **0.68.3 (core 7.18.0):** two more keys ride the same projection duty as `parks` itself: `originUnconfirmed: true` on a row (never `false`; absence is the confirmed state) and `resumeAdmissionIncomplete: true` on the run (presence means "not a resume base"). Dropping either would turn a refused record back into an admissible one, so the guard pins both, including that neither folds into the other **0.69.1 (CC-12):** both keys now also ride the projected `WorkflowRunState`, so a host that only sees the projection can render them |
381
386
  | `scripts/run-retired-vocabulary-census-test.mjs` | Whether a retirement really happened. When upstream removes a family, a downstream package can cut it out or keep a courteous alias — and the alias is the worse outcome: three clients keep writing branches for something nobody emits, and a status line advertises a state it can never reach. Choosing the clean cut only means something if a guard holds it, since a comment saying *retired* is not an exit code. Each registered entry is held two ways: the name must be gone from **code positions** in this package (comments stripped first, because the explanation is supposed to stay) and off the published surface, and — the half that keeps this from being self-congratulation — it must really be gone **upstream**, since that is the entire reason it was removed here; if it comes back, the disposition deserves reconsideration rather than silence. The scanner proves it can speak by finding a symbol that is genuinely present before any absence is believed, and distinguishes a mention inside a comment from one in a string literal, which is exactly the form being cleared. A closing check runs the other way: the retirement **story** must remain in the comments, including a promise this package made earlier and has now had to withdraw — deleting the history alongside the code is a bad way to satisfy *zero hits*, and leaves the next reader with code that has no reason |
382
387
  | `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
@@ -388,6 +393,11 @@ public-surface guard checks that last one).
388
393
  | `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
389
394
  | `scripts/run-text-segment-authority-test.mjs` | The **authoritative segment replacement** on `text_end` (L-310, server >=7.75.3). `text_end.content` now goes through the same redactor as `result` and the ledger while `text_delta` stays verbatim, so the two **may differ** — an answer that quoted a credential used to be committed to the local transcript in its unredacted form, because the arm only forwarded the boundary signal. Six timing shapes are pinned, two of which an adversarial review reproduced against the installed engine's real bytes and which the first design got wrong in both directions: a second boundary in the same turn (the per-block case on one provider lane) used to make the first segment's prose vanish, and a boundary that arrives *after* the tool card (the other lane emits it at finalize) used to be read as "this package never handled that segment" and reported nothing at all. Three additive keys, all never-false; the two shapes that look alike are told apart by the second one, because the host's action in them is the opposite. The end-to-end legs drive the real pipeline without hand-inserting a segment commit — doing so is exactly what hid the first defect. A second review round then found two combination timings on top of the first fix — a tool card followed by *more* deltas in the same segment, and a byte count that had been documented as a message count — and both are pinned here too. A third round caught a length that the prose called bytes while the code returned UTF-16 units — harmless in ASCII, and on CJK text enough to leave the credential on screen — plus a backfill ledger that had to be kept in step, so the terminal frame does not re-render the segment a second time — kept in step only where the whole stretch sits in one message, because those ledgers are per-message and a fourth round showed that writing across them charges one message's prose to another. A fifth round settled the whole class into one invariant the guard now checks against the previous release's behaviour: this package only rewrites bytes it is still holding in the current message — once a segment has crossed a package-side boundary it emits the three keys and changes nothing else **0.68.2 (CC-01):** the segment identity is now minted here, not by the host: every committed assistant text row carries a top-level `_sema_segment_id` (stamped once at the `adapt()` exit, so the durable whole-message leg and the streamed-segment leg are covered alike; thinking blocks, tool_use-tailed rows and chrome events are left byte-for-byte), `text_segment_end` carries the same value as `segmentId` before rotating, subagent boundaries never rotate, and a replayed stream yields the same identities. Three mutations (no rotation / no stamping / stamping tool_use rows) each turn the guard red **0.69.0 (CC-02):** the same authority replacement now covers the reasoning face (`reasoning_end`, server >=7.77.0): a thinking block still buffered is swapped whole and its live tail recomputed; one already committed at a boundary (the usual timing, since the first text delta commits it) is left untouched and the host is told the row to replace by its uuid, never re-emitted. Subagent boundaries are ignored and the text-segment identity does not rotate **0.69.1 (CC-09):** the run-stream replay guard still drops a frame whose event id was already seen, but it now reports the drop through the host's dropped-frame sink as `duplicate_seq` instead of vanishing silently (server 7.77.0 reuses the first reasoning delta's id for `reasoning_end`, so that authoritative segment is lost on the print lane until 7.78.1); the interactive adapter has no such guard and keeps receiving it **0.69.1 (CC-10/CC-11):** subagent segment-end frames are fenced on all three identity keys (a frame carrying only `sourceTaskId` no longer masquerades as the leader's), and a reasoning segment that spans tool cards now hands the host every committed row it covers (`committedUuids`) so nothing unredacted is left behind |
390
395
  | `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the 74 suites above actually turn red when the material they check really breaks — a census had found 16 of them clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. Seven guards of the same shape and 51 behaviour/projection suites are catalogued rather than rehearsed this round — see `docs/GATE-NEGATIVE-CONTROLS.md` for the full table, the reasons, and a one-minute manual replay recipe for each blind one. The suite cross-checks its own case count against that document's row counts in both directions, so a case quietly dropped from the array without the document following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
396
+ | `scripts/run-engine-cap-reader-factory-test.mjs` | The one shared implementation behind every capability reader's four ports (cache, generation gate, probe tee, invalidation), exercised as a table: every reader in the table runs the *same* criteria (the table length is the source of truth, and a roster check fails the gate if any source file calls the factory without having a row) — the four states, a throwing projection treated exactly like an unreadable one (and never escaping the tee), the generation rules (a stale generation is dropped before the projection even runs; a projection that changes the generation mid-flight cannot overwrite the newer value, whether it returns or throws; the caller's `opts` is snapshotted once; and omitting the generation still writes, because that supply is additive and this refactor does not quietly tighten it), the top-level key's getter being read exactly once, a freshly minted "unobserved" reading on every miss (two misses are never the same object, so a consumer that mutates one cannot taint another base URL), one independent table per reader that never take each other down, invalidating one base URL leaving every other base URL's reading untouched, `forget`/tee being no-ops on an empty or non-string base URL, the read anchor being resolved dynamically, and the deliberate split in how presence is judged per reader. |
397
+ | `scripts/run-mcp-liveness-test.mjs` | The engine's **liveness observation** about each MCP server it hosts (`wiring_manifest.mcp[].liveness`, engine-side from core 7.24.3 / server 7.91.2): one reader, one word list, one leg-level verdict. The cell answers *can this server still be reached* — it is not the connect-time verdict beside it, which the engine deliberately freezes (a server that died mid-run still reads `connected`), and it is not a re-dial's judgement either, so `status: "failed"` next to `liveness.state: "reachable"` is a **real row**: the server answered the handshake and answered with a protocol error — up, and misconfigured. The three words are read as a closed set (an exchange completed / it was lost in transport or the clock / what came back does not answer the question), and a word outside it is malformed rather than rendered, because a word nobody upstream has defined is not a sentence worth putting on screen. **Absence is the fourth reading and is not one of the words**: it means *no liveness record is available*, which on the wire covers a server this leg never reached, a declaration that could not be dialled, an older engine, and a record the projection ahead of us dropped — all indistinguishable, so it is never read as "we looked and could not tell" (a strictly stronger claim), never as healthy and never as off. The failure-class footnote rides the unreachable word only, and one that turns up anywhere else, or that is malformed, loses **just the footnote** while the word and its timestamp stay: the honest reading is then "cannot be reached, reason not given", not "this record is broken". Malformed never becomes healthy: a cell that is present but unreadable marks its row and pulls the leg-level verdict back to *cannot tell*, since letting it sit beside a reachable row would report a leg as reachable on the strength of a record that may well have said the opposite. The verdict takes the worst fact first rather than a majority or the newest reading, carries no server count — so there is no fabricated zero to be read as "no problems" — and its timestamp belongs to **that leg's** observation, not to now: the engine runs no probe and adds no traffic of its own, so this is a per-leg snapshot rather than a heartbeat. The replayed roster on the session panel and the live leg go through the same reader, and the panel's own deployment-side rows carry no liveness position today, so the verdict does not borrow a word from that face. The word list is bitten in both directions where an installed witness exists and the absence of one is itself asserted against the installed engine's version, so the day it ships the comparison starts on its own; the day the wire types declare the cell, the guard turns red and asks for the anchor to move there |
398
+ | `scripts/run-peer-lane-rules-write-capability-test.mjs` | Two more engine self-descriptions read the same four-state way as their seven sibling capability readers (`capabilities.peerLane`, `capabilities.permissionRulesWrite`): an absent key is not reported (an older engine that predates the position, never folded into `false`), `true` is present, `false` is a positive absent (the cross-session lane not being mounted on this deployment, or this particular call not being able to reach the tightening-direction write entry), and anything else is unreadable and drops the cell. Each carries its own single-source verdict (`peerLaneAvailable` returns `yes`/`no`/`unknown`; `permissionRulesWriteAvailable` collapses to a plain boolean, present being the only `true`). The write-entry position pairs with a boolean convenience port in the persisted-rules module, and this guard pins that port to derive from nothing but this one reader's own reading — never a conjunction with the lane-reachable position, and never a second read of the deployment-level existence signal the revoke surface uses (the two are documented as reading differently on purpose): a deployment where the lane answers true but the write entry's key is simply absent (an older binary) must still come back `false`, a deployment where the write entry answers true while the lane key is entirely unseen must still come back `true` (proving no silent conjunction crept in), seeding only the general capabilities cache — never this reader's own feed — must still come back `false` (proving the convenience port cannot be satisfied by the wrong table), and passing an explicit `undefined` base URL must still come back `false` even while a different, already-installed engine target answers `true` for the same position (an adversarial pass found the naive forward of that parameter falls through to the reader's own convenience default, silently answering for whichever engine happens to be installed rather than the caller's absent target — the fix routes an explicit absence through the same empty-string path the reader treats as unobserved). |
399
+ | `scripts/run-lane-proof-identity-test.mjs` | The **instance identity of a lane proof**: the main-lane proof is minted fresh on every emission. Previously a single module-level constant object was handed both to `laneOf(an unregistered task id)` and to some fifty main-lane emission points, so two unrelated consumers — across adapter instances, across streams, across turns — held the same object: writing a card id onto one of them was readable on the other, and the four opening main-lane events changed together. Nothing in this package writes to a lane proof and the known consumers only read it, so this is an **aliasing hazard on a published output surface** rather than an observed corruption — a consumer that uses the proof as an identity key, for dedup, or as a view-layer identity would conflate two unrelated rows without writing a single byte, which is precisely the half that freezing the object would not solve. The gate therefore anchors on instance identity: two independent adapter instances, two rows inside one instance, the same id read twice, and two arms in one beat are each distinct references; mutating one leaves the others byte-identical; and the subagent lane, which already minted fresh, is the control that proves the criterion discriminates. The main-lane **value** is unchanged — an unregistered id still answers `{lane:"main"}` with exactly one own key and still emits its events, so absence is not turned into a second kind of absence — with ordering pinned three ways (registered-then-read, read-then-registered with no retroactive edit of an already delivered proof, the same id twice) and the id failure classes pinned four ways (unregistered, empty string, absent, non-string, the last two emitting no panel event at all rather than an ownerless proof). Where one row emits **two** events — the terminal-tick and card-close legs, which each yield a lifecycle stop and a panel end — the attribution is decided once (a consumer binding a card between the two yields must not split one row across two lanes) while each event still gets its own proof, so a host consuming them one at a time cannot poison the second before it is even yielded. The run stream leg is covered as the same shape, and a syntax-tree check forbids reintroducing a module-level lane-proof object literal or a module-level `LaneProof`-annotated binding (judged on the type node, not on text, so a compile-time pin tuple that merely mentions the type is not miscaught), backed by a type-checker pass that also catches an un-annotated module-level cache such as `const x = mainLane()` while letting the callable factory itself through, while the module-private three-state sentinels of the untrusted read are frozen instead — only `Object.freeze` counts, never `Object.seal`, which still permits writes to existing keys — their exposure being confined to one module |
400
+ | `scripts/run-memory-entries-wire-test.mjs` | The two memory-governance capability bits and the three memory-entry response readers. Each bit (`capabilities.memoryCompliance`, for the entry-provenance and erasure endpoints; `capabilities.memoryOrigin`, for the external-origin listing and clearance endpoints) is read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `false` — an older engine simply does not answer, and the right next step is to try the endpoint and read its 501), `true` is the face being mounted, `false` is a positive "not on this deployment" (the wire does not distinguish a backend without control-plane ownership from an empty operator roster, so the wording never guesses which), any non-boolean value is unreadable and drops the cell instead of being folded into "absent", and a capabilities body that is not an object at all is unreadable rather than "not reported". The two bits deliberately stay **two** readers with two separate per-engine tables, because the engine deliberately keeps them two separate product faces even while they happen to carry the same value today: feeding one an unreadable body, or invalidating one, leaves the other's reading untouched, and a body where one is on and the other off is answered one bit at a time. The entry-export reader narrows each row on its own (an empty array really is zero rows, a non-empty array with nothing readable in it is reported as unreadable rather than as "no rows", and partly bad rows are kept with a dropped count), reads the external-origin marker as three states rather than a boolean (the two structural carriers mark a row; a row whose frontmatter cannot be read, or which carries the third, suspended-form carrier, is undecidable, because the judge for that carrier lives in the engine and this package refuses to mint a second copy of it), and treats an unreadable "is this the whole scope" flag as "not the whole scope". Its verdict port implements — in code, not in a comment — the rule that an empty answer is never a clean store: the caller must state whether the request declared origin-awareness, because this endpoint withholds marked entries by default and the two bodies are shaped identically, so without that statement an empty answer is only ever "unknown"; the affirmative answer is scoped to the one named scope and carries that scope with it, and the type has no store-wide arm at all. The erasure receipt reader keeps three things apart that are easy to collapse: "this call erased nothing" (a real receipt whose erased list is empty and whose not-found list explains why, per id), "a 200 with an empty body", and "a body that could not be read" — at the reading, the counting and the verdict layer alike; it refuses a version envelope it does not recognise instead of reinterpreting it, treats the three closed vocabularies as closed (an unknown word is unreadable, never folded into a known arm), keeps an unreadable binding as unknown instead of claiming "unbound", passes the "history cannot be judged" flag through as four states (set, explicitly unset, absent, and present-but-unreadable — an unreadable flag is kept distinct from an absent one, and the history verdict then answers "unknown" rather than the stronger claim), and answers the replay question as three states so that the degraded lane is never retried automatically. The clearance receipt reader carries the cleared marker through verbatim and says separately whether it was reported at all. Every array in every response is snapshotted once — the length is read exactly once and each index exactly once, rather than iterating the caller's own iterator — because an array that reports one length while being walked and another afterwards could otherwise have a marked row quietly dropped while the "was anything unreadable" check saw nothing, which ends in calling the scope clean; an array that reports an absurd length is reported as unreadable rather than silently truncated to its first rows. All three readers never throw. |
391
401
 
392
402
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
393
403
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**