@sema-agent/client-core 0.68.2 → 0.69.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -49,6 +49,97 @@
49
49
  > 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
50
50
  > commit 漏转在发布前就红,不再靠人记。
51
51
 
52
+ ## 0.69.1(2026-09-16)
53
+
54
+ > 主题:0.69.0 的**安全面收尾**(cli 接入回执打回 + 对抗复审在合并树上复现):CC-09 runStream 去重**留痕** / CC-10 段末帧子流断闸改按三键 / CC-11 跨工具卡推理段归属交全 / CC-12 park 两新键进视图 + 文档批。CC-04 技能帽退役改预算式改随 server 7.78.1 进 **0.69.2**。接入面 `docs/INTEGRATION-CLIENTS.md` **§37**。
55
+
56
+ ### Fixed
57
+
58
+ - **`runStream` 的 seq 去重丢帧从此留痕**(CC-09)—— 修前 `adapter/runStream.ts` 对「已见过的 durable seq」是一条裸 `continue`
59
+ (零痕)。server 7.77.0 给 `reasoning_end` 的 SSE `id:` **复用**该段首枚 `reasoning_delta` 的 id,于是 `-p` / runStream 链上
60
+ 权威推理段被当作重放帧**静默**丢掉(`text_end` 自带 id 故不受影响);归口 = server 改铸唯一 id(7.78.1)。本层**不改去重律**
61
+ (契约 02 §1.1),只把这次丢弃交给 `ctx.onDroppedFrame({ type, why: 'duplicate_seq' })`(开集 why 词,措辞表加行)。
62
+ 🔴 已知边界:**交互车道 `adapt()` 不经这只去重**(源码直证),`thinking_segment_end` 在交互车道照到;`-p` / runStream 链在
63
+ server <7.78.1 上权威推理段不到端,宿主只能靠留痕知道。门:`run-text-segment-authority-test.mjs` U 段。
64
+ - **段末帧子流断闸改按身份三键**(CC-10;cli 接入回执 ① 实抓的 fail-open)—— 修前 `segmentEndProjection` 只透
65
+ `parentToolCallId`、两臂只按它断闸,而 sdk `EventIdentity` 的子代身份是三键(`parentToolCallId` / `sourceTaskId` / `bgAgentId`,
66
+ core 顶注「stamped ONLY on … AS A SUB-AGENT」):只带 `sourceTaskId` 的子代 `reasoning_end` / `text_end` 被当 leader 帧投上来,
67
+ 端据它换 leader 的思考/文本行 = 子代覆盖 leader。现在三键原样穿上内部臂(`text_end` / `reasoning_end`),任一在场即断,
68
+ 在场非串 ⇒ malformed。门 TSA V1–V3。
69
+ - **跨工具卡的推理段归属交全**(CC-11;回执 ②)—— 修前 `feedThinking` 开新块清对照物,一段推理跨工具卡时 `thinking_segment_end`
70
+ 只交末条已提交行的 `committedUuid`,前面的行无人说属于这一段(端只能整块换末条,其余未脱敏原样留盘)。现在段内已提交思考行
71
+ 按序累积(turn 内有界,`reasoning_end` 用过即清),事件新带 **`committedUuids[]`**(全部按序;`committedUuid` = 末条兼容位),
72
+ `committedPrefixLen` = 全部之和;判分歧按 `concat(rows) + buffer` 对 `content`:前缀对得上 ⇒ 只换还押着的后段(b1),对不上 ⇒
73
+ 整段归宿主(b2,缓冲清空不再交)。**宿主义务改口**:`committedPrefixDiverged` ⇒ 删 `committedUuids` 全部行、首行位置放 `content`。
74
+ 门 TSA V4–V10;臂表义务同文。
75
+ - **park 两新键进 `WorkflowRunState`**(CC-12;回执 ④)—— 0.69.0 只出了 `readWorkflowParks` / `readWorkflowResumeAdmissionIncomplete`
76
+ 两读器,`projectWorkflowRun` 逐键铸视图时丢了它们、`LiveWorkflowController` 不出裸体 ⇒ 端渲染面结构性造不出。现在
77
+ `WorkflowRunState.parks?: readonly WorkflowParkRowView[]` / `resumeAdmissionIncomplete?: true`(additive,由两读器铸,缺席两义同读器)。
78
+ 门 park H 段。
79
+ - 三件文档批(回执 ③⑤⑥):`mcpPanel.ts` jsdoc「四句」订正为五句;`sdkWireTransit` 型面转口加 `SessionPermissionRules`
80
+ (`loosenReasons` 两参的型);§36d 补「`hasCoreValuePorts()`(整袋)与 `coreValuePortMisses()`(逐件)在部分装载时分歧」与
81
+ 「镜像按 core 7.18.0 写,7.17.x 同形」;§37 注明 §35z #15「候 sdk 9.5」已由 §36z #18 收口。
82
+
83
+ ## 0.69.0(2026-09-15)
84
+
85
+ > 主题:**提货批**(原排 0.68.3;因 peer 地板抬升改 minor)—— sdk 9.4.0 / core 7.18.0 / settings-schema 2.0.0 三家换钉,
86
+ > 推理面接上 `text_end` 那一套权威替换律(CC-02),core 7.18.0 的通告码与 park / 分类器新键逐字投影(CC-05)。
87
+ > 接入面详见 `docs/INTEGRATION-CLIENTS.md` **§36**(逐键处置表 §36z)。
88
+
89
+ ### BREAKING
90
+
91
+ - **peer `@sema-agent/sdk` 地板 `>=8.8.0` → `>=9.4.0`**(peerDependencies / lockfile / README 地板句 / `run-sdk-floor-test.mjs` 的 `FLOOR` 四处同批抬齐;本包 devDep `^9.4.0`)。理由:`reasoning_end` 帧型与 `McpStatusPanel.lastLegMcp?` 自 9.4.0 起才在型面上,本包按 9.4.0 的 `SegmentEndFields` 单源接 `text_end` / `reasoning_end` 两枚帧。**运行期零 BREAKING**:老 server(<7.77.0)不发 `reasoning_end`,本包一字不改照跑。
92
+ - devDep `@sema-agent/core` `~7.17.1` → `~7.18.0`(零值级直连,BREAKING 的 `ReadFace` 族改名对本包零影响);devDep `@sema-agent/settings-schema` `^1.10.0` → `^2.0.0`(本包零 src import,两只门按真字节读,提货零动作)。
93
+
94
+ ### Added
95
+
96
+ - **chrome 臂 `thinking_segment_end`**(CC-02;server ≥7.77.0 S-310 `reasoning_end`)—— 思考块的**权威段替换**:
97
+ `content` 经与 `text_end` / `result` 同一只脱敏器,本包在臂上整段替换思考缓冲 + 活体尾巴;三个 additive 键与
98
+ `text_segment_end` **同名同律**(never-false),外加 `committedUuid`(思考块无 idle-flush,已提交即整块 ⇒
99
+ `committedPrefixDiverged` 在场时给那条已 committed 思考行的 `uuid`,宿主整块换)。`CHROME_ARM_TABLE` 本臂
100
+ `required: true`(不接 = 未脱敏的推理字节留在本地转录,类目同 §34e ①)。内部臂 `reasoning_end` 进
101
+ `INTERNAL_SDK_ARM_TYPES`(⚠️ print 车道对内部臂「要么吞、要么转成真 CC 帧」,壳按 `text_end` 同律接)。
102
+ - **`readWorkflowResumeAdmissionIncomplete(run)`**(CC-05;core 7.18.0 #755)—— `resumeAdmissionIncomplete` 在场 = 不是
103
+ resume 基;与 `readWorkflowParks` 同一条投影义务(丢了会把已拒记录变回可 resume)。`WorkflowParkRowView` 新增
104
+ never-false 第四键 `originUnconfirmed?: true`(缺席 = 已确认;凭据仍结构性不可达)。
105
+ - **`GateClassifierRoundView` / `GateClassifierFallbackView`**(CC-05;core 7.18.0 #742)—— `gate.disposition.classifier?`
106
+ 两臂可带的分类器轮记录(`modelRequested` / `modelUsed` / `fallback{from,to,cause}`,`cause` 开集);缺席 = 未知 ≠ 席自答;
107
+ 半只 ⇒ 整只缺席。
108
+ - **`fetchMcpPanel(client, sessionId, { signal? })`**(CC-03 补件)—— 经 sdk `client.sessions.mcp`(自 0.0.75 在场)取
109
+ `GET /v1/sessions/:id/mcp` 体并投成 `McpPanelView`;传输失败 / 体坏 / 空 sessionId 同返 `undefined`(分辨用
110
+ `mcpPanelLastLegDetail` 的 `reachable` 位);不裸 fetch、不做第二套鉴权铸点。收口 §35z #15。
111
+ - **`installCoreValuePorts` + 十个诚实缺席读口**(CC-07;cli L-249 到期桶 core-value 清零)—— core **值级**面十件
112
+ (`resolveAutonomousLoopPrompt` / `loosenReasons` / `combinePolicies` / `createAllowDenyPolicy` / `MCP_NAMESPACE` /
113
+ `RETIRED_TOOL_NAMES` / `protocolOf` / `ruleToolGrammarOf` / `validatePermissionRules` / `DISCUSSION_WORKFLOW_NAME`)的端口
114
+ 注入口,与 `editedRuleTextPrecheck` 同形:本包不能值级 re-export core(barrel 拉 Node 内建,浏览器打包面焊进整台引擎),
115
+ 于是包声明端口 + 缺席读口,Node 宿主在装配根把 core 真实现原样装进来;型面全部结构镜像(零 core import)。
116
+ `hasCoreValuePorts()` / `hasCoreValuePort(name)` / `coreValuePortMisses()` 三存在性口。接入面 §36d。
117
+ - **`ENGINE_NOTICE_CODES` 镜像 +4**(core 7.18.0):`config.write_face_swapped` / `config.write_face_deployment_clamped` /
118
+ `classifier.fallback` / `route.classifier_unroutable`,audience 均 `operator`(52 → 56 码,顺序同源)。
119
+
120
+ ### Changed
121
+
122
+ - 投影口 `eventToSdkMessage`:`text_end` 从 raw 预分派**搬进 `switch`**(sdk 9.4.0 起在 union 里;到期复核按原定计划兑现,
123
+ 行为一字不改),与 `reasoning_end` 共用 `segmentEndProjection`(两处字面量铸点);`tool_roster_delta` 从 `unknown_arm`
124
+ 改为**有痕 `dropped('unsupported_arm')`**(L-315 候 `wiring_manifest.tools` 进切片那一批一起接)。
125
+
126
+ ### Guards
127
+
128
+ - `run-text-segment-authority-test.mjs` **T 段**(31 格):思考面两形(押着整块换 + 活体尾巴重算 / 已整块提交 ⇒ 不再交 +
129
+ `committedUuid`)、零经手、端到端迟到/先到/子流断闸/段身份不轮换/臂表义务;`run-engine-notice-catalog-test.mjs` 换钉当天
130
+ 红抓「上游 56 / 本包 52」→ 绿;`run-workflow-park-truth-projection-test.mjs` G 段 22 格;`run-decide-receipt-test.mjs` H 段 15 格;
131
+ `run-sdk-floor-test.mjs` 地板 9.4.0;typeshape unknown 出境 338 → 339(`readWorkflowResumeAdmissionIncomplete(run: unknown)`)。
132
+
133
+ ### 已知局限(本版新增)
134
+
135
+ - **`tool_progress` / `ProgressMessage<BashProgress>` 投影候 clay 裁定 DV-741**(帧形是非 CC 锚定);`tool_disclosure` /
136
+ `context_usage.sections?` / `WiringManifest.hooks?`·`lsp?` 三项与 L-315 同批设计 —— 四者今走投影口有痕 dropped,不静默。
137
+ - **同名型影子**:sdk 9.4.0 新导出类型别名 `WiringManifestMcpEntry`(`errorCode` 闭集)与本包同名 interface(开集)同名不同形;
138
+ 本包**不**改名(改名对 cli 是导出 BREAKING),消费端两边都 import 时请用 `import type { WiringManifestMcpEntry } from
139
+ '@sema-agent/client-core'`;收敛排 0.70(候 sdk 侧改名或本包加 `…View` 别名)。
140
+ - core 7.18.0 的 wire 键账本 `wire-changes/core@7.18.0.json` **不随 7.18.0 包出**(core #811,结构性:tag 后才算清单),本批按
141
+ core main `d3bc298a` 的权威副本提货;7.18.1 起包内自带。
142
+
52
143
  ## 0.68.2(2026-09-15)
53
144
 
54
145
  > 主题:**归层** —— 三件「本包一份、端各一份」的东西收成一份。都不是新能力,只是把同一句话
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.68.2
38
+ **Version:** 0.69.1
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -67,7 +67,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
67
67
  against — the tables live upstream precisely so this package does not keep a second copy that can
68
68
  fall behind. The browser bundle really bundles the SDK through (the portability guard would
69
69
  exit 3 rather than quietly mark it external).
70
- - The declared floor is `>=8.8.0`, and it is *witnessed*: the guard checks that an actually
70
+ - The declared floor is `>=9.4.0` (raised from `>=8.8.0` in 0.69.0: the `reasoning_end` frame and `McpStatusPanel.lastLegMcp` are typed from 9.4.0 on), and it is *witnessed*: the guard checks that an actually
71
71
  installed SDK at that line still exports every value-level symbol this package imports and still
72
72
  declares `TaskStats.costMicroUsd` (the key `costOrNull` reads). A floor nobody ever ran is a
73
73
  promise, not a contract.
@@ -306,7 +306,8 @@ public-surface guard checks that last one).
306
306
  | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor |
307
307
  | `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
308
308
  | `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
309
- | `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes |
309
+ | `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes **0.69.0:** `fetchMcpPanel` fetches the panel through the SDK client's own `sessions.mcp` call (same transport and auth as every other read) and projects it; transport failure, an unreadable body and an empty session id all come back as `undefined`, never as a fabricated empty panel |
310
+ | `scripts/run-core-value-ports-test.mjs` | The port-injection seam for ten **engine value-level** facilities (autonomous-loop prompt assembly, permission-rule loosening, tool-policy composition, protocol/retired-name/grammar lookups, rule compilation, the discussion workflow name). This package cannot re-export them (the engine barrel drags Node built-ins into the browser bundle), so it declares the ports and honest-absence readers; a Node host installs the engine's own functions verbatim. The guard pins: every reader returns `undefined` when nothing is installed (never a fabricated empty array or default policy), arguments and results pass through by reference, engine errors propagate unchanged, partial installs read partially, restore functions unwind to the previous bag, and the module source has zero engine imports |
310
311
  | `scripts/run-read-face-posture-projection-test.mjs` | The operator-face `readFace: ReadFacePosture` reader (server >=7.65.0). Three ways of "can't say" are pinned to three different, literal sentences, and none of them may read as "nothing is pinned" — that statement belongs to exactly one case, `face: null`, which is a positive fact reported by the engine, not an absence: not having read an operator response yet, having read one from an engine too old to report the key, and the engine actually saying nothing is pinned are three different next steps for an operator and must not collapse into each other. `source` is read as an open set (the server's closed four words plus an escape hatch) rather than narrowed to an enum, so a new word added upstream is not silently turned into a bad reading. The free-text `note` is sanitized and length-capped before it is ever rendered. A companion pure function flags disagreement between this face and the tenant-facing `capabilities.readFace` — silent only when the two actually agree, honest-absent when either side cannot be read at all, never asserting agreement as a fact. The gate's last leg reads the installed SDK's own `openapi.yaml` directly rather than restating the schema in prose, so the package's leniency cannot quietly drift from the real contract |
311
312
  | `scripts/run-display-cap-order-test.mjs` | The order in which untrusted text is sanitised and length-capped, across every mint point that puts an engine- or database-supplied string on a screen. The sanitiser rewrites each invisible character as a six-character escape, so capping the **raw** string first and escaping afterwards hands the screen six times the width that was budgeted — a forty-character allowance becomes two hundred and forty. The guard does not hardcode that allowance, because each mint point wraps its field in different fixed prose and the prose moves: it anchors on the deciding quantity instead, feeding one benign and one control-character input of the same length through the same mint and requiring the second not to come out longer. That criterion is immune to wording changes and stays sensitive to the expansion, and it is `<=` rather than `==` on purpose — a correct escape-then-cap backs the cut off a partially-consumed escape token, so the control-character line is legitimately the shorter of the two, and demanding equality would score that avoidance as a regression. Each mint is bracketed by two positive controls (the input really reaches the screen; the cap really engages) and the expansion predicate is shown to turn red against a deliberately cap-then-escape reference, so an all-green run cannot mean the guard simply measured nothing. The shared mint point is checked directly for the two avoidances it owes — never splitting an escape token in half, which would leave something on screen that looks like the beginning of a complete answer, and never splitting a legal surrogate pair, which would manufacture the very lone surrogate the sanitiser exists to catch |
312
313
  | `scripts/run-seat-task-request-origin-test.mjs` | Where every field of the seat lane's send-message payload comes from, and whether it actually lands anywhere. The seat payload is a closed interface this package mints itself, and most of its fields are meant to ride verbatim onto the engine's request body — two facts nothing used to connect, so both directions could drift in silence. A seat field could be named after a request position that does not exist, in which case a client writes to it, the wire carries it, the engine ignores the whole key, and the screen shows a switch that does nothing; conversely a new request position could arrive with no seat to sit in, which is **structural** absence — the closed set *is* the carrier, so a decision missing from it has nowhere to be put at all, the same shape logged when the effort dial had no seat. The guard turns each field's origin into data: either it names the request position it forwards to, or it is declared seat-local with a written reason, and the two are mutually exclusive. Forwarding claims are then checked against the **installed** SDK's type declarations, parsed rather than restated — a hand-copied list of position names would only ever prove that two transcriptions agree. The parser is held to reading top-level positions only, since a nested option object's inner keys would otherwise be mistaken for positions of the request itself, and it proves that discrimination on synthetic input before any verdict is given. The two subagent fields carry a standing regression pin, and the retention window's inner keys are read from the declaration the same way, so a seat that offers a tunable window cannot offer one the wire has no room for |
@@ -352,11 +353,11 @@ public-surface guard checks that last one).
352
353
  | `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one |
353
354
  | `scripts/run-resume-retry-later-test.mjs` | The two resume refusals that carry a **wait quantity** — the only members of that refusal family that do, which is the whole reason they form a closed set. Carrying a wait is not the same as being the only ones worth waiting on: a sibling refusal in the same family clears on its own and the engine says so in words, it just cannot put a number on it, so *not recognised here* must never be read as *waiting will not help*. One of the two also has a *terminal* upstream branch that arrives under the same code with the distinguishing detail only in prose, so recognition alone is not permission to say "try again": the disposition is decided by **positive evidence** and pinned from both directions — the quota-window code is evidence in itself, the preflight code counts only when the server really supplied a wait (an upstream fact, not a convention: the terminal branch throws with no detail at all, so a wait value cannot reach the client on that path), and a preflight refusal with no wait reads as *undecidable* (say what is true of both branches — nothing was consumed — and leave redeemability to the engine's own line) rather than being rendered as either a retry or an ending. Every other member means waiting will not help (change a setting, relaunch, the retained session is gone), so the recognition is a **closed set of two codes**: widening it to a family prefix would tell half the users to wait and the other half to keep waiting for something that will never arrive, and the negative controls drive exactly those codes through it, plus a same-named code on a different door (the submission-side quota refusal), the two underscore-form siblings, and a code merely quoted inside a message body. The wait value is narrowed to the same domain the server mints it in (a whole number of seconds, at least one): zero, a negative, a fraction and a non-number all read as **no window given** rather than as zero, because a zero tells the caller to retry immediately and the wake-up it would retry is an at-most-once action with real side effects. Reading is structural rather than `instanceof`, since the client is host-injected and the same class name across two bundles is two classes, and a null-prototype plain object must still be recognised. The failure classifier gains this one disposition without any existing one moving, an unknown code still falls to the honest open-set arm and its wait value is **not** believed, and an end-to-end call proves the disposition and the window reach the host while the call itself is still attempted exactly once. The recognised code set is a **frozen array**, not a type-level readonly set: the latter is a plain mutable collection at runtime and the decision reads the same instance, so one `.add` from any consumer would turn a refusal that waiting cannot fix into one that claims it can — the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift |
354
355
  | `scripts/run-model-capability-probe-test.mjs` | Whether a model on the OpenAI-completions lane **thinks**, and whether that thinking can be **turned off** — a question nobody can answer by looking at a model name, and one whose wrong answer costs every later call. The probe is judgement only: the network half arrives as an injected port, so the package mints no URL, reads no credential and never calls `fetch` — pinned by a source-level assertion, because a package that reaches the network once has changed what every host must trust it with. The seven dialect words are a **copy**, reconciled element-wise against the installed engine’s own bytes in both directions, since the words belong upstream and a private table drifts the day a dialect is added; the settings package deliberately declines to restate them, so the table cannot be imported from there and this guard is what stands in for the import. The **order** the dialects are tried in is a public promise rather than an implementation detail — each extra attempt is real money and real latency against someone’s gateway — so the guard pins the exact call sequence a stub records, and reversing it reds on the wasted round trip; the template-parameter spelling leads because an observed gateway keeps thinking, and answers with an empty body, when handed the top-level switch instead. That observation is also why an empty answer is **not** accepted as *thinking is off*: a knob that deletes the reply is not a knob that disabled reasoning, and accepting it would write a spelling into the catalogue that the gateway does not honour. Two dialect words whose request bytes are identical to another’s do not each burn an attempt. The two verdicts that look alike are held apart from both directions: *tried everything, still thinking* requires at least one attempt to have **cleanly answered**, and when every attempt was refused the verdict is *could not tell* instead — and on the unanswerable path the result carries **no** thinking flag at all rather than a fabricated `false`, while the pure write-back returns the very same entry object untouched. A verdict that reasoning cannot be disabled **removes** a previously declared spelling rather than leaving it, since a refuted spelling keeps the engine sending bytes the gateway ignores while the catalogue still renders it as already off. Evidence is lengths, finish positions and status codes only — a planted secret in both the answer and the reasoning channel must appear nowhere in the result, so the record can go into a log or a ticket whole One cross-package premise is checked by really running the other package’s parser rather than quoting its documentation: everything this probe writes eventually passes through the settings schema on its way into a catalogue, and that field is declared parse-transparent precisely so the vocabulary can live on the consuming side. If it ever narrows, the spelling is stripped **silently** — indistinguishable from the probe never having run — so the guard feeds the probe’s real output through the real parser, checks the compat object comes back key for key, and checks a dialect word this client has never heard of survives too. Two shapes that must be rejected really are rejected, since otherwise the survival checks would hold on a parser that accepts anything, and a bare entry is asserted valid first, because the first run of this section reddened on a space in a fixture’s name — a fixture that cannot pass would disguise the real alarm as already having fired |
355
- | `scripts/run-decide-receipt-test.mjs` | What a decision verb actually **answered** — and, more importantly, what it did not. A success response on the newest lane is only an acknowledgement that the decision was accepted for delivery: the approval is still pending, and a client that clears the card on it shows either a ghost card that was already approved or a card that vanished while the decision was lost. So the package deliberately has **no** "was it resolved" predicate — nothing in that body can answer it — only the opposite one, whose `false` is likewise not evidence of resolution; resolution is only ever the next running arm on the stream. The guard pins that inversion in the product source too: the success path must no longer clear the latched gate, while the stream-observing path that really clears it must still be there. The body has four shapes with **no** key common to all of them, so every position is read as honestly absent, and the handoff handle — which run to watch from here on — requires **two** facts together, since either one alone would either point the stream at the run it already had or mint an empty handle. The record of what finally happened to an already-decided action is read through the **same** reader as every other gate record rather than a second copy, and its absence means **unknown**, never *it was allowed* — the two can even contradict each other, so the card says nothing at all when it is missing. The three refusals on that lane each get one distinct sentence and a disposition taken from **why** each was refused rather than from severity: one cannot be helped by re-sending at all, one waits on the host, one just drops an option — and none of them carries a countdown, because the server never mints a wait for them. Recognition is a **closed set**: an unrecognised code on the same prefix returns nothing rather than a guess, since that prefix also houses a safety signal whose whole rule is never to retry automatically, and the recovery handle is read as absent when unreadable rather than substituted from a different identifier that no longer appears on that lane |
356
+ | `scripts/run-decide-receipt-test.mjs` | What a decision verb actually **answered** — and, more importantly, what it did not. A success response on the newest lane is only an acknowledgement that the decision was accepted for delivery: the approval is still pending, and a client that clears the card on it shows either a ghost card that was already approved or a card that vanished while the decision was lost. So the package deliberately has **no** "was it resolved" predicate — nothing in that body can answer it — only the opposite one, whose `false` is likewise not evidence of resolution; resolution is only ever the next running arm on the stream. The guard pins that inversion in the product source too: the success path must no longer clear the latched gate, while the stream-observing path that really clears it must still be there. The body has four shapes with **no** key common to all of them, so every position is read as honestly absent, and the handoff handle — which run to watch from here on — requires **two** facts together, since either one alone would either point the stream at the run it already had or mint an empty handle. The record of what finally happened to an already-decided action is read through the **same** reader as every other gate record rather than a second copy, and its absence means **unknown**, never *it was allowed* — the two can even contradict each other, so the card says nothing at all when it is missing. The three refusals on that lane each get one distinct sentence and a disposition taken from **why** each was refused rather than from severity: one cannot be helped by re-sending at all, one waits on the host, one just drops an option — and none of them carries a countdown, because the server never mints a wait for them. Recognition is a **closed set**: an unrecognised code on the same prefix returns nothing rather than a guess, since that prefix also houses a safety signal whose whole rule is never to retry automatically, and the recovery handle is read as absent when unreadable rather than substituted from a different identifier that no longer appears on that lane **0.68.3 (core 7.18.0):** `gate.disposition.classifier` is read key by key into a named view (requested model, model that answered, and the ladder fallback when one happened); a half-shaped record yields no classifier at all rather than a half view, and absence stays "unknown", never "the seat answered itself" |
356
357
  | `scripts/run-approval-frame-chrome-arms-test.mjs` | The two in-stream approval frames finally reaching every host through the shared pipeline instead of one shell's private branch — the shape of a layering defect: hosts that only consume the package could not rebuild their pending cards after a reconnect, and did not clear a card the engine had withdrawn. The payload is deliberately carried as the **envelope** the upstream types declare rather than the first-version card: the stream parser applies no predicate, so narrowing here would let a legitimately newer frame pass as the older shape and invite consumers to read keys a newer card never promised. The guard therefore pins that every open key survives untouched, that an unknown version still passes through, and that narrowing is left to the host's own predicates — with the fallback being a generic card and a person, **never** an automatic denial. A frame whose version cannot be read at all is reported as malformed rather than dropped in silence, because both frames carry user-visible decisions and state changes. Both arms are registered as **required** host duties, and their duty text names the load-bearing rules a host would otherwise have to rediscover: which predicate to narrow with, that the reconnect preamble — not a replayed historical frame — is the authority on which cards exist, and that a withdrawal frame can be lost entirely. Unlike the sibling arms, these carry **no** sub-stream cutoff: an approval raised under a delegated call still has to reach a person, and filtering it by ownership is the host's job, not a reason to discard it. Finally the upstream bytes that justify the envelope discipline are checked to still be there, since the whole design rests on them |
357
358
  | `scripts/run-terminal-status-vocabulary-test.mjs` | One place that decides whether a run has **ended** and whether it ended badly — written because that judgement had already been hand-copied three times, so the day the engine added a word for *the agent itself reported it cannot continue*, every copy missed it and a panel settled a self-reported failure as a success. The distinction the table exists for is pinned from both sides: that word belongs in it, while the two words meaning *waiting for a person to decide* deliberately do **not** — reading those as endings would bury a run that is actively waiting on the reader. A word this client does not know answers *no*, and the guard states plainly that *no* is not evidence of success: proving success means reading the positive side, so negating this predicate is the very mistake that caused two earlier incidents. The fleet lane gets the same treatment from the other direction: a workflow parked on a durable approval used to fall through to *running*, leaving the person with no hint that a card was waiting, and it now lands on the same rendered word the task lane already used — same fact, same word, checked end to end on a real row. Why the word was added directly rather than carried as a private superset key is checked mechanically against the upstream declaration being open, so the day it closes this reds and the decision gets revisited. The residue sweep is the point: the source tree must contain **no** further inlined copy of the judgement, each of the three former sites is checked to really read the single predicate, and the one reviewed exemption carries its reason **and** a liveness assertion, so an exemption whose justification expires cannot quietly keep standing |
358
359
  | `scripts/run-terminal-word-source-test.mjs` | Two tables of ending words, kept apart by **who owns them** — because they used to be one. The engine's own closed set of reasons a run ended, and the server's set of row states a run can finish in, overlap in three words but not in all of them: one word for *something outside stopped it* exists only on the server side, and one for *it paused and can be resumed* exists only on the engine side and means very nearly the opposite of an ending. Merged into a single list, those two sources became indistinguishable, so a new word on either side looked the same as a new word on the other, and the safest-looking move — folding the unknown word into a known one — is the exact mistake that has caused incidents here before. The engine-owned table is checked as a **copy, not an opinion**: it is reconciled word-for-word and in order against the installed engine package, read from both its declaration and its runtime bytes with the two required to agree, so the day upstream adds a fifth reason this reds before anything ships. The two dividing words are each pinned from both sides, including against the upstream declaration directly rather than only against this package's own list. Why the table is copied rather than re-exported is itself an assertion with an expiry: the day upstream publishes the set as a value, this guard reds and the decision gets revisited. The renamed tables leave **no alias** behind, since an alias would let a reader keep consuming the merged list and the split would have bought nothing |
359
- | `scripts/run-workflow-park-truth-projection-test.mjs` | The read face for *which approvals a workflow run left parked* — and the credential that must never ride along with it. Upstream strips the redemption token from that response, and this package's reader is built so the token **cannot** come back: each row is assembled field by field from the three identity keys, never copied wholesale, so an extra key appearing upstream is structurally unable to reach anything this package hands a UI. The guard proves that rather than asserting it — a poisoned row carrying a secret is read, and the secret is searched for across the **entire** serialized result, with the same search proven to find it in the input so a blind search cannot pass; renaming the credential key does not help it through, because the rule is *only these three*, not a blocklist; and the reader's own source is checked to contain no object spread, since one such line would quietly void all of it. The other half is an absence distinction with opposite consequences: a record with **no** parks field at all was written by an older engine and proves nothing about whether approvals are waiting, while an empty list is a positive statement that none are — collapsing those two would let a run whose parked approvals cannot be proven be resumed anyway, so they are kept literally distinguishable, and a payload whose rows are all unreadable answers *unknown* rather than *none*. The four refusal codes for this family are checked code by code against the engine's real bytes, never matched by name prefix, and the older umbrella code they were split out of is asserted to still be **alive** — treating the whole code as retired would make a family of real refusals vanish silently |
360
+ | `scripts/run-workflow-park-truth-projection-test.mjs` | The read face for *which approvals a workflow run left parked* — and the credential that must never ride along with it. Upstream strips the redemption token from that response, and this package's reader is built so the token **cannot** come back: each row is assembled field by field from the three identity keys, never copied wholesale, so an extra key appearing upstream is structurally unable to reach anything this package hands a UI. The guard proves that rather than asserting it — a poisoned row carrying a secret is read, and the secret is searched for across the **entire** serialized result, with the same search proven to find it in the input so a blind search cannot pass; renaming the credential key does not help it through, because the rule is *only these three*, not a blocklist; and the reader's own source is checked to contain no object spread, since one such line would quietly void all of it. The other half is an absence distinction with opposite consequences: a record with **no** parks field at all was written by an older engine and proves nothing about whether approvals are waiting, while an empty list is a positive statement that none are — collapsing those two would let a run whose parked approvals cannot be proven be resumed anyway, so they are kept literally distinguishable, and a payload whose rows are all unreadable answers *unknown* rather than *none*. The four refusal codes for this family are checked code by code against the engine's real bytes, never matched by name prefix, and the older umbrella code they were split out of is asserted to still be **alive** — treating the whole code as retired would make a family of real refusals vanish silently **0.68.3 (core 7.18.0):** two more keys ride the same projection duty as `parks` itself: `originUnconfirmed: true` on a row (never `false`; absence is the confirmed state) and `resumeAdmissionIncomplete: true` on the run (presence means "not a resume base"). Dropping either would turn a refused record back into an admissible one, so the guard pins both, including that neither folds into the other **0.69.1 (CC-12):** both keys now also ride the projected `WorkflowRunState`, so a host that only sees the projection can render them |
360
361
  | `scripts/run-retired-vocabulary-census-test.mjs` | Whether a retirement really happened. When upstream removes a family, a downstream package can cut it out or keep a courteous alias — and the alias is the worse outcome: three clients keep writing branches for something nobody emits, and a status line advertises a state it can never reach. Choosing the clean cut only means something if a guard holds it, since a comment saying *retired* is not an exit code. Each registered entry is held two ways: the name must be gone from **code positions** in this package (comments stripped first, because the explanation is supposed to stay) and off the published surface, and — the half that keeps this from being self-congratulation — it must really be gone **upstream**, since that is the entire reason it was removed here; if it comes back, the disposition deserves reconsideration rather than silence. The scanner proves it can speak by finding a symbol that is genuinely present before any absence is believed, and distinguishes a mention inside a comment from one in a string literal, which is exactly the form being cleared. A closing check runs the other way: the retirement **story** must remain in the comments, including a promise this package made earlier and has now had to withdraw — deleting the history alongside the code is a bad way to satisfy *zero hits*, and leaves the next reader with code that has no reason |
361
362
  | `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
362
363
  | `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
@@ -365,7 +366,7 @@ public-surface guard checks that last one).
365
366
  | `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
366
367
  | `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
367
368
  | `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
368
- | `scripts/run-text-segment-authority-test.mjs` | The **authoritative segment replacement** on `text_end` (L-310, server >=7.75.3). `text_end.content` now goes through the same redactor as `result` and the ledger while `text_delta` stays verbatim, so the two **may differ** — an answer that quoted a credential used to be committed to the local transcript in its unredacted form, because the arm only forwarded the boundary signal. Six timing shapes are pinned, two of which an adversarial review reproduced against the installed engine's real bytes and which the first design got wrong in both directions: a second boundary in the same turn (the per-block case on one provider lane) used to make the first segment's prose vanish, and a boundary that arrives *after* the tool card (the other lane emits it at finalize) used to be read as "this package never handled that segment" and reported nothing at all. Three additive keys, all never-false; the two shapes that look alike are told apart by the second one, because the host's action in them is the opposite. The end-to-end legs drive the real pipeline without hand-inserting a segment commit — doing so is exactly what hid the first defect. A second review round then found two combination timings on top of the first fix — a tool card followed by *more* deltas in the same segment, and a byte count that had been documented as a message count — and both are pinned here too. A third round caught a length that the prose called bytes while the code returned UTF-16 units — harmless in ASCII, and on CJK text enough to leave the credential on screen — plus a backfill ledger that had to be kept in step, so the terminal frame does not re-render the segment a second time — kept in step only where the whole stretch sits in one message, because those ledgers are per-message and a fourth round showed that writing across them charges one message's prose to another. A fifth round settled the whole class into one invariant the guard now checks against the previous release's behaviour: this package only rewrites bytes it is still holding in the current message — once a segment has crossed a package-side boundary it emits the three keys and changes nothing else **0.68.2 (CC-01):** the segment identity is now minted here, not by the host: every committed assistant text row carries a top-level `_sema_segment_id` (stamped once at the `adapt()` exit, so the durable whole-message leg and the streamed-segment leg are covered alike; thinking blocks, tool_use-tailed rows and chrome events are left byte-for-byte), `text_segment_end` carries the same value as `segmentId` before rotating, subagent boundaries never rotate, and a replayed stream yields the same identities. Three mutations (no rotation / no stamping / stamping tool_use rows) each turn the guard red |
369
+ | `scripts/run-text-segment-authority-test.mjs` | The **authoritative segment replacement** on `text_end` (L-310, server >=7.75.3). `text_end.content` now goes through the same redactor as `result` and the ledger while `text_delta` stays verbatim, so the two **may differ** — an answer that quoted a credential used to be committed to the local transcript in its unredacted form, because the arm only forwarded the boundary signal. Six timing shapes are pinned, two of which an adversarial review reproduced against the installed engine's real bytes and which the first design got wrong in both directions: a second boundary in the same turn (the per-block case on one provider lane) used to make the first segment's prose vanish, and a boundary that arrives *after* the tool card (the other lane emits it at finalize) used to be read as "this package never handled that segment" and reported nothing at all. Three additive keys, all never-false; the two shapes that look alike are told apart by the second one, because the host's action in them is the opposite. The end-to-end legs drive the real pipeline without hand-inserting a segment commit — doing so is exactly what hid the first defect. A second review round then found two combination timings on top of the first fix — a tool card followed by *more* deltas in the same segment, and a byte count that had been documented as a message count — and both are pinned here too. A third round caught a length that the prose called bytes while the code returned UTF-16 units — harmless in ASCII, and on CJK text enough to leave the credential on screen — plus a backfill ledger that had to be kept in step, so the terminal frame does not re-render the segment a second time — kept in step only where the whole stretch sits in one message, because those ledgers are per-message and a fourth round showed that writing across them charges one message's prose to another. A fifth round settled the whole class into one invariant the guard now checks against the previous release's behaviour: this package only rewrites bytes it is still holding in the current message — once a segment has crossed a package-side boundary it emits the three keys and changes nothing else **0.68.2 (CC-01):** the segment identity is now minted here, not by the host: every committed assistant text row carries a top-level `_sema_segment_id` (stamped once at the `adapt()` exit, so the durable whole-message leg and the streamed-segment leg are covered alike; thinking blocks, tool_use-tailed rows and chrome events are left byte-for-byte), `text_segment_end` carries the same value as `segmentId` before rotating, subagent boundaries never rotate, and a replayed stream yields the same identities. Three mutations (no rotation / no stamping / stamping tool_use rows) each turn the guard red **0.69.0 (CC-02):** the same authority replacement now covers the reasoning face (`reasoning_end`, server >=7.77.0): a thinking block still buffered is swapped whole and its live tail recomputed; one already committed at a boundary (the usual timing, since the first text delta commits it) is left untouched and the host is told the row to replace by its uuid, never re-emitted. Subagent boundaries are ignored and the text-segment identity does not rotate **0.69.1 (CC-09):** the run-stream replay guard still drops a frame whose event id was already seen, but it now reports the drop through the host's dropped-frame sink as `duplicate_seq` instead of vanishing silently (server 7.77.0 reuses the first reasoning delta's id for `reasoning_end`, so that authoritative segment is lost on the print lane until 7.78.1); the interactive adapter has no such guard and keeps receiving it **0.69.1 (CC-10/CC-11):** subagent segment-end frames are fenced on all three identity keys (a frame carrying only `sourceTaskId` no longer masquerades as the leader's), and a reasoning segment that spans tool cards now hands the host every committed row it covers (`committedUuids`) so nothing unredacted is left behind |
369
370
  | `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the 74 suites above actually turn red when the material they check really breaks — a census had found 16 of them clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. Seven guards of the same shape and 51 behaviour/projection suites are catalogued rather than rehearsed this round — see `docs/GATE-NEGATIVE-CONTROLS.md` for the full table, the reasons, and a one-minute manual replay recipe for each blind one. The suite cross-checks its own case count against that document's row counts in both directions, so a case quietly dropped from the array without the document following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
370
371
 
371
372
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
@@ -446,6 +446,8 @@ const approvalFrameArm = (kind) => function* (m) {
446
446
  * · 第二道 `content` 在场判(投影层已判 malformed/empty):同 `engine_notice` 的 carrier 二道判,
447
447
  * 防的是**非投影口喂进来的帧**(宿主自建管线 / 重放存量转录),不是重复判据。
448
448
  */
449
+ /** CC-10:子流身份三键(sdk `EventIdentity`)任一在场 = 这条段边界不是 leader 的;按「键在不在」判(坏值已在投影口 malformed)。 */
450
+ const isSubFlowSegmentEnd = (m) => m.parentToolCallId !== undefined || m.sourceTaskId !== undefined || m.bgAgentId !== undefined;
449
451
  const textSegmentEndArm = function* (m, { text }) {
450
452
  // 🔴 断闸按「**键在不在**」判,不按「是不是串」判(异源对抗复审第三轮 [medium] 采纳)。
451
453
  // 投影层已对坏 lane 位整帧 fail-closed;这一道是给**非投影口**喂进来的帧(宿主自建管线 /
@@ -453,7 +455,8 @@ const textSegmentEndArm = function* (m, { text }) {
453
455
  // 子代的段边界**擦成 leader 的**上到宿主面。同族臂的 `typeof` 写法在它们那里只影响一行装饰,
454
456
  // 在本臂上决定的是「这条边界算谁的」,方向必须更严。
455
457
  // ⚠️ `null` 也算在场(不给它开口子):wire schema 只允许缺席或 string,`null` 是坏值不是缺席。
456
- if (m.parentToolCallId !== undefined)
458
+ // CC-10(0.69.1):三键任一在场即断(只带 `sourceTaskId` 的子代帧是合法形,修前漏断 = 子代覆盖 leader)。
459
+ if (isSubFlowSegmentEnd(m))
457
460
  return;
458
461
  const content = typeof m.content === 'string' ? m.content : '';
459
462
  if (content.length === 0)
@@ -475,6 +478,32 @@ const textSegmentEndArm = function* (m, { text }) {
475
478
  // CC-01:段界到此为止 —— 事件已交到宿主手上之后才轮换(宿主按帧上的身份认行时,行上盖的还是同一个)。
476
479
  text.rotateSegmentIdentity(m);
477
480
  };
481
+ /**
482
+ * `reasoning_end` 内部臂 → 思考块**权威段替换** + chrome `thinking_segment_end`(CC-02,0.68.3;
483
+ * server ≥7.77.0 S-310,sdk ≥9.4.0 `SegmentEndFields` 与 `text_end` 单源)。
484
+ * 与 {@link textSegmentEndArm} 逐条同律:子流断闸(`parentToolCallId` 键在不在)/ 第二道 `content` 在场判 /
485
+ * 替换在**发事件之前**。差别只在两形(见 `TextStream.replaceThinkingSegment` 头注)与多一把 `committedUuid`。
486
+ * 🔴 **不轮换段身份**:段身份是文本行的(CC-01),思考块的收口与它无关。
487
+ */
488
+ const thinkingSegmentEndArm = function* (m, { text }) {
489
+ if (isSubFlowSegmentEnd(m))
490
+ return;
491
+ const content = typeof m.content === 'string' ? m.content : '';
492
+ if (content.length === 0)
493
+ return;
494
+ const r = text.replaceThinkingSegment(content, m);
495
+ yield chrome({
496
+ kind: 'thinking_segment_end',
497
+ laneProof: MAIN,
498
+ content,
499
+ ...(r.diverged ? { diverged: true } : {}),
500
+ ...(r.committedPrefixDiverged ? { committedPrefixDiverged: true } : {}),
501
+ ...(r.committedPrefixLen > 0 ? { committedPrefixLen: r.committedPrefixLen } : {}),
502
+ ...(r.committedUuid !== undefined ? { committedUuid: r.committedUuid } : {}),
503
+ ...(r.committedUuids !== undefined && r.committedUuids.length > 0 ? { committedUuids: r.committedUuids } : {}),
504
+ ...(typeof m.eventId === 'string' && m.eventId.length > 0 ? { eventId: m.eventId } : {}),
505
+ });
506
+ };
478
507
  const promptSuggestionsArm = function* (m) {
479
508
  // T57 第五处断闸(#47 矩阵 #5 同族):子流(parentToolCallId 标)的建议绝不骑主 composer。
480
509
  if (typeof m.parentToolCallId === 'string')
@@ -1111,6 +1140,7 @@ export const ARMS = new Map([
1111
1140
  ['message_committed', messageCommittedArm],
1112
1141
  ['engine_notice', engineNoticeArm],
1113
1142
  ['text_end', textSegmentEndArm],
1143
+ ['reasoning_end', thinkingSegmentEndArm],
1114
1144
  ['wiring_manifest', wiringManifestArm],
1115
1145
  // B-078 / L-208(0.65.0):design/172 流内审批两帧 —— 修前它们在投影层就 `dropped`,
1116
1146
  // 壳自己另接一份 ⇒ desktop/web 拿不到卡集与撤卡(归层违例)。
@@ -97,6 +97,19 @@ export interface TextSegmentReplacement {
97
97
  */
98
98
  committedPrefixDiverged: boolean;
99
99
  }
100
+ /** {@link TextStream.replaceThinkingSegment} 的回执 —— 与 {@link TextSegmentReplacement} 同名同律,多一把 uuid。 */
101
+ export interface ThinkingSegmentReplacement {
102
+ /** 本包缝出来的那一块(已提交的 ∪ 还押着的)与权威全文不逐字节相等。 */
103
+ diverged: boolean;
104
+ /** 已 committed 思考行的全文长度(整块;0 = 没提交过)。 */
105
+ committedPrefixLen: number;
106
+ /** 已 committed 的那条思考行与权威全文不相等 ⇒ 本包不再交字节,宿主整块换。 */
107
+ committedPrefixDiverged: boolean;
108
+ /** 有已提交行时在场:**末条**行的 `uuid`(0.69.0 起;兼容位)。 */
109
+ committedUuid?: string;
110
+ /** CC-11(0.69.1):有已提交行时在场:段内**全部**已提交思考行的 `uuid`(按提交序;跨工具卡时多条)。 */
111
+ committedUuids?: readonly string[];
112
+ }
100
113
  /** M1 对外的七个动作 + 两个读位(矩阵 §3.2 的 #2/#3 两条跨模块接口就是最后那三件)。 */
101
114
  export interface TextStream {
102
115
  /** A5 `text_delta` 半场:MOD-1 思考→回答边界(思考在场就先 committed 上屏)+ 段锚 + 三缓冲累加。 */
@@ -232,6 +245,18 @@ export interface TextStream {
232
245
  * 🔴 它**不动** `emittedAssistantText`:那一位问的是「这一轮产没产过正文」,与消息边界无关。
233
246
  */
234
247
  beginAssistantMessage(): void;
248
+ /**
249
+ * CC-02(0.68.3;server ≥7.77.0 `reasoning_end`):思考块的**权威段替换** —— `text_end` 那一套
250
+ * ({@link TextStream.replaceAnswerSegment})在推理面的孪生,但**只有两形**:
251
+ * (a) 思考还押在缓冲里(没到任何提交边界)⇒ 缓冲整块换成 `content`,活体尾巴按已泄前缀重算;
252
+ * (b) 思考已在边界**整块**提交(思考→回答边界最常见:`text_delta` 先到、server 的 `reasoning_end`
253
+ * 后到)⇒ 撤不回,一个字节都不再交;立 `committedPrefixDiverged` + 那条行的 `uuid`,由宿主整块换。
254
+ * (f) 本包一个字节都没经手(durable 整块思考臂 / 宿主只喂 end 不喂 delta)⇒ 什么都不做。
255
+ * CC-11(0.69.1)补上 (b1):一段推理**跨工具卡**时已提交的是多条行,前缀 = 全部行按序拼接 —— 对得上就只换还押着的
256
+ * 后段,对不上就整段归宿主换;两形都把全部行的 uuid 交出去(`committedUuids`,`committedUuid` = 末条兼容位)。
257
+ * 无 idle-flush、无封存(思考块只在正常边界整块提交)。
258
+ */
259
+ replaceThinkingSegment(content: string, anchor?: Frame): ThinkingSegmentReplacement;
235
260
  /**
236
261
  * 当前**段身份**(语义与铸法见 {@link SEMA_SEGMENT_ID_KEY} 头注)。同一窗口内多次读**同值**;
237
262
  * 首次读时窗口还没有帧锚(durable 整条消息先于任何增量到达)⇒ 退到 `ctx.uuid()` 并缓存。
@@ -136,6 +136,13 @@ export function createTextStream(ctx, idOf) {
136
136
  /** committed 消息的 id 锚 —— 该段/该思考块的**第一帧**(重放确定性:同流同 id)。 */
137
137
  let segmentAnchor = null;
138
138
  let thinkingAnchor = null;
139
+ /**
140
+ * CC-02:**当前窗口里最近一次整块提交的思考行**(正文 + 本包铸的 uuid)—— `reasoning_end` 迟到时
141
+ * (`text_delta` 先触发了思考→回答边界的提交)唯一能拿来算分歧、指认行的对照物。
142
+ * 写于 `takeThinking` 提交那一拍;清于 `replaceThinkingSegment` 用过之后、以及下一块思考开始
143
+ * (`feedThinking` 首条)—— 上一块的提交不是这一块的对照物。
144
+ */
145
+ let committedThinking = [];
139
146
  /**
140
147
  * 段身份窗口(见 {@link SEMA_SEGMENT_ID_KEY} 头注):窗口的第一帧锚 + 序号 + 派生结果缓存。
141
148
  * 🔴 缓存不是优化:锚缺席那一形派生落到 `ctx.uuid()`,不缓存的话同一窗口两次读会得到两个身份。
@@ -162,6 +169,9 @@ export function createTextStream(ctx, idOf) {
162
169
  session_id: ctx.sessionId,
163
170
  parent_tool_use_id: null,
164
171
  };
172
+ // CC-11(0.69.1):**按序累积**,不覆盖 —— 一段推理可以跨工具卡(openai 车道 finalize 时序),每一块提交都是
173
+ // 这一段的一部分;`reasoning_end` 到达时把它们全部交出去(`committedUuids`),宿主才有整段的归属。
174
+ committedThinking.push({ text: thinking, uuid: msg.uuid });
165
175
  thinking = '';
166
176
  thinkingAnchor = null;
167
177
  thinkingBlockOpen = false;
@@ -275,6 +285,8 @@ export function createTextStream(ctx, idOf) {
275
285
  segmentWindowAnchor = frame;
276
286
  if (thinking.length === 0) {
277
287
  thinkingAnchor = frame;
288
+ // CC-11:新一块开始**不**清对照物(修前这里清了 —— 跨工具卡的推理段只剩末条归属,cli 接入回执 ② 实抓);
289
+ // 对照物只在 `reasoning_end` 用过之后清(段界由引擎说了算),turn 结束整只丢,天然有界。
278
290
  // P2d:elapsed 锚在**首条** leader 推理增量(token 累计由 stream_delta.estimatedTokens
279
291
  // 承载,宿主自加;>30s 无增量的 stall 提示同样由宿主按增量时间戳判)。
280
292
  yield chrome({ kind: 'thinking_activity', laneProof: MAIN, active: true });
@@ -420,6 +432,44 @@ export function createTextStream(ctx, idOf) {
420
432
  if (typeof committed === 'string')
421
433
  committedText += committed;
422
434
  },
435
+ replaceThinkingSegment: (content) => {
436
+ // CC-11:段 = 段内**全部**已提交思考行(按序)+ 还押着的缓冲。分歧按整段拼文比;归属一次交出全部 uuid。
437
+ const rows = committedThinking;
438
+ committedThinking = [];
439
+ const prefix = rows.map((r) => r.text).join('');
440
+ const uuids = rows.map((r) => r.uuid);
441
+ const last = uuids.length > 0 ? uuids[uuids.length - 1] : undefined;
442
+ // (f) 本包没经手过这一段(durable 整块臂铸过 transcript 行 / 宿主只喂 end)。
443
+ if (prefix.length === 0 && thinking.length === 0) {
444
+ return { diverged: false, committedPrefixLen: 0, committedPrefixDiverged: false };
445
+ }
446
+ const live = prefix + thinking;
447
+ const diverged = live !== content;
448
+ if (prefix.length === 0) {
449
+ // (a) 全部还押在缓冲里:整块换,活体尾巴按已泄前缀重算。
450
+ const emitted = thinking.slice(0, thinking.length - thinkingPending.length);
451
+ thinkingPending = content.startsWith(emitted) ? content.slice(emitted.length) : '';
452
+ thinking = content;
453
+ return { diverged, committedPrefixLen: 0, committedPrefixDiverged: false };
454
+ }
455
+ // (b) 有已提交行(一条或跨卡多条):已提交那截撤不回。
456
+ const committedPrefixDiverged = !content.startsWith(prefix);
457
+ if (!committedPrefixDiverged) {
458
+ // (b1) 已提交前缀逐字对得上 ⇒ 只换还押着的尾段(跨卡后段);`thinking` 为空时零字节可换,只报归属。
459
+ if (thinking.length > 0) {
460
+ const tail = content.slice(prefix.length);
461
+ const emitted = thinking.slice(0, thinking.length - thinkingPending.length);
462
+ thinkingPending = tail.startsWith(emitted) ? tail.slice(emitted.length) : '';
463
+ thinking = tail;
464
+ }
465
+ return { diverged, committedPrefixLen: prefix.length, committedPrefixDiverged: false, ...(last !== undefined ? { committedUuid: last } : {}), committedUuids: uuids };
466
+ }
467
+ // (b2) 已提交那截自己也过期 ⇒ 一个字节都不再交(还押着的尾段清掉,别让它落成第二份),整段归宿主换。
468
+ thinking = '';
469
+ thinkingPending = '';
470
+ thinkingAnchor = null;
471
+ return { diverged: true, committedPrefixLen: prefix.length, committedPrefixDiverged: true, ...(last !== undefined ? { committedUuid: last } : {}), committedUuids: uuids };
472
+ },
423
473
  segmentId: segmentIdNow,
424
474
  rotateSegmentIdentity: (anchor) => {
425
475
  segmentWindowSerial += 1;
@@ -49,8 +49,12 @@ export const INTERNAL_SDK_ARM_TYPES = new Set([
49
49
  'turn_usage',
50
50
  // #310 / #318 件①:引擎结构化通告的会话面(raw 预分派铸点,见 eventToSdkMessage 顶部)。
51
51
  'engine_notice',
52
- // #323 / core #447:assistant 流式**散文段边界**(raw 预分派铸点,见 `textEndProjection` 头注)。
52
+ // #323 / core #447:assistant 流式**散文段边界**(sdk 9.4.0 起进 union,`case 'text_end'`;见 `textEndProjection` 头注)。
53
53
  'text_end',
54
+ // CC-02(0.68.3;server ≥7.77.0 S-310 / sdk ≥9.4.0 `SegmentEndFields`):**推理段边界** + 权威全文,
55
+ // `text_end` 的孪生(形逐字同,空段判据归 server)。消费口 = `adapt/arms.ts` 的 `thinkingSegmentEndArm`。
56
+ // ⚠️ 三端:print 车道对内部臂「要么吞、要么转成真 CC 帧」—— 本臂在 CC stdout 帧全集零同位帧,壳按 `text_end` 同律接。
57
+ 'reasoning_end',
54
58
  // L-70 / L-108②③(core #524 + core 147③,server ≥7.58.0):`wiring_manifest` 帧上**两个新段**的
55
59
  // 超集投影(`_sema_modelGate` / `_sema_autoMode`)。整份 manifest 的其余段仍不在本切片里 ——
56
60
  // 射程写在 `case 'wiring_manifest'` 头注,别读成「manifest 接上了」。
@@ -150,10 +154,27 @@ export function eventToSdkMessage(ev, ctx) {
150
154
  // 🔴 **到期复核(自退休,不靠人记)**:预分派用 `(ev as {type?:unknown})` 形读判别键,**不收窄** `ev`
151
155
  // ⇒ 臂一进 union,switch 的 `default` 仍看得见它,B5 穷举断言 `assertNeverArm` **编译期真红**,
152
156
  // 逼下一棒把它搬进 switch。搬进去时行为一字不改(下面的投影函数原样复用)。
153
- if (ev.type === 'text_end') {
154
- return textEndProjection(ev, ctx);
155
- }
157
+ // 🔴 **到期复核已兑现(sdk 9.4.0 提货,0.68.3 CC-02)**:`text_end` 进了 union(与 `reasoning_end` 共用
158
+ // `SegmentEndFields`),按原定计划搬进 switch;`textEndProjection` 原样复用,行为一字不改。
156
159
  switch (ev.type) {
160
+ case 'text_end':
161
+ return textEndProjection(ev, ctx);
162
+ /**
163
+ * CC-02(0.68.3):`reasoning_end` —— `text_end` 的推理面孪生(server ≥7.77.0 自铸,core 不发;sdk 9.4.0
164
+ * 两枚帧一份形 `SegmentEndFields`)。处置逐条同 `text_end`:中性内部臂、绝不铸 transcript 行、绝不
165
+ * `not_in_slice`、`content` 非串/lane 位坏 ⇒ malformed、空串 ⇒ empty_payload、身份两键原样透传。
166
+ * 差别只在消费口:`adapt/arms.ts` 的 `thinkingSegmentEndArm` 做**思考块**的整段替换 + chrome
167
+ * `thinking_segment_end`(两形,见 `TextStream.replaceThinkingSegment`)。
168
+ */
169
+ case 'reasoning_end':
170
+ return segmentEndProjection(ev, ctx, 'reasoning_end');
171
+ /**
172
+ * L-315(候 0.68.4):`tool_roster_delta` —— 名册增量。本包今天不投 `wiring_manifest.tools`(切片只投
173
+ * 三段),名册增量没有可挂的基线 ⇒ **有痕 dropped**,与 0.65.0 前的 `approval_request` 同形(不静默、
174
+ * 不 not_in_slice);接上 `tools` 段的那一批把它一起接。
175
+ */
176
+ case 'tool_roster_delta':
177
+ return dropped('unsupported_arm', 'tool_roster_delta');
157
178
  // ── `human_input`(core 5.14.0 design/171 / server 7.4.0 SSE,[3017]/[3020])────────────
158
179
  // 🔴 **到期复核已兑现(sdk 6.9.0 提货,2026-08-08)**:本臂此前是 switch **之前**的一条 raw
159
180
  // 预分派(理由 = 它还没进 SDK 6.3.0 的 `AgentEvent` union,`case 'human_input'` 在 `ev.type`
@@ -922,9 +943,16 @@ export function eventToSdkMessage(ev, ctx) {
922
943
  * ⚠️ **UNTRUSTED、仅展示**:`content` 是模型输出,契约与 `text_delta` 同 —— 渲染,绝不回喂模型。
923
944
  */
924
945
  function textEndProjection(ev, ctx) {
946
+ return segmentEndProjection(ev, ctx, 'text_end');
947
+ }
948
+ /**
949
+ * 段末帧(`text_end` / `reasoning_end`)的**共用投影**(sdk 9.4.0 `SegmentEndFields` 单源 ⇒ 本侧也只写一份;
950
+ * 判据与留痕见 {@link textEndProjection} 头注,两枚帧逐字同)。`arm` 只决定内部臂名与留痕里的帧名。
951
+ */
952
+ function segmentEndProjection(ev, ctx, arm) {
925
953
  const content = ev.content;
926
954
  if (typeof content !== 'string')
927
- return dropped('malformed', 'text_end');
955
+ return dropped('malformed', arm);
928
956
  // 🔴 **lane 身份位坏了就整帧 fail-closed**(异源对抗复审第三轮 [medium] 采纳;与本文件
929
957
  // `applyBgNotification` 对 `parentTaskId` 的 B3-DIRTY 处置逐字同族):`parentToolCallId` 是
930
958
  // **子流断闸的锚**。键在场却不是串时,若按「不是串就当没有」处理,这条**子代**的段边界会被
@@ -932,19 +960,27 @@ function textEndProjection(ev, ctx) {
932
960
  // 正是本臂头注点名要防的跨 lane 状态破坏,而且是 fail-**open** 方向。
933
961
  // 三态与 B3-DIRTY 同:键缺席/`undefined` ⇒ 本来就是 leader 帧(照旧);键在场却非串(含 `null`,
934
962
  // wire schema 只允许缺席或 string)⇒ `malformed` 丢弃并留痕,坏值不许买路。
935
- if (ev.parentToolCallId !== undefined && typeof ev.parentToolCallId !== 'string') {
936
- return dropped('malformed', 'text_end');
963
+ // CC-10(0.69.1;cli 接入回执 ① 实抓的 fail-open):子流身份是**三键**(sdk `EventIdentity`:`parentToolCallId` /
964
+ // `sourceTaskId` / `bgAgentId`),core 顶注逐字「stamped ONLY on the content events of a task running AS A SUB-AGENT」——
965
+ // 只带 `sourceTaskId`(无 `parentToolCallId`)是合法的子代帧形。修前只透 `parentToolCallId`,臂只按它断闸 ⇒ 那一形的
966
+ // 子代段末帧被当 leader 投上来,端据它换 leader 的思考/文本行 = 子代推理覆盖 leader 推理面。⇒ 三键**原样穿上内部臂**,
967
+ // 任一键在场却非串 ⇒ 整帧 malformed(与 `parentToolCallId` 同一条 B3-DIRTY 律:坏值不许买路)。
968
+ for (const k of ['parentToolCallId', 'sourceTaskId', 'bgAgentId']) {
969
+ if (ev[k] !== undefined && typeof ev[k] !== 'string')
970
+ return dropped('malformed', arm);
937
971
  }
938
972
  if (content.length === 0)
939
973
  return nothing('empty_payload');
940
- return projected(stamp(ctx, armBody({
941
- type: 'text_end',
942
- content,
974
+ const identity = {
943
975
  ...(typeof ev.eventId === 'string' ? { eventId: ev.eventId } : {}),
944
- ...(typeof ev.parentToolCallId === 'string'
945
- ? { parentToolCallId: ev.parentToolCallId }
946
- : {}),
947
- })));
976
+ ...(typeof ev.parentToolCallId === 'string' ? { parentToolCallId: ev.parentToolCallId } : {}),
977
+ ...(typeof ev.sourceTaskId === 'string' ? { sourceTaskId: ev.sourceTaskId } : {}),
978
+ ...(typeof ev.bgAgentId === 'string' ? { bgAgentId: ev.bgAgentId } : {}),
979
+ };
980
+ // 两处字面量铸点(不是一处参数化):内部臂词表门按 `armBody({ type: '<字面量>'` 抽铸点,参数化会让两臂在账上「多登记」。
981
+ return arm === 'text_end'
982
+ ? projected(stamp(ctx, armBody({ type: 'text_end', content, ...identity })))
983
+ : projected(stamp(ctx, armBody({ type: 'reasoning_end', content, ...identity })));
948
984
  }
949
985
  /**
950
986
  * `wiring_manifest.mcp` 的形校验(S-124 / core 7.5.0,server ≥7.60.0)。
@@ -255,6 +255,8 @@ function sanitizeFrameType(type) {
255
255
  * 去看引擎那一侧。
256
256
  */
257
257
  const DROPPED_WHY_SENTENCE = Object.freeze({
258
+ duplicate_seq: 'a frame with this event id was already consumed on this run stream, so it was dropped as a replay (durable idempotency, contract 02 §1.1). ' +
259
+ 'If the engine reuses an id across two DIFFERENT frames (server 7.77.0 does this for `reasoning_end`, fixed upstream in 7.78.1), the second one is lost here — this line is the only trace.',
258
260
  turn_end_usage_absent: "the engine declares `turn_end.usage` as always present (core >= 7.17.0), and this frame has none. " +
259
261
  'The turn is counted as UNKNOWN spend (the run total is reported as a lower bound), and the frame itself renders NOWHERE.',
260
262
  });
@@ -354,8 +356,14 @@ async function* runStreamInner(events, ctx, handle = {}) {
354
356
  // event-id idempotency — drop a re-seen durable seq (contract 02 §1.1).
355
357
  const seq = eventSeq(ev);
356
358
  if (seq !== undefined) {
357
- if (seen.has(seq))
359
+ if (seen.has(seq)) {
360
+ // CC-09(0.69.1;cli L-319 `-p` 真根因的包侧半场):**丢也要留痕**。server 7.77.0 给 `reasoning_end` 的
361
+ // SSE `id:` 复用了该段首枚 `reasoning_delta` 的 id ⇒ 这条去重把**权威段**当重放帧丢了,而修前这里是一条裸
362
+ // `continue` —— 比 0.68.2 的 `unknown_arm`(至少有痕)更静默。归口 = server 改铸唯一 id(7.78.1);本层
363
+ // 不猜「这是真重放还是 id 复用」(去重律不变,契约 02 §1.1),只把这次丢弃交给宿主的丢帧留痕口。
364
+ reportDroppedFrame('duplicate_seq', String(ev.type ?? 'unknown'), ctx);
358
365
  continue;
366
+ }
359
367
  seen.add(seq);
360
368
  }
361
369
  // C1 — SUBAGENT content divert (service 1.89 forwardSubagentEvents): a CONTENT event stamped with