@sema-agent/client-core 0.65.1 → 0.67.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +139 -0
- package/README.md +5 -2
- package/dist/adapter/activeRunSelfHeal.js +3 -2
- package/dist/adapter/downstream/eventToSdkMessage.d.ts +14 -0
- package/dist/adapter/downstream/eventToSdkMessage.js +13 -1
- package/dist/adapter/downstream/terminalToSdkResult.d.ts +219 -1
- package/dist/adapter/downstream/terminalToSdkResult.js +317 -7
- package/dist/adapter/downstream/turnUsageToModelUsage.d.ts +20 -1
- package/dist/adapter/downstream/turnUsageToModelUsage.js +7 -2
- package/dist/adapter/runStream.js +118 -4
- package/dist/adapter/types.d.ts +27 -0
- package/dist/autoModeUnavailable.d.ts +95 -70
- package/dist/autoModeUnavailable.js +130 -97
- package/dist/classifierStatus.d.ts +7 -5
- package/dist/classifierStatus.js +62 -20
- package/dist/engineNoticeCodes.d.ts +19 -2
- package/dist/engineNoticeCodes.js +25 -3
- package/dist/gateOutcome.d.ts +14 -0
- package/dist/gateOutcome.js +3 -1
- package/dist/gateVocabulary.d.ts +37 -0
- package/dist/gateVocabulary.js +69 -8
- package/dist/hitl/toolApprovalWire.d.ts +44 -43
- package/dist/hitl/toolApprovalWire.js +47 -59
- package/dist/seam.d.ts +89 -0
- package/dist/seam.js +22 -0
- package/docs/INTEGRATION-CLIENTS.md +776 -11
- package/package.json +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,145 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.67.0(2026-09-12)
|
|
53
|
+
|
|
54
|
+
**core 7.14.0 → 7.16.0 提货批**(逐键处置表见 `docs/INTEGRATION-CLIENTS.md` §32z;每件的键名 / 形 /
|
|
55
|
+
缺席语义 / 三端待办 / 黑盒判据 G32-01 ~ G32-27 见 §32b–§32h)。
|
|
56
|
+
|
|
57
|
+
### BREAKING
|
|
58
|
+
|
|
59
|
+
- **删公面导出 `classifierUnavailableOf`** —— core 7.14.0 的 Retired keys 表把 `classifierUnavailable`
|
|
60
|
+
从六个载体上整族删掉(ask 那几处自 #661 起引擎就没写过);继任者是 `classifierDenyCauseOf(门记录)`。
|
|
61
|
+
clean-cut,**不留别名**。
|
|
62
|
+
- **删公面导出 `classifierUnavailableDetail`** —— 同上;继任者是 `classifierDenyCauseDetail(cause)`。
|
|
63
|
+
- **删公面类型 `ClassifierUnavailableView`** —— 同上(它是那只读器的读数形,也是卡位与帧位的型)。
|
|
64
|
+
- **删卡面键 `ApprovalCardRequest.classifierUnavailable` 与帧面键
|
|
65
|
+
`ToolApprovalFrame.classifierUnavailable`** —— 同一个装配位换成 `ruleStoreUnreadable`;旧耐久行带旧键时
|
|
66
|
+
投影**忽略**(不读、不渲、不崩)。
|
|
67
|
+
- **`ASK_ORIGIN_WORDS` 换词:`unresolvable` → `ancestor_marked`(无 alias)** —— core 7.14.0 的
|
|
68
|
+
`ASK_ORIGINS` BREAKING 换词。旧值在每一道成员筛上都是**非成员** ⇒ 投影省略该 origin、措辞走开集兜底句。
|
|
69
|
+
- **`ENGINE_NOTICE_AUDIENCE` 值域变更:`memory.consolidation_withheld` `user` → `operator`** ——
|
|
70
|
+
core 7.14.0 value-domain BREAKING(这条通告没有 `sessionId`,一次固化跑在任何会话之外)。
|
|
71
|
+
- **`classifierStatusOf` 的第二参语义变更**(签名与三态词表**未改**,不是编译期 BREAKING):
|
|
72
|
+
从「本轮那只 ask」变成「本轮那一次**观测**」(ask / 耐久行 / 带 `gate` 的 `tool_end` 帧 / 门记录本体);
|
|
73
|
+
喂旧形的端那一态会恒不出现。
|
|
74
|
+
|
|
75
|
+
### Added
|
|
76
|
+
|
|
77
|
+
- `CLASSIFIER_DENY_CAUSES` / `isClassifierDenyCause` / `classifierDenyCauseOf` /
|
|
78
|
+
`classifierDenyCauseDetail` —— core 7.14.0 `GateDisposition.denied.cause`(闭二词 `unavailable` /
|
|
79
|
+
`parse_error`)的**拒绝面**读器与唯一措辞铸点;`GateOutcomeView.disposition` 的 denied 臂同批多一格
|
|
80
|
+
`cause?`(**开集透传**,判成员留给公面读器,两层分工与退役前逐字同形)。
|
|
81
|
+
- `RULE_STORE_UNREADABLE_KINDS` / `isRuleStoreUnreadableKind` / `ruleStoreUnreadableDetail` ——
|
|
82
|
+
core 7.14.0 #688 C3:`origin: "rule_store_unavailable"` 底下的**机制位**(`store` = 规则店整体读不出来 /
|
|
83
|
+
`call` = 这条调用对不上规则行),两句人话的下一步**相反**;卡面两腿(活卡帧 + durable park 行)
|
|
84
|
+
经**同一把**窄读器投到 `ApprovalCardRequest.ruleStoreUnreadable`。
|
|
85
|
+
- `WiringManifestAutoMode.deniedSource?: string`(core 7.14.0 `AutoModeArmFact`)—— 哪一层设置面的棘轮
|
|
86
|
+
关掉了 auto(`org` / `local` / `settings`,**开集读**);**只在 `reason === "denied"` 旁收**(段内自洽,
|
|
87
|
+
与本段既有的 `armed === (reason === 'armed')` 互证同族),缺席不铸。
|
|
88
|
+
- 终帧 `_sema_usage_lower_bound: true` 与 `RunCostReconcile.usageLowerBound`(core 7.14.0
|
|
89
|
+
`TaskResult.stats.usageMissing`)—— 「这一面的数字是**下界**不是一次测量」的判别位;
|
|
90
|
+
🔴 CC 同名键 `usage` / `modelUsage` / `total_cost_usd` **语义零改**,never false。
|
|
91
|
+
- **per-subagent usage 分表**(L-228):新 chrome 臂 `subagent_turn_usage`(`required: false`,`laneProof`
|
|
92
|
+
恒是子流那条)+ 终帧 `_sema_nested_usage_by_task` 与诚实位 `_sema_nested_usage_by_task_partial`。
|
|
93
|
+
🔴 §E2 的两处主臂断闸**一个字节不动**;`partial` 的判据锚在引擎自己的权威合计 `stats.nested` 上
|
|
94
|
+
(turns 之和 + 行数都对得上才算可证完整),失效方向只会**多**铸 partial。
|
|
95
|
+
- 码册加员 `task.interrupt_unconsumed`(audience `user`)—— **合同外加员**,由
|
|
96
|
+
`run-engine-notice-catalog-test.mjs` 的码数双向对账当天红抓出,如实登记在 §32z ⑤。
|
|
97
|
+
|
|
98
|
+
### Changed
|
|
99
|
+
|
|
100
|
+
- devDep `@sema-agent/core` `~7.12.0` → `~7.16.0`。
|
|
101
|
+
- `run-gate-vocabulary-test.mjs` 的 `AskOrigin` **对账基准从 sdk 换成 core**(词表属主是 core,
|
|
102
|
+
而 sdk 8.8.0 的联合滞后一代);sdk 的滞后改成一条**带退出条件的登记**(追平那天自红逼删)。
|
|
103
|
+
- 同批**退役**一条恒真的型面钉 `_askOriginWordsPin`(sdk 的 `AskOrigin` 带 `(string & {})` 逃生口
|
|
104
|
+
⇒ 任何字符串都满足它,换词当天一声不响);判据整只移交那道门的 B 段。
|
|
105
|
+
- `run-retired-vocabulary-census-test.mjs` 的剥注释器修一处**词法跑偏**:单/双引号串遇裸换行即收口
|
|
106
|
+
(JS 词法本来如此)—— 修前一次误判会让其后的每一段注释都被当成串体保留,而本门的全部意义正是
|
|
107
|
+
「注释里写了不算数,码才算数」。
|
|
108
|
+
|
|
109
|
+
### 已知局限(本版新增)
|
|
110
|
+
|
|
111
|
+
- 🔴 **F-6 的供给面未端到端实证**(异源对抗复审 [high] 提出,本批如实登记而不是销格):core 里子代
|
|
112
|
+
事件有两条外送腿 —— 工具 ctx 的 `forwardEvent` **白名单**(`prepare-run-refs.js`:无条件过
|
|
113
|
+
`task_progress`,开了 `forwardSubagentEvents` 再加五类转录事件,**`turn_end` 不在其中**;`tool-spec.d.ts`
|
|
114
|
+
顶注逐字「other event types never cross it」)与**委派车道自己的 tap**(同一段顶注逐字「forwards the
|
|
115
|
+
child's FULL event stream」)。本包消费的是后者,而本批**没有**跑通 core→server→父流的真机供给实证
|
|
116
|
+
(本树无 server fixture)。⇒ 只走白名单腿的部署上分表**恒零行**,两个终帧键按设计**整键不铸**
|
|
117
|
+
(端读到的是「说不出来」,不是一个假的零)。本件因此按「包边界承诺」交付:子流 `turn_end` 到得了这条流
|
|
118
|
+
就有行,到不了就两个键都不在场;端**不许**把键缺席渲成「这条 run 没委派子代」。真机供给与集成门另立。
|
|
119
|
+
|
|
120
|
+
- `WorkflowRun.errorCode`(core 7.14.0)**本包今天到不了它**:本包投影的两条腿(core `formatWorkflowRun`
|
|
121
|
+
的 TaskOutput JSON / `summarizeWorkflowRun` 的列表行)都不发 run 级 `errorCode`,第三条腿
|
|
122
|
+
`GET /v1/workflows/:id` 的 sdk 型面无此声明且本树未装 server fixture ⇒ 证不出 service 有没有转投。
|
|
123
|
+
按「core 定型 ≠ 壳可消费」办:**不猜载体**,`pending`(§32z ①-4)。
|
|
124
|
+
- 首请求 `tools[]` 形(core 7.15.0)对自报面有一处影响未修:`ENGINE_RUNNER_FACE` 把 `ToolSearch` 写成
|
|
125
|
+
**无条件**成员,而 7.15.0 之后它**有条件**了(延迟集为空 ⇒ 无 `ToolSearch`)。阈值只有引擎算得出,
|
|
126
|
+
正解是让自报面改读引擎的 roster 面 ⇒ `pending`(§32z ③-2)。
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## 0.66.0(2026-09-11)
|
|
131
|
+
|
|
132
|
+
> **终局真值三件**(cli 台账 L-192①② / B-068 · DEBTS L-198):三件都是同一条病形 ——
|
|
133
|
+
> 上游把事实摆在 wire 上,而包边界用一个**字面量**把它答成了常数。处置也只有一条:CC 形上必填的位
|
|
134
|
+
> **值不动**(旧消费者逐字不变),「不知道」交给同行的 `_sema_` 判别位;CC 形上没有的事实另开超集座位,
|
|
135
|
+
> **绝不改 CC 同名键的语义**。接入文档 §31(速览 / 逐件键表与缺席语义 / 黑盒判据 G31-1..G31-16 /
|
|
136
|
+
> 三端待办 / 本批的门)。判据编号 G31-1..G31-18。
|
|
137
|
+
|
|
138
|
+
- **① L-192① 终帧拒绝清单不再硬编 `[]`**(`adapter/downstream/terminalToSdkResult.ts`,成功臂与错误
|
|
139
|
+
信封**两处**):按 `TaskStats.humanReview.gates[]` 里 `decision === "deny"` 的行**逐条**铸记录,
|
|
140
|
+
顺序 = wire 顺序、条数 = 被拒次数,键按 wire 能兑现的铸(`tool_name` ⇐ `toolName`;入参摘要
|
|
141
|
+
`_sema_tool_arg` ⇐ `toolArg`,引擎侧已脱敏截短)。🔴 **两条清单**:CC 的 `SDKPermissionDenial`
|
|
142
|
+
三键**全是必填**,而 core 的 gate 账本给不出 `tool_use_id` / `tool_input`(逐字:耐久 resume 的 gate
|
|
143
|
+
只带 `toolName`)—— 补零补空是**编造**,把半条记录塞进 CC 数组则**破坏元素契约**(严格消费方
|
|
144
|
+
`safeParse` 会把**整条 result 帧**判非法)⇒ CC 数组只收三键齐全的记录(今天恒空,上游补齐后
|
|
145
|
+
**自动**开始填,包侧一行不用改),wire 上真有的每一条走超集载体 `_sema_permission_denials`。
|
|
146
|
+
顶层判别位 `_sema_permission_denials_absent: true` = 「CC 那条清单不可声称完整」,四条路径:
|
|
147
|
+
账本读不出(老引擎 / 409 拒绝信封 / `failed` 事件帧)/ 有读不出的行 / 有**认不出的判词**
|
|
148
|
+
(两张判词表之外的词、以及缺判词的耐久 wake 行)/ 有记录只在超集载体上。⇒「零拒绝」与
|
|
149
|
+
「清单不完整」从此可分。门 `run-permission-denial-projection-test.mjs`(新,68 checks,含**缺席证据**
|
|
150
|
+
与**自动升级腿**的正控:上游哪天把两格串上账本当场红/当场填)。
|
|
151
|
+
- **② B-068 / L-198 终局成本对账**(同文件 + `adapter/runStream.ts` + `seam.ts`):`TaskStats.costBreakdown`
|
|
152
|
+
与 `nested` 两段投上终帧超集键 `_sema_cost_breakdown` / `_sema_nested_usage`,**micro-USD 原值不折 USD**,
|
|
153
|
+
逐键窄读、坏键剥掉不连坐、整段读不出不铸空对象、上游新类目原样过境;nested 的 `costMicroUsd` 缺席 =
|
|
154
|
+
委派花费**没定价** ⇒ 键缺席,绝不铸 0。新增 chrome 臂 **`run_cost_reconciled`**(`required: false`):
|
|
155
|
+
终局那一拍交出 own / nested / compaction 三段 + `reconciledMicroUsd = own + nested`,任一段不知道
|
|
156
|
+
(或和本身非有限)则只铸 `costAbsent` / `nestedCostAbsent` 判别位而**不铸**对账值;`stats` 读不出的
|
|
157
|
+
终帧(409 / park 体 / `failed` 事件帧 / 畸形载体)一条都不发,非成功终局照发。🔴 「委派过」的判据是
|
|
158
|
+
`stats.nested` **这个载体在不在**,不是它里面有没有读得出的数 —— 否则 `nested: {}` 会被当成「没委派」
|
|
159
|
+
而把 own 铸成一个**确定的总额**。🔴 **CC 同名键 `total_cost_usd` 语义一字未改**(仍是 own 花费,
|
|
160
|
+
nested 不折进去);对账值由消费方按超集键自己加。🔴 子流 `turn_end` 仍不发 `last_turn_usage`(既有断闸
|
|
161
|
+
一字未动)—— 子代花费只经终局 `nested` 到账,端**不要**把流中增量与终局总账相加。
|
|
162
|
+
新增公面导出 **`readRunCostFacts`**(终帧两个超集键与 chrome 臂**共用**它 ⇒ 两面不会各算各的)。
|
|
163
|
+
门 `run-cost-reconcile-projection-test.mjs`(新,76 checks)。
|
|
164
|
+
- **③ L-192② 终帧扁平 `usage` 三格**(同文件 + `adapter/downstream/turnUsageToModelUsage.ts`):
|
|
165
|
+
`webSearchRequests` 从**字面量 0** 改按 CC 的名字开集宽读,读不出时值仍 0(CC 形必填 number)+ 同行
|
|
166
|
+
判别位 `_sema_web_search_requests_absent` —— **两个 mint 点**(终帧扁平 usage 与 CC `ModelUsage` 镜像)
|
|
167
|
+
同扫,同形不留第二处。新增 cache-INCLUSIVE 总量座位 `usage._sema_total_input_tokens`
|
|
168
|
+
(⇐ `TaskStats.totalInputTokens`);终帧此前**没有任何载体**能说出「这条 run 摆了多少上下文」。
|
|
169
|
+
同批按**同一条理由**给**终局 per-model 行**补三座位(`_sema_total_input_tokens` /
|
|
170
|
+
`_sema_usage_basis` 口径标记 / `_sema_cache_write_tokens_long`)—— 一个全局总量答不了「哪个模型摆了
|
|
171
|
+
多少」,而那张分表同样没有逐字通道;加挂只发生在两条**终局**腿上,per-turn 与终局共用的那个 mint 点
|
|
172
|
+
一个字不动(门里有反向钉)。扁平 usage 同批补 `_sema_cache_write_tokens_long`:🔴 **只另给不相加** ——
|
|
173
|
+
core 对 `cacheWriteTokens` 与 `cacheWriteTokensLong` 的说法互相矛盾(前者自述是协议侧那个**含 1h 子项**
|
|
174
|
+
的量,而总量式又把两格并列相加),相加与不加各有一种错法,**证据不足不猜**,CC 那一格逐字不动。
|
|
175
|
+
🔴 **`inputTokens` 逐字不动**,仍是 cache-**MISS** 分量:台账原句要求改读 `totalInputTokens`,核合同后
|
|
176
|
+
**不采** —— core RB-457-a 把 `promptTokens` 翻成 MISS 分量而 CC/Anthropic 的 `input_tokens` 本义正是 MISS,
|
|
177
|
+
灌总量会与同帧两个 cache 格**双算**(core 逐字:98% 命中率被渲成 49.5%),那是改 CC 同名键语义不是超集。
|
|
178
|
+
理由与「本批刻意不投的位」(`contextWindow`/`maxOutputTokens`、`cacheHitRate`/`toolCalls`/`mechanisms` 等)
|
|
179
|
+
逐条登记在 §31d。门 `run-cost-absence-projection-test.mjs` 扩 E/F/G/H 四段(144 checks)。
|
|
180
|
+
- **④ 异源对抗复审(查漏轨)采纳四件**:①**[high]** 终局 chrome 回调的**异步**拒绝此前会外溢成
|
|
181
|
+
`unhandledRejection`(端口契约是 `void | Promise<void>`,而 `try/catch` 只接得住同步抛)——
|
|
182
|
+
把未处理拒绝当致命的宿主会整只退出;修 = 新增共用发射口 `emitChromeFireAndForget`(同步抛在
|
|
183
|
+
`try` 里吞、异步拒绝挂 `.catch`),**同形族扫**把既有的 `last_turn_usage` 那条腿一并收编,语义
|
|
184
|
+
逐字不变(不 await、不改时序);②nested 存在性口径(见上);③CC 元素契约(见 ①);
|
|
185
|
+
④两处同形漏键(长 TTL 分量 / 终局 per-model 三座位)。
|
|
186
|
+
- **消费方影响**:只读既有键的端**零改**(CC 必填位的值在同样输入下一字不变);要新事实的端按 §31e
|
|
187
|
+
三端待办接。超集键 10 行登记进 `docs/type-superset.json`(常驻门双向对账);`unknown` 出境棘轮
|
|
188
|
+
320→321(逐条记账在 `run-client-core-typeshape-test.mjs`:`SemaPermissionDenial.tool_input` 是
|
|
189
|
+
CC 契约**本身**的形 `Record<string, unknown>`)。
|
|
190
|
+
|
|
52
191
|
## 0.65.1(2026-09-11)
|
|
53
192
|
|
|
54
193
|
test [6961] 对 0.65.0 十件五路交叉验证全 PASS,对抗复审轨另抓两处带坐标的实现缺口(cli 亲核源码成立,两处都是 patch):
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.
|
|
38
|
+
**Version:** 0.67.0
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -230,7 +230,7 @@ public-surface guard checks that last one).
|
|
|
230
230
|
| `scripts/run-client-core-portability-test.mjs` | Kernel / A-layer / index import closures, the runtime-dependency equality gate, barrel reachability, and a real esbuild `--platform=browser` bundle |
|
|
231
231
|
| `scripts/run-client-core-diff-test.mjs` | Differential equivalence against the CLI reference bridge + replay-id invariant + ledger round-trip |
|
|
232
232
|
| `scripts/run-seat-contract-keys-test.mjs` | The seat IPC contract: verb list ↔ SPEC ↔ types, element-wise |
|
|
233
|
-
| `scripts/run-approval-frame-keys-test.mjs` | The tool-approval frame key mirror, element-wise against the SDK's runtime anchor (one carve-out: AHEAD_OF_ANCHOR entries — keys the server already emits but the SDK anchor has not caught up to — may lead by one generation; the gate turns red the day the SDK catches up, forcing the entry's removal — the register is occupied again by the
|
|
233
|
+
| `scripts/run-approval-frame-keys-test.mjs` | The tool-approval frame key mirror, element-wise against the SDK's runtime anchor (one carve-out: AHEAD_OF_ANCHOR entries — keys the server already emits but the SDK anchor has not caught up to — may lead by one generation; the gate turns red the day the SDK catches up, forcing the entry's removal — the register is occupied again — this time by the rule-store-unreadable key the engine now defines, carrying both the release that minted it and the byte coordinates that prove it, so the lead is a dated record rather than an exemption; its predecessor left the register the other way, by being retired upstream rather than by the anchor catching up) |
|
|
234
234
|
| `scripts/run-print-bash-iserror-test.mjs` | The print lane's Bash `is_error` authority (structured over regex) |
|
|
235
235
|
| `scripts/run-bash-benign-exit-interpretation-test.mjs` | Benign non-zero Bash exits (`returnCodeInterpretation`) stay non-errors across all three derivation arms, and the annotation transits to the card |
|
|
236
236
|
| `scripts/run-sdk-floor-test.mjs` | The SDK version floor — and, more to the point, that the *installed* type declarations still carry the keys this package reads |
|
|
@@ -250,6 +250,7 @@ public-surface guard checks that last one).
|
|
|
250
250
|
| `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header |
|
|
251
251
|
| `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters |
|
|
252
252
|
| `scripts/run-subagent-content-budget-test.mjs` | The **byte** budget on the sub-agent transcript ledger. It used to be bounded only by *counts* — so many entries per child, so many children — and a count is not a budget when a single entry has no ceiling of its own: one tool result carrying an inlined attachment, or one long model answer, and a single slot sits on tens of megabytes. The guard anchors on how many bytes are **still held** after over-filling, not on whether truncation fired, because an implementation that flags the overflow without actually dropping anything satisfies the second and not the first. Dropping is required to leave a record — how much went and where the retained content now starts — and that record has to reach the render plan, because content that vanishes with no marker gives the reader a transcript shorter than what happened with nothing to say so; the record is one per child, updated in place, pinned to the front, and excluded from the budget it describes. Order matters and is checked: oldest entries go first and the live tail is trimmed only as a last resort, since taking the text the user is watching stream while older history survives is the wrong end. The total budget evicts a whole least-recently-used child rather than shaving every child, and the configuration surface is fail-loud on zero, negatives, non-finite and non-integer values — a silently ignored budget is the exact failure this exists to remove — with the rejection proven atomic so a bad second field cannot leave half a configuration behind. The defaults are checked to be a magnitude that can really be reached, since a number too large to hit is a field rather than a budget |
|
|
253
|
+
| `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other |
|
|
253
254
|
| `scripts/run-result-text-backfill-test.mjs` | What happens when the terminal frame's answer text and the text already on screen do not match. A turn's answer normally streams in and the terminal frame carries the same words again, so the two agree — but when the connection drops mid-answer and the reconnect brings the finished version, "this turn already produced assistant text" is true, the terminal fallback is skipped entirely, and the screen stays permanently short of whatever arrived while the stream was down, with nothing to say so. Four cases are pinned. Nothing on screen yet: render the terminal text whole, byte for byte the previous behaviour. On-screen text is a **prefix** of the terminal text: emit only the missing tail, and the guard measures the deciding quantity — the total bytes that reached the screen must equal the terminal text, which fails both for a missing tail and for a re-render that would print the first half twice; when the two are already equal, nothing is emitted at all. Terminal text is a prefix of what is on screen (an engine-side trim): touch nothing, since there is nothing missing and overwriting with the shorter version would erase what the reader already saw. Neither is a prefix of the other: emit **nothing** and raise a fact instead — which version counts is the engine's to say, and appending the terminal version after the streamed one composes a passage nobody ever wrote. That fact carries lengths rather than text, so a renderer is not handed a third version to choose from, and its declared duty is to *reword* the transcript line, never to render more. A cross-segment case proves the comparison reads the whole committed answer rather than the last segment, and the whole thing is driven through the real two-stage path rather than hand-built messages |
|
|
254
255
|
| `scripts/run-engine-vocab-floor-test.mjs` | Engine-mirrored vocabularies (structured card whitelist, self-reported tool face, control verbs, recogniser sets) against the *installed* `@sema-agent/core` |
|
|
255
256
|
| `scripts/run-limits-env-failloud-test.mjs` | `SEMA_HEADLESS_*` env-lane limits reject invalid values as loudly as the flag lane (no silent "no budget" runs) |
|
|
@@ -295,6 +296,8 @@ public-surface guard checks that last one).
|
|
|
295
296
|
| `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
|
|
296
297
|
| `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
|
|
297
298
|
| `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
|
|
299
|
+
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift |
|
|
300
|
+
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end |
|
|
298
301
|
| `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
|
|
299
302
|
| `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
|
|
300
303
|
|
|
@@ -723,8 +723,9 @@ async function staleParkArm(taskId, busy, runs, deps) {
|
|
|
723
723
|
// —— 「不知道」和「行又回来了」都不构成销毁一条 run 的授权。
|
|
724
724
|
// 🔴 **在册边界(P-44,异源对抗复审 R2 [high] 如实登记)**:这是「查了再做」,**不是原子条件取消**。
|
|
725
725
|
// 复证与那一枪之间仍有毫秒级窗口,行恰在此间恢复的话那一枪照样落下去。客户端关不死它 ——
|
|
726
|
-
//
|
|
727
|
-
//
|
|
726
|
+
// 真正的关法是 **server 读写面**的条件取消(core [7006] 定谳:core 无席——`TaskSpec.signal` 无条件中止、
|
|
727
|
+
// `claimTerminal` 是 ask 行 CAS 非 run lease;server [7007] 认领:cancel 已是两臂状态 CAS,7.73.0 契约明写 + 可选
|
|
728
|
+
// 版本前置,S-122 车),那是 wire 能力不是壳能自造的语义。这里能做的是把窗口从「人看卡的任意长时间」压到最小,
|
|
728
729
|
// 并把剩余风险登记在册(docs/INTEGRATION-CLIENTS.md §7b P-44),不假装它不存在。
|
|
729
730
|
const stillGone = await readOwnedPendingCount(recheckOwnedPending, deps);
|
|
730
731
|
if (stillGone !== 0)
|
|
@@ -169,6 +169,20 @@ export interface WiringManifestAutoMode {
|
|
|
169
169
|
armed: boolean;
|
|
170
170
|
/** core **六词逐字透传**。🔴 不映射 `/v1/capabilities.permissionModeAuto.reason` —— 见投影函数头注。 */
|
|
171
171
|
reason: string;
|
|
172
|
+
/**
|
|
173
|
+
* 0.67.0(core 7.14.0 `AutoModeArmFact.deniedSource`,`wiring-manifest.d.ts:87`)——
|
|
174
|
+
* **哪一层设置面**的棘轮把 auto 模式关掉了(core `AUTO_MODE_DENY_SOURCES`:`org` / `local` /
|
|
175
|
+
* `settings`)。解析器**自己的词,原样带过来,从不推断**(合同顶注 `:79-82` 逐字)。
|
|
176
|
+
*
|
|
177
|
+
* 🔴 **只在 `reason === "denied"` 旁在场**;`denied` 却没有它 = 那个解析器**不记来源**
|
|
178
|
+
* (一句正面事实,不是「不知道」的同义词);`resolver_fault` / `armed` / `no_intent` /
|
|
179
|
+
* `no_face` 上**恒不在场**(屏掉一个它拒绝了的来源是 core 自己做的)。
|
|
180
|
+
* 🔴 **开集读**(与同段 `reason` 逐字同规):词表属主在 core,包在边界抄一份闭集只会在 core
|
|
181
|
+
* 加词那天把一个合法值判没。端的 `switch` 必须带 `default`。
|
|
182
|
+
* 🔴 **缺席不铸**:绝不折成空串,更不折成 `"settings"` 这种看起来最像的默认值 —— 那是替引擎
|
|
183
|
+
* 指认一个它没点名的设置面,而用户会照着去改错的那一层。
|
|
184
|
+
*/
|
|
185
|
+
deniedSource?: string;
|
|
172
186
|
}
|
|
173
187
|
/**
|
|
174
188
|
* {@link wiringManifestSupersetBody} 的 `mcp[]` 一条目(S-124 / core 7.5.0,server ≥7.60.0 的形)。
|
|
@@ -1036,7 +1036,19 @@ function projectAutoModeSection(raw) {
|
|
|
1036
1036
|
return undefined;
|
|
1037
1037
|
if (armed !== (reason === 'armed'))
|
|
1038
1038
|
return undefined;
|
|
1039
|
-
|
|
1039
|
+
// ── 0.67.0(core 7.14.0):`deniedSource` 补位 ───────────────────────────────────────────────
|
|
1040
|
+
// 见 {@link WiringManifestAutoMode.deniedSource}。本层**只在 `reason === "denied"` 上收** ——
|
|
1041
|
+
// core 的段内规矩是它只站在 `denied` 旁边,而这一条与上面那条 `armed === (reason === 'armed')`
|
|
1042
|
+
// 互证判据**同族**:非投影口(宿主自建管线 / 重放存量转录)喂进来的帧不过 server,一个
|
|
1043
|
+
// `{reason:'armed', deniedSource:'org'}` 会让消费端同时读到「武装了」和「被 org 关掉了」。
|
|
1044
|
+
// ⚠️ 与 `errorCode`/`origin` 的「不校配对」不同裁,理由也是上游自己给的:那些是**跨系统的
|
|
1045
|
+
// 不变量**(server 已按它铸),而本条是**同一条帧上的段内自洽**(gateOutcome.ts 顶注点名的
|
|
1046
|
+
// 那条反向先例,逐字就是本函数)。
|
|
1047
|
+
// 🔴 开集读 + 缺席不铸键(绝不折成空串/默认来源)。
|
|
1048
|
+
const deniedSource = reason === 'denied' && typeof a.deniedSource === 'string' && a.deniedSource.length > 0
|
|
1049
|
+
? a.deniedSource
|
|
1050
|
+
: undefined;
|
|
1051
|
+
return { armed, reason, ...(deniedSource !== undefined ? { deniedSource } : {}) };
|
|
1040
1052
|
}
|
|
1041
1053
|
/**
|
|
1042
1054
|
* `wiring_manifest` 帧 → 两个超集键的**纯投影**(公面导出;三端共用,壳侧绝不自抄一份形校验)。
|
|
@@ -75,8 +75,226 @@
|
|
|
75
75
|
* 部署级「这台 worker 到底配没配价表」仍可另问 `Capabilities.pricingConfigured`(contract 02
|
|
76
76
|
* §2.10 VERIFY)。本投影器只做单位换算(/1e6)并原样过境。
|
|
77
77
|
*/
|
|
78
|
-
import type { AgentEvent } from '@sema-agent/sdk';
|
|
78
|
+
import type { AgentEvent, TaskStats } from '@sema-agent/sdk';
|
|
79
79
|
import { type SDKMessage, type EmitContext } from '../types.js';
|
|
80
|
+
import { type SemaModelUsage } from './turnUsageToModelUsage.js';
|
|
81
|
+
/**
|
|
82
|
+
* CC 的 `NonNullableUsage` 占位形(coreSchemas 里是 `z.unknown()`,这里给它一个名字)
|
|
83
|
+
* + 0.66.0 / D-2 的两位 sema 超集(两位都**只在该说话时在场**,绝不铸 `false` / 假 0)。
|
|
84
|
+
*/
|
|
85
|
+
export interface SemaFlatUsage {
|
|
86
|
+
readonly inputTokens: number;
|
|
87
|
+
readonly outputTokens: number;
|
|
88
|
+
readonly cacheReadInputTokens: number;
|
|
89
|
+
readonly cacheCreationInputTokens: number;
|
|
90
|
+
readonly webSearchRequests: number;
|
|
91
|
+
/**
|
|
92
|
+
* D-2 / L-192② —— 在场且为 `true` ⇒ 同行的 `webSearchRequests: 0` 是「**wire 上没有这本账**」,
|
|
93
|
+
* 不是「这条 run 一次网搜都没做」。core 的 `TaskStats` 面上**没有**这一格(网搜只在散文里出现),
|
|
94
|
+
* 而 CC 的这一格型面是必填 `number` ⇒ 「不知道」只能以判别位在场(与 `_sema_cost_absent`
|
|
95
|
+
* 同一条两键合读纪律)。上游哪天真发这一格,值照读、本位退场(门里有正反两控)。
|
|
96
|
+
*/
|
|
97
|
+
readonly _sema_web_search_requests_absent?: true;
|
|
98
|
+
/**
|
|
99
|
+
* D-2 / L-192② —— `TaskStats.totalInputTokens`:**cache-INCLUSIVE** 的输入总量
|
|
100
|
+
* (core 逐字「the quantity cost is computed from … 'how much context did this task present'」)。
|
|
101
|
+
*
|
|
102
|
+
* 🔴 **为什么另开一格而不是灌进 `inputTokens`**:core RB-457-a(3.0.0 BREAKING)把
|
|
103
|
+
* `promptTokens` 的语义翻成 cache **MISS** 分量,而 CC / Anthropic 的 `input_tokens` 本义
|
|
104
|
+
* 正是 MISS —— 同帧已经另有 `cacheReadInputTokens` / `cacheCreationInputTokens` 两格,把总量
|
|
105
|
+
* 灌进 `inputTokens` 就是 core 逐字点名的那笔双算(「a 98% hit rate surfaced as 49.5%」)。
|
|
106
|
+
* 🔴 **为什么这一位必须长在扁平 usage 上**(与 [2295]「CC 形状不承载非 CC 语义」不矛盾,
|
|
107
|
+
* 差别与 `_sema_cost_absent` 那条同源):per-turn 那一面有**逐字通道**
|
|
108
|
+
* (`EngineTurnUsage` / chrome `last_turn_usage.engineUsage`),镜像不必长第二个座位;而
|
|
109
|
+
* **终帧这一面没有任何别的载体** —— 不给座位,这条 run 到底摆了多少上下文在终帧上就问不出来。
|
|
110
|
+
* 缺席 = wire 没报(老引擎 / 网关没报 usage),**绝不**拿 `promptTokens` 冒充总量。
|
|
111
|
+
*/
|
|
112
|
+
readonly _sema_total_input_tokens?: number;
|
|
113
|
+
/**
|
|
114
|
+
* D-2 族扫(0.66.0;异源对抗复审 [medium])—— `TaskStats.cacheWriteTokensLong`:**1 小时 TTL**
|
|
115
|
+
* 那一档的缓存写入分量。缺席 = wire 没报;显式 `0` 是**读数**(core 逐字「0 unless 1h caching
|
|
116
|
+
* is in use」),不是缺席。
|
|
117
|
+
* 🔴 **刻意不求和进 `cacheCreationInputTokens`**:core 对这两格的说法互相矛盾 ——
|
|
118
|
+
* `cacheWriteTokens` 自述是「Anthropic `cache_creation_input_tokens`」(协议侧那个量本就含 1h
|
|
119
|
+
* 子项),而 `totalInputTokens` 的求和式又把两格**并列相加**(那要求它们互不重叠)。相加与不加
|
|
120
|
+
* 各有一种错法,证据不足**不猜**:CC 那一格逐字保持 `cacheWriteTokens`(既有值一字不动),长 TTL
|
|
121
|
+
* 分量原样另给,对账由消费方按两个数自己做。上游澄清后再定(登记见 INTEGRATION §31d)。
|
|
122
|
+
*/
|
|
123
|
+
readonly _sema_cache_write_tokens_long?: number;
|
|
124
|
+
}
|
|
125
|
+
/**
|
|
126
|
+
* D-1 / L-192①(0.66.0)—— 一条被拒记录的**全可选**形:CC `SDKPermissionDenial`
|
|
127
|
+
* (agent-types `permissions.d.ts`,真形 `{tool_name: string; tool_use_id: string;
|
|
128
|
+
* tool_input: Record<string, unknown>}` **三键必填**)的三键 + 一位 sema 超集,**每一键都按
|
|
129
|
+
* wire 能不能兑现决定在不在**。
|
|
130
|
+
*
|
|
131
|
+
* 🔴 **今天三键里只兑现得出一键**:这本账在 wire 上是 `TaskStats.humanReview.gates[]`,而 core
|
|
132
|
+
* 的 gate 记录只有五格(`kind` / `waitMs` / `decision?` / `toolName?` / `toolArg?`,
|
|
133
|
+
* `task-result.d.ts` 真字节)—— **没有** `toolCallId`,**没有**被拒时的入参对象(core 逐字:
|
|
134
|
+
* 「a durable-resume gate carries `toolName` only (its input is not threaded onto the persisted
|
|
135
|
+
* gate — a documented follow-on)」;sdk `types.d.ts` 的同一格也逐字记着「`tool_input`/
|
|
136
|
+
* `toolInput` **不在** gate ledger 上」)。⇒ 那两键**缺席**,绝不铸 `""` / `{}`:一个空对象在
|
|
137
|
+
* CC 形上读起来是「这次调用的入参是空的」,那是编的。
|
|
138
|
+
* 🔴 **两键仍然声明在这里**(不是假 affordance):mint 点对它们是**开集宽读** —— 上游哪天把
|
|
139
|
+
* 两格串上 gate 账本,这一形与 CC 那条清单**自动**开始带值(见 {@link permissionDenialParts}
|
|
140
|
+
* 的自动升级腿),门里同时钉着今天的缺席证据与那一天的正控。消费端读它们**必须按可选位读**。
|
|
141
|
+
* 🔴 `_sema_tool_arg` = core 已经**脱敏并截短**的一行入参摘要(`primaryActivityArg` 同一道口),
|
|
142
|
+
* 它是 CC「denied: Bash(rm …)」那行显示唯一拿得到的材料。UNTRUSTED-for-display:只渲染,
|
|
143
|
+
* 绝不回喂模型、绝不当鉴权判据。
|
|
144
|
+
*/
|
|
145
|
+
export interface SemaPermissionDenial {
|
|
146
|
+
/** 被拒的工具名(⇐ `gates[].toolName`);wire 没报 ⇒ 键缺席,绝不编一个名字。 */
|
|
147
|
+
readonly tool_name?: string;
|
|
148
|
+
/** 被拒的**那一次调用**(⇐ `gates[].toolCallId`,今天 wire 上没有 ⇒ 恒缺席,见下方 mint 点头注)。 */
|
|
149
|
+
readonly tool_use_id?: string;
|
|
150
|
+
/** 被拒调用的**完整入参**(⇐ `gates[].toolInput`,今天 wire 上没有 ⇒ 恒缺席;非对象一律不铸)。 */
|
|
151
|
+
readonly tool_input?: Record<string, unknown>;
|
|
152
|
+
/** 被拒调用的一行入参摘要(⇐ `gates[].toolArg`,core 侧已脱敏截短);缺席 = 这条腿没串入参。 */
|
|
153
|
+
readonly _sema_tool_arg?: string;
|
|
154
|
+
}
|
|
155
|
+
/**
|
|
156
|
+
* D-3 / B-068 · L-198(0.66.0)—— core `TaskStats.costBreakdown`(`task-result.d.ts` 的
|
|
157
|
+
* finance taxonomy)的**窄读投影**,单位 = **micro-USD 原值**(键名即单位,包不折 USD:
|
|
158
|
+
* 折一次就多一次浮点漂移,而这一面正是账单面)。
|
|
159
|
+
*
|
|
160
|
+
* 每一格都是**可选**的:读不出的键不铸(绝不补 0 —— core 明说 unpriced 时整段与 `costMicroUsd`
|
|
161
|
+
* 一起省略,而一个补出来的 0 在账单面上就是一句「这一段没花钱」的假话)。**开集**:core 往这段
|
|
162
|
+
* 里加新类目时原样过境(sdk 的型面是 `costBreakdown?: unknown`,词表属主在 core)。
|
|
163
|
+
*/
|
|
164
|
+
export interface SemaCostBreakdown {
|
|
165
|
+
/** 根 agent 的 LLM 花费 = `costMicroUsd − compactionMicroUsd`(**不减 nested**)。 */
|
|
166
|
+
readonly llmRootMicroUsd?: number;
|
|
167
|
+
/** 委派子代的 LLM 花费 = `nested?.costMicroUsd ?? 0`;它**在 `costMicroUsd` 之外**。 */
|
|
168
|
+
readonly nestedSubagentMicroUsd?: number;
|
|
169
|
+
readonly memoryConsolidationMicroUsd?: number;
|
|
170
|
+
readonly suggestionsMicroUsd?: number;
|
|
171
|
+
/** 任务内压缩的 LLM 花费 —— 它**在 `costMicroUsd` 里面**(所以从 root 里减掉)。 */
|
|
172
|
+
readonly compactionMicroUsd?: number;
|
|
173
|
+
/** 开集:core 新加的类目原样过境(消费方 switch 必须带 default)。 */
|
|
174
|
+
readonly [k: string]: number | undefined;
|
|
175
|
+
}
|
|
176
|
+
/**
|
|
177
|
+
* D-3 —— core `NestedUsage`(`tool-spec.d.ts`)的窄读投影:这条 run **委派出去**的那本账。
|
|
178
|
+
* 🔴 `costMicroUsd` **缺席 = 委派花费没定价**(core 逐字 `ABSENT when the delegated spend was
|
|
179
|
+
* unpriced (RB-368) — never a fabricated 0`)⇒ 键缺席,绝不铸 0。
|
|
180
|
+
*/
|
|
181
|
+
export interface SemaNestedUsage {
|
|
182
|
+
readonly tokens?: number;
|
|
183
|
+
readonly turns?: number;
|
|
184
|
+
readonly tasks?: number;
|
|
185
|
+
readonly costMicroUsd?: number;
|
|
186
|
+
}
|
|
187
|
+
/**
|
|
188
|
+
* D-3 —— 终局**对账三段 + 两个判别位**。这正是 chrome 臂 `run_cost_reconciled` 的载荷本体:
|
|
189
|
+
* 两面共用**同一个**读器,所以「终帧超集键」与「chrome 对账臂」永远不会各算各的
|
|
190
|
+
* ([paired-mechanisms-must-share-premise])。
|
|
191
|
+
*
|
|
192
|
+
* core 的两条对账式(`task-result.d.ts` 逐字):
|
|
193
|
+
* · `llmRootMicroUsd + compactionMicroUsd === costMicroUsd`(压缩在 own 里面);
|
|
194
|
+
* · fully-reconciled spend = `costMicroUsd + nested.costMicroUsd`(子代在 own **外面**)。
|
|
195
|
+
* 🔴 本包**不当第二个会计**:三段照实过境,不改数、不补差;`reconciledMicroUsd` 只在**两段都
|
|
196
|
+
* 读得出**时才铸 —— 少了任何一边,总额就是不知道,而「不知道」只能以判别位在场。
|
|
197
|
+
*/
|
|
198
|
+
export interface RunCostReconcile {
|
|
199
|
+
/** 本任务 own 花费(⇐ `costMicroUsd`),**不含**子代。缺席 ⇒ 没定价,见 `costAbsent`。 */
|
|
200
|
+
readonly ownMicroUsd?: number;
|
|
201
|
+
/** 委派子代花费(⇐ `nested.costMicroUsd`)。缺席 ⇒ 没委派、或委派花费没定价(见判别位)。 */
|
|
202
|
+
readonly nestedMicroUsd?: number;
|
|
203
|
+
/** 任务内压缩花费(⇐ `costBreakdown.compactionMicroUsd`);它已含在 `ownMicroUsd` 里。 */
|
|
204
|
+
readonly compactionMicroUsd?: number;
|
|
205
|
+
/** 根 agent 花费(⇐ `costBreakdown.llmRootMicroUsd`);`own − compaction`。 */
|
|
206
|
+
readonly llmRootMicroUsd?: number;
|
|
207
|
+
/** `own + nested` —— core 逐字的 fully-reconciled spend;**任一段不知道就不铸**。 */
|
|
208
|
+
readonly reconciledMicroUsd?: number;
|
|
209
|
+
/** `true` ⇒ own 花费**没定价**(不是 0)。绝不铸 `false`。 */
|
|
210
|
+
readonly costAbsent?: true;
|
|
211
|
+
/** `true` ⇒ 委派过,但那本账**没定价**(不是 0)。绝不铸 `false`。 */
|
|
212
|
+
readonly nestedCostAbsent?: true;
|
|
213
|
+
/**
|
|
214
|
+
* 0.67.0(core 7.14.0 `TaskResult.stats.usageMissing`)—— `true` ⇒ 这条 run 上**至少有一轮**
|
|
215
|
+
* 没报 usage,所以本对账里的每一个数字(以及终帧 `usage` / `modelUsage` / `_sema_nested_usage`
|
|
216
|
+
* 的每一个 token 数)都是**下界**,不是一笔已知的账。
|
|
217
|
+
*
|
|
218
|
+
* 🔴 **`tokens: 0` 在这一位在场时读作「不知道」,不是「免费」**(core 合同逐字)。
|
|
219
|
+
* 🔴 **never false**:core 只在真缺 usage 时铸 `true`,缺席 ⇔ 每一轮都报了 usage ⇒ 本包同律
|
|
220
|
+
* **缺席不铸**(铸一个 `false` 就是替引擎说「我全都数到了」)。
|
|
221
|
+
* ⚠️ 它**与成本位正交**:`costAbsent` 说的是「没定价」,本位说的是「数得不全」——
|
|
222
|
+
* 一笔定了价、但少数了几轮的账,两位可以同时在场,渲染面要分别说。
|
|
223
|
+
*/
|
|
224
|
+
readonly usageLowerBound?: true;
|
|
225
|
+
}
|
|
226
|
+
/** {@link readRunCostFacts} 的产物:两面(终帧超集键 / chrome 对账臂)各取所需。 */
|
|
227
|
+
export interface RunCostFacts {
|
|
228
|
+
/** 终帧 `_sema_cost_breakdown` 的值;一个键都读不出 ⇒ `undefined`(不铸空对象)。 */
|
|
229
|
+
readonly breakdown?: SemaCostBreakdown;
|
|
230
|
+
/** 终帧 `_sema_nested_usage` 的值;同上。 */
|
|
231
|
+
readonly nested?: SemaNestedUsage;
|
|
232
|
+
/** chrome 臂 `run_cost_reconciled` 的载荷本体。 */
|
|
233
|
+
readonly reconcile: RunCostReconcile;
|
|
234
|
+
}
|
|
235
|
+
/**
|
|
236
|
+
* D-3 / B-068 · L-198 —— 终局成本事实的**唯一读器**(终帧超集键与 chrome 对账臂共用)。
|
|
237
|
+
*
|
|
238
|
+
* 🔴 `stats` 不是可读对象(409 拒绝信封 / park 体 / `failed` 事件帧)⇒ 返 `undefined` =
|
|
239
|
+
* **这条帧没有账**,调用方据此「不说话」(不发臂、不铸键),而不是发一条全缺席的空账。
|
|
240
|
+
*/
|
|
241
|
+
export declare function readRunCostFacts(stats: TaskStats | undefined): RunCostFacts | undefined;
|
|
242
|
+
/**
|
|
243
|
+
* L-228(0.67.0)—— 一只子任务在**这条流上被看见的**那本账(`_sema_nested_usage_by_task` 的行形)。
|
|
244
|
+
*
|
|
245
|
+
* 🔴 **它是流内累加的产物,不是引擎报的一个字段**:core 终局只有合计 `stats.nested`
|
|
246
|
+
* (`{tokens, turns, tasks, costMicroUsd}`),**没有 per-task 分项**(`task-result.d.ts` #594-660 亲验)。
|
|
247
|
+
* ⇒ 这几个数的唯一来源是子流 `turn_end` 的逐轮累加,而「这条流看见了多少」与「这条 run 一共有多少」
|
|
248
|
+
* 可以不相等 —— 那正是 {@link SemaNestedUsageByTask.partial} 存在的理由。
|
|
249
|
+
* 🔴 **没有 `costMicroUsd`**:`turn_end.usage` 上没有钱这一格(定价在终局做),编一个出来就是造账。
|
|
250
|
+
*/
|
|
251
|
+
export interface SemaSubagentUsageRow {
|
|
252
|
+
/** 这条流上看见的该子任务 `turn_end` 条数(**不是** core 的 `nested.turns`,见顶注)。 */
|
|
253
|
+
readonly turns: number;
|
|
254
|
+
/** 逐轮 `inputTokens` 求和。`usageMissing` 在场时它是**下界**。 */
|
|
255
|
+
readonly inputTokens: number;
|
|
256
|
+
/** 逐轮 `outputTokens` 求和。同上。 */
|
|
257
|
+
readonly outputTokens: number;
|
|
258
|
+
/** 逐轮 `cacheReadTokens` 求和;一轮都没报过 ⇒ **键不在**(绝不铸 0 冒充「零命中」)。 */
|
|
259
|
+
readonly cacheReadTokens?: number;
|
|
260
|
+
/** `true` ⇒ 这只子任务**至少有一轮**没报 usage,本行三个数是**下界**。never false。 */
|
|
261
|
+
readonly usageMissing?: true;
|
|
262
|
+
}
|
|
263
|
+
/** {@link SemaSubagentUsageRow} 的累加中间态(runStream 持有;`readonly` 在收口那一拍才加)。 */
|
|
264
|
+
export interface MutableSubagentUsageRow {
|
|
265
|
+
turns: number;
|
|
266
|
+
inputTokens: number;
|
|
267
|
+
outputTokens: number;
|
|
268
|
+
cacheReadTokens?: number;
|
|
269
|
+
usageMissing?: true;
|
|
270
|
+
}
|
|
271
|
+
/** 终帧两个超集键的产物形(见 {@link nestedUsageByTaskParts})。 */
|
|
272
|
+
export interface SemaNestedUsageByTask {
|
|
273
|
+
readonly rows: Readonly<Record<string, SemaSubagentUsageRow>>;
|
|
274
|
+
/** 见 `_sema_nested_usage_by_task_partial`。 */
|
|
275
|
+
readonly partial: boolean;
|
|
276
|
+
}
|
|
277
|
+
/**
|
|
278
|
+
* D-2 族扫(0.66.0;异源对抗复审 [medium])—— **终局** per-model 行的 sema 超集位。
|
|
279
|
+
*
|
|
280
|
+
* 🔴 为什么只长在终局这一面:per-turn 那一面有**逐字通道**(`EngineTurnUsage` /
|
|
281
|
+
* chrome `last_turn_usage.engineUsage`,整对象原形过境),镜像不必长第二个座位([2295] 裁 ②);
|
|
282
|
+
* 而**终局的 per-model 分表没有任何逐字通道** —— 不给座位,多模型部署下「哪个模型摆了多少上下文 /
|
|
283
|
+
* 那一行的分量口径可不可信」在终帧上就问不出来。故本形只由 `mapModelUsage` / `modelUsageFor`
|
|
284
|
+
* (两条终局腿)加挂,`toCcModelUsage` 那个共用 mint 点一个字不动。
|
|
285
|
+
*/
|
|
286
|
+
export interface SemaTerminalModelUsage extends SemaModelUsage {
|
|
287
|
+
/** ⇐ 行上的 `totalInputTokens`(cache-INCLUSIVE 总量);缺席 = 这一行没报。 */
|
|
288
|
+
readonly _sema_total_input_tokens?: number;
|
|
289
|
+
/**
|
|
290
|
+
* ⇐ 行上的 `usageBasis` —— **口径版本标记**(sdk 逐字:`"uncached-components-v1"` = 三个输入
|
|
291
|
+
* 分量键互不重叠;缺席 = 口径不可保证,可能是存量行、也可能是跨阶段混窗聚合)。**开集串**:
|
|
292
|
+
* 原样过境,消费方 `switch` 必须带 `default`;🔴 **缺席不许倒推口径**(sdk 明令)。
|
|
293
|
+
*/
|
|
294
|
+
readonly _sema_usage_basis?: string;
|
|
295
|
+
/** ⇐ 行上的 `cacheWriteTokensLong`;语义与不求和的理由见 `SemaFlatUsage` 的同名位。 */
|
|
296
|
+
readonly _sema_cache_write_tokens_long?: number;
|
|
297
|
+
}
|
|
80
298
|
/** `done` → SDKResultSuccess (contract 02 §2.10 / 08 CS-10). */
|
|
81
299
|
export declare function doneToSdkResult(ev: Extract<AgentEvent, {
|
|
82
300
|
type: 'done';
|