@sema-agent/client-core 0.65.0 → 0.66.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/README.md +3 -1
- package/dist/adapt/panelTasks.d.ts +2 -2
- package/dist/adapt/panelTasks.js +20 -15
- package/dist/adapter/downstream/terminalToSdkResult.d.ts +172 -1
- package/dist/adapter/downstream/terminalToSdkResult.js +239 -7
- package/dist/adapter/downstream/turnUsageToModelUsage.d.ts +20 -1
- package/dist/adapter/downstream/turnUsageToModelUsage.js +7 -2
- package/dist/adapter/runStream.js +43 -5
- package/dist/seam.d.ts +29 -0
- package/dist/seam.js +9 -0
- package/docs/INTEGRATION-CLIENTS.md +287 -7
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,83 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.66.0(2026-09-11)
|
|
53
|
+
|
|
54
|
+
> **终局真值三件**(cli 台账 L-192①② / B-068 · DEBTS L-198):三件都是同一条病形 ——
|
|
55
|
+
> 上游把事实摆在 wire 上,而包边界用一个**字面量**把它答成了常数。处置也只有一条:CC 形上必填的位
|
|
56
|
+
> **值不动**(旧消费者逐字不变),「不知道」交给同行的 `_sema_` 判别位;CC 形上没有的事实另开超集座位,
|
|
57
|
+
> **绝不改 CC 同名键的语义**。接入文档 §31(速览 / 逐件键表与缺席语义 / 黑盒判据 G31-1..G31-16 /
|
|
58
|
+
> 三端待办 / 本批的门)。判据编号 G31-1..G31-18。
|
|
59
|
+
|
|
60
|
+
- **① L-192① 终帧拒绝清单不再硬编 `[]`**(`adapter/downstream/terminalToSdkResult.ts`,成功臂与错误
|
|
61
|
+
信封**两处**):按 `TaskStats.humanReview.gates[]` 里 `decision === "deny"` 的行**逐条**铸记录,
|
|
62
|
+
顺序 = wire 顺序、条数 = 被拒次数,键按 wire 能兑现的铸(`tool_name` ⇐ `toolName`;入参摘要
|
|
63
|
+
`_sema_tool_arg` ⇐ `toolArg`,引擎侧已脱敏截短)。🔴 **两条清单**:CC 的 `SDKPermissionDenial`
|
|
64
|
+
三键**全是必填**,而 core 的 gate 账本给不出 `tool_use_id` / `tool_input`(逐字:耐久 resume 的 gate
|
|
65
|
+
只带 `toolName`)—— 补零补空是**编造**,把半条记录塞进 CC 数组则**破坏元素契约**(严格消费方
|
|
66
|
+
`safeParse` 会把**整条 result 帧**判非法)⇒ CC 数组只收三键齐全的记录(今天恒空,上游补齐后
|
|
67
|
+
**自动**开始填,包侧一行不用改),wire 上真有的每一条走超集载体 `_sema_permission_denials`。
|
|
68
|
+
顶层判别位 `_sema_permission_denials_absent: true` = 「CC 那条清单不可声称完整」,四条路径:
|
|
69
|
+
账本读不出(老引擎 / 409 拒绝信封 / `failed` 事件帧)/ 有读不出的行 / 有**认不出的判词**
|
|
70
|
+
(两张判词表之外的词、以及缺判词的耐久 wake 行)/ 有记录只在超集载体上。⇒「零拒绝」与
|
|
71
|
+
「清单不完整」从此可分。门 `run-permission-denial-projection-test.mjs`(新,68 checks,含**缺席证据**
|
|
72
|
+
与**自动升级腿**的正控:上游哪天把两格串上账本当场红/当场填)。
|
|
73
|
+
- **② B-068 / L-198 终局成本对账**(同文件 + `adapter/runStream.ts` + `seam.ts`):`TaskStats.costBreakdown`
|
|
74
|
+
与 `nested` 两段投上终帧超集键 `_sema_cost_breakdown` / `_sema_nested_usage`,**micro-USD 原值不折 USD**,
|
|
75
|
+
逐键窄读、坏键剥掉不连坐、整段读不出不铸空对象、上游新类目原样过境;nested 的 `costMicroUsd` 缺席 =
|
|
76
|
+
委派花费**没定价** ⇒ 键缺席,绝不铸 0。新增 chrome 臂 **`run_cost_reconciled`**(`required: false`):
|
|
77
|
+
终局那一拍交出 own / nested / compaction 三段 + `reconciledMicroUsd = own + nested`,任一段不知道
|
|
78
|
+
(或和本身非有限)则只铸 `costAbsent` / `nestedCostAbsent` 判别位而**不铸**对账值;`stats` 读不出的
|
|
79
|
+
终帧(409 / park 体 / `failed` 事件帧 / 畸形载体)一条都不发,非成功终局照发。🔴 「委派过」的判据是
|
|
80
|
+
`stats.nested` **这个载体在不在**,不是它里面有没有读得出的数 —— 否则 `nested: {}` 会被当成「没委派」
|
|
81
|
+
而把 own 铸成一个**确定的总额**。🔴 **CC 同名键 `total_cost_usd` 语义一字未改**(仍是 own 花费,
|
|
82
|
+
nested 不折进去);对账值由消费方按超集键自己加。🔴 子流 `turn_end` 仍不发 `last_turn_usage`(既有断闸
|
|
83
|
+
一字未动)—— 子代花费只经终局 `nested` 到账,端**不要**把流中增量与终局总账相加。
|
|
84
|
+
新增公面导出 **`readRunCostFacts`**(终帧两个超集键与 chrome 臂**共用**它 ⇒ 两面不会各算各的)。
|
|
85
|
+
门 `run-cost-reconcile-projection-test.mjs`(新,76 checks)。
|
|
86
|
+
- **③ L-192② 终帧扁平 `usage` 三格**(同文件 + `adapter/downstream/turnUsageToModelUsage.ts`):
|
|
87
|
+
`webSearchRequests` 从**字面量 0** 改按 CC 的名字开集宽读,读不出时值仍 0(CC 形必填 number)+ 同行
|
|
88
|
+
判别位 `_sema_web_search_requests_absent` —— **两个 mint 点**(终帧扁平 usage 与 CC `ModelUsage` 镜像)
|
|
89
|
+
同扫,同形不留第二处。新增 cache-INCLUSIVE 总量座位 `usage._sema_total_input_tokens`
|
|
90
|
+
(⇐ `TaskStats.totalInputTokens`);终帧此前**没有任何载体**能说出「这条 run 摆了多少上下文」。
|
|
91
|
+
同批按**同一条理由**给**终局 per-model 行**补三座位(`_sema_total_input_tokens` /
|
|
92
|
+
`_sema_usage_basis` 口径标记 / `_sema_cache_write_tokens_long`)—— 一个全局总量答不了「哪个模型摆了
|
|
93
|
+
多少」,而那张分表同样没有逐字通道;加挂只发生在两条**终局**腿上,per-turn 与终局共用的那个 mint 点
|
|
94
|
+
一个字不动(门里有反向钉)。扁平 usage 同批补 `_sema_cache_write_tokens_long`:🔴 **只另给不相加** ——
|
|
95
|
+
core 对 `cacheWriteTokens` 与 `cacheWriteTokensLong` 的说法互相矛盾(前者自述是协议侧那个**含 1h 子项**
|
|
96
|
+
的量,而总量式又把两格并列相加),相加与不加各有一种错法,**证据不足不猜**,CC 那一格逐字不动。
|
|
97
|
+
🔴 **`inputTokens` 逐字不动**,仍是 cache-**MISS** 分量:台账原句要求改读 `totalInputTokens`,核合同后
|
|
98
|
+
**不采** —— core RB-457-a 把 `promptTokens` 翻成 MISS 分量而 CC/Anthropic 的 `input_tokens` 本义正是 MISS,
|
|
99
|
+
灌总量会与同帧两个 cache 格**双算**(core 逐字:98% 命中率被渲成 49.5%),那是改 CC 同名键语义不是超集。
|
|
100
|
+
理由与「本批刻意不投的位」(`contextWindow`/`maxOutputTokens`、`cacheHitRate`/`toolCalls`/`mechanisms` 等)
|
|
101
|
+
逐条登记在 §31d。门 `run-cost-absence-projection-test.mjs` 扩 E/F/G/H 四段(144 checks)。
|
|
102
|
+
- **④ 异源对抗复审(查漏轨)采纳四件**:①**[high]** 终局 chrome 回调的**异步**拒绝此前会外溢成
|
|
103
|
+
`unhandledRejection`(端口契约是 `void | Promise<void>`,而 `try/catch` 只接得住同步抛)——
|
|
104
|
+
把未处理拒绝当致命的宿主会整只退出;修 = 新增共用发射口 `emitChromeFireAndForget`(同步抛在
|
|
105
|
+
`try` 里吞、异步拒绝挂 `.catch`),**同形族扫**把既有的 `last_turn_usage` 那条腿一并收编,语义
|
|
106
|
+
逐字不变(不 await、不改时序);②nested 存在性口径(见上);③CC 元素契约(见 ①);
|
|
107
|
+
④两处同形漏键(长 TTL 分量 / 终局 per-model 三座位)。
|
|
108
|
+
- **消费方影响**:只读既有键的端**零改**(CC 必填位的值在同样输入下一字不变);要新事实的端按 §31e
|
|
109
|
+
三端待办接。超集键 10 行登记进 `docs/type-superset.json`(常驻门双向对账);`unknown` 出境棘轮
|
|
110
|
+
320→321(逐条记账在 `run-client-core-typeshape-test.mjs`:`SemaPermissionDenial.tool_input` 是
|
|
111
|
+
CC 契约**本身**的形 `Record<string, unknown>`)。
|
|
112
|
+
|
|
113
|
+
## 0.65.1(2026-09-11)
|
|
114
|
+
|
|
115
|
+
test [6961] 对 0.65.0 十件五路交叉验证全 PASS,对抗复审轨另抓两处带坐标的实现缺口(cli 亲核源码成立,两处都是 patch):
|
|
116
|
+
|
|
117
|
+
- **B-087 通知腿判重次序**(`adapt/panelTasks.ts` `settleFromNotification`):拆分时逐字保留了「先发 `panel_task{kind:'end'}`
|
|
118
|
+
再判重」的旧次序,判重只护着 if-started stop;B-074 让终态 tick 也能结 bg 行之后,`tick(completed) → task_notification`
|
|
119
|
+
这个顺序在面板上落**两条** end(反序只一条)。修=判重移到该腿唯一的 yield 之前,与 tick 腿 / 关卡腿同形;常驻账照清。
|
|
120
|
+
§30d G30-11 补注;门 `run-task-progress-terminal-projection` +B5a/B5b(两个到达顺序都恰一条)。
|
|
121
|
+
- **B-088 裸 `usageMissing` 帧静默**(`adapter/runStream.ts` turn_end 折叠点):`turn_usage` 发臂条件只认「数值
|
|
122
|
+
`outputTokens` ∨ 非空 `stopReason`」,判别位自己不算话 ⇒ core 真铸形(`run-harness-handlers.ts:491`,该轮零 usage
|
|
123
|
+
帧时无 `usage` 字段;core [6962] 证实)`{type:'turn_end', usageMissing:true}` 整条静默,G30-23 那一档在最诚实的帧上
|
|
124
|
+
反而拿不到 `_sema_usage_missing`。修=三者任一在场即发;三者皆缺席仍不发(F4 逐字不变)。§30e 情形表 +一行、
|
|
125
|
+
G30-24b;门 `run-assistant-arm-identity` +F5a–c。
|
|
126
|
+
- 消费方零改:新发的那条臂只带判别位,adapt `turnUsageArm` 以 `typeof outputTokens === 'number'` 开门 ⇒ no-op;
|
|
127
|
+
面板 end 少一条重复,端侧无需处置。
|
|
128
|
+
|
|
52
129
|
## 0.65.0(2026-09-11)
|
|
53
130
|
|
|
54
131
|
> 两半场同版:**① 投影臂族归层修**(B-071 / B-072 / B-073 / B-074 / L-215③)与
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.
|
|
38
|
+
**Version:** 0.66.0
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -295,6 +295,8 @@ public-surface guard checks that last one).
|
|
|
295
295
|
| `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
|
|
296
296
|
| `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
|
|
297
297
|
| `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
|
|
298
|
+
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift |
|
|
299
|
+
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end |
|
|
298
300
|
| `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
|
|
299
301
|
| `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
|
|
300
302
|
|
|
@@ -12,8 +12,8 @@
|
|
|
12
12
|
*
|
|
13
13
|
* 🔴 拆分定案两条(矩阵 §3.2 #4/#6):
|
|
14
14
|
* ① `endedPanelTasks` 拆分前被 `settlePanelTasks` **和** task_notification 臂两处直写。
|
|
15
|
-
* 现在通知臂那半场收成 `settleFromNotification()`(clearResident +
|
|
16
|
-
*
|
|
15
|
+
* 现在通知臂那半场收成 `settleFromNotification()`(clearResident + endedPanelTasks 判重 +
|
|
16
|
+
* panel end + if-started stop 整段;0.65.1 起判重在 end 之前),通知臂只调一次 ⇒ 跨臂直写消失。
|
|
17
17
|
* ② `firedSubagentStartHookTaskIds` 的读者(通知臂)随之进了本模块内部 —— 通知臂**不再直读**
|
|
18
18
|
* 实例台账,那条跨模块读同批消失。
|
|
19
19
|
*
|
package/dist/adapt/panelTasks.js
CHANGED
|
@@ -163,31 +163,36 @@ export function createPanelTaskLedger(ctx, cards, inst) {
|
|
|
163
163
|
/**
|
|
164
164
|
* 通知臂(task_notification)那半场,整段收进来 —— 拆分前它在臂里直写 `endedPanelTasks`
|
|
165
165
|
* 与直读 `firedSubagentStartHookTaskIds`,是矩阵点名的两条跨模块面。
|
|
166
|
-
*
|
|
166
|
+
* 🔴 **判重在唯一的 yield 之前**(0.65.1 / B-087;test [6961] 对抗复审轨实抓):拆分时这里
|
|
167
|
+
* 「逐字保持」了拆分前的次序 —— 先发 end 行、再判重 —— 判重只护着 if-started stop,end
|
|
168
|
+
* 行本身无条件发。B-074 让终态 tick 也能结 bg 行之后,`tick(completed) → task_notification`
|
|
169
|
+
* 这个顺序就在面板上落**两条** end(反序只有一条,因为 tick 腿判重在前)。三条腿共用
|
|
170
|
+
* `endedPanelTasks` 的承诺(§30d G30-11)要在**每条腿**的 yield 之前成立;常驻账照清
|
|
171
|
+
* (清一次幂等,真终态权威性不变)。
|
|
167
172
|
*/
|
|
168
173
|
*settleFromNotification(taskId, isError) {
|
|
169
174
|
// #6 通知-settle 边(推送半场):session 常驻的 bg 子代行不再被 turn sweep 假结,
|
|
170
175
|
// 真终态在此落行(failed/killed ⇒ isError 真归因)。
|
|
171
176
|
clearEnginePanelTaskResident(taskId);
|
|
177
|
+
if (endedPanelTasks.has(taskId))
|
|
178
|
+
return;
|
|
179
|
+
endedPanelTasks.add(taskId);
|
|
172
180
|
yield chrome({
|
|
173
181
|
kind: 'panel_task',
|
|
174
182
|
laneProof: MAIN,
|
|
175
183
|
event: { kind: 'end', taskId, isError },
|
|
176
184
|
});
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
guard: 'if-started',
|
|
189
|
-
});
|
|
190
|
-
}
|
|
185
|
+
// SubagentStop 只对**真 fire 过 Start** 的子代(workflow runId / 外来 id 不配对乱 fire)。
|
|
186
|
+
// cli 同门 `firedSubagentStartHookTaskIds.has(taskId)` —— B4 前本臂无条件发 stop,
|
|
187
|
+
// 宿主若照 guard 注释自建台账才不会乱 fire;台账在库里就该库来判。
|
|
188
|
+
if (inst.hasFiredSubagentStart(taskId)) {
|
|
189
|
+
yield chrome({
|
|
190
|
+
kind: 'subagent_lifecycle',
|
|
191
|
+
laneProof: MAIN,
|
|
192
|
+
phase: 'stop',
|
|
193
|
+
taskId,
|
|
194
|
+
guard: 'if-started',
|
|
195
|
+
});
|
|
191
196
|
}
|
|
192
197
|
},
|
|
193
198
|
/**
|
|
@@ -75,8 +75,179 @@
|
|
|
75
75
|
* 部署级「这台 worker 到底配没配价表」仍可另问 `Capabilities.pricingConfigured`(contract 02
|
|
76
76
|
* §2.10 VERIFY)。本投影器只做单位换算(/1e6)并原样过境。
|
|
77
77
|
*/
|
|
78
|
-
import type { AgentEvent } from '@sema-agent/sdk';
|
|
78
|
+
import type { AgentEvent, TaskStats } from '@sema-agent/sdk';
|
|
79
79
|
import { type SDKMessage, type EmitContext } from '../types.js';
|
|
80
|
+
import { type SemaModelUsage } from './turnUsageToModelUsage.js';
|
|
81
|
+
/**
|
|
82
|
+
* CC 的 `NonNullableUsage` 占位形(coreSchemas 里是 `z.unknown()`,这里给它一个名字)
|
|
83
|
+
* + 0.66.0 / D-2 的两位 sema 超集(两位都**只在该说话时在场**,绝不铸 `false` / 假 0)。
|
|
84
|
+
*/
|
|
85
|
+
export interface SemaFlatUsage {
|
|
86
|
+
readonly inputTokens: number;
|
|
87
|
+
readonly outputTokens: number;
|
|
88
|
+
readonly cacheReadInputTokens: number;
|
|
89
|
+
readonly cacheCreationInputTokens: number;
|
|
90
|
+
readonly webSearchRequests: number;
|
|
91
|
+
/**
|
|
92
|
+
* D-2 / L-192② —— 在场且为 `true` ⇒ 同行的 `webSearchRequests: 0` 是「**wire 上没有这本账**」,
|
|
93
|
+
* 不是「这条 run 一次网搜都没做」。core 的 `TaskStats` 面上**没有**这一格(网搜只在散文里出现),
|
|
94
|
+
* 而 CC 的这一格型面是必填 `number` ⇒ 「不知道」只能以判别位在场(与 `_sema_cost_absent`
|
|
95
|
+
* 同一条两键合读纪律)。上游哪天真发这一格,值照读、本位退场(门里有正反两控)。
|
|
96
|
+
*/
|
|
97
|
+
readonly _sema_web_search_requests_absent?: true;
|
|
98
|
+
/**
|
|
99
|
+
* D-2 / L-192② —— `TaskStats.totalInputTokens`:**cache-INCLUSIVE** 的输入总量
|
|
100
|
+
* (core 逐字「the quantity cost is computed from … 'how much context did this task present'」)。
|
|
101
|
+
*
|
|
102
|
+
* 🔴 **为什么另开一格而不是灌进 `inputTokens`**:core RB-457-a(3.0.0 BREAKING)把
|
|
103
|
+
* `promptTokens` 的语义翻成 cache **MISS** 分量,而 CC / Anthropic 的 `input_tokens` 本义
|
|
104
|
+
* 正是 MISS —— 同帧已经另有 `cacheReadInputTokens` / `cacheCreationInputTokens` 两格,把总量
|
|
105
|
+
* 灌进 `inputTokens` 就是 core 逐字点名的那笔双算(「a 98% hit rate surfaced as 49.5%」)。
|
|
106
|
+
* 🔴 **为什么这一位必须长在扁平 usage 上**(与 [2295]「CC 形状不承载非 CC 语义」不矛盾,
|
|
107
|
+
* 差别与 `_sema_cost_absent` 那条同源):per-turn 那一面有**逐字通道**
|
|
108
|
+
* (`EngineTurnUsage` / chrome `last_turn_usage.engineUsage`),镜像不必长第二个座位;而
|
|
109
|
+
* **终帧这一面没有任何别的载体** —— 不给座位,这条 run 到底摆了多少上下文在终帧上就问不出来。
|
|
110
|
+
* 缺席 = wire 没报(老引擎 / 网关没报 usage),**绝不**拿 `promptTokens` 冒充总量。
|
|
111
|
+
*/
|
|
112
|
+
readonly _sema_total_input_tokens?: number;
|
|
113
|
+
/**
|
|
114
|
+
* D-2 族扫(0.66.0;异源对抗复审 [medium])—— `TaskStats.cacheWriteTokensLong`:**1 小时 TTL**
|
|
115
|
+
* 那一档的缓存写入分量。缺席 = wire 没报;显式 `0` 是**读数**(core 逐字「0 unless 1h caching
|
|
116
|
+
* is in use」),不是缺席。
|
|
117
|
+
* 🔴 **刻意不求和进 `cacheCreationInputTokens`**:core 对这两格的说法互相矛盾 ——
|
|
118
|
+
* `cacheWriteTokens` 自述是「Anthropic `cache_creation_input_tokens`」(协议侧那个量本就含 1h
|
|
119
|
+
* 子项),而 `totalInputTokens` 的求和式又把两格**并列相加**(那要求它们互不重叠)。相加与不加
|
|
120
|
+
* 各有一种错法,证据不足**不猜**:CC 那一格逐字保持 `cacheWriteTokens`(既有值一字不动),长 TTL
|
|
121
|
+
* 分量原样另给,对账由消费方按两个数自己做。上游澄清后再定(登记见 INTEGRATION §31d)。
|
|
122
|
+
*/
|
|
123
|
+
readonly _sema_cache_write_tokens_long?: number;
|
|
124
|
+
}
|
|
125
|
+
/**
|
|
126
|
+
* D-1 / L-192①(0.66.0)—— 一条被拒记录的**全可选**形:CC `SDKPermissionDenial`
|
|
127
|
+
* (agent-types `permissions.d.ts`,真形 `{tool_name: string; tool_use_id: string;
|
|
128
|
+
* tool_input: Record<string, unknown>}` **三键必填**)的三键 + 一位 sema 超集,**每一键都按
|
|
129
|
+
* wire 能不能兑现决定在不在**。
|
|
130
|
+
*
|
|
131
|
+
* 🔴 **今天三键里只兑现得出一键**:这本账在 wire 上是 `TaskStats.humanReview.gates[]`,而 core
|
|
132
|
+
* 的 gate 记录只有五格(`kind` / `waitMs` / `decision?` / `toolName?` / `toolArg?`,
|
|
133
|
+
* `task-result.d.ts` 真字节)—— **没有** `toolCallId`,**没有**被拒时的入参对象(core 逐字:
|
|
134
|
+
* 「a durable-resume gate carries `toolName` only (its input is not threaded onto the persisted
|
|
135
|
+
* gate — a documented follow-on)」;sdk `types.d.ts` 的同一格也逐字记着「`tool_input`/
|
|
136
|
+
* `toolInput` **不在** gate ledger 上」)。⇒ 那两键**缺席**,绝不铸 `""` / `{}`:一个空对象在
|
|
137
|
+
* CC 形上读起来是「这次调用的入参是空的」,那是编的。
|
|
138
|
+
* 🔴 **两键仍然声明在这里**(不是假 affordance):mint 点对它们是**开集宽读** —— 上游哪天把
|
|
139
|
+
* 两格串上 gate 账本,这一形与 CC 那条清单**自动**开始带值(见 {@link permissionDenialParts}
|
|
140
|
+
* 的自动升级腿),门里同时钉着今天的缺席证据与那一天的正控。消费端读它们**必须按可选位读**。
|
|
141
|
+
* 🔴 `_sema_tool_arg` = core 已经**脱敏并截短**的一行入参摘要(`primaryActivityArg` 同一道口),
|
|
142
|
+
* 它是 CC「denied: Bash(rm …)」那行显示唯一拿得到的材料。UNTRUSTED-for-display:只渲染,
|
|
143
|
+
* 绝不回喂模型、绝不当鉴权判据。
|
|
144
|
+
*/
|
|
145
|
+
export interface SemaPermissionDenial {
|
|
146
|
+
/** 被拒的工具名(⇐ `gates[].toolName`);wire 没报 ⇒ 键缺席,绝不编一个名字。 */
|
|
147
|
+
readonly tool_name?: string;
|
|
148
|
+
/** 被拒的**那一次调用**(⇐ `gates[].toolCallId`,今天 wire 上没有 ⇒ 恒缺席,见下方 mint 点头注)。 */
|
|
149
|
+
readonly tool_use_id?: string;
|
|
150
|
+
/** 被拒调用的**完整入参**(⇐ `gates[].toolInput`,今天 wire 上没有 ⇒ 恒缺席;非对象一律不铸)。 */
|
|
151
|
+
readonly tool_input?: Record<string, unknown>;
|
|
152
|
+
/** 被拒调用的一行入参摘要(⇐ `gates[].toolArg`,core 侧已脱敏截短);缺席 = 这条腿没串入参。 */
|
|
153
|
+
readonly _sema_tool_arg?: string;
|
|
154
|
+
}
|
|
155
|
+
/**
|
|
156
|
+
* D-3 / B-068 · L-198(0.66.0)—— core `TaskStats.costBreakdown`(`task-result.d.ts` 的
|
|
157
|
+
* finance taxonomy)的**窄读投影**,单位 = **micro-USD 原值**(键名即单位,包不折 USD:
|
|
158
|
+
* 折一次就多一次浮点漂移,而这一面正是账单面)。
|
|
159
|
+
*
|
|
160
|
+
* 每一格都是**可选**的:读不出的键不铸(绝不补 0 —— core 明说 unpriced 时整段与 `costMicroUsd`
|
|
161
|
+
* 一起省略,而一个补出来的 0 在账单面上就是一句「这一段没花钱」的假话)。**开集**:core 往这段
|
|
162
|
+
* 里加新类目时原样过境(sdk 的型面是 `costBreakdown?: unknown`,词表属主在 core)。
|
|
163
|
+
*/
|
|
164
|
+
export interface SemaCostBreakdown {
|
|
165
|
+
/** 根 agent 的 LLM 花费 = `costMicroUsd − compactionMicroUsd`(**不减 nested**)。 */
|
|
166
|
+
readonly llmRootMicroUsd?: number;
|
|
167
|
+
/** 委派子代的 LLM 花费 = `nested?.costMicroUsd ?? 0`;它**在 `costMicroUsd` 之外**。 */
|
|
168
|
+
readonly nestedSubagentMicroUsd?: number;
|
|
169
|
+
readonly memoryConsolidationMicroUsd?: number;
|
|
170
|
+
readonly suggestionsMicroUsd?: number;
|
|
171
|
+
/** 任务内压缩的 LLM 花费 —— 它**在 `costMicroUsd` 里面**(所以从 root 里减掉)。 */
|
|
172
|
+
readonly compactionMicroUsd?: number;
|
|
173
|
+
/** 开集:core 新加的类目原样过境(消费方 switch 必须带 default)。 */
|
|
174
|
+
readonly [k: string]: number | undefined;
|
|
175
|
+
}
|
|
176
|
+
/**
|
|
177
|
+
* D-3 —— core `NestedUsage`(`tool-spec.d.ts`)的窄读投影:这条 run **委派出去**的那本账。
|
|
178
|
+
* 🔴 `costMicroUsd` **缺席 = 委派花费没定价**(core 逐字 `ABSENT when the delegated spend was
|
|
179
|
+
* unpriced (RB-368) — never a fabricated 0`)⇒ 键缺席,绝不铸 0。
|
|
180
|
+
*/
|
|
181
|
+
export interface SemaNestedUsage {
|
|
182
|
+
readonly tokens?: number;
|
|
183
|
+
readonly turns?: number;
|
|
184
|
+
readonly tasks?: number;
|
|
185
|
+
readonly costMicroUsd?: number;
|
|
186
|
+
}
|
|
187
|
+
/**
|
|
188
|
+
* D-3 —— 终局**对账三段 + 两个判别位**。这正是 chrome 臂 `run_cost_reconciled` 的载荷本体:
|
|
189
|
+
* 两面共用**同一个**读器,所以「终帧超集键」与「chrome 对账臂」永远不会各算各的
|
|
190
|
+
* ([paired-mechanisms-must-share-premise])。
|
|
191
|
+
*
|
|
192
|
+
* core 的两条对账式(`task-result.d.ts` 逐字):
|
|
193
|
+
* · `llmRootMicroUsd + compactionMicroUsd === costMicroUsd`(压缩在 own 里面);
|
|
194
|
+
* · fully-reconciled spend = `costMicroUsd + nested.costMicroUsd`(子代在 own **外面**)。
|
|
195
|
+
* 🔴 本包**不当第二个会计**:三段照实过境,不改数、不补差;`reconciledMicroUsd` 只在**两段都
|
|
196
|
+
* 读得出**时才铸 —— 少了任何一边,总额就是不知道,而「不知道」只能以判别位在场。
|
|
197
|
+
*/
|
|
198
|
+
export interface RunCostReconcile {
|
|
199
|
+
/** 本任务 own 花费(⇐ `costMicroUsd`),**不含**子代。缺席 ⇒ 没定价,见 `costAbsent`。 */
|
|
200
|
+
readonly ownMicroUsd?: number;
|
|
201
|
+
/** 委派子代花费(⇐ `nested.costMicroUsd`)。缺席 ⇒ 没委派、或委派花费没定价(见判别位)。 */
|
|
202
|
+
readonly nestedMicroUsd?: number;
|
|
203
|
+
/** 任务内压缩花费(⇐ `costBreakdown.compactionMicroUsd`);它已含在 `ownMicroUsd` 里。 */
|
|
204
|
+
readonly compactionMicroUsd?: number;
|
|
205
|
+
/** 根 agent 花费(⇐ `costBreakdown.llmRootMicroUsd`);`own − compaction`。 */
|
|
206
|
+
readonly llmRootMicroUsd?: number;
|
|
207
|
+
/** `own + nested` —— core 逐字的 fully-reconciled spend;**任一段不知道就不铸**。 */
|
|
208
|
+
readonly reconciledMicroUsd?: number;
|
|
209
|
+
/** `true` ⇒ own 花费**没定价**(不是 0)。绝不铸 `false`。 */
|
|
210
|
+
readonly costAbsent?: true;
|
|
211
|
+
/** `true` ⇒ 委派过,但那本账**没定价**(不是 0)。绝不铸 `false`。 */
|
|
212
|
+
readonly nestedCostAbsent?: true;
|
|
213
|
+
}
|
|
214
|
+
/** {@link readRunCostFacts} 的产物:两面(终帧超集键 / chrome 对账臂)各取所需。 */
|
|
215
|
+
export interface RunCostFacts {
|
|
216
|
+
/** 终帧 `_sema_cost_breakdown` 的值;一个键都读不出 ⇒ `undefined`(不铸空对象)。 */
|
|
217
|
+
readonly breakdown?: SemaCostBreakdown;
|
|
218
|
+
/** 终帧 `_sema_nested_usage` 的值;同上。 */
|
|
219
|
+
readonly nested?: SemaNestedUsage;
|
|
220
|
+
/** chrome 臂 `run_cost_reconciled` 的载荷本体。 */
|
|
221
|
+
readonly reconcile: RunCostReconcile;
|
|
222
|
+
}
|
|
223
|
+
/**
|
|
224
|
+
* D-3 / B-068 · L-198 —— 终局成本事实的**唯一读器**(终帧超集键与 chrome 对账臂共用)。
|
|
225
|
+
*
|
|
226
|
+
* 🔴 `stats` 不是可读对象(409 拒绝信封 / park 体 / `failed` 事件帧)⇒ 返 `undefined` =
|
|
227
|
+
* **这条帧没有账**,调用方据此「不说话」(不发臂、不铸键),而不是发一条全缺席的空账。
|
|
228
|
+
*/
|
|
229
|
+
export declare function readRunCostFacts(stats: TaskStats | undefined): RunCostFacts | undefined;
|
|
230
|
+
/**
|
|
231
|
+
* D-2 族扫(0.66.0;异源对抗复审 [medium])—— **终局** per-model 行的 sema 超集位。
|
|
232
|
+
*
|
|
233
|
+
* 🔴 为什么只长在终局这一面:per-turn 那一面有**逐字通道**(`EngineTurnUsage` /
|
|
234
|
+
* chrome `last_turn_usage.engineUsage`,整对象原形过境),镜像不必长第二个座位([2295] 裁 ②);
|
|
235
|
+
* 而**终局的 per-model 分表没有任何逐字通道** —— 不给座位,多模型部署下「哪个模型摆了多少上下文 /
|
|
236
|
+
* 那一行的分量口径可不可信」在终帧上就问不出来。故本形只由 `mapModelUsage` / `modelUsageFor`
|
|
237
|
+
* (两条终局腿)加挂,`toCcModelUsage` 那个共用 mint 点一个字不动。
|
|
238
|
+
*/
|
|
239
|
+
export interface SemaTerminalModelUsage extends SemaModelUsage {
|
|
240
|
+
/** ⇐ 行上的 `totalInputTokens`(cache-INCLUSIVE 总量);缺席 = 这一行没报。 */
|
|
241
|
+
readonly _sema_total_input_tokens?: number;
|
|
242
|
+
/**
|
|
243
|
+
* ⇐ 行上的 `usageBasis` —— **口径版本标记**(sdk 逐字:`"uncached-components-v1"` = 三个输入
|
|
244
|
+
* 分量键互不重叠;缺席 = 口径不可保证,可能是存量行、也可能是跨阶段混窗聚合)。**开集串**:
|
|
245
|
+
* 原样过境,消费方 `switch` 必须带 `default`;🔴 **缺席不许倒推口径**(sdk 明令)。
|
|
246
|
+
*/
|
|
247
|
+
readonly _sema_usage_basis?: string;
|
|
248
|
+
/** ⇐ 行上的 `cacheWriteTokensLong`;语义与不求和的理由见 `SemaFlatUsage` 的同名位。 */
|
|
249
|
+
readonly _sema_cache_write_tokens_long?: number;
|
|
250
|
+
}
|
|
80
251
|
/** `done` → SDKResultSuccess (contract 02 §2.10 / 08 CS-10). */
|
|
81
252
|
export declare function doneToSdkResult(ev: Extract<AgentEvent, {
|
|
82
253
|
type: 'done';
|