@sema-agent/client-core 0.67.1 → 0.67.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +38 -0
- package/README.md +4 -3
- package/dist/adapter/downstream/terminalToSdkResult.d.ts +46 -6
- package/dist/adapter/downstream/terminalToSdkResult.js +98 -29
- package/dist/adapter/runStream.js +51 -14
- package/dist/adapter/types.d.ts +0 -27
- package/docs/INTEGRATION-CLIENTS.md +103 -2
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,44 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.67.2(2026-09-12)
|
|
53
|
+
|
|
54
|
+
**异源对抗复审轮一三条 [medium] 的修复批**(patch;零公面导出
|
|
55
|
+
增删、零 wire 键增删、零 BREAKING;逐项的形 / 判据 / 三端待办见 `docs/INTEGRATION-CLIENTS.md` §32f、
|
|
56
|
+
§32g 两段「0.67.2 订正」)。
|
|
57
|
+
|
|
58
|
+
- **I-1 流内 usage 缺口观测位是 per-stream,不再住在共享 `EmitContext` 上** —— 那一位此前写在
|
|
59
|
+
`ctx.usageMissingObserved`、且**只置 `true` 永不清**,而 ctx 是**调用方的对象**、可以复用给多条流
|
|
60
|
+
(`startedAtMs` 本来就是这么用的)⇒ 顺序复用时上一条流的缺口把**下一条账数得全**的流的终帧与
|
|
61
|
+
`run_cost_reconciled` 一起标成下界,并发复用时两条互串。修形:观测位落在 `runStreamInner` 的局部量上,
|
|
62
|
+
终局把**本流快照**按值同时交给两个投影口(`terminalToSdkResult` 第三参 / `readRunCostFacts` 第二参)——
|
|
63
|
+
两面仍读同一次计算,但谁都读不到别人那份。`EmitContext.usageMissingObserved` **退役**(该位原本就逐字
|
|
64
|
+
写着「宿主不要自己填」)。反钉:顺序复用第二条流两面都不铸 + 同 ctx 上真有缺口的第三条流照铸 +
|
|
65
|
+
并发交错互不串 + ctx 上不再留位。
|
|
66
|
+
- **I-2 两种身份的键空间碰撞不再绕过混合出身检测** —— 行键的两个命名空间(真身份 `sourceTaskId` /
|
|
67
|
+
回落 `parentToolCallId`)**字面可相等**而上游不保证互斥;修前裸 id 当累加键,撞字面的两行在累加层
|
|
68
|
+
就并掉、行上的出身位只剩一种 ⇒ 0.67.1 那道「出身混合 ⇒ partial 恒立」的闸读不出混合被绕过
|
|
69
|
+
(三帧反例:行数与轮数两条对账同时成立、`partial` 不铸,而 A 实际 40 / B 20 被错并成 30 / 30)。
|
|
70
|
+
修形:累加键按出身前缀隔离(`s:` / `p:`)、逐事件记来源;交付面的键**仍是裸 id**(端零改),跨空间
|
|
71
|
+
同字面合并成一行并立新的行级判别位 **`keyCollision: true`**(never false)、`partial` 恒立。
|
|
72
|
+
合并臂照走 `putOwn`(`__proto__` 纪律不因第二条路被绕开)。
|
|
73
|
+
- **I-2b 子代分表本身也不再住在共享 `EmitContext` 上**(件 I-1 的同形存量,轮二复审实抓)——
|
|
74
|
+
`ctx.nestedUsageByTask` 流结束后留在调用方对象上,而三只终帧投影器是**公面导出**:端「A 走
|
|
75
|
+
`runStream`、B 直调终帧投影」共用一个 ctx 时,B 的终帧带出 **A 的**分表,且 B 的 `nested` 计数恰好
|
|
76
|
+
对得上时连 `partial` 都不铸(一张属于别人的表被标成「可证完整」)。修形与 I-1 同一条:分表随终帧
|
|
77
|
+
第三参按值传,`EmitContext.nestedUsageByTask` **退役**;没传快照 ⇒ 两个分表键都不铸(诚实缺席)。
|
|
78
|
+
- **I-3 门负控的备份改独占创建** —— `scripts/run-gate-negative-controls-test.mjs` 的
|
|
79
|
+
`existsSync` + `copyFileSync` 之间没有互斥且复制默认允许覆盖,两实例交错即可把「唯一复原依据」
|
|
80
|
+
换成**已篡改**的内容(源复原不回来、备份也被删),而本套的全部立论就是「演练不留痕」。改
|
|
81
|
+
`copyFileSync(..., COPYFILE_EXCL)`(检查与创建同一次系统调用),已存在 ⇒ 响亮拒绝不覆盖;
|
|
82
|
+
`EEXIST` 之外的 errno 原样抛出。新增自证④ 在 `mkdtemp` 隔离目录里调度那次交错。
|
|
83
|
+
|
|
84
|
+
**消费方待办**:无。wire 键零增删;`_sema_nested_usage_by_task` 的行上多一个 additive 判别位
|
|
85
|
+
`keyCollision`(它在场时 `_sema_nested_usage_by_task_partial` 必然同时在场 ⇒ 只读 `partial` 的端行为
|
|
86
|
+
逐位不变);`EmitContext` 上退役两格(`usageMissingObserved` / `nestedUsageByTask`),三端零读写点、经
|
|
87
|
+
`runStream` 的路径逐位不变。cli 侧只需把 lock 抬到 0.67.2。
|
|
88
|
+
|
|
89
|
+
|
|
52
90
|
## 0.67.1(2026-09-12)
|
|
53
91
|
|
|
54
92
|
**test [7055] 对 0.67.0 的 G32-01~30 验证批回帖三处真实发现的修复批**(patch;零公面导出增删、零
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.67.
|
|
38
|
+
**Version:** 0.67.2
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -250,7 +250,7 @@ public-surface guard checks that last one).
|
|
|
250
250
|
| `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header |
|
|
251
251
|
| `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters |
|
|
252
252
|
| `scripts/run-subagent-content-budget-test.mjs` | The **byte** budget on the sub-agent transcript ledger. It used to be bounded only by *counts* — so many entries per child, so many children — and a count is not a budget when a single entry has no ceiling of its own: one tool result carrying an inlined attachment, or one long model answer, and a single slot sits on tens of megabytes. The guard anchors on how many bytes are **still held** after over-filling, not on whether truncation fired, because an implementation that flags the overflow without actually dropping anything satisfies the second and not the first. Dropping is required to leave a record — how much went and where the retained content now starts — and that record has to reach the render plan, because content that vanishes with no marker gives the reader a transcript shorter than what happened with nothing to say so; the record is one per child, updated in place, pinned to the front, and excluded from the budget it describes. Order matters and is checked: oldest entries go first and the live tail is trimmed only as a last resort, since taking the text the user is watching stream while older history survives is the wrong end. The total budget evicts a whole least-recently-used child rather than shaving every child, and the configuration surface is fail-loud on zero, negatives, non-finite and non-integer values — a silently ignored budget is the exact failure this exists to remove — with the rejection proven atomic so a bad second field cannot leave half a configuration behind. The defaults are checked to be a magnitude that can really be reached, since a number too large to hit is a field rather than a budget |
|
|
253
|
-
| `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other |
|
|
253
|
+
| `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other. One more thing the totals cannot settle: the row key has **two namespaces** — the originating-task id and the parent call id it falls back to — and nothing upstream promises they are disjoint, so the same literal can name one child's identity and another child's parent call. Accumulation therefore keys on the origin as well as the id; the delivered table still keys on the bare id, and a cross-namespace clash is merged into one row that says so, with the partial verdict forced, because a row count and a turn count can both reconcile while the attribution behind them is wrong. The table itself is likewise a **per-stream snapshot** handed to the terminal projector by value rather than left on the caller's context: the three terminal projectors are public, so a host may drive one run through the stream and project another's terminal directly on the same context, and a table left behind would be attributed to whoever projects next — silently called complete whenever that run's own totals happen to match. Without a snapshot, both table keys are simply absent |
|
|
254
254
|
| `scripts/run-result-text-backfill-test.mjs` | What happens when the terminal frame's answer text and the text already on screen do not match. A turn's answer normally streams in and the terminal frame carries the same words again, so the two agree — but when the connection drops mid-answer and the reconnect brings the finished version, "this turn already produced assistant text" is true, the terminal fallback is skipped entirely, and the screen stays permanently short of whatever arrived while the stream was down, with nothing to say so. Four cases are pinned. Nothing on screen yet: render the terminal text whole, byte for byte the previous behaviour. On-screen text is a **prefix** of the terminal text: emit only the missing tail, and the guard measures the deciding quantity — the total bytes that reached the screen must equal the terminal text, which fails both for a missing tail and for a re-render that would print the first half twice; when the two are already equal, nothing is emitted at all. Terminal text is a prefix of what is on screen (an engine-side trim): touch nothing, since there is nothing missing and overwriting with the shorter version would erase what the reader already saw. Neither is a prefix of the other: emit **nothing** and raise a fact instead — which version counts is the engine's to say, and appending the terminal version after the streamed one composes a passage nobody ever wrote. That fact carries lengths rather than text, so a renderer is not handed a third version to choose from, and its declared duty is to *reword* the transcript line, never to render more. A cross-segment case proves the comparison reads the whole committed answer rather than the last segment, and the whole thing is driven through the real two-stage path rather than hand-built messages |
|
|
255
255
|
| `scripts/run-engine-vocab-floor-test.mjs` | Engine-mirrored vocabularies (structured card whitelist, self-reported tool face, control verbs, recogniser sets) against the *installed* `@sema-agent/core` |
|
|
256
256
|
| `scripts/run-limits-env-failloud-test.mjs` | `SEMA_HEADLESS_*` env-lane limits reject invalid values as loudly as the flag lane (no silent "no budget" runs) |
|
|
@@ -297,9 +297,10 @@ public-surface guard checks that last one).
|
|
|
297
297
|
| `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
|
|
298
298
|
| `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
|
|
299
299
|
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift |
|
|
300
|
-
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end |
|
|
300
|
+
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
|
|
301
301
|
| `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
|
|
302
302
|
| `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
|
|
303
|
+
| `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the 74 suites above actually turn red when the material they check really breaks — a census had found 16 of them clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. Seven guards of the same shape and 51 behaviour/projection suites are catalogued rather than rehearsed this round — see `docs/GATE-NEGATIVE-CONTROLS.md` for the full table, the reasons, and a one-minute manual replay recipe for each blind one. The suite cross-checks its own case count against that document's row counts in both directions, so a case quietly dropped from the array without the document following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
|
|
303
304
|
|
|
304
305
|
Each suite carries a floor that only moves up — a refactor that stops executing a group of
|
|
305
306
|
assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
|
|
@@ -238,13 +238,15 @@ export interface RunCostFacts {
|
|
|
238
238
|
* 🔴 `stats` 不是可读对象(409 拒绝信封 / park 体 / `failed` 事件帧)⇒ 返 `undefined` =
|
|
239
239
|
* **这条帧没有账**,调用方据此「不说话」(不发臂、不铸键),而不是发一条全缺席的空账。
|
|
240
240
|
*
|
|
241
|
-
* 🔴 **第二参 `
|
|
241
|
+
* 🔴 **第二参 `observed`(0.67.1 / B-091,additive)**:**这一条流**的流内观测快照。给了它,对账段上的
|
|
242
242
|
* {@link RunCostReconcile.usageLowerBound} 就是**取并后**的读数(见 {@link usageLowerBoundOf});
|
|
243
243
|
* 不给(旧签名)⇒ 只读 `stats` 那一半,既有端逐位不变。
|
|
244
244
|
* ⚠️ 包内的两个调用点(`costFactParts` 与 `runStream` 的 `run_cost_reconciled` 铸点)**都必须**
|
|
245
245
|
* 传它 —— 少传一处就是把本件修的那条不对称原样种回去。
|
|
246
|
+
* 🔴 0.67.2 / 车 I 件 I-1:它是**按次调用的入参**,不再是 `EmitContext` 上的一格 —— 共享 ctx 被复用
|
|
247
|
+
* 给多条流时,那一格会把别的流的缺口串进这一条(见 {@link usageLowerBoundOf} 的第三段 🔴)。
|
|
246
248
|
*/
|
|
247
|
-
export declare function readRunCostFacts(stats: TaskStats | undefined,
|
|
249
|
+
export declare function readRunCostFacts(stats: TaskStats | undefined, observed?: {
|
|
248
250
|
readonly usageMissingObserved?: boolean;
|
|
249
251
|
}): RunCostFacts | undefined;
|
|
250
252
|
/**
|
|
@@ -267,9 +269,24 @@ export interface SemaSubagentUsageRow {
|
|
|
267
269
|
readonly cacheReadTokens?: number;
|
|
268
270
|
/** `true` ⇒ 这只子任务**至少有一轮**没报 usage,本行三个数是**下界**。never false。 */
|
|
269
271
|
readonly usageMissing?: true;
|
|
272
|
+
/**
|
|
273
|
+
* 0.67.2(车 I 件 I-2)—— `true` ⇒ **这一行是两个身份命名空间的同字面碰撞合并出来的**:
|
|
274
|
+
* 一半轮次的行键来自真身份 `sourceTaskId`、另一半来自回落的 `parentToolCallId`,而两者的**字面相等**。
|
|
275
|
+
* 本层没有任何读数能把它们拆回去(core 的合同不保证两个命名空间互斥),所以这一行的数字是**几只
|
|
276
|
+
* 子任务混在一起**的和。⇒ 这一位在场时**必然**伴随 `_sema_nested_usage_by_task_partial`,消费方
|
|
277
|
+
* 不许把本行当成某一只子代的账去渲。**never false**(缺席 = 这一行的每一轮都来自同一个命名空间)。
|
|
278
|
+
*/
|
|
279
|
+
readonly keyCollision?: true;
|
|
270
280
|
}
|
|
271
281
|
/** {@link SemaSubagentUsageRow} 的累加中间态(runStream 持有;`readonly` 在收口那一拍才加)。 */
|
|
272
282
|
export interface MutableSubagentUsageRow {
|
|
283
|
+
/**
|
|
284
|
+
* 0.67.2(车 I 件 I-2)—— **交付面的裸 id**(`sourceTaskId`,缺席时是回落的 `parentToolCallId`)。
|
|
285
|
+
* 🔴 累加表的**键**从 0.67.2 起带出身前缀(`s:` / `p:`,见 `runStream` 的 `rowKey`),因为两个身份
|
|
286
|
+
* 命名空间的字面可以相等、并在累加那一层就把两只子任务并掉(并掉之后出身位只剩一种,「出身混合」
|
|
287
|
+
* 那道闸读不出混合);交付面的键仍然必须是**裸 id**(端零改),所以原值存在行上,前缀不出本层。
|
|
288
|
+
*/
|
|
289
|
+
taskId: string;
|
|
273
290
|
turns: number;
|
|
274
291
|
inputTokens: number;
|
|
275
292
|
outputTokens: number;
|
|
@@ -317,14 +334,37 @@ export interface SemaTerminalModelUsage extends SemaModelUsage {
|
|
|
317
334
|
/** `done` → SDKResultSuccess (contract 02 §2.10 / 08 CS-10). */
|
|
318
335
|
export declare function doneToSdkResult(ev: Extract<AgentEvent, {
|
|
319
336
|
type: 'done';
|
|
320
|
-
}>, ctx: EmitContext
|
|
337
|
+
}>, ctx: EmitContext,
|
|
338
|
+
/**
|
|
339
|
+
* 0.67.2:**这一条流**的收口快照 —— usage 缺口观测(见 {@link usageLowerBoundOf})与 per-subagent
|
|
340
|
+
* 累加表(见 {@link nestedUsageByTaskParts})。两格都是 per-stream 的事实,缺席 ⇒ 这次投影没有流内面
|
|
341
|
+
* (只读 `stats` 那一半、两个分表键都不铸),**绝不**回头去读调用方对象上可能残留的上一条流。
|
|
342
|
+
*/
|
|
343
|
+
observed?: {
|
|
344
|
+
readonly usageMissingObserved?: boolean;
|
|
345
|
+
readonly nestedUsageByTask?: ReadonlyMap<string, MutableSubagentUsageRow>;
|
|
346
|
+
}): SDKMessage;
|
|
321
347
|
/** `failed` → SDKResultError (contract 02 §2.11 / 08 CS-11). */
|
|
322
348
|
export declare function failedToSdkResult(ev: Extract<AgentEvent, {
|
|
323
349
|
type: 'failed';
|
|
324
|
-
}>, ctx: EmitContext
|
|
325
|
-
/**
|
|
350
|
+
}>, ctx: EmitContext,
|
|
351
|
+
/** 0.67.2:同 {@link doneToSdkResult} 的第三参。 */
|
|
352
|
+
observed?: {
|
|
353
|
+
readonly usageMissingObserved?: boolean;
|
|
354
|
+
readonly nestedUsageByTask?: ReadonlyMap<string, MutableSubagentUsageRow>;
|
|
355
|
+
}): SDKMessage;
|
|
356
|
+
/**
|
|
357
|
+
* Dispatch a terminal AgentEvent to its SDKResult arm.
|
|
358
|
+
*
|
|
359
|
+
* 🔴 第三参(0.67.2 / 车 I 件 I-1):**这一条流**的 usage 观测快照。`runStream` 在终帧那一拍**按值**
|
|
360
|
+
* 交给它与 `run_cost_reconciled` 铸臂 —— 两个投影口读的是**同一份本流快照**,而不是一个可能被别的流
|
|
361
|
+
* 写过的共享位。缺席(端直调)⇒ 与 0.67.1 的旧签名逐位同行为。
|
|
362
|
+
*/
|
|
326
363
|
export declare function terminalToSdkResult(ev: Extract<AgentEvent, {
|
|
327
364
|
type: 'done';
|
|
328
365
|
} | {
|
|
329
366
|
type: 'failed';
|
|
330
|
-
}>, ctx: EmitContext
|
|
367
|
+
}>, ctx: EmitContext, observed?: {
|
|
368
|
+
readonly usageMissingObserved?: boolean;
|
|
369
|
+
readonly nestedUsageByTask?: ReadonlyMap<string, MutableSubagentUsageRow>;
|
|
370
|
+
}): SDKMessage;
|
|
@@ -185,21 +185,26 @@ function finiteOrAbsent(v) {
|
|
|
185
185
|
*
|
|
186
186
|
* 🔴 它为什么必须是一只函数、而不是两处各写一遍的表达式:下界位有**两个来源**——
|
|
187
187
|
* · `stats.usageMissing`:只在带得出 `TaskResult` 的终帧上有;
|
|
188
|
-
* ·
|
|
188
|
+
* · **本条流**的流内观测(`failed` 事件帧 / 409 拒绝信封 / park 体**没有 stats**,
|
|
189
189
|
* 只读 stats 的话一条已经观测到缺口的 run 会在终帧上被读成「每一轮都报了 usage」)。
|
|
190
190
|
* 修前取并只包在**终帧**那一面({@link costFactParts}),而 chrome 对账臂的铸点直接展开
|
|
191
191
|
* `readRunCostFacts(stats).reconcile` ⇒ 「流内观测到缺口、终局 stats 对此缄默」那一形上两面各说
|
|
192
192
|
* 各的(终帧铸了 `_sema_usage_lower_bound`、同一拍的臂上没有 `usageLowerBound`)——
|
|
193
193
|
* [paired-mechanisms-must-share-premise] 的教科书形。⇒ 取并**下沉到这里**,两面共用。
|
|
194
194
|
*
|
|
195
|
+
* 🔴 **第二参是「这一条流」的观测快照,不是一个跨流留存的状态**(0.67.2 / 车 I 件 I-1):它此前住在
|
|
196
|
+
* `EmitContext` 上、只写 `true` 永不清,而 ctx 是调用方的对象、可以复用给多条流 ⇒ 上一条流的缺口
|
|
197
|
+
* 会把下一条账数得全的流标成下界(顺序复用),或两条并发流互相串。现在由 `runStream` 按流持有、
|
|
198
|
+
* 终局**按值**交给两个投影口;端自建管线同样是**按次调用**传入。
|
|
199
|
+
*
|
|
195
200
|
* 🔴 **严格 true 才认**(core 契约:`true` 或缺席,恒不写 `false`/`null`);`stats` 不是可读对象时
|
|
196
201
|
* 它那一半读作「没说」,而流内观测那一半**照旧成立**。
|
|
197
202
|
*/
|
|
198
|
-
function usageLowerBoundOf(stats,
|
|
203
|
+
function usageLowerBoundOf(stats, observed) {
|
|
199
204
|
const statsSaid = stats !== null && typeof stats === 'object' && !Array.isArray(stats)
|
|
200
205
|
? stats.usageMissing === true
|
|
201
206
|
: false;
|
|
202
|
-
return statsSaid ||
|
|
207
|
+
return statsSaid || observed?.usageMissingObserved === true;
|
|
203
208
|
}
|
|
204
209
|
/**
|
|
205
210
|
* D-3 / B-068 · L-198 —— 终局成本事实的**唯一读器**(终帧超集键与 chrome 对账臂共用)。
|
|
@@ -207,13 +212,15 @@ function usageLowerBoundOf(stats, ctx) {
|
|
|
207
212
|
* 🔴 `stats` 不是可读对象(409 拒绝信封 / park 体 / `failed` 事件帧)⇒ 返 `undefined` =
|
|
208
213
|
* **这条帧没有账**,调用方据此「不说话」(不发臂、不铸键),而不是发一条全缺席的空账。
|
|
209
214
|
*
|
|
210
|
-
* 🔴 **第二参 `
|
|
215
|
+
* 🔴 **第二参 `observed`(0.67.1 / B-091,additive)**:**这一条流**的流内观测快照。给了它,对账段上的
|
|
211
216
|
* {@link RunCostReconcile.usageLowerBound} 就是**取并后**的读数(见 {@link usageLowerBoundOf});
|
|
212
217
|
* 不给(旧签名)⇒ 只读 `stats` 那一半,既有端逐位不变。
|
|
213
218
|
* ⚠️ 包内的两个调用点(`costFactParts` 与 `runStream` 的 `run_cost_reconciled` 铸点)**都必须**
|
|
214
219
|
* 传它 —— 少传一处就是把本件修的那条不对称原样种回去。
|
|
220
|
+
* 🔴 0.67.2 / 车 I 件 I-1:它是**按次调用的入参**,不再是 `EmitContext` 上的一格 —— 共享 ctx 被复用
|
|
221
|
+
* 给多条流时,那一格会把别的流的缺口串进这一条(见 {@link usageLowerBoundOf} 的第三段 🔴)。
|
|
215
222
|
*/
|
|
216
|
-
export function readRunCostFacts(stats,
|
|
223
|
+
export function readRunCostFacts(stats, observed) {
|
|
217
224
|
// 数组也不是「一份账」:`typeof [] === 'object'` 会把一条畸形载体放进来,然后它的每一格都读不出
|
|
218
225
|
// ⇒ 发出一条「own 没定价」的臂,而真相是**这条帧根本没有账**(两句话又折成一句)。
|
|
219
226
|
if (stats === null || typeof stats !== 'object' || Array.isArray(stats))
|
|
@@ -260,7 +267,7 @@ export function readRunCostFacts(stats, ctx) {
|
|
|
260
267
|
// 0.67.0(core 7.14.0):usage 下界位。**严格 true 才铸**(core 契约:`true` 或缺席,恒不写
|
|
261
268
|
// `false`/`null`;认宽了就会把一个 falsy 值渲成「数得不全」)。它与成本三段正交,所以读在这里、
|
|
262
269
|
// 与三段同一只读器出 —— 两面(终帧超集键 / chrome 对账臂)因此永远不会各算各的。
|
|
263
|
-
const usageLowerBound = usageLowerBoundOf(stats,
|
|
270
|
+
const usageLowerBound = usageLowerBoundOf(stats, observed);
|
|
264
271
|
const reconcile = {
|
|
265
272
|
...(ownMicroUsd !== undefined ? { ownMicroUsd } : { costAbsent: true }),
|
|
266
273
|
...(usageLowerBound ? { usageLowerBound: true } : {}),
|
|
@@ -283,8 +290,8 @@ export function readRunCostFacts(stats, ctx) {
|
|
|
283
290
|
* (超集键纪律:CC 形上已有的位不许塞我们自己的含义)。「fully-reconciled spend」由消费方按
|
|
284
291
|
* 这两个超集键自己加 —— 包给的是**可对账的事实**,不是一个改了口径的数。
|
|
285
292
|
*/
|
|
286
|
-
function costFactParts(stats,
|
|
287
|
-
const facts = readRunCostFacts(stats,
|
|
293
|
+
function costFactParts(stats, observed) {
|
|
294
|
+
const facts = readRunCostFacts(stats, observed);
|
|
288
295
|
// 🔴 **下界位与成本三段分开算**(异源对抗复审 [medium] 实抓):`stats` 读不出(`failed` 事件帧 /
|
|
289
296
|
// 409 拒绝信封 / park 体)时**成本**那三段确实没有账、一条都不该说;但「这条流观测到过一轮
|
|
290
297
|
// 没有 usage」这件事**照旧成立**,而那种终帧的 `usage` 恰恰是 `flattenUsage(undefined)` 的
|
|
@@ -293,7 +300,7 @@ function costFactParts(stats, ctx) {
|
|
|
293
300
|
// 🔴 0.67.1 / B-091:取并本身已经**下沉**到 {@link usageLowerBoundOf} —— 这里与读器内部、与
|
|
294
301
|
// chrome 对账臂读的是**同一只函数**,三面不会各算各的(`facts === undefined` 时读器整只不
|
|
295
302
|
// 返回,所以这一行必须自己再调一次那只判据,而不是回头读 `facts`)。
|
|
296
|
-
const lowerBound = usageLowerBoundOf(stats,
|
|
303
|
+
const lowerBound = usageLowerBoundOf(stats, observed);
|
|
297
304
|
const lowerBoundPart = lowerBound ? { _sema_usage_lower_bound: true } : {};
|
|
298
305
|
if (facts === undefined)
|
|
299
306
|
return lowerBoundPart;
|
|
@@ -332,31 +339,74 @@ function costFactParts(stats, ctx) {
|
|
|
332
339
|
* 本判据只会**多**铸 partial(把一张其实完整的表说成不完整),**永远不会**把一张残表说成完整。
|
|
333
340
|
* 🔴 **两个键不互证、也不相加**:`_sema_nested_usage`(合计,引擎报的)与本表(流内看见的)是
|
|
334
341
|
* 两份独立的账;`partial` 在场时两者**本来就该不等**,消费方不许拿其中一份去「修正」另一份。
|
|
342
|
+
* 🔴 **0.67.2(车 I 件 I-2)身份键碰撞**:累加表的键按出身隔离(`s:` / `p:`),交付面按裸 id 归并 ——
|
|
343
|
+
* 跨命名空间的同字面合并成一行、立 {@link SemaSubagentUsageRow.keyCollision},并让 `partial` 恒立。
|
|
344
|
+
* 🔴 **0.67.2(车 I 件 I-2b)这张表是「本流快照」,不再从 `EmitContext` 上读**(与件 I-1 逐字同形的
|
|
345
|
+
* 同形存量,轮二异源复审实抓):流结束后它此前**留在调用方对象上**,而本文件这三只终帧投影器是
|
|
346
|
+
* **公面导出** —— 端「A 走 runStream、B 直调终帧投影」共用一个 ctx 时,B 的终帧会带出 **A 的**分表;
|
|
347
|
+
* B 的 `nested` 计数若恰好与那张表对得上(`tasks`/`turns` 相等),`partial` 还不铸 = 一张属于别人的
|
|
348
|
+
* 表被标成「可证完整」。⇒ 表随终帧那一拍**按值**传进来;没传 ⇒ 两个键都不铸(诚实缺席:这次投影
|
|
349
|
+
* 一条子流都没看见),**绝不**回头读残留。
|
|
335
350
|
*/
|
|
336
351
|
function nestedUsageByTaskParts(stats, rollup) {
|
|
337
352
|
// 🔴 一行都没有 ⇒ **什么都不说**:空表会被读成「这条 run 一个子代都没委派」,而真相可能是
|
|
338
353
|
// 「委派了,但这条流没看见任何一轮」(重连车道)。两句话不许折成一句。
|
|
339
354
|
if (rollup === undefined || rollup.size === 0)
|
|
340
355
|
return {};
|
|
341
|
-
const rows = {};
|
|
342
356
|
let turnsSeen = 0;
|
|
343
357
|
// 🔴 0.67.1:行键的**出身**统计(见 {@link MutableSubagentUsageRow.keyFromParentFallback})。
|
|
344
|
-
// 出身按行是均匀的 ——
|
|
345
|
-
// (
|
|
358
|
+
// 出身按行是均匀的 —— 0.67.2 起这件事由**键空间**保证:累加表的键带 `s:` / `p:` 前缀,所以
|
|
359
|
+
// 一行的每一轮必然来自同一个命名空间(修前靠的是「`taskId = sourceTaskId ?? parent`」那条推理,
|
|
360
|
+
// 而那条推理在两个空间撞字面时不成立)。
|
|
346
361
|
let fallbackKeyedRows = 0;
|
|
347
362
|
let sourceKeyedRows = 0;
|
|
348
|
-
|
|
363
|
+
// ── 🔴 0.67.2(车 I 件 I-2)—— 交付面按**裸 id** 归并,跨命名空间的同字面 = **碰撞** ──────────
|
|
364
|
+
// 交付面的键必须是裸 id(端零改),而累加表的键带出身前缀 ⇒ 两个命名空间里字面相同的两行会在这里
|
|
365
|
+
// 落到同一个交付键上。本层**没有任何读数**能把它们拆回去(core 的合同不保证两个命名空间互斥),
|
|
366
|
+
// 于是:合并数字(不偷偷丢一边)+ 立 `keyCollision`(如实说这一行是混的)+ `partial` 恒立
|
|
367
|
+
// (证不出逐任务归属就不许说完整)。
|
|
368
|
+
// 🔴 为什么不「两行分列」:交付形是 `Record<taskId, row>`,同一个裸 id 不可能占两格 —— 要分列就得
|
|
369
|
+
// 改交付键的形(把出身前缀推到 wire 上),那是三端 BREAKING,而且把本层的内部编码变成端要解的
|
|
370
|
+
// 身份语义。合并 + 显式标记是**同一批事实**下唯一不撒谎的交付形。
|
|
371
|
+
const merged = new Map();
|
|
372
|
+
for (const r of rollup.values()) {
|
|
349
373
|
turnsSeen += r.turns;
|
|
350
374
|
if (r.keyFromParentFallback === true)
|
|
351
375
|
fallbackKeyedRows += 1;
|
|
352
376
|
else
|
|
353
377
|
sourceKeyedRows += 1;
|
|
378
|
+
const prev = merged.get(r.taskId);
|
|
379
|
+
if (prev === undefined) {
|
|
380
|
+
merged.set(r.taskId, {
|
|
381
|
+
turns: r.turns,
|
|
382
|
+
inputTokens: r.inputTokens,
|
|
383
|
+
outputTokens: r.outputTokens,
|
|
384
|
+
...(r.cacheReadTokens !== undefined ? { cacheReadTokens: r.cacheReadTokens } : {}),
|
|
385
|
+
...(r.usageMissing === true ? { usageMissing: true } : {}),
|
|
386
|
+
});
|
|
387
|
+
continue;
|
|
388
|
+
}
|
|
389
|
+
prev.turns += r.turns;
|
|
390
|
+
prev.inputTokens += r.inputTokens;
|
|
391
|
+
prev.outputTokens += r.outputTokens;
|
|
392
|
+
// `cacheReadTokens` 的缺席语义在合并处也守住:两边都没报过 ⇒ 键仍不铸(0 会被读成「零命中」)。
|
|
393
|
+
if (r.cacheReadTokens !== undefined)
|
|
394
|
+
prev.cacheReadTokens = (prev.cacheReadTokens ?? 0) + r.cacheReadTokens;
|
|
395
|
+
if (r.usageMissing === true)
|
|
396
|
+
prev.usageMissing = true;
|
|
397
|
+
prev.keyCollision = true;
|
|
398
|
+
}
|
|
399
|
+
const keyCollisions = [...merged.values()].filter((v) => v.keyCollision === true).length;
|
|
400
|
+
const rows = {};
|
|
401
|
+
for (const [taskId, v] of merged) {
|
|
402
|
+
// 🔴 落键仍走 `putOwn`(`__proto__` 陷阱):合并臂不是绕过那条纪律的第二条路(见 `putOwn` 头注)。
|
|
354
403
|
putOwn(rows, taskId, {
|
|
355
|
-
turns:
|
|
356
|
-
inputTokens:
|
|
357
|
-
outputTokens:
|
|
358
|
-
...(
|
|
359
|
-
...(
|
|
404
|
+
turns: v.turns,
|
|
405
|
+
inputTokens: v.inputTokens,
|
|
406
|
+
outputTokens: v.outputTokens,
|
|
407
|
+
...(v.cacheReadTokens !== undefined ? { cacheReadTokens: v.cacheReadTokens } : {}),
|
|
408
|
+
...(v.usageMissing === true ? { usageMissing: true } : {}),
|
|
409
|
+
...(v.keyCollision === true ? { keyCollision: true } : {}),
|
|
360
410
|
});
|
|
361
411
|
}
|
|
362
412
|
const nested = stats !== null && typeof stats === 'object' && !Array.isArray(stats)
|
|
@@ -369,7 +419,11 @@ function nestedUsageByTaskParts(stats, rollup) {
|
|
|
369
419
|
// 出身一致时既有两条对账才是充分的:全真身份 ⇒ 一行一任务;全回落 ⇒ 多任务共父会让行数 <
|
|
370
420
|
// `nested.tasks`,那一格自己会翻。失效方向仍是安全的那一侧(只会**多**铸 partial)。
|
|
371
421
|
const mixedKeyOrigin = fallbackKeyedRows > 0 && sourceKeyedRows > 0;
|
|
372
|
-
|
|
422
|
+
// 🔴 0.67.2:碰撞也单独挡一格。按构造 `keyCollisions > 0 ⇒ mixedKeyOrigin`(同一个命名空间里的键
|
|
423
|
+
// 本来就唯一,撞字面只能跨空间发生),所以这一格今天是**冗余的第二道**;留着是因为两条判据问的
|
|
424
|
+
// 不是同一件事(「这条流上有两种出身」vs「这一行里混了两种出身」),而上游哪天再加一个身份来源
|
|
425
|
+
// 时,前者的判法要改、后者不用。
|
|
426
|
+
const complete = !mixedKeyOrigin && keyCollisions === 0 && authTurns !== undefined && authTasks !== undefined &&
|
|
373
427
|
authTurns === turnsSeen && authTasks === Object.keys(rows).length;
|
|
374
428
|
return {
|
|
375
429
|
_sema_nested_usage_by_task: rows,
|
|
@@ -628,10 +682,10 @@ function errorResult(ctx, parts) {
|
|
|
628
682
|
// D-1 / L-192①:两句话不再折成一句 —— 清单 + 「有没有这本账」的判别位,见 permissionDenialParts。
|
|
629
683
|
...permissionDenialParts(parts.stats),
|
|
630
684
|
// D-3 / B-068:失败/到限/park 的 run 一样花过钱,账不因结局不好就不报。
|
|
631
|
-
...costFactParts(parts.stats,
|
|
685
|
+
...costFactParts(parts.stats, parts.observed),
|
|
632
686
|
// L-228(0.67.0):流内 per-subagent 分表的收口快照(判据本体在 `nestedUsageByTaskParts`)。
|
|
633
687
|
// 🔴 与成本三段同理 —— 失败的 run 一样委派过,账不因结局不好就不报。
|
|
634
|
-
...nestedUsageByTaskParts(parts.stats,
|
|
688
|
+
...nestedUsageByTaskParts(parts.stats, parts.observed?.nestedUsageByTask),
|
|
635
689
|
errors: [...parts.errors],
|
|
636
690
|
...(parts.errorCode !== undefined && parts.errorCode.length > 0 ? { errorCode: parts.errorCode } : {}),
|
|
637
691
|
...(parts.degraded !== undefined ? { degraded: parts.degraded } : {}),
|
|
@@ -641,7 +695,13 @@ function errorResult(ctx, parts) {
|
|
|
641
695
|
});
|
|
642
696
|
}
|
|
643
697
|
/** `done` → SDKResultSuccess (contract 02 §2.10 / 08 CS-10). */
|
|
644
|
-
export function doneToSdkResult(ev, ctx
|
|
698
|
+
export function doneToSdkResult(ev, ctx,
|
|
699
|
+
/**
|
|
700
|
+
* 0.67.2:**这一条流**的收口快照 —— usage 缺口观测(见 {@link usageLowerBoundOf})与 per-subagent
|
|
701
|
+
* 累加表(见 {@link nestedUsageByTaskParts})。两格都是 per-stream 的事实,缺席 ⇒ 这次投影没有流内面
|
|
702
|
+
* (只读 `stats` 那一半、两个分表键都不铸),**绝不**回头去读调用方对象上可能残留的上一条流。
|
|
703
|
+
*/
|
|
704
|
+
observed) {
|
|
645
705
|
// sdk 4.1.0([2395]E 调和):`done.result` 是判别联合 `TaskResult | ActiveRunConflictDoneResult`。
|
|
646
706
|
// 🔄 全窗复审 ADAPTER-1/-2 收口(2026-08-03):首版接线用 `'stats' in r` 铸 tr 并让 result/model/
|
|
647
707
|
// degraded 全走 tr?.——把「stats 在不在」错当成了这些键的门控(anchor-on-the-deciding-quantity
|
|
@@ -693,7 +753,7 @@ export function doneToSdkResult(ev, ctx) {
|
|
|
693
753
|
// failed + limits.max_{tokens,walltime}_exceeded 两族共用)。
|
|
694
754
|
const degraded = degradedOf(r);
|
|
695
755
|
/** 四个 error 臂共享的固定位(信封其余 13 位见 `errorResult`)。 */
|
|
696
|
-
const errorBase = { durationMs, stats, model: r.model, errorCode, degraded };
|
|
756
|
+
const errorBase = { durationMs, stats, model: r.model, errorCode, degraded, observed };
|
|
697
757
|
if (terminal?.kind === 'failed') {
|
|
698
758
|
// [909]B1 — failed 臂 subtype 语义化(见文件头对表);core 5.8.0([2489])起到限码全部改名,
|
|
699
759
|
// 映射本身已收进单点 `subtypeForErrorCode`(退役批后只认新码)。
|
|
@@ -824,10 +884,10 @@ export function doneToSdkResult(ev, ctx) {
|
|
|
824
884
|
// D-1 / L-192①:同形第二处 —— 与错误信封共用**同一个** mint 点(修前两处各一个字面量 [])。
|
|
825
885
|
...permissionDenialParts(stats),
|
|
826
886
|
// D-3 / B-068:成本明细与子代那本账(micro-USD 原值);`total_cost_usd` 语义一字不动。
|
|
827
|
-
...costFactParts(stats,
|
|
887
|
+
...costFactParts(stats, observed),
|
|
828
888
|
// L-228(0.67.0):流内 per-subagent 分表的收口快照;与 `_sema_nested_usage`(引擎报的合计)
|
|
829
889
|
// 是**两份独立的账**,不相加、不互证(见 `nestedUsageByTaskParts` 顶注)。
|
|
830
|
-
...nestedUsageByTaskParts(stats,
|
|
890
|
+
...nestedUsageByTaskParts(stats, observed?.nestedUsageByTask),
|
|
831
891
|
// MF-25 — the effective served model id (`done.result.model`, e.g. "deepseek-v4-pro"). The CC
|
|
832
892
|
// SDKResultSuccess schema has no `model` field, so this rides as an additive seam field a cost/overview
|
|
833
893
|
// consumer reads (it is ALSO surfaced as the `modelUsage` key). Omitted when the wire didn't carry it.
|
|
@@ -835,7 +895,9 @@ export function doneToSdkResult(ev, ctx) {
|
|
|
835
895
|
});
|
|
836
896
|
}
|
|
837
897
|
/** `failed` → SDKResultError (contract 02 §2.11 / 08 CS-11). */
|
|
838
|
-
export function failedToSdkResult(ev, ctx
|
|
898
|
+
export function failedToSdkResult(ev, ctx,
|
|
899
|
+
/** 0.67.2:同 {@link doneToSdkResult} 的第三参。 */
|
|
900
|
+
observed) {
|
|
839
901
|
// Flatten the 4 CC error subtypes onto the single neutral errorCode.
|
|
840
902
|
// error_max_budget_usd is a non-error "budget exceeded" notice, not a crash
|
|
841
903
|
// (contract 02 §2.11) — the renderer branches on subtype.
|
|
@@ -851,6 +913,7 @@ export function failedToSdkResult(ev, ctx) {
|
|
|
851
913
|
// `failed` 事件帧本体不带 stats/model/degraded,所以这三个位如实缺席 —— 不是「这里少算了」。
|
|
852
914
|
return errorResult(ctx, {
|
|
853
915
|
subtype,
|
|
916
|
+
observed,
|
|
854
917
|
durationMs: elapsedMs(ctx),
|
|
855
918
|
// [909]B1 — errorCode 透传(additive seam 字段;done 臂同款):`failed` 事件的 cancelled/
|
|
856
919
|
// limits.* 等引擎值原样给集成面。
|
|
@@ -866,7 +929,13 @@ export function failedToSdkResult(ev, ctx) {
|
|
|
866
929
|
],
|
|
867
930
|
});
|
|
868
931
|
}
|
|
869
|
-
/**
|
|
870
|
-
|
|
871
|
-
|
|
932
|
+
/**
|
|
933
|
+
* Dispatch a terminal AgentEvent to its SDKResult arm.
|
|
934
|
+
*
|
|
935
|
+
* 🔴 第三参(0.67.2 / 车 I 件 I-1):**这一条流**的 usage 观测快照。`runStream` 在终帧那一拍**按值**
|
|
936
|
+
* 交给它与 `run_cost_reconciled` 铸臂 —— 两个投影口读的是**同一份本流快照**,而不是一个可能被别的流
|
|
937
|
+
* 写过的共享位。缺席(端直调)⇒ 与 0.67.1 的旧签名逐位同行为。
|
|
938
|
+
*/
|
|
939
|
+
export function terminalToSdkResult(ev, ctx, observed) {
|
|
940
|
+
return ev.type === 'done' ? doneToSdkResult(ev, ctx, observed) : failedToSdkResult(ev, ctx, observed);
|
|
872
941
|
}
|
|
@@ -319,6 +319,17 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
319
319
|
// 开流时挂等于让后开的那条把先开的那条的表顶掉,先开的终帧于是报出别人的账。
|
|
320
320
|
// 挂在终帧那一拍 + 与 `terminalToSdkResult(...)` 在**同一个同步步**里,那个窗按构造不存在。
|
|
321
321
|
const nestedUsageByTask = new Map();
|
|
322
|
+
// ── 🔴 0.67.2 / 车 I 件 I-1(异源对抗复审 [medium])—— usage 缺口观测位是 **per-stream** ──
|
|
323
|
+
// 修前它写在 `ctx.usageMissingObserved` 上:`EmitContext` 是**调用方的对象**、可以被复用给多条流
|
|
324
|
+
// (`startedAtMs` 本来就是这么用的;`run-subagent-usage-projection-test.mjs` 的 J3 段明确支持这一形),
|
|
325
|
+
// 而那一位**只置 true、永不清** ⇒
|
|
326
|
+
// · 顺序复用:上一条流里的一轮缺口,让**下一条 usage 完整的流**的终帧铸出 `_sema_usage_lower_bound`、
|
|
327
|
+
// 同一拍的 `run_cost_reconciled` 也被标成下界 —— 宿主据此把一条账数得全的 run 渲成「≥」并持久化;
|
|
328
|
+
// · 并发复用:两条流互相串这一位。
|
|
329
|
+
// 「别人那条流有缺口」不是「这条流的数字是下界」的证据。⇒ 观测位落在**本函数的局部量**上,终局
|
|
330
|
+
// 把**本流快照**同时交给两个投影口(终帧与对账臂),两面读同一份、且谁都读不到别人那份。
|
|
331
|
+
// 🔴 **不是**靠「终帧那一拍把 ctx 上那一位清掉」修的:并发的两条流会互相覆盖那次清除。
|
|
332
|
+
let usageMissingObserved = false;
|
|
322
333
|
for await (const ev of events) {
|
|
323
334
|
// event-id idempotency — drop a re-seen durable seq (contract 02 §1.1).
|
|
324
335
|
const seq = eventSeq(ev);
|
|
@@ -395,7 +406,7 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
395
406
|
// (而 `usage` 那几格恰好是 `flattenUsage(undefined)` 的全零)——「不知道」渲成了精确零。
|
|
396
407
|
// ⇒ 流内观测到就记下来,终帧那一拍与 stats 的读数**取并**(见 costFactParts)。
|
|
397
408
|
if (usageMissing)
|
|
398
|
-
|
|
409
|
+
usageMissingObserved = true;
|
|
399
410
|
const stopReasonRaw = ev.stopReason;
|
|
400
411
|
const stopWord = typeof stopReasonRaw === 'string' && stopReasonRaw.length > 0 ? stopReasonRaw : undefined;
|
|
401
412
|
const usage = turnEndUsage(ev);
|
|
@@ -476,6 +487,19 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
476
487
|
// ⚠️ 回落**有损**:同父调用多子任务会并成一行 —— 那时终帧那张表的 `partial` 判别位会因
|
|
477
488
|
// 行数对不上 `nested.tasks` 而立起来(诚实缺席优先于假装分得开)。
|
|
478
489
|
const taskId = sourceTaskId ?? parent;
|
|
490
|
+
// ── 🔴 0.67.2 / 车 I 件 I-2(异源对抗复审 [medium])—— **累加表的键按出身隔离** ──
|
|
491
|
+
// 病:`taskId` 有**两个命名空间**(真身份 `sourceTaskId` / 回落 `parentToolCallId`),而 core 的
|
|
492
|
+
// 合同**没有**保证两者互斥 —— 一只子任务的 id 与另一只子任务的父调用 id 完全可以撞字面。修前
|
|
493
|
+
// 直接拿裸 `taskId` 当累加键,于是撞字面的两行**在累加那一层就已经并掉**,而出身位
|
|
494
|
+
// (`keyFromParentFallback`)是**按行**记的 ⇒ 并掉之后那一行只剩一种出身,「出身混合」这道闸
|
|
495
|
+
// (`nestedUsageByTaskParts` 的 `mixedKeyOrigin`)当场读不出混合,于是被绕过。
|
|
496
|
+
// 三帧反例(异源复审逐字复现):`(sourceTaskId,parentToolCallId,inputTokens)` = ('p','a',10) / (缺席,'p',20) /
|
|
497
|
+
// (缺席,'a',30),终局 `nested={tasks:2,turns:3}` ⇒ 修前输出 `{p:30, a:30}`、行数与轮数两条对账
|
|
498
|
+
// **同时成立** ⇒ `partial` 不铸,而真相是父调用 `a` 那只花了 40、父调用 `p` 那只花了 20。
|
|
499
|
+
// ⇒ 累加键前缀化(`s:` = 真身份 / `p:` = 回落),**逐事件**把出身记进键本身;交付面的键仍是
|
|
500
|
+
// 裸 id(端零改),同字面的跨空间碰撞由 `nestedUsageByTaskParts` 判出来并如实标记。
|
|
501
|
+
// 🔴 前缀只活在**本层的累加表**里:它不是身份的一部分,wire 上自带 `s:`/`p:` 前缀的 id 因此
|
|
502
|
+
// 不会与别人串(两个空间的键各带自己的前缀,`s:` + `"p:x"` ≠ `p:` + `"s:x"`)。
|
|
479
503
|
// 🔴 0.67.1 / B-090:入表条件从「行键 **且** 父调用 id 都读得出」放宽到「**行键**读得出」。
|
|
480
504
|
// 修前那个 `&& parent !== undefined` 把「只带 `sourceTaskId`」的子代整条挡在表外 ——
|
|
481
505
|
// 而 `parent` 在这里的唯一用处是**铸 chrome 增量臂的车道证明**,不是行的身份。
|
|
@@ -509,13 +533,17 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
509
533
|
...(stopWord !== undefined ? { stopReason: stopWord } : {}),
|
|
510
534
|
});
|
|
511
535
|
}
|
|
512
|
-
|
|
536
|
+
// 累加键 = 出身前缀 + 裸 id(见上面那段 🔴)。`sourceTaskId` 缺席时 `taskId === parent`
|
|
537
|
+
// (它就是 `sourceTaskId ?? parent`),所以这里不必再写一次回落判据。
|
|
538
|
+
const rowKey = sourceTaskId !== undefined ? `s:${sourceTaskId}` : `p:${taskId}`;
|
|
539
|
+
const row = nestedUsageByTask.get(rowKey) ?? { taskId, turns: 0, inputTokens: 0, outputTokens: 0 };
|
|
513
540
|
// 🔴 0.67.1(异源复审 [medium] 实抓):记下**这一行的键是回落来的**(真身份缺席)。
|
|
514
541
|
// B-090 放宽入表条件之后行键可以有两种出身,而混合出身时「拆一行 + 并一行」的计数
|
|
515
542
|
// 误差方向相反、可以恰好抵消 ⇒ 终帧那张表的 `partial` 判据要看得见出身
|
|
516
543
|
// (理由与反例逐字在 `MutableSubagentUsageRow.keyFromParentFallback` 的头注里)。
|
|
517
|
-
//
|
|
518
|
-
//
|
|
544
|
+
// 🔴 0.67.2 / 件 I-2:出身按行均匀这件事现在由**键空间**保证(累加键带 `s:`/`p:` 前缀),
|
|
545
|
+
// 不再依赖「`taskId = sourceTaskId ?? parent` 所以真身份在场的轮不会落到回落行上」这条
|
|
546
|
+
// 推理 —— 那条推理在两个空间**撞字面**时不成立(见上面 `rowKey` 的头注)。
|
|
519
547
|
if (sourceTaskId === undefined)
|
|
520
548
|
row.keyFromParentFallback = true;
|
|
521
549
|
row.turns += 1;
|
|
@@ -531,7 +559,7 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
531
559
|
}
|
|
532
560
|
if (usageMissing)
|
|
533
561
|
row.usageMissing = true;
|
|
534
|
-
nestedUsageByTask.set(
|
|
562
|
+
nestedUsageByTask.set(rowKey, row);
|
|
535
563
|
}
|
|
536
564
|
}
|
|
537
565
|
const outputTokens = ev.usage?.outputTokens;
|
|
@@ -701,19 +729,28 @@ async function* runStreamInner(events, ctx, handle = {}) {
|
|
|
701
729
|
// fail-soft 同 plan_review_park:sink 抛错不影响终帧照常投影。
|
|
702
730
|
if (ev.type === 'done' && ctx.emitChrome) {
|
|
703
731
|
const doneStats = ev.result?.stats;
|
|
704
|
-
// 🔴 0.67.1 / B-091:**第二参必须传** —— 下界位的取并(`stats.usageMissing` ∪ 流内观测
|
|
705
|
-
//
|
|
706
|
-
//
|
|
707
|
-
|
|
732
|
+
// 🔴 0.67.1 / B-091:**第二参必须传** —— 下界位的取并(`stats.usageMissing` ∪ 流内观测)
|
|
733
|
+
// 已经下沉进读器;不传就是把「臂绕开取并」那条不对称原样种回去(终帧铸了
|
|
734
|
+
// `_sema_usage_lower_bound`、同一拍的臂上却没有 `usageLowerBound`)。
|
|
735
|
+
// 🔴 0.67.2 / 件 I-1:传的是**本流快照**(局部量按值包一层),不再是共享 ctx —— 与下面那行
|
|
736
|
+
// `terminalToSdkResult(..., observed)` 是**同一份**读数。
|
|
737
|
+
const costFacts = readRunCostFacts(doneStats, { usageMissingObserved });
|
|
708
738
|
if (costFacts !== undefined) {
|
|
709
739
|
emitChromeFireAndForget(ctx, { kind: 'run_cost_reconciled', laneProof: MAIN, ...costFacts.reconcile });
|
|
710
740
|
}
|
|
711
741
|
}
|
|
712
|
-
// L-228
|
|
713
|
-
//
|
|
714
|
-
//
|
|
715
|
-
ctx
|
|
716
|
-
|
|
742
|
+
// L-228 / 🔴 0.67.2 件 I-2b 订正:分表**不再挂到 `ctx` 上**,而是与观测位一起随第三参按值交给
|
|
743
|
+
// 终帧投影器。此前那句「挂表与铸终帧在同一个同步步里,复用 ctx 的并发流不会串账」只覆盖了
|
|
744
|
+
// **两边都经 runStream** 的路径 —— 而这三只终帧投影器是**公面导出**,端完全可以「A 走
|
|
745
|
+
// runStream、B 直调终帧投影」共用一个 ctx,那时 B 的终帧带出的是 A 的分表(B 的 `nested`
|
|
746
|
+
// 计数恰好对得上时连 `partial` 都不铸)。与件 I-1 逐字同形,所以同批一起摘掉。
|
|
747
|
+
// 🔴 0.67.2 / 件 I-1 订正:这里此前还写着「`usageMissingObserved` 是 per-ctx 的单调布尔,
|
|
748
|
+
// 复用 ctx 的两条流里只要有一条观测到缺口,两条都该按下界读 —— 取并是安全的那一侧」。
|
|
749
|
+
// **那句话是错的**:下界位问的是「**这一条 run** 的账数全了没有」,别的流的缺口对它一个
|
|
750
|
+
// 字节的证据都不是;按那句话办,一条账数得全的 run 会被渲成「≥」并被宿主持久化成不完整状态
|
|
751
|
+
// (失效方向在这里**不是**安全的那一侧,它是在断言一件没发生的事)。
|
|
752
|
+
// ⇒ 观测位改为 per-stream 局部量,终帧按值收(第三参),与上面对账臂读的是同一份。
|
|
753
|
+
yield terminalToSdkResult(ev, ctx, { usageMissingObserved, nestedUsageByTask });
|
|
717
754
|
return;
|
|
718
755
|
}
|
|
719
756
|
// Every other arm → typed three-state projection(REF-CC-058)。
|
package/dist/adapter/types.d.ts
CHANGED
|
@@ -36,7 +36,6 @@
|
|
|
36
36
|
import type { AgentEvent } from '@sema-agent/sdk';
|
|
37
37
|
import type { ModelUsage, SDKMessage } from '@sema-agent/agent-types';
|
|
38
38
|
import type { ChromeEvent } from '../seam.js';
|
|
39
|
-
import type { MutableSubagentUsageRow } from './downstream/terminalToSdkResult.js';
|
|
40
39
|
export type { ModelUsage, SDKMessage };
|
|
41
40
|
export type StampedAgentEvent = AgentEvent & {
|
|
42
41
|
id?: string;
|
|
@@ -101,32 +100,6 @@ export interface EmitContext {
|
|
|
101
100
|
* · sink 抛错**绝不影响流**,并且痕迹落回 console —— 让位的前提是它真接住了。
|
|
102
101
|
*/
|
|
103
102
|
onDroppedFrame?(info: DroppedFrameInfo): void;
|
|
104
|
-
/**
|
|
105
|
-
* L-228(0.67.0)—— **这条流上看见的 per-subagent turn 用量分表**,键 = 子任务 id
|
|
106
|
-
* (wire 缺席时回落 `parentToolCallId`)。终帧的两个超集键
|
|
107
|
-
* `_sema_nested_usage_by_task` / `_sema_nested_usage_by_task_partial` 由它铸出
|
|
108
|
-
* (唯一 mint 点 = `terminalToSdkResult.ts` 的 `nestedUsageByTaskParts`)。
|
|
109
|
-
*
|
|
110
|
-
* 🔴 **由 `runStream` 在流内写,宿主不要自己填** —— 与同接口的 `startedAtMs` 同一类
|
|
111
|
-
* (「请求是宿主构造的,这一格是驱动自己攒的」)。宿主塞一份进来 = 把一份**不是这条流看见的**
|
|
112
|
-
* 账当成这条流的,而下游那个 `partial` 判别位恰恰是靠「这条流看见了多少」才成立的。
|
|
113
|
-
* 🔴 **缺席 / 空表 ⇒ 终帧两个键都不铸**:空表会被读成「一个子代都没委派」,而真相可能是
|
|
114
|
-
* 「委派了但这条流没看见任何一轮」。
|
|
115
|
-
*/
|
|
116
|
-
nestedUsageByTask?: ReadonlyMap<string, MutableSubagentUsageRow>;
|
|
117
|
-
/**
|
|
118
|
-
* 0.67.0 —— **这条流上观测到过「某一轮没有 usage」**(core 的 `turn_end.usageMissing`)。
|
|
119
|
-
* 终帧的 `_sema_usage_lower_bound` 与 chrome 对账臂的 `usageLowerBound` 与 `stats.usageMissing`
|
|
120
|
-
* **取并**读它。
|
|
121
|
-
*
|
|
122
|
-
* 🔴 **为什么非有它不可**:`stats.usageMissing` 只在带得出 `TaskResult` 的终帧上有,而
|
|
123
|
-
* `failed` 事件帧 / 409 拒绝信封 / park 体**根本没有 stats** ⇒ 只读 stats 的话,一条**已经
|
|
124
|
-
* 观测到缺口**的 run 会在终帧上落成「判别位缺席」,而按新合同那读作「每一轮都报了 usage」——
|
|
125
|
-
* 偏偏那种终帧的 `usage` 是 `flattenUsage(undefined)` 的**全零**:「不知道」被渲成了精确零。
|
|
126
|
-
* 🔴 **由 `runStream` 在流内写,宿主不要自己填**(同 `nestedUsageByTask` / `startedAtMs`)。
|
|
127
|
-
* 🔴 **只置 `true`,从不置回 false**:一轮不知道,整条流的数字就是下界,后面的轮补不回来。
|
|
128
|
-
*/
|
|
129
|
-
usageMissingObserved?: true;
|
|
130
103
|
}
|
|
131
104
|
/** Stamp `uuid` + `session_id` onto a freshly-built arm body. */
|
|
132
105
|
export declare function stamp<T extends {
|
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
|
|
20
20
|
| 项 | 值 | 真源 |
|
|
21
21
|
|---|---|---|
|
|
22
|
-
| 本包 | `@sema-agent/client-core` **0.67.
|
|
22
|
+
| 本包 | `@sema-agent/client-core` **0.67.2**(本批发布版 = patch:内容批 a021307,异源对抗复审三条 [medium] + 轮二同形:流内观测位 per-stream / 身份键按出身隔离 + 行级 `keyCollision` / 分表随第三参按值传 / 负控备份独占创建,`EmitContext` 退役两格,零公面增删零 wire 键增删零 BREAKING,详见 §32f/§32g 0.67.2 订正段;`CHANGELOG.md` `## 0.67.2(2026-09-12)`;0.67.1 patch = 内容批 a5ea52c,test [7055] 三修 B-090/B-091/__proto__ + 出身混合 partial 恒立,详见 §32f/§32g 0.67.1 订正段;`CHANGELOG.md` `## 0.67.1(2026-09-12)`;0.67.0 minor = 内容批 df6b2dc,core 7.14.0→7.16.0 提货七件 F-1…F-7,🔴 七条 BREAKING(分类器卡面族 clean-cut、`unresolvable`→`ancestor_marked`、audience 换档、`classifierStatusOf` 第二参语义),详见 §32a–§32z;0.66.0 见 §31;`CHANGELOG.md` `## 0.67.0(2026-09-12)`;bump 与冻结账两阶段由发包批做) | `package.json` `version` |
|
|
23
23
|
| peer:wire 契约 | `@sema-agent/sdk` **>=8.4.0**(value-level,非 type-only;**0.60.0 抬版**,四条硬理由见 §24a 与 `scripts/run-sdk-floor-test.mjs` 的 `FLOOR` 注;上一次是 0.59.0 的 `>=8.3.0`)。🔴 支持窗同批收到 **engine ≥7.64.0**:sdk 8.4.0 与 7.63.0 及以前的 wire **不同窗** | `package.json` `peerDependencies` |
|
|
24
24
|
| peer:会话词汇表 | `@sema-agent/agent-types` **>=0.2.0**(type-only,零运行时) | 同上 |
|
|
25
25
|
| runtime dep | `diff` ^9.0.0(**唯一**一条;portability 门按**等值**钉死) | `package.json` `dependencies` |
|
|
@@ -6573,6 +6573,42 @@ prepare 期拒绝 / 合成 abort)。**数字位仍是必填、仍是数到的那
|
|
|
6573
6573
|
`usageLowerBound === true`;两面的**在场性逐位相等**。反向:同一条路上没观测到缺口 ⇒ 两面**都**键缺席
|
|
6574
6574
|
(never false)。子流那一轮报的缺口**同样算数**(下界是这条流的性质,不分车道)。
|
|
6575
6575
|
|
|
6576
|
+
#### 🔧 0.67.2 订正(车 I 件 I-1,异源对抗复审 [medium] 实抓):流内观测位是 **per-stream**
|
|
6577
|
+
|
|
6578
|
+
0.67.1 把取并下沉到唯一读器,但**那一半来源自己住错了地方** —— 它写在
|
|
6579
|
+
`ctx.usageMissingObserved` 上,而 `EmitContext` 是**调用方的对象**、可以被复用给多条流
|
|
6580
|
+
(`startedAtMs` 本来就是这么用的;本包分表门的 J3 段明确支持「同一个 ctx 跑两条并发流」),
|
|
6581
|
+
且那一位**只置 `true`、永不清**:
|
|
6582
|
+
|
|
6583
|
+
- **顺序复用**:上一条流里的一轮缺口,让**下一条 usage 完整的流**的终帧铸出
|
|
6584
|
+
`_sema_usage_lower_bound`、同一拍的 `run_cost_reconciled` 也被标成下界 —— 宿主据此把一条
|
|
6585
|
+
账数得全的 run 渲成「≥」,并把不完整状态持久化;
|
|
6586
|
+
- **并发复用**:两条流互相串这一位。
|
|
6587
|
+
|
|
6588
|
+
「别人那条流有缺口」不是「这条流的数字是下界」的证据 —— 这一位问的是**这一条 run** 的账数全了没有。
|
|
6589
|
+
0.67.0 的头注曾写「复用 ctx 的两条流里只要有一条观测到缺口,两条就都该按下界读 —— 取并是安全的那一侧」,
|
|
6590
|
+
**那句话是错的**:失效方向在这里不是「多说一句」,是**断言一件没发生的事**。
|
|
6591
|
+
|
|
6592
|
+
**修形**:观测位落在 `runStreamInner` 的**局部量**上(per-stream),终局把**本流快照**按值同时交给
|
|
6593
|
+
两个投影口 —— 终帧(`terminalToSdkResult(ev, ctx, observed)` 第三参,additive)与 `run_cost_reconciled`
|
|
6594
|
+
铸臂(`readRunCostFacts(stats, observed)` 第二参),两面读的仍是**同一次计算**。
|
|
6595
|
+
🔴 **不是**靠「终帧那一拍把 ctx 上那一位清掉」修的:并发的两条流会互相覆盖那次清除,而且在共享对象上
|
|
6596
|
+
留一个会漂的位本身就是下一次串账的入口。
|
|
6597
|
+
|
|
6598
|
+
**型面变化**:`EmitContext.usageMissingObserved` **退役**(该位此前逐字写着「由 `runStream` 在流内写,
|
|
6599
|
+
宿主不要自己填」⇒ 端上不该有任何读者/写者)。端自建管线照旧用
|
|
6600
|
+
`readRunCostFacts(stats, { usageMissingObserved })` —— 那是一个**按次调用的入参**,不是跨流留存的状态。
|
|
6601
|
+
|
|
6602
|
+
**消费方待办**:无(cli / web / desktop 均无 `usageMissingObserved` 读写点;wire 键零增删)。
|
|
6603
|
+
|
|
6604
|
+
**黑盒判据(新增,不改既有号)**
|
|
6605
|
+
|
|
6606
|
+
- **G32-21c**:同一个 `EmitContext` 对象**顺序**跑两条流 —— 第一条 `turn_end{usageMissing:true}` + 终帧,
|
|
6607
|
+
第二条每一轮都报 usage、`stats` 也不说 ⇒ 第二条的终帧上**没有** `_sema_usage_lower_bound`、同一拍的
|
|
6608
|
+
`run_cost_reconciled` 上**没有** `usageLowerBound`(两面同时不在场);同一个 ctx 上第三条**自己真有**
|
|
6609
|
+
缺口的流照旧两面都铸(反向自证)。**并发**形:两条交错的流共用一个 ctx,有缺口的那条铸、没缺口的那条
|
|
6610
|
+
不铸。且跑完之后 ctx 上**没有** `usageMissingObserved` 这一格;宿主手填一个也不再影响投影。
|
|
6611
|
+
|
|
6576
6612
|
---
|
|
6577
6613
|
|
|
6578
6614
|
### 32g. F-6(L-228)per-subagent usage 分表(**ADDITIVE**:新 chrome 臂 + 两个终帧超集键)
|
|
@@ -6590,7 +6626,7 @@ prepare 期拒绝 / 合成 abort)。**数字位仍是必填、仍是数到的那
|
|
|
6590
6626
|
| 键 | 载体 | 形 | 缺席语义 |
|
|
6591
6627
|
|---|---|---|---|
|
|
6592
6628
|
| chrome 臂 `subagent_turn_usage` | 每条**子流** `turn_end` 一条(`required: false`) | `{ taskId, parentToolCallId, usage, engineUsage?, usageMissing?, stopReason? }`;`laneProof` 恒是 `{lane:'subagent', parentToolCallId}` | 不接本臂 = 子代用量面在那个宿主上看不见(不是报错) |
|
|
6593
|
-
| `_sema_nested_usage_by_task` | CC result 帧顶层(成功臂 + 错误信封) | `Record<taskId, { turns, inputTokens, outputTokens, cacheReadTokens?, usageMissing? }>` | **一行都没有 ⇒ 整键不铸** —— 空表会被读成「一个子代都没委派」,而真相可能是「委派了但这条流没看见任何一轮」 |
|
|
6629
|
+
| `_sema_nested_usage_by_task` | CC result 帧顶层(成功臂 + 错误信封) | `Record<taskId, { turns, inputTokens, outputTokens, cacheReadTokens?, usageMissing?, keyCollision? }>`(`keyCollision` 见 0.67.2 订正段) | **一行都没有 ⇒ 整键不铸** —— 空表会被读成「一个子代都没委派」,而真相可能是「委派了但这条流没看见任何一轮」 |
|
|
6594
6630
|
| `_sema_nested_usage_by_task_partial: true` | 同上 | 判别位 | 见下;**never false** |
|
|
6595
6631
|
|
|
6596
6632
|
**🔴 供给面的射程边界(如实记,异源对抗复审逼出)**:core 里子代事件有**两条**外送腿,而它们
|
|
@@ -6761,6 +6797,71 @@ setter ——**不产生自有属性**(值是对象时还顺手改了表自己
|
|
|
6761
6797
|
判别力自证三形:全真身份对得上 ⇒ 不铸;全回落且父调用与任务一一对应 ⇒ 不铸;全回落且多任务共父
|
|
6762
6798
|
⇒ 照旧铸(本条不是那一格的替身)。
|
|
6763
6799
|
|
|
6800
|
+
#### 🔧 0.67.2 订正(车 I 件 I-2,异源对抗复审 [medium] 实抓):两种身份的**键空间碰撞**
|
|
6801
|
+
|
|
6802
|
+
订正 ③(G32-25b)漏了一格:出身混合那道闸读的是**行上的**出身位,而行是按**裸 `taskId`** 攒的 ——
|
|
6803
|
+
行键的两个命名空间(真身份 `sourceTaskId` / 回落 `parentToolCallId`)**字面可以相等**,core 的合同
|
|
6804
|
+
并**没有**保证两者互斥。撞字面的两行在**累加那一层**就已经并掉,并掉之后那一行只剩一种出身 ⇒
|
|
6805
|
+
`mixedKeyOrigin` 当场读不出混合,订正 ③ 被绕过。
|
|
6806
|
+
|
|
6807
|
+
反例(异源复审逐字复现):三条子流 `turn_end` 的 `(sourceTaskId, parentToolCallId, inputTokens)` 依次为
|
|
6808
|
+
`('p','a',10)` / `(缺席,'p',20)` / `(缺席,'a',30)`,终局 `nested = {tasks: 2, turns: 3}`。
|
|
6809
|
+
真相是父调用 `a` 下那只任务花了 `10 + 30 = 40`、父调用 `p` 下那只花了 `20`;
|
|
6810
|
+
修前输出 `{p: 30, a: 30}`,**行数 2 === tasks 2**、**turns 之和 3 === turns 3**、出身位读不出混合
|
|
6811
|
+
⇒ `partial` **不铸** —— 一张归属错误的表被标成「可证完整」。
|
|
6812
|
+
|
|
6813
|
+
**修形**(三层,交付面零 BREAKING):
|
|
6814
|
+
|
|
6815
|
+
1. **累加表的键按出身隔离** —— `s:<sourceTaskId>` / `p:<parentToolCallId>`,**逐事件**把来源记进键本身。
|
|
6816
|
+
「一行的每一轮出身一致」从此是**构造保证**,不再是一条会失效的推理。前缀只活在包内:wire 上自带
|
|
6817
|
+
`s:` / `p:` 前缀的 id 因此不会与别人串。
|
|
6818
|
+
2. **交付面的键仍是裸 id**(端零改)。跨命名空间的同字面因此落到同一个交付键上 ⇒ **合并**成一行
|
|
6819
|
+
(数字与轮数按两边之和,不偷偷丢一边;`cacheReadTokens` 的「没报过 ⇒ 键不铸」在合并处照守),
|
|
6820
|
+
并在该行上立 **`keyCollision: true`**(**never false**)。
|
|
6821
|
+
3. **`partial` 恒立**:证不出逐任务归属就不许说完整。
|
|
6822
|
+
|
|
6823
|
+
🔴 **为什么不「两行分列」**:交付形是 `Record<taskId, row>`,同一个裸 id 不可能占两格 —— 要分列就得把
|
|
6824
|
+
出身前缀推到 wire 上,那是三端 BREAKING,而且把包内的编码变成端要解的身份语义。
|
|
6825
|
+
「合并 + 显式标记」是同一批事实下唯一不撒谎的交付形。
|
|
6826
|
+
|
|
6827
|
+
**消费方待办**:`keyCollision === true` 的行**不许**当成某一只子代的账去渲(它是几只混在一起的和);
|
|
6828
|
+
该位在场时 `_sema_nested_usage_by_task_partial` 必然同时在场,只读 `partial` 的端**行为不变**
|
|
6829
|
+
(失效方向仍是只会多说一句「数不全」)。
|
|
6830
|
+
|
|
6831
|
+
**黑盒判据(新增,不改既有号)**
|
|
6832
|
+
|
|
6833
|
+
- **G32-25c**:上面那三帧反例 ⇒ 终帧 `_sema_nested_usage_by_task_partial === true`,`p` 那一行
|
|
6834
|
+
`keyCollision === true`(`turns === 2`、`inputTokens === 30`),`a` 那一行**不受连坐**
|
|
6835
|
+
(`inputTokens === 30`、无 `keyCollision`),交付键仍是裸 `a` / `p`。
|
|
6836
|
+
判别力自证四形:全真身份且计数对得上 ⇒ 两位都不铸;**出身混合但不撞字面** ⇒ 无 `keyCollision`
|
|
6837
|
+
(那一形由 G32-25b 的 `partial` 管,两条判据不混用);**同一命名空间**里的同一个 id(同一只子代
|
|
6838
|
+
两轮)⇒ 正常并行、不是碰撞;wire 上自带 `s:` / `p:` 前缀的 id ⇒ 原样交付、互不并行。
|
|
6839
|
+
碰撞行的键是 `"__proto__"` 时仍产生自有属性、`JSON.stringify` 之后仍可见(G32-23c 那条纪律不因
|
|
6840
|
+
合并臂被绕开)。
|
|
6841
|
+
|
|
6842
|
+
#### 🔧 0.67.2 订正 ②(车 I 件 I-2b,轮二异源复审实抓):分表本身是**本流快照**,不再住在 `EmitContext` 上
|
|
6843
|
+
|
|
6844
|
+
与 §32f 的件 I-1 **逐字同形**的同形存量:「这条流看见了哪几只子代的哪几轮」同样是 per-stream 的事实,
|
|
6845
|
+
而它此前流结束后**留在调用方对象上**(`ctx.nestedUsageByTask`)。`terminalToSdkResult` /
|
|
6846
|
+
`doneToSdkResult` / `failedToSdkResult` 三只都是**公面导出** —— 端自建管线「A 走 `runStream`、
|
|
6847
|
+
B 直调终帧投影」共用一个 ctx 是它们存在的理由 —— 于是 B 的终帧带出 **A 的**分表;B 的 `nested` 计数
|
|
6848
|
+
若恰好与那张表对得上(`tasks` / `turns` 相等),`partial` 还**不铸** = 一张属于别人的表被标成
|
|
6849
|
+
「可证完整」。
|
|
6850
|
+
|
|
6851
|
+
**修形**:分表与观测位一起走终帧投影的**第三参**(`terminalToSdkResult(ev, ctx, { usageMissingObserved,
|
|
6852
|
+
nestedUsageByTask })`),`EmitContext.nestedUsageByTask` **退役**(该位此前同样逐字写着「由 `runStream`
|
|
6853
|
+
在流内写,宿主不要自己填」)。**没传快照 ⇒ 两个分表键都不铸**(诚实缺席:这次投影一条子流都没看见),
|
|
6854
|
+
绝不回头读残留。
|
|
6855
|
+
|
|
6856
|
+
**消费方待办**:无(cli / web / desktop 均无 `nestedUsageByTask` 读写点;经 `runStream` 的路径逐位不变)。
|
|
6857
|
+
自建管线若直调三只终帧投影器且想要分表,把本流的表放进第三参。
|
|
6858
|
+
|
|
6859
|
+
- **G32-25d**:同一个 `EmitContext` 上,A 经 `runStream` 跑完一条带子代的流(`nested` 为 `{tasks:1,turns:1}`)
|
|
6860
|
+
之后,B **直调** `terminalToSdkResult`(自己一条子流都没看见、`nested` 同为 `{tasks:1,turns:1}`)⇒
|
|
6861
|
+
B 的终帧上 `_sema_nested_usage_by_task` 与 `…_partial` **两个键都不在场**;`failedToSdkResult` 同理。
|
|
6862
|
+
跑完之后 ctx 上没有 `nestedUsageByTask` 这一格;宿主手填一份也不再影响投影。
|
|
6863
|
+
判别力自证:经 `runStream` 的下一条流照旧带**自己**那张表。
|
|
6864
|
+
|
|
6764
6865
|
#### ⚠️ 0.67.1 未修的**已知局限**(同族、已立案,本批**刻意不动**):子代**正文**分流仍只认 `parentToolCallId`
|
|
6765
6866
|
|
|
6766
6867
|
`src/adapter/runStream.ts` 的 C1 子代内容分流(`text_delta` / `reasoning_delta` / `text` / `reasoning` /
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@sema-agent/client-core",
|
|
3
|
-
"version": "0.67.
|
|
3
|
+
"version": "0.67.2",
|
|
4
4
|
"description": "Client-side session runtime shared by every sema human client (TUI / web / desktop): sema wire frames (AgentEvent) -> CC session vocabulary (SDKMessage) with dual-plane output (transcript/chrome), deterministic transcript ids, lane discipline as a type, and the notification/dedup ledgers. Every CC-skin shape is collected here so the wire itself stays neutral. Renamed from @sema-agent/wire-cc-adapter (0.1.x).",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|