@sema-agent/client-core 0.84.0 → 0.84.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +98 -0
- package/README.md +12 -5
- package/dist/abortableSleep.d.ts +10 -0
- package/dist/abortableSleep.js +37 -0
- package/dist/adapter/downstream/terminalToSdkResult.d.ts +1 -0
- package/dist/adapter/downstream/terminalToSdkResult.js +10 -5
- package/dist/agentsWireCaps.js +8 -0
- package/dist/decideFailureNote.d.ts +2 -1
- package/dist/decideFailureNote.js +5 -3
- package/dist/detachWire.d.ts +1 -0
- package/dist/detachWire.js +13 -5
- package/dist/displayUntrusted.js +114 -18
- package/dist/engineErrorCodes.d.ts +2 -0
- package/dist/engineErrorCodes.js +3 -0
- package/dist/engineNoticeCodes.d.ts +30 -4
- package/dist/engineNoticeCodes.js +73 -5
- package/dist/engineWireSdk.d.ts +2 -0
- package/dist/hitl/askGateWire.d.ts +1 -1
- package/dist/hitl/askGateWire.js +1 -1
- package/dist/hitl/hitlHostSurface.js +1 -1
- package/dist/hitl/parkResolver.d.ts +1 -0
- package/dist/hitl/parkResolver.js +9 -3
- package/dist/hitl/planReviewWire.d.ts +30 -3
- package/dist/hitl/planReviewWire.js +139 -32
- package/dist/hitl/sessionPolicyDeliverable.d.ts +2 -0
- package/dist/hitl/sessionPolicyDeliverable.js +14 -3
- package/dist/hitl/sessionPolicyWire.d.ts +40 -0
- package/dist/hitl/sessionPolicyWire.js +210 -0
- package/dist/hitl/toolApprovalWire.js +2 -2
- package/dist/index.d.ts +2 -1
- package/dist/index.js +2 -1
- package/dist/liveInitToolFace.js +28 -10
- package/dist/resumeRefusalCopy.d.ts +11 -0
- package/dist/resumeRefusalCopy.js +39 -1
- package/dist/systemReminderTag.d.ts +5 -0
- package/dist/systemReminderTag.js +20 -9
- package/docs/INTEGRATION-CLIENTS.md +461 -16
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,104 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.84.1(2026-09-28)
|
|
53
|
+
|
|
54
|
+
> 主题:**patch** —— 九件同发,零 BREAKING。① `<system-reminder>` 开标签**定位口**(与剥离 / 解包同一个判定);② detach durable-off 400 的**纯判定**(不带进程门,任何宿主都能调);③ 服务端自铸通告码进码册与受众表(`instructions.source_changed` 从运维面回到用户面)+ 它的事实读器 + **按码派发**的 typed 通告事实口;④ 停泊审批决断撞 422 `parked_resume.startup_failed` 时的读口、成因分说的一句话与「同一个决定不重发、deny 照开」判定,决断失败先按码分类再扫词;⑤ 会话规则记录里有引擎拒启名字的会话:**窄撤销动词**(必交这台部署一跑的工具名册)与启动失败**出路句**;⑥ plan-review 编排三口与两只能力探针收**注入 client**(浏览器同源中继宿主零凭证形);⑦ 终帧 CC 形 `permission_denials` 收入参只被传输层脱敏过的拒绝(带脱敏标记);⑧ 凭据读法两处放行收口;⑨ 门夹具第二批换现役形 + 扁平终帧棘轮 + 权威信封标签对账门。根公面运行期导出 1260 → 1275(+15);公面类型 +18;超集键 +1 `_sema_tool_input_redacted`;peer sdk 地板 `>=12.0.1` 不动;开发依赖引擎 `~7.33.1` 不变。🔴 编译期零改动,但已发导出有几处**可观察的行为变化**,Changed 里逐条写了谁要跟:① `noticeAudienceOf('instructions.source_changed')` `operator` → `user`;② `engineNoticeInCatalog` 对两枚服务端自铸码 `false` → `true`;③ `CONFIG_REFUSAL_CODES` 七员 → 八员;④ 终帧 `permission_denials` 收只被传输层脱敏的行、`_sema_permission_denials_absent` 在只剩这一类时不再落下;⑤ `detachDurableOffHint` 读属性抛错时不再外抛、改回 `null`;⑥ `displayUntrusted` 系地址豁免的判法(输出字节与 0.83.4 不同);⑦ 停泊腿遇 422 `parked_resume.startup_failed` 不再判「门已决」(码优先);⑧ `decidePlanReview` 决断撞上请求时限那一臂的结局正文(两条路同一句「已送出、时限内没有答复」,不再说「连不上」;注入 client 的 `timeoutMs` 要求见接入文档 §108a-9)。另 `readReadRootGrantNotice` 入参型放宽为 `unknown`(非破坏)。按原样字节断言过的测试请改按「不含值」写。
|
|
55
|
+
|
|
56
|
+
### Added
|
|
57
|
+
|
|
58
|
+
- **`findSystemReminderOpenTag(text, from?)`** ⇒ `SystemReminderOpenTagSpan | null`(型 `SystemReminderOpenTagSpan = { readonly start: number; readonly end: number }`)(CC-210):引擎 `<system-reminder>` 信封的**开标签定位口**,给自己按开标签把消息切成「信封块 / 正文」的宿主(历史回放拆分、转录搜索这一类),不必再自持一份开标签文法。接入文档 **§108a O-1 / 108a-1**。
|
|
59
|
+
- 从 `from`(缺席 = 0)起找下一枚开标签,回它的 UTF-16 起止(`end` 为开标签后一位,`text.slice(start, end)` 恰是那枚开标签);没有回 `null`。
|
|
60
|
+
- 只认引擎铸的两形:裸 `<system-reminder>`,或恰一个 22 位 base64url `mark` 属性的 `<system-reminder mark="…">`。与 `stripSystemReminderBlocks` / `unwrapSystemReminder` 走**同一个**判定(本版起三只口认开标签是一份实现),所以三口对「哪一处是开标签」永远同答。
|
|
61
|
+
- 串尾截断的开标签(标签名半截、属性值半截、少收引号或 `>`)不认;只认 `start ≥ from` 的那一枚(`from` 落在一枚开标签内部时,那一枚不认)。
|
|
62
|
+
- 与块上下文无关:块正文里的字面开标签、没闭合的开标签、嵌套的内层开标签都认 —— 配闭标签、处置截断块是宿主的循环(剥离口的块语义见接入文档 §104 的 104a-1)。
|
|
63
|
+
- 线性:单次调用一趟扫到命中或串尾;`from` 前移逐枚找完整串,总开销线性。
|
|
64
|
+
- 入参不抛:`text` 非串 ⇒ `null`(不强转);`from` 不是 0 到 `text.length` 之间的整数(负数、小数、`NaN`、`Infinity`、非数)⇒ `null`。
|
|
65
|
+
- **`isDetachDurableOff400(err)`** ⇒ `boolean`(CC-219):这一枚错误是不是对 `x-detach-on-disconnect` 头的拒绝 —— 部署没有耐久 run 账本(没有 run store 或这次提交没有会话)时,提交答 400。判据只有两条:`status === 400`(数字),`message` 是串且含本包的锚句常量 `DETACH_DURABLE_OFF_400_ANCHOR`。接入文档 **§108a DT-1 / 108a-2**。
|
|
66
|
+
- **不认错误码**:这条拒绝带的码是通用前置条件码,同一条提交路上与 detach 头无关的拒绝也用它(比如 device 部署首轮没有会话、cascade 没配梯子);按码判会把那些提交当成「去头重发」的对象。
|
|
67
|
+
- 不读进程台账:浏览器 / 桌面这类从不 arm 终端 detach 台账的宿主也能用同一个判定。`detachDurableOffHint(err)` 本版起就是「头确已发出(`isDetachArmed()`)∧ 本判定」(差别见 Changed)。
|
|
68
|
+
- **一律不抛**:非对象、缺字段、`message` 不是串 ⇒ `false`;读 `status` / `message` 时 getter 抛错、Proxy 陷阱抛错、已撤销的 Proxy ⇒ `false`(不外抛、不重试读)。每个属性至多读一次,先读 `status`,不是 400 就不读 `message`。它通常跑在宿主的 `catch` 里 —— 判定自己再抛,会把原本那一枚错误遮掉。
|
|
69
|
+
- **服务端自铸通告码进码册与受众表**(CC-213):服务端自 7.101.0 起经同一条通告通道投它自己铸的码,这些码不在引擎的码册里 —— 修前 `engineNoticeInCatalog` 对它们答「册外」、`noticeAudienceOf` 回落 `operator`,按受众分发的端把一条给用户看的通告投进了运维面。新公面表 `SERVER_NOTICE_AUDIENCE`(冻结对象,码 → 受众),两行:`instructions.source_changed`(`user`)/ `memory.layer_locked_legacy`(`operator`);`engineNoticeInCatalog` 与 `noticeAudienceOf` 同认(先查引擎镜像表,再查本表,两表之外仍保守判 `operator`;查表按自有属性)。本表**不并进**引擎镜像的 `ENGINE_NOTICE_CODES` / `ENGINE_NOTICE_AUDIENCE`(那两张照旧与引擎码册逐码相等),两表零交集。接入文档 **§108a N-1 / N-2**。
|
|
70
|
+
- **`readInstructionsSourceChanged(notice)`**(CC-213):`instructions.source_changed`(同一会话这一次运行读的项目指令文件与上一次运行读的不是同一份)的事实读器,视图 `InstructionsSourceChangedFactsView { previous: string | null; current: string | null; sessionId: string }`。`previous` / `current` 是指令文件的相对路径,`null` = 那一次运行没有项目指令文件(真读数);三格必填,键缺席、空串、非串非 `null` ⇒ 整只 `undefined`(「不知道」不折成 `null` 或空串)。只认自己那一个码,敌意取值器不抛。接入文档 **§108a N-3**。
|
|
71
|
+
- **按码派发的通告事实口 `readEngineNoticeFacts(notice)`**(CC-214):一条通告 → 按 `code` 判别的联合 `EngineNoticeFacts`(每臂 `{ code, audience, facts }`,`facts` 是那个码的读器视图,`audience` = `noticeAudienceOf(code)`);表里没有这个码 / 读器判不成形 / 坏入参 ⇒ `undefined`,不抛。派发表 `ENGINE_NOTICE_FACT_READERS`(冻结对象,码 → 读器)覆盖本包全部七只事实读器、八个码:`mcp.injection_dropped` / `delegation.ask_unresolvable` / `config.durable_gate_unavailable` / `instructions.source_changed` / `approval.read_root_granted` / `approval.read_root_grant_rejected`(受众 `user`,**只有这六个属用户面**;后两码同是 `readReadRootGrantNotice`,视图按 `outcome` 判别)与 `skills.listing_truncated` / `mcp.server_redialed`(受众 `operator`,只属运维面)。`code` 与 `detail` 各只读一次,读器拿到的是这两只读数的快照。新具名型 `EngineNoticeFactsCode` / `EngineNoticeFactsOf<K>` / `EngineNoticeFacts`。接入文档 **§108a D-1 / D-2**。
|
|
72
|
+
- **停泊审批决断 422 `parked_resume.startup_failed` 的读口、一句话与判定**(CC-225):码常量 `PARKED_RESUME_STARTUP_FAILED`;读口 `parkedResumeStartupFromError(err, decision?)` → `ParkedResumeStartupDetail { code; cause; taskId?; resendable: false; denyStaysOpen: true } | null`(只认 `errorCode` 串,不看状态数字、不读正文、不要求 SDK 错误类);措辞 `parkedResumeStartupContent(detail)`。这枚码有两种成因:续跑没起来(被停泊的代理会话没了,或引擎拒了这次决断本身,如答案挂错了问题);停在一只「代理从发起它的任务继承来的工具」上的批准,而这台服务端交不出那只工具(同一个批准永远不会成,只能拒掉或等服务端升级)。**服务端回体上没有能分开两者的字段** ⇒ 本包只按调用方自己发的决断方向排除:`decision: 'deny'` ⇒ `cause: 'resume_rejected'`(第二种成因只对批准出);批准或未声明 ⇒ `cause: 'undetermined'`,那一句把两条出路都摆出来。判定恒定:同一个决定不自动重发(`resendable: false`),deny 照开(`denyStaysOpen: true`)。两句都不替引擎断言卡还在(第一种成因里一形行已不可再决),都指路「重取待决列表」。新具名型 `ParkedResumeStartupDecision` / `ParkedResumeStartupCause` / `ParkedResumeStartupDetail`。接入文档 **§108a P-1–P-3 / 108a-3**。
|
|
73
|
+
- **会话规则记录里有引擎拒启名字的会话:出路**(CC-229):会话规则记录里一旦有一条引擎会拒的名字(引擎退役名,或含 `__` 而不以协议前缀开头的名字),这条会话此后每一跑都在准备阶段以 `config.legacy_tool_name` 失败、模型一次都调不到。0.83.6 起本包在**写之前**扣下这两类;本版补上**已经写进去**那一侧。接入文档 **§108a E-1–E-6 / 108a-4–108a-6**。
|
|
74
|
+
- **`removeRefusedSessionRules(facade, sessionId, roster, opts?)`** —— 窄撤销动词:读当前记录 → 只拿掉 `toolDeny` / `toolAllow` 里引擎会拒启、**且名字不在 `roster` 里**的条目 → 带版本号整份回写。`roster` = 这台部署上一跑的工具名册(`wiring_manifest` 经 `projectToolRoster` 读出的那一份):引擎那道审计先看名字在不在这一跑的名册里(每行的名 ∪ 别名,原样比、分大小写),在就放行 —— 部署自己挂了一只恰叫退役名的工具(或别名)时,那一条是**生效的**限制,撤了就是把 deny 撤成 allow,所以名册里有的名字一条不碰。名册缺席(`undefined` / `null`)或读不了 ⇒ 结局 `roster_required`,一条不撤、一个请求都不发;读得懂的空名册照常判。因这一码启动失败的那一跑在交出装配清单之前就停了,名册要取自这条会话 / 这台部署上过了准备阶段的一跑。其余桶与条目逐字节原样、序不变;白名单被拿空时留成在场的空列表(一个也不放行 —— 丢掉它才是放宽)。命令桶与目录桶不进引擎那道名字审计,一字不动。撞上并发键只重读重写一次;不读能力位、不按版本猜。结局八臂:`roster_required`(`why`:`absent` / `unreadable`)/ `removed`(带新版本号与撤掉的条目)/ `nothing_to_remove` / `loosen_forbidden`(引擎把这次删除判成放宽:这台引擎上只有运营方能从会话规则记录里删条目,记录不变)/ `conflict` / `read_failed` / `write_refused` / `write_unconfirmed`(写发出去了但裁决没到手或回执对不上 —— 可能已经生效;再调一次就是核对,生效了会答 `nothing_to_remove`)。
|
|
75
|
+
- **`sessionPolicyRemovalNotice(outcome)`** —— 八臂各一句用户面话(唯一措辞真源):放宽被拒那句点明「这台引擎上只有运营方能删」,写不定那句说「可能已经生效」,名册那两句说清为什么要名册,其余失败句说「什么都没改」;撤成那一句**按撤掉的桶分说**(撤 deny:不解开任何挂着的工具,要按现名拦就新加一条 deny;撤 allow:不放行任何新东西,白名单拿空即一个都不放行,要按现名放行就把现名加回白名单);不回显任何规则串或引擎原文。
|
|
76
|
+
- **`legacyToolNameFailureNoteOf(code)`** —— 一跑以 `config.legacy_tool_name` 启动失败时的出路句;只按码判(别的码 ⇒ `undefined`)。句子列出那一条可能在的四处(这条会话的规则记录 / 随跑发送的设置 / 父会话的规则记录 / 部署自己的工具规则)—— 引擎的拒因里没有层位,句子不断言一定在会话规则记录里;不点名谁能删(新旧引擎答案不同,撤销动词的结局会如实说出这一台的答案);不回显名字(引擎自己的错误原文已点名)。
|
|
77
|
+
- **`isEngineRefusedToolName(name)`** —— 引擎准备期按名字拒启的判据本体(退役名表按原字节查;含 `__` 且不以任一协议前缀开头,区分大小写)。它比 `sessionPolicyDeliverable` 的 `legacy_tool_name` 判词**宽**:` foo__bar`(首尾空白)、`foo__*`(带通配)在判词里先落了别的成因,可引擎照样拒启 —— 撤销动词用的是本体。适合给「本会话规则」可见段标坏行。
|
|
78
|
+
- **`CONFIG_LEGACY_TOOL_NAME`**(`'config.legacy_tool_name'`)码常量。
|
|
79
|
+
- 公面类型 +3:`SessionPolicyRemovalOutcome` / `SessionPolicyRemovedEntry` / `SessionPolicyRosterRequiredWhy`。
|
|
80
|
+
- **plan-review 编排收注入 client**(CC-209;接入文档 §7 P-28 的 plan-review 半场)。浏览器同源中继宿主没有合法的引擎目标可装(`EngineWireTarget.token` 只收串,装不进 `{ mode: 'same-origin-relay' }`),此前只能采本包的纯判定、在端侧重建整条决断编排。现在三口各收一个可选的注入连线,宿主把自己构造的 `AgentClient`(中继形即可)交进来。接入文档 **§108a W-1–W-6**:
|
|
81
|
+
- `decidePlanReview(taskId, decision, permissionModeAfter?, opts?)` 第四参 `DecidePlanReviewOptions`:`wire: { client, capsBaseUrl? }` —— 给了就**只**用它(决断 POST 与决后回拉都经它;已装的引擎目标一个字不读,混用 = 把决断发到另一个部署);`reason` —— 随决断上送的理由(`PlanReviewRequest.reason`;任一决断都可带;trim,空白 / 非串 ⇒ 键不落,超 4096 字符截到 4096 并留一行 debug —— 超长会被服务端 413 打回,丢的是整次决断);`onOutcome` —— 结局交回宿主(见下)。请求体键序 `decision` → `reason` → `permissionModeAfter`;不带理由时体逐字节同旧。`default` 撞 400 `request.field_conflict` 的去键重发保留 `reason`。
|
|
82
|
+
- `armPlanReviewApproval(result, sessionKey?, opts?)` 的 `opts` +2 可选位 `wire` / `onOutcome`:给了 `wire` ⇒ 在场判看它(client 不可用 ⇒ 不出卡,`false` + error 留痕 —— 不出一张答不出去的卡),CC-46 三选卡的版本证据从 `wire.capsBaseUrl`(= 宿主 `kickEngineCapsProbe` 用的那把 baseUrl)读,不给 ⇒ `unknown` ⇒ 老两选卡,**不**回落已装目标的 baseUrl;responder 递交决断时把 `wire` / `onOutcome` 原样带给 `decidePlanReview`。
|
|
83
|
+
- `reopenPlanReviewCard` 的 `ReopenPlanReviewOpts` +2 可选位 `wire` / `onOutcome`:缺省投递口从此经注入 client 投递;于是**非默认会话槽给了 `wire` 就不必再自带 `deliverDecision`**(两者都缺席时照旧拒开)。canonical 复用臂(`mintFreshQuestionId: false` 且首呈卡的作答口仍绑着)沿用首呈那一次给的连线,本位不生效(与 `deliverDecision` 同一成文例外,命中留一行 debug)。
|
|
84
|
+
- 注入的 client 不可用(不是对象 / 缺 `assistant.planReview` 或 `runs.get`)⇒ 一个字节都没送出去:结局 `not_sent`,专句「the engine client supplied by the host cannot send it」,绝不回落已装目标。
|
|
85
|
+
- 新型 `PlanReviewWire` / `PlanReviewWireClient`(结构型:`assistant.planReview` + `runs.get`,`AgentClient` 直接满足)/ `PlanReviewDeliveryOptions` / `DecidePlanReviewOptions` / `PlanReviewDecisionOutcome` / `ArmPlanReviewOptions`(`armPlanReviewApproval` 第三参的命名形,字段与此前的内联对象逐字相同 + 两个新可选位)。运行期导出零新增。
|
|
86
|
+
- **决断结局交回宿主**(CC-209):`onOutcome(outcome)` 与队列口 `enqueueMetaPrompt` **并行**(不是二选一):每次被在飞闩放行的决断恰回调一次,载荷 `{ taskId, dispatchNo, decision, effect, detail, tail, prompt, status? }` —— 前四键与队列项机读位 `_sema_planReviewOutcome` 同源,`prompt` 与队列项 `value` 逐字节相同,`detail` 是结局正文、`tail` 是尾句,`status` 是措辞所依据的任务状态(回拉读到的,次之决断 200 体自带的;都没读到 = 键不落)。被闩拒的那次不回调(与「不铸号、不投结局」同律);本地拒发(零 POST)也回调(`effect: 'not_sent'`);回调抛只留一行 debug,不影响队列投递与 promise 落定。**队列口缺席而宿主接了回调 ⇒ 不再打 error 级「MISSED」留痕**(改一行 debug);两条通道都没有才照旧 error 级。没有模型提示注入通道的宿主(浏览器)据此渲自己的回执。
|
|
87
|
+
- **两只能力探针收注入 client**(CC-209):`EngineProbeOpts.client`(`EngineProbeClient` = `capabilities` + `scenarioCapabilities` 两只动词,`AgentClient` 直接满足)。给了 ⇒ `engineSupportsTaskAgents` / `probeScenarioTools` 直接用它,`baseUrl` / `authToken` / `principal` / `fetchImpl` 都不用来构造(不读);`timeoutMs`(3 s / 2 s 缺省)是**独立落定**的截止:到点即答「不知道」并中止底层请求 —— 注入 client 的读重试在退避 / `Retry-After` 睡眠期间不看取消信号,只中止会被拖过预算。失败分类不变(`undefined` / `null`)。建议给注入 client `maxRetries: 0`(本包自构造的 client 同此)。
|
|
88
|
+
- **`_sema_permission_denials[]._sema_tool_input_redacted?: true`**(CC-235,超集位;`SemaPermissionDenial` +1 可选成员,CC 数组同一条目同值):在场 ⇒ 该条 `tool_input` 是同一条流 `tool_start` 帧上的入参,但到达客户端之前已被传输层脱敏 —— 它是那次调用入参的**脱敏视图,不是原值**;拿 `tool_input` 重跑、比对或生成放行规则之前先读它。只在 `true` 时在场,绝不铸 `false`;在场时 `_sema_tool_input_source` 恒为 `'tool_start'`。判据是字面记号(`«redacted` 前缀 / 叶值恰等于两个占位串;多认不少认 —— 入参里本来就含这些字面的调用也会被盖,方向只多一句「可能不是原值」)。接入文档 **§108a PD-2**。
|
|
89
|
+
|
|
90
|
+
### Changed
|
|
91
|
+
|
|
92
|
+
- `stripSystemReminderBlocks` / `unwrapSystemReminder` 的开标签识别改走与定位口同一个判定函数。行为不变:与 0.83.3 冻结实现的定种子对拍、§104 的正负样本照旧逐字节同。
|
|
93
|
+
- `detachDurableOffHint` 改为包 `isDetachDurableOff400`。
|
|
94
|
+
- 🔴 **`detachDurableOffHint` 遇到读属性会抛错的错误对象不再抛**(CC-219):`status` / `message` 的 getter 抛错、Proxy 陷阱抛错、已撤销的 Proxy,此前(已 arm 时)原样抛出,本版起回 `null`(「不是那一枚 400」)。其余输入的返回值逐字不变(门里冻结改前实现逐条对拍);`message` 此前读两次、本版起至多读一次(只有带计数 getter 的对象看得出)。**谁要跟**:按「会抛」写的测试改锚为回 `null`;端上包着它的 try/catch 从此不再因这类对象触发。四端产品源码零改动。
|
|
95
|
+
- 🔴 **`readReadRootGrantNotice(notice)` 的入参型放宽为 `unknown`**(CC-214;此前 `{ code?, detail? } | null | undefined`):与本模块其余事实读器同型,派发口因此对派发表逐行做编译期入参检查。读法与返回值不变。**谁要跟**:零改动 —— 任何实参照旧可传;只有拿它的形参型做类型推导的代码会看到 `unknown`。
|
|
96
|
+
- 🔴 **`noticeAudienceOf('instructions.source_changed')` 由 `operator`(回落)改为 `user`**(CC-213)。**谁要跟**:按 `noticeAudienceOf` 分发的端不改代码即跟着走 —— 这条通告从运维面挪到用户面;若为它写过运维面专属行,那一行跟着挪。按「受众 `operator`」写的判据改锚。`memory.layer_locked_legacy` 的受众仍是 `operator`(此前是回落得来,本版起是表上写明)。接入文档 **§108a N-2 / 108a′**。
|
|
97
|
+
- 🔴 **`engineNoticeInCatalog` 对两枚服务端自铸码(`instructions.source_changed` / `memory.layer_locked_legacy`)由 `false` 改为 `true`**(CC-213)。**谁要跟**:按 `engineNoticeInCatalog` 决定「渲不渲」的端,这两枚码从「只落调试」变成「册内通用行」,不改代码即跟着走;按「这两枚码册外 / 只落调试」写的判据改锚。接入文档 **§108a N-2 / 108a′**。
|
|
98
|
+
- 🔴 **决断失败码优先于「门早已决」扫词臂**(CC-225):修前,停泊审批决断撞 422 `parked_resume.startup_failed` 时,包内停泊审批腿按失败文字扫「门已决」词表(`no pending checkpoint` / `resolved` / `already` / `not found`)——被停泊的代理会话没了那一形,引擎原句恰是「Parked resume did not complete: parked resume failed — the approval is no longer redeemable: Session not found: <id>」,命中 `not found` ⇒ 判成「这张卡早被别处决了」、静默重连一轮,失败文字(连同本码那一句)不上失败面,连续命中后以一句坐标失配的话收场。本版起这枚码先按码分类:不重连,失败文字带本码那一句收场。没有机读码、或别的码的失败,扫词臂照旧。新公面判定 `isCodeClassifiedGateFailure(code)`:这枚决断失败码自己说清了处置、失败文字不许再按那张词表扫(= `isGateStandingErrorCode` 的两码 ∪ 本码);本码**不**进 `isGateStandingErrorCode`(会话没了那一形门已不在)。**谁要跟**:走包内停泊腿的终端零改动即得;端侧自有「门已决」判据链的,在扫词之前问这一口;按「这一形会重连一轮」写的判据作废。接入文档 **§108a P-5**。
|
|
99
|
+
- **四个决断出口的失败文字补上 422 那一句**(CC-225):审批卡两条腿(批准 / 拒绝)、提问卡腿(作答 = 带答案的批准 / 空作答 = 拒绝)、中断撤卡(拒绝)撞 `parked_resume.startup_failed` 时,失败文字(结局 `reason`、宿主日志、中断撤卡上屏的那一行)在原句之后补 ` — <那一句>`,按这一次发的决断方向选句;别的码 / 无码时失败文字逐字节不变。补句不含本包「门早已决」扫词表里的词。**谁要跟**:按失败文字整串断言本码的格多出那一段;别的码零改动。接入文档 **§108a P-4**。
|
|
100
|
+
- `sessionPolicyDeliverable` 的 `legacy_tool_name` 臂改为调用 `isEngineRefusedToolName`(同一张镜像表、同一判序,判决与话一字不变)。
|
|
101
|
+
- 🔴 **配置拒绝识别表 `CONFIG_REFUSAL_CODES` +1 员 `config.legacy_tool_name`(七员 → 八员,排在最后)**(CC-229),与前缀判定 `isConfigRefusalCode` 一致(前缀谓词本来就认它)。**谁要跟**:按成员数或逐员逐序断言这张表的测试改锚;按 `.has` 判的端零改动。
|
|
102
|
+
- 🔴 **`decidePlanReview` 决断撞上请求时限那一臂改说「已送出、时限内没有答复」**(CC-209;两条路同一句):sdk 的每请求时限到点(`TimeoutError`)时这一发其实已经送出、引擎多半正在驱动续跑 —— 此前两条路都落「could not reach the engine」。本版起结局正文是「The plan_review <决断> was sent, but no answer came back from the engine within the time limit this host allows for it — it may have taken effect (…)」,`effect` 仍是 `unconfirmed`;真连不上(连接被拒、断网)照旧是「could not reach the engine」。已装目标那条路的时限是 6 h;注入形用宿主 client 的 `timeoutMs`,🔴 要求不低于 6 h(sdk 的逐请求选项只有取消信号,本包改不了注入 client 的时限;沿用 sdk 缺省 60 s 的 client,续跑超过一分钟的批准会落这一臂),见接入文档 **§108a-9**。**谁要跟**:按「时限到点 ⇒ could not reach」整串断言的格改锚;按 `effect` 分支的端零改动;网页端给中继 client 设 `timeoutMs`(108d)。
|
|
103
|
+
- 决后回拉在注入形上是 15 s 的**独立落定**截止:到点即当回拉失败(回落决断 200 体自带的状态措辞)并中止底层请求,不被注入 client 的读重试退避拖长;已装目标那条路照旧是独立的 15 s 短超时 client。
|
|
104
|
+
- 注入连线本身坏形(`wire` 为 `null` / 非对象 / 读 `client` 就抛 —— JS 调用方)与 client 不可用同一处置:结局 `not_sent`、`onOutcome` 照回、零请求,不抛(此前这一形在决断入口之外抛出,`void decidePlanReview(…)` 留下一只未处理的拒绝、结局与回调都没有)。
|
|
105
|
+
- **不变**:不带新位的三口(请求体 / 出站头 / 结局文字〔时限那一臂除外〕/ 队列项键集 / 投递闩 / 投递序号 / 重开判决)逐字节同 0.84.0;`EngineWireTarget` 型不动;`makeEngineWireClient` 不动;子代族动词(tail / steer / output / compact / task handle / delegated prompt / resume)仍只吃已装目标(见接入文档 §7 P-28)。
|
|
106
|
+
- 🔴 **终帧 CC 形 `permission_denials` 收入参被传输层脱敏过的拒绝**(CC-235):服务端把 `tool_start.args` 转发给客户端之前做一道保形脱敏(结构与键名不动,字符串叶里的凭据形子串换成 `«redacted…»` 记号;环与超深子树换成 `[circular]` / `[depth-limit]` 占位串)。0.73.4–0.84.0 把这一类当成「拿不出可信入参」:那一行只进 `_sema_permission_denials`,并让 `_sema_permission_denials_absent: true` 落下 ⇒ 只读 CC 三键的消费方把一次真实的拒绝读成没发生(入参里写了 `scheme://user:token@host` 形 URL 的命令被拒时就是这一形,与哪条拒绝路径无关)。而同一条输出流里,这只对象早已作为转录 `tool_use.input` 交给了同一批消费方;参照形的记录点手里恒有入参、三键必填,从不因入参形态拒记一条。本版起:入参是普通对象、在扫描预算内、只是被传输层动过 ⇒ 照进 CC 数组,`tool_input` 就是帧上**同一只**对象(与转录 `tool_use.input` 同一个引用;不复制、不反解、不剔记号),条目带 `_sema_tool_input_redacted: true`(见 Added)。CC 数组是 `_sema_permission_denials` 的过滤子集,两处是**同一个条目**、同在同值。清单里只有这一类不完整时,`_sema_permission_denials_absent` 不再落下(它的语义不变 —— 仍是「CC 那条清单不可声称完整」;条数完整了,入参不是原值这件事由标记说)。账本行与流内补行(0.82.1 对账补上的那一半)同一条律。
|
|
107
|
+
- **不变**:真没有入参对象的几类 —— 行无 `toolCallId`、本流没见过那只 `tool_start`(重连后的流、宿主自建管线没交快照、同一个 id 见到两份不同入参)、入参不是普通对象、入参超出扫描预算(节点 / 深度)、本流快照溢出、流内补行的工具名读不出 —— 仍只在 `_sema_permission_denials`,`tool_input` 缺席(绝不铸 `{}`),`_absent: true` 照落。「被传输层动过 ∧ 超预算」按超预算判,不因见到记号先放行。账本自己带入参时账本那份恒优先,不盖标记。
|
|
108
|
+
- 这是**读法变化**:已发的 §66 G2(「入参带会被传输层替换的凭据形串 ⇒ `permission_denials` 空、`_absent: true`」)与 G3 在本版上判红是预期行为。成文改口见接入文档 **§108a′**。**谁要跟**:终端 —— 「这条 run 真的拒过」的判定结论不变(这一类此前经超集载体判「拒了」,现在经 CC 数组判「拒了」),`-p` 结果帧整帧过境零改动即得;凡拿 `tool_input` 重跑 / 比对 / 生成放行规则的读点先读标记(今天没有这类读点)。网页端 / 桌面端 / 管理台今天零读点。验收方按 §66 G2 / G3 写的判据改锚。
|
|
109
|
+
|
|
110
|
+
- 🔴 **凭据读法:`displayUntrusted` 系地址豁免的判法改了,两处「值 / 口令原样上屏」收口(CC-234;输出字节与 0.83.4 不同,断言按「不含值」写)**。机读口(合成终局行、结果帧 `errors[]`)、`displayUntrusted`(全部载体与开关)、`displayUntrustedMarks` 同批。接入文档 **§108a K-1 / K-2 / 108a-8**:
|
|
111
|
+
- **方案词 / 凭据标签与地址形的值之间、紧贴值夹着不可见单字**(格式类字符 —— 零宽空格、软连字符、标签字符等 —— 孤代理项、非空白控制符、没接成序列的 ESC / C1;原字,或本出口画出的可见转义形 `\uD800` / `` 这一类)时,值不再按地址豁免(豁免 = 主机与路径原样、只遮 userinfo / query / 片段)。`scheme://` 留,主机与路径换 `«redacted:secret»`,query / 片段照旧换 `«redacted:query»` / `«redacted:fragment»`。例:`Authorization: Bearer` + U+D800 + `https://secret.example/x` ⇒ `Authorization: Bearer` + U+D800 + `https://«redacted:secret»`。0.83.5–0.84.0 上这一形在正文 / 单行 / 字段三种载体都原样露出值;同一位置换成零宽空格的一形各版都露。
|
|
112
|
+
- **`user:<口令>@` 后面的主机位被可剥单元占住、再往后没有主机**(串尾 / 空白 / `/ ? #` 等右界、`:<端口>`、或下一枚 `@`)⇒ 按 userinfo 遮成 `«redacted:userinfo»@…`;这一形里用户名一段跨过可剥单元(与点形载体上屏的读法一致)。此前只有点形 / 空格形载体顺带遮住(单元画成的 `.` 恰像主机),机读口、默认形、转义形原样上屏。候选挂在一枚凭据标签 / 方案词上 —— 是它的值(分隔一直吃到候选起点,含引号值的开引号:`password: "user:<口令>@` + 零宽空格 + `/ <尾>"`),或把它吞了一截(`token=u:<口令>@…`、`"token":"u:<口令>@…`)—— 时不走这一条,整只值照旧由标签那一遍换记号(与 0.84.0 同答);标签 / 方案词被可剥单元拆开(`to` + 软连字符 + `ken=`、`Bea` + 零宽空格 + `rer `、`tok` + 着色序列 + `en=`)也按剥掉单元的读法认。标签的值是别的词、与候选之间隔着空白(`token=abc user:<口令>@` + 零宽空格)⇒ 候选照认 userinfo。
|
|
113
|
+
- **不变的**:值左邻是真空白、分隔符或引号时照旧按地址豁免(`Bearer https://docs.example.com/x`、`api_key: https://console.example.com/keys` 逐字节不变);紧贴值的是完整转义序列(着色等)时照旧豁免(标签后给值上色是正当排版);`@` 后面直接是真空白 / 串尾的无主机形照旧不算 userinfo;没有冒号的 `a@<单元>` 不动;凭据之外的不可见字符逐字节不动;孤代理项与零宽空格同位同答照旧成立。
|
|
114
|
+
- 🔴 **可见字节差异**(按「输出不含凭据」写的判据零改动;按原样字节写的改锚):不可见单字 / 它的转义形留在原处、不进记号(`Bearer \uD800https://«redacted:secret»`);`Authorization:` 头值上 0.83.4 是整段一枚记号,本版是 `Bearer` + 单字 + `https://«redacted:secret»`;点形载体上「方案词 + 空格 + 单字 + 地址」一形为 `Bearer «redacted:secret»`。**谁要跟**:管理台 —— 提货测试里钉「三载体露值」与零宽空格对照的格按预期翻红,翻面为「不含值 ∧ 含 `«redacted:`」;凡按字节断言凭据输出的格一律改按「不含值」写,不按 0.83.4 原样字节写。终端 / 网页端经本包适配层的机读口升级即净,零改动。
|
|
115
|
+
|
|
116
|
+
### Gates
|
|
117
|
+
|
|
118
|
+
- 新门 `run-system-reminder-open-tag-test.mjs`(30 格):引擎两形逐下标 · 非引擎形 15 形不认、不吞其后真形 · 串尾截断六形不认 · `from` 域外十一形 / 非串八形 `null` 不抛 · 与块上下文无关(嵌套 / 未闭合 / 块正文里的字面开标签)· 2 万条定种子随机串上,定位口找出的开标签集合 ≡ 只用剥离口、解包口的可观察答案反推的集合(不读内部),用定位口重写的剥离 / 解包与两口逐字节同 · 线性两格(单次扫大量近似开标签;`from` 前移走完整串)。
|
|
119
|
+
- 新门 `run-detach-durable-off-verdict-test.mjs`(34 格,其中服务端铸点见证一段只在给了服务端源码树时跑):锚句在场 ⇒ 真(含假 fetch 回 400 体、经 SDK 提交流真抛出的那一枚)· 同码另三句原文 ⇒ 假 · 非 400 七形 / 畸形输入十八形 ⇒ 假不抛 · 不读进程台账 · 冻结改前 `detachDurableOffHint`,除抛错形外逐条对拍返回值(armed / 未 arm × 22 条)· 抛错形五形(getter 抛错、Proxy 陷阱抛错、已撤销 Proxy)纯判定回 `false`、包装层回 `null`,改前两者都抛 · 读次数:`status` 恰一次、`message` 至多一次 · 服务端源码树在场时:通用码的全部 400 铸点里恰一处含锚句且以它起头,其余逐句判假。
|
|
120
|
+
- 按错误文案分支的门:具名豁免那一处从 `detachDurableOffHint` 搬进 `isDetachDurableOff400`(同一处锚句判断,站点数不变)。
|
|
121
|
+
- 扩门 `run-engine-notice-catalog-test.mjs`:I 段(服务端自铸码表两行冻结、与引擎镜像码册和实装引擎码册零交集、引擎镜像不被污染;在册判据与受众读口同认、原型键与脏码不冒充;`readInstructionsSourceChanged` 三格与缺席 / 空串 / 非串 / 原型链 / 敌意取值器各形)· IX 段(可选腿:指向一份装了服务端发布包的目录时,对它的服务端自铸码受众表双向等值、受众逐行同值、与引擎码册零交集、判 `user` 的码在用户流白名单而判 `operator` 的不在、派发口的用户面码全在白名单;没指时如实标为未跑,必跑档缺席即红)· J 段(派发表冻结;本模块每只导出的事实读器都是表里的同一只函数、表里每只都是本模块的导出、表里的码都在册;按码分家交叉矩阵 —— 每只读器恰读得出它挂的那几个码的夹具、别的码一概不给读;表里每个码都必须有门夹具,派发结果与直接读逐键同值、受众与受众读口同答;用户面恰六码、运维面两码;读目录授权两码经派发口视图按 `outcome` 判别、`covers` 坏词经派发口仍 `undefined`;在册无读器 / 册外 / 原型键 / 坏入参 / 敌意取值器 ⇒ `undefined`;`code` 与 `detail` 各只读一次)。另有编译期钉:派发表与本模块导出的事实读器两向相等(删表里一行构建即红)。
|
|
122
|
+
- 新门 `run-parked-resume-startup-test.mjs`(89 格):读口两成因与分不出(正文点名哪种成因都不分臂)、决断方向只认严格 `'deny'`、`taskId` 顶层优先附加键兜底 / 缺席不编;判定两位严格恒值;上游前提钉(实装引擎产物里「继承工具」那一因的判据行只对批准出);两句三段、互异、都指路重取待决列表、分不出那一句摆两条出路、不带扫词表里的词、零内部词;措辞单源(源码全树逐文件扫,两句只在一处铸);别码 / 无码 / 大小写 / 空白 / 兄弟码 / 退役键位 / 敌意取值器 ⇒ `null`,不看状态数字;真 SDK `toApiError(422)` 实例照读;五条出口补对应那一句(整串相等),别的码逐字节同旧;码优先:引擎真正文(会话没了那一形,先自证它命中「门已决」词表)上,停泊提问卡腿作答 / 空作答、审批卡腿批准 / 拒绝都不重连、失败文字带对应那一句,同一正文无码 / 别码照旧走扫词臂重连;`isCodeClassifiedGateFailure` 认本码与两只「门还在」码,本码不在「门还在」判定里。
|
|
123
|
+
- 新门 `run-session-policy-refused-removal-test.mjs`(111 格):判据本体对开发依赖引擎的真退役名表与协议前缀表独立判官逐名两向相等(向量 + 退役名表全员 + 前缀大小写变体 + 定种子随机名),且比 `legacy_tool_name` 判词宽;真 `Runner` + 内存会话规则店,单条名字分别写进 `toolDeny` / `toolAllow` 跑一次 —— 判据本体说「会拒」⟺ 引擎以 `config.legacy_tool_name` 拒启且模型零调用;撤销动词只撤会拒条目(其余桶与条目逐字节原样、序不变、在场空桶留空桶、一读一写、版本号 +1)、撤后同一会话下一跑过准备阶段(语料全员逐名跑,两只桶)、没有要撤的零写;老引擎真店上普通调用方删 `toolDeny` ⇒ `loosen_forbidden`、恰一次写、记录逐字节不变、下一跑仍拒启,只有 `toolAllow` 里有 ⇒ 撤得成;并发键撞一次恰重读重写一次、撞两次 ⇒ `conflict`;读失败 / 写被拒 / 写不定逐类(写不定含裸抛 / 5xx / 无码 4xx / 落盘后抛 / 回执读不懂 / 回执对不上),写不定之后再调一次 ⇒ `nothing_to_remove`;名册:真 `Runner` 挂一只叫 `Task` 的自定义工具(与一只别名是 `KillShell` 的)⇒ 引擎认那一条 deny,撤销动词拿这一跑的真名册不撤它、下一跑照过准备阶段且 deny 还在,名册只按原样比(小写 `task` 不豁免 `Task`,与引擎同),名册缺席 / 读不了十一形 ⇒ `roster_required`、零请求、不抛,读得懂的空名册照撤;八臂措辞两两互异、撤成那句按撤掉的桶分说、零用户字节、表外值不抛;出路句只认一码、列四处位置、不点名谁能删,真 `Runner` 拒启交出的正是这枚码,同一枚码也从调用方策略那一层来(记录为空)⇒ 撤销动词答 `nothing_to_remove`。
|
|
124
|
+
- 新门 `run-plan-review-injected-wire-test.mjs`(119 格):真 dist + 真 sdk 中继 client(spy fetch 路由到假引擎 A),已装目标故意指向另一只假引擎 B —— P 正控(spy 记到 Bearer 头)· W1 十三格场景矩阵注入形 vs 目标形结局 / 体 / 次数 / 机读位逐字节同答且 B 零命中 · W2 投递闩三形 · W3 `field_conflict` 去键重发恰一次 / `acceptEdits` 不重发,CC-46 双闸只从 `capsBaseUrl` 读证据(不回落已装目标)· W4 中继形每一发零 `authorization`(判别力:同一 spy 此刻仍记到带 token client 的 Bearer)· W5 植入的假 token 在四路结局 + 拒连路上零出现于日志 / 结局 / 队列项 / 回调 · W6 reason 七形 · W7 onOutcome 九形 · W8 reopen + wire 三形 · W9 五形坏 client · W10 两只探针 · W11 结构化最小 client 端到端 + 回拉带 signal · W12 决断时限两路同答(真 sdk + 缩放计时器:按接入段要求构造的注入 client 与已装目标对同一只慢回答逐字节同答;注入 client 时限短于回答 ⇒ `unconfirmed`「已送出、时限内没有答复」,已装目标撞上每请求时限 ⇒ 同一句,连接被拒仍是「could not reach」)· W13 读口截止独立落定(真 sdk、读口答 429 + `Retry-After: 60`、缺省读重试:两只探针在 `timeoutMs` 附近落定答「不知道」,决后回拉在截止处落定并回落 200 体状态,截止同时中止底层请求)· W14 坏连线四形(`null` / 串 / 数 / 读 `client` 就抛)不抛、无未处理拒绝、`not_sent` + 回调恰一次。
|
|
125
|
+
- 扩门 `run-permission-denial-projection-test.mjs`(349 → 375 格):F5 改判 —— 五种被动过的形(顶层记号 / 嵌套残片记号 / 裸记号 / 环占位 / 深度占位)各进 CC 数组恰一条、三键齐、`tool_input` 是同一只对象、两载体同一条目严格 `true` 标记 + 来源位、判别位不落、账本摘要位照过境;F1e 没被动过的入参 ⇒ 标记键缺席;F5b 正文里提到占位串不盖标记;F5g / F5g2 被动过 ∧ 超节点预算(记号先被扫到 / 预算先耗尽两序)、F5h 被动过 ∧ 超深度、F5i 入参挂在别的 id 上、F5j 带记号但不是普通对象 ⇒ 仍不进、不盖标记;F6b 账本自带入参 ⇒ 不盖标记;F10b 公开入口 `runStream` 端到端:`tool_input` 与同一条输出流里转录 `tool_use.input` 是同一只对象;L4 改判 —— 流内补行同一条律,L4b 没被动过不盖标记。变异反证七枚逐格见红(账本行标记丢失 / 流内补行标记丢失 / 两条路径各把「没见过 `tool_start`」按空对象收进来 / 见到记号即提前收 / 没动过也铸 `false` / 标记盖到账本自带入参上)。
|
|
126
|
+
- 扩门 `run-display-untrusted-projection-test.mjs`(402 → 414 格):X41 地址形值左邻是不可见单字 14 形 × 12 种出口(机读口、十种呈前载体、标记位置读口)凭据不露 · X41b 原形实得值 · X41c 主机与路径只剩被剥的单元时不凭空铸记号 · X42 空主机 userinfo 10 形 × 12 种出口不露 · X43 / X43b 真空白 / 着色序列 / 无冒号形逐字节不变 · X44 幂等与标记位置读口同字节 · X45 线性 · X46 与上一发布产物差分抽出的最小形(上一版遮住的出口上本版不露)· X47 空主机 userinfo 候选挂在标签 / 方案词上的 15 形(引号值夹单元 3 形 + 标签 / 方案词被单元拆开 12 形)× 12 种出口,口令与尾巴都不露(0.84.0 全遮)· X48 引号内空白池 / 拆标签池各取的样本 5 串 × 12 种出口 · X49 标签不挂在候选上时 userinfo 照认。另对 0.84.0 发布产物跑差分(宽池、拆标签定向池与引号内空白池各三到五组种子,全池五组种子 × 13 种出口;对 0.83.4 发布产物跑两形池):「上一版遮 / 本版露」只剩一类 3 格(点形单行载体上控制空白前的上一个词,KL-181 同类),其余计数 0。
|
|
127
|
+
- **终帧夹具第二批换成现役因由形**(CC-233;只改门,出包面零变化)。只换「结论依赖终态读法」的套:`run-permission-denial-projection-test.mjs`(15 处:成功臂与错误信封是否同一个铸点,哪一臂由终态读法判出)、`run-streamjson-timing-honesty-test.mjs`(13 处:到限码 → `subtype`、合成终局行、抢救正文)、`run-terminal-identity-copy-test.mjs`(18 处:终局行说哪一句)、`run-result-text-backfill-test.mjs`(1 处:补差读的是成功臂上的终答)、`run-parked-resume-startup-test.mjs`(3 处:停泊审批 / 提问的 park 终帧;问答门按引擎真实铸形补 `kind: 'human'`),以及 `run-client-core-pure-test.mjs` 余下的 41 处(终态投影、成本单轨、409 负控里的普通失败、降级链各臂、抢救腿、错误信封单一铸点等)—— `done` 帧从退役平面形(`status` + `errorCode` / `errorMessage` / `blockedReason`)换成 `result.terminal`(`completed` / `failed{code, message}` / `blocked{reason}` / `paused{gate}`)。只换形不改判据:六套换形前后检查数逐一相同且全绿(375 / 41 / 987 / 105 / 89 / 4112)。换形之前,把终态读法的因由分支改坏(成功读成「认不出」、失败从平面键名取码与原话、被挡读成成功),前五套照样全绿;换形之后同一组改动五套都红(权限拒绝投影停在错误信封的正控格,exit 9)。停泊决断 422 那一套另做一次:把因由形的 park 读成「认不出」,换形前全绿、换形后在「作答真走到决断」的自证格停下(exit 9)。每套另做一次夹具负控:把换进来的终态词改成表外词 ⇒ `run-streamjson-timing-honesty-test.mjs`、`run-terminal-identity-copy-test.mjs`(失败 / 被挡各一次)、`run-result-text-backfill-test.mjs`、`run-client-core-pure-test.mjs`(失败一次、终态投影成功臂一次)红,`run-parked-resume-startup-test.mjs` 在自证格停下(exit 9:认不出 park 就走不到决断);`run-permission-denial-projection-test.mjs` 对表外词**不红** —— 表外词落「认不出的终态」臂,那一臂同样是错误信封,而本套的结论只分「成功臂 / 错误信封」;对它改做「被挡 → 成功」的换词负控,当场停在错误信封的正控格。
|
|
128
|
+
- **合法的扁平帧逐格注明**:仍是扁平形的 29 处都是现役 wire 上真会出现的扁平字节,帧上或帧正上方各注一行 `// 409: …`(服务端自己的 active-run 拒绝信封,含老服务端只带人话的那一形)或 `// 回放面: …`(升级前落盘、经读盘面逐字回放的历史行;同步提交 park 体经幂等重放上 SSE;只可能出现在回放行上的退役词 / 表外词防御格)。涉及 `run-client-core-pure-test.mjs`、`run-streamjson-timing-honesty-test.mjs`、`run-terminal-identity-copy-test.mjs`、`run-selfheal-reopen-test.mjs`、`run-result-frame-projection-test.mjs`、`run-display-untrusted-projection-test.mjs`、`run-terminal-cause-projection-test.mjs`、`run-hitl-gate-honesty-test.mjs`、`run-shell-gate-durable-allow-test.mjs`;这些套的检查数不变。
|
|
129
|
+
- 新门 `run-fixture-flat-done-ratchet-test.mjs`(CC-233,12 格):扫 `scripts/` 下每一只 `{ type: 'done', result }` 字面量,把 `result` 追到它真正的对象字面量 —— 就地、经变量、经助手参数(同文件每个调用点各算一处)、经展开、经迭代(`for … of` 的数组行、`[…].map` 回调)五形都追 —— 再判因由形 / 扁平形;扁平位点没有 `// 回放面:` / `// 409:` 注(注释以标记开头、冒号后有理由;字符串 / 模板 / 正则里的 `//` 不算)的,计数不许超过登记物里的上限(只降;低于上限打印 `RATCHET-SLACK`)。本版实测 143 → 23:剩下 23 处都在结论与终态读法无关的九套里(见 Known limits)。尺子先植后量:植入语料里六形扁平、字符串里的假注、没写理由的注、两处真注、三处因由形与一只 `turn_end` 干扰项逐形逐数对拍,对不上 exit 9;真树另有人口地板(文件数、done 字面量数、因由形位点数)。
|
|
130
|
+
- 新门 `run-authority-envelope-mirror-test.mjs`(CC-230 过渡对账,37 格):同事正文写进转录行之前要拆火的权威信封标签表,是引擎那张表的本包镜像;此前没有门对它对账。现在对开发依赖引擎包双向对账 —— 引擎有、本包缺 ⇒ 红(那个标签的伪造信封会原样进转录);本包有、引擎没有 ⇒ 红(同事正文里的普通标签被改写);并核引擎那张表仍由它的信封登记表派生。引擎包根不导出这张表,门按路径读引擎模块里的值,门头注写明这是过渡读法(候镜像表改为构建期生成物)。消费面:三条同事车道渲出的转录行里,引擎每个权威标签(开 / 闭 / 带属性 / 全大写)全被拆火,引擎表里的非权威信封标签原样不动。
|
|
131
|
+
- 登记物:根公面基线 1260 → 1275(+15);型面门 unknown 出境棘轮 417 → 422(新口收宿主 catch 到的原值 / 通告原值,派发口入参放宽一格);`export-liveness` 登记 46 → 44 行(`DETACH_HEADER` / `engineSupportsTaskAgents` 两条 contract 行因有门按名引用而按退出条件删,棘轮 `maxRows` 同批 46 → 44);超集键台账 89 → 90(`_sema_tool_input_redacted`);单例清单 526 → 529(服务端自铸码受众表、派发表、422 两句措辞表入册);同名影子对账门豁免 25 → 26(一端本地同名同义的 detach 判定,到 0.85.0 为止);门数 144 → 151,README「Guards」表 151 行(`run-registry-test.mjs` 那一行改为不写死读它的门数);负控文档「自动化」表 25 行 = 负控套 `CASES` 25(新增两枚:棘轮上限改小 1 ⇒ 门必须红且点名;引擎信封登记表里把一个 framing 行改判 authority ⇒ 门必须红且点名缺的那个标签);棘轮新格 `fixtureFlatDone.unmarkedCeiling` = 23(附一条沿革账);可移植闭包 kernel 18 / adapt 35 / index 213 不变;裸定时器登记 `agentsWireCaps.ts` 2 → 1、`liveInitToolFace.ts` / `hitl/planReviewWire.ts` 出表(三处读口截止收进包内共享等待叶,该叶登记数不变)。
|
|
132
|
+
|
|
133
|
+
### Known limits(本版新增)
|
|
134
|
+
|
|
135
|
+
- 开标签文法常量本身仍不上公面:定位口回「在哪」,不回一段拿去拼正则的源码(与上一版同一取舍,见文末 KL-135 那一条)。
|
|
136
|
+
- `isDetachDurableOff400` 只认锚句:那条拒绝的原文改了,判定答 `false`(方向:不退让、不出提示行,原 400 照常交给宿主),在这条拒绝配上专码之前没有更稳的判据(KL-176)。
|
|
137
|
+
- 422 `parked_resume.startup_failed` 的两种成因在回体上分不开:对批准,本包只能说「分不出」那一句(两条出路都摆);继承工具那一形里,端若照旧把同一张卡连同批准再呈,用户会再撞一次同一个 422。等服务端在回体上带出机读成因位后再分句(KL-177)。
|
|
138
|
+
- 端侧自有「门早已决」判据链(按失败文字扫词)的,要在扫词之前问 `isCodeClassifiedGateFailure`,否则会话没了那一形仍会在端上被认成「门已决」;失败文字被拍平成串、机读码读不到的端路径,只能按 SDK 错误类名兜底(KL-178)。
|
|
139
|
+
- 服务端自铸码表只镜像服务端当前两行;服务端再加自铸码时,本包在对账腿(需指向服务端发布包)上先红,没指发布包的日常跑不出这一红(KL-179)。
|
|
140
|
+
- 拒启会话的出路:撤销动词按宿主交来的**一份**工具名册判「名字在不在名册里」—— 因这一码启动失败的那一跑交不出名册,名册要取自同一部署上过了准备阶段的另一跑;两跑挂的工具不同(子代腿、按场景卸载)时按交来的那一份判(KL-182);退役名表是本包构建时对过账的一份,引擎扩表后新加的名字在本包跟进前撤不掉(撤销动词答 `nothing_to_remove`)(KL-183);老引擎上一份记录同时在 `toolDeny` 与 `toolAllow` 里有坏行时,整份写被判放宽而拒,白名单那几条也一起没撤(KL-184);写口直接拒收这两类名字的 400 尚无专码,到货前落在收紧结局的 `request_rejected`(KL-185);出路句说不出是哪一层、也不点名谁能删(KL-186)。
|
|
141
|
+
- plan-review 注入形的决断腿时限 = 宿主那只 client 的 `timeoutMs`(sdk 缺省 60 s),本包改不了(sdk 的逐请求选项只有取消信号):短于 6 h 的 client,分钟级的续跑会落 `unconfirmed`「已送出、时限内没有答复」;注入 client 的读重试在退避睡眠期间不看取消信号,本包的截止落定之后它仍可能再睡一轮(至多 60 s)才收手 —— 给注入 client `timeoutMs` ≥ 6 h 与 `maxRetries: 0` 即无这两形(KL-191)。子代族动词在中继部署下仍只吃已装目标(接入文档 §7 P-28 余半场)。
|
|
142
|
+
- 拒绝清单上的 `tool_input`(0.73.4 起在 CC 数组里的条目与本版新收的脱敏条目同样)取自 `tool_start` 帧 —— 那是调用**开始时**的入参,先于 hook / 策略在判定前的改写;被拒的是改写后的最终入参,两者可以不同。上游在被拒调用的收口帧上带出最终入参之后改读它(KL-175)。「入参被传输层动过」今天仍按字面记号认:上游没有记录级的机读「已脱敏」位;入参里本来就含记号字面的调用会被多盖标记(KL-11,本版改写)。
|
|
143
|
+
- 凭据读法:地址形的值左邻只有完整转义序列(着色等)、没有不可见单字时,照旧按地址豁免,与真空白同答(KL-180);点形单行载体上,空主机 userinfo 前面隔着换行 / 制表符的上一个词不再随用户名一起遮 —— 上一版那一形是把换行画成点之后的副作用,本版与有主机形两版同答(KL-181)。另两形既有、本版不修:用户名与 `:` 之间紧贴 C1 引导符 U+009B 时机读口与转义形露口令、点形遮住(KL-192);无 scheme 的 userinfo 左界不在 `"` 处断开,主机在场时会把 JSON 键或引号值的开引号吞进记号,引号里空白之后的尾巴原样上屏(KL-193)。
|
|
144
|
+
- 九套(23 处)终帧夹具仍是扁平形、未注:`run-cost-reconcile-projection-test.mjs` 8 / `run-assistant-arm-identity-test.mjs` 4 / `run-subagent-usage-projection-test.mjs` 3 / `run-cost-absence-projection-test.mjs` 2 / `run-tool-roster-projection-test.mjs` 2 / `run-lane-proof-identity-test.mjs` 1 / `run-selfheal-reopen-test.mjs` 1 / `run-subagent-durable-divert-test.mjs` 1 / `run-terminal-facts-projection-test.mjs` 1。它们的终帧只是流的终止符,或只断两臂共有的量(成本、用量、名册、车道证明、子代分表),结论与终态读法无关;由棘轮封顶、只许减少(KL-187)。
|
|
145
|
+
- 棘轮只管一格总数:同一批里删掉一处未注扁平帧、别处新增一处,总数不变 ⇒ 绿。注是声明不是证明,门不判注了「回放面」的帧是不是真的回放面;追不到对象字面量的形(调用返回值、成员读)打印但不计(KL-188)。三处格对「表外终态词」夹具负控不响(那几格只分成功 / 非成功或与终态无关,认不出的终态词按设计落错误信封)(KL-190)。
|
|
146
|
+
- 权威信封标签对账读的是开发依赖引擎包的模块内部值:端连的引擎新于本包对账的那一版、且新版加了权威标签时,新标签在本包跟进之前不被拆火(对账门在抬开发依赖当天红)(KL-189)。
|
|
147
|
+
- 上一版登记的「开标签文法不上公面、按开标签切块的端仍需自持文法」一条本版销(定位口上公面;KL-135);「孤代理项与格式字符同处置之后两形露出」一条本版销(凭据读法收口;残余见 KL-180 / KL-181;KL-143)。
|
|
148
|
+
- 完整台账见接入文档 §108 末行「包侧缺口」。
|
|
149
|
+
|
|
52
150
|
## 0.84.0(2026-09-27)
|
|
53
151
|
|
|
54
152
|
> 主题:🔴 **minor** —— peer sdk 地板 `>=11.3.0` → `>=12.0.1`,同版四件 additive 与一批陈旧逻辑清扫。① 后台代理**登记读数**(缺席行的下半场):读口 + 取代判定 + 归类口 + 登记键桥 + 补行谓词,宿主终于有真读数可以填进缺席行回收入参的 `registry` 位;② 读目录授权两枚**结论通告**的事实读口;③ 接线回执 **`hands` 段**(引擎对「这条腿有没有挂上它自带的文件 / shell 工具」的正面声明);④ 审批词表与云控制面子路径**追平 sdk 12**(读法零改);⑤ **清扫**:退役键 `rewind.rewindFiles` 构造期拒收(专句点名去处)、`WorkflowRunState.agentCount` 改可选、26 个内部件退出包根、九张公面判定表换成只读的 Set 子类、包内判定用的十五张数组运行期冻结、`resume_at` 文本兼容腿退役。根公面运行期导出 1278 → 1260(+8 −26);公面类型 +10;超集键 +1 `_sema_hands`;开发依赖引擎 `~7.33.1` 不变。🔴 换钉前先读下面 BREAKING 五条(终端有一处编译期会红:`agentCount` 改可选)。
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.84.
|
|
38
|
+
**Version:** 0.84.1
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -303,7 +303,7 @@ guard still cross-checks the table by name).
|
|
|
303
303
|
| `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
|
|
304
304
|
| `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there. From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. |
|
|
305
305
|
| `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
|
|
306
|
-
| `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. |
|
|
306
|
+
| `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails. From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. The server also mints a few notices of its own through the same channel; those codes live in a second table with their own audiences, kept apart from the engine mirror (which must stay equal to the engine's catalog) and reconciled against the server's published package when one is supplied, so a user-facing server notice is no longer filed under operations. One dispatcher returns the typed facts for every code that has a reader, discriminated by code and tagged with its audience — only the user-audience codes belong on a user surface — and the guard ties the dispatch table to the module's own exported readers in both directions, so a reader cannot be exported without a row and a row cannot be dropped without the guard failing. |
|
|
307
307
|
| `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever. One reading here answers a question that the terminal state structurally cannot: whether this run was assembled with any file-and-shell tools at all. The engine's terminal vocabulary says a run finished, not whether the work got done, so an orchestrator that waits for the end and then guesses has nothing to guess from — while the assembly manifest already said it at the start, one row per mounted instance with the single condition that mounted it. The reading is three-state and both folds are refused: a roster that is readable and carries no such row is the engine stating a fact, while no roster at all is not that fact — the static half of a manifest never carries one, and an older engine reports rosters without naming the mount condition at all, where an empty count would be a statement about the reader rather than about the run. Those two are kept apart in the reason the reading carries, and the wording for every unknown case is checked never to claim the run had no tools. The same roster now decides the tool list on the first line of a non-interactive run: the host holds that line until the roster arrives and lists exactly what the engine mounted at the start of the run, in mount order. The guard runs a real assembly frame through the projection into the decision, and pins that the host falls back to the estimate only once the roster is known not to be coming — a manifest without one, an unreadable one, model output or the run's end arriving first — rather than on a timer alone (model activity counts, including a model call that is still waiting or retrying; an error line the stream synthesizes when a run fails before assembly counts as the run ending), that a sub-run's manifest is never mistaken for the run's own, that an empty roster is taken as the engine's answer rather than as silence, and that the wait bound covers both sequential default budgets the engine gives an external tool server to connect and list its tools. The holding logic itself lives in the package as a small per-run gate — buffer, decide once, release the held messages in arrival order, then pass through — and the guard drives real stream output through it to pin that the release happens exactly once, at the manifest, releasing exactly the held prefix. The ordering itself also lives in the package as a stream wrapper, and the guard checks the final output a consumer reads: the first line is always the tool-list line, a message that arrives while that line is still being built comes after it, a timer firing races nothing out of order, a source that ends or fails before the decision still gets its first line and held messages out before the error, and an early exit closes the source. From 0.84.0 the roster-derived sentence source no longer throws on a value it does not recognise, including a reading of the manifest's `hands` section passed by mistake: it answers the same "not stated" sentence as the `hands` reader, from one shared source, and its six known sentences do not change. |
|
|
308
308
|
| `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
|
|
309
309
|
| `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, ten words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because an unknown word means different things in each: a denial layer this build does not know may have been added by a newer engine or may come from a damaged record, so its sentence says it cannot tell which instead of asserting damage; the asker vocabulary is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine). Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer; it must not say the question is asked every time — an answer for this one call may come from the person, a hook or an automatic check the deployment runs — and its wording is checked against the engine package's own description of the mandate. A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. Since 0.83.2 a sixth list covers the word a mandated question stands on (`APPROVAL_MANDATE_WORDS`, six words): it must equal the engine's own list word for word and in order, membership is exact, the card reader `readApprovalMandate` answers only for an own key holding one of the six words, and each word has one fixed sentence explaining why the question must be confirmed — six distinct sentences that never point the reader at writing a rule, never promise a question every time, never claim only a person may answer, and repeat no other sentence the package mints. The list is also pinned against the engine's type at compile time in both directions, while the published build references no engine package at all: every `.js` and `.d.ts` file in the build is scanned, and the same scan is first shown to fire on references planted in a scratch directory. |
|
|
@@ -356,7 +356,7 @@ guard still cross-checks the table by name).
|
|
|
356
356
|
| `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
|
|
357
357
|
| `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
|
|
358
358
|
| `scripts/run-client-core-message-branching-test.mjs` | §B8 (branching on error **text**) and §B10 (truthiness standing in for existence when the value can be `0`). AST + type-checker census over `src/`, a named ALLOW list carrying owner and expiry, a known-site floor, and two fixed corpora with a known verdict judged by the same classifier on every run |
|
|
359
|
-
| `scripts/run-registry-test.mjs` | Every ratchet number the suites compare against lives in exactly one place, `scripts/registry.json`: each cell must be a finite non-negative integer, and each one carries a dated ledger of every change (`from` / `to` / `when` / `why`) whose last entry for that cell must equal the number in force — a number moved without an entry, or an entry written without moving the number, is red in both directions.
|
|
359
|
+
| `scripts/run-registry-test.mjs` | Every ratchet number the suites compare against lives in exactly one place, `scripts/registry.json`: each cell must be a finite non-negative integer, and each one carries a dated ledger of every change (`from` / `to` / `when` / `why`) whose last entry for that cell must equal the number in force — a number moved without an entry, or an entry written without moving the number, is red in both directions. Every suite that reads it is checked to read it: no ratchet name may sit next to a numeric literal in their own source (an AST check, so an accounting comment or a test string mentioning the number is fine), each names the cell it reads, and each keeps the step-down channel — when a measurement comes in *under* a ceiling the suite prints one `RATCHET-SLACK` line instead of passing in silence, so slack cannot quietly accumulate under a ceiling nobody lowered. Negative control: a cell rewritten as a string, a fractional or negative cell, a missing cell, a cell with no ledger entry, a ledger entry missing a field, and a ledger tail that disagrees with the number in force each have to make the check speak |
|
|
360
360
|
| `scripts/run-client-core-failloud-test.mjs` | §C1/§C2: an empty `catch` with no comment anywhere inside it, a pure-swallow `catch` nobody reasoned about, and `void <write>` that really returns a Promise with no `.catch`. The exemption instrument is a comment saying why *this* failure may die; the documented-swallow count is a ratchet that only goes down |
|
|
361
361
|
| `scripts/run-client-core-typeshape-test.mjs` | Type discipline as a guard rather than a build side effect: the set of enabled strict knobs (one silently switched off is red), `tsc --noEmit`, and export-surface ratchets for inline anonymous shapes (≥3 members), `unknown` leaving the surface, and bare `unknown` returns — zero slack in either direction |
|
|
362
362
|
| `scripts/run-client-core-singleton-test.mjs` | Module-level singletons ⇄ `docs/refactor/p1-scan/singleton-manifest.json`, **both directions**: an unregistered singleton is red (registering it forces someone to answer "what if this got duplicated"), a stale entry is red, and the `dupRisk: high` count only goes down |
|
|
@@ -392,7 +392,7 @@ guard still cross-checks the table by name).
|
|
|
392
392
|
| `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
|
|
393
393
|
| `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
|
|
394
394
|
| `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
|
|
395
|
-
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and
|
|
395
|
+
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and the scan over them finishes within budget. Since 0.84.1 arguments the transport redacted on the way (a string leaf carrying a replacement token, or a cycle / depth placeholder) join as well: the reference list receives the very object the transcript's tool call shows, and that entry, on both lists, carries a superset flag saying the input is a redacted view rather than the original — redacted arguments that are over budget, not a plain object, or never seen on this stream still stay off the reference list, and the flag is never minted as false; both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false. From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. A later section reconciles the terminal list against the denied calls seen in the same stream, so denials that never reached a human (rules, hooks, classifiers, write protection) are listed too: complete rows join the CC list, rows missing a field stay on the extended list and mark it incomplete. A further section feeds the same raw tool-result frame through the projector into both lanes and requires the denial category on the interactive transcript record, on the non-interactive frame and from the reader on the raw frame to agree, including frames whose settlement, classifier cause or classifier attribution is present but unreadable — a shape the narrowed gate view on the internal frame cannot show — and pins that the reader gives the same answer on the internal frame as on the raw one. |
|
|
396
396
|
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
|
|
397
397
|
| `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
|
|
398
398
|
| `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
|
|
@@ -428,11 +428,18 @@ guard still cross-checks the table by name).
|
|
|
428
428
|
| `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
|
|
429
429
|
| `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed six-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule that names a server or agent peer without naming one of its tools (for every protocol namespace the engine knows, checked against the engine package's own table) or contains a wildcard (*) anywhere (an engine that compares exact names would block nothing), an entry that is not a tool name, and a name the engine refuses at the start of every run — a retired tool name, or one containing "__" without a protocol prefix, where the prefix check is case-sensitive (once such a name is in the record, every later run of the session fails at startup until that entry is removed; the retired-name list is checked entry by entry against the engine package's own list, and a withheld retired name carries its current name when the engine says it was renamed) — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. An end-to-end check runs the engine package itself: every batch this function would deliver — the recorded vectors and a fixed-seed sample of generated names — is written into an in-memory session rule store and the next run must get past its start-up checks and reach the model, while every name withheld as refused — every retired name included — must indeed make that run fail at start-up, and every string literal in the judgement source that it withholds as refused must be one the engine package's own tables refuse. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for four deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; peer-wide rules outside the MCP namespace, which it did not recognise; and names the engine refuses at start-up, which it sent as ordinary names), whose disagreements are counted per class and must match an independent count exactly. |
|
|
430
430
|
| `scripts/run-plugin-hooks-projection-test.mjs` | Plugin hooks: each command hook an enabled plugin declares is decided one by one as running in the engine, running in this client, or not running at all, and the page of hooks sent with a request is built from the same per-turn plan the client uses to skip its own copies, so one hook never runs in two places. Governance is judged first and always wins — a managed hooks switch-off, an untrusted workspace, safe or bare mode, or a governance read that fails sends no plugin hook and does not list it as a gap; managed-hooks-only (set directly, or through a merged non-managed hooks switch-off) keeps only managed plugins; the plugin-only customization lock does not touch plugin hooks. A hook reaches the engine only when this client started the engine on this machine, the engine reports plugin-hook support, the entry is a command, the plugin declares no sensitive option, and the event still fits the engine's per-event limits; the gate walks that matrix cell by cell, including the limit boundaries and a session goal hook counting toward them. A fact that was never read is reported as not known rather than as a fact: a host that does not say where the engine runs gets a "not known whether this client started the engine" reason, an engine whose capabilities have not been read yet gets a "not known yet whether it supports plugin hooks" reason, and the plan's two engine facts are null in those cases, not false. Events the engine never fires run only if the client says it fires them itself, and hooks the upstream behaviour itself refuses (option references in a shell-form command, an unset option in exec form, malformed entries) run nowhere. Exec-form arguments are passed element by element with only saved non-sensitive option references filled in; path placeholders are left for the executor. Sensitive option values never reach the request: with a host that wrongly supplies one, every string in the plan, the request body, the notice, the labels and the log is searched for it across eight cells. A host without the plugin reader keeps the previous request body and gets exactly one warning per settings port; plugin data that throws while it is being read (a throwing getter, a revoked proxy) is treated like a failing reader — no plugin hooks this turn, settings hooks still sent, nothing thrown; the not-running notice names the plugin and events, never a command or an option value, and escapes control characters in names. Command hooks from settings that carry arguments (a non-empty `args` array, which is the exec form, or any other non-null value) are removed from the request until the engine reports support for arguments, because the engine would otherwise drop the arguments and run the bare command through a shell; an empty `args` array is not treated as carrying arguments when the command is made only of letters, digits and `_ . / : + -` (the shell runs the same executable), so such a guard still reaches the engine, while an empty array on a command with spaces or shell characters is removed; `args` on a prompt or http entry, a null `args`, or an entry with no type is left alone, and those go out unchanged. MCP tool hooks, which the engine cannot parse, are removed only from a request built from a plan, whose not-running notice the host shows; a request built without a plan still carries them on engine-fired events, so the engine rejects the whole request loudly instead of a guard hook silently not running — the gate checks both request bodies against the engine's own hooks schema. Without a plan, every removed hook of that kind on an engine-fired event produces one warning per settings port, event and reason. Malformed entries still pass through for the engine to reject loudly, and passing null where the options object goes behaves like passing nothing; a `plan` option that is not a plan is ignored rather than turning the whole page into nothing, and a plan passed directly in place of the options object is recognised and used. A `plugin` key written by hand on a settings hook is stripped before sending (even when its value is undefined), because only hooks that come from the plugin reader may carry plugin context; the settings document itself is left untouched and a debug line records the count. The two hand-copied tables, the engine-fired event list and the engine limits, are checked against their owners. |
|
|
431
|
-
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. |
|
|
431
|
+
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. Since 0.84.1 an address-shaped value (`scheme://…`) that sits right after an invisible single character — a format character, a lone surrogate, a non-whitespace control character, or the outlet's own escape token for one — is no longer let through as an address: the scheme stays and the host and path are replaced, with the query and fragment masked as before; `user:<password>@` followed only by stripped units and then a boundary, a port or another `@` is masked as userinfo. A value after real whitespace or after a colour sequence is still treated as an address, and clean text stays unchanged; forms that the previous release masked are checked not to leak on the same outlets. A `user:<password>@` candidate that is the value of a credential label or scheme word — including a label or scheme word split by stripped characters, and a quoted value that goes on past whitespace — is left to the label pass, so the whole value is masked as in the previous release; both reported shapes and frozen samples from the targeted pools are checked on every outlet. |
|
|
432
432
|
| `scripts/run-ask-survives-posture-test.mjs` | The single posture predicate `askSurvivesPosture(card, facts)` for sessions whose standing mode would otherwise answer approval cards on the user's behalf (bypass-style modes). It reads two facts and returns one of three verdicts. The first is the ask origin stamped on the card: the question tool (`content_question`), an organization rule (`org_rule`), a hook (`hook`), an explicit ask rule (`ask_rule`), an organization policy or rule store that could not be read (`org_unavailable`, `rule_store_unavailable`) and the classifier's hand-off after its denial limit (`denial_limit_fallback`) must still be asked (the engine requires a real person to answer all three) and every other origin this build knows is left to the posture only once the host has also reported that its own ask rules did not match. The second is the host's own reading of its settings ask rules for this call: a positive match must be asked, and a command the host could not fully parse counts as no match. When the host reported no reading, every card outside those seven origins gets `unknown`, because an origin says who asked and not that the user's own ask rules did not match; an origin this build does not know gets `unknown` even after a reported non-match. `unknown` is never an approval: the host falls back to its own settings rules. The guard checks the verdict for every origin word, both with no host reading and with a reported non-match, against an independent table whose word set must equal the package's origin list, so a new upstream word fails the guard until it is classified; it covers the combinations of both facts, malformed inputs (non-boolean readings, empty or non-string origins, prototype keys, a different letter case), the fact that the predicate does not read the stronger bits on the card (those stay with the host's earlier checks), real card requests produced by the live-frame, parked-row and suspended-ask paths, and a closed, frozen verdict shape. Since 0.84.0 the product table itself is also checked against the SDK's runtime list of known origin words, so a word the upstream does not know cannot sit in the table. |
|
|
433
433
|
| `scripts/run-engine-agent-absence-projection-test.mjs` | Absent background agents: when the engine stops reporting a background agent and no final state has arrived, the row is marked absent and this package owns every decision about it, so all clients agree. One predicate says whether a row is absent (the mark, not the status, decides). An absent row keeps its last known status, never counts as running, and is never counted as completed, failed or stopped; its elapsed time stops at the last moment it was seen, and its sentence says it may still be running. The end-of-turn sweep never settles an absent row (or a resident one). A row that comes back, or a real final state for the current cycle, clears the mark; a late final state from an earlier cycle does not. Absent rows are never removed at the short grace window. After the hard limit (30 minutes from the last time they were seen) the host is asked for the background-agent registry reading of each row: only a reading that the agent has ended or is not listed lets the row go, and each removal is returned as a fact the host must act on and announce; a reading of running, unknown, missing or unrecognised keeps the row and schedules nothing, so no standing poll is created. Until a registry reading is available every absent row stays. A row someone is viewing is held and reported separately only once the registry confirms it is gone. The row sentence, the removal sentence and the late-result sentence come from one place, never state an outcome or that the agent finished, and escape control characters in names, in the engine's removal word and in the late-result status. An end-to-end cell drives the real fleet projection and the real absence channel through every decision. From 0.84.0 the package also reads the registry itself: a reader lists the session's background registry only when the server advertises the listing, classifies its failures as unavailable (a failed read keeps the HTTP status and error code when the error carries them, and never invents either), and refuses a partly readable listing as a whole; a classifier turns the listing into the per-row reading, treating a registry status it cannot place as unknown and reading "not listed" only for keys shaped like registry handles and only when the host states the engine is a single process and names the row's session, because a key from another identity space proves nothing by being absent and, on a multi-replica deployment, a listing answered by another replica does not contain this session's agents at all (without that statement a missing row reads unknown and the row stays); and a predicate returns the agents the registry still counts as live that the host has no row for. A listing at the server's 500-row cap cannot show that an agent is gone either, so a missing row there reads unknown; the cap is checked against the engine's own list clamp and, when a server build is supplied, against the server's route. Each listing carries its session and a per-client sequence number, and a listing superseded by a later delivered read, or read for another session than the one the host names, counts for nothing. Only the exact listing object the reader returned can authorize removing a row or adding one: a spread copy, a structured clone, a filtered or a hand-built listing reads unknown and adds nothing, the listing, its row array and every row are frozen, the cap check uses the row count recorded when the page was read, and rewriting a listing's sequence number fools nothing. The fill predicate holds back agents first seen within the absence settle window — timed on one monotonic local clock from when this client's reader first saw the id, so a skew between the server's clock and this one, or reconnecting to a long-running agent, cannot skip the window — agents without a readable registration time and agents the host saw end since the read went out; the registry key of a row, progress or absence event is whichever of its two ids is shaped like a registry handle; and none of the reader, the classifier or the predicate throws. |
|
|
434
434
|
| `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. |
|
|
435
435
|
| `scripts/run-gate-interrupt-safety-test.mjs` | The two safety properties of running the suites themselves. The collecting runner no longer kills a suite outright when its time limit expires: it sends a termination signal first, waits a grace period for the suite to clean up, and only then kills it — and it records the suite as timed out however it exits, so a suite that exits cleanly after the signal is still not counted as passing. The limit and the grace period can be widened for a single suite in `gates-manifest.json` (an optional entry with a written reason; a malformed entry, including one with a misspelled field, stops the runner before any suite starts instead of silently falling back to the default, and an entry filed under a misspelled name is reported by this guard rather than ignored), and the negative-control suite is widened there. Every property is exercised on byte-for-byte copies of the real runner and the real negative-control suite in a throwaway directory: a suite that honours the signal finishes within the grace period, one that ignores it is killed when the grace period ends, and the summary lines are unchanged; once a suite has exited the runner waits at most a short drain window for its output pipes, so a child process that inherited them and outlives the suite neither turns a passing suite into a timeout nor holds the runner past the limit, the grace period and that window; the negative-control suite, interrupted in the middle of a rehearsal by any of the three signals or by the runner's own time limit, restores the file byte-for-byte, leaves no backup behind, stops the rehearsed guard together with anything it started, regenerates the build output (checked on a small project in the throwaway directory), and exits with 128 plus the signal number; a rebuild that would overrun the grace period is cut short and its compiler stopped; a backup left behind by an earlier run — next to a later target, loose in the source tree, or inside a symlinked dependency directory — makes it refuse to start without touching anything, naming the file and how to restore it. |
|
|
436
|
+
| `scripts/run-system-reminder-open-tag-test.mjs` | The opening-tag locator for the engine's `<system-reminder>` envelope, for hosts that split a message into envelopes and body text themselves rather than only stripping or unwrapping it. `findSystemReminderOpenTag(text, from?)` returns the start and end of the next opening tag at or after `from` (UTF-16 indices, `end` one past the tag) or `null`, and it recognises exactly the two shapes the engine mints — a bare tag, or a tag with a single `mark` attribute whose value is 22 base64url characters — through the very same matcher the strip and unwrap entry points use, so there is no second grammar to drift. The guard pins both engine shapes to the exact index, refuses the non-engine shapes listed in the integration guide plus near relatives (and does not let them swallow a real tag that follows), refuses an opening tag truncated at the end of the text, answers `null` without throwing for a non-string text and for a `from` outside the integer range 0..length, and shows that the locator ignores block context (a tag inside a block body, an unclosed tag and a nested inner tag are all located — pairing with a close is the caller's loop, as in the strip entry point). On twenty thousand seeded random strings mixing both shapes, truncations, near relatives and nesting, the set of tags the locator finds equals the set derived purely from the observable answers of the strip and unwrap entry points, and a strip and an unwrap rebuilt on the locator agree with the real ones byte for byte. Two timing cells show a single call over many near-miss tags and a full walk with `from` moving forward both stay linear |
|
|
437
|
+
| `scripts/run-detach-durable-off-verdict-test.mjs` | The verdict behind the durable-off hint, now a public entry point without the process-level gate. `isDetachDurableOff400(err)` answers whether an error is the engine's refusal of the detach-on-disconnect header on a deployment with no durable run ledger: status 400 and the engine's own refusal sentence in the message, and nothing else — in particular not the error code that refusal carries, because the engine uses the same generic precondition code for unrelated refusals on the same submit path, and treating those as this one would silently resubmit a turn without the header. Browser and desktop hosts, which never arm the terminal's detach ledger, can now ask the same question instead of keeping their own copy. The guard pins the positive case on the real sentence (with or without the code, with surrounding text, and on the error the SDK actually throws when a stubbed engine answers 400), the negative case on the engine's other refusals that share the code (verbatim, and again through the SDK), non-400 statuses, and eighteen malformed inputs that must answer `false` without throwing. It shows the verdict does not read the process ledger and that the hint still returns nothing before the header has been sent. The verdict never throws: an error object whose `status` or `message` getter throws, a proxy whose trap throws and a revoked proxy all answer `false`, each property is read at most once (and `message` not at all unless the status is 400), and the hint therefore answers nothing for those objects instead of throwing as it did before. Against a frozen copy of the previous hint, every other input gives the same answer byte for byte, armed or not. When the engine's source tree is available, it also checks that exactly one refusal site with that code carries the sentence and that every other one is answered `false` |
|
|
438
|
+
| `scripts/run-parked-resume-startup-test.mjs` | The one reading and the one sentence for a parked resume that fails while starting (`422 parked_resume.startup_failed`). The engine has two unrelated reasons for it — the parked agent's session is gone or the engine rejected the decision itself, or the approval is for a tool the agent inherits from the task that started it and this server cannot hand that tool over — and the response carries **no field that tells them apart**; the reason lives only in the engine's prose, which this package does not branch on. So the reading narrows the cause only by what the caller itself sent: a deny can only be the first reason (the inherited-tool refusal is minted for approvals alone, and the guard pins that upstream premise), while an approval, or a caller that does not say, gets one sentence that lays out both ways forward instead of guessing. Either way the verdict is fixed: sending the same decision again is not a recovery path, and deny stays available. The two sentences are minted in one place, never claim the card is still waiting (in one of the two shapes it is not), and avoid every word the package's own "already decided" scan looks for, since they are appended to failure text that scan reads; all five decision exits append exactly that sentence, and any other failure keeps its text byte for byte. The code also takes precedence over the package's "already decided" word scan: the engine's own text for the gone-session shape contains "not found", which that scan used to read as "someone already settled this card" and answer with a silent reconnect; now a failure carrying this code is classified by the code (no reconnect, the sentence reaches the failure text), while failures with no code or another code are scanned exactly as before. The predicate that says which codes classify themselves is public, so a client with its own "already decided" chain can ask it before scanning words. |
|
|
439
|
+
| `scripts/run-plan-review-injected-wire-test.mjs` | The plan-review orchestration accepts a host-supplied engine client, so a browser host behind a same-origin relay — which cannot install a token-bearing engine target — runs the package's own decision path instead of rebuilding it. With a client injected, the decision, its in-flight latch, the resend-once-without-the-key handling of a `request.field_conflict` refusal, the post-decide status re-pull, the six-state effect classification and the outcome text are byte-for-byte what the installed-target path produces, proven against a real relay-form SDK client and a fake engine while the installed target points at a different fake engine that must receive nothing. The engine-version evidence behind the three-choice card is read only under the cache key the host names, never from the installed target; every request of the relay-form client leaves without an `authorization` header, checked beside a token-bearing client on the same spy so the absence is not the spy's blindness; and a planted token never reaches a log line, an outcome sentence, a queue item or the host callback on any of the failure paths. A decision may carry a `reason` (trimmed, dropped when blank, capped at the server's limit), an `onOutcome` callback hands the outcome back to hosts that have no model-prompt queue (called once per admitted decision, never for a latched duplicate, and its own failure never affects delivery), a reopen with an injected client no longer needs a host delivery function on a non-default session, and the two capability probes take the same injected client with a bounded wait. A decision that the injected client's own time limit cuts off is reported as sent without an answer within the time limit (it may have taken effect), in the same sentence the installed-target path uses when its limit is reached, while a connection that cannot be made is still reported as unreachable; a client built with the documented decision budget answers a slow approval exactly as the installed target does. The package's own reads through an injected client — the two probes and the post-decide re-pull — settle at their deadlines even while the client sleeps through a retry back-off, and abort the request underneath; a malformed injected connection is reported as not sent, never thrown. |
|
|
440
|
+
| `scripts/run-session-policy-refused-removal-test.mjs` | The way out for a session that already carries a tool name the engine refuses. The engine refuses to start a run when any rule applied to it names a retired tool, or a name containing "__" without a protocol prefix, so a session whose own rule record holds such a name fails at startup on every run. The guard pins three pieces. First, the refusal test itself: it is compared name by name, in both directions, with an independent reading of the installed engine's own tables (retired names and protocol prefixes) over a corpus that includes padded, wildcard and parenthesised forms, and it is shown to be wider than the verdict used before a write — a padded or wildcard form is withheld there for another reason yet still refuses the run, so a removal driven by that verdict would leave the session broken. Every corpus name is then written alone into the deny list and into the allow list of a real engine run: the test says "refused" exactly when the run fails at startup with the refusal code and the model is never called, while the same shapes in the command lists are not audited at all. Second, the removal verb: it takes only the refused entries out of those two lists and leaves every other entry, list and ordering byte-for-byte, keeping an emptied allow list as a present empty list; after it runs, the same session's next real engine run gets past startup, for every refused name in the corpus and in both lists. Against the installed engine as an ordinary caller, taking a deny entry out is refused as a loosening — reported as such, written exactly once, the record unchanged and the next run still refused, never a false success — while an allow-list-only removal, which narrows, succeeds; against a model of the newer contract the engine has announced, the removal succeeds, and the same model still refuses a removal of anything else. A concurrent writer causes exactly one re-read, judged again on the other writer's record, and a second collision is reported rather than retried; read failures, coded write refusals and unconfirmed writes (no verdict, a receipt that cannot be read, or a receipt that does not account for the removal, including one that still holds a removed entry) each land in their own arm, and running the verb again after an unconfirmed write reports nothing left to remove. Third, the wording: one sentence per outcome; the loosening refusal names operators and says the record is unchanged, the unconfirmed one says the change may already be in effect, and none carries a rule string or engine text. The startup-failure sentence answers to that one code only, lists every place the entry can be set without naming who may remove it, and a real run with the name in the caller's own policy rather than in the session record yields the same code — which is why the sentence may not claim the entry is in the session record. The verb also requires the tool roster reported by a run on the deployment: a name listed there, as a tool name or an alias and compared exactly as the engine compares it, is left in place even when it is a retired name, because the engine accepts that rule — with a real engine run that mounts a deployment tool under a retired name (or with that name as an alias), the removal keeps the deny rule and the next run still starts. Without a roster, or with one that cannot be read, nothing is read or written and the outcome says why; a readable empty roster is still a roster. The removal sentence speaks separately about names taken out of the deny list and out of the allow list. |
|
|
441
|
+
| `scripts/run-fixture-flat-done-ratchet-test.mjs` | Test fixtures that feed the engine's terminal `done` frame are kept in the shape the live engine actually sends. Since the terminal record became a tagged cause (`result.terminal`), the older flat shape (`result.status` plus loose keys) reaches a client only in two ways: a stored row replayed verbatim from before the upgrade, and the server's own refusal envelope — so a fixture written in the flat shape tests the replay path while claiming to test the live one. The suite finds every `{ type: 'done', result }` literal under `scripts/`, follows `result` to the object literal it really comes from (in place, through a variable, through a helper's parameter at each of its call sites, through a spread, or through the rows of an iterated array) and counts the flat ones that carry no one-line note saying they model a replayed row or a refusal envelope. That count may only go down: it is a ceiling kept in the ratchet registry, and a count under the ceiling prints a step-down line instead of passing in silence. A planted corpus with a known number of flat, noted and cause-shaped frames in each of those forms must be counted exactly, and a note that sits inside a string or gives no reason does not count |
|
|
442
|
+
| `scripts/run-authority-envelope-mirror-test.mjs` | The authority-envelope tags that a peer's message body is defused against before it is written into a transcript line. Tags the engine uses to speak with its own authority (reminders, completion notices and the like) must never survive inside a body a peer wrote, or a forged completion notice could be read back on resume as a real one and poison the dedup ledger. The package keeps its own copy of the engine's list, so the copy is reconciled against the installed engine package in both directions: a tag the engine treats as authority and the package lacks is red, and so is a tag the package defuses that the engine does not, since that rewrites ordinary text in a peer's message. The engine does not export this list from its package entry yet, so the check reads it from the engine's own module by path and says so; it also confirms the list is still derived from the engine's envelope registry. At the rendered output, every engine authority tag placed in a body (opening, closing, with attributes, upper case) comes out defused on all three peer lanes, and the engine's non-authority envelope tags come out untouched |
|
|
436
443
|
|
|
437
444
|
Each suite carries a floor that only moves up — a refactor that stops executing a group of
|
|
438
445
|
assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
|
package/dist/abortableSleep.d.ts
CHANGED
|
@@ -1 +1,11 @@
|
|
|
1
1
|
export declare function abortableSleep(ms: number, signal: AbortSignal): Promise<void>;
|
|
2
|
+
export type DeadlineSettled<T> = {
|
|
3
|
+
kind: 'value';
|
|
4
|
+
value: T;
|
|
5
|
+
} | {
|
|
6
|
+
kind: 'error';
|
|
7
|
+
detail: string;
|
|
8
|
+
} | {
|
|
9
|
+
kind: 'deadline';
|
|
10
|
+
};
|
|
11
|
+
export declare function settleWithinDeadline<T>(ms: number, run: (signal: AbortSignal) => Promise<T>): Promise<DeadlineSettled<T>>;
|
package/dist/abortableSleep.js
CHANGED
|
@@ -13,3 +13,40 @@ export function abortableSleep(ms, signal) {
|
|
|
13
13
|
signal.addEventListener('abort', onAbort, { once: true });
|
|
14
14
|
});
|
|
15
15
|
}
|
|
16
|
+
function errorDetail(error) {
|
|
17
|
+
try {
|
|
18
|
+
return String(error);
|
|
19
|
+
}
|
|
20
|
+
catch {
|
|
21
|
+
return '';
|
|
22
|
+
}
|
|
23
|
+
}
|
|
24
|
+
export function settleWithinDeadline(ms, run) {
|
|
25
|
+
const ac = new AbortController();
|
|
26
|
+
return new Promise((resolve) => {
|
|
27
|
+
let settled = false;
|
|
28
|
+
const timer = setTimeout(() => {
|
|
29
|
+
if (settled)
|
|
30
|
+
return;
|
|
31
|
+
settled = true;
|
|
32
|
+
resolve({ kind: 'deadline' });
|
|
33
|
+
ac.abort();
|
|
34
|
+
}, ms);
|
|
35
|
+
const finish = (out) => {
|
|
36
|
+
if (settled)
|
|
37
|
+
return;
|
|
38
|
+
settled = true;
|
|
39
|
+
clearTimeout(timer);
|
|
40
|
+
resolve(out);
|
|
41
|
+
};
|
|
42
|
+
let pending;
|
|
43
|
+
try {
|
|
44
|
+
pending = Promise.resolve(run(ac.signal));
|
|
45
|
+
}
|
|
46
|
+
catch (error) {
|
|
47
|
+
finish({ kind: 'error', detail: errorDetail(error) });
|
|
48
|
+
return;
|
|
49
|
+
}
|
|
50
|
+
pending.then((value) => finish({ kind: 'value', value }), (error) => finish({ kind: 'error', detail: errorDetail(error) }));
|
|
51
|
+
});
|
|
52
|
+
}
|
|
@@ -30,6 +30,7 @@ export interface SemaPermissionDenial {
|
|
|
30
30
|
readonly tool_use_id?: string;
|
|
31
31
|
readonly tool_input?: Record<string, unknown>;
|
|
32
32
|
readonly _sema_tool_input_source?: 'tool_start';
|
|
33
|
+
readonly _sema_tool_input_redacted?: true;
|
|
33
34
|
readonly _sema_tool_arg?: string;
|
|
34
35
|
readonly deniedBy?: DeniedBy;
|
|
35
36
|
readonly _sema_denied_by_source?: 'tool_end';
|