@sema-agent/client-core 0.82.7 → 0.83.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +55 -0
- package/README.md +11 -6
- package/dist/adapt/arms.d.ts +3 -2
- package/dist/adapt/arms.js +16 -11
- package/dist/adapt/ids.js +12 -8
- package/dist/adapt/textStream.d.ts +3 -2
- package/dist/adapt/toolCards.d.ts +2 -3
- package/dist/adapt/toolCards.js +1 -1
- package/dist/adapt.d.ts +1 -1
- package/dist/adapt.js +13 -6
- package/dist/adapter/activeRunSelfHeal.js +34 -17
- package/dist/adapter/downstream/eventToSdkMessage.js +1 -1
- package/dist/adapter/downstream/terminalToSdkResult.d.ts +2 -4
- package/dist/adapter/downstream/terminalToSdkResult.js +43 -5
- package/dist/adapter/runStream.js +6 -6
- package/dist/clientSlice.d.ts +0 -5
- package/dist/engineNoticeCodes.js +4 -0
- package/dist/fleet/fleetRowAgentType.d.ts +0 -1
- package/dist/fleet/fleetRowAgentType.js +0 -3
- package/dist/hitl/approvalResolution.js +2 -0
- package/dist/hitl/frameRouter.d.ts +3 -2
- package/dist/hitl/toolApprovalWire.d.ts +1 -0
- package/dist/hitl/toolApprovalWire.js +56 -10
- package/dist/hooksWireCaps.js +27 -1
- package/dist/host.d.ts +1 -0
- package/dist/liveInitToolFace.d.ts +4 -2
- package/dist/liveInitToolFace.js +2 -2
- package/dist/model/catalogLoader.d.ts +1 -6
- package/dist/panelRunningHistory.d.ts +3 -2
- package/dist/peerFrames.d.ts +0 -1
- package/dist/printInitToolFace.js +1 -1
- package/dist/request/printNotification.js +26 -26
- package/dist/seam.d.ts +8 -2
- package/dist/seam.js +18 -7
- package/dist/systemReminderTag.d.ts +0 -1
- package/dist/systemReminderTag.js +0 -3
- package/dist/toolResult.js +1 -0
- package/dist/uuidV5.d.ts +3 -0
- package/dist/uuidV5.js +178 -0
- package/dist/webSearchWireCaps.d.ts +0 -3
- package/dist/webSearchWireCaps.js +21 -51
- package/dist/wireFailureShape.d.ts +2 -1
- package/docs/INTEGRATION-CLIENTS.md +421 -11
- package/package.json +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,61 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.83.0(2026-09-26)
|
|
53
|
+
|
|
54
|
+
> 主题:🔴 **minor**(行为面与型面都有 BREAKING)—— CC 形消息上 24 项非 CC 键按裁定 C-R103([8200] / [8232])改 `_sema_` 名、删除或改走 chrome 臂,转录 id 改成 UUID 形,0.82.7 标过渡的三只联网搜索旧读口到期删除,公面类型 −4;同版另有结果帧 CC 键 `terminal_reason`、用户层 `disableAllHooks` 在引擎腿上生效、退化审批卡三形拒收改写、注入件自愈句整族重写与引擎通告码册 +2;根公面运行期导出 1240 → **1238**,peer sdk 地板 `>=11.3.0` 不动。
|
|
55
|
+
|
|
56
|
+
### BREAKING
|
|
57
|
+
|
|
58
|
+
- **结果帧(`type:'result'`)上四个非 CC 键改名**([8200] 第 2–5 项):降级链 `degraded` → `_sema_model_degraded`(成功臂与错误信封都改;不叫 `_sema_degraded`,那是工具结果记录上另一个意思的键)、错误码 `errorCode` → `_sema_error_code`、成功臂上引擎选定的模型 `model` → `_sema_selected_model`、错误信封上从写出窗抢救回来的正文 `result` → `_sema_salvaged_result`(`result` 只留给成功臂)。值与缺席语义逐位不变(唯一例外:成功臂上的空串 `model` 此前原样铸出,`_sema_selected_model` 对空串缺席,见 Added),旧名自本版起零铸,旧转录不迁移、不改写。接入文档 **§100a K-2–K-5 / 100b**。
|
|
59
|
+
- **合成终局行的行类旗 `isApiErrorMessage` → `_sema_api_error_message`**([8200] 第 10 项):终态错误行与「结局不知道」行两处合成点同改,只在合成行上在场、恒 `true`;包内 print 车道 init 判定闸同批改按新名判「run 已结束」。🔴 端上凭 `isApiErrorMessage === true` 认出合成行、再对正文做凭据洗消的,必须同批改读新名 —— 不改读则洗消整段跳过,报错正文里的凭据(URL 里的用户名口令、`api_key=…` 这类值)原样上屏与落盘。接入文档 **§100a K-10 / 100d**。
|
|
60
|
+
- **工具结果记录删掉过渡名 `toolUseResult`**([8200] 第 1 项,0.81.0 起预告):只留 SDK 面的 `tool_use_result`,值与缺席语义不变,交互记录与 `-p` 帧在结构化结果这一位上从此同形。端内部渲染链沿用驼峰名的,请在自己的入口从 `tool_use_result` 同值起别名。接入文档 **§100a K-1**。
|
|
61
|
+
- **system 行三处退役**([8200] 第 6 / 8 / 9 项):压缩分割线 `compact_boundary` 的附带文件键 `attachedFiles` → `_sema_attached_files`(值原样);本包不再给 system 行补到达时戳 `timestamp`(`user` / `assistant` 行照旧恒带,入参自带的原样过境),也不再在任何 CC 形消息上铸 `isMeta`。包内把附带文件投成「Referenced file」附件行的读点本版两名并读(新名优先),0.84.0 起只认新名。接入文档 **§100a K-6 / K-8 / K-9**。
|
|
62
|
+
- **print 车道完成通知帧的十三个归因位平铺成 `_sema_` 蛇形名**([8200] 第 11–23 项):`taskNotificationToPrintFrame` 出口帧上 `stoppedBy` / `resumable` / `partial` / `exitCode` / `diagnostics` / `result` / `lines` / `recentSteps` / `editedFiles` / `task_type` / `source` / `seq` / `injected` 依次改为 `_sema_stopped_by` / `_sema_resumable` / `_sema_partial` / `_sema_exit_code` / `_sema_diagnostics` / `_sema_result` / `_sema_lines` / `_sema_recent_steps` / `_sema_edited_files` / `_sema_task_type` / `_sema_source` / `_sema_seq` / `_sema_injected`,取值与缺席纪律逐位不变。CC 已命名的六位(`task_id` / `status` / `summary` / `tool_use_id` / `output_file` / `usage`)、交互车道的完成通知与 wire 读法都不动,`PRINT_NOTIFICATION_FIELD_MATRIX` 的 `field` 列同批换名(`from` 列仍是 wire 原名)。接入文档 **§100a K-11–K-23**。
|
|
63
|
+
- **记忆写入从转录面撤下,改发 chrome 臂 `memory_saved`**([8200] 第 7 / 24 项):一次成功的 `Remember` 写入此前在转录面铸一条 `system{subtype:'memory_saved', writtenPaths:[…]}`,本版起转录面零行,改发 `{kind:'memory_saved', laneProof, id, notes}`(新型 `MemorySavedChromeEvent`;`notes` 是笔记原文、不是路径;`id` 由这次写入的 wire 稳定键派生,同一组 wire 事件重投同值、不与同一张卡的工具结果行撞 id)。`CHROME_ARMS` 把这一臂登记为 `required: true`:要保留「Saved N memories」那一行的宿主必须接它,否则那一行换钉当天从转录里消失 —— 按 kind 分派 chrome 事件、带兜底臂的端(兜底臂丢弃未知 kind)与只放行几只渲染位臂的端,不接这一臂就静默丢。接入文档 **§100a K-7 / K-24 / 100a-2 / 100d**。
|
|
64
|
+
- **转录 id 改成 UUID 形**(CC-147):`deriveTranscriptId` 与适配器铸出的转录行 `uuid` 此前是 `wid_<帧 id>` / `wseq_<seq>` / `wtc_<toolCallId>` 前缀串,现在是由同一个稳定键确定性派生的 UUID 形(小写 8-4-4-4-12);同一条派生路径上的 tool_use 卡 `message.id` 后半段(仍以 `msg_sema_` 开头)、段身份 `_sema_segment_id` / `text_segment_end.segmentId`、段末事件的 `committedUuid` / `committedUuids` 与 `attachment` 事件的 `id` 同批换形。只承诺 UUID 形、同一份适配器输入重放得到同一串 id、同一稳定键得到同一 id(不读时钟、不读随机源);同一组 wire 事件经投影口重投时,只有按 wire `eventId` 取键的行(带 `eventId` 的工具结果行、`memory_saved`)跨重投稳定(0.82.x 同样如此),派生参数不在承诺内,旧转录不迁移、新旧两形共存。按前缀形认「是不是本包铸的」的判据作废,改为比相等。接入文档 **§100a U-1 / 100a-6**。
|
|
65
|
+
|
|
66
|
+
### Added
|
|
67
|
+
|
|
68
|
+
- **`deriveTranscriptId(frame, ctx, suffix?)` 多一个可选第三参**:同一个稳定键派生多条消息时的去撞后缀,三个键空间都生效(此前只有帧 id 空间生效,seq / toolCallId 空间把它丢掉 ⇒ 同一帧派生的两条消息同 id);不传时与此前逐字同。接入文档 **§100a U-1**。
|
|
69
|
+
- **耐久 park 卡透传待决行的 `mandated`**(CC-183):待决列表行顶层带 `mandated: true` 时(引擎 7.31.0 在 park 行铸这一位,服务端 7.101.0 起投到行顶层),卡请求严格透传 `mandated: true`,判据口 `approvalIsMandated` 在 park 卡上从此能答 `true`;其它值与嵌套同名键不认,永不铸 `false`;对更早的服务端行为不变。接入文档 **§100a A-2**。
|
|
70
|
+
- **结果帧铸 CC 键 `terminal_reason`**(CC-175):只铸能从引擎终局推出的四个词 —— 成功 ⇒ `completed`,`limits.max_turns_exceeded` ⇒ `max_turns`,`limits.max_cost_exceeded` ⇒ `budget_exhausted`,`output.invalid` ⇒ `structured_output_retry_exhausted`;其余情形(墙钟 / token 预算到限、分类器拒到限额、取消、未知码、被挡、停泊、结局不知道、409 拒绝信封)键缺席,绝不铸 CC 闭集外的词。判据从已选定的 CC `subtype` 单源派生(success ⇒ completed、三个到限 subtype 各对一词),扁平形回放行上不会与 subtype 自相矛盾。新增公开读口 `terminalReasonForResult(msg)`,与铸点同一只判据(按 `subtype` 读、不读错误码),给端自拼的结果帧与旧转录用。接入文档 **§100a R-1 / 100a-1**。
|
|
71
|
+
- **错误信封也带 `_sema_selected_model`**(C-R103 Q5):done 帧的结果记录带非空串 `model` 时,错误信封同样铸出引擎选定的模型;`failed` 事件帧没有结果记录,这一位如实缺席。它与供应商自报名 `_sema_response_model` 是两位,不合并、不互相回落;空串 / 非串的 `model` 两臂都不铸(此前成功臂会把空串原样铸出)。接入文档 **§100a K-4′**。
|
|
72
|
+
- **`SettingsPort` 新增可选方法 `mergedDisableAllHooks()`,用户层 `disableAllHooks` 在引擎腿上生效**(CC-162):宿主返回设置合并之后的 `disableAllHooks`(与本机 hooks 执行器读的同一个值),为 `true` 时 `hooksForWire()` 只发 managed 设置里的 hooks —— 非 managed 设置关得掉自己的 hooks,关不掉组织下发的;会话目标的 Stop 钩子随之不发,显式开启的终局校验不再因为被挡掉的用户 Stop hook 让位。宿主没实现这个方法 ⇒ 行为同 0.82.x,每个已装的 settings 口经宿主日志口告警一次(`warn`);方法抛错 ⇒ 按合并值为真处置(只发 managed 的 hooks,不是一条都不发)。接入文档 **§100a H-1 / 100a-5**。
|
|
73
|
+
- **`resolveLiveInitToolFace` 新增可选 `opts.webSearchStamped`**:宿主明说这次请求盖没盖 `settings.webSearch`,明说的一律优先(从 settings 段采纳、env 缺席的那一形只有宿主知道);缺省按新判决推(`webSearchVerdictFromEnv(env).kind === 'honored'`),print 车道 init 帧回落面不再跟随旧读口。选项形命名导出为 `LiveInitToolFaceOptions`。接入文档 **§100a W-2**。
|
|
74
|
+
- **引擎通告码册 +2**(CC-186):`ENGINE_NOTICE_CODES` 加 `config.context_settings_swapped`(引擎 7.30.0:上下文覆写表随模型目录换档)与 `config.secret_env_scrubbed`(引擎 7.31.0:子进程环境里按名剥掉了凭据形变量,只列名、不带值),audience 都是 `operator`,码册 69 → 71、顺序与上游同源。开发依赖钉引擎 `~7.31.0`;引擎 7.31.0 的 BREAKING 项 `deprecatedLayers` 本包零读点,peer 地板不动。结构化卡白名单 `STRUCTURED_DETAIL_TYPES` 同批 +`read-inbox`(引擎 7.30.0 的 `ReadInbox` 工具卡,只在挂了跨会话收件车道的部署上出现;漏了这一词那一类卡会退回按模型面文本解析)。接入文档 **§100a C-1**。
|
|
75
|
+
- **`ADAPTER_DIVERGENCES` 多两条并改写两条**:新增 DIVERGENCE-11(system 行到达时戳只由宿主接收口补)与 DIVERGENCE-12(记忆写入改走 chrome 臂);DIVERGENCE-9(工具结果行 uuid 改为 UUID 形)与 DIVERGENCE-10(只铸 `tool_use_result`、驼峰名删除)同批改写文案,常量由 10 条变 12 条。接入文档 **§100a X-2**。
|
|
76
|
+
|
|
77
|
+
### Changed
|
|
78
|
+
|
|
79
|
+
- **忙碌会话自愈里「注入件」一族上屏句整族重写**:`activeRunSelfHealRow(…, 'injected')` 的十二句不再用「A system notification … the run …」这类内部词,改说「A follow-up message sema sent on its own (not one you typed)」与「the reply already in progress / an earlier reply」,每句只讲这是什么、当时什么挡住了它、要不要你动手;句柄从 `(run <id>)` 改为 `(id <id>)`。「不用你动手」只在两档说(宿主声明已跟上那一轮回复;会话已放开、正在重发),「会发」只在候决断那一档说,其余一律如实说没送达 / 模型没看到、不承诺重投。接入文档 **§100a E-1 / 100a-3**。
|
|
80
|
+
- 交接按 `delivery` 分句,修正两处不实:`queued` 时那一轮其实停在一道决断上(按重开判决再分卡已呈 / 答案在路上 / 卡呈不出三句),`parked_for_wake` 时那一轮早已结束、消息没到模型 —— 此前这两形都说「交给了正在工作的那一轮,请看着它」。
|
|
81
|
+
- `resending` 此前说「前一轮被取消了」,对「权限规则自己决了、那一轮自己跑完」那一形不成立,改说「会话已放开」;`not-delivered` 此前说「没能把决断卡摆上屏」,对仍在跑 / 引擎不认得 / 已结束 / 读不出状态几形不成立,改说「屏上没有一张答了就能放开会话的卡」。
|
|
82
|
+
- 处置分类(`selfHealSubmissionDisposition`)、续跟意图(`steerFollowIntent`)与用户形文案逐字节不变,`copy.rowFor` 整行覆盖照旧优先;按旧句做文本匹配的端需改锚。
|
|
83
|
+
- **联网搜索判决跟随服务端 7.101.0**(CC-184):`endpoint` 带 userinfo(用户名或口令任一非空,含只带用户名、含省了 `//` 的 `https:u:p@host`)⇒ 判形错(`malformed`,出错字段 `endpoint`),env 车道同判;空 userinfo(`https://@host`)照收、值原样。原因句逐字镜像服务端的 wire 句:provider 三形「(it is missing | it is not a string | it is not one of them)」、searxngParams「— entry #N …」(重名点出先出现那一项的序号)与 endpoint 的 userinfo 句,都不回显任何用户配的值。对 7.100.x 服务端:带 userinfo 的 endpoint 本包先拒、服务端仍采纳(本包更严);原因句已与服务端 7.101.0 的判官逐字对拍一致(对拍用 7.101.0 发布标签源码构建的判官)。接入文档 **§100a W-3**。
|
|
84
|
+
|
|
85
|
+
### Fixed
|
|
86
|
+
|
|
87
|
+
- **合成终局行的 `uuid` 改成 UUID 形、不再读墙钟**(转录 id 换形的同形存量):终态错误行与「结局不知道」行此前以 `err-<毫秒时间戳>` 当 `uuid`,同一毫秒里合成两行就撞 id(按 `uuid` 去重的宿主会吞掉第二行),也不是 UUID 形;现在与结果帧同一铸法。接入文档 **§100a U-2**。
|
|
88
|
+
- **入参不是工具真实入参的审批卡:三形一律拒收改写、如实说明**(CC-135):流内审批帧的入参超上限被省略或根本没带、且事件流上也没有这只调用的入参时,以及耐久待决行上的入参缺席或被换成超限标记 `{ truncated, bytes }` 时,卡请求带 `argsUnavailable: true`,卡回「编辑后批准」一次都不发(respond / decide / 会话级放行都不发),上屏 `EDIT_REFUSED_ON_BLIND_ASK_WARN_TEXT` 一次,这只审批仍挂着、可照原样批准或拒绝。修前这道闸只覆盖悬挂审批一形,另两形上的改写会原样上 wire,把工具的真实入参换成卡上那份残缺入参;纯批准与拒绝不受影响,有真实入参时改写照常转发。接入文档 **§100a A-1 / 100a-4**。
|
|
89
|
+
- 结局:流内帧腿 `{ decision: 'unresolved', editRefused: true }`(既有形);耐久腿新结局 `{ kind: 'failed', stage: 'card', editRefused: true }`(`FsApprovalOutcome` 的 `failed` 臂 +1 可选位),`approvalResolutionOf` 把它读成 `not_sent` / `edit_refused`(此前读成 `card_unavailable`)。
|
|
90
|
+
- 卡上说明句如实:流内帧入参超上限时按「path 解出来没有」分两句,都不再说「下面的 diff 是从问句重建的」;耐久待决行入参缺席 / 超限各一句,超限标记不再当入参渲上卡。只认与上游标记完全同形的对象(自有键恰好 `truncated: true` 与非负数 `bytes` 两个),近似形照旧当入参渲。
|
|
91
|
+
|
|
92
|
+
### Removed
|
|
93
|
+
|
|
94
|
+
- 🔴 **联网搜索旧三读口 `webSearchFromEnv` / `webSearchFromSettings` / `resolveWebSearch` 退出公面**(0.82.7 标过渡、本版到期):对 7.100.0 及以后的服务端,它们把拼错或缺席的 `provider` 整段丢掉,服务端看到「没带」、改用部署后端。换读 `resolveWebSearchVerdict(env, 原始段, apiKeyFor?)` / `webSearchVerdictFromEnv(env)` / `judgeWebSearchSettings(raw, apiKeyFor?)`(§99),还在调旧读口的端编译期即红。接入文档 **§100a W-1**。
|
|
95
|
+
- 🔴 **公面类型 −4**(CC-171):孤儿类型 `SemaNestedUsageByTask` 删除(从没有产出点真正返回过这个具名形;读终帧上的 `_sema_nested_usage_by_task` / `_sema_nested_usage_by_task_partial` 两个 wire 键即可,键本身不变);`ClientVerbSpec` / `CatalogCacheEnvelope` / `PeerFrameLane` 三个纯类型不再从包根导出(形状不变)。四名在已知各端源码里零取用,运行期导出零变化;同批删掉两个从未导出、零引用的内部函数,并给 11 个本就不在公面上的类型去掉多余的 `export`。接入文档 **§100a X-1**。
|
|
96
|
+
|
|
97
|
+
### Gates
|
|
98
|
+
|
|
99
|
+
- **包侧归层对账门**(CC-176,不改出包面):新门 `run-layering-shadow-export-test.mjs` 从本包一侧同时看四个消费本包的端:终端、桌面端、网页端、管理台。它用 TypeScript 语法树取出每个端**自己声明**的顶层运行时导出,与本包公面运行时导出名求交 —— 同名就是「同一件公共逻辑住了两处」,端上那份不在豁免表(`scripts/layering-shadow-exemptions.json`)⇒ 红。对本包的转口(`export { X } from '@sema-agent/client-core'`)是正形,不算。
|
|
100
|
+
- 读源:本机该端克隆的 `origin/main`,没有这个 ref 时退回已提交的 HEAD。只读,不 fetch,不看工作区。每端打印一行所读 ref、提交号、提交日期;所读提交早于 7 天时多打一行陈旧告警(不判红):本机克隆陈旧,这一端的结果只代表那一刻。`origin/main` 在却读不成提交(ref 悬空)按环境故障处理,不会静默改读 HEAD。
|
|
101
|
+
- 豁免行不常驻:每行必带到期版本(本包版本到了即红;最远只许写到当前 minor 之后第三条 minor 线)与本包票号;端已删掉那份而行还在 ⇒ 红。某个端的源码树不在场时该端打 `SKIPPED-SECTION`(部分跑,不算通过)。
|
|
102
|
+
- 尺子自证:内存里的假端植入四种形状的影子必须恰被抓到、五种合法形一条不抓;读源选择两形、陈旧告警(用构造的旧日期)各有一格;在一个临时仓上验证真实的 ref 解析(不在 ⇒ 退回 HEAD,悬空 ⇒ 故障);某端扫描面为零 ⇒ 报工具故障,不给「零命中」。
|
|
103
|
+
- 基线(09-25,各端本机 `origin/main`):终端 16 条、桌面端 0 条(本机克隆陈旧,只代表 08-13 那一刻)、网页端 6 条、管理台 3 条,全部登记,到期 0.85.0(逐条清单与迁移建议见接入文档 **§100e**)。0.85.0 之前端上既不删、也不改名的,本包门在 0.85.0 当天红。
|
|
104
|
+
- **联网搜索判决门的服务端对拍段**:对拍用的服务端判官可由环境变量 `SEMA_WEBSEARCH_ORACLE_ROOT` 指向一份发布包目录;默认读的服务端构建若早于判官源码最后一次提交,这一段打 `SKIPPED-SECTION` 并点名「构建陈旧」(此前按版本号去比陈旧构建,会误报不一致);读不了该树的提交史按环境故障退出。末行总结如实写这一段跑了没有、对的是哪一版。
|
|
105
|
+
- 登记物:`gates-manifest.json` 136 → 137、README「Guards」表 137 行、负控文档「自动化」表 23 行 = 负控套 `CASES` 23(新增一枚:注入一行已到期的豁免 ⇒ 门必须红且点名);公面基线不动。
|
|
106
|
+
|
|
52
107
|
## 0.82.7(2026-09-24)
|
|
53
108
|
|
|
54
109
|
### Added
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.
|
|
38
|
+
**Version:** 0.83.0
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -81,7 +81,8 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
81
81
|
pure-transient goes to chrome."*
|
|
82
82
|
- **Deterministic transcript ids** — derived from wire stable keys (frame id > seq >
|
|
83
83
|
toolCallId); `ctx.uuid()` only for keyless synthetic frames. Invariant (guarded):
|
|
84
|
-
same stream replayed ⇒ same id sequence.
|
|
84
|
+
same stream replayed ⇒ same id sequence. Since 0.83.0 the ids are UUID-shaped; only the
|
|
85
|
+
shape, same-stream determinism and same-key-same-id are promised, not the derivation.
|
|
85
86
|
- **Lane discipline as a type** — chrome events require a `LaneProof`; a historical ghost-row bug
|
|
86
87
|
family is structurally impossible to reintroduce.
|
|
87
88
|
- **Field-level round-trip guard** — every semantic wire field either maps into
|
|
@@ -323,7 +324,7 @@ guard still cross-checks the table by name).
|
|
|
323
324
|
| `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
|
|
324
325
|
| `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string \| null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
|
|
325
326
|
| `scripts/run-approvals-feed-unknown-test.mjs` | The approvals feed tells three states apart: **N items waiting**, **nothing waiting**, and **this fetch did not come back, so we do not know**. Every way a fetch can fail (the call throwing or rejecting, a body that is not an object, a `livePending` section that is not an array — including the `null` seen in the field, a `pending` that is not an array or holds a malformed row) publishes `{kind:'unknown', why, at, mode}` on the subscription — never an empty snapshot and never silence. Real snapshots carry `kind:'snapshot'`; `snapshot()` still answers only with the last real one (a fact about the past) while `reading()` answers whether it is current (`unobserved` / `present` / `unknown`). Recovery always publishes a real snapshot again, even when the contents are byte-identical to before the failure. An unknown reading is never counted as zero: the awaiting-decision counts read `null`, the view is empty, and the tracker reports no removals, so cards on screen are not retracted for a failed fetch. Retry, backoff and circuit-breaking are unchanged. |
|
|
326
|
-
| `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. From 0.80.0 that last pin reads the syntax tree instead of scanning lines: the same return had been rewritten across several lines with conditional spreads, a shape a line-wise search misses entirely, which would have quietly turned the pin into a check of nothing. |
|
|
327
|
+
| `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. From 0.80.0 that last pin reads the syntax tree instead of scanning lines: the same return had been rewritten across several lines with conditional spreads, a shape a line-wise search misses entirely, which would have quietly turned the pin into a check of nothing. From 0.83.0 the parked leg's card-stage failure carrying `editRefused` (an edited approval on a card whose tool arguments were not available) reads as not sent with the `edit_refused` cause, the same cause the live leg already had; the flag counts only as a strict `true` on the card stage. |
|
|
327
328
|
| `scripts/run-panel-cycle-identity-test.mjs` | The **cycle identity** on agent-panel events and the fleet ledger's **departure read-out**: a background agent may be revived under the same id, so `fleet-row` and `end` events now carry the wire's own `cycleSeq` / `startedAt` when present (absent means the row has no notion of generations, never "generation one"), `isStaleEngineAgentPanelEnd` is the single rule for ignoring a late `end` from a previous cycle (only when both sides carry a comparable identity; absence never drops a real terminal), a changed `cycleSeq` is a new cycle for usage stickiness and buffer coalescing, the notification lane carries `cycleSeq` only when the wire really sent `seq`, and `task_remove` frames reach the host through `onTaskRemoved` with `removeReason` / `cycleSeq` verbatim, a stale previous-generation removal leaving the newer row in place. A terminal row held back because the consumer has no such row yet is also released by the keys it carries itself (its transcript id, or the delegating call id of a subagent already on screen under its wire id), since the key tables are only written once a row has actually been published — the release still goes through the one funnel, so the subagent stays one row; a running frame arriving after such a held terminal row is a stale snapshot when both sides carry a comparable generation and it matches (no event, the held row keeps its final usage), a revival when the frame is provably newer (the held row is dropped), and is treated as a revival when neither side can be compared. A subagent lifecycle event carries `agentType` only from an honest source — the fleet row's own agent type, recorded before it is folded into the row label — and omits the key when there is none, never substituting the display name |
|
|
328
329
|
| `scripts/run-workflow-size-warning-test.mjs` | The **workflow size warning** verdict shared by every host footer / panel: a three-state result (`warn` / `ok` / `unknown`) read off the optional fleet view keys, where an unknown size is never reported as a normal one (absent `totalCount` / `tokens` without positive over-cap evidence is `unknown`, naming the missing keys), positive evidence on either axis wins regardless of absent keys, the per-agent denominator uses the engine's started count only when it is not below done+failed (a smaller value is a stale reading), otherwise falls back to the done+failed lower bound only when both keys are present — and a lower-bound denominator only yields an upper bound of the projection, which can prove *within cap* (`ok`, flagged) but never *over cap* (`unknown`, with the upper bound exposed) — and the prior is used only when the engine itself reports zero started agents; cap precedence env > explicit guideline > default, prototype keys never act as a guideline, the env reader is pure and does not fall through to the second name on a bad first value; caps, guideline table, env names and the three copy variants are single-sourced |
|
|
329
330
|
| `scripts/run-approvals-stream-live-capability-test.mjs` | The engine's live-approval-push self-description (`capabilities.approvalsStreamLive`, engine ≥7.87.1), read the same four-state way as its four sibling capability readers: an absent key is reported as not reported (never folded into `false`), the value must be a strict boolean, and the one decision the feed consumer needs — whether it must keep pulling suspended asks itself — is answered by `livePendingNeedsReconcile`, which only says no when the engine explicitly says it pushes. |
|
|
@@ -350,7 +351,7 @@ guard still cross-checks the table by name).
|
|
|
350
351
|
| `scripts/run-usage-verbatim-channel-test.mjs` | The two complementary usage disciplines (core 3.0.0 metering semantics): the CC `ModelUsage` mirror stays pure (five pinned keys, `totalInputTokens` has no seat), while the sema-owned channel forwards the engine `turn_end.usage` object **verbatim** (six keys, incl. `totalInputTokens`) via `last_turn_usage.engineUsage` / `handle.latestEngineUsage` — honest absence on pre-3.0.0 engines, no fabricated zeros |
|
|
351
352
|
| `scripts/run-plan-review-decide-verify-test.mjs` | `decidePlanReview`'s post-decide honesty ([2315]/[2316], engine RB-471 family): a 2xx from the decide endpoint is **not** a terminal — the wire re-pulls the task status and words the outcome by the real shape (still-locked / legal new gate / genuinely left park / unverified), never claiming success it hasn't earned; when the engine answers that the session's stored resume context cannot be read, the outcome names the unreadable row and says the decision was not applied. Driven against a real fake-engine HTTP server through the shipped dist. The outcome queue item also carries a machine-readable `_sema_planReviewOutcome` (task id, a package-minted dispatch number, decision, effect) so a host can tell which in-flight decision an outcome belongs to without searching the prose; the prose is byte-identical, the number is minted only for a decision the in-flight latch admits, and a caller that passes no metadata gets no key. |
|
|
352
353
|
| `scripts/run-shell-gate-durable-allow-test.mjs` | #110: the durable approval leg for **shell** gates. The tool_end HOLD/REJECT predicate must cover Bash the same way park detection already does (otherwise the park poison frame `Operation aborted` hits the transcript, `endedCalls` swallows the real replayed result, and the user who pressed Yes watches a command that really ran be reported as aborted); a replayed, already-decided park must resume reading the stream instead of being reported as a failed turn; `lastEventId` must track numeric `seq` too. Mutation-proven: each of the three fixes reverted turns the gate red |
|
|
353
|
-
| `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings. A final section pins the decide-operation observer: one `start` in the same tick as the first request and exactly one `end` after the last attempt has settled, across success, retried timeouts, exhausted transient failures, semantic refusal, binding mismatch and both kinds of caller abort, with nothing between retries; an observer that throws or rejects — even when logging that fault fails — never changes what is sent or returned, and the stream-level dependency reaches both durable park legs. A package-internal re-delivery of the same decision (plain approve after an older server rejects the session-scope flag, or a re-send without the attribution key) is reported as one operation with a single start and end. An operation handle only groups sends for the same session and bound call — a send for another gate through the same handle is its own operation — and closing a handle while a send is still in flight defers the end until that send settles. When a person picks "allow for this session" on a parked card and the grant is known not to have been stored — the server answers so, which newer servers do for the gates they can recognise from the parked row as needing a person each time, or the server refuses the session-wide grant with one of the refusal codes that are fixed by the row or the deployment and the package falls back to a plain approval — the parked path now says so with the same line the live path uses, and the receipt carries the server's bit for hosts that call that path directly; an answer without the bit, or any other failure — including a conflict that an internal retry can hit after an earlier send already stored the grant — is treated as unknown and says nothing, a plain approval that never asked for the grant says nothing, and a host logger that throws after a successful decision, on either card path or in the bridge's retry step, can no longer turn it into a second send or a failure. |
|
|
354
|
+
| `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings. A final section pins the decide-operation observer: one `start` in the same tick as the first request and exactly one `end` after the last attempt has settled, across success, retried timeouts, exhausted transient failures, semantic refusal, binding mismatch and both kinds of caller abort, with nothing between retries; an observer that throws or rejects — even when logging that fault fails — never changes what is sent or returned, and the stream-level dependency reaches both durable park legs. A package-internal re-delivery of the same decision (plain approve after an older server rejects the session-scope flag, or a re-send without the attribution key) is reported as one operation with a single start and end. An operation handle only groups sends for the same session and bound call — a send for another gate through the same handle is its own operation — and closing a handle while a send is still in flight defers the end until that send settles. When a person picks "allow for this session" on a parked card and the grant is known not to have been stored — the server answers so, which newer servers do for the gates they can recognise from the parked row as needing a person each time, or the server refuses the session-wide grant with one of the refusal codes that are fixed by the row or the deployment and the package falls back to a plain approval — the parked path now says so with the same line the live path uses, and the receipt carries the server's bit for hosts that call that path directly; an answer without the bit, or any other failure — including a conflict that an internal retry can hit after an earlier send already stored the grant — is treated as unknown and says nothing, a plain approval that never asked for the grant says nothing, and a host logger that throws after a successful decision, on either card path or in the bridge's retry step, can no longer turn it into a second send or a failure. From 0.83.0 it also pins approval cards whose arguments are not the tool's real input: a live frame whose arguments were omitted over the size cap or never sent (the card holds at most a path recovered from the question), and a parked approval whose stored input is missing or replaced by the upstream size marker. On every such card an edited approval sends nothing at all — no respond, no decide, no session-grant attempt — ends as edit-refused, raises the same notice once and flags the card so hosts do not offer editing; both production entry points are driven end to end, and the parked path neither mistakes it for an already-decided gate nor reconnects. The explanatory line is pinned word for word in its four forms: a recovered path is the only thing it claims to have recovered, an unrecoverable one says nothing was recovered, neither claims a reconstructed diff, and the parked forms each say which of the two gaps it is, with near-miss shapes of the size marker still rendered as ordinary input. Real input on the stream, the frame or the record keeps edits flowing, and plain approvals and denials are untouched. |
|
|
354
355
|
| `scripts/run-park-hop-progress-test.mjs` | L-80: the park re-attach loop budgets **stalled** rounds, not parks. A turn where the model keeps hitting gates and every one of them is really decided (a card was answered, the engine really moved on) must never be cut off by the hop budget — the budget counts consecutive rounds that produced no progress, and "the engine revived and immediately parked again on the same coordinates" is not progress. The three non-progress arms (already-resolved, decide-transport-exhausted, and a re-scan that was adopted but led nowhere) share one same-cause limit instead of one arm having a limit and the others having none, and every non-progress re-attach is announced once through the host callback rather than only to the debug log. When the limit is spent the resolver reads the approval queue once more and puts whatever is decidable in front of the user before it gives up; only when there is genuinely nothing to show does it fail soft, and the terminal message then carries the real cause and a real way out instead of a sentence about a budget. On the self-heal side, a reopen verdict that reports `decidedWithoutCard` — the chain settled the gate by rule, so there was no card to present — is progress, not a reopen failure, and the user is not told their message was NOT sent. Negative control: a genuinely empty queue with a run that never moves still fails soft |
|
|
355
356
|
| `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
|
|
356
357
|
| `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
|
|
@@ -371,7 +372,7 @@ guard still cross-checks the table by name).
|
|
|
371
372
|
| `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" |
|
|
372
373
|
| `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
|
|
373
374
|
| `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited. The two "card could not be reopened" rows can carry a host-declared way to keep the conversation, which says the card comes back on resume only if it is still waiting: it is placed before the route that abandons it, never offered for an injected submission, while a decision is still on its way, or once the pending approval has been proven gone (the outcome then carries a flag saying so; the proof only counts before the cleanup card is shown, so a fallback after the card carries no flag unless a fresh read finds the run finished, and a recheck that finds the approval back clears it), the host function is not even called in those cases, and it is treated as unavailable when it throws or returns an empty value; a host can also switch off the engine decide route on the interactive rows while the cancel route stays, and with neither given all four rows are pinned byte-for-byte to the text the previous release produced. |
|
|
374
|
-
| `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `isApiErrorMessage` class flag so brief-mode visibility filtering does not silently drop it. The second half is who the rejected submission belonged to: the self-heal copy told every caller "Your message was NOT sent … send it again", which is three separate untruths for a system injection (a plan-review outcome, a cron wake-up, a task notification) — not the user's message, and not re-sendable, since a host queue marks those non-editable and non-recallable. The injected form says so instead, and the one sentence that promises re-delivery is pinned to the single disposition that earns it: `selfHealSubmissionDisposition` is the same function the host consults before putting the item back on its queue, so the promise and the behaviour cannot drift apart, and the arms where no card could be surfaced state plainly that nothing was delivered and nothing will retry. Since 0.72.6 the same gate pins the **follow intent** after a steer (): a message handed to a live run only pays off if someone tails that run's own event stream, so `steerFollowIntent` decides from the delivery word whether to tail now, after the pending decision, or only after a wake — and the "watch that run" sentence turns into a factual "sema is following that run" **only** when the host declares it attached that tail, so a shell that did not wire it can never claim it did |
|
|
375
|
+
| `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `isApiErrorMessage` class flag so brief-mode visibility filtering does not silently drop it. The second half is who the rejected submission belonged to: the self-heal copy told every caller "Your message was NOT sent … send it again", which is three separate untruths for a system injection (a plan-review outcome, a cron wake-up, a task notification) — not the user's message, and not re-sendable, since a host queue marks those non-editable and non-recallable. The injected form says so instead, and the one sentence that promises re-delivery is pinned to the single disposition that earns it: `selfHealSubmissionDisposition` is the same function the host consults before putting the item back on its queue, so the promise and the behaviour cannot drift apart, and the arms where no card could be surfaced state plainly that nothing was delivered and nothing will retry. Since 0.72.6 the same gate pins the **follow intent** after a steer (): a message handed to a live run only pays off if someone tails that run's own event stream, so `steerFollowIntent` decides from the delivery word whether to tail now, after the pending decision, or only after a wake — and the "watch that run" sentence ("watch that reply" on the rows about follow-up messages sema sent on its own) turns into a factual "sema is following that run" ("… that reply") **only** when the host declares it attached that tail, so a shell that did not wire it can never claim it did. Since 0.83.0 the rows about follow-up messages sema sent on its own carry no engine-internal words and say what happened per delivery shape, promising a resend or "nothing for you to do" only where the code guarantees it |
|
|
375
376
|
| `scripts/run-additive-key-passthrough-test.mjs` | The one disease shape behind two legs: a **closed whitelist / flattening arm** dropping a fact that is already on the wire, while both sides of the seam look correct. (1) The `task_progress` projection carries a registered **key ledger** — a frame populated with every key the service really projects is pushed through the shipped `eventToSdkMessage`, and the set of wire keys that survive must equal the registered pass-through list **name for name in both directions**, so quietly forwarding one more key is as red as quietly dropping one. `model` (the child run's model id, minted by core as `prepared.model.id` and projected by the server since 7.52.1) is the key this batch adds, with the same conditional the server itself applies: a non-empty string or no key at all — an empty string is neither a model id nor "unknown". The ledger is also checked against the fenced list in `docs/INTEGRATION-CLIENTS.md` §3d, so a doc that still says seven keys while the code forwards eight is red rather than merely stale. (2) The decide-failure arms carry the server's S-02 `currentPending` pointer key from a 409 `approval_stale` refusal onto the outcome the host reads. The reader is structural rather than `instanceof`, because the client is host-injected and the class identity is not this package's to assume; a half triple never mints (half a pointer cannot relocate anything), an empty string is not presence, and `checkpointToken` never transits. Both the allow and the deny leg are driven end to end through the real durable approval path — as is the accept-session leg, where a refusal carrying the pointer key must now re-raise instead of silently re-sending the human's answer for the **old** card as a plain approve (one decide call, pointer preserved), while a legacy 400 still falls back exactly as before — and all three flattening points must call the one shared reader — the same-shape residue check that makes "fixed one arm and left the twin" red instead of invisible. (3) The same disease growing on the REQUEST side: the `.mcp.json` → server-spec projection rebuilds each server key by key, and the settings schema deliberately leaves some keys parse-transparent — whatever JSON the file carries reaches the engine untouched, because validating them where the whole domain parses all-or-nothing would let one bad declaration take every server down silently. The whitelist had no row for the newest of them, so an operator's per-tool declarations — the ones the write fence reads — were stripped at the package boundary while both sides looked correct. The criterion is not "is that key handled" but the transparent-key table read out of the INSTALLED schema at runtime, reconciled name-for-name against this leg's ledger, so the day upstream adds a third one this turns red and forces an explicit decision. Behaviour is pinned on both transports, by object identity rather than deep equality (a rebuild would be a second judge), and malformed values must transit UNCHANGED rather than be refused here — the engine refuses them loudly and names the server, whereas a package-side judge can only swallow a declared protection quietly. Absence still mints no key, unknown keys still never reach the wire (the fix is the dropped key, not the gate), and the one transparent key this leg deliberately does not forward is a ledger entry with its own exit condition: it belongs to the deployment plane, and the day the request-plane type declares it the entry's premise is gone and the gate says so |
|
|
376
377
|
| `scripts/run-esc-halt-plan-test.mjs` | The Esc stop decision every client shares: fire the **turn-level** halt first, and escalate to a **run-level** cancel in exactly two cases — the engine itself answered with a 409 from the closed code set (it is saying "there is no in-flight turn here; use cancel for a run-level stop"), or that shot came back with no verdict at all *and* the shell can independently prove a permission card was on screen. Everything else does not escalate. The asymmetry is the whole point and every negative control guards the same direction — deciding *not* to escalate costs the user one more choice on a busy-session card (recoverable), deciding to escalate wrongly tears down a run that was alive and takes every in-flight tool with it (not). So: the closed code set is a **frozen** value, not a `ReadonlySet` — type-level immutability does not stop a consumer's `.add()`, and the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift; the escalation gate is the **conjunction** of that closed set and the 409 status, since honouring the code alone lets a 500 that merely quotes it drive a destructive call; `interrupt.not_held` and `steering.not_running` are deliberately outside the set (the first means *this replica* has no live face — the run may be perfectly alive on another); an unreadable code falls to the no-escalation side; a `parked` flag never overrides a verdict the engine did give, and only strict `true` counts when it did not. The first shot is unconditional by construction — it does not consult `parked`, because the 409 it earns is exactly the verdict the gate wants — and the verdict itself is a closed machine-readable reason word, not display copy. A third escalating case was added once tearing the stream stopped reaping the run: with detach armed, a shot that never lands leaves the run going all the way to the end of the turn, so the Esc the user pressed has no effect at all and nothing on screen says so — the old behaviour had a silent backstop (tearing the stream ended the run) and that backstop is gone. The new fact is held to the same three disciplines as `parked`: it is read only where the engine gave no verdict, it is judged **after** `parked` so an existing host's reason word does not change under it, and only strict `true` counts. Absence is proven to be a no-op rather than asserted — the guard carries its own reference implementation of the previous version's table, runs the full grid through both, requires zero divergence when the new field is omitted, and first shows the comparison really does report a difference on the one cell where the two versions are meant to differ |
|
|
377
378
|
| `scripts/run-peer-frame-projection-test.mjs` | The three engine-injected lanes design/385 puts on the **one** `task_notification` carrier, which are not the same kind of thing at all: a delegated child's uplink (`agentMessage`), another session's message drained from this session's own box (`crossSessionMessage`), and a receipt about one of *this* session's own outbound messages (`crossSessionNotice`). The engine renders none of them inside a `<task-notification>` shell, so a client that projects them as the generic completion card shows "background task finished" while the model read a colleague's sentence — two faces describing different events. The discriminator is pinned to the **typed carrier being present**, never to the `summary` text: those carriers can only be minted by the engine's injection legs (the external `notify()` input is a strict subset of the payload and can wear none of them), while `summary` is filled by every notification there is — so anchoring on text would let any background task impersonate a colleague's message by writing `<agent-message from="…">` into its own summary, and a positive control asserts exactly that payload still projects as the generic card. Fail-closed has two tiers rather than one: a broken **required** field (empty `from`, a non-string `body`, a notice `kind` outside the closed set) returns absence so the caller falls back to the generic card — an honest downgrade where the user still sees the notification — while a broken **optional** field drops only itself, because losing an attribution note and losing a colleague's whole message are not the same magnitude. The provenance side record is **required and must agree on four points** (`kind` matches the lane; `from`/`taskId`/`seq` are present and equal the carrier/payload — each equality is anchored on a core mint site and pinned by the cli wire-anchor A-K24), so a carrier signed with a trusted name but a disagreeing provenance falls back to the generic card; peer bodies pass the same authority-envelope neutralization core applies (`<task-notification>` etc. are defused) so a colleague's text can never seed the resume dedup ledger. Lane precedence copies the engine renderer's own order, because the model already read the frame in that order and a client ordering of its own would put a card on screen that disagrees with the frame the model saw. Rendering and parsing of the transcript line live in the same module and are round-tripped in both directions, including a body carrying a forged closing tag (a parser fooled there hands half a message to the next row) and a quote inside the sender label (which must not forge a second attribute); the notice lane is deliberately kept **out** of the parser, since recognising it would mean anchoring the `[Cross-session …]` prefix and a user typing that same line would be rendered as engine speech. Hostile carriers are read as own **data** descriptors only and accessors are never invoked at all — `catch` catches throwing, not never returning — proven by a counting getter that must stay at zero calls, alongside a revoked proxy and a prototype-only carrier; and four legacy payload shapes assert the no-carrier path is byte-identical to before, which is the executable form of "zero difference for an older host". A re-supplied cross-session message — same task id, status and sequence as the first delivery, handed to the model again after compaction — is not rendered a second time, because the first delivery is still on the user's screen; a record with the next sequence number still renders |
|
|
@@ -420,7 +421,11 @@ guard still cross-checks the table by name).
|
|
|
420
421
|
| `scripts/run-file-history-capture-capability-test.mjs` | The engine's file-history-capture self-description (`capabilities.fileHistoryCapture`), read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `off`), words are taken as an open set so a newer mode is not mistaken for a malformed answer, `fileHistoryCaptureMode` recognises only `off` and `on-always`, and the wording for `off` speaks about capture only — whether code can be rewound is left to the rewind readings. |
|
|
421
422
|
| `scripts/run-model-identity-resolvability-test.mjs` | The model-identity judgement a client makes before letting anyone in: can the engine it is about to use start with a model name? Each end reports what it read from each place that can feed a model name to a local engine (complete, partial — a gateway address or a credential but no model name —, absent, or unreadable), or, for an engine that runs elsewhere, whether that engine has been seen answering; `modelIdentityResolvability` answers resolvable, unresolvable or unknown. Having part of an upstream configuration is not having enough of one, so partial lanes never add up to resolvable; a lane that was not reported or could not be read makes the answer unknown rather than unresolvable; an engine that runs elsewhere is never judged unresolvable and local lanes are never consulted for it (it does not start without a model name, so seeing it answer is enough to call it resolvable). `modelSetupDecision` combines that answer with whether this end can configure a model at all: setup is offered only for unresolvable on an end that can configure one, an end that cannot says so and points at whoever runs the engine, and unknown never opens setup. The detail and notice sentences are checked to be pairwise distinct, unknown sentences neither claim a model is configured nor that it is not, and the module is checked to import no platform I/O |
|
|
422
423
|
| `scripts/run-cloud-effective-projection-test.mjs` | The cloud control plane's effective-configuration response beyond its four configuration domains, and what a locally started engine does with the models document derived from it. Three top-level keys are read with the meaning their producer gives them: `warnings` (degradation warnings from the build that produced the served view — an empty list is a clean build, an absent key is an older server that cannot tell), and `budget` / `runtimeCaps`, where `null` means two different things: nothing resolves for this principal when you view yourself, and values withheld when the response previews another principal. A missing key or a wrong type is a third state, unknown, and none of the three is ever folded into a zero, a `false` or "no budget". A malformed warning row, budget field or cap costs only itself, and a known budget field of the wrong type is named as not shown rather than silently read as "no limit on this axis"; warning kinds are an open set, so a kind this client does not recognise still produces a warning line. The budget and cap readers are reconciled against the installed settings schema. The single wording source puts degradation warnings first and keeps every "not set / withheld / unknown" sentence free of digits, while a zero the server really sent is shown as a zero. On the models side, when the host injects an entry check the models document carries the default model, the tier groups and the active tier group (without the check none of the three is written), and every catalog reference the local engine's schema would reject — a default, role, @-mention entry, tier binding or active group that names something outside the catalog served to this principal — is dropped and recorded, because one dangling reference makes the local engine discard the whole models domain and fall back to its environment catalog; this is proven by reading the produced document with the installed file store. An @-mention allowlist that would be pruned to empty is kept as sent, since an empty list means "everything may be mentioned"; that case is recorded, produces its own warning that a locally started engine will reject the cloud model settings and use its environment catalog instead, and the gate reads the document with the installed file store to confirm exactly that outcome, so the sentence turns red the day the local reader becomes lenient. A per-model budget in which no field could be read is never described as having no limits. Registry annotation keys on model entries (`origin`, `overridesTeam`) are removed before the models document is written: the local engine's schema does not accept them, and a configuration refresh would otherwise be rejected as a whole. A model entry the host-injected entry check rejects is left out of the document and references to it are dropped: the local engine drops such an entry at startup, but a refresh rejects the whole configuration over it, so the gate requires a clean read of the produced document; when no entry passes the check, the catalog is kept as sent and gets its own warning, which the gate proves by reading the document back. The check receives a copy, so it cannot alter what is written. Tier words outside the local schema's closed set are dropped and recorded as unsupported, and the package's tier word list is reconciled against the installed schema in both directions; an active tier group is judged against group names, never model names. |
|
|
423
|
-
| `scripts/run-websearch-verdict-test.mjs` | The per-request web-search configuration a host puts on the wire. Newer servers refuse the whole request when that section is malformed — and a missing or misspelled search provider now counts as malformed, because the section names where the searches go and an unknown destination is refused rather than silently swapped for the deployment's own backend. The old readers in this package dropped such a section without a word, which let the server swap destinations after all. The guard pins the new three-way verdict (absent, honoured, malformed with the field that is wrong) against the server's own judge, vector by vector, whenever that judge is available next to this package; it pins that a half-configured environment is malformed rather than ignored, that a malformed environment never falls back to the settings file (that would change the destination too), and that neither the sentence shown to the user nor the recorded reason repeats an endpoint, a key, or a search-parameter name or value — an unrecognised provider is never echoed either (the sentence lists the valid words instead), so a URL or key pasted into the wrong field does not come back out, even when it happens to be all letters. A host's key store is plugged in through a callback the package calls only after the provider has been recognised, so the precedence between environment and settings stays inside the package. |
|
|
424
|
+
| `scripts/run-websearch-verdict-test.mjs` | The per-request web-search configuration a host puts on the wire. Newer servers refuse the whole request when that section is malformed — and a missing or misspelled search provider now counts as malformed, because the section names where the searches go and an unknown destination is refused rather than silently swapped for the deployment's own backend. The old readers in this package dropped such a section without a word, which let the server swap destinations after all. The guard pins the new three-way verdict (absent, honoured, malformed with the field that is wrong) against the server's own judge, vector by vector, whenever that judge is available next to this package; it pins that a half-configured environment is malformed rather than ignored, that a malformed environment never falls back to the settings file (that would change the destination too), and that neither the sentence shown to the user nor the recorded reason repeats an endpoint, a key, or a search-parameter name or value — an unrecognised provider is never echoed either (the sentence lists the valid words instead), so a URL or key pasted into the wrong field does not come back out, even when it happens to be all letters. A host's key store is plugged in through a callback the package calls only after the provider has been recognised, so the precedence between environment and settings stays inside the package. Since 0.83.0 an endpoint that carries a user name or password is malformed as well (the server judges the same way from 7.101.0), the reason sentences match the server's own word for word, and the three older readers that dropped a misspelled provider are gone. |
|
|
425
|
+
| `scripts/run-hooks-merged-disable-projection-test.mjs` | The fourth governance leg of the hooks projection: `disableAllHooks` set in a non-managed settings source. The value that counts is the **merged** one, read from the host through the optional `SettingsPort.mergedDisableAllHooks()`, because a per-source approximation ("any source says true") reads user `true` with local `false` backwards — the merged value there is `false` and every source's hooks ship. When the merged value is `true`, only managed-settings hooks are sent to the engine: non-managed settings can switch off their own hooks, never the managed ones, and a managed `disableAllHooks` is still judged first and sends nothing at all. A full matrix over the four sources, each true, false or absent, is merged with the reference rule (later sources override earlier ones, managed settings last) and every cell's projection is asserted. The reader is the only authority: per-source values never second-guess it, and only a strict `true` counts. A host that does not implement it keeps the previous behaviour and gets exactly one warning per installed settings port, never one per request, and none on paths where the reader would not have been consulted; a reader that throws is treated as `true`, so managed hooks still ship. The session goal's Stop hook and the final-verification yield rule, which reads the projected hooks, follow the same verdict, and the trust gate and the three managed gates are evaluated before the reader is ever called. The last leg pins the member's declared shape in the built declarations: optional, no parameters, returning a boolean or `undefined`. |
|
|
426
|
+
| `scripts/run-memory-saved-projection-test.mjs` | Engine memory writes (a successful `Remember` tool call) moved off the transcript onto the additive `memory_saved` chrome event, driven through the real pipeline: zero transcript rows for the write (the transcript is byte-identical to the same frames with the write reported as not successful), exactly one event whose `notes` carry the note text verbatim (notes, not file paths) and whose key set is exactly kind / laneProof / id / notes; no event for a missing, empty or non-string note, a non-`true` `ok`, a tool error, a missing result or another tool name; two writes give two events in order with distinct ids; a sub-agent write rides the sub-agent lane and an empty parent id emits nothing rather than falling back to the main lane; the `id` is derived from the write's wire key (the tool-end event id, else the tool-start event id, else the call id; empty ids count as absent), so projecting the same wire events twice gives the same id, and it never collides with the tool result row of the same or another call; the event sits right after the tool result row; the arm is registered as required. |
|
|
427
|
+
| `scripts/run-result-frame-projection-test.mjs` | Result frames and the synthesized terminal rows. The CC key `terminal_reason` is minted on result frames only where it follows from what the engine reported: `completed` on success, `max_turns`, `budget_exhausted` and `structured_output_retry_exhausted` for the three matching engine codes, on both the done-frame path and the failed-event path. Every other outcome leaves the key absent as an own property rather than present with an undefined value: wall-clock and token-budget limits, the classifier denial limit, cancellation, unknown codes, blocked, paused, unreadable or missing terminal records, and the busy-session refusal. The public reader `terminalReasonForResult` shares the minting predicate and is checked to agree with the minted key on every frame the gate produces. Both the minted key and the reader derive the word from the frame's CC subtype (success with `is_error` strictly false, and the three limit subtypes), not from the error code, so a replayed row whose status is paused, blocked or unrecognised never carries a word that contradicts its subtype. The four words are checked against the mirrored CC union, and the minting file is checked to hold no hand-copied code literals. The renamed superset keys (`_sema_error_code`, `_sema_salvaged_result`, `_sema_model_degraded`, `_sema_selected_model`, and the row flag `_sema_api_error_message`) are driven through the real stream pipeline. Each must be present under its new name, the old name must be absent, and every frame the gate saw is swept for old names. The selected model appears on error envelopes whenever the terminal record carries it, and never on a failed event, which has no record. It stays separate from the provider-reported model name. The two in-package readers still work: the interactive result arm reads the salvaged text under its new name (and old-shape frames under the old one), and the print init gate treats the renamed flag as the run having ended. |
|
|
428
|
+
| `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
|
|
424
429
|
|
|
425
430
|
Each suite carries a floor that only moves up — a refactor that stops executing a group of
|
|
426
431
|
assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
|
package/dist/adapt/arms.d.ts
CHANGED
|
@@ -5,7 +5,7 @@ import type { PanelTaskLedger } from './panelTasks.js';
|
|
|
5
5
|
import type { TextStream } from './textStream.js';
|
|
6
6
|
import type { ToolCardLedger } from './toolCards.js';
|
|
7
7
|
import type { TurnFlags } from './turnFlags.js';
|
|
8
|
-
|
|
8
|
+
interface ProjectionArmDeps {
|
|
9
9
|
readonly ctx: AdapterContext;
|
|
10
10
|
readonly idOf: IdOf;
|
|
11
11
|
}
|
|
@@ -17,5 +17,6 @@ export interface ArmDeps extends ProjectionArmDeps {
|
|
|
17
17
|
readonly inst: AdapterInstanceLedger;
|
|
18
18
|
}
|
|
19
19
|
export type ProjectionArmFn = (m: Frame, deps: ProjectionArmDeps) => Generator<AdapterOutput>;
|
|
20
|
-
|
|
20
|
+
type ArmFn = (m: Frame, deps: ArmDeps) => Generator<AdapterOutput>;
|
|
21
21
|
export declare const ARMS: ReadonlyMap<string, ArmFn>;
|
|
22
|
+
export {};
|
package/dist/adapt/arms.js
CHANGED
|
@@ -140,7 +140,7 @@ const userArm = function* (m, { ctx, idOf }) {
|
|
|
140
140
|
};
|
|
141
141
|
const systemArm = function* (m, { ctx, idOf }) {
|
|
142
142
|
yield transcript(m, ctx.now());
|
|
143
|
-
const attachedFiles = m.attachedFiles;
|
|
143
|
+
const attachedFiles = m._sema_attached_files ?? m.attachedFiles;
|
|
144
144
|
if (Array.isArray(attachedFiles)) {
|
|
145
145
|
let i = 0;
|
|
146
146
|
for (const f of attachedFiles) {
|
|
@@ -491,7 +491,7 @@ const promptSuggestionsArm = function* (m) {
|
|
|
491
491
|
yield chrome({ kind: 'prompt_suggestions', laneProof: mainLane(), suggestions: list });
|
|
492
492
|
}
|
|
493
493
|
};
|
|
494
|
-
const toolEndResultArm = function* (m, {
|
|
494
|
+
const toolEndResultArm = function* (m, { idOf, cards, panel }) {
|
|
495
495
|
const callId = typeof m.toolCallId === 'string' ? m.toolCallId : undefined;
|
|
496
496
|
recordEngineToolLabel(callId, m.label);
|
|
497
497
|
if (callId !== undefined) {
|
|
@@ -526,14 +526,18 @@ const toolEndResultArm = function* (m, { ctx, idOf, cards, panel }) {
|
|
|
526
526
|
: undefined;
|
|
527
527
|
if (p.name === 'Remember' && m.isError !== true) {
|
|
528
528
|
if (s && s.ok === true && typeof s.note === 'string' && s.note.length > 0) {
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
|
|
532
|
-
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
529
|
+
const memParent = closeParent ?? p.parentToolCallId;
|
|
530
|
+
if (memParent !== '') {
|
|
531
|
+
yield chrome({
|
|
532
|
+
kind: 'memory_saved',
|
|
533
|
+
laneProof: memParent === undefined ? mainLane() : { lane: 'subagent', parentToolCallId: memParent },
|
|
534
|
+
id: (() => {
|
|
535
|
+
const key = [closeEventId, p.eventId].find((v) => typeof v === 'string' && v.length > 0);
|
|
536
|
+
return idOf(key !== undefined ? { id: key } : { toolCallId: p.id }, 'memory-saved');
|
|
537
|
+
})(),
|
|
538
|
+
notes: [s.note],
|
|
539
|
+
});
|
|
540
|
+
}
|
|
537
541
|
}
|
|
538
542
|
}
|
|
539
543
|
if (s && (s.type === 'task' || s.type === 'task-list')) {
|
|
@@ -692,7 +696,8 @@ const resultArm = function* (m, { ctx, idOf, text, cards, panel, flags }) {
|
|
|
692
696
|
return lowerBound && ot === 0 ? undefined : ot;
|
|
693
697
|
});
|
|
694
698
|
yield* text.takeAnswerSegment();
|
|
695
|
-
const
|
|
699
|
+
const salvaged = m._sema_salvaged_result;
|
|
700
|
+
const terminalText = typeof salvaged === 'string' ? salvaged : typeof m.result === 'string' ? m.result : '';
|
|
696
701
|
const committed = text.lastCommittedAnswerText;
|
|
697
702
|
const terminalIdentity = messageIdentityOf(m, ctx);
|
|
698
703
|
const emitTerminal = function* (body, tag) {
|
package/dist/adapt/ids.js
CHANGED
|
@@ -14,12 +14,16 @@ export function messageIdentityOf(frame, ctx) {
|
|
|
14
14
|
};
|
|
15
15
|
}
|
|
16
16
|
export const chrome = (event) => ({ plane: 'chrome', event });
|
|
17
|
-
export const transcript = (message, nowMs) =>
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
:
|
|
22
|
-
|
|
17
|
+
export const transcript = (message, nowMs) => {
|
|
18
|
+
const arm = message.type;
|
|
19
|
+
const stampsArrival = arm === 'user' || arm === 'assistant';
|
|
20
|
+
return {
|
|
21
|
+
plane: 'transcript',
|
|
22
|
+
message: (!stampsArrival || typeof message.timestamp === 'string'
|
|
23
|
+
? message
|
|
24
|
+
: { ...message, timestamp: new Date(nowMs).toISOString() }),
|
|
25
|
+
};
|
|
26
|
+
};
|
|
23
27
|
export function makeIdOf(ctx) {
|
|
24
28
|
return (frame, suffix) => {
|
|
25
29
|
const base = typeof frame.id === 'string' && frame.id.length > 0
|
|
@@ -28,9 +32,9 @@ export function makeIdOf(ctx) {
|
|
|
28
32
|
? frame.uuid
|
|
29
33
|
: undefined;
|
|
30
34
|
return deriveTranscriptId({
|
|
31
|
-
...(base !== undefined ? { id:
|
|
35
|
+
...(base !== undefined ? { id: base } : {}),
|
|
32
36
|
...(typeof frame.seq === 'number' ? { seq: frame.seq } : {}),
|
|
33
37
|
...(typeof frame.toolCallId === 'string' ? { toolCallId: frame.toolCallId } : {}),
|
|
34
|
-
}, ctx);
|
|
38
|
+
}, ctx, suffix);
|
|
35
39
|
};
|
|
36
40
|
}
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
import type { AdapterContext, AdapterOutput, TextSegmentCommittedRow } from '../seam.js';
|
|
2
2
|
import { type Frame, type IdOf } from './ids.js';
|
|
3
3
|
export declare const SEMA_SEGMENT_ID_KEY = "_sema_segment_id";
|
|
4
|
-
|
|
4
|
+
interface TextSegmentReplacement {
|
|
5
5
|
diverged: boolean;
|
|
6
6
|
committedUuids?: readonly string[];
|
|
7
7
|
committedUuid?: string;
|
|
@@ -9,7 +9,7 @@ export interface TextSegmentReplacement {
|
|
|
9
9
|
committedPrefixLen: number;
|
|
10
10
|
committedPrefixDiverged: boolean;
|
|
11
11
|
}
|
|
12
|
-
|
|
12
|
+
interface ThinkingSegmentReplacement {
|
|
13
13
|
diverged: boolean;
|
|
14
14
|
committedPrefixLen: number;
|
|
15
15
|
committedPrefixDiverged: boolean;
|
|
@@ -40,3 +40,4 @@ export interface TextStream {
|
|
|
40
40
|
readonly answerLength: number;
|
|
41
41
|
}
|
|
42
42
|
export declare function createTextStream(ctx: AdapterContext, idOf: IdOf): TextStream;
|
|
43
|
+
export {};
|
|
@@ -10,18 +10,17 @@ export type ToolResultOwner = 'leader' | 'subagent';
|
|
|
10
10
|
export declare function toolResultOutcomeOf(toolName: string, end: ToolEndPayload | undefined, rawInput: PendingToolUse['rawInput'], owner: ToolResultOwner, ledgerTier?: (isError: boolean) => RichToolUseResult | null): ToolResultOutcome;
|
|
11
11
|
export interface ToolResultEnvelopeKeys {
|
|
12
12
|
readonly tool_use_result?: ToolResultBody['toolUseResult'];
|
|
13
|
-
readonly toolUseResult?: ToolResultBody['toolUseResult'];
|
|
14
13
|
readonly _sema_degraded?: ToolResultBody['degraded'];
|
|
15
14
|
}
|
|
16
15
|
export declare function toolResultEnvelopeParts(o: ToolResultOutcome): ToolResultEnvelopeKeys;
|
|
17
|
-
|
|
16
|
+
interface PendingToolUse {
|
|
18
17
|
id: string;
|
|
19
18
|
name: string;
|
|
20
19
|
rawInput: unknown;
|
|
21
20
|
eventId?: string;
|
|
22
21
|
parentToolCallId?: string;
|
|
23
22
|
}
|
|
24
|
-
|
|
23
|
+
interface ToolEndPayload {
|
|
25
24
|
output?: unknown;
|
|
26
25
|
structured?: unknown;
|
|
27
26
|
truncated?: boolean;
|
package/dist/adapt/toolCards.js
CHANGED
|
@@ -4,7 +4,7 @@ import { chrome, mainLane, transcript } from './ids.js';
|
|
|
4
4
|
import { flattenWireOutput } from './wireShapes.js';
|
|
5
5
|
import { isCcToolDenialKind } from '../gateVocabulary.js';
|
|
6
6
|
function toolUseResultParts(v) {
|
|
7
|
-
return v !== undefined ? { tool_use_result: v
|
|
7
|
+
return v !== undefined ? { tool_use_result: v } : {};
|
|
8
8
|
}
|
|
9
9
|
export function toolResultOutcomeOf(toolName, end, rawInput, owner, ledgerTier) {
|
|
10
10
|
const isError = end?.isError ?? false;
|
package/dist/adapt.d.ts
CHANGED
|
@@ -14,7 +14,7 @@ export declare const ADAPTER_COVERAGE: {
|
|
|
14
14
|
readonly b5: readonly ["B 层 structuredToToolUseResult 14 case(ask-user-question/bash/edit/create/update/text/grep/glob/notebook-edit/cron-create/cron-delete/cron-list/agent/task-output)+ workflow-run 登记副作用", "D 层 wireOutputToBody 3 臂(bash / taskoutput / 泛化);structured 白名单在场时 T11+T14 正则退位,缺席保回落", "E 层 4 臂:#117a bg Bash 回执 + #158 Monitor 回执(同臂按工具名分,registration.source 判别;判定在库→chrome bgshell_register)· TodoWrite oldTodos 富卡 · ReportFindings 按名认领 · Remember→memory_saved 系统消息", "tool_end_result → user tool_result 正常路径(卡本体);result 与 turn 末两处开卡兜底关闭(abort 不兜底,cli 同)", "task/task-list structured → chrome task_ledger_sync{source:\"structured\"}(判定归一在库、落盘留宿主)", "T7 诚实缺席(§8-4:不搬 mock 合成器,改 degraded 标记)· T20 diff hunks 进包(src/diff/patch.ts)"];
|
|
15
15
|
readonly todo: readonly ["T13 的 **D 层** 半场(taskoutput block content)未随 structured 退位:退位会改 block content 形状(cli 那里是裸对象 = 件1 隐患),属行为面改动,无真实需求驱动 ⇒ 见 DIVERGENCE-8", "macrotask 让渡(cli 在提交前让出宏任务给 Ink;宿主渲染策略,不进库)"];
|
|
16
16
|
};
|
|
17
|
-
export declare const ADAPTER_DIVERGENCES: readonly [string, string, string, string, string, string, string, string, string, string];
|
|
17
|
+
export declare const ADAPTER_DIVERGENCES: readonly [string, string, string, string, string, string, string, string, string, string, string, string];
|
|
18
18
|
export interface WireToCcAdapterWithLedger extends WireToCcAdapter {
|
|
19
19
|
exportLedger(): AdapterLedgerState & Record<string, unknown>;
|
|
20
20
|
importLedger(state: Partial<AdapterLedgerState> | Record<string, unknown>): number;
|
package/dist/adapt.js
CHANGED
|
@@ -69,7 +69,7 @@ export const ADAPTER_DIVERGENCES = [
|
|
|
69
69
|
'DIVERGENCE-1:`user` 帧透传——cli 的 switch 无 user 臂(default 丢弃),adapt 原样透传。' +
|
|
70
70
|
'现役 wire 不产 user 帧(eventToSdkMessage 无该出口),故对真流是惰性的;保留是为了 desktop 重放存量转录。',
|
|
71
71
|
'DIVERGENCE-2:transcript uuid 取值形——cli 铸随机 v4 / 原样透传帧 uuid,adapt 一律走 ' +
|
|
72
|
-
'deriveTranscriptId(
|
|
72
|
+
'deriveTranscriptId(由稳定键确定性派生的 UUID 形,同一稳定键恒得同一值)。差分守卫按「别名结构」比对(同一侧内的复用/区分关系),不比字面值。',
|
|
73
73
|
'DIVERGENCE-3:合并节流的**分片边界**不比对——cli 用 Date.now() 真时钟,adapt 用 ctx.now();' +
|
|
74
74
|
'且 B3 起节拍可由宿主 `ctx.coalesceIntervalMs` 覆盖(缺席=cli 同值 100ms)。' +
|
|
75
75
|
'守卫比对的是 chrome 增量按 channel 拼接后的全文与顺序(语义不变量),不是 chunk 切分。',
|
|
@@ -101,11 +101,18 @@ export const ADAPTER_DIVERGENCES = [
|
|
|
101
101
|
'(框架里的尾换行等是渲染副产物)。T13/T24 的 **B 层**半场同款惰性退位(structured 带齐 ' +
|
|
102
102
|
'status+content 就根本不解模型面文本);T13 的 **D 层** block-content 半场本批**未**退位 —— ' +
|
|
103
103
|
'那一步会改 block content 形状(cli 在这里是裸对象,件1 隐患),无真实需求驱动,如实留白。',
|
|
104
|
-
'DIVERGENCE-9(B5):`tool_result` 消息的 uuid —— cli 在缺 `eventId` 时铸随机 v4
|
|
105
|
-
'
|
|
106
|
-
'DIVERGENCE-10(0.81.0
|
|
107
|
-
'CC 内部转录面驼峰 `toolUseResult`
|
|
108
|
-
'
|
|
104
|
+
'DIVERGENCE-9(B5):`tool_result` 消息的 uuid —— cli 在缺 `eventId` 时铸随机 v4,本包由 toolUseId ' +
|
|
105
|
+
'确定性派生 UUID 形(DIVERGENCE-2 同族:同流重放 ⇒ 同 id 序列,[1653])。守卫按别名结构比,不比字面。',
|
|
106
|
+
'DIVERGENCE-10(0.81.0 起,0.83.0 定形,CC-131):`tool_result` 消息上本包只铸 SDK 面键 `tool_use_result` ' +
|
|
107
|
+
'(0.81.0–0.82.x 曾与 CC 内部转录面驼峰 `toolUseResult` 同引用并铸,0.83.0 删驼峰);cli 侧仍带驼峰 —— 所钉的本包' +
|
|
108
|
+
'旧版本两键同铸,升级后在转录入口从 `tool_use_result` 同值起别名。守卫在两键同在场且同引用时略过驼峰键、' +
|
|
109
|
+
'按 snake 比(两侧对称);cli 侧不再带驼峰后本条自然零触发 ⇒ 自退役。',
|
|
110
|
+
'DIVERGENCE-11(0.83.0,C-R103):system 转录行的到达时戳 —— 本包只给 user / assistant 行补 `timestamp`' +
|
|
111
|
+
'(agent-types 只在这两臂声明),system 行不补(入参自带的照旧原样过境);壳的接收口给每条转录行补到达时刻(只补缺席)。' +
|
|
112
|
+
'守卫在逐字比对前只从壳侧 system 行上去掉这一位;本包侧 system 行若带 `timestamp`,照样判不等。',
|
|
113
|
+
'DIVERGENCE-12(0.83.0,C-R103):记忆写入 —— 本包发 chrome `memory_saved` 臂(`notes` = 笔记原文,`id` 由这次写入的 wire ' +
|
|
114
|
+
'稳定键派生),转录面零行;壳侧是一条 system/memory_saved 转录行(`writtenPaths` 装同一份笔记)。守卫把壳侧这类行移出逐字比对,' +
|
|
115
|
+
'改比:事件 `notes` ⇔ 行 `writtenPaths` 逐项相等、位置相等(之前已覆盖转录行的条数);`id` 不与旧行 uuid 比(本版不承诺相等)。',
|
|
109
116
|
];
|
|
110
117
|
class WireToCcAdapterImpl {
|
|
111
118
|
ledger = createAdapterInstanceLedger();
|
|
@@ -631,35 +631,52 @@ export function activeRunSelfHealRow(outcome, signal, copy, origin, follow) {
|
|
|
631
631
|
return base + governanceOriginClause(signal);
|
|
632
632
|
}
|
|
633
633
|
}
|
|
634
|
+
const INJECTED_LEAD = 'A follow-up message sema sent on its own (not one you typed)';
|
|
634
635
|
function injectedSubmissionRow(outcome, following = false) {
|
|
635
|
-
const
|
|
636
|
+
const lead = INJECTED_LEAD;
|
|
637
|
+
const id = 'taskId' in outcome && outcome.taskId ? ` (id ${outcome.taskId})` : '';
|
|
636
638
|
if (outcome.kind === 'running-steer-failed') {
|
|
637
639
|
return outcome.delivery === 'unknown'
|
|
638
|
-
?
|
|
639
|
-
`(${outcome.detail}) — that
|
|
640
|
-
`
|
|
641
|
-
:
|
|
642
|
-
`
|
|
640
|
+
? `${lead} was passed to the reply already in progress${id}, but its delivery could not be confirmed ` +
|
|
641
|
+
`(${outcome.detail}) — that reply may or may not have received it. sema did not retry, because sending it twice ` +
|
|
642
|
+
`is not safe; watch that reply for what it does next.`
|
|
643
|
+
: `${lead} was not delivered: passing it to the reply already in progress${id} was refused (${outcome.detail}). ` +
|
|
644
|
+
`sema did not retry, so the model has not seen it.`;
|
|
643
645
|
}
|
|
644
646
|
if ((outcome.kind === 'ask-reopen-failed' || outcome.kind === 'plan-review-reopen-failed') && outcome.decisionInFlight === true) {
|
|
645
|
-
return (
|
|
646
|
-
`is still on its way
|
|
647
|
+
return (`${lead} was not delivered: this session is still busy with an earlier reply${id}, and your answer to the decision ` +
|
|
648
|
+
`it is waiting on is still on its way. sema did not retry, so the model has not seen it.`);
|
|
649
|
+
}
|
|
650
|
+
if (outcome.kind === 'running-steered' && outcome.delivery === 'queued') {
|
|
651
|
+
const queuedOn = `${lead} is queued on an earlier reply${id} that is paused waiting for a decision`;
|
|
652
|
+
if (outcome.reopened !== null && reopenDelivered(outcome.reopened)) {
|
|
653
|
+
return `${queuedOn}; sema has shown that decision card, and the message is picked up once you answer it.`;
|
|
654
|
+
}
|
|
655
|
+
if (reopenRefusedForDecisionInFlight(outcome.reopened)) {
|
|
656
|
+
return `${queuedOn}; your answer to it is still on its way, and the message is picked up once that answer is applied.`;
|
|
657
|
+
}
|
|
658
|
+
return `${queuedOn} sema could not show here; the message is picked up only once that decision is made.`;
|
|
659
|
+
}
|
|
660
|
+
if (outcome.kind === 'running-steered' && outcome.delivery === 'parked_for_wake') {
|
|
661
|
+
return (`${lead} has not reached the model: the earlier reply${id} had already finished, so the message was set aside on it ` +
|
|
662
|
+
`instead of starting a new reply, and sema did not retry it.`);
|
|
647
663
|
}
|
|
648
664
|
switch (selfHealSubmissionDisposition(outcome)) {
|
|
649
665
|
case 'held-for-decision':
|
|
650
|
-
return (
|
|
651
|
-
`
|
|
666
|
+
return (`${lead} is on hold: this session is busy with an earlier reply${id} that is waiting on a decision. ` +
|
|
667
|
+
`sema kept the message queued and will send it after you answer the open card.`);
|
|
652
668
|
case 'resending':
|
|
653
|
-
return
|
|
669
|
+
return (`${lead} was held back because this session was busy with an earlier reply${id}; the session is free again, ` +
|
|
670
|
+
`so sema is sending the message now — nothing for you to do.`);
|
|
654
671
|
case 'handed-off':
|
|
655
672
|
return following
|
|
656
|
-
?
|
|
657
|
-
`
|
|
658
|
-
:
|
|
659
|
-
`
|
|
673
|
+
? `${lead} was passed to the reply already in progress${id} instead of starting a new one. ` +
|
|
674
|
+
`sema is following that reply and will show whatever it asks for next; nothing for you to do now.`
|
|
675
|
+
: `${lead} was passed to the reply already in progress${id} instead of starting a new one — ` +
|
|
676
|
+
`watch that reply for what it does with it.`;
|
|
660
677
|
case 'not-delivered':
|
|
661
|
-
return (
|
|
662
|
-
`
|
|
678
|
+
return (`${lead} was not delivered: this session was still busy with an earlier reply${id}, and there was no open card ` +
|
|
679
|
+
`you could answer to free it. sema did not retry, so the model has not seen it.`);
|
|
663
680
|
}
|
|
664
681
|
}
|
|
665
682
|
function decisionInFlightRow(parkedOn) {
|
|
@@ -161,7 +161,7 @@ export function eventToSdkMessage(ev, ctx) {
|
|
|
161
161
|
: {}),
|
|
162
162
|
...(freedTokens !== undefined ? { _sema_freed_tokens: freedTokens } : {}),
|
|
163
163
|
},
|
|
164
|
-
...(Array.isArray(attachedFiles) && attachedFiles.length > 0 ? { attachedFiles } : {}),
|
|
164
|
+
...(Array.isArray(attachedFiles) && attachedFiles.length > 0 ? { _sema_attached_files: attachedFiles } : {}),
|
|
165
165
|
})));
|
|
166
166
|
}
|
|
167
167
|
case 'tool_end': {
|