@sema-agent/client-core 0.83.0 → 0.83.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -49,6 +49,48 @@
49
49
  > 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
50
50
  > commit 漏转在发布前就红,不再靠人记。
51
51
 
52
+ ## 0.83.1(2026-09-26)
53
+
54
+ > 主题:patch —— 三件公共逻辑归包,只增公面、另有几处行为订正。① 展示层安全出口 `displayUntrusted` 三端单源,本包合成的终局行与结果帧 `errors[]` 在铸点洗掉凭据(CC-112 / CC-187);② 会话规则记录的无损判定 `sessionPolicyDeliverable`(CC-117);③ 插件 hook 逐条判「投给引擎 / 本客户端执行 / 如实不跑」,开机「不会执行」清单与 `/hooks` 标注措辞单源,同批不再把带参数的设置来源 command hook 条目送上请求、传了计划时也不再送 `mcp_tool` 条目(CC-174)。根公面运行期导出 1238 → **1254**(+16),公面类型 +26,`SettingsPort` +1 可选成员,`hooksForWire` +1 可选参;peer sdk 地板 `>=11.3.0` 不动。
55
+
56
+ ### Added
57
+
58
+ - **`displayUntrusted(text, opts?)`**(CC-112):wire 派生文本上屏前的合成出口,三端一只。五张网默认全开 —— 凭据 URL 结构面(userinfo 整段、query 的每个值、fragment、路径段参数值、以已知凭据前缀打头的路径段)、凭据词级面(`Authorization: Bearer <值>`、`Authorization: token <值>` 一类非 bearer / basic 方案(方案词留、凭据段换记号;参数表形方案 —— Digest、签名算法形、OAuth 1.0 —— 参数名一律保留,`response` / `signature` / `oauth_signature` 的值换记号)、`api_key=<值>` / JSON 引号形 / `OPENAI_API_KEY=<值>` 这类键值对的值,分隔符两侧的空白不设上限;散文里不带标签的已知凭据字面形 —— `sk-…` / `ghp_…` / `AKIA…` / JWT 三段形 —— 整词换记号(随机段不足 16 位的 `sk-video` 这类名字不算);PEM 私钥块(含 PGP 私钥块)的主体换成一枚记号,头尾两行与它们旁边的换行保留,没有 END 行时遮到块尾,凭据标签后面紧跟私钥块时同样整块处理;诊断词与 `max_tokens` 一族计量单位不洗,方案词后跟的是一张窄词表里的常见英文词(`Invalid bearer token`;表外的词照遮)、反引号里点名的凭据变量名、裸复数 `tokens:` 后的纯整数计数、裸复数 `keys:` 后的键名清单也不洗)、控制符(C0 / DEL / C1 / 孤代理项)、双向与格式字符(格式类整类,ZWNJ / ZWJ 除外;含行 / 段分隔符)、行折平;顺序为凭据 → 字符 →(字符面改动过字节时)凭据 → 字符 … 到不动点。凭据两网在一份「判别视图」上匹配:ANSI 转义序列(按 ECMA-48 整族认:CSI、`ESC ( B` 一类字符集指定、`ESC 7` / `ESC M` 一类单字功能、已结尾的 OSC / DCS 控制串)、格式字符、控制符(以及本出口自己铸的 `\uXXXX` 转义)夹在标签、分隔符与值之间照样认得出,遮盖映回原文 —— 序列留在记号两侧,不会留下半截转义加记号的误导形;已结尾控制串的内容(窗口标题、超链接地址)另作一段文本洗;序列的最后一个字符恰是凭据词的首字母时(`<ESC>Bearer <值>`),按「终端吞掉这个字符」与「这个字符是词首」两种读法各判一遍,任一种认出就遮。整段文本线性处理(十万字符量级的标签密排 / 控制串引导符密排文本在百毫秒量级完成)。RGI 地区旗(黑旗加标签字符的整串)原样保留。参数表形方案的参数名与等号之间、等号与值之间带空白(`username = "bob"`)照认,头值折行续写照认(Negotiate / NTLM / Basic / Bearer 一类的首段照旧整段换记号);无 scheme 的 `user:pass@host` 与凭据标签后的转义反斜杠都不设长度上限(此前口令超过 256 字符 —— 例如把 JWT 当口令 —— 或 JSON 转义套三层以上时整段认不出、原样上屏)。选项:五张网各自可关(`credentialUrls` / `credentialWords` / `controls` / `bidi` / `foldLines`)、`keepLayout`(多行正文保 `\t` `\n`)、`mark`(危险字符呈现为可见转义 `\uXXXX`〔astral 为 `\u{…}`〕/ 点 / 空格,默认可见转义)、`max`(按输出封长,截点不劈转义记号、不劈代理对;不是有限数时不封)。幂等(带 `max` 时同样);干净文本原样返回(默认形开着行折平:换行、连续空白与首尾空白仍会折平);只管呈现、不参与任何判定。新增公面类型 `DisplayUntrustedOptions` / `DisplayMark`。接入文档 **§101a D-1 / 101a-1**。
59
+ - **会话规则记录的无损判定 `sessionPolicyDeliverable(behavior, rules)`**(CC-117):给一批同一 behavior 的规则串,答「能不能原样写进这条会话的规则记录」。只有一类能无损对上 —— 不带限定、恰好是一个工具名的 deny,写进 `toolDeny`(逐字、序不变、重复保留);其余每一条带一个成因词,闭集 `SESSION_POLICY_WITHHELD_WHY` 五词:`ask_no_bucket` / `qualified_deny` / `peer_wide` / `wildcard` / `not_a_tool_name`。🔴 **整批可送才送**:批里只要有一条送不了,`deliverable` 就是 `{}`,绝不给子集。坏形入参(表外 behavior、不是数组、读的时候抛错)不抛、一条都不送。每个成因词一句用户面话,单源 `sessionPolicyWithheldNotice(why)`,只收成因词、结构上回显不了规则串。判据与整批语义与此前宿主侧那一份相同,只有三处更严(都是此前判能写、而会话规则记录按原字节比一条都拦不住的形,本版扣下):① 首尾带空白的名字(如 ` Read`,`not_a_tool_name`);② 串里任何位置带 `*` 的规则(如 `*` / `Bash*` / `Web*` / `mcp__*__get_user` / `mcp__srv__get_*`,此前只扣下 MCP 工具段恰为 `*` 那一形;`wildcard`)—— 带 `*` 的规则在 CC 的规则语义里是通配(终端现行的匹配器只认 MCP 工具段恰为 `*` 那一形),而会话规则记录按精确名比、一条都拦不住;没有哪只工具的名字带 `*`,扣下的代价只是晚一拍生效。带括号限定的规则(如 `Bash(git push:*)`)仍报 `qualified_deny`:括号里的 `*` 是参数样式,不是工具名通配。③ 覆盖一个服务器或对端全部工具的规则不只认 MCP:引擎认得的每个协议命名空间(今天是 `mcp__` 与 `a2a__`)的 `<命名空间>__<服务器或对端>` 一律扣下(`peer_wide`)—— 此前 `a2a__payroll` 这类被判能写,而引擎把它当作覆盖该对端全部工具的规则,会话规则记录按精确名比一条都拦不住。成因词的用户面话 `sessionPolicyWithheldNotice(why)` 对任何入参都不抛(包括转成字符串时会抛错的值),表外的值一律回通用句。整批读不懂的入参(表外 behavior、不是数组、读的时候抛错)结果带 `unreadable: true`,与空批可分。新增公面类型 `SessionPolicyDeliverability` / `SessionPolicyWithheldRule` / `SessionPolicyWithheldWhy`。接入文档 **§101a P-1 / P-2 / 101a-3**。
60
+ - **插件 hook 每轮计划 `hooksWirePlan(opts?)`**(CC-174):设置来源投影 + 每一只插件 hook 的判定(`engine` / `shell` / `not_run` 加原因)+ 被治理筛掉的插件 hook + 被拿掉的设置来源 hook + 这一轮请求体 `settings.hooks` 的确切内容(`plan.wire`)。入参是宿主报的三条读数:`baseUrl`(读引擎的 `taskSettings.pluginHooks` 能力位)、`engineOwnedByThisShell`(引擎是不是本机由本客户端自起的)、`shellHookEvents`(本客户端自己的本地执行器为插件 hook 触发哪些事件)。宿主没报 `engineOwnedByThisShell`、或引擎能力还没探到时,判定原因与计划上的引擎事实都如实写「不知道」(原因 `engine_locality_unknown` / `engine_capability_unknown`,`plan.engine` 上对应位为 `null`),不说成「不是本机起的」「不支持」。传 `null` 与不传同处置。宿主交来的插件数据在读的时候抛错(会抛的 getter、已撤销的代理)与读口抛错同处置:这一轮 `pluginReader: 'failed'`、零插件判定与插件条目,设置来源照发,不外抛。🔴 每轮算一次、同一份用到底:先按它决定本地跳过哪些插件 hook,再把同一份交给 `hooksForWire({ plan })` 组请求体。接入文档 **§101a H-2 / 101a-4 / 101a-5**。
61
+ - **`SettingsPort.enabledPluginHooks?(): PluginHooksReading`**(可选,CC-174):已启用插件的 hooks(`hooks/hooks.json` 与 manifest `hooks` 按宿主加载后的形)、插件 id / 名 / 根目录 / 数据目录 / 是否由 managed 设置启用 / 声明的选项(敏感选项**只报声明,型面上没有值位**),加安全·bare 模式位。不实现 ⇒ 插件条目照旧不投,每个已装的端口告警一次。接入文档 **§101a H-1 / 101b**。
62
+ - **`hooksForWire(opts?)` 新增可选 `{ plan }`**:给了就原样返回 `plan.wire`(同一个对象);不给时不读插件口、不投插件条目。传 `null`(或 `{ plan: null }`)与不传同处置,不抛;`plan` 位上不是计划的值(空对象、数、串等)按不传处置 —— 设置来源照发,不会整份变空;把计划直接当第一个参数传入(`hooksForWire(plan)`)认得出,按 `{ plan }` 处置。接入文档 **§101a H-2**。
63
+ - **插件 hook 的查询口、措辞单源与闭集**(CC-174):查询口 `pluginHookVerdictOf(plan, pluginId, event, groupIndex, hookIndex)`(本地跳过同一只用);开机一次性「不会执行」清单 `hooksNotRunNotice(plan)`(按插件 × 原因一行,只含插件名、事件名与固定短句;名字过 `displayUntrusted` 的字符面,零宽 / 标签字符 / 软连字符渲成可见转义,引号转义成 `\"`,超过 120 个字符截断加省略号,转义与封长单遍完成)、`/hooks` 执行方标注 `pluginHookExecutorLabel(verdict)`、原因短句 `pluginHookReasonText(reason)` / `settingsHookDropText(reason)`;事件判定口 `isEngineFiredHookEvent(event)`;闭集 `ENGINE_FIRED_HOOK_EVENTS` / `PLUGIN_HOOK_DISPOSITIONS` / `PLUGIN_HOOK_REASONS` / `PLUGIN_HOOK_EXCLUSION_REASONS` / `SETTINGS_HOOK_DROP_REASONS` 与对应类型。判定原因闭集 `PLUGIN_HOOK_REASONS` 共 13 词,其中「引擎不是本机由本客户端起的」(`engine_not_local`)与「引擎不支持插件 hook」(`engine_capability_absent`)只在读数确实这么说时用;宿主没报 / 能力没探到走「不知道」两词(`engine_locality_unknown` / `engine_capability_unknown`)。接入文档 **§101a H-3 / H-4 / H-8 / 101a-7**。
64
+ - **插件 hook 投给引擎的投影形**(CC-174):本机自起的引擎报 `taskSettings.pluginHooks === true`、条目是 command、插件没声明敏感选项、并进去不超引擎每事件上限时,条目带 `plugin: { name, root, dataDir, options? }` 投出(`options` 只带已存的非敏感值)。🔴 截至本版,引擎尚未报出这一能力位,`plugin` / `args` 两键在引擎契约上也还没有定形 ⇒ 实际零投;判定、告知与本地跳过今天就生效。能力位与契约形在引擎侧同版出现时,本包在那一版核对键名与值形之后再放开;形与这里不同则按引擎的改。接入文档 **§101a-6**。
65
+
66
+ ### Changed
67
+
68
+ - 🔴 **结果帧错误信封 `errors[]` 洗掉凭据 —— wire 可见的行为变化**(CC-187):`errors[]` 逐条过凭据两网,凭据位换成闭形记号(`«redacted:userinfo»` / `«redacted:query»` / `«redacted:fragment»` / `«redacted:secret»`),条数、顺序与其余字节不变。方向只会更安全(原来原样带出的凭据值不再带出);按 `errors[0]` 渲失败原因的端零改动即得净文本。射程按载体划:被挡原因不论是引擎、模型(经报告被挡的工具)还是 hook 反馈写的,进了合成终局行与 `errors[]` 就照洗;assistant 正文行、成功臂的 `result` 与错误信封的 `_sema_salvaged_result` 一个字节不碰。`errors[]` 里的非字符串元素原样放回。要在错误文本里抠 URL / 键值的消费方请按记号处理。接入文档 **§101a D-3 / 101a-2**。
69
+ - **本包合成的终局行正文在铸点洗掉凭据**(CC-187):终态错误行(`API Error:` / `Run stopped:` / `Run blocked…` / `Model output error:` 各形与「会话被占」那一句)与 `Outcome unknown:` 行,正文里的凭据位换成同一组记号;主机、端口、路径、query 键名、行首身份与其余文字逐字节不动。print 车道与交互车道是同一份产物;身份判定仍按引擎原话判完再洗,洗消只改呈现字节,不做字符面(那归呈现边界)。此前本包不洗,洗消只在终端自己的出口做,且只认 0.83.0 已改名的旧行类旗。`Outcome unknown:` 行与 `errors[0]` 洗后仍是同一句。接入文档 **§101a D-2 / 101a-2**。
70
+ - **设置来源 hook 条目上手写的 `plugin` 键在发出前剥掉**(CC-174):值为 `undefined` 也剥,条目其余部分照发;设置里写的 hook 不是插件,不许借这一键冒充插件上下文,请求体上的 `plugin` 位只由本包按插件读口交来的身份铸。宿主交来的设置文档本身不改,剥了会留一行 debug。接入文档 **§101a H-7 / 101a-6**。
71
+ - **字符面三只旧口改由同一只引擎实现,输出逐字节不变**(CC-112):`escapeDisplayControlChars` / `collapseLabel` / `capForDisplay`、同伴消息署名规范化与 hook 故障横幅共用 `displayUntrusted` 的字符面引擎,字符集保持各自旧集(双向族按枚举,不含零宽 / 软连字符 / 标签字符;署名规范化只折 C0 / DEL / 行段分隔符),全部 BMP 码元与代理组合对拍逐字节相同。它们窄于 `displayUntrusted` 的默认,新代码请用新口。接入文档 **§101a D-5**。
72
+
73
+ - **强制审批卡的那一句说明改口**(`mandatedApprovalDetail()`):此前那一句说「…so it is asked every time」,与引擎契约不符 —— 强制卡的约束是「存下的 allow 规则与记住的回答都清不掉,每一次调用要各自被回答」,而这一次的回答者可以是人,也可以是 hook 或部署运行的自动裁决(全域放行席、自动模式分类器、沙箱准入),所以这张卡未必每次都摆到人面前。新句:`no saved allow rule and no remembered answer can retire this question; each call is settled on its own — by you here, by a hook, or by an automatic check the deployment runs — and an answer covers only that call`。仍是零参数,仍不指人去写规则。按旧句逐字断言的测试改锚。接入文档 **§101a G-1**。
74
+
75
+ ### Fixed
76
+
77
+ - **`failed` 帧上不是字符串的 `errorMessage` / `errorCode` 让流在合成终局行这一步抛错**:现在按缺席处理(回落到下一位,两位都缺席时正文为 `run failed`),流照常收尾。接入文档 **§101a D-2**。
78
+ - **交互车道重建的合成终局行丢了行类旗**(CC-187 同批):适配器把整条 assistant 消息重建成转录行时,上游帧带 `_sema_api_error_message: true` 的,重建行此前把它丢掉(合成行在交互车道上只剩 `model: '<synthetic>'` 可认),现在同样带上;模型行不新增任何键。接入文档 **§101a D-4**。
79
+ - **审批回决备注与子代续跑收据放过了双向 / 格式字符**(CC-112 残留收编):`readDecisionNoteAudit(ack).note` 与 `decisionNoteAuditLine(...)` 引用的备注,此前清掉控制字符却放过双向重排 / 格式字符(U+202E、U+2066、零宽、标签字符)、行 / 段分隔符与孤代理项,能把一条拒绝理由在屏上重排成另一句,现在一并折成空格(清洗后为空的备注与纯空白同处置:不算回显正文,不再渲一对空引号);`resumeSettledSubagent(...)` 成功时的 `receipt` 与失败时的调试日志行,此前只清 C0 / DEL,现在 C1(如 U+0085)、双向 / 格式字符与孤代理项一并换成 `.`,收据封长时不再截出半个代理对。除这几类字符外输出逐字节不变。接入文档 **§101a D-6 / D-7**。
80
+ - **设置来源的 `mcp_tool` hook 条目让整个请求被拒**(CC-174):引擎的 hooks 契约不认这一型,原样发出时整份 hooks 解析失败、整个请求被拒。现在:① 按计划组请求体(`hooksForWire({ plan })`)时不发,逐条记进 `plan.settingsDropped`,因此哪都不跑的进「不会执行」清单;② 不传计划的 `hooksForWire()` 在引擎点亮的事件上**照旧发出** —— 这类宿主看不到清单,整个请求被拒是它们唯一看得见的信号(被静默拿掉的守卫 hook = 用户以为在生效、其实没跑),有意保留;③ 引擎不点亮的事件上一律不发(引擎本来不跑这些事件,发出去只换来一次整份被拒)。接入文档 **§101a H-5**。
81
+ - **设置来源带参数的 command hook 条目被引擎只拿可执行名去跑**(CC-174):`type` 为 `command` 且带参数的条目 —— 非空数组(exec 形),或串 / 对象 / 数这些非数组值(上游本身不认,但照发同样会被剥)—— 引擎报出 `taskSettings.pluginHooks` 之前会静默丢掉 `args`、把 `command` 交给 shell 去跑(参数丢了,`command: "bash"` 这类会把 hook 的输入当脚本执行);现在能力位到货之前不发这类条目(不传计划时按未报判),到货之后原样发(非数组形由引擎响亮拒)。prompt / http 等条目上的 `args`、值为 `null` 的 `args` 不算,照旧原样发出;空数组 `args: []` 也不算带参数、照旧原样发出 —— 没有参数可丢,引擎按 shell 形跑的就是同一个可执行文件(如托管设置里的 `{ command: "/opt/guard/deny-dangerous", args: [] }`);只有命令串含空白、引号或 shell 特殊字符(shell 会拆词、展开或串接,与直接执行不等价)时才按带参数拿掉,判定只认由字母、数字与 `_ . / : + -` 组成的命令串为等价;其余条目(含形状不对的)照旧原样发出,由引擎响亮拒。拿掉原因与用户面话同一句(`exec_form_unsupported`)。不传计划的 `hooksForWire()` 拿掉这类条目、且那个事件由引擎点亮时,经日志口告警 —— 每个 settings 端口 × 事件 × 原因恰一次。接入文档 **§101a H-6**。
82
+
83
+ ### Gates
84
+
85
+ - 新增常驻门三道:`run-display-untrusted-projection-test.mjs`(凭据两网的判据与反例样本、合成出口、旧口收编逐字节回归、残留收编、两车道合成行洗消、单源普查)、`run-session-policy-deliverable-test.mjs`(向量表、成因闭集、整批语义、坏形不抛、措辞纪律;在场时与宿主侧那一份逐例差分,分歧只许是首尾空白、带 `*` 的通配与非 MCP 命名空间整对端三类(逐类计数);另对实装引擎包的协议命名空间表双向对账)、`run-plugin-hooks-projection-test.mjs`(真端口夹具 + 真能力缓存进真 dist 的计划、请求体与措辞口)。`gates-manifest.json` 137 → 140,README「Guards」表同批 +3 行。
86
+
87
+ ### Known limits(本版新增)
88
+
89
+ - 展示层:fragment 被整段换成记号之后,值里未编码的停字符(`|` `"` `<` 等)右边那一截落在洗法外(`#api_key=AB|<尾>` ⇒ `#«redacted:fragment»|<尾>`;真实 URL 进文本前已百分号编码);旧三口不跟随 `displayUntrusted` 的宽字符集(放宽是另一次行为变更);服务端逐串脱敏后投出的透传文本(引擎通告、审批卡正文与入参、回决备注、续跑收据)本包不再洗第二遍凭据;凭据面认不出的形 —— 驼峰名标签(`secretAccessKey`)、无前缀的 40 位 AWS 秘钥串、字面反斜杠写法的转义(`\x1B[33m`)夹在标签与值之间、括号或 YAML 块标记包住的值(块标记那一形里真值在下一行原样可见)、口令里含未编码 `/` `#` 的 userinfo、`Authorization` 方案词后换行再跟的普通值、标签与分隔符之间隔着空白又紧贴在上一只值后面的内层标签;仍会多遮的形 —— 标签后的普通词(`key: model`、`Missing key: ANTHROPIC_API_KEY`、`x-ratelimit-reset-tokens: 6m0s`)、方案词后的表外英文词(`basic subscription`)、引号里的键名、文档地址的 query / fragment 值、标签在行尾隔空行后的下一段首词(功能词 / 诊断词开头的除外)、`Bearer realm="…"` 挑战形里的 `realm=`;机读面(`errors[]` / 合成终局行)在记号两侧保留原有的控制符 / 格式字符(字符面归呈现边界)。与终端此前自带的那一只相比,凭据面上本包几乎只在更严一侧有差(控制符 / 格式字符 / 着色序列夹带、分隔符后长空白、`Authorization` 的非 bearer / basic 方案、散文里的已知凭据字面形),随机语料差分里更松的个例均为对方靠记号里的冒号误吃,或剥掉不可见字符后两边同样不遮。
90
+ - 会话规则:CC 旧工具名(如 `Task`)与带转义括号 / 反斜杠的名字会被判「能写」而系统里没有一层拦得住(与此前宿主侧同答);`mcp__<server>`、以及带 `*` 的规则(`mcp__<server>__*` / `mcp__<server>__get_*` / `Bash*` 等)在较新的引擎上有的其实能按集合命中,但 wire 上没有位说出对面是哪一代引擎,照旧不写(代价是晚一拍生效)。
91
+ - 插件 hook:引擎侧能力位与执行器尚未到货,插件 hook 在引擎腿上仍然不跑,本版做到的是如实说与判定就绪;守卫类插件 hook 投不出去时只如实告知,不改成逐次询问;插件 hook 模块(`register(on)` 形)与 skill / frontmatter hook 不在判定范围;exec 形里 `${user_config.KEY}` 由本包先替换、路径占位由执行方后替换,与上游次序相反;由引擎执行的插件 hook 没有进度与成功输出的展示面;设置来源的 hooks 在安全模式下的处置本包仍没有读口;不传计划的宿主若在引擎点亮的事件上配了 `mcp_tool` hook,整个请求照旧被拒(有意保留的响亮失败,改按计划组请求体即可);设置来源 command 条目上值为 `null` 的 `args` 照旧原样发出(与不写 `args` 同跑,上游本身不认这一形);不传计划的宿主上,被拿掉的 exec 形设置来源 hook(包括守卫类)只经日志口告警一次 —— 宿主若把日志只写进调试记录,用户看不到这一句,换钉前请改按计划组请求体并渲「不会执行」清单(见接入文档 §101y);空参数数组的 command 条目只在命令串由字母、数字与 `_ . / : + -` 组成时照发,引擎在 Windows 上改用 PowerShell 执行时这一等价判据未实测。
92
+ - 完整台账见接入文档 §101 末行「包侧缺口」。
93
+
52
94
  ## 0.83.0(2026-09-26)
53
95
 
54
96
  > 主题:🔴 **minor**(行为面与型面都有 BREAKING)—— CC 形消息上 24 项非 CC 键按裁定 C-R103([8200] / [8232])改 `_sema_` 名、删除或改走 chrome 臂,转录 id 改成 UUID 形,0.82.7 标过渡的三只联网搜索旧读口到期删除,公面类型 −4;同版另有结果帧 CC 键 `terminal_reason`、用户层 `disableAllHooks` 在引擎腿上生效、退化审批卡三形拒收改写、注入件自愈句整族重写与引擎通告码册 +2;根公面运行期导出 1240 → **1238**,peer sdk 地板 `>=11.3.0` 不动。
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.83.0
38
+ **Version:** 0.83.1
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -306,7 +306,7 @@ guard still cross-checks the table by name).
306
306
  | `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails |
307
307
  | `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever. One reading here answers a question that the terminal state structurally cannot: whether this run was assembled with any file-and-shell tools at all. The engine's terminal vocabulary says a run finished, not whether the work got done, so an orchestrator that waits for the end and then guesses has nothing to guess from — while the assembly manifest already said it at the start, one row per mounted instance with the single condition that mounted it. The reading is three-state and both folds are refused: a roster that is readable and carries no such row is the engine stating a fact, while no roster at all is not that fact — the static half of a manifest never carries one, and an older engine reports rosters without naming the mount condition at all, where an empty count would be a statement about the reader rather than about the run. Those two are kept apart in the reason the reading carries, and the wording for every unknown case is checked never to claim the run had no tools. The same roster now decides the tool list on the first line of a non-interactive run: the host holds that line until the roster arrives and lists exactly what the engine mounted at the start of the run, in mount order. The guard runs a real assembly frame through the projection into the decision, and pins that the host falls back to the estimate only once the roster is known not to be coming — a manifest without one, an unreadable one, model output or the run's end arriving first — rather than on a timer alone (model activity counts, including a model call that is still waiting or retrying; an error line the stream synthesizes when a run fails before assembly counts as the run ending), that a sub-run's manifest is never mistaken for the run's own, that an empty roster is taken as the engine's answer rather than as silence, and that the wait bound covers both sequential default budgets the engine gives an external tool server to connect and list its tools. The holding logic itself lives in the package as a small per-run gate — buffer, decide once, release the held messages in arrival order, then pass through — and the guard drives real stream output through it to pin that the release happens exactly once, at the manifest, releasing exactly the held prefix. The ordering itself also lives in the package as a stream wrapper, and the guard checks the final output a consumer reads: the first line is always the tool-list line, a message that arrives while that line is still being built comes after it, a timer firing races nothing out of order, a source that ends or fails before the decision still gets its first line and held messages out before the error, and an early exit closes the source |
308
308
  | `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
309
- | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. |
309
+ | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer; it must not say the question is asked every time — an answer for this one call may come from the person, a hook or an automatic check the deployment runs — and its wording is checked against the engine package's own description of the mandate A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. |
310
310
  | `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
311
311
  | `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
312
312
  | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key on the success result: the CC-spelled `structured_output` is the only home for the value the wire calls `structuredOutput`. The camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rode alongside it for exactly one release (0.79.1) and is **absent from 0.80.0 on**, pinned both by own-key and by `in`, so a consumer still reading the old name sees `undefined` rather than a stale copy. The wire position is read exactly once, so a value-changing accessor is only ever asked for its first answer; absence stays absence; a wire key that is present but `undefined` mints nothing, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` are still minted, and so are shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope carries no such key, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
@@ -426,6 +426,9 @@ guard still cross-checks the table by name).
426
426
  | `scripts/run-memory-saved-projection-test.mjs` | Engine memory writes (a successful `Remember` tool call) moved off the transcript onto the additive `memory_saved` chrome event, driven through the real pipeline: zero transcript rows for the write (the transcript is byte-identical to the same frames with the write reported as not successful), exactly one event whose `notes` carry the note text verbatim (notes, not file paths) and whose key set is exactly kind / laneProof / id / notes; no event for a missing, empty or non-string note, a non-`true` `ok`, a tool error, a missing result or another tool name; two writes give two events in order with distinct ids; a sub-agent write rides the sub-agent lane and an empty parent id emits nothing rather than falling back to the main lane; the `id` is derived from the write's wire key (the tool-end event id, else the tool-start event id, else the call id; empty ids count as absent), so projecting the same wire events twice gives the same id, and it never collides with the tool result row of the same or another call; the event sits right after the tool result row; the arm is registered as required. |
427
427
  | `scripts/run-result-frame-projection-test.mjs` | Result frames and the synthesized terminal rows. The CC key `terminal_reason` is minted on result frames only where it follows from what the engine reported: `completed` on success, `max_turns`, `budget_exhausted` and `structured_output_retry_exhausted` for the three matching engine codes, on both the done-frame path and the failed-event path. Every other outcome leaves the key absent as an own property rather than present with an undefined value: wall-clock and token-budget limits, the classifier denial limit, cancellation, unknown codes, blocked, paused, unreadable or missing terminal records, and the busy-session refusal. The public reader `terminalReasonForResult` shares the minting predicate and is checked to agree with the minted key on every frame the gate produces. Both the minted key and the reader derive the word from the frame's CC subtype (success with `is_error` strictly false, and the three limit subtypes), not from the error code, so a replayed row whose status is paused, blocked or unrecognised never carries a word that contradicts its subtype. The four words are checked against the mirrored CC union, and the minting file is checked to hold no hand-copied code literals. The renamed superset keys (`_sema_error_code`, `_sema_salvaged_result`, `_sema_model_degraded`, `_sema_selected_model`, and the row flag `_sema_api_error_message`) are driven through the real stream pipeline. Each must be present under its new name, the old name must be absent, and every frame the gate saw is swept for old names. The selected model appears on error envelopes whenever the terminal record carries it, and never on a failed event, which has no record. It stays separate from the provider-reported model name. The two in-package readers still work: the interactive result arm reads the salvaged text under its new name (and old-shape frames under the old one), and the print init gate treats the renamed flag as the run having ended. |
428
428
  | `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
429
+ | `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed five-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule covering every tool of one server or agent peer (for every protocol namespace the engine knows, checked against the engine package's own table) or containing a wildcard (*) anywhere (an engine that compares exact names would block nothing), and an entry that is not a tool name — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for three deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; and peer-wide rules outside the MCP namespace, which it did not recognise), whose disagreements are counted per class and must match an independent count exactly. |
430
+ | `scripts/run-plugin-hooks-projection-test.mjs` | Plugin hooks: each command hook an enabled plugin declares is decided one by one as running in the engine, running in this client, or not running at all, and the page of hooks sent with a request is built from the same per-turn plan the client uses to skip its own copies, so one hook never runs in two places. Governance is judged first and always wins — a managed hooks switch-off, an untrusted workspace, safe or bare mode, or a governance read that fails sends no plugin hook and does not list it as a gap; managed-hooks-only (set directly, or through a merged non-managed hooks switch-off) keeps only managed plugins; the plugin-only customization lock does not touch plugin hooks. A hook reaches the engine only when this client started the engine on this machine, the engine reports plugin-hook support, the entry is a command, the plugin declares no sensitive option, and the event still fits the engine's per-event limits; the gate walks that matrix cell by cell, including the limit boundaries and a session goal hook counting toward them. A fact that was never read is reported as not known rather than as a fact: a host that does not say where the engine runs gets a "not known whether this client started the engine" reason, an engine whose capabilities have not been read yet gets a "not known yet whether it supports plugin hooks" reason, and the plan's two engine facts are null in those cases, not false. Events the engine never fires run only if the client says it fires them itself, and hooks the upstream behaviour itself refuses (option references in a shell-form command, an unset option in exec form, malformed entries) run nowhere. Exec-form arguments are passed element by element with only saved non-sensitive option references filled in; path placeholders are left for the executor. Sensitive option values never reach the request: with a host that wrongly supplies one, every string in the plan, the request body, the notice, the labels and the log is searched for it across eight cells. A host without the plugin reader keeps the previous request body and gets exactly one warning per settings port; plugin data that throws while it is being read (a throwing getter, a revoked proxy) is treated like a failing reader — no plugin hooks this turn, settings hooks still sent, nothing thrown; the not-running notice names the plugin and events, never a command or an option value, and escapes control characters in names. Command hooks from settings that carry arguments (a non-empty `args` array, which is the exec form, or any other non-null value) are removed from the request until the engine reports support for arguments, because the engine would otherwise drop the arguments and run the bare command through a shell; an empty `args` array is not treated as carrying arguments when the command is made only of letters, digits and `_ . / : + -` (the shell runs the same executable), so such a guard still reaches the engine, while an empty array on a command with spaces or shell characters is removed; `args` on a prompt or http entry, a null `args`, or an entry with no type is left alone, and those go out unchanged. MCP tool hooks, which the engine cannot parse, are removed only from a request built from a plan, whose not-running notice the host shows; a request built without a plan still carries them on engine-fired events, so the engine rejects the whole request loudly instead of a guard hook silently not running — the gate checks both request bodies against the engine's own hooks schema. Without a plan, every removed hook of that kind on an engine-fired event produces one warning per settings port, event and reason. Malformed entries still pass through for the engine to reject loudly, and passing null where the options object goes behaves like passing nothing; a `plan` option that is not a plan is ignored rather than turning the whole page into nothing, and a plan passed directly in place of the options object is recognised and used. A `plugin` key written by hand on a settings hook is stripped before sending (even when its value is undefined), because only hooks that come from the plugin reader may carry plugin context; the settings document itself is left untouched and a debug line records the count. The two hand-copied tables, the engine-fired event list and the engine limits, are checked against their owners. |
431
+ | `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. |
429
432
 
430
433
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
431
434
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
@@ -119,6 +119,7 @@ const assistantArm = function* (m, { ctx, idOf, text, cards, inst }) {
119
119
  uuid: idOf(m),
120
120
  session_id: ctx.sessionId,
121
121
  parent_tool_use_id: null,
122
+ ...(m._sema_api_error_message === true ? { _sema_api_error_message: true } : {}),
122
123
  }, ctx.now());
123
124
  text.markEmittedText(textBlock.text);
124
125
  }
@@ -1,6 +1,7 @@
1
1
  import { isCcToolDenialKind, isCcToolDenialKindADenial, isGateDeniedByWord } from '../../gateVocabulary.js';
2
2
  import { stamp } from '../types.js';
3
3
  import { putOwnKey } from '../../ownKey.js';
4
+ import { REDACTION_MARKER_PREFIX, washCredentials } from '../../displayUntrusted.js';
4
5
  import { readEffectiveReasoning, readEffectiveMemoryScopes } from '../../effectiveFacts.js';
5
6
  import { isReviewPark, readRunTerminal, runTerminalCode, runTerminalGateToolName, } from '../../runTerminal.js';
6
7
  import { toCcModelUsage } from './turnUsageToModelUsage.js';
@@ -30,7 +31,6 @@ function nonEmptyStr(v) {
30
31
  return typeof v === 'string' && v.length > 0 ? v : undefined;
31
32
  }
32
33
  export const TOOL_INPUT_JOIN_MAX_CALLS = 4096;
33
- const REDACTION_TOKEN_PREFIX = '\u00abredacted';
34
34
  function isTransportPlaceholderLeaf(v) {
35
35
  return v === '[circular]' || v === '[depth-limit]';
36
36
  }
@@ -46,7 +46,7 @@ function joinableToolInput(args) {
46
46
  if (++nodes > TOOL_INPUT_SCAN_MAX_NODES || depth > TOOL_INPUT_SCAN_MAX_DEPTH)
47
47
  return undefined;
48
48
  if (typeof v === 'string') {
49
- if (v.includes(REDACTION_TOKEN_PREFIX) || isTransportPlaceholderLeaf(v))
49
+ if (v.includes(REDACTION_MARKER_PREFIX) || isTransportPlaceholderLeaf(v))
50
50
  return undefined;
51
51
  continue;
52
52
  }
@@ -483,7 +483,7 @@ function errorResult(ctx, parts) {
483
483
  ...costFactParts(parts.stats, parts.observed),
484
484
  ...reemitEffectiveFacts(parts.facts),
485
485
  ...nestedUsageByTaskParts(parts.stats, parts.observed?.nestedUsageByTask),
486
- errors: [...parts.errors],
486
+ errors: parts.errors.map(washCredentials),
487
487
  ...terminalReasonParts(parts.subtype, true),
488
488
  ...(parts.errorCode !== undefined && parts.errorCode.length > 0 ? { _sema_error_code: parts.errorCode } : {}),
489
489
  ...(parts.degraded !== undefined ? { _sema_model_degraded: parts.degraded } : {}),
@@ -5,6 +5,7 @@ import { readRunCostFacts, terminalToSdkResult, TOOL_INPUT_JOIN_MAX_CALLS } from
5
5
  import { coerceOutput, publishSubagentContentEvent } from '../subagentContentStore.js';
6
6
  import { ACTIVE_RUN_BUSY_ERROR_CODE, OUTPUT_INVALID, isLimitsExceededCode } from '../engineErrorCodes.js';
7
7
  import { isReviewPark, readRunTerminal, runTerminalCode } from '../runTerminal.js';
8
+ import { washCredentials } from '../displayUntrusted.js';
8
9
  import { ccToolDenialKindForToolEnd, gateDeniedBy, gateOutcomeOf } from '../gateOutcome.js';
9
10
  import { isCcToolDenialKind, isCcToolDenialKindADenial, isGateDeniedByWord } from '../gateVocabulary.js';
10
11
  const mainLane = () => ({ lane: 'main' });
@@ -101,7 +102,7 @@ function syntheticTerminalRow(ctx, text) {
101
102
  uuid: undefined,
102
103
  session_id: undefined,
103
104
  type: 'assistant',
104
- message: { role: 'assistant', model: '<synthetic>', content: [{ type: 'text', text }] },
105
+ message: { role: 'assistant', model: '<synthetic>', content: [{ type: 'text', text: washCredentials(text) }] },
105
106
  parent_tool_use_id: null,
106
107
  _sema_api_error_message: true,
107
108
  });
@@ -158,6 +159,21 @@ export function _resetDroppedFrameReportForTest() {
158
159
  export function _droppedFrameMemoSizeForTest() {
159
160
  return reportedDroppedTypes.size;
160
161
  }
162
+ function withStringFailureFields(ev) {
163
+ if (ev.type !== 'failed')
164
+ return ev;
165
+ const got = ev;
166
+ const badMessage = got.errorMessage !== undefined && typeof got.errorMessage !== 'string';
167
+ const badCode = got.errorCode !== undefined && typeof got.errorCode !== 'string';
168
+ if (!badMessage && !badCode)
169
+ return ev;
170
+ const { errorMessage, errorCode, ...rest } = ev;
171
+ return {
172
+ ...rest,
173
+ ...(badMessage ? {} : { errorMessage }),
174
+ ...(badCode ? {} : { errorCode }),
175
+ };
176
+ }
161
177
  let inFlightTurns = 0;
162
178
  export function isRunStreamActive() {
163
179
  return inFlightTurns > 0;
@@ -199,7 +215,8 @@ async function* runStreamInner(events, ctx, handle = {}) {
199
215
  return false;
200
216
  }
201
217
  };
202
- for await (const ev of events) {
218
+ for await (const rawEv of events) {
219
+ const ev = withStringFailureFields(rawEv);
203
220
  if (ev.type === 'tool_end' && ev.isError !== true)
204
221
  successfulToolEndObserved = true;
205
222
  const seq = eventSeq(ev);
@@ -0,0 +1,24 @@
1
+ export declare const REDACTION_MARKER_PREFIX = "\u00ABredacted";
2
+ export declare function looksLikeKnownSecret(s: string): boolean;
3
+ export declare function washCredentials(text: string): string;
4
+ export type DisplayMark = 'escape' | 'dot' | 'space';
5
+ export interface DisplayFace {
6
+ readonly unsafe: RegExp | null;
7
+ readonly fold: boolean;
8
+ readonly mark: DisplayMark;
9
+ }
10
+ export declare const LEGACY_ESCAPE_FACE: DisplayFace;
11
+ export declare const LEGACY_LABEL_FACE: DisplayFace;
12
+ export declare const LEGACY_NAME_FACE: DisplayFace;
13
+ export declare function applyDisplayFace(text: string, face: DisplayFace): string;
14
+ export interface DisplayUntrustedOptions {
15
+ readonly credentialUrls?: boolean;
16
+ readonly credentialWords?: boolean;
17
+ readonly controls?: boolean;
18
+ readonly bidi?: boolean;
19
+ readonly foldLines?: boolean;
20
+ readonly keepLayout?: boolean;
21
+ readonly mark?: DisplayMark;
22
+ readonly max?: number;
23
+ }
24
+ export declare function displayUntrusted(text: string, opts?: DisplayUntrustedOptions): string;