@sema-agent/client-core 0.85.0 → 0.85.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +91 -0
- package/README.md +5 -5
- package/dist/adapter/activeRunSelfHeal.d.ts +6 -0
- package/dist/adapter/activeRunSelfHeal.js +272 -50
- package/dist/hitl/armedGateRegistry.d.ts +1 -0
- package/dist/hitl/armedGateRegistry.js +1 -1
- package/dist/hitl/planReviewWire.js +14 -3
- package/dist/host.d.ts +1 -1
- package/dist/host.js +11 -2
- package/dist/index.d.ts +1 -1
- package/dist/index.js +1 -1
- package/dist/request/taskRequest.d.ts +2 -0
- package/dist/request/taskRequest.js +12 -0
- package/dist/subagent/subagentOwnerAbsence.js +2 -5
- package/docs/INTEGRATION-CLIENTS.md +335 -14
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -49,6 +49,97 @@
|
|
|
49
49
|
> 挡住 ⇒ 本批把它机械化——④a0 对 `pending` 行**要求段头已是日期形**(`(未发布)` 直接红),阶段一
|
|
50
50
|
> commit 漏转在发布前就红,不再靠人记。
|
|
51
51
|
|
|
52
|
+
## 0.85.2(2026-09-30)
|
|
53
|
+
|
|
54
|
+
> 主题:**patch**,一件 + 一处同族修复 —— 自愈重开的两只失败结局(`plan-review-reopen-failed` / `ask-reopen-failed`)带上**这一次重开放弃所依据的那一发失败请求的外层 wire 码**(CC-260,终端的请托);另把自愈腿读**错误对象与成功应答记录**上的各位收进保护(见 Fixed),注入形交接行在投递方式未知时改说实话(见 Changed)。零型面 BREAKING;根公面运行期导出 1310 不变;测试钩 60 不变;公面类型名不变(两个已导出型各多一个可选位:`ReopenCardVerdict` 不成功臂 `_sema_errorCode` · `SelfHealOutcome` 两只失败臂 `errorCode`);超集键 +1(`_sema_errorCode`);`activeRunSelfHealRow` 不引用 `errorCode`,只在注入形交接行、投递方式未知那一形换两句(见 Changed);零投影臂;peer sdk 地板 `>=12.0.1` 不动;开发依赖不变。
|
|
55
|
+
|
|
56
|
+
### Added
|
|
57
|
+
|
|
58
|
+
- **自愈失败结局的 `errorCode`**(CC-260):`SelfHealOutcome` 的 `plan-review-reopen-failed` 与 `ask-reopen-failed` 各 +1 可选键 `errorCode?: string` —— 这一次重开**因为一发请求失败而放弃**时,那一发的外层 wire 码。present-iff(只在有码时在场,值恒为非空串)。读法与本包按码分诊的各处同一只:折叠码 `resume_blocked_by_policy` 在时答 `blockedBy` 那个原码,否则答 `errorCode`(服务端没给具体码时是按状态派生的粗码,例 `request.rejected`)。接入文档 **§111 EC-1 / 111a-1**。
|
|
59
|
+
- 口径是「放弃所依据的那一发」:带的是**放弃所依据**那一发的码;之后的请求若**改变了放弃的依据**(例:轮询终于答上了、待决行出现了)⇒ 不带;只作确认的读(例:决断 POST 失败之后确认待决行仍在)不改依据,照带那一发 POST 的码。放弃所依据的那一发没有 wire 码(断连 / 截止 / 中止 / 客户端自己的错)⇒ 不带,不回溯更早的码;调用方已中止(放弃的依据是中止)⇒ 不带;因判定拒开(决断在路上 / 卡是用户关的 / 会话没有接界面)⇒ 不带 —— 按判决本身答的拒开,两只失败结局在这三形上恒不带 `errorCode`(本包保证;ask 臂认判决上的用户关卡位来丢码,结局不加这一位);ask 判决带 `pendingRowGone` 且不在飞时改道陈旧 park 那条路,本包自己那一发读 run 状态失败照口径带码(放弃所依据的那一发是它,判决上别的判定位不拦)。本包自己的来路上,`pendingRowGone` 与它同在只有一形:待决项已证不在、紧接着本包复核那条 run 的状态时请求失败;宿主判决自带 `pendingRowGone` 与 `_sema_errorCode`(不合口径的组合)时,计划臂原样透传两位。
|
|
60
|
+
- 码的三条来路:宿主重开口(`reopenPlanReview` / `reopenAskPark`)抛错 ⇒ 抛出物的码;宿主重开判决上的 `_sema_errorCode`(下一条);待决项已证不在时本包自己的三发 —— 呈卡前、呈卡后两次读那条 run 的状态,开枪前再问一次属主待决行数 —— 失败即收口,带那一发的码(只认请求本身失败:请求答了、答回的记录上 `status` 读不出不算,按读不到状态处置、不铸码)。
|
|
61
|
+
- 出句不引用它:`activeRunSelfHealRow` 的每一句与不带码时逐字节同;要说成因的端自己读这一键。
|
|
62
|
+
- **重开判决的失败码位**(CC-260):`ReopenCardVerdict` 不成功臂 +1 可选位 `_sema_errorCode?: string`,宿主自带的重开口按同一口径**可选**铸(只在 `reopened: false` 上读,只认非空串;读不出按缺席,与判决其余各位同一读法)。宿主重开口抛错时(调用方未中止),本包在保护内把抛出物的码铸到判决上 —— 两只失败结局与 steer `queued` 结局的 `reopened` 读的是同一个位;`queued` 结局的 `reopened` 是宿主判决的原样拷贝,宿主判决上判定位与码两样都给了就两样都在(丢码只在两只失败结局上做)。包内计划复核重开口 `reopenPlanReviewCard` 在重开这一步不发任何 wire 请求,不铸这一位。接入文档 **§111 EC-2 / 111a-1**。
|
|
63
|
+
|
|
64
|
+
### Changed
|
|
65
|
+
|
|
66
|
+
- CC-260 在已发导出上的可观察变化:宿主重开口**抛出一个带 wire 码的错误**(调用方没中止)时,两只失败结局多 `errorCode`,steer `queued` 结局的 `reopened` 从 `{ reopened: false }` 变为 `{ reopened: false, _sema_errorCode: <码> }`;抛出物没有 wire 码(普通 `Error`、断连、非对象)时与 0.85.1 逐字节同。另:待决项已证不在、本包复核请求失败的那几形,`ask-reopen-failed` 多 `errorCode`。其余结局、判决、出句与 0.85.1 逐字节同(错误对象 / 应答记录读不出的那几形见 Fixed)。
|
|
67
|
+
- 注入形(`activeRunSelfHealRow(…, 'injected', follow)`)的交接行:`running-steered` 的投递方式缺席 / 读不出 / 表外词时,修前与 `applied` 同句(「was passed to the reply already in progress … instead of starting a new one」,宿主声明续听时还说「nothing for you to do now」,替引擎断言消息已交给正在进行的那一轮),本版改说投递方式未知 —— 两种续听声明下都不说「不用动手」:
|
|
68
|
+
- 未声明续听 —— `A follow-up message sema sent on its own (not one you typed) was accepted for the reply already in progress (id <id>), but how it will be delivered was not reported, so sema cannot tell whether that reply will pick it up — watch that reply before deciding whether anything needs to be sent again.`
|
|
69
|
+
- 声明续听 —— `A follow-up message sema sent on its own (not one you typed) was accepted for the reply already in progress (id <id>), but how it will be delivered was not reported, so sema cannot tell whether that reply will pick it up. sema is following that reply and will show whatever it asks for next; check what it does before deciding whether anything needs to be sent again.`
|
|
70
|
+
- `applied` 那两句、用户形(本就说「the engine did not say how it will be delivered … before sending it again」)、处置档(`handed-off`)与续听意图(投递方式未知仍按保守当活 ⇒ `tail`)都不变。接入文档 **§111 EC-7 / 111a-2**。
|
|
71
|
+
|
|
72
|
+
### Fixed
|
|
73
|
+
|
|
74
|
+
- `attemptActiveRunSelfHeal` 读**错误对象与成功应答记录上的**各位收进保护(CC-260 同族修复;0.85.1 把宿主重开判决各位收进了保护,这是另外两半)。错误对象:自愈腿里 `runs.get` / `runs.steer` / `runs.cancel` 抛出的错误对象,`status` / `message` 是会抛的取值器、或整个是读就抛的 Proxy(`get` / `has` 陷阱)⇒ 按那一位缺席处置,结局照常落定。修前(0.85.1 实测)这几形 `attemptActiveRunSelfHeal` 直接 reject,抛出的就是取值器那个错:待决项已证不在时本包读 run 状态、`running` 三选卡出卡前的存活对账(两处都读 `status === 404`);kind 与状态都缺席时补读 run 状态失败、steer / cancel 失败腿(细节句读 `'message' in e` / `.message`);steer / cancel 的至多一次失败分类(读 `status`)。这几处读法分别自 0.29.0 / 0.30.0 / 0.37.0 起就是裸读(按源码历史;早于 0.85.1 的产物没有实跑)。修后:读不出 `status` ⇒ 不当 404(不判幽灵)、至多一次失败分类答 `unknown`(送达未知,不断言「没送到」);读不出 `message` ⇒ 细节句 `the engine gave no detail`;`errorCode` / `blockedBy` 读不出按缺席同样落定(取码口原本就在保护内,本版补格钉回归)。公面 `atMostOnceFailureClass` 对这类对象答 `'unknown'`、不再抛;公面 `readSteerDelivery` / `readSteerReceiptStatus` 同样收进保护:回执上那一位读不出(取值器 / Proxy 抛)⇒ 答 `null`(与缺那一位同答)、不再抛(0.85.1 及更早直接抛)—— 自愈 steer 腿随之:回执读不出 ⇒ 落 `running-steered`(投递方式读作缺席),用户形说「sema handed your message to run …, but the engine did not say how it will be delivered … watch that run before sending it again」,注入形说投递方式未知那一句(见 Changed)。修前(0.85.1 实测)落 `running-steer-failed`:取值器抛普通错误 ⇒ 送达未知那一句,细节句是那条取值异常;取值器抛出的是带 422 形状态码的对象 ⇒ `delivery: 'rejected'`,用户形说「Your message was NOT sent; send it again if you still want it.」—— 而引擎其实已经收下了那一枪,这句话诱导用户把一条非幂等指令再发一次。成功应答记录:kind 与状态都缺席时那一次补读 run 状态,`runs.get` 答回的记录上 `status` 读不出(会抛的取值器 / 读 `status` 就抛的 Proxy)⇒ 与读不出状态同处置,落 `state-unknown`「the engine reported no status for that run」(不把那条异常当成引擎给的细节);修前(0.85.1 实测)reject,这一读自 0.29.0 起就在保护外(按源码历史)。待决项已证不在时两次读状态与 `running` 存活对账:请求答了、记录上 `status` 读不出 ⇒ 同样按读不到状态处置(不铸码、不判幽灵);修前(0.85.1)这一读与请求共用一处保护,读抛按「那一发没答上来」处置 —— 结局相同,只差一形:取值器抛出的恰是 404 形错误时,修前被当成引擎说「查无此 run」(落 `ask-run-not-found` / `running-not-found`),修后按读不到状态收口。cancel 之后的释放轮询不变(读抛按那一发没答上来)。普通错误对象与普通记录(SDK 给的那一类)的结局与 0.85.1 逐字节同。接入文档 **§111 EC-6**。
|
|
75
|
+
|
|
76
|
+
### Gates
|
|
77
|
+
|
|
78
|
+
- 扩门 `scripts/run-selfheal-reopen-test.mjs`(595 → 704 格)。G14(+22,CC-260):宿主重开口抛 503 / 409 ⇒ 结局带那个码、其余键逐字节同;折叠码按 `blockedBy`(顶层 / `extra` 两处),`blockedBy` 缺席答折叠码本身;客户端错 / 断连 / 空串码 / 只有状态码 / 抛非对象 ⇒ 无键、与 0.85.1 逐字节同;steer `queued` 时重开口抛错 ⇒ 结局的判决带 `_sema_errorCode`、结局形不加键、两形行句不变;宿主判决位透传,坏形(非串 / 空串 / 数组 / `null` / `undefined`)不带,只在 `reopened` 严格 `false` 上读,挂在原型链上的取值器与不可枚举的自有键照读;带码结局的用户形 / 注入形 / 带保会话出路三句与处置档与去掉码的同一结局逐字同;判定形(在飞 / 用户关卡 / 无界面)宿主带了码也丢;包内计划复核重开口无界面、关卡后门实例读失败 ⇒ 判定形、零码;待决项已证不在时本包三发:呈卡前读状态失败 ⇒ `pendingRowGone` 与码同在、宿主早先的码被这一发的码盖掉,断连 ⇒ 不带且不回溯,读状态成功后按判定收口(无呈卡口 / 表外状态词 / 奇形记录 / 读口缺席)⇒ 不带(宿主判决上的码不沿用),呈卡后读状态 / 开枪前复证失败 ⇒ 带码、不带 `pendingRowGone`、零 cancel,复证答非 0 / 非 wire 错 / 呈卡口抛错 ⇒ 不带,404 ⇒ `ask-run-not-found` 不加键,读到计划停驻转计划臂 ⇒ 只看计划重开口的判决。G15(+47,同族修复):错误对象七形(`status` / `errorCode` / `blockedBy` / `message` 取值器抛、Proxy `get` / `has` / 全陷阱抛)× 五条路(陈旧 park 卡前复核 / `running` 存活对账 / kind 与状态都缺席时的补读 / steer 失败腿 / cancel 失败腿)⇒ 不 reject、结局照常落定;公面至多一次失败分类口对七形不抛、答 `'unknown'`(折叠码那一形 409 照答 `'rejected'`),正常形逐格同旧;宿主重开口抛这些对象 ⇒ 两臂照常落定;成功应答记录两形(`status` 取值器抛 / 只在读 `status` 时抛的 Proxy)× kind 与状态都缺席时的补读 ⇒ 落 `state-unknown`、细节句是「读不出状态」那一句,另三条读记录的路(陈旧 park 复核 / `running` 存活对账 / cancel 后释放轮询)与全陷阱都抛的 Proxy 记录(`await` 探 `then` 即抛,按读口失败处置)钉回归。G16(+16):成功答回的记录上 `status` 取值器 / Proxy 抛带码错 ⇒ 卡前 / 卡后复核都不带码(卡后零 cancel)、抛 404 形错误不判幽灵,请求本身 reject 照旧铸码 / 404 照旧判幽灵;调用方中止后重开口抛带码错 ⇒ 两臂与 steer `queued` 路径都不带码,没中止照旧带;ask 判决带用户关卡位与码 ⇒ 不带码、结局不加位;用户关卡那一刻门实例读失败 ⇒ 此刻那一次读不发、答用户关卡、不带码;steer 回执两只公面读口遇读不出的回执答 `null`、不抛,自愈 steer 腿落 `running-steered`。G17(+18):回执整只不可读 / 只投递方式不可读 × 两种续听声明 ⇒ 注入形走投递方式未知那一句、不说「nothing for you to do」,用户形同说「引擎没说怎样送达、先看那一轮」;投递方式缺席 / 表外词 / 空串同句;`applied` 两句照旧,续听意图与处置档不变。判决读取保护那一组(G2⑧u–⑧x)的位名单加第九位 `_sema_errorCode`(+6)。
|
|
79
|
+
- 登记物:超集台账 `docs/type-superset.json` 92 → 93 条(`_sema_errorCode`:判决形声明、读口与抛错铸点同在一处);扩门 `scripts/run-terminal-identity-copy-test.mjs`(1041 → 1077 条):注入形事实表 15 → 17 句(投递方式未知的两句),样本补缺席 / 空串 / 表外词三形;「nothing for you to do」只许在续听已挂且投递方式为 `applied` 的交接档与重投档出现。门清单 `scripts/gates-manifest.json` 自愈重开门的说明与 README「Guards」派生表同批更新;门数 158 不变。
|
|
80
|
+
|
|
81
|
+
### Known limits(本版新增)
|
|
82
|
+
|
|
83
|
+
- 包内计划复核重开口在自动触发、撞上用户关卡记号时要读两次门实例(用户关卡那一刻起读的那一次、此刻那一次),任一次失败都按「证不出是新门」答用户关卡:结局带 `dismissedByUser`、不带 `errorCode` —— 由请求失败推出的判定形(0.83.4 起如此);读失败若是因为引擎在排空,端拿不到那个码(KL-287)。
|
|
84
|
+
- 完整台账见接入文档 §111 末行「包侧缺口」。
|
|
85
|
+
|
|
86
|
+
## 0.85.1(2026-09-30)
|
|
87
|
+
|
|
88
|
+
> 主题:**patch**,三件 —— 请求装配回执的省略成因补一张**给终端用户看**的句表与取句口(CC-248);plan-review 重开口在「这个会话没有接能显示卡的界面」那一形上带机读成因位、自愈行补一句成因(CC-258);多会话宿主在自己的会话键上登记呈现回执的口 `registerArmedGateFromQuestionIdFor` 回到根入口(CC-259)。零型面 BREAKING;已发导出的可观察变化:「没有接界面」那一形(见 Changed),另有两处修复 —— 宿主重开判决某一位读不出时自愈不再 reject、宿主日志 / 探针口抛错不再改变判决(见 Fixed);根公面运行期导出 1307 → 1310(+3:`taskRequestOmissionCauseNotice` / `TASK_REQUEST_OMISSION_CAUSE_NOTICES` / `registerArmedGateFromQuestionIdFor`);测试钩 60 不变;公面类型名不变(两个已导出型各多可选位:`ReopenCardVerdict` 不成功臂 `_sema_noPresentationSurface` · `SelfHealOutcome` 两臂 `noPresentationSurface`;一个已导出函数的返回型 `void` → `boolean`:`hostLog`,见 Fixed);超集键 +1(`_sema_noPresentationSurface`);成因闭集 `TASK_REQUEST_OMISSION_CAUSES` 不变;接入方那只 `taskRequestOmissionCauseDetail` 逐字不变;零投影臂;peer sdk 地板 `>=12.0.1` 不动;开发依赖不变。
|
|
89
|
+
|
|
90
|
+
### Added
|
|
91
|
+
|
|
92
|
+
- **`taskRequestOmissionCauseNotice(cause)`** ⇒ `string`(CC-248):装配回执(`assembleTaskRequest(…).omitted`)一行的成因词 → 一句**给终端用户看**的话:这一项没发出去、为什么、用户能做什么(有才说)。三端要把同一份回执说给用户时不必再各写一张成因句表。接入文档 **§110 R-1 / 110a-1**。
|
|
93
|
+
- 与 `taskRequestOmissionCauseDetail`(写给接入方:带改法,表外词原样回显供对账)读**同一个**成因闭集 `TASK_REQUEST_OMISSION_CAUSES`;两张句表各自穷举这一个闭集,逐词不同句、零共句。
|
|
94
|
+
- 四句逐字:
|
|
95
|
+
- `upstream_absent` —— `Not sent with this request: this version has no way to send this setting in any request, so giving it or leaving it out makes no difference to what is sent.`
|
|
96
|
+
- `other_channel` —— `Not sent from here: that part of the request belongs to a separate step, so what goes out for it, if anything, is decided there rather than by the value given here.`
|
|
97
|
+
- `off_lane` —— `Not sent with this request: this kind of request does not carry this setting, so giving it or leaving it out makes no difference to what this request sends.`
|
|
98
|
+
- `not_live` —— `Not sent with this request: the request was marked as not going to a live engine (for example an offline or test run), and this setting is only sent on requests that go to one.`
|
|
99
|
+
- 句子只说这一次带没带,不说「生效 / 不生效」(带了不等于引擎兑现了);不含车道、回执、上游、字段这类接入方用词。`not_live` 那一句只说「这一发被标成不发往在线引擎」—— 本包只知道调用方交来的标记,不替宿主断言它连没连引擎;`other_channel` 那一句不说「后面会设上」—— 那一步落不落、落什么,本包不知道。
|
|
100
|
+
- 表外词、空串、非串、包装串对象、数组、带 `toString` 的对象、原型链上的名字(`constructor` / `__proto__` …)⇒ 一句通用句 `Not sent with this request, for a reason this version does not recognize.` —— 不抛(读属性会抛的对象、已撤销的 Proxy 同样不抛)、不回显那个词、不借任何一个成因。只认自身就是串的值,不强转。
|
|
101
|
+
- **`TASK_REQUEST_OMISSION_CAUSE_NOTICES`**(冻结对象,型 `Readonly<Record<TaskRequestOmissionCause, string>>`):上面那四句的表本身。给要按 `TaskRequestOmissionCause` **编译期穷举**、自建呈现表(比如逐词决定渲不渲)的端:句子从这里取,成因闭集加一个词时句子随包一起到。按一个外来的串取句请走取句口 —— 直接下标会撞到 `Object.prototype` 上的成员。
|
|
102
|
+
- **重开判决的「没有接能显示卡的界面」成因位**(CC-258):`ReopenCardVerdict` 不成功臂 +1 可选位 `_sema_noPresentationSurface?: true`。`reopenPlanReviewCard` 出卡 = 把一帧问题帧投给宿主**为这个会话键订阅**的问题帧端口(`onQuestionFrameFor(sessionKey, handler)`;默认会话 `onQuestionFrame`);这个会话没有订阅者 ⇒ 卡没有地方渲、拒开,判决从此带这一位:`{ reopened: false, _sema_noPresentationSurface: true }`。同步形与带 `presentationReceiptMs` 的回执形同答;异步判定(`trigger: 'automatic'` 撞关卡记号、读门实例)读回来之后界面没了,同答 —— 不论读出的是不是新门(读后复核与入口同序:先问界面,再判用户关卡)。只认严格 `true`。接入文档 **§110 PS-1 / PS-2 / 110a-2**。
|
|
103
|
+
- **自愈结局的同名键**(CC-258):`SelfHealOutcome` 的 `plan-review-reopen-failed` 与 `ask-reopen-failed` 各 +1 可选键 `noPresentationSurface?: true` —— `attemptActiveRunSelfHeal` 拿到的重开判决带上面那一位时在场(kind 不变;两臂同判,宿主自带的 ask 重开口铸这一位时 ask 臂同样带)。接入文档 **§110 PS-3**。
|
|
104
|
+
- **`registerArmedGateFromQuestionIdFor(sessionKey, questionId)`**(CC-259;0.72.0 退出公面的名字复活):多会话宿主按自己的会话键调 `reopenPlanReviewCard(taskId, { sessionKey, presentationReceiptMs })` 时,回执形等的是**那个会话键**上的呈现回执;宿主在卡真进界面那一刻用这一口登记(原始帧 id 的回执发在这个会话键上,归一后的卡身份记进这个会话的呈现台账)。修前根入口只有默认会话口 `registerArmedGateFromQuestionId`:非默认会话键的宿主照它登记,回执窗等不到、窗尽答 `{ reopened: false }`,而卡其实已经上屏;改用 `registerArmedGateFor(sessionKey, 原始 id)` 能让回执落定,但记进台账的是未归一的原始 id,同一张卡再开也读成首见。入参语义与默认会话口逐字同:`questionId` 不是非空串 ⇒ 什么都不做(不抛)。接入文档 **§110 RG-1**。
|
|
105
|
+
|
|
106
|
+
### Changed
|
|
107
|
+
|
|
108
|
+
- 请求装配一族(CC-248)的已发导出零可观察变化:`taskRequestOmissionCauseDetail` 的四句与兜底句、`TASK_REQUEST_OMISSION_CAUSES` 的内容与序、`assembleTaskRequest` 同一入参的 `{ request, omitted }` 都与 0.85.0 逐字节同。接入方那张句表的字面量补挂 `satisfies Record<TaskRequestOmissionCause, string>`(只改构建期的类型检查,出包字节不变;见 Gates K4c)。
|
|
109
|
+
- CC-258 的可观察变化**只在「没有接界面」这一形上**(两处修复见 Fixed):
|
|
110
|
+
- `reopenPlanReviewCard`:这一形的判决从 `{ reopened: false }` 变为 `{ reopened: false, _sema_noPresentationSurface: true }`(多一键,`reopened` 照旧 `false`)。其余拒开形(回执窗没落定 / 决断在飞 / 界面在场时的用户关卡 / 非默认会话键没有投递口 / 陈旧评估时关卡记号已不在)与成功形逐字节同 0.85.0。按「恰 `{ reopened: false }`」逐键断言无界面形的测试请改锚(§110 110a′)。读后复核证出新门、读在路上时界面没了那一形,关卡账与 0.85.0 同:用户对旧门的记号撤掉(新门从未被用户关过),之后对新门的自动重开照常出卡 —— 变的只有判决多一键;入口那一形与同一道门那一形不动关卡记号。
|
|
111
|
+
- `reopenPlanReviewCard` 读后复核的一形:自动触发撞上同一枚关卡记号、读门实例证不出新门,而读在路上时界面没了 —— 0.85.0 及更早(0.84.1 实测同答)答 `{ reopened: false, dismissedByUser: true }`(「是不是新门」那一判排在界面复核之前;自愈行因此说「卡是你关的、发条消息叫回来」,而会话没有界面、发消息也叫不回卡),本版改答 `{ reopened: false, _sema_noPresentationSurface: true }`:读后复核与入口同序,先问界面、再判用户关卡。界面在场时同一形照旧答 `dismissedByUser`。
|
|
112
|
+
- `attemptActiveRunSelfHeal`:上面那一形的结局多 `noPresentationSurface: true`。
|
|
113
|
+
- `activeRunSelfHealRow`:五句在这一形上补同一句成因 `this session has no view attached that can show decision cards`,其余字节不变 —— plan / ask 用户形在 `could not reopen that approval card` / `could not reopen that card here` 之后接 `: <这一句>`;steer 回执 `queued` 的用户形在 `could not surface that decision card here` 之后、注入形在 `could not show here` 之后、注入形未送达在 `no open card you could answer to free it` 之后各接 ` (<这一句>)`。插入点之前与之后的两段各自逐字不变:整段落在插入点一侧的子串照旧命中;跨过插入点的子串在这一形上不再命中 —— 整句逐字、带句点的 `could not reopen that approval card.` / `could not reopen that card here.`、`could not surface that decision card here, so`、`sema could not show here;`、`free it. sema did not retry` 这类(§100 那张注入件句表的第 5 / 12 句在这一形上随之分句,见 §110 110a′ ⑤)。逐字表见 §110 110a-2。
|
|
114
|
+
- 首呈口 `armPlanReviewApproval` 不变(CC-258):它不问有没有界面,没有订阅者时照旧答 `true`、卡帧不会送到任何地方(见 Known limits)。
|
|
115
|
+
|
|
116
|
+
### Fixed
|
|
117
|
+
|
|
118
|
+
- `attemptActiveRunSelfHeal` 读宿主重开口交来的判决(`reopenPlanReview` / `reopenAskPark` 的返回值)收进保护:判决在保护内逐键拷一次、成一份拷贝,之后两臂、steer `queued` 结局与出句口读的都是这一份;某一位读不出(会抛的取值器 / Proxy)⇒ 按这一位缺席处置,结局照常落定,与判决只缺这一位时逐字节同(修前自愈 Promise reject,`queued` 结局的出句口同步抛)。这是把本版新读的 `_sema_noPresentationSurface` 与既有七位(`reopened` / `firstSight` / `presented` / `decidedWithoutCard` / `pendingRowGone` / `_sema_decisionInFlight` / `dismissedByUser`)一并收进保护 —— 既有七位的取值器抛错时 0.85.0 及更早同样 reject(`decidedWithoutCard` 只在 ask 臂、`dismissedByUser` 只在计划臂,其余五位两臂都);出句口会同步抛的是其中 `reopened` / `presented` / `_sema_decisionInFlight` / `dismissedByUser` 四位(`firstSight` / `decidedWithoutCard` / `pendingRowGone` 出句口不读)—— 按 0.85.0 与 0.84.1 产物实跑核过,不是本版引入。拷贝拷的是宿主判决的全部自有可枚举键(字符串键与 Symbol 键),宿主加的超集键照旧原样随行,值原样、键序照宿主给的;null 原型的判决拷成 null 原型,其余是普通对象;读不出的那一键(含可枚举性判定本身读不出的)跳过;列不出宿主判决的键时只按上面八位读。普通判决的结局与行句与 0.85.0 一致 —— 键、键序、值、原型都同(`running-steered` 结局的 `reopened` 是这份拷贝,不再是宿主交来的同一个对象)。`activeRunSelfHealRow` 读宿主自铸结局里的判决同样不抛:读不出的位按缺席,`reopened` 整个缺席与 `null` 同句;结局上 `reopened` 这一键本身读不出(会抛的取值器 / 读这一键就抛的 Proxy)⇒ 同样按缺席出句(修前两处 `queued` 出句分支同步抛,0.85.0 及更早)。接入文档 **§110 PS-6**。
|
|
119
|
+
- 宿主经 `installHost({ log, probe })` 装的日志口 / 探针口抛错,不再改变任何判决:`hostLog` / `hostProbe` 自己兜住口抛出的错(诊断面坏了只丢那一行诊断)。这是 0.85.0 及更早就有的形(0.84.1 实测同形;入口「决断在飞」与读后复核「读在路上时交出决断」两形要 0.85.0 才有的交出登记口,0.84.1 上测不了)—— 修前日志口抛错时 `reopenPlanReviewCard` 实测:入口「决断在飞」答 `{ reopened: false }`(丢了 `_sema_decisionInFlight`)· 同步形用户关卡答 `{ reopened: false }`(丢了 `dismissedByUser`)· 入口铸卡段答 `{ reopened: false }`,而这道门旧身份的卡已经撤下、新卡没有铸出来 · 读后复核里「宿主交出谓词抛」「读在路上时交出决断」「读在路上时会话换代」「新门铸卡」四形都改答 `{ reopened: false, dismissedByUser: true }`(自愈行随之说「卡是你关的」)。本版起这七形的判决与卡帧序列与日志口不抛时逐字节同。`hostLog` 的返回值随之由 `void` 改为 `boolean`:`false` = 宿主日志口这一次抛了错(已吞,这一行没递到);没装口或正常递到 ⇒ `true` —— 口抛错原是调用方唯一看得见的递送失败信号,吞掉之后改由返回值交出(本包「同一件事只警告一次」的去重记号据此撤回,留痕机会不被一次坏口烧掉);`hostProbe` 仍返回 `void`。射程:只兜宿主口本身抛出的错,调用点拼诊断句的实参不在内;投影内核的上下文 `AdapterContext` 上的那一对 `log` / `probe` 是另一对口,本版未动。接入文档 **§110 PS-7**。
|
|
120
|
+
|
|
121
|
+
### Gates
|
|
122
|
+
|
|
123
|
+
- 扩门 `scripts/run-task-request-omission-receipt-test.mjs`(445 → 508 格;与终端成因句表对照的那一段只在给了终端源码树时跑,读不到时 506 格、总结行标为未跑):L 段 —— 导出口在场且句表冻结、键 ≡ 成因闭集、取句口答的就是表里那一句、每词一句句句不同、句子挂对了词(每句带自己那个成因的语义锚、不带别的成因的锚)、与接入方那张逐词不同句;措辞恒以 `Not sent` 起头、零接入方用词(判别式先对接入方四句句句响作正控)、不说生效、过出包卫生词表;表外词 / 近形 / 非串 / 包装串 / 数组 / 会抛的 `toString` 与 `Symbol.toPrimitive` / 每个陷阱都抛的 Proxy / 已撤销的 Proxy / 十个原型链名 / 被污染的原型名 ⇒ 通用句、不抛、不回显;就地改写句表被拒;源码级 —— `as const satisfies Record<TaskRequestOmissionCause, string>` 挂在表字面量自己身上(删一句 / 多一句即编译红),源码里以成因词为键的对象字面量恰三处(判官表 + 两张句表,第三张即红);终端成因句表的每个成因词都在本包成因闭集里、且不重复(终端有、本包无即红;本包有、终端无只出读数;措辞只出读数不判;那张表退役或改名后该段如实标为未跑)。K4c —— 接入方那张句表的字面量同样挂 `satisfies Record<TaskRequestOmissionCause, string>`:此前它只有声明注解,少一句编译红、多一句却照样编译通过(只被运行期格抓到);本版起两个方向都在构建期红。出包字节不变。
|
|
124
|
+
- 同名影子对账门 `scripts/run-layering-shadow-export-test.mjs`:终端那张成因句表的语义差分格改为对位终端用户向取句口 `taskRequestOmissionCauseNotice`(此前对位接入方那只 —— 那一只写给接入方、带改法,不是端句的对位);登记行同批改名,登记的分歧类不变。
|
|
125
|
+
- 扩门 `scripts/run-plan-review-dismissal-test.mjs`(94 → 133 格)N 组(CC-258):会话没有订阅者 ⇒ 同步形 / 回执形 / 带触发者都答 `{ reopened: false, _sema_noPresentationSurface: true }`(键集恰两位)、零帧;按会话键问(别的会话有订阅不算;本会话装上即照开,判决形与旧版同);不带位的六形逐格恰等(有订阅的成功形 / 回执窗没落定 / 决断在飞〔有无订阅同答〕/ 用户关卡 / 非默认会话键没有投递口 / 关卡记号已不在);同一枚关卡记号上订阅没了 ⇒ 答没有界面(不说「卡是你关的」)—— 入口判时,与异步判定读在路上时订阅没了(读后复核与入口同序,0.85.0 在这一形答「卡是你关的」),订阅装回后同一道门照旧答用户关卡;异步判定读门实例证出新门、读在路上时订阅没了 ⇒ 读后复核带位、零新卡,订阅装回后同一道新门照开,且用户对旧门的关卡记号已随那一次撤掉(装回后同步形自动重开 / 回执形门实例读口答 null ⇒ 照开出卡、零次判定读);读后复核的判定序(在飞 → 陈旧评估 → 界面 → 用户关卡 / 新门):读在路上时交出了决断(台账登记 / 宿主交出谓词翻真)又退订 ⇒ 仍答在飞,关卡记号换了一枚 / 没了又退订 ⇒ 与不退订时逐字节同答陈旧评估那一形;宿主装的日志口对每一级都抛时,同步形 / 回执形 / 读后复核三形的无界面判决不变、零帧,既有七形(入口决断在飞 / 同步形用户关卡 / 入口铸卡段 / 读后复核四形)的判决与卡帧序列与日志口不抛时逐字节同,公面 `hostLog` / `hostProbe` 在宿主口抛错时不抛;首呈口无订阅的今日读数(答 `true`、零帧,不是承诺)。
|
|
126
|
+
- 扩门 `scripts/run-selfheal-reopen-test.mjs`(513 → 595 格)。判决读取收进保护(+59):判决上本包读的八位逐位换成会抛的取值器 × plan 臂 / ask 臂 / steer `queued` 两臂 / 宿主自铸的 `queued` 结局 ⇒ 结局照常落定、与判决只缺这一位时逐字节同,用户形 / 注入形行句不抛且逐字同;两种 Proxy(取值全抛 / 列键抛)同律;宿主自铸结局漏了 `reopened` ⇒ 与 `null` 同句;普通对象判决(含宿主加的超集键)进 `queued` 结局逐字节同(含键序、值为 `undefined` 的自有键);超集键的取值器抛 ⇒ 只跳过那一键;自有键恰叫 `__proto__` 的判决照原样拷、不改拷贝的原型;带 Symbol 键 / null 原型的判决进 `queued` 结局,按 `Reflect.ownKeys`、逐键值与原型对拍与原物一致;不可枚举的自有键不拷、可枚举性判定本身抛的那一键跳过;宿主自铸 `queued` 结局上 `reopened` 这一键本身读不出(取值器 / Proxy)⇒ 两形行句不抛、与 `null` 同句。CC-258(+15):包内重开口无订阅 ⇒ 结局带 `noPresentationSurface`;五句各与「不带位那一句只差这一句成因」逐字相等;带宿主「保会话」出路时同律;ask 臂同判;含糊值(`'true'` / `1` / `false`)与在飞 / 用户关卡 / 不带位的拒开既不带键也不含这一句;装上订阅 ⇒ 结局 `plan-review-reopened`。CC-259 G7b(+8):非默认会话键端到端 —— 订阅 `onQuestionFrameFor(K, …)` → 回执形重开 → 收帧后用新口登记 ⇒ `{ reopened: true, firstSight: true }`,卡帧只落在这个会话、默认会话零帧,呈现台账记在这个会话的卡身份上;未作答再开 ⇒ `firstSight: false`;决断性作答之后(宿主不另报记账)再开 ⇒ `firstSight` 回到 `true`;不作答地拿开卡(Esc)之后再开 ⇒ 仍 `false`、零投递;对照:同一会话键改用默认会话口登记 ⇒ 恰 `{ reopened: false }`(卡已交给订阅者)。另打一行读数:改用 `registerArmedGateFor(K, 原始 id)` 时两次重开都答首见(只作读数,不当判据)。
|
|
127
|
+
- 扩门 `scripts/run-terminal-identity-copy-test.mjs`(987 → 1041 条;CC-258):注入形事实表 13 → 15 句(无界面形的未送达与 `queued` 各一句),逐字钉、零内部词扫、承诺对账覆盖新样本。
|
|
128
|
+
- 导出存活门 `scripts/run-export-liveness-test.mjs` 新增 G′(CC-259):已退出公面的名字重新上公面只许经 `scripts/export-liveness.json` 的 `revived` 点名账,逐行同批成立(退出那一版的账里真有它 / 复活版晚于退出版且不晚于本文件顶段 / 票号与理由齐全 / 已回到公面基线);成立的名字才从「不许回到基线」那一格里减掉,退出账照原样保留(判官先对合成账逐形自证)。冻结账门 `scripts/run-integration-doc-freshness-test.mjs` ⑧c 读同一份点名账,把复活的名字从「已退役名」集合里减掉。同门新增 G″:退出账只增不删有了对历史的校验 —— 与**预期上一发布版**(冻结账门那本冻结账里最新一条已转正的版本)那枚 tag 上的同一份账逐版逐名比,旧账每个版本都在、每个名字仍在同一版本下(不许删名、不许挪版本、不许删整版),新增只许追加;修前只查结构,删掉一个历史名(连同它的复活行)再把它加回基线,G / G′ 都不红。基线只认预期那一枚 tag:它不在本地(浅克隆没取 tag / 发布后还没打)、或它上面没有这份账 ⇒ 这一格如实标未核(总结行 PARTIAL,不当绿),不退到别的 tag;不在 git 仓同样标未核;git 读不了 ⇒ 退出码 2;判官(选基线与判缩水两段)先对合成输入逐形自证。负控套 `scripts/run-gate-negative-controls-test.mjs` 为这道门补两枚常驻负控:从退出账最早一版、最近一版各摘掉一个没有复活行的名字 ⇒ G″ 必红且点名(删最近一版那一枚只有基线确是预期上一发布版时才红)。
|
|
129
|
+
- 登记物:根公面基线 1307 → 1310(+3);型面门 unknown 出境棘轮 430 → 432(CC-248 取句口入参与接入方那只同为 `unknown`,在函数内窄化、出境是串;CC-259 新口的 `questionId` 与同文件默认会话口同为 `unknown`,函数内先判非空串再归一,出境 void;沿革账 +2 条);失败显式门的已注释豁免棘轮 44 → 43(宿主诊断口改由唯一调用出口带值返回之后,子代台账缺席留痕那一处包住日志调用、靠捕获异常撤回去重记号的吞随之退役;沿革账 +1 条);单例清单 536 → 538 条(CC-248 句表;自愈腿判决快照按它逐位读的位名表);超集台账 `docs/type-superset.json` 90 → 92 条(CC-258 `_sema_noPresentationSurface`:判决形声明处与唯一铸点处各一);`scripts/export-liveness.json` 新增 `revived` 账一行(CC-259);门清单 `scripts/gates-manifest.json` 里回执省略、plan-review 关卡、自愈重开、导出存活四门的说明(导出存活门另有负控格)与 README「Guards」派生表同批更新;门数 158 不变。
|
|
130
|
+
|
|
131
|
+
### Known limits(本版新增)
|
|
132
|
+
|
|
133
|
+
- 取句口的通用句仍以「Not sent with this request」起头:它信任调用方交来的是回执省略行的成因,拿一个不是回执成因的值来问,它照样说这一项没发出去(真回执每一行必带四词之一,走不到通用句)(KL-259)。
|
|
134
|
+
- `other_channel` 那一句说不出另一步最终发没发、发了什么:本包不知道宿主在构造之后调没调补位那一层、给没给值(KL-260)。
|
|
135
|
+
- `off_lane` 那一句不点名哪一种请求带这一项:车道名是接入方词,端知道自己的入口形,可在句前后补上(KL-261)。
|
|
136
|
+
- 四句与通用句只有英文一份,不随界面语言切换(KL-262)。
|
|
137
|
+
- 门那一侧:与终端成因句表的对照只读终端源码树上的一个文件、只比词集;别的端不读;终端改名或挪走那只文件时如实标为未跑,分不出是退役还是搬走(KL-263)。
|
|
138
|
+
- 首呈口 `armPlanReviewApproval` 不问有没有界面:没有订阅者时照旧答 `true`(布尔,带不了成因位),卡帧不会送到任何地方,作答口照登记;重开口此时带位,首呈口没有对位。宿主要判可在调用前问 `hasQuestionOverlayFor`;之后装上订阅,重开口照常出卡(KL-283)。
|
|
139
|
+
- 陈旧评估(异步判定读在路上时关卡记号换了一枚)按此刻的记号答、不问界面:同一任务另一次自动重开因订阅没了没上屏、记号随之撤掉时,这一次答不带位的 `{ reopened: false }` —— 这一位缺席不证明会话有界面(KL-284)。
|
|
140
|
+
- 门那一侧:导出存活门的「退出账只增不删」只对预期上一发布版那枚 tag 上的同一份账比,同一个未发版本里先追加又删掉的名字看不见;不在 git 仓、本地缺那枚 tag(浅克隆 / 发布后还没打)、或那枚 tag 上没有这份账时如实标未核,不退到别的 tag(KL-286)。
|
|
141
|
+
- 完整台账见接入文档 §110 末行「包侧缺口」。
|
|
142
|
+
|
|
52
143
|
## 0.85.0(2026-09-29)
|
|
53
144
|
|
|
54
145
|
> 主题:**minor** —— 十件同发。① 引擎事实表(通告码册 / 受众表 / MCP 注入丢弃原因,以及会话规则翻译判定内部用的协议前缀与退役工具名)改为**构建期生成**,不再手抄;出身词与强制词两张词表直接就是 SDK 导出的那两个数组;包内同一事实的多份字面收成一份。② 子代族八个入口与两只行停止门收**注入 client**(浏览器同源中继宿主的正位入口;plan-review 那一半 0.84.1 已到),三只写动词的失败结局加「可能已送到、却没有答复」机读位。③ 会话规则无损判定收**可选工具名册**(只收紧)。④ 过渡物**退役登记**与到期门,本版执行其中三条;按 0.84.0 的预告**三个运行期名字退出根入口**,第四个撤回退役预告。⑤ plan-review **决断台账**收进包,重开口按台账拒开。⑥ 会话后台任务**停止口**收进包转调(此前登为「端直调」,而端不直连 sdk),附受据三态归类口。⑦ 服务端 7.104.0 **会话侧**提货:会话存档读不出(`corrupt_session`)三处载体同一判定与一句人话、跳过记录的通告、删会话 409 按码读、`/decide` 停驻挪动族按码读。⑧ 服务端 7.104.0 **能力 / 同步体侧**提货:同步 park 新体与 park 行 `result`、文件历史捕获两代词与「这一次别捕获」请求词、公钥发现段、折叠粗码。⑨ 同名影子对账门的豁免改按语义分类(只改门)。⑩ 导出存活门的端消费证据改读已提交历史(只改门)。根公面运行期导出 1275 → 1307(+35 −3);测试钩 59 → 60;公面类型 857 → 891(+34);超集键零增减;零新投影臂;请求键 +1 `fileHistory`(只在宿主经片段口给值时上 wire);peer sdk 地板 `>=12.0.1` 不动(服务端 7.104.0 与 sdk 13 的新形一律按结构读,不 import sdk 13 才有的型 / 值);开发依赖引擎 `~7.33.1` 不变。🔴 **型面 BREAKING 一处**(三个运行期名字退出根入口,见 Removed);另有几处**已发导出的可观察变化**,Changed 开头逐条单列、写明谁要跟。
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.85.
|
|
38
|
+
**Version:** 0.85.2
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -310,7 +310,7 @@ guard still cross-checks the table by name).
|
|
|
310
310
|
| `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
|
|
311
311
|
| `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
|
|
312
312
|
| `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key on the success result: the CC-spelled `structured_output` is the only home for the value the wire calls `structuredOutput`. The camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rode alongside it for exactly one release (0.79.1) and is **absent from 0.80.0 on**, pinned both by own-key and by `in`, so a consumer still reading the old name sees `undefined` rather than a stale copy. The wire position is read exactly once, so a value-changing accessor is only ever asked for its first answer; absence stays absence; a wire key that is present but `undefined` mints nothing, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` are still minted, and so are shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope carries no such key, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
|
|
313
|
-
| `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the client repositories are on disk the consumption evidence is checked by name, from committed history and never from a working tree: each client's `origin/main` (or its HEAD when there is no such ref), the checked-out HEAD however old it is, and every local branch head committed within the last 14 days, so an import that so far exists only on an unmerged branch still blocks the retirement. Two readings are kept apart. The registry states facts, so a `contract` row's claimed consumers must equal the clients that really bind the name — a named import from the package or one of its subpaths, a re-export from it, a member read on a namespace or dynamically imported module object, a destructured dynamic import — and an `internal` name that a client binds must be re-registered as `contract`; a client-side declaration that merely shares the name is not consumption. The retirement check asks only whether removing the name could break a client, so it also counts a plain identifier match in any file importing the package, and, when a client re-exports the whole package, any appearance of the name in that client's sources; the guard prints both counts. The guard prints the ref, commit and commit date of everything it read, and it shares one repository reader with the same-name shadow guard, so both guards judge the same snapshot of a client. A client that is absent or is not a git repository is reported as a skipped section, never as a client with no consumers, and a dangling or unreadable branch ref is a fault rather than a branch quietly left out. The reader proves itself on a throwaway repository: an import on a recent branch that is not checked out must be found and attributed to its branch and line, while one on a branch older than the window, in an uncommitted file, or outside the scanned folders must not count. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. |
|
|
313
|
+
| `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the client repositories are on disk the consumption evidence is checked by name, from committed history and never from a working tree: each client's `origin/main` (or its HEAD when there is no such ref), the checked-out HEAD however old it is, and every local branch head committed within the last 14 days, so an import that so far exists only on an unmerged branch still blocks the retirement. Two readings are kept apart. The registry states facts, so a `contract` row's claimed consumers must equal the clients that really bind the name — a named import from the package or one of its subpaths, a re-export from it, a member read on a namespace or dynamically imported module object, a destructured dynamic import — and an `internal` name that a client binds must be re-registered as `contract`; a client-side declaration that merely shares the name is not consumption. The retirement check asks only whether removing the name could break a client, so it also counts a plain identifier match in any file importing the package, and, when a client re-exports the whole package, any appearance of the name in that client's sources; the guard prints both counts. The guard prints the ref, commit and commit date of everything it read, and it shares one repository reader with the same-name shadow guard, so both guards judge the same snapshot of a client. A client that is absent or is not a git repository is reported as a skipped section, never as a client with no consumers, and a dangling or unreadable branch ref is a fault rather than a branch quietly left out. The reader proves itself on a throwaway repository: an import on a recent branch that is not checked out must be found and attributed to its branch and line, while one on a branch older than the window, in an uncommitted file, or outside the scanned folders must not count. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. A name may come back only through a separate `revived` ledger that names it, the version it left in, the version it returns in (later than the one it left in and not ahead of the changelog), a ticket and a reason, and only if it is really back in the baseline; the `removed` ledger itself is never rewritten, and the documentation freshness check reads the same `revived` ledger so a returned name is no longer treated as retired. The `removed` ledger is also checked against history: compared with the same file at the tag of the expected previous release (the newest released version in the documentation freshness check's frozen ledger), every version it listed must still be there, with every name it listed still under that same version — entries can only be added. Only that tag is used: outside a git checkout, when that tag is not present locally, or when it does not carry the ledger, the check is reported as skipped rather than passed, never run against some other tag. |
|
|
314
314
|
| `scripts/run-wire-refusal-copy-test.mjs` | Two wire refusals read the same way on every client: a cancel's 409 carries one of two codes with opposite dispositions (`conflict.approval_settled` — someone else already decided, go read the result; `conflict.run_not_running` — nothing changed, send the cancel again), an unrecognised or codeless 409 is reported as such rather than guessed, and the submit-side 429 `usage.window_exhausted` is read as a waitable refusal whose wait is stated only when the engine supplied one. `ControlRouter.cancel` raises a distinct safety code for the retry-directly case. A third family covers the refusals that a deny's *settler note* can draw: a deployment that signs the decisions it accepts but does not sign that note, a body whose decision and note contradict each other, and a word the engine does not recognise. All three are refused **before** the approval is judged, so each sentence states plainly that nothing was decided and the approval is still waiting, and each names a different next step — none of them “send it again unchanged”, which would simply be refused again. Two of the three codes already mean something else in this package: one is shared with the plan-review leg, so the reader refuses to claim it unless the caller states that the note really was sent, and the other is split by the field the engine names, because without that field the same code means a review outcome was rejected for its content. The third sentence deliberately does not say the word was misspelled: upstream mints that same field for at least four different reasons, so the only thing it proves is that the engine would not take the settlement details and named which part — which is what the sentence says, and where it points. The three sentences are pinned verbatim rather than by keyword, because a keyword check passes a sentence that tells the reader to send the same body again unchanged, which is the one next step that is certainly wrong. The field itself is read from the bag the SDK keeps additional response keys in, not off the top of the error, since only the hand-built shapes a test would write carry it there. Both decision legs hand the reading back on their outcome, and a leg that never sent the note claims nothing. When a decision fails because the engine cannot read a stored row, the row the engine names is read (from the error's top level, or from the SDK's bag of additional response keys where the real SDK puts it) and carried into the failure reason on all three decision legs, escaped for display. |
|
|
315
315
|
| `scripts/run-tool-disclosure-progress-projection-test.mjs` | The two wire arms sdk 9.6.0 adds — `tool_disclosure` (name-only tool census: open-set `policy`, `thresholdPercent` absent ≠ default, `deferred`/`activated` full snapshots) and `tool_progress` (one frame, two beats: Bash ticks carry an output tail with `totalLines`/`totalBytes` that come and go together; other tools carry only `elapsedSeconds`) — project to neutral internal arms plus chrome arms. Required keys missing ⇒ `malformed`; bad optional keys drop only themselves; the sub-flow three-key gate keeps child frames off the leader lane; both arms are `required: false` in the arm table with duties stated (the output tail is untrusted raw and must never be fed back to the model). |
|
|
316
316
|
| `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes **0.69.0:** `fetchMcpPanel` fetches the panel through the SDK client's own `sessions.mcp` call (same transport and auth as every other read) and projects it; transport failure, an unreadable body and an empty session id all come back as `undefined`, never as a fabricated empty panel 0.71.0 adds section K: `mcpEngineLegPresence(view)` — the engine-side MCP presence tri-state read only off the panel view (`unknown` when the view could not be read, never rendered as "no MCP configured") |
|
|
@@ -371,7 +371,7 @@ guard still cross-checks the table by name).
|
|
|
371
371
|
| `scripts/run-type-superset-ledger-test.mjs` | The type/wire **superset ledger** (`docs/type-superset.json`): positions this package adds on top of a CC-shaped contract, each carrying the evidence for what CC's own type surface does or does not have there. Completeness is deliberately uneven and the ledger says so. The `_sema_*` private-key class is checked in **both** directions (a key in the source that never entered the ledger is red, naming key and file; a ledger row whose key left the source is red) — but only for keys written as literals, which is the convention the ledger mandates. A key assembled by string arithmetic is beyond what any static rule can enumerate, so the guard fails closed on every shape it *can* decide (a bare `_sema_` prefix is red wherever it appears, save one pinned guard site) and leaves the rest as a convention violation for review to catch, rather than claiming a completeness it does not have. The two hand-surveyed classes are only checked for coordinate and evidence integrity, never discovered. Both directions read the source through the **TypeScript AST**, not a text scan, and they read two different sets out of it. A *key site* is an identifier, or a string whose whole value is the key — so `'_sema_decision-v2'` is carried whole rather than truncated at the first non-identifier character into some *other* key that happens to be registered. A *mention* is the key appearing inside a longer string, which is prose, not usage. The staleness direction counts key sites only: a comment or a doc sentence left behind after the last real mint site is deleted must not keep the row alive (mutation-proven — with both the comment and the prose string untouched, removing the one real site turns the guard red). And because a prefix can be concatenated or interpolated into a key no static set will ever see, the bare `_sema_` literal is refused outright rather than traced: every occurrence is red except the single inline `startsWith` guard the sanitizer needs, because the set of expressions a bare prefix can travel through on its way to a concatenation is open-ended and enumerating it is always one form behind. Every row's `host` must still resolve, with the key being a real **member of that declaration** rather than a string occurring somewhere in the same file — `governanceForced`/`delegation` each live on two different shapes in one file, and a member commented out is a member deleted, which a text-shaped check happily reads as still present. And the direction worth the most: each machine-form `ccAbsenceEvidence` is re-derived from the row's own `key` — the ledger's recorded string must match that derivation verbatim, since a row quietly witnessing `\bnever_present\b` is green forever while watching nothing (mutation-proven: the same edit passes the unbound form and is caught by the bound one) — and the check runs against the names the installed `@sema-agent/agent-types` `.d.ts` set actually declares, parsed with the TypeScript AST rather than grepped, so a name CC merely mentions in a comment cannot force the row into the manual escape hatch and thereby retire the very witness that was supposed to fire the day CC declares that name for real. That escape hatch is gated by an allowlist living **in the guard**, not the ledger, so claiming it costs a reviewed diff. Missing material never reads as a pass, and the verdict splits by *why* it is missing: no TypeScript parser skips the suite before it starts; a missing `agent-types` still runs and prints the first three directions, then exits **1** when `package.json` declares the mirror but it is not installed — a broken install must not retire the repository's only "the day CC declares this name" alarm, and reporting it as a skip would leave "never evaluated" and "evaluated, no drift" indistinguishable to the runner — and exits 3 only when nothing declares the mirror at all, which is the one case where the direction genuinely does not apply. Either way a run that evaluated no witness is never counted as one that did. When the mirror *is* present its **installed version** is witnessed too (the two declared floors must agree with each other and the installed copy must meet them), since four preflight probes are satisfied by an arbitrarily stale mirror — they prove the extractor speaks, not that it is current. Every direction carries a positive control — known-present CC symbols, a comment-only sample proving the extractor distinguishes declaration from mention, and synthetic corpora fed through the **same** discriminator function the real verdict uses, so a verdict quietly rewritten to return nothing takes its own control down with it |
|
|
372
372
|
| `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" `listAllPersistedRules` needs only `list` (`RulesListFacade`; the other four methods are known to the parameter type as optional members): a synthetic consumer compiled against the built declarations passes a two-method object, a list-only literal, a `Pick` slice, the named type, the full facade, and the full or two-method facade written as an inline object literal — including arrow functions with untyped parameters — all with zero diagnostics, while a misspelled method name in such a literal is still reported. |
|
|
373
373
|
| `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
|
|
374
|
-
| `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited. The two "card could not be reopened" rows can carry a host-declared way to keep the conversation, which says the card comes back on resume only if it is still waiting: it is placed before the route that abandons it, never offered for an injected submission, while a decision is still on its way, or once the pending approval has been proven gone (the outcome then carries a flag saying so; the proof only counts before the cleanup card is shown, so a fallback after the card carries no flag unless a fresh read finds the run finished, and a recheck that finds the approval back clears it), the host function is not even called in those cases, and it is treated as unavailable when it throws or returns an empty value; a host can also switch off the engine decide route on the interactive rows while the cancel route stays, and with neither given all four rows are pinned byte-for-byte to the text the previous release produced. When the plan reopen refuses because the user closed that card (`dismissedByUser`), the outcome carries the flag and one package-owned line says so — you closed it, a message brings it back — in the typed, the injected and the queued-steer forms, with no way-to-keep and no engine decide route; the reopen port receives `{ trigger }` mapped from the submission origin (typed ⇒ `user`, injected ⇒ `automatic`, absent ⇒ the old single-argument call). |
|
|
374
|
+
| `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited. The two "card could not be reopened" rows can carry a host-declared way to keep the conversation, which says the card comes back on resume only if it is still waiting: it is placed before the route that abandons it, never offered for an injected submission, while a decision is still on its way, or once the pending approval has been proven gone (the outcome then carries a flag saying so; the proof only counts before the cleanup card is shown, so a fallback after the card carries no flag unless a fresh read finds the run finished, and a recheck that finds the approval back clears it), the host function is not even called in those cases, and it is treated as unavailable when it throws or returns an empty value; a host can also switch off the engine decide route on the interactive rows while the cancel route stays, and with neither given all four rows are pinned byte-for-byte to the text the previous release produced. When the plan reopen refuses because the user closed that card (`dismissedByUser`), the outcome carries the flag and one package-owned line says so — you closed it, a message brings it back — in the typed, the injected and the queued-steer forms, with no way-to-keep and no engine decide route; the reopen port receives `{ trigger }` mapped from the submission origin (typed ⇒ `user`, injected ⇒ `automatic`, absent ⇒ the old single-argument call). When the reopen refuses because no view is subscribed to question frames for that session (`_sema_noPresentationSurface`), the plan and ask outcomes carry `noPresentationSurface`, and the five lines that talk about the card not coming back — both arms' typed line, the queued-steer typed and injected lines, and the injected not-delivered line — add one package-owned clause saying the session has no view that can show decision cards, byte-identical to the plain refusal otherwise; loose values and the other refusals (answer in flight, user-closed card, plain refusal) carry neither the flag nor the clause. For a host that reopens under its own session key with a presentation receipt, the suite drives the whole sequence end to end: subscribing for that key, reopening, and registering the receipt through the per-session call `registerArmedGateFromQuestionIdFor` gives `reopened: true` with the right first-sight reading, a second reopen reads as seen before, a decisive answer (with no extra bookkeeping from the host) makes the next reopen read as first sight again, and a dismissal does not; registering through the default-session call instead leaves the receipt window to run out with `{ reopened: false }`. Every bit the package reads off a host's reopen verdict is read once, inside the guard: a bit whose getter throws (or a proxy that throws) is treated as absent, so the self-heal still settles — with exactly the outcome it would give if that bit were missing — and the rows never throw, including for an outcome the host built itself; a plain verdict lands in the queued-steer outcome byte for byte, key order and any extra keys the host added included, a key whose getter throws is the only one left out, and an own key literally named `__proto__` is copied as a key rather than changing the copy's prototype; symbol keys are copied too and a verdict with a null prototype stays one (compared key by key, value by value and by prototype, not through JSON); and an outcome whose `reopened` property itself cannot be read is phrased as if it had none. The two reopen-failed outcomes carry `errorCode` only when the reopen gave up because a request failed — the outer wire code of that request (a folded policy code is read through `blockedBy`), taken from a reopen port that threw, from the host verdict's `_sema_errorCode`, or from the stale-park arm's own three requests; a later request that changed what the reopen gave up on (the poll finally answered, the row appeared), a failure with no wire code, a reopen port that failed after the caller had already aborted, and the three judgment refusals (answer in flight, user-closed card, no view) all leave it absent, while a read that only confirms (the row is still pending after a failed decision) keeps it; only a request that failed counts, never a field that cannot be read on a record that did come back; and every row is byte-identical with or without it. The error objects the self-heal leg reads (thrown by the run lookup, the steer and the cancel) are read inside the guard too: a `status` or `message` whose getter throws, or a proxy that throws on any read, is treated as absent, so the self-heal still settles instead of rejecting, and the at-most-once failure class reads such an error as delivery unknown; the same goes for the run record a successful lookup returns — when its `status` cannot be read, the self-heal treats it as a record without a status rather than rejecting — and for a steer receipt, whose two readers answer `null` for a field they cannot read. When the delivery of a steer is missing, unreadable or an unknown word, the injected-submission line says the message was accepted but its delivery was not reported and to watch that reply before deciding whether anything needs sending again, instead of the handed-over wording, and it never says nothing is needed, with or without a follow declaration. |
|
|
375
375
|
| `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `_sema_api_error_message` class flag (spelled `isApiErrorMessage` before 0.83.0) so brief-mode visibility filtering does not silently drop it. The second half is who the rejected submission belonged to: the self-heal copy told every caller "Your message was NOT sent … send it again", which is three separate untruths for a system injection (a plan-review outcome, a cron wake-up, a task notification) — not the user's message, and not re-sendable, since a host queue marks those non-editable and non-recallable. The injected form says so instead, and the one sentence that promises re-delivery is pinned to the single disposition that earns it: `selfHealSubmissionDisposition` is the same function the host consults before putting the item back on its queue, so the promise and the behaviour cannot drift apart, and the arms where no card could be surfaced state plainly that nothing was delivered and nothing will retry. Since 0.72.6 the same gate pins the **follow intent** after a steer (): a message handed to a live run only pays off if someone tails that run's own event stream, so `steerFollowIntent` decides from the delivery word whether to tail now, after the pending decision, or only after a wake — and the "watch that run" sentence ("watch that reply" on the rows about follow-up messages sema sent on its own) turns into a factual "sema is following that run" ("… that reply") **only** when the host declares it attached that tail, so a shell that did not wire it can never claim it did. Since 0.83.0 the rows about follow-up messages sema sent on its own carry no engine-internal words and say what happened per delivery shape, promising a resend or "nothing for you to do" only where the code guarantees it |
|
|
376
376
|
| `scripts/run-additive-key-passthrough-test.mjs` | The one disease shape behind two legs: a **closed whitelist / flattening arm** dropping a fact that is already on the wire, while both sides of the seam look correct. (1) The `task_progress` projection carries a registered **key ledger** — a frame populated with every key the service really projects is pushed through the shipped `eventToSdkMessage`, and the set of wire keys that survive must equal the registered pass-through list **name for name in both directions**, so quietly forwarding one more key is as red as quietly dropping one. `model` (the child run's model id, minted by core as `prepared.model.id` and projected by the server since 7.52.1) is the key this batch adds, with the same conditional the server itself applies: a non-empty string or no key at all — an empty string is neither a model id nor "unknown". The ledger is also checked against the fenced list in `docs/INTEGRATION-CLIENTS.md` §3d, so a doc that still says seven keys while the code forwards eight is red rather than merely stale. (2) The decide-failure arms carry the server's S-02 `currentPending` pointer key from a 409 `approval_stale` refusal onto the outcome the host reads. The reader is structural rather than `instanceof`, because the client is host-injected and the class identity is not this package's to assume; a half triple never mints (half a pointer cannot relocate anything), an empty string is not presence, and `checkpointToken` never transits. Both the allow and the deny leg are driven end to end through the real durable approval path — as is the accept-session leg, where a refusal carrying the pointer key must now re-raise instead of silently re-sending the human's answer for the **old** card as a plain approve (one decide call, pointer preserved), while a legacy 400 still falls back exactly as before — and all three flattening points must call the one shared reader — the same-shape residue check that makes "fixed one arm and left the twin" red instead of invisible. (3) The same disease growing on the REQUEST side: the `.mcp.json` → server-spec projection rebuilds each server key by key, and the settings schema deliberately leaves some keys parse-transparent — whatever JSON the file carries reaches the engine untouched, because validating them where the whole domain parses all-or-nothing would let one bad declaration take every server down silently. The whitelist had no row for the newest of them, so an operator's per-tool declarations — the ones the write fence reads — were stripped at the package boundary while both sides looked correct. The criterion is not "is that key handled" but the transparent-key table read out of the INSTALLED schema at runtime, reconciled name-for-name against this leg's ledger, so the day upstream adds a third one this turns red and forces an explicit decision. Behaviour is pinned on both transports, by object identity rather than deep equality (a rebuild would be a second judge), and malformed values must transit UNCHANGED rather than be refused here — the engine refuses them loudly and names the server, whereas a package-side judge can only swallow a declared protection quietly. Absence still mints no key, unknown keys still never reach the wire (the fix is the dropped key, not the gate), and the one transparent key this leg deliberately does not forward is a ledger entry with its own exit condition: it belongs to the deployment plane, and the day the request-plane type declares it the entry's premise is gone and the gate says so |
|
|
377
377
|
| `scripts/run-esc-halt-plan-test.mjs` | The Esc stop decision every client shares: fire the **turn-level** halt first, and escalate to a **run-level** cancel in exactly two cases — the engine itself answered with a 409 from the closed code set (it is saying "there is no in-flight turn here; use cancel for a run-level stop"), or that shot came back with no verdict at all *and* the shell can independently prove a permission card was on screen. Everything else does not escalate. The asymmetry is the whole point and every negative control guards the same direction — deciding *not* to escalate costs the user one more choice on a busy-session card (recoverable), deciding to escalate wrongly tears down a run that was alive and takes every in-flight tool with it (not). So: the closed code set is a **frozen** value, not a `ReadonlySet` — type-level immutability does not stop a consumer's `.add()`, and the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift; the escalation gate is the **conjunction** of that closed set and the 409 status, since honouring the code alone lets a 500 that merely quotes it drive a destructive call; `interrupt.not_held` and `steering.not_running` are deliberately outside the set (the first means *this replica* has no live face — the run may be perfectly alive on another); an unreadable code falls to the no-escalation side; a `parked` flag never overrides a verdict the engine did give, and only strict `true` counts when it did not. The first shot is unconditional by construction — it does not consult `parked`, because the 409 it earns is exactly the verdict the gate wants — and the verdict itself is a closed machine-readable reason word, not display copy. A third escalating case was added once tearing the stream stopped reaping the run: with detach armed, a shot that never lands leaves the run going all the way to the end of the turn, so the Esc the user pressed has no effect at all and nothing on screen says so — the old behaviour had a silent backstop (tearing the stream ended the run) and that backstop is gone. The new fact is held to the same three disciplines as `parked`: it is read only where the engine gave no verdict, it is judged **after** `parked` so an existing host's reason word does not change under it, and only strict `true` counts. Absence is proven to be a no-op rather than asserted — the guard carries its own reference implementation of the previous version's table, runs the full grid through both, requires zero divergence when the new field is omitted, and first shows the comparison really does report a difference on the one cell where the two versions are meant to differ |
|
|
@@ -406,7 +406,7 @@ guard still cross-checks the table by name).
|
|
|
406
406
|
| `scripts/run-memory-verbs-wire-test.mjs` | The five memory-governance verbs as call ports — entry provenance, compliance erasure, the external-origin listing, the clearance ledger and the un-mark valve — on top of the readers above. One failure judge serves all five, and its first question is **provenance, not status**: the engine stamps a machine code on every refusal it mints, so a 501, 405, 409, 404 or 400 that carries **no code** proves nothing about who answered — a proxy or gateway returning the same status may well have passed the request on first — and every such answer is reported as "no verdict" rather than as "nothing happened". Twenty-one coded refusals each get their own arm, branched on the code alone: the status cannot tell them apart (nine different operator actions ride the same 409 here), and conjoining the status would silently demote a refusal the day the engine moved it. A coded 5xx, a coded answer with no status at all, and a coded 4xx this version does not recognise all land in the "cannot tell" arm, because on a non-idempotent verb the default for "could not classify" must be "do not know", never "did not happen". The judge is called from catch blocks, so each of its own property reads is guarded: an error object whose accessors throw is classified, not re-thrown. The capability gate runs before the call and reads the two bits the engine keeps deliberately separate (one for provenance and erasure, one for the origin faces); only an engine that positively says the face is off stops the request, while "this binary does not report that bit" and "this process never saw a capabilities body" both still send — folding "cannot say" into "is not there" is the dishonest-absence shape this package refuses, and these routes answer the capability gate before touching anything. A missing capability reading is a named, explainable error rather than a silent default that would answer for whichever engine happens to be installed. Each verb returns its own discriminated union whose success arm, refusal arm and cannot-tell arm share no keys, so a consumer cannot express "could not read it" as "it worked". The two write legs reuse the erasure and clearance receipt readers rather than minting a second copy, which keeps "this call erased nothing", "a 200 with an empty body" and "a body that could not be read" three separate things here too; the provenance account is narrowed only to its envelope and discriminants and otherwise passes through verbatim, and an account stamped with a newer envelope version is refused rather than reinterpreted. All three receipt-bearing verbs additionally reconcile identity — the account id, the attestation request id and the clearance receipt entry id must be the ones that were sent — because a readable receipt is not yet a receipt about this call. The origin listing carries the server's own echo of the scopes it actually audited, the type has no store-wide arm, and the coverage port takes a mandatory second argument and has no "clean" arm at all: the strongest thing it will say is which scopes were audited. Neither write leg is ever retried, including the refusal whose documented recovery is to send again, because that resend completes whichever clearance row is already open and the audit attribution on it is a person's signature. Request bodies are handed over verbatim — the degraded-erasure authorization is never injected — while the scope list is sent as the snapshot this port validated, so an array that reports one length while being read and another afterwards cannot make "the list I checked" and "the list I sent" two different things. |
|
|
407
407
|
| `scripts/run-dist-orphan-test.mjs` | Every `.js` / `.d.ts` under `dist/` must have a same-named source under `src/`, and every source must have its build output — because the compiler only writes and never deletes, so a module removed from the sources keeps shipping from the previous build (the whole `dist/` directory is on the publish whitelist) while the public-surface gate only looks at what the barrel exports and the hygiene gate only looks at forbidden words. Orphans are named one by one; the pre-publish posture is a clean rebuild, and this gate is the check that the posture was actually followed. |
|
|
408
408
|
| `scripts/run-dist-comments-test.mjs` | **dist ships zero comments.** Since 0.77.2 the build strips comments (`removeComments`); this gate walks every shipped `dist/**/*.js` / `*.d.ts` and counts comment trivia with the TypeScript scanner (string literals containing `//` and generator methods are not comments), failing on the first one (`DIST-COMMENT-FAIL`). Source comments are an internal surface; what still ships is code, string literals and type-level text, which the hygiene gate screens. Negative control: one plain comment appended to `dist/index.js` turns it red. |
|
|
409
|
-
| `scripts/run-task-request-omission-receipt-test.mjs` | Where every key a client hands to the request constructor ends up. The constructor used to answer "not stamped" the same way for four different reasons — value absent, no such row, wrong lane, live gate closed — and a key it had never heard of did not even get that: an unattended run could pass a system prompt, an output schema and a spend cap and receive a body holding the objective and the session id, with nothing anywhere saying what was left out or why. The guard pins the three answers apart. **Seated** keys reach the body verbatim on the unattended lane. Keys the package **knows but did not carry** never throw, never reach the body, and each gets a receipt row with one word from a frozen cause list — every present key is on the body or on the receipt, never both and never neither, checked across both lanes with the live gate open and closed against a key-by-key table written independently of the package's own routing. Keys the package **does not know** are refused loudly and are a separate cell, not a fourth cause: the cause list has no word that could hold them, and the seat reader answers `unknown`, not `none`. The cause list is bitten from both sides (exact, every word producible, nothing produced outside it, the judge table's keys read from source through the TypeScript parser) and no second hand-copied list may exist in `src/`. Upstream claims are read straight off the installed SDK typings: a key seated in this release must be a named request field, a key registered as having no upstream counterpart must not be — the day it appears the guard turns red — and the index signature counts as evidence for nothing. The per-run file-history opt-out word is covered the same way: seated on the two user lanes behind the live gate and absent from the side-channel lane, stamped only for its one legal word, refused loudly for any other value, and read after every existing declaration so that it cannot erase them. |
|
|
409
|
+
| `scripts/run-task-request-omission-receipt-test.mjs` | Where every key a client hands to the request constructor ends up. The constructor used to answer "not stamped" the same way for four different reasons — value absent, no such row, wrong lane, live gate closed — and a key it had never heard of did not even get that: an unattended run could pass a system prompt, an output schema and a spend cap and receive a body holding the objective and the session id, with nothing anywhere saying what was left out or why. The guard pins the three answers apart. **Seated** keys reach the body verbatim on the unattended lane. Keys the package **knows but did not carry** never throw, never reach the body, and each gets a receipt row with one word from a frozen cause list — every present key is on the body or on the receipt, never both and never neither, checked across both lanes with the live gate open and closed against a key-by-key table written independently of the package's own routing. Keys the package **does not know** are refused loudly and are a separate cell, not a fourth cause: the cause list has no word that could hold them, and the seat reader answers `unknown`, not `none`. The cause list is bitten from both sides (exact, every word producible, nothing produced outside it, the judge table's keys read from source through the TypeScript parser) and no second hand-copied list may exist in `src/`. Upstream claims are read straight off the installed SDK typings: a key seated in this release must be a named request field, a key registered as having no upstream counterpart must not be — the day it appears the guard turns red — and the index signature counts as evidence for nothing. The per-run file-history opt-out word is covered the same way: seated on the two user lanes behind the live gate and absent from the side-channel lane, stamped only for its one legal word, refused loudly for any other value, and read after every existing declaration so that it cannot erase them. The wording shown to end users is pinned alongside the wording for integrators: a second sentence table, keyed by the same frozen cause list and exhaustive over it at compile time (the check sits on the table literal itself), gives each cause one plain sentence that says the item was not sent and why (the integrator table now carries the same literal-level check — before, a missing sentence failed the build but an extra one did not), with no integrator vocabulary and no claim about whether anything took effect; every sentence differs from the others and from the integrator wording and passes the published-text hygiene list, and an unknown word, a non-string, a throwing object or a prototype-chain name gets one honest generic sentence instead of a throw, an echo or an invented cause. When a consuming client's own copy of the cause table is on disk, the guard also checks that every word in it is one of the package's causes, with no word listed twice (a word only the client's table has fails; a word only the package has is reported, not judged), and reports that check as skipped rather than passed when the table is not there. |
|
|
410
410
|
| `scripts/run-session-policy-wire-test.mjs` | The per-session tool-rule face: the capability bit that says whether an engine keeps such rules at all, and the narrow read plus tightening orchestration built on it. The bit is read the same four-state way as its sibling capability readers — an absent key is not reported (this binary predates the position itself, which says nothing about whether the face exists), `true` is present, `false` is a positive absent (this deployment keeps no per-session rules), any non-boolean value is unreadable and drops the cell rather than being folded into "absent", and a capabilities body that is not an object at all (an array included) is unreadable rather than "not reported". Its single-source verdict answers whether to show the tightening entry: only an engine that says yes is `yes`, both a positive no and a binary too old to answer are `no`, and never having observed a capabilities body is `unknown`. Whether to put a request on the wire is deliberately a **different** question with a different answer for that last state, and lives with the orchestration. The read narrows three ways that must not collapse into each other: a record that really is empty (present, version zero — what an engine answers for a session no rules were ever written for), a record that cannot be read, and a call that failed with a typed disposition. A half-bad record — one rule bucket well-formed and another the wrong shape — counts as unreadable in full, because the write verb replaces the whole record: dropping the bad bucket and writing the rest back would empty it, which relaxes the rules while the caller sees a 200. An unreadable version stamp is never filled in with a zero, a bucket that reports an implausible number of entries is unreadable rather than walked or truncated, each array's length and each of its indices are read exactly once, and a throwing accessor is unreadable rather than propagated. Every load-bearing key is read as an own property — the envelope, the version stamp, each of the five buckets and each array index — because a prototype lookup would let a polluted prototype put a bucket into the reading that the wire never carried, and since the write replaces the whole record the union would then write that invented restriction back as a real one; a guard pollutes the object and array prototypes in place and proves all four shapes stay out. Because the write replaces the whole record, adding a restriction means writing "what is already there, plus the new entries": the union only ever adds, de-duplicates verbatim, keeps a bucket that is present but empty (present-and-empty and absent are opposite meanings, and dropping it would relax the rules), mints no bucket neither side had, copies entry bytes as they came (no trimming, sorting or path rewriting — those judgements belong to the engine), and takes its bucket names from the engine's own type surface rather than a hand-copied list, so a new bucket upstream is a compile error instead of a silently dropped one. Which differences count as relaxing is the engine's judgement and is never re-implemented here: a refusal on those grounds is reported verbatim, never swallowed and never retried. The orchestration is guarded on three axes. Timing: when the record moves between the read and the write, it re-reads and re-writes **exactly once** — two reads and two writes, no more — and the second attempt's union carries the other writer's entries, which is the entire point of re-reading; a second collision is reported rather than retried a third time, and an uncontended write makes exactly one round trip. Concurrency: an explicit barrier holds both orchestrations first reads at the same version before either may write, and the criterion is how many times the store actually rejected a stale version rather than how many writes it saw — the latter is equally true of two serial successes, so it would stop detecting contention the day the interleaving changed. Under real contention the store rejects exactly once, both writers land on strictly different versions, both writers entries survive in the final record, and the round trips are exactly three reads and three writes; the same two orchestrations run serially are asserted to reject zero times in two reads and two writes, which is what proves those numbers are discriminating. On a store where every write loses the race both report a collision having written exactly twice each. Failure classification: a relaxation refusal, a refusal to stamp a version the store cannot establish (the same status code as the relaxation refusal but a different machine code, and folding it into that arm would send the caller off to edit entries that are not the problem), a missing session, a deployment without the face, an ownerless session, a collision code, a bare conflict with no machine code, a rejected body and an unauthorized call each land on their own arm — the collision arm is matched on the machine code verbatim rather than on the status, because two different situations share that status and only one of them is worth retrying. The remaining split is not "which code is this" but "did the engine answer at all": an answered client-side refusal is allowed to say nothing was written, because every such refusal on this endpoint is emitted before the record is touched, while a throw with no answer at all — a dropped connection, a timeout, a response body the transport itself could not decode, a server fault — can only say "unknown", since that throw may well have happened after the record was already saved. Two guards prove that is not theoretical: a write whose receipt cannot be read, and a write that throws after the fixture store has committed, both leave the record changed. Neither is success nor failure: the only honest answer is "unknown", it carries no version, and it is never retried. The direction of the change is likewise never claimed. The engine’s tighten-only rule is an identity gate, not a field gate — for a principal the deployment treats as an operator it does not run at all, so a union that adds a name to an existing allowlist is accepted and really does widen it. This package does not mint a second copy of that rule, so what it reports is the fact it can stand behind: the record now holds what it already had plus the entries sent here. The sentence for a saved write is pinned to contain no claim of tightening, narrowing or restriction, and a guard reproduces the operator case to prove the widening is real while the wording stays honest. Every sentence the module mints is checked pairwise distinct, with the receipt-unreadable one required to keep its "may already be in effect" and the record-unreadable one required to say nothing was written. |
|
|
411
411
|
| `scripts/run-persisted-rule-write-test.mjs` | The **single-step tightening write** for persisted permission rules — the dual of the revoke surface, and the half where a hopeful reading is expensive. The two behaviours this entry accepts are **derived** from the three-state vocabulary by subtracting the widening one, never hand-copied: the guard bites in both directions (every word in the derived table is really accepted; every constructed outsider — casing variants, trailing whitespace, the widening word itself — is refused before a single round trip), keeps a word-count canary against the parent table, and pins that the source file contains **exactly one** array literal carrying two or more behaviour words, so a second hand-written table shows up as a boundary failure rather than as drift nobody reads. A standing approval is minted by answering a permission question or by importing settings; this entry is not a third route, and the widening word is unspellable in the type. The outcome is a discriminated union whose two failure arms are **not** interchangeable: ten refusal causes each promise the same single thing — not one byte reached the store — while three separate words say the opposite, that the outcome could not be read at all. The service's own "I cannot tell" (a write that could not be confirmed as standing: store wobble, a redemption leg with no decidable ending, or a write that landed and was revoked concurrently before the read-back) stays in the second group, because announcing "nothing was written" invites a clean retry that is not clean, and announcing success misreports a tightening that may already be gone. Anything the shared failure classifier does not recognise defaults to the same place — this is a non-idempotent verb, so "unclassified" must mean "unknown", never "no write": a 500 can happen after the store commits. A 2xx whose body cannot be read is pinned in the same direction and from both sides: it reads as unknown, and the unknown arm carries **neither** the revision nor the written row, so a consumer cannot even spell the shape that would let "unreadable" pass for "written". `persisted` is guarded against the reading everyone reaches for first: it says *this call wrote*, not *a new rule now exists* — an equivalent rule already in the store still mints a fresh causal point, so the lane honestly reports `persisted`, and the material for judging whether the **logical** rule is new (the approval ledger on the returned row) is handed to the caller rather than folded into the discriminant, since the package does not have the one fact that judgement needs. The returned row goes through the **same single narrower** the listing surface uses — proven by running one row corpus through both legs and asserting the two verdicts agree entry for entry (an adversarial pass that forks the listing leg back into an inline copy reds here immediately), plus a source pin that the predicate is defined once and called from exactly the two legs. That sharing is what keeps a row whose behaviour cell is unreadable **visible in both places** rather than hidden by one of them — and the shared narrower is deliberately followed by a second, *different* question only the write leg can ask: is the row that came back **the rule that was just sent**? A rule's identity is a triple, so a receipt missing its behaviour cell, carrying the sibling state, carrying the widening one, or naming another text or another scope is not evidence that the requested tightening is standing; it reads as unknown with its own word, kept distinct from "unreadable" so the two stay tellable apart, while display and derived cells may vary freely. An adversarial pass found both of the gaps this pins: the receipt check that only looked at whether the row was renderable, and a subtler one — pulling the verb off the injected port and calling it bare drops the receiver, so a host that hands over a real resource object (a class instance whose verbs reach the transport through `this`) would see every write throw and be reported as "could not tell", retry after retry, while the package's three other ports call their verbs as methods and work fine. Both are pinned from the failing side: a shorthand-method facade and a class-instance facade must reach the transport and return a real outcome, with the bare-call throw proven to be a real failure mode first. A 405 is split in two, because only the engine's own bare code is evidence about **the engine**: with it, the path exists and this verb does not, so this worker predates the verb; without it — an absent, empty, or foreign code, which is what a proxy or gateway blocking the method typically returns as HTML or an empty body — what was seen is that the verb was refused, while **who** refused it and **at which hop** is unknown, so it lands in the unreadable-outcome arm with its own word rather than sending someone to upgrade a worker that is fine, hiding an entry that is live, or — the part a second adversarial pass insisted on — promising that nothing was written. That promise is what the refusal group means, and a middlebox is free to forward the request and only then answer 405 on its own policy, so a caller who skipped reconciliation on that word would leave behind a standing refusal the user believes never took effect; the guard pins exactly that shape, with a double that writes the rule and *then* answers 405, and with the engine's own bare code still landing in the refusal group beside it. (The same passes caught the naive status-only reading and the asymmetry where an empty-string code fell through to a different bucket.) The documented recovery for an unreadable outcome — retry, then reconcile — is pinned to be **ledger-safe** rather than merely asserted: against a double that models the engine's own "is this identity already standing?" question, re-sending the same identity comes back as a no-op with the approval ledger and the bucket revision both unmoved, however many times it is repeated, while two concurrent writers each landing a causal point are both honestly reported as having written. Finally the four local gates are pinned to be free: a missing write verb on the injected port, an unwritable direction, an unreadable identity pair, and a principal key that is present but cannot name anyone all refuse **without sending anything** — the last of those because silently degrading a blank target into absence would land a tightening aimed at someone else in the caller's own bucket and return a 200 |
|
|
412
412
|
| `scripts/run-registrar-tables-test.mjs` | The four registrar table bodies — this Guards table and the three census tables in the repository's negative-control record — are **generated** from `scripts/gates-manifest.json`, the one file that describes a suite. Every row's text must equal what the generator emits, the manifest's suite set must equal the suites on disk, each entry must declare how it is negative-controlled (rehearsed, blind, or behavioural, with the census taker itself declared as such since it does not appear in its own tables), and each row's outward prose is scanned against the published-surface word list — the same list the packaging-hygiene guard uses, shared rather than copied — before the generator may write it into this file. Adding a guard is therefore one manifest entry plus one generator run instead of six hand edits across three files, and a description that drifts in one place and not the others stops being expressible. The row count is no longer what is compared: the earlier arrangement checked the census tables by **length**, so rows naming the wrong suites reconciled green. Positive controls run entirely on in-memory copies — a changed description, a dropped entry, an added entry, a changed class and a hand-edited row on disk each have to make the same judgement speak — and the quieter halves are pinned too: nothing outside a table body may move, a line inside one that is not a recognisable row makes the generator refuse rather than drop it, byte equality is backed by a column-count check (a cell holding a bare pipe splits a row into extra columns, and a code span does not protect it), and the malformed rows kept byte-for-byte as they are found are registered individually, so the registration turns red the day it stops being needed rather than outliving its reason — audited in both directions, since a registration pointing at a row that is no longer malformed and one pointing at a guard that was reclassified or deleted are both exemptions nobody reads |
|
|
@@ -431,7 +431,7 @@ guard still cross-checks the table by name).
|
|
|
431
431
|
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. Since 0.84.1 an address-shaped value (`scheme://…`) that sits right after an invisible single character — a format character, a lone surrogate, a non-whitespace control character, or the outlet's own escape token for one — is no longer let through as an address: the scheme stays and the host and path are replaced, with the query and fragment masked as before; `user:<password>@` followed only by stripped units and then a boundary, a port or another `@` is masked as userinfo. A value after real whitespace or after a colour sequence is still treated as an address, and clean text stays unchanged; forms that the previous release masked are checked not to leak on the same outlets. A `user:<password>@` candidate that is the value of a credential label or scheme word — including a label or scheme word split by stripped characters, and a quoted value that goes on past whitespace — is left to the label pass, so the whole value is masked as in the previous release; both reported shapes and frozen samples from the targeted pools are checked on every outlet. |
|
|
432
432
|
| `scripts/run-ask-survives-posture-test.mjs` | The single posture predicate `askSurvivesPosture(card, facts)` for sessions whose standing mode would otherwise answer approval cards on the user's behalf (bypass-style modes). It reads two facts and returns one of three verdicts. The first is the ask origin stamped on the card: the question tool (`content_question`), an organization rule (`org_rule`), a hook (`hook`), an explicit ask rule (`ask_rule`), an organization policy or rule store that could not be read (`org_unavailable`, `rule_store_unavailable`) and the classifier's hand-off after its denial limit (`denial_limit_fallback`) must still be asked (the engine requires a real person to answer all three) and every other origin this build knows is left to the posture only once the host has also reported that its own ask rules did not match. The second is the host's own reading of its settings ask rules for this call: a positive match must be asked, and a command the host could not fully parse counts as no match. When the host reported no reading, every card outside those seven origins gets `unknown`, because an origin says who asked and not that the user's own ask rules did not match; an origin this build does not know gets `unknown` even after a reported non-match. `unknown` is never an approval: the host falls back to its own settings rules. The guard checks the verdict for every origin word, both with no host reading and with a reported non-match, against an independent table whose word set must equal the package's origin list, so a new upstream word fails the guard until it is classified; it covers the combinations of both facts, malformed inputs (non-boolean readings, empty or non-string origins, prototype keys, a different letter case), the fact that the predicate does not read the stronger bits on the card (those stay with the host's earlier checks), real card requests produced by the live-frame, parked-row and suspended-ask paths, and a closed, frozen verdict shape. Since 0.84.0 the product table itself is also checked against the SDK's runtime list of known origin words, so a word the upstream does not know cannot sit in the table. |
|
|
433
433
|
| `scripts/run-engine-agent-absence-projection-test.mjs` | Absent background agents: when the engine stops reporting a background agent and no final state has arrived, the row is marked absent and this package owns every decision about it, so all clients agree. One predicate says whether a row is absent (the mark, not the status, decides). An absent row keeps its last known status, never counts as running, and is never counted as completed, failed or stopped; its elapsed time stops at the last moment it was seen, and its sentence says it may still be running. The end-of-turn sweep never settles an absent row (or a resident one). A row that comes back, or a real final state for the current cycle, clears the mark; a late final state from an earlier cycle does not. Absent rows are never removed at the short grace window. After the hard limit (30 minutes from the last time they were seen) the host is asked for the background-agent registry reading of each row: only a reading that the agent has ended or is not listed lets the row go, and each removal is returned as a fact the host must act on and announce; a reading of running, unknown, missing or unrecognised keeps the row and schedules nothing, so no standing poll is created. Until a registry reading is available every absent row stays. A row someone is viewing is held and reported separately only once the registry confirms it is gone. The row sentence, the removal sentence and the late-result sentence come from one place, never state an outcome or that the agent finished, and escape control characters in names, in the engine's removal word and in the late-result status. An end-to-end cell drives the real fleet projection and the real absence channel through every decision. From 0.84.0 the package also reads the registry itself: a reader lists the session's background registry only when the server advertises the listing, classifies its failures as unavailable (a failed read keeps the HTTP status and error code when the error carries them, and never invents either), and refuses a partly readable listing as a whole; a classifier turns the listing into the per-row reading, treating a registry status it cannot place as unknown and reading "not listed" only for keys shaped like registry handles and only when the host states the engine is a single process and names the row's session, because a key from another identity space proves nothing by being absent and, on a multi-replica deployment, a listing answered by another replica does not contain this session's agents at all (without that statement a missing row reads unknown and the row stays); and a predicate returns the agents the registry still counts as live that the host has no row for. A listing at the server's 500-row cap cannot show that an agent is gone either, so a missing row there reads unknown; the cap is checked against the engine's own list clamp and, when a server build is supplied, against the server's route. Each listing carries its session and a per-client sequence number, and a listing superseded by a later delivered read, or read for another session than the one the host names, counts for nothing. Only the exact listing object the reader returned can authorize removing a row or adding one: a spread copy, a structured clone, a filtered or a hand-built listing reads unknown and adds nothing, the listing, its row array and every row are frozen, the cap check uses the row count recorded when the page was read, and rewriting a listing's sequence number fools nothing. The fill predicate holds back agents first seen within the absence settle window — timed on one monotonic local clock from when this client's reader first saw the id, so a skew between the server's clock and this one, or reconnecting to a long-running agent, cannot skip the window — agents without a readable registration time and agents the host saw end since the read went out; the registry key of a row, progress or absence event is whichever of its two ids is shaped like a registry handle; and none of the reader, the classifier or the predicate throws. |
|
|
434
|
-
| `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. |
|
|
434
|
+
| `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. A reopen for a session with no view subscribed to question frames (`onQuestionFrameFor(sessionKey, handler)`; `onQuestionFrame` for the default session) refuses with `{ reopened: false, _sema_noPresentationSurface: true }` — in the synchronous and the receipt form alike, per session key, and also when the view goes away while the asynchronous check is reading — and that answer comes before the user-closed check, at the entry and after the asynchronous read alike, whether or not the read proved a newer gate (after the read, a decision handed over while the read was in progress, and a dismissal record that was replaced or cleared meanwhile, still take precedence over the missing view); a host log or probe port that throws changes none of the reopen answers (the suite runs the same arrangement with a quiet and with a throwing log port and requires the same answer and the same card frames); every other refusal (receipt window lapsed, decision in flight, user-closed card, non-default slot without a delivery path, record gone) keeps its own keys without the flag. The first-presentation call does not check for a view: without one it still answers `true` and the card frame goes nowhere, which the suite records as today's behaviour, not as a guarantee. |
|
|
435
435
|
| `scripts/run-gate-interrupt-safety-test.mjs` | The two safety properties of running the suites themselves. The collecting runner no longer kills a suite outright when its time limit expires: it sends a termination signal first, waits a grace period for the suite to clean up, and only then kills it — and it records the suite as timed out however it exits, so a suite that exits cleanly after the signal is still not counted as passing. The limit and the grace period can be widened for a single suite in `gates-manifest.json` (an optional entry with a written reason; a malformed entry, including one with a misspelled field, stops the runner before any suite starts instead of silently falling back to the default, and an entry filed under a misspelled name is reported by this guard rather than ignored), and the negative-control suite is widened there. Every property is exercised on byte-for-byte copies of the real runner and the real negative-control suite in a throwaway directory: a suite that honours the signal finishes within the grace period, one that ignores it is killed when the grace period ends, and the summary lines are unchanged; once a suite has exited the runner waits at most a short drain window for its output pipes, so a child process that inherited them and outlives the suite neither turns a passing suite into a timeout nor holds the runner past the limit, the grace period and that window; the negative-control suite, interrupted in the middle of a rehearsal by any of the three signals or by the runner's own time limit, restores the file byte-for-byte, leaves no backup behind, stops the rehearsed guard together with anything it started, regenerates the build output (checked on a small project in the throwaway directory), and exits with 128 plus the signal number; a rebuild that would overrun the grace period is cut short and its compiler stopped; a backup left behind by an earlier run — next to a later target, loose in the source tree, or inside a symlinked dependency directory — makes it refuse to start without touching anything, naming the file and how to restore it. |
|
|
436
436
|
| `scripts/run-system-reminder-open-tag-test.mjs` | The opening-tag locator for the engine's `<system-reminder>` envelope, for hosts that split a message into envelopes and body text themselves rather than only stripping or unwrapping it. `findSystemReminderOpenTag(text, from?)` returns the start and end of the next opening tag at or after `from` (UTF-16 indices, `end` one past the tag) or `null`, and it recognises exactly the two shapes the engine mints — a bare tag, or a tag with a single `mark` attribute whose value is 22 base64url characters — through the very same matcher the strip and unwrap entry points use, so there is no second grammar to drift. The guard pins both engine shapes to the exact index, refuses the non-engine shapes listed in the integration guide plus near relatives (and does not let them swallow a real tag that follows), refuses an opening tag truncated at the end of the text, answers `null` without throwing for a non-string text and for a `from` outside the integer range 0..length, and shows that the locator ignores block context (a tag inside a block body, an unclosed tag and a nested inner tag are all located — pairing with a close is the caller's loop, as in the strip entry point). On twenty thousand seeded random strings mixing both shapes, truncations, near relatives and nesting, the set of tags the locator finds equals the set derived purely from the observable answers of the strip and unwrap entry points, and a strip and an unwrap rebuilt on the locator agree with the real ones byte for byte. Two timing cells show a single call over many near-miss tags and a full walk with `from` moving forward both stay linear |
|
|
437
437
|
| `scripts/run-detach-durable-off-verdict-test.mjs` | The verdict behind the durable-off hint, now a public entry point without the process-level gate. `isDetachDurableOff400(err)` answers whether an error is the engine's refusal of the detach-on-disconnect header on a deployment with no durable run ledger: status 400 and the engine's own refusal sentence in the message, and nothing else — in particular not the error code that refusal carries, because the engine uses the same generic precondition code for unrelated refusals on the same submit path, and treating those as this one would silently resubmit a turn without the header. Browser and desktop hosts, which never arm the terminal's detach ledger, can now ask the same question instead of keeping their own copy. The guard pins the positive case on the real sentence (with or without the code, with surrounding text, and on the error the SDK actually throws when a stubbed engine answers 400), the negative case on the engine's other refusals that share the code (verbatim, and again through the SDK), non-400 statuses, and eighteen malformed inputs that must answer `false` without throwing. It shows the verdict does not read the process ledger and that the hint still returns nothing before the header has been sent. The verdict never throws: an error object whose `status` or `message` getter throws, a proxy whose trap throws and a revoked proxy all answer `false`, each property is read at most once (and `message` not at all unless the status is 400), and the hint therefore answers nothing for those objects instead of throwing as it did before. Against a frozen copy of the previous hint, every other input gives the same answer byte for byte, armed or not. When the engine's source tree is available, it also checks that exactly one refusal site with that code carries the sentence and that every other one is answered `false` |
|
|
@@ -35,6 +35,8 @@ export type ReopenCardVerdict = {
|
|
|
35
35
|
pendingRowGone?: true;
|
|
36
36
|
_sema_decisionInFlight?: true;
|
|
37
37
|
dismissedByUser?: true;
|
|
38
|
+
_sema_noPresentationSurface?: true;
|
|
39
|
+
_sema_errorCode?: string;
|
|
38
40
|
} | {
|
|
39
41
|
reopened: true;
|
|
40
42
|
firstSight: boolean;
|
|
@@ -86,6 +88,8 @@ export type SelfHealOutcome = {
|
|
|
86
88
|
decisionInFlight?: true;
|
|
87
89
|
pendingRowGone?: true;
|
|
88
90
|
dismissedByUser?: true;
|
|
91
|
+
noPresentationSurface?: true;
|
|
92
|
+
errorCode?: string;
|
|
89
93
|
} | {
|
|
90
94
|
kind: 'not-parked';
|
|
91
95
|
taskId: string;
|
|
@@ -108,6 +112,8 @@ export type SelfHealOutcome = {
|
|
|
108
112
|
decidePath: string | null;
|
|
109
113
|
decisionInFlight?: true;
|
|
110
114
|
pendingRowGone?: true;
|
|
115
|
+
noPresentationSurface?: true;
|
|
116
|
+
errorCode?: string;
|
|
111
117
|
} | {
|
|
112
118
|
kind: 'ask-decided-without-card';
|
|
113
119
|
taskId: string;
|