@sema-agent/client-core 0.78.2 → 0.79.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/CHANGELOG.md +51 -0
  2. package/README.md +24 -19
  3. package/dist/adapt/arms.js +7 -0
  4. package/dist/adapt/textStream.d.ts +1 -0
  5. package/dist/adapt/textStream.js +10 -0
  6. package/dist/adapt/toolCards.d.ts +1 -0
  7. package/dist/adapt/toolCards.js +2 -0
  8. package/dist/adapter/downstream/eventToSdkMessage.js +4 -0
  9. package/dist/adapter/downstream/terminalToSdkResult.d.ts +9 -1
  10. package/dist/adapter/downstream/terminalToSdkResult.js +34 -7
  11. package/dist/adapter/runStream.js +60 -3
  12. package/dist/agentSession/contract.d.ts +2 -1
  13. package/dist/engineNoticeCodes.js +2 -0
  14. package/dist/gateVocabulary.d.ts +6 -0
  15. package/dist/gateVocabulary.js +21 -0
  16. package/dist/hitl/askGateWire.d.ts +3 -0
  17. package/dist/hitl/askGateWire.js +2 -0
  18. package/dist/hitl/frameRouter.d.ts +1 -0
  19. package/dist/hitl/frameRouter.js +39 -8
  20. package/dist/hitl/gateLedger.d.ts +9 -4
  21. package/dist/hitl/gateLedger.js +29 -5
  22. package/dist/hitl/parkResolver.js +6 -1
  23. package/dist/hitl/persistedRulesWire.d.ts +12 -1
  24. package/dist/hitl/persistedRulesWire.js +20 -5
  25. package/dist/hitl/toolApprovalWire.d.ts +15 -1
  26. package/dist/hitl/toolApprovalWire.js +56 -4
  27. package/dist/index.d.ts +2 -0
  28. package/dist/index.js +2 -0
  29. package/dist/mcpLiveness.js +3 -0
  30. package/dist/mcpProbeCapability.d.ts +18 -0
  31. package/dist/mcpProbeCapability.js +55 -0
  32. package/dist/mcpProbeWire.d.ts +51 -0
  33. package/dist/mcpProbeWire.js +217 -0
  34. package/dist/printToolResultFrame.d.ts +1 -0
  35. package/dist/printToolResultFrame.js +2 -0
  36. package/dist/request/taskRequest.d.ts +20 -1
  37. package/dist/request/taskRequest.js +198 -52
  38. package/dist/sdkRegistryTransit.d.ts +6 -0
  39. package/dist/sdkRegistryTransit.js +5 -0
  40. package/dist/sdkWireTransit.d.ts +2 -1
  41. package/dist/sdkWireTransit.js +1 -0
  42. package/dist/toolResult.d.ts +15 -1
  43. package/dist/toolResult.js +61 -4
  44. package/dist/wireFailureShape.d.ts +1 -0
  45. package/dist/wireFailureShape.js +44 -10
  46. package/docs/INTEGRATION-CLIENTS.md +209 -15
  47. package/package.json +8 -4
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.78.2
38
+ **Version:** 0.79.1
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -67,7 +67,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
67
67
  against — the tables live upstream precisely so this package does not keep a second copy that can
68
68
  fall behind. The browser bundle really bundles the SDK through (the portability guard would
69
69
  exit 3 rather than quietly mark it external).
70
- - The declared floor is `>=11.0.1` (raised from `>=9.8.1` in 0.78.0: `DeniedBy` carries its tenth word `read_boundary`, `rules.write` answers a `stillLive`-discriminated body, `RemovalLiveness` / `RuleWriteRequest` / `RuleWriteResult` / `RuleWriteBehavior` are exported from the SDK root and `TaskRequest.excludeAllTools` is typed from 11.x on, the package now compiles against 11.0.1, and no consumer ships 9.8.x any more, so the older floor lost its witness; before that raised from `>=9.7.1` in 0.75.0: `Capabilities.deviceExecutor.management` is typed from 9.8.x on, the package now compiles against 9.8.1, and no consumer ships 9.7.x any more, so the older floor lost its witness; before that raised from `>=9.6.0` in 0.74.0: `Capabilities.approvalsStreamLive` / `.executionLane`, `LivePendingRow.frame`, the `live_*` approval-stream events and `gates[].toolCallId` are typed from 9.7.x on, and no consumer ships 9.6.0 any more, so the older floor lost its witness; before that raised from `>=9.4.0` in 0.71.0: the `tool_disclosure` / `tool_progress` frames and `ToolApprovalFrame.readRootCandidate` are typed there; earlier: raised from `>=8.8.0` in 0.69.0: the `reasoning_end` frame and `McpStatusPanel.lastLegMcp` are typed from 9.4.0 on), and it is *witnessed*: the guard checks that an actually
70
+ - The declared floor is `>=11.2.1` (raised from `>=11.0.1` in 0.79.0: `capabilities.mcpProbe`, the MCP probe face (`mcpCapabilities` / `probeMcp` with `McpProbeFace`) and the write receipt's third liveness arm (`stillLive: "unknown"`) are typed from 11.2.x on, 11.2.1 adds `TaskRequest.approverPosture` and the `mandated` approval-frame key to the types and the runtime key anchor, the package now compiles against 11.2.1, and no consumer ships 11.0.x / 11.1.x any more, so the older floor lost its witness; before that raised from `>=9.8.1` in 0.78.0: `DeniedBy` carries its tenth word `read_boundary`, `rules.write` answers a `stillLive`-discriminated body, `RemovalLiveness` / `RuleWriteRequest` / `RuleWriteResult` / `RuleWriteBehavior` are exported from the SDK root and `TaskRequest.excludeAllTools` is typed from 11.x on, the package now compiles against 11.0.1, and no consumer ships 9.8.x any more, so the older floor lost its witness; before that raised from `>=9.7.1` in 0.75.0: `Capabilities.deviceExecutor.management` is typed from 9.8.x on, the package now compiles against 9.8.1, and no consumer ships 9.7.x any more, so the older floor lost its witness; before that raised from `>=9.6.0` in 0.74.0: `Capabilities.approvalsStreamLive` / `.executionLane`, `LivePendingRow.frame`, the `live_*` approval-stream events and `gates[].toolCallId` are typed from 9.7.x on, and no consumer ships 9.6.0 any more, so the older floor lost its witness; before that raised from `>=9.4.0` in 0.71.0: the `tool_disclosure` / `tool_progress` frames and `ToolApprovalFrame.readRootCandidate` are typed there; earlier: raised from `>=8.8.0` in 0.69.0: the `reasoning_end` frame and `McpStatusPanel.lastLegMcp` are typed from 9.4.0 on), and it is *witnessed*: the guard checks that an actually
71
71
  installed SDK at that line still exports every value-level symbol this package imports and still
72
72
  declares `TaskStats.costMicroUsd` (the key `costOrNull` reads). A floor nobody ever ran is a
73
73
  promise, not a contract.
@@ -281,9 +281,10 @@ itself (FAILED names + skipped names + arithmetic reconciliation). All-SKIP repo
281
281
  The suite set is a **name-equality gate**, not a lower bound: the runner cross-checks what it finds
282
282
  on disk against `scripts/gates-manifest.json` in both directions. A registered suite that is missing
283
283
  is red (a guard was deleted or renamed); a suite on disk that is not registered is red (whoever
284
- added it skipped the registration). Adding a guard is therefore three actions in one commit — the
285
- `run-*-test.mjs` file, its row in `gates-manifest.json`, and its row in the table below (the
286
- public-surface guard checks that last one).
284
+ added it skipped the registration). Adding a guard is therefore two actions in one commit — the `run-*-test.mjs` file and its entry in
285
+ `gates-manifest.json` (the row below is generated from that entry by `scripts/gen-registrar-tables.mjs`, and
286
+ `scripts/run-registrar-tables-test.mjs` reds when the table on disk and the manifest disagree; the public-surface
287
+ guard still cross-checks the table by name).
287
288
 
288
289
  | Suite | What it guards |
289
290
  |---|---|
@@ -291,28 +292,28 @@ public-surface guard checks that last one).
291
292
  | `scripts/run-client-core-portability-test.mjs` | Kernel / A-layer / index import closures, the runtime-dependency equality gate, barrel reachability, and a real esbuild `--platform=browser` bundle |
292
293
  | `scripts/run-client-core-diff-test.mjs` | Differential equivalence against the CLI reference bridge + replay-id invariant + ledger round-trip |
293
294
  | `scripts/run-seat-contract-keys-test.mjs` | The seat IPC contract: verb list ↔ SPEC ↔ types, element-wise |
294
- | `scripts/run-approval-frame-keys-test.mjs` | The tool-approval frame key mirror, element-wise against the SDK's runtime anchor (one carve-out: AHEAD_OF_ANCHOR entries — keys the server already emits but the SDK anchor has not caught up to — may lead by one generation; the gate turns red the day the SDK catches up, forcing the entry's removal — the register is occupied again — this time by the rule-store-unreadable key the engine now defines, carrying both the release that minted it and the byte coordinates that prove it, so the lead is a dated record rather than an exemption; its predecessor left the register the other way, by being retired upstream rather than by the anchor catching up) |
295
+ | `scripts/run-approval-frame-keys-test.mjs` | The tool-approval frame key mirror, element-wise against the SDK's runtime anchor (one carve-out: AHEAD_OF_ANCHOR entries — keys the server already emits but the SDK anchor has not caught up to — may lead by one generation; the gate turns red the day the SDK catches up, forcing the entry's removal — the register is occupied again — this time by the bit that says a saved allow rule cannot retire a given approval card, carrying both the release that minted it and the byte coordinates that prove it, so the lead is a dated record rather than an exemption; its predecessor left the register the other way, by being retired upstream rather than by the anchor catching up). Beside the key mirror it now guards three further faces of that bit: the closed word table the durable leg reads it through must be **the very array object the SDK exports**, not a same-looking copy — reference identity, because an equal-contents check still permits a second table that diverges the day upstream adds a member; the one predicate a client is meant to call answers over both legs — the live card's bit and the parked row's absence word, which is all the row carries, since the row has no such bit at all — and answers `false` for a malformed value exactly as its presence-only siblings do, a strictness the upstream mint shares; and the one sentence minted for it must never point the reader at writing a rule, since a rule written in answer to a mandated question can never take effect where it was written. The same bit's key also has to reach the card request itself, which the SDK's card anchor does not list — a fact that arrives at the package boundary and stops there is the shape of defect this file's guards exist to catch |
295
296
  | `scripts/run-segment-authority-single-source-test.mjs` | The authoritative-segment replacement verdict, single-sourced. `text_end.content` and the `text_delta` stream stopped being byte-identical the day the engine started redacting the former through the same filter as the result, so every consumer now has to decide six ways what to do with the segment it has half-emitted — and until this release that decision existed **twice**: once here for the transcript lane, once in the shell for the print lane, hot-fixed a version apart. The verdict is now one pure function both lanes call, and the guard pins it on the quantity that actually decides the outcome: whether the authoritative text still *starts with* the bytes that already left, not whether a flush has happened — the latter is a precondition, and anchoring on it withholds a perfectly ordinary answer. Each of the six forms is checked with its counter-case, the prefix length is pinned to UTF-16 code units against a non-ASCII sample whose UTF-8 byte count differs (slicing by bytes leaves the very thing being redacted on screen), and the withheld-segment ledger is compared by normalised equality rather than substring, because a short redaction marker quoted in an unrelated later answer would otherwise suppress that answer entirely. The same file pins the session-level memory-capture declaration to one mint point — the wire value is a single-member closed set, and a consumer that spells it wrong gets a loud refusal rather than a silently dropped privacy request — and pins the SDK URL/health transit to be the **same function reference**, since wrapping it would discard the one guarantee the transit exists for. A last section strips comments with the TypeScript parser and asserts the second expression has not grown back |
296
- | `scripts/run-print-bash-iserror-test.mjs` | The print lane's Bash `is_error` authority (structured over regex) |
297
+ | `scripts/run-print-bash-iserror-test.mjs` | The print lane's Bash `is_error` authority (structured over regex). A second section pins where the denial classification word lands on this lane: on the message envelope, never inside the tool-result block, because that block is forwarded verbatim to the provider on compaction and a self-minted key there is the shape of an old, real defect. A word outside the upstream table — or an empty string, a non-string, or nothing at all — mints no key rather than a guess, and the word never moves the error flag, because attribution does not decide anything |
297
298
  | `scripts/run-bash-benign-exit-interpretation-test.mjs` | Benign non-zero Bash exits (`returnCodeInterpretation`) stay non-errors across all three derivation arms, and the annotation transits to the card |
298
299
  | `scripts/run-sdk-floor-test.mjs` | The SDK version floor — and, more to the point, that the *installed* type declarations still carry the keys this package reads |
299
300
  | `scripts/run-engine-caps-ledger-test.mjs` | A per-key disposition ledger for `GET /v1/capabilities`. The SDK's `Capabilities` grew from 74 keys to 93 in one release and nothing on the board could see it: this package consumes that table through four synchronous readers, and *nineteen new positions arriving while the package does not move* is exactly the disease shape this repo keeps logging on other axes — the fact is already on the wire, the package boundary is the cell that swallows it, and no client can read it however they write their side. So the ledger is reconciled **element-wise against the SDK interface in both directions**: a key the SDK added with no ledger row is red (someone must classify it), and a row for a key the SDK removed is red too (a registration that no longer does anything). Each row then has to survive its own claim — a `read` row names the source file, and the **code** there (comments stripped) must really mention the key, because prose asserting an alignment is the classic way these guards go hollow; a `not_read` row must have **zero** read sites in the tree, so wiring one up while the ledger still says the package ignores it is red rather than invisible. The census behind those two directions recognises five call shapes, each of which really occurs here — a reader whose base argument carries its own parentheses, a direct `caps.<key>`, a narrowing cast, an own-property read helper, and a `*_CAP` constant — and proves it on fabricated samples first, since a census that recognises one shape reports "nothing here" for the other four. What the guard deliberately does **not** judge is whether a position *ought* to be read: that is a design call, and the ledger only pins that every capability was looked at once by a person and that what they wrote down does not contradict the code |
300
301
  | `scripts/run-sql-engine-capability-test.mjs` | The SQL-posture read face and the four-state capability reader underneath it. One capability cell here carries **four different things**, and each one points an operator somewhere else: nothing has been observed yet in this process (a one-shot doctor run is always in that state), the response arrived but carries no such key (an older engine), the engine explicitly answered `null` — *this deployment has no SQL backend*, which is a **positive fact** rather than an absence — and a full reading. Fold any two together and the screen states something flatly, confidently, and wrongly, so every positive control here is paired with a control pointing the opposite way, and the four sentences the doctor row can print are checked to be pairwise distinct and non-implying. The reading itself is narrowed no tighter than the mint: `txnMode: null` is a **legal value** — two of the three engines always report it that way, and the upstream type note names reading it as "optimistic" as the error — so treating it as malformed would throw away the entire reading for ordinary deployments, which is the same disease this repo logged when a consumer's domain was narrower than the producer's. A response that cannot be parsed **clears** the cell rather than leaving the previous engine's answer in place, and a separate invalidation port exists for the case the generation latch cannot catch — a same-port respawn whose new probe never succeeded, where the stale reading would otherwise be answered as current fact. Untrusted values (the isolation string is read back from a database server variable) are sanitised and bounded before display, and the bound is applied **before** escaping so a visible escape never gets cut in half. Finally the export names are themselves a guard: the shell still carries a copy that is meant to go red on the package's same-named export and be swapped out, so renaming anything here would silently disarm that lock |
301
302
  | `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
302
- | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed | failed | blocked | paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there |
303
+ | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there |
303
304
  | `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
304
305
  | `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails |
305
306
  | `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever |
306
307
  | `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
307
- | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor |
308
+ | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, nine words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because the sets differ in kind: one is genuinely closed on the wire (an out-of-set record is withheld by the engine, so reading one means the record is damaged) while the other is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine) Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer |
308
309
  | `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
309
310
  | `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
310
- | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. |
311
+ | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key pair on the success result: the CC-spelled `structured_output` is the authoritative home for the value the wire calls `structuredOutput`, and the camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rides alongside it for one release with the **same value and the same reference**, so a consumer reading either name gets the same object. The wire position is read exactly once, because two reads let a value-changing accessor mint the two names as two different objects; absence is absence on both names; a wire key that is present but `undefined` mints neither, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` still mint both, and so do shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope mints neither, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
311
312
  | `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the sibling client trees are on disk the consumption evidence is checked by name — a `contract` row's claimed consumers must equal the real set, and a `retire` name must not be imported by any client. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. |
312
313
  | `scripts/run-wire-refusal-copy-test.mjs` | Two wire refusals read the same way on every client: a cancel's 409 carries one of two codes with opposite dispositions (`conflict.approval_settled` — someone else already decided, go read the result; `conflict.run_not_running` — nothing changed, send the cancel again), an unrecognised or codeless 409 is reported as such rather than guessed, and the submit-side 429 `usage.window_exhausted` is read as a waitable refusal whose wait is stated only when the engine supplied one. `ControlRouter.cancel` raises a distinct safety code for the retry-directly case. |
313
314
  | `scripts/run-tool-disclosure-progress-projection-test.mjs` | The two wire arms sdk 9.6.0 adds — `tool_disclosure` (name-only tool census: open-set `policy`, `thresholdPercent` absent ≠ default, `deferred`/`activated` full snapshots) and `tool_progress` (one frame, two beats: Bash ticks carry an output tail with `totalLines`/`totalBytes` that come and go together; other tools carry only `elapsedSeconds`) — project to neutral internal arms plus chrome arms. Required keys missing ⇒ `malformed`; bad optional keys drop only themselves; the sub-flow three-key gate keeps child frames off the leader lane; both arms are `required: false` in the arm table with duties stated (the output tail is untrusted raw and must never be fed back to the model). |
314
315
  | `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes **0.69.0:** `fetchMcpPanel` fetches the panel through the SDK client's own `sessions.mcp` call (same transport and auth as every other read) and projects it; transport failure, an unreadable body and an empty session id all come back as `undefined`, never as a fabricated empty panel 0.71.0 adds section K: `mcpEngineLegPresence(view)` — the engine-side MCP presence tri-state read only off the panel view (`unknown` when the view could not be read, never rendered as "no MCP configured") |
315
- | `scripts/run-absence-fold-census-test.mjs` | A package-wide census of the "absence folded into a positive outcome" defect shape, so that fixing the six sites this release does not merely move the shape somewhere else. The defect is defined by position, not syntax: a fallback position (the unconditional tail return, the `default:` arm, the literal minted when there is nothing to pass on, the value returned from an error path) may only say `unknown` or stay absent, never a positive word. Detection walks the syntax tree of every source file, so comments, strings and multi-line spellings cannot hide or fake a hit, and covers five forms: the right arm of `??` / `||`, the else arm of a ternary, the first return of an explicit `default:`, a `catch` block or `.catch(() => …)` arrow returning a healthy value, and a function whose last statement returns a positive word after other returns. Every remaining hit must be registered with a written reason, an unregistered hit fails the gate naming the file and line, the registered count must equal the real count so a cleared site cannot leave a spare allowance behind, and the gate proves its own teeth behind a fence (a failed self-proof refuses to report any count): each form injected into an in-memory copy must add exactly one hit, two correct spellings are pinned as non-hits, and samples inside comments or strings do not count. It also pins the headline site: the fleet panel projection no longer mints an `end` with `isError: false` on absence |
316
+ | `scripts/run-absence-fold-census-test.mjs` | A package-wide census of the "absence folded into a positive outcome" defect shape, so that fixing the six sites this release does not merely move the shape somewhere else. The defect is defined by position, not syntax: a fallback position (the unconditional tail return, the `default:` arm, the literal minted when there is nothing to pass on, the value returned from an error path) may only say `unknown` or stay absent, never a positive word. Detection walks the syntax tree of every source file, so comments, strings and multi-line spellings cannot hide or fake a hit, and covers five forms: the right arm of `??` / `\|\|`, the else arm of a ternary, the first return of an explicit `default:`, a `catch` block or `.catch(() => …)` arrow returning a healthy value, and a function whose last statement returns a positive word after other returns. Every remaining hit must be registered with a written reason, an unregistered hit fails the gate naming the file and line, the registered count must equal the real count so a cleared site cannot leave a spare allowance behind, and the gate proves its own teeth behind a fence (a failed self-proof refuses to report any count): each form injected into an in-memory copy must add exactly one hit, two correct spellings are pinned as non-hits, and samples inside comments or strings do not count. It also pins the headline site: the fleet panel projection no longer mints an `end` with `isError: false` on absence |
316
317
  | `scripts/run-device-executor-management-capability-test.mjs` | The engine's device-management self-description (`capabilities.deviceExecutor.management`, engine ≥7.88.0), read the same four-state way as its five sibling capability readers: an absent `management` key is reported as not reported (never folded into `false`; an older engine really ships the lane object without it), an absent `deviceExecutor` key is likewise not reported, `deviceExecutor: false` is the lane being absent, presence is judged by own-property not truthiness, the value must be a strict boolean, the tee never throws and drops stale generations, and the package owns the verdict on whether the `/v1/devices` management verbs are usable (`yes` only when present and true, `no` when present-false or lane-absent, otherwise `unknown`) |
317
318
  | `scripts/run-run-cancel-context-test.mjs` | The run record's `cancelContext` side-note (engine ≥7.87.3) read structurally, and the cause of a `turn_aborted{engine_error}` classified from machine-readable evidence only: `cancelled` (code `cancelled`, with the cancel-time context when present) / `engine_error` (any other failure code, passed through verbatim) / `run_still_live` (the record is not terminal — a dropped stream is a client-side fact, not the run's cause) / `unknown` (never guessed). An absent `cancelContext` reads as *not reported*, never as "not cancelled"; `elapsedMs` is never folded to 0. |
318
319
  | `scripts/run-suspended-reopen-projection-test.mjs` | The durable `suspended` event's `reopened` key read as three distinct states — `reopened` (with the engine's code, verbatim), `not_reopened` (an explicit `null`), `unstated` (key absent or unreadable) — and carried on the HITL bridge's active gate (`currentGateReopen()`), re-read on every `suspended` and cleared with the gate. |
@@ -320,7 +321,7 @@ public-surface guard checks that last one).
320
321
  | `scripts/run-prompt-assembled-projection-test.mjs` | The `prompt_assembled` frame (one prepare's prompt-assembly manifest) projected to an internal arm and then to the additive `prompt_assembled` chrome event — the per-section / per-block **character** counts, the mounted tool names and `totalChars`, each key present only when the engine really sent it (the frame's `constitution` is deliberately not carried: no consumer asks for it today, and every published key is a contract to keep). The manifest carries **no token counts** anywhere upstream, so this projection mints none: a token figure derived from characters would be an invented number, and the engine's own estimate lives on `context_usage.sections[].tokens` (same id wordlist, joinable). Bad rows are dropped one by one, and a face that loses every row reads as an absent key rather than an empty array — so an absent face means only "this event carries no readable view of it" (an absent upstream key, an empty array and a fully filtered list all land on the same shape) and is never reported as a diagnosis about the engine. `blocks[].id` and `sections[].id` are two different wordlists with a many-to-one relation, and the token join against `context_usage.sections[].tokens` only holds when both sides carry a section view. A frame with no readable composition key at all is malformed, ids and slots are read as an open set, one chrome event per frame with zero transcript rows, several prepares per task are all handed over (de-duplication — "take the last one" — is the host's move), and the lane is told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). Both entry points obey the same rule: the adapt layer rebuilds every row too, so a host pipeline (or a replayed transcript) that feeds the raw frame straight into `adapt()` cannot smuggle extra keys (`tokens`, digests, aliases), a negative `chars` or a `null` row into the chrome payload, an empty array does not count as a composition face, the identity keys are snapshotted once on both paths (read exactly once each, a throwing accessor rejects the whole frame — reading one twice is what lets an accessor frame land on a different lane on each path), and the two paths are compared verbatim so the two readers cannot drift. |
321
322
  | `scripts/run-compaction-outcome-projection-test.mjs` | The `compaction_outcome` frame (a compaction that did **not** end as compacted: mooted by the task ending, failed, …) projected to an internal arm and then to the additive `compaction_outcome` chrome event — `outcome` required and verbatim (open set), `trigger` / `reason` present only when the engine sent a non-empty string, malformed frames dropped, zero transcript rows, the lane told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). |
322
323
  | `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
323
- | `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string | null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
324
+ | `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string \| null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
324
325
  | `scripts/run-approvals-feed-unknown-test.mjs` | The approvals feed tells three states apart: **N items waiting**, **nothing waiting**, and **this fetch did not come back, so we do not know**. Every way a fetch can fail (the call throwing or rejecting, a body that is not an object, a `livePending` section that is not an array — including the `null` seen in the field, a `pending` that is not an array or holds a malformed row) publishes `{kind:'unknown', why, at, mode}` on the subscription — never an empty snapshot and never silence. Real snapshots carry `kind:'snapshot'`; `snapshot()` still answers only with the last real one (a fact about the past) while `reading()` answers whether it is current (`unobserved` / `present` / `unknown`). Recovery always publishes a real snapshot again, even when the contents are byte-identical to before the failure. An unknown reading is never counted as zero: the awaiting-decision counts read `null`, the view is empty, and the tracker reports no removals, so cards on screen are not retracted for a failed fetch. Retry, backoff and circuit-breaking are unchanged. |
325
326
  | `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. |
326
327
  | `scripts/run-panel-cycle-identity-test.mjs` | The **cycle identity** on agent-panel events and the fleet ledger's **departure read-out**: a background agent may be revived under the same id, so `fleet-row` and `end` events now carry the wire's own `cycleSeq` / `startedAt` when present (absent means the row has no notion of generations, never "generation one"), `isStaleEngineAgentPanelEnd` is the single rule for ignoring a late `end` from a previous cycle (only when both sides carry a comparable identity; absence never drops a real terminal), a changed `cycleSeq` is a new cycle for usage stickiness and buffer coalescing, the notification lane carries `cycleSeq` only when the wire really sent `seq`, and `task_remove` frames reach the host through `onTaskRemoved` with `removeReason` / `cycleSeq` verbatim, a stale previous-generation removal leaving the newer row in place. A terminal row held back because the consumer has no such row yet is also released by the keys it carries itself (its transcript id, or the delegating call id of a subagent already on screen under its wire id), since the key tables are only written once a row has actually been published — the release still goes through the one funnel, so the subagent stays one row; a running frame arriving after such a held terminal row is a stale snapshot when both sides carry a comparable generation and it matches (no event, the held row keeps its final usage), a revival when the frame is provably newer (the held row is dropped), and is treated as a revival when neither side can be compared. A subagent lifecycle event carries `agentType` only from an honest source — the fleet row's own agent type, recorded before it is folded into the row label — and omits the key when there is none, never substituting the display name |
@@ -333,14 +334,14 @@ public-surface guard checks that last one).
333
334
  | `scripts/run-mcp-reconnect-test.mjs` | The in-session MCP re-dial verb (`POST /v1/sessions/:id/mcp/reconnect`, engine ≥7.85.0), consumed. The single discriminant is `outcome` and all three answers are HTTP 200, so the reader branches on the word and never on the status; the `unsupported` answer carries exactly five keys and the reader refuses to invent a zero or an empty list for the four fields the engine did not produce, while `accepted` / `refused` treat those four as required and go malformed when one is missing. The tool roster follows the **presence** of `toolNames` (absent = untouched, empty = withdrawn), the connection record passes `errorCode` through as an open set, and every remote-authored string is sanitised and bounded before display. Failures are classified by `errorCode` alone, a missing code is reported as unknown rather than guessed, the verb never throws, and the request-side guard (non-empty name, ≤190 chars) stops a call that the contract would reject anyway. The capability bit reads absent as "cannot tell" rather than "unavailable", and the one sentence the contract insists every UI carries — that re-dialing is a transaction, not a refresh — is minted here once |
334
335
  | `scripts/run-core-value-ports-test.mjs` | The port-injection seam for ten **engine value-level** facilities (autonomous-loop prompt assembly, permission-rule loosening, tool-policy composition, protocol/retired-name/grammar lookups, rule compilation, the discussion workflow name). This package cannot re-export them (the engine barrel drags Node built-ins into the browser bundle), so it declares the ports and honest-absence readers; a Node host installs the engine's own functions verbatim. The guard pins: every reader returns `undefined` when nothing is installed (never a fabricated empty array or default policy), arguments and results pass through by reference, engine errors propagate unchanged, partial installs read partially, restore functions unwind to the previous bag, and the module source has zero engine imports |
335
336
  | `scripts/run-read-face-posture-projection-test.mjs` | The operator-face `readFace: ReadFacePosture` reader (server >=7.65.0). Three ways of "can't say" are pinned to three different, literal sentences, and none of them may read as "nothing is pinned" — that statement belongs to exactly one case, `face: null`, which is a positive fact reported by the engine, not an absence: not having read an operator response yet, having read one from an engine too old to report the key, and the engine actually saying nothing is pinned are three different next steps for an operator and must not collapse into each other. `source` is read as an open set (the server's closed four words plus an escape hatch) rather than narrowed to an enum, so a new word added upstream is not silently turned into a bad reading. The free-text `note` is sanitized and length-capped before it is ever rendered. A companion pure function flags disagreement between this face and the tenant-facing `capabilities.readFace` — silent only when the two actually agree, honest-absent when either side cannot be read at all, never asserting agreement as a fact. The gate's last leg reads the installed SDK's own `openapi.yaml` directly rather than restating the schema in prose, so the package's leniency cannot quietly drift from the real contract |
336
- | `scripts/run-display-body-test.mjs` | The engine wraps text it hands a model in a fence — an opening marker naming the payload, the payload itself, and a closing marker — so the model reads it as data and not as instructions. That fence is minted and read in one place here, which makes stripping it for a human reader this package's job rather than each shell's: a shell that renders the envelope verbatim is showing a person a defence that was written for a model. The reader answers with a discriminated union — fenced, with the label and the payload, or not fenced, with the text as it came in — and it reaches that answer through the **same** matcher the mint side registers, never a second copy of it; the guard proves that by walking the syntax tree of every source file and requiring exactly one literal carrying the marker text, and by requiring the reader's own body to contain no matcher of its own. Eighteen shapes are run through both entry points and required to agree line for line. Anything the package does not recognise — a near-miss in the wording, a hyphen where the marker has a dash, a different case, an opening marker with no close, a close before an open, a truncated close, or any non-whitespace byte outside the pair — comes back unfenced with the input returned **verbatim**: no guessing, no trimming, no repair, because a half-stripped envelope puts a sentence on screen that nobody wrote. Only the outermost layer is removed, so a nested fence, or one forged inside the payload, survives byte-for-byte in the body — those bytes are part of what the engine said, not part of this protocol. Nothing else is washed: control characters, leading and trailing whitespace and a twenty-thousand-character payload all pass through untouched, and so does the label, because sanitising and length-capping belong to the mint point that puts a string on a screen and a passage of text must not have two launderers. A value that is not text is answered with **nothing at all** rather than with an empty payload: the reader never stringifies it, never calls its `toString`, and never emits `[object Object]`, and it does not hand back a body of zero length either — an empty payload is a real reading (a fence can legitimately wrap nothing, and an empty string is an empty string), so folding "there was no readable text" into it would leave a caller unable to show a degraded line at all. Those three stay apart: no text yields nothing, an empty string yields an unfenced empty payload, and an empty fenced payload yields a fenced one with its label. The `fenced` discriminator is always present on a reading, and the label key exists only on the fenced arm, so a missing label is never rendered as an empty one |
337
+ | `scripts/run-display-body-test.mjs` | The engine wraps text it hands a model in a fence — an opening marker naming the payload, the payload itself, and a closing marker — so the model reads it as data and not as instructions. That fence is minted and read in one place here, which makes stripping it for a human reader this package's job rather than each shell's: a shell that renders the envelope verbatim is showing a person a defence that was written for a model. The reader answers with a discriminated union — fenced, with the label and the payload, or not fenced, with the text as it came in — and it reaches that answer through the **same** matcher the mint side registers, never a second copy of it; the guard proves that by walking the syntax tree of every source file and requiring exactly one literal carrying the marker text, and by requiring the reader's own body to contain no matcher of its own. Eighteen shapes are run through both entry points and required to agree line for line. Anything the package does not recognise — a near-miss in the wording, a hyphen where the marker has a dash, a different case, an opening marker with no close, a close before an open, a truncated close, or any non-whitespace byte outside the pair — comes back unfenced with the input returned **verbatim**: no guessing, no trimming, no repair, because a half-stripped envelope puts a sentence on screen that nobody wrote. Only the outermost layer is removed, so a nested fence, or one forged inside the payload, survives byte-for-byte in the body — those bytes are part of what the engine said, not part of this protocol. Nothing else is washed: control characters, leading and trailing whitespace and a twenty-thousand-character payload all pass through untouched, and so does the label, because sanitising and length-capping belong to the mint point that puts a string on a screen and a passage of text must not have two launderers. A value that is not text is answered with **nothing at all** rather than with an empty payload: the reader never stringifies it, never calls its `toString`, and never emits `[object Object]`, and it does not hand back a body of zero length either — an empty payload is a real reading (a fence can legitimately wrap nothing, and an empty string is an empty string), so folding "there was no readable text" into it would leave a caller unable to show a degraded line at all. Those three stay apart: no text yields nothing, an empty string yields an unfenced empty payload, and an empty fenced payload yields a fenced one with its label. The `fenced` discriminator is always present on a reading, and the label key exists only on the fenced arm, so a missing label is never rendered as an empty one. One shape needed more than the whole-string match this started with. When the engine reports back from a delegated run, the fence is only **one section** of the report: ahead of it sit a frame header, a handful of optional field lines and a section label, behind it a closing instruction addressed to the model, an optional internal identifier and a usage block. A matcher anchored to both ends of the input answers *not fenced* on that, and the whole scaffold - written for a model - goes on screen. The reader therefore also locates the fence **inside** a recognised report frame, using four anchors that are always present and always byte-for-byte fixed, and hands the located slice back to the same single matcher rather than a second one; the guard assembles its corpus from the engine package actually installed (the fence from that package's own constructor, the frame lines read structurally out of the minting file and then compared byte-for-byte with what this package registers), so a rewording or a reordering upstream turns the guard red the day it lands. Three readings are pinned one cell each: a fenced result section yields the payload the delegate actually wrote; a partial-findings section yields that text and says which of the two it is, with the prefix line kept out of the payload; and a section the engine filled with its own no-text sentinel yields an empty payload with a reason, which stays distinguishable from an input that was simply an empty string. Whatever surrounded the fence is returned alongside rather than dropped - the bytes before it, the slice itself and the bytes after it reassemble into the input exactly - and the payload never contains a line of the frame. The criterion deliberately does **not** enumerate the lines outside the fence: several of those field lines carry interpolated untrusted text and cannot be told apart from prose, so requiring every one of them to be recognised would mean that a single new field line upstream sends every failed report back to being unreadable, and a failed report is exactly when a person most needs to read what the delegate said. Fourteen negative shapes hold the line against the easy widening, strip anything that looks like a fence: a fence sitting in ordinary text, a missing frame header, a different sentence where the closing instruction belongs, a report with no closing instruction at all while the fence sits at the very end, a header and label in the wrong order, mismatched open and close tags, an open with no close, a bare unfenced section, a stray line between the label and the fence or between the fence and the closing instruction, an entirely absent section, a no-text sentinel with another line after it, a partial-findings prefix followed by something that is not a fence, and the whole report in carriage-return line endings all come back unfenced with the input verbatim and no frame at all. Above all, **provenance is not in the text**: locating a fence inside a frame happens only when the caller states where the bytes came from, because four anchors can only recognise a shape and never prove an origin. The default reading is byte-for-byte what it was before, so a passage of ordinary prose that happens to quote a report - with real warnings on either side of the quoted part - is returned untouched and those warnings stay on screen; a caller that does state the origin gets the located reading, and even then every surrounding byte comes back alongside. Twenty-three malformed origin values fall back to the narrower default without throwing: a near-miss in case, a camel-cased spelling, the right word padded with spaces or tabs or a newline or a zero-width character, a string wrapper object, an object whose `toString` or `valueOf` reports the right word, and an object whose converters both throw. The guard compares the caller’s argument for **exact equality** and nothing else — no trimming, no stringifying, no calling the value’s own converters, because that would let a value of unknown provenance choose its own lane — and a source-level cell requires that the argument reach the comparison unreassigned and unnormalised, with its own two-way check that those patterns speak. The two accepted words are read off the published type rather than copied into the guard; adding a third word later is a type-compatibility change for any caller that switches exhaustively on them. Each of the three anchors must also be **unique** in the text, and ambiguity means the reader declines. The reason is not hypothetical: the frame header interpolates the task's own description verbatim, and the engine only requires that description to be a string, so it can carry newlines and a complete set of protocol lines. Any rule that picks one candidate out of several can therefore be made to pick the planted one, hiding the real result among the surrounding bytes - a guard cell reproduces exactly that, with a planted section and a real one, and requires the reader to decline and the real text to stay on screen. Two reports back to back are the same ambiguity and are declined the same way, with both payloads left visible; a delegate that quotes any one of the three anchor lines inside its own answer also falls back to the input verbatim, which is the registered cost of the rule, and a discrimination cell shows the same corpus reads cleanly once the quoted line is gone. The two legs are not the same shape either: the forked one ends at its closing instruction with no trailing bytes at all, and that real shape has its own cell. The witness arm reads each leg only inside its own array of lines, decodes every extracted literal to its **runtime** value rather than trusting the spelling in the source, and fails loudly if it cannot - an escape rewrite upstream leaves the runtime label unchanged while the spelling diverges, and since that same extracted value builds the corpus and serves as the expectation, trusting the spelling would close a self-proving loop. The two payload labels are therefore also pinned in the guard and compared against what was extracted, so an equivalent rewrite stays green while a real rename turns red the day it lands. A delegate that quotes the frame lines inside its own answer does not move the location, and those quoted lines survive in the payload byte-for-byte |
337
338
  | `scripts/run-display-cap-order-test.mjs` | The order in which untrusted text is sanitised and length-capped, across every mint point that puts an engine- or database-supplied string on a screen. The sanitiser rewrites each invisible character as a six-character escape, so capping the **raw** string first and escaping afterwards hands the screen six times the width that was budgeted — a forty-character allowance becomes two hundred and forty. The guard does not hardcode that allowance, because each mint point wraps its field in different fixed prose and the prose moves: it anchors on the deciding quantity instead, feeding one benign and one control-character input of the same length through the same mint and requiring the second not to come out longer. That criterion is immune to wording changes and stays sensitive to the expansion, and it is `<=` rather than `==` on purpose — a correct escape-then-cap backs the cut off a partially-consumed escape token, so the control-character line is legitimately the shorter of the two, and demanding equality would score that avoidance as a regression. Each mint is bracketed by two positive controls (the input really reaches the screen; the cap really engages) and the expansion predicate is shown to turn red against a deliberately cap-then-escape reference, so an all-green run cannot mean the guard simply measured nothing. The shared mint point is checked directly for the two avoidances it owes — never splitting an escape token in half, which would leave something on screen that looks like the beginning of a complete answer, and never splitting a legal surrogate pair, which would manufacture the very lone surrogate the sanitiser exists to catch |
338
339
  | `scripts/run-seat-task-request-origin-test.mjs` | Where every field of the seat lane's send-message payload comes from, and whether it actually lands anywhere. The seat payload is a closed interface this package mints itself, and most of its fields are meant to ride verbatim onto the engine's request body — two facts nothing used to connect, so both directions could drift in silence. A seat field could be named after a request position that does not exist, in which case a client writes to it, the wire carries it, the engine ignores the whole key, and the screen shows a switch that does nothing; conversely a new request position could arrive with no seat to sit in, which is **structural** absence — the closed set *is* the carrier, so a decision missing from it has nowhere to be put at all, the same shape logged when the effort dial had no seat. The guard turns each field's origin into data: either it names the request position it forwards to, or it is declared seat-local with a written reason, and the two are mutually exclusive. Forwarding claims are then checked against the **installed** SDK's type declarations, parsed rather than restated — a hand-copied list of position names would only ever prove that two transcriptions agree. The parser is held to reading top-level positions only, since a nested option object's inner keys would otherwise be mistaken for positions of the request itself, and it proves that discrimination on synthetic input before any verdict is given. The two subagent fields carry a standing regression pin, and the retention window's inner keys are read from the declaration the same way, so a seat that offers a tunable window cannot offer one the wire has no room for |
339
340
  | `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header |
340
341
  | `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters |
341
342
  | `scripts/run-subagent-content-budget-test.mjs` | The **byte** budget on the sub-agent transcript ledger. It used to be bounded only by *counts* — so many entries per child, so many children — and a count is not a budget when a single entry has no ceiling of its own: one tool result carrying an inlined attachment, or one long model answer, and a single slot sits on tens of megabytes. The guard anchors on how many bytes are **still held** after over-filling, not on whether truncation fired, because an implementation that flags the overflow without actually dropping anything satisfies the second and not the first. Dropping is required to leave a record — how much went and where the retained content now starts — and that record has to reach the render plan, because content that vanishes with no marker gives the reader a transcript shorter than what happened with nothing to say so; the record is one per child, updated in place, pinned to the front, and excluded from the budget it describes. Order matters and is checked: oldest entries go first and the live tail is trimmed only as a last resort, since taking the text the user is watching stream while older history survives is the wrong end. The total budget evicts a whole least-recently-used child rather than shaving every child, and the configuration surface is fail-loud on zero, negatives, non-finite and non-integer values — a silently ignored budget is the exact failure this exists to remove — with the rejection proven atomic so a bad second field cannot leave half a configuration behind. The defaults are checked to be a magnitude that can really be reached, since a number too large to hit is a field rather than a budget |
342
343
  | `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other. One more thing the totals cannot settle: the row key has **two namespaces** — the originating-task id and the parent call id it falls back to — and nothing upstream promises they are disjoint, so the same literal can name one child's identity and another child's parent call. Accumulation therefore keys on the origin as well as the id; the delivered table still keys on the bare id, and a cross-namespace clash is merged into one row that says so, with the partial verdict forced, because a row count and a turn count can both reconcile while the attribution behind them is wrong. The table itself is likewise a **per-stream snapshot** handed to the terminal projector by value rather than left on the caller's context: the three terminal projectors are public, so a host may drive one run through the stream and project another's terminal directly on the same context, and a table left behind would be attributed to whoever projects next — silently called complete whenever that run's own totals happen to match. Without a snapshot, both table keys are simply absent |
343
- | `scripts/run-result-text-backfill-test.mjs` | What happens when the terminal frame's answer text and the text already on screen do not match. A turn's answer normally streams in and the terminal frame carries the same words again, so the two agree — but when the connection drops mid-answer and the reconnect brings the finished version, "this turn already produced assistant text" is true, the terminal fallback is skipped entirely, and the screen stays permanently short of whatever arrived while the stream was down, with nothing to say so. Four cases are pinned. Nothing on screen yet: render the terminal text whole, byte for byte the previous behaviour. On-screen text is a **prefix** of the terminal text: emit only the missing tail, and the guard measures the deciding quantity — the total bytes that reached the screen must equal the terminal text, which fails both for a missing tail and for a re-render that would print the first half twice; when the two are already equal, nothing is emitted at all. Terminal text is a prefix of what is on screen (an engine-side trim): touch nothing, since there is nothing missing and overwriting with the shorter version would erase what the reader already saw. Neither is a prefix of the other: emit **nothing** and raise a fact instead — which version counts is the engine's to say, and appending the terminal version after the streamed one composes a passage nobody ever wrote. That fact carries lengths rather than text, so a renderer is not handed a third version to choose from, and its declared duty is to *reword* the transcript line, never to render more. A cross-segment case proves the comparison reads the whole committed answer rather than the last segment, and the whole thing is driven through the real two-stage path rather than hand-built messages |
344
+ | `scripts/run-result-text-backfill-test.mjs` | What happens when the terminal frame's answer text and the text already on screen do not match. A turn's answer normally streams in and the terminal frame carries the same words again, so the two agree — but when the connection drops mid-answer and the reconnect brings the finished version, "this turn already produced assistant text" is true, the terminal fallback is skipped entirely, and the screen stays permanently short of whatever arrived while the stream was down, with nothing to say so. Four cases are pinned. Nothing on screen yet: render the terminal text whole, byte for byte the previous behaviour. On-screen text is a **prefix** of the terminal text: emit only the missing tail, and the guard measures the deciding quantity — the total bytes that reached the screen must equal the terminal text, which fails both for a missing tail and for a re-render that would print the first half twice; when the two are already equal, nothing is emitted at all. Terminal text is a prefix of what is on screen (an engine-side trim): touch nothing, since there is nothing missing and overwriting with the shorter version would erase what the reader already saw. Neither is a prefix of the other: emit **nothing** and raise a fact instead — which version counts is the engine's to say, and appending the terminal version after the streamed one composes a passage nobody ever wrote. That fact carries lengths rather than text, so a renderer is not handed a third version to choose from, and its declared duty is to *reword* the transcript line, never to render more. A cross-segment case proves the comparison reads the whole committed answer rather than the last segment, and the whole thing is driven through the real two-stage path rather than hand-built messages. The comparison only holds if both sides are the same kind of thing — one assistant message against one assistant message — and that depends on the client knowing where a message ends. A tool card is one place a message ends, but not the only one: when a model round finishes and the next one starts writing prose straight away, with no tool call in between, the two passages belong to two different assistant messages even though nothing visible separates them on the wire. Reading them as one used to glue the two passages into a single transcript line, running the end of one sentence into the start of the next, and then handed the terminal comparison a concatenation to check against the engine's **last** message — so a perfectly ordinary multi-part answer was reported to the reader as *the final answer does not match the text streamed above*. The end of a model round is therefore treated as the end of an assistant message: the pending prose is committed and a new message begins. That holds even when the round reports no spend at all, because *this round is over* and *this is what it cost* are two different facts and only the first one decides a boundary. A round belonging to a **sub-flow** decides nothing for the leader, and the test pins all three identity members, empty strings included, against a control that proves the same shape really does divide when no identity is present. The rest of the section is regression: a boundary landing immediately after a tool card must not lose the fallback that lets a card-ends-the-turn answer compare against the prose before the card; repeated boundaries with no prose between them must not mint empty messages or lose track of which message the comparison should read; a reconnect that brings the finished answer still backfills only the missing tail; the leg that already carries whole messages is not divided twice; and a segment boundary that arrives **after** the round ended still replaces the segment authoritatively and hands back the row attribution, without minting a second copy of the passage. A last group covers where a message boundary meets a segment that the engine announces late or not at all, and it pins only the half that is unambiguous. A round whose prose the engine never announces, followed by one it does, used to have the first passage **silently overwritten** by the second one's authoritative text — no transcript line for it and an empty attribution list, so a host had nothing to correct; it now keeps its own line and the attribution is handed over in full, with the offset measured on that passage rather than on the authoritative text. Which of the two passages survives a compliant host's rewrite is deliberately **not** asserted: the row named belongs to an earlier message, exactly as it already did at a tool-call boundary, and the underlying cause — the client splitting messages on its own boundaries while the engine announces segments on content blocks — is recorded as a known limit rather than pinned as a desired outcome. Where the second round's prose arrives as a whole block instead, there is no segment announcement at all and both passages keep their own line **in the order they happened** — previously the whole block was written first and the earlier passage only landed at the close, so the reader saw them reversed. An announcement that arrives after its round has already ended, and whose authoritative text merely extends what was shown, is pinned on the two things that are not in question: the bytes on screen add up to the authoritative text exactly once, and the divergence signal with its offset is still emitted so a host can reword. That ordering is unreachable on the installed engine — the announcement is pushed while the message is still being assembled and the round end only after it is complete, both through one synchronous dispatcher onto one queue — so the case is defensive; the late-announcement path exists because a tool call can be admitted before assembly finishes, which a round end cannot. Finally, one guard names a layer seam rather than a behaviour: the round-end frame is folded into a neutral usage arm that carries **none** of the three identity members, so a round belonging to a forwarded background child arrives with nothing to judge and the sub-flow cutoff cannot reach it. That guard reddening is the signal that identity now survives the fold and the cutoff has become effective |
344
345
  | `scripts/run-engine-vocab-floor-test.mjs` | Engine-mirrored vocabularies (structured card whitelist, self-reported tool face, control verbs, recogniser sets) against the *installed* `@sema-agent/core` |
345
346
  | `scripts/run-limits-env-failloud-test.mjs` | `SEMA_HEADLESS_*` env-lane limits reject invalid values as loudly as the flag lane (no silent "no budget" runs) |
346
347
  | `scripts/run-streamjson-timing-honesty-test.mjs` | Stream timing & terminal honesty ([2084]): held errored fs-write results release on model progress; a wall-clock stop maps to `error_during_execution` with a truthful salvage note; the synthetic API-error assistant row carries the `<synthetic>` in-message sentinel. Also ([2489], core 5.8.0): the run-limit `errorCode` -> CC subtype map is pinned code by code (`limits.max_{cost,turns,tokens,walltime}_exceeded`), token/wall-clock stops keep the text the engine already produced, and the `failed` event arm shares the one mapping point. The 5.7 dual-vocabulary legs retired with server 6.0.0 (which bundles core 5.8.0); four **retirement negative controls** stand in their place — the retired `status:'timeout'` and the retired codes must fall to the honest fallback subtype and must never drop back to an empty success, so putting any of them back turns the gate red |
@@ -349,7 +350,7 @@ public-surface guard checks that last one).
349
350
  | `scripts/run-usage-verbatim-channel-test.mjs` | The two complementary usage disciplines (core 3.0.0 metering semantics): the CC `ModelUsage` mirror stays pure (five pinned keys, `totalInputTokens` has no seat), while the sema-owned channel forwards the engine `turn_end.usage` object **verbatim** (six keys, incl. `totalInputTokens`) via `last_turn_usage.engineUsage` / `handle.latestEngineUsage` — honest absence on pre-3.0.0 engines, no fabricated zeros |
350
351
  | `scripts/run-plan-review-decide-verify-test.mjs` | `decidePlanReview`'s post-decide honesty ([2315]/[2316], engine RB-471 family): a 2xx from the decide endpoint is **not** a terminal — the wire re-pulls the task status and words the outcome by the real shape (still-locked / legal new gate / genuinely left park / unverified), never claiming success it hasn't earned. Driven against a real fake-engine HTTP server through the shipped dist |
351
352
  | `scripts/run-shell-gate-durable-allow-test.mjs` | #110: the durable approval leg for **shell** gates. The tool_end HOLD/REJECT predicate must cover Bash the same way park detection already does (otherwise the park poison frame `Operation aborted` hits the transcript, `endedCalls` swallows the real replayed result, and the user who pressed Yes watches a command that really ran be reported as aborted); a replayed, already-decided park must resume reading the stream instead of being reported as a failed turn; `lastEventId` must track numeric `seq` too. Mutation-proven: each of the three fixes reverted turns the gate red |
352
- | `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage |
353
+ | `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings |
353
354
  | `scripts/run-park-hop-progress-test.mjs` | L-80: the park re-attach loop budgets **stalled** rounds, not parks. A turn where the model keeps hitting gates and every one of them is really decided (a card was answered, the engine really moved on) must never be cut off by the hop budget — the budget counts consecutive rounds that produced no progress, and "the engine revived and immediately parked again on the same coordinates" is not progress. The three non-progress arms (already-resolved, decide-transport-exhausted, and a re-scan that was adopted but led nowhere) share one same-cause limit instead of one arm having a limit and the others having none, and every non-progress re-attach is announced once through the host callback rather than only to the debug log. When the limit is spent the resolver reads the approval queue once more and puts whatever is decidable in front of the user before it gives up; only when there is genuinely nothing to show does it fail soft, and the terminal message then carries the real cause and a real way out instead of a sentence about a budget. On the self-heal side, a reopen verdict that reports `decidedWithoutCard` — the chain settled the gate by rule, so there was no card to present — is progress, not a reopen failure, and the user is not told their message was NOT sent. Negative control: a genuinely empty queue with a run that never moves still fails soft |
354
355
  | `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
355
356
  | `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
@@ -365,7 +366,7 @@ public-surface guard checks that last one).
365
366
  | `scripts/run-crash-converged-projection-test.mjs` | The `crashConverged` read face on `GET /v1/approvals` (L-38): what the *previous life* of a crashed local engine left behind, projected for every client. Three judgements are pinned. First, **absence is not an empty list** — a missing key (an older server, deps not present, or a carrier that is not an array at all) returns `undefined`, and the client renders nothing; an empty array returns a present zero-count object, which is the server actually saying "none". Folding the first into `{total:0}` would have the client assert "nothing was left behind" on a surface a person uses to decide whether it is safe to re-run something — the worst possible direction for a false statement — so the two cases are pinned to different **return shapes** and a test asserts the two verdicts are unequal. Second, bucketing is a **four-term conjunction**: `orphanState === 'pending'` *and* `resumeSafe === true` *and* both approval-evidence keys (`originalDecision`, `decidedAtMs`) absent. A fifth term rejects any row carrying an **accessor**, and accessors are never invoked at all — reading one means synchronously running someone else's code, and `catch` catches throwing, not *never returning*, so a looping getter would pin the startup thread forever (the row cap does nothing against that shape). The same rule covers the three untrusted reads outside the row as well — the envelope's `crashConverged` key, the carrier's `length`, and every numeric index are read as own property *descriptors* and only data descriptors are used, so accessors and prototype entries read as absent and are never invoked. Such a key is treated as absent: if it was a required field the row is counted as dropped, if it was optional or additive the row survives without it. That also closes the ordering attack, since spreading runs getters in property order and an earlier one could `delete` the approval evidence before it is ever copied (measured before the fix: such a row reached the resume-safe bucket), and the check therefore moves ahead of the read, onto the property descriptors — from which the snapshot is then built directly, because checking descriptors and *then* spreading is two independent observations of the same row, and a non-throwing proxy can make the two `ownKeys` calls disagree (first showing `originalDecision: 'approve'` so the row reads as plain data, then omitting that configurable key so the snapshot loses the evidence; measured before the fix: the dangerous row reached the resume-safe bucket after exactly two enumerations, and after it, one). Keys are written with `Object.defineProperty` rather than plain assignment, because `'__proto__'` is a legal own enumerable key and `o['__proto__'] = x` does not store a value — it calls the prototype setter, letting a row whose own properties are all plain data (so the accessor gate never fires) inject a prototype whose `sessionId` getter deletes the approval evidence from the snapshot during validation; `defineProperty` fires no setter, so the key survives as ordinary additive data and the snapshot keeps `Object.prototype`. A row that simply arrives with a custom prototype is treated the same way, since the snapshot only enumerates own properties: approval evidence sitting on the prototype would never reach it, and a perfectly ordinary object with no proxy and no accessors could otherwise be called safe to re-run — real bodies come from `JSON.parse` and always carry `Object.prototype`, so nothing genuine trips it). Validation itself runs on a **null-prototype** dictionary and the bucketing verdict is carried out of that same pass rather than re-read from the delivered row, because every property lookup on an ordinary `{}` reaches `Object.prototype`: a polluted `sessionId` getter there would delete the approval evidence from the snapshot mid-validation and send the row to the safe bucket (measured before the fix). The row handed to the client is still an ordinary object — the null prototype is an implementation detail of the check, not of the value) — real JSON bodies are all data properties, so only a middle-layer-synthesised payload ever trips it, and it too lands in the human bucket rather than being dropped. The `decided` arm means the human had already approved and side effects may be half-landed, so it always goes to the human bucket, as does `resumeSafe === false` and — the last two terms — any row whose own fields contradict each other, since `pending` claims nothing ran while that evidence says somebody pressed approve. Deciding "not safe" costs one extra question (recoverable); deciding "safe" wrongly has somebody re-run work that already partly happened (not). A 2x2 truth table pins that exactly one cell is resume-safe, so reading either key alone turns red, and the contradictory rows are routed to the human bucket rather than dropped — they are real orphans, and the ones most worth showing. Third, unreadable rows are **dropped and counted**, never thrown and never passed through: the product is declared as `CrashConvergedRow`, so letting a row missing a required field — or carrying one of the wrong type — past would be a lie at the type level, and the closed literal discriminators (`decision` / `cause` / `orphanState`) decide family membership rather than being an open vocabulary. The measuring stick stops at the **type** floor, though: degenerate-but-well-typed values (`ts: NaN`, an empty `toolName`) are kept, because swallowing a real orphan over a decorative field is the worse direction, and the one deliberate exception is `approvalId`, which must be non-empty to be a row identity at all. `dropped` is kept separate from `total` so unreadable rows never inflate "N approvals were affected"; each row is a **one-shot snapshot** — every own enumerable key is read exactly once, and validation, bucketing and the handed-back value all read that same snapshot, so additive upstream keys survive while a **non-idempotent** getter (one that never throws, just answers differently on a second read) can no longer erase the approval evidence between the check and the bucketing (measured before the fix: such a row landed in the resume-safe bucket while its checked value was `"approve"`). Hostile carriers are counted rather than allowed to reject: **every** touch of the carrier is guarded — envelope property reads, `Array.isArray` itself (it throws on a revoked proxy), the `length` read, each indexed read and each row's property reads — and a traversal that dies halfway returns absence rather than a half-counted total. A row that cannot be read never takes the batch with it: its own shape check is inside its own guard, so one revoked-proxy row costs a `dropped` tick rather than collapsing the whole projection to absence — which a client would have read as "this deployment does not offer the surface". Traversal goes by **numeric index, never the carrier's own iterator protocol**, because `for...of` hands the carrier the question of which rows exist: an array carrying an overridden `Symbol.iterator` can yield nothing (measured before the fix: a real orphan became `{total:0}`, which a client reads as "the server said there are none") or swap a dangerous `decided` row for a safe-looking one (measured: `fake-safe` was returned in place of `real-danger`). Row count is capped at 100000 and the cap is checked **before** the walk: requiring only a non-negative integer `length` does not stop a proxy trap reporting a billion, and this surface runs on the startup / `--resume` path, where a synchronous spin freezes the thread (measured before the cap: twenty million rows took 18.3 seconds and twenty million index reads; a billion does not come back). The honest boundary is stated rather than overclaimed — a proxy can still lie in its `length` or index traps, which is the same thing as a host injecting a lying transport — and the widening of `ApprovalsResourceLike.list()` is proven **additive** by really running tsc over a legacy `{ pending }` mock *and* over the real `AgentClient` path — the projector takes `unknown` precisely because a parameter shaped as "an object with an optional `crashConverged`" is a TypeScript weak type that the installed SDK's own `list()` return shape shares no property with, which only a real-client compile would have caught — with a known-red control so a clean run means the checker spoke |
366
367
  | `scripts/run-self-orchestration-denial-test.mjs` | The three judgements behind a **denied self-orchestration request** (server 7.57.0), each of which all three clients would otherwise get wrong on their own. First, whether to retry at all is a **conjunction that may not be loosened**: HTTP 501 *and* an `errorCode` that is **exactly** `capability.self_orchestration_required`. That code shares its shape with every other `capability.*` 501, so dispatching on the prefix would drag "some other capability is not wired up" into the retry arm — those requests do not become acceptable once the two keys are gone, so the client would spend a request and then tell the user the wrong reason. Negative controls cover all four directions: a sibling `capability.*` code, a truncated or suffixed variant of the right one, a codeless 501 (it decides nothing, so it decides nothing — no guessing), and the right code under 500 / 400 / 503 or a string `"501"`. The classifier reads structurally rather than by `instanceof` (a host may inject its own transport; across realms or duplicate SDK instances an understandable error would read as unreadable), so a class instance, a bare `{status, errorCode}` literal and an error carrying those fields on its **prototype** all reach the same verdict — and a hostile proxy or a throwing getter yields `null` instead of throwing, because this classifier runs inside a `catch` block where anything it throws escapes the caller's own guard. Second, removing the intent is a **structural** operation, not wording: `selfOrchestration` sits at the top level while `ultracode` sits under `settings` — two different stamping legs — and a client hand-writing `delete` will miss the second one, which costs the user the same failure twice. The single stripper is pinned to touch exactly those two: other `settings` sub-keys and their values survive byte for byte, `deferTools` is left alone (pulling `Workflow` out would be a behaviour change, not a removal of intent), additive unknown keys survive at both levels, the input object is never mutated, `settings` is only dropped entirely when `ultracode` was really there and nothing else remains (an already-empty one is left as is), a non-object `settings` is not touched at all, an `ultracode` that only exists on the prototype does not count, and the whole thing is idempotent. The end-to-end leg runs a real `buildTaskRequest` product through it and asserts the stripped body still passes the registration gate key by key. Third, on the capabilities body, **absence is not "switched off"**: a pre-7.57 server has no `workflowsGate` key at all, so reading absence as "the engine says no" asserts something the server never said, and the mirror-image disease is folding an **unrecognised** `denial` into `null`, which would have the client render "nothing was denied" when the truth is "denied, for a reason I do not recognise". Five shapes are pinned — caps unreadable, gate absent, closed-set member, unknown value, accessor — with the unknown arm carrying the raw token (or an empty one when the value is not even a string) and never collapsing to `null`. All four untrusted reads go through own **data descriptors** only, and the guard pins the getter invocation count at zero, since `catch` catches throwing but not *never returning*; a descriptor trap that throws and a revoked proxy both yield honest absence rather than an exception — though *what* absence means differs by field, and the guard pins that split rather than a blanket rule: an accessor on `workflows`, `workflowsGate` or `engineCan` reads as absent, while an accessor on `denial` reads as `{unknown:''}`, because a key that is **not there** is the gate saying "nothing was denied" whereas a key that is there but cannot be read is "denied, and I could not read why" — folding the second into the first is exactly the false statement this face exists to prevent. Two further pins came out of an adversarial review. The exported retry list is **frozen at runtime**, not merely `as const`: the verdict hands out that same reference, so any consumer splicing it once would poison every later verdict in the process — the guard asserts `Object.isFrozen`, that four different mutation attempts leave it byte-identical, and that a verdict issued *after* those attempts still carries the original two entries. And the classifier reads `denial` only **after** both criteria have passed, since it is not a criterion but an extra field on the verdict: the guard pins the getter invocation count at zero for any error that does not match and at most one for an error that does. The scope line is drawn explicitly rather than overclaimed — "no getter ever runs" holds for `projectWorkflowsGate`, which reads **wire JSON** where every field is an own data property by definition, but not for the classifier, which reads a **thrown value** that may well be an SDK `APIError` class instance carrying `status` and `errorCode` on its prototype; insisting on own data descriptors there would report a perfectly readable error as unreadable, so that side promises only that it never throws. A final pin covers the **integration document's own worked example** rather than the library: the shipped SDK's `tasks.stream()` is an `async` generator, so calling it issues no request at all — the POST happens inside `streamRaw` on the first iteration, and a `try` wrapped around the `stream(...)` call itself can never catch the 501. A client following a submit-shaped recipe on the streaming leg would never run the classifier, and the whole strip-and-retry path would silently do nothing. The guard drives the **real** `TasksResource` against a fake transport, offline, and pins both halves: the synchronous leg is in flight the moment it is called, the streaming leg has issued zero requests after the call and raises on the first `next()` — and it does so through the **real** error path, with `openStream` returning an actual 501 `Response` that the SDK's own `errorFromResponse` turns into the typed error, pinning the `openStream`→`errorFrom` call order so a transport that stops minting `errorCode` cannot pass. The documented recipe is then **executed** rather than keyword-counted: exactly one retry, a second body that really lost both keys while every other setting survives byte for byte, the caller's own request object left untouched, one disclosure and only one, a second 501 propagating with the request count still at two, and — after the first 501 — an abort leaving the count at one with nothing disclosed. A last leg is type-level: `stripSelfOrchestrationIntent` carries an SDK `TaskRequest` overload, because the wide `Record<string, unknown>` form erases the caller's type and the document's "strip and resubmit" line would not compile without an unsafe cast; a real tsc run over a virtual file proves both the narrow and the wide path, with a known-red control — and it compiles the document's two recipes **verbatim**, extracted from the section itself, because a recipe that does not compile is a recipe that was never given: `{ transientOk: true, signal }` is a TS2379 under `exactOptionalPropertyTypes`, which no amount of prose review had caught. The last thing pinned is the one that would have been quietest of all: the SDK's `stream()` returns only on a `done` or `failed` frame, so a stream truncated mid-run — or yielding nothing at all — ends the `for await` just as normally as a completed one. The documented `runOnce` therefore tracks whether it ever saw a terminal frame and raises when it did not, the guard's success fixture emits a real terminal and asserts the handler received it, and a truncated-stream control asserts that shape is reported as a failure with no retry and nothing disclosed. That terminal-frame rule then needed one more turn of its own: the underlying reader returns *normally* when the signal is aborted, so the check as first written rewrote a user's cancellation into a generic stream fault — a client keying off `AbortError` to suppress the error would instead have shown a failure, or resubmitted. Cancellation is therefore checked first, a real-SDK case aborts from inside the handler and asserts the original `AbortError` survives with no retry and nothing disclosed, and the document is checked for that ordering. The harness runs the documented `handle` and `transcript.note` as real spies rather than pushing frames itself, the drive loop rethrows exactly as the document does, and the disclosure ledger is proven to be the caller's own array by a positive identity assertion — without which the cancellation leg's "nothing disclosed" would have been vacuously true. Each recipe is compiled **on its own**, with a preamble that declares only what a host supplies and injects no library symbol, since compiling them together let the second one borrow the first one's imports, and the preamble's own types are decoupled from what the recipes import so the "remove the imports and it must fail" control fails for the right reason — which is checked by attribution, not merely by redness. Ordering is the last thing to get right: the cancellation check must come before the truncation error but **both** must sit behind the terminal-frame test, because a cancellation that lands after the run already reported `done` would otherwise overwrite a real outcome — one that may have already had effects — with "cancelled", and a person reading that will run it again. Aborting from inside `handle(done)` and `handle(failed)` are both pinned to still report success, and the ordering assertion is anchored inside the streaming `runOnce` body rather than the section, since the section's first `throwIfAborted` belongs to the synchronous recipe and would have made a reversed streaming recipe pass — and that ordering check is now anchored on the TypeScript AST rather than on text, since a comment reproducing the two statements in the right order let a genuinely reversed body pass. One more timing fact had to be written into the recipe: a single SSE read buffers several frames and the SDK yields them back to back, so checking the signal only after the loop lets a cancelled run keep consuming the rest of the chunk — measured, an abort inside `handle(turn_start)` still swallowed the `done` that followed and reported success. The recipe therefore re-checks after every non-terminal frame. Finally, the behavioural matrix is no longer run against a copy of the recipe: both recipes are extracted from the document, transpiled, and **executed** with injected host objects, so the disclosure assertion really exercises the document's own `transcript.note(disclose(...))` line, and the synchronous leg gets the same full matrix the streaming one does |
367
368
  | `scripts/run-package-hygiene-test.mjs` | Everything `package.json` `files` ships — dist JS/typings and the Markdown docs — is screened line-by-line against a deny-list of strings that must never reach a public tarball (internal hostnames, codenames, person names, collaboration-process words, other repos' ledger ids and repo names; opaque ticket ids `CC-nnn` and post numbers `[nnnn]` are allowed as traceability references). Since 0.77.2 the build strips comments (`removeComments`; enforced by `run-dist-comments-test.mjs`), so what this gate screens in dist is code, string literals and type-level text. Markdown docs are enforced forward-only (CHANGELOG from 0.77.2, the integration doc from §81) because published sections are frozen.
368
- | `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T | null`, the `installHost` family returns `T | undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. |
369
+ | `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T \| null`, the `installHost` family returns `T \| undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. |
369
370
  | `scripts/run-type-superset-ledger-test.mjs` | The type/wire **superset ledger** (`docs/type-superset.json`): positions this package adds on top of a CC-shaped contract, each carrying the evidence for what CC's own type surface does or does not have there. Completeness is deliberately uneven and the ledger says so. The `_sema_*` private-key class is checked in **both** directions (a key in the source that never entered the ledger is red, naming key and file; a ledger row whose key left the source is red) — but only for keys written as literals, which is the convention the ledger mandates. A key assembled by string arithmetic is beyond what any static rule can enumerate, so the guard fails closed on every shape it *can* decide (a bare `_sema_` prefix is red wherever it appears, save one pinned guard site) and leaves the rest as a convention violation for review to catch, rather than claiming a completeness it does not have. The two hand-surveyed classes are only checked for coordinate and evidence integrity, never discovered. Both directions read the source through the **TypeScript AST**, not a text scan, and they read two different sets out of it. A *key site* is an identifier, or a string whose whole value is the key — so `'_sema_decision-v2'` is carried whole rather than truncated at the first non-identifier character into some *other* key that happens to be registered. A *mention* is the key appearing inside a longer string, which is prose, not usage. The staleness direction counts key sites only: a comment or a doc sentence left behind after the last real mint site is deleted must not keep the row alive (mutation-proven — with both the comment and the prose string untouched, removing the one real site turns the guard red). And because a prefix can be concatenated or interpolated into a key no static set will ever see, the bare `_sema_` literal is refused outright rather than traced: every occurrence is red except the single inline `startsWith` guard the sanitizer needs, because the set of expressions a bare prefix can travel through on its way to a concatenation is open-ended and enumerating it is always one form behind. Every row's `host` must still resolve, with the key being a real **member of that declaration** rather than a string occurring somewhere in the same file — `governanceForced`/`delegation` each live on two different shapes in one file, and a member commented out is a member deleted, which a text-shaped check happily reads as still present. And the direction worth the most: each machine-form `ccAbsenceEvidence` is re-derived from the row's own `key` — the ledger's recorded string must match that derivation verbatim, since a row quietly witnessing `\bnever_present\b` is green forever while watching nothing (mutation-proven: the same edit passes the unbound form and is caught by the bound one) — and the check runs against the names the installed `@sema-agent/agent-types` `.d.ts` set actually declares, parsed with the TypeScript AST rather than grepped, so a name CC merely mentions in a comment cannot force the row into the manual escape hatch and thereby retire the very witness that was supposed to fire the day CC declares that name for real. That escape hatch is gated by an allowlist living **in the guard**, not the ledger, so claiming it costs a reviewed diff. Missing material never reads as a pass, and the verdict splits by *why* it is missing: no TypeScript parser skips the suite before it starts; a missing `agent-types` still runs and prints the first three directions, then exits **1** when `package.json` declares the mirror but it is not installed — a broken install must not retire the repository's only "the day CC declares this name" alarm, and reporting it as a skip would leave "never evaluated" and "evaluated, no drift" indistinguishable to the runner — and exits 3 only when nothing declares the mirror at all, which is the one case where the direction genuinely does not apply. Either way a run that evaluated no witness is never counted as one that did. When the mirror *is* present its **installed version** is witnessed too (the two declared floors must agree with each other and the installed copy must meet them), since four preflight probes are satisfied by an arbitrarily stale mirror — they prove the extractor speaks, not that it is current. Every direction carries a positive control — known-present CC symbols, a comment-only sample proving the extractor distinguishes declaration from mention, and synthetic corpora fed through the **same** discriminator function the real verdict uses, so a verdict quietly rewritten to return nothing takes its own control down with it |
370
371
  | `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" |
371
372
  | `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
@@ -390,14 +391,14 @@ public-surface guard checks that last one).
390
391
  | `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
391
392
  | `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
392
393
  | `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
393
- | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins |
394
+ | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false |
394
395
  | `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
395
396
  | `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
396
397
  | `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
397
398
  | `scripts/run-text-segment-authority-test.mjs` | The **authoritative segment replacement** on `text_end` (server >=7.75.3). `text_end.content` now goes through the same redactor as `result` and the ledger while `text_delta` stays verbatim, so the two **may differ** — an answer that quoted a credential used to be committed to the local transcript in its unredacted form, because the arm only forwarded the boundary signal. Six timing shapes are pinned, two of which an adversarial review reproduced against the installed engine's real bytes and which the first design got wrong in both directions: a second boundary in the same turn (the per-block case on one provider lane) used to make the first segment's prose vanish, and a boundary that arrives *after* the tool card (the other lane emits it at finalize) used to be read as "this package never handled that segment" and reported nothing at all. Three additive keys, all never-false; the two shapes that look alike are told apart by the second one, because the host's action in them is the opposite. The end-to-end legs drive the real pipeline without hand-inserting a segment commit — doing so is exactly what hid the first defect. A second review round then found two combination timings on top of the first fix — a tool card followed by *more* deltas in the same segment, and a byte count that had been documented as a message count — and both are pinned here too. A third round caught a length that the prose called bytes while the code returned UTF-16 units — harmless in ASCII, and on CJK text enough to leave the credential on screen — plus a backfill ledger that had to be kept in step, so the terminal frame does not re-render the segment a second time — kept in step only where the whole stretch sits in one message, because those ledgers are per-message and a fourth round showed that writing across them charges one message's prose to another. A fifth round settled the whole class into one invariant the guard now checks against the previous release's behaviour: this package only rewrites bytes it is still holding in the current message — once a segment has crossed a package-side boundary it emits the three keys and changes nothing else **0.68.2 (CC-01):** the segment identity is now minted here, not by the host: every committed assistant text row carries a top-level `_sema_segment_id` (stamped once at the `adapt()` exit, so the durable whole-message leg and the streamed-segment leg are covered alike; thinking blocks, tool_use-tailed rows and chrome events are left byte-for-byte), `text_segment_end` carries the same value as `segmentId` before rotating, subagent boundaries never rotate, and a replayed stream yields the same identities. Three mutations (no rotation / no stamping / stamping tool_use rows) each turn the guard red **0.69.0 (CC-02):** the same authority replacement now covers the reasoning face (`reasoning_end`, server >=7.77.0): a thinking block still buffered is swapped whole and its live tail recomputed; one already committed at a boundary (the usual timing, since the first text delta commits it) is left untouched and the host is told the row to replace by its uuid, never re-emitted. Subagent boundaries are ignored and the text-segment identity does not rotate **0.69.1 (CC-09):** the run-stream replay guard still drops a frame whose event id was already seen, but it now reports the drop through the host's dropped-frame sink as `duplicate_seq` instead of vanishing silently (server 7.77.0 reuses the first reasoning delta's id for `reasoning_end`, so that authoritative segment is lost on the print lane until 7.78.1); the interactive adapter has no such guard and keeps receiving it **0.69.1 (CC-10/CC-11):** subagent segment-end frames are fenced on all three identity keys (a frame carrying only `sourceTaskId` no longer masquerades as the leader's), and a reasoning segment that spans tool cards now hands the host every committed row it covers (`committedUuids`) so nothing unredacted is left behind |
398
- | `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the 74 suites above actually turn red when the material they check really breaks — a census had found 16 of them clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. Seven guards of the same shape and 51 behaviour/projection suites are catalogued rather than rehearsed this round — see `docs/GATE-NEGATIVE-CONTROLS.md` for the full table, the reasons, and a one-minute manual replay recipe for each blind one. The suite cross-checks its own case count against that document's row counts in both directions, so a case quietly dropped from the array without the document following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
399
+ | `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the suites above actually turn red when the material they check really breaks — a census found the ones clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. The guards of the same shape that are too costly to rehearse, and the behaviour/projection suites that red on their own assertions, are catalogued rather than rehearsed — see `docs/GATE-NEGATIVE-CONTROLS.md` for the three tables (all generated from `gates-manifest.json`, which carries each suite's classification), the reasons, and a one-minute manual replay recipe for each blind one. The suite reconciles its case list against that classification by name in both directions, so a case quietly dropped from the array without the manifest following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
399
400
  | `scripts/run-engine-cap-reader-factory-test.mjs` | The one shared implementation behind every capability reader's four ports (cache, generation gate, probe tee, invalidation), exercised as a table: every reader in the table runs the *same* criteria (the table length is the source of truth, and a roster check fails the gate if any source file calls the factory without having a row) — the four states, a throwing projection treated exactly like an unreadable one (and never escaping the tee), the generation rules (a stale generation is dropped before the projection even runs; a projection that changes the generation mid-flight cannot overwrite the newer value, whether it returns or throws; the caller's `opts` is snapshotted once; and omitting the generation still writes, because that supply is additive and this refactor does not quietly tighten it), the top-level key's getter being read exactly once, a freshly minted "unobserved" reading on every miss (two misses are never the same object, so a consumer that mutates one cannot taint another base URL), one independent table per reader that never take each other down, invalidating one base URL leaving every other base URL's reading untouched, `forget`/tee being no-ops on an empty or non-string base URL, the read anchor being resolved dynamically, and the deliberate split in how presence is judged per reader. |
400
- | `scripts/run-mcp-liveness-test.mjs` | The engine's **liveness observation** about each MCP server it hosts (`wiring_manifest.mcp[].liveness`, engine-side from core 7.24.3 / server 7.91.2): one reader, one word list, one leg-level verdict. The cell answers *can this server still be reached* — it is not the connect-time verdict beside it, which the engine deliberately freezes (a server that died mid-run still reads `connected`), and it is not a re-dial's judgement either, so `status: "failed"` next to `liveness.state: "reachable"` is a **real row**: the server answered the handshake and answered with a protocol error — up, and misconfigured. The three words are read as a closed set (an exchange completed / it was lost in transport or the clock / what came back does not answer the question), and a word outside it is malformed rather than rendered, because a word nobody upstream has defined is not a sentence worth putting on screen. **Absence is the fourth reading and is not one of the words**: it means *no liveness record is available*, which on the wire covers a server this leg never reached, a declaration that could not be dialled, an older engine, and a record the projection ahead of us dropped — all indistinguishable, so it is never read as "we looked and could not tell" (a strictly stronger claim), never as healthy and never as off. The failure-class footnote rides the unreachable word only, and one that turns up anywhere else, or that is malformed, loses **just the footnote** while the word and its timestamp stay: the honest reading is then "cannot be reached, reason not given", not "this record is broken". Malformed never becomes healthy: a cell that is present but unreadable marks its row and pulls the leg-level verdict back to *cannot tell*, since letting it sit beside a reachable row would report a leg as reachable on the strength of a record that may well have said the opposite. The verdict takes the worst fact first rather than a majority or the newest reading, carries no server count — so there is no fabricated zero to be read as "no problems" — and its timestamp belongs to **that leg's** observation, not to now: the engine runs no probe and adds no traffic of its own, so this is a per-leg snapshot rather than a heartbeat. The replayed roster on the session panel and the live leg go through the same reader, and the panel's own deployment-side rows carry no liveness position today, so the verdict does not borrow a word from that face. The word list is bitten in both directions where an installed witness exists and the absence of one is itself asserted against the installed engine's version, so the day it ships the comparison starts on its own; the day the wire types declare the cell, the guard turns red and asks for the anchor to move there |
401
+ | `scripts/run-mcp-liveness-test.mjs` | The engine's **liveness observation** about each MCP server it hosts (`wiring_manifest.mcp[].liveness`, engine-side from core 7.24.3 / server 7.91.2): one reader, one word list, one leg-level verdict. The cell answers *can this server still be reached* — it is not the connect-time verdict beside it, which the engine deliberately freezes (a server that died mid-run still reads `connected`), and it is not a re-dial's judgement either, so `status: "failed"` next to `liveness.state: "reachable"` is a **real row**: the server answered the handshake and answered with a protocol error — up, and misconfigured. The three words are read as a closed set (an exchange completed / it was lost in transport or the clock / what came back does not answer the question), and a word outside it is malformed rather than rendered, because a word nobody upstream has defined is not a sentence worth putting on screen. **Absence is the fourth reading and is not one of the words**: it means *no liveness record is available*, which on the wire covers a server this leg never reached, a declaration that could not be dialled, an older engine, and a record the projection ahead of us dropped — all indistinguishable, so it is never read as "we looked and could not tell" (a strictly stronger claim), never as healthy and never as off. The failure-class footnote rides the unreachable word only, and one that turns up anywhere else, or that is malformed, loses **just the footnote** while the word and its timestamp stay: the honest reading is then "cannot be reached, reason not given", not "this record is broken". Malformed never becomes healthy: a cell that is present but unreadable marks its row and pulls the leg-level verdict back to *cannot tell*, since letting it sit beside a reachable row would report a leg as reachable on the strength of a record that may well have said the opposite. The verdict takes the worst fact first rather than a majority or the newest reading, carries no server count — so there is no fabricated zero to be read as "no problems" — and its timestamp belongs to **that leg's** observation, not to now: the engine runs no probe and adds no traffic of its own, so this is a per-leg snapshot rather than a heartbeat. The replayed roster on the session panel and the live leg go through the same reader, and the panel's own deployment-side rows carry no liveness position today, so the verdict does not borrow a word from that face. The word list is bitten in both directions where an installed witness exists and the absence of one is itself asserted against the installed engine's version, so the day it ships the comparison starts on its own; since sdk 11.2.0 the wire types declare the cell, and the word list and the presence test are pinned to those types at compile time in both directions, so the two can never drift apart |
401
402
  | `scripts/run-peer-lane-rules-write-capability-test.mjs` | Two more engine self-descriptions read the same four-state way as their seven sibling capability readers (`capabilities.peerLane`, `capabilities.permissionRulesWrite`): an absent key is not reported (an older engine that predates the position, never folded into `false`), `true` is present, `false` is a positive absent (the cross-session lane not being mounted on this deployment, or this particular call not being able to reach the tightening-direction write entry), and anything else is unreadable and drops the cell. Each carries its own single-source verdict (`peerLaneAvailable` returns `yes`/`no`/`unknown`; `permissionRulesWriteAvailable` collapses to a plain boolean, present being the only `true`). The write-entry position pairs with a boolean convenience port in the persisted-rules module, and this guard pins that port to derive from nothing but this one reader's own reading — never a conjunction with the lane-reachable position, and never a second read of the deployment-level existence signal the revoke surface uses (the two are documented as reading differently on purpose): a deployment where the lane answers true but the write entry's key is simply absent (an older binary) must still come back `false`, a deployment where the write entry answers true while the lane key is entirely unseen must still come back `true` (proving no silent conjunction crept in), seeding only the general capabilities cache — never this reader's own feed — must still come back `false` (proving the convenience port cannot be satisfied by the wrong table), and passing an explicit `undefined` base URL must still come back `false` even while a different, already-installed engine target answers `true` for the same position (an adversarial pass found the naive forward of that parameter falls through to the reader's own convenience default, silently answering for whichever engine happens to be installed rather than the caller's absent target — the fix routes an explicit absence through the same empty-string path the reader treats as unobserved). |
402
403
  | `scripts/run-lane-proof-identity-test.mjs` | The **instance identity of a lane proof**: the main-lane proof is minted fresh on every emission. Previously a single module-level constant object was handed both to `laneOf(an unregistered task id)` and to some fifty main-lane emission points, so two unrelated consumers — across adapter instances, across streams, across turns — held the same object: writing a card id onto one of them was readable on the other, and the four opening main-lane events changed together. Nothing in this package writes to a lane proof and the known consumers only read it, so this is an **aliasing hazard on a published output surface** rather than an observed corruption — a consumer that uses the proof as an identity key, for dedup, or as a view-layer identity would conflate two unrelated rows without writing a single byte, which is precisely the half that freezing the object would not solve. The gate therefore anchors on instance identity: two independent adapter instances, two rows inside one instance, the same id read twice, and two arms in one beat are each distinct references; mutating one leaves the others byte-identical; and the subagent lane, which already minted fresh, is the control that proves the criterion discriminates. The main-lane **value** is unchanged — an unregistered id still answers `{lane:"main"}` with exactly one own key and still emits its events, so absence is not turned into a second kind of absence — with ordering pinned three ways (registered-then-read, read-then-registered with no retroactive edit of an already delivered proof, the same id twice) and the id failure classes pinned four ways (unregistered, empty string, absent, non-string, the last two emitting no panel event at all rather than an ownerless proof). Where one row emits **two** events — the terminal-tick and card-close legs, which each yield a lifecycle stop and a panel end — the attribution is decided once (a consumer binding a card between the two yields must not split one row across two lanes) while each event still gets its own proof, so a host consuming them one at a time cannot poison the second before it is even yielded. The run stream leg is covered as the same shape, and a syntax-tree check forbids reintroducing a module-level lane-proof object literal or a module-level `LaneProof`-annotated binding (judged on the type node, not on text, so a compile-time pin tuple that merely mentions the type is not miscaught), backed by a type-checker pass that also catches an un-annotated module-level cache such as `const x = mainLane()` while letting the callable factory itself through, while the module-private three-state sentinels of the untrusted read are frozen instead — only `Object.freeze` counts, never `Object.seal`, which still permits writes to existing keys — their exposure being confined to one module |
403
404
  | `scripts/run-memory-entries-wire-test.mjs` | The two memory-governance capability bits and the three memory-entry response readers. Each bit (`capabilities.memoryCompliance`, for the entry-provenance and erasure endpoints; `capabilities.memoryOrigin`, for the external-origin listing and clearance endpoints) is read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `false` — an older engine simply does not answer, and the right next step is to try the endpoint and read its 501), `true` is the face being mounted, `false` is a positive "not on this deployment" (the wire does not distinguish a backend without control-plane ownership from an empty operator roster, so the wording never guesses which), any non-boolean value is unreadable and drops the cell instead of being folded into "absent", and a capabilities body that is not an object at all is unreadable rather than "not reported". The two bits deliberately stay **two** readers with two separate per-engine tables, because the engine deliberately keeps them two separate product faces even while they happen to carry the same value today: feeding one an unreadable body, or invalidating one, leaves the other's reading untouched, and a body where one is on and the other off is answered one bit at a time. The entry-export reader narrows each row on its own (an empty array really is zero rows, a non-empty array with nothing readable in it is reported as unreadable rather than as "no rows", and partly bad rows are kept with a dropped count), reads the external-origin marker as three states rather than a boolean (the two structural carriers mark a row; a row whose frontmatter cannot be read, or which carries the third, suspended-form carrier, is undecidable, because the judge for that carrier lives in the engine and this package refuses to mint a second copy of it), and treats an unreadable "is this the whole scope" flag as "not the whole scope". Its verdict port implements — in code, not in a comment — the rule that an empty answer is never a clean store: the caller must state whether the request declared origin-awareness, because this endpoint withholds marked entries by default and the two bodies are shaped identically, so without that statement an empty answer is only ever "unknown"; the affirmative answer is scoped to the one named scope and carries that scope with it, and the type has no store-wide arm at all. The erasure receipt reader keeps three things apart that are easy to collapse: "this call erased nothing" (a real receipt whose erased list is empty and whose not-found list explains why, per id), "a 200 with an empty body", and "a body that could not be read" — at the reading, the counting and the verdict layer alike; it refuses a version envelope it does not recognise instead of reinterpreting it, treats the three closed vocabularies as closed (an unknown word is unreadable, never folded into a known arm), keeps an unreadable binding as unknown instead of claiming "unbound", passes the "history cannot be judged" flag through as four states (set, explicitly unset, absent, and present-but-unreadable — an unreadable flag is kept distinct from an absent one, and the history verdict then answers "unknown" rather than the stronger claim), and answers the replay question as three states so that the degraded lane is never retried automatically. The clearance receipt reader carries the cleared marker through verbatim and says separately whether it was reported at all. Every array in every response is snapshotted once — the length is read exactly once and each index exactly once, rather than iterating the caller's own iterator — because an array that reports one length while being walked and another afterwards could otherwise have a marked row quietly dropped while the "was anything unreadable" check saw nothing, which ends in calling the scope clean; an array that reports an absurd length is reported as unreadable rather than silently truncated to its first rows. All three readers never throw. |
@@ -407,6 +408,10 @@ public-surface guard checks that last one).
407
408
  | `scripts/run-task-request-omission-receipt-test.mjs` | Where every key a client hands to the request constructor ends up. The constructor used to answer "not stamped" the same way for four different reasons — value absent, no such row, wrong lane, live gate closed — and a key it had never heard of did not even get that: an unattended run could pass a system prompt, an output schema and a spend cap and receive a body holding the objective and the session id, with nothing anywhere saying what was left out or why. The guard pins the three answers apart. **Seated** keys reach the body verbatim on the unattended lane. Keys the package **knows but did not carry** never throw, never reach the body, and each gets a receipt row with one word from a frozen cause list — every present key is on the body or on the receipt, never both and never neither, checked across both lanes with the live gate open and closed against a key-by-key table written independently of the package's own routing. Keys the package **does not know** are refused loudly and are a separate cell, not a fourth cause: the cause list has no word that could hold them, and the seat reader answers `unknown`, not `none`. The cause list is bitten from both sides (exact, every word producible, nothing produced outside it, the judge table's keys read from source through the TypeScript parser) and no second hand-copied list may exist in `src/`. Upstream claims are read straight off the installed SDK typings: a key seated in this release must be a named request field, a key registered as having no upstream counterpart must not be — the day it appears the guard turns red — and the index signature counts as evidence for nothing. |
408
409
  | `scripts/run-session-policy-wire-test.mjs` | The per-session tool-rule face: the capability bit that says whether an engine keeps such rules at all, and the narrow read plus tightening orchestration built on it. The bit is read the same four-state way as its sibling capability readers — an absent key is not reported (this binary predates the position itself, which says nothing about whether the face exists), `true` is present, `false` is a positive absent (this deployment keeps no per-session rules), any non-boolean value is unreadable and drops the cell rather than being folded into "absent", and a capabilities body that is not an object at all (an array included) is unreadable rather than "not reported". Its single-source verdict answers whether to show the tightening entry: only an engine that says yes is `yes`, both a positive no and a binary too old to answer are `no`, and never having observed a capabilities body is `unknown`. Whether to put a request on the wire is deliberately a **different** question with a different answer for that last state, and lives with the orchestration. The read narrows three ways that must not collapse into each other: a record that really is empty (present, version zero — what an engine answers for a session no rules were ever written for), a record that cannot be read, and a call that failed with a typed disposition. A half-bad record — one rule bucket well-formed and another the wrong shape — counts as unreadable in full, because the write verb replaces the whole record: dropping the bad bucket and writing the rest back would empty it, which relaxes the rules while the caller sees a 200. An unreadable version stamp is never filled in with a zero, a bucket that reports an implausible number of entries is unreadable rather than walked or truncated, each array's length and each of its indices are read exactly once, and a throwing accessor is unreadable rather than propagated. Every load-bearing key is read as an own property — the envelope, the version stamp, each of the five buckets and each array index — because a prototype lookup would let a polluted prototype put a bucket into the reading that the wire never carried, and since the write replaces the whole record the union would then write that invented restriction back as a real one; a guard pollutes the object and array prototypes in place and proves all four shapes stay out. Because the write replaces the whole record, adding a restriction means writing "what is already there, plus the new entries": the union only ever adds, de-duplicates verbatim, keeps a bucket that is present but empty (present-and-empty and absent are opposite meanings, and dropping it would relax the rules), mints no bucket neither side had, copies entry bytes as they came (no trimming, sorting or path rewriting — those judgements belong to the engine), and takes its bucket names from the engine's own type surface rather than a hand-copied list, so a new bucket upstream is a compile error instead of a silently dropped one. Which differences count as relaxing is the engine's judgement and is never re-implemented here: a refusal on those grounds is reported verbatim, never swallowed and never retried. The orchestration is guarded on three axes. Timing: when the record moves between the read and the write, it re-reads and re-writes **exactly once** — two reads and two writes, no more — and the second attempt's union carries the other writer's entries, which is the entire point of re-reading; a second collision is reported rather than retried a third time, and an uncontended write makes exactly one round trip. Concurrency: an explicit barrier holds both orchestrations first reads at the same version before either may write, and the criterion is how many times the store actually rejected a stale version rather than how many writes it saw — the latter is equally true of two serial successes, so it would stop detecting contention the day the interleaving changed. Under real contention the store rejects exactly once, both writers land on strictly different versions, both writers entries survive in the final record, and the round trips are exactly three reads and three writes; the same two orchestrations run serially are asserted to reject zero times in two reads and two writes, which is what proves those numbers are discriminating. On a store where every write loses the race both report a collision having written exactly twice each. Failure classification: a relaxation refusal, a refusal to stamp a version the store cannot establish (the same status code as the relaxation refusal but a different machine code, and folding it into that arm would send the caller off to edit entries that are not the problem), a missing session, a deployment without the face, an ownerless session, a collision code, a bare conflict with no machine code, a rejected body and an unauthorized call each land on their own arm — the collision arm is matched on the machine code verbatim rather than on the status, because two different situations share that status and only one of them is worth retrying. The remaining split is not "which code is this" but "did the engine answer at all": an answered client-side refusal is allowed to say nothing was written, because every such refusal on this endpoint is emitted before the record is touched, while a throw with no answer at all — a dropped connection, a timeout, a response body the transport itself could not decode, a server fault — can only say "unknown", since that throw may well have happened after the record was already saved. Two guards prove that is not theoretical: a write whose receipt cannot be read, and a write that throws after the fixture store has committed, both leave the record changed. Neither is success nor failure: the only honest answer is "unknown", it carries no version, and it is never retried. The direction of the change is likewise never claimed. The engine’s tighten-only rule is an identity gate, not a field gate — for a principal the deployment treats as an operator it does not run at all, so a union that adds a name to an existing allowlist is accepted and really does widen it. This package does not mint a second copy of that rule, so what it reports is the fact it can stand behind: the record now holds what it already had plus the entries sent here. The sentence for a saved write is pinned to contain no claim of tightening, narrowing or restriction, and a guard reproduces the operator case to prove the widening is real while the wording stays honest. Every sentence the module mints is checked pairwise distinct, with the receipt-unreadable one required to keep its "may already be in effect" and the record-unreadable one required to say nothing was written. |
409
410
  | `scripts/run-persisted-rule-write-test.mjs` | The **single-step tightening write** for persisted permission rules — the dual of the revoke surface, and the half where a hopeful reading is expensive. The two behaviours this entry accepts are **derived** from the three-state vocabulary by subtracting the widening one, never hand-copied: the guard bites in both directions (every word in the derived table is really accepted; every constructed outsider — casing variants, trailing whitespace, the widening word itself — is refused before a single round trip), keeps a word-count canary against the parent table, and pins that the source file contains **exactly one** array literal carrying two or more behaviour words, so a second hand-written table shows up as a boundary failure rather than as drift nobody reads. A standing approval is minted by answering a permission question or by importing settings; this entry is not a third route, and the widening word is unspellable in the type. The outcome is a discriminated union whose two failure arms are **not** interchangeable: ten refusal causes each promise the same single thing — not one byte reached the store — while three separate words say the opposite, that the outcome could not be read at all. The service's own "I cannot tell" (a write that could not be confirmed as standing: store wobble, a redemption leg with no decidable ending, or a write that landed and was revoked concurrently before the read-back) stays in the second group, because announcing "nothing was written" invites a clean retry that is not clean, and announcing success misreports a tightening that may already be gone. Anything the shared failure classifier does not recognise defaults to the same place — this is a non-idempotent verb, so "unclassified" must mean "unknown", never "no write": a 500 can happen after the store commits. A 2xx whose body cannot be read is pinned in the same direction and from both sides: it reads as unknown, and the unknown arm carries **neither** the revision nor the written row, so a consumer cannot even spell the shape that would let "unreadable" pass for "written". `persisted` is guarded against the reading everyone reaches for first: it says *this call wrote*, not *a new rule now exists* — an equivalent rule already in the store still mints a fresh causal point, so the lane honestly reports `persisted`, and the material for judging whether the **logical** rule is new (the approval ledger on the returned row) is handed to the caller rather than folded into the discriminant, since the package does not have the one fact that judgement needs. The returned row goes through the **same single narrower** the listing surface uses — proven by running one row corpus through both legs and asserting the two verdicts agree entry for entry (an adversarial pass that forks the listing leg back into an inline copy reds here immediately), plus a source pin that the predicate is defined once and called from exactly the two legs. That sharing is what keeps a row whose behaviour cell is unreadable **visible in both places** rather than hidden by one of them — and the shared narrower is deliberately followed by a second, *different* question only the write leg can ask: is the row that came back **the rule that was just sent**? A rule's identity is a triple, so a receipt missing its behaviour cell, carrying the sibling state, carrying the widening one, or naming another text or another scope is not evidence that the requested tightening is standing; it reads as unknown with its own word, kept distinct from "unreadable" so the two stay tellable apart, while display and derived cells may vary freely. An adversarial pass found both of the gaps this pins: the receipt check that only looked at whether the row was renderable, and a subtler one — pulling the verb off the injected port and calling it bare drops the receiver, so a host that hands over a real resource object (a class instance whose verbs reach the transport through `this`) would see every write throw and be reported as "could not tell", retry after retry, while the package's three other ports call their verbs as methods and work fine. Both are pinned from the failing side: a shorthand-method facade and a class-instance facade must reach the transport and return a real outcome, with the bare-call throw proven to be a real failure mode first. A 405 is split in two, because only the engine's own bare code is evidence about **the engine**: with it, the path exists and this verb does not, so this worker predates the verb; without it — an absent, empty, or foreign code, which is what a proxy or gateway blocking the method typically returns as HTML or an empty body — what was seen is that the verb was refused, while **who** refused it and **at which hop** is unknown, so it lands in the unreadable-outcome arm with its own word rather than sending someone to upgrade a worker that is fine, hiding an entry that is live, or — the part a second adversarial pass insisted on — promising that nothing was written. That promise is what the refusal group means, and a middlebox is free to forward the request and only then answer 405 on its own policy, so a caller who skipped reconciliation on that word would leave behind a standing refusal the user believes never took effect; the guard pins exactly that shape, with a double that writes the rule and *then* answers 405, and with the engine's own bare code still landing in the refusal group beside it. (The same passes caught the naive status-only reading and the asymmetry where an empty-string code fell through to a different bucket.) The documented recovery for an unreadable outcome — retry, then reconcile — is pinned to be **ledger-safe** rather than merely asserted: against a double that models the engine's own "is this identity already standing?" question, re-sending the same identity comes back as a no-op with the approval ledger and the bucket revision both unmoved, however many times it is repeated, while two concurrent writers each landing a causal point are both honestly reported as having written. Finally the four local gates are pinned to be free: a missing write verb on the injected port, an unwritable direction, an unreadable identity pair, and a principal key that is present but cannot name anyone all refuse **without sending anything** — the last of those because silently degrading a blank target into absence would land a tightening aimed at someone else in the caller's own bucket and return a 200 |
411
+ | `scripts/run-registrar-tables-test.mjs` | The four registrar table bodies — this Guards table and the three census tables in the repository's negative-control record — are **generated** from `scripts/gates-manifest.json`, the one file that describes a suite. Every row's text must equal what the generator emits, the manifest's suite set must equal the suites on disk, each entry must declare how it is negative-controlled (rehearsed, blind, or behavioural, with the census taker itself declared as such since it does not appear in its own tables), and each row's outward prose is scanned against the published-surface word list — the same list the packaging-hygiene guard uses, shared rather than copied — before the generator may write it into this file. Adding a guard is therefore one manifest entry plus one generator run instead of six hand edits across three files, and a description that drifts in one place and not the others stops being expressible. The row count is no longer what is compared: the earlier arrangement checked the census tables by **length**, so rows naming the wrong suites reconciled green. Positive controls run entirely on in-memory copies — a changed description, a dropped entry, an added entry, a changed class and a hand-edited row on disk each have to make the same judgement speak — and the quieter halves are pinned too: nothing outside a table body may move, a line inside one that is not a recognisable row makes the generator refuse rather than drop it, byte equality is backed by a column-count check (a cell holding a bare pipe splits a row into extra columns, and a code span does not protect it), and the malformed rows kept byte-for-byte as they are found are registered individually, so the registration turns red the day it stops being needed rather than outliving its reason — audited in both directions, since a registration pointing at a row that is no longer malformed and one pointing at a guard that was reclassified or deleted are both exemptions nobody reads |
412
+ | `scripts/run-mcp-probe-face-test.mjs` | The engine's **run-free MCP status face** (server ≥7.93.0, sdk 11.2.0): the `capabilities.mcpProbe` bit read the same four-state way as its eleven sibling capability readers, and two call ports on top of it — read the deployment's own declared servers, or probe a caller-supplied list. Presence of the bit is carried by the engine version, so an **absent key means an older engine** (that route answers a coded 404) and is read as *not reported*, never as *no*; an explicit `false` is the engine's own no and is reported without inventing a reason (whether caller-supplied declarations are accepted is a different bit's question); a non-boolean is malformed and is dropped rather than folded into a no. The availability verdict answers only *should this call go on the wire*: an explicit no means zero requests, and both kinds of *cannot tell* are sent anyway, so an older engine answers with its own coded refusal instead of being silently skipped. One failure judge serves both ports and asks **provenance before status**: a 4xx, 501 or 503 that carries no machine code proves nothing about who answered and is reported as *no verdict*; twelve coded refusals each get their own arm (three identity codes, one of which sits outside the `auth.` family so a prefix fallback would miss it; three different treatments ride the same 400), and a coded answer this version does not recognise lands in *cannot tell*, never in *this engine has no such face*. Retry-after seconds ride 429 and 503 and are absent rather than 0 when the engine gave none. The 200 body is narrowed through the **same per-row narrower** as the streaming `wiring_manifest.mcp[]` leg and the session panel's replay leg, liveness cell included; on this face rows are paired with the submitted declarations **by index**, so a dropped row makes the whole answer unreadable rather than a half table, a row count that differs from the submitted count is reported as misaligned, and an honest empty `servers: []` is kept apart from *non-empty but nothing readable*. `probedAt` and `ttlSec` pass through untouched — this package mints no freshness verdict — the declaration list is handed over as-is with its length snapshotted once, a list that serialises itself differently from what was counted is refused locally with zero requests, and no port ever retries a dial. |
413
+ | `scripts/run-sdk-wire-transit-test.mjs` | The package's pass-through of a few SDK names (`sdkWireTransit`), pinned: every value re-export is the **same reference** as the SDK's own (a client class the host recognises with `instanceof`, the two approval-frame predicates, the three session-bundle calls and the SDK error class), not a look-alike wrapper — wrapping would discard the one anti-drift guarantee a pass-through has; every type re-export is present by name in the emitted declaration file; the gate's list and the source file's export lists are compared in both directions so a name added to one without the other turns red; and a name the SDK does not export must fail the same test, so the gate is not vacuously green. |
414
+ | `scripts/run-sdk-registry-transit-test.mjs` | The package's cloud control-plane surface (`sdkRegistryTransit`), which lives behind its own `./registry` subpath entry point rather than on the root barrel, and this guard holds both halves of that decision. Upstream publishes the same surface behind a subpath of its own, because the subject differs: the engine-wire surface speaks for one engine's service credential, this one for a person's rotating token, and their refresh and error semantics were deliberately never merged. Keeping it behind a second entry point means a client that never touches the control plane neither resolves nor type-checks it. Every value re-export is therefore read from the file the subpath entry actually resolves to, and must be the **same reference** as upstream's own (the control-plane client class, the three config reads, the health probe, the feedback call, the auth-path constant, the content-address helper, and the two typed error classes a host recognises with `instanceof`); every type re-export is compared with the emitted declaration file in both directions; the gate's list and the source file's export lists are likewise compared both ways; and the root entry is checked to carry none of these names, with the root barrel's own source checked to reference the entry file nowhere — a name leaking onto the root would put the cost of this surface back on clients that never asked for it. The subpath is then verified end to end: the installed SDK must really publish its own `./registry` entry and declare every transited name inside it, and this package's own `exports` must point that subpath at exactly the files the gate just judged. Portability is two checks rather than one, done with a parser rather than a text search: the upstream subpath's emitted JavaScript, walked recursively, must contain neither a `node:` specifier (static, side-effect, dynamic, `require` and re-export forms all exercised) nor a Node **global** — because the same package's third entry point is a Node-only surface that imports nothing at all and reaches for the `Buffer` global, so a specifier check alone would pass it as isomorphic, while a byte-level search of it reports two `node:` hits that live entirely inside a documentation example. A text scanner that merely strips comments first gets both directions wrong on ordinary JavaScript — a regular-expression literal containing a slash pair swallows the rest of its line, and the word in a string reads as a reference — so both scanners run off the syntax tree and are checked against fixtures for each failure direction as well as against that real material — including a dynamic import written with a template literal, which a check that accepts only quoted strings misses entirely, and a dynamic import whose target cannot be determined statically, which is refused rather than read as no edge at all. Zero-processing is likewise enforced with a syntax-tree allowlist rather than a keyword search: every top-level statement must be a named re-export carrying that one specifier, so an import followed by an in-place edit of the upstream prototype is refused with a file and line — that shape leaves the name lists untouched and even keeps the same-reference check green, since both sides are then the one object that was damaged. The last section measures, rather than merely notes, one declaration-level gap: two of the upstream subpath's declaration files reference a package the SDK lists only among its own dev dependencies. Moving such a check to a scratch directory is not isolation, because package resolution walks the ancestor directories, so the gate builds a sandbox served by a restricted compiler host and proves the isolation both ways — a decoy copy of the missing package placed one level above the sandbox must silence the errors for an unrestricted host and must not silence them for the restricted one. Then it installs this package into that same sandbox as a real consumer would, and pins the two readings that justify the entry-point split: a consumer that imports only from the root sees no unresolved-module errors at all, while a consumer that imports the subpath sees exactly the two, reported honestly rather than swallowed by this layer. The day upstream ships those declarations, that section turns red and the note comes out with it. The separation itself rests on the root closure being computed correctly, so the portability guard that computes it was extended in the same change: a template-literal dynamic import is followed like any other edge, and an edge whose target cannot be resolved statically is refused on every one of the four graphs — without that, a single line in a third file already reachable from the root would put this surface back into the root runtime while every guard stayed green. |
410
415
 
411
416
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
412
417
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**