@sema-agent/client-core 0.84.0 → 0.85.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (101) hide show
  1. package/CHANGELOG.md +247 -0
  2. package/README.md +27 -13
  3. package/dist/abortableSleep.d.ts +10 -0
  4. package/dist/abortableSleep.js +37 -0
  5. package/dist/adapt/arms.js +1 -1
  6. package/dist/adapt.d.ts +1 -1
  7. package/dist/adapt.js +1 -1
  8. package/dist/adapter/downstream/terminalToSdkResult.d.ts +1 -0
  9. package/dist/adapter/downstream/terminalToSdkResult.js +19 -10
  10. package/dist/adapter/runStream.js +6 -2
  11. package/dist/agentSession/engineSessionBackgroundStop.d.ts +50 -0
  12. package/dist/agentSession/engineSessionBackgroundStop.js +291 -0
  13. package/dist/agentsWireCaps.d.ts +2 -1
  14. package/dist/agentsWireCaps.js +8 -0
  15. package/dist/cloudConfigWireCaps.js +2 -1
  16. package/dist/corruptSession.d.ts +3 -0
  17. package/dist/corruptSession.js +29 -0
  18. package/dist/corruptSessionCopy.d.ts +1 -0
  19. package/dist/corruptSessionCopy.js +4 -0
  20. package/dist/decideFailureNote.d.ts +2 -1
  21. package/dist/decideFailureNote.js +8 -3
  22. package/dist/decideReceipt.d.ts +6 -0
  23. package/dist/decideReceipt.js +30 -1
  24. package/dist/detachWire.d.ts +1 -0
  25. package/dist/detachWire.js +13 -5
  26. package/dist/displayUntrusted.js +114 -18
  27. package/dist/effortWire.d.ts +2 -1
  28. package/dist/effortWire.js +2 -4
  29. package/dist/engineErrorCodes.d.ts +11 -0
  30. package/dist/engineErrorCodes.js +32 -0
  31. package/dist/engineNoticeCodes.d.ts +39 -4
  32. package/dist/engineNoticeCodes.js +117 -161
  33. package/dist/engineWireFor.d.ts +43 -0
  34. package/dist/engineWireFor.js +399 -0
  35. package/dist/engineWireSdk.d.ts +6 -0
  36. package/dist/fileHistoryCaptureCapability.d.ts +4 -1
  37. package/dist/fileHistoryCaptureCapability.js +43 -9
  38. package/dist/gateVocabulary.d.ts +1 -2
  39. package/dist/gateVocabulary.js +1 -22
  40. package/dist/generated/engineNoticeTables.d.ts +3 -0
  41. package/dist/generated/engineNoticeTables.js +156 -0
  42. package/dist/generated/toolNameTables.d.ts +2 -0
  43. package/dist/generated/toolNameTables.js +43 -0
  44. package/dist/headlessPermissionModeWire.js +2 -4
  45. package/dist/hitl/askGateWire.d.ts +1 -1
  46. package/dist/hitl/askGateWire.js +1 -1
  47. package/dist/hitl/frameRouter.js +4 -2
  48. package/dist/hitl/hitlHostSurface.js +1 -1
  49. package/dist/hitl/parkResolver.d.ts +1 -0
  50. package/dist/hitl/parkResolver.js +11 -3
  51. package/dist/hitl/planReviewWire.d.ts +41 -3
  52. package/dist/hitl/planReviewWire.js +264 -49
  53. package/dist/hitl/rosterMountedNames.d.ts +2 -0
  54. package/dist/hitl/rosterMountedNames.js +38 -0
  55. package/dist/hitl/sessionPolicyDeliverable.d.ts +8 -2
  56. package/dist/hitl/sessionPolicyDeliverable.js +39 -48
  57. package/dist/hitl/sessionPolicyWire.d.ts +40 -0
  58. package/dist/hitl/sessionPolicyWire.js +173 -0
  59. package/dist/hitl/toolApprovalWire.d.ts +4 -2
  60. package/dist/hitl/toolApprovalWire.js +6 -4
  61. package/dist/index.d.ts +10 -3
  62. package/dist/index.js +7 -2
  63. package/dist/liveInitToolFace.js +28 -10
  64. package/dist/permissionWireCaps.d.ts +2 -1
  65. package/dist/permissionWireCaps.js +2 -7
  66. package/dist/request/taskRequest.d.ts +2 -0
  67. package/dist/request/taskRequest.js +18 -1
  68. package/dist/resumeRefusalCopy.d.ts +11 -0
  69. package/dist/resumeRefusalCopy.js +42 -2
  70. package/dist/runTerminal.js +4 -0
  71. package/dist/sandboxWire.js +1 -1
  72. package/dist/sealedKeyCapability.d.ts +17 -0
  73. package/dist/sealedKeyCapability.js +42 -0
  74. package/dist/seatContract.d.ts +2 -1
  75. package/dist/seatContract.js +5 -21
  76. package/dist/subagent/engineCompactWire.d.ts +13 -2
  77. package/dist/subagent/engineCompactWire.js +48 -37
  78. package/dist/subagent/engineDelegatedPrompt.d.ts +9 -1
  79. package/dist/subagent/engineDelegatedPrompt.js +16 -16
  80. package/dist/subagent/engineRowStopGate.d.ts +2 -2
  81. package/dist/subagent/engineRowStopGate.js +4 -4
  82. package/dist/subagent/engineSubagentOutput.d.ts +8 -0
  83. package/dist/subagent/engineSubagentOutput.js +37 -18
  84. package/dist/subagent/engineSubagentResume.d.ts +11 -2
  85. package/dist/subagent/engineSubagentResume.js +28 -14
  86. package/dist/subagent/engineSubagentSteer.d.ts +10 -1
  87. package/dist/subagent/engineSubagentSteer.js +16 -11
  88. package/dist/subagent/engineSubagentTail.d.ts +10 -1
  89. package/dist/subagent/engineSubagentTail.js +35 -17
  90. package/dist/subagent/engineTaskHandleWire.d.ts +13 -0
  91. package/dist/subagent/engineTaskHandleWire.js +41 -20
  92. package/dist/systemReminderTag.d.ts +5 -0
  93. package/dist/systemReminderTag.js +20 -9
  94. package/dist/taskRequestWords.d.ts +5 -0
  95. package/dist/taskRequestWords.js +4 -0
  96. package/dist/wireErrorTriage.d.ts +1 -0
  97. package/dist/wireErrorTriage.js +3 -1
  98. package/dist/wireRefusalCopy.d.ts +10 -0
  99. package/dist/wireRefusalCopy.js +62 -2
  100. package/docs/INTEGRATION-CLIENTS.md +1128 -22
  101. package/package.json +1 -1
package/README.md CHANGED
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
35
35
 
36
36
  ## Scope
37
37
 
38
- **Version:** 0.84.0
38
+ **Version:** 0.85.0
39
39
 
40
40
  - **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
41
41
  B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
@@ -301,16 +301,16 @@ guard still cross-checks the table by name).
301
301
  | `scripts/run-engine-caps-ledger-test.mjs` | A per-key disposition ledger for `GET /v1/capabilities`. The SDK's `Capabilities` grew from 74 keys to 93 in one release and nothing on the board could see it: this package consumes that table through four synchronous readers, and *nineteen new positions arriving while the package does not move* is exactly the disease shape this repo keeps logging on other axes — the fact is already on the wire, the package boundary is the cell that swallows it, and no client can read it however they write their side. So the ledger is reconciled **element-wise against the SDK interface in both directions**: a key the SDK added with no ledger row is red (someone must classify it), and a row for a key the SDK removed is red too (a registration that no longer does anything). Each row then has to survive its own claim — a `read` row names the source file, and the **code** there (comments stripped) must really mention the key, because prose asserting an alignment is the classic way these guards go hollow; a `not_read` row must have **zero** read sites in the tree, so wiring one up while the ledger still says the package ignores it is red rather than invisible. The census behind those two directions recognises five call shapes, each of which really occurs here — a reader whose base argument carries its own parentheses, a direct `caps.<key>`, a narrowing cast, an own-property read helper, and a `*_CAP` constant — and proves it on fabricated samples first, since a census that recognises one shape reports "nothing here" for the other four. What the guard deliberately does **not** judge is whether a position *ought* to be read: that is a design call, and the ledger only pins that every capability was looked at once by a person and that what they wrote down does not contradict the code |
302
302
  | `scripts/run-sql-engine-capability-test.mjs` | The SQL-posture read face and the four-state capability reader underneath it. One capability cell here carries **four different things**, and each one points an operator somewhere else: nothing has been observed yet in this process (a one-shot doctor run is always in that state), the response arrived but carries no such key (an older engine), the engine explicitly answered `null` — *this deployment has no SQL backend*, which is a **positive fact** rather than an absence — and a full reading. Fold any two together and the screen states something flatly, confidently, and wrongly, so every positive control here is paired with a control pointing the opposite way, and the four sentences the doctor row can print are checked to be pairwise distinct and non-implying. The reading itself is narrowed no tighter than the mint: `txnMode: null` is a **legal value** — two of the three engines always report it that way, and the upstream type note names reading it as "optimistic" as the error — so treating it as malformed would throw away the entire reading for ordinary deployments, which is the same disease this repo logged when a consumer's domain was narrower than the producer's. A response that cannot be parsed **clears** the cell rather than leaving the previous engine's answer in place, and a separate invalidation port exists for the case the generation latch cannot catch — a same-port respawn whose new probe never succeeded, where the stale reading would otherwise be answered as current fact. Untrusted values (the isolation string is read back from a database server variable) are sanitised and bounded before display, and the bound is applied **before** escaping so a visible escape never gets cut in half. Finally the export names are themselves a guard: the shell still carries a copy that is meant to go red on the package's same-named export and be swapped out, so renaming anything here would silently disarm that lock |
303
303
  | `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
304
- | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there. From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. |
304
+ | `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there. From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. From engine `7.104` a synchronous submit that stops at a gate returns the engine result itself plus a three-key receipt: it now carries the tagged cause, so it reads as `paused` on the cause generation (older engines still send the flat three-key body, which keeps reading as before), and the top-level park word is looked at first, the same order the SDK documents for all three generations, so a body the server says is parked is never read as finished. A parked run row now carries its result too, and the headless reconnect path turns it into the parked terminal frame on the first attempt instead of spending its whole retry budget — both generations are pinned, including an end-to-end drive through the public reconnect entry point. |
305
305
  | `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
306
- | `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. |
306
+ | `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails. From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. The server also mints a few notices of its own through the same channel; those codes live in a second table with their own audiences, kept apart from the engine mirror (which must stay equal to the engine's catalog) and reconciled against the server's published package when one is supplied, so a user-facing server notice is no longer filed under operations. One dispatcher returns the typed facts for every code that has a reader, discriminated by code and tagged with its audience — only the user-audience codes belong on a user surface — and the guard ties the dispatch table to the module's own exported readers in both directions, so a reader cannot be exported without a row and a row cannot be dropped without the guard failing. From 0.85.0 it also covers a third server-minted notice, the one saying that part of a session's saved history could not be read when the session was reopened: it is a user-audience notice, its two counts are read one by one and anything that is not a non-negative integer reads as `unknown` rather than zero (neither "nothing was skipped" nor "nothing is left" may be invented), and one extra sentence — the context is empty but the turn runs — is given only when the count of entries left is exactly zero. When the server package is supplied, the guard drives the server's own emitter for that notice and reads what it emits back through the dispatcher. |
307
307
  | `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever. One reading here answers a question that the terminal state structurally cannot: whether this run was assembled with any file-and-shell tools at all. The engine's terminal vocabulary says a run finished, not whether the work got done, so an orchestrator that waits for the end and then guesses has nothing to guess from — while the assembly manifest already said it at the start, one row per mounted instance with the single condition that mounted it. The reading is three-state and both folds are refused: a roster that is readable and carries no such row is the engine stating a fact, while no roster at all is not that fact — the static half of a manifest never carries one, and an older engine reports rosters without naming the mount condition at all, where an empty count would be a statement about the reader rather than about the run. Those two are kept apart in the reason the reading carries, and the wording for every unknown case is checked never to claim the run had no tools. The same roster now decides the tool list on the first line of a non-interactive run: the host holds that line until the roster arrives and lists exactly what the engine mounted at the start of the run, in mount order. The guard runs a real assembly frame through the projection into the decision, and pins that the host falls back to the estimate only once the roster is known not to be coming — a manifest without one, an unreadable one, model output or the run's end arriving first — rather than on a timer alone (model activity counts, including a model call that is still waiting or retrying; an error line the stream synthesizes when a run fails before assembly counts as the run ending), that a sub-run's manifest is never mistaken for the run's own, that an empty roster is taken as the engine's answer rather than as silence, and that the wait bound covers both sequential default budgets the engine gives an external tool server to connect and list its tools. The holding logic itself lives in the package as a small per-run gate — buffer, decide once, release the held messages in arrival order, then pass through — and the guard drives real stream output through it to pin that the release happens exactly once, at the manifest, releasing exactly the held prefix. The ordering itself also lives in the package as a stream wrapper, and the guard checks the final output a consumer reads: the first line is always the tool-list line, a message that arrives while that line is still being built comes after it, a timer firing races nothing out of order, a source that ends or fails before the decision still gets its first line and held messages out before the error, and an early exit closes the source. From 0.84.0 the roster-derived sentence source no longer throws on a value it does not recognise, including a reading of the manifest's `hands` section passed by mistake: it answers the same "not stated" sentence as the `hands` reader, from one shared source, and its six known sentences do not change. |
308
308
  | `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
309
309
  | `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, ten words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because an unknown word means different things in each: a denial layer this build does not know may have been added by a newer engine or may come from a damaged record, so its sentence says it cannot tell which instead of asserting damage; the asker vocabulary is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine). Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer; it must not say the question is asked every time — an answer for this one call may come from the person, a hook or an automatic check the deployment runs — and its wording is checked against the engine package's own description of the mandate. A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. Since 0.83.2 a sixth list covers the word a mandated question stands on (`APPROVAL_MANDATE_WORDS`, six words): it must equal the engine's own list word for word and in order, membership is exact, the card reader `readApprovalMandate` answers only for an own key holding one of the six words, and each word has one fixed sentence explaining why the question must be confirmed — six distinct sentences that never point the reader at writing a rule, never promise a question every time, never claim only a person may answer, and repeat no other sentence the package mints. The list is also pinned against the engine's type at compile time in both directions, while the published build references no engine package at all: every `.js` and `.d.ts` file in the build is scanned, and the same scan is first shown to fire on references planted in a scratch directory. |
310
310
  | `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
311
311
  | `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
312
312
  | `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key on the success result: the CC-spelled `structured_output` is the only home for the value the wire calls `structuredOutput`. The camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rode alongside it for exactly one release (0.79.1) and is **absent from 0.80.0 on**, pinned both by own-key and by `in`, so a consumer still reading the old name sees `undefined` rather than a stale copy. The wire position is read exactly once, so a value-changing accessor is only ever asked for its first answer; absence stays absence; a wire key that is present but `undefined` mints nothing, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` are still minted, and so are shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope carries no such key, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
313
- | `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the sibling client trees are on disk the consumption evidence is checked by name — a `contract` row's claimed consumers must equal the real set, and a `retire` name must not be imported by any client. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. |
313
+ | `scripts/run-export-liveness-test.mjs` | Every runtime export in the public baseline must be **alive**: referenced by some gate, or explicitly registered in `scripts/export-liveness.json` as `contract` (consumed by a client with no gate yet), `internal` (an internal helper amplified onto the public surface by `export *`), `candidate` (with ticket + retire-by) or `retire` (dead; retire-by version). Registration is accounting, not exemption: a row for a name a gate already references is stale and must go, a row for a name no longer exported is red, `retire`/`candidate` rows go red the moment `package.json` reaches their retire-by version, and the row count only ratchets down. When the client repositories are on disk the consumption evidence is checked by name, from committed history and never from a working tree: each client's `origin/main` (or its HEAD when there is no such ref), the checked-out HEAD however old it is, and every local branch head committed within the last 14 days, so an import that so far exists only on an unmerged branch still blocks the retirement. Two readings are kept apart. The registry states facts, so a `contract` row's claimed consumers must equal the clients that really bind the name — a named import from the package or one of its subpaths, a re-export from it, a member read on a namespace or dynamically imported module object, a destructured dynamic import — and an `internal` name that a client binds must be re-registered as `contract`; a client-side declaration that merely shares the name is not consumption. The retirement check asks only whether removing the name could break a client, so it also counts a plain identifier match in any file importing the package, and, when a client re-exports the whole package, any appearance of the name in that client's sources; the guard prints both counts. The guard prints the ref, commit and commit date of everything it read, and it shares one repository reader with the same-name shadow guard, so both guards judge the same snapshot of a client. A client that is absent or is not a git repository is reported as a skipped section, never as a client with no consumers, and a dangling or unreadable branch ref is a fault rather than a branch quietly left out. The reader proves itself on a throwaway repository: an import on a recent branch that is not checked out must be found and attributed to its branch and line, while one on a branch older than the window, in an uncommitted file, or outside the scanned folders must not count. Names that have already left the surface are kept in a per-version `removed` ledger: they must never reappear in the baseline or the registry, and the ledger's versions must not run ahead of the changelog. |
314
314
  | `scripts/run-wire-refusal-copy-test.mjs` | Two wire refusals read the same way on every client: a cancel's 409 carries one of two codes with opposite dispositions (`conflict.approval_settled` — someone else already decided, go read the result; `conflict.run_not_running` — nothing changed, send the cancel again), an unrecognised or codeless 409 is reported as such rather than guessed, and the submit-side 429 `usage.window_exhausted` is read as a waitable refusal whose wait is stated only when the engine supplied one. `ControlRouter.cancel` raises a distinct safety code for the retry-directly case. A third family covers the refusals that a deny's *settler note* can draw: a deployment that signs the decisions it accepts but does not sign that note, a body whose decision and note contradict each other, and a word the engine does not recognise. All three are refused **before** the approval is judged, so each sentence states plainly that nothing was decided and the approval is still waiting, and each names a different next step — none of them “send it again unchanged”, which would simply be refused again. Two of the three codes already mean something else in this package: one is shared with the plan-review leg, so the reader refuses to claim it unless the caller states that the note really was sent, and the other is split by the field the engine names, because without that field the same code means a review outcome was rejected for its content. The third sentence deliberately does not say the word was misspelled: upstream mints that same field for at least four different reasons, so the only thing it proves is that the engine would not take the settlement details and named which part — which is what the sentence says, and where it points. The three sentences are pinned verbatim rather than by keyword, because a keyword check passes a sentence that tells the reader to send the same body again unchanged, which is the one next step that is certainly wrong. The field itself is read from the bag the SDK keeps additional response keys in, not off the top of the error, since only the hand-built shapes a test would write carry it there. Both decision legs hand the reading back on their outcome, and a leg that never sent the note claims nothing. When a decision fails because the engine cannot read a stored row, the row the engine names is read (from the error's top level, or from the SDK's bag of additional response keys where the real SDK puts it) and carried into the failure reason on all three decision legs, escaped for display. |
315
315
  | `scripts/run-tool-disclosure-progress-projection-test.mjs` | The two wire arms sdk 9.6.0 adds — `tool_disclosure` (name-only tool census: open-set `policy`, `thresholdPercent` absent ≠ default, `deferred`/`activated` full snapshots) and `tool_progress` (one frame, two beats: Bash ticks carry an output tail with `totalLines`/`totalBytes` that come and go together; other tools carry only `elapsedSeconds`) — project to neutral internal arms plus chrome arms. Required keys missing ⇒ `malformed`; bad optional keys drop only themselves; the sub-flow three-key gate keeps child frames off the leader lane; both arms are `required: false` in the arm table with duties stated (the output tail is untrusted raw and must never be fed back to the model). |
316
316
  | `scripts/run-mcp-panel-projection-test.mjs` | The `GET /v1/sessions/:id/mcp` panel reader (`projectMcpPanel`; server >=7.77.0 adds the optional `lastLegMcp` key) and the single wording mint for its "last leg" line. Absence of `lastLegMcp` is one literal sentence that never blames the engine version (a new session, a leg outside the retention window, a leg without a manifest and an older engine all look the same on the wire); a key that is present but unreadable is a different sentence plus a `lastLegMcpUnreadable: true` mark, never folded into absence. The `mcp[]` roster goes through the same reader as the live `wiring_manifest` third section, so a replayed roster and a live one have one shape. The two faces of the panel (`servers[]` and the last-leg roster) may legitimately differ, so the view carries no agreement flag and none of the five sentences mentions `servers`. Required keys are pinned to the SDK `openapi.yaml` component bytes **0.69.0:** `fetchMcpPanel` fetches the panel through the SDK client's own `sessions.mcp` call (same transport and auth as every other read) and projects it; transport failure, an unreadable body and an empty session id all come back as `undefined`, never as a fabricated empty panel 0.71.0 adds section K: `mcpEngineLegPresence(view)` — the engine-side MCP presence tri-state read only off the panel view (`unknown` when the view could not be read, never rendered as "no MCP configured") |
@@ -356,7 +356,7 @@ guard still cross-checks the table by name).
356
356
  | `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
357
357
  | `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
358
358
  | `scripts/run-client-core-message-branching-test.mjs` | §B8 (branching on error **text**) and §B10 (truthiness standing in for existence when the value can be `0`). AST + type-checker census over `src/`, a named ALLOW list carrying owner and expiry, a known-site floor, and two fixed corpora with a known verdict judged by the same classifier on every run |
359
- | `scripts/run-registry-test.mjs` | Every ratchet number the suites compare against lives in exactly one place, `scripts/registry.json`: each cell must be a finite non-negative integer, and each one carries a dated ledger of every change (`from` / `to` / `when` / `why`) whose last entry for that cell must equal the number in force — a number moved without an entry, or an entry written without moving the number, is red in both directions. The five suites that read it are checked to read it: no ratchet name may sit next to a numeric literal in their own source (an AST check, so an accounting comment or a test string mentioning the number is fine), each names the cell it reads, and each keeps the step-down channel — when a measurement comes in *under* a ceiling the suite prints one `RATCHET-SLACK` line instead of passing in silence, so slack cannot quietly accumulate under a ceiling nobody lowered. Negative control: a cell rewritten as a string, a fractional or negative cell, a missing cell, a cell with no ledger entry, a ledger entry missing a field, and a ledger tail that disagrees with the number in force each have to make the check speak |
359
+ | `scripts/run-registry-test.mjs` | Every ratchet number the suites compare against lives in exactly one place, `scripts/registry.json`: each cell must be a finite non-negative integer, and each one carries a dated ledger of every change (`from` / `to` / `when` / `why`) whose last entry for that cell must equal the number in force — a number moved without an entry, or an entry written without moving the number, is red in both directions. Every suite that reads it is checked to read it: no ratchet name may sit next to a numeric literal in their own source (an AST check, so an accounting comment or a test string mentioning the number is fine), each names the cell it reads, and each keeps the step-down channel — when a measurement comes in *under* a ceiling the suite prints one `RATCHET-SLACK` line instead of passing in silence, so slack cannot quietly accumulate under a ceiling nobody lowered. Negative control: a cell rewritten as a string, a fractional or negative cell, a missing cell, a cell with no ledger entry, a ledger entry missing a field, and a ledger tail that disagrees with the number in force each have to make the check speak |
360
360
  | `scripts/run-client-core-failloud-test.mjs` | §C1/§C2: an empty `catch` with no comment anywhere inside it, a pure-swallow `catch` nobody reasoned about, and `void <write>` that really returns a Promise with no `.catch`. The exemption instrument is a comment saying why *this* failure may die; the documented-swallow count is a ratchet that only goes down |
361
361
  | `scripts/run-client-core-typeshape-test.mjs` | Type discipline as a guard rather than a build side effect: the set of enabled strict knobs (one silently switched off is red), `tsc --noEmit`, and export-surface ratchets for inline anonymous shapes (≥3 members), `unknown` leaving the surface, and bare `unknown` returns — zero slack in either direction |
362
362
  | `scripts/run-client-core-singleton-test.mjs` | Module-level singletons ⇄ `docs/refactor/p1-scan/singleton-manifest.json`, **both directions**: an unregistered singleton is red (registering it forces someone to answer "what if this got duplicated"), a stale entry is red, and the `dupRisk: high` count only goes down |
@@ -367,7 +367,7 @@ guard still cross-checks the table by name).
367
367
  | `scripts/run-crash-converged-projection-test.mjs` | The `crashConverged` read face on `GET /v1/approvals` (L-38): what the *previous life* of a crashed local engine left behind, projected for every client. Three judgements are pinned. First, **absence is not an empty list** — a missing key (an older server, deps not present, or a carrier that is not an array at all) returns `undefined`, and the client renders nothing; an empty array returns a present zero-count object, which is the server actually saying "none". Folding the first into `{total:0}` would have the client assert "nothing was left behind" on a surface a person uses to decide whether it is safe to re-run something — the worst possible direction for a false statement — so the two cases are pinned to different **return shapes** and a test asserts the two verdicts are unequal. Second, bucketing is a **four-term conjunction**: `orphanState === 'pending'` *and* `resumeSafe === true` *and* both approval-evidence keys (`originalDecision`, `decidedAtMs`) absent. A fifth term rejects any row carrying an **accessor**, and accessors are never invoked at all — reading one means synchronously running someone else's code, and `catch` catches throwing, not *never returning*, so a looping getter would pin the startup thread forever (the row cap does nothing against that shape). The same rule covers the three untrusted reads outside the row as well — the envelope's `crashConverged` key, the carrier's `length`, and every numeric index are read as own property *descriptors* and only data descriptors are used, so accessors and prototype entries read as absent and are never invoked. Such a key is treated as absent: if it was a required field the row is counted as dropped, if it was optional or additive the row survives without it. That also closes the ordering attack, since spreading runs getters in property order and an earlier one could `delete` the approval evidence before it is ever copied (measured before the fix: such a row reached the resume-safe bucket), and the check therefore moves ahead of the read, onto the property descriptors — from which the snapshot is then built directly, because checking descriptors and *then* spreading is two independent observations of the same row, and a non-throwing proxy can make the two `ownKeys` calls disagree (first showing `originalDecision: 'approve'` so the row reads as plain data, then omitting that configurable key so the snapshot loses the evidence; measured before the fix: the dangerous row reached the resume-safe bucket after exactly two enumerations, and after it, one). Keys are written with `Object.defineProperty` rather than plain assignment, because `'__proto__'` is a legal own enumerable key and `o['__proto__'] = x` does not store a value — it calls the prototype setter, letting a row whose own properties are all plain data (so the accessor gate never fires) inject a prototype whose `sessionId` getter deletes the approval evidence from the snapshot during validation; `defineProperty` fires no setter, so the key survives as ordinary additive data and the snapshot keeps `Object.prototype`. A row that simply arrives with a custom prototype is treated the same way, since the snapshot only enumerates own properties: approval evidence sitting on the prototype would never reach it, and a perfectly ordinary object with no proxy and no accessors could otherwise be called safe to re-run — real bodies come from `JSON.parse` and always carry `Object.prototype`, so nothing genuine trips it). Validation itself runs on a **null-prototype** dictionary and the bucketing verdict is carried out of that same pass rather than re-read from the delivered row, because every property lookup on an ordinary `{}` reaches `Object.prototype`: a polluted `sessionId` getter there would delete the approval evidence from the snapshot mid-validation and send the row to the safe bucket (measured before the fix). The row handed to the client is still an ordinary object — the null prototype is an implementation detail of the check, not of the value) — real JSON bodies are all data properties, so only a middle-layer-synthesised payload ever trips it, and it too lands in the human bucket rather than being dropped. The `decided` arm means the human had already approved and side effects may be half-landed, so it always goes to the human bucket, as does `resumeSafe === false` and — the last two terms — any row whose own fields contradict each other, since `pending` claims nothing ran while that evidence says somebody pressed approve. Deciding "not safe" costs one extra question (recoverable); deciding "safe" wrongly has somebody re-run work that already partly happened (not). A 2x2 truth table pins that exactly one cell is resume-safe, so reading either key alone turns red, and the contradictory rows are routed to the human bucket rather than dropped — they are real orphans, and the ones most worth showing. Third, unreadable rows are **dropped and counted**, never thrown and never passed through: the product is declared as `CrashConvergedRow`, so letting a row missing a required field — or carrying one of the wrong type — past would be a lie at the type level, and the closed literal discriminators (`decision` / `cause` / `orphanState`) decide family membership rather than being an open vocabulary. The measuring stick stops at the **type** floor, though: degenerate-but-well-typed values (`ts: NaN`, an empty `toolName`) are kept, because swallowing a real orphan over a decorative field is the worse direction, and the one deliberate exception is `approvalId`, which must be non-empty to be a row identity at all. `dropped` is kept separate from `total` so unreadable rows never inflate "N approvals were affected"; each row is a **one-shot snapshot** — every own enumerable key is read exactly once, and validation, bucketing and the handed-back value all read that same snapshot, so additive upstream keys survive while a **non-idempotent** getter (one that never throws, just answers differently on a second read) can no longer erase the approval evidence between the check and the bucketing (measured before the fix: such a row landed in the resume-safe bucket while its checked value was `"approve"`). Hostile carriers are counted rather than allowed to reject: **every** touch of the carrier is guarded — envelope property reads, `Array.isArray` itself (it throws on a revoked proxy), the `length` read, each indexed read and each row's property reads — and a traversal that dies halfway returns absence rather than a half-counted total. A row that cannot be read never takes the batch with it: its own shape check is inside its own guard, so one revoked-proxy row costs a `dropped` tick rather than collapsing the whole projection to absence — which a client would have read as "this deployment does not offer the surface". Traversal goes by **numeric index, never the carrier's own iterator protocol**, because `for...of` hands the carrier the question of which rows exist: an array carrying an overridden `Symbol.iterator` can yield nothing (measured before the fix: a real orphan became `{total:0}`, which a client reads as "the server said there are none") or swap a dangerous `decided` row for a safe-looking one (measured: `fake-safe` was returned in place of `real-danger`). Row count is capped at 100000 and the cap is checked **before** the walk: requiring only a non-negative integer `length` does not stop a proxy trap reporting a billion, and this surface runs on the startup / `--resume` path, where a synchronous spin freezes the thread (measured before the cap: twenty million rows took 18.3 seconds and twenty million index reads; a billion does not come back). The honest boundary is stated rather than overclaimed — a proxy can still lie in its `length` or index traps, which is the same thing as a host injecting a lying transport — and the widening of `ApprovalsResourceLike.list()` is proven **additive** by really running tsc over a legacy `{ pending }` mock *and* over the real `AgentClient` path — the projector takes `unknown` precisely because a parameter shaped as "an object with an optional `crashConverged`" is a TypeScript weak type that the installed SDK's own `list()` return shape shares no property with, which only a real-client compile would have caught — with a known-red control so a clean run means the checker spoke |
368
368
  | `scripts/run-self-orchestration-denial-test.mjs` | The three judgements behind a **denied self-orchestration request** (server 7.57.0), each of which all three clients would otherwise get wrong on their own. First, whether to retry at all is a **conjunction that may not be loosened**: HTTP 501 *and* an `errorCode` that is **exactly** `capability.self_orchestration_required`. That code shares its shape with every other `capability.*` 501, so dispatching on the prefix would drag "some other capability is not wired up" into the retry arm — those requests do not become acceptable once the two keys are gone, so the client would spend a request and then tell the user the wrong reason. Negative controls cover all four directions: a sibling `capability.*` code, a truncated or suffixed variant of the right one, a codeless 501 (it decides nothing, so it decides nothing — no guessing), and the right code under 500 / 400 / 503 or a string `"501"`. The classifier reads structurally rather than by `instanceof` (a host may inject its own transport; across realms or duplicate SDK instances an understandable error would read as unreadable), so a class instance, a bare `{status, errorCode}` literal and an error carrying those fields on its **prototype** all reach the same verdict — and a hostile proxy or a throwing getter yields `null` instead of throwing, because this classifier runs inside a `catch` block where anything it throws escapes the caller's own guard. Second, removing the intent is a **structural** operation, not wording: `selfOrchestration` sits at the top level while `ultracode` sits under `settings` — two different stamping legs — and a client hand-writing `delete` will miss the second one, which costs the user the same failure twice. The single stripper is pinned to touch exactly those two: other `settings` sub-keys and their values survive byte for byte, `deferTools` is left alone (pulling `Workflow` out would be a behaviour change, not a removal of intent), additive unknown keys survive at both levels, the input object is never mutated, `settings` is only dropped entirely when `ultracode` was really there and nothing else remains (an already-empty one is left as is), a non-object `settings` is not touched at all, an `ultracode` that only exists on the prototype does not count, and the whole thing is idempotent. The end-to-end leg runs a real `buildTaskRequest` product through it and asserts the stripped body still passes the registration gate key by key. Third, on the capabilities body, **absence is not "switched off"**: a pre-7.57 server has no `workflowsGate` key at all, so reading absence as "the engine says no" asserts something the server never said, and the mirror-image disease is folding an **unrecognised** `denial` into `null`, which would have the client render "nothing was denied" when the truth is "denied, for a reason I do not recognise". Five shapes are pinned — caps unreadable, gate absent, closed-set member, unknown value, accessor — with the unknown arm carrying the raw token (or an empty one when the value is not even a string) and never collapsing to `null`. All four untrusted reads go through own **data descriptors** only, and the guard pins the getter invocation count at zero, since `catch` catches throwing but not *never returning*; a descriptor trap that throws and a revoked proxy both yield honest absence rather than an exception — though *what* absence means differs by field, and the guard pins that split rather than a blanket rule: an accessor on `workflows`, `workflowsGate` or `engineCan` reads as absent, while an accessor on `denial` reads as `{unknown:''}`, because a key that is **not there** is the gate saying "nothing was denied" whereas a key that is there but cannot be read is "denied, and I could not read why" — folding the second into the first is exactly the false statement this face exists to prevent. Two further pins came out of an adversarial review. The exported retry list is **frozen at runtime**, not merely `as const`: the verdict hands out that same reference, so any consumer splicing it once would poison every later verdict in the process — the guard asserts `Object.isFrozen`, that four different mutation attempts leave it byte-identical, and that a verdict issued *after* those attempts still carries the original two entries. And the classifier reads `denial` only **after** both criteria have passed, since it is not a criterion but an extra field on the verdict: the guard pins the getter invocation count at zero for any error that does not match and at most one for an error that does. The scope line is drawn explicitly rather than overclaimed — "no getter ever runs" holds for `projectWorkflowsGate`, which reads **wire JSON** where every field is an own data property by definition, but not for the classifier, which reads a **thrown value** that may well be an SDK `APIError` class instance carrying `status` and `errorCode` on its prototype; insisting on own data descriptors there would report a perfectly readable error as unreadable, so that side promises only that it never throws. A final pin covers the **integration document's own worked example** rather than the library: the shipped SDK's `tasks.stream()` is an `async` generator, so calling it issues no request at all — the POST happens inside `streamRaw` on the first iteration, and a `try` wrapped around the `stream(...)` call itself can never catch the 501. A client following a submit-shaped recipe on the streaming leg would never run the classifier, and the whole strip-and-retry path would silently do nothing. The guard drives the **real** `TasksResource` against a fake transport, offline, and pins both halves: the synchronous leg is in flight the moment it is called, the streaming leg has issued zero requests after the call and raises on the first `next()` — and it does so through the **real** error path, with `openStream` returning an actual 501 `Response` that the SDK's own `errorFromResponse` turns into the typed error, pinning the `openStream`→`errorFrom` call order so a transport that stops minting `errorCode` cannot pass. The documented recipe is then **executed** rather than keyword-counted: exactly one retry, a second body that really lost both keys while every other setting survives byte for byte, the caller's own request object left untouched, one disclosure and only one, a second 501 propagating with the request count still at two, and — after the first 501 — an abort leaving the count at one with nothing disclosed. A last leg is type-level: `stripSelfOrchestrationIntent` carries an SDK `TaskRequest` overload, because the wide `Record<string, unknown>` form erases the caller's type and the document's "strip and resubmit" line would not compile without an unsafe cast; a real tsc run over a virtual file proves both the narrow and the wide path, with a known-red control — and it compiles the document's two recipes **verbatim**, extracted from the section itself, because a recipe that does not compile is a recipe that was never given: `{ transientOk: true, signal }` is a TS2379 under `exactOptionalPropertyTypes`, which no amount of prose review had caught. The last thing pinned is the one that would have been quietest of all: the SDK's `stream()` returns only on a `done` or `failed` frame, so a stream truncated mid-run — or yielding nothing at all — ends the `for await` just as normally as a completed one. The documented `runOnce` therefore tracks whether it ever saw a terminal frame and raises when it did not, the guard's success fixture emits a real terminal and asserts the handler received it, and a truncated-stream control asserts that shape is reported as a failure with no retry and nothing disclosed. That terminal-frame rule then needed one more turn of its own: the underlying reader returns *normally* when the signal is aborted, so the check as first written rewrote a user's cancellation into a generic stream fault — a client keying off `AbortError` to suppress the error would instead have shown a failure, or resubmitted. Cancellation is therefore checked first, a real-SDK case aborts from inside the handler and asserts the original `AbortError` survives with no retry and nothing disclosed, and the document is checked for that ordering. The harness runs the documented `handle` and `transcript.note` as real spies rather than pushing frames itself, the drive loop rethrows exactly as the document does, and the disclosure ledger is proven to be the caller's own array by a positive identity assertion — without which the cancellation leg's "nothing disclosed" would have been vacuously true. Each recipe is compiled **on its own**, with a preamble that declares only what a host supplies and injects no library symbol, since compiling them together let the second one borrow the first one's imports, and the preamble's own types are decoupled from what the recipes import so the "remove the imports and it must fail" control fails for the right reason — which is checked by attribution, not merely by redness. Ordering is the last thing to get right: the cancellation check must come before the truncation error but **both** must sit behind the terminal-frame test, because a cancellation that lands after the run already reported `done` would otherwise overwrite a real outcome — one that may have already had effects — with "cancelled", and a person reading that will run it again. Aborting from inside `handle(done)` and `handle(failed)` are both pinned to still report success, and the ordering assertion is anchored inside the streaming `runOnce` body rather than the section, since the section's first `throwIfAborted` belongs to the synchronous recipe and would have made a reversed streaming recipe pass — and that ordering check is now anchored on the TypeScript AST rather than on text, since a comment reproducing the two statements in the right order let a genuinely reversed body pass. One more timing fact had to be written into the recipe: a single SSE read buffers several frames and the SDK yields them back to back, so checking the signal only after the loop lets a cancelled run keep consuming the rest of the chunk — measured, an abort inside `handle(turn_start)` still swallowed the `done` that followed and reported success. The recipe therefore re-checks after every non-terminal frame. Finally, the behavioural matrix is no longer run against a copy of the recipe: both recipes are extracted from the document, transpiled, and **executed** with injected host objects, so the disclosure assertion really exercises the document's own `transcript.note(disclose(...))` line, and the synchronous leg gets the same full matrix the streaming one does |
369
369
  | `scripts/run-package-hygiene-test.mjs` | Everything `package.json` `files` ships — dist JS/typings and the Markdown docs — is screened line-by-line against a deny-list of strings that must never reach a public tarball (internal hostnames, codenames, person names, collaboration-process words, other repos' ledger ids and repo names; opaque ticket ids `CC-nnn` and post numbers `[nnnn]` are allowed as traceability references). Since 0.77.2 the build strips comments (`removeComments`; enforced by `run-dist-comments-test.mjs`), so what this gate screens in dist is code, string literals and type-level text. Markdown docs are enforced forward-only (CHANGELOG from 0.77.2, the integration doc from §81) because published sections are frozen.
370
- | `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T \| null`, the `installHost` family returns `T \| undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. |
370
+ | `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T \| null`, the `installHost` family returns `T \| undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the live sections of the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. Frozen sections (a numbered section whose heading carries a version already recorded as released) are historical and never rewritten, so a retired name there is allowed only if the live retirement catalog has a row for it: a retirement label, the name, the version it left in (matching the ledger) and what to use instead, next to the sentence stating that frozen sections are historical records and the catalog is authoritative. A section whose version cannot be read, or is not yet released, is judged as live. |
371
371
  | `scripts/run-type-superset-ledger-test.mjs` | The type/wire **superset ledger** (`docs/type-superset.json`): positions this package adds on top of a CC-shaped contract, each carrying the evidence for what CC's own type surface does or does not have there. Completeness is deliberately uneven and the ledger says so. The `_sema_*` private-key class is checked in **both** directions (a key in the source that never entered the ledger is red, naming key and file; a ledger row whose key left the source is red) — but only for keys written as literals, which is the convention the ledger mandates. A key assembled by string arithmetic is beyond what any static rule can enumerate, so the guard fails closed on every shape it *can* decide (a bare `_sema_` prefix is red wherever it appears, save one pinned guard site) and leaves the rest as a convention violation for review to catch, rather than claiming a completeness it does not have. The two hand-surveyed classes are only checked for coordinate and evidence integrity, never discovered. Both directions read the source through the **TypeScript AST**, not a text scan, and they read two different sets out of it. A *key site* is an identifier, or a string whose whole value is the key — so `'_sema_decision-v2'` is carried whole rather than truncated at the first non-identifier character into some *other* key that happens to be registered. A *mention* is the key appearing inside a longer string, which is prose, not usage. The staleness direction counts key sites only: a comment or a doc sentence left behind after the last real mint site is deleted must not keep the row alive (mutation-proven — with both the comment and the prose string untouched, removing the one real site turns the guard red). And because a prefix can be concatenated or interpolated into a key no static set will ever see, the bare `_sema_` literal is refused outright rather than traced: every occurrence is red except the single inline `startsWith` guard the sanitizer needs, because the set of expressions a bare prefix can travel through on its way to a concatenation is open-ended and enumerating it is always one form behind. Every row's `host` must still resolve, with the key being a real **member of that declaration** rather than a string occurring somewhere in the same file — `governanceForced`/`delegation` each live on two different shapes in one file, and a member commented out is a member deleted, which a text-shaped check happily reads as still present. And the direction worth the most: each machine-form `ccAbsenceEvidence` is re-derived from the row's own `key` — the ledger's recorded string must match that derivation verbatim, since a row quietly witnessing `\bnever_present\b` is green forever while watching nothing (mutation-proven: the same edit passes the unbound form and is caught by the bound one) — and the check runs against the names the installed `@sema-agent/agent-types` `.d.ts` set actually declares, parsed with the TypeScript AST rather than grepped, so a name CC merely mentions in a comment cannot force the row into the manual escape hatch and thereby retire the very witness that was supposed to fire the day CC declares that name for real. That escape hatch is gated by an allowlist living **in the guard**, not the ledger, so claiming it costs a reviewed diff. Missing material never reads as a pass, and the verdict splits by *why* it is missing: no TypeScript parser skips the suite before it starts; a missing `agent-types` still runs and prints the first three directions, then exits **1** when `package.json` declares the mirror but it is not installed — a broken install must not retire the repository's only "the day CC declares this name" alarm, and reporting it as a skip would leave "never evaluated" and "evaluated, no drift" indistinguishable to the runner — and exits 3 only when nothing declares the mirror at all, which is the one case where the direction genuinely does not apply. Either way a run that evaluated no witness is never counted as one that did. When the mirror *is* present its **installed version** is witnessed too (the two declared floors must agree with each other and the installed copy must meet them), since four preflight probes are satisfied by an arbitrarily stale mirror — they prove the extractor speaks, not that it is current. Every direction carries a positive control — known-present CC symbols, a comment-only sample proving the extractor distinguishes declaration from mention, and synthetic corpora fed through the **same** discriminator function the real verdict uses, so a verdict quietly rewritten to return nothing takes its own control down with it |
372
372
  | `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" `listAllPersistedRules` needs only `list` (`RulesListFacade`; the other four methods are known to the parameter type as optional members): a synthetic consumer compiled against the built declarations passes a two-method object, a list-only literal, a `Pick` slice, the named type, the full facade, and the full or two-method facade written as an inline object literal — including arrow functions with untyped parameters — all with zero diagnostics, while a misspelled method name in such a literal is still reported. |
373
373
  | `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
@@ -379,7 +379,7 @@ guard still cross-checks the table by name).
379
379
  | `scripts/run-wiring-manifest-projection-test.mjs` | The two end-user facts carried on the engine's `wiring_manifest` frame (`modelGate`: which tools this run's model gate removed and the verbatim restore hint; `autoMode`: whether auto mode is actually armed and the engine's own reason word). Projection: both sections ride as `_sema_`-prefixed superset keys, verbatim, and no SDK-named key is minted; a frame where neither section is well-formed projects to `none/not_in_slice` (no empty arm); `modelGate` needs all three keys and treats `removed: []` as a bad value rather than a reading; `autoMode` needs a boolean plus a non-empty reason that agrees with it, and the reason word is never mapped onto the capabilities vocabulary; the frame is flat (a nested `manifest:{}` wrapper is not a supply); `eventId` rides like every other arm. Adapter: exactly one chrome event on the main lane, a sub-flow frame (any `parentToolCallId`, `null` included) yields nothing, and an absent `eventId` leaves the key absent. Added at receiving time because the shell-side gate could not see this package's behaviour: two mutations (empty `removed` accepted, sub-flow gate removed) had passed the package suite untouched 0.71.0 adds sections F–I: the fourth/fifth/sixth manifest sections (`tools` via the roster reader, `hooks[]` rows dropped one by one when malformed, `lsp` absent unless `mounted` is a boolean), the `tool_roster_delta` arm (narrowed `delta`, `malformed` when `fromDigest`/`roster` cannot be read, host applies it against its own digest), the `context_usage` arm (finite-gated scalars plus `sections[]` rows dropped one by one), and the `WiringManifestMcpEntryView` rename with `MAX_AGENT_SKILLS` gone from the surface From 0.84.0 the `hands` section (whether this leg was assembled with the engine's built-in file and shell tools) is projected as well: a strict boolean `mounted` plus the engine's reason word passed through as written, absence kept as "not reported" rather than folded to either answer; a three-state reader and one sentence source share their opening and closing with the roster-derived sentence, whose bytes do not change. The sentence source answers the not-reported sentence for anything it does not recognise, including a roster-derived reading, and never throws. |
380
380
  | `scripts/run-submit-wiring-manifest-test.mjs` | The non-streaming submit receipt can carry the run's opening wiring manifest (`TaskResult.wiringManifest`, additive on newer servers). `readSubmitWiringManifest` answers one of three: the key is absent on the receipt itself (older server, or a deployment whose engine never produced that frame) — not the same as unreadable; the key is present but cannot be read (not an object, or none of the nine sections survive); or a manifest view. The view is the same shape the streaming lane's chrome event carries (minus its two envelope keys) and is assembled by the same code path, so both lanes agree byte for byte on the same object. Liveness fields ride through untouched — this reader never mints a liveness verdict — and an operator-shaped receipt with extra governance sections reads to the same view as a tenant-shaped one. A zero-tool roster is a real reading, not an absence. |
381
381
  | `scripts/run-rule-offers-reader-test.mjs` | The narrowing reader behind the "don't ask again" options, now a public entry point rather than a card-port-only one. Hosts that render the frame themselves (a browser has no three-way terminal card) previously had to rebuild this reader on their side, and what it carries is a **redemption-safety** judgement, not a convenience: the batch arm is redeemed by **index**, so a reader that compacts the array after dropping a malformed entry makes the k-th option a person clicked and the k-th rule the server writes two different rules. So: a bad entry is dropped **on its own** (one bad option must not make a real one disappear) while every surviving entry keeps its **original wire index** — pinned from both ends, with the bad entries leading and trailing. A batch's *members* are the opposite: any malformed member drops the whole batch, because a conjunctive batch is one "yes" to all of them and a batch missing a member is a different grant; its honest-remainder count is a reading, not decoration, so a non-integer or negative value drops the batch rather than rendering a fabricated zero. An empty array, a non-array, an over-cap array and an all-bad array all read as **absence** rather than an empty list, because an empty list renders as "there is an option lane with nothing in it". The two wire generations are ordered by a rule, not a preference: the newer key wins outright, a newer key that is **present but unreadable** does not fall back to the retired key (borrowing the older material would pass someone else's options off as this request's), and a `null` newer key reads as absence so a relaying layer that serialises "missing" as null cannot delete the whole lane on older engines. The public entry is finally reconciled against **both** card-port legs on the same material, byte for byte, so the exported reader and the one the card sees can never become two. Two upstream vocabularies used to be **hand-copied** here, and both had fallen behind: a match word outside the copied pair dropped an otherwise valid option outright, and a batch carrying a directory-read member — a member kind the copy did not know — dropped the whole batch. Both tables now come from one place upstream and are re-exported verbatim, pinned in both directions: every word in the table must be accepted (a narrower copy reds on the words it never learned) and a word constructed to be outside it must still be refused (a reader widened to "any string" reds too), with the retired-key normalising leg sharing the same narrowing so the fix cannot land on one leg only. A member whose kind is genuinely unknown still drops **the whole batch and only that batch** — never one member, because a conjunctive batch one member short renders "yes to N" as "yes to N−1", and never the card, because the honest single beside it is intact — while a member from before the discriminant existed normalises to the historical kind rather than being refused. The additive per-segment reasons ride through verbatim, drop only the row that is malformed, and stay **absent rather than empty** when nothing survives, since an empty list would read as "confirmed nothing uncovered" while the count remains the only source of truth |
382
- | `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one. The engine's folded resume refusal (`resume_blocked_by_policy`) gets its own reading — the original code as named, absent or unreadable — and a wording that neither claims nothing was consumed nor predicts whether a retry would pass. |
382
+ | `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one. The engine's folded resume refusal (`resume_blocked_by_policy`) gets its own reading — the original code as named, absent or unreadable — and a wording that neither claims nothing was consumed nor predicts whether a retry would pass. From engine `7.104` a folded admission refusal always names a code: when the underlying refusal carried none, the server fills in the generic code for its HTTP status. Those generic codes are not a specific check, so the sentence says the refusal did not name a more specific check (and shows the generic code in brackets) instead of presenting it as the check that refused the resume; the generic-code list is compared word for word, in both directions, with the server dispatcher that mints them. |
383
383
  | `scripts/run-resume-retry-later-test.mjs` | The two resume refusals that carry a **wait quantity** — the only members of that refusal family that do, which is the whole reason they form a closed set. Carrying a wait is not the same as being the only ones worth waiting on: a sibling refusal in the same family clears on its own and the engine says so in words, it just cannot put a number on it, so *not recognised here* must never be read as *waiting will not help*. One of the two also has a *terminal* upstream branch that arrives under the same code with the distinguishing detail only in prose, so recognition alone is not permission to say "try again": the disposition is decided by **positive evidence** and pinned from both directions — the quota-window code is evidence in itself, the preflight code counts only when the server really supplied a wait (an upstream fact, not a convention: the terminal branch throws with no detail at all, so a wait value cannot reach the client on that path), and a preflight refusal with no wait reads as *undecidable* (say what is true of both branches — nothing was consumed — and leave redeemability to the engine's own line) rather than being rendered as either a retry or an ending. Every other member means waiting will not help (change a setting, relaunch, the retained session is gone), so the recognition is a **closed set of two codes**: widening it to a family prefix would tell half the users to wait and the other half to keep waiting for something that will never arrive, and the negative controls drive exactly those codes through it, plus a same-named code on a different door (the submission-side quota refusal), the two underscore-form siblings, and a code merely quoted inside a message body. The wait value is narrowed to the same domain the server mints it in (a whole number of seconds, at least one): zero, a negative, a fraction and a non-number all read as **no window given** rather than as zero, because a zero tells the caller to retry immediately and the wake-up it would retry is an at-most-once action with real side effects. Reading is structural rather than `instanceof`, since the client is host-injected and the same class name across two bundles is two classes, and a null-prototype plain object must still be recognised. The failure classifier gains this one disposition without any existing one moving, an unknown code still falls to the honest open-set arm and its wait value is **not** believed, and an end-to-end call proves the disposition and the window reach the host while the call itself is still attempted exactly once. The recognised code set is a **frozen array**, not a type-level readonly set: the latter is a plain mutable collection at runtime and the decision reads the same instance, so one `.add` from any consumer would turn a refusal that waiting cannot fix into one that claims it can — the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift |
384
384
  | `scripts/run-model-capability-probe-test.mjs` | Whether a model on the OpenAI-completions lane **thinks**, and whether that thinking can be **turned off** — a question nobody can answer by looking at a model name, and one whose wrong answer costs every later call. The probe is judgement only: the network half arrives as an injected port, so the package mints no URL, reads no credential and never calls `fetch` — pinned by a source-level assertion, because a package that reaches the network once has changed what every host must trust it with. The seven dialect words are a **copy**, reconciled element-wise against the installed engine’s own bytes in both directions, since the words belong upstream and a private table drifts the day a dialect is added; the settings package deliberately declines to restate them, so the table cannot be imported from there and this guard is what stands in for the import. The **order** the dialects are tried in is a public promise rather than an implementation detail — each extra attempt is real money and real latency against someone’s gateway — so the guard pins the exact call sequence a stub records, and reversing it reds on the wasted round trip; the template-parameter spelling leads because an observed gateway keeps thinking, and answers with an empty body, when handed the top-level switch instead. That observation is also why an empty answer is **not** accepted as *thinking is off*: a knob that deletes the reply is not a knob that disabled reasoning, and accepting it would write a spelling into the catalogue that the gateway does not honour. Two dialect words whose request bytes are identical to another’s do not each burn an attempt. The two verdicts that look alike are held apart from both directions: *tried everything, still thinking* requires at least one attempt to have **cleanly answered**, and when every attempt was refused the verdict is *could not tell* instead — and on the unanswerable path the result carries **no** thinking flag at all rather than a fabricated `false`, while the pure write-back returns the very same entry object untouched. A verdict that reasoning cannot be disabled **removes** a previously declared spelling rather than leaving it, since a refuted spelling keeps the engine sending bytes the gateway ignores while the catalogue still renders it as already off. Evidence is lengths, finish positions and status codes only — a planted secret in both the answer and the reasoning channel must appear nowhere in the result, so the record can go into a log or a ticket whole. One cross-package premise is checked by really running the other package’s parser rather than quoting its documentation: everything this probe writes eventually passes through the settings schema on its way into a catalogue, and that field is declared parse-transparent precisely so the vocabulary can live on the consuming side. If it ever narrows, the spelling is stripped **silently** — indistinguishable from the probe never having run — so the guard feeds the probe’s real output through the real parser, checks the compat object comes back key for key, and checks a dialect word this client has never heard of survives too. Two shapes that must be rejected really are rejected, since otherwise the survival checks would hold on a parser that accepts anything, and a bare entry is asserted valid first, because the first run of this section reddened on a space in a fixture’s name — a fixture that cannot pass would disguise the real alarm as already having fired |
385
385
  | `scripts/run-decide-receipt-test.mjs` | What a decision verb actually **answered** — and, more importantly, what it did not. A success response on the newest lane is only an acknowledgement that the decision was accepted for delivery: the approval is still pending, and a client that clears the card on it shows either a ghost card that was already approved or a card that vanished while the decision was lost. So the package deliberately has **no** "was it resolved" predicate — nothing in that body can answer it — only the opposite one, whose `false` is likewise not evidence of resolution; resolution is only ever the next running arm on the stream. The guard pins that inversion in the product source too: the success path must no longer clear the latched gate, while the stream-observing path that really clears it must still be there. The body has four shapes with **no** key common to all of them, so every position is read as honestly absent, and the handoff handle — which run to watch from here on — requires **two** facts together, since either one alone would either point the stream at the run it already had or mint an empty handle. The record of what finally happened to an already-decided action is read through the **same** reader as every other gate record rather than a second copy, and its absence means **unknown**, never *it was allowed* — the two can even contradict each other, so the card says nothing at all when it is missing. The three refusals on that lane each get one distinct sentence and a disposition taken from **why** each was refused rather than from severity: one cannot be helped by re-sending at all, one waits on the host, one just drops an option — and none of them carries a countdown, because the server never mints a wait for them. Recognition is a **closed set**: an unrecognised code on the same prefix returns nothing rather than a guess, since that prefix also houses a safety signal whose whole rule is never to retry automatically, and the recovery handle is read as absent when unreadable rather than substituted from a different identifier that no longer appears on that lane **0.68.3 (core 7.18.0):** `gate.disposition.classifier` is read key by key into a named view (requested model, model that answered, and the ladder fallback when one happened); a half-shaped record yields no classifier at all rather than a half view, and absence stays "unknown", never "the seat answered itself" |
@@ -392,7 +392,7 @@ guard still cross-checks the table by name).
392
392
  | `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
393
393
  | `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
394
394
  | `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
395
- | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false. From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. A later section reconciles the terminal list against the denied calls seen in the same stream, so denials that never reached a human (rules, hooks, classifiers, write protection) are listed too: complete rows join the CC list, rows missing a field stay on the extended list and mark it incomplete. A further section feeds the same raw tool-result frame through the projector into both lanes and requires the denial category on the interactive transcript record, on the non-interactive frame and from the reader on the raw frame to agree, including frames whose settlement, classifier cause or classifier attribution is present but unreadable — a shape the narrowed gate view on the internal frame cannot show — and pins that the reader gives the same answer on the internal frame as on the raw one. |
395
+ | `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and the scan over them finishes within budget. Since 0.84.1 arguments the transport redacted on the way (a string leaf carrying a replacement token, or a cycle / depth placeholder) join as well: the reference list receives the very object the transcript's tool call shows, and that entry, on both lists, carries a superset flag saying the input is a redacted view rather than the original — redacted arguments that are over budget, not a plain object, or never seen on this stream still stay off the reference list, and the flag is never minted as false; both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false. From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. A later section reconciles the terminal list against the denied calls seen in the same stream, so denials that never reached a human (rules, hooks, classifiers, write protection) are listed too: complete rows join the CC list, rows missing a field stay on the extended list and mark it incomplete. A further section feeds the same raw tool-result frame through the projector into both lanes and requires the denial category on the interactive transcript record, on the non-interactive frame and from the reader on the raw frame to agree, including frames whose settlement, classifier cause or classifier attribution is present but unreadable — a shape the narrowed gate view on the internal frame cannot show — and pins that the reader gives the same answer on the internal frame as on the raw one. |
396
396
  | `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
397
397
  | `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
398
398
  | `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
@@ -406,7 +406,7 @@ guard still cross-checks the table by name).
406
406
  | `scripts/run-memory-verbs-wire-test.mjs` | The five memory-governance verbs as call ports — entry provenance, compliance erasure, the external-origin listing, the clearance ledger and the un-mark valve — on top of the readers above. One failure judge serves all five, and its first question is **provenance, not status**: the engine stamps a machine code on every refusal it mints, so a 501, 405, 409, 404 or 400 that carries **no code** proves nothing about who answered — a proxy or gateway returning the same status may well have passed the request on first — and every such answer is reported as "no verdict" rather than as "nothing happened". Twenty-one coded refusals each get their own arm, branched on the code alone: the status cannot tell them apart (nine different operator actions ride the same 409 here), and conjoining the status would silently demote a refusal the day the engine moved it. A coded 5xx, a coded answer with no status at all, and a coded 4xx this version does not recognise all land in the "cannot tell" arm, because on a non-idempotent verb the default for "could not classify" must be "do not know", never "did not happen". The judge is called from catch blocks, so each of its own property reads is guarded: an error object whose accessors throw is classified, not re-thrown. The capability gate runs before the call and reads the two bits the engine keeps deliberately separate (one for provenance and erasure, one for the origin faces); only an engine that positively says the face is off stops the request, while "this binary does not report that bit" and "this process never saw a capabilities body" both still send — folding "cannot say" into "is not there" is the dishonest-absence shape this package refuses, and these routes answer the capability gate before touching anything. A missing capability reading is a named, explainable error rather than a silent default that would answer for whichever engine happens to be installed. Each verb returns its own discriminated union whose success arm, refusal arm and cannot-tell arm share no keys, so a consumer cannot express "could not read it" as "it worked". The two write legs reuse the erasure and clearance receipt readers rather than minting a second copy, which keeps "this call erased nothing", "a 200 with an empty body" and "a body that could not be read" three separate things here too; the provenance account is narrowed only to its envelope and discriminants and otherwise passes through verbatim, and an account stamped with a newer envelope version is refused rather than reinterpreted. All three receipt-bearing verbs additionally reconcile identity — the account id, the attestation request id and the clearance receipt entry id must be the ones that were sent — because a readable receipt is not yet a receipt about this call. The origin listing carries the server's own echo of the scopes it actually audited, the type has no store-wide arm, and the coverage port takes a mandatory second argument and has no "clean" arm at all: the strongest thing it will say is which scopes were audited. Neither write leg is ever retried, including the refusal whose documented recovery is to send again, because that resend completes whichever clearance row is already open and the audit attribution on it is a person's signature. Request bodies are handed over verbatim — the degraded-erasure authorization is never injected — while the scope list is sent as the snapshot this port validated, so an array that reports one length while being read and another afterwards cannot make "the list I checked" and "the list I sent" two different things. |
407
407
  | `scripts/run-dist-orphan-test.mjs` | Every `.js` / `.d.ts` under `dist/` must have a same-named source under `src/`, and every source must have its build output — because the compiler only writes and never deletes, so a module removed from the sources keeps shipping from the previous build (the whole `dist/` directory is on the publish whitelist) while the public-surface gate only looks at what the barrel exports and the hygiene gate only looks at forbidden words. Orphans are named one by one; the pre-publish posture is a clean rebuild, and this gate is the check that the posture was actually followed. |
408
408
  | `scripts/run-dist-comments-test.mjs` | **dist ships zero comments.** Since 0.77.2 the build strips comments (`removeComments`); this gate walks every shipped `dist/**/*.js` / `*.d.ts` and counts comment trivia with the TypeScript scanner (string literals containing `//` and generator methods are not comments), failing on the first one (`DIST-COMMENT-FAIL`). Source comments are an internal surface; what still ships is code, string literals and type-level text, which the hygiene gate screens. Negative control: one plain comment appended to `dist/index.js` turns it red. |
409
- | `scripts/run-task-request-omission-receipt-test.mjs` | Where every key a client hands to the request constructor ends up. The constructor used to answer "not stamped" the same way for four different reasons — value absent, no such row, wrong lane, live gate closed — and a key it had never heard of did not even get that: an unattended run could pass a system prompt, an output schema and a spend cap and receive a body holding the objective and the session id, with nothing anywhere saying what was left out or why. The guard pins the three answers apart. **Seated** keys reach the body verbatim on the unattended lane. Keys the package **knows but did not carry** never throw, never reach the body, and each gets a receipt row with one word from a frozen cause list — every present key is on the body or on the receipt, never both and never neither, checked across both lanes with the live gate open and closed against a key-by-key table written independently of the package's own routing. Keys the package **does not know** are refused loudly and are a separate cell, not a fourth cause: the cause list has no word that could hold them, and the seat reader answers `unknown`, not `none`. The cause list is bitten from both sides (exact, every word producible, nothing produced outside it, the judge table's keys read from source through the TypeScript parser) and no second hand-copied list may exist in `src/`. Upstream claims are read straight off the installed SDK typings: a key seated in this release must be a named request field, a key registered as having no upstream counterpart must not be — the day it appears the guard turns red — and the index signature counts as evidence for nothing. |
409
+ | `scripts/run-task-request-omission-receipt-test.mjs` | Where every key a client hands to the request constructor ends up. The constructor used to answer "not stamped" the same way for four different reasons — value absent, no such row, wrong lane, live gate closed — and a key it had never heard of did not even get that: an unattended run could pass a system prompt, an output schema and a spend cap and receive a body holding the objective and the session id, with nothing anywhere saying what was left out or why. The guard pins the three answers apart. **Seated** keys reach the body verbatim on the unattended lane. Keys the package **knows but did not carry** never throw, never reach the body, and each gets a receipt row with one word from a frozen cause list — every present key is on the body or on the receipt, never both and never neither, checked across both lanes with the live gate open and closed against a key-by-key table written independently of the package's own routing. Keys the package **does not know** are refused loudly and are a separate cell, not a fourth cause: the cause list has no word that could hold them, and the seat reader answers `unknown`, not `none`. The cause list is bitten from both sides (exact, every word producible, nothing produced outside it, the judge table's keys read from source through the TypeScript parser) and no second hand-copied list may exist in `src/`. Upstream claims are read straight off the installed SDK typings: a key seated in this release must be a named request field, a key registered as having no upstream counterpart must not be — the day it appears the guard turns red — and the index signature counts as evidence for nothing. The per-run file-history opt-out word is covered the same way: seated on the two user lanes behind the live gate and absent from the side-channel lane, stamped only for its one legal word, refused loudly for any other value, and read after every existing declaration so that it cannot erase them. |
410
410
  | `scripts/run-session-policy-wire-test.mjs` | The per-session tool-rule face: the capability bit that says whether an engine keeps such rules at all, and the narrow read plus tightening orchestration built on it. The bit is read the same four-state way as its sibling capability readers — an absent key is not reported (this binary predates the position itself, which says nothing about whether the face exists), `true` is present, `false` is a positive absent (this deployment keeps no per-session rules), any non-boolean value is unreadable and drops the cell rather than being folded into "absent", and a capabilities body that is not an object at all (an array included) is unreadable rather than "not reported". Its single-source verdict answers whether to show the tightening entry: only an engine that says yes is `yes`, both a positive no and a binary too old to answer are `no`, and never having observed a capabilities body is `unknown`. Whether to put a request on the wire is deliberately a **different** question with a different answer for that last state, and lives with the orchestration. The read narrows three ways that must not collapse into each other: a record that really is empty (present, version zero — what an engine answers for a session no rules were ever written for), a record that cannot be read, and a call that failed with a typed disposition. A half-bad record — one rule bucket well-formed and another the wrong shape — counts as unreadable in full, because the write verb replaces the whole record: dropping the bad bucket and writing the rest back would empty it, which relaxes the rules while the caller sees a 200. An unreadable version stamp is never filled in with a zero, a bucket that reports an implausible number of entries is unreadable rather than walked or truncated, each array's length and each of its indices are read exactly once, and a throwing accessor is unreadable rather than propagated. Every load-bearing key is read as an own property — the envelope, the version stamp, each of the five buckets and each array index — because a prototype lookup would let a polluted prototype put a bucket into the reading that the wire never carried, and since the write replaces the whole record the union would then write that invented restriction back as a real one; a guard pollutes the object and array prototypes in place and proves all four shapes stay out. Because the write replaces the whole record, adding a restriction means writing "what is already there, plus the new entries": the union only ever adds, de-duplicates verbatim, keeps a bucket that is present but empty (present-and-empty and absent are opposite meanings, and dropping it would relax the rules), mints no bucket neither side had, copies entry bytes as they came (no trimming, sorting or path rewriting — those judgements belong to the engine), and takes its bucket names from the engine's own type surface rather than a hand-copied list, so a new bucket upstream is a compile error instead of a silently dropped one. Which differences count as relaxing is the engine's judgement and is never re-implemented here: a refusal on those grounds is reported verbatim, never swallowed and never retried. The orchestration is guarded on three axes. Timing: when the record moves between the read and the write, it re-reads and re-writes **exactly once** — two reads and two writes, no more — and the second attempt's union carries the other writer's entries, which is the entire point of re-reading; a second collision is reported rather than retried a third time, and an uncontended write makes exactly one round trip. Concurrency: an explicit barrier holds both orchestrations first reads at the same version before either may write, and the criterion is how many times the store actually rejected a stale version rather than how many writes it saw — the latter is equally true of two serial successes, so it would stop detecting contention the day the interleaving changed. Under real contention the store rejects exactly once, both writers land on strictly different versions, both writers entries survive in the final record, and the round trips are exactly three reads and three writes; the same two orchestrations run serially are asserted to reject zero times in two reads and two writes, which is what proves those numbers are discriminating. On a store where every write loses the race both report a collision having written exactly twice each. Failure classification: a relaxation refusal, a refusal to stamp a version the store cannot establish (the same status code as the relaxation refusal but a different machine code, and folding it into that arm would send the caller off to edit entries that are not the problem), a missing session, a deployment without the face, an ownerless session, a collision code, a bare conflict with no machine code, a rejected body and an unauthorized call each land on their own arm — the collision arm is matched on the machine code verbatim rather than on the status, because two different situations share that status and only one of them is worth retrying. The remaining split is not "which code is this" but "did the engine answer at all": an answered client-side refusal is allowed to say nothing was written, because every such refusal on this endpoint is emitted before the record is touched, while a throw with no answer at all — a dropped connection, a timeout, a response body the transport itself could not decode, a server fault — can only say "unknown", since that throw may well have happened after the record was already saved. Two guards prove that is not theoretical: a write whose receipt cannot be read, and a write that throws after the fixture store has committed, both leave the record changed. Neither is success nor failure: the only honest answer is "unknown", it carries no version, and it is never retried. The direction of the change is likewise never claimed. The engine’s tighten-only rule is an identity gate, not a field gate — for a principal the deployment treats as an operator it does not run at all, so a union that adds a name to an existing allowlist is accepted and really does widen it. This package does not mint a second copy of that rule, so what it reports is the fact it can stand behind: the record now holds what it already had plus the entries sent here. The sentence for a saved write is pinned to contain no claim of tightening, narrowing or restriction, and a guard reproduces the operator case to prove the widening is real while the wording stays honest. Every sentence the module mints is checked pairwise distinct, with the receipt-unreadable one required to keep its "may already be in effect" and the record-unreadable one required to say nothing was written. |
411
411
  | `scripts/run-persisted-rule-write-test.mjs` | The **single-step tightening write** for persisted permission rules — the dual of the revoke surface, and the half where a hopeful reading is expensive. The two behaviours this entry accepts are **derived** from the three-state vocabulary by subtracting the widening one, never hand-copied: the guard bites in both directions (every word in the derived table is really accepted; every constructed outsider — casing variants, trailing whitespace, the widening word itself — is refused before a single round trip), keeps a word-count canary against the parent table, and pins that the source file contains **exactly one** array literal carrying two or more behaviour words, so a second hand-written table shows up as a boundary failure rather than as drift nobody reads. A standing approval is minted by answering a permission question or by importing settings; this entry is not a third route, and the widening word is unspellable in the type. The outcome is a discriminated union whose two failure arms are **not** interchangeable: ten refusal causes each promise the same single thing — not one byte reached the store — while three separate words say the opposite, that the outcome could not be read at all. The service's own "I cannot tell" (a write that could not be confirmed as standing: store wobble, a redemption leg with no decidable ending, or a write that landed and was revoked concurrently before the read-back) stays in the second group, because announcing "nothing was written" invites a clean retry that is not clean, and announcing success misreports a tightening that may already be gone. Anything the shared failure classifier does not recognise defaults to the same place — this is a non-idempotent verb, so "unclassified" must mean "unknown", never "no write": a 500 can happen after the store commits. A 2xx whose body cannot be read is pinned in the same direction and from both sides: it reads as unknown, and the unknown arm carries **neither** the revision nor the written row, so a consumer cannot even spell the shape that would let "unreadable" pass for "written". `persisted` is guarded against the reading everyone reaches for first: it says *this call wrote*, not *a new rule now exists* — an equivalent rule already in the store still mints a fresh causal point, so the lane honestly reports `persisted`, and the material for judging whether the **logical** rule is new (the approval ledger on the returned row) is handed to the caller rather than folded into the discriminant, since the package does not have the one fact that judgement needs. The returned row goes through the **same single narrower** the listing surface uses — proven by running one row corpus through both legs and asserting the two verdicts agree entry for entry (an adversarial pass that forks the listing leg back into an inline copy reds here immediately), plus a source pin that the predicate is defined once and called from exactly the two legs. That sharing is what keeps a row whose behaviour cell is unreadable **visible in both places** rather than hidden by one of them — and the shared narrower is deliberately followed by a second, *different* question only the write leg can ask: is the row that came back **the rule that was just sent**? A rule's identity is a triple, so a receipt missing its behaviour cell, carrying the sibling state, carrying the widening one, or naming another text or another scope is not evidence that the requested tightening is standing; it reads as unknown with its own word, kept distinct from "unreadable" so the two stay tellable apart, while display and derived cells may vary freely. An adversarial pass found both of the gaps this pins: the receipt check that only looked at whether the row was renderable, and a subtler one — pulling the verb off the injected port and calling it bare drops the receiver, so a host that hands over a real resource object (a class instance whose verbs reach the transport through `this`) would see every write throw and be reported as "could not tell", retry after retry, while the package's three other ports call their verbs as methods and work fine. Both are pinned from the failing side: a shorthand-method facade and a class-instance facade must reach the transport and return a real outcome, with the bare-call throw proven to be a real failure mode first. A 405 is split in two, because only the engine's own bare code is evidence about **the engine**: with it, the path exists and this verb does not, so this worker predates the verb; without it — an absent, empty, or foreign code, which is what a proxy or gateway blocking the method typically returns as HTML or an empty body — what was seen is that the verb was refused, while **who** refused it and **at which hop** is unknown, so it lands in the unreadable-outcome arm with its own word rather than sending someone to upgrade a worker that is fine, hiding an entry that is live, or — the part a second adversarial pass insisted on — promising that nothing was written. That promise is what the refusal group means, and a middlebox is free to forward the request and only then answer 405 on its own policy, so a caller who skipped reconciliation on that word would leave behind a standing refusal the user believes never took effect; the guard pins exactly that shape, with a double that writes the rule and *then* answers 405, and with the engine's own bare code still landing in the refusal group beside it. (The same passes caught the naive status-only reading and the asymmetry where an empty-string code fell through to a different bucket.) The documented recovery for an unreadable outcome — retry, then reconcile — is pinned to be **ledger-safe** rather than merely asserted: against a double that models the engine's own "is this identity already standing?" question, re-sending the same identity comes back as a no-op with the approval ledger and the bucket revision both unmoved, however many times it is repeated, while two concurrent writers each landing a causal point are both honestly reported as having written. Finally the four local gates are pinned to be free: a missing write verb on the injected port, an unwritable direction, an unreadable identity pair, and a principal key that is present but cannot name anyone all refuse **without sending anything** — the last of those because silently degrading a blank target into absence would land a tightening aimed at someone else in the caller's own bucket and return a 200 |
412
412
  | `scripts/run-registrar-tables-test.mjs` | The four registrar table bodies — this Guards table and the three census tables in the repository's negative-control record — are **generated** from `scripts/gates-manifest.json`, the one file that describes a suite. Every row's text must equal what the generator emits, the manifest's suite set must equal the suites on disk, each entry must declare how it is negative-controlled (rehearsed, blind, or behavioural, with the census taker itself declared as such since it does not appear in its own tables), and each row's outward prose is scanned against the published-surface word list — the same list the packaging-hygiene guard uses, shared rather than copied — before the generator may write it into this file. Adding a guard is therefore one manifest entry plus one generator run instead of six hand edits across three files, and a description that drifts in one place and not the others stops being expressible. The row count is no longer what is compared: the earlier arrangement checked the census tables by **length**, so rows naming the wrong suites reconciled green. Positive controls run entirely on in-memory copies — a changed description, a dropped entry, an added entry, a changed class and a hand-edited row on disk each have to make the same judgement speak — and the quieter halves are pinned too: nothing outside a table body may move, a line inside one that is not a recognisable row makes the generator refuse rather than drop it, byte equality is backed by a column-count check (a cell holding a bare pipe splits a row into extra columns, and a code span does not protect it), and the malformed rows kept byte-for-byte as they are found are registered individually, so the registration turns red the day it stops being needed rather than outliving its reason — audited in both directions, since a registration pointing at a row that is no longer malformed and one pointing at a guard that was reclassified or deleted are both exemptions nobody reads |
@@ -418,21 +418,35 @@ guard still cross-checks the table by name).
418
418
  | `scripts/run-cc-message-key-census-test.mjs` | The one rule behind CC-shaped messages this package emits: every top-level key on a `user` / `assistant` / `result` / `system` message must be a member the CC SDK mirror (`@sema-agent/agent-types`) declares for that same arm, or one of this package's `_sema_`-prefixed additive keys, or a row in a dated transition table that goes red the day its retire version arrives. Mint sites are found syntactically (identifier, quoted and computed `type` names alike) and their key sets are resolved syntactically too — inline literals, both branches of a conditional, the nullish-coalescing and logical or/and operators, `const` initializers and every `return` of a helper — so a type assertion, a `Partial<Pick<…>>` narrowing or a computed name cannot launder a key past the check, while a spread the tool cannot follow (a parameter, a `let`, a member access) fails the tool, never the product. Arms that carry a `subtype` discriminator are checked against that subtype's own member set, so a key declared only for another subtype does not pass on the strength of the arm-wide union. The runtime section drives the real adapter pipeline and pins the tool-result record's SDK-spelled `tool_use_result` (the transitional camelCase twin rides along by reference until 0.83.0). A second runtime section feeds the same tool-result frame to both the interactive adapter and the non-interactive frame builder and requires the structured-result key and the no-output marker to be present together, absent together and equal on the two outputs, while the transitional camelCase name stays on the transcript record only; a static section checks a per-key routing table for the internal tool-result frame against the key sets resolved at the three mint sites, in both directions, and requires the non-interactive frame's structured-result value to trace back to the very same syntax node the transcript record uses — so recomputing or copying that logic on the non-interactive path fails even when the values happen to agree. |
419
419
  | `scripts/run-rewind-archive-capability-test.mjs` | The read face for whether a rewind can restore the code archive, and the honest three-state answer the ends render from it. The shells used to decide this from a local backup table that only their own in-process tools ever fill, so on any session where the engine runs the tools it stayed empty and the two code-restoring rewind modes simply never appeared, while the engine had been keeping a file history the whole time. The judgement now comes from what the engine itself advertises, and each of the four bits it advertises answers a different question: whether the conversation can be forked at a message at all, whether a fork can restore the tracked set, whether a code-only restore is possible, and whether this deployment understands the current spelling of the request key. The mode that rewinds the conversation and restores the code together needs both of the first two, and the engine offers no single bit for that combination, so the combination is made here: one bit stated off is enough to rule the mode out, both stated on make it available, anything else stays unknown — reading the file-history bit alone would offer that mode on a deployment that keeps file history but has no conversation anchors, where the request can only fail. A bit that is absent means an older engine that never spoke about it, which is not the same as an engine that said no, and a bit reported in a shape this reader cannot read is a third thing again — it is recorded as unreadable rather than quietly filed under "not reported", because those two send an operator to different places. One unreadable bit does not discard its siblings; only a response that is not a capability object at all clears the cell. What the ends get is available, unavailable or unknown, and unknown stays unknown: folding it into unavailable would hide the mode again, which is the mirror image of the bug this replaces. The sentences the doctor row can print are checked to be pairwise distinct and to avoid implying a refusal the engine never made, and the spelling of the outgoing request is deliberately not made to follow the epoch bit, since the older spelling is rejected outright by current engines |
420
420
  | `scripts/run-registry-quota-usage-test.mjs` | The projection of the cloud control plane's quota reading, and the three different things a missing number can mean there. This response says `null` in two places and means something different each time: no token quota is configured for this principal on this instance, and this window has no cap at all. Both are **facts the server is asserting**, not gaps in the reading — while a key that is absent or carries the wrong type is a genuine gap. All three have to survive to the screen separately, because folding them is how a user ends up staring at a confident `0`: an uncapped window rendered as if nothing were left, or a deployment that simply never configured quotas rendered as if the quota were exhausted. The complaint that started this was the opposite direction — a centrally configured quota that the command line could not see at all — so the reading also refuses to let an unreadable response masquerade as "no quota configured". Two fields deliberately do not share one signal: whether a window is exhausted and whether there is a recovery time, since the recovery time is only ever populated in the exhausted case and reading its absence as "not exhausted" would answer a question the response never answered. Counts that the server always provides are narrowed no further than the mint: a used counter has no uncapped state, so a null there is unreadable rather than zero. The wording helper carries the only human-facing phrasing, and the sentences for "no cap" and "unknown" are checked to contain no digits at all |
421
- | `scripts/run-file-history-capture-capability-test.mjs` | The engine's file-history-capture self-description (`capabilities.fileHistoryCapture`), read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `off`), words are taken as an open set so a newer mode is not mistaken for a malformed answer, `fileHistoryCaptureMode` recognises only `off` and `on-always`, and the wording for `off` speaks about capture only — whether code can be rewound is left to the rewind readings. |
421
+ | `scripts/run-file-history-capture-capability-test.mjs` | The engine's file-history-capture self-description (`capabilities.fileHistoryCapture`), read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `off`), words are taken as an open set so a newer mode is not mistaken for a malformed answer, `fileHistoryCaptureMode` recognises `off`, `on-optional` (captured, and a run may opt out) and the older `on-always` (captured, cannot be turned off) that pre-`7.104` engines still send, and the wording for `off` speaks about capture only — whether code can be rewound is left to the rewind readings; a malformed reading makes neither the verdict nor the wording throw. The request side is covered too: the fragment helper yields the per-run opt-out word only when the reading is `on-optional` and the host asked not to capture — never for an older engine that would refuse the unknown key — and the fragment is driven through the real request constructor on every lane. The recognised words are checked both ways against the capture vocabulary of a `7.104` engine package. |
422
422
  | `scripts/run-model-identity-resolvability-test.mjs` | The model-identity judgement a client makes before letting anyone in: can the engine it is about to use start with a model name? Each end reports what it read from each place that can feed a model name to a local engine (complete, partial — a gateway address or a credential but no model name —, absent, or unreadable), or, for an engine that runs elsewhere, whether that engine has been seen answering; `modelIdentityResolvability` answers resolvable, unresolvable or unknown. Having part of an upstream configuration is not having enough of one, so partial lanes never add up to resolvable; a lane that was not reported or could not be read makes the answer unknown rather than unresolvable; an engine that runs elsewhere is never judged unresolvable and local lanes are never consulted for it (it does not start without a model name, so seeing it answer is enough to call it resolvable). `modelSetupDecision` combines that answer with whether this end can configure a model at all: setup is offered only for unresolvable on an end that can configure one, an end that cannot says so and points at whoever runs the engine, and unknown never opens setup. The detail and notice sentences are checked to be pairwise distinct, unknown sentences neither claim a model is configured nor that it is not, and the module is checked to import no platform I/O. The catalog lane counts as complete only when the host reports that the local engine accepts a catalog default; a report of `false` reads the catalog as partial, and no report (an explicit `undefined` included) or an unusable one leaves that lane undetermined — the same answer as before the catalog could be complete. Every sentence about a complete catalog names the catalog, including the undetermined sentence when the report is `false` and another lane cannot be read, and a 25-cell matrix pins that the report changes nothing when the catalog is not complete. |
423
423
  | `scripts/run-cloud-effective-projection-test.mjs` | The cloud control plane's effective-configuration response beyond its four configuration domains, and what a locally started engine does with the models document derived from it. Three top-level keys are read with the meaning their producer gives them: `warnings` (degradation warnings from the build that produced the served view — an empty list is a clean build, an absent key is an older server that cannot tell), and `budget` / `runtimeCaps`, where `null` means two different things: nothing resolves for this principal when you view yourself, and values withheld when the response previews another principal. A missing key or a wrong type is a third state, unknown, and none of the three is ever folded into a zero, a `false` or "no budget". A malformed warning row, budget field or cap costs only itself, and a known budget field of the wrong type is named as not shown rather than silently read as "no limit on this axis"; warning kinds are an open set, so a kind this client does not recognise still produces a warning line. The budget and cap readers are reconciled against the installed settings schema. The single wording source puts degradation warnings first and keeps every "not set / withheld / unknown" sentence free of digits, while a zero the server really sent is shown as a zero. On the models side, when the host injects an entry check the models document carries the default model, the tier groups and the active tier group (without the check none of the three is written), and every catalog reference the local engine's schema would reject — a default, role, @-mention entry, tier binding or active group that names something outside the catalog served to this principal — is dropped and recorded, because one dangling reference makes the local engine discard the whole models domain and fall back to its environment catalog; this is proven by reading the produced document with the installed file store. An @-mention allowlist that would be pruned to empty is kept as sent, since an empty list means "everything may be mentioned"; that case is recorded, produces its own warning that a locally started engine will reject the cloud model settings and use its environment catalog instead, and the gate reads the document with the installed file store to confirm exactly that outcome, so the sentence turns red the day the local reader becomes lenient. A per-model budget in which no field could be read is never described as having no limits. Registry annotation keys on model entries (`origin`, `overridesTeam`) are removed before the models document is written: the local engine's schema does not accept them, and a configuration refresh would otherwise be rejected as a whole. A model entry the host-injected entry check rejects is left out of the document and references to it are dropped: the local engine drops such an entry at startup, but a refresh rejects the whole configuration over it, so the gate requires a clean read of the produced document; when no entry passes the check, the catalog is kept as sent and gets its own warning, which the gate proves by reading the document back. The check receives a copy, so it cannot alter what is written. Tier words outside the local schema's closed set are dropped and recorded as unsupported, and the package's tier word list is reconciled against the installed schema in both directions; an active tier group is judged against group names, never model names. |
424
424
  | `scripts/run-websearch-verdict-test.mjs` | The per-request web-search configuration a host puts on the wire. Newer servers refuse the whole request when that section is malformed — and a missing or misspelled search provider now counts as malformed, because the section names where the searches go and an unknown destination is refused rather than silently swapped for the deployment's own backend. The old readers in this package dropped such a section without a word, which let the server swap destinations after all. The guard pins the new three-way verdict (absent, honoured, malformed with the field that is wrong) against the server's own judge, vector by vector, whenever that judge is available next to this package; it pins that a half-configured environment is malformed rather than ignored, that a malformed environment never falls back to the settings file (that would change the destination too), and that neither the sentence shown to the user nor the recorded reason repeats an endpoint, a key, or a search-parameter name or value — an unrecognised provider is never echoed either (the sentence lists the valid words instead), so a URL or key pasted into the wrong field does not come back out, even when it happens to be all letters. A host's key store is plugged in through a callback the package calls only after the provider has been recognised, so the precedence between environment and settings stays inside the package. Since 0.83.0 an endpoint that carries a user name or password is malformed as well (the server judges the same way from 7.101.0), the reason sentences match the server's own word for word, and the three older readers that dropped a misspelled provider are gone. |
425
425
  | `scripts/run-hooks-merged-disable-projection-test.mjs` | The fourth governance leg of the hooks projection: `disableAllHooks` set in a non-managed settings source. The value that counts is the **merged** one, read from the host through the optional `SettingsPort.mergedDisableAllHooks()`, because a per-source approximation ("any source says true") reads user `true` with local `false` backwards — the merged value there is `false` and every source's hooks ship. When the merged value is `true`, only managed-settings hooks are sent to the engine: non-managed settings can switch off their own hooks, never the managed ones, and a managed `disableAllHooks` is still judged first and sends nothing at all. A full matrix over the four sources, each true, false or absent, is merged with the reference rule (later sources override earlier ones, managed settings last) and every cell's projection is asserted. The reader is the only authority: per-source values never second-guess it, and only a strict `true` counts. A host that does not implement it keeps the previous behaviour and gets exactly one warning per installed settings port, never one per request, and none on paths where the reader would not have been consulted; a reader that throws is treated as `true`, so managed hooks still ship. The session goal's Stop hook and the final-verification yield rule, which reads the projected hooks, follow the same verdict, and the trust gate and the three managed gates are evaluated before the reader is ever called. The last leg pins the member's declared shape in the built declarations: optional, no parameters, returning a boolean or `undefined`. The flag settings source (a settings file or inline settings given at startup) is projected after the local source, and settings hooks stay concatenated managed first (managed → user → project → local → flag): the engine runs hooks in request order, and three of its budgets go to whoever comes first — the per-event wall-clock budget, the per-event cap on model-backed entries, and the total cap on added context — so a managed hook placed after the others could be crowded out and silently not run. The flag source moves with user, project and local under every managed or merged gate; its exec-form entries are dropped and warned about like any other settings source, and the not-run notice names it `Flag settings`. An optional leg feeds the request body to an installed server's hook runner and requires the managed deny, block and context to take effect while non-managed hooks exhaust each of those budgets, and pins the known cost of that order: a non-managed `PreToolUse` hook placed after the managed one can still rewrite the input after the managed check allowed it. Without an installed server the leg reports that it did not run. Safe and bare mode: the host reports its startup mode through the optional `SettingsPort.hooksStartupMode()`, which both the plain and the plan path read and which answers `'safe'` or `'bare'`; the flag already carried by the plugin hooks reading, which only the plan path reads and which cannot tell the two modes apart, counts as safe mode unless the reader named a mode. In either mode only managed-settings hooks are sent from the settings sources: bare mode is handled like safe mode, so managed hooks are still sent, also under the managed-only gates. In both modes plugin hooks are excluded as before and the session goal's Stop hook is still sent. Only the two exact words count; a host that does not implement the reader, returns `undefined`, returns any other value or throws keeps the previous request byte for byte, checked for every such reply, both flag states, both paths and with or without a session goal; any other value and a throw each leave one debug line, and so does a reader written as a property instead of a method, which narrows nothing. The reader is called at most once per request and not at all when hooks are already switched off entirely. The optional server leg requires the left-out sources' hooks not to run while the managed ones still run and still deny, in either mode. A hook written in more than one settings place is sent once: entries on the same event, in groups that are identical apart from their hooks (same matcher), with the same identity — command, shell, arguments and condition for command hooks; type, prompt and condition for prompt and agent hooks; URL for HTTP hooks; server, tool and input for MCP tool hooks — are sent once, at the position of their first occurrence in the concatenation order and with the fields of the copy from the highest-precedence source (managed, then flag, local, project, user; within one source the later copy), so a managed copy keeps both its place and its values and a user copy duplicated in flag settings keeps its place but uses the flag copy's timeout; the other entries keep their order, groups without duplicates are sent as they were, and a group left empty is not sent. Different matchers, events, a one-byte command difference, a different condition, an absent versus explicit shell, different types or extra group keys are not merged; malformed entries and entries whose identity cannot be computed are left alone; plugin groups and the session goal hook do not take part, and hooks removed by safe or bare mode never come back. The optional server leg requires such a duplicate to run once with its added context appearing once, and a user copy with a one-second timeout duplicated by a flag copy with a ten-second timeout to run to completion on a two-second command. |
426
426
  | `scripts/run-memory-saved-projection-test.mjs` | Engine memory writes (a successful `Remember` tool call) moved off the transcript onto the additive `memory_saved` chrome event, driven through the real pipeline: zero transcript rows for the write (the transcript is byte-identical to the same frames with the write reported as not successful), exactly one event whose `notes` carry the note text verbatim (notes, not file paths) and whose key set is exactly kind / laneProof / id / notes; no event for a missing, empty or non-string note, a non-`true` `ok`, a tool error, a missing result or another tool name; two writes give two events in order with distinct ids; a sub-agent write rides the sub-agent lane and an empty parent id emits nothing rather than falling back to the main lane; the `id` is derived from the write's wire key (the tool-end event id, else the tool-start event id, else the call id; empty ids count as absent), so projecting the same wire events twice gives the same id, and it never collides with the tool result row of the same or another call; the event sits right after the tool result row; the arm is registered as required. |
427
427
  | `scripts/run-result-frame-projection-test.mjs` | Result frames and the synthesized terminal rows. The CC key `terminal_reason` is minted on result frames only where it follows from what the engine reported: `completed` on success, `max_turns`, `budget_exhausted` and `structured_output_retry_exhausted` for the three matching engine codes, on both the done-frame path and the failed-event path. Every other outcome leaves the key absent as an own property rather than present with an undefined value: wall-clock and token-budget limits, the classifier denial limit, cancellation, unknown codes, blocked, paused, unreadable or missing terminal records, and the busy-session refusal. The public reader `terminalReasonForResult` shares the minting predicate and is checked to agree with the minted key on every frame the gate produces. Both the minted key and the reader derive the word from the frame's CC subtype (success with `is_error` strictly false, and the three limit subtypes), not from the error code, so a replayed row whose status is paused, blocked or unrecognised never carries a word that contradicts its subtype. The four words are checked against the mirrored CC union, and the minting file is checked to hold no hand-copied code literals. The renamed superset keys (`_sema_error_code`, `_sema_salvaged_result`, `_sema_model_degraded`, `_sema_selected_model`, and the row flag `_sema_api_error_message`) are driven through the real stream pipeline. Each must be present under its new name, the old name must be absent, and every frame the gate saw is swept for old names. The selected model appears on error envelopes whenever the terminal record carries it, and never on a failed event, which has no record. It stays separate from the provider-reported model name. The two in-package readers still work: the interactive result arm reads the salvaged text under its new name (and old-shape frames under the old one), and the print init gate treats the renamed flag as the run having ended. |
428
- | `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
429
- | `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed six-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule that names a server or agent peer without naming one of its tools (for every protocol namespace the engine knows, checked against the engine package's own table) or contains a wildcard (*) anywhere (an engine that compares exact names would block nothing), an entry that is not a tool name, and a name the engine refuses at the start of every run — a retired tool name, or one containing "__" without a protocol prefix, where the prefix check is case-sensitive (once such a name is in the record, every later run of the session fails at startup until that entry is removed; the retired-name list is checked entry by entry against the engine package's own list, and a withheld retired name carries its current name when the engine says it was renamed) — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. An end-to-end check runs the engine package itself: every batch this function would deliver — the recorded vectors and a fixed-seed sample of generated names — is written into an in-memory session rule store and the next run must get past its start-up checks and reach the model, while every name withheld as refused — every retired name included — must indeed make that run fail at start-up, and every string literal in the judgement source that it withholds as refused must be one the engine package's own tables refuse. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for four deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; peer-wide rules outside the MCP namespace, which it did not recognise; and names the engine refuses at start-up, which it sent as ordinary names), whose disagreements are counted per class and must match an independent count exactly. |
428
+ | `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Every row also names what kind of duplicate it is, and carries the evidence that kind requires. A client copy that is this package's own function object behind a type assertion, or a thin wrapper that only passes the client's own dependencies into this package's function, must be proven so on every run by reading the client's syntax tree; such rows are due for review within two minor versions, and the guard fails as soon as the proof no longer holds. A wrapper whose extra logic the client has confirmed to be host-specific carries that confirmation (the client's own wording and where it was stated) together with a pinned review record that is due within two minor versions. A wrapper that adds its own decisions, or a same-name function that does something else, carries a review record pinned to a hash of the normalized client declaration (comments and formatting do not count); once the client's copy changes, the guard fails until it is reviewed again. A second implementation of the same logic carries a semantic diff: the client's copy is taken from the commit being read, the named declarations and the client modules they import are extracted with the TypeScript parser, transpiled and run in memory against this package's build on the same inputs, and any disagreement outside the classes registered on that row fails the guard (each class is a fixed predicate over both answers). A row can follow a client-side rename to a new name, and can be registered ahead of a client change for a limited number of versions before the client code exists. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
429
+ | `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed seven-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule that names a server or agent peer without naming one of its tools (for every protocol namespace the engine knows, checked against the engine package's own table) or contains a wildcard (*) anywhere (an engine that compares exact names would block nothing), an entry that is not a tool name, and a name the engine refuses at the start of every run — a retired tool name, or one containing "__" without a protocol prefix, where the prefix check is case-sensitive (once such a name is in the record, every later run of the session fails at startup until that entry is removed; the retired-name list is checked entry by entry against the engine package's own list, and a withheld retired name carries its current name when the engine says it was renamed), and — only when the caller passes the tool roster of a run — a name that roster does not list as a tool name or alias, compared exactly with letter case counted (recorded as it is it would block nothing on that deployment) — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. An end-to-end check runs the engine package itself: every batch this function would deliver — the recorded vectors and a fixed-seed sample of generated names — is written into an in-memory session rule store and the next run must get past its start-up checks and reach the model, while every name withheld as refused — every retired name included — must indeed make that run fail at start-up, and every string literal in the judgement source that it withholds as refused must be one the engine package's own tables refuse. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for four deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; peer-wide rules outside the MCP namespace, which it did not recognise; and names the engine refuses at start-up, which it sent as ordinary names), whose disagreements are counted per class and must match an independent count exactly. The optional roster only ever withholds more: with no roster, or an absent or null one, every answer is byte-identical to the previous release, checked against that release's published file over tens of thousands of inputs; a name withheld without a roster stays withheld with the same word, and the same current name, when the roster lists it as a name or an alias; a name delivered without a roster is withheld as not in the roster exactly when the roster lacks it; the engine's own namespace-covering forms keep their word; and a roster or options value that cannot be read marks the whole batch unreadable instead of falling back to no roster, while a readable empty roster is a roster. An end-to-end check takes the roster from a real run of the engine package that mounts a tool with an alias: every name delivered with that roster passes the engine's start-up name audit without being reported, and every name withheld as not in the roster is one that audit reports as matching no mounted tool. When the previous in-client implementation carries its own roster check, the two are compared under a given roster and may differ only for an alias (delivered here, withheld there) and for a refused name that the roster lists (withheld here, delivered there). |
430
430
  | `scripts/run-plugin-hooks-projection-test.mjs` | Plugin hooks: each command hook an enabled plugin declares is decided one by one as running in the engine, running in this client, or not running at all, and the page of hooks sent with a request is built from the same per-turn plan the client uses to skip its own copies, so one hook never runs in two places. Governance is judged first and always wins — a managed hooks switch-off, an untrusted workspace, safe or bare mode, or a governance read that fails sends no plugin hook and does not list it as a gap; managed-hooks-only (set directly, or through a merged non-managed hooks switch-off) keeps only managed plugins; the plugin-only customization lock does not touch plugin hooks. A hook reaches the engine only when this client started the engine on this machine, the engine reports plugin-hook support, the entry is a command, the plugin declares no sensitive option, and the event still fits the engine's per-event limits; the gate walks that matrix cell by cell, including the limit boundaries and a session goal hook counting toward them. A fact that was never read is reported as not known rather than as a fact: a host that does not say where the engine runs gets a "not known whether this client started the engine" reason, an engine whose capabilities have not been read yet gets a "not known yet whether it supports plugin hooks" reason, and the plan's two engine facts are null in those cases, not false. Events the engine never fires run only if the client says it fires them itself, and hooks the upstream behaviour itself refuses (option references in a shell-form command, an unset option in exec form, malformed entries) run nowhere. Exec-form arguments are passed element by element with only saved non-sensitive option references filled in; path placeholders are left for the executor. Sensitive option values never reach the request: with a host that wrongly supplies one, every string in the plan, the request body, the notice, the labels and the log is searched for it across eight cells. A host without the plugin reader keeps the previous request body and gets exactly one warning per settings port; plugin data that throws while it is being read (a throwing getter, a revoked proxy) is treated like a failing reader — no plugin hooks this turn, settings hooks still sent, nothing thrown; the not-running notice names the plugin and events, never a command or an option value, and escapes control characters in names. Command hooks from settings that carry arguments (a non-empty `args` array, which is the exec form, or any other non-null value) are removed from the request until the engine reports support for arguments, because the engine would otherwise drop the arguments and run the bare command through a shell; an empty `args` array is not treated as carrying arguments when the command is made only of letters, digits and `_ . / : + -` (the shell runs the same executable), so such a guard still reaches the engine, while an empty array on a command with spaces or shell characters is removed; `args` on a prompt or http entry, a null `args`, or an entry with no type is left alone, and those go out unchanged. MCP tool hooks, which the engine cannot parse, are removed only from a request built from a plan, whose not-running notice the host shows; a request built without a plan still carries them on engine-fired events, so the engine rejects the whole request loudly instead of a guard hook silently not running — the gate checks both request bodies against the engine's own hooks schema. Without a plan, every removed hook of that kind on an engine-fired event produces one warning per settings port, event and reason. Malformed entries still pass through for the engine to reject loudly, and passing null where the options object goes behaves like passing nothing; a `plan` option that is not a plan is ignored rather than turning the whole page into nothing, and a plan passed directly in place of the options object is recognised and used. A `plugin` key written by hand on a settings hook is stripped before sending (even when its value is undefined), because only hooks that come from the plugin reader may carry plugin context; the settings document itself is left untouched and a debug line records the count. The two hand-copied tables, the engine-fired event list and the engine limits, are checked against their owners. |
431
- | `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. |
431
+ | `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. Since 0.84.1 an address-shaped value (`scheme://…`) that sits right after an invisible single character — a format character, a lone surrogate, a non-whitespace control character, or the outlet's own escape token for one — is no longer let through as an address: the scheme stays and the host and path are replaced, with the query and fragment masked as before; `user:<password>@` followed only by stripped units and then a boundary, a port or another `@` is masked as userinfo. A value after real whitespace or after a colour sequence is still treated as an address, and clean text stays unchanged; forms that the previous release masked are checked not to leak on the same outlets. A `user:<password>@` candidate that is the value of a credential label or scheme word — including a label or scheme word split by stripped characters, and a quoted value that goes on past whitespace — is left to the label pass, so the whole value is masked as in the previous release; both reported shapes and frozen samples from the targeted pools are checked on every outlet. |
432
432
  | `scripts/run-ask-survives-posture-test.mjs` | The single posture predicate `askSurvivesPosture(card, facts)` for sessions whose standing mode would otherwise answer approval cards on the user's behalf (bypass-style modes). It reads two facts and returns one of three verdicts. The first is the ask origin stamped on the card: the question tool (`content_question`), an organization rule (`org_rule`), a hook (`hook`), an explicit ask rule (`ask_rule`), an organization policy or rule store that could not be read (`org_unavailable`, `rule_store_unavailable`) and the classifier's hand-off after its denial limit (`denial_limit_fallback`) must still be asked (the engine requires a real person to answer all three) and every other origin this build knows is left to the posture only once the host has also reported that its own ask rules did not match. The second is the host's own reading of its settings ask rules for this call: a positive match must be asked, and a command the host could not fully parse counts as no match. When the host reported no reading, every card outside those seven origins gets `unknown`, because an origin says who asked and not that the user's own ask rules did not match; an origin this build does not know gets `unknown` even after a reported non-match. `unknown` is never an approval: the host falls back to its own settings rules. The guard checks the verdict for every origin word, both with no host reading and with a reported non-match, against an independent table whose word set must equal the package's origin list, so a new upstream word fails the guard until it is classified; it covers the combinations of both facts, malformed inputs (non-boolean readings, empty or non-string origins, prototype keys, a different letter case), the fact that the predicate does not read the stronger bits on the card (those stay with the host's earlier checks), real card requests produced by the live-frame, parked-row and suspended-ask paths, and a closed, frozen verdict shape. Since 0.84.0 the product table itself is also checked against the SDK's runtime list of known origin words, so a word the upstream does not know cannot sit in the table. |
433
433
  | `scripts/run-engine-agent-absence-projection-test.mjs` | Absent background agents: when the engine stops reporting a background agent and no final state has arrived, the row is marked absent and this package owns every decision about it, so all clients agree. One predicate says whether a row is absent (the mark, not the status, decides). An absent row keeps its last known status, never counts as running, and is never counted as completed, failed or stopped; its elapsed time stops at the last moment it was seen, and its sentence says it may still be running. The end-of-turn sweep never settles an absent row (or a resident one). A row that comes back, or a real final state for the current cycle, clears the mark; a late final state from an earlier cycle does not. Absent rows are never removed at the short grace window. After the hard limit (30 minutes from the last time they were seen) the host is asked for the background-agent registry reading of each row: only a reading that the agent has ended or is not listed lets the row go, and each removal is returned as a fact the host must act on and announce; a reading of running, unknown, missing or unrecognised keeps the row and schedules nothing, so no standing poll is created. Until a registry reading is available every absent row stays. A row someone is viewing is held and reported separately only once the registry confirms it is gone. The row sentence, the removal sentence and the late-result sentence come from one place, never state an outcome or that the agent finished, and escape control characters in names, in the engine's removal word and in the late-result status. An end-to-end cell drives the real fleet projection and the real absence channel through every decision. From 0.84.0 the package also reads the registry itself: a reader lists the session's background registry only when the server advertises the listing, classifies its failures as unavailable (a failed read keeps the HTTP status and error code when the error carries them, and never invents either), and refuses a partly readable listing as a whole; a classifier turns the listing into the per-row reading, treating a registry status it cannot place as unknown and reading "not listed" only for keys shaped like registry handles and only when the host states the engine is a single process and names the row's session, because a key from another identity space proves nothing by being absent and, on a multi-replica deployment, a listing answered by another replica does not contain this session's agents at all (without that statement a missing row reads unknown and the row stays); and a predicate returns the agents the registry still counts as live that the host has no row for. A listing at the server's 500-row cap cannot show that an agent is gone either, so a missing row there reads unknown; the cap is checked against the engine's own list clamp and, when a server build is supplied, against the server's route. Each listing carries its session and a per-client sequence number, and a listing superseded by a later delivered read, or read for another session than the one the host names, counts for nothing. Only the exact listing object the reader returned can authorize removing a row or adding one: a spread copy, a structured clone, a filtered or a hand-built listing reads unknown and adds nothing, the listing, its row array and every row are frozen, the cap check uses the row count recorded when the page was read, and rewriting a listing's sequence number fools nothing. The fill predicate holds back agents first seen within the absence settle window — timed on one monotonic local clock from when this client's reader first saw the id, so a skew between the server's clock and this one, or reconnecting to a long-running agent, cannot skip the window — agents without a readable registration time and agents the host saw end since the read went out; the registry key of a row, progress or absence event is whichever of its two ids is shaped like a registry handle; and none of the reader, the classifier or the predicate throws. |
434
434
  | `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. |
435
435
  | `scripts/run-gate-interrupt-safety-test.mjs` | The two safety properties of running the suites themselves. The collecting runner no longer kills a suite outright when its time limit expires: it sends a termination signal first, waits a grace period for the suite to clean up, and only then kills it — and it records the suite as timed out however it exits, so a suite that exits cleanly after the signal is still not counted as passing. The limit and the grace period can be widened for a single suite in `gates-manifest.json` (an optional entry with a written reason; a malformed entry, including one with a misspelled field, stops the runner before any suite starts instead of silently falling back to the default, and an entry filed under a misspelled name is reported by this guard rather than ignored), and the negative-control suite is widened there. Every property is exercised on byte-for-byte copies of the real runner and the real negative-control suite in a throwaway directory: a suite that honours the signal finishes within the grace period, one that ignores it is killed when the grace period ends, and the summary lines are unchanged; once a suite has exited the runner waits at most a short drain window for its output pipes, so a child process that inherited them and outlives the suite neither turns a passing suite into a timeout nor holds the runner past the limit, the grace period and that window; the negative-control suite, interrupted in the middle of a rehearsal by any of the three signals or by the runner's own time limit, restores the file byte-for-byte, leaves no backup behind, stops the rehearsed guard together with anything it started, regenerates the build output (checked on a small project in the throwaway directory), and exits with 128 plus the signal number; a rebuild that would overrun the grace period is cut short and its compiler stopped; a backup left behind by an earlier run — next to a later target, loose in the source tree, or inside a symlinked dependency directory — makes it refuse to start without touching anything, naming the file and how to restore it. |
436
+ | `scripts/run-system-reminder-open-tag-test.mjs` | The opening-tag locator for the engine's `<system-reminder>` envelope, for hosts that split a message into envelopes and body text themselves rather than only stripping or unwrapping it. `findSystemReminderOpenTag(text, from?)` returns the start and end of the next opening tag at or after `from` (UTF-16 indices, `end` one past the tag) or `null`, and it recognises exactly the two shapes the engine mints — a bare tag, or a tag with a single `mark` attribute whose value is 22 base64url characters — through the very same matcher the strip and unwrap entry points use, so there is no second grammar to drift. The guard pins both engine shapes to the exact index, refuses the non-engine shapes listed in the integration guide plus near relatives (and does not let them swallow a real tag that follows), refuses an opening tag truncated at the end of the text, answers `null` without throwing for a non-string text and for a `from` outside the integer range 0..length, and shows that the locator ignores block context (a tag inside a block body, an unclosed tag and a nested inner tag are all located — pairing with a close is the caller's loop, as in the strip entry point). On twenty thousand seeded random strings mixing both shapes, truncations, near relatives and nesting, the set of tags the locator finds equals the set derived purely from the observable answers of the strip and unwrap entry points, and a strip and an unwrap rebuilt on the locator agree with the real ones byte for byte. Two timing cells show a single call over many near-miss tags and a full walk with `from` moving forward both stay linear |
437
+ | `scripts/run-detach-durable-off-verdict-test.mjs` | The verdict behind the durable-off hint, now a public entry point without the process-level gate. `isDetachDurableOff400(err)` answers whether an error is the engine's refusal of the detach-on-disconnect header on a deployment with no durable run ledger: status 400 and the engine's own refusal sentence in the message, and nothing else — in particular not the error code that refusal carries, because the engine uses the same generic precondition code for unrelated refusals on the same submit path, and treating those as this one would silently resubmit a turn without the header. Browser and desktop hosts, which never arm the terminal's detach ledger, can now ask the same question instead of keeping their own copy. The guard pins the positive case on the real sentence (with or without the code, with surrounding text, and on the error the SDK actually throws when a stubbed engine answers 400), the negative case on the engine's other refusals that share the code (verbatim, and again through the SDK), non-400 statuses, and eighteen malformed inputs that must answer `false` without throwing. It shows the verdict does not read the process ledger and that the hint still returns nothing before the header has been sent. The verdict never throws: an error object whose `status` or `message` getter throws, a proxy whose trap throws and a revoked proxy all answer `false`, each property is read at most once (and `message` not at all unless the status is 400), and the hint therefore answers nothing for those objects instead of throwing as it did before. Against a frozen copy of the previous hint, every other input gives the same answer byte for byte, armed or not. When the engine's source tree is available, it also checks that exactly one refusal site with that code carries the sentence and that every other one is answered `false` |
438
+ | `scripts/run-parked-resume-startup-test.mjs` | The one reading and the one sentence for a parked resume that fails while starting (`422 parked_resume.startup_failed`). The engine has two unrelated reasons for it — the parked agent's session is gone or the engine rejected the decision itself, or the approval is for a tool the agent inherits from the task that started it and this server cannot hand that tool over — and the response carries **no field that tells them apart**; the reason lives only in the engine's prose, which this package does not branch on. So the reading narrows the cause only by what the caller itself sent: a deny can only be the first reason (the inherited-tool refusal is minted for approvals alone, and the guard pins that upstream premise), while an approval, or a caller that does not say, gets one sentence that lays out both ways forward instead of guessing. Either way the verdict is fixed: sending the same decision again is not a recovery path, and deny stays available. The two sentences are minted in one place, never claim the card is still waiting (in one of the two shapes it is not), and avoid every word the package's own "already decided" scan looks for, since they are appended to failure text that scan reads; all five decision exits append exactly that sentence, and any other failure keeps its text byte for byte. The code also takes precedence over the package's "already decided" word scan: the engine's own text for the gone-session shape contains "not found", which that scan used to read as "someone already settled this card" and answer with a silent reconnect; now a failure carrying this code is classified by the code (no reconnect, the sentence reaches the failure text), while failures with no code or another code are scanned exactly as before. The predicate that says which codes classify themselves is public, so a client with its own "already decided" chain can ask it before scanning words. |
439
+ | `scripts/run-plan-review-injected-wire-test.mjs` | The plan-review orchestration accepts a host-supplied engine client, so a browser host behind a same-origin relay — which cannot install a token-bearing engine target — runs the package's own decision path instead of rebuilding it. With a client injected, the decision, its in-flight latch, the resend-once-without-the-key handling of a `request.field_conflict` refusal, the post-decide status re-pull, the six-state effect classification and the outcome text are byte-for-byte what the installed-target path produces, proven against a real relay-form SDK client and a fake engine while the installed target points at a different fake engine that must receive nothing. The engine-version evidence behind the three-choice card is read only under the cache key the host names, never from the installed target; every request of the relay-form client leaves without an `authorization` header, checked beside a token-bearing client on the same spy so the absence is not the spy's blindness; and a planted token never reaches a log line, an outcome sentence, a queue item or the host callback on any of the failure paths. A decision may carry a `reason` (trimmed, dropped when blank, capped at the server's limit), an `onOutcome` callback hands the outcome back to hosts that have no model-prompt queue (called once per admitted decision, never for a latched duplicate, and its own failure never affects delivery), a reopen with an injected client no longer needs a host delivery function on a non-default session, and the two capability probes take the same injected client with a bounded wait. A decision that the injected client's own time limit cuts off is reported as sent without an answer within the time limit (it may have taken effect), in the same sentence the installed-target path uses when its limit is reached, while a connection that cannot be made is still reported as unreachable and one that drops after the decision went out is reported as possibly applied (the same no-answer judgement the background-agent verbs use); a client built with the documented decision budget answers a slow approval exactly as the installed target does, and a client whose namespaces are callable functions is accepted as before. The package's own reads through an injected client — the two probes and the post-decide re-pull — settle at their deadlines even while the client sleeps through a retry back-off, and abort the request underneath; a malformed injected connection is reported as not sent, never thrown. |
440
+ | `scripts/run-session-policy-refused-removal-test.mjs` | The way out for a session that already carries a tool name the engine refuses. The engine refuses to start a run when any rule applied to it names a retired tool, or a name containing "__" without a protocol prefix, so a session whose own rule record holds such a name fails at startup on every run. The guard pins three pieces. First, the refusal test itself: it is compared name by name, in both directions, with an independent reading of the installed engine's own tables (retired names and protocol prefixes) over a corpus that includes padded, wildcard and parenthesised forms, and it is shown to be wider than the verdict used before a write — a padded or wildcard form is withheld there for another reason yet still refuses the run, so a removal driven by that verdict would leave the session broken. Every corpus name is then written alone into the deny list and into the allow list of a real engine run: the test says "refused" exactly when the run fails at startup with the refusal code and the model is never called, while the same shapes in the command lists are not audited at all. Second, the removal verb: it takes only the refused entries out of those two lists and leaves every other entry, list and ordering byte-for-byte, keeping an emptied allow list as a present empty list; after it runs, the same session's next real engine run gets past startup, for every refused name in the corpus and in both lists. Against the installed engine as an ordinary caller, taking a deny entry out is refused as a loosening — reported as such, written exactly once, the record unchanged and the next run still refused, never a false success — while an allow-list-only removal, which narrows, succeeds; against a model of the newer contract the engine has announced, the removal succeeds, and the same model still refuses a removal of anything else. A concurrent writer causes exactly one re-read, judged again on the other writer's record, and a second collision is reported rather than retried; read failures, coded write refusals and unconfirmed writes (no verdict, a receipt that cannot be read, or a receipt that does not account for the removal, including one that still holds a removed entry) each land in their own arm, and running the verb again after an unconfirmed write reports nothing left to remove. Third, the wording: one sentence per outcome; the loosening refusal names operators and says the record is unchanged, the unconfirmed one says the change may already be in effect, and none carries a rule string or engine text. The startup-failure sentence answers to that one code only, lists every place the entry can be set without naming who may remove it, and a real run with the name in the caller's own policy rather than in the session record yields the same code — which is why the sentence may not claim the entry is in the session record. The verb also requires the tool roster reported by a run on the deployment: a name listed there, as a tool name or an alias and compared exactly as the engine compares it, is left in place even when it is a retired name, because the engine accepts that rule — with a real engine run that mounts a deployment tool under a retired name (or with that name as an alias), the removal keeps the deny rule and the next run still starts. Without a roster, or with one that cannot be read, nothing is read or written and the outcome says why; a readable empty roster is still a roster. The removal sentence speaks separately about names taken out of the deny list and out of the allow list. |
441
+ | `scripts/run-fixture-flat-done-ratchet-test.mjs` | Test fixtures that feed the engine's terminal `done` frame are kept in the shape the live engine actually sends. Since the terminal record became a tagged cause (`result.terminal`), the older flat shape (`result.status` plus loose keys) reaches a client only in two ways: a stored row replayed verbatim from before the upgrade, and the server's own refusal envelope — so a fixture written in the flat shape tests the replay path while claiming to test the live one. The suite finds every `{ type: 'done', result }` literal under `scripts/`, follows `result` to the object literal it really comes from (in place, through a variable, through a helper's parameter at each of its call sites, through a spread, or through the rows of an iterated array) and counts the flat ones that carry no one-line note saying they model a replayed row or a refusal envelope. That count may only go down: it is a ceiling kept in the ratchet registry, and a count under the ceiling prints a step-down line instead of passing in silence. A planted corpus with a known number of flat, noted and cause-shaped frames in each of those forms must be counted exactly, and a note that sits inside a string or gives no reason does not count |
442
+ | `scripts/run-authority-envelope-mirror-test.mjs` | The authority-envelope tags that a peer's message body is defused against before it is written into a transcript line. Tags the engine uses to speak with its own authority (reminders, completion notices and the like) must never survive inside a body a peer wrote, or a forged completion notice could be read back on resume as a real one and poison the dedup ledger. The package keeps its own copy of the engine's list, so the copy is reconciled against the installed engine package in both directions: a tag the engine treats as authority and the package lacks is red, and so is a tag the package defuses that the engine does not, since that rewrites ordinary text in a peer's message. The engine does not export this list from its package entry yet, so the check reads it from the engine's own module by path and says so; it also confirms the list is still derived from the engine's envelope registry. At the rendered output, every engine authority tag placed in a body (opening, closing, with attributes, upper case) comes out defused on all three peer lanes, and the engine's non-authority envelope tags come out untouched |
443
+ | `scripts/run-upstream-tables-test.mjs` | Engine facts the package may not import at run time (the notice code register and its audience table, the reasons an injected MCP server is dropped, the protocol namespace prefixes and the retired tool names with their current names) are generated at build time from the installed engine package's own public entry point into plain literal modules, which are committed. The generated files must be byte-identical to what the installed engine produces today, no other file may sit beside them, each names the generator and the engine version it came from, none imports anything, and the package's public names are the generated objects themselves. The generator refuses rather than guesses: a table that is missing or empty, a notice code without an audience row (or the reverse), a third audience value, or a retirement note in a shape it does not know all stop generation instead of producing an empty or partial table, and it reads only the engine package installed in this package's own dependency folder. Two word lists the SDK already exports as values (who asked a question, and which word a mandated approval stands on) are now the SDK's own arrays rather than copies; because those arrays are not frozen, the package's own decisions are shown not to change when they are modified in place. Inside the package, the thinking-level and permission-mode vocabularies, the tier words, the rejection sentence and the cancelled code each exist exactly once in the source; until the next engine or SDK upgrade, the values of every table touched are also checked element by element, in order, against the previous release |
444
+ | `scripts/run-retirement-ledger-test.mjs` | Transitional code whose retirement condition used to live only in a comment now has a machine-readable trigger. `scripts/retirement-ledger.json` lists each piece (a compatibility read, a mirrored table, a normalizer inside the differential check) with an anchor into its own code, the retirement condition quoted from the source comment, and one of three trigger kinds: this package's version reaching a retire-by version; the installed upstream package reaching a version, exporting a name from its entry, or declaring a member on a given type (read with the TypeScript checker); or the reference client tree used by the differential check having moved to a given version of this package. The check fails when a trigger has fired and the transitional code is still there, when a row's anchor can no longer be found (the code is gone but the row stayed), when the quoted condition no longer matches the source comment, when a row has no machine trigger at all, when a retire-by version sits more than three minor lines ahead, and when the anchor of an already retired piece reappears. A reading that cannot be taken is reported as skipped when the absence is legitimate (no reference tree on disk, an upstream declared only as a peer and not installed) and is a harness fault otherwise. Before giving any verdict the judge proves itself on in-memory fixtures and on one known-present and one known-absent name in the installed upstream types. |
445
+ | `scripts/run-plan-review-decision-ledger-test.mjs` | The plan-review reopen path refuses on one decision ledger shared by every delivery path, not only the package's own POST window: a decision the host reports as handed over (before the card's answer reaches the package, or on the host's own delivery path) counts as in flight until it settles; a pending-row read that started before a decision settled is a stale snapshot of the gate just answered, so a reopen carrying that read mark refuses; and a decision proven not to have taken effect (nothing sent, a refusal that by contract changed nothing, or the engine proven to still hold the very same gate instance — a status alone does not prove it, since a run that moved on to a new plan gate reports the same status) neither counts as in flight nor makes such a snapshot stale, so a failed decision does not consume a legitimate reopen. The ledger's capacity bound never evicts the package's own in-flight decisions, so a duplicate approval stays latched however many decisions are in flight. The package's own in-flight window now runs from the POST to the moment the outcome is emitted (settlement happens before the outcome is queued, so a read the host starts from the queue port is not stale), while the duplicate-delivery latch keeps its POST-only window. A handed-over decision the package's delivery path then sends is the same ledger entry (the host's early settle does not release it); two decisions on one run each settle on their own; positive evidence that the run left the gate settles what is open without releasing the latch; bad inputs neither throw nor record. The ledger is checked against the terminal's current implementation vector by vector when that tree is on disk |
446
+ | `scripts/run-subagent-injected-wire-test.mjs` | The background-agent verbs — the child report read, the live tail, the task-handle output read and stop, steer, the resume assembly, manual compaction with its capability warm-up, and the delegated-prompt read — accept a host-supplied engine client, and the plan-review orchestration now resolves its client through the same single construction point, so a browser host behind a same-origin relay runs every one of these paths without an installed engine target. With a client injected, each entry sends only through it: proven against a real relay-form SDK client and a fake engine while the installed target (default and keyed slots alike) points at a different fake engine that must receive nothing, with the row's own session parameter carried, no `authorization` header on any request (checked beside a token-bearing client on the same spy), and a planted token never reaching a log line, a return value or a failure detail on the injected, installed-target and unreachable paths. The capability evidence behind the task-handle verbs and the row stop gate is read only under the key the host names, never from the installed target; the report, tail and compaction warm-up cache their own probe under that key and re-probe when it is absent. A malformed injected client or connection lands on each entry's existing cannot-send outcome (`null`, `no-wire`, `unavailable`, `offline`, `unreadable`) with zero requests, never a throw or an unhandled rejection. Every request through an injected client settles at the package's own deadline even while the client sleeps through a retry back-off and aborts the request underneath; the caller's own signal still applies, and the old signatures answer as before apart from the no-answer outcomes described next. A steer, resume or stop that may have reached the engine but got no answer — the package's own deadline on an injected client, the client's per-request limit on either path, or a transport failure that cannot be shown to have happened before anything was sent (the connection dropped after the request went out) — is reported with a machine-readable `unconfirmed: true` beside the unchanged `reason`, so a host can say it may have taken effect instead of calling it a failure; a cancellation by the caller after the request went out is reported the same way (only the wait was cancelled), while a real 4xx or 5xx answer, a connection that could not be made at all, a verb that threw before returning a promise, and a cancellation already in place before the call (nothing is sent) carry no such key. An installed target is copied by value when a call resolves it, so a delayed or retried compaction still goes to the engine and credentials it was resolved against even if the host's target object changes meanwhile, and a verb that throws synchronously is classified exactly like one that rejects. The resume failure classifier that hosts reuse for their own direct calls makes the same judgement from the same single source, so every client renders it the same way. |
447
+ | `scripts/run-session-background-stop-test.mjs` | The session-level stop for background work (`stopEngineSessionBackground`): one call asks the engine to stop the background tasks of a session, with `includeRetained` passed through as the required true / false choice. It resolves its client through the same single construction point as the other engine verbs (a host-supplied client, or the installed engine target for the chosen session slot) and sends nothing unless the engine has declared the stop face (`background.exitFaces` strictly `true`) under the capability key that applies: evidence that is missing or unread, an older engine without the face, a malformed flag, a flag found only on the prototype chain, or evidence recorded for a different deployment all end in `unavailable` with zero requests. The request body on the wire is exactly `{"includeRetained": <boolean>}` whatever object the caller passes. A 200 answer returns the receipts verbatim (identifier and outcome word; a row whose word it does not recognise is kept with that word as it is, so the host can still name it and the companion below sorts it as may still be running), counts rows whose identifier or word it cannot read instead of inventing them, and leaves the receipt list absent rather than empty when the body cannot be read. Refusals are sorted by machine code before HTTP status (missing service credential, no ownership face, unknown session versus missing route on the same 404, request shape, not permitted, rate limited with the server's wait hint, anything else), a server failure is an `error`, and each call sends exactly one request even through a client configured to retry. A request that may have reached the engine but got no answer — a time limit, a connection that dropped after the request went out, or a cancellation by the caller after the request went out — carries `unconfirmed: true`, while a real answer, a connection that never opened, or a cancellation already in place before the call (nothing is sent) does not. Malformed arguments, options, connections, clients, capability records, answer bodies and error objects never throw and never leave an unhandled rejection, and a planted token never reaches a log line or a result. A pure companion (`engineSessionBackgroundReceiptReadingOf`, with the frozen state list `ENGINE_SESSION_BACKGROUND_RECEIPT_STATES` derived from its wording table) sorts each receipt word into stopped, still running (kept on purpose) or may still be running, with one fixed sentence per state so every client says the same thing: each of the four words lands in its state, while any word it does not recognise — a future word, a different letter case, stray whitespace, a near spelling, a prototype-chain name — lands in may still be running with the word carried back verbatim, and non-string input (boxed strings and objects with their own `toString` included) lands there too without being coerced and without throwing; it never reports stopped for anything it cannot confirm. The three sentences differ, the two that are not stopped never use the word stop, and the state list and the states actually produced match in both directions. |
448
+ | `scripts/run-sealed-key-capability-test.mjs` | The engine's sealed-key discovery segment (`capabilities.sealedKey`), read the same four-state way as the other capability readers: exactly three non-empty strings (`alg`, `publicKeyId`, `publicKey`) make a usable reading and are passed through as they are; an absent key means *do not seal* (an older engine, or an engine that could not set up its key custody — the same action either way) and is never treated as a key; a present but malformed segment is dropped rather than turned into a half reading; and anything the segment carries beyond the three public fields — a private key, a creation time, anything else — never reaches the reading, so this reader cannot become a credential channel. The reading is checked against the projection a `7.104` engine package actually mints, both with a key store and without one. |
449
+ | `scripts/run-session-integrity-refusal-test.mjs` | Three refusals a session can meet on newer servers, read the same way on every client. When a session's saved history cannot be read, the server answers with one code whether the refusal arrives as an HTTP error, as the failure recorded at the end of a streamed run, or on the background run's record; one predicate decides it for all three carriers, it keys on the code alone — never on the status or on the engine's wording — and one sentence explains it: the history is damaged, retrying will not help, whoever operates the engine has to repair it, and meanwhile a new session works. The package's own end-of-run error line and the error of the result frame both carry that sentence, identically, and nothing changes for any other code. When a delete is refused because background agents the session started are still running somewhere this server cannot stop them, the refusal is read with the ids it names — each id checked on its own, a bad entry dropped rather than invented, and a list with no usable entry reported as absent rather than as an empty list, which would read as "no agents left"; a refusal because a run is still in progress is read with that run's id, and any other conflict falls back to a sentence that does not guess. When a decision reaches an approval that moved in the meantime (stopped, decided elsewhere, parked again, claimed by another decision, or — newly — belonging to a session that was deleted), the package reads the code rather than the error class, says the decision was not applied and that the pending approvals should be fetched again, and the approval legs treat it as "the approval is no longer where it was" and keep reading the run, instead of letting the wording of the server's sentence decide. When the server package is supplied, the guard drives the server's own code for each of these answers and pushes the result through the SDK's error mapping before the package reads it. |
436
450
 
437
451
  Each suite carries a floor that only moves up — a refactor that stops executing a group of
438
452
  assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**