@sema-agent/client-core 0.83.4 → 0.83.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +97 -0
- package/README.md +16 -15
- package/dist/adapter/activeRunSelfHeal.d.ts +3 -0
- package/dist/adapter/activeRunSelfHeal.js +39 -12
- package/dist/compensations.js +1 -1
- package/dist/displayUntrusted.d.ts +13 -0
- package/dist/displayUntrusted.js +751 -125
- package/dist/hooksWireCaps.js +78 -8
- package/dist/host.d.ts +2 -0
- package/dist/index.d.ts +2 -2
- package/dist/index.js +1 -1
- package/dist/pluginHooksWire.d.ts +9 -0
- package/dist/pluginHooksWire.js +119 -0
- package/dist/request/taskRequest.js +1 -1
- package/docs/INTEGRATION-CLIENTS.md +280 -66
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.83.
|
|
38
|
+
**Version:** 0.83.5
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -301,12 +301,12 @@ guard still cross-checks the table by name).
|
|
|
301
301
|
| `scripts/run-engine-caps-ledger-test.mjs` | A per-key disposition ledger for `GET /v1/capabilities`. The SDK's `Capabilities` grew from 74 keys to 93 in one release and nothing on the board could see it: this package consumes that table through four synchronous readers, and *nineteen new positions arriving while the package does not move* is exactly the disease shape this repo keeps logging on other axes — the fact is already on the wire, the package boundary is the cell that swallows it, and no client can read it however they write their side. So the ledger is reconciled **element-wise against the SDK interface in both directions**: a key the SDK added with no ledger row is red (someone must classify it), and a row for a key the SDK removed is red too (a registration that no longer does anything). Each row then has to survive its own claim — a `read` row names the source file, and the **code** there (comments stripped) must really mention the key, because prose asserting an alignment is the classic way these guards go hollow; a `not_read` row must have **zero** read sites in the tree, so wiring one up while the ledger still says the package ignores it is red rather than invisible. The census behind those two directions recognises five call shapes, each of which really occurs here — a reader whose base argument carries its own parentheses, a direct `caps.<key>`, a narrowing cast, an own-property read helper, and a `*_CAP` constant — and proves it on fabricated samples first, since a census that recognises one shape reports "nothing here" for the other four. What the guard deliberately does **not** judge is whether a position *ought* to be read: that is a design call, and the ledger only pins that every capability was looked at once by a person and that what they wrote down does not contradict the code |
|
|
302
302
|
| `scripts/run-sql-engine-capability-test.mjs` | The SQL-posture read face and the four-state capability reader underneath it. One capability cell here carries **four different things**, and each one points an operator somewhere else: nothing has been observed yet in this process (a one-shot doctor run is always in that state), the response arrived but carries no such key (an older engine), the engine explicitly answered `null` — *this deployment has no SQL backend*, which is a **positive fact** rather than an absence — and a full reading. Fold any two together and the screen states something flatly, confidently, and wrongly, so every positive control here is paired with a control pointing the opposite way, and the four sentences the doctor row can print are checked to be pairwise distinct and non-implying. The reading itself is narrowed no tighter than the mint: `txnMode: null` is a **legal value** — two of the three engines always report it that way, and the upstream type note names reading it as "optimistic" as the error — so treating it as malformed would throw away the entire reading for ordinary deployments, which is the same disease this repo logged when a consumer's domain was narrower than the producer's. A response that cannot be parsed **clears** the cell rather than leaving the previous engine's answer in place, and a separate invalidation port exists for the case the generation latch cannot catch — a same-port respawn whose new probe never succeeded, where the stale reading would otherwise be answered as current fact. Untrusted values (the isolation string is read back from a database server variable) are sanitised and bounded before display, and the bound is applied **before** escaping so a visible escape never gets cut in half. Finally the export names are themselves a guard: the shell still carries a copy that is meant to go red on the package's same-named export and be swapped out, so renaming anything here would silently disarm that lock |
|
|
303
303
|
| `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
|
|
304
|
-
| `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. |
|
|
304
|
+
| `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there. From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. |
|
|
305
305
|
| `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
|
|
306
306
|
| `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails |
|
|
307
307
|
| `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever. One reading here answers a question that the terminal state structurally cannot: whether this run was assembled with any file-and-shell tools at all. The engine's terminal vocabulary says a run finished, not whether the work got done, so an orchestrator that waits for the end and then guesses has nothing to guess from — while the assembly manifest already said it at the start, one row per mounted instance with the single condition that mounted it. The reading is three-state and both folds are refused: a roster that is readable and carries no such row is the engine stating a fact, while no roster at all is not that fact — the static half of a manifest never carries one, and an older engine reports rosters without naming the mount condition at all, where an empty count would be a statement about the reader rather than about the run. Those two are kept apart in the reason the reading carries, and the wording for every unknown case is checked never to claim the run had no tools. The same roster now decides the tool list on the first line of a non-interactive run: the host holds that line until the roster arrives and lists exactly what the engine mounted at the start of the run, in mount order. The guard runs a real assembly frame through the projection into the decision, and pins that the host falls back to the estimate only once the roster is known not to be coming — a manifest without one, an unreadable one, model output or the run's end arriving first — rather than on a timer alone (model activity counts, including a model call that is still waiting or retrying; an error line the stream synthesizes when a run fails before assembly counts as the run ending), that a sub-run's manifest is never mistaken for the run's own, that an empty roster is taken as the engine's answer rather than as silence, and that the wait bound covers both sequential default budgets the engine gives an external tool server to connect and list its tools. The holding logic itself lives in the package as a small per-run gate — buffer, decide once, release the held messages in arrival order, then pass through — and the guard drives real stream output through it to pin that the release happens exactly once, at the manifest, releasing exactly the held prefix. The ordering itself also lives in the package as a stream wrapper, and the guard checks the final output a consumer reads: the first line is always the tool-list line, a message that arrives while that line is still being built comes after it, a timer firing races nothing out of order, a source that ends or fails before the decision still gets its first line and held messages out before the error, and an early exit closes the source |
|
|
308
308
|
| `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
|
|
309
|
-
| `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`,
|
|
309
|
+
| `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, ten words) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the gate parses the members straight out of the installed SDK's declarations and reconciles them against the package's tables in both directions, so a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because an unknown word means different things in each: a denial layer this build does not know may have been added by a newer engine or may come from a damaged record, so its sentence says it cannot tell which instead of asserting damage; the asker vocabulary is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine). Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer; it must not say the question is asked every time — an answer for this one call may come from the person, a hook or an automatic check the deployment runs — and its wording is checked against the engine package's own description of the mandate. A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. Since 0.83.2 a sixth list covers the word a mandated question stands on (`APPROVAL_MANDATE_WORDS`, six words): it must equal the engine's own list word for word and in order, membership is exact, the card reader `readApprovalMandate` answers only for an own key holding one of the six words, and each word has one fixed sentence explaining why the question must be confirmed — six distinct sentences that never point the reader at writing a rule, never promise a question every time, never claim only a person may answer, and repeat no other sentence the package mints. The list is also pinned against the engine's type at compile time in both directions, while the published build references no engine package at all: every `.js` and `.d.ts` file in the build is scanned, and the same scan is first shown to fire on references planted in a scratch directory. |
|
|
310
310
|
| `scripts/run-engine-identity-test.mjs` | The engine generation anchors on `/health` (`pid`, `instanceId`, `startedAt`; engine >=7.67.0). `/health` is the one unauthenticated door and its heartbeat is always green, so "another host restarted the shared engine" used to be discoverable only by having some authenticated request hit a 401 first — a path that misreads a restart as a network fault. The reader narrows each anchor independently (one malformed field never hides the other two) and always hands back a reading object rather than an absence, because the caller is asking which anchors answered, not whether there was a response. The comparison is a three-word verdict, not a boolean: `unknown` when the two readings share no comparable anchor at all — an empty intersection means nothing could be compared, never that nothing changed — and the boolean convenience is pinned so that only `true` is an assertion. Any comparable anchor differing decides `changed`, so a reading whose `startedAt` matches while its `instanceId` does not cannot be waved through as the same life; precedence only decides which anchor gets named in the diagnosis |
|
|
311
311
|
| `scripts/run-posture-knob-projection-test.mjs` | The three deployment knobs on the operator face (`serverGates.durableApproval` / `streamAskWindowMs` / `sessionAutoTitle`, engine >=7.67.0), each read as a value **plus who set it plus one operator-facing pointer** rather than a bare value — a bare boolean cannot answer why this particular machine is on this setting or how to pin it back, and a default that flips with the deployment shape is invisible without that. A worker too old to report readings still sends a bare boolean; the reader folds it into the same shell so consumers keep one branch, but raises a `legacy` bit, answers `undefined` from the machine-readable source accessor, and mints a sentence that contains no source word at all — claiming a source nobody reported is worse than admitting the worker cannot say. The other two knobs are honestly absent on such a worker rather than defaulted, a malformed side knob drops only itself while the anchor knob drops the whole reading, and the four sentences are pinned literally distinct so an operator can tell "not observed" from "not reported" from a real value. The last leg reads the installed SDK's `openapi.yaml` and `types.d.ts` directly, including a pin that exactly one knob on this face is numeric — the premise the millisecond-to-prose rendering rests on |
|
|
312
312
|
| `scripts/run-terminal-facts-projection-test.mjs` | The four unconsumed terminal-receipt facts: `TaskResult.effectiveReasoning` / `effectiveMemoryScopes` are narrowed into `_sema_effective_reasoning` / `_sema_effective_memory_scopes` on the CC-shaped `result` (success and error envelopes alike; a malformed value mints nothing, never a default tier), the resume **reopen** family (`resume.env_failed` / `tool_unavailable` / `tool_contract_mismatch`) is a frozen closed set with a reader and three-sentence copy that is disjoint from the refusal and retry-later sets, and `routePairingVerdict` reads `ModelInfo.routePairing` as ok / broken / unknown without policing the open set. A fifth section pins the structured-output key on the success result: the CC-spelled `structured_output` is the only home for the value the wire calls `structuredOutput`. The camelCase spelling this package used to mint on its own — a misspelling of the CC field, not an additive field of our own — rode alongside it for exactly one release (0.79.1) and is **absent from 0.80.0 on**, pinned both by own-key and by `in`, so a consumer still reading the old name sees `undefined` rather than a stale copy. The wire position is read exactly once, so a value-changing accessor is only ever asked for its first answer; absence stays absence; a wire key that is present but `undefined` mints nothing, since a key whose value is `undefined` makes a consumer that tests presence read "the engine produced nothing" as "the engine produced an empty result"; falsy-but-present values such as `null`, `0`, `""` and `false` are still minted, and so are shapes that are not records at all — an empty array, a populated array, a string, a number, a boolean — each carried through by the same reference, because the shape of that value is decided by the caller's own schema and the package does not get to filter it; and the error envelope carries no such key, because the CC error arm has no such field. Which spelling CC itself declares is witnessed from the mirror's own syntax tree rather than a constant copied into the guard, so the day that field is renamed upstream the guard says so. |
|
|
@@ -321,7 +321,7 @@ guard still cross-checks the table by name).
|
|
|
321
321
|
| `scripts/run-panel-identity-normalization-test.mjs` | One background subagent has two ids on the wire — the fleet row id tail and the `task_progress` task id (its transcript id). Every panel event goes through one funnel that rewrites the `tick` / `end` task id onto the fleet row's id once a `fleet-row` has registered the key (`transcriptId` first, `parentToolCallId` as the fallback), carrying the original as `wireTaskId` and marking `taskIdOrigin`; an unbound tick whose row has not arrived yet waits one beat (bounded) and is released verbatim on the next tick / `end`, when the buffer is full, or after `MAX_HELD_WIRE_TICK_BEATS` other fleet-row / end / sweep events (a `sweep` itself leaves it alone: there is no row to settle yet); a normalized `end` that carries no cycle identity borrows the registering row's, so a late close of a revived task is recognized as stale; the key table is an LRU (a task that keeps ticking is never evicted by newer registrations); the residency mark migrates with the id and both keys are cleared on settle — except that a stale (previous-cycle) terminal never clears the revived row's mark — so the notification lane can clear it. Once a tick has been delivered verbatim under its UUID, that UUID is the subagent's key: later fleet rows and fleet-side ends are rewritten onto it (the tail kept in `wireTaskId`), so a consumer sees one row in every arrival order; a late tick from a previous cycle is dropped rather than folded into the revived row. |
|
|
322
322
|
| `scripts/run-prompt-assembled-projection-test.mjs` | The `prompt_assembled` frame (one prepare's prompt-assembly manifest) projected to an internal arm and then to the additive `prompt_assembled` chrome event — the per-section / per-block **character** counts, the mounted tool names and `totalChars`, each key present only when the engine really sent it (the frame's `constitution` is deliberately not carried: no consumer asks for it today, and every published key is a contract to keep). The manifest carries **no token counts** anywhere upstream, so this projection mints none: a token figure derived from characters would be an invented number, and the engine's own estimate lives on `context_usage.sections[].tokens` (same id wordlist, joinable). Bad rows are dropped one by one, and a face that loses every row reads as an absent key rather than an empty array — so an absent face means only "this event carries no readable view of it" (an absent upstream key, an empty array and a fully filtered list all land on the same shape) and is never reported as a diagnosis about the engine. `blocks[].id` and `sections[].id` are two different wordlists with a many-to-one relation, and the token join against `context_usage.sections[].tokens` only holds when both sides carry a section view. A frame with no readable composition key at all is malformed, ids and slots are read as an open set, one chrome event per frame with zero transcript rows, several prepares per task are all handed over (de-duplication — "take the last one" — is the host's move), and the lane is told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). Both entry points obey the same rule: the adapt layer rebuilds every row too, so a host pipeline (or a replayed transcript) that feeds the raw frame straight into `adapt()` cannot smuggle extra keys (`tokens`, digests, aliases), a negative `chars` or a `null` row into the chrome payload, an empty array does not count as a composition face, the identity keys are snapshotted once on both paths (read exactly once each, a throwing accessor rejects the whole frame — reading one twice is what lets an accessor frame land on a different lane on each path), and the two paths are compared verbatim so the two readers cannot drift. |
|
|
323
323
|
| `scripts/run-compaction-outcome-projection-test.mjs` | The `compaction_outcome` frame (a compaction that did **not** end as compacted: mooted by the task ending, failed, …) projected to an internal arm and then to the additive `compaction_outcome` chrome event — `outcome` required and verbatim (open set), `trigger` / `reason` present only when the engine sent a non-empty string, malformed frames dropped, zero transcript rows, the lane told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). |
|
|
324
|
-
| `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
|
|
324
|
+
| `scripts/run-approval-card-retract-test.mjs` | The approval card's **decision-free retraction** and the in-stream frame leg's **outcome hand-back**: a host that must withdraw a card that no longer has a decision channel (session switch, engine switch, a tracker reporting the ask gone) answers `{ kind: 'retracted' }` and the package sends nothing on any of the three legs (in-stream frame, suspended ask, durable park), reporting `decision: 'unresolved'` with a `retracted` flag; `aborted` / `failed` / `deny` keep their meaning (a real deny is still posted), and `onToolApprovalOutcome` hands every in-stream outcome back to the host exactly once, tolerating a throwing or rejecting callback. Also the single source for the host-side approval-outcome note (`approvalOutcomeNoteOf`): `settled` is whether the decision was delivered, `retracted` is an independent key present only when the card was retracted, and `detail` is the retraction / edit-refused sentence or the refusal code and message — never a fabricated sentence. |
|
|
325
325
|
| `scripts/run-memory-spec-wire-test.mjs` | The per-agent **memory spec** (`agents[].memory`) read once for every client, plus the judge for the engine's **closed** key list. Two states are kept apart that clients habitually collapse: an absent `scopes` means *no layers were specified*, never "zero layers", and an explicit `writeScope: null` is a positive fact — this run has memory **read-only** (no remember tool, no consolidation write; recall still works) — which is neither "unspecified" nor "memory off". Each of the four keys is read once, on own properties only (an inherited key never reaches the wire, so reading one would report a value the engine cannot see), and a key that is present but unreadable stays in its own slot instead of collapsing into "unspecified"; `enabled` must be a strict boolean and `scopeContract` is an open-set verbatim word. A spec that cannot be read at all answers *undefined*, kept distinct from an agent that simply has no spec. The judge earns its keep on the consequence: the engine checks this spec against a closed list, so one unlisted key — most often the retired singular `scope` — is refused together with the **whole agent definition**, not just that key, and the single sentence minted here says so. What counts as "on the wire" is decided by the bytes, not by the shape of the in-process object: both the reader and the judge work off a `JSON` snapshot of the spec taken **in its property position** (wrapped under the same key, never serialized as a root value — otherwise a `toJSON(key)` that branches on the key hands us one shape and the engine another: one such input made the snapshot say *read-only, no violations* while the real bytes carried the retired key and a writable scope), because `Object.keys` and `hasOwn` disagree with the serializer in ways that change the answer — a non-enumerable `writeScope: null` would otherwise be reported as "memory is read-only for this run" while the engine receives *unspecified* and may still write; a key whose value is `undefined` would be reported as a violation that never leaves the process; a `toJSON` (even inherited) adds keys that `Object.keys` cannot see, including the retired singular one; and a throwing getter would let the judge claim it had looked when the spec cannot be serialized at all. A spec that fails to serialize is reported as unreadable by both ports, and the snapshot is taken once, so every getter runs exactly once. The judge answers in three states, never two: `[]` is an assertion (*looked, nothing unlisted* — including an agent that carries no spec at all), a non-empty list is what it saw, and *undefined* means it could not read the spec (a non-object item, an array, an unreadable `memory`, a throwing getter) — an unreadable spec never poses as a clean one, and an array is not a spec so its index keys are noise rather than findings. It reports only the snapshot's string keys, sorted and bounded, so a prototype, symbol, non-enumerable or `undefined`-valued key is never blamed while a `__proto__` that really does serialize is; the empty-string key is kept rather than dismissed as noise, because it does serialize and dropping it left a non-empty violation list with nothing said about the consequence; every key name in the sentence is quoted and escaped one code point at a time, so no escape is ever cut in half (a half-cut escape used to make the closing quote itself look escaped) and an empty name, a key literally named `""`, a key containing a backslash and a real control character versus a literal `\uXXXX` all read as different violations; a name too long to show is marked `(truncated)` outside the quotes with a pointer to the judge's verbatim list, so a prefix is never presented as the whole key — two long names sharing a prefix do show the same, which is why the mark and the pointer are there; key names are sanitized and bounded on the way into the sentence while the judge itself hands back the verbatim key, because sanitizing belongs in prose and never in a verdict. The announced future key `projectKey` is still unlisted today and is reported as such, with a sentence saying it is not a typo. The two construction-time refusals (`config.memory_project_key_spelling` — a spelling, 400; `config.memory_write_scope_mismatch` — a conflict with the scope already in force, 409) join the existing `config.` recognition table rather than a second word list, and each gets one sentence stating that the refusal landed **before the run started**, so nothing ran; the engine owns the triage and an unrecognised code gets no sentence at all. The write face is widened in the same batch so the package can actually mint what the reader can read: `TaskAgentWireMemory.writeScope` is now an optional `string \| null`, since a reader that understands "memory is read-only for this run" while the writer cannot express it is worse than no reader at all — it makes the support look real. Minting `null` survives serialization and reads back as read-only, minting `undefined` drops the key and reads back as unspecified, and the projector still pins `writeScope` explicitly every time. The accepted key list is reconciled against an upstream witness rather than a second local copy: the guard reads the SDK's own declaration comment for this key, requires the two sets to match in both directions, requires that comment to still name the singular `scope` as retired, and requires it to still not mention `projectKey` — so the day upstream admits that key, the guard goes red instead of the package quietly continuing to promise a 400. Each port takes its own snapshot, so a consumer that wants one self-consistent answer about a spec that can still change under it should read `spec.unknownKeys` off the reader — which comes from the same snapshot as the four slots — and send that materialized data rather than the live object. |
|
|
326
326
|
| `scripts/run-approvals-feed-unknown-test.mjs` | The approvals feed tells three states apart: **N items waiting**, **nothing waiting**, and **this fetch did not come back, so we do not know**. Every way a fetch can fail (the call throwing or rejecting, a body that is not an object, a `livePending` section that is not an array — including the `null` seen in the field, a `pending` that is not an array or holds a malformed row) publishes `{kind:'unknown', why, at, mode}` on the subscription — never an empty snapshot and never silence. Real snapshots carry `kind:'snapshot'`; `snapshot()` still answers only with the last real one (a fact about the past) while `reading()` answers whether it is current (`unobserved` / `present` / `unknown`). Recovery always publishes a real snapshot again, even when the contents are byte-identical to before the failure. An unknown reading is never counted as zero: the awaiting-decision counts read `null`, the view is empty, and the tracker reports no removals, so cards on screen are not retracted for a failed fetch. Retry, backoff and circuit-breaking are unchanged. |
|
|
327
327
|
| `scripts/run-approval-resolution-test.mjs` | The single discriminated union for **how an approval decision ended** (`ApprovalResolution`: `decided` / `not_sent` / `unsettled`) and its one mapping entry `approvalResolutionOf`: every outcome of the durable-park leg (12 shapes) and of the in-stream frame / suspended-ask leg (4 shapes) lands on exactly one arm and cause; the three meanings of `decision: 'unresolved'` (retracted card, refused edit, respond that never settled) land on three different arms, with `retracted` winning when both flags are set; an interrupted durable card really posts a deny, so it is `unsettled` (`interrupted`), never `not_sent`; a safety stop never claims the decision left the package, and a refusal is only attributed to the engine when the outcome carries positive evidence (a wire error code, or the pointer key the engine mints on a rejection body) — an aborted or code-less decide failure is reported as a plain decide failure; the decision word is passed through without re-validating the closed set; an unreadable outcome is `unsettled` (`unreadable`), never guessed as `decided`; both cause vocabularies are frozen tuples with every word covered by a case, plus the three predicates; the approval-outcome note (`approvalOutcomeNoteOf`) is now derived from the union and compared key-by-key against a reference copy of its previous logic over the released inputs, with a self-check that the comparison can fail; a source-text pin asserts every `return` carrying `respondRefusal` also carries `'unresolved'`. No behaviour change: the existing outcome types and keys are untouched. From 0.80.0 that last pin reads the syntax tree instead of scanning lines: the same return had been rewritten across several lines with conditional spreads, a shape a line-wise search misses entirely, which would have quietly turned the pin into a check of nothing. From 0.83.0 the parked leg's card-stage failure carrying `editRefused` (an edited approval on a card whose tool arguments were not available) reads as not sent with the `edit_refused` cause, the same cause the live leg already had; the flag counts only as a strict `true` on the card stage. |
|
|
@@ -339,7 +339,7 @@ guard still cross-checks the table by name).
|
|
|
339
339
|
| `scripts/run-display-cap-order-test.mjs` | The order in which untrusted text is sanitised and length-capped, across every mint point that puts an engine- or database-supplied string on a screen. The sanitiser rewrites each invisible character as a six-character escape, so capping the **raw** string first and escaping afterwards hands the screen six times the width that was budgeted — a forty-character allowance becomes two hundred and forty. The guard does not hardcode that allowance, because each mint point wraps its field in different fixed prose and the prose moves: it anchors on the deciding quantity instead, feeding one benign and one control-character input of the same length through the same mint and requiring the second not to come out longer. That criterion is immune to wording changes and stays sensitive to the expansion, and it is `<=` rather than `==` on purpose — a correct escape-then-cap backs the cut off a partially-consumed escape token, so the control-character line is legitimately the shorter of the two, and demanding equality would score that avoidance as a regression. Each mint is bracketed by two positive controls (the input really reaches the screen; the cap really engages) and the expansion predicate is shown to turn red against a deliberately cap-then-escape reference, so an all-green run cannot mean the guard simply measured nothing. The shared mint point is checked directly for the two avoidances it owes — never splitting an escape token in half, which would leave something on screen that looks like the beginning of a complete answer, and never splitting a legal surrogate pair, which would manufacture the very lone surrogate the sanitiser exists to catch |
|
|
340
340
|
| `scripts/run-seat-task-request-origin-test.mjs` | Where every field of the seat lane's send-message payload comes from, and whether it actually lands anywhere. The seat payload is a closed interface this package mints itself, and most of its fields are meant to ride verbatim onto the engine's request body — two facts nothing used to connect, so both directions could drift in silence. A seat field could be named after a request position that does not exist, in which case a client writes to it, the wire carries it, the engine ignores the whole key, and the screen shows a switch that does nothing; conversely a new request position could arrive with no seat to sit in, which is **structural** absence — the closed set *is* the carrier, so a decision missing from it has nowhere to be put at all, the same shape logged when the effort dial had no seat. The guard turns each field's origin into data: either it names the request position it forwards to, or it is declared seat-local with a written reason, and the two are mutually exclusive. Forwarding claims are then checked against the **installed** SDK's type declarations, parsed rather than restated — a hand-copied list of position names would only ever prove that two transcriptions agree. The parser is held to reading top-level positions only, since a nested option object's inner keys would otherwise be mistaken for positions of the request itself, and it proves that discrimination on synthetic input before any verdict is given. The two subagent fields carry a standing regression pin, and the retention window's inner keys are read from the declaration the same way, so a seat that offers a tunable window cannot offer one the wire has no room for |
|
|
341
341
|
| `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header |
|
|
342
|
-
| `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters When the host declares the non-interactive lane, a sub-agent's tool calls and results are also forwarded into the main output with their parent tool-use id (the sub-agent's text and thinking still stay out, as in the reference CLI); without that declaration the output is unchanged. |
|
|
342
|
+
| `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters. When the host declares the non-interactive lane, a sub-agent's tool calls and results are also forwarded into the main output with their parent tool-use id (the sub-agent's text and thinking still stay out, as in the reference CLI); without that declaration the output is unchanged. |
|
|
343
343
|
| `scripts/run-subagent-content-budget-test.mjs` | The **byte** budget on the sub-agent transcript ledger. It used to be bounded only by *counts* — so many entries per child, so many children — and a count is not a budget when a single entry has no ceiling of its own: one tool result carrying an inlined attachment, or one long model answer, and a single slot sits on tens of megabytes. The guard anchors on how many bytes are **still held** after over-filling, not on whether truncation fired, because an implementation that flags the overflow without actually dropping anything satisfies the second and not the first. Dropping is required to leave a record — how much went and where the retained content now starts — and that record has to reach the render plan, because content that vanishes with no marker gives the reader a transcript shorter than what happened with nothing to say so; the record is one per child, updated in place, pinned to the front, and excluded from the budget it describes. Order matters and is checked: oldest entries go first and the live tail is trimmed only as a last resort, since taking the text the user is watching stream while older history survives is the wrong end. The total budget evicts a whole least-recently-used child rather than shaving every child, and the configuration surface is fail-loud on zero, negatives, non-finite and non-integer values — a silently ignored budget is the exact failure this exists to remove — with the rejection proven atomic so a bad second field cannot leave half a configuration behind. The defaults are checked to be a magnitude that can really be reached, since a number too large to hit is a field rather than a budget |
|
|
344
344
|
| `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other. One more thing the totals cannot settle: the row key has **two namespaces** — the originating-task id and the parent call id it falls back to — and nothing upstream promises they are disjoint, so the same literal can name one child's identity and another child's parent call. Accumulation therefore keys on the origin as well as the id; the delivered table still keys on the bare id, and a cross-namespace clash is merged into one row that says so, with the partial verdict forced, because a row count and a turn count can both reconcile while the attribution behind them is wrong. The table itself is likewise a **per-stream snapshot** handed to the terminal projector by value rather than left on the caller's context: the three terminal projectors are public, so a host may drive one run through the stream and project another's terminal directly on the same context, and a table left behind would be attributed to whoever projects next — silently called complete whenever that run's own totals happen to match. Without a snapshot, both table keys are simply absent |
|
|
345
345
|
| `scripts/run-result-text-backfill-test.mjs` | What happens when the terminal frame's answer text and the text already on screen do not match. A turn's answer normally streams in and the terminal frame carries the same words again, so the two agree — but when the connection drops mid-answer and the reconnect brings the finished version, "this turn already produced assistant text" is true, the terminal fallback is skipped entirely, and the screen stays permanently short of whatever arrived while the stream was down, with nothing to say so. Four cases are pinned. Nothing on screen yet: render the terminal text whole, byte for byte the previous behaviour. On-screen text is a **prefix** of the terminal text: emit only the missing tail, and the guard measures the deciding quantity — the total bytes that reached the screen must equal the terminal text, which fails both for a missing tail and for a re-render that would print the first half twice; when the two are already equal, nothing is emitted at all. Terminal text is a prefix of what is on screen (an engine-side trim): touch nothing, since there is nothing missing and overwriting with the shorter version would erase what the reader already saw. Neither is a prefix of the other: emit **nothing** and raise a fact instead — which version counts is the engine's to say, and appending the terminal version after the streamed one composes a passage nobody ever wrote. That fact carries lengths rather than text, so a renderer is not handed a third version to choose from, and its declared duty is to *reword* the transcript line, never to render more. A cross-segment case proves the comparison reads the whole committed answer rather than the last segment, and the whole thing is driven through the real two-stage path rather than hand-built messages. The comparison only holds if both sides are the same kind of thing — one assistant message against one assistant message — and that depends on the client knowing where a message ends. A tool card is one place a message ends, but not the only one: when a model round finishes and the next one starts writing prose straight away, with no tool call in between, the two passages belong to two different assistant messages even though nothing visible separates them on the wire. Reading them as one used to glue the two passages into a single transcript line, running the end of one sentence into the start of the next, and then handed the terminal comparison a concatenation to check against the engine's **last** message — so a perfectly ordinary multi-part answer was reported to the reader as *the final answer does not match the text streamed above*. The end of a model round is therefore treated as the end of an assistant message: the pending prose is committed and a new message begins. That holds even when the round reports no spend at all, because *this round is over* and *this is what it cost* are two different facts and only the first one decides a boundary. A round belonging to a **sub-flow** decides nothing for the leader, and the test pins all three identity members, empty strings included, against a control that proves the same shape really does divide when no identity is present. The rest of the section is regression: a boundary landing immediately after a tool card must not lose the fallback that lets a card-ends-the-turn answer compare against the prose before the card; repeated boundaries with no prose between them must not mint empty messages or lose track of which message the comparison should read; a reconnect that brings the finished answer still backfills only the missing tail; the leg that already carries whole messages is not divided twice; and a segment boundary that arrives **after** the round ended still replaces the segment authoritatively and hands back the row attribution, without minting a second copy of the passage. A last group covers where a message boundary meets a segment that the engine announces late or not at all, and it pins only the half that is unambiguous. A round whose prose the engine never announces, followed by one it does, used to have the first passage **silently overwritten** by the second one's authoritative text — no transcript line for it and an empty attribution list, so a host had nothing to correct; it now keeps its own line and the attribution is handed over in full, with the offset measured on that passage rather than on the authoritative text. Which of the two passages survives a compliant host's rewrite is deliberately **not** asserted: the row named belongs to an earlier message, exactly as it already did at a tool-call boundary, and the underlying cause — the client splitting messages on its own boundaries while the engine announces segments on content blocks — is recorded as a known limit rather than pinned as a desired outcome. Where the second round's prose arrives as a whole block instead, there is no segment announcement at all and both passages keep their own line **in the order they happened** — previously the whole block was written first and the earlier passage only landed at the close, so the reader saw them reversed. An announcement that arrives after its round has already ended, and whose authoritative text merely extends what was shown, is pinned on the two things that are not in question: the bytes on screen add up to the authoritative text exactly once, and the divergence signal with its offset is still emitted so a host can reword. That ordering is unreachable on the installed engine — the announcement is pushed while the message is still being assembled and the round end only after it is complete, both through one synchronous dispatcher onto one queue — so the case is defensive; the late-announcement path exists because a tool call can be admitted before assembly finishes, which a round end cannot. Finally, one guard names a layer seam rather than a behaviour: the round-end frame is folded into a neutral usage arm that carries **none** of the three identity members, so a round belonging to a forwarded background child arrives with nothing to judge and the sub-flow cutoff cannot reach it. That guard reddening is the signal that identity now survives the fold and the cutoff has become effective |
|
|
@@ -351,7 +351,7 @@ guard still cross-checks the table by name).
|
|
|
351
351
|
| `scripts/run-usage-verbatim-channel-test.mjs` | The two complementary usage disciplines (core 3.0.0 metering semantics): the CC `ModelUsage` mirror stays pure (five pinned keys, `totalInputTokens` has no seat), while the sema-owned channel forwards the engine `turn_end.usage` object **verbatim** (six keys, incl. `totalInputTokens`) via `last_turn_usage.engineUsage` / `handle.latestEngineUsage` — honest absence on pre-3.0.0 engines, no fabricated zeros |
|
|
352
352
|
| `scripts/run-plan-review-decide-verify-test.mjs` | `decidePlanReview`'s post-decide honesty ([2315]/[2316], engine RB-471 family): a 2xx from the decide endpoint is **not** a terminal — the wire re-pulls the task status and words the outcome by the real shape (still-locked / legal new gate / genuinely left park / unverified), never claiming success it hasn't earned; when the engine answers that the session's stored resume context cannot be read, the outcome names the unreadable row and says the decision was not applied. Driven against a real fake-engine HTTP server through the shipped dist. The outcome queue item also carries a machine-readable `_sema_planReviewOutcome` (task id, a package-minted dispatch number, decision, effect) so a host can tell which in-flight decision an outcome belongs to without searching the prose; the prose is byte-identical, the number is minted only for a decision the in-flight latch admits, and a caller that passes no metadata gets no key. |
|
|
353
353
|
| `scripts/run-shell-gate-durable-allow-test.mjs` | #110: the durable approval leg for **shell** gates. The tool_end HOLD/REJECT predicate must cover Bash the same way park detection already does (otherwise the park poison frame `Operation aborted` hits the transcript, `endedCalls` swallows the real replayed result, and the user who pressed Yes watches a command that really ran be reported as aborted); a replayed, already-decided park must resume reading the stream instead of being reported as a failed turn; `lastEventId` must track numeric `seq` too. Mutation-proven: each of the three fixes reverted turns the gate red |
|
|
354
|
-
| `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings. A final section pins the decide-operation observer: one `start` in the same tick as the first request and exactly one `end` after the last attempt has settled, across success, retried timeouts, exhausted transient failures, semantic refusal, binding mismatch and both kinds of caller abort, with nothing between retries; an observer that throws or rejects — even when logging that fault fails — never changes what is sent or returned, and the stream-level dependency reaches both durable park legs. A package-internal re-delivery of the same decision (plain approve after an older server rejects the session-scope flag, or a re-send without the attribution key) is reported as one operation with a single start and end. An operation handle only groups sends for the same session and bound call — a send for another gate through the same handle is its own operation — and closing a handle while a send is still in flight defers the end until that send settles. When a person picks "allow for this session" on a parked card and the grant is known not to have been stored — the server answers so, which newer servers do for the gates they can recognise from the parked row as needing a person each time, or the server refuses the session-wide grant with one of the refusal codes that are fixed by the row or the deployment and the package falls back to a plain approval — the parked path now says so with the same line the live path uses, and the receipt carries the server's bit for hosts that call that path directly; an answer without the bit, or any other failure — including a conflict that an internal retry can hit after an earlier send already stored the grant — is treated as unknown and says nothing, a plain approval that never asked for the grant says nothing, and a host logger that throws after a successful decision, on either card path or in the bridge's retry step, can no longer turn it into a second send or a failure. From 0.83.0 it also pins approval cards whose arguments are not the tool's real input: a live frame whose arguments were omitted over the size cap or never sent (the card holds at most a path recovered from the question), and a parked approval whose stored input is missing or replaced by the upstream size marker. On every such card an edited approval sends nothing at all — no respond, no decide, no session-grant attempt — ends as edit-refused, raises the same notice once and flags the card so hosts do not offer editing; both production entry points are driven end to end, and the parked path neither mistakes it for an already-decided gate nor reconnects. The explanatory line is pinned word for word in its four forms: a recovered path is the only thing it claims to have recovered, an unrecoverable one says nothing was recovered, neither claims a reconstructed diff, and the parked forms each say which of the two gaps it is, with near-miss shapes of the size marker still rendered as ordinary input. Real input on the stream, the frame or the record keeps edits flowing, and plain approvals and denials are untouched. Since 0.83.2 the live frame's `mandate` word reaches the card only when it is one of the six known words, read back identically by `readApprovalMandate`; it never adds the separate mandated key to the card, but a known word on its own makes `approvalIsMandated` answer true, because the engine only sends the word as the whole reason for that bit. The frame's `origin` is read from its own keys only, so a value inherited through the prototype chain never reaches the card. |
|
|
354
|
+
| `scripts/run-hitl-gate-honesty-test.mjs` | [2393] the four HITL disciplines that a passing type-check cannot see. (1) The park predicate and the `tool_end` predicate must cover the **same** set — the park side admits a first-class `kind:'tool_approval'` gate for *any* tool name, and a `tool_end` frame carries no `kind`, so the frame-level judge falls back to the engine's exact abort marker; otherwise the poison frame hits the transcript and `markEnded` swallows the real replayed result (the #110 disease, reopened on kind-only gates). (2) The already-decided identity criterion is **one-shot**: its two inputs are monotonic, so without consumption one successful decide makes every later park failure — including a real `approvals.list` outage — read as "already resolved" until the 24-hop budget runs out and reports a cause that has nothing to do with what happened. (3) A `plan_review` card dismissed without an answer must be re-presentable: the idempotent re-arm short-circuit re-publishes the still-armed card, and a stale armed id (responder gone) re-arms from scratch rather than presenting a card nobody can answer. (4) `HitlSafetyError` is a safety signal — the `remember` fallback arm must re-raise it instead of auto-retrying the decide, while a plain unknown-key 400 still falls back. (5) The polling leg reschedules after an escaping throw and flips `mode()` to `idle` once it consistently fails, so the honesty surface stops reporting a dead feed as live. (6) The live-frame leg carries the fact behind "you are being asked because the auto-mode classifier could not run" all the way to the card port. Transit narrows on SHAPE only — a non-empty cause string is taken verbatim, an open set, because the word table's owner is the engine and re-checking a closed table at the package boundary would drop a legal value the day a new cause word appears, which is exactly the information worth keeping. A malformed carrier degrades to absence rather than half-minting, and absence stays absence: it covers "the classifier answered", "this ask never qualified" and "this deployment has no classifier" at once, so nothing may render it as reassurance. The guard also pins the division of labour that makes the open set safe — the same word that transits is judged again by the public display reader, which narrows to the availability axis, so a word the engine says it never stamps on this fact renders no sentence while still being visible on the card for triage. A later section pins the split this release introduced on the deny close-out frame. Until now every denied tool call was stamped with the same sentence — the one that says *the user* does not want to proceed — including the calls denied automatically on a lane that has no approval surface at all, where nobody was ever asked. The guard drives all three shapes (a person pressed No, a rule settled it, nobody said which) through both close-out arms and the durable park leg, and pins that the third shape is byte-identical to the previous release: an attribution nobody supplied is not evidence for either answer. The rule-settled shape carries the shell's own reason on a second line when there is one and stands alone when there is not, because a blank line where a reason should be reads worse than no line at all. The attribution is read from own data properties only, so neither a polluted prototype nor a getter can make an automatic denial claim a person made it — and the getter case is pinned to never run at all. The transcript classification word is minted only on the two paths where the upstream transcript format really carries one; the three classifier words and the two abort words are left absent, with the abort words pinned against the strings this package actually normalises interruptions to, which are different strings. A final section pins the decide-operation observer: one `start` in the same tick as the first request and exactly one `end` after the last attempt has settled, across success, retried timeouts, exhausted transient failures, semantic refusal, binding mismatch and both kinds of caller abort, with nothing between retries; an observer that throws or rejects — even when logging that fault fails — never changes what is sent or returned, and the stream-level dependency reaches both durable park legs. A package-internal re-delivery of the same decision (plain approve after an older server rejects the session-scope flag, or a re-send without the attribution key) is reported as one operation with a single start and end. An operation handle only groups sends for the same session and bound call — a send for another gate through the same handle is its own operation — and closing a handle while a send is still in flight defers the end until that send settles. When a person picks "allow for this session" on a parked card and the grant is known not to have been stored — the server answers so, which newer servers do for the gates they can recognise from the parked row as needing a person each time, or the server refuses the session-wide grant with one of the refusal codes that are fixed by the row or the deployment and the package falls back to a plain approval — the parked path now says so with the same line the live path uses, and the receipt carries the server's bit for hosts that call that path directly; an answer without the bit, or any other failure — including a conflict that an internal retry can hit after an earlier send already stored the grant — is treated as unknown and says nothing, a plain approval that never asked for the grant says nothing, and a host logger that throws after a successful decision, on either card path or in the bridge's retry step, can no longer turn it into a second send or a failure. From 0.83.0 it also pins approval cards whose arguments are not the tool's real input: a live frame whose arguments were omitted over the size cap or never sent (the card holds at most a path recovered from the question), and a parked approval whose stored input is missing or replaced by the upstream size marker. On every such card an edited approval sends nothing at all — no respond, no decide, no session-grant attempt — ends as edit-refused, raises the same notice once and flags the card so hosts do not offer editing; both production entry points are driven end to end, and the parked path neither mistakes it for an already-decided gate nor reconnects. The explanatory line is pinned word for word in its four forms: a recovered path is the only thing it claims to have recovered, an unrecoverable one says nothing was recovered, neither claims a reconstructed diff, and the parked forms each say which of the two gaps it is, with near-miss shapes of the size marker still rendered as ordinary input. Real input on the stream, the frame or the record keeps edits flowing, and plain approvals and denials are untouched. Since 0.83.2 the live frame's `mandate` word reaches the card only when it is one of the six known words, read back identically by `readApprovalMandate`; it never adds the separate mandated key to the card, but a known word on its own makes `approvalIsMandated` answer true, because the engine only sends the word as the whole reason for that bit. The frame's `origin` is read from its own keys only, so a value inherited through the prototype chain never reaches the card. |
|
|
355
355
|
| `scripts/run-park-hop-progress-test.mjs` | L-80: the park re-attach loop budgets **stalled** rounds, not parks. A turn where the model keeps hitting gates and every one of them is really decided (a card was answered, the engine really moved on) must never be cut off by the hop budget — the budget counts consecutive rounds that produced no progress, and "the engine revived and immediately parked again on the same coordinates" is not progress. The three non-progress arms (already-resolved, decide-transport-exhausted, and a re-scan that was adopted but led nowhere) share one same-cause limit instead of one arm having a limit and the others having none, and every non-progress re-attach is announced once through the host callback rather than only to the debug log. When the limit is spent the resolver reads the approval queue once more and puts whatever is decidable in front of the user before it gives up; only when there is genuinely nothing to show does it fail soft, and the terminal message then carries the real cause and a real way out instead of a sentence about a budget. On the self-heal side, a reopen verdict that reports `decidedWithoutCard` — the chain settled the gate by rule, so there was no card to present — is progress, not a reopen failure, and the user is not told their message was NOT sent. Negative control: a genuinely empty queue with a run that never moves still fails soft |
|
|
356
356
|
| `scripts/run-notif-fleet-honesty-test.mjs` | [2393] the five notification/fleet disciplines a green type-check cannot see, each proven by reverting the fix. (1) The workflow-side dedup `return` keeps a count and a trace — without it "suppressed by design" and "a real completion swallowed because the runId minting changed" are the same observation. (2) `seq` normalisation has exactly one mint point, so a 0-based or fractional wire `seq` cannot make the watcher lane and the frame lane key the same completion differently (which would feed the model twice). (3) The TTL sweep defers to a probe arm that is still inside its own deadline — an entry recorded as "abandoned" must not be delivered a moment later — while an arm that has outlived its deadline never blocks the sweep, so the headless exit gate keeps its liveness. (4) The reset hook really clears every ledger it claims to (the sticky `prompt` ledger leaked across cases). (5) The fleet ledger counts all three drop paths (malformed / unknown frame type / isolation drop), and the panel projection's settled recycling is anchored on the settle instant and skips still-present rows, so the dedup token is never carried off with the entry (which would re-emit `end`) |
|
|
357
357
|
| `scripts/run-public-surface-test.mjs` | The outward promises: the npm export surface baseline (an **exact set**, both directions — a new export that never entered the baseline is one nobody watched leave, and deleting it later would not be red), the peer floor witness, and this README's claims |
|
|
@@ -362,7 +362,7 @@ guard still cross-checks the table by name).
|
|
|
362
362
|
| `scripts/run-client-core-singleton-test.mjs` | Module-level singletons ⇄ `docs/refactor/p1-scan/singleton-manifest.json`, **both directions**: an unregistered singleton is red (registering it forces someone to answer "what if this got duplicated"), a stale entry is red, and the `dupRisk: high` count only goes down |
|
|
363
363
|
| `scripts/run-catalog-loader-gates-test.mjs` | The model-catalog candidate chain (`loadCatalogWithSources`) and the provider device-code seam: offline ⇒ `bundled` with an honest `online.reason`, a good source ⇒ `online` plus a cache write, a second offline run ⇒ `cacheHit`; the three hostile source shapes (malformed JSON, `schemaVersion: 99`, off-domain `http`) each fall through to the bundled table, and an off-allowlist target is **never dialled** — including a `302` to another host, proven by a real loopback server's hit counter staying at zero; a one-byte edit to `catalog.sha256` drops that source while an unavailable sidecar only warns; and the device-code poller's `pending → ok` / `expired` arms run against a real loopback HTTP server with an injected clock |
|
|
364
364
|
| `scripts/run-abortable-sleep-test.mjs` | The shared `abortableSleep(ms, signal)` leaf (consumed by `workflowClient.ts` and `agentSession/backgroundView.ts`'s poll backoff): normal timeout resolution, immediate wake-up on `abort` mid-wait, `clearTimeout` really firing on that path, and a post-resolve late abort staying a no-op |
|
|
365
|
-
| `scripts/run-durable-card-display-keys-test.mjs` | The durable approval row's two display keys survive the row→card recast in `surfaceFsApprovalAndDecide`: `governanceForced` stamps on strict `true` only (absence is "no evidence", never `false`), `ruleSuggestions`
|
|
365
|
+
| `scripts/run-durable-card-display-keys-test.mjs` | The durable approval row's two display keys survive the row→card recast in `surfaceFsApprovalAndDecide`: `governanceForced` stamps on strict `true` only (absence is "no evidence", never `false`), the row's rule offers (`ruleOffers`, or the older `ruleSuggestions` key that earlier servers send) pass through the same shape-narrowing reader as the live-frame leg and land on the **read-only** card key `ruleOffersReadOnly` — plus a standing pin that the durable leg never stamps the redeemable `ruleOffers` card position (the `/decide` body has no rule slot; offering a "don't ask again" option there would be an affordance nothing can honour), and a section for the parked twin of the classifier-unavailable fact: the upstream declares that key on the parked action itself, verbatim and under the same name as the synchronous ask, so this leg reads it rather than guessing a carrier name the way the deliberately unprojected keys must. The guard drives both legs with the same cause and asserts the card ends up byte-identical either way — the observable consequence of one reader serving two key paths, and the thing that silently diverges the day someone writes a second copy. Its own reach is printed rather than implied: what is proven is the package-boundary promise "on the row ⇒ on the card", not that today's engine flattens that key onto the pending row. A further section covers the two display facts the recast had been dropping for far longer. One of them the row has carried all along under a DIFFERENT NAME than the live frame uses — the frame puts it at the top level, the row nests it under the risk descriptor — and that difference in name is exactly why it went unnoticed; unlike the keys this leg deliberately refuses to project, its carrier is witnessed in the engine's own artefact rather than guessed. Neither is decoration: the shell's stand-aside arm reads them, so a call that matched a remembered allow rule which could NOT silence it looked like an ordinary ask on the durable path and was auto-approved with no card at all. Both land on the SAME card slot the live leg uses (one shape for the ends), verbatim bytes, present only when non-blank, never folded into an empty string — and the guard pins the discipline in both directions, including that a top-level key the upstream row does not actually have must still not grow this position. The security-class approval bit (`requiresRealApproval`) rides a parked row's card when the row carries it at top level, on strict `true` only, while look-alike nested carriers are ignored; this is pinned with a constructed row, because today's pending list does not carry the bit yet. A second, separate bit (`irreversibleParkGate`) marks a parked card whose row sits on the irreversible-ask gate kind — a gate-kind fact that covers asks the engine flagged for real approval at the first decision plus safety-tightened gates, with a known engine gap for approval demands raised only on a storage recheck — on the park path only, on the exact gate word only, and never in place of the real bit. Since 0.83.2 the recast also carries the row's ask origin (`origin`) exactly as the live-frame leg does — a non-empty string, verbatim, open vocabulary, never invented or defaulted — and the word a mandated question stands on (`mandate`), accepted only when it is one of the six known words and read back through `readApprovalMandate`; a word outside that set, or a malformed value, leaves the card without it. The word never adds the separate mandated key to the card, but a known word on its own makes `approvalIsMandated` answer true, because the engine only sends the word as the whole reason for that bit. An `origin` inherited through the row's prototype chain is not read, and since 0.83.4 neither is an inherited `mandated` on either leg — both legs stamp it from an own strict `true`, the same reading the seat crossing uses; the single judge `approvalIsMandated` reads `mandated` and `ruleOffersAbsence` the same way, so an inherited key no longer makes it answer true. |
|
|
366
366
|
| `scripts/run-session-memory-status-test.mjs` | The session **memory-status** read face (S-53): the two judgements three clients would otherwise each get wrong. First, *same status, different code* — this route's 404 carries two unrelated meanings (`not_found.session` = unknown or non-owned session; `not_found.route` = a pre-7.53 server that has no such route at all), so dispatching on the **status** would report "your deployment lacks this surface" as "your session does not exist". The verdict is anchored on `errorCode`, the two 404s are pinned to **different** verdicts, and — the load-bearing negative control — a 404 carrying **no** code falls to `failed` rather than guessing either way, since a wrong guess in either direction is a false statement a user would act on. 501 is allowed a codeless fallback because both of its arms mean the same thing here, and `capability.*` stays split from `feature.*` because those two share a status while their dispositions are opposite. Second, *absence means something different per key*: `optOutSource` and `lastCaptureAt` are legitimately absent on a **healthy** session (a zero-history session really is `{captureOptedOut:false, committedCount:0, foldedCount:0}` with no degradation at all), so reading absence as "off/none/0" asserts something unprovable. Two combined readers are pinned: capture opt-out is read from **both** its keys (a record-store fault yields `indeterminate`, never `active` — the difference between "your conversation is being remembered" and "nobody knows"), and last-capture is a **three-state** read whose discriminator is the *other* key, because `lastCaptureAt`'s absence alone covers both "ledger unreadable" and "genuinely no contributions" and therefore decides nothing; the two shapes are pinned to different verdicts so a single-key read turns red. The thin wrapper is the only IO: it never throws, drops malformed keys to absence rather than trusting them (an unreadable value must answer "don't know", never render as truth), refuses to spend a request on an empty `sessionId`, and passes `signal` through untouched |
|
|
367
367
|
| `scripts/run-crash-converged-projection-test.mjs` | The `crashConverged` read face on `GET /v1/approvals` (L-38): what the *previous life* of a crashed local engine left behind, projected for every client. Three judgements are pinned. First, **absence is not an empty list** — a missing key (an older server, deps not present, or a carrier that is not an array at all) returns `undefined`, and the client renders nothing; an empty array returns a present zero-count object, which is the server actually saying "none". Folding the first into `{total:0}` would have the client assert "nothing was left behind" on a surface a person uses to decide whether it is safe to re-run something — the worst possible direction for a false statement — so the two cases are pinned to different **return shapes** and a test asserts the two verdicts are unequal. Second, bucketing is a **four-term conjunction**: `orphanState === 'pending'` *and* `resumeSafe === true` *and* both approval-evidence keys (`originalDecision`, `decidedAtMs`) absent. A fifth term rejects any row carrying an **accessor**, and accessors are never invoked at all — reading one means synchronously running someone else's code, and `catch` catches throwing, not *never returning*, so a looping getter would pin the startup thread forever (the row cap does nothing against that shape). The same rule covers the three untrusted reads outside the row as well — the envelope's `crashConverged` key, the carrier's `length`, and every numeric index are read as own property *descriptors* and only data descriptors are used, so accessors and prototype entries read as absent and are never invoked. Such a key is treated as absent: if it was a required field the row is counted as dropped, if it was optional or additive the row survives without it. That also closes the ordering attack, since spreading runs getters in property order and an earlier one could `delete` the approval evidence before it is ever copied (measured before the fix: such a row reached the resume-safe bucket), and the check therefore moves ahead of the read, onto the property descriptors — from which the snapshot is then built directly, because checking descriptors and *then* spreading is two independent observations of the same row, and a non-throwing proxy can make the two `ownKeys` calls disagree (first showing `originalDecision: 'approve'` so the row reads as plain data, then omitting that configurable key so the snapshot loses the evidence; measured before the fix: the dangerous row reached the resume-safe bucket after exactly two enumerations, and after it, one). Keys are written with `Object.defineProperty` rather than plain assignment, because `'__proto__'` is a legal own enumerable key and `o['__proto__'] = x` does not store a value — it calls the prototype setter, letting a row whose own properties are all plain data (so the accessor gate never fires) inject a prototype whose `sessionId` getter deletes the approval evidence from the snapshot during validation; `defineProperty` fires no setter, so the key survives as ordinary additive data and the snapshot keeps `Object.prototype`. A row that simply arrives with a custom prototype is treated the same way, since the snapshot only enumerates own properties: approval evidence sitting on the prototype would never reach it, and a perfectly ordinary object with no proxy and no accessors could otherwise be called safe to re-run — real bodies come from `JSON.parse` and always carry `Object.prototype`, so nothing genuine trips it). Validation itself runs on a **null-prototype** dictionary and the bucketing verdict is carried out of that same pass rather than re-read from the delivered row, because every property lookup on an ordinary `{}` reaches `Object.prototype`: a polluted `sessionId` getter there would delete the approval evidence from the snapshot mid-validation and send the row to the safe bucket (measured before the fix). The row handed to the client is still an ordinary object — the null prototype is an implementation detail of the check, not of the value) — real JSON bodies are all data properties, so only a middle-layer-synthesised payload ever trips it, and it too lands in the human bucket rather than being dropped. The `decided` arm means the human had already approved and side effects may be half-landed, so it always goes to the human bucket, as does `resumeSafe === false` and — the last two terms — any row whose own fields contradict each other, since `pending` claims nothing ran while that evidence says somebody pressed approve. Deciding "not safe" costs one extra question (recoverable); deciding "safe" wrongly has somebody re-run work that already partly happened (not). A 2x2 truth table pins that exactly one cell is resume-safe, so reading either key alone turns red, and the contradictory rows are routed to the human bucket rather than dropped — they are real orphans, and the ones most worth showing. Third, unreadable rows are **dropped and counted**, never thrown and never passed through: the product is declared as `CrashConvergedRow`, so letting a row missing a required field — or carrying one of the wrong type — past would be a lie at the type level, and the closed literal discriminators (`decision` / `cause` / `orphanState`) decide family membership rather than being an open vocabulary. The measuring stick stops at the **type** floor, though: degenerate-but-well-typed values (`ts: NaN`, an empty `toolName`) are kept, because swallowing a real orphan over a decorative field is the worse direction, and the one deliberate exception is `approvalId`, which must be non-empty to be a row identity at all. `dropped` is kept separate from `total` so unreadable rows never inflate "N approvals were affected"; each row is a **one-shot snapshot** — every own enumerable key is read exactly once, and validation, bucketing and the handed-back value all read that same snapshot, so additive upstream keys survive while a **non-idempotent** getter (one that never throws, just answers differently on a second read) can no longer erase the approval evidence between the check and the bucketing (measured before the fix: such a row landed in the resume-safe bucket while its checked value was `"approve"`). Hostile carriers are counted rather than allowed to reject: **every** touch of the carrier is guarded — envelope property reads, `Array.isArray` itself (it throws on a revoked proxy), the `length` read, each indexed read and each row's property reads — and a traversal that dies halfway returns absence rather than a half-counted total. A row that cannot be read never takes the batch with it: its own shape check is inside its own guard, so one revoked-proxy row costs a `dropped` tick rather than collapsing the whole projection to absence — which a client would have read as "this deployment does not offer the surface". Traversal goes by **numeric index, never the carrier's own iterator protocol**, because `for...of` hands the carrier the question of which rows exist: an array carrying an overridden `Symbol.iterator` can yield nothing (measured before the fix: a real orphan became `{total:0}`, which a client reads as "the server said there are none") or swap a dangerous `decided` row for a safe-looking one (measured: `fake-safe` was returned in place of `real-danger`). Row count is capped at 100000 and the cap is checked **before** the walk: requiring only a non-negative integer `length` does not stop a proxy trap reporting a billion, and this surface runs on the startup / `--resume` path, where a synchronous spin freezes the thread (measured before the cap: twenty million rows took 18.3 seconds and twenty million index reads; a billion does not come back). The honest boundary is stated rather than overclaimed — a proxy can still lie in its `length` or index traps, which is the same thing as a host injecting a lying transport — and the widening of `ApprovalsResourceLike.list()` is proven **additive** by really running tsc over a legacy `{ pending }` mock *and* over the real `AgentClient` path — the projector takes `unknown` precisely because a parameter shaped as "an object with an optional `crashConverged`" is a TypeScript weak type that the installed SDK's own `list()` return shape shares no property with, which only a real-client compile would have caught — with a known-red control so a clean run means the checker spoke |
|
|
368
368
|
| `scripts/run-self-orchestration-denial-test.mjs` | The three judgements behind a **denied self-orchestration request** (server 7.57.0), each of which all three clients would otherwise get wrong on their own. First, whether to retry at all is a **conjunction that may not be loosened**: HTTP 501 *and* an `errorCode` that is **exactly** `capability.self_orchestration_required`. That code shares its shape with every other `capability.*` 501, so dispatching on the prefix would drag "some other capability is not wired up" into the retry arm — those requests do not become acceptable once the two keys are gone, so the client would spend a request and then tell the user the wrong reason. Negative controls cover all four directions: a sibling `capability.*` code, a truncated or suffixed variant of the right one, a codeless 501 (it decides nothing, so it decides nothing — no guessing), and the right code under 500 / 400 / 503 or a string `"501"`. The classifier reads structurally rather than by `instanceof` (a host may inject its own transport; across realms or duplicate SDK instances an understandable error would read as unreadable), so a class instance, a bare `{status, errorCode}` literal and an error carrying those fields on its **prototype** all reach the same verdict — and a hostile proxy or a throwing getter yields `null` instead of throwing, because this classifier runs inside a `catch` block where anything it throws escapes the caller's own guard. Second, removing the intent is a **structural** operation, not wording: `selfOrchestration` sits at the top level while `ultracode` sits under `settings` — two different stamping legs — and a client hand-writing `delete` will miss the second one, which costs the user the same failure twice. The single stripper is pinned to touch exactly those two: other `settings` sub-keys and their values survive byte for byte, `deferTools` is left alone (pulling `Workflow` out would be a behaviour change, not a removal of intent), additive unknown keys survive at both levels, the input object is never mutated, `settings` is only dropped entirely when `ultracode` was really there and nothing else remains (an already-empty one is left as is), a non-object `settings` is not touched at all, an `ultracode` that only exists on the prototype does not count, and the whole thing is idempotent. The end-to-end leg runs a real `buildTaskRequest` product through it and asserts the stripped body still passes the registration gate key by key. Third, on the capabilities body, **absence is not "switched off"**: a pre-7.57 server has no `workflowsGate` key at all, so reading absence as "the engine says no" asserts something the server never said, and the mirror-image disease is folding an **unrecognised** `denial` into `null`, which would have the client render "nothing was denied" when the truth is "denied, for a reason I do not recognise". Five shapes are pinned — caps unreadable, gate absent, closed-set member, unknown value, accessor — with the unknown arm carrying the raw token (or an empty one when the value is not even a string) and never collapsing to `null`. All four untrusted reads go through own **data descriptors** only, and the guard pins the getter invocation count at zero, since `catch` catches throwing but not *never returning*; a descriptor trap that throws and a revoked proxy both yield honest absence rather than an exception — though *what* absence means differs by field, and the guard pins that split rather than a blanket rule: an accessor on `workflows`, `workflowsGate` or `engineCan` reads as absent, while an accessor on `denial` reads as `{unknown:''}`, because a key that is **not there** is the gate saying "nothing was denied" whereas a key that is there but cannot be read is "denied, and I could not read why" — folding the second into the first is exactly the false statement this face exists to prevent. Two further pins came out of an adversarial review. The exported retry list is **frozen at runtime**, not merely `as const`: the verdict hands out that same reference, so any consumer splicing it once would poison every later verdict in the process — the guard asserts `Object.isFrozen`, that four different mutation attempts leave it byte-identical, and that a verdict issued *after* those attempts still carries the original two entries. And the classifier reads `denial` only **after** both criteria have passed, since it is not a criterion but an extra field on the verdict: the guard pins the getter invocation count at zero for any error that does not match and at most one for an error that does. The scope line is drawn explicitly rather than overclaimed — "no getter ever runs" holds for `projectWorkflowsGate`, which reads **wire JSON** where every field is an own data property by definition, but not for the classifier, which reads a **thrown value** that may well be an SDK `APIError` class instance carrying `status` and `errorCode` on its prototype; insisting on own data descriptors there would report a perfectly readable error as unreadable, so that side promises only that it never throws. A final pin covers the **integration document's own worked example** rather than the library: the shipped SDK's `tasks.stream()` is an `async` generator, so calling it issues no request at all — the POST happens inside `streamRaw` on the first iteration, and a `try` wrapped around the `stream(...)` call itself can never catch the 501. A client following a submit-shaped recipe on the streaming leg would never run the classifier, and the whole strip-and-retry path would silently do nothing. The guard drives the **real** `TasksResource` against a fake transport, offline, and pins both halves: the synchronous leg is in flight the moment it is called, the streaming leg has issued zero requests after the call and raises on the first `next()` — and it does so through the **real** error path, with `openStream` returning an actual 501 `Response` that the SDK's own `errorFromResponse` turns into the typed error, pinning the `openStream`→`errorFrom` call order so a transport that stops minting `errorCode` cannot pass. The documented recipe is then **executed** rather than keyword-counted: exactly one retry, a second body that really lost both keys while every other setting survives byte for byte, the caller's own request object left untouched, one disclosure and only one, a second 501 propagating with the request count still at two, and — after the first 501 — an abort leaving the count at one with nothing disclosed. A last leg is type-level: `stripSelfOrchestrationIntent` carries an SDK `TaskRequest` overload, because the wide `Record<string, unknown>` form erases the caller's type and the document's "strip and resubmit" line would not compile without an unsafe cast; a real tsc run over a virtual file proves both the narrow and the wide path, with a known-red control — and it compiles the document's two recipes **verbatim**, extracted from the section itself, because a recipe that does not compile is a recipe that was never given: `{ transientOk: true, signal }` is a TS2379 under `exactOptionalPropertyTypes`, which no amount of prose review had caught. The last thing pinned is the one that would have been quietest of all: the SDK's `stream()` returns only on a `done` or `failed` frame, so a stream truncated mid-run — or yielding nothing at all — ends the `for await` just as normally as a completed one. The documented `runOnce` therefore tracks whether it ever saw a terminal frame and raises when it did not, the guard's success fixture emits a real terminal and asserts the handler received it, and a truncated-stream control asserts that shape is reported as a failure with no retry and nothing disclosed. That terminal-frame rule then needed one more turn of its own: the underlying reader returns *normally* when the signal is aborted, so the check as first written rewrote a user's cancellation into a generic stream fault — a client keying off `AbortError` to suppress the error would instead have shown a failure, or resubmitted. Cancellation is therefore checked first, a real-SDK case aborts from inside the handler and asserts the original `AbortError` survives with no retry and nothing disclosed, and the document is checked for that ordering. The harness runs the documented `handle` and `transcript.note` as real spies rather than pushing frames itself, the drive loop rethrows exactly as the document does, and the disclosure ledger is proven to be the caller's own array by a positive identity assertion — without which the cancellation leg's "nothing disclosed" would have been vacuously true. Each recipe is compiled **on its own**, with a preamble that declares only what a host supplies and injects no library symbol, since compiling them together let the second one borrow the first one's imports, and the preamble's own types are decoupled from what the recipes import so the "remove the imports and it must fail" control fails for the right reason — which is checked by attribution, not merely by redness. Ordering is the last thing to get right: the cancellation check must come before the truncation error but **both** must sit behind the terminal-frame test, because a cancellation that lands after the run already reported `done` would otherwise overwrite a real outcome — one that may have already had effects — with "cancelled", and a person reading that will run it again. Aborting from inside `handle(done)` and `handle(failed)` are both pinned to still report success, and the ordering assertion is anchored inside the streaming `runOnce` body rather than the section, since the section's first `throwIfAborted` belongs to the synchronous recipe and would have made a reversed streaming recipe pass — and that ordering check is now anchored on the TypeScript AST rather than on text, since a comment reproducing the two statements in the right order let a genuinely reversed body pass. One more timing fact had to be written into the recipe: a single SSE read buffers several frames and the SDK yields them back to back, so checking the signal only after the loop lets a cancelled run keep consuming the rest of the chunk — measured, an abort inside `handle(turn_start)` still swallowed the `done` that followed and reported success. The recipe therefore re-checks after every non-terminal frame. Finally, the behavioural matrix is no longer run against a copy of the recipe: both recipes are extracted from the document, transpiled, and **executed** with injected host objects, so the disclosure assertion really exercises the document's own `transcript.note(disclose(...))` line, and the synchronous leg gets the same full matrix the streaming one does |
|
|
@@ -372,16 +372,16 @@ guard still cross-checks the table by name).
|
|
|
372
372
|
| `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" `listAllPersistedRules` needs only `list` (`RulesListFacade`; the other four methods are known to the parameter type as optional members): a synthetic consumer compiled against the built declarations passes a two-method object, a list-only literal, a `Pick` slice, the named type, the full facade, and the full or two-method facade written as an inline object literal — including arrow functions with untyped parameters — all with zero diagnostics, while a misspelled method name in such a literal is still reported. |
|
|
373
373
|
| `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
|
|
374
374
|
| `scripts/run-selfheal-reopen-test.mjs` | The 409 active-run self-heal decision chain: `governanceForced` narrows on strict `true` only; triage prefers the wire's `pendingGate.kind` and falls back to the status table (an off-table kind is never guessed into a card arm — hands-off plus the honest wording); a first-sight card makes zero closed/reopened claims and a host presentation receipt of `presented: false` demotes the outcome to reopen-failed; park-row ownership is a fail-closed positive proof (own-run ledger or session id — unprovable is not owned); the three gate-identity key literals live in exactly one mint (`hitl/gateIdentity.ts`, AST string-token scan); the armed-gate presentation ledger is per-session; and the `plan_review` reopen arm shares the arm arm's card body, three-state verdict and delivery pipe, consuming the presentation history once a decision is delivered. The same chain also carries the `running` three-way card: both plan-family gate kinds route to the plan arm and all four ask-family kinds to the ask arm (an off-table kind still never gets guessed into either); the card is offered only for verbs that can actually be honoured and a missing presenter means zero action rather than a silent cancel; a steer is sent **exactly once** with its three delivery outcomes worded apart (a `queued` receipt is the wire correcting the triage input, so the named park word decides which card gets reopened, and an unrecognised park word drives neither arm), and a steer failure is split into *provably not delivered* (4xx) and *delivery unknown*, because telling a user to resend a non-idempotent instruction that may already have landed is how duplicates get made. After a user-chosen cancel, "the session is free" is asserted only from a whitelist of terminal states — park states hold the claim, an unrecognised state word is not a release, a failed read is *unknown* rather than a release, and only a 404 counts as one — and the honest timeout line quotes how long it really waited. The two "card could not be reopened" rows can carry a host-declared way to keep the conversation, which says the card comes back on resume only if it is still waiting: it is placed before the route that abandons it, never offered for an injected submission, while a decision is still on its way, or once the pending approval has been proven gone (the outcome then carries a flag saying so; the proof only counts before the cleanup card is shown, so a fallback after the card carries no flag unless a fresh read finds the run finished, and a recheck that finds the approval back clears it), the host function is not even called in those cases, and it is treated as unavailable when it throws or returns an empty value; a host can also switch off the engine decide route on the interactive rows while the cancel route stays, and with neither given all four rows are pinned byte-for-byte to the text the previous release produced. When the plan reopen refuses because the user closed that card (`dismissedByUser`), the outcome carries the flag and one package-owned line says so — you closed it, a message brings it back — in the typed, the injected and the queued-steer forms, with no way-to-keep and no engine decide route; the reopen port receives `{ trigger }` mapped from the submission origin (typed ⇒ `user`, injected ⇒ `automatic`, absent ⇒ the old single-argument call). |
|
|
375
|
-
| `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `
|
|
375
|
+
| `scripts/run-terminal-identity-copy-test.mjs` | Terminal-state **identity**, in both lanes where a stop gets a name. A run stopped by this deployment's own governance knobs — the open-set `limits.*` family, `output.invalid`, and the `blocked` contract terminal a ReportBlocked agent produces — is not a provider failure, and labelling it `API Error:` sends the reader to check the network, the key and the quota when the handle is the `--max-turns` they passed themselves. Those terminals now render a neutral row; the reverse direction is guarded just as hard, because asserting "this is *not* an API error" on a code the package does not recognise is the same misfiling pointed the other way — a real `gateway HTTP 502`, a `conflict.session_active_run` and any unknown code all keep the `API Error:` prefix, and the row keeps its `_sema_api_error_message` class flag (spelled `isApiErrorMessage` before 0.83.0) so brief-mode visibility filtering does not silently drop it. The second half is who the rejected submission belonged to: the self-heal copy told every caller "Your message was NOT sent … send it again", which is three separate untruths for a system injection (a plan-review outcome, a cron wake-up, a task notification) — not the user's message, and not re-sendable, since a host queue marks those non-editable and non-recallable. The injected form says so instead, and the one sentence that promises re-delivery is pinned to the single disposition that earns it: `selfHealSubmissionDisposition` is the same function the host consults before putting the item back on its queue, so the promise and the behaviour cannot drift apart, and the arms where no card could be surfaced state plainly that nothing was delivered and nothing will retry. Since 0.72.6 the same gate pins the **follow intent** after a steer (): a message handed to a live run only pays off if someone tails that run's own event stream, so `steerFollowIntent` decides from the delivery word whether to tail now, after the pending decision, or only after a wake — and the "watch that run" sentence ("watch that reply" on the rows about follow-up messages sema sent on its own) turns into a factual "sema is following that run" ("… that reply") **only** when the host declares it attached that tail, so a shell that did not wire it can never claim it did. Since 0.83.0 the rows about follow-up messages sema sent on its own carry no engine-internal words and say what happened per delivery shape, promising a resend or "nothing for you to do" only where the code guarantees it |
|
|
376
376
|
| `scripts/run-additive-key-passthrough-test.mjs` | The one disease shape behind two legs: a **closed whitelist / flattening arm** dropping a fact that is already on the wire, while both sides of the seam look correct. (1) The `task_progress` projection carries a registered **key ledger** — a frame populated with every key the service really projects is pushed through the shipped `eventToSdkMessage`, and the set of wire keys that survive must equal the registered pass-through list **name for name in both directions**, so quietly forwarding one more key is as red as quietly dropping one. `model` (the child run's model id, minted by core as `prepared.model.id` and projected by the server since 7.52.1) is the key this batch adds, with the same conditional the server itself applies: a non-empty string or no key at all — an empty string is neither a model id nor "unknown". The ledger is also checked against the fenced list in `docs/INTEGRATION-CLIENTS.md` §3d, so a doc that still says seven keys while the code forwards eight is red rather than merely stale. (2) The decide-failure arms carry the server's S-02 `currentPending` pointer key from a 409 `approval_stale` refusal onto the outcome the host reads. The reader is structural rather than `instanceof`, because the client is host-injected and the class identity is not this package's to assume; a half triple never mints (half a pointer cannot relocate anything), an empty string is not presence, and `checkpointToken` never transits. Both the allow and the deny leg are driven end to end through the real durable approval path — as is the accept-session leg, where a refusal carrying the pointer key must now re-raise instead of silently re-sending the human's answer for the **old** card as a plain approve (one decide call, pointer preserved), while a legacy 400 still falls back exactly as before — and all three flattening points must call the one shared reader — the same-shape residue check that makes "fixed one arm and left the twin" red instead of invisible. (3) The same disease growing on the REQUEST side: the `.mcp.json` → server-spec projection rebuilds each server key by key, and the settings schema deliberately leaves some keys parse-transparent — whatever JSON the file carries reaches the engine untouched, because validating them where the whole domain parses all-or-nothing would let one bad declaration take every server down silently. The whitelist had no row for the newest of them, so an operator's per-tool declarations — the ones the write fence reads — were stripped at the package boundary while both sides looked correct. The criterion is not "is that key handled" but the transparent-key table read out of the INSTALLED schema at runtime, reconciled name-for-name against this leg's ledger, so the day upstream adds a third one this turns red and forces an explicit decision. Behaviour is pinned on both transports, by object identity rather than deep equality (a rebuild would be a second judge), and malformed values must transit UNCHANGED rather than be refused here — the engine refuses them loudly and names the server, whereas a package-side judge can only swallow a declared protection quietly. Absence still mints no key, unknown keys still never reach the wire (the fix is the dropped key, not the gate), and the one transparent key this leg deliberately does not forward is a ledger entry with its own exit condition: it belongs to the deployment plane, and the day the request-plane type declares it the entry's premise is gone and the gate says so |
|
|
377
377
|
| `scripts/run-esc-halt-plan-test.mjs` | The Esc stop decision every client shares: fire the **turn-level** halt first, and escalate to a **run-level** cancel in exactly two cases — the engine itself answered with a 409 from the closed code set (it is saying "there is no in-flight turn here; use cancel for a run-level stop"), or that shot came back with no verdict at all *and* the shell can independently prove a permission card was on screen. Everything else does not escalate. The asymmetry is the whole point and every negative control guards the same direction — deciding *not* to escalate costs the user one more choice on a busy-session card (recoverable), deciding to escalate wrongly tears down a run that was alive and takes every in-flight tool with it (not). So: the closed code set is a **frozen** value, not a `ReadonlySet` — type-level immutability does not stop a consumer's `.add()`, and the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift; the escalation gate is the **conjunction** of that closed set and the 409 status, since honouring the code alone lets a 500 that merely quotes it drive a destructive call; `interrupt.not_held` and `steering.not_running` are deliberately outside the set (the first means *this replica* has no live face — the run may be perfectly alive on another); an unreadable code falls to the no-escalation side; a `parked` flag never overrides a verdict the engine did give, and only strict `true` counts when it did not. The first shot is unconditional by construction — it does not consult `parked`, because the 409 it earns is exactly the verdict the gate wants — and the verdict itself is a closed machine-readable reason word, not display copy. A third escalating case was added once tearing the stream stopped reaping the run: with detach armed, a shot that never lands leaves the run going all the way to the end of the turn, so the Esc the user pressed has no effect at all and nothing on screen says so — the old behaviour had a silent backstop (tearing the stream ended the run) and that backstop is gone. The new fact is held to the same three disciplines as `parked`: it is read only where the engine gave no verdict, it is judged **after** `parked` so an existing host's reason word does not change under it, and only strict `true` counts. Absence is proven to be a no-op rather than asserted — the guard carries its own reference implementation of the previous version's table, runs the full grid through both, requires zero divergence when the new field is omitted, and first shows the comparison really does report a difference on the one cell where the two versions are meant to differ |
|
|
378
378
|
| `scripts/run-peer-frame-projection-test.mjs` | The three engine-injected lanes design/385 puts on the **one** `task_notification` carrier, which are not the same kind of thing at all: a delegated child's uplink (`agentMessage`), another session's message drained from this session's own box (`crossSessionMessage`), and a receipt about one of *this* session's own outbound messages (`crossSessionNotice`). The engine renders none of them inside a `<task-notification>` shell, so a client that projects them as the generic completion card shows "background task finished" while the model read a colleague's sentence — two faces describing different events. The discriminator is pinned to the **typed carrier being present**, never to the `summary` text: those carriers can only be minted by the engine's injection legs (the external `notify()` input is a strict subset of the payload and can wear none of them), while `summary` is filled by every notification there is — so anchoring on text would let any background task impersonate a colleague's message by writing `<agent-message from="…">` into its own summary, and a positive control asserts exactly that payload still projects as the generic card. Fail-closed has two tiers rather than one: a broken **required** field (empty `from`, a non-string `body`, a notice `kind` outside the closed set) returns absence so the caller falls back to the generic card — an honest downgrade where the user still sees the notification — while a broken **optional** field drops only itself, because losing an attribution note and losing a colleague's whole message are not the same magnitude. The provenance side record is **required and must agree on four points** (`kind` matches the lane; `from`/`taskId`/`seq` are present and equal the carrier/payload — each equality is anchored on a core mint site and pinned by the cli wire-anchor A-K24), so a carrier signed with a trusted name but a disagreeing provenance falls back to the generic card; peer bodies pass the same authority-envelope neutralization core applies (`<task-notification>` etc. are defused) so a colleague's text can never seed the resume dedup ledger. Lane precedence copies the engine renderer's own order, because the model already read the frame in that order and a client ordering of its own would put a card on screen that disagrees with the frame the model saw. Rendering and parsing of the transcript line live in the same module and are round-tripped in both directions, including a body carrying a forged closing tag (a parser fooled there hands half a message to the next row) and a quote inside the sender label (which must not forge a second attribute); the notice lane is deliberately kept **out** of the parser, since recognising it would mean anchoring the `[Cross-session …]` prefix and a user typing that same line would be rendered as engine speech. Hostile carriers are read as own **data** descriptors only and accessors are never invoked at all — `catch` catches throwing, not never returning — proven by a counting getter that must stay at zero calls, alongside a revoked proxy and a prototype-only carrier; and four legacy payload shapes assert the no-carrier path is byte-identical to before, which is the executable form of "zero difference for an older host". A re-supplied cross-session message — same task id, status and sequence as the first delivery, handed to the model again after compaction — is not rendered a second time, because the first delivery is still on the user's screen; a record with the next sequence number still renders |
|
|
379
379
|
| `scripts/run-wiring-manifest-projection-test.mjs` | The two end-user facts carried on the engine's `wiring_manifest` frame (`modelGate`: which tools this run's model gate removed and the verbatim restore hint; `autoMode`: whether auto mode is actually armed and the engine's own reason word). Projection: both sections ride as `_sema_`-prefixed superset keys, verbatim, and no SDK-named key is minted; a frame where neither section is well-formed projects to `none/not_in_slice` (no empty arm); `modelGate` needs all three keys and treats `removed: []` as a bad value rather than a reading; `autoMode` needs a boolean plus a non-empty reason that agrees with it, and the reason word is never mapped onto the capabilities vocabulary; the frame is flat (a nested `manifest:{}` wrapper is not a supply); `eventId` rides like every other arm. Adapter: exactly one chrome event on the main lane, a sub-flow frame (any `parentToolCallId`, `null` included) yields nothing, and an absent `eventId` leaves the key absent. Added at receiving time because the shell-side gate could not see this package's behaviour: two mutations (empty `removed` accepted, sub-flow gate removed) had passed the package suite untouched 0.71.0 adds sections F–I: the fourth/fifth/sixth manifest sections (`tools` via the roster reader, `hooks[]` rows dropped one by one when malformed, `lsp` absent unless `mounted` is a boolean), the `tool_roster_delta` arm (narrowed `delta`, `malformed` when `fromDigest`/`roster` cannot be read, host applies it against its own digest), the `context_usage` arm (finite-gated scalars plus `sections[]` rows dropped one by one), and the `WiringManifestMcpEntryView` rename with `MAX_AGENT_SKILLS` gone from the surface |
|
|
380
380
|
| `scripts/run-submit-wiring-manifest-test.mjs` | The non-streaming submit receipt can carry the run's opening wiring manifest (`TaskResult.wiringManifest`, additive on newer servers). `readSubmitWiringManifest` answers one of three: the key is absent on the receipt itself (older server, or a deployment whose engine never produced that frame) — not the same as unreadable; the key is present but cannot be read (not an object, or none of the nine sections survive); or a manifest view. The view is the same shape the streaming lane's chrome event carries (minus its two envelope keys) and is assembled by the same code path, so both lanes agree byte for byte on the same object. Liveness fields ride through untouched — this reader never mints a liveness verdict — and an operator-shaped receipt with extra governance sections reads to the same view as a tenant-shaped one. A zero-tool roster is a real reading, not an absence. |
|
|
381
381
|
| `scripts/run-rule-offers-reader-test.mjs` | The narrowing reader behind the "don't ask again" options, now a public entry point rather than a card-port-only one. Hosts that render the frame themselves (a browser has no three-way terminal card) previously had to rebuild this reader on their side, and what it carries is a **redemption-safety** judgement, not a convenience: the batch arm is redeemed by **index**, so a reader that compacts the array after dropping a malformed entry makes the k-th option a person clicked and the k-th rule the server writes two different rules. So: a bad entry is dropped **on its own** (one bad option must not make a real one disappear) while every surviving entry keeps its **original wire index** — pinned from both ends, with the bad entries leading and trailing. A batch's *members* are the opposite: any malformed member drops the whole batch, because a conjunctive batch is one "yes" to all of them and a batch missing a member is a different grant; its honest-remainder count is a reading, not decoration, so a non-integer or negative value drops the batch rather than rendering a fabricated zero. An empty array, a non-array, an over-cap array and an all-bad array all read as **absence** rather than an empty list, because an empty list renders as "there is an option lane with nothing in it". The two wire generations are ordered by a rule, not a preference: the newer key wins outright, a newer key that is **present but unreadable** does not fall back to the retired key (borrowing the older material would pass someone else's options off as this request's), and a `null` newer key reads as absence so a relaying layer that serialises "missing" as null cannot delete the whole lane on older engines. The public entry is finally reconciled against **both** card-port legs on the same material, byte for byte, so the exported reader and the one the card sees can never become two. Two upstream vocabularies used to be **hand-copied** here, and both had fallen behind: a match word outside the copied pair dropped an otherwise valid option outright, and a batch carrying a directory-read member — a member kind the copy did not know — dropped the whole batch. Both tables now come from one place upstream and are re-exported verbatim, pinned in both directions: every word in the table must be accepted (a narrower copy reds on the words it never learned) and a word constructed to be outside it must still be refused (a reader widened to "any string" reds too), with the retired-key normalising leg sharing the same narrowing so the fix cannot land on one leg only. A member whose kind is genuinely unknown still drops **the whole batch and only that batch** — never one member, because a conjunctive batch one member short renders "yes to N" as "yes to N−1", and never the card, because the honest single beside it is intact — while a member from before the discriminant existed normalises to the historical kind rather than being refused. The additive per-segment reasons ride through verbatim, drop only the row that is malformed, and stay **absent rather than empty** when nothing survives, since an empty list would read as "confirmed nothing uncovered" while the count remains the only source of truth |
|
|
382
|
-
| `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one The engine's folded resume refusal (`resume_blocked_by_policy`) gets its own reading — the original code as named, absent or unreadable — and a wording that neither claims nothing was consumed nor predicts whether a retry would pass. |
|
|
382
|
+
| `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one. The engine's folded resume refusal (`resume_blocked_by_policy`) gets its own reading — the original code as named, absent or unreadable — and a wording that neither claims nothing was consumed nor predicts whether a retry would pass. |
|
|
383
383
|
| `scripts/run-resume-retry-later-test.mjs` | The two resume refusals that carry a **wait quantity** — the only members of that refusal family that do, which is the whole reason they form a closed set. Carrying a wait is not the same as being the only ones worth waiting on: a sibling refusal in the same family clears on its own and the engine says so in words, it just cannot put a number on it, so *not recognised here* must never be read as *waiting will not help*. One of the two also has a *terminal* upstream branch that arrives under the same code with the distinguishing detail only in prose, so recognition alone is not permission to say "try again": the disposition is decided by **positive evidence** and pinned from both directions — the quota-window code is evidence in itself, the preflight code counts only when the server really supplied a wait (an upstream fact, not a convention: the terminal branch throws with no detail at all, so a wait value cannot reach the client on that path), and a preflight refusal with no wait reads as *undecidable* (say what is true of both branches — nothing was consumed — and leave redeemability to the engine's own line) rather than being rendered as either a retry or an ending. Every other member means waiting will not help (change a setting, relaunch, the retained session is gone), so the recognition is a **closed set of two codes**: widening it to a family prefix would tell half the users to wait and the other half to keep waiting for something that will never arrive, and the negative controls drive exactly those codes through it, plus a same-named code on a different door (the submission-side quota refusal), the two underscore-form siblings, and a code merely quoted inside a message body. The wait value is narrowed to the same domain the server mints it in (a whole number of seconds, at least one): zero, a negative, a fraction and a non-number all read as **no window given** rather than as zero, because a zero tells the caller to retry immediately and the wake-up it would retry is an at-most-once action with real side effects. Reading is structural rather than `instanceof`, since the client is host-injected and the same class name across two bundles is two classes, and a null-prototype plain object must still be recognised. The failure classifier gains this one disposition without any existing one moving, an unknown code still falls to the honest open-set arm and its wait value is **not** believed, and an end-to-end call proves the disposition and the window reach the host while the call itself is still attempted exactly once. The recognised code set is a **frozen array**, not a type-level readonly set: the latter is a plain mutable collection at runtime and the decision reads the same instance, so one `.add` from any consumer would turn a refusal that waiting cannot fix into one that claims it can — the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift |
|
|
384
|
-
| `scripts/run-model-capability-probe-test.mjs` | Whether a model on the OpenAI-completions lane **thinks**, and whether that thinking can be **turned off** — a question nobody can answer by looking at a model name, and one whose wrong answer costs every later call. The probe is judgement only: the network half arrives as an injected port, so the package mints no URL, reads no credential and never calls `fetch` — pinned by a source-level assertion, because a package that reaches the network once has changed what every host must trust it with. The seven dialect words are a **copy**, reconciled element-wise against the installed engine’s own bytes in both directions, since the words belong upstream and a private table drifts the day a dialect is added; the settings package deliberately declines to restate them, so the table cannot be imported from there and this guard is what stands in for the import. The **order** the dialects are tried in is a public promise rather than an implementation detail — each extra attempt is real money and real latency against someone’s gateway — so the guard pins the exact call sequence a stub records, and reversing it reds on the wasted round trip; the template-parameter spelling leads because an observed gateway keeps thinking, and answers with an empty body, when handed the top-level switch instead. That observation is also why an empty answer is **not** accepted as *thinking is off*: a knob that deletes the reply is not a knob that disabled reasoning, and accepting it would write a spelling into the catalogue that the gateway does not honour. Two dialect words whose request bytes are identical to another’s do not each burn an attempt. The two verdicts that look alike are held apart from both directions: *tried everything, still thinking* requires at least one attempt to have **cleanly answered**, and when every attempt was refused the verdict is *could not tell* instead — and on the unanswerable path the result carries **no** thinking flag at all rather than a fabricated `false`, while the pure write-back returns the very same entry object untouched. A verdict that reasoning cannot be disabled **removes** a previously declared spelling rather than leaving it, since a refuted spelling keeps the engine sending bytes the gateway ignores while the catalogue still renders it as already off. Evidence is lengths, finish positions and status codes only — a planted secret in both the answer and the reasoning channel must appear nowhere in the result, so the record can go into a log or a ticket whole One cross-package premise is checked by really running the other package’s parser rather than quoting its documentation: everything this probe writes eventually passes through the settings schema on its way into a catalogue, and that field is declared parse-transparent precisely so the vocabulary can live on the consuming side. If it ever narrows, the spelling is stripped **silently** — indistinguishable from the probe never having run — so the guard feeds the probe’s real output through the real parser, checks the compat object comes back key for key, and checks a dialect word this client has never heard of survives too. Two shapes that must be rejected really are rejected, since otherwise the survival checks would hold on a parser that accepts anything, and a bare entry is asserted valid first, because the first run of this section reddened on a space in a fixture’s name — a fixture that cannot pass would disguise the real alarm as already having fired |
|
|
384
|
+
| `scripts/run-model-capability-probe-test.mjs` | Whether a model on the OpenAI-completions lane **thinks**, and whether that thinking can be **turned off** — a question nobody can answer by looking at a model name, and one whose wrong answer costs every later call. The probe is judgement only: the network half arrives as an injected port, so the package mints no URL, reads no credential and never calls `fetch` — pinned by a source-level assertion, because a package that reaches the network once has changed what every host must trust it with. The seven dialect words are a **copy**, reconciled element-wise against the installed engine’s own bytes in both directions, since the words belong upstream and a private table drifts the day a dialect is added; the settings package deliberately declines to restate them, so the table cannot be imported from there and this guard is what stands in for the import. The **order** the dialects are tried in is a public promise rather than an implementation detail — each extra attempt is real money and real latency against someone’s gateway — so the guard pins the exact call sequence a stub records, and reversing it reds on the wasted round trip; the template-parameter spelling leads because an observed gateway keeps thinking, and answers with an empty body, when handed the top-level switch instead. That observation is also why an empty answer is **not** accepted as *thinking is off*: a knob that deletes the reply is not a knob that disabled reasoning, and accepting it would write a spelling into the catalogue that the gateway does not honour. Two dialect words whose request bytes are identical to another’s do not each burn an attempt. The two verdicts that look alike are held apart from both directions: *tried everything, still thinking* requires at least one attempt to have **cleanly answered**, and when every attempt was refused the verdict is *could not tell* instead — and on the unanswerable path the result carries **no** thinking flag at all rather than a fabricated `false`, while the pure write-back returns the very same entry object untouched. A verdict that reasoning cannot be disabled **removes** a previously declared spelling rather than leaving it, since a refuted spelling keeps the engine sending bytes the gateway ignores while the catalogue still renders it as already off. Evidence is lengths, finish positions and status codes only — a planted secret in both the answer and the reasoning channel must appear nowhere in the result, so the record can go into a log or a ticket whole. One cross-package premise is checked by really running the other package’s parser rather than quoting its documentation: everything this probe writes eventually passes through the settings schema on its way into a catalogue, and that field is declared parse-transparent precisely so the vocabulary can live on the consuming side. If it ever narrows, the spelling is stripped **silently** — indistinguishable from the probe never having run — so the guard feeds the probe’s real output through the real parser, checks the compat object comes back key for key, and checks a dialect word this client has never heard of survives too. Two shapes that must be rejected really are rejected, since otherwise the survival checks would hold on a parser that accepts anything, and a bare entry is asserted valid first, because the first run of this section reddened on a space in a fixture’s name — a fixture that cannot pass would disguise the real alarm as already having fired |
|
|
385
385
|
| `scripts/run-decide-receipt-test.mjs` | What a decision verb actually **answered** — and, more importantly, what it did not. A success response on the newest lane is only an acknowledgement that the decision was accepted for delivery: the approval is still pending, and a client that clears the card on it shows either a ghost card that was already approved or a card that vanished while the decision was lost. So the package deliberately has **no** "was it resolved" predicate — nothing in that body can answer it — only the opposite one, whose `false` is likewise not evidence of resolution; resolution is only ever the next running arm on the stream. The guard pins that inversion in the product source too: the success path must no longer clear the latched gate, while the stream-observing path that really clears it must still be there. The body has four shapes with **no** key common to all of them, so every position is read as honestly absent, and the handoff handle — which run to watch from here on — requires **two** facts together, since either one alone would either point the stream at the run it already had or mint an empty handle. The record of what finally happened to an already-decided action is read through the **same** reader as every other gate record rather than a second copy, and its absence means **unknown**, never *it was allowed* — the two can even contradict each other, so the card says nothing at all when it is missing. The three refusals on that lane each get one distinct sentence and a disposition taken from **why** each was refused rather than from severity: one cannot be helped by re-sending at all, one waits on the host, one just drops an option — and none of them carries a countdown, because the server never mints a wait for them. Recognition is a **closed set**: an unrecognised code on the same prefix returns nothing rather than a guess, since that prefix also houses a safety signal whose whole rule is never to retry automatically, and the recovery handle is read as absent when unreadable rather than substituted from a different identifier that no longer appears on that lane **0.68.3 (core 7.18.0):** `gate.disposition.classifier` is read key by key into a named view (requested model, model that answered, and the ladder fallback when one happened); a half-shaped record yields no classifier at all rather than a half view, and absence stays "unknown", never "the seat answered itself" |
|
|
386
386
|
| `scripts/run-approval-frame-chrome-arms-test.mjs` | The two in-stream approval frames finally reaching every host through the shared pipeline instead of one shell's private branch — the shape of a layering defect: hosts that only consume the package could not rebuild their pending cards after a reconnect, and did not clear a card the engine had withdrawn. The payload is deliberately carried as the **envelope** the upstream types declare rather than the first-version card: the stream parser applies no predicate, so narrowing here would let a legitimately newer frame pass as the older shape and invite consumers to read keys a newer card never promised. The guard therefore pins that every open key survives untouched, that an unknown version still passes through, and that narrowing is left to the host's own predicates — with the fallback being a generic card and a person, **never** an automatic denial. A frame whose version cannot be read at all is reported as malformed rather than dropped in silence, because both frames carry user-visible decisions and state changes. Both arms are registered as **required** host duties, and their duty text names the load-bearing rules a host would otherwise have to rediscover: which predicate to narrow with, that the reconnect preamble — not a replayed historical frame — is the authority on which cards exist, and that a withdrawal frame can be lost entirely. Unlike the sibling arms, these carry **no** sub-stream cutoff: an approval raised under a delegated call still has to reach a person, and filtering it by ownership is the host's job, not a reason to discard it. Finally the upstream bytes that justify the envelope discipline are checked to still be there, since the whole design rests on them |
|
|
387
387
|
| `scripts/run-terminal-status-vocabulary-test.mjs` | One place that decides whether a run has **ended** and whether it ended badly — written because that judgement had already been hand-copied three times, so the day the engine added a word for *the agent itself reported it cannot continue*, every copy missed it and a panel settled a self-reported failure as a success. The distinction the table exists for is pinned from both sides: that word belongs in it, while the two words meaning *waiting for a person to decide* deliberately do **not** — reading those as endings would bury a run that is actively waiting on the reader. A word this client does not know answers *no*, and the guard states plainly that *no* is not evidence of success: proving success means reading the positive side, so negating this predicate is the very mistake that caused two earlier incidents. The fleet lane gets the same treatment from the other direction: a workflow parked on a durable approval used to fall through to *running*, leaving the person with no hint that a card was waiting, and it now lands on the same rendered word the task lane already used — same fact, same word, checked end to end on a real row. Why the word was added directly rather than carried as a private superset key is checked mechanically against the upstream declaration being open, so the day it closes this reds and the decision gets revisited. The residue sweep is the point: the source tree must contain **no** further inlined copy of the judgement, each of the three former sites is checked to really read the single predicate, and the one reviewed exemption carries its reason **and** a liveness assertion, so an exemption whose justification expires cannot quietly keep standing |
|
|
@@ -392,12 +392,12 @@ guard still cross-checks the table by name).
|
|
|
392
392
|
| `scripts/run-classifier-status-test.mjs` | What state the auto-mode classifier is in **on this session** — the question a doctor line, a model settings page and a permission card’s status row all ask, and a different question from the one the approval card asks (*why am I being asked right now*), so the sentences are pinned mutually distinct from that face’s as well as from each other. The session-level half of this reading — a breaker record the engine used to keep — was **retired upstream**, and the guard now holds that retirement from **both** sides: the engine's own declarations must really no longer carry it (a fact coming back would mean the removal here was the wrong disposition, and that deserves a conversation rather than silence), and this package must carry no alias, no state word and no leftover narrowing for it — a reading kept alive for something nobody emits any more is a promise the interface cannot keep, and it left the doctor line advertising a state it can never reach. What remains is ordered by the quantity that actually decides whether the classifier is running: the fact from **this round** first, then whether this leg is armed — a decider is minted per run, so a later leg can be armed again. Not armed, and a section that never arrived, both answer **undefined** rather than *available*; that arming question has its own field and answering it twice grows a second ledger. Arming and availability are also **two words, not one**: the engine says a decider was minted *for this leg*, which is an assembly-time fact, while whether that decider answers any given round is a **per-call** one — so an armed leg reads `armed` and only a positive per-call fact (an ask whose origin is the classifier's own denial-bound fallback, which by construction stands *after* the classifier ran) reads `available`. Every other ask origin is refused as evidence and for a stated reason rather than out of caution: several are ones the classifier is structurally forbidden to answer, and for the rest a surviving ask is precisely the case where it did **not** resolve one — so reading availability off them would be a guess. The projection is a **whitelist**, so an older engine still sending the retired member loses it at the boundary while the two live facts beside it ride through untouched. Rendering never throws and never impersonates: a state word this client does not know — including the retired one, which a restored view can still carry — reaches an honest fallback that names it verbatim, carries no invented explanation of a mechanism that no longer exists, and is proven distinct from all three real sentences; prototype keys reach that same fallback rather than a function body, checked against a real out-of-table word so the comparison cannot hold vacuously |
|
|
393
393
|
| `scripts/run-compaction-boundary-projection-test.mjs` | The compaction divider and the one frame that makes its anchor resolvable. The trigger word is passed through as an **open set** instead of being folded to two: the engine deliberately stopped flattening its third value (a compaction that was not optional — a prompt-too-long recovery or trim pressure) and carries what the hook layer saw, so folding it again at the package boundary re-introduces exactly what upstream had just removed, while a consumer branching on *is it manual* keeps its behaviour byte for byte. Only an unreadable word (absent, empty, non-string) falls back — that is *could not read it*, not *read it and did not recognise it*. Two superset keys ride the metadata and neither fabricates: the preserved-segment anchor is minted only when its id really reads out, because half an anchor sends the host looking up an empty string in its map, and the clamp ratio is a **disclosure** whose real zero is a fact rather than an absence. The clamp ratio also carries a registered exit condition — the service really sends it while the SDK arm has no seat for it yet, so the read is defensive and this guard reds the day that seat appears, forcing a re-check instead of leaving a cast to rot. The committed-message frame moves out of *deliberately not projected*: that classification was true about transcript rows and false about **positioning**, since the engine states that consumers build their own id-to-message map from this frame to place the divider — projecting the anchor without it hands the host something it cannot resolve. It becomes a neutral internal arm and an optional chrome ledger event, never a transcript row (the frame carries no body, so minting one would put words in the engine's mouth), with both required ids narrowed and a malformed frame recorded rather than half-minted |
|
|
394
394
|
| `scripts/run-cost-absence-projection-test.mjs` | Telling **declared free** apart from **never priced**, in both directions, because the package was getting each one wrong in the opposite way. The engine separates them on the wire — an absent cost means some spend had no price table, an explicit zero means the model declared itself free — and the result projector used to require a *positive* number, so a genuinely free run could not say so; while the per-model mirror folded absence to zero, so an unpriced run told a billing consumer it cost nothing. The total is now reported as the engine stated it, with absence and non-finite values alone reading as unknown, and a negative passed through rather than corrected, since a refund is a legal figure and the package is not a second accountant. The per-model figure keeps the CC shape intact — that field is a required number and *unknown* is simply not expressible in it — so the value stays zero and a **companion superset bit** carries the distinction, which means the two are read together and a reader that only ever looked at the number is unchanged; the bit is minted only in the absent case and never as `false`, since a key present with a false value reads as a third state. The same mint point serves both the wire's per-model split and the synthesised current-model row, so neither can drift. Alongside it the cache-write figure stops being a hardcoded zero and reads the field the wire has always carried, in both the flat usage and the synthesised row, and all four flat token slots move from a null-coalesce to a finite-number guard — the stats object has an open index signature and the wire is JSON, so a string or an infinity would otherwise land in a slot the types promise is a number, compiling green and surfacing only when something sums it |
|
|
395
|
-
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. A later section reconciles the terminal list against the denied calls seen in the same stream, so denials that never reached a human (rules, hooks, classifiers, write protection) are listed too: complete rows join the CC list, rows missing a field stay on the extended list and mark it incomplete. A further section feeds the same raw tool-result frame through the projector into both lanes and requires the denial category on the interactive transcript record, on the non-interactive frame and from the reader on the raw frame to agree, including frames whose settlement, classifier cause or classifier attribution is present but unreadable — a shape the narrowed gate view on the internal frame cannot show — and pins that the reader gives the same answer on the internal frame as on the raw one. |
|
|
395
|
+
| `scripts/run-permission-denial-projection-test.mjs` | The terminal result's **permission-denial list** being the wire's real one rather than a hardcoded empty array. The session vocabulary carries a list of tool calls that were denied; the projector used to mint `[]` in both the success arm and the error envelope, which folded two different statements into one — *nothing was denied on this run* and *this frame carries no such ledger at all* (an older engine, a rejection envelope, a failure event that arrives without stats) looked identical. Each denied gate on the wire's human-review ledger now becomes one record, in wire order, carrying the keys the wire can actually honour: the tool name when it reported one, and a superset field with the engine's own short, redacted one-line summary of the call's input. **Two lists, deliberately.** The reference shape requires three fields on every element — tool name, call id, and the full input object — and the wire's ledger carries only the first. Filling the other two with an empty string and an empty object would be invention; putting a half-filled element into the reference array would break the element contract, and a strict consumer validating the stream drops the *whole* result message rather than one field. So the reference array admits only fully-formed records — empty today, and filling itself the day the wire grows the two missing fields, with no code change — while every record the wire really has rides a superset carrier beside it. A contract check pins today's absence, so that day turns this guard red on purpose. The companion bit means *this reference list cannot be claimed complete*: no ledger, an unreadable row, an unrecognised decision word (a rejected plan is not a denied tool call, and a row with no decision at all is not a judgement), or a record that could not be fully formed. Only its absence lets a reader say *zero denials*; it is never minted as `false`. Rows that cannot be read drop themselves rather than the whole ledger, and both arms go through one mint point so they cannot drift. Since 0.73.4 the third CC key is sourced from the same stream's `tool_start` frame, joined by call id: a row joins only when the frame was seen on this stream, its arguments are a plain object, and no string leaf carries a transport replacement token or a cycle / depth placeholder (scan budgeted); both halves have positive controls (a fully joined list drops the discriminator, a partially joined one keeps it), the ledger's own input wins when present, the snapshot is per-stream and capped with a one-way overflow latch, and an id seen with two different argument objects never joins. Later sections add the second stream-local join and the two discriminators the headless exit-code rule needs. "Which layer denied this" is not on the denial ledger at all — it is on the gate record of the same call's close-out frame, so it is joined by call id under the same law as the arguments: the closed word table is checked on the collecting side, the ledger's own value wins if it ever arrives, a word from outside the table is not stamped, and a row that cannot be joined keeps the key absent rather than claiming nobody denied it. The classification word is carried on both lists under the same name and the same value, so a consumer needs one reader, not two. The "this run produced no tool output and was denied" flag is present only when three independent things hold at once — the denial evidence is read from the full list rather than the strict one, which can be empty for reasons that have nothing to do with denials; this stream saw no successful tool close-out; and this stream can honestly claim to have watched the run from its first frame. A stream that reconnected mid-run cannot make the last claim, so it mints nothing rather than a false negative, and the flag is never minted as false. From 0.80.0 that classification has a **second source**. It used to come only from this package's own decision path, so a refusal the engine settled on its own — a deployment policy answering the card on an unattended lane, with no client involved — left the field empty even though the same stream's gate record said exactly what had happened. The engine's own settlement word now fills it when, and only when, the local one is absent: the package's own attribution always wins, because letting a replayed frame overwrite it would let the wire change what the host itself said. The word is read literally in both directions and never reverse-engineered, and the separate field naming *which layer* refused is left exactly as the wire wrote it — the two answer different questions, and rewriting one to match the other would make them say the same thing twice. A later section reconciles the terminal list against the denied calls seen in the same stream, so denials that never reached a human (rules, hooks, classifiers, write protection) are listed too: complete rows join the CC list, rows missing a field stay on the extended list and mark it incomplete. A further section feeds the same raw tool-result frame through the projector into both lanes and requires the denial category on the interactive transcript record, on the non-interactive frame and from the reader on the raw frame to agree, including frames whose settlement, classifier cause or classifier attribution is present but unreadable — a shape the narrowed gate view on the internal frame cannot show — and pins that the reader gives the same answer on the internal frame as on the raw one. |
|
|
396
396
|
| `scripts/run-cost-reconcile-projection-test.mjs` | The **end-of-run cost reconciliation** reaching consumers at all. The engine splits a run's spend on the wire — the task's own cost, which deliberately excludes delegated sub-agents, the delegated total itself, and the within-task compaction subtotal that sits inside the own figure — and states two reconciliation identities for them. The package used to project none of it, so a cost view could only ever see one number and under-reported both delegated and compaction spend. Both structures are now projected onto the result as superset fields in the wire's integer micro-currency unit, read key by key, with unreadable keys dropped individually, an entirely unreadable structure omitted rather than emitted empty, and unknown categories passed through since the vocabulary belongs upstream. The delegated cost stays **absent when it was never priced**, never a fabricated zero. The same reader also feeds a terminal chrome arm carrying the three parts plus the reconciled total, so the two faces can never compute different answers; the reconciled total is minted only when both sides are known, and otherwise a discriminator bit says which side is unknown. **The reference field for total cost keeps its meaning** — it remains the task's own spend and the delegated total is not folded into it — because that is a shape the wider ecosystem reads; the reconciled figure is offered beside it, not in place of it. A frame that carries no stats emits no arm at all, and the existing rule that in-stream per-turn usage is not published for sub-flows is pinned unchanged, since delegated spend arrives once, at the end. The bit that says those figures are a lower bound is **per stream**, not per context: the emit context belongs to the caller and may be reused across streams, so a gap observed on one run is no evidence at all about the next one — the observation is held for the duration of one stream and handed to both projection faces by value, and the guard drives a reused context both sequentially and concurrently to prove neither direction leaks |
|
|
397
397
|
| `scripts/run-task-progress-terminal-projection-test.mjs` | The one tick that says a delegated child **finished**. The engine fires exactly one final beat carrying a terminal face, and says in the same breath why it exists — so a consumer sees the row finish instead of watching it vanish after the last running beat — but the package's projection whitelist had no seat for that field and its adapter still carried the older premise in a comment, so the terminal beat arrived byte-identical to another running one: the panel row stayed up waiting for a defensive sweep (which only ever settles rows bound to a card still open this turn) or for a separate notification frame. The status now rides through as an **open set** with the vocabulary left upstream, while the question *which words are terminal* is answered by a closed pair on the adapter side — an unrecognised new word takes the running path, because guessing it terminal ends a row that is still working whereas one extra running beat merely renders late. A terminal beat settles the row directly under the lane proof its binding gives it (not the main lane a notification would use, and not by card id, since the engine is naming a child rather than closing a card), freezes the inline group-row twin in the same beat so a later sweep cannot reset the real tool count, clears the session-resident ledger, and fires the stop hook only for a child whose start really fired. It does not mark the row live or emit a second progress beat, and it shares the settled-row ledger with the other two settle legs so a replay or a double-delivery cannot produce a second end. Three things are pinned **unchanged**: a running beat, an absent status (older engines never send the field, and reading absence as terminal would make every child row disappear on its first beat), and the workflow lane gate, which still runs before any of this |
|
|
398
398
|
| `scripts/run-assistant-arm-identity-test.mjs` | The identity keys on an assistant row, and an explicit account of the two that are **deliberately not** there. What the renderer received was a bare role-and-content object, so a dozen consumer sites downstream were each estimating what the message envelope should have told them. The id is taken from the engine's own event id rather than minted locally, because it has to be **the same value** on the live leg and on a durable replay — a freshly minted one would make a replayed message look new to a host's dedup and to rewind — and when the wire carries none the key is simply absent rather than filled with a random stand-in wearing an identity it does not have; it is also kept distinct from the envelope's own local render key, which is a different identity. The model name comes from what the host pinned when it opened the stream (the request was the host's to build) and is never guessed, since a wrong model name is worse than none once a billing or capability face looks it up. Usage and stop reason are **not** minted on this arm, and the reason is frame order rather than effort: content arms arrive before the turn's closing frame, so at the moment the arm is emitted the engine has not yet said what the round cost — anything put there would be an estimate, which is the very thing this work exists to remove — and synthesising a follow-up assistant update when the real figure lands is also refused, because that shape does not exist upstream and would place a message in the transcript the engine never sent. Their real values leave through the turn's own neutral arm as two superset keys, the usage one reusing the **same single mint point** the footer rollup already folds so the two faces cannot diverge, and the stop reason passed through verbatim as an open set — the machine signal for *was this turn cut short*, previously blind on both the stream and the trace. The existing behaviours beside them are pinned too: no arm at all when usage is wholly absent, and the sub-flow cut-out that keeps a child's turn from driving the leader's face |
|
|
399
399
|
| `scripts/run-text-segment-authority-test.mjs` | The **authoritative segment replacement** on `text_end` (server >=7.75.3). `text_end.content` now goes through the same redactor as `result` and the ledger while `text_delta` stays verbatim, so the two **may differ** — an answer that quoted a credential used to be committed to the local transcript in its unredacted form, because the arm only forwarded the boundary signal. Six timing shapes are pinned, two of which an adversarial review reproduced against the installed engine's real bytes and which the first design got wrong in both directions: a second boundary in the same turn (the per-block case on one provider lane) used to make the first segment's prose vanish, and a boundary that arrives *after* the tool card (the other lane emits it at finalize) used to be read as "this package never handled that segment" and reported nothing at all. Three additive keys, all never-false; the two shapes that look alike are told apart by the second one, because the host's action in them is the opposite. The end-to-end legs drive the real pipeline without hand-inserting a segment commit — doing so is exactly what hid the first defect. A second review round then found two combination timings on top of the first fix — a tool card followed by *more* deltas in the same segment, and a byte count that had been documented as a message count — and both are pinned here too. A third round caught a length that the prose called bytes while the code returned UTF-16 units — harmless in ASCII, and on CJK text enough to leave the credential on screen — plus a backfill ledger that had to be kept in step, so the terminal frame does not re-render the segment a second time — kept in step only where the whole stretch sits in one message, because those ledgers are per-message and a fourth round showed that writing across them charges one message's prose to another. A fifth round settled the whole class into one invariant the guard now checks against the previous release's behaviour: this package only rewrites bytes it is still holding in the current message — once a segment has crossed a package-side boundary it emits the three keys and changes nothing else **0.68.2 (CC-01):** the segment identity is now minted here, not by the host: every committed assistant text row carries a top-level `_sema_segment_id` (stamped once at the `adapt()` exit, so the durable whole-message leg and the streamed-segment leg are covered alike; thinking blocks, tool_use-tailed rows and chrome events are left byte-for-byte), `text_segment_end` carries the same value as `segmentId` before rotating, subagent boundaries never rotate, and a replayed stream yields the same identities. Three mutations (no rotation / no stamping / stamping tool_use rows) each turn the guard red **0.69.0 (CC-02):** the same authority replacement now covers the reasoning face (`reasoning_end`, server >=7.77.0): a thinking block still buffered is swapped whole and its live tail recomputed; one already committed at a boundary (the usual timing, since the first text delta commits it) is left untouched and the host is told the row to replace by its uuid, never re-emitted. Subagent boundaries are ignored and the text-segment identity does not rotate **0.69.1 (CC-09):** the run-stream replay guard still drops a frame whose event id was already seen, but it now reports the drop through the host's dropped-frame sink as `duplicate_seq` instead of vanishing silently (server 7.77.0 reuses the first reasoning delta's id for `reasoning_end`, so that authoritative segment is lost on the print lane until 7.78.1); the interactive adapter has no such guard and keeps receiving it **0.69.1 (CC-10/CC-11):** subagent segment-end frames are fenced on all three identity keys (a frame carrying only `sourceTaskId` no longer masquerades as the leader's), and a reasoning segment that spans tool cards now hands the host every committed row it covers (`committedUuids`) so nothing unredacted is left behind |
|
|
400
|
-
| `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the suites above actually turn red when the material they check really breaks — a census found the ones clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. The guards of the same shape that are too costly to rehearse, and the behaviour/projection suites that red on their own assertions, are catalogued rather than rehearsed — see `docs/GATE-NEGATIVE-CONTROLS.md` for the three tables (all generated from `gates-manifest.json`, which carries each suite's classification), the reasons, and a one-minute manual replay recipe for each blind one. The suite reconciles its case list against that classification by name in both directions, so a case quietly dropped from the array without the manifest following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own |
|
|
400
|
+
| `scripts/run-gate-negative-controls-test.mjs` | Whether the registry-shaped guards among the suites above actually turn red when the material they check really breaks — a census found the ones clean enough to rehearse safely (closed sets, mirrors, baselines, floors, a type-shape ratchet) without touching any judgement code. Each is exercised by tampering a disk copy of the real material, spawning the guard's own unmodified script, asserting it exits non-zero and names the disease, then restoring the file byte-for-byte. The guards of the same shape that are too costly to rehearse, and the behaviour/projection suites that red on their own assertions, are catalogued rather than rehearsed — see `docs/GATE-NEGATIVE-CONTROLS.md` for the three tables (all generated from `gates-manifest.json`, which carries each suite's classification), the reasons, and a one-minute manual replay recipe for each blind one. The suite reconciles its case list against that classification by name in both directions, so a case quietly dropped from the array without the manifest following is itself an undeclared blind guard. The backup that makes the restore possible is taken by **exclusive create**: checking for it and then copying are otherwise two steps, and two instances can pass the check together — the later one overwrites the only clean copy with material the earlier one has already tampered, and the rehearsal that promises to leave no trace leaves a permanently corrupted file instead. That interleaving is rehearsed too, in a throwaway directory of its own. A run that is interrupted — by the suite runner's time limit, by Ctrl-C or by its terminal closing — restores what it was in the middle of tampering with before it exits: the rehearsed guard's whole process group is stopped first, so nothing that guard started can write the tampered state back afterwards; the file is restored byte-for-byte from its backup and the backup removed only once the bytes match; the build output is regenerated from the restored sources within the grace period the runner allows (a rebuild that would overrun it is cut short, its compiler stopped, and the run says the build output was not regenerated); and the run exits with 128 plus the signal number rather than passing. Before the first rehearsal the suite also looks for backups left behind by an earlier run that could not clean up — in the trees its targets live in, following a symlinked dependency directory — and if it finds one it refuses to start, names the file and how to check and restore it, and leaves the backup untouched rather than overwriting what may be the only clean copy. |
|
|
401
401
|
| `scripts/run-engine-cap-reader-factory-test.mjs` | The one shared implementation behind every capability reader's four ports (cache, generation gate, probe tee, invalidation), exercised as a table: every reader in the table runs the *same* criteria (the table length is the source of truth, and a roster check fails the gate if any source file calls the factory without having a row) — the four states, a throwing projection treated exactly like an unreadable one (and never escaping the tee), the generation rules (a stale generation is dropped before the projection even runs; a projection that changes the generation mid-flight cannot overwrite the newer value, whether it returns or throws; the caller's `opts` is snapshotted once; and omitting the generation still writes, because that supply is additive and this refactor does not quietly tighten it), the top-level key's getter being read exactly once, a freshly minted "unobserved" reading on every miss (two misses are never the same object, so a consumer that mutates one cannot taint another base URL), one independent table per reader that never take each other down, invalidating one base URL leaving every other base URL's reading untouched, `forget`/tee being no-ops on an empty or non-string base URL, the read anchor being resolved dynamically, and the deliberate split in how presence is judged per reader. |
|
|
402
402
|
| `scripts/run-mcp-liveness-test.mjs` | The engine's **liveness observation** about each MCP server it hosts (`wiring_manifest.mcp[].liveness`, engine-side from core 7.24.3 / server 7.91.2): one reader, one word list, one leg-level verdict. The cell answers *can this server still be reached* — it is not the connect-time verdict beside it, which the engine deliberately freezes (a server that died mid-run still reads `connected`), and it is not a re-dial's judgement either, so `status: "failed"` next to `liveness.state: "reachable"` is a **real row**: the server answered the handshake and answered with a protocol error — up, and misconfigured. The three words are read as a closed set (an exchange completed / it was lost in transport or the clock / what came back does not answer the question), and a word outside it is malformed rather than rendered, because a word nobody upstream has defined is not a sentence worth putting on screen. **Absence is the fourth reading and is not one of the words**: it means *no liveness record is available*, which on the wire covers a server this leg never reached, a declaration that could not be dialled, an older engine, and a record the projection ahead of us dropped — all indistinguishable, so it is never read as "we looked and could not tell" (a strictly stronger claim), never as healthy and never as off. The failure-class footnote rides the unreachable word only, and one that turns up anywhere else, or that is malformed, loses **just the footnote** while the word and its timestamp stay: the honest reading is then "cannot be reached, reason not given", not "this record is broken". Malformed never becomes healthy: a cell that is present but unreadable marks its row and pulls the leg-level verdict back to *cannot tell*, since letting it sit beside a reachable row would report a leg as reachable on the strength of a record that may well have said the opposite. The verdict takes the worst fact first rather than a majority or the newest reading, carries no server count — so there is no fabricated zero to be read as "no problems" — and its timestamp belongs to **that leg's** observation, not to now: the engine runs no probe and adds no traffic of its own, so this is a per-leg snapshot rather than a heartbeat. The replayed roster on the session panel and the live leg go through the same reader, and the panel's own deployment-side rows carry no liveness position today, so the verdict does not borrow a word from that face. The word list is bitten in both directions where an installed witness exists and the absence of one is itself asserted against the installed engine's version, so the day it ships the comparison starts on its own; since sdk 11.2.0 the wire types declare the cell, and the word list and the presence test are pinned to those types at compile time in both directions, so the two can never drift apart |
|
|
403
403
|
| `scripts/run-peer-lane-rules-write-capability-test.mjs` | Two more engine self-descriptions read the same four-state way as their seven sibling capability readers (`capabilities.peerLane`, `capabilities.permissionRulesWrite`): an absent key is not reported (an older engine that predates the position, never folded into `false`), `true` is present, `false` is a positive absent (the cross-session lane not being mounted on this deployment, or this particular call not being able to reach the tightening-direction write entry), and anything else is unreadable and drops the cell. Each carries its own single-source verdict (`peerLaneAvailable` returns `yes`/`no`/`unknown`; `permissionRulesWriteAvailable` collapses to a plain boolean, present being the only `true`). The write-entry position pairs with a boolean convenience port in the persisted-rules module, and this guard pins that port to derive from nothing but this one reader's own reading — never a conjunction with the lane-reachable position, and never a second read of the deployment-level existence signal the revoke surface uses (the two are documented as reading differently on purpose): a deployment where the lane answers true but the write entry's key is simply absent (an older binary) must still come back `false`, a deployment where the write entry answers true while the lane key is entirely unseen must still come back `true` (proving no silent conjunction crept in), seeding only the general capabilities cache — never this reader's own feed — must still come back `false` (proving the convenience port cannot be satisfied by the wrong table), and passing an explicit `undefined` base URL must still come back `false` even while a different, already-installed engine target answers `true` for the same position (an adversarial pass found the naive forward of that parameter falls through to the reader's own convenience default, silently answering for whichever engine happens to be installed rather than the caller's absent target — the fix routes an explicit absence through the same empty-string path the reader treats as unobserved). |
|
|
@@ -419,19 +419,20 @@ guard still cross-checks the table by name).
|
|
|
419
419
|
| `scripts/run-rewind-archive-capability-test.mjs` | The read face for whether a rewind can restore the code archive, and the honest three-state answer the ends render from it. The shells used to decide this from a local backup table that only their own in-process tools ever fill, so on any session where the engine runs the tools it stayed empty and the two code-restoring rewind modes simply never appeared, while the engine had been keeping a file history the whole time. The judgement now comes from what the engine itself advertises, and each of the four bits it advertises answers a different question: whether the conversation can be forked at a message at all, whether a fork can restore the tracked set, whether a code-only restore is possible, and whether this deployment understands the current spelling of the request key. The mode that rewinds the conversation and restores the code together needs both of the first two, and the engine offers no single bit for that combination, so the combination is made here: one bit stated off is enough to rule the mode out, both stated on make it available, anything else stays unknown — reading the file-history bit alone would offer that mode on a deployment that keeps file history but has no conversation anchors, where the request can only fail. A bit that is absent means an older engine that never spoke about it, which is not the same as an engine that said no, and a bit reported in a shape this reader cannot read is a third thing again — it is recorded as unreadable rather than quietly filed under "not reported", because those two send an operator to different places. One unreadable bit does not discard its siblings; only a response that is not a capability object at all clears the cell. What the ends get is available, unavailable or unknown, and unknown stays unknown: folding it into unavailable would hide the mode again, which is the mirror image of the bug this replaces. The sentences the doctor row can print are checked to be pairwise distinct and to avoid implying a refusal the engine never made, and the spelling of the outgoing request is deliberately not made to follow the epoch bit, since the older spelling is rejected outright by current engines |
|
|
420
420
|
| `scripts/run-registry-quota-usage-test.mjs` | The projection of the cloud control plane's quota reading, and the three different things a missing number can mean there. This response says `null` in two places and means something different each time: no token quota is configured for this principal on this instance, and this window has no cap at all. Both are **facts the server is asserting**, not gaps in the reading — while a key that is absent or carries the wrong type is a genuine gap. All three have to survive to the screen separately, because folding them is how a user ends up staring at a confident `0`: an uncapped window rendered as if nothing were left, or a deployment that simply never configured quotas rendered as if the quota were exhausted. The complaint that started this was the opposite direction — a centrally configured quota that the command line could not see at all — so the reading also refuses to let an unreadable response masquerade as "no quota configured". Two fields deliberately do not share one signal: whether a window is exhausted and whether there is a recovery time, since the recovery time is only ever populated in the exhausted case and reading its absence as "not exhausted" would answer a question the response never answered. Counts that the server always provides are narrowed no further than the mint: a used counter has no uncapped state, so a null there is unreadable rather than zero. The wording helper carries the only human-facing phrasing, and the sentences for "no cap" and "unknown" are checked to contain no digits at all |
|
|
421
421
|
| `scripts/run-file-history-capture-capability-test.mjs` | The engine's file-history-capture self-description (`capabilities.fileHistoryCapture`), read the same four-state way as its sibling capability readers: an absent key is reported as not reported (never folded into `off`), words are taken as an open set so a newer mode is not mistaken for a malformed answer, `fileHistoryCaptureMode` recognises only `off` and `on-always`, and the wording for `off` speaks about capture only — whether code can be rewound is left to the rewind readings. |
|
|
422
|
-
| `scripts/run-model-identity-resolvability-test.mjs` | The model-identity judgement a client makes before letting anyone in: can the engine it is about to use start with a model name? Each end reports what it read from each place that can feed a model name to a local engine (complete, partial — a gateway address or a credential but no model name —, absent, or unreadable), or, for an engine that runs elsewhere, whether that engine has been seen answering; `modelIdentityResolvability` answers resolvable, unresolvable or unknown. Having part of an upstream configuration is not having enough of one, so partial lanes never add up to resolvable; a lane that was not reported or could not be read makes the answer unknown rather than unresolvable; an engine that runs elsewhere is never judged unresolvable and local lanes are never consulted for it (it does not start without a model name, so seeing it answer is enough to call it resolvable). `modelSetupDecision` combines that answer with whether this end can configure a model at all: setup is offered only for unresolvable on an end that can configure one, an end that cannot says so and points at whoever runs the engine, and unknown never opens setup. The detail and notice sentences are checked to be pairwise distinct, unknown sentences neither claim a model is configured nor that it is not, and the module is checked to import no platform I/O The catalog lane counts as complete only when the host reports that the local engine accepts a catalog default; a report of `false` reads the catalog as partial, and no report (an explicit `undefined` included) or an unusable one leaves that lane undetermined — the same answer as before the catalog could be complete. Every sentence about a complete catalog names the catalog, including the undetermined sentence when the report is `false` and another lane cannot be read, and a 25-cell matrix pins that the report changes nothing when the catalog is not complete. |
|
|
422
|
+
| `scripts/run-model-identity-resolvability-test.mjs` | The model-identity judgement a client makes before letting anyone in: can the engine it is about to use start with a model name? Each end reports what it read from each place that can feed a model name to a local engine (complete, partial — a gateway address or a credential but no model name —, absent, or unreadable), or, for an engine that runs elsewhere, whether that engine has been seen answering; `modelIdentityResolvability` answers resolvable, unresolvable or unknown. Having part of an upstream configuration is not having enough of one, so partial lanes never add up to resolvable; a lane that was not reported or could not be read makes the answer unknown rather than unresolvable; an engine that runs elsewhere is never judged unresolvable and local lanes are never consulted for it (it does not start without a model name, so seeing it answer is enough to call it resolvable). `modelSetupDecision` combines that answer with whether this end can configure a model at all: setup is offered only for unresolvable on an end that can configure one, an end that cannot says so and points at whoever runs the engine, and unknown never opens setup. The detail and notice sentences are checked to be pairwise distinct, unknown sentences neither claim a model is configured nor that it is not, and the module is checked to import no platform I/O. The catalog lane counts as complete only when the host reports that the local engine accepts a catalog default; a report of `false` reads the catalog as partial, and no report (an explicit `undefined` included) or an unusable one leaves that lane undetermined — the same answer as before the catalog could be complete. Every sentence about a complete catalog names the catalog, including the undetermined sentence when the report is `false` and another lane cannot be read, and a 25-cell matrix pins that the report changes nothing when the catalog is not complete. |
|
|
423
423
|
| `scripts/run-cloud-effective-projection-test.mjs` | The cloud control plane's effective-configuration response beyond its four configuration domains, and what a locally started engine does with the models document derived from it. Three top-level keys are read with the meaning their producer gives them: `warnings` (degradation warnings from the build that produced the served view — an empty list is a clean build, an absent key is an older server that cannot tell), and `budget` / `runtimeCaps`, where `null` means two different things: nothing resolves for this principal when you view yourself, and values withheld when the response previews another principal. A missing key or a wrong type is a third state, unknown, and none of the three is ever folded into a zero, a `false` or "no budget". A malformed warning row, budget field or cap costs only itself, and a known budget field of the wrong type is named as not shown rather than silently read as "no limit on this axis"; warning kinds are an open set, so a kind this client does not recognise still produces a warning line. The budget and cap readers are reconciled against the installed settings schema. The single wording source puts degradation warnings first and keeps every "not set / withheld / unknown" sentence free of digits, while a zero the server really sent is shown as a zero. On the models side, when the host injects an entry check the models document carries the default model, the tier groups and the active tier group (without the check none of the three is written), and every catalog reference the local engine's schema would reject — a default, role, @-mention entry, tier binding or active group that names something outside the catalog served to this principal — is dropped and recorded, because one dangling reference makes the local engine discard the whole models domain and fall back to its environment catalog; this is proven by reading the produced document with the installed file store. An @-mention allowlist that would be pruned to empty is kept as sent, since an empty list means "everything may be mentioned"; that case is recorded, produces its own warning that a locally started engine will reject the cloud model settings and use its environment catalog instead, and the gate reads the document with the installed file store to confirm exactly that outcome, so the sentence turns red the day the local reader becomes lenient. A per-model budget in which no field could be read is never described as having no limits. Registry annotation keys on model entries (`origin`, `overridesTeam`) are removed before the models document is written: the local engine's schema does not accept them, and a configuration refresh would otherwise be rejected as a whole. A model entry the host-injected entry check rejects is left out of the document and references to it are dropped: the local engine drops such an entry at startup, but a refresh rejects the whole configuration over it, so the gate requires a clean read of the produced document; when no entry passes the check, the catalog is kept as sent and gets its own warning, which the gate proves by reading the document back. The check receives a copy, so it cannot alter what is written. Tier words outside the local schema's closed set are dropped and recorded as unsupported, and the package's tier word list is reconciled against the installed schema in both directions; an active tier group is judged against group names, never model names. |
|
|
424
424
|
| `scripts/run-websearch-verdict-test.mjs` | The per-request web-search configuration a host puts on the wire. Newer servers refuse the whole request when that section is malformed — and a missing or misspelled search provider now counts as malformed, because the section names where the searches go and an unknown destination is refused rather than silently swapped for the deployment's own backend. The old readers in this package dropped such a section without a word, which let the server swap destinations after all. The guard pins the new three-way verdict (absent, honoured, malformed with the field that is wrong) against the server's own judge, vector by vector, whenever that judge is available next to this package; it pins that a half-configured environment is malformed rather than ignored, that a malformed environment never falls back to the settings file (that would change the destination too), and that neither the sentence shown to the user nor the recorded reason repeats an endpoint, a key, or a search-parameter name or value — an unrecognised provider is never echoed either (the sentence lists the valid words instead), so a URL or key pasted into the wrong field does not come back out, even when it happens to be all letters. A host's key store is plugged in through a callback the package calls only after the provider has been recognised, so the precedence between environment and settings stays inside the package. Since 0.83.0 an endpoint that carries a user name or password is malformed as well (the server judges the same way from 7.101.0), the reason sentences match the server's own word for word, and the three older readers that dropped a misspelled provider are gone. |
|
|
425
|
-
| `scripts/run-hooks-merged-disable-projection-test.mjs` | The fourth governance leg of the hooks projection: `disableAllHooks` set in a non-managed settings source. The value that counts is the **merged** one, read from the host through the optional `SettingsPort.mergedDisableAllHooks()`, because a per-source approximation ("any source says true") reads user `true` with local `false` backwards — the merged value there is `false` and every source's hooks ship. When the merged value is `true`, only managed-settings hooks are sent to the engine: non-managed settings can switch off their own hooks, never the managed ones, and a managed `disableAllHooks` is still judged first and sends nothing at all. A full matrix over the four sources, each true, false or absent, is merged with the reference rule (later sources override earlier ones, managed settings last) and every cell's projection is asserted. The reader is the only authority: per-source values never second-guess it, and only a strict `true` counts. A host that does not implement it keeps the previous behaviour and gets exactly one warning per installed settings port, never one per request, and none on paths where the reader would not have been consulted; a reader that throws is treated as `true`, so managed hooks still ship. The session goal's Stop hook and the final-verification yield rule, which reads the projected hooks, follow the same verdict, and the trust gate and the three managed gates are evaluated before the reader is ever called. The last leg pins the member's declared shape in the built declarations: optional, no parameters, returning a boolean or `undefined`. The flag settings source (a settings file or inline settings given at startup) is projected after the local source, and settings hooks stay concatenated managed first (managed → user → project → local → flag): the engine runs hooks in request order, and three of its budgets go to whoever comes first — the per-event wall-clock budget, the per-event cap on model-backed entries, and the total cap on added context — so a managed hook placed after the others could be crowded out and silently not run. The flag source moves with user, project and local under every managed or merged gate; its exec-form entries are dropped and warned about like any other settings source, and the not-run notice names it `Flag settings`. An optional leg feeds the request body to an installed server's hook runner and requires the managed deny, block and context to take effect while non-managed hooks exhaust each of those budgets, and pins the known cost of that order: a non-managed `PreToolUse` hook placed after the managed one can still rewrite the input after the managed check allowed it. Without an installed server the leg reports that it did not run. |
|
|
425
|
+
| `scripts/run-hooks-merged-disable-projection-test.mjs` | The fourth governance leg of the hooks projection: `disableAllHooks` set in a non-managed settings source. The value that counts is the **merged** one, read from the host through the optional `SettingsPort.mergedDisableAllHooks()`, because a per-source approximation ("any source says true") reads user `true` with local `false` backwards — the merged value there is `false` and every source's hooks ship. When the merged value is `true`, only managed-settings hooks are sent to the engine: non-managed settings can switch off their own hooks, never the managed ones, and a managed `disableAllHooks` is still judged first and sends nothing at all. A full matrix over the four sources, each true, false or absent, is merged with the reference rule (later sources override earlier ones, managed settings last) and every cell's projection is asserted. The reader is the only authority: per-source values never second-guess it, and only a strict `true` counts. A host that does not implement it keeps the previous behaviour and gets exactly one warning per installed settings port, never one per request, and none on paths where the reader would not have been consulted; a reader that throws is treated as `true`, so managed hooks still ship. The session goal's Stop hook and the final-verification yield rule, which reads the projected hooks, follow the same verdict, and the trust gate and the three managed gates are evaluated before the reader is ever called. The last leg pins the member's declared shape in the built declarations: optional, no parameters, returning a boolean or `undefined`. The flag settings source (a settings file or inline settings given at startup) is projected after the local source, and settings hooks stay concatenated managed first (managed → user → project → local → flag): the engine runs hooks in request order, and three of its budgets go to whoever comes first — the per-event wall-clock budget, the per-event cap on model-backed entries, and the total cap on added context — so a managed hook placed after the others could be crowded out and silently not run. The flag source moves with user, project and local under every managed or merged gate; its exec-form entries are dropped and warned about like any other settings source, and the not-run notice names it `Flag settings`. An optional leg feeds the request body to an installed server's hook runner and requires the managed deny, block and context to take effect while non-managed hooks exhaust each of those budgets, and pins the known cost of that order: a non-managed `PreToolUse` hook placed after the managed one can still rewrite the input after the managed check allowed it. Without an installed server the leg reports that it did not run. Safe and bare mode: the host reports its startup mode through the optional `SettingsPort.hooksStartupMode()`, which both the plain and the plan path read and which answers `'safe'` or `'bare'`; the flag already carried by the plugin hooks reading, which only the plan path reads and which cannot tell the two modes apart, counts as safe mode unless the reader named a mode. In either mode only managed-settings hooks are sent from the settings sources: bare mode is handled like safe mode, so managed hooks are still sent, also under the managed-only gates. In both modes plugin hooks are excluded as before and the session goal's Stop hook is still sent. Only the two exact words count; a host that does not implement the reader, returns `undefined`, returns any other value or throws keeps the previous request byte for byte, checked for every such reply, both flag states, both paths and with or without a session goal; any other value and a throw each leave one debug line, and so does a reader written as a property instead of a method, which narrows nothing. The reader is called at most once per request and not at all when hooks are already switched off entirely. The optional server leg requires the left-out sources' hooks not to run while the managed ones still run and still deny, in either mode. A hook written in more than one settings place is sent once: entries on the same event, in groups that are identical apart from their hooks (same matcher), with the same identity — command, shell, arguments and condition for command hooks; type, prompt and condition for prompt and agent hooks; URL for HTTP hooks; server, tool and input for MCP tool hooks — are sent once, at the position of their first occurrence in the concatenation order and with the fields of the copy from the highest-precedence source (managed, then flag, local, project, user; within one source the later copy), so a managed copy keeps both its place and its values and a user copy duplicated in flag settings keeps its place but uses the flag copy's timeout; the other entries keep their order, groups without duplicates are sent as they were, and a group left empty is not sent. Different matchers, events, a one-byte command difference, a different condition, an absent versus explicit shell, different types or extra group keys are not merged; malformed entries and entries whose identity cannot be computed are left alone; plugin groups and the session goal hook do not take part, and hooks removed by safe or bare mode never come back. The optional server leg requires such a duplicate to run once with its added context appearing once, and a user copy with a one-second timeout duplicated by a flag copy with a ten-second timeout to run to completion on a two-second command. |
|
|
426
426
|
| `scripts/run-memory-saved-projection-test.mjs` | Engine memory writes (a successful `Remember` tool call) moved off the transcript onto the additive `memory_saved` chrome event, driven through the real pipeline: zero transcript rows for the write (the transcript is byte-identical to the same frames with the write reported as not successful), exactly one event whose `notes` carry the note text verbatim (notes, not file paths) and whose key set is exactly kind / laneProof / id / notes; no event for a missing, empty or non-string note, a non-`true` `ok`, a tool error, a missing result or another tool name; two writes give two events in order with distinct ids; a sub-agent write rides the sub-agent lane and an empty parent id emits nothing rather than falling back to the main lane; the `id` is derived from the write's wire key (the tool-end event id, else the tool-start event id, else the call id; empty ids count as absent), so projecting the same wire events twice gives the same id, and it never collides with the tool result row of the same or another call; the event sits right after the tool result row; the arm is registered as required. |
|
|
427
427
|
| `scripts/run-result-frame-projection-test.mjs` | Result frames and the synthesized terminal rows. The CC key `terminal_reason` is minted on result frames only where it follows from what the engine reported: `completed` on success, `max_turns`, `budget_exhausted` and `structured_output_retry_exhausted` for the three matching engine codes, on both the done-frame path and the failed-event path. Every other outcome leaves the key absent as an own property rather than present with an undefined value: wall-clock and token-budget limits, the classifier denial limit, cancellation, unknown codes, blocked, paused, unreadable or missing terminal records, and the busy-session refusal. The public reader `terminalReasonForResult` shares the minting predicate and is checked to agree with the minted key on every frame the gate produces. Both the minted key and the reader derive the word from the frame's CC subtype (success with `is_error` strictly false, and the three limit subtypes), not from the error code, so a replayed row whose status is paused, blocked or unrecognised never carries a word that contradicts its subtype. The four words are checked against the mirrored CC union, and the minting file is checked to hold no hand-copied code literals. The renamed superset keys (`_sema_error_code`, `_sema_salvaged_result`, `_sema_model_degraded`, `_sema_selected_model`, and the row flag `_sema_api_error_message`) are driven through the real stream pipeline. Each must be present under its new name, the old name must be absent, and every frame the gate saw is swept for old names. The selected model appears on error envelopes whenever the terminal record carries it, and never on a failed event, which has no record. It stays separate from the provider-reported model name. The two in-package readers still work: the interactive result arm reads the salvaged text under its new name (and old-shape frames under the old one), and the print init gate treats the renamed flag as the run having ended. |
|
|
428
428
|
| `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
|
|
429
429
|
| `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed five-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule covering every tool of one server or agent peer (for every protocol namespace the engine knows, checked against the engine package's own table) or containing a wildcard (*) anywhere (an engine that compares exact names would block nothing), and an entry that is not a tool name — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for three deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; and peer-wide rules outside the MCP namespace, which it did not recognise), whose disagreements are counted per class and must match an independent count exactly. |
|
|
430
430
|
| `scripts/run-plugin-hooks-projection-test.mjs` | Plugin hooks: each command hook an enabled plugin declares is decided one by one as running in the engine, running in this client, or not running at all, and the page of hooks sent with a request is built from the same per-turn plan the client uses to skip its own copies, so one hook never runs in two places. Governance is judged first and always wins — a managed hooks switch-off, an untrusted workspace, safe or bare mode, or a governance read that fails sends no plugin hook and does not list it as a gap; managed-hooks-only (set directly, or through a merged non-managed hooks switch-off) keeps only managed plugins; the plugin-only customization lock does not touch plugin hooks. A hook reaches the engine only when this client started the engine on this machine, the engine reports plugin-hook support, the entry is a command, the plugin declares no sensitive option, and the event still fits the engine's per-event limits; the gate walks that matrix cell by cell, including the limit boundaries and a session goal hook counting toward them. A fact that was never read is reported as not known rather than as a fact: a host that does not say where the engine runs gets a "not known whether this client started the engine" reason, an engine whose capabilities have not been read yet gets a "not known yet whether it supports plugin hooks" reason, and the plan's two engine facts are null in those cases, not false. Events the engine never fires run only if the client says it fires them itself, and hooks the upstream behaviour itself refuses (option references in a shell-form command, an unset option in exec form, malformed entries) run nowhere. Exec-form arguments are passed element by element with only saved non-sensitive option references filled in; path placeholders are left for the executor. Sensitive option values never reach the request: with a host that wrongly supplies one, every string in the plan, the request body, the notice, the labels and the log is searched for it across eight cells. A host without the plugin reader keeps the previous request body and gets exactly one warning per settings port; plugin data that throws while it is being read (a throwing getter, a revoked proxy) is treated like a failing reader — no plugin hooks this turn, settings hooks still sent, nothing thrown; the not-running notice names the plugin and events, never a command or an option value, and escapes control characters in names. Command hooks from settings that carry arguments (a non-empty `args` array, which is the exec form, or any other non-null value) are removed from the request until the engine reports support for arguments, because the engine would otherwise drop the arguments and run the bare command through a shell; an empty `args` array is not treated as carrying arguments when the command is made only of letters, digits and `_ . / : + -` (the shell runs the same executable), so such a guard still reaches the engine, while an empty array on a command with spaces or shell characters is removed; `args` on a prompt or http entry, a null `args`, or an entry with no type is left alone, and those go out unchanged. MCP tool hooks, which the engine cannot parse, are removed only from a request built from a plan, whose not-running notice the host shows; a request built without a plan still carries them on engine-fired events, so the engine rejects the whole request loudly instead of a guard hook silently not running — the gate checks both request bodies against the engine's own hooks schema. Without a plan, every removed hook of that kind on an engine-fired event produces one warning per settings port, event and reason. Malformed entries still pass through for the engine to reject loudly, and passing null where the options object goes behaves like passing nothing; a `plan` option that is not a plan is ignored rather than turning the whole page into nothing, and a plan passed directly in place of the options object is recognised and used. A `plugin` key written by hand on a settings hook is stripped before sending (even when its value is undefined), because only hooks that come from the plugin reader may carry plugin context; the settings document itself is left untouched and a debug line records the count. The two hand-copied tables, the engine-fired event list and the engine limits, are checked against their owners. |
|
|
431
|
-
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. |
|
|
431
|
+
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. |
|
|
432
432
|
| `scripts/run-ask-survives-posture-test.mjs` | The single posture predicate `askSurvivesPosture(card, facts)` for sessions whose standing mode would otherwise answer approval cards on the user's behalf (bypass-style modes). It reads two facts and returns one of three verdicts. The first is the ask origin stamped on the card: the question tool (`content_question`), an organization rule (`org_rule`), a hook (`hook`), an explicit ask rule (`ask_rule`), an organization policy or rule store that could not be read (`org_unavailable`, `rule_store_unavailable`) and the classifier's hand-off after its denial limit (`denial_limit_fallback`) must still be asked (the engine requires a real person to answer all three) and every other origin this build knows is left to the posture only once the host has also reported that its own ask rules did not match. The second is the host's own reading of its settings ask rules for this call: a positive match must be asked, and a command the host could not fully parse counts as no match. When the host reported no reading, every card outside those seven origins gets `unknown`, because an origin says who asked and not that the user's own ask rules did not match; an origin this build does not know gets `unknown` even after a reported non-match. `unknown` is never an approval: the host falls back to its own settings rules. The guard checks the verdict for every origin word, both with no host reading and with a reported non-match, against an independent table whose word set must equal the package's origin list, so a new upstream word fails the guard until it is classified; it covers the combinations of both facts, malformed inputs (non-boolean readings, empty or non-string origins, prototype keys, a different letter case), the fact that the predicate does not read the stronger bits on the card (those stay with the host's earlier checks), real card requests produced by the live-frame, parked-row and suspended-ask paths, and a closed, frozen verdict shape. |
|
|
433
433
|
| `scripts/run-engine-agent-absence-projection-test.mjs` | Absent background agents: when the engine stops reporting a background agent and no final state has arrived, the row is marked absent and this package owns every decision about it, so all clients agree. One predicate says whether a row is absent (the mark, not the status, decides). An absent row keeps its last known status, never counts as running, and is never counted as completed, failed or stopped; its elapsed time stops at the last moment it was seen, and its sentence says it may still be running. The end-of-turn sweep never settles an absent row (or a resident one). A row that comes back, or a real final state for the current cycle, clears the mark; a late final state from an earlier cycle does not. Absent rows are never removed at the short grace window. After the hard limit (30 minutes from the last time they were seen) the host is asked for the background-agent registry reading of each row: only a reading that the agent has ended or is not listed lets the row go, and each removal is returned as a fact the host must act on and announce; a reading of running, unknown, missing or unrecognised keeps the row and schedules nothing, so no standing poll is created. Until a registry reading is available every absent row stays. A row someone is viewing is held and reported separately only once the registry confirms it is gone. The row sentence, the removal sentence and the late-result sentence come from one place, never state an outcome or that the agent finished, and escape control characters in names, in the engine's removal word and in the late-result status. An end-to-end cell drives the real fleet projection and the real absence channel through every decision. |
|
|
434
434
|
| `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. |
|
|
435
|
+
| `scripts/run-gate-interrupt-safety-test.mjs` | The two safety properties of running the suites themselves. The collecting runner no longer kills a suite outright when its time limit expires: it sends a termination signal first, waits a grace period for the suite to clean up, and only then kills it — and it records the suite as timed out however it exits, so a suite that exits cleanly after the signal is still not counted as passing. The limit and the grace period can be widened for a single suite in `gates-manifest.json` (an optional entry with a written reason; a malformed entry, including one with a misspelled field, stops the runner before any suite starts instead of silently falling back to the default, and an entry filed under a misspelled name is reported by this guard rather than ignored), and the negative-control suite is widened there. Every property is exercised on byte-for-byte copies of the real runner and the real negative-control suite in a throwaway directory: a suite that honours the signal finishes within the grace period, one that ignores it is killed when the grace period ends, and the summary lines are unchanged; once a suite has exited the runner waits at most a short drain window for its output pipes, so a child process that inherited them and outlives the suite neither turns a passing suite into a timeout nor holds the runner past the limit, the grace period and that window; the negative-control suite, interrupted in the middle of a rehearsal by any of the three signals or by the runner's own time limit, restores the file byte-for-byte, leaves no backup behind, stops the rehearsed guard together with anything it started, regenerates the build output (checked on a small project in the throwaway directory), and exits with 128 plus the signal number; a rebuild that would overrun the grace period is cut short and its compiler stopped; a backup left behind by an earlier run — next to a later target, loose in the source tree, or inside a symlinked dependency directory — makes it refuse to start without touching anything, naming the file and how to restore it. |
|
|
435
436
|
|
|
436
437
|
Each suite carries a floor that only moves up — a refactor that stops executing a group of
|
|
437
438
|
assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
|