@sema-agent/client-core 0.86.0 → 0.88.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +128 -0
- package/README.md +18 -12
- package/dist/adapt/arms.js +9 -0
- package/dist/adapter/downstream/eventToSdkMessage.js +14 -1
- package/dist/adapter/downstream/terminalToSdkResult.js +4 -1
- package/dist/adapter/downstream/wiringManifestView.d.ts +2 -0
- package/dist/adapter/downstream/wiringManifestView.js +1 -0
- package/dist/adapter/runStream.js +75 -2
- package/dist/adapter/types.d.ts +3 -0
- package/dist/decideReceipt.d.ts +1 -1
- package/dist/decideReceipt.js +1 -1
- package/dist/detachWire.d.ts +6 -1
- package/dist/detachWire.js +22 -1
- package/dist/displayUntrusted.d.ts +4 -0
- package/dist/displayUntrusted.js +1657 -573
- package/dist/engineErrorCodes.d.ts +2 -0
- package/dist/engineErrorCodes.js +1 -0
- package/dist/engineNoticeCodes.d.ts +20 -0
- package/dist/engineNoticeCodes.js +71 -0
- package/dist/engineWireFor.d.ts +1 -0
- package/dist/engineWireFor.js +1 -1
- package/dist/gateOutcome.d.ts +4 -0
- package/dist/gateOutcome.js +5 -0
- package/dist/generated/engineFactTables.d.ts +1 -0
- package/dist/generated/engineFactTables.js +3 -0
- package/dist/generated/engineNoticeTables.js +2 -0
- package/dist/hitl/crashConverged.d.ts +2 -1
- package/dist/hitl/crashConverged.js +23 -3
- package/dist/hitl/hitlBridge.d.ts +4 -0
- package/dist/hitl/hitlBridge.js +6 -0
- package/dist/hitl/parkResolver.js +17 -12
- package/dist/hitl/planReviewWire.js +30 -10
- package/dist/hitl/suspendedReopen.d.ts +7 -0
- package/dist/hitl/suspendedReopen.js +38 -7
- package/dist/hitl/toolApprovalWire.js +4 -5
- package/dist/index.d.ts +7 -0
- package/dist/index.js +4 -0
- package/dist/promptDelivery.d.ts +20 -0
- package/dist/promptDelivery.js +153 -0
- package/dist/registryConflict.d.ts +3 -0
- package/dist/registryConflict.js +44 -0
- package/dist/request/taskRequest.d.ts +1 -0
- package/dist/request/taskRequest.js +25 -1
- package/dist/resumeRefusalCopy.d.ts +8 -6
- package/dist/resumeRefusalCopy.js +41 -19
- package/dist/rewindArchiveCapability.js +5 -0
- package/dist/seam.d.ts +12 -1
- package/dist/seam.js +7 -0
- package/dist/sessionMemoryErase.d.ts +76 -0
- package/dist/sessionMemoryErase.js +432 -0
- package/dist/subagent/engineSubagentResume.js +2 -1
- package/dist/toolHistoryMismatch.d.ts +63 -0
- package/dist/toolHistoryMismatch.js +425 -0
- package/dist/workflowClient.d.ts +3 -2
- package/dist/workflowClient.js +3 -3
- package/docs/INTEGRATION-CLIENTS.md +598 -13
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ Renamed from **`@sema-agent/wire-cc-adapter`** (0.1.x, deprecated — see *Migra
|
|
|
35
35
|
|
|
36
36
|
## Scope
|
|
37
37
|
|
|
38
|
-
**Version:** 0.
|
|
38
|
+
**Version:** 0.88.0
|
|
39
39
|
|
|
40
40
|
- **Today** — the adapter seam, the whole `adapt()` pipeline (all 14 A-layer arms plus the
|
|
41
41
|
B/D/E tool-card layers), the notification/caps/model families, the adapter kernel (stream driver
|
|
@@ -303,7 +303,7 @@ guard still cross-checks the table by name).
|
|
|
303
303
|
| `scripts/run-web-search-backend-capability-test.mjs` | The deployment-default WebSearch backend read face (`capabilities.webSearch.backend`, engine ≥7.82.1). Same four-state discipline as the SQL and write-protection cells, with two things that are specific here and therefore guarded: a **missing key** (an older engine) and an explicit **`"none"`** (the engine says this deployment has no default search backend) point an operator in opposite directions — "cannot tell" versus "not configured" — and must never be folded; and the `none` sentence has to say both halves of the contract at once: the default scenario mounts no WebSearch tool, **and** a caller-supplied `webSearch` setting can still mount it on a single-user lane, because the capability advertises the deployment default, not whether this request has search. The backend word is read as an **open set** — the engine's closed set is typed from its own provider tuple and grows with it, so hand-copying three words here would turn a newly configured backend into "unreadable" (the narrower-than-the-mint disease this repo already logged once). `webSearch: null` is malformed rather than `none` (the mint never emits `null`), extra members never cross, an unparseable response clears the cell, a stale probe generation is dropped, the invalidation port clears to "not observed", and the open-set word is sanitised and bounded before display |
|
|
304
304
|
| `scripts/run-terminal-cause-projection-test.mjs` | The `7.64.0` wire reshape, projected. A run's ending stopped being eight parallel flat keys and became **one tagged cause** (`completed \| failed \| blocked \| paused`), and a tool call's gate stopped being four orthogonal words and became **one record** (`disposition` / `settlement?` / `origin?`). Both are read in exactly one place in this package, and this guard pins them at **two levels**, because the dangerous seam is "the reader was updated, the consumer was not": each terminal arm is checked on the reader *and* on the `subtype` / `is_error` / `errors[]` the projector actually emits. Two properties carry most of the weight. First, a terminal word this reader does not know is **never** laundered into an empty success — it lands on an `unknown` arm carrying the word verbatim, while a payload with no terminal word at all (the mock lane) keeps the success arm exactly as before, which is the one and only case the reader answers `null`. Second, the three window words (`approval_window_expired`, `denial_limit_window_expired`, `park_sla_expired`) must each be told apart by a different predicate: the previous generation collapsed all three onto one `timeout`, and re-merging them would throw away the discrimination this reshape just restored. Two byte generations are read by one reader, keyed on the discriminator upstream nailed (`"terminal" in result`): the current cause form, and the **flat** form that a current engine still emits on two lanes — replayed persisted bytes, which the service passes through verbatim rather than back-filling, and the service's own rejection envelope. A cause-form payload that also carries stale flat keys must ignore them entirely: keeping one compatibility read is what gives a single fact two sources. The same file also pins the MCP delivery verdict and HTTP status riding the wiring manifest, the four-state write-protection reading (where three of the four states mean *cannot tell*, and none of them may be printed as "there is no table"), and the park-reopen fetch identity: that predicate is asserted through the **real entry point**, since the defect being fixed was precisely a call site wired to a different predicate than the one that routed the row there. From 0.80.0 one of those three boundaries flips: the key naming **who settled a refusal** stopped being a dead byte and became part of the wire, so the check stopped scanning the build output for the word and started reading the request bodies the two decision legs actually send. A refusal attributed to the deployment's own policy carries the word; one attributed to a person, one with no attribution at all, and one carrying a word the vocabulary does not hold carry nothing — the wire has no slot for “a person decided this” other than the key's absence, so inventing one would be minting a word upstream does not have. The allow family never carries it on any of its routes, because that combination is refused before the approval is judged while the side effects of allowing have already landed, and the three refusals nobody was asked about (a card that failed, a user who walked away, an interruption) carry nothing either. A deployment that signs the bodies it accepts does not sign that word, and there is no capability bit to ask beforehand, so a refusal on exactly that ground is answered by re-sending the same decision once with that one key removed — byte-for-byte the same otherwise — rather than letting an optional note take the whole denial down with it. The guard measures that along three axes: the decision still lands and is reported as decided with the attribution handed back and a separate flag saying it never reached the wire; a caller who aborted in between gets no second request; every other refusal code, and every decision that never carried the key, send exactly once. The classification of a second failure is made from what the second body actually carried, not from what the card asked for. From engine `7.104` a synchronous submit that stops at a gate returns the engine result itself plus a three-key receipt: it now carries the tagged cause, so it reads as `paused` on the cause generation (older engines still send the flat three-key body, which keeps reading as before), and the top-level park word is looked at first, the same order the SDK documents for all three generations, so a body the server says is parked is never read as finished. A parked run row now carries its result too, and the headless reconnect path turns it into the parked terminal frame on the first attempt instead of spending its whole retry budget — both generations are pinned, including an end-to-end drive through the public reconnect entry point. |
|
|
305
305
|
| `scripts/run-auto-mode-unavailable-test.mjs` | The fact behind "you are being asked because the auto-mode classifier could not run", and the one place its sentence is minted. The cause table is a **copy**, reconciled word for word in both directions against the installed engine's own bytes — it narrowed upstream, and the guard follows rather than keeping the old shape: a table checked against something nobody ships any more is the oldest way for a guard to be green and wrong. The retirement is held from both sides — the removed table must really be gone upstream, and the removed reader and word must really be gone here — while the word that left keeps arriving cleanly from an older engine, because the reader takes the cause as an **open set**: the vocabulary belongs upstream, so a copied list here would discard a legal value the day one is added, and the value discarded is precisely "this outage is a NEW kind". The reader's one exclusion is the word the engine says it never stamps here — the classifier did run and did answer, just outside its contract, so reading it as a failure would invent an event the engine denies. That exclusion used to be derived from a second table which no longer exists; the reason for it never lived in that table, so it is now stated where it actually comes from, pinned as a **named** set (a magic literal scattered through the reader reds) and cross-checked against the engine's own verdict declaration and against the reader having exactly one such comparison. One reader serves both the live ask and its durable parked twin, since the two carry the same key path and a second copy is how two ledgers drift apart. Absence is pinned as absence — most asks never consulted a classifier at all — and the sentences are checked mutually distinct, prototype-safe, and walked end to end: an unknown word reaches the sentence a person reads (the fallback that names it verbatim) and the status reading (unavailable for this round, never a fallback to "available"), with counter-controls proving neither assertion is vacuous |
|
|
306
|
-
| `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails. From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. The server also mints a few notices of its own through the same channel; those codes live in a second table with their own audiences, kept apart from the engine mirror (which must stay equal to the engine's catalog) and reconciled against the server's published package when one is supplied, so a user-facing server notice is no longer filed under operations. One dispatcher returns the typed facts for every code that has a reader, discriminated by code and tagged with its audience — only the user-audience codes belong on a user surface — and the guard ties the dispatch table to the module's own exported readers in both directions, so a reader cannot be exported without a row and a row cannot be dropped without the guard failing. From 0.85.0 it also covers a third server-minted notice, the one saying that part of a session's saved history could not be read when the session was reopened: it is a user-audience notice, its two counts are read one by one and anything that is not a non-negative integer reads as `unknown` rather than zero (neither "nothing was skipped" nor "nothing is left" may be invented), and one extra sentence — the context is empty but the turn runs — is given only when the count of entries left is exactly zero. When the server package is supplied, the guard drives the server's own emitter for that notice and reads what it emits back through the dispatcher. From 0.86.0 the catalog grows by six engine notices — a personal rule store that could not be read and was skipped, an auto-mode classification that could not decide and where the call went, a requested permission mode that did not take effect, a permission mode answering a question on the person's behalf, a mode release past a read-deny table, and a legacy checkpoint's mode keys being migrated — each with its own typed reader: branching words (the classifier's cause and destination, the reason a mode did not take effect) are read as closed sets, descriptive words pass through as written, and a malformed notice reads as absent; two of them are checked against notices produced by the engine's own emitters. The unresolvable-ask notice's optional remedy sentence is carried verbatim when it is a non-empty string — never trimmed, rewritten or filtered by cause — and its absence (older engines, other causes, an empty or non-string value) leaves the rest of the view unchanged; a notice without it reads byte-for-byte as before, and the engine's own mint function drives three cases through the dispatcher. |
|
|
306
|
+
| `scripts/run-engine-notice-catalog-test.mjs` | The engine-notice catalog and its audience table. Whether a notice deserves a person's attention is not decided by whether this end happens to have a phrasing for it — that drifts with each client's build order — but by whether the engine minted the code into its own written catalog; the audience row answers the separate question of *who* the fact is for, since an operations fact pushed at an end user is noise and a user-facing fact buried in an operator log is something withheld from the person who could act on it. Both tables are reconciled against the installed engine's own artefacts in both directions and pinned in lockstep with each other, unknown codes fall back to the conservative operator side, and catalog membership is tested on the raw value so a code carrying control characters cannot impersonate a registered one after sanitizing. The reader for a dropped MCP injection keys on its own code alone and treats a missing session, server or reason as absence rather than throwing at a read site. A reverse pin enforces the upstream's single-mint contract: the engine composes those sentences from the host's facts, so a copy of them appearing in this package's source or build is a second source that would drift, and fails. From 0.84.0 it also covers the reader for the two read-directory grant notices: it recognises only those two codes, needs the tool call id to match a card, passes the rejection reason through as written, and treats only the granted notice as evidence that a directory was added; a granted notice without both the directory and the spelling the engine now holds, or with a scope other than `exact`, is not read at all, and the scope word is pinned to the engine's type at compile time. The server also mints a few notices of its own through the same channel; those codes live in a second table with their own audiences, kept apart from the engine mirror (which must stay equal to the engine's catalog) and reconciled against the server's published package when one is supplied, so a user-facing server notice is no longer filed under operations. One dispatcher returns the typed facts for every code that has a reader, discriminated by code and tagged with its audience — only the user-audience codes belong on a user surface — and the guard ties the dispatch table to the module's own exported readers in both directions, so a reader cannot be exported without a row and a row cannot be dropped without the guard failing. From 0.85.0 it also covers a third server-minted notice, the one saying that part of a session's saved history could not be read when the session was reopened: it is a user-audience notice, its two counts are read one by one and anything that is not a non-negative integer reads as `unknown` rather than zero (neither "nothing was skipped" nor "nothing is left" may be invented), and one extra sentence — the context is empty but the turn runs — is given only when the count of entries left is exactly zero. When the server package is supplied, the guard drives the server's own emitter for that notice and reads what it emits back through the dispatcher. From 0.86.0 the catalog grows by six engine notices — a personal rule store that could not be read and was skipped, an auto-mode classification that could not decide and where the call went, a requested permission mode that did not take effect, a permission mode answering a question on the person's behalf, a mode release past a read-deny table, and a legacy checkpoint's mode keys being migrated — each with its own typed reader: branching words (the classifier's cause and destination, the reason a mode did not take effect) are read as closed sets, descriptive words pass through as written, and a malformed notice reads as absent; two of them are checked against notices produced by the engine's own emitters. The unresolvable-ask notice's optional remedy sentence is carried verbatim when it is a non-empty string — never trimmed, rewritten or filtered by cause — and its absence (older engines, other causes, an empty or non-string value) leaves the rest of the view unchanged; a notice without it reads byte-for-byte as before, and the engine's own mint function drives three cases through the dispatcher. A fourth server-minted code, the user-facing notice that a model endpoint could not be connected to, is registered with its audience and has its own reader: the failure code, the endpoint, the remedy sentence and the session are required, the name of the proxy variable is optional, the code and remedy pass through as written, and the server's sentences are never copied into the package. With server fixtures, the server's own decorator mints the notice for a direct and a proxied connection and for every connect-failure code it knows, and stays silent when a response arrived or the call did not end in error; the audience table is reconciled against both fixture generations, a row newer than the older fixture being required to be absent there. From 0.88.0 the catalog grows by one operator notice — a provider request whose tool turn broke the pairing rules was repaired on its way out — with a typed reader: each repair's form word passes through as written (the engine does not publish that word list from its entry point and has announced more words), the call ids are read as a non-empty list, any malformed entry makes the whole notice read as absent, and notices produced by the engine's own repair law and emitter read back field for field. The model-unreachable reader reads only the notice detail's own properties, like the other server-minted readers. |
|
|
307
307
|
| `scripts/run-tool-roster-projection-test.mjs` | The leg's tool roster — what the engine says it actually mounted and what face each tool wears — replacing three word lists that were only ever an estimate taken from one traffic capture against one pinned engine. The reader copies the engine's own all-or-nothing discipline: a roster whose row cannot be read, or whose declared count disagrees with the rows, is dropped whole rather than handed over short, because a consumer reading a short roster concludes the missing tools are not mounted — the upstream says in as many words that this is worse than sending nothing. A malformed *face* on a row (path target, render hints) drops only that face, since a face is not an identity. Shims are built strictly from roster rows and never guessed from a tool's name, and an axis that cannot be read stays absent rather than defaulting to `false` or `never`, which would render "unknown" as "safe". For run-time changes the guard pins the one hard rule in the contract: a digest that does not match is **not** a rejection — the carried roster is the new state regardless and only the summary becomes unusable, because refusing the swap would leave the consumer holding a stale roster forever. One reading here answers a question that the terminal state structurally cannot: whether this run was assembled with any file-and-shell tools at all. The engine's terminal vocabulary says a run finished, not whether the work got done, so an orchestrator that waits for the end and then guesses has nothing to guess from — while the assembly manifest already said it at the start, one row per mounted instance with the single condition that mounted it. The reading is three-state and both folds are refused: a roster that is readable and carries no such row is the engine stating a fact, while no roster at all is not that fact — the static half of a manifest never carries one, and an older engine reports rosters without naming the mount condition at all, where an empty count would be a statement about the reader rather than about the run. Those two are kept apart in the reason the reading carries, and the wording for every unknown case is checked never to claim the run had no tools. The same roster now decides the tool list on the first line of a non-interactive run: the host holds that line until the roster arrives and lists exactly what the engine mounted at the start of the run, in mount order. The guard runs a real assembly frame through the projection into the decision, and pins that the host falls back to the estimate only once the roster is known not to be coming — a manifest without one, an unreadable one, model output or the run's end arriving first — rather than on a timer alone (model activity counts, including a model call that is still waiting or retrying; an error line the stream synthesizes when a run fails before assembly counts as the run ending), that a sub-run's manifest is never mistaken for the run's own, that an empty roster is taken as the engine's answer rather than as silence, and that the wait bound covers both sequential default budgets the engine gives an external tool server to connect and list its tools. The holding logic itself lives in the package as a small per-run gate — buffer, decide once, release the held messages in arrival order, then pass through — and the guard drives real stream output through it to pin that the release happens exactly once, at the manifest, releasing exactly the held prefix. The ordering itself also lives in the package as a stream wrapper, and the guard checks the final output a consumer reads: the first line is always the tool-list line, a message that arrives while that line is still being built comes after it, a timer firing races nothing out of order, a source that ends or fails before the decision still gets its first line and held messages out before the error, and an early exit closes the source. From 0.84.0 the roster-derived sentence source no longer throws on a value it does not recognise, including a reading of the manifest's `hands` section passed by mistake: it answers the same "not stated" sentence as the `hands` reader, from one shared source, and its six known sentences do not change. From 0.86.0 that first line also says where its tool list came from, in two added keys the wrapper writes onto the host's own line object in place: a source of `engine-roster`, `estimate` or `host-static`, and, only for an estimate, which of the four reasons made the roster known not to be coming. The guard checks every decision path end to end, that the host's object keeps its identity and every other key byte for byte, that stale values of the same two keys left by the host are replaced (a stale reason is removed rather than left behind), and that a host line that cannot be written is passed through untouched instead of failing the first line. |
|
|
308
308
|
| `scripts/run-permission-rule-issue-codes-test.mjs` | The rule-lint refusal codes an engine reports when it will not compile a permission rule. The SDK publishes neither a schema nor a type for them, so the package mints the table from the engine's own bytes and the guard pays the cost of that copy instead of leaving it to somebody remembering: it parses the codes the engine actually mints and reconciles them against the table in both directions, so a code added upstream (the user would see a bare code) and a code only the package believes in (a branch that can never fire) both fail. It also reconciles the table plus a small retired ledger against the engine's declared union, which is deliberately not the same set — one member was renamed and its old name is still declared — so reviving a code the engine will never mint again is impossible and a future stale member shows up immediately. Sentences are pinned one per code, mutually distinct, and split by family: a rule that is wrong and a rule that is legal but unsupported on this lane are different next steps and may not share a sentence. The engine's own message rides along as prose — sanitized and capped after escaping, never matched on |
|
|
309
309
|
| `scripts/run-gate-vocabulary-test.mjs` | The two gate vocabularies — who denied a call (`DeniedBy`, eleven words across two engine generations) and who asked about it (`AskOrigin`, eleven) — together with the one place their sentences are minted, so the same denial does not read three different ways across three clients. The tables are copies, not opinions: the asker table is reconciled in both directions against the lists the installed SDK and engine packages publish, and the denier table is the installed engine's own list followed by the one word only older engines still send, checked so that every word the SDK declares is present, the retired word really was a word of an earlier generation, and every word ahead of the SDK comes from the engine — a word added upstream (nobody renders it, the user sees a bare code) and a word only the package believes in (a branch that can never fire) both fail. Every word must carry its own literal sentence and no two may collide, including the sibling pairs the upstream deliberately split apart — an organization store and a personal rule store being unreadable send you to different people, and the two tighten origins exist precisely to name which layer of engine logic asked. The two fallbacks are pinned distinct because an unknown word means different things in each: a denial layer this build does not know may have been added by a newer engine or may come from a damaged record, so its sentence says it cannot tell which instead of asserting damage; the asker vocabulary is genuinely open (the server only checks for a non-empty string, so an unknown word just means the client is older than the engine). Alongside them sits an **uplift anchor** rather than a third table: the reason a call was decided the way it was is a distinct semantic face from who denied it and who asked, one upstream has not mirrored into the SDK at all, and one whose newest member — a shell command allowed because it only reads — has no sentence anywhere yet. Minting the union here would create the second drifting source the day upstream publishes it, so the guard instead asserts the **absence** from both ends: the SDK declarations carry no such union near that word, and the installed engine’s own list does not carry the word either. The engine end fires first, on the batch that raises the dependency, which is exactly when the ownership question should be answered; the SDK end fires when the mirror lands. Either red is the work order to mint the sentence, never a reason to delete the anchor. A fourth mint now sits beside the three tables and is not a table at all: a single presence-only fact — that no saved rule and no standing posture can retire this question — earns one sentence, taking no argument precisely so a caller cannot mistake it for a second kind of mandate, pinned distinct from every sentence the tables mint, pinned never to point at rule-writing, and pinned not to overclaim the stronger neighbouring demand that a person rather than a configuration must answer; it must not say the question is asked every time — an answer for this one call may come from the person, a hook or an automatic check the deployment runs — and its wording is checked against the engine package's own description of the mandate. A fifth table joins them from 0.80.0: the thirteen words for **how a wait ended**, mirrored in both directions from the engine's own declarations — the table's owner — with the wire SDK's copy held alongside as a second witness that must match it word for word and in order, so the day the SDK falls a generation behind, that is what turns red rather than the mirror silently following the wrong source. The newest of them says a deployment's own policy answered the card — not a person, and not “nobody could be asked” — so the guard pins it apart from both neighbours by behaviour, feeding every one of the thirteen words through all five named predicates and checking which word makes which one speak, rather than what any predicate returns. Two of the thirteen also decide how a refusal is filed in the session transcript; that mapping is minted once and reused by both of the package's own entry points, and anything outside those two words yields nothing rather than a guess. Since 0.83.2 a sixth list covers the word a mandated question stands on (`APPROVAL_MANDATE_WORDS`, six words): it must equal the engine's own list word for word and in order, membership is exact, the card reader `readApprovalMandate` answers only for an own key holding one of the six words, and each word has one fixed sentence explaining why the question must be confirmed — six distinct sentences that never point the reader at writing a rule, never promise a question every time, never claim only a person may answer, and repeat no other sentence the package mints. The list is also pinned against the engine's type at compile time in both directions, while the published build references no engine package at all: every `.js` and `.d.ts` file in the build is scanned, and the same scan is first shown to fire on references planted in a scratch directory. |
|
|
@@ -317,7 +317,7 @@ guard still cross-checks the table by name).
|
|
|
317
317
|
| `scripts/run-absence-fold-census-test.mjs` | A package-wide census of the "absence folded into a positive outcome" defect shape, so that fixing the six sites this release does not merely move the shape somewhere else. The defect is defined by position, not syntax: a fallback position (the unconditional tail return, the `default:` arm, the literal minted when there is nothing to pass on, the value returned from an error path) may only say `unknown` or stay absent, never a positive word. Detection walks the syntax tree of every source file, so comments, strings and multi-line spellings cannot hide or fake a hit, and covers five forms: the right arm of `??` / `\|\|`, the else arm of a ternary, the first return of an explicit `default:`, a `catch` block or `.catch(() => …)` arrow returning a healthy value, and a function whose last statement returns a positive word after other returns. Every remaining hit must be registered with a written reason, an unregistered hit fails the gate naming the file and line, the registered count must equal the real count so a cleared site cannot leave a spare allowance behind, and the gate proves its own teeth behind a fence (a failed self-proof refuses to report any count): each form injected into an in-memory copy must add exactly one hit, two correct spellings are pinned as non-hits, and samples inside comments or strings do not count. It also pins the headline site: the fleet panel projection no longer mints an `end` with `isError: false` on absence |
|
|
318
318
|
| `scripts/run-device-executor-management-capability-test.mjs` | The engine's device-management self-description (`capabilities.deviceExecutor.management`, engine ≥7.88.0), read the same four-state way as its five sibling capability readers: an absent `management` key is reported as not reported (never folded into `false`; an older engine really ships the lane object without it), an absent `deviceExecutor` key is likewise not reported, `deviceExecutor: false` is the lane being absent, presence is judged by own-property not truthiness, the value must be a strict boolean, the tee never throws and drops stale generations, and the package owns the verdict on whether the `/v1/devices` management verbs are usable (`yes` only when present and true, `no` when present-false or lane-absent, otherwise `unknown`) |
|
|
319
319
|
| `scripts/run-run-cancel-context-test.mjs` | The run record's `cancelContext` side-note (engine ≥7.87.3) read structurally, and the cause of a `turn_aborted{engine_error}` classified from machine-readable evidence only: `cancelled` (code `cancelled`, with the cancel-time context when present) / `engine_error` (any other failure code, passed through verbatim) / `run_still_live` (the record is not terminal — a dropped stream is a client-side fact, not the run's cause) / `unknown` (never guessed). An absent `cancelContext` reads as *not reported*, never as "not cancelled"; `elapsedMs` is never folded to 0. |
|
|
320
|
-
| `scripts/run-suspended-reopen-projection-test.mjs` | The durable `suspended` event's `reopened` key read as three distinct states — `reopened` (with the engine's code, verbatim), `not_reopened` (an explicit `null`), `unstated` (key absent or unreadable) — and carried on the HITL bridge's active gate (`currentGateReopen()`), re-read on every `suspended` and cleared with the gate. |
|
|
320
|
+
| `scripts/run-suspended-reopen-projection-test.mjs` | The durable `suspended` event's `reopened` key read as three distinct states — `reopened` (with the engine's code, verbatim), `not_reopened` (an explicit `null`), `unstated` (key absent or unreadable) — and carried on the HITL bridge's active gate (`currentGateReopen()`), re-read on every `suspended` and cleared with the gate. The stream driver also turns a `suspended` frame that reads `reopened` into the chrome arm `suspended_reopened` (the reading, the single-source reopen sentence when the code is in the reopen family, the frame's event id, and the run id only when the frame itself carries one), pinned end to end: one notice per event id on a stream, emitted in stream order on a replay from the first frame, nothing for a cursor that starts after it, and an unchanged transcript when the host has no chrome sink or the sink fails. |
|
|
321
321
|
| `scripts/run-panel-identity-normalization-test.mjs` | One background subagent has two ids on the wire — the fleet row id tail and the `task_progress` task id (its transcript id). Every panel event goes through one funnel that rewrites the `tick` / `end` task id onto the fleet row's id once a `fleet-row` has registered the key (`transcriptId` first, `parentToolCallId` as the fallback), carrying the original as `wireTaskId` and marking `taskIdOrigin`; an unbound tick whose row has not arrived yet waits one beat (bounded) and is released verbatim on the next tick / `end`, when the buffer is full, or after `MAX_HELD_WIRE_TICK_BEATS` other fleet-row / end / sweep events (a `sweep` itself leaves it alone: there is no row to settle yet); a normalized `end` that carries no cycle identity borrows the registering row's, so a late close of a revived task is recognized as stale; the key table is an LRU (a task that keeps ticking is never evicted by newer registrations); the residency mark migrates with the id and both keys are cleared on settle — except that a stale (previous-cycle) terminal never clears the revived row's mark — so the notification lane can clear it. Once a tick has been delivered verbatim under its UUID, that UUID is the subagent's key: later fleet rows and fleet-side ends are rewritten onto it (the tail kept in `wireTaskId`), so a consumer sees one row in every arrival order; a late tick from a previous cycle is dropped rather than folded into the revived row. |
|
|
322
322
|
| `scripts/run-prompt-assembled-projection-test.mjs` | The `prompt_assembled` frame (one prepare's prompt-assembly manifest) projected to an internal arm and then to the additive `prompt_assembled` chrome event — the per-section / per-block **character** counts, the mounted tool names and `totalChars`, each key present only when the engine really sent it (the frame's `constitution` is deliberately not carried: no consumer asks for it today, and every published key is a contract to keep). The manifest carries **no token counts** anywhere upstream, so this projection mints none: a token figure derived from characters would be an invented number, and the engine's own estimate lives on `context_usage.sections[].tokens` (same id wordlist, joinable). Bad rows are dropped one by one, and a face that loses every row reads as an absent key rather than an empty array — so an absent face means only "this event carries no readable view of it" (an absent upstream key, an empty array and a fully filtered list all land on the same shape) and is never reported as a diagnosis about the engine. `blocks[].id` and `sections[].id` are two different wordlists with a many-to-one relation, and the token join against `context_usage.sections[].tokens` only holds when both sides carry a section view. A frame with no readable composition key at all is malformed, ids and slots are read as an open set, one chrome event per frame with zero transcript rows, several prepares per task are all handed over (de-duplication — "take the last one" — is the host's move), and the lane is told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). Both entry points obey the same rule: the adapt layer rebuilds every row too, so a host pipeline (or a replayed transcript) that feeds the raw frame straight into `adapt()` cannot smuggle extra keys (`tokens`, digests, aliases), a negative `chars` or a `null` row into the chrome payload, an empty array does not count as a composition face, the identity keys are snapshotted once on both paths (read exactly once each, a throwing accessor rejects the whole frame — reading one twice is what lets an accessor frame land on a different lane on each path), and the two paths are compared verbatim so the two readers cannot drift. |
|
|
323
323
|
| `scripts/run-compaction-outcome-projection-test.mjs` | The `compaction_outcome` frame (a compaction that did **not** end as compacted: mooted by the task ending, failed, …) projected to an internal arm and then to the additive `compaction_outcome` chrome event — `outcome` required and verbatim (open set), `trigger` / `reason` present only when the engine sent a non-empty string, malformed frames dropped, zero transcript rows, the lane told honestly (`parentToolCallId` ⇒ subagent lane; a frame attributable only by `sourceTaskId` / `bgAgentId` is not surfaced on the main lane). |
|
|
@@ -338,7 +338,7 @@ guard still cross-checks the table by name).
|
|
|
338
338
|
| `scripts/run-display-body-test.mjs` | The engine wraps text it hands a model in a fence — an opening marker naming the payload, the payload itself, and a closing marker — so the model reads it as data and not as instructions. That fence is minted and read in one place here, which makes stripping it for a human reader this package's job rather than each shell's: a shell that renders the envelope verbatim is showing a person a defence that was written for a model. The reader answers with a discriminated union — fenced, with the label and the payload, or not fenced, with the text as it came in — and it reaches that answer through the **same** matcher the mint side registers, never a second copy of it; the guard proves that by walking the syntax tree of every source file and requiring exactly one literal carrying the marker text, and by requiring the reader's own body to contain no matcher of its own. Eighteen shapes are run through both entry points and required to agree line for line. Anything the package does not recognise — a near-miss in the wording, a hyphen where the marker has a dash, a different case, an opening marker with no close, a close before an open, a truncated close, or any non-whitespace byte outside the pair — comes back unfenced with the input returned **verbatim**: no guessing, no trimming, no repair, because a half-stripped envelope puts a sentence on screen that nobody wrote. Only the outermost layer is removed, so a nested fence, or one forged inside the payload, survives byte-for-byte in the body — those bytes are part of what the engine said, not part of this protocol. Nothing else is washed: control characters, leading and trailing whitespace and a twenty-thousand-character payload all pass through untouched, and so does the label, because sanitising and length-capping belong to the mint point that puts a string on a screen and a passage of text must not have two launderers. A value that is not text is answered with **nothing at all** rather than with an empty payload: the reader never stringifies it, never calls its `toString`, and never emits `[object Object]`, and it does not hand back a body of zero length either — an empty payload is a real reading (a fence can legitimately wrap nothing, and an empty string is an empty string), so folding "there was no readable text" into it would leave a caller unable to show a degraded line at all. Those three stay apart: no text yields nothing, an empty string yields an unfenced empty payload, and an empty fenced payload yields a fenced one with its label. The `fenced` discriminator is always present on a reading, and the label key exists only on the fenced arm, so a missing label is never rendered as an empty one. One shape needed more than the whole-string match this started with. When the engine reports back from a delegated run, the fence is only **one section** of the report: ahead of it sit a frame header, a handful of optional field lines and a section label, behind it a closing instruction addressed to the model, an optional internal identifier and a usage block. A matcher anchored to both ends of the input answers *not fenced* on that, and the whole scaffold - written for a model - goes on screen. The reader therefore also locates the fence **inside** a recognised report frame, using four anchors that are always present and always byte-for-byte fixed, and hands the located slice back to the same single matcher rather than a second one; the guard assembles its corpus from the engine package actually installed (the fence from that package's own constructor, the frame lines read structurally out of the minting file and then compared byte-for-byte with what this package registers), so a rewording or a reordering upstream turns the guard red the day it lands. Three readings are pinned one cell each: a fenced result section yields the payload the delegate actually wrote; a partial-findings section yields that text and says which of the two it is, with the prefix line kept out of the payload; and a section the engine filled with its own no-text sentinel yields an empty payload with a reason, which stays distinguishable from an input that was simply an empty string. Whatever surrounded the fence is returned alongside rather than dropped - the bytes before it, the slice itself and the bytes after it reassemble into the input exactly - and the payload never contains a line of the frame. The criterion deliberately does **not** enumerate the lines outside the fence: several of those field lines carry interpolated untrusted text and cannot be told apart from prose, so requiring every one of them to be recognised would mean that a single new field line upstream sends every failed report back to being unreadable, and a failed report is exactly when a person most needs to read what the delegate said. Fourteen negative shapes hold the line against the easy widening, strip anything that looks like a fence: a fence sitting in ordinary text, a missing frame header, a different sentence where the closing instruction belongs, a report with no closing instruction at all while the fence sits at the very end, a header and label in the wrong order, mismatched open and close tags, an open with no close, a bare unfenced section, a stray line between the label and the fence or between the fence and the closing instruction, an entirely absent section, a no-text sentinel with another line after it, a partial-findings prefix followed by something that is not a fence, and the whole report in carriage-return line endings all come back unfenced with the input verbatim and no frame at all. Above all, **provenance is not in the text**: locating a fence inside a frame happens only when the caller states where the bytes came from, because four anchors can only recognise a shape and never prove an origin. The default reading is byte-for-byte what it was before, so a passage of ordinary prose that happens to quote a report - with real warnings on either side of the quoted part - is returned untouched and those warnings stay on screen; a caller that does state the origin gets the located reading, and even then every surrounding byte comes back alongside. Twenty-three malformed origin values fall back to the narrower default without throwing: a near-miss in case, a camel-cased spelling, the right word padded with spaces or tabs or a newline or a zero-width character, a string wrapper object, an object whose `toString` or `valueOf` reports the right word, and an object whose converters both throw. The guard compares the caller’s argument for **exact equality** and nothing else — no trimming, no stringifying, no calling the value’s own converters, because that would let a value of unknown provenance choose its own lane — and a source-level cell requires that the argument reach the comparison unreassigned and unnormalised, with its own two-way check that those patterns speak. The two accepted words are read off the published type rather than copied into the guard; adding a third word later is a type-compatibility change for any caller that switches exhaustively on them. Each of the three anchors must also be **unique** in the text, and ambiguity means the reader declines. The reason is not hypothetical: the frame header interpolates the task's own description verbatim, and the engine only requires that description to be a string, so it can carry newlines and a complete set of protocol lines. Any rule that picks one candidate out of several can therefore be made to pick the planted one, hiding the real result among the surrounding bytes - a guard cell reproduces exactly that, with a planted section and a real one, and requires the reader to decline and the real text to stay on screen. Two reports back to back are the same ambiguity and are declined the same way, with both payloads left visible; a delegate that quotes any one of the three anchor lines inside its own answer also falls back to the input verbatim, which is the registered cost of the rule, and a discrimination cell shows the same corpus reads cleanly once the quoted line is gone. The two legs are not the same shape either: the forked one ends at its closing instruction with no trailing bytes at all, and that real shape has its own cell. The witness arm reads each leg only inside its own array of lines, decodes every extracted literal to its **runtime** value rather than trusting the spelling in the source, and fails loudly if it cannot - an escape rewrite upstream leaves the runtime label unchanged while the spelling diverges, and since that same extracted value builds the corpus and serves as the expectation, trusting the spelling would close a self-proving loop. The two payload labels are therefore also pinned in the guard and compared against what was extracted, so an equivalent rewrite stays green while a real rename turns red the day it lands. A delegate that quotes the frame lines inside its own answer does not move the location, and those quoted lines survive in the payload byte-for-byte |
|
|
339
339
|
| `scripts/run-display-cap-order-test.mjs` | The order in which untrusted text is sanitised and length-capped, across every mint point that puts an engine- or database-supplied string on a screen. The sanitiser rewrites each invisible character as a six-character escape, so capping the **raw** string first and escaping afterwards hands the screen six times the width that was budgeted — a forty-character allowance becomes two hundred and forty. The guard does not hardcode that allowance, because each mint point wraps its field in different fixed prose and the prose moves: it anchors on the deciding quantity instead, feeding one benign and one control-character input of the same length through the same mint and requiring the second not to come out longer. That criterion is immune to wording changes and stays sensitive to the expansion, and it is `<=` rather than `==` on purpose — a correct escape-then-cap backs the cut off a partially-consumed escape token, so the control-character line is legitimately the shorter of the two, and demanding equality would score that avoidance as a regression. Each mint is bracketed by two positive controls (the input really reaches the screen; the cap really engages) and the expansion predicate is shown to turn red against a deliberately cap-then-escape reference, so an all-green run cannot mean the guard simply measured nothing. The shared mint point is checked directly for the two avoidances it owes — never splitting an escape token in half, which would leave something on screen that looks like the beginning of a complete answer, and never splitting a legal surrogate pair, which would manufacture the very lone surrogate the sanitiser exists to catch |
|
|
340
340
|
| `scripts/run-seat-task-request-origin-test.mjs` | Where every field of the seat lane's send-message payload comes from, and whether it actually lands anywhere. The seat payload is a closed interface this package mints itself, and most of its fields are meant to ride verbatim onto the engine's request body — two facts nothing used to connect, so both directions could drift in silence. A seat field could be named after a request position that does not exist, in which case a client writes to it, the wire carries it, the engine ignores the whole key, and the screen shows a switch that does nothing; conversely a new request position could arrive with no seat to sit in, which is **structural** absence — the closed set *is* the carrier, so a decision missing from it has nowhere to be put at all, the same shape logged when the effort dial had no seat. The guard turns each field's origin into data: either it names the request position it forwards to, or it is declared seat-local with a written reason, and the two are mutually exclusive. Forwarding claims are then checked against the **installed** SDK's type declarations, parsed rather than restated — a hand-copied list of position names would only ever prove that two transcriptions agree. The parser is held to reading top-level positions only, since a nested option object's inner keys would otherwise be mistaken for positions of the request itself, and it proves that discrimination on synthetic input before any verdict is given. The two subagent fields carry a standing regression pin, and the retention window's inner keys are read from the declaration the same way, so a seat that offers a tunable window cannot offer one the wire has no room for |
|
|
341
|
-
| `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header |
|
|
341
|
+
| `scripts/run-wire-auth-source-test.mjs` | **When** the outbound credential is read. A literal string is consumed at construction — the transport captures it in a closure and every later request reuses that one copy — so once the engine is replaced by another session and the credential rotates, a long-lived client keeps presenting the old one and the only way out is to rebuild the client along with everything hanging off it. The credential position now also accepts a getter that is called **once per outbound request**. The guard anchors on the deciding quantity, which is not "was the getter called" — reading once at construction and reusing the result would satisfy that too, and is exactly the shape being removed — but *which read produced the value on the wire*: it changes the getter's answer between two requests through the same client and requires the second request to carry the new one, and it requires construction to read the getter **zero** times. The three-state credential semantics are replayed per request rather than assumed: on loopback an unavailable credential sends **no** authorization header at all rather than a fabricated one, off loopback it sends the fail-closed anonymous identity so the deployment answers with an honest 401, and the guard shows a single client moving between those states across successive requests. A getter that throws is fail-soft — the request still goes out under the no-credential branch, because a broken credential port should not take the whole wire down, and the exception may itself carry credential material. The same-origin relay form is checked to stay out of the getter path entirely, and every request is checked to keep the credential in the authorization header only — never in the URL, never in another header. The same getter form is also accepted by two further entry points, and each is judged by its end result. For the workflow activity ledger, connection identity now follows the credential's *source* rather than the value read from it: a string is compared by value, a getter by reference, the same-origin relay declaration by its mode, and a change of form counts as a new source. The guard rotates the getter's answer between two warm-up calls and requires that no second stream opens and that the ledger keeps every frame it had already collected — before, a host could only pass a freshly read string, so each rotation looked like a different connection and the ledger was replaced by an empty one. A reconnect of that same ledger and the monitor's detail reads must then carry the rotated value, and the getter must be read exactly as many times as requests go out, so comparing identities never reads the credential. As the negative control, a different getter returning the same value opens its own stream, replaces the ledger, and the monitor holding the previous getter stands down instead of taking the slot back. For the detach cancel fallback, arming with a getter reads it zero times; the pick-up the host calls on its signal path reads it at that moment and still hands out only the two established forms, so an unchanged copy of the host's cancel leg sends the credential that is current at send time. The three-state semantics match the wire client (no authorization header on loopback when nothing is available, the anonymous identity elsewhere), a throwing getter never makes the pick-up throw, and string or loopback arms are handed back as the very same object. The widened inputs are checked with the compiler: the getter form is assignable, every previous form still is, the pick-up still narrows to the two forms, and illegal forms are rejected |
|
|
342
342
|
| `scripts/run-subagent-durable-divert-test.mjs` | The side-channel that keeps a **sub-agent's** content out of the leader's transcript, on the replay leg. A content frame stamped with a parent tool-call id belongs to a child, and rendering a child's tokens as the leader's own text is the pollution this divert exists to prevent — but the predicate only listed the four **live** frame shapes, while the durable leg replays the same segment in its **aggregated** form. Those frames fell straight through onto the main projection path, which is how a reconnect or a resumed session ended up with the child's answer printed as the leader's. The anchor is unchanged and shared: the parent tool-call id is what says whose frame this is, and whether the frame is an increment or a whole segment has nothing to do with whose it is — judging the two shapes separately is exactly how one of them got missed. Folding the aggregate into a synthetic increment would have been the smaller diff and the wrong one: an increment means *append*, so a segment that already streamed live and then replays whole would be counted **twice**. The two are kept distinct and the aggregate absorbs instead — a whole segment whose prefix is what the buffer already holds replaces it, which also makes a redelivery of the same frame idempotent, and a prefix that does not match falls back to appending both rather than deciding on the engine's behalf which version counts. Segment boundaries stay with the tool frames rather than moving into the aggregate arm, since closing there would turn a second replay of one segment into a second entry, and the increment arm is pinned to keep appending so a token run that happens to be a prefix of the next does not silently lose characters. When the host declares the non-interactive lane, a sub-agent's tool calls and results are also forwarded into the main output with their parent tool-use id (the sub-agent's text and thinking still stay out, as in the reference CLI); without that declaration the output is unchanged. |
|
|
343
343
|
| `scripts/run-subagent-content-budget-test.mjs` | The **byte** budget on the sub-agent transcript ledger. It used to be bounded only by *counts* — so many entries per child, so many children — and a count is not a budget when a single entry has no ceiling of its own: one tool result carrying an inlined attachment, or one long model answer, and a single slot sits on tens of megabytes. The guard anchors on how many bytes are **still held** after over-filling, not on whether truncation fired, because an implementation that flags the overflow without actually dropping anything satisfies the second and not the first. Dropping is required to leave a record — how much went and where the retained content now starts — and that record has to reach the render plan, because content that vanishes with no marker gives the reader a transcript shorter than what happened with nothing to say so; the record is one per child, updated in place, pinned to the front, and excluded from the budget it describes. Order matters and is checked: oldest entries go first and the live tail is trimmed only as a last resort, since taking the text the user is watching stream while older history survives is the wrong end. The total budget evicts a whole least-recently-used child rather than shaving every child, and the configuration surface is fail-loud on zero, negatives, non-finite and non-integer values — a silently ignored budget is the exact failure this exists to remove — with the rejection proven atomic so a bad second field cannot leave half a configuration behind. The defaults are checked to be a magnitude that can really be reached, since a number too large to hit is a field rather than a budget |
|
|
344
344
|
| `scripts/run-subagent-usage-projection-test.mjs` | Per-subagent usage, split by task. The engine's final accounting carries the delegated spend as **one total** — tokens, turns, task count — and no per-task breakdown, while every sub-flow turn on the stream carries its own usage. This package used to fold that away at the leader/sub-flow divide (a child's output tokens must never reconcile the leader's response length), so a client showing a subagent's detail pane had nothing to print. The split table can therefore only be accumulated from the stream, and this guard pins what that costs. The two existing leader-only arms stay **byte-for-byte unchanged** — the new arm is additive and always carries the sub-flow's own lane proof, so a host cannot mistake a child's numbers for the session window. Attribution is by the engine's own originating-task id — deliberately not a second `taskId`, which the event identity does not carry and whose absence would silently collapse every child under one parent call — falling back to the parent call id; a turn that answers neither is dropped rather than filed under an invented row, because merging two children's ledgers is worse than missing one. Cache-read tokens are read from the **engine's own shape** rather than the mirrored one, since the mirror fills that member with zero when the wire omits it and reading it there would erase the difference between *not reported* and *no cache hit*. A turn that reported no usage at all still counts as a turn and still adds its zeros — the numbers are a lower bound, and dropping the round would make the bound less true, so the honesty bit rides on the row instead and is never spelled `false`; such a round still emits its live arm, because the frame that says "this round has no account" is the one a real-time consumer most needs and the easiest one to drop. The same honesty bit also survives a terminal that carries no statistics at all: what the stream observed is unioned with what the final record says, so a run that already reported an unmeasured round cannot come out the other end looking like an exact zero. Finally the table says whether it is **partial**, and that verdict is anchored on the quantity that actually decides it: the engine's own totals. Turn count and row count must both reconcile before the table claims to cover the whole run; anything else — including totals that cannot be read — marks it partial, so the failure direction is always the safe one (a complete table called partial, never the reverse). The two accounts are kept separate and are never added together or used to correct each other. One more thing the totals cannot settle: the row key has **two namespaces** — the originating-task id and the parent call id it falls back to — and nothing upstream promises they are disjoint, so the same literal can name one child's identity and another child's parent call. Accumulation therefore keys on the origin as well as the id; the delivered table still keys on the bare id, and a cross-namespace clash is merged into one row that says so, with the partial verdict forced, because a row count and a turn count can both reconcile while the attribution behind them is wrong. The table itself is likewise a **per-stream snapshot** handed to the terminal projector by value rather than left on the caller's context: the three terminal projectors are public, so a host may drive one run through the stream and project another's terminal directly on the same context, and a table left behind would be attributed to whoever projects next — silently called complete whenever that run's own totals happen to match. Without a snapshot, both table keys are simply absent |
|
|
@@ -364,10 +364,10 @@ guard still cross-checks the table by name).
|
|
|
364
364
|
| `scripts/run-abortable-sleep-test.mjs` | The shared `abortableSleep(ms, signal)` leaf (consumed by `workflowClient.ts` and `agentSession/backgroundView.ts`'s poll backoff): normal timeout resolution, immediate wake-up on `abort` mid-wait, `clearTimeout` really firing on that path, and a post-resolve late abort staying a no-op |
|
|
365
365
|
| `scripts/run-durable-card-display-keys-test.mjs` | The durable approval row's two display keys survive the row→card recast in `surfaceFsApprovalAndDecide`: `governanceForced` stamps on strict `true` only (absence is "no evidence", never `false`), the row's rule offers (`ruleOffers`, or the older `ruleSuggestions` key that earlier servers send) pass through the same shape-narrowing reader as the live-frame leg and land on the **read-only** card key `ruleOffersReadOnly` — plus a standing pin that the durable leg never stamps the redeemable `ruleOffers` card position (the `/decide` body has no rule slot; offering a "don't ask again" option there would be an affordance nothing can honour), and a section for the parked twin of the classifier-unavailable fact: the upstream declares that key on the parked action itself, verbatim and under the same name as the synchronous ask, so this leg reads it rather than guessing a carrier name the way the deliberately unprojected keys must. The guard drives both legs with the same cause and asserts the card ends up byte-identical either way — the observable consequence of one reader serving two key paths, and the thing that silently diverges the day someone writes a second copy. Its own reach is printed rather than implied: what is proven is the package-boundary promise "on the row ⇒ on the card", not that today's engine flattens that key onto the pending row. A further section covers the two display facts the recast had been dropping for far longer. One of them the row has carried all along under a DIFFERENT NAME than the live frame uses — the frame puts it at the top level, the row nests it under the risk descriptor — and that difference in name is exactly why it went unnoticed; unlike the keys this leg deliberately refuses to project, its carrier is witnessed in the engine's own artefact rather than guessed. Neither is decoration: the shell's stand-aside arm reads them, so a call that matched a remembered allow rule which could NOT silence it looked like an ordinary ask on the durable path and was auto-approved with no card at all. Both land on the SAME card slot the live leg uses (one shape for the ends), verbatim bytes, present only when non-blank, never folded into an empty string — and the guard pins the discipline in both directions, including that a top-level key the upstream row does not actually have must still not grow this position. The security-class approval bit (`requiresRealApproval`) rides a parked row's card when the row carries it at top level, on strict `true` only, while look-alike nested carriers are ignored; this is pinned with a constructed row, because today's pending list does not carry the bit yet. A second, separate bit (`irreversibleParkGate`) marks a parked card whose row sits on the irreversible-ask gate kind — a gate-kind fact that covers asks the engine flagged for real approval at the first decision plus safety-tightened gates, with a known engine gap for approval demands raised only on a storage recheck — on the park path only, on the exact gate word only, and never in place of the real bit. Since 0.83.2 the recast also carries the row's ask origin (`origin`) exactly as the live-frame leg does — a non-empty string, verbatim, open vocabulary, never invented or defaulted — and the word a mandated question stands on (`mandate`), accepted only when it is one of the six known words and read back through `readApprovalMandate`; a word outside that set, or a malformed value, leaves the card without it. The word never adds the separate mandated key to the card, but a known word on its own makes `approvalIsMandated` answer true, because the engine only sends the word as the whole reason for that bit. An `origin` inherited through the row's prototype chain is not read, and since 0.83.4 neither is an inherited `mandated` on either leg — both legs stamp it from an own strict `true`, the same reading the seat crossing uses; the single judge `approvalIsMandated` reads `mandated` and `ruleOffersAbsence` the same way, so an inherited key no longer makes it answer true. |
|
|
366
366
|
| `scripts/run-session-memory-status-test.mjs` | The session **memory-status** read face (S-53): the two judgements three clients would otherwise each get wrong. First, *same status, different code* — this route's 404 carries two unrelated meanings (`not_found.session` = unknown or non-owned session; `not_found.route` = a pre-7.53 server that has no such route at all), so dispatching on the **status** would report "your deployment lacks this surface" as "your session does not exist". The verdict is anchored on `errorCode`, the two 404s are pinned to **different** verdicts, and — the load-bearing negative control — a 404 carrying **no** code falls to `failed` rather than guessing either way, since a wrong guess in either direction is a false statement a user would act on. 501 is allowed a codeless fallback because both of its arms mean the same thing here, and `capability.*` stays split from `feature.*` because those two share a status while their dispositions are opposite. Second, *absence means something different per key*: `optOutSource` and `lastCaptureAt` are legitimately absent on a **healthy** session (a zero-history session really is `{captureOptedOut:false, committedCount:0, foldedCount:0}` with no degradation at all), so reading absence as "off/none/0" asserts something unprovable. Two combined readers are pinned: capture opt-out is read from **both** its keys (a record-store fault yields `indeterminate`, never `active` — the difference between "your conversation is being remembered" and "nobody knows"), and last-capture is a **three-state** read whose discriminator is the *other* key, because `lastCaptureAt`'s absence alone covers both "ledger unreadable" and "genuinely no contributions" and therefore decides nothing; the two shapes are pinned to different verdicts so a single-key read turns red. The thin wrapper is the only IO: it never throws, drops malformed keys to absence rather than trusting them (an unreadable value must answer "don't know", never render as truth), refuses to spend a request on an empty `sessionId`, and passes `signal` through untouched |
|
|
367
|
-
| `scripts/run-crash-converged-projection-test.mjs` | The `crashConverged` read face on `GET /v1/approvals` (L-38): what the *previous life* of a crashed local engine left behind, projected for every client. Three judgements are pinned. First, **absence is not an empty list** — a missing key (an older server, deps not present, or a carrier that is not an array at all) returns `undefined`, and the client renders nothing; an empty array returns a present zero-count object, which is the server actually saying "none". Folding the first into `{total:0}` would have the client assert "nothing was left behind" on a surface a person uses to decide whether it is safe to re-run something — the worst possible direction for a false statement — so the two cases are pinned to different **return shapes** and a test asserts the two verdicts are unequal. Second, bucketing is a **four-term conjunction**: `orphanState === 'pending'` *and* `resumeSafe === true` *and* both approval-evidence keys (`originalDecision`, `decidedAtMs`) absent. A fifth term rejects any row carrying an **accessor**, and accessors are never invoked at all — reading one means synchronously running someone else's code, and `catch` catches throwing, not *never returning*, so a looping getter would pin the startup thread forever (the row cap does nothing against that shape). The same rule covers the three untrusted reads outside the row as well — the envelope's `crashConverged` key, the carrier's `length`, and every numeric index are read as own property *descriptors* and only data descriptors are used, so accessors and prototype entries read as absent and are never invoked. Such a key is treated as absent: if it was a required field the row is counted as dropped, if it was optional or additive the row survives without it. That also closes the ordering attack, since spreading runs getters in property order and an earlier one could `delete` the approval evidence before it is ever copied (measured before the fix: such a row reached the resume-safe bucket), and the check therefore moves ahead of the read, onto the property descriptors — from which the snapshot is then built directly, because checking descriptors and *then* spreading is two independent observations of the same row, and a non-throwing proxy can make the two `ownKeys` calls disagree (first showing `originalDecision: 'approve'` so the row reads as plain data, then omitting that configurable key so the snapshot loses the evidence; measured before the fix: the dangerous row reached the resume-safe bucket after exactly two enumerations, and after it, one). Keys are written with `Object.defineProperty` rather than plain assignment, because `'__proto__'` is a legal own enumerable key and `o['__proto__'] = x` does not store a value — it calls the prototype setter, letting a row whose own properties are all plain data (so the accessor gate never fires) inject a prototype whose `sessionId` getter deletes the approval evidence from the snapshot during validation; `defineProperty` fires no setter, so the key survives as ordinary additive data and the snapshot keeps `Object.prototype`. A row that simply arrives with a custom prototype is treated the same way, since the snapshot only enumerates own properties: approval evidence sitting on the prototype would never reach it, and a perfectly ordinary object with no proxy and no accessors could otherwise be called safe to re-run — real bodies come from `JSON.parse` and always carry `Object.prototype`, so nothing genuine trips it). Validation itself runs on a **null-prototype** dictionary and the bucketing verdict is carried out of that same pass rather than re-read from the delivered row, because every property lookup on an ordinary `{}` reaches `Object.prototype`: a polluted `sessionId` getter there would delete the approval evidence from the snapshot mid-validation and send the row to the safe bucket (measured before the fix). The row handed to the client is still an ordinary object — the null prototype is an implementation detail of the check, not of the value) — real JSON bodies are all data properties, so only a middle-layer-synthesised payload ever trips it, and it too lands in the human bucket rather than being dropped. The `decided` arm means the human had already approved and side effects may be half-landed, so it always goes to the human bucket, as does `resumeSafe === false` and — the last two terms — any row whose own fields contradict each other, since `pending` claims nothing ran while that evidence says somebody pressed approve. Deciding "not safe" costs one extra question (recoverable); deciding "safe" wrongly has somebody re-run work that already partly happened (not). A 2x2 truth table pins that exactly one cell is resume-safe, so reading either key alone turns red, and the contradictory rows are routed to the human bucket rather than dropped — they are real orphans, and the ones most worth showing. Third, unreadable rows are **dropped and counted**, never thrown and never passed through: the product is declared as `CrashConvergedRow`, so letting a row missing a required field — or carrying one of the wrong type — past would be a lie at the type level, and the closed literal discriminators (`decision` / `cause` / `orphanState`) decide family membership rather than being an open vocabulary. The measuring stick stops at the **type** floor, though: degenerate-but-well-typed values (`ts: NaN`, an empty `toolName`) are kept, because swallowing a real orphan over a decorative field is the worse direction, and the one deliberate exception is `approvalId`, which must be non-empty to be a row identity at all. `dropped` is kept separate from `total` so unreadable rows never inflate "N approvals were affected"; each row is a **one-shot snapshot** — every own enumerable key is read exactly once, and validation, bucketing and the handed-back value all read that same snapshot, so additive upstream keys survive while a **non-idempotent** getter (one that never throws, just answers differently on a second read) can no longer erase the approval evidence between the check and the bucketing (measured before the fix: such a row landed in the resume-safe bucket while its checked value was `"approve"`). Hostile carriers are counted rather than allowed to reject: **every** touch of the carrier is guarded — envelope property reads, `Array.isArray` itself (it throws on a revoked proxy), the `length` read, each indexed read and each row's property reads — and a traversal that dies halfway returns absence rather than a half-counted total. A row that cannot be read never takes the batch with it: its own shape check is inside its own guard, so one revoked-proxy row costs a `dropped` tick rather than collapsing the whole projection to absence — which a client would have read as "this deployment does not offer the surface". Traversal goes by **numeric index, never the carrier's own iterator protocol**, because `for...of` hands the carrier the question of which rows exist: an array carrying an overridden `Symbol.iterator` can yield nothing (measured before the fix: a real orphan became `{total:0}`, which a client reads as "the server said there are none") or swap a dangerous `decided` row for a safe-looking one (measured: `fake-safe` was returned in place of `real-danger`). Row count is capped at 100000 and the cap is checked **before** the walk: requiring only a non-negative integer `length` does not stop a proxy trap reporting a billion, and this surface runs on the startup / `--resume` path, where a synchronous spin freezes the thread (measured before the cap: twenty million rows took 18.3 seconds and twenty million index reads; a billion does not come back). The honest boundary is stated rather than overclaimed — a proxy can still lie in its `length` or index traps, which is the same thing as a host injecting a lying transport — and the widening of `ApprovalsResourceLike.list()` is proven **additive** by really running tsc over a legacy `{ pending }` mock *and* over the real `AgentClient` path — the projector takes `unknown` precisely because a parameter shaped as "an object with an optional `crashConverged`" is a TypeScript weak type that the installed SDK's own `list()` return shape shares no property with, which only a real-client compile would have caught — with a known-red control so a clean run means the checker spoke |
|
|
367
|
+
| `scripts/run-crash-converged-projection-test.mjs` | The `crashConverged` read face on `GET /v1/approvals` (L-38): what the *previous life* of a crashed local engine left behind, projected for every client. Three judgements are pinned. First, **absence is not an empty list** — a missing key (an older server, deps not present, or a carrier that is not an array at all) returns `undefined`, and the client renders nothing; an empty array returns a present zero-count object, which is the server actually saying "none". Folding the first into `{total:0}` would have the client assert "nothing was left behind" on a surface a person uses to decide whether it is safe to re-run something — the worst possible direction for a false statement — so the two cases are pinned to different **return shapes** and a test asserts the two verdicts are unequal. Second, bucketing is a **four-term conjunction**: `orphanState === 'pending'` *and* `resumeSafe === true` *and* both approval-evidence keys (`originalDecision`, `decidedAtMs`) absent. A fifth term rejects any row carrying an **accessor**, and accessors are never invoked at all — reading one means synchronously running someone else's code, and `catch` catches throwing, not *never returning*, so a looping getter would pin the startup thread forever (the row cap does nothing against that shape). The same rule covers the three untrusted reads outside the row as well — the envelope's `crashConverged` key, the carrier's `length`, and every numeric index are read as own property *descriptors* and only data descriptors are used, so accessors and prototype entries read as absent and are never invoked. Such a key is treated as absent: if it was a required field the row is counted as dropped, if it was optional or additive the row survives without it. That also closes the ordering attack, since spreading runs getters in property order and an earlier one could `delete` the approval evidence before it is ever copied (measured before the fix: such a row reached the resume-safe bucket), and the check therefore moves ahead of the read, onto the property descriptors — from which the snapshot is then built directly, because checking descriptors and *then* spreading is two independent observations of the same row, and a non-throwing proxy can make the two `ownKeys` calls disagree (first showing `originalDecision: 'approve'` so the row reads as plain data, then omitting that configurable key so the snapshot loses the evidence; measured before the fix: the dangerous row reached the resume-safe bucket after exactly two enumerations, and after it, one). Keys are written with `Object.defineProperty` rather than plain assignment, because `'__proto__'` is a legal own enumerable key and `o['__proto__'] = x` does not store a value — it calls the prototype setter, letting a row whose own properties are all plain data (so the accessor gate never fires) inject a prototype whose `sessionId` getter deletes the approval evidence from the snapshot during validation; `defineProperty` fires no setter, so the key survives as ordinary additive data and the snapshot keeps `Object.prototype`. A row that simply arrives with a custom prototype is treated the same way, since the snapshot only enumerates own properties: approval evidence sitting on the prototype would never reach it, and a perfectly ordinary object with no proxy and no accessors could otherwise be called safe to re-run — real bodies come from `JSON.parse` and always carry `Object.prototype`, so nothing genuine trips it). Validation itself runs on a **null-prototype** dictionary and the bucketing verdict is carried out of that same pass rather than re-read from the delivered row, because every property lookup on an ordinary `{}` reaches `Object.prototype`: a polluted `sessionId` getter there would delete the approval evidence from the snapshot mid-validation and send the row to the safe bucket (measured before the fix). The row handed to the client is still an ordinary object — the null prototype is an implementation detail of the check, not of the value) — real JSON bodies are all data properties, so only a middle-layer-synthesised payload ever trips it, and it too lands in the human bucket rather than being dropped. The `decided` arm means the human had already approved and side effects may be half-landed, so it always goes to the human bucket, as does `resumeSafe === false` and — the last two terms — any row whose own fields contradict each other, since `pending` claims nothing ran while that evidence says somebody pressed approve. Deciding "not safe" costs one extra question (recoverable); deciding "safe" wrongly has somebody re-run work that already partly happened (not). A 2x2 truth table pins that exactly one cell is resume-safe, so reading either key alone turns red, and the contradictory rows are routed to the human bucket rather than dropped — they are real orphans, and the ones most worth showing. Third, unreadable rows are **dropped and counted**, never thrown and never passed through: the product is declared as `CrashConvergedRow`, so letting a row missing a required field — or carrying one of the wrong type — past would be a lie at the type level, and the closed literal discriminators (`decision` / `cause` / `orphanState`) decide family membership rather than being an open vocabulary. The measuring stick stops at the **type** floor, though: degenerate-but-well-typed values (`ts: NaN`, an empty `toolName`) are kept, because swallowing a real orphan over a decorative field is the worse direction, and the one deliberate exception is `approvalId`, which must be non-empty to be a row identity at all. `dropped` is kept separate from `total` so unreadable rows never inflate "N approvals were affected"; each row is a **one-shot snapshot** — every own enumerable key is read exactly once, and validation, bucketing and the handed-back value all read that same snapshot, so additive upstream keys survive while a **non-idempotent** getter (one that never throws, just answers differently on a second read) can no longer erase the approval evidence between the check and the bucketing (measured before the fix: such a row landed in the resume-safe bucket while its checked value was `"approve"`). Hostile carriers are counted rather than allowed to reject: **every** touch of the carrier is guarded — envelope property reads, `Array.isArray` itself (it throws on a revoked proxy), the `length` read, each indexed read and each row's property reads — and a traversal that dies halfway returns absence rather than a half-counted total. A row that cannot be read never takes the batch with it: its own shape check is inside its own guard, so one revoked-proxy row costs a `dropped` tick rather than collapsing the whole projection to absence — which a client would have read as "this deployment does not offer the surface". Traversal goes by **numeric index, never the carrier's own iterator protocol**, because `for...of` hands the carrier the question of which rows exist: an array carrying an overridden `Symbol.iterator` can yield nothing (measured before the fix: a real orphan became `{total:0}`, which a client reads as "the server said there are none") or swap a dangerous `decided` row for a safe-looking one (measured: `fake-safe` was returned in place of `real-danger`). Row count is capped at 100000 and the cap is checked **before** the walk: requiring only a non-negative integer `length` does not stop a proxy trap reporting a billion, and this surface runs on the startup / `--resume` path, where a synchronous spin freezes the thread (measured before the cap: twenty million rows took 18.3 seconds and twenty million index reads; a billion does not come back). The honest boundary is stated rather than overclaimed — a proxy can still lie in its `length` or index traps, which is the same thing as a host injecting a lying transport — and the widening of `ApprovalsResourceLike.list()` is proven **additive** by really running tsc over a legacy `{ pending }` mock *and* over the real `AgentClient` path — the projector takes `unknown` precisely because a parameter shaped as "an object with an optional `crashConverged`" is a TypeScript weak type that the installed SDK's own `list()` return shape shares no property with, which only a real-client compile would have caught — with a known-red control so a clean run means the checker spoke Every row in the `needsHuman` bucket also carries a closed-set `needsHumanReason` (`approved_then_interrupted` / `pending_unsafe` / `unstable_row` / `contradictory`) derived from the same single read that bucketed it; `resumeSafe` rows never carry it, a same-named key on the supplied row is overwritten or removed, and the bucketing itself is unchanged. The `cause` discriminant is a closed set of two words — a crash, or an orderly shutdown that cut an unfinished step off — and neither the bucket nor the reason ever reads it: every combination buckets identically under both words, a word outside the set still drops the row, and the row reaches the client with its word unchanged and no sentence attached. An exhaustive switch over `cause` that only knows the older word fails to compile, by design. When server fixtures are present, rows minted by the server's own audit store are projected for both generations: the newer one yields the shutdown rows and an untracked step's pending orphan that is no longer marked safe to re-run, the older one yields neither. |
|
|
368
368
|
| `scripts/run-self-orchestration-denial-test.mjs` | The three judgements behind a **denied self-orchestration request** (server 7.57.0), each of which all three clients would otherwise get wrong on their own. First, whether to retry at all is a **conjunction that may not be loosened**: HTTP 501 *and* an `errorCode` that is **exactly** `capability.self_orchestration_required`. That code shares its shape with every other `capability.*` 501, so dispatching on the prefix would drag "some other capability is not wired up" into the retry arm — those requests do not become acceptable once the two keys are gone, so the client would spend a request and then tell the user the wrong reason. Negative controls cover all four directions: a sibling `capability.*` code, a truncated or suffixed variant of the right one, a codeless 501 (it decides nothing, so it decides nothing — no guessing), and the right code under 500 / 400 / 503 or a string `"501"`. The classifier reads structurally rather than by `instanceof` (a host may inject its own transport; across realms or duplicate SDK instances an understandable error would read as unreadable), so a class instance, a bare `{status, errorCode}` literal and an error carrying those fields on its **prototype** all reach the same verdict — and a hostile proxy or a throwing getter yields `null` instead of throwing, because this classifier runs inside a `catch` block where anything it throws escapes the caller's own guard. Second, removing the intent is a **structural** operation, not wording: `selfOrchestration` sits at the top level while `ultracode` sits under `settings` — two different stamping legs — and a client hand-writing `delete` will miss the second one, which costs the user the same failure twice. The single stripper is pinned to touch exactly those two: other `settings` sub-keys and their values survive byte for byte, `deferTools` is left alone (pulling `Workflow` out would be a behaviour change, not a removal of intent), additive unknown keys survive at both levels, the input object is never mutated, `settings` is only dropped entirely when `ultracode` was really there and nothing else remains (an already-empty one is left as is), a non-object `settings` is not touched at all, an `ultracode` that only exists on the prototype does not count, and the whole thing is idempotent. The end-to-end leg runs a real `buildTaskRequest` product through it and asserts the stripped body still passes the registration gate key by key. Third, on the capabilities body, **absence is not "switched off"**: a pre-7.57 server has no `workflowsGate` key at all, so reading absence as "the engine says no" asserts something the server never said, and the mirror-image disease is folding an **unrecognised** `denial` into `null`, which would have the client render "nothing was denied" when the truth is "denied, for a reason I do not recognise". Five shapes are pinned — caps unreadable, gate absent, closed-set member, unknown value, accessor — with the unknown arm carrying the raw token (or an empty one when the value is not even a string) and never collapsing to `null`. All four untrusted reads go through own **data descriptors** only, and the guard pins the getter invocation count at zero, since `catch` catches throwing but not *never returning*; a descriptor trap that throws and a revoked proxy both yield honest absence rather than an exception — though *what* absence means differs by field, and the guard pins that split rather than a blanket rule: an accessor on `workflows`, `workflowsGate` or `engineCan` reads as absent, while an accessor on `denial` reads as `{unknown:''}`, because a key that is **not there** is the gate saying "nothing was denied" whereas a key that is there but cannot be read is "denied, and I could not read why" — folding the second into the first is exactly the false statement this face exists to prevent. Two further pins came out of an adversarial review. The exported retry list is **frozen at runtime**, not merely `as const`: the verdict hands out that same reference, so any consumer splicing it once would poison every later verdict in the process — the guard asserts `Object.isFrozen`, that four different mutation attempts leave it byte-identical, and that a verdict issued *after* those attempts still carries the original two entries. And the classifier reads `denial` only **after** both criteria have passed, since it is not a criterion but an extra field on the verdict: the guard pins the getter invocation count at zero for any error that does not match and at most one for an error that does. The scope line is drawn explicitly rather than overclaimed — "no getter ever runs" holds for `projectWorkflowsGate`, which reads **wire JSON** where every field is an own data property by definition, but not for the classifier, which reads a **thrown value** that may well be an SDK `APIError` class instance carrying `status` and `errorCode` on its prototype; insisting on own data descriptors there would report a perfectly readable error as unreadable, so that side promises only that it never throws. A final pin covers the **integration document's own worked example** rather than the library: the shipped SDK's `tasks.stream()` is an `async` generator, so calling it issues no request at all — the POST happens inside `streamRaw` on the first iteration, and a `try` wrapped around the `stream(...)` call itself can never catch the 501. A client following a submit-shaped recipe on the streaming leg would never run the classifier, and the whole strip-and-retry path would silently do nothing. The guard drives the **real** `TasksResource` against a fake transport, offline, and pins both halves: the synchronous leg is in flight the moment it is called, the streaming leg has issued zero requests after the call and raises on the first `next()` — and it does so through the **real** error path, with `openStream` returning an actual 501 `Response` that the SDK's own `errorFromResponse` turns into the typed error, pinning the `openStream`→`errorFrom` call order so a transport that stops minting `errorCode` cannot pass. The documented recipe is then **executed** rather than keyword-counted: exactly one retry, a second body that really lost both keys while every other setting survives byte for byte, the caller's own request object left untouched, one disclosure and only one, a second 501 propagating with the request count still at two, and — after the first 501 — an abort leaving the count at one with nothing disclosed. A last leg is type-level: `stripSelfOrchestrationIntent` carries an SDK `TaskRequest` overload, because the wide `Record<string, unknown>` form erases the caller's type and the document's "strip and resubmit" line would not compile without an unsafe cast; a real tsc run over a virtual file proves both the narrow and the wide path, with a known-red control — and it compiles the document's two recipes **verbatim**, extracted from the section itself, because a recipe that does not compile is a recipe that was never given: `{ transientOk: true, signal }` is a TS2379 under `exactOptionalPropertyTypes`, which no amount of prose review had caught. The last thing pinned is the one that would have been quietest of all: the SDK's `stream()` returns only on a `done` or `failed` frame, so a stream truncated mid-run — or yielding nothing at all — ends the `for await` just as normally as a completed one. The documented `runOnce` therefore tracks whether it ever saw a terminal frame and raises when it did not, the guard's success fixture emits a real terminal and asserts the handler received it, and a truncated-stream control asserts that shape is reported as a failure with no retry and nothing disclosed. That terminal-frame rule then needed one more turn of its own: the underlying reader returns *normally* when the signal is aborted, so the check as first written rewrote a user's cancellation into a generic stream fault — a client keying off `AbortError` to suppress the error would instead have shown a failure, or resubmitted. Cancellation is therefore checked first, a real-SDK case aborts from inside the handler and asserts the original `AbortError` survives with no retry and nothing disclosed, and the document is checked for that ordering. The harness runs the documented `handle` and `transcript.note` as real spies rather than pushing frames itself, the drive loop rethrows exactly as the document does, and the disclosure ledger is proven to be the caller's own array by a positive identity assertion — without which the cancellation leg's "nothing disclosed" would have been vacuously true. Each recipe is compiled **on its own**, with a preamble that declares only what a host supplies and injects no library symbol, since compiling them together let the second one borrow the first one's imports, and the preamble's own types are decoupled from what the recipes import so the "remove the imports and it must fail" control fails for the right reason — which is checked by attribution, not merely by redness. Ordering is the last thing to get right: the cancellation check must come before the truncation error but **both** must sit behind the terminal-frame test, because a cancellation that lands after the run already reported `done` would otherwise overwrite a real outcome — one that may have already had effects — with "cancelled", and a person reading that will run it again. Aborting from inside `handle(done)` and `handle(failed)` are both pinned to still report success, and the ordering assertion is anchored inside the streaming `runOnce` body rather than the section, since the section's first `throwIfAborted` belongs to the synchronous recipe and would have made a reversed streaming recipe pass — and that ordering check is now anchored on the TypeScript AST rather than on text, since a comment reproducing the two statements in the right order let a genuinely reversed body pass. One more timing fact had to be written into the recipe: a single SSE read buffers several frames and the SDK yields them back to back, so checking the signal only after the loop lets a cancelled run keep consuming the rest of the chunk — measured, an abort inside `handle(turn_start)` still swallowed the `done` that followed and reported success. The recipe therefore re-checks after every non-terminal frame. Finally, the behavioural matrix is no longer run against a copy of the recipe: both recipes are extracted from the document, transpiled, and **executed** with injected host objects, so the disclosure assertion really exercises the document's own `transcript.note(disclose(...))` line, and the synchronous leg gets the same full matrix the streaming one does |
|
|
369
|
-
| `scripts/run-package-hygiene-test.mjs` | Everything `package.json` `files` ships — dist JS/typings and the Markdown docs — is screened line-by-line against a deny-list of strings that must never reach a public tarball (internal hostnames, codenames, person names, collaboration-process words, other repos' ledger ids and repo names; opaque ticket ids `CC-nnn` and post numbers `[nnnn]` are allowed as traceability references). Since 0.77.2 the build strips comments (`removeComments`; enforced by `run-dist-comments-test.mjs`), so what this gate screens in dist is code, string literals and type-level text. Markdown docs are enforced forward-only (CHANGELOG from 0.77.2, the integration doc from §81) because published sections are frozen.
|
|
370
|
-
| `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T \| null`, the `installHost` family returns `T \| undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the live sections of the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. Frozen sections (a numbered section whose heading carries a version already recorded as released) are historical and never rewritten, so a retired name there is allowed only if the live retirement catalog has a row for it: a retirement label, the name, the version it left in (matching the ledger) and what to use instead, next to the sentence stating that frozen sections are historical records and the catalog is authoritative. A section whose version cannot be read, or is not yet released, is judged as live. |
|
|
369
|
+
| `scripts/run-package-hygiene-test.mjs` | Everything `package.json` `files` ships — dist JS/typings and the Markdown docs — is screened line-by-line against a deny-list of strings that must never reach a public tarball (internal hostnames, codenames, person names, collaboration-process words, other repos' ledger ids and repo names; opaque ticket ids `CC-nnn` and post numbers `[nnnn]` are allowed as traceability references). Since 0.77.2 the build strips comments (`removeComments`; enforced by `run-dist-comments-test.mjs`), so what this gate screens in dist is code, string literals and type-level text. Markdown docs are enforced forward-only (CHANGELOG from 0.77.2, the integration doc from §81) because published sections are frozen. From 0.88.0 every shipped file is also scanned in full, frozen sections included, for merge-conflict markers at the start of a line (`<<<<<<< `, a bare `=======`, `>>>>>>> `, and the seven-bar base marker of three-way conflicts). From 0.88.0 a short local list of release-process phrasings (integration-branch and review-round wording) is screened as well, forward-only from the 0.88.0 sections.
|
|
370
|
+
| `scripts/run-integration-doc-freshness-test.mjs` | The **integration contract** (`docs/INTEGRATION-CLIENTS.md`) and the **changelog** (`CHANGELOG.md`) checked against the code, because a document with no guard rots — this one had a whole nest of drift found on it within a day of being written. Five directions, each a claim a machine can actually evaluate. (1) *Counting discipline*: the version-anchor row for the guard count may no longer carry a hand-copied number at all — it changes every time a guard is added, and writing it down is planting a timer; the export counts that are still hand-copied (the surface total, the test-hook count, the sentence describing the surface's internal composition, the sum of the sixteen domain rows, and the three sub-counts) are each compared against a value **derived** from `public-export-baseline.json`, which is the drift a human reviewer caught last time. (2) *Coordinates alive*: every `src/` `scripts/` `docs/` path the doc quotes must be on disk **and tracked by git** — on disk is not in the repo, and a doc that points readers at a file living only in its author's working tree sends every clone to nothing. A file landing in the same commit takes a named carve-out that **stops applying** the moment the file is really tracked (it can no longer let anything through, and the guard prints a line asking for it to be deleted) — deliberately not a red, since turning red on the very commit that lands the file would just manufacture a break that only a follow-up commit could clear. (3) *Arm tables*: the `hitl_out_of_slice` row and the `not_in_slice` fenced list must equal, name for name and in **both** directions, the case labels that really fall into those two buckets — read through the **TypeScript AST**, since which bucket an arm lands in is decided by the argument to `nothing(...)` and by nothing a comment says. The extractor is anchored to the one production projector: exactly one function named `eventToSdkMessage`, exactly one `switch (ev.type)` inside it, and no repeated case label — anything else is a broken anchor rather than a verdict, because a second same-shaped switch elsewhere in the file would otherwise overwrite the real one's conclusions and leave the doc agreeing with a switch nobody runs. The list is delimited by a machine-readable fence rather than by section headings, because the same section also names the terminal arms as a counter-example and prose boundaries cannot tell a member from a foil. (4) *Released sections are frozen*: an **append-only ledger** carries every version ever published — its number, the commit it was published from, and the sha256 of its section — and each one is checked, not just the current release, since pinning only the latest would set every earlier version free the moment the next one ships. The ledger cannot vouch for itself either: each recorded hash is **re-derived from that release commit** through git, so editing an old section and its constant together no longer passes — and the commit the row names is in turn checked against the `gitHead` npm recorded at publish time, which is the one value this repository cannot rewrite, so pointing an old version at a freshly written commit does not pass either. The *set* of versions that must be frozen comes from the registry too, so deleting an old row together with its section — which would otherwise remove that version from every set the guard looks at — is red rather than invisible. A failed registry call is classified rather than swallowed, and the classification consults the registry's own status code *before* it considers connection-level symptoms, so an auth refusal whose body happens to mention the network is still red rather than a skip. The version set is compared as full SemVer including prereleases — matching only `x.y.z` would silently drop a published `0.30.0-beta.1` and reopen the very hole this direction closes — and section headings are matched on a whole-version boundary so a stable release cannot bind itself to the release-candidate section sitting above it. Publishing itself is a two-phase protocol rather than a paradox: before a release, exactly one row may be marked pending and must name the current `package.json` version, exempt from the checks whose inputs do not exist yet; once the registry has that version the row must be promoted, so the temporary state cannot survive its own release. And because the pending exemption rests entirely on "this version is not out yet," it is refused outright when the registry cannot be reached to confirm that — an unverifiable premise is not a licence. Three reverse directions close the rest: a section claiming to be released but absent from the ledger, a ledger entry whose section has vanished, and a `package.json` version that was never frozen. Publishing appends a row; it never rewrites one. (6) *Sentinels*: the readers §5a hands hosts for "is this port installed" are checked against what the source actually declares it returns — `hasXxx()` is a `boolean`, the card port / HITL surface / wire target return `T \| null`, the `installHost` family returns `T \| undefined`. Testing a `null`-returning reader for `!== undefined` is *always true*, and a self-check that passes whether or not the port is installed is worse than none, because hosts retire their own fallback on the strength of it. Both directions are red: an implementation that changes its sentinel without the doc following, and a doc that names the wrong one. The roster covers the zero-argument readers and their `*For` variants alike — a multi-session host reads the variants, so leaving them off would let exactly the surface desktop depends on drift unwatched — and the §5a table and the §8-B checklist line are each checked against the source, because hosts tick the checklist, and a guard that only watches the prose table misses the line people actually follow. (5) *Packaging*: the README ships with the package and opens by pointing hosts at the integration doc, and the checklist names two more files as required reading before an upgrade — all three must really appear in the `npm pack` manifest, or an npm consumer follows a relative link that npmjs rewrites onto a private repository. Missing tooling never takes the whole verdict down with it: when git, npm or the registry is unreachable those legs print the `SKIPPED-SECTION` marker and the rest still judges, while a release commit the ledger names but git cannot resolve is red rather than skipped. The guard says in its own header what it does **not** do: it judges counts, coordinates, arm sets, released bytes and the packing list — whether a sentence is *right* is still for review and for the hosts to report (7) *Retired names*: every name in the per-version `removed` ledger of `scripts/export-liveness.json` may appear in the live sections of the integration doc only where a retirement note follows the name inside the same clause (or the table row's label cell is itself a retirement label); the scan is by identifier boundary after invisible text (HTML comments, link targets, reference-link labels, tag attributes) has been stripped, so a signature line in a code block, an inline `NAME = 4096`, a hidden note, or a note that belongs to a neighbouring name all count as a bare recommendation and go red. Frozen sections (a numbered section whose heading carries a version already recorded as released) are historical and never rewritten, so a retired name there is allowed only if the live retirement catalog has a row for it: a retirement label, the name, the version it left in (matching the ledger) and what to use instead, next to the sentence stating that frozen sections are historical records and the catalog is authoritative. A section whose version cannot be read, or is not yet released, is judged as live. (9) *Known-limits tables*: every row of every table in `docs/KNOWN-LIMITS.md` must split into as many cells as its header — a pipe character inside inline code still splits a Markdown table cell unless it is escaped, which silently shifts the trigger and upgrade columns. |
|
|
371
371
|
| `scripts/run-type-superset-ledger-test.mjs` | The type/wire **superset ledger** (`docs/type-superset.json`): positions this package adds on top of a CC-shaped contract, each carrying the evidence for what CC's own type surface does or does not have there. Completeness is deliberately uneven and the ledger says so. The `_sema_*` private-key class is checked in **both** directions (a key in the source that never entered the ledger is red, naming key and file; a ledger row whose key left the source is red) — but only for keys written as literals, which is the convention the ledger mandates. A key assembled by string arithmetic is beyond what any static rule can enumerate, so the guard fails closed on every shape it *can* decide (a bare `_sema_` prefix is red wherever it appears, save one pinned guard site) and leaves the rest as a convention violation for review to catch, rather than claiming a completeness it does not have. The two hand-surveyed classes are only checked for coordinate and evidence integrity, never discovered. Both directions read the source through the **TypeScript AST**, not a text scan, and they read two different sets out of it. A *key site* is an identifier, or a string whose whole value is the key — so `'_sema_decision-v2'` is carried whole rather than truncated at the first non-identifier character into some *other* key that happens to be registered. A *mention* is the key appearing inside a longer string, which is prose, not usage. The staleness direction counts key sites only: a comment or a doc sentence left behind after the last real mint site is deleted must not keep the row alive (mutation-proven — with both the comment and the prose string untouched, removing the one real site turns the guard red). And because a prefix can be concatenated or interpolated into a key no static set will ever see, the bare `_sema_` literal is refused outright rather than traced: every occurrence is red except the single inline `startsWith` guard the sanitizer needs, because the set of expressions a bare prefix can travel through on its way to a concatenation is open-ended and enumerating it is always one form behind. Every row's `host` must still resolve, with the key being a real **member of that declaration** rather than a string occurring somewhere in the same file — `governanceForced`/`delegation` each live on two different shapes in one file, and a member commented out is a member deleted, which a text-shaped check happily reads as still present. And the direction worth the most: each machine-form `ccAbsenceEvidence` is re-derived from the row's own `key` — the ledger's recorded string must match that derivation verbatim, since a row quietly witnessing `\bnever_present\b` is green forever while watching nothing (mutation-proven: the same edit passes the unbound form and is caught by the bound one) — and the check runs against the names the installed `@sema-agent/agent-types` `.d.ts` set actually declares, parsed with the TypeScript AST rather than grepped, so a name CC merely mentions in a comment cannot force the row into the manual escape hatch and thereby retire the very witness that was supposed to fire the day CC declares that name for real. That escape hatch is gated by an allowlist living **in the guard**, not the ledger, so claiming it costs a reviewed diff. Missing material never reads as a pass, and the verdict splits by *why* it is missing: no TypeScript parser skips the suite before it starts; a missing `agent-types` still runs and prints the first three directions, then exits **1** when `package.json` declares the mirror but it is not installed — a broken install must not retire the repository's only "the day CC declares this name" alarm, and reporting it as a skip would leave "never evaluated" and "evaluated, no drift" indistinguishable to the runner — and exits 3 only when nothing declares the mirror at all, which is the one case where the direction genuinely does not apply. Either way a run that evaluated no witness is never counted as one that did. When the mirror *is* present its **installed version** is witnessed too (the two declared floors must agree with each other and the installed copy must meet them), since four preflight probes are satisfied by an arbitrarily stale mirror — they prove the extractor speaks, not that it is current. Every direction carries a positive control — known-present CC symbols, a comment-only sample proving the extractor distinguishes declaration from mention, and synthetic corpora fed through the **same** discriminator function the real verdict uses, so a verdict quietly rewritten to return nothing takes its own control down with it |
|
|
372
372
|
| `scripts/run-rules-side-test.mjs` | The persisted-permission-rules lane's shared decision half. The two capability bits are checked as **two independent gates** — a worker can honestly advertise the rules lane while predating the revoke routes, and that shape must *hide* the governance surface rather than render a dead entry. Failure classification is by **disposition, not cause**: the two 404s (route missing vs. dead ticket) never share a bucket, a 503 `rule_import_retry` means *the ticket is still alive* (the opposite handling of a dead one), and a stale-cursor 400 drops the cursor and re-lists from the top exactly once — never resuming a stale keyset, never surfacing a partial governance list, and never paging past the hard cap. The persist-ack reader is **merged into** `readToolApprovalRespondAck`: the three-state verdict (`persisted` / `refused` / `unknown`) is derived only from an ack that passed the package's structural narrowing, and a half-shaped object such as `{rulePersisted: true}` with no `delivery` reads as `unknown` — the pre-merge shell read would have said `persisted`, which is precisely the double-ledger drift this file closes, so that case is pinned in reverse. The local-allow-rule skeleton pins all five narrowings (whole-tool, tool-name match, literal anchor with the escaped-star counter-example, bare interpreter prefix consulted only for Bash, and the canonical dangerous-pattern overlay) **with their refusal strings byte-for-byte** — the cli's 128-assertion suite anchors the same strings, so a one-character edit here changes observable behaviour on three clients — and asserts the parse is a pure function of its input, because the same call backs both "render the option" and "resolve the selected value" `listAllPersistedRules` needs only `list` (`RulesListFacade`; the other four methods are known to the parameter type as optional members): a synthetic consumer compiled against the built declarations passes a two-method object, a list-only literal, a `Pick` slice, the named type, the full facade, and the full or two-method facade written as an inline object literal — including arrow functions with untyped parameters — all with zero diagnostics, while a misspelled method name in such a literal is still reported. |
|
|
373
373
|
| `scripts/run-park-decision-layer-test.mjs` | The decision layer behind the "stuck behind a card" family, shared by every client. A pending row that is **not in the queue** is three states, not one: a bounded, interruptible re-probe loop distinguishes *a decidable row*, *not born yet* (no positive evidence that anything settled — an empty queue proves nothing) and *settled elsewhere*, always probes at least once so a zero budget keeps the pre-fix semantics verbatim, cuts a hung read face off at the window rather than only noticing afterwards, and reports the honest failure when the window is spent instead of inventing a decision. The decision-note reader is likewise three-state: an explicit `noteRecorded: false` outranks an echoed note body, absence renders **no line at all**, and untrusted note text is flattened and bounded before it ever reaches a renderer. Row routing anchors on the deciding quantity — a row carrying `gateKind: "human"` with `toolName: "Write"` is a tool gate, because `human` is the engine's *generic* "someone must decide", not a synonym for a question — and the queue scan refuses to surface a row it cannot positively prove belongs to this session. A chain that fails after the row vanished is split by whether a card was ever presented: decided-elsewhere, or not-its-turn-yet. A row-level single-flight makes "at most one card per pending item" structural rather than incidental. The resume three-way card pins the option **order** (the zero-effect choice sits at index 0, because the frame carries no default-focus field and a stray Enter must not attach or cancel), renders only options the wired verbs can honour, collapses every ambiguous answer to zero action, omits the liveness line entirely when the engine gave no evidence, and — when there is no card lane at all — prints three real routes and exits on a dedicated code rather than reporting success |
|
|
@@ -376,7 +376,7 @@ guard still cross-checks the table by name).
|
|
|
376
376
|
| `scripts/run-additive-key-passthrough-test.mjs` | The one disease shape behind two legs: a **closed whitelist / flattening arm** dropping a fact that is already on the wire, while both sides of the seam look correct. (1) The `task_progress` projection carries a registered **key ledger** — a frame populated with every key the service really projects is pushed through the shipped `eventToSdkMessage`, and the set of wire keys that survive must equal the registered pass-through list **name for name in both directions**, so quietly forwarding one more key is as red as quietly dropping one. `model` (the child run's model id, minted by core as `prepared.model.id` and projected by the server since 7.52.1) is the key this batch adds, with the same conditional the server itself applies: a non-empty string or no key at all — an empty string is neither a model id nor "unknown". The ledger is also checked against the fenced list in `docs/INTEGRATION-CLIENTS.md` §3d, so a doc that still says seven keys while the code forwards eight is red rather than merely stale. (2) The decide-failure arms carry the server's S-02 `currentPending` pointer key from a 409 `approval_stale` refusal onto the outcome the host reads. The reader is structural rather than `instanceof`, because the client is host-injected and the class identity is not this package's to assume; a half triple never mints (half a pointer cannot relocate anything), an empty string is not presence, and `checkpointToken` never transits. Both the allow and the deny leg are driven end to end through the real durable approval path — as is the accept-session leg, where a refusal carrying the pointer key must now re-raise instead of silently re-sending the human's answer for the **old** card as a plain approve (one decide call, pointer preserved), while a legacy 400 still falls back exactly as before — and all three flattening points must call the one shared reader — the same-shape residue check that makes "fixed one arm and left the twin" red instead of invisible. (3) The same disease growing on the REQUEST side: the `.mcp.json` → server-spec projection rebuilds each server key by key, and the settings schema deliberately leaves some keys parse-transparent — whatever JSON the file carries reaches the engine untouched, because validating them where the whole domain parses all-or-nothing would let one bad declaration take every server down silently. The whitelist had no row for the newest of them, so an operator's per-tool declarations — the ones the write fence reads — were stripped at the package boundary while both sides looked correct. The criterion is not "is that key handled" but the transparent-key table read out of the INSTALLED schema at runtime, reconciled name-for-name against this leg's ledger, so the day upstream adds a third one this turns red and forces an explicit decision. Behaviour is pinned on both transports, by object identity rather than deep equality (a rebuild would be a second judge), and malformed values must transit UNCHANGED rather than be refused here — the engine refuses them loudly and names the server, whereas a package-side judge can only swallow a declared protection quietly. Absence still mints no key, unknown keys still never reach the wire (the fix is the dropped key, not the gate), and the one transparent key this leg deliberately does not forward is a ledger entry with its own exit condition: it belongs to the deployment plane, and the day the request-plane type declares it the entry's premise is gone and the gate says so |
|
|
377
377
|
| `scripts/run-esc-halt-plan-test.mjs` | The Esc stop decision every client shares: fire the **turn-level** halt first, and escalate to a **run-level** cancel in exactly two cases — the engine itself answered with a 409 from the closed code set (it is saying "there is no in-flight turn here; use cancel for a run-level stop"), or that shot came back with no verdict at all *and* the shell can independently prove a permission card was on screen. Everything else does not escalate. The asymmetry is the whole point and every negative control guards the same direction — deciding *not* to escalate costs the user one more choice on a busy-session card (recoverable), deciding to escalate wrongly tears down a run that was alive and takes every in-flight tool with it (not). So: the closed code set is a **frozen** value, not a `ReadonlySet` — type-level immutability does not stop a consumer's `.add()`, and the guard proves it by really trying to mutate the exported value and then checking the verdict did not drift; the escalation gate is the **conjunction** of that closed set and the 409 status, since honouring the code alone lets a 500 that merely quotes it drive a destructive call; `interrupt.not_held` and `steering.not_running` are deliberately outside the set (the first means *this replica* has no live face — the run may be perfectly alive on another); an unreadable code falls to the no-escalation side; a `parked` flag never overrides a verdict the engine did give, and only strict `true` counts when it did not. The first shot is unconditional by construction — it does not consult `parked`, because the 409 it earns is exactly the verdict the gate wants — and the verdict itself is a closed machine-readable reason word, not display copy. A third escalating case was added once tearing the stream stopped reaping the run: with detach armed, a shot that never lands leaves the run going all the way to the end of the turn, so the Esc the user pressed has no effect at all and nothing on screen says so — the old behaviour had a silent backstop (tearing the stream ended the run) and that backstop is gone. The new fact is held to the same three disciplines as `parked`: it is read only where the engine gave no verdict, it is judged **after** `parked` so an existing host's reason word does not change under it, and only strict `true` counts. Absence is proven to be a no-op rather than asserted — the guard carries its own reference implementation of the previous version's table, runs the full grid through both, requires zero divergence when the new field is omitted, and first shows the comparison really does report a difference on the one cell where the two versions are meant to differ |
|
|
378
378
|
| `scripts/run-peer-frame-projection-test.mjs` | The three engine-injected lanes design/385 puts on the **one** `task_notification` carrier, which are not the same kind of thing at all: a delegated child's uplink (`agentMessage`), another session's message drained from this session's own box (`crossSessionMessage`), and a receipt about one of *this* session's own outbound messages (`crossSessionNotice`). The engine renders none of them inside a `<task-notification>` shell, so a client that projects them as the generic completion card shows "background task finished" while the model read a colleague's sentence — two faces describing different events. The discriminator is pinned to the **typed carrier being present**, never to the `summary` text: those carriers can only be minted by the engine's injection legs (the external `notify()` input is a strict subset of the payload and can wear none of them), while `summary` is filled by every notification there is — so anchoring on text would let any background task impersonate a colleague's message by writing `<agent-message from="…">` into its own summary, and a positive control asserts exactly that payload still projects as the generic card. Fail-closed has two tiers rather than one: a broken **required** field (empty `from`, a non-string `body`, a notice `kind` outside the closed set) returns absence so the caller falls back to the generic card — an honest downgrade where the user still sees the notification — while a broken **optional** field drops only itself, because losing an attribution note and losing a colleague's whole message are not the same magnitude. The provenance side record is **required and must agree on four points** (`kind` matches the lane; `from`/`taskId`/`seq` are present and equal the carrier/payload — each equality is anchored on a core mint site and pinned by the cli wire-anchor A-K24), so a carrier signed with a trusted name but a disagreeing provenance falls back to the generic card; peer bodies pass the same authority-envelope neutralization core applies (`<task-notification>` etc. are defused) so a colleague's text can never seed the resume dedup ledger. Lane precedence copies the engine renderer's own order, because the model already read the frame in that order and a client ordering of its own would put a card on screen that disagrees with the frame the model saw. Rendering and parsing of the transcript line live in the same module and are round-tripped in both directions, including a body carrying a forged closing tag (a parser fooled there hands half a message to the next row) and a quote inside the sender label (which must not forge a second attribute); the notice lane is deliberately kept **out** of the parser, since recognising it would mean anchoring the `[Cross-session …]` prefix and a user typing that same line would be rendered as engine speech. Hostile carriers are read as own **data** descriptors only and accessors are never invoked at all — `catch` catches throwing, not never returning — proven by a counting getter that must stay at zero calls, alongside a revoked proxy and a prototype-only carrier; and four legacy payload shapes assert the no-carrier path is byte-identical to before, which is the executable form of "zero difference for an older host". A re-supplied cross-session message — same task id, status and sequence as the first delivery, handed to the model again after compaction — is not rendered a second time, because the first delivery is still on the user's screen; a record with the next sequence number still renders |
|
|
379
|
-
| `scripts/run-wiring-manifest-projection-test.mjs` | The two end-user facts carried on the engine's `wiring_manifest` frame (`modelGate`: which tools this run's model gate removed and the verbatim restore hint; `autoMode`: whether auto mode is actually armed and the engine's own reason word). Projection: both sections ride as `_sema_`-prefixed superset keys, verbatim, and no SDK-named key is minted; a frame where neither section is well-formed projects to `none/not_in_slice` (no empty arm); `modelGate` needs all three keys and treats `removed: []` as a bad value rather than a reading; `autoMode` needs a boolean plus a non-empty reason that agrees with it, and the reason word is never mapped onto the capabilities vocabulary; the frame is flat (a nested `manifest:{}` wrapper is not a supply); `eventId` rides like every other arm. Adapter: exactly one chrome event on the main lane, a sub-flow frame (any `parentToolCallId`, `null` included) yields nothing, and an absent `eventId` leaves the key absent. Added at receiving time because the shell-side gate could not see this package's behaviour: two mutations (empty `removed` accepted, sub-flow gate removed) had passed the package suite untouched 0.71.0 adds sections F–I: the fourth/fifth/sixth manifest sections (`tools` via the roster reader, `hooks[]` rows dropped one by one when malformed, `lsp` absent unless `mounted` is a boolean), the `tool_roster_delta` arm (narrowed `delta`, `malformed` when `fromDigest`/`roster` cannot be read, host applies it against its own digest), the `context_usage` arm (finite-gated scalars plus `sections[]` rows dropped one by one), and the `WiringManifestMcpEntryView` rename with `MAX_AGENT_SKILLS` gone from the surface From 0.84.0 the `hands` section (whether this leg was assembled with the engine's built-in file and shell tools) is projected as well: a strict boolean `mounted` plus the engine's reason word passed through as written, absence kept as "not reported" rather than folded to either answer; a three-state reader and one sentence source share their opening and closing with the roster-derived sentence, whose bytes do not change. The sentence source answers the not-reported sentence for anything it does not recognise, including a roster-derived reading, and never throws. |
|
|
379
|
+
| `scripts/run-wiring-manifest-projection-test.mjs` | The two end-user facts carried on the engine's `wiring_manifest` frame (`modelGate`: which tools this run's model gate removed and the verbatim restore hint; `autoMode`: whether auto mode is actually armed and the engine's own reason word). Projection: both sections ride as `_sema_`-prefixed superset keys, verbatim, and no SDK-named key is minted; a frame where neither section is well-formed projects to `none/not_in_slice` (no empty arm); `modelGate` needs all three keys and treats `removed: []` as a bad value rather than a reading; `autoMode` needs a boolean plus a non-empty reason that agrees with it, and the reason word is never mapped onto the capabilities vocabulary; the frame is flat (a nested `manifest:{}` wrapper is not a supply); `eventId` rides like every other arm. Adapter: exactly one chrome event on the main lane, a sub-flow frame (any `parentToolCallId`, `null` included) yields nothing, and an absent `eventId` leaves the key absent. Added at receiving time because the shell-side gate could not see this package's behaviour: two mutations (empty `removed` accepted, sub-flow gate removed) had passed the package suite untouched 0.71.0 adds sections F–I: the fourth/fifth/sixth manifest sections (`tools` via the roster reader, `hooks[]` rows dropped one by one when malformed, `lsp` absent unless `mounted` is a boolean), the `tool_roster_delta` arm (narrowed `delta`, `malformed` when `fromDigest`/`roster` cannot be read, host applies it against its own digest), the `context_usage` arm (finite-gated scalars plus `sections[]` rows dropped one by one), and the `WiringManifestMcpEntryView` rename with `MAX_AGENT_SKILLS` gone from the surface From 0.84.0 the `hands` section (whether this leg was assembled with the engine's built-in file and shell tools) is projected as well: a strict boolean `mounted` plus the engine's reason word passed through as written, absence kept as "not reported" rather than folded to either answer; a three-state reader and one sentence source share their opening and closing with the roster-derived sentence, whose bytes do not change. The sentence source answers the not-reported sentence for anything it does not recognise, including a roster-derived reading, and never throws. From 0.88.0 the `fileHistory` section (whether this leg records file history; `on` / `disabled` today, passed through as written for unknown words) is projected as well: present only as a non-empty own string, absence kept as "not reported" rather than folded to `disabled`, malformed values dropped like the server drops them, the same view on the non-streaming submit receipt, and a leg built and published by an installed server package checked end to end when one is available. A `fileHistory` value that throws when read leaves only that section absent; the other sections of the same frame are still delivered on both the streaming and the submit-receipt paths. |
|
|
380
380
|
| `scripts/run-submit-wiring-manifest-test.mjs` | The non-streaming submit receipt can carry the run's opening wiring manifest (`TaskResult.wiringManifest`, additive on newer servers). `readSubmitWiringManifest` answers one of three: the key is absent on the receipt itself (older server, or a deployment whose engine never produced that frame) — not the same as unreadable; the key is present but cannot be read (not an object, or none of the nine sections survive); or a manifest view. The view is the same shape the streaming lane's chrome event carries (minus its two envelope keys) and is assembled by the same code path, so both lanes agree byte for byte on the same object. Liveness fields ride through untouched — this reader never mints a liveness verdict — and an operator-shaped receipt with extra governance sections reads to the same view as a tenant-shaped one. A zero-tool roster is a real reading, not an absence. |
|
|
381
381
|
| `scripts/run-rule-offers-reader-test.mjs` | The narrowing reader behind the "don't ask again" options, now a public entry point rather than a card-port-only one. Hosts that render the frame themselves (a browser has no three-way terminal card) previously had to rebuild this reader on their side, and what it carries is a **redemption-safety** judgement, not a convenience: the batch arm is redeemed by **index**, so a reader that compacts the array after dropping a malformed entry makes the k-th option a person clicked and the k-th rule the server writes two different rules. So: a bad entry is dropped **on its own** (one bad option must not make a real one disappear) while every surviving entry keeps its **original wire index** — pinned from both ends, with the bad entries leading and trailing. A batch's *members* are the opposite: any malformed member drops the whole batch, because a conjunctive batch is one "yes" to all of them and a batch missing a member is a different grant; its honest-remainder count is a reading, not decoration, so a non-integer or negative value drops the batch rather than rendering a fabricated zero. An empty array, a non-array, an over-cap array and an all-bad array all read as **absence** rather than an empty list, because an empty list renders as "there is an option lane with nothing in it". The two wire generations are ordered by a rule, not a preference: the newer key wins outright, a newer key that is **present but unreadable** does not fall back to the retired key (borrowing the older material would pass someone else's options off as this request's), and a `null` newer key reads as absence so a relaying layer that serialises "missing" as null cannot delete the whole lane on older engines. The public entry is finally reconciled against **both** card-port legs on the same material, byte for byte, so the exported reader and the one the card sees can never become two. Two upstream vocabularies used to be **hand-copied** here, and both had fallen behind: a match word outside the copied pair dropped an otherwise valid option outright, and a batch carrying a directory-read member — a member kind the copy did not know — dropped the whole batch. Both tables now come from one place upstream and are re-exported verbatim, pinned in both directions: every word in the table must be accepted (a narrower copy reds on the words it never learned) and a word constructed to be outside it must still be refused (a reader widened to "any string" reds too), with the retired-key normalising leg sharing the same narrowing so the fix cannot land on one leg only. A member whose kind is genuinely unknown still drops **the whole batch and only that batch** — never one member, because a conjunctive batch one member short renders "yes to N" as "yes to N−1", and never the card, because the honest single beside it is intact — while a member from before the discriminant existed normalises to the historical kind rather than being refused. The additive per-segment reasons ride through verbatim, drop only the row that is malformed, and stay **absent rather than empty** when nothing survives, since an empty list would read as "confirmed nothing uncovered" while the count remains the only source of truth |
|
|
382
382
|
| `scripts/run-resume-refusal-copy-test.mjs` | The **words** a client says when a resume is refused, minted once here instead of three times. The facts behind them already lived in this package; the sentences did not, so each client wrote its own — and those sentences answer a safety question (was my decision consumed, can this token still be redeemed), which is exactly the kind of answer that must not vary by client. Two closed sets meet here and the guard pins their relationship in both directions, because it is a premise rather than a coincidence: one set answers *can waiting help* (the codes the server mints a wait on), the other answers *what should a person be told*, they **intersect in exactly one code**, and each keeps a member the other must not have — a placement mismatch is never waitable no matter what arrives on the response, since its remedy is a changed argument rather than elapsed time, and a full governance window needs no prose because "you can wait" is the whole message. The overlapping code delegates its wait and its disposition to the existing reading rather than judging again: nine shapes of input drive both entry points and the two readings must agree byte for byte, the absent case included, because two judges always diverge somewhere. The wait is narrowed to the domain the server mints it in, which is **stricter than the shell's own copy was** — a zero now reads as no window rather than as "retry now", and the wake-up it would retry is an at-most-once action with real side effects. The third sentence is chosen by the disposition, never by the engine's prose: rewriting the message to either upstream branch's exact wording, with the window untouched, must leave all three sentences unchanged, while adding a window must change the third one and only the third one. The engine's folded resume refusal (`resume_blocked_by_policy`) gets its own reading — the original code as named, absent or unreadable — and a wording that neither claims nothing was consumed nor predicts whether a retry would pass. From engine `7.104` a folded admission refusal always names a code: when the underlying refusal carried none, the server fills in the generic code for its HTTP status. Those generic codes are not a specific check, so the sentence says the refusal did not name a more specific check (and shows the generic code in brackets) instead of presenting it as the check that refused the resume; the generic-code list is compared word for word, in both directions, with the server dispatcher that mints them. |
|
|
@@ -428,7 +428,7 @@ guard still cross-checks the table by name).
|
|
|
428
428
|
| `scripts/run-layering-shadow-export-test.mjs` | Same-name shadows across the first-party clients that consume this package (terminal, desktop, web and the admin console). Each client's product sources are read at the local clone's `origin/main` (or its HEAD when there is no such ref), without fetching, and parsed with the TypeScript parser; every top-level runtime export the client declares itself is compared with this package's public runtime exports. The guard prints which ref, commit and commit date it read for each client, and warns (without failing) when that commit is more than seven days old, because the result then only describes that older snapshot. A client-side declaration carrying the name of a package export means a piece of shared logic now lives in two places and can drift apart. It fails the guard unless it is listed in `scripts/layering-shadow-exemptions.json`, and a listed row must carry a retire-by version no more than three minor lines ahead (it fails once the package reaches it). It also fails once the client has removed the shadow and the row still stands. Every row also names what kind of duplicate it is, and carries the evidence that kind requires. A client copy that is this package's own function object behind a type assertion, or a thin wrapper that only passes the client's own dependencies into this package's function, must be proven so on every run by reading the client's syntax tree; such rows are due for review within two minor versions, and the guard fails as soon as the proof no longer holds. A wrapper whose extra logic the client has confirmed to be host-specific carries that confirmation (the client's own wording and where it was stated) together with a pinned review record that is due within two minor versions. A wrapper that adds its own decisions, or a same-name function that does something else, carries a review record pinned to a hash of the normalized client declaration (comments and formatting do not count); once the client's copy changes, the guard fails until it is reviewed again. A second implementation of the same logic carries a semantic diff: the client's copy is taken from the commit being read, the named declarations and the client modules they import are extracted with the TypeScript parser, transpiled and run in memory against this package's build on the same inputs, and any disagreement outside the classes registered on that row fails the guard (each class is a fixed predicate over both answers). A row can follow a client-side rename to a new name, and can be registered ahead of a client change for a limited number of versions before the client code exists. Re-exports of this package's own exports are the intended form and never count. A client tree that is not present is reported as a skipped section, not as a pass. The ruler proves itself on an in-memory fake client (planted shadows must be caught, legal forms must not), on a throwaway repository (a missing `origin/main` falls back to HEAD, a broken one is a fault rather than a silent fallback), and refuses to report zero on a client whose scan surface is empty. |
|
|
429
429
|
| `scripts/run-session-policy-deliverable-test.mjs` | Which of a batch of user-written permission rules can be written into a session’s own rule record without changing their meaning, and why each of the others cannot. The record holds whole tool names and command names only, so exactly one class maps across losslessly: a deny rule that names one tool with no qualifier. Everything else is withheld with one word from a closed seven-word list — an ask rule (the record has no ask tier), a deny rule with a parenthesised qualifier (recording just the name could block more), a rule that names a server or agent peer without naming one of its tools (for every protocol namespace the engine knows, checked against the engine package's own table) or contains a wildcard (*) anywhere (an engine that compares exact names would block nothing), an entry that is not a tool name, and a name the engine refuses at the start of every run — a retired tool name, or one containing "__" without a protocol prefix, where the prefix check is case-sensitive (once such a name is in the record, every later run of the session fails at startup until that entry is removed; the retired-name list is checked entry by entry against the engine package's own list, and a withheld retired name carries its current name when the engine says it was renamed), and — only when the caller passes the tool roster of a run — a name that roster does not list as a tool name or alias, compared exactly with letter case counted (recorded as it is it would block nothing on that deployment) — and each word has one sentence, which never echoes the rule itself; asking for the sentence never throws, even with a value that throws when turned into a string. A name with leading or trailing whitespace counts as not a tool name: the record compares exact bytes, so it would block nothing. The guard pins the batch semantics: the deliverable part is either the whole batch or empty, never a subset, so a caller cannot send half a change and report it as saved. An end-to-end check runs the engine package itself: every batch this function would deliver — the recorded vectors and a fixed-seed sample of generated names — is written into an in-memory session rule store and the next run must get past its start-up checks and reach the model, while every name withheld as refused — every retired name included — must indeed make that run fail at start-up, and every string literal in the judgement source that it withholds as refused must be one the engine package's own tables refuse. It also checks that malformed input never throws and never delivers anything (non-arrays, non-string entries, holes, a polluted array prototype, a length or index that throws, a changing index read once), that a batch which cannot be read at all is marked `unreadable: true` while an empty batch is not, so the two stay tellable apart, that each word is produced by some vector and nothing outside the list is produced, and — when a checkout of the previous in-client implementation is present — that this function gives the same answer on every recorded vector and on tens of thousands of generated rules and pairs, except for four deliberately stricter classes (whitespace-padded names; rules with a wildcard anywhere, which the previous implementation sent as exact names unless the wildcard was the whole tool part of a server rule; peer-wide rules outside the MCP namespace, which it did not recognise; and names the engine refuses at start-up, which it sent as ordinary names), whose disagreements are counted per class and must match an independent count exactly. The optional roster only ever withholds more: with no roster, or an absent or null one, every answer is byte-identical to the previous release, checked against that release's published file over tens of thousands of inputs; a name withheld without a roster stays withheld with the same word, and the same current name, when the roster lists it as a name or an alias; a name delivered without a roster is withheld as not in the roster exactly when the roster lacks it; the engine's own namespace-covering forms keep their word; and a roster or options value that cannot be read marks the whole batch unreadable instead of falling back to no roster, while a readable empty roster is a roster. An end-to-end check takes the roster from a real run of the engine package that mounts a tool with an alias: every name delivered with that roster passes the engine's start-up name audit without being reported, and every name withheld as not in the roster is one that audit reports as matching no mounted tool. When the previous in-client implementation carries its own roster check, the two are compared under a given roster and may differ only for an alias (delivered here, withheld there) and for a refused name that the roster lists (withheld here, delivered there). |
|
|
430
430
|
| `scripts/run-plugin-hooks-projection-test.mjs` | Plugin hooks: each command hook an enabled plugin declares is decided one by one as running in the engine, running in this client, or not running at all, and the page of hooks sent with a request is built from the same per-turn plan the client uses to skip its own copies, so one hook never runs in two places. Governance is judged first and always wins — a managed hooks switch-off, an untrusted workspace, safe or bare mode, or a governance read that fails sends no plugin hook and does not list it as a gap; managed-hooks-only (set directly, or through a merged non-managed hooks switch-off) keeps only managed plugins; the plugin-only customization lock does not touch plugin hooks. A hook reaches the engine only when this client started the engine on this machine, the engine reports plugin-hook support, the entry is a command, the plugin declares no sensitive option, and the event still fits the engine's per-event limits; the gate walks that matrix cell by cell, including the limit boundaries and a session goal hook counting toward them. A fact that was never read is reported as not known rather than as a fact: a host that does not say where the engine runs gets a "not known whether this client started the engine" reason, an engine whose capabilities have not been read yet gets a "not known yet whether it supports plugin hooks" reason, and the plan's two engine facts are null in those cases, not false. Events the engine never fires run only if the client says it fires them itself, and hooks the upstream behaviour itself refuses (option references in a shell-form command, an unset option in exec form, malformed entries) run nowhere. Exec-form arguments are passed element by element with only saved non-sensitive option references filled in; path placeholders are left for the executor. Sensitive option values never reach the request: with a host that wrongly supplies one, every string in the plan, the request body, the notice, the labels and the log is searched for it across eight cells. A host without the plugin reader keeps the previous request body and gets exactly one warning per settings port; plugin data that throws while it is being read (a throwing getter, a revoked proxy) is treated like a failing reader — no plugin hooks this turn, settings hooks still sent, nothing thrown; the not-running notice names the plugin and events, never a command or an option value, and escapes control characters in names. Command hooks from settings that carry arguments (a non-empty `args` array, which is the exec form, or any other non-null value) are removed from the request until the engine reports support for arguments, because the engine would otherwise drop the arguments and run the bare command through a shell; an empty `args` array is not treated as carrying arguments when the command is made only of letters, digits and `_ . / : + -` (the shell runs the same executable), so such a guard still reaches the engine, while an empty array on a command with spaces or shell characters is removed; `args` on a prompt or http entry, a null `args`, or an entry with no type is left alone, and those go out unchanged. MCP tool hooks, which the engine cannot parse, are removed only from a request built from a plan, whose not-running notice the host shows; a request built without a plan still carries them on engine-fired events, so the engine rejects the whole request loudly instead of a guard hook silently not running — the gate checks both request bodies against the engine's own hooks schema. Without a plan, every removed hook of that kind on an engine-fired event produces one warning per settings port, event and reason. Malformed entries still pass through for the engine to reject loudly, and passing null where the options object goes behaves like passing nothing; a `plan` option that is not a plan is ignored rather than turning the whole page into nothing, and a plan passed directly in place of the options object is recognised and used. A `plugin` key written by hand on a settings hook is stripped before sending (even when its value is undefined), because only hooks that come from the plugin reader may carry plugin context; the settings document itself is left untouched and a debug line records the count. The two hand-copied tables, the engine-fired event list and the engine limits, are checked against their owners. |
|
|
431
|
-
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. Since 0.84.1 an address-shaped value (`scheme://…`) that sits right after an invisible single character — a format character, a lone surrogate, a non-whitespace control character, or the outlet's own escape token for one — is no longer let through as an address: the scheme stays and the host and path are replaced, with the query and fragment masked as before; `user:<password>@` followed only by stripped units and then a boundary, a port or another `@` is masked as userinfo. A value after real whitespace or after a colour sequence is still treated as an address, and clean text stays unchanged; forms that the previous release masked are checked not to leak on the same outlets. A `user:<password>@` candidate that is the value of a credential label or scheme word — including a label or scheme word split by stripped characters, and a quoted value that goes on past whitespace — is left to the label pass, so the whole value is masked as in the previous release; both reported shapes and frozen samples from the targeted pools are checked on every outlet. |
|
|
431
|
+
| `scripts/run-display-untrusted-projection-test.mjs` | The single display-safety outlet (`displayUntrusted`) and the credential wash on the end-of-run rows this package mints. The outlet composes two credential nets (URL structure: userinfo, every query value, the fragment, path parameters and path segments that start with a known secret prefix; key/value words such as `Authorization: Bearer ...`, `Authorization: token ...` or `api_key=...`, plus well-known secret literals that appear without a label, such as `sk-...`, `ghp_...`, `AKIA...`, JWTs and the body of a PEM private key) with three character nets (control characters, bidirectional and format characters, whitespace folding). The credential nets match on a view of the text with ANSI sequences, format characters, control characters and the outlet's own escape tokens stripped, and map the result back onto the original, so colouring or an invisible character wedged between a label, its separator and its value cannot hide the value, and no stray marker is left behind. Whitespace of any length around the separator is accepted. Hosts, ports, paths, query key names and surrounding prose stay byte-for-byte, clean text comes back unchanged, the result is idempotent (also with a length cap), a length cap never splits an escape token or a surrogate pair, an invalid cap means no cap, and every net can be switched off on its own. A few narrow shapes are left alone because they name something rather than carry a value (a plain English word after `bearer` or `basic`, a back-quoted credential variable name, a plain integer after `tokens:`, a list of key names after `keys:`), each with a counter-example that is still washed. Regional flag emoji built from tag characters are kept whole. The existing single-line helpers (`escapeDisplayControlChars`, `collapseLabel`, `capForDisplay`, peer sender names and the hook failure banner) now run on the same engine and are held byte-identical to their previous output over every BMP code unit plus random strings. The approval decision-note echo, the subagent resume receipt (and its failure debug line) and the startup list of plugin hooks that will not run now also drop bidirectional and format characters (and, for the receipt, C1 controls); a note that is empty after cleaning is treated as absent. The synthetic end-of-run rows (`API Error:`, `Run stopped:`, `Model output error:`, `Outcome unknown:`) and the result frame's `errors[]` pass both credential nets before they leave the package, on the print lane and on the interactive lane (which also keeps the row-class flag); this covers a blocked reason whoever wrote it, while assistant text rows, a successful `result` and salvaged output are never touched, and a non-string `errors[]` entry is passed through unchanged. The known-secret-prefix check is a local copy of the configuration package's detector and is compared with the installed one entry by entry. Since 0.83.5 the outlet has two opt-in switches and a position read-out. `escapeBackslashes` (escape form only) writes every literal backslash as a pair, so each output decodes back to exactly one input (a real invisible character and its literal six-character spelling no longer look alike); a reference decoder round-trips thousands of random strings, the output is byte-identical to the default when the input has no backslash, a length cap is measured on the paired output and keeps the longest fitting prefix, credentials are washed exactly as in the default, and the switch is not idempotent by design (use it only at the final render). Zero-width joiners and non-joiners are kept only inside emoji sequences drawn as emoji (so not between symbols such as © or ™ that display as text) and between letters of scripts where they change the shaping (joining scripts such as Arabic, and the Brahmic family), each listed script checked both ways; the one other place a joiner is kept is right after a virama at the end of a word, the older spelling still found in Malayalam and Bengali text. Next to Latin, Cyrillic, CJK and other letters, next to modifier letters shared across scripts, at the start of a word, at the end of a word without a virama before it, or on their own they are now marked. `blanks` marks characters that look like a space but are not an ASCII space (no-break and other width spaces, the ideographic space, Hangul fillers, the blank Braille pattern) before whitespace folding, for names that must never look alike. `displayUntrustedMarks` returns the same text plus the position of every character mark; its text is compared with the outlet over thousands of inputs. With both credential nets off, the character face stays byte-identical to the previous release for input without joiners. The credential nets read escape sequences the way a terminal would when one is cut short: an unfinished colouring or character-set sequence interrupted by another one is dropped as a whole, a sequence never takes the `@` of an address as its final character, and a final character that starts a well-known secret literal (`sk-`, `ghp_`, `AKIA`, a JWT) is also read as the start of that literal; escape tokens this outlet writes are read as one unit, while look-alike text it never writes (an upper-case `\U`, or a code point it never marks) is read as plain text. A URL is cut before a run of non-ASCII blanks followed by a credential label or scheme word, Hangul fillers and the blank Braille pattern count as spaces around a label's separator, and a value that itself starts with a quoted label (`token= "password":"..."`) is left to that inner label. With `blanks` on, a blank written as an escape token right after a label is read as a blank when the value is judged, so `password:` followed by a no-break space and `missing` stays unmasked and an empty value gets no marker; the credential nets also read the text the way it looks after default whitespace folding and combine what each reading masks, so whatever the default form masks stays masked with `blanks` on (checked over a seeded corpus for the escape form, with paired backslashes and without folding; the exceptions are text that itself contains a literal six-character blank escape, which cannot be told apart from one the outlet wrote, and the dot and space marks, which cannot tell a marked blank from a real dot or space). A lone surrogate wedged between a label, its separator and its value no longer hides the value: the credential nets treat it exactly like a format character, in the machine-readable wash, in a single pass of the display outlet, and on the end-of-run rows and the result frame's `errors[]` on both lanes, while lone surrogates anywhere else are left byte-for-byte. Since 0.84.1 an address-shaped value (`scheme://…`) that sits right after an invisible single character — a format character, a lone surrogate, a non-whitespace control character, or the outlet's own escape token for one — is no longer let through as an address: the scheme stays and the host and path are replaced, with the query and fragment masked as before; `user:<password>@` followed only by stripped units and then a boundary, a port or another `@` is masked as userinfo. A value after real whitespace or after a colour sequence is still treated as an address, and clean text stays unchanged; forms that the previous release masked are checked not to leak on the same outlets. A `user:<password>@` candidate that is the value of a credential label or scheme word — including a label or scheme word split by stripped characters, and a quoted value that goes on past whitespace — is left to the label pass, so the whole value is masked as in the previous release; both reported shapes and frozen samples from the targeted pools are checked on every outlet. A scheme-less `user:password@host` whose password holds a double quote, backtick or angle bracket is still masked when a credential label or scheme word sits right before the user name, or when the user name is a plain word that is not itself a label, the host has a real shape and the address is not wrapped by that quote; quoted JSON values, quoted or bracketed e-mail addresses and quoted arrays stay byte-identical. When a label inside a quoted value ends right at that value's closing quote, the quote and separator stay outside the mask, and an address that begins the rest of a quoted value counts as an address only up to the next blank, so washing the output again never changes it. The credential stage now repeats its single scan on its own output until nothing changes (the cap counts rewrites, at most eight; clean text takes exactly one pass), so masking the output again never changes it for every input that settles within the cap — the random, exhaustive and grammar pools and the two known step-by-step constructions settle in at most three passes counting the confirming one, and the searched constructions kept in the pool in at most six; a deliberately constructed input can need more rewrites as it grows, and at the cap the last rewrite is returned, which a further pass can only mask more of. The checks cover the pass cap and how reaching it is reported, single-pass clean text, a grammar sample of the busiest shapes, constructions that once advanced one step per pass, the mark positions on the final output and how positions compose across rewrites. A username that is itself a credential label with a quote or angle bracket in the password is masked as userinfo (also with only the URL net on) unless the label's own value really reaches the `@`; a URL parameter such as `;token=...` or `,token=...` stays inside the URL piece; a private-key header the outlet cannot recognise does not hide the labels inside its block; the tail glued after the host of a label's address value is masked to the value's end; and a value this package's own transcript wash split at a carriage return is masked when the display joins it back. |
|
|
432
432
|
| `scripts/run-ask-survives-posture-test.mjs` | The single posture predicate `askSurvivesPosture(card, facts)` for sessions whose standing mode would otherwise answer approval cards on the user's behalf (bypass-style modes). It reads two facts and returns one of three verdicts. The first is the ask origin stamped on the card: the question tool (`content_question`), an organization rule (`org_rule`), a hook (`hook`), an explicit ask rule (`ask_rule`), an organization policy or rule store that could not be read (`org_unavailable`, `rule_store_unavailable`) and the classifier's hand-off after its denial limit (`denial_limit_fallback`) must still be asked (the engine requires a real person to answer all three) and every other origin this build knows is left to the posture only once the host has also reported that its own ask rules did not match. The second is the host's own reading of its settings ask rules for this call: a positive match must be asked, and a command the host could not fully parse counts as no match. When the host reported no reading, every card outside those seven origins gets `unknown`, because an origin says who asked and not that the user's own ask rules did not match; an origin this build does not know gets `unknown` even after a reported non-match. `unknown` is never an approval: the host falls back to its own settings rules. The guard checks the verdict for every origin word, both with no host reading and with a reported non-match, against an independent table whose word set must equal the package's origin list, so a new upstream word fails the guard until it is classified; it covers the combinations of both facts, malformed inputs (non-boolean readings, empty or non-string origins, prototype keys, a different letter case), the fact that the predicate does not read the stronger bits on the card (those stay with the host's earlier checks), real card requests produced by the live-frame, parked-row and suspended-ask paths, and a closed, frozen verdict shape. Since 0.84.0 the product table itself is also checked against the SDK's runtime list of known origin words, so a word the upstream does not know cannot sit in the table. |
|
|
433
433
|
| `scripts/run-engine-agent-absence-projection-test.mjs` | Absent background agents: when the engine stops reporting a background agent and no final state has arrived, the row is marked absent and this package owns every decision about it, so all clients agree. One predicate says whether a row is absent (the mark, not the status, decides). An absent row keeps its last known status, never counts as running, and is never counted as completed, failed or stopped; its elapsed time stops at the last moment it was seen, and its sentence says it may still be running. The end-of-turn sweep never settles an absent row (or a resident one). A row that comes back, or a real final state for the current cycle, clears the mark; a late final state from an earlier cycle does not. Absent rows are never removed at the short grace window. After the hard limit (30 minutes from the last time they were seen) the host is asked for the background-agent registry reading of each row: only a reading that the agent has ended or is not listed lets the row go, and each removal is returned as a fact the host must act on and announce; a reading of running, unknown, missing or unrecognised keeps the row and schedules nothing, so no standing poll is created. Until a registry reading is available every absent row stays. A row someone is viewing is held and reported separately only once the registry confirms it is gone. The row sentence, the removal sentence and the late-result sentence come from one place, never state an outcome or that the agent finished, and escape control characters in names, in the engine's removal word and in the late-result status. An end-to-end cell drives the real fleet projection and the real absence channel through every decision. From 0.84.0 the package also reads the registry itself: a reader lists the session's background registry only when the server advertises the listing, classifies its failures as unavailable (a failed read keeps the HTTP status and error code when the error carries them, and never invents either), and refuses a partly readable listing as a whole; a classifier turns the listing into the per-row reading, treating a registry status it cannot place as unknown and reading "not listed" only for keys shaped like registry handles and only when the host states the engine is a single process and names the row's session, because a key from another identity space proves nothing by being absent and, on a multi-replica deployment, a listing answered by another replica does not contain this session's agents at all (without that statement a missing row reads unknown and the row stays); and a predicate returns the agents the registry still counts as live that the host has no row for. A listing at the server's 500-row cap cannot show that an agent is gone either, so a missing row there reads unknown; the cap is checked against the engine's own list clamp and, when a server build is supplied, against the server's route. Each listing carries its session and a per-client sequence number, and a listing superseded by a later delivered read, or read for another session than the one the host names, counts for nothing. Only the exact listing object the reader returned can authorize removing a row or adding one: a spread copy, a structured clone, a filtered or a hand-built listing reads unknown and adds nothing, the listing, its row array and every row are frozen, the cap check uses the row count recorded when the page was read, and rewriting a listing's sequence number fools nothing. The fill predicate holds back agents first seen within the absence settle window — timed on one monotonic local clock from when this client's reader first saw the id, so a skew between the server's clock and this one, or reconnecting to a long-running agent, cannot skip the window — agents without a readable registration time and agents the host saw end since the read went out; the registry key of a row, progress or absence event is whichever of its two ids is shaped like a registry handle; and none of the reader, the classifier or the predicate throws. |
|
|
434
434
|
| `scripts/run-plan-review-dismissal-test.mjs` | An automatic reopen does not put back a plan-review card the user closed (the first-presentation path neither checks nor clears that record, so a replayed park frame still presents its card). A plan-review card the user dismissed (Esc, abort, or any answer that is not approve or reject) is recorded per session and run at the moment of dismissal, synchronously, before anything queued behind the card can be released; a reopen marked `trigger: 'automatic'` then refuses with `{ reopened: false, dismissedByUser: true }` instead of minting a new card the user's next keystroke would land on, while the user's own next action (`trigger: 'user'`) reopens it and clears the record. The record is keyed by gate instance when the host supplies an instance reader: a new plan gate on the same run is still surfaced, and anything that cannot prove the gate is new (no reader, a failed, empty, thrown or timed-out read) refuses on the conservative side. The asynchronous form re-checks after its reads and before minting — a decision handed over meanwhile (seen by the package, or reported by the host's optional hand-over predicate) answers as "your answer is on its way"; a record that changed meanwhile makes the stale evaluation mint nothing and answer from the current record: another close refuses as the user's close and keeps the newer record, a card already back on screen (the user's own action or a concurrent automatic reopen, waited for within the receipt window and re-read once the wait is over) answers `reopened: true`, and a record that is gone (session change, ledger overflow) answers a plain refusal without `dismissedByUser`; a hand-over predicate that throws refuses with a plain `{ reopened: false }`. Each run has at most one reopen on its way: a reopen that arrives while an earlier one's card is published but not yet settled joins it instead of minting a second card and retiring the first card's answer path. Instance readers are snapshotted when they resolve, so a host that hands over its own live set still gets a new gate recognised; and a decisive-looking host answer note does not clear the record for the very card the package's own responder already judged non-decisive (the label was not on that card). A successful reopen replaces only the record taken before the card was minted (with an on-screen marker, not a deletion), a decisive answer clears it, the per-session ledger is bounded, and a session change clears its bucket. Without `trigger` the reopen answers exactly as before, apart from joining a reopen already on its way. A reopen for a session with no view subscribed to question frames (`onQuestionFrameFor(sessionKey, handler)`; `onQuestionFrame` for the default session) refuses with `{ reopened: false, _sema_noPresentationSurface: true }` — in the synchronous and the receipt form alike, per session key, and also when the view goes away while the asynchronous check is reading — and that answer comes before the user-closed check, at the entry and after the asynchronous read alike, whether or not the read proved a newer gate (after the read, a decision handed over while the read was in progress, and a dismissal record that was replaced or cleared meanwhile, still take precedence over the missing view); a host log or probe port that throws changes none of the reopen answers (the suite runs the same arrangement with a quiet and with a throwing log port and requires the same answer and the same card frames); every other refusal (receipt window lapsed, decision in flight, user-closed card, non-default slot without a delivery path, record gone) keeps its own keys without the flag. The first-presentation call does not check for a view: without one it still answers `true` and the card frame goes nowhere, which the suite records as today's behaviour, not as a guarantee. |
|
|
@@ -440,20 +440,26 @@ guard still cross-checks the table by name).
|
|
|
440
440
|
| `scripts/run-session-policy-refused-removal-test.mjs` | The way out for a session that already carries a tool name the engine refuses. The engine refuses to start a run when any rule applied to it names a retired tool, or a name containing "__" without a protocol prefix, so a session whose own rule record holds such a name fails at startup on every run. The guard pins three pieces. First, the refusal test itself: it is compared name by name, in both directions, with an independent reading of the installed engine's own tables (retired names and protocol prefixes) over a corpus that includes padded, wildcard and parenthesised forms, and it is shown to be wider than the verdict used before a write — a padded or wildcard form is withheld there for another reason yet still refuses the run, so a removal driven by that verdict would leave the session broken. Every corpus name is then written alone into the deny list and into the allow list of a real engine run: the test says "refused" exactly when the run fails at startup with the refusal code and the model is never called, while the same shapes in the command lists are not audited at all. Second, the removal verb: it takes only the refused entries out of those two lists and leaves every other entry, list and ordering byte-for-byte, keeping an emptied allow list as a present empty list; after it runs, the same session's next real engine run gets past startup, for every refused name in the corpus and in both lists. Against the installed engine as an ordinary caller, taking a deny entry out is refused as a loosening — reported as such, written exactly once, the record unchanged and the next run still refused, never a false success — while an allow-list-only removal, which narrows, succeeds; against a model of the newer contract the engine has announced, the removal succeeds, and the same model still refuses a removal of anything else. A concurrent writer causes exactly one re-read, judged again on the other writer's record, and a second collision is reported rather than retried; read failures, coded write refusals and unconfirmed writes (no verdict, a receipt that cannot be read, or a receipt that does not account for the removal, including one that still holds a removed entry) each land in their own arm, and running the verb again after an unconfirmed write reports nothing left to remove. Third, the wording: one sentence per outcome; the loosening refusal names operators and says the record is unchanged, the unconfirmed one says the change may already be in effect, and none carries a rule string or engine text. The startup-failure sentence answers to that one code only, lists every place the entry can be set without naming who may remove it, and a real run with the name in the caller's own policy rather than in the session record yields the same code — which is why the sentence may not claim the entry is in the session record. The verb also requires the tool roster reported by a run on the deployment: a name listed there, as a tool name or an alias and compared exactly as the engine compares it, is left in place even when it is a retired name, because the engine accepts that rule — with a real engine run that mounts a deployment tool under a retired name (or with that name as an alias), the removal keeps the deny rule and the next run still starts. Without a roster, or with one that cannot be read, nothing is read or written and the outcome says why; a readable empty roster is still a roster. The removal sentence speaks separately about names taken out of the deny list and out of the allow list. |
|
|
441
441
|
| `scripts/run-fixture-flat-done-ratchet-test.mjs` | Test fixtures that feed the engine's terminal `done` frame are kept in the shape the live engine actually sends. Since the terminal record became a tagged cause (`result.terminal`), the older flat shape (`result.status` plus loose keys) reaches a client only in two ways: a stored row replayed verbatim from before the upgrade, and the server's own refusal envelope — so a fixture written in the flat shape tests the replay path while claiming to test the live one. The suite finds every `{ type: 'done', result }` literal under `scripts/`, follows `result` to the object literal it really comes from (in place, through a variable, through a helper's parameter at each of its call sites, through a spread, or through the rows of an iterated array) and counts the flat ones that carry no one-line note saying they model a replayed row or a refusal envelope. That count may only go down: it is a ceiling kept in the ratchet registry, and a count under the ceiling prints a step-down line instead of passing in silence. A planted corpus with a known number of flat, noted and cause-shaped frames in each of those forms must be counted exactly, and a note that sits inside a string or gives no reason does not count |
|
|
442
442
|
| `scripts/run-authority-envelope-mirror-test.mjs` | The authority-envelope tags that a peer's message body is defused against before it is written into a transcript line. Tags the engine uses to speak with its own authority (reminders, completion notices and the like) must never survive inside a body a peer wrote, or a forged completion notice could be read back on resume as a real one and poison the dedup ledger. The package's list is generated at build time from the engine package's public entry point, and the guard checks that the source binds that generated list rather than a hand-written copy and that it matches the installed engine package in both directions: a tag the engine treats as authority and the package lacks is red, and so is a tag the package defuses that the engine does not, since that rewrites ordinary text in a peer's message. It also confirms the published list is the engine module's own value and is still derived from the engine's envelope registry, which it reads from the engine's module by path because the registry itself is not published. At the rendered output, every engine authority tag placed in a body (opening, closing, with attributes, upper case) comes out defused on all three peer lanes, and the engine's non-authority envelope tags come out untouched |
|
|
443
|
-
| `scripts/run-upstream-tables-test.mjs` | Engine facts the package may not import at run time (the notice code register and its audience table, the reasons an injected MCP server is dropped, the protocol namespace prefixes and the retired tool names with their current names) are generated at build time from the installed engine package's own public entry point into plain literal modules, which are committed. The generated files must be byte-identical to what the installed engine produces today, no other file may sit beside them, each names the generator and the engine version it came from, none imports anything, and the package's public names are the generated objects themselves. The generator refuses rather than guesses: a table that is missing or empty, a notice code without an audience row (or the reverse), a third audience value, or a retirement note in a shape it does not know all stop generation instead of producing an empty or partial table, and it reads only the engine package installed in this package's own dependency folder. Two word lists the SDK already exports as values (who asked a question, and which word a mandated approval stands on) are now the SDK's own arrays rather than copies; because those arrays are not frozen, the package's own decisions are shown not to change when they are modified in place. Inside the package, the thinking-level and permission-mode vocabularies, the tier words, the rejection sentence and the cancelled code each exist exactly once in the source. A second batch covers engine fact tables that newer engine packages publish from their entry point (the notice cause and source lists, the run-ending reasons, the reserved control-verb names, the workflow size caps, the authority envelope tags, the denial layers, the reasons a requested permission mode did not take effect, and the structured tool-card types): each is generated the same way and bound where it is used, no hand-written copy of any generated table may remain in the source, and an engine package too old to publish them makes the generator stop and name every missing export rather than fall back to reading private files. A third batch covers the closed set of codes the engine refuses a declared caller tool with: it is generated the same way and folded into the package's configuration-refusal recognition set, which must contain every one of those codes; the set only recognises them (no code path branches on them), an unexpected code shape stops generation, and an engine package that does not yet publish the set makes the generator stop and name it |
|
|
443
|
+
| `scripts/run-upstream-tables-test.mjs` | Engine facts the package may not import at run time (the notice code register and its audience table, the reasons an injected MCP server is dropped, the protocol namespace prefixes and the retired tool names with their current names) are generated at build time from the installed engine package's own public entry point into plain literal modules, which are committed. The generated files must be byte-identical to what the installed engine produces today, no other file may sit beside them, each names the generator and the engine version it came from, none imports anything, and the package's public names are the generated objects themselves. The generator refuses rather than guesses: a table that is missing or empty, a notice code without an audience row (or the reverse), a third audience value, or a retirement note in a shape it does not know all stop generation instead of producing an empty or partial table, and it reads only the engine package installed in this package's own dependency folder. Two word lists the SDK already exports as values (who asked a question, and which word a mandated approval stands on) are now the SDK's own arrays rather than copies; because those arrays are not frozen, the package's own decisions are shown not to change when they are modified in place. Inside the package, the thinking-level and permission-mode vocabularies, the tier words, the rejection sentence and the cancelled code each exist exactly once in the source. A second batch covers engine fact tables that newer engine packages publish from their entry point (the notice cause and source lists, the run-ending reasons, the reserved control-verb names, the workflow size caps, the authority envelope tags, the denial layers, the reasons a requested permission mode did not take effect, and the structured tool-card types): each is generated the same way and bound where it is used, no hand-written copy of any generated table may remain in the source, and an engine package too old to publish them makes the generator stop and name every missing export rather than fall back to reading private files. A third batch covers the closed set of codes the engine refuses a declared caller tool with: it is generated the same way and folded into the package's configuration-refusal recognition set, which must contain every one of those codes; the set only recognises them (no code path branches on them), an unexpected code shape stops generation, and an engine package that does not yet publish the set makes the generator stop and name it. A fourth batch (from 0.88.0) covers the closed set of reasons a permission mode's release of a question may carry: it is generated the same way, the package's public word list is that generated object, no hand-written copy of its words remains in the source, and an engine package that does not yet publish it makes the generator stop and name it |
|
|
444
444
|
| `scripts/run-retirement-ledger-test.mjs` | Transitional code whose retirement condition used to live only in a comment now has a machine-readable trigger. `scripts/retirement-ledger.json` lists each piece (a compatibility read, a mirrored table, a normalizer inside the differential check) with an anchor into its own code, the retirement condition quoted from the source comment, and one of three trigger kinds: this package's version reaching a retire-by version; the installed upstream package reaching a version, exporting a name from its entry, or declaring a member on a given type (read with the TypeScript checker); or the reference client tree used by the differential check having moved to a given version of this package. The check fails when a trigger has fired and the transitional code is still there, when a row's anchor can no longer be found (the code is gone but the row stayed), when the quoted condition no longer matches the source comment, when a row has no machine trigger at all, when a retire-by version sits more than three minor lines ahead, and when the anchor of an already retired piece reappears. A reading that cannot be taken is reported as skipped when the absence is legitimate (no reference tree on disk, an upstream declared only as a peer and not installed) and is a harness fault otherwise. Before giving any verdict the judge proves itself on in-memory fixtures and on one known-present and one known-absent name in the installed upstream types. |
|
|
445
|
-
| `scripts/run-plan-review-decision-ledger-test.mjs` | The plan-review reopen path refuses on one decision ledger shared by every delivery path, not only the package's own POST window: a decision the host reports as handed over (before the card's answer reaches the package, or on the host's own delivery path) counts as in flight until it settles; a pending-row read that started before a decision settled is a stale snapshot of the gate just answered, so a reopen carrying that read mark refuses; and a decision proven not to have taken effect (nothing sent, a refusal that by contract changed nothing, or the engine proven to still hold the very same gate instance — a status alone does not prove it, since a run that moved on to a new plan gate reports the same status) neither counts as in flight nor makes such a snapshot stale, so a failed decision does not consume a legitimate reopen. The ledger's capacity bound never evicts the package's own in-flight decisions, so a duplicate approval stays latched however many decisions are in flight. The package's own in-flight window now runs from the POST to the moment the outcome is emitted (settlement happens before the outcome is queued, so a read the host starts from the queue port is not stale), while the duplicate-delivery latch keeps its POST-only window. A handed-over decision the package's delivery path then sends is the same ledger entry (the host's early settle does not release it); two decisions on one run each settle on their own; positive evidence that the run left the gate settles what is open without releasing the latch; bad inputs neither throw nor record. The ledger is checked against the terminal's current implementation vector by vector when that tree is on disk |
|
|
445
|
+
| `scripts/run-plan-review-decision-ledger-test.mjs` | The plan-review reopen path refuses on one decision ledger shared by every delivery path, not only the package's own POST window: a decision the host reports as handed over (before the card's answer reaches the package, or on the host's own delivery path) counts as in flight until it settles; a pending-row read that started before a decision settled is a stale snapshot of the gate just answered, so a reopen carrying that read mark refuses; and a decision proven not to have taken effect (nothing sent, a refusal that by contract changed nothing, or the engine proven to still hold the very same gate instance — a status alone does not prove it, since a run that moved on to a new plan gate reports the same status) neither counts as in flight nor makes such a snapshot stale, so a failed decision does not consume a legitimate reopen. The ledger's capacity bound never evicts the package's own in-flight decisions, so a duplicate approval stays latched however many decisions are in flight. The package's own in-flight window now runs from the POST to the moment the outcome is emitted (settlement happens before the outcome is queued, so a read the host starts from the queue port is not stale), while the duplicate-delivery latch keeps its POST-only window. A handed-over decision the package's delivery path then sends is the same ledger entry (the host's early settle does not release it); two decisions on one run each settle on their own; positive evidence that the run left the gate settles what is open without releasing the latch; bad inputs neither throw nor record. The ledger is checked against the terminal's current implementation vector by vector when that tree is on disk. Positive evidence that the run left the gate can arrive after the host has handed a decision over and before the package's delivery path picks it up (while the host is still waiting on its gate snapshot, or in the microtasks between that wait ending and the delivery call); the delivery path then still picks up that same entry instead of minting a new one, so the latest decision mark the host keyed its snapshot and its not-applied notice on does not move, the duplicate-delivery latch and the capacity rule hold for it, and the three failure outcomes read byte for byte as they would without the hand-over. Only an entry settled by such evidence is picked up again: one the host settled itself, one it revoked, and one that already produced an outcome are not, and a new decision gets a new mark |
|
|
446
446
|
| `scripts/run-subagent-injected-wire-test.mjs` | The background-agent verbs — the child report read, the live tail, the task-handle output read and stop, steer, the resume assembly, manual compaction with its capability warm-up, and the delegated-prompt read — accept a host-supplied engine client, and the plan-review orchestration now resolves its client through the same single construction point, so a browser host behind a same-origin relay runs every one of these paths without an installed engine target. With a client injected, each entry sends only through it: proven against a real relay-form SDK client and a fake engine while the installed target (default and keyed slots alike) points at a different fake engine that must receive nothing, with the row's own session parameter carried, no `authorization` header on any request (checked beside a token-bearing client on the same spy), and a planted token never reaching a log line, a return value or a failure detail on the injected, installed-target and unreachable paths. The capability evidence behind the task-handle verbs and the row stop gate is read only under the key the host names, never from the installed target; the report, tail and compaction warm-up cache their own probe under that key and re-probe when it is absent. A malformed injected client or connection lands on each entry's existing cannot-send outcome (`null`, `no-wire`, `unavailable`, `offline`, `unreadable`) with zero requests, never a throw or an unhandled rejection. Every request through an injected client settles at the package's own deadline even while the client sleeps through a retry back-off and aborts the request underneath; the caller's own signal still applies, and the old signatures answer as before apart from the no-answer outcomes described next. A steer, resume or stop that may have reached the engine but got no answer — the package's own deadline on an injected client, the client's per-request limit on either path, or a transport failure that cannot be shown to have happened before anything was sent (the connection dropped after the request went out) — is reported with a machine-readable `unconfirmed: true` beside the unchanged `reason`, so a host can say it may have taken effect instead of calling it a failure; a cancellation by the caller after the request went out is reported the same way (only the wait was cancelled), while a real 4xx or 5xx answer, a connection that could not be made at all, a verb that threw before returning a promise, and a cancellation already in place before the call (nothing is sent) carry no such key. An installed target is copied by value when a call resolves it, so a delayed or retried compaction still goes to the engine and credentials it was resolved against even if the host's target object changes meanwhile, and a verb that throws synchronously is classified exactly like one that rejects. The resume failure classifier that hosts reuse for their own direct calls makes the same judgement from the same single source, so every client renders it the same way. |
|
|
447
447
|
| `scripts/run-session-background-stop-test.mjs` | The session-level stop for background work (`stopEngineSessionBackground`): one call asks the engine to stop the background tasks of a session, with `includeRetained` passed through as the required true / false choice. It resolves its client through the same single construction point as the other engine verbs (a host-supplied client, or the installed engine target for the chosen session slot) and sends nothing unless the engine has declared the stop face (`background.exitFaces` strictly `true`) under the capability key that applies: evidence that is missing or unread, an older engine without the face, a malformed flag, a flag found only on the prototype chain, or evidence recorded for a different deployment all end in `unavailable` with zero requests. The request body on the wire is exactly `{"includeRetained": <boolean>}` whatever object the caller passes. A 200 answer returns the receipts verbatim (identifier and outcome word; a row whose word it does not recognise is kept with that word as it is, so the host can still name it and the companion below sorts it as may still be running), counts rows whose identifier or word it cannot read instead of inventing them, and leaves the receipt list absent rather than empty when the body cannot be read. Refusals are sorted by machine code before HTTP status (missing service credential, no ownership face, unknown session versus missing route on the same 404, request shape, not permitted, rate limited with the server's wait hint, anything else), a server failure is an `error`, and each call sends exactly one request even through a client configured to retry. A request that may have reached the engine but got no answer — a time limit, a connection that dropped after the request went out, or a cancellation by the caller after the request went out — carries `unconfirmed: true`, while a real answer, a connection that never opened, or a cancellation already in place before the call (nothing is sent) does not. Malformed arguments, options, connections, clients, capability records, answer bodies and error objects never throw and never leave an unhandled rejection, and a planted token never reaches a log line or a result. A pure companion (`engineSessionBackgroundReceiptReadingOf`, with the frozen state list `ENGINE_SESSION_BACKGROUND_RECEIPT_STATES` derived from its wording table) sorts each receipt word into stopped, still running (kept on purpose) or may still be running, with one fixed sentence per state so every client says the same thing: each of the four words lands in its state, while any word it does not recognise — a future word, a different letter case, stray whitespace, a near spelling, a prototype-chain name — lands in may still be running with the word carried back verbatim, and non-string input (boxed strings and objects with their own `toString` included) lands there too without being coerced and without throwing; it never reports stopped for anything it cannot confirm. The three sentences differ, the two that are not stopped never use the word stop, and the state list and the states actually produced match in both directions. |
|
|
448
448
|
| `scripts/run-sealed-key-capability-test.mjs` | The engine's sealed-key discovery segment (`capabilities.sealedKey`), read the same four-state way as the other capability readers: exactly three non-empty strings (`alg`, `publicKeyId`, `publicKey`) make a usable reading and are passed through as they are; an absent key means *do not seal* (an older engine, or an engine that could not set up its key custody — the same action either way) and is never treated as a key; a present but malformed segment is dropped rather than turned into a half reading; and anything the segment carries beyond the three public fields — a private key, a creation time, anything else — never reaches the reading, so this reader cannot become a credential channel. The reading is checked against the projection a `7.104` engine package actually mints, both with a key store and without one. |
|
|
449
449
|
| `scripts/run-session-integrity-refusal-test.mjs` | Three refusals a session can meet on newer servers, read the same way on every client. When a session's saved history cannot be read, the server answers with one code whether the refusal arrives as an HTTP error, as the failure recorded at the end of a streamed run, or on the background run's record; one predicate decides it for all three carriers, it keys on the code alone — never on the status or on the engine's wording — and one sentence explains it: the history is damaged, retrying will not help, whoever operates the engine has to repair it, and meanwhile a new session works. The package's own end-of-run error line and the error of the result frame both carry that sentence, identically, and nothing changes for any other code. When a delete is refused because background agents the session started are still running somewhere this server cannot stop them, the refusal is read with the ids it names — each id checked on its own, a bad entry dropped rather than invented, and a list with no usable entry reported as absent rather than as an empty list, which would read as "no agents left"; a refusal because a run is still in progress is read with that run's id, and any other conflict falls back to a sentence that does not guess. When a decision reaches an approval that moved in the meantime (stopped, decided elsewhere, parked again, claimed by another decision, or — newly — belonging to a session that was deleted), the package reads the code rather than the error class, says the decision was not applied and that the pending approvals should be fetched again, and the approval legs treat it as "the approval is no longer where it was" and keep reading the run, instead of letting the wording of the server's sentence decide. When the server package is supplied, the guard drives the server's own code for each of these answers and pushes the result through the SDK's error mapping before the package reads it. |
|
|
450
|
-
| `scripts/run-mode-release-projection-test.mjs` | The permission-mode fields on a tool call's gate record. Newer engines record, on an allowed call, that a permission mode answered a question nobody was asked (which mode, which question class, the whole class set, and the origin and mandate words the card would have shown), and on a refused call that the session's permission mode itself refused it (which mode and which question class). The package's gate reader carries both into its view — the allowed-call account whole or not at all, the refusal fields one by one, unfamiliar mode or class words passed through as written — and exposes a named reader for the allowed-call account. Absence is not a negative: an older engine, or a server that does not forward these fields, produces a record without them, so the reader answers `undefined` rather than claiming no mode was involved; when a server package is supplied, the guard runs the same records through that server's own projection to show exactly that. The fixtures are first passed through the engine's own record screen, the view's fields are checked against the engine's declared shapes in both directions, the fields survive the projection onto the transcript arm unchanged, and a mode refusal is filed in the transcript as a permission-rule denial. |
|
|
450
|
+
| `scripts/run-mode-release-projection-test.mjs` | The permission-mode fields on a tool call's gate record. Newer engines record, on an allowed call, that a permission mode answered a question nobody was asked (which mode, which question class, the whole class set, and the origin and mandate words the card would have shown), and on a refused call that the session's permission mode itself refused it (which mode and which question class). The package's gate reader carries both into its view — the allowed-call account whole or not at all, the refusal fields one by one, unfamiliar mode or class words passed through as written — and exposes a named reader for the allowed-call account. Absence is not a negative: an older engine, or a server that does not forward these fields, produces a record without them, so the reader answers `undefined` rather than claiming no mode was involved; when a server package is supplied, the guard runs the same records through that server's own projection to show exactly that. The fixtures are first passed through the engine's own record screen, the view's fields are checked against the engine's declared shapes in both directions, the fields survive the projection onto the transcript arm unchanged, and a mode refusal is filed in the transcript as a permission-rule denial. From 0.88.0 the release reason word has a closed-set mirror: the public word list is the build-time generated table taken from the engine package's entry point (the same object, frozen, equal to the engine's list word for word), the membership predicate answers exactly as the engine's own predicate on members, unknown words, case and whitespace variants, empty strings, non-strings and prototype keys, and the reader still passes an unfamiliar reason through unchanged; a release minted by the installed engine's own mode arbitration reads back field for field with its reason recognised. |
|
|
451
451
|
| `scripts/run-read-credential-refusal-projection-test.mjs` | The read-credential refusal on the engine's pre-gate read routes (workflow reads and the live fleet stream), projected as a named state instead of a silent retry loop or a generic failure. On a deployment that requires a service credential, these reads now answer 401 when the client presents none it accepts, and before this change the workflow monitor kept reporting "still loading" and retried every few seconds forever, while the background view folded the refusal into "unavailable". The guard drives the real monitor through a scripted transport: a 401 with the unauthorized code, or with no code, on either the detail read or the list read, yields the new `unauthorized` state carrying the last error and — when the run had been live — the last snapshot; the loop then waits on a long back-off (the timer values it schedules are recorded) rather than the short one, opens no activity stream, and returns to live once the credential is accepted. A 403, a 407 and the two other 401 codes (a missing principal, an unverifiable signed principal) do not enter this state, and for every other status and for connection errors the outcome is compared, case by case, with the previous judgement frozen inside the guard — a deployment without a service credential never reaches the new state. The background view's fleet source gets the matching `unauthorized` health word for both 401 shapes, keeps its rows out of the view, leaves the assistant source untouched and recovers to `ok`; every other failure shape is again compared with the frozen previous judgement, and the assistant source's own 401 staying `unavailable` is pinned as a documented limit. One exported sentence serves both places: it says the engine requires a credential the client did not present, and never that there is nothing to show or that the engine is down. |
|
|
452
452
|
| `scripts/run-rules-legacy-tool-name-projection-test.mjs` | The session rule write refusal for tool names the engine would refuse at run startup, projected as a named outcome. A session rule write whose tool-name lists carry a retired tool name, or a name containing "__" without a protocol prefix, is refused before anything is stored, with a dedicated code and the offending list and name at the top of the error body. The guard pins the new code constant and its family predicate (an open prefix test shaped like the configuration family's, false for anything that is not a string of that family), then feeds the failure classifier the error exactly as two SDK generations deliver it — a plain API error on the older one, a bad-request error on the newer one — and requires both to land in the named arm by code, ahead of the status. The list and name are read only from the body slot the SDK provides, narrowed to the two tool-name lists and a non-empty name, and never from the error object's own `name`, which is its class name. An unknown member of the same family lands in the generic answered-refusal arm rather than being folded into the named one or described as a malformed body. Against a model of the refusing write gate that uses the installed engine's own tables and order, the tightening flow reports the named outcome, writes exactly once and leaves the stored record untouched — both when a host sends such a name directly and when the stored record already holds one from before the gate existed (the outcome then says so), and when this package's mirror of the retired-name table is older than the engine's. The removal verb, which deliberately keeps a name the supplied roster mounts, gets the same named outcome when the gate refuses to store that name. One sentence covers the tightening outcome and the code lookup; it shares its description of the two name classes with the startup-refusal sentence and its fix with the before-write sentence, word for word, carries no rule text, and differs from every other outcome sentence. |
|
|
453
453
|
| `scripts/run-store-center-wiring-capability-test.mjs` | The two cloud-mode capability positions (engine ≥7.106.0), read the same four-state way as the other capability readers. The store posture (`capabilities.store`: declared backend, assembled backend, whether it is durable, an optional degradation reason, and the SQL engine seat) is always present on a current engine, so an absent key means an older engine and is never read as “not durable”; words are carried as reported rather than folded into known ones, the SQL seat is judged by the very same rule as the SQL-posture reader, and one malformed seat discards the whole reading. The center wiring (`capabilities.centerWiring`) is present only when the engine is connected to a configuration center: with no quota lease configured its three lease-dependent keys are legitimately absent, and an absent key is split by the same response — a current engine (it carries the store posture) is reported as not connected, an older one as not reported. Two predicates live in the package (does this replica persist; is the caller's budget lease-enforced — known words matched exactly, everything else unknown), and each reader has one table of doctor sentences in which no two states read alike, the older-engine sentence never claims “not durable” or “not connected”, the two failure postures cannot be said the wrong way round, and unrecognised words are shown escaped and bounded. Both are checked against the engine's own projections: every form they mint reads back exactly, every word in their vocabularies lands on a known arm, and the member lists the engine declares for these two positions equal the reading's own keys. |
|
|
454
454
|
| `scripts/run-inherit-env-wire-test.mjs` | The request field that lets the model-driven shell inherit this machine's whole environment (`"all"`) or keep the default scrubbed environment (`"scrub"`). Both words go out exactly as given on the two user lanes, an absent value is never filled in, and any other value (including a list of variable names) is refused before anything is sent. The value reaches a request only through a helper that checks the engine's version: an engine too old to know the field, or one whose version cannot be read, gets nothing, and the caller receives the reason plus one sentence saying what the run's shell gets instead, including that leaving the field out also clears an earlier whole-environment choice on engines that know it. An invalid-field refusal of a request that asked for `"all"` is read by code with one conditionally worded sentence, because the same code has other causes. |
|
|
455
455
|
| `scripts/run-submit-refusal-projection-test.mjs` | Four kinds of refusal that mean the request itself needs fixing, read by code with one sentence each. Permission rule lists that are too long or hold a bad entry are named precisely (which list, how many entries against the limit, or which entry by position) for a new request and, with different wording, when parked work cannot continue with the request stored for it; the same sentence is appended to a decision that fails for that reason. A tool declaration the engine refuses at the start of a run is read by its code rather than its HTTP status, so a status that looks like a temporary outage is not presented as one: the sentence says the same request will be refused the same way. A deployment refusing a single-user-only setting is attributed from the request the caller sent (the whole-environment shell setting, bypassPermissions, both, or unknown) and always says nothing ran and the local settings are not at fault. A body-shape refusal exposes its two lists of unrecognized and unsupported keys exactly as sent (only well-formed lists are read; a missing list is not treated as empty), with a three-way check of whether any listed key is a settings path. |
|
|
456
456
|
| `scripts/run-mandated-card-values-test.mjs` | Approval cards that newer servers mark as mandatory for four kinds of question (a deployment command policy, an operator's never-auto list, an exhausted durable budget, and supervisor routing) are built with the server's own code and fed through both card paths: the card carries the mandatory mark and the matching reason word, the shared check reports it as mandatory, no mandate word is invented, and a session-wide allow that the server declines to remember produces the single not-remembered notice. An ordinary approval-list card stays non-mandatory and remembered, and the older card shapes are shown to read the other way. This check needs an installed server package to run and reports itself as skipped otherwise. |
|
|
457
|
+
| `scripts/run-park-failed-reason-projection-test.mjs` | The fail-soft true cause when a parked approval cannot be decided and the engine's own `done{suspended}` terminal is the one handed back: that terminal is still returned as-is (no second terminal), but it carries `_sema_park_failed_reason` — the same sentence, byte for byte, that the durable leg puts on its synthesized terminal — and the result frame carries it through with credentials washed the same way `errors[]` is (value only; same sentence as the durable leg after washing). The key is left off when this host has no decision surface at all (no approval card port, a port installed under another session key, or no question overlay): that sentence's way out is to decide on the card, which such a host never shows. Pinned with a policy-fold refusal on decide, both stall exits, a host-retracted card, credential samples through the full bridge on both legs, the no-decision-surface forms, and the absent forms (user interrupt, malformed values, ordinary terminals). |
|
|
458
|
+
| `scripts/run-tool-history-mismatch-projection-test.mjs` | When a provider rejects every request of a session because a tool call in the conversation history no longer pairs up with its result, the terminal error row carries a machine-readable key naming the kind of mismatch and the first tool call id the provider named; a host that offers the rewind command gets the same recovery sentence CC shows (unless that host names the engine and the engine reports it cannot rewind the conversation), plus a count from the second time in a row, while the print and utility lanes, hosts without that command, and the result frame's error text stay unchanged. Only the run's own failure is read, never assistant text, and only when the provider sentence opens the provider's message: other statuses, other 400s and quoted copies do not match. The count is per session and per run, ignores replays, resets on any other outcome and never overstates when it cannot tell; the row-resolving helper that finds the prompt to go back before is checked for the target, the first-turn case and every undecidable case. The check also feeds outcomes built and published by an installed server package when one is available and reports that part as skipped otherwise. |
|
|
459
|
+
| `scripts/run-display-untrusted-baseline-diff-test.mjs` | The credential side of the display outlet (`displayUntrusted` and `washCredentials`), diffed byte for byte against the previous published release: five input pools (seeded random strings, exhaustive shapes around labels, quoted values, userinfo, schemes and invisible units, and URL parameters separated by `;` or `,`) run through fifteen outlets. A secret the previous release masked must stay masked wherever it sits in a credential position; the outlet with only the URL net enabled is judged within that net's own reach. Clean lines from the package's own documents pass through unchanged, and every outlet is idempotent (masking the output again changes nothing). The previous release tarball is supplied through `SEMA_DU_BASELINE_TGZ`; without it the diff sections report as skipped and the run counts as partial, not green. The minimal-shape pool also carries scheme-less userinfo whose password holds a quote or angle bracket after a credential label, and the shapes that once washed differently on a second pass. A full-product grammar pool of 64,512 strings joins the idempotence check, and every pool is also checked for the credential stage reaching its pass cap (it must never do so). |
|
|
460
|
+
| `scripts/run-registry-conflict-projection-test.mjs` | How a client tells the two meanings of a registry write conflict apart. The registry answers the same 409 when a write merely lost an optimistic-lock race (re-read, re-base, write once more) and when the space being written was deleted, or deleted and re-created under the same name, while the write was in flight (nothing was written, and a re-read would return the replacement space, so a retry would put an old draft into a space nobody reviewed). The package ships the same three-word classification the newer SDK uses under the same name, plus a predicate for caught errors, because the SDK version this package supports does not have them yet. The guard checks every status and body combination against the newer SDK's own function when that SDK is available, checks the predicate against the newer SDK's own error classes, reads every 409 the registry service mints from its source when that tree is available and replays the resulting bodies, and checks that unreadable bodies never throw and never classify as the retryable kind. The parts that need the newer SDK or the registry source report as skipped when those are absent. |
|
|
461
|
+
| `scripts/run-prompt-delivery-projection-test.mjs` | Whether a user message actually reached the engine, judged once for every client. A ledger records, per message id and per submitted batch, that a submission entered the submit path and whether the engine created a run for it: a stream frame from the new run counts, a later accepted steer into the running run counts, while the session-busy refusal does not, and neither do the frames a server writes before creating the run (the stream-head `meta` frame and replays of approval cards still pending from earlier runs), which arrive ahead of that refusal. Both sets only grow, so the answers do not depend on arrival order, replays or repeats; a message never seen on the submit path is reported as unknown rather than undelivered. Two row-level readers filter side-channel payloads (only messages known to be undelivered are removed, everything else is kept in order). Records serialize one per line in the shape hosts already persist. The check also replays the exact frames an installed server package writes for a busy session through the real SDK stream parser when one is available and reports that part as skipped otherwise. The idempotent-replay form of the busy-session refusal (the raw conflict body without a status, sent when a host reuses an idempotency key while the first submission is still in flight) does not count as delivered either. |
|
|
462
|
+
| `scripts/run-session-memory-erase-projection-test.mjs` | The session owner's memory erase (`eraseSessionMemory`, with the pure judge `sessionMemoryEraseOutcome`): one call sends exactly one request to erase what a session contributed to memory, with a body of exactly the request id plus the unevidenced-erase consent only when the caller gives it as a boolean (never added when absent, and no other key the caller passes reaches the wire); a dot-segment session id is refused locally. The connection takes the same configuration as the engine client factory, and the outbound headers are compared byte for byte with a client that factory builds for every credential form (string, per-request value function, a value function that returns nothing or throws, empty, same-origin relay, loopback without a token, blank or missing principal); a token-bearing connection in a browser host sends nothing. Malformed arguments, options and connections never throw and send nothing. A request that may have reached the engine without a readable answer (time limit, a connection dropped after the request went out or while the body was read, a cancellation after sending, an unreadable response) is reported as no answer, while a connection that never opened, a cancellation before sending or a synchronous fetch failure is reported as not sent. The judge sorts each answer into erased, partly erased, refused (nothing was erased: not sent, or a named engine code issued before any change) or not confirmed (no answer, an unreadable receipt or one for a different request or session, a refusal with two possible meanings, an unrecognised code, or a non-success status without the engine's own error code; a code found only on the prototype chain does not count), by machine code before HTTP status, with one fixed sentence per outcome and a retry posture; a caller who gave the unevidenced consent always gets a person decides on the not confirmed outcomes. Receipts are read with the existing erase receipt reader. The closed lists are frozen, derived from the wording tables and equal to what the judge produces; sentences differ, refusals always say nothing was erased, and none copies a server sentence. The check also boots installed server packages of the current and the previous generation, when available, with their real dispatcher and handler and the engine's real file memory store, and drives the call against them: real receipts and replays, a request id reused on another session, a store without erase evidence with and without consent, unknown and foreign sessions, missing credentials, missing faces, malformed bodies and the coded refusals; the previous generation's untyped internal errors are judged not confirmed. The unevidenced-erase consent counts only when it is the caller's own property (an inherited one is never sent), a session id that cannot be encoded into the path is refused locally without a request and without a rejection, and a signal-shaped object that is not a real abort signal is refused as bad input. |
|
|
457
463
|
|
|
458
464
|
Each suite carries a floor that only moves up — a refactor that stops executing a group of
|
|
459
465
|
assertions is a failure, not a quieter pass. Guards anchor on the **installed artefact's content**
|