@openwop/openwop-conformance 2.45.3 → 2.45.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,30 @@
1
1
  # `@openwop/openwop-conformance` Changelog
2
2
 
3
+ ## [2.45.4] — 2026-10-01 — v2 ports of production-backpressure and budget-enforcement, on a per-major profile
4
+
5
+ - **The gate runs every host-free scenario** (not packed; no scenario changes). New `scripts/list-host-free-scenarios.mjs` lists the scenarios whose import closure never reaches `lib/driver`, `lib/sse` or `lib/env` and never calls `fetch(`: 70 files today. `scripts/openwop-check.sh` runs them with `OPENWOP_BASE_URL` and `OPENWOP_API_KEY` unset, after the nine hand-listed files.
6
+ - **Why.** `host-callback-declaration` and `runner-ledger` failed on every host in the 2.45.3 candidate and nothing in CI said so (#1835).
7
+ - **Proof.** Clean tree: 66 files pass and 4 skip, no failure. With `REQUIRES_HOST_CALLBACK` removed from `v2-tool-authorization-fail-closed`, the step fails on `host-callback-declaration`.
8
+ - **Limits.** The walk reads static relative imports. `--check` fails if the set is empty or loses a known member. `ai-envelope-shape` is on the hand list and outside the derived set, so the hand list stays.
9
+ - **The 2.45.4 cycle opens.** Two new scenarios and one new fixture change the packed content.
10
+ - **Fix: an unclaimed any-of floor no longer blocks a major-1 bundle** (a 2.45.3 defect). The runner wrote the `openwop.floor.anyof.byok-roundtrip+secrets-run-witness` summary row for every floor in the v1 table, claimed or not. A host that does not advertise secrets records both members `inapplicable`, so it got one `blocked` row for `openwop-secrets`, a profile it never claimed, and a v3 bundle with any `blocked` row certifies nothing (RFC 0168 §E.1). The row is now written only when the profile is claimed. A claimed profile is unchanged: unwitnessed members still record `blocked` and deny it.
11
+ - **Who was affected:** a major-1 host cutting a v3 bundle on 2.45.3 without claiming `openwop-secrets`. Major-2 bundles never carried the row.
12
+ - **Proof.** `floor-any-of.test.ts`: the same evidence yields no row unclaimed and a `blocked` row claimed. Found by running the derivation, after reading the code had suggested the row was harmless.
13
+ - **What differs between majors is now data** (`src/lib/major-profile.ts`). One table row per major holds the run path, the version headers, the event names (from the codemap), where retry timing lives, how a family is advertised, and how a run budget rides on `createRun`. A shared witness takes a profile and never names a major, so a port to a later major is one row plus a thin scenario file. A major with no row throws; it never falls back to an older major's rules.
14
+ - **New scenario `v2-production-backpressure`** (major 2, unaided; `conformance.md` §Production profile). Gated on `production.backpressure.inflightCap`. The suite holds that many event streams open, waits for each to open, and sends one more request.
15
+ - `openwop.requirement.production.backpressure-refusal`: the extra request answers `503 service_unavailable` with `Retry-After`, equal to `retryAfterSeconds` where advertised.
16
+ - `openwop.requirement.0171.error-registry.no-retry-details`: the refusal carries no `details.retryAfter*`. v1 required `details.retryAfter`; v2 forbids it.
17
+ - **Dispositions.** No `production`, `backpressure` or `inflightCap`, a cap above 64, or an unadvertised hold or probe fixture: `inapplicable`. A slot that cannot be held, or a probe with no response: `blocked`.
18
+ - **Not ported:** v1's "discovery is exempt from the cap" leg. No v2 document states it.
19
+ - **New scenario `v2-budget-enforcement`** (major 2, unaided; `runs.md` §`budget` section). At major 2 a budget rides on `createRun` (`configurable.budget`), so no seam is needed. The suite runs the new fixture with `{ maxToolCalls: 2, thresholdPercent: 50, onExhaustion: "fail" }` and reads the log through the poll.
20
+ - `openwop.requirement.runs.budget-lifecycle`: `budget.reserved`, `budget.threshold-crossed` and `budget.exhausted`, in order.
21
+ - `openwop.requirement.runs.budget-enforcement`: a `hard` host emits `cap.breached` (`budget-tool-calls`) and fails the run `budget_exhausted`; an `advisory` host does not stop it.
22
+ - `openwop.requirement.runs.budget-content-free`: no `budget.*` or `cap.breached` payload carries pricing or a credential.
23
+ - **Dispositions.** No `budget`, no `toolCalls` dimension, or the fixture unadvertised: `inapplicable`. The fixture is the opt-in, so a host that advertises `budget` without seeding it is unwitnessed, not blocked. A refused valid create: `executed-fail`.
24
+ - **Not ported:** v1's `budget_model_denied` leg. It needs a fixture that resolves a model without a seam.
25
+ - **New fixture `conformance-budget-tool-calls`**: one `core.conformance.mock-agent` with three scripted tool calls.
26
+ - **Proof.** No host serves either family unaided at major 2, so both witnesses are proven against `src/lib/scratch-host.ts`, a profile-driven test double: 29 self-tests, one defect each (`backpressure-witness.test.ts`, `budget-witness.test.ts`, `major-profile.test.ts`). The scenario files were also run end to end against a conforming and a defective scratch host.
27
+
3
28
  ## [2.45.3] — 2026-10-01 — RFC 0229 (Active): a production host can witness secret resolution without an oracle; v2 tenant-isolation witnesses for storage, fs, memory, workspace, queues and secrets
4
29
 
5
30
  - **Two self-checks that failed on every host are fixed before the cut.** Found by running the candidate on openwop-app at major 1 against 2.45.2: four rows went `executed-pass` → `executed-fail`, and both files failed the same way with no host.
package/README.md CHANGED
@@ -135,7 +135,7 @@ Exit code is non-zero on any failed assertion. `--certify` distinguishes: `0`
135
135
 
136
136
  ## What's Covered
137
137
 
138
- The current suite has 587 scenario files under `src/scenarios/`.
138
+ The current suite has 589 scenario files under `src/scenarios/`.
139
139
  - 2026-09-29 (suite 2.45.0, openwop#1763): NEW `v2-cors-preflight.test.ts` — `headers.md` §Cross-origin preflight. For every operation it preflights with `Origin`, the method, and the operation's declared request headers (plus `authorization` / `content-type` where they apply, from `spec/v2/path-manifest.json`), and reads the answer as the Fetch standard's CORS-preflight check does. Self-gated: `inapplicable` when no operation's preflight grants `OPENWOP_CORS_ORIGIN` (default `https://conformance.invalid`).
140
140
  - 2026-09-23 (suite 2.37.0 cycle, RFC 0213): NEW `v2-sse-last-event-id-cursor.test.ts` (a `Last-Event-ID` past the log is an exclusive cursor; a malformed id, when refused, is `400 validation_error`; the cursor never changes the answer for an unknown or foreign-tenant run — public test of `event-cursor-after-authorization`), `v2-idempotency-in-flight.test.ts` (five concurrent same-key creates yield one run; each loser is a marked replay or `409 idempotency_in_flight` with no retry timing in `details`; `partial-witness` when no loser was refused in flight) and `v2-interrupt-resolve-terminal.test.ts` (a run-scoped resolve after cancel or completion is `409 interrupt_already_resolved`, never `interrupt_cancelled`). All three sit off the core-standard floor until measured on the three bundle hosts.
141
141
  - 2026-09-03 (suite `1.157.0 -> 1.158.0`, gap G17): NEW `idempotency-concurrent-claim.test.ts` — drives the new `host-sample-test-seams.md` §25 concurrent duplicate-delivery seam for the RFC 0150 §B / `idempotency.md` §"Concurrent duplicates (Layer 2)" atomic-claim MUST, which is unconditional and had no witness of any kind. Asserts every executor mints the SAME `logicalInvocationId` **before** asserting `delivered === 1` — without the identity check a host passes by minting different ids and never colliding, one effect because nothing raced. Not profile-gated and so not opt-out-able (the obligation is unconditional); an unmounted seam records `blocked`, which is not certifiable. Graduates `layer2-invocation-claim-atomic` reference-impl -> protocol.
@@ -482,7 +482,7 @@ Server-required (added in 1.7.0):
482
482
  | ------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
483
483
  | **Redaction** | [`capabilities.md`](../spec/v1/capabilities.md) §"Secrets" + NFR-7 + §"aiProviders" | Vendor-neutral assertions that the server doesn't leak secret material. Three scenario groups: (a) discovery shape contract — `secrets` + `aiProviders` advertisements are well-formed regardless of `secrets.supported`; when `supported === true`, scopes MUST be non-empty + `resolution === 'host-managed'`; `byok ⊆ supported`. (b) bearer-token redaction — invalid Bearer canary in `Authorization` header is not echoed in the 401 response body. (c) credentialRef echo control — gated on `secrets.supported === true`; canary planted in `configurable.ai.credentialRef` MUST NOT appear in any RunEvent payload (poll-based capture; transport-agnostic). Uses runtime-built canary fixtures (`lib/canaries.ts`) that defeat static secret scanners. 6 scenarios. |
484
484
 
485
- Current source tree: 587 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
485
+ Current source tree: 589 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
486
486
 
487
487
  ## Remaining Gaps
488
488
 
@@ -26,7 +26,7 @@
26
26
  * host or a vitest subprocess.
27
27
  */
28
28
  import { scenarioFileOfItId } from './requirement-ids.js';
29
- import { PROFILE_FLOOR_SCENARIOS, floorMemberFiles } from './profiles.js';
29
+ import { DEPRECATED_PROFILE_ALIASES, PROFILE_FLOOR_SCENARIOS, floorMemberFiles } from './profiles.js';
30
30
  import { targetMajor } from './seams.js';
31
31
  import { PKG_ROOT_PATH } from './paths.js';
32
32
  import { v2ProfileFloorFiles } from './requirement-registry.js';
@@ -399,10 +399,20 @@ document) {
399
399
  // no member failing; `inapplicable`/`skipped` members never satisfy it. v1
400
400
  // hand table only, like the prefix groups (the major-2 floors have none).
401
401
  const groups = new Map();
402
+ //
403
+ // Only for a CLAIMED profile's floor (2.45.4). The row is `blocked` when no
404
+ // member was witnessed, and a v1 floor reads `inapplicable` as certifiable, so
405
+ // the row cannot simply say `inapplicable`. But 2.45.3 wrote it for every
406
+ // floor in the table: a host that does not advertise secrets records both
407
+ // members `inapplicable`, got one `blocked` row for a profile it never
408
+ // claimed, and a bundle with any `blocked` row certifies nothing (RFC 0168
409
+ // §E.1). A group nobody claims is not a requirement on this host.
410
+ const claimed = new Set(claimedProfiles.map((p) => DEPRECATED_PROFILE_ALIASES[p] ?? p));
402
411
  if (!v2FloorsActive())
403
- for (const floor of Object.values(PROFILE_FLOOR_SCENARIOS))
404
- for (const g of floor.requiredAnyOf ?? [])
405
- groups.set(requirementIdForAnyOf(g), g);
412
+ for (const [profile, floor] of Object.entries(PROFILE_FLOOR_SCENARIOS))
413
+ if (claimed.has(profile))
414
+ for (const g of floor.requiredAnyOf ?? [])
415
+ groups.set(requirementIdForAnyOf(g), g);
406
416
  for (const [id, members] of [...groups.entries()].sort((a, b) => a[0].localeCompare(b[0]))) {
407
417
  const scenarioId = `anyof:${members.join('|')}`;
408
418
  const matching = members.map((f) => perFile.get(f)).filter((r) => r !== undefined);
@@ -1,5 +1,5 @@
1
1
  {
2
2
  "package": "@openwop/spec-artifacts",
3
- "version": "2.45.3",
4
- "stampSha256": "2b7ce94326d2cb0fa88b311a7d6d89c0a0b68f03e66934ad49a19722904a75e7"
3
+ "version": "2.45.4",
4
+ "stampSha256": "1b559bef6691e393c6dc1cdd29e661a8bbec495e43529135d56aba30e1d0a704"
5
5
  }
@@ -0,0 +1,81 @@
1
+ {
2
+ "id": "conformance-budget-tool-calls",
3
+ "name": "Conformance: Budget (tool calls)",
4
+ "version": "1.0",
5
+ "description": "Spends a run budget without a model or a seam. One `core.conformance.mock-agent` node makes three scripted tool calls, so a run created with `budget.maxToolCalls: 2` crosses its threshold on the first call and is exhausted before the third. See v2-budget-enforcement.test.ts.",
6
+ "nodes": [
7
+ {
8
+ "id": "spender",
9
+ "typeId": "core.conformance.mock-agent",
10
+ "name": "Budget Spender",
11
+ "position": {
12
+ "x": 0,
13
+ "y": 0
14
+ },
15
+ "config": {
16
+ "mockToolCalls": [
17
+ {
18
+ "toolId": "openwop.echo:echo",
19
+ "arguments": {
20
+ "x": 1
21
+ },
22
+ "result": {
23
+ "x": 1
24
+ },
25
+ "durationMs": 1
26
+ },
27
+ {
28
+ "toolId": "openwop.echo:echo",
29
+ "arguments": {
30
+ "x": 2
31
+ },
32
+ "result": {
33
+ "x": 2
34
+ },
35
+ "durationMs": 1
36
+ },
37
+ {
38
+ "toolId": "openwop.echo:echo",
39
+ "arguments": {
40
+ "x": 3
41
+ },
42
+ "result": {
43
+ "x": 3
44
+ },
45
+ "durationMs": 1
46
+ }
47
+ ],
48
+ "mockDecision": {
49
+ "decision": {
50
+ "next": "done"
51
+ },
52
+ "confidence": 1
53
+ }
54
+ },
55
+ "inputs": {},
56
+ "agent": {
57
+ "agentId": "core.conformance.budget-spender",
58
+ "modelClass": "reasoning"
59
+ }
60
+ }
61
+ ],
62
+ "edges": [],
63
+ "triggers": [
64
+ {
65
+ "id": "manual",
66
+ "type": "manual",
67
+ "enabled": true
68
+ }
69
+ ],
70
+ "variables": [],
71
+ "metadata": {
72
+ "tags": [
73
+ "conformance",
74
+ "budget",
75
+ "rfc-0084"
76
+ ]
77
+ },
78
+ "settings": {
79
+ "timeout": 15000
80
+ }
81
+ }
package/fixtures.md CHANGED
@@ -73,6 +73,7 @@ All fixtures MUST advertise:
73
73
  | Agent Identity | `conformance-agent-identity` | Phase 1 — `RunSnapshot.agent` / `runOrchestrator` AgentRef wire-shape | `completed` | ≤ 10s |
74
74
  | Agent Reasoning | `conformance-agent-reasoning` | Phase 1 / RFC 0023 — `agent.*` event family emission + `callId` pairing on `core.conformance.mock-agent` | `completed` | ≤ 15s |
75
75
  | Agent Reasoning Streaming | `conformance-agent-reasoning-streaming` | RFC 0024 — `core.conformance.mock-agent` with `mockReasoning.streamChunks` drives incremental `agent.reasoning.delta` events (sequence 0..N-1) followed by exactly one closing `agent.reasoned` whose `reasoning` equals the concatenation. Gated on `capabilities.agents.reasoning.streaming: true`. | `completed` | ≤ 15s |
76
+ | Budget (tool calls) | `conformance-budget-tool-calls` | `runs.md` §`budget` section — one `core.conformance.mock-agent` makes three scripted tool calls, so a run budget is spent with no model and no seam. Seeding it opts the host in to `v2-budget-enforcement` | `completed`; `failed` (`budget_exhausted`) under a hard budget below 3 calls | ≤ 15s |
76
77
  | Agent Low-Confidence | `conformance-agent-low-confidence` | Phase 1 / CP-1 / RFC 0023 — `core.conformance.mock-agent` emits `agent.decided` with confidence < threshold; host MUST follow with `node.suspended { reason: 'low-confidence' }` | `waiting-approval` (suspends) | unbounded (suspends) |
77
78
  | Message Reducer | `conformance-message-reducer` | Phase 1 — `message` reducer idempotency on duplicate `messageId` | `completed` | ≤ 10s |
78
79
  | Agent Pack Install | `conformance-agent-pack-install` | Phase 2 — pack `agents[]` surface as AgentManifest at `GET /v1/packs` | `completed` | ≤ 5s |
@@ -371,6 +372,16 @@ The `messages`-mode stream fixture (AI token streaming) is covered by the determ
371
372
 
372
373
  ---
373
374
 
375
+
376
+ ### `conformance-budget-tool-calls`
377
+
378
+ - **Purpose**: spend a run budget unaided, so `budget` can be witnessed without a model or a seam (`v2-budget-enforcement`).
379
+ - **Shape**: one `core.conformance.mock-agent` node with three `mockToolCalls`. Each call counts once against the `toolCalls` dimension.
380
+ - **Inputs**: none.
381
+ - **Expected behavior**: created with `configurable.budget = { maxToolCalls: 2, thresholdPercent: 50, onExhaustion: "fail" }`, the run emits `budget.reserved`, then `budget.threshold-crossed`, then `budget.exhausted`. A host advertising `enforce: "hard"` then emits `cap.breached` (`kind: "budget-tool-calls"`) and fails the run `budget_exhausted`; an `advisory` host lets it complete.
382
+ - **Terminal status**: `completed` with no budget; `failed` under a hard budget below three calls.
383
+ - **Opt-in**: a host that advertises `budget` without seeding this fixture records `inapplicable` for the witness.
384
+
374
385
  ## `conformance-version-fold` (closes F5)
375
386
 
376
387
  - **Consuming scenario**: `conformance/src/scenarios/version-fold.test.ts` (added 2026-06-11; previously this fixture had no consuming scenario).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@openwop/openwop-conformance",
3
- "version": "2.45.3",
3
+ "version": "2.45.4",
4
4
  "description": "Production-ready black-box conformance suite for OpenWOP v1.0 compliant servers.",
5
5
  "repository": {
6
6
  "type": "git",
@@ -56,6 +56,6 @@
56
56
  "@openwop/spec-artifacts": "file:../spec-artifacts"
57
57
  },
58
58
  "peerDependencies": {
59
- "@openwop/spec-artifacts": "2.45.3"
59
+ "@openwop/spec-artifacts": "2.45.4"
60
60
  }
61
61
  }
package/requirements.json CHANGED
@@ -2,11 +2,11 @@
2
2
  "$comment": "GENERATED by conformance/scripts/generate-requirement-registry.mjs — do not edit. One record per it()/test() in src/scenarios. Ids: openwop.it.<file-stem>.<title-slug>[~n] (src/lib/requirement-ids.ts). A record with id null has an interpolated title; its run-time row is keyed by the rendered title and maps here by file+line only. Renamed ids need a row in requirement-aliases.json.",
3
3
  "generatedFrom": "src/scenarios/*.test.ts",
4
4
  "counts": {
5
- "files": 662,
6
- "tests": 2535,
7
- "withStableId": 2535,
5
+ "files": 664,
6
+ "tests": 2540,
7
+ "withStableId": 2540,
8
8
  "interpolatedTitles": 0,
9
- "explicitIds": 2411
9
+ "explicitIds": 2416
10
10
  },
11
11
  "records": [
12
12
  {
@@ -23601,7 +23601,7 @@
23601
23601
  {
23602
23602
  "id": "openwop.it.runner-ledger.a-floor-file-that-vitest-passed-but-that-recorded-nothing-is-unclassified-when-a",
23603
23603
  "file": "runner-ledger.test.ts",
23604
- "line": 146,
23604
+ "line": 143,
23605
23605
  "title": "a floor file that vitest passed but that recorded NOTHING is unclassified when a ledger exists — silence is not a witness",
23606
23606
  "explicitId": "openwop.it.runner-ledger.a-floor-file-that-vitest-passed-but-that-recorded-nothing-is-unclassified-when-a",
23607
23607
  "citations": [
@@ -23614,7 +23614,7 @@
23614
23614
  {
23615
23615
  "id": "openwop.it.runner-ledger.a-vacuous-pass-executed-pass-with-assertioncount-0-on-a-claimed-floor-is-unclass",
23616
23616
  "file": "runner-ledger.test.ts",
23617
- "line": 158,
23617
+ "line": 155,
23618
23618
  "title": "a VACUOUS pass (executed-pass with assertionCount 0) on a claimed floor is unclassified and REJECTS certification",
23619
23619
  "explicitId": "openwop.it.runner-ledger.a-vacuous-pass-executed-pass-with-assertioncount-0-on-a-claimed-floor-is-unclass",
23620
23620
  "citations": [
@@ -23627,7 +23627,7 @@
23627
23627
  {
23628
23628
  "id": "openwop.it.runner-ledger.a-skipped-file-with-no-ledger-entry-is-blocked-unclassified-and-it-never-reads-a",
23629
23629
  "file": "runner-ledger.test.ts",
23630
- "line": 167,
23630
+ "line": 164,
23631
23631
  "title": "a skipped file with no ledger entry is blocked (unclassified) — and it never reads as skipped/inapplicable",
23632
23632
  "explicitId": "openwop.it.runner-ledger.a-skipped-file-with-no-ledger-entry-is-blocked-unclassified-and-it-never-reads-a",
23633
23633
  "citations": [
@@ -23640,7 +23640,7 @@
23640
23640
  {
23641
23641
  "id": "openwop.it.runner-ledger.a-recorded-inapplicable-skipped-is-certifiable-and-not-unclassified-a-recorded-e",
23642
23642
  "file": "runner-ledger.test.ts",
23643
- "line": 179,
23643
+ "line": 176,
23644
23644
  "title": "a recorded inapplicable/skipped is certifiable and NOT unclassified; a recorded executed-fail blocks but is classified",
23645
23645
  "explicitId": "openwop.it.runner-ledger.a-recorded-inapplicable-skipped-is-certifiable-and-not-unclassified-a-recorded-e",
23646
23646
  "citations": [
@@ -23653,7 +23653,7 @@
23653
23653
  {
23654
23654
  "id": "openwop.it.runner-ledger.a-discoveryonly-profile-certifies-with-an-empty-floor-an-unknown-profile-does-no",
23655
23655
  "file": "runner-ledger.test.ts",
23656
- "line": 192,
23656
+ "line": 189,
23657
23657
  "title": "a discoveryOnly profile certifies with an empty floor; an unknown profile does not",
23658
23658
  "explicitId": "openwop.it.runner-ledger.a-discoveryonly-profile-certifies-with-an-empty-floor-an-unknown-profile-does-no",
23659
23659
  "citations": [
@@ -23666,7 +23666,7 @@
23666
23666
  {
23667
23667
  "id": "openwop.it.runner-ledger.a-runtime-derived-profile-openwop-node-packs-is-held-only-when-every-floor-row-i",
23668
23668
  "file": "runner-ledger.test.ts",
23669
- "line": 199,
23669
+ "line": 196,
23670
23670
  "title": "a RUNTIME-DERIVED profile (openwop-node-packs) is HELD only when every floor row is a witnessed pass; otherwise it is not held — not rejected, not blocked",
23671
23671
  "explicitId": "openwop.it.runner-ledger.a-runtime-derived-profile-openwop-node-packs-is-held-only-when-every-floor-row-i",
23672
23672
  "citations": [
@@ -23683,7 +23683,7 @@
23683
23683
  {
23684
23684
  "id": "openwop.it.runner-ledger.the-runner-s-a-resolution-of-a-silent-zero-assertion-file-blocked-marker-is-stil",
23685
23685
  "file": "runner-ledger.test.ts",
23686
- "line": 234,
23686
+ "line": 231,
23687
23687
  "title": "the runner's §A resolution of a silent zero-assertion file (blocked + marker) is STILL unclassified for a claimed floor",
23688
23688
  "explicitId": "openwop.it.runner-ledger.the-runner-s-a-resolution-of-a-silent-zero-assertion-file-blocked-marker-is-stil",
23689
23689
  "citations": [
@@ -23696,7 +23696,7 @@
23696
23696
  {
23697
23697
  "id": "openwop.it.runner-ledger.with-no-ledger-at-all-every-skipped-file-is-blocked-and-the-runner-says-so-the-p",
23698
23698
  "file": "runner-ledger.test.ts",
23699
- "line": 247,
23699
+ "line": 244,
23700
23700
  "title": "with no ledger at all, every skipped file is blocked and the runner says so (the pre-S6 honest reading)",
23701
23701
  "explicitId": "openwop.it.runner-ledger.with-no-ledger-at-all-every-skipped-file-is-blocked-and-the-runner-says-so-the-p",
23702
23702
  "citations": [
@@ -23709,7 +23709,7 @@
23709
23709
  {
23710
23710
  "id": "openwop.it.runner-ledger.notes-are-keyed-to-the-current-test-file-worst-first-when-mixed-and-read-back-wi",
23711
23711
  "file": "runner-ledger.test.ts",
23712
- "line": 257,
23712
+ "line": 254,
23713
23713
  "title": "notes are keyed to the current test file, worst-first when mixed, and read back with joined reasons",
23714
23714
  "explicitId": "openwop.it.runner-ledger.notes-are-keyed-to-the-current-test-file-worst-first-when-mixed-and-read-back-wi",
23715
23715
  "citations": [
@@ -23722,7 +23722,7 @@
23722
23722
  {
23723
23723
  "id": "openwop.it.runner-ledger.files-that-finished-before-this-one-appear-in-the-ledger-file-with-an-assertion",
23724
23724
  "file": "runner-ledger.test.ts",
23725
- "line": 279,
23725
+ "line": 276,
23726
23726
  "title": "files that finished before this one appear in the ledger file with an assertion count",
23727
23727
  "explicitId": "openwop.it.runner-ledger.files-that-finished-before-this-one-appear-in-the-ledger-file-with-an-assertion",
23728
23728
  "citations": [
@@ -28618,6 +28618,48 @@
28618
28618
  }
28619
28619
  ]
28620
28620
  },
28621
+ {
28622
+ "id": "openwop.it.v2-budget-enforcement.a-budgeted-run-emits-budget-reserved-budget-threshold-crossed-and-budget-exhaust",
28623
+ "file": "v2-budget-enforcement.test.ts",
28624
+ "line": 76,
28625
+ "title": "a budgeted run emits budget.reserved, budget.threshold-crossed and budget.exhausted in order",
28626
+ "explicitId": "openwop.requirement.runs.budget-lifecycle",
28627
+ "citations": [
28628
+ {
28629
+ "section": null,
28630
+ "requirement": null,
28631
+ "interpolated": true
28632
+ }
28633
+ ]
28634
+ },
28635
+ {
28636
+ "id": "openwop.it.v2-budget-enforcement.a-hard-host-stops-the-run-budget-exhausted-after-cap-breached-and-an-advisory-ho",
28637
+ "file": "v2-budget-enforcement.test.ts",
28638
+ "line": 82,
28639
+ "title": "a hard host stops the run budget_exhausted after cap.breached, and an advisory host does not stop it",
28640
+ "explicitId": "openwop.requirement.runs.budget-enforcement",
28641
+ "citations": [
28642
+ {
28643
+ "section": null,
28644
+ "requirement": null,
28645
+ "interpolated": true
28646
+ }
28647
+ ]
28648
+ },
28649
+ {
28650
+ "id": "openwop.it.v2-budget-enforcement.no-budget-or-cap-breached-payload-carries-pricing-or-a-credential",
28651
+ "file": "v2-budget-enforcement.test.ts",
28652
+ "line": 89,
28653
+ "title": "no budget.* or cap.breached payload carries pricing or a credential",
28654
+ "explicitId": "openwop.requirement.runs.budget-content-free",
28655
+ "citations": [
28656
+ {
28657
+ "section": null,
28658
+ "requirement": null,
28659
+ "interpolated": true
28660
+ }
28661
+ ]
28662
+ },
28621
28663
  {
28622
28664
  "id": "openwop.it.v2-bundle-v3-signed.a-signed-v3-bundle-validates-against-the-closed-root-schema-an-extra-root-key-is",
28623
28665
  "file": "v2-bundle-v3-signed.test.ts",
@@ -33153,6 +33195,34 @@
33153
33195
  }
33154
33196
  ]
33155
33197
  },
33198
+ {
33199
+ "id": "openwop.it.v2-production-backpressure.with-inflightcap-slots-held-the-next-request-is-refused-503-service-unavailable",
33200
+ "file": "v2-production-backpressure.test.ts",
33201
+ "line": 58,
33202
+ "title": "with inflightCap slots held, the next request is refused 503 service_unavailable with Retry-After",
33203
+ "explicitId": "openwop.requirement.production.backpressure-refusal",
33204
+ "citations": [
33205
+ {
33206
+ "section": null,
33207
+ "requirement": null,
33208
+ "interpolated": true
33209
+ }
33210
+ ]
33211
+ },
33212
+ {
33213
+ "id": "openwop.it.v2-production-backpressure.the-refusal-carries-its-retry-timing-in-the-retry-after-header-only",
33214
+ "file": "v2-production-backpressure.test.ts",
33215
+ "line": 66,
33216
+ "title": "the refusal carries its retry timing in the Retry-After header only",
33217
+ "explicitId": "openwop.requirement.0171.error-registry.no-retry-details",
33218
+ "citations": [
33219
+ {
33220
+ "section": null,
33221
+ "requirement": null,
33222
+ "interpolated": true
33223
+ }
33224
+ ]
33225
+ },
33156
33226
  {
33157
33227
  "id": "openwop.it.v2-profiles-derived-only.the-v2-document-carries-no-root-profiles",
33158
33228
  "file": "v2-profiles-derived-only.test.ts",
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "$comment": "GENERATED by conformance/scripts/generate-scenario-majors.mjs (RFC 0168 §D.3). Do not edit; add a file to BOTH_MAJORS in the generator to target both majors.",
3
3
  "counts": {
4
- "files": 587,
4
+ "files": 589,
5
5
  "v1": 454,
6
- "v2": 148
6
+ "v2": 150
7
7
  },
8
8
  "majors": {
9
9
  "a2a-1-0-agent-card.test.ts": [
@@ -1317,6 +1317,9 @@
1317
1317
  "v2-bound-id-path-projection.test.ts": [
1318
1318
  2
1319
1319
  ],
1320
+ "v2-budget-enforcement.test.ts": [
1321
+ 2
1322
+ ],
1320
1323
  "v2-bundle-v3-signed.test.ts": [
1321
1324
  2
1322
1325
  ],
@@ -1521,6 +1524,9 @@
1521
1524
  "v2-preferred-version-default.test.ts": [
1522
1525
  2
1523
1526
  ],
1527
+ "v2-production-backpressure.test.ts": [
1528
+ 2
1529
+ ],
1524
1530
  "v2-profiles-derived-only.test.ts": [
1525
1531
  2
1526
1532
  ],
@@ -1,17 +1,17 @@
1
1
  {
2
2
  "_comment": "Provenance of @openwop/spec-artifacts (RFC 0168 §D.2). files: SHA-256 per file; the conformance suite compares the installed peer against dist/spec-artifacts.lock.json at start.",
3
3
  "package": "@openwop/spec-artifacts",
4
- "version": "2.45.3",
5
- "corpusTag": "v2.45.3",
4
+ "version": "2.45.4",
5
+ "corpusTag": "v2.45.4",
6
6
  "files": {
7
7
  "api/.redocly.lint-ignore.yaml": "bf5a8350b88a72fa43f59605ed8d903ed24b6cfccda5e45509c9f6ed9ee4e712",
8
8
  "api/asyncapi.yaml": "d5ecb9ee6114582be3b1f662c84bfac9ae96dae7bacb853e461168f70a8e1c7d",
9
9
  "api/grpc/openwop.proto": "c3e72bb17cba514ee98feb6434e6c9b6ea6795bfd086489ec69fd882dd1ad977",
10
10
  "api/openapi.yaml": "ad37a5a5c732ad46100fc03c3f4243fdd4f2976e330f02a8791b6b5b10a8a0c1",
11
11
  "api/redocly.yaml": "b0604c89b2ca6d5076ec25725c539dad44a741a811fe524439ee6daef8baa09f",
12
- "api/seams-v2.yaml": "1f22ab7223019080332312b0fc9f7eb4c97bca2eebd57cab886acd1d6e3081d0",
13
- "api/v2/asyncapi.yaml": "5b74151b1867c3c234203aa1f2dfb1ceeea216e013395e8696da02d4a8630412",
14
- "api/v2/openapi.yaml": "05a07d9cd1e7f80da1e7044aace1073bf4ab584186cdd67979ba2664f90a80d0",
12
+ "api/seams-v2.yaml": "e4a8e19b494d12c0d105ed8d1d27337995564a6443cb8a747521a9d67977b52b",
13
+ "api/v2/asyncapi.yaml": "d53ca2c06a47282a0cc39d9e0b6b488068da0ad2c91c3166598b005dcf72aa21",
14
+ "api/v2/openapi.yaml": "24bce91f02de6527121f12ab23916c4f96bab2cbf8f2203208c1d3a444c632c5",
15
15
  "api/v2/redocly.yaml": "1e66b60e6118ad11a823bb620678be464d99dfe50a40e3e6f93ec9429b88b34c",
16
16
  "schemas/README.md": "0c0b737ffcf8f30e7d2809cec8a498232de710f41443212922ad8337cdde0b51",
17
17
  "schemas/a2a-task-state.schema.json": "75d5049dea8bd873ff0e7546f1c60c8d36c219bec7084264360be54a8be30ae1",
@@ -204,7 +204,7 @@
204
204
  "schemas/workspace-file.schema.json": "464de85c2a068243084ee9c1d969bc7cd5d8f7948574e58450d6493c38a0e1e4",
205
205
  "spec/v1/alias-detectors.json": "0b2f28808ffe0b9d98a187775038c3406bb70919747a788a6e4f866e293ae6d3",
206
206
  "spec/v1/capability-declaration-classes.json": "e7729aed5c4b4e1dd02abab0530f14cc95f5d4070fe51fb139e7f5cccefa00c6",
207
- "spec/v1/core-standard-manifest.json": "746278bbf9cf46840ba3bd5a85683251ddbc379466585f54e582754ce684d637",
207
+ "spec/v1/core-standard-manifest.json": "aa1ecb9c1f1c1c3039f948b137e4672031fad29103bb9ac4166eb4ff89b7c85e",
208
208
  "spec/v1/deprecations.json": "a3293072ea951857395d05f0437264e88965f1e9019ff61c55bacfa9c1aaa17c",
209
209
  "spec/v1/deprecations.schema.json": "3e393c405d2a41b467d8c5e3c468549078df1ce6a6d2588fc488097b95d9b55b",
210
210
  "spec/v1/event-codemap.json": "3da60d884157793a360da532a9fcbbfb5285636db325a74cec94b34622186d97",
@@ -303,9 +303,9 @@
303
303
  "spec/v2/path-manifest.json": "27c2307b84376b2f22bea4e2213fc7e33ef3bf8fd1d18039f7955edb00dc9c9a",
304
304
  "spec/v2/peer-dependency-aliases.json": "d10299280abee08258502925bc327293ee413e0108cd6e6ec75ff6110653308d",
305
305
  "spec/v2/profiles.json": "180beb9ef2d161766fa23b2987e7e829851e3d316273147cf6a4a3f324befd6b",
306
- "spec/v2/release.json": "0befa21e6634c82cd3d6be9e487f79b7bf77fd3143ce3e0aed301230753b21bc",
306
+ "spec/v2/release.json": "d1b495b89fc43030e607d7dc312732500f0433e84656ccc88ad8c80c2ac444b1",
307
307
  "spec/v2/retention-floors.json": "eaf3722d95c79947af1d4269ef85117e126518c588cfcf1a2b21b97269f51624",
308
- "spec/v2/surface-baseline.json": "fd44c977c6646c0bcbd85f9f5c4dce620bba29ee56b10cd077e3f71437e0276a"
308
+ "spec/v2/surface-baseline.json": "9e17340cea86146899703c91246ddd5a5403c68e40e749f0ad1958801f60b619"
309
309
  },
310
- "corpusCommit": "f84011f0ce7ba068e53c1aa2b92d17fa97f5df5f"
310
+ "corpusCommit": "5a4ad1ab0a3584c7ea30f4b32f4599e53ce1f75a"
311
311
  }
@@ -0,0 +1,163 @@
1
+ /**
2
+ * The backpressure witness, shared by every major (`major-profile.ts`).
3
+ *
4
+ * A host that advertises `production.backpressure.inflightCap` names a number
5
+ * the suite can saturate: `cap` long-lived requests hold the slots and the
6
+ * next request must be refused `503 service_unavailable` with `Retry-After`.
7
+ *
8
+ * Two halves, so each can be proven alone:
9
+ * - {@link saturate} drives the host and returns what it saw (or why it
10
+ * could not see anything). It asserts nothing.
11
+ * - {@link judge} is pure: observation in, findings out. The scenario turns
12
+ * findings into `expect` calls; the self-test feeds it a conforming and a
13
+ * defective observation and checks the verdicts differ.
14
+ *
15
+ * Run with `--no-file-parallelism`: saturating the cap leaves no headroom for
16
+ * a neighbouring scenario's requests.
17
+ */
18
+
19
+ import { driver, type OpenWOPResponse } from './driver.js';
20
+ import { loadEnv } from './env.js';
21
+ import { readErrorCode } from './error-envelope.js';
22
+ import { isFixtureAdvertised } from './fixtures.js';
23
+ import type { MajorProfile } from './major-profile.js';
24
+
25
+ /** The long-running fixture that holds a slot, and the cheap one that probes. */
26
+ export const HOLD_FIXTURE = 'conformance-delay';
27
+ export const PROBE_FIXTURE = 'conformance-noop';
28
+ /** Long enough that the first slot is still held when the last is filled; the runs are cancelled afterwards. */
29
+ const HOLD_MS = 15_000;
30
+ /** The most streams the suite will hold open. A larger advertised cap is not saturated. */
31
+ export const MAX_SATURABLE_CAP = 64;
32
+ const STREAM_OPEN_MS = 5_000;
33
+ const RETRY_DETAIL_KEYS = ['retryAfter', 'retryAfterMs', 'retryAfterSeconds'] as const;
34
+
35
+ export interface Refusal {
36
+ readonly cap: number;
37
+ readonly status: number;
38
+ readonly code: string | undefined;
39
+ readonly retryAfterHeader: string | null;
40
+ readonly details: Record<string, unknown> | null;
41
+ readonly advertisedRetryAfterSeconds: number | undefined;
42
+ }
43
+ export type Saturation =
44
+ | { readonly kind: 'skip'; readonly disposition: 'inapplicable' | 'blocked'; readonly reason: string }
45
+ | { readonly kind: 'refused'; readonly refusal: Refusal };
46
+
47
+ export interface Finding {
48
+ /** Which rule the finding is about; the scenario maps it to a requirement id. */
49
+ readonly rule: 'refusal' | 'retry-after-advertised' | 'retry-timing';
50
+ readonly ok: boolean;
51
+ readonly doc: string;
52
+ readonly message: string;
53
+ }
54
+
55
+ const isRecord = (v: unknown): v is Record<string, unknown> => v !== null && typeof v === 'object' && !Array.isArray(v);
56
+ async function http(fn: () => Promise<OpenWOPResponse>): Promise<OpenWOPResponse | null> {
57
+ try { return await fn(); } catch { return null; }
58
+ }
59
+
60
+ /** Hold `cap` slots with long-lived event streams, send one more request, and report its answer. */
61
+ export async function saturate(profile: MajorProfile, discovery: unknown): Promise<Saturation> {
62
+ const production = profile.family(discovery, 'production');
63
+ if (production === null) return { kind: 'skip', disposition: 'inapplicable', reason: 'the host does not advertise production' };
64
+ const bp = production['backpressure'];
65
+ if (!isRecord(bp)) return { kind: 'skip', disposition: 'inapplicable', reason: 'the host does not advertise production.backpressure' };
66
+ const cap = bp['inflightCap'];
67
+ if (typeof cap !== 'number' || !Number.isInteger(cap) || cap < 1) {
68
+ return { kind: 'skip', disposition: 'inapplicable', reason: 'the host does not advertise production.backpressure.inflightCap — there is no cap the suite can saturate' };
69
+ }
70
+ if (cap > MAX_SATURABLE_CAP) return { kind: 'skip', disposition: 'inapplicable', reason: `inflightCap ${cap} is above the ${MAX_SATURABLE_CAP} streams the suite will hold open — not saturated` };
71
+ for (const f of [HOLD_FIXTURE, PROBE_FIXTURE]) {
72
+ if (!isFixtureAdvertised(f)) return { kind: 'skip', disposition: 'inapplicable', reason: `${f} fixture not advertised — the suite cannot hold or probe a slot` };
73
+ }
74
+
75
+ const env = loadEnv();
76
+ const streams: AbortController[] = [];
77
+ const pending: Promise<number | null>[] = [];
78
+ const runIds: string[] = [];
79
+ try {
80
+ for (let i = 0; i < cap; i++) {
81
+ const created = await http(() => driver.post(profile.runsPath, { workflowId: HOLD_FIXTURE, inputs: { delayMs: HOLD_MS } }));
82
+ if (created === null) return { kind: 'skip', disposition: 'blocked', reason: `POST ${profile.runsPath} unreachable while filling slot ${i + 1} of ${cap}` };
83
+ const runId = (created.json as { runId?: unknown } | undefined)?.runId;
84
+ if (created.status !== 201 || typeof runId !== 'string') {
85
+ return { kind: 'skip', disposition: 'blocked', reason: `slot ${i + 1} of ${cap} could not be filled: POST ${profile.runsPath} answered ${created.status} ${readErrorCode(created.json) ?? ''} (already saturated by a parallel file? run with --no-file-parallelism)`.trim() };
86
+ }
87
+ runIds.push(runId);
88
+ const ctl = new AbortController();
89
+ streams.push(ctl);
90
+ // `Accept: text/event-stream` keeps the request in flight; without it a
91
+ // negotiating host answers a one-shot snapshot and the slot drops.
92
+ const stream = fetch(`${env.baseUrl}${profile.runsPath}/${encodeURIComponent(runId)}/events`, {
93
+ headers: { Authorization: `Bearer ${env.apiKey}`, Accept: 'text/event-stream', ...profile.versionHeaders },
94
+ signal: ctl.signal,
95
+ }).then((r) => r.status, () => null);
96
+ pending.push(stream);
97
+ // The slot is held once the stream's headers are back, not after a
98
+ // guessed delay. A probe sent before that would convict a host for a slot
99
+ // the suite had not yet taken.
100
+ const opened = await Promise.race([stream, new Promise<'slow'>((r) => setTimeout(() => r('slow'), STREAM_OPEN_MS))]);
101
+ if (opened !== 200) {
102
+ return { kind: 'skip', disposition: 'blocked', reason: `slot ${i + 1} of ${cap} was not held: the event stream ${opened === 'slow' ? `did not open within ${STREAM_OPEN_MS} ms` : opened === null ? 'failed to connect' : `answered ${opened}`}` };
103
+ }
104
+ }
105
+
106
+ const probe = await http(() => driver.post(profile.runsPath, { workflowId: PROBE_FIXTURE }));
107
+ if (probe === null) return { kind: 'skip', disposition: 'blocked', reason: `the cap+1 probe got no response (POST ${profile.runsPath})` };
108
+ const accepted = (probe.json as { runId?: unknown } | undefined)?.runId;
109
+ if (probe.status === 201 && typeof accepted === 'string') runIds.push(accepted);
110
+ const body = isRecord(probe.json) ? probe.json : {};
111
+ const advertised = bp['retryAfterSeconds'];
112
+ return {
113
+ kind: 'refused',
114
+ refusal: {
115
+ cap,
116
+ status: probe.status,
117
+ code: readErrorCode(probe.json) ?? undefined,
118
+ retryAfterHeader: probe.headers.get('retry-after'),
119
+ details: isRecord(body['details']) ? body['details'] : null,
120
+ advertisedRetryAfterSeconds: typeof advertised === 'number' ? advertised : undefined,
121
+ },
122
+ };
123
+ } finally {
124
+ // The held runs outlive their streams. Cancel them so the next file does
125
+ // not meet a saturated host.
126
+ for (const c of streams) c.abort();
127
+ if (runIds.length > 0) {
128
+ const bulk = await http(() => driver.post(`${profile.runsPath}:bulk-cancel`, { runIds }));
129
+ if (bulk === null || bulk.status >= 400) {
130
+ for (const id of runIds) await http(() => driver.post(`${profile.runsPath}/${encodeURIComponent(id)}/cancel`, {}));
131
+ }
132
+ }
133
+ await Promise.allSettled(pending);
134
+ }
135
+ }
136
+
137
+ /** What the refusal must look like at this major. Pure. */
138
+ export function judge(profile: MajorProfile, r: Refusal): Finding[] {
139
+ const home = `${profile.specRoot}/${profile.major === 1 ? 'production-profile.md §Backpressure' : 'conformance.md §Production profile'}`;
140
+ const out: Finding[] = [];
141
+ out.push({ rule: 'refusal', ok: r.status === 503, doc: home, message: `with inflightCap ${r.cap} slots held, the next request MUST be refused 503 (got ${r.status})` });
142
+ out.push({ rule: 'refusal', ok: r.code === 'service_unavailable', doc: home, message: `the refusal MUST carry the code service_unavailable (got ${r.code ?? 'none'})` });
143
+ const header = (r.retryAfterHeader ?? '').trim();
144
+ out.push({ rule: 'refusal', ok: header.length > 0, doc: home, message: 'the 503 MUST set Retry-After' });
145
+
146
+ if (r.advertisedRetryAfterSeconds !== undefined) {
147
+ out.push({
148
+ rule: 'retry-after-advertised',
149
+ ok: /^\d+$/.test(header) && Number(header) === r.advertisedRetryAfterSeconds,
150
+ doc: 'capabilities.schema.json production.backpressure.retryAfterSeconds',
151
+ message: `Retry-After MUST equal the advertised retryAfterSeconds ${r.advertisedRetryAfterSeconds} (got "${header}")`,
152
+ });
153
+ }
154
+
155
+ if (profile.retryTiming === 'header-only') {
156
+ const spelled = RETRY_DETAIL_KEYS.filter((k) => r.details !== null && k in r.details);
157
+ out.push({ rule: 'retry-timing', ok: spelled.length === 0, doc: `${profile.specRoot}/errors.md §Retry timing`, message: `retry timing lives in the Retry-After header only; details MUST NOT carry ${spelled.join(', ') || 'retryAfter*'}` });
158
+ } else {
159
+ const d = r.details?.['retryAfter'];
160
+ out.push({ rule: 'retry-timing', ok: typeof d === 'number' && /^\d+$/.test(header) && d === Number(header), doc: home, message: 'details.retryAfter MUST be numeric and equal the Retry-After header in seconds' });
161
+ }
162
+ return out;
163
+ }
@@ -0,0 +1,165 @@
1
+ /**
2
+ * The run-budget witness, shared by every major that gives a budget a run wire
3
+ * surface (`major-profile.ts` `runBudget`).
4
+ *
5
+ * Unaided: the suite creates a run of `conformance-budget-tool-calls` (three
6
+ * scripted tool calls, no model, no seam) with a two-call budget, waits for it
7
+ * to end, and reads its event log through the poll.
8
+ *
9
+ * Two halves, as in `backpressure-witness.ts`: {@link drive} observes and
10
+ * asserts nothing; {@link judge} is pure.
11
+ *
12
+ * The fixture is the opt-in. A host that advertises `budget` without seeding it
13
+ * records `inapplicable`: the family is then unwitnessed at this major, which
14
+ * the coverage report shows, and the host is not denied certification for a
15
+ * fixture it was never asked to seed.
16
+ */
17
+
18
+ import { driver, type OpenWOPResponse } from './driver.js';
19
+ import { readErrorCode } from './error-envelope.js';
20
+ import { isFixtureAdvertised } from './fixtures.js';
21
+ import type { MajorProfile } from './major-profile.js';
22
+
23
+ export const BUDGET_FIXTURE = 'conformance-budget-tool-calls';
24
+ /** The fixture makes 3 calls: 50% is crossed on the first, and the budget cannot cover the third. */
25
+ export const BUDGET_POLICY = { maxToolCalls: 2, thresholdPercent: 50, onExhaustion: 'fail' } as const;
26
+ const DIMENSION = 'toolCalls';
27
+ const CAP_KIND = 'budget-tool-calls';
28
+ /** Keys a `budget.*` or `cap.breached` payload must never carry (`budget-no-pricing-leak`). */
29
+ export const PRICING_KEYS = ['pricing', 'priceTable', 'prices', 'rate', 'rates', 'unitPrice', 'costModel', 'tokenPrice', 'secret', 'apiKey'] as const;
30
+ const TERMINAL = new Set(['completed', 'failed', 'cancelled']);
31
+
32
+ export interface RunEvent { readonly type: string; readonly sequence: number; readonly payload: Record<string, unknown> }
33
+ export interface BudgetObservation {
34
+ readonly enforce: 'hard' | 'advisory' | undefined;
35
+ readonly status: string;
36
+ readonly errorCode: string | undefined;
37
+ readonly events: readonly RunEvent[];
38
+ /** This major's names for the four events the witness reads. */
39
+ readonly names: { readonly reserved: string; readonly threshold: string; readonly exhausted: string; readonly capBreached: string };
40
+ }
41
+ export type BudgetRun =
42
+ | { readonly kind: 'skip'; readonly disposition: 'inapplicable' | 'blocked'; readonly reason: string }
43
+ | { readonly kind: 'refused'; readonly status: number; readonly code: string | undefined }
44
+ | { readonly kind: 'observed'; readonly observation: BudgetObservation };
45
+
46
+ export interface Finding {
47
+ readonly rule: 'create' | 'lifecycle' | 'hard-stop' | 'advisory' | 'content-free';
48
+ readonly ok: boolean;
49
+ readonly doc: string;
50
+ readonly message: string;
51
+ }
52
+
53
+ const isRecord = (v: unknown): v is Record<string, unknown> => v !== null && typeof v === 'object' && !Array.isArray(v);
54
+ async function http(fn: () => Promise<OpenWOPResponse>): Promise<OpenWOPResponse | null> {
55
+ try { return await fn(); } catch { return null; }
56
+ }
57
+
58
+ async function readEvents(profile: MajorProfile, runId: string): Promise<RunEvent[] | null> {
59
+ const out: RunEvent[] = [];
60
+ let after = 0;
61
+ for (let page = 0; page < 50; page++) {
62
+ const res = await http(() => driver.get(`${profile.runsPath}/${encodeURIComponent(runId)}/events/poll?timeout=1&afterSequence=${after}`));
63
+ if (res === null || res.status !== 200) return page === 0 ? null : out;
64
+ const batch = (res.json as { events?: unknown } | undefined)?.events;
65
+ if (!Array.isArray(batch) || batch.length === 0) return out;
66
+ for (const e of batch) {
67
+ if (!isRecord(e) || typeof e['type'] !== 'string' || typeof e['sequence'] !== 'number') continue;
68
+ out.push({ type: e['type'], sequence: e['sequence'], payload: isRecord(e['payload']) ? e['payload'] : {} });
69
+ after = Math.max(after, e['sequence']);
70
+ }
71
+ }
72
+ return out;
73
+ }
74
+
75
+ /** Create the budgeted run, wait for it to end, and return its log. */
76
+ export async function drive(profile: MajorProfile, discovery: unknown, timeoutMs = 20_000): Promise<BudgetRun> {
77
+ const family = profile.family(discovery, 'budget');
78
+ if (family === null) return { kind: 'skip', disposition: 'inapplicable', reason: 'the host does not advertise budget' };
79
+ const fragment = profile.runBudget(BUDGET_POLICY);
80
+ if (fragment === null) return { kind: 'skip', disposition: 'inapplicable', reason: `major ${profile.major} gives a run budget no createRun surface` };
81
+ const dims = family['dimensions'];
82
+ if (!Array.isArray(dims) || !dims.includes(DIMENSION)) return { kind: 'skip', disposition: 'inapplicable', reason: `budget.dimensions does not list ${DIMENSION} — the host does not enforce the dimension this witness spends` };
83
+ if (!isFixtureAdvertised(BUDGET_FIXTURE)) return { kind: 'skip', disposition: 'inapplicable', reason: `${BUDGET_FIXTURE} fixture not advertised — the host has not opted in to the unaided budget witness` };
84
+
85
+ const names = {
86
+ reserved: profile.eventType('budget.reserved'),
87
+ threshold: profile.eventType('budget.threshold.crossed'),
88
+ exhausted: profile.eventType('budget.exhausted'),
89
+ capBreached: profile.eventType('cap.breached'),
90
+ };
91
+ if (names.reserved === undefined || names.threshold === undefined || names.exhausted === undefined || names.capBreached === undefined) {
92
+ return { kind: 'skip', disposition: 'blocked', reason: `the event name map for major ${profile.major} is not on disk in this layout — the witness will not guess event names` };
93
+ }
94
+
95
+ const created = await http(() => driver.post(profile.runsPath, { workflowId: BUDGET_FIXTURE, ...fragment }));
96
+ if (created === null) return { kind: 'skip', disposition: 'blocked', reason: `POST ${profile.runsPath} unreachable (fetch failed)` };
97
+ if (created.status === 429) return { kind: 'skip', disposition: 'blocked', reason: `POST ${profile.runsPath} answered 429 — the run budget of the host, not the wire` };
98
+ const runId = (created.json as { runId?: unknown } | undefined)?.runId;
99
+ if (created.status !== 201 || typeof runId !== 'string') return { kind: 'refused', status: created.status, code: readErrorCode(created.json) ?? undefined };
100
+
101
+ const deadline = Date.now() + timeoutMs;
102
+ let snap: Record<string, unknown> | null = null;
103
+ for (;;) {
104
+ const res = await http(() => driver.get(`${profile.runsPath}/${encodeURIComponent(runId)}`));
105
+ snap = res?.status === 200 && isRecord(res.json) ? res.json : null;
106
+ if (snap !== null && TERMINAL.has(String(snap['status']))) break;
107
+ if (Date.now() > deadline) {
108
+ await http(() => driver.post(`${profile.runsPath}/${encodeURIComponent(runId)}/cancel`, {}));
109
+ return { kind: 'skip', disposition: 'blocked', reason: `the budgeted run did not reach a terminal status within ${timeoutMs} ms (last: ${snap === null ? 'unreadable' : String(snap['status'])})` };
110
+ }
111
+ await new Promise((r) => setTimeout(r, 250));
112
+ }
113
+ const events = await readEvents(profile, runId);
114
+ if (events === null) return { kind: 'skip', disposition: 'blocked', reason: `GET ${profile.runsPath}/{runId}/events/poll unreadable — the run ended but its log could not be read` };
115
+
116
+ const enforce = family['enforce'];
117
+ const error = snap['error'];
118
+ return {
119
+ kind: 'observed',
120
+ observation: {
121
+ enforce: enforce === 'hard' || enforce === 'advisory' ? enforce : undefined,
122
+ status: String(snap['status']),
123
+ errorCode: isRecord(error) && typeof error['code'] === 'string' ? error['code'] : undefined,
124
+ events,
125
+ names: names as BudgetObservation['names'],
126
+ },
127
+ };
128
+ }
129
+
130
+ /** What the log must show at this major. Pure. */
131
+ export function judge(profile: MajorProfile, o: BudgetObservation): Finding[] {
132
+ const home = `${profile.specRoot}/runs.md §budget section`;
133
+ const first = (type: string): RunEvent | undefined => o.events.find((e) => e.type === type);
134
+ const reserved = first(o.names.reserved);
135
+ const threshold = first(o.names.threshold);
136
+ const exhausted = first(o.names.exhausted);
137
+ const breach = o.events.find((e) => e.type === o.names.capBreached && typeof e.payload['kind'] === 'string' && (e.payload['kind'] as string).startsWith('budget-'));
138
+ const out: Finding[] = [];
139
+
140
+ out.push({ rule: 'lifecycle', ok: reserved !== undefined, doc: home, message: `a budgeted run MUST emit ${o.names.reserved}` });
141
+ out.push({ rule: 'lifecycle', ok: threshold !== undefined && typeof threshold.payload['percent'] === 'number', doc: home, message: `spending past thresholdPercent MUST emit ${o.names.threshold} with a numeric percent` });
142
+ out.push({ rule: 'lifecycle', ok: exhausted !== undefined, doc: home, message: `a budget that cannot cover the run MUST emit ${o.names.exhausted}` });
143
+ if (reserved !== undefined && threshold !== undefined && exhausted !== undefined) {
144
+ out.push({ rule: 'lifecycle', ok: reserved.sequence < threshold.sequence && threshold.sequence < exhausted.sequence, doc: home, message: `the log MUST order ${o.names.reserved} < ${o.names.threshold} < ${o.names.exhausted}` });
145
+ }
146
+
147
+ if (o.enforce === 'hard') {
148
+ out.push({ rule: 'hard-stop', ok: breach !== undefined && breach.payload['kind'] === CAP_KIND, doc: home, message: `hard exhaustion under onExhaustion: fail MUST emit ${o.names.capBreached} with kind ${CAP_KIND} (got ${breach === undefined ? 'none' : String(breach.payload['kind'])})` });
149
+ if (breach !== undefined && exhausted !== undefined) {
150
+ out.push({ rule: 'hard-stop', ok: exhausted.sequence <= breach.sequence, doc: home, message: `${o.names.capBreached} MUST NOT precede ${o.names.exhausted}` });
151
+ }
152
+ out.push({ rule: 'hard-stop', ok: o.status === 'failed' && o.errorCode === 'budget_exhausted', doc: home, message: `hard exhaustion MUST fail the run budget_exhausted (got ${o.status} ${o.errorCode ?? ''})`.trim() });
153
+ }
154
+ if (o.enforce === 'advisory') {
155
+ out.push({ rule: 'advisory', ok: breach === undefined && o.errorCode !== 'budget_exhausted', doc: home, message: 'an advisory host MUST NOT stop the run' });
156
+ }
157
+
158
+ const leaks: string[] = [];
159
+ for (const e of o.events) {
160
+ if (!e.type.startsWith('budget.') && e !== breach) continue;
161
+ for (const k of PRICING_KEYS) if (k in e.payload) leaks.push(`${e.type}.${k}`);
162
+ }
163
+ out.push({ rule: 'content-free', ok: leaks.length === 0, doc: home, message: `budget.* and cap.breached MUST NOT carry rate cards, unit prices or credentials (found ${leaks.join(', ') || 'none'})` });
164
+ return out;
165
+ }
@@ -0,0 +1,93 @@
1
+ /**
2
+ * What differs between protocol majors, as data.
3
+ *
4
+ * A scenario that is "the same requirement at another major" used to be a
5
+ * copy of the earlier file with the paths, event names and envelope rules
6
+ * edited by hand. The differences are few and regular, so they live in one
7
+ * table here. A shared witness (`backpressure-witness.ts`) takes a profile and
8
+ * never names a major; the scenario file for each major is a thin wrapper.
9
+ *
10
+ * **Adding a major** is one row in {@link MAJOR_PROFILES} plus one thin
11
+ * scenario file per ported witness. `majorProfile()` throws on a major with no
12
+ * row, so a missing row is a loud failure and never a silent fall back to an
13
+ * older major's rules.
14
+ *
15
+ * Keep a field here only when a shared witness reads it. A rule one major
16
+ * states and another does not is not a field: the witness asks the profile
17
+ * whether the rule binds (see `retryTiming`), and cites the document that
18
+ * states it.
19
+ */
20
+
21
+ import { capabilityFamily } from './discovery-capabilities.js';
22
+ import { codemapV1toV2 } from './era2-seed.js';
23
+
24
+ export interface MajorProfile {
25
+ readonly major: number;
26
+ /** The run collection, e.g. `/v1/runs` or `/runs`. */
27
+ readonly runsPath: string;
28
+ /** Headers a raw `fetch` must add to speak this major (the driver adds its own). */
29
+ readonly versionHeaders: Readonly<Record<string, string>>;
30
+ /** Where the corpus states this major's rules, for requirement citations. */
31
+ readonly specRoot: string;
32
+ /**
33
+ * Where retry timing lives on an error response. `header-and-details`: the
34
+ * envelope repeats `Retry-After` in `details.retryAfter`. `header-only`: the
35
+ * header is the only place, and `details.retryAfter*` is forbidden.
36
+ */
37
+ readonly retryTiming: 'header-and-details' | 'header-only';
38
+ /** The advertised record for a capability family, or `null` when the host does not advertise it. */
39
+ family(doc: unknown, key: string): Record<string, unknown> | null;
40
+ /**
41
+ * The `createRun` body fragment that carries a run budget policy, or `null`
42
+ * when this major has no run wire surface for one (a witness that needs it
43
+ * is then `inapplicable` at this major).
44
+ */
45
+ runBudget(policy: Readonly<Record<string, unknown>>): Record<string, unknown> | null;
46
+ /**
47
+ * This major's name for an event, given its era-1 name. `undefined` when the
48
+ * mapping is not on disk in this layout: the caller records that as an
49
+ * unread observation and never guesses a name.
50
+ */
51
+ eventType(era1Name: string): string | undefined;
52
+ }
53
+
54
+ const isRecord = (v: unknown): v is Record<string, unknown> => v !== null && typeof v === 'object' && !Array.isArray(v);
55
+
56
+ export const MAJOR_PROFILES: Readonly<Record<number, MajorProfile>> = {
57
+ 1: {
58
+ major: 1,
59
+ runsPath: '/v1/runs',
60
+ versionHeaders: {},
61
+ specRoot: 'spec/v1',
62
+ retryTiming: 'header-and-details',
63
+ // v1: a family is advertised when its record says `supported: true`.
64
+ family: (doc, key) => {
65
+ const rec = capabilityFamily(doc, key);
66
+ return isRecord(rec) && rec['supported'] === true ? rec : null;
67
+ },
68
+ // v1 budgets are driven through a host seam, not `createRun`.
69
+ runBudget: () => null,
70
+ eventType: (era1Name) => era1Name,
71
+ },
72
+ 2: {
73
+ major: 2,
74
+ runsPath: '/runs',
75
+ versionHeaders: { 'OpenWOP-Version': '2.0' },
76
+ specRoot: 'spec/v2/core',
77
+ retryTiming: 'header-only',
78
+ // v2: presence of the record is the advertisement.
79
+ family: (doc, key) => {
80
+ const rec = isRecord(doc) ? doc[key] : undefined;
81
+ return isRecord(rec) ? rec : null;
82
+ },
83
+ runBudget: (policy) => ({ configurable: { version: 1, budget: { ...policy } } }),
84
+ eventType: (era1Name) => codemapV1toV2().get(era1Name),
85
+ },
86
+ };
87
+
88
+ /** The profile for a major. Throws when the table has no row for it. */
89
+ export function majorProfile(major: number): MajorProfile {
90
+ const p = MAJOR_PROFILES[major];
91
+ if (p === undefined) throw new Error(`no MajorProfile for major ${major} — add a row to MAJOR_PROFILES in lib/major-profile.ts`);
92
+ return p;
93
+ }
@@ -27,7 +27,7 @@
27
27
  */
28
28
 
29
29
  import { scenarioFileOfItId } from './requirement-ids.js';
30
- import { PROFILE_FLOOR_SCENARIOS, floorMemberFiles } from './profiles.js';
30
+ import { DEPRECATED_PROFILE_ALIASES, PROFILE_FLOOR_SCENARIOS, floorMemberFiles } from './profiles.js';
31
31
  import { targetMajor } from './seams.js';
32
32
  import { PKG_ROOT_PATH } from './paths.js';
33
33
  import { v2ProfileFloorFiles } from './requirement-registry.js';
@@ -478,7 +478,16 @@ export function deriveRequirementDispositions(
478
478
  // no member failing; `inapplicable`/`skipped` members never satisfy it. v1
479
479
  // hand table only, like the prefix groups (the major-2 floors have none).
480
480
  const groups = new Map<string, readonly string[]>();
481
- if (!v2FloorsActive()) for (const floor of Object.values(PROFILE_FLOOR_SCENARIOS)) for (const g of floor.requiredAnyOf ?? []) groups.set(requirementIdForAnyOf(g), g);
481
+ //
482
+ // Only for a CLAIMED profile's floor (2.45.4). The row is `blocked` when no
483
+ // member was witnessed, and a v1 floor reads `inapplicable` as certifiable, so
484
+ // the row cannot simply say `inapplicable`. But 2.45.3 wrote it for every
485
+ // floor in the table: a host that does not advertise secrets records both
486
+ // members `inapplicable`, got one `blocked` row for a profile it never
487
+ // claimed, and a bundle with any `blocked` row certifies nothing (RFC 0168
488
+ // §E.1). A group nobody claims is not a requirement on this host.
489
+ const claimed = new Set(claimedProfiles.map((p) => DEPRECATED_PROFILE_ALIASES[p] ?? p));
490
+ if (!v2FloorsActive()) for (const [profile, floor] of Object.entries(PROFILE_FLOOR_SCENARIOS)) if (claimed.has(profile)) for (const g of floor.requiredAnyOf ?? []) groups.set(requirementIdForAnyOf(g), g);
482
491
  for (const [id, members] of [...groups.entries()].sort((a, b) => a[0].localeCompare(b[0]))) {
483
492
  const scenarioId = `anyof:${members.join('|')}`;
484
493
  const matching = members.map((f) => perFile.get(f)).filter((r): r is DerivedRequirement => r !== undefined);
@@ -0,0 +1,130 @@
1
+ /**
2
+ * A scratch host: the smallest HTTP server a shared witness can be proven
3
+ * against, in both directions, inside the suite's own self-tests.
4
+ *
5
+ * Some families are advertised by no reference host, so a new witness for them
6
+ * cannot be sabotage-proved against a real one. A witness that was never seen
7
+ * to fail is not a witness. The scratch host serves a discovery document and a
8
+ * scripted run surface, and each self-test turns on ONE defect to show the
9
+ * witness convicts it.
10
+ *
11
+ * It is a test double, not a host: it executes nothing, and nothing it does is
12
+ * evidence about any implementation. It takes a {@link MajorProfile}, so the
13
+ * same double proves the next major's port.
14
+ */
15
+
16
+ import { createServer, type IncomingMessage, type Server, type ServerResponse } from 'node:http';
17
+ import type { AddressInfo } from 'node:net';
18
+ import type { MajorProfile } from './major-profile.js';
19
+
20
+ export interface ScriptedEvent { readonly type: string; readonly payload?: Record<string, unknown> | undefined }
21
+ export interface ScriptedRun {
22
+ readonly status: 'completed' | 'failed' | 'running';
23
+ readonly error?: { readonly code: string; readonly message: string };
24
+ readonly events: readonly ScriptedEvent[];
25
+ }
26
+ export interface ScratchRefusal {
27
+ readonly status: number;
28
+ readonly headers?: Readonly<Record<string, string>>;
29
+ readonly body: Record<string, unknown>;
30
+ }
31
+ export interface ScratchOptions {
32
+ readonly profile: MajorProfile;
33
+ /** The discovery document, served as given. */
34
+ readonly discovery: Readonly<Record<string, unknown>>;
35
+ /** What a created run looks like, by workflowId and request body. Default: an empty running run. */
36
+ readonly script?: ((workflowId: string, body: Readonly<Record<string, unknown>>) => ScriptedRun) | undefined;
37
+ /** Concurrent in-flight requests admitted before {@link refusal} answers. Unset: no cap. */
38
+ readonly inflightCap?: number | undefined;
39
+ readonly refusal?: ScratchRefusal | undefined;
40
+ }
41
+
42
+ interface StoredRun { readonly id: string; run: ScriptedRun }
43
+
44
+ export class ScratchHost {
45
+ private server: Server | undefined;
46
+ private readonly runs = new Map<string, StoredRun>();
47
+ private inflight = 0;
48
+ private seq = 0;
49
+ url = '';
50
+
51
+ constructor(private opts: ScratchOptions) {}
52
+
53
+ /** Swap behaviour between cases. The address stays the same, since the driver caches its base URL per file. */
54
+ reconfigure(next: Partial<ScratchOptions>): void {
55
+ this.opts = { ...this.opts, ...next };
56
+ this.runs.clear();
57
+ }
58
+
59
+ async start(): Promise<string> {
60
+ this.server = createServer((req, res) => { void this.handle(req, res); });
61
+ await new Promise<void>((r) => this.server!.listen(0, '127.0.0.1', r));
62
+ this.url = `http://127.0.0.1:${(this.server.address() as AddressInfo).port}`;
63
+ return this.url;
64
+ }
65
+
66
+ async stop(): Promise<void> {
67
+ const s = this.server;
68
+ if (s === undefined) return;
69
+ s.closeAllConnections();
70
+ await new Promise<void>((r) => s.close(() => r()));
71
+ }
72
+
73
+ private send(res: ServerResponse, status: number, body: unknown, headers: Readonly<Record<string, string>> = {}): void {
74
+ res.writeHead(status, { 'content-type': 'application/json', ...headers });
75
+ res.end(JSON.stringify(body));
76
+ }
77
+
78
+ private async body(req: IncomingMessage): Promise<Record<string, unknown>> {
79
+ const chunks: Buffer[] = [];
80
+ for await (const c of req) chunks.push(c as Buffer);
81
+ try {
82
+ const v: unknown = JSON.parse(Buffer.concat(chunks).toString('utf8') || '{}');
83
+ return v !== null && typeof v === 'object' && !Array.isArray(v) ? (v as Record<string, unknown>) : {};
84
+ } catch { return {}; }
85
+ }
86
+
87
+ private async handle(req: IncomingMessage, res: ServerResponse): Promise<void> {
88
+ const path = (req.url ?? '/').split('?')[0] ?? '/';
89
+ const runs = this.opts.profile.runsPath;
90
+ // Discovery is never counted against the cap.
91
+ if (req.method === 'GET' && path === '/.well-known/openwop') return this.send(res, 200, this.opts.discovery);
92
+
93
+ const cap = this.opts.inflightCap;
94
+ if (cap !== undefined && this.inflight >= cap) {
95
+ const r = this.opts.refusal ?? { status: 503, headers: { 'retry-after': '1' }, body: { error: 'service_unavailable', message: 'at capacity' } };
96
+ return this.send(res, r.status, r.body, r.headers);
97
+ }
98
+ this.inflight++;
99
+ res.on('close', () => { this.inflight--; });
100
+
101
+ if (req.method === 'POST' && path === runs) {
102
+ const body = await this.body(req);
103
+ const workflowId = String(body['workflowId'] ?? '');
104
+ const id = `scratch/run-${++this.seq}`;
105
+ this.runs.set(id, { id, run: this.opts.script?.(workflowId, body) ?? { status: 'running', events: [] } });
106
+ return this.send(res, 201, { runId: id, status: 'running' });
107
+ }
108
+ if (req.method === 'POST' && path === `${runs}:bulk-cancel`) return this.send(res, 200, { results: [] });
109
+
110
+ const m = new RegExp(`^${runs}/([^/]+)(/events/poll|/events|/cancel)?$`).exec(path);
111
+ const stored = m ? this.runs.get(decodeURIComponent(m[1] as string)) : undefined;
112
+ if (!m || stored === undefined) return this.send(res, 404, { error: 'not_found', message: 'no such route or run' });
113
+ const tail = m[2];
114
+ if (tail === '/cancel') return this.send(res, 200, { status: 'cancelled' });
115
+ if (tail === '/events') {
116
+ // Held open until the client aborts: this is what occupies a slot.
117
+ res.writeHead(200, { 'content-type': 'text/event-stream' });
118
+ res.write(': held\n\n');
119
+ return;
120
+ }
121
+ if (tail === '/events/poll') {
122
+ const after = Number(new URL(req.url ?? '/', 'http://x').searchParams.get('afterSequence') ?? '0');
123
+ const events = stored.run.events
124
+ .map((e, i) => ({ eventId: `e${i + 1}`, runId: stored.id, type: e.type, payload: e.payload ?? {}, sequence: i + 1, schemaVersion: 2, timestamp: new Date(0).toISOString() }))
125
+ .filter((e) => e.sequence > after);
126
+ return this.send(res, 200, { events });
127
+ }
128
+ return this.send(res, 200, { runId: stored.id, workflowId: 'scratch', status: stored.run.status, ...(stored.run.error === undefined ? {} : { error: stored.run.error }) });
129
+ }
130
+ }
@@ -135,12 +135,9 @@ describe('RFC 0148 §A (S6) — the runner derivation', () => {
135
135
  expect(p.disposition).toBe('executed-pass');
136
136
  expect(p.scenarioId).toBe(`${prefix}*`);
137
137
  }
138
- // An any-of group belongs to its own profile's floor (RFC 0229 §E). None of
139
- // its members is in this synthetic report, so its summary row is honestly
140
- // `blocked`, and it is not one of the requirements this floor certifies on.
141
- const foreignAnyOf = d.requirements.filter((r) => r.scenarioId.startsWith('anyof:') && !(floor.requiredAnyOf ?? []).some((g) => g.join('|') === r.scenarioId.slice('anyof:'.length)));
142
- for (const r of foreignAnyOf) expect(r.disposition, r.requirementId).toBe('blocked');
143
- expect(d.totals.executedPass).toBe(d.requirements.length - foreignAnyOf.length);
138
+ // An any-of group's summary row is written only for a claimed profile's floor
139
+ // (2.45.4), so this floor's rows are all there is and every one passed.
140
+ expect(d.totals.executedPass).toBe(d.requirements.length);
144
141
  });
145
142
 
146
143
  it('a floor file that vitest passed but that recorded NOTHING is unclassified when a ledger exists — silence is not a witness', () => {
@@ -0,0 +1,94 @@
1
+ /**
2
+ * v2 — a run budget is reserved, crossed, exhausted and enforced
3
+ * (`spec/v2/core/runs.md` §`budget` section). The v1 twin is
4
+ * `budget-enforcement`, which drives a host seam; at major 2 the budget rides
5
+ * on `createRun` (`configurable.budget`), so this witness is unaided. The
6
+ * logic is `lib/budget-witness.ts`; what differs between majors is the profile
7
+ * row in `lib/major-profile.ts`.
8
+ *
9
+ * The suite creates a run of `conformance-budget-tool-calls` (three scripted
10
+ * tool calls) with `{ maxToolCalls: 2, thresholdPercent: 50, onExhaustion:
11
+ * "fail" }` and reads its log through the poll.
12
+ *
13
+ * lifecycle `budget.reserved`, `budget.threshold-crossed` (numeric
14
+ * `percent`) and `budget.exhausted`, in that order. Both
15
+ * enforce modes owe these.
16
+ * enforcement `enforce: hard`: `cap.breached` with kind
17
+ * `budget-tool-calls`, not before `budget.exhausted`, and the
18
+ * run fails `budget_exhausted`. `enforce: advisory`: the run is
19
+ * not stopped.
20
+ * content-free no `budget.*` or `cap.breached` payload carries a rate card,
21
+ * a unit price or a credential.
22
+ *
23
+ * Dispositions: no `budget`, `toolCalls` not in `budget.dimensions`, or the
24
+ * fixture unadvertised ⇒ `inapplicable`. The fixture is the opt-in: a host
25
+ * that advertises `budget` and has not seeded it is unwitnessed here, not
26
+ * blocked. A valid budgeted create that is refused ⇒ `executed-fail`.
27
+ *
28
+ * Not ported: v1's model-denied leg (`budget_model_denied`). It needs a
29
+ * fixture that resolves a model, and none does so without a seam.
30
+ *
31
+ * Proven both ways against the scratch host in `lib/budget-witness.test.ts`.
32
+ *
33
+ * @see spec/v2/core/runs.md §`budget` section
34
+ * @see conformance/fixtures.md §conformance-budget-tool-calls
35
+ */
36
+
37
+ import { describe, it, expect } from 'vitest';
38
+ import { v2Discovery } from '../lib/v2.js';
39
+ import { softSkip } from '../lib/soft-skip.js';
40
+ import { req } from '../lib/requirement-ids.js';
41
+ import { majorProfile } from '../lib/major-profile.js';
42
+ import { drive, judge, type BudgetRun, type Finding } from '../lib/budget-witness.js';
43
+
44
+ const PROFILE = majorProfile(2);
45
+ const DOC = 'spec/v2/core/runs.md §budget section';
46
+ const ID_LIFECYCLE = 'openwop.requirement.runs.budget-lifecycle';
47
+ const ID_ENFORCEMENT = 'openwop.requirement.runs.budget-enforcement';
48
+ const ID_CONTENT_FREE = 'openwop.requirement.runs.budget-content-free';
49
+
50
+ /** One budgeted run per file: three legs read the same log. */
51
+ let once: Promise<BudgetRun> | undefined;
52
+ function budgetRun(): Promise<BudgetRun> {
53
+ once ??= (async (): Promise<BudgetRun> => {
54
+ let doc: Record<string, unknown> | null;
55
+ try { doc = await v2Discovery(); } catch { doc = null; }
56
+ if (!doc) return { kind: 'skip', disposition: 'blocked', reason: 'v2 discovery unreachable — /.well-known/openwop did not answer 200 with a JSON body under OpenWOP-Version: 2.0' };
57
+ return drive(PROFILE, doc);
58
+ })();
59
+ return once;
60
+ }
61
+
62
+ type Leg = { readonly skip: { readonly disposition: 'inapplicable' | 'blocked'; readonly reason: string } } | { readonly findings: Finding[] };
63
+
64
+ /** The findings for one leg, or why there are none. A refused create fails the leg here. */
65
+ async function leg(id: string, rules: ReadonlyArray<Finding['rule']>): Promise<Leg> {
66
+ const r = await budgetRun();
67
+ if (r.kind === 'skip') return { skip: { disposition: r.disposition, reason: r.reason } };
68
+ if (r.kind === 'refused') {
69
+ expect(r.status, req(id, DOC, `a host that advertises budget and the fixture MUST accept a valid configurable.budget (got ${r.status} ${r.code ?? ''})`.trim())).toBe(201);
70
+ return { findings: [] }; // unreachable: a refusal is never 201
71
+ }
72
+ return { findings: judge(PROFILE, r.observation).filter((f) => rules.includes(f.rule)) };
73
+ }
74
+
75
+ describe('v2 budget enforcement (runs.md §budget section)', () => {
76
+ it('a budgeted run emits budget.reserved, budget.threshold-crossed and budget.exhausted in order', async () => {
77
+ const l = await leg(ID_LIFECYCLE, ['lifecycle']);
78
+ if ('skip' in l) return softSkip(l.skip.disposition, l.skip.reason);
79
+ for (const f of l.findings) expect(f.ok, req(ID_LIFECYCLE, f.doc, f.message)).toBe(true);
80
+ });
81
+
82
+ it('a hard host stops the run budget_exhausted after cap.breached, and an advisory host does not stop it', async () => {
83
+ const l = await leg(ID_ENFORCEMENT, ['hard-stop', 'advisory']);
84
+ if ('skip' in l) return softSkip(l.skip.disposition, l.skip.reason);
85
+ if (l.findings.length === 0) return softSkip('inapplicable', 'budget.enforce is not advertised — neither the hard stop nor the advisory rule binds');
86
+ for (const f of l.findings) expect(f.ok, req(ID_ENFORCEMENT, f.doc, f.message)).toBe(true);
87
+ });
88
+
89
+ it('no budget.* or cap.breached payload carries pricing or a credential', async () => {
90
+ const l = await leg(ID_CONTENT_FREE, ['content-free']);
91
+ if ('skip' in l) return softSkip(l.skip.disposition, l.skip.reason);
92
+ for (const f of l.findings) expect(f.ok, req(ID_CONTENT_FREE, f.doc, f.message)).toBe(true);
93
+ });
94
+ });
@@ -0,0 +1,74 @@
1
+ /**
2
+ * v2 — a host at capacity answers `503 service_unavailable` with `Retry-After`
3
+ * (`spec/v2/core/conformance.md` §Production profile, `backpressure` facet).
4
+ * The v1 twin is `production-backpressure`; the logic both share is
5
+ * `lib/backpressure-witness.ts`, and what differs between majors is the
6
+ * profile row in `lib/major-profile.ts`.
7
+ *
8
+ * Unaided. Gated on `production.backpressure.inflightCap`, the number a host
9
+ * advertises so the suite can saturate it: `inflightCap` event streams hold
10
+ * the slots and one more request is sent.
11
+ *
12
+ * refusal the extra request answers `503`, code `service_unavailable`,
13
+ * with `Retry-After`; where `retryAfterSeconds` is advertised
14
+ * the header equals it.
15
+ * retry timing the refusal carries no `details.retryAfter*`
16
+ * (`errors.md` §Retry timing). This is where v2 differs from
17
+ * v1, which required `details.retryAfter`.
18
+ *
19
+ * Dispositions: no `production`, no `backpressure`, no `inflightCap`, an
20
+ * `inflightCap` above 64 (the most streams the suite holds open), or the hold
21
+ * or probe fixture unadvertised ⇒ `inapplicable`. A slot that cannot be
22
+ * filled, or a probe with no response ⇒ `blocked`.
23
+ *
24
+ * v1's "discovery is exempt from the cap" leg is not ported: no v2 document
25
+ * states it.
26
+ *
27
+ * Run with `--no-file-parallelism`. Proven both ways against the scratch host
28
+ * in `lib/backpressure-witness.test.ts`.
29
+ *
30
+ * @see spec/v2/core/conformance.md §Production profile
31
+ * @see spec/v2/core/errors.md §Retry timing
32
+ */
33
+
34
+ import { describe, it, expect } from 'vitest';
35
+ import { v2Discovery } from '../lib/v2.js';
36
+ import { softSkip } from '../lib/soft-skip.js';
37
+ import { req } from '../lib/requirement-ids.js';
38
+ import { majorProfile } from '../lib/major-profile.js';
39
+ import { judge, saturate, type Saturation } from '../lib/backpressure-witness.js';
40
+
41
+ const PROFILE = majorProfile(2);
42
+ const ID_REFUSAL = 'openwop.requirement.production.backpressure-refusal';
43
+ const ID_TIMING = 'openwop.requirement.0171.error-registry.no-retry-details';
44
+
45
+ /** One saturation per file: two legs read the same refusal. */
46
+ let once: Promise<Saturation> | undefined;
47
+ function saturation(): Promise<Saturation> {
48
+ once ??= (async (): Promise<Saturation> => {
49
+ let doc: Record<string, unknown> | null;
50
+ try { doc = await v2Discovery(); } catch { doc = null; }
51
+ if (!doc) return { kind: 'skip', disposition: 'blocked', reason: 'v2 discovery unreachable — /.well-known/openwop did not answer 200 with a JSON body under OpenWOP-Version: 2.0' };
52
+ return saturate(PROFILE, doc);
53
+ })();
54
+ return once;
55
+ }
56
+
57
+ describe('v2 production backpressure (conformance.md §Production profile)', () => {
58
+ it('with inflightCap slots held, the next request is refused 503 service_unavailable with Retry-After', async () => {
59
+ const s = await saturation();
60
+ if (s.kind === 'skip') return softSkip(s.disposition, s.reason);
61
+ for (const f of judge(PROFILE, s.refusal).filter((x) => x.rule !== 'retry-timing')) {
62
+ expect(f.ok, req(ID_REFUSAL, f.doc, f.message)).toBe(true);
63
+ }
64
+ });
65
+
66
+ it('the refusal carries its retry timing in the Retry-After header only', async () => {
67
+ const s = await saturation();
68
+ if (s.kind === 'skip') return softSkip(s.disposition, s.reason);
69
+ if (s.refusal.status !== 503) return softSkip('blocked', `no 503 was observed at inflightCap + 1 (got ${s.refusal.status}) — there is no refusal whose details could be read`);
70
+ for (const f of judge(PROFILE, s.refusal).filter((x) => x.rule === 'retry-timing')) {
71
+ expect(f.ok, req(ID_TIMING, f.doc, f.message)).toBe(true);
72
+ }
73
+ });
74
+ });