@openwop/openwop-conformance 2.1.6 → 2.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/CHANGELOG.md +98 -0
  2. package/README.md +2 -2
  3. package/dist/spec-artifacts.lock.json +2 -2
  4. package/fixtures/conformance-approval-refine.json +45 -0
  5. package/fixtures/conformance-envelope-nl-to-format-engaged.json +1 -1
  6. package/fixtures/conformance-envelope-recovery-applied.json +2 -2
  7. package/fixtures/conformance-envelope-refusal.json +1 -1
  8. package/fixtures/conformance-envelope-retry-attempted.json +2 -2
  9. package/fixtures/conformance-envelope-retry-exhausted.json +1 -1
  10. package/fixtures/conformance-envelope-truncated.json +1 -1
  11. package/fixtures/conformance-envelope-truncation-cap-exhaustion.json +1 -1
  12. package/fixtures/conformance-phase4-nondet-tool.json +2 -2
  13. package/fixtures/conformance-phase4-replay-divergence.json +2 -2
  14. package/fixtures.md +16 -0
  15. package/package.json +2 -2
  16. package/requirements.json +93 -12
  17. package/scenario-majors.json +5 -2
  18. package/schemas/CORPUS-STAMP.json +21 -21
  19. package/src/lib/bound-id.ts +69 -0
  20. package/src/scenarios/envelope-completion-distinguishes-truncation.test.ts +20 -11
  21. package/src/scenarios/envelope-nl-to-format-engaged.test.ts +1 -1
  22. package/src/scenarios/envelope-recovery-applied.test.ts +1 -1
  23. package/src/scenarios/envelope-refusal-shape.test.ts +1 -1
  24. package/src/scenarios/envelope-retry-attempted.test.ts +1 -1
  25. package/src/scenarios/envelope-retry-exhausted.test.ts +1 -1
  26. package/src/scenarios/envelope-truncated.test.ts +1 -1
  27. package/src/scenarios/envelope-truncation-cap-exhaustion.test.ts +1 -1
  28. package/src/scenarios/interrupt-approval.test.ts +87 -0
  29. package/src/scenarios/replay-divergence-at-refusal.test.ts +2 -2
  30. package/src/scenarios/v2-bound-id-path-projection.test.ts +121 -0
  31. package/src/scenarios/v2-created-run-readable.test.ts +1 -1
  32. package/src/scenarios/v2-id-grammar.test.ts +1 -1
  33. package/src/scenarios/v2-stream-sse-projection.test.ts +1 -1
  34. package/src/setup.ts +38 -0
package/CHANGELOG.md CHANGED
@@ -1,5 +1,103 @@
1
1
  # `@openwop/openwop-conformance` Changelog
2
2
 
3
+ ## [2.2.0] — 2026-09-16 — the suite wipes the host's mock program store between scenario files, and a bound id travels as one segment
4
+
5
+ **Why a minor and not the 2.1.8 that was pinned for two days.** `PUBLISHING.md`
6
+ §"Versioning alignment": a conformance scenario addition is a minor bump, and
7
+ #1359 added `v2-bound-id-path-projection` (its header says so). No 2.1.x patch
8
+ ever added a scenario file; this release does, so it is 2.2.0. Nothing was
9
+ published as 2.1.8.
10
+
11
+ This entry was first written on 2026-09-14 and said *"No scenario behaviour
12
+ changes."* That was true when written and false by the time the tag was cut, so
13
+ the entry is rewritten rather than appended to: two behaviour changes rode this
14
+ version after the pin was bumped, and a package whose changelog denies them is a
15
+ package lying about itself.
16
+
17
+ ### Fixed
18
+
19
+ - **The suite now wipes the host's conformance-mock program store between
20
+ scenario FILES** (openwop #1357; half two of two, paired with openwop-app
21
+ #3870, which adds the seam). `src/setup.ts` posts
22
+ `POST /v1/host/openwop-app/test/mock-ai/reset` in a per-file `afterAll`.
23
+ Why: a host that keeps mock programs in a module-level map keyed by `nodeId`
24
+ retains any program a scenario does not fully drain, and the four
25
+ envelope-truncation scenarios seed `finishReason: 'length'` programs on
26
+ purpose — so a later file dispatching on a colliding node consumed the
27
+ leftovers and failed `envelope_truncation_unrecoverable` with nothing in that
28
+ file to explain it. Measured on `replay-observable-sequence-determinism`: red
29
+ in-suite, green alone, five runs on unchanged bases, one red at load 3.4 on an
30
+ idle box — not contention, and a wrong value rather than a missing result.
31
+ 2.1.7's per-fixture node ids narrowed the collision; this removes the
32
+ leftover. **Best-effort by design:** a host without the seam answers 404 and
33
+ the call is swallowed, because a suite must never fail a compliant host for
34
+ lacking a TEST seam — so on such a host the leak, if it has one, persists
35
+ silently. Per file rather than per test because programs are seeded for a
36
+ whole scenario's attempt sequence.
37
+
38
+ ### Changed
39
+
40
+ - **A tenant-bound id travels as one `~`-escaped path segment** (RFC 0184 §A.1,
41
+ openwop #1359; §A.5 exactly-once projection, #1360). A bound id is two
42
+ segments joined by `/`; a path parameter is one. The corpus said `%2F`
43
+ carried the separator, and a tier-1 host's front door decoded it back to `/`
44
+ before forwarding, so every bound id was unreachable through its own front
45
+ door. The escape marker is now `~` (RFC 3986 unreserved — an intermediary has
46
+ no license to rewrite it): every byte outside `[A-Za-z0-9._-]` becomes
47
+ `~XX`, total over bytes so a later grammar widening cannot invalidate a
48
+ projection already on the wire. New `src/lib/bound-id.ts` codec;
49
+ `v2-id-grammar` asserts the new form; new scenario
50
+ `v2-bound-id-path-projection` creates one run and asserts the projection is
51
+ applied exactly once (the codec is deliberately not idempotent — `a~3Ab` →
52
+ `a~7E3Ab` — so a double projection corrupts silently, and both reporting hosts
53
+ found a read path that already composed one). **Hosts: a v2 path parameter
54
+ carrying a bound id is now expected in the `~` form. `%2F` is no longer the
55
+ contract.**
56
+ - **RFC 0183 requirement rows** (openwop #1358): `requirements.json` gains the
57
+ ids for the resume actions `interruptResolved` cannot record, and
58
+ `fixtures.md` documents `conformance-approval-refine`. The RFC is `Draft`; no
59
+ scenario asserts these rows yet.
60
+
61
+ ### Packaging
62
+
63
+ - `spec/v2/declaration.json` gains an optional `normativeText` on family
64
+ entries (RFC 0169, openwop #1351): the path(s) where a family's behaviour is
65
+ actually written, as distinct from `section`, which names the declaration
66
+ site. The declaration and the schemas generated from it ship in the tarball,
67
+ which is what first required the identity bump.
68
+
69
+ ## [2.1.7] — 2026-09-13 — nine fixtures shared one programmable node id
70
+
71
+ The mock-AI program seam is keyed by `nodeId` alone — no run, no workflow, no
72
+ tenant. `host-sample-test-seams.md` §5 stated the isolation that relies on:
73
+ *"each conformance scenario uses a unique fixture (and therefore unique
74
+ `nodeId`)"*. The corpus did not honour it. Nine fixtures declared a node called
75
+ `structured-call`, and vitest runs scenario files in parallel, so any two of
76
+ them could overwrite each other's program mid-run.
77
+
78
+ That is the long-standing `replay-observable-sequence-determinism` flake.
79
+ That scenario never programs the mock: it runs `conformance-phase4-nondet-tool`
80
+ and inherits whatever the last writer left. When an envelope scenario had just
81
+ staged `{ stopReason: 'safety' }`, the replay fixture's own node refused, the
82
+ source run reached `failed`, and the scenario reported
83
+
84
+ expected 'failed' to be 'completed'
85
+
86
+ — an assertion about replay determinism failing for a reason that has nothing
87
+ to do with replay. It was green in isolation, red under the full gate, and did
88
+ not track load, because contention was never the variable; overlap was.
89
+
90
+ Each of the nine nodes is now named for its fixture
91
+ (`refusal-structured-call`, `nondet-structured-call`, and so on), which is what
92
+ §5 always claimed. `conformance/scripts/check-mock-ai-node-ids-unique.mjs`
93
+ holds the line: every fixture node dispatching to the mock provider must be
94
+ owned by exactly one fixture.
95
+
96
+ **Hosts: re-register the conformance fixtures when you pin 2.1.7.** The node
97
+ ids inside those nine workflow definitions changed. A host still serving the
98
+ old definitions will program a node its fixture no longer has, and the envelope
99
+ scenarios will fail loudly rather than silently — but they will fail.
100
+
3
101
  ## [2.1.6] — 2026-09-13 — the unknown-run probe is tenant-bound
4
102
 
5
103
  `v2-compensation-read-projection` asserted `404 not_found` for an unknown run
package/README.md CHANGED
@@ -117,7 +117,7 @@ Exit code is non-zero on any failed assertion. `--certify` distinguishes: `0`
117
117
 
118
118
  ## What's Covered
119
119
 
120
- The current suite has 517 scenario files under `src/scenarios/`.
120
+ The current suite has 518 scenario files under `src/scenarios/`.
121
121
  - 2026-09-03 (suite `1.157.0 -> 1.158.0`, gap G17): NEW `idempotency-concurrent-claim.test.ts` — drives the new `host-sample-test-seams.md` §25 concurrent duplicate-delivery seam for the RFC 0150 §B / `idempotency.md` §"Concurrent duplicates (Layer 2)" atomic-claim MUST, which is unconditional and had no witness of any kind. Asserts every executor mints the SAME `logicalInvocationId` **before** asserting `delivered === 1` — without the identity check a host passes by minting different ids and never colliding, one effect because nothing raced. Not profile-gated and so not opt-out-able (the obligation is unconditional); an unmounted seam records `blocked`, which is not certifiable. Graduates `layer2-invocation-claim-atomic` reference-impl -> protocol.
122
122
  - 2026-08-19 (suite `1.137.0 → 1.138.0`): NEW `durability-poison-exhaustion.test.ts` — RFC 0158 §C.8, the FIRST row of that RFC's conformance table to land. Asserts what `failure-path.test.ts` cannot: not just that deterministically failing work reaches terminal, but that attempts STOP — counted on the log, re-counted after a scaled quiet window, asserted unchanged. A host still redelivering records more. Seam-gated on the existing event-log seam (`blocked` = unobservable, not unmet) and outside every profile floor.
123
123
  - 2026-08-19 (suite `1.136.15 → 1.137.0`): NEW `replay-fanout-suppression.test.ts` — capability-gated on `webhooks.supported`, **outside every profile floor**; witnesses `replay.md` §"Host-initiated fan-out is an external effect", which was the largest normative MUST NOT on the replay surface with no scenario and no SECURITY invariant. Three legs in ONE `it` against ONE receiver and ONE subscription — a positive control, the MUST NOT, and a `branch` boundary leg — because "no delivery arrived" passes identically when delivery never worked, so absence is asserted only after presence is proven on that exact wiring. A host with an SSRF guard correctly refuses the loopback receiver and records `blocked`: **unobservable, not unmet.**
@@ -462,7 +462,7 @@ Server-required (added in 1.7.0):
462
462
  | ------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
463
463
  | **Redaction** | [`capabilities.md`](../spec/v1/capabilities.md) §"Secrets" + NFR-7 + §"aiProviders" | Vendor-neutral assertions that the server doesn't leak secret material. Three scenario groups: (a) discovery shape contract — `secrets` + `aiProviders` advertisements are well-formed regardless of `secrets.supported`; when `supported === true`, scopes MUST be non-empty + `resolution === 'host-managed'`; `byok ⊆ supported`. (b) bearer-token redaction — invalid Bearer canary in `Authorization` header is not echoed in the 401 response body. (c) credentialRef echo control — gated on `secrets.supported === true`; canary planted in `configurable.ai.credentialRef` MUST NOT appear in any RunEvent payload (poll-based capture; transport-agnostic). Uses runtime-built canary fixtures (`lib/canaries.ts`) that defeat static secret scanners. 6 scenarios. |
464
464
 
465
- Current source tree: 517 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
465
+ Current source tree: 518 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
466
466
 
467
467
  ## Remaining Gaps
468
468
 
@@ -1,5 +1,5 @@
1
1
  {
2
2
  "package": "@openwop/spec-artifacts",
3
- "version": "2.1.6",
4
- "stampSha256": "366973a92d7c7b3a59f4c872e35c720a6e8075d3f8e2da6920eebca507609b42"
3
+ "version": "2.2.0",
4
+ "stampSha256": "b5398e169f00362c4f6821354b1746c79fff01a341488ebfb2cf2cbd20f19b4a"
5
5
  }
@@ -0,0 +1,45 @@
1
+ {
2
+ "id": "conformance-approval-refine",
3
+ "name": "Conformance: Approval (refine)",
4
+ "version": "1.0",
5
+ "description": "Suspends on an approval gate offering accept|reject|refine. Resume with {action:'refine', refineFeedback:{...}} exercises RFC 0183 \u00a7A.1/\u00a7A.2 \u2014 the resolved payload must carry the action and its structured feedback.",
6
+ "nodes": [
7
+ {
8
+ "id": "gate",
9
+ "typeId": "core.approvalGate",
10
+ "name": "Approval Gate",
11
+ "position": {
12
+ "x": 0,
13
+ "y": 0
14
+ },
15
+ "config": {
16
+ "title": "Conformance approval (refine-capable)",
17
+ "description": "Conformance suite \u2014 please accept to complete the run.",
18
+ "actions": [
19
+ "accept",
20
+ "reject",
21
+ "refine"
22
+ ]
23
+ },
24
+ "inputs": {}
25
+ }
26
+ ],
27
+ "edges": [],
28
+ "triggers": [
29
+ {
30
+ "id": "manual",
31
+ "type": "manual",
32
+ "enabled": true
33
+ }
34
+ ],
35
+ "variables": [],
36
+ "metadata": {
37
+ "tags": [
38
+ "conformance",
39
+ "hitl"
40
+ ]
41
+ },
42
+ "settings": {
43
+ "timeout": 0
44
+ }
45
+ }
@@ -5,7 +5,7 @@
5
5
  "description": "Drives `core.ai.structuredOutput` against the conformance-only `mock` provider. The conformance scenario POSTs a 3-entry program: all three attempts return natural-language prose (no JSON sigil at the start). The host's `dispatchStructured` retry loop exhausts on parse-error, detects the NL shape, and fires ONE additional dispatch with a corrective coercion fragment — emitting `envelope.nlToFormat.engaged { originalEnvelopeType, fallbackCalls: 1 }` BEFORE the secondary call. The pre-seeded 4th program entry returns valid JSON, the schema validates, and the run terminates `completed`.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "nl-to-format-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider (NL responses)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -2,10 +2,10 @@
2
2
  "id": "conformance-envelope-recovery-applied",
3
3
  "name": "Conformance: envelope.recovery.applied (RFC 0032 §B.6)",
4
4
  "version": "1.0",
5
- "description": "Drives `core.ai.structuredOutput` against the conformance-only `mock` provider. The conformance scenario POSTs a 1-entry program to `/v1/host/sample/test/mock-ai/program` (keyed by `nodeId: 'structured-call'`) BEFORE starting the run: the mock returns a markdown-fenced JSON envelope (e.g., ```json\\n{\"result\":\"ok\"}\\n```). The host's `dispatchStructured` lenient-parse fallback strips the fence via `tryLenientParse(text)`, emits exactly one `envelope.recovery.applied` with `path: 'markdown-fence'`, and accepts the parsed value WITHOUT counting against the retry budget per RFC 0033 §D. Run terminates `completed`.",
5
+ "description": "Drives `core.ai.structuredOutput` against the conformance-only `mock` provider. The conformance scenario POSTs a 1-entry program to `/v1/host/sample/test/mock-ai/program` (keyed by `nodeId: 'recovery-applied-structured-call'`) BEFORE starting the run: the mock returns a markdown-fenced JSON envelope (e.g., ```json\\n{\"result\":\"ok\"}\\n```). The host's `dispatchStructured` lenient-parse fallback strips the fence via `tryLenientParse(text)`, emits exactly one `envelope.recovery.applied` with `path: 'markdown-fence'`, and accepts the parsed value WITHOUT counting against the retry budget per RFC 0033 §D. Run terminates `completed`.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "recovery-applied-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider (markdown-fenced)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -5,7 +5,7 @@
5
5
  "description": "Single `core.ai.structuredOutput` node against the conformance `mock` provider with a pre-seeded program returning `stopReason: 'safety'` + `refusalText: '...'` on attempt 1. Host's `dispatchStructured()` MUST: (a) emit exactly one `envelope.refusal` event with the canonical payload shape; (b) NOT retry (RFC 0032 §B.3 + RFC 0033 §D — refusal is terminal); (c) fail the node with `error.code: 'envelope_refused_by_provider'` per RFC 0033 §F; (d) NOT echo the refusal text in `RunSnapshot.error.message` (SECURITY invariant `envelope-refusal-no-prompt-leak` — refusal text lives only on the event-log entry, scrubbed via the existing SR-1 redaction harness).",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "refusal-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider (refusal)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -2,10 +2,10 @@
2
2
  "id": "conformance-envelope-retry-attempted",
3
3
  "name": "Conformance: envelope.retry.attempted (RFC 0032 §B.1)",
4
4
  "version": "1.0",
5
- "description": "Drives `core.ai.structuredOutput` against the conformance-only `mock` provider. The conformance scenario POSTs a 2-entry program to `/v1/host/sample/test/mock-ai/program` (keyed by `nodeId: 'structured-call'`) BEFORE starting the run: attempt 1 returns invalid JSON, attempt 2 returns a valid envelope. The host's `dispatchStructured` retry loop MUST emit exactly one `envelope.retry.attempted` event with `attempt: 2` between the two provider calls (RFC 0032 §B.1). The run terminates `completed` after the second attempt succeeds.",
5
+ "description": "Drives `core.ai.structuredOutput` against the conformance-only `mock` provider. The conformance scenario POSTs a 2-entry program to `/v1/host/sample/test/mock-ai/program` (keyed by `nodeId: 'retry-attempted-structured-call'`) BEFORE starting the run: attempt 1 returns invalid JSON, attempt 2 returns a valid envelope. The host's `dispatchStructured` retry loop MUST emit exactly one `envelope.retry.attempted` event with `attempt: 2` between the two provider calls (RFC 0032 §B.1). The run terminates `completed` after the second attempt succeeds.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "retry-attempted-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider",
11
11
  "position": { "x": 0, "y": 0 },
@@ -5,7 +5,7 @@
5
5
  "description": "Drives `core.ai.structuredOutput` against the conformance `mock` provider with a program that returns invalid JSON on EVERY attempt. The host's `dispatchStructured` retry loop MUST exhaust its budget and emit exactly one `envelope.retry.exhausted` event with `finalReason: 'schema-violation'` (or `'parse-error'` per RFC 0032 §B.1 reason enum). The node MUST fail with `error.code: 'envelope_payload_invalid'` (existing RFC 0021 code per RFC 0033 §C). Pairs with `envelope-retry-exhausted.test.ts`.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "retry-exhausted-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider (always invalid)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -5,7 +5,7 @@
5
5
  "description": "Drives `core.ai.structuredOutput` against the conformance `mock` provider with a program: attempt 1 returns `stopReason: 'max_tokens'` (truncation); attempt 2 returns a valid envelope. The host's `dispatchStructured` retry loop MUST: (a) emit exactly one `envelope.truncated` event with `stopReason: 'max_tokens'`; (b) retry with an INCREASED output budget per RFC 0033 §B (the host's `truncationBudgetMultiplier` — default 2×); (c) NOT inject the corrective schema fragment on the truncation retry (truncation is an output-size problem, not a schema problem). Eventually completes after attempt 2 succeeds. Pairs with `envelope-truncated.test.ts`.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "truncated-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output (truncation then success)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -5,7 +5,7 @@
5
5
  "description": "Drives `core.ai.structuredOutput` against the conformance `mock` provider with a program that returns `stopReason: 'max_tokens'` on EVERY attempt. The host's `dispatchStructured` retry loop MUST: (a) emit `envelope.truncated` on each attempt (or at least the first one — RFC 0032 §B.4 is per-attempt); (b) double the budget each retry (RFC 0033 §B); (c) exhaust retries after `maxRetryAttempts` (default 3); (d) emit exactly one `envelope.retry.exhausted` with `finalReason: 'truncation'`; (e) emit `cap.breached` with `kind: 'schema'`; (f) fail the node with `error.code: 'envelope_truncation_unrecoverable'` per RFC 0033 §F. The run does NOT exceed `maxRetryAttempts` total LLM calls — DoS-bound assertion. Pairs with `envelope-truncation-cap-exhaustion.test.ts`.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "truncation-cap-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output (perpetual truncation → cap exhaustion)",
11
11
  "position": { "x": 0, "y": 0 },
@@ -16,7 +16,7 @@
16
16
  "inputs": {}
17
17
  },
18
18
  {
19
- "id": "structured-call",
19
+ "id": "nondet-structured-call",
20
20
  "typeId": "core.ai.structuredOutput",
21
21
  "name": "Structured output via mock provider (consumes nondet-tool result)",
22
22
  "position": { "x": 200, "y": 0 },
@@ -40,7 +40,7 @@
40
40
  }
41
41
  ],
42
42
  "edges": [
43
- { "id": "e1", "sourceNodeId": "nondet-tool", "targetNodeId": "structured-call" }
43
+ { "id": "e1", "sourceNodeId": "nondet-tool", "targetNodeId": "nondet-structured-call" }
44
44
  ],
45
45
  "triggers": [
46
46
  { "id": "manual", "type": "manual", "enabled": true }
@@ -2,10 +2,10 @@
2
2
  "id": "conformance-phase4-replay-divergence",
3
3
  "name": "Conformance: RFC 0041 §B replay-divergence-at-refusal (Phase 4)",
4
4
  "version": "1.0",
5
- "description": "Single `core.ai.structuredOutput` node against the conformance `mock` provider. Conformance scenario `replay-divergence-at-refusal.test.ts` pre-seeds the mock with a two-entry program via `POST /v1/host/sample/test/mock-ai/program` keyed on the structured-call nodeId: entry [0] returns a valid envelope (consumed by the original run); entry [1] returns `stopReason: 'safety'` + `refusalText` (consumed by the `:fork mode: replay`). Phase 4 hosts advertising `multiAgent.executionModel.replayDeterminism.refusalDivergenceEmission: true` MUST detect the divergence at replay time, emit a `replay.divergedAtRefusal` event with `originalEnvelopeKind: 'valid'` + `replayEnvelopeKind: 'refusal'`, and fail the replay with HTTP `422` + `error.code: 'replay_diverged_at_refusal'` per `spec/v1/rest-endpoints.md §\"Common error codes\"`. Silent substitution of the refusal for the original envelope is non-conformant.",
5
+ "description": "Single `core.ai.structuredOutput` node against the conformance `mock` provider. Conformance scenario `replay-divergence-at-refusal.test.ts` pre-seeds the mock with a two-entry program via `POST /v1/host/sample/test/mock-ai/program` keyed on the divergence-structured-call nodeId: entry [0] returns a valid envelope (consumed by the original run); entry [1] returns `stopReason: 'safety'` + `refusalText` (consumed by the `:fork mode: replay`). Phase 4 hosts advertising `multiAgent.executionModel.replayDeterminism.refusalDivergenceEmission: true` MUST detect the divergence at replay time, emit a `replay.divergedAtRefusal` event with `originalEnvelopeKind: 'valid'` + `replayEnvelopeKind: 'refusal'`, and fail the replay with HTTP `422` + `error.code: 'replay_diverged_at_refusal'` per `spec/v1/rest-endpoints.md §\"Common error codes\"`. Silent substitution of the refusal for the original envelope is non-conformant.",
6
6
  "nodes": [
7
7
  {
8
- "id": "structured-call",
8
+ "id": "divergence-structured-call",
9
9
  "typeId": "core.ai.structuredOutput",
10
10
  "name": "Structured output via mock provider (Phase 4 replay-divergence probe)",
11
11
  "position": { "x": 0, "y": 0 },
package/fixtures.md CHANGED
@@ -47,6 +47,7 @@ All fixtures MUST advertise:
47
47
  | Delay | `conformance-delay` | Verifies poll/SSE behavior over time | `completed` | ≤ 30s (input-controlled) |
48
48
  | Failure | `conformance-failure` | Verifies error-event surface | `failed` | ≤ 5s |
49
49
  | Approval | `conformance-approval` | Verifies HITL approval interrupt + resume | `completed` after resolve | unbounded (suspends) |
50
+ | Approval (refine) | `conformance-approval-refine` | RFC 0183 — refine resolution carries action + refineFeedback |
50
51
  | Clarification | `conformance-clarification` | Verifies HITL clarification interrupt + resume | `completed` after resolve | unbounded (suspends) |
51
52
  | Multi-node | `conformance-multi-node` | Verifies edge ordering + per-node events | `completed` | ≤ 10s |
52
53
  | Idempotent | `conformance-idempotent` | Verifies `Idempotency-Key` cache | `completed` | ≤ 5s |
@@ -177,6 +178,20 @@ The `messages`-mode stream fixture (AI token streaming) is covered by the determ
177
178
  - **Terminal status (after accept)**: `completed`.
178
179
  - **Resolve schema**: `{action: "accept" | "reject"}`. Server MUST reject any other shape with 400.
179
180
 
181
+ ### `conformance-approval-refine`
182
+
183
+ - **Purpose**: verify a `refine` resolution round-trips its action and structured feedback (RFC 0183 §A.1/§A.2).
184
+ - **Inputs**: none.
185
+ - **Behavior**:
186
+ 1. Run starts and reaches an `approvalGate` offering `accept | reject | refine`.
187
+ 2. Run status MUST be `waiting-approval`.
188
+ 3. Client POSTs `{action: 'refine', refineFeedback: {scope: 'whole', ...}}` to the interrupt.
189
+ 4. The resolved event payload MUST carry `action: 'refine'` and the `refineFeedback` supplied.
190
+ 5. A `refine` resolution supplying no `refineFeedback` MUST be refused.
191
+ - **Why it is separate from `conformance-approval`**: adding `refine` to that fixture's `actions` would change a
192
+ registered workflow definition every host already serves, forcing a re-registration for a test-only widening.
193
+ A new id costs one catalog row and no host churn.
194
+
180
195
  ### `conformance-clarification`
181
196
 
182
197
  - **Purpose**: verify HITL clarification interrupt + resume.
@@ -487,6 +502,7 @@ conformance/
487
502
  conformance-delay.json
488
503
  conformance-failure.json
489
504
  conformance-approval.json
505
+ conformance-approval-refine.json
490
506
  conformance-clarification.json
491
507
  conformance-multi-node.json
492
508
  conformance-idempotent.json
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@openwop/openwop-conformance",
3
- "version": "2.1.6",
3
+ "version": "2.2.0",
4
4
  "description": "Production-ready black-box conformance suite for OpenWOP v1.0 compliant servers.",
5
5
  "repository": {
6
6
  "type": "git",
@@ -56,6 +56,6 @@
56
56
  "@openwop/spec-artifacts": "file:../spec-artifacts"
57
57
  },
58
58
  "peerDependencies": {
59
- "@openwop/spec-artifacts": "2.1.6"
59
+ "@openwop/spec-artifacts": "2.2.0"
60
60
  }
61
61
  }
package/requirements.json CHANGED
@@ -2,11 +2,11 @@
2
2
  "$comment": "GENERATED by conformance/scripts/generate-requirement-registry.mjs — do not edit. One record per it()/test() in src/scenarios. Ids: openwop.it.<file-stem>.<title-slug>[~n] (src/lib/requirement-ids.ts). A record with id null has an interpolated title; its run-time row is keyed by the rendered title and maps here by file+line only. Renamed ids need a row in requirement-aliases.json.",
3
3
  "generatedFrom": "src/scenarios/*.test.ts",
4
4
  "counts": {
5
- "files": 566,
6
- "tests": 2156,
7
- "withStableId": 2156,
5
+ "files": 567,
6
+ "tests": 2159,
7
+ "withStableId": 2159,
8
8
  "interpolatedTitles": 0,
9
- "explicitIds": 2111
9
+ "explicitIds": 2114
10
10
  },
11
11
  "records": [
12
12
  {
@@ -11218,7 +11218,7 @@
11218
11218
  {
11219
11219
  "id": "openwop.it.envelope-completion-distinguishes-truncation.capabilities-envelopes-reliability-completion-when-present-conforms-to-rfc-0033",
11220
11220
  "file": "envelope-completion-distinguishes-truncation.test.ts",
11221
- "line": 94,
11221
+ "line": 103,
11222
11222
  "title": "capabilities.envelopes.reliability.completion (when present) conforms to RFC 0033 §E",
11223
11223
  "explicitId": "openwop.it.envelope-completion-distinguishes-truncation.capabilities-envelopes-reliability-completion-when-present-conforms-to-rfc-0033",
11224
11224
  "citations": [
@@ -11235,7 +11235,7 @@
11235
11235
  {
11236
11236
  "id": "openwop.it.envelope-completion-distinguishes-truncation.truncation-emits-envelope-truncated-envelope-retry-attempted-with-reason-truncat",
11237
11237
  "file": "envelope-completion-distinguishes-truncation.test.ts",
11238
- "line": 117,
11238
+ "line": 126,
11239
11239
  "title": "truncation: emits envelope.truncated + envelope.retry.attempted with reason: \"truncation\"",
11240
11240
  "explicitId": "openwop.it.envelope-completion-distinguishes-truncation.truncation-emits-envelope-truncated-envelope-retry-attempted-with-reason-truncat",
11241
11241
  "citations": [
@@ -11256,7 +11256,7 @@
11256
11256
  {
11257
11257
  "id": "openwop.it.envelope-completion-distinguishes-truncation.truncation-retry-budget-strictly-greater-than-initial-rfc-0033-b-truncationbudge",
11258
11258
  "file": "envelope-completion-distinguishes-truncation.test.ts",
11259
- "line": 142,
11259
+ "line": 151,
11260
11260
  "title": "truncation: retry budget strictly greater than initial (RFC 0033 §B truncationBudgetMultiplier)",
11261
11261
  "explicitId": "openwop.it.envelope-completion-distinguishes-truncation.truncation-retry-budget-strictly-greater-than-initial-rfc-0033-b-truncationbudge",
11262
11262
  "citations": [
@@ -11269,7 +11269,7 @@
11269
11269
  {
11270
11270
  "id": "openwop.it.envelope-completion-distinguishes-truncation.schema-violation-no-envelope-truncated-envelope-retry-attempted-reason-schema-vi",
11271
11271
  "file": "envelope-completion-distinguishes-truncation.test.ts",
11272
- "line": 166,
11272
+ "line": 175,
11273
11273
  "title": "schema-violation: NO envelope.truncated; envelope.retry.attempted reason ∈ {schema-violation, parse-error}",
11274
11274
  "explicitId": "openwop.it.envelope-completion-distinguishes-truncation.schema-violation-no-envelope-truncated-envelope-retry-attempted-reason-schema-vi",
11275
11275
  "citations": [
@@ -11286,7 +11286,7 @@
11286
11286
  {
11287
11287
  "id": "openwop.it.envelope-completion-distinguishes-truncation.schema-violation-retry-budget-unchanged-from-initial-no-budget-multiplication-on",
11288
11288
  "file": "envelope-completion-distinguishes-truncation.test.ts",
11289
- "line": 196,
11289
+ "line": 205,
11290
11290
  "title": "schema-violation: retry budget UNCHANGED from initial (no budget multiplication on this path)",
11291
11291
  "explicitId": "openwop.it.envelope-completion-distinguishes-truncation.schema-violation-retry-budget-unchanged-from-initial-no-budget-multiplication-on",
11292
11292
  "citations": [
@@ -14544,7 +14544,7 @@
14544
14544
  {
14545
14545
  "id": "openwop.it.interrupt-approval.run-suspends-at-gate-accept-resolution-drives-terminal-completed",
14546
14546
  "file": "interrupt-approval.test.ts",
14547
- "line": 21,
14547
+ "line": 24,
14548
14548
  "title": "run suspends at gate, accept resolution drives terminal completed",
14549
14549
  "explicitId": "openwop.it.interrupt-approval.run-suspends-at-gate-accept-resolution-drives-terminal-completed",
14550
14550
  "citations": [
@@ -14565,7 +14565,7 @@
14565
14565
  {
14566
14566
  "id": "openwop.it.interrupt-approval.400-or-422-when-action-is-not-in-accept-reject",
14567
14567
  "file": "interrupt-approval.test.ts",
14568
- "line": 50,
14568
+ "line": 53,
14569
14569
  "title": "400 (or 422) when action is not in {accept, reject}",
14570
14570
  "explicitId": "openwop.it.interrupt-approval.400-or-422-when-action-is-not-in-accept-reject",
14571
14571
  "citations": [
@@ -14578,7 +14578,7 @@
14578
14578
  {
14579
14579
  "id": "openwop.it.interrupt-approval.400-404-when-nodeid-does-not-match-an-active-interrupt",
14580
14580
  "file": "interrupt-approval.test.ts",
14581
- "line": 78,
14581
+ "line": 81,
14582
14582
  "title": "400/404 when nodeId does not match an active interrupt",
14583
14583
  "explicitId": "openwop.it.interrupt-approval.400-404-when-nodeid-does-not-match-an-active-interrupt",
14584
14584
  "citations": [
@@ -14588,6 +14588,44 @@
14588
14588
  }
14589
14589
  ]
14590
14590
  },
14591
+ {
14592
+ "id": "openwop.it.interrupt-approval.a-refine-resolve-round-trips-action-and-the-structured-feedback-it-requires",
14593
+ "file": "interrupt-approval.test.ts",
14594
+ "line": 117,
14595
+ "title": "a refine resolve round-trips `action` and the structured feedback it requires",
14596
+ "explicitId": "openwop.it.interrupt-approval.a-refine-resolve-round-trips-action-and-the-structured-feedback-it-requires",
14597
+ "citations": [
14598
+ {
14599
+ "section": "RFCS/0183-interrupt-resolved-action-fidelity.md §A.1",
14600
+ "requirement": "a gate offering `refine` MUST accept a refine resolution carrying refineFeedback"
14601
+ },
14602
+ {
14603
+ "section": "RFCS/0183-interrupt-resolved-action-fidelity.md §A.1",
14604
+ "requirement": "the resolution MUST be recorded as an event"
14605
+ },
14606
+ {
14607
+ "section": "RFCS/0183-interrupt-resolved-action-fidelity.md §A.1",
14608
+ "requirement": "the resolved payload MUST carry `action: \"refine\"` — the seat RFC 0183 adds"
14609
+ },
14610
+ {
14611
+ "section": "RFCS/0183-interrupt-resolved-action-fidelity.md §A.2",
14612
+ "requirement": "a refine resolution MUST carry the refineFeedback it was given — an action whose meaning is incomplete without it"
14613
+ }
14614
+ ]
14615
+ },
14616
+ {
14617
+ "id": "openwop.it.interrupt-approval.refuses-a-refine-resolution-that-supplies-no-refinefeedback",
14618
+ "file": "interrupt-approval.test.ts",
14619
+ "line": 167,
14620
+ "title": "refuses a refine resolution that supplies no refineFeedback",
14621
+ "explicitId": "openwop.it.interrupt-approval.refuses-a-refine-resolution-that-supplies-no-refinefeedback",
14622
+ "citations": [
14623
+ {
14624
+ "section": "RFCS/0183-interrupt-resolved-action-fidelity.md §A.2",
14625
+ "requirement": "refine without refineFeedback MUST be refused — the feedback is what the action means"
14626
+ }
14627
+ ]
14628
+ },
14591
14629
  {
14592
14630
  "id": "openwop.it.interrupt-approver-routing.capabilities-schema-declares-interrupt-approverrouting",
14593
14631
  "file": "interrupt-approver-routing.test.ts",
@@ -25930,6 +25968,49 @@
25930
25968
  }
25931
25969
  ]
25932
25970
  },
25971
+ {
25972
+ "id": "openwop.it.v2-bound-id-path-projection.a-tenant-bound-id-is-readable-at-its-escaped-path-segment-links-carry-that-form",
25973
+ "file": "v2-bound-id-path-projection.test.ts",
25974
+ "line": 51,
25975
+ "title": "a tenant-bound id is readable at its ~-escaped path segment, links carry that form, and a malformed escape is refused",
25976
+ "explicitId": "openwop.requirement.0184.bound-id-path-projection",
25977
+ "citations": [
25978
+ {
25979
+ "section": "spec/v2/core/identity.md §5",
25980
+ "requirement": null,
25981
+ "interpolated": true
25982
+ },
25983
+ {
25984
+ "section": "spec/v2/core/identity.md §5",
25985
+ "requirement": null,
25986
+ "interpolated": true
25987
+ },
25988
+ {
25989
+ "section": "spec/v2/core/identity.md §5",
25990
+ "requirement": "the run read at the projected segment MUST be the run that was created — a host that decodes the escape to a DIFFERENT id has a non-injective decoder"
25991
+ },
25992
+ {
25993
+ "section": "spec/v2/core/identity.md §5",
25994
+ "requirement": null,
25995
+ "interpolated": true
25996
+ },
25997
+ {
25998
+ "section": "spec/v2/core/identity.md §5",
25999
+ "requirement": null,
26000
+ "interpolated": true
26001
+ },
26002
+ {
26003
+ "section": "spec/v2/core/identity.md §5",
26004
+ "requirement": null,
26005
+ "interpolated": true
26006
+ },
26007
+ {
26008
+ "section": "spec/v2/core/identity.md §5",
26009
+ "requirement": null,
26010
+ "interpolated": true
26011
+ }
26012
+ ]
26013
+ },
25933
26014
  {
25934
26015
  "id": "openwop.it.v2-bundle-v3-signed.a-signed-v3-bundle-validates-against-the-closed-root-schema-an-extra-root-key-is",
25935
26016
  "file": "v2-bundle-v3-signed.test.ts",
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "$comment": "GENERATED by conformance/scripts/generate-scenario-majors.mjs (RFC 0168 §D.3). Do not edit; add a file to BOTH_MAJORS in the generator to target both majors.",
3
3
  "counts": {
4
- "files": 517,
4
+ "files": 518,
5
5
  "v1": 445,
6
- "v2": 74
6
+ "v2": 75
7
7
  },
8
8
  "majors": {
9
9
  "a2a-1-0-agent-card.test.ts": [
@@ -1244,6 +1244,9 @@
1244
1244
  "v2-assurance-downgrade-audited.test.ts": [
1245
1245
  2
1246
1246
  ],
1247
+ "v2-bound-id-path-projection.test.ts": [
1248
+ 2
1249
+ ],
1247
1250
  "v2-bundle-v3-signed.test.ts": [
1248
1251
  2
1249
1252
  ],