@openwop/openwop-conformance 2.43.1 → 2.43.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -0
- package/README.md +3 -3
- package/dist/lib/scenario-disposition.js +30 -0
- package/dist/spec-artifacts.lock.json +2 -2
- package/fixtures/conformance-replay-ordinal-loop.json +31 -0
- package/fixtures.md +12 -0
- package/package.json +2 -2
- package/requirements.json +54 -4
- package/scenario-majors.json +8 -2
- package/schemas/CORPUS-STAMP.json +18 -18
- package/src/lib/scenario-disposition.ts +40 -0
- package/src/scenarios/v2-replay-suppression-ordinal.test.ts +108 -0
- package/src/scenarios/v2-run-fork-ancestry.test.ts +83 -0
- package/src/setup.ts +25 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,13 @@
|
|
|
1
1
|
# `@openwop/openwop-conformance` Changelog
|
|
2
2
|
|
|
3
|
+
## [2.43.2] — 2026-09-28 — three Class 3 corrections, and a failing scenario hook no longer hides its test
|
|
4
|
+
|
|
5
|
+
- **The 2.43.2 cycle opens.** #1724 changes the packed `@openwop/spec-artifacts` tree (`spec/v2/core/interrupt.md`, `spec/v1/gaps.json`) after `v2.43.1` was tagged. No scenario changes in this entry: `approval.rejected` goes from SHOULD to MAY (Class 3), and no leg asserted the SHOULD.
|
|
6
|
+
- **The replay outcome key uses the execution ordinal** (#1718, Class 3). New `v2-replay-suppression-ordinal.test.ts` (target major 2), row `openwop.requirement.replay.suppression-execution-ordinal`, fixture `conformance-replay-ordinal-loop` (start → effect → wait → effect). The suite cancels the source after `effect`'s first execution, then forks `mode: "replay"`. `effect`'s second execution has no recorded outcome for n = 2 and MUST fail closed with `replay_source_missing`; a host keyed on `(nodeId, attempt)`, or on the latest outcome, completes it instead. `inapplicable` unless the host runs cycles and advertises the fixture. No host witnesses it yet.
|
|
7
|
+
- **Suite `2.43.2`**. `@openwop/spec-artifacts` moves in lockstep at the same exact pin.
|
|
8
|
+
- **A scenario hook that fails no longer hides its test** (#1753). vitest runs a file's `afterEach` hooks in stack order, so a scenario's own cleanup runs before the runner's recorder. If it threw or timed out, the recorder never saw the test and the file row read "no test executed and no disposition recorded" (`blocked`). `setup.ts`'s `afterAll` now recovers such tests from vitest's task results (`unrecordedTests` in `lib/scenario-disposition.ts`, self-test `lib/unrecorded-tests.test.ts`), so the row is `executed-fail` and names the error.
|
|
9
|
+
- **A fork has no ancestry parent** (Class 3). New `v2-run-fork-ancestry.test.ts` (target major 2), row `openwop.requirement.0040.fork-has-no-ancestry-parent`. It is gated on `replay` and on `multiAgent.executionModel.crossHostCausation.ancestryEndpointSupported`, and records `inapplicable` without them. It creates a `conformance-noop` run and checks its ancestry `parent` is `null` as a control. It then makes a branch fork at `fromSeq: 0`, checks the `201` names the source as `sourceRunId`, and asserts the fork's ancestry `parent` is `null` (`runs.md` §Diff and ancestry). It is a file of its own so that most v2 hosts, which do not advertise ancestry, do not get a `partial-witness` row on `v2-run-fork-refusals`. Measured on the v2 reference host (openwop-examples #124) with a temporary ancestry advertisement: `executed-pass`. With its previous `cause: "core.subWorkflow"` line restored the row fails, and a host reporting a parent for every run fails the control. As shipped, with nothing advertised, it records `inapplicable`. openwop-app also derives ancestry from the `parentRunId` its forks set, so it fails this row wherever it advertises ancestry.
|
|
10
|
+
|
|
3
11
|
## [2.43.1] — 2026-09-28 — major-1 certification works again, RFC 0223's routed and timed-out rejects are witnessed, and four Class 3 corrections
|
|
4
12
|
|
|
5
13
|
- **`v2-manifest-hatch-carried`'s agent fixture is schema-valid.** Its `agents[0]` had neither `systemPrompt` nor `systemPromptRef`, and `agent-manifest` (v1 and v2) requires exactly one. A conforming host therefore had to refuse the pack for a reason unrelated to the `x-` hatch, and `0177.manifest-hatch-carried.agents-x-field` could not pass on any host. The fixture now carries `systemPrompt`. Validated against `schemas/v2/agent-manifest.schema.json`: invalid before, valid after. Reported by MyndHyve.
|
package/README.md
CHANGED
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
# --legacy-peer-deps is REQUIRED, not optional: the exact peer pin is what npm's
|
|
12
12
|
# default resolver refuses. npm 10.9 fails outright with
|
|
13
13
|
# "Cannot read properties of null (reading 'edgesOut')" — use npm >= 11.
|
|
14
|
-
npm install --legacy-peer-deps @openwop/openwop-conformance@2.43.
|
|
14
|
+
npm install --legacy-peer-deps @openwop/openwop-conformance@2.43.2 @openwop/spec-artifacts@2.43.2
|
|
15
15
|
# or run without install:
|
|
16
16
|
npx @openwop/openwop-conformance --base-url https://api.example.com --api-key hk_test_...
|
|
17
17
|
```
|
|
@@ -135,7 +135,7 @@ Exit code is non-zero on any failed assertion. `--certify` distinguishes: `0`
|
|
|
135
135
|
|
|
136
136
|
## What's Covered
|
|
137
137
|
|
|
138
|
-
The current suite has
|
|
138
|
+
The current suite has 573 scenario files under `src/scenarios/`.
|
|
139
139
|
- 2026-09-23 (suite 2.37.0 cycle, RFC 0213): NEW `v2-sse-last-event-id-cursor.test.ts` (a `Last-Event-ID` past the log is an exclusive cursor; a malformed id, when refused, is `400 validation_error`; the cursor never changes the answer for an unknown or foreign-tenant run — public test of `event-cursor-after-authorization`), `v2-idempotency-in-flight.test.ts` (five concurrent same-key creates yield one run; each loser is a marked replay or `409 idempotency_in_flight` with no retry timing in `details`; `partial-witness` when no loser was refused in flight) and `v2-interrupt-resolve-terminal.test.ts` (a run-scoped resolve after cancel or completion is `409 interrupt_already_resolved`, never `interrupt_cancelled`). All three sit off the core-standard floor until measured on the three bundle hosts.
|
|
140
140
|
- 2026-09-03 (suite `1.157.0 -> 1.158.0`, gap G17): NEW `idempotency-concurrent-claim.test.ts` — drives the new `host-sample-test-seams.md` §25 concurrent duplicate-delivery seam for the RFC 0150 §B / `idempotency.md` §"Concurrent duplicates (Layer 2)" atomic-claim MUST, which is unconditional and had no witness of any kind. Asserts every executor mints the SAME `logicalInvocationId` **before** asserting `delivered === 1` — without the identity check a host passes by minting different ids and never colliding, one effect because nothing raced. Not profile-gated and so not opt-out-able (the obligation is unconditional); an unmounted seam records `blocked`, which is not certifiable. Graduates `layer2-invocation-claim-atomic` reference-impl -> protocol.
|
|
141
141
|
- 2026-08-19 (suite `1.137.0 → 1.138.0`): NEW `durability-poison-exhaustion.test.ts` — RFC 0158 §C.8, the FIRST row of that RFC's conformance table to land. Asserts what `failure-path.test.ts` cannot: not just that deterministically failing work reaches terminal, but that attempts STOP — counted on the log, re-counted after a scaled quiet window, asserted unchanged. A host still redelivering records more. Seam-gated on the existing event-log seam (`blocked` = unobservable, not unmet) and outside every profile floor.
|
|
@@ -481,7 +481,7 @@ Server-required (added in 1.7.0):
|
|
|
481
481
|
| ------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
482
482
|
| **Redaction** | [`capabilities.md`](../spec/v1/capabilities.md) §"Secrets" + NFR-7 + §"aiProviders" | Vendor-neutral assertions that the server doesn't leak secret material. Three scenario groups: (a) discovery shape contract — `secrets` + `aiProviders` advertisements are well-formed regardless of `secrets.supported`; when `supported === true`, scopes MUST be non-empty + `resolution === 'host-managed'`; `byok ⊆ supported`. (b) bearer-token redaction — invalid Bearer canary in `Authorization` header is not echoed in the 401 response body. (c) credentialRef echo control — gated on `secrets.supported === true`; canary planted in `configurable.ai.credentialRef` MUST NOT appear in any RunEvent payload (poll-based capture; transport-agnostic). Uses runtime-built canary fixtures (`lib/canaries.ts`) that defeat static secret scanners. 6 scenarios. |
|
|
483
483
|
|
|
484
|
-
Current source tree:
|
|
484
|
+
Current source tree: 573 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
|
|
485
485
|
|
|
486
486
|
## Remaining Gaps
|
|
487
487
|
|
|
@@ -142,6 +142,36 @@ testName) {
|
|
|
142
142
|
return { disposition: 'blocked', detail: 'unclassified return: the test passed with zero assertions and recorded no reason — RFC 0148 §A resolves it to blocked, never to a pass' };
|
|
143
143
|
return { disposition: 'skipped', detail: 'vitest skipped the test (ctx.skip / it.skip) without a recorded gate reason' };
|
|
144
144
|
}
|
|
145
|
+
/**
|
|
146
|
+
* The tests the per-test recorder never saw, as file states and failures.
|
|
147
|
+
*
|
|
148
|
+
* vitest runs a file's `afterEach` hooks in stack order, so a scenario's own
|
|
149
|
+
* `afterEach` (unregistering webhooks, closing a receiver) runs BEFORE the
|
|
150
|
+
* recorder `setup.ts` registers. If that hook throws or times out, vitest
|
|
151
|
+
* fails the test and skips the remaining hooks, so the recorder never runs.
|
|
152
|
+
* The file then has no recorded state, and its row read "no test executed and
|
|
153
|
+
* no disposition recorded" (`blocked`). That row gives no reason and hides a
|
|
154
|
+
* real failure. One tier-2 cut of `v2-webhook-message-id-stable` produced it
|
|
155
|
+
* on 2026-09-28.
|
|
156
|
+
*
|
|
157
|
+
* Only tests that finished `pass` or `fail` are recovered. A skipped test
|
|
158
|
+
* never reaches the recorder either way, and the existing rules already cover it.
|
|
159
|
+
*/
|
|
160
|
+
export function unrecordedTests(tests, recordedIds) {
|
|
161
|
+
const states = [];
|
|
162
|
+
const failures = [];
|
|
163
|
+
for (const t of tests) {
|
|
164
|
+
if (recordedIds.has(t.id))
|
|
165
|
+
continue;
|
|
166
|
+
if (t.state === 'pass')
|
|
167
|
+
states.push('pass');
|
|
168
|
+
else if (t.state === 'fail') {
|
|
169
|
+
states.push('fail');
|
|
170
|
+
failures.push({ name: t.name, message: `the test or its afterEach hook failed before the runner recorded it: ${t.message ?? 'no message'}` });
|
|
171
|
+
}
|
|
172
|
+
}
|
|
173
|
+
return { states, failures };
|
|
174
|
+
}
|
|
145
175
|
/**
|
|
146
176
|
* The `executed-fail` detail for a file row: WHICH cases failed, and what the
|
|
147
177
|
* first one said.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-replay-ordinal-loop",
|
|
3
|
+
"name": "Conformance: Replay outcome keyed on the execution ordinal",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "openwop#1718. start → effect (a side-effecting node) → wait (core.delay) → back to effect, bounded by settings.maxLoopbackIterations 2, so effect executes twice. The suite cancels the source run inside wait, after effect's first execution completed and before its second, then forks mode:\"replay\" at effect's first node.started. In the fork, effect's first execution resolves from the recorded outcome for n = 1, and its second execution has no recorded outcome for n = 2, so it MUST fail closed with replay_source_missing (replay.md §Suppression rule 2). A host keying on (nodeId, attempt) or selecting the latest outcome resolves the second execution from the first and completes it. core.conformance.side-effect is the reserved side-effecting typeId (see conformance-replay-side-effect). Advertised only by a host that runs cycles and suppresses side effects in replay.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{ "id": "start", "typeId": "core.noop", "name": "Start", "position": { "x": 0, "y": 0 }, "config": {}, "inputs": {} },
|
|
8
|
+
{ "id": "effect", "typeId": "core.conformance.side-effect", "name": "Side Effect", "position": { "x": 200, "y": 0 }, "config": {}, "inputs": {} },
|
|
9
|
+
{
|
|
10
|
+
"id": "wait",
|
|
11
|
+
"typeId": "core.delay",
|
|
12
|
+
"name": "Wait",
|
|
13
|
+
"position": { "x": 400, "y": 0 },
|
|
14
|
+
"config": {},
|
|
15
|
+
"inputs": { "delayMs": { "type": "variable", "variableName": "delayMs" } }
|
|
16
|
+
}
|
|
17
|
+
],
|
|
18
|
+
"edges": [
|
|
19
|
+
{ "id": "start-effect", "sourceNodeId": "start", "targetNodeId": "effect", "triggerRule": "any_success" },
|
|
20
|
+
{ "id": "effect-wait", "sourceNodeId": "effect", "targetNodeId": "wait" },
|
|
21
|
+
{ "id": "wait-effect", "sourceNodeId": "wait", "targetNodeId": "effect", "triggerRule": "any_success" }
|
|
22
|
+
],
|
|
23
|
+
"triggers": [
|
|
24
|
+
{ "id": "manual", "type": "manual", "enabled": true }
|
|
25
|
+
],
|
|
26
|
+
"variables": [
|
|
27
|
+
{ "name": "delayMs", "type": "number", "defaultValue": 3000 }
|
|
28
|
+
],
|
|
29
|
+
"metadata": { "tags": ["conformance", "replay"] },
|
|
30
|
+
"settings": { "timeout": 0, "maxLoopbackIterations": 2 }
|
|
31
|
+
}
|
package/fixtures.md
CHANGED
|
@@ -61,6 +61,7 @@ All fixtures MUST advertise:
|
|
|
61
61
|
| Idempotent | `conformance-idempotent` | Verifies `Idempotency-Key` cache | `completed` | ≤ 5s |
|
|
62
62
|
| Cancellable | `conformance-cancellable` | Verifies `:cancel` endpoint mid-run | `cancelled` after cancel | ≤ 60s (input-controlled) |
|
|
63
63
|
| Replay Side-Effect Suppression | `conformance-replay-side-effect` | RFC 0140. Verifies a `mode:"replay"` fork does NOT re-fire a side-effecting node | fork `failed` (`error.code='replay_source_missing'`) | ≤ 60s (input-controlled) |
|
|
64
|
+
| Replay ordinal loop | `conformance-replay-ordinal-loop` | openwop#1718 — a side-effecting node executed twice in a loop; a replay fork keys each execution by its ordinal `n` | `cancelled` (source); fork: execution 2 fails `replay_source_missing` | ~3 s + fork |
|
|
64
65
|
| Capability Missing | `conformance-capability-missing` | Verifies dispatch refusal on unsatisfied `requires` | `failed` (`error.code='capability_not_provided'`) | ≤ 5s |
|
|
65
66
|
| Prompt End-to-End | `conformance-prompt-end-to-end` | RFC 0027 + RFC 0029 end-to-end. Single `mock-ai` node with `config.systemPromptRef` set; host MUST emit `agent.promptResolved` + `prompt.composed` events during dispatch, then complete. Capability-gated on `capabilities.prompts.supported`. | `completed` | ≤ 10s |
|
|
66
67
|
| Prompt All Four Kinds | `conformance-prompt-all-four-kinds` | RFC 0027 §A four-kind dispatch coverage with a MULTI-ENTRY `fewShotPromptRefs[]` array. Single `mock-ai` node with one ref per singular-kind slot (`systemPromptRef`, `userPromptRef`, `schemaHintPromptRef`) + two distinct templateIds in `fewShotPromptRefs[]`; host MUST emit 5 `agent.promptResolved` events (one per slot) AND 5 `prompt.composed` events. Multi-entry few-shot is the regression pin for `fewShotPromptRefs[slotIndex]` per-index resolution — a host that hard-codes `[0]` would emit the same template twice in the few-shot events and fail the per-templateId assertion. Capability-gated on `capabilities.prompts.supported`. | `completed` | ≤ 10s |
|
|
@@ -303,6 +304,17 @@ The `messages`-mode stream fixture (AI token streaming) is covered by the determ
|
|
|
303
304
|
3. Server emits `run.cancelled` within 5s.
|
|
304
305
|
4. Subsequent `GET /v1/runs/{runId}` MUST return `status: "cancelled"`.
|
|
305
306
|
|
|
307
|
+
### `conformance-replay-ordinal-loop`
|
|
308
|
+
|
|
309
|
+
- **Purpose**: witness that a replay fork keys a side-effecting node's recorded outcome on `(sourceRunId, nodeId, n)`, where `n` counts the node's `node.started` events through this execution (`spec/v2/core/replay.md` §Suppression rule 2, openwop#1718).
|
|
310
|
+
- **Shape**: `start` (`core.noop`) → `effect` (`core.conformance.side-effect`, `any_success`) → `wait` (`core.delay`) → `effect` (`any_success`); `settings.maxLoopbackIterations: 2`, so `effect` executes twice.
|
|
311
|
+
- **Inputs**: `delayMs` (integer, default 3000), long enough to cancel inside `wait`.
|
|
312
|
+
- **Expected behavior**:
|
|
313
|
+
1. The source run completes `effect`'s first execution. The client cancels it inside `wait`, before the second execution.
|
|
314
|
+
2. The client forks `mode: "replay"` at `effect`'s first `node.started`.
|
|
315
|
+
3. In the fork, execution 1 (n = 1) resolves from the recorded outcome. Execution 2 (n = 2) has none, so it MUST fail closed with `replay_source_missing` and MUST NOT complete.
|
|
316
|
+
- **Advertise only if** the host runs cycles and maps `core.conformance.side-effect` to a side-effecting node (see `conformance-replay-side-effect`).
|
|
317
|
+
|
|
306
318
|
### `conformance-replay-side-effect`
|
|
307
319
|
|
|
308
320
|
- **Purpose**: verify RFC 0140 — a `mode:"replay"` fork MUST NOT re-perform a side-effecting node's effect.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@openwop/openwop-conformance",
|
|
3
|
-
"version": "2.43.
|
|
3
|
+
"version": "2.43.2",
|
|
4
4
|
"description": "Production-ready black-box conformance suite for OpenWOP v1.0 compliant servers.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -56,6 +56,6 @@
|
|
|
56
56
|
"@openwop/spec-artifacts": "file:../spec-artifacts"
|
|
57
57
|
},
|
|
58
58
|
"peerDependencies": {
|
|
59
|
-
"@openwop/spec-artifacts": "2.43.
|
|
59
|
+
"@openwop/spec-artifacts": "2.43.2"
|
|
60
60
|
}
|
|
61
61
|
}
|
package/requirements.json
CHANGED
|
@@ -2,11 +2,11 @@
|
|
|
2
2
|
"$comment": "GENERATED by conformance/scripts/generate-requirement-registry.mjs — do not edit. One record per it()/test() in src/scenarios. Ids: openwop.it.<file-stem>.<title-slug>[~n] (src/lib/requirement-ids.ts). A record with id null has an interpolated title; its run-time row is keyed by the rendered title and maps here by file+line only. Renamed ids need a row in requirement-aliases.json.",
|
|
3
3
|
"generatedFrom": "src/scenarios/*.test.ts",
|
|
4
4
|
"counts": {
|
|
5
|
-
"files":
|
|
6
|
-
"tests":
|
|
7
|
-
"withStableId":
|
|
5
|
+
"files": 648,
|
|
6
|
+
"tests": 2490,
|
|
7
|
+
"withStableId": 2490,
|
|
8
8
|
"interpolatedTitles": 0,
|
|
9
|
-
"explicitIds":
|
|
9
|
+
"explicitIds": 2372
|
|
10
10
|
},
|
|
11
11
|
"records": [
|
|
12
12
|
{
|
|
@@ -33075,6 +33075,24 @@
|
|
|
33075
33075
|
}
|
|
33076
33076
|
]
|
|
33077
33077
|
},
|
|
33078
|
+
{
|
|
33079
|
+
"id": "openwop.it.v2-replay-suppression-ordinal.a-replay-fork-resolves-each-execution-of-a-looped-side-effecting-node-by-its-own",
|
|
33080
|
+
"file": "v2-replay-suppression-ordinal.test.ts",
|
|
33081
|
+
"line": 73,
|
|
33082
|
+
"title": "a replay fork resolves each execution of a looped side-effecting node by its own ordinal, failing closed where none was recorded",
|
|
33083
|
+
"explicitId": "openwop.requirement.replay.suppression-execution-ordinal",
|
|
33084
|
+
"citations": [
|
|
33085
|
+
{
|
|
33086
|
+
"section": "spec/v2/core/replay.md §Suppression",
|
|
33087
|
+
"requirement": "only the execution with a recorded outcome (n = 1) may complete in the fork; execution 2 has none and MUST NOT complete"
|
|
33088
|
+
},
|
|
33089
|
+
{
|
|
33090
|
+
"section": null,
|
|
33091
|
+
"requirement": null,
|
|
33092
|
+
"interpolated": true
|
|
33093
|
+
}
|
|
33094
|
+
]
|
|
33095
|
+
},
|
|
33078
33096
|
{
|
|
33079
33097
|
"id": "openwop.it.v2-revocation-honored.a-revoked-next-request-credential-is-refused-on-the-next-request-with-credential",
|
|
33080
33098
|
"file": "v2-revocation-honored.test.ts",
|
|
@@ -33322,6 +33340,38 @@
|
|
|
33322
33340
|
}
|
|
33323
33341
|
]
|
|
33324
33342
|
},
|
|
33343
|
+
{
|
|
33344
|
+
"id": "openwop.it.v2-run-fork-ancestry.a-branch-fork-is-not-a-composition-child-its-ancestry-parent-is-null",
|
|
33345
|
+
"file": "v2-run-fork-ancestry.test.ts",
|
|
33346
|
+
"line": 51,
|
|
33347
|
+
"title": "a branch fork is not a composition child: its ancestry parent is null",
|
|
33348
|
+
"explicitId": "openwop.requirement.0040.fork-has-no-ancestry-parent",
|
|
33349
|
+
"citations": [
|
|
33350
|
+
{
|
|
33351
|
+
"section": "spec/v2/core/runs.md §Diff and ancestry",
|
|
33352
|
+
"requirement": null,
|
|
33353
|
+
"interpolated": true
|
|
33354
|
+
},
|
|
33355
|
+
{
|
|
33356
|
+
"section": "spec/v2/core/runs.md §Diff and ancestry",
|
|
33357
|
+
"requirement": "the control: a run created by POST /runs was dispatched by no parent, so its ancestry parent MUST be null"
|
|
33358
|
+
},
|
|
33359
|
+
{
|
|
33360
|
+
"section": "spec/v2/core/runs.md §Fork",
|
|
33361
|
+
"requirement": "the fork's lineage is carried by the 201: sourceRunId MUST name the source"
|
|
33362
|
+
},
|
|
33363
|
+
{
|
|
33364
|
+
"section": "spec/v2/core/runs.md §Diff and ancestry",
|
|
33365
|
+
"requirement": null,
|
|
33366
|
+
"interpolated": true
|
|
33367
|
+
},
|
|
33368
|
+
{
|
|
33369
|
+
"section": "spec/v2/core/runs.md §Diff and ancestry",
|
|
33370
|
+
"requirement": null,
|
|
33371
|
+
"interpolated": true
|
|
33372
|
+
}
|
|
33373
|
+
]
|
|
33374
|
+
},
|
|
33325
33375
|
{
|
|
33326
33376
|
"id": "openwop.it.v2-run-fork-prefix.a-replay-fork-inherits-exactly-0-fromseq-the-event-at-fromseq-is-re-executed-nev",
|
|
33327
33377
|
"file": "v2-run-fork-prefix.test.ts",
|
package/scenario-majors.json
CHANGED
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$comment": "GENERATED by conformance/scripts/generate-scenario-majors.mjs (RFC 0168 §D.3). Do not edit; add a file to BOTH_MAJORS in the generator to target both majors.",
|
|
3
3
|
"counts": {
|
|
4
|
-
"files":
|
|
4
|
+
"files": 573,
|
|
5
5
|
"v1": 452,
|
|
6
|
-
"v2":
|
|
6
|
+
"v2": 136
|
|
7
7
|
},
|
|
8
8
|
"majors": {
|
|
9
9
|
"a2a-1-0-agent-card.test.ts": [
|
|
@@ -1521,6 +1521,9 @@
|
|
|
1521
1521
|
"v2-relaxation-recorded.test.ts": [
|
|
1522
1522
|
2
|
|
1523
1523
|
],
|
|
1524
|
+
"v2-replay-suppression-ordinal.test.ts": [
|
|
1525
|
+
2
|
|
1526
|
+
],
|
|
1524
1527
|
"v2-revocation-honored.test.ts": [
|
|
1525
1528
|
2
|
|
1526
1529
|
],
|
|
@@ -1539,6 +1542,9 @@
|
|
|
1539
1542
|
"v2-run-diff-identical.test.ts": [
|
|
1540
1543
|
2
|
|
1541
1544
|
],
|
|
1545
|
+
"v2-run-fork-ancestry.test.ts": [
|
|
1546
|
+
2
|
|
1547
|
+
],
|
|
1542
1548
|
"v2-run-fork-prefix.test.ts": [
|
|
1543
1549
|
2
|
|
1544
1550
|
],
|
|
@@ -1,17 +1,17 @@
|
|
|
1
1
|
{
|
|
2
2
|
"_comment": "Provenance of @openwop/spec-artifacts (RFC 0168 §D.2). files: SHA-256 per file; the conformance suite compares the installed peer against dist/spec-artifacts.lock.json at start.",
|
|
3
3
|
"package": "@openwop/spec-artifacts",
|
|
4
|
-
"version": "2.43.
|
|
5
|
-
"corpusTag": "v2.43.
|
|
4
|
+
"version": "2.43.2",
|
|
5
|
+
"corpusTag": "v2.43.2",
|
|
6
6
|
"files": {
|
|
7
7
|
"api/.redocly.lint-ignore.yaml": "bf5a8350b88a72fa43f59605ed8d903ed24b6cfccda5e45509c9f6ed9ee4e712",
|
|
8
8
|
"api/asyncapi.yaml": "d5ecb9ee6114582be3b1f662c84bfac9ae96dae7bacb853e461168f70a8e1c7d",
|
|
9
9
|
"api/grpc/openwop.proto": "c3e72bb17cba514ee98feb6434e6c9b6ea6795bfd086489ec69fd882dd1ad977",
|
|
10
10
|
"api/openapi.yaml": "69464bc0e1e71ef3b543a73d4fa4f67202343d2603c80af8033f6c3619ef1b30",
|
|
11
11
|
"api/redocly.yaml": "b0604c89b2ca6d5076ec25725c539dad44a741a811fe524439ee6daef8baa09f",
|
|
12
|
-
"api/seams-v2.yaml": "
|
|
13
|
-
"api/v2/asyncapi.yaml": "
|
|
14
|
-
"api/v2/openapi.yaml": "
|
|
12
|
+
"api/seams-v2.yaml": "987123ecfbfbc906c5655538b9380021ac0bcbfd93a467a67818e855d624ed0e",
|
|
13
|
+
"api/v2/asyncapi.yaml": "76474704b7a8e829588abec30d31f09e11557ee5daacad301c3be7e77affe7ce",
|
|
14
|
+
"api/v2/openapi.yaml": "f4af9e7fddad438d3a70753af29e5738ea48d15d1abd214a3fbeec97260cea9d",
|
|
15
15
|
"api/v2/redocly.yaml": "1e66b60e6118ad11a823bb620678be464d99dfe50a40e3e6f93ec9429b88b34c",
|
|
16
16
|
"schemas/README.md": "0c0b737ffcf8f30e7d2809cec8a498232de710f41443212922ad8337cdde0b51",
|
|
17
17
|
"schemas/a2a-task-state.schema.json": "75d5049dea8bd873ff0e7546f1c60c8d36c219bec7084264360be54a8be30ae1",
|
|
@@ -204,13 +204,13 @@
|
|
|
204
204
|
"schemas/workspace-file.schema.json": "464de85c2a068243084ee9c1d969bc7cd5d8f7948574e58450d6493c38a0e1e4",
|
|
205
205
|
"spec/v1/alias-detectors.json": "29f21e0fb2c7a232a92a5c19b4cba3a86bc3b5e16f06780452f0421c66c6ec38",
|
|
206
206
|
"spec/v1/capability-declaration-classes.json": "e7729aed5c4b4e1dd02abab0530f14cc95f5d4070fe51fb139e7f5cccefa00c6",
|
|
207
|
-
"spec/v1/core-standard-manifest.json": "
|
|
207
|
+
"spec/v1/core-standard-manifest.json": "0be2c8da837ddb170866303ab4640c8e8a8b2554d8c2bb9ac3d1bd9780bede5a",
|
|
208
208
|
"spec/v1/deprecations.json": "46df807fb4adf0aae24671c20f69975cf3496e3bc3a521cb7e6ea556bdce563c",
|
|
209
209
|
"spec/v1/deprecations.schema.json": "3e393c405d2a41b467d8c5e3c468549078df1ce6a6d2588fc488097b95d9b55b",
|
|
210
210
|
"spec/v1/event-codemap.json": "3da60d884157793a360da532a9fcbbfb5285636db325a74cec94b34622186d97",
|
|
211
211
|
"spec/v1/event-codemap.schema.json": "d05933b2e88103aff51a2f774df97b9bda065e0fc88dbd9b114cc47a32189174",
|
|
212
212
|
"spec/v1/extensions.json": "79a60754aa16cbdbdb604f8e5c00038af535a1690a4d07bd5ff7c38f0256a7bc",
|
|
213
|
-
"spec/v1/gaps.json": "
|
|
213
|
+
"spec/v1/gaps.json": "7eefa3aa02e978a41e1e5948818961718d23cb0cbe29a4a491524360feb9fd6f",
|
|
214
214
|
"spec/v1/gaps.schema.json": "8fd83259f556553c9df0f53e7a82ca8c2a4771d471197f69b4e2ae0ceeacfd66",
|
|
215
215
|
"spec/v1/migrations.json": "2efcd9064b2bc65a84a3b3d281aa86a91efd002690c1f4f14e8335c4a18194d1",
|
|
216
216
|
"spec/v1/migrations.schema.json": "886779aa6c22e646db097f5df210adb018a4dd14a7b815465a18c8a7056c8f72",
|
|
@@ -230,15 +230,15 @@
|
|
|
230
230
|
"spec/v2/core/idempotency.md": "2084d756789c1faf5121593fb39e67ca36a4c227178e175004742e403178ea49",
|
|
231
231
|
"spec/v2/core/identity.md": "2dceb1c2697015e4a4fa8a9b28f4fbd25fa44a76ca5a919c752e2a8163c819db",
|
|
232
232
|
"spec/v2/core/interop.md": "e72e25dff864c83b2980d7366b34054cffa072af0225821199f260f6228d260e",
|
|
233
|
-
"spec/v2/core/interrupt.md": "
|
|
233
|
+
"spec/v2/core/interrupt.md": "2c8a4026930ec40ba47d7415dd02b6eda997a03d7a084e5084fec6f1261aba1c",
|
|
234
234
|
"spec/v2/core/node-pack-runtimes.md": "e1088cb69b8c3c67256c494063328916c2980ab1aa97a7063de25bab82fdb1e5",
|
|
235
235
|
"spec/v2/core/oauth.md": "c96483a44b62978c596945fbb7d7c863cf6056cd980fdb4cf4e97555e3e10b72",
|
|
236
236
|
"spec/v2/core/overview.md": "4087f171b00a3907551f508fec32fb9b0b3ccec8c0ac1b0c7c6d32cd2c0e3ff8",
|
|
237
237
|
"spec/v2/core/packs.md": "b472a583211d466d6ab8e00c414c225b246e678043fe6283a710c899fadde913",
|
|
238
238
|
"spec/v2/core/persistence.md": "711874f74cb08b03cd118d683056db10a063517762fcea73ec70feb5d983e6a4",
|
|
239
239
|
"spec/v2/core/portability.md": "87c4327f674f714fda6217fd79b99d2ba219ad47dc02f9751bb56c200fb2c17a",
|
|
240
|
-
"spec/v2/core/replay.md": "
|
|
241
|
-
"spec/v2/core/runs.md": "
|
|
240
|
+
"spec/v2/core/replay.md": "4d379494057746c6883f6047972e93e621863b998ed291b9a691d3aa57c1ff43",
|
|
241
|
+
"spec/v2/core/runs.md": "55e21292ae4b41d0d104199a6c2634ecb3f787c4a671f9ebc37084be38ceef03",
|
|
242
242
|
"spec/v2/core/security-defaults.md": "153884da7a2d54c30c81b5e649cfdef31e151b3a2a19dee1bc5a16c35d60a3aa",
|
|
243
243
|
"spec/v2/core/tool-catalog.md": "d16a9004418719c48301804aefa4d7ab6f41112a4ed83727f552aa1b92b418b0",
|
|
244
244
|
"spec/v2/core/versioning.md": "e3b3ab6d76e51035039a32173e9b412628a52c8fe7b16687ff58a2d0d983faa0",
|
|
@@ -246,7 +246,7 @@
|
|
|
246
246
|
"spec/v2/core/workflow-chain-packs.md": "fdce7a8477e269950a2bac05ee6eddbd66eb0b670b464d00054f7196ec05d29c",
|
|
247
247
|
"spec/v2/corrections.json": "59f647f4d22477f6c2b06dcf913b682dde5d51dd5e607a5d0cf648ab60864d6e",
|
|
248
248
|
"spec/v2/corrections.schema.json": "46bbffb954d13bd9995aebc912ceeb2544f7064331abf25fd6ab85c81a5cec94",
|
|
249
|
-
"spec/v2/declaration.json": "
|
|
249
|
+
"spec/v2/declaration.json": "184b83919e56d2f12bda73c7772905810a441cbab53fc42ad0b3135937fac3e5",
|
|
250
250
|
"spec/v2/declaration.schema.json": "ef951633e4070899f64828f5cf1b5f4870270a0360f81dbb49b5b4d6cd407550",
|
|
251
251
|
"spec/v2/errors.json": "dc1a846abe3e888acb0ce4a33cf9d1fb98e41fcc8b49f42b3682fd65deca2eb2",
|
|
252
252
|
"spec/v2/errors.schema.json": "47404ad53e6147cf9017b813483712466573e419fb33e6c4b6c12d3e7deba1bc",
|
|
@@ -254,7 +254,7 @@
|
|
|
254
254
|
"spec/v2/event-codemap.schema.json": "b173f2a9bcc0bf9b62a474ce3b91e1431d597e64fd1560fc58662cd8606eaf9c",
|
|
255
255
|
"spec/v2/ext/README.md": "e2731ebb5156ece1840ed3b96e7a347f6345e73071ff4e54df0f21c1585c5027",
|
|
256
256
|
"spec/v2/ext/a2uiSurface/README.md": "5518c7f355262eb7cc584422f4ff8e548f066b211878f969274a3a578127622d",
|
|
257
|
-
"spec/v2/ext/brand/README.md": "
|
|
257
|
+
"spec/v2/ext/brand/README.md": "82989c018d7965bab029bcce7a67b1d76ffc433d08971a4cac91befcccc09438",
|
|
258
258
|
"spec/v2/ext/canvas/README.md": "0e91b68566d2be90319feedfaebe2cd2db0bd8f72fb59c958fa723c05e38023f",
|
|
259
259
|
"spec/v2/ext/chat/README.md": "f59b6706162abbfff96f8bccd4c0e1aeed721047cf62e9af58c875727a38d716",
|
|
260
260
|
"spec/v2/ext/coordination/README.md": "c7fe0884f52405fdd02c5f9ec309ac6b585685ee06ec5f338fff7acf03c07148",
|
|
@@ -263,12 +263,12 @@
|
|
|
263
263
|
"spec/v2/ext/grpc-transport/README.md": "6ba378e3c6d1f2e3269b392cfdf0c40565ee9bfa259d1baaa9ee1d350e38e792",
|
|
264
264
|
"spec/v2/ext/kanban/README.md": "bfc386dc0f874a4afa1e3ec6a2f99084b547dcce9bd363da3b8ebd8786b9c184",
|
|
265
265
|
"spec/v2/ext/knowledge/README.md": "cbe1f945a757791dedac0fd542a68c86c5b457778b831d58d9ae44a2baa270f2",
|
|
266
|
-
"spec/v2/ext/launchStudio/README.md": "
|
|
267
|
-
"spec/v2/ext/messaging/README.md": "
|
|
266
|
+
"spec/v2/ext/launchStudio/README.md": "a941ef1d2401438a412eaea48d2d6ca807a032b44ea08719fea33120daf088a0",
|
|
267
|
+
"spec/v2/ext/messaging/README.md": "949f350e35294f52ad490e42699f8c560cd154549516fbaa1805f06312b7f342",
|
|
268
268
|
"spec/v2/ext/portability/README.md": "c59423ee446638b748a911b9e97e803c4786157be6120f656f0a060586bc7b82",
|
|
269
269
|
"spec/v2/ext/provider-idempotency/README.md": "9d1bccef0ec16a19a3370ae79c2f5a1726ce0d1d25a51352907f288a2ccb15d7",
|
|
270
270
|
"spec/v2/ext/provider-idempotency/registry.json": "5b5fecad604fb4ab38da39f5abe697a4e57b05406c398f8afeb377331b2bd965",
|
|
271
|
-
"spec/v2/ext/restTransport/README.md": "
|
|
271
|
+
"spec/v2/ext/restTransport/README.md": "a8115faefeab01b7c85d516de60ebfc99b8e89552d16ba654456224518447b28",
|
|
272
272
|
"spec/v2/ext/sandbox-runtime-notes/README.md": "0d1290fb1e20b2e35c0b15051541ab6546e4e2a6f693b4f912cb3f7d6a65e92c",
|
|
273
273
|
"spec/v2/ext/webResearch/README.md": "b675e6f6646128d781b661f754c2393eed86e5935d16c3a22df582a5ce025b9b",
|
|
274
274
|
"spec/v2/facets/a2a.schema.json": "b68bf5e29e455e46c2ad7cfcccf5963730ec2310ff9db81884550d8df30fd6ef",
|
|
@@ -298,9 +298,9 @@
|
|
|
298
298
|
"spec/v2/path-manifest.json": "17298a245c5da9a70aa7ab0b7f2db4277e06209d6fcdd7187a8b4cea179c08f3",
|
|
299
299
|
"spec/v2/peer-dependency-aliases.json": "d10299280abee08258502925bc327293ee413e0108cd6e6ec75ff6110653308d",
|
|
300
300
|
"spec/v2/profiles.json": "180beb9ef2d161766fa23b2987e7e829851e3d316273147cf6a4a3f324befd6b",
|
|
301
|
-
"spec/v2/release.json": "
|
|
301
|
+
"spec/v2/release.json": "ec99d76b53fba89d3d8788206c649a782774b55acf5284b75cab38b79ecf1d61",
|
|
302
302
|
"spec/v2/retention-floors.json": "eaf3722d95c79947af1d4269ef85117e126518c588cfcf1a2b21b97269f51624",
|
|
303
|
-
"spec/v2/surface-baseline.json": "
|
|
303
|
+
"spec/v2/surface-baseline.json": "3d3d0fc7c0b48e7d7339e5d8ada0640487908f37220b2de65ccc9bff093f876a"
|
|
304
304
|
},
|
|
305
|
-
"corpusCommit": "
|
|
305
|
+
"corpusCommit": "6612868966f2ecb9672d0a64382cb7067491c004"
|
|
306
306
|
}
|
|
@@ -153,6 +153,46 @@ export interface TestFailure {
|
|
|
153
153
|
readonly message?: string;
|
|
154
154
|
}
|
|
155
155
|
|
|
156
|
+
/** A finished `it` as vitest's own task tree reports it. */
|
|
157
|
+
export interface FinishedTest {
|
|
158
|
+
readonly id: string;
|
|
159
|
+
readonly name: string;
|
|
160
|
+
readonly state: string | undefined;
|
|
161
|
+
readonly message?: string;
|
|
162
|
+
}
|
|
163
|
+
|
|
164
|
+
/**
|
|
165
|
+
* The tests the per-test recorder never saw, as file states and failures.
|
|
166
|
+
*
|
|
167
|
+
* vitest runs a file's `afterEach` hooks in stack order, so a scenario's own
|
|
168
|
+
* `afterEach` (unregistering webhooks, closing a receiver) runs BEFORE the
|
|
169
|
+
* recorder `setup.ts` registers. If that hook throws or times out, vitest
|
|
170
|
+
* fails the test and skips the remaining hooks, so the recorder never runs.
|
|
171
|
+
* The file then has no recorded state, and its row read "no test executed and
|
|
172
|
+
* no disposition recorded" (`blocked`). That row gives no reason and hides a
|
|
173
|
+
* real failure. One tier-2 cut of `v2-webhook-message-id-stable` produced it
|
|
174
|
+
* on 2026-09-28.
|
|
175
|
+
*
|
|
176
|
+
* Only tests that finished `pass` or `fail` are recovered. A skipped test
|
|
177
|
+
* never reaches the recorder either way, and the existing rules already cover it.
|
|
178
|
+
*/
|
|
179
|
+
export function unrecordedTests(
|
|
180
|
+
tests: readonly FinishedTest[],
|
|
181
|
+
recordedIds: ReadonlySet<string>,
|
|
182
|
+
): { states: FileTestState[]; failures: TestFailure[] } {
|
|
183
|
+
const states: FileTestState[] = [];
|
|
184
|
+
const failures: TestFailure[] = [];
|
|
185
|
+
for (const t of tests) {
|
|
186
|
+
if (recordedIds.has(t.id)) continue;
|
|
187
|
+
if (t.state === 'pass') states.push('pass');
|
|
188
|
+
else if (t.state === 'fail') {
|
|
189
|
+
states.push('fail');
|
|
190
|
+
failures.push({ name: t.name, message: `the test or its afterEach hook failed before the runner recorded it: ${t.message ?? 'no message'}` });
|
|
191
|
+
}
|
|
192
|
+
}
|
|
193
|
+
return { states, failures };
|
|
194
|
+
}
|
|
195
|
+
|
|
156
196
|
/**
|
|
157
197
|
* The `executed-fail` detail for a file row: WHICH cases failed, and what the
|
|
158
198
|
* first one said.
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* `spec/v2/core/replay.md` §Suppression rule 2 — a replay fork resolves a
|
|
3
|
+
* side-effecting node's outcome keyed on `(sourceRunId, nodeId, n)`, where `n`
|
|
4
|
+
* counts the node's `node.started` events through this execution, including
|
|
5
|
+
* retries, later visits and the fork's inherited prefix (openwop#1718, suite
|
|
6
|
+
* 2.43.2, target major 2).
|
|
7
|
+
*
|
|
8
|
+
* The rule said `(nodeId, attempt)`. In a run where a node executes more than
|
|
9
|
+
* once (a loop back over an edge), that key names no single execution: a host
|
|
10
|
+
* counting attempts per execution resolves every visit from the first visit's
|
|
11
|
+
* outcome, and one selecting the latest recorded outcome does the same when
|
|
12
|
+
* only one exists.
|
|
13
|
+
*
|
|
14
|
+
* Construction (`conformance-replay-ordinal-loop`): start → effect → wait →
|
|
15
|
+
* effect, so `effect` executes twice. The suite cancels the source inside
|
|
16
|
+
* `wait`, after `effect`'s first execution completed and before its second, and
|
|
17
|
+
* forks `mode: "replay"` at `effect`'s first `node.started`. In the fork:
|
|
18
|
+
* - `effect` execution 1 (n = 1) has a recorded outcome and resolves from it;
|
|
19
|
+
* - `effect` execution 2 (n = 2) has none, so it MUST fail closed with
|
|
20
|
+
* `replay_source_missing` (rule 3) and MUST NOT complete.
|
|
21
|
+
* Only the fail-closed half is asserted as the discriminator: a host keyed on
|
|
22
|
+
* the wrong ordinal completes execution 2 from execution 1's outcome.
|
|
23
|
+
*
|
|
24
|
+
* Gated on `replay` and on the fixture, which only a host that runs cycles and
|
|
25
|
+
* suppresses side effects advertises. No such host is known today (the v2
|
|
26
|
+
* reference host does not loop), so the row has a negative control only — see
|
|
27
|
+
* RFC 0140's gap register.
|
|
28
|
+
*
|
|
29
|
+
* @see spec/v2/core/replay.md §Suppression
|
|
30
|
+
* @see RFCS/0140-replay-side-effect-suppression.md
|
|
31
|
+
*/
|
|
32
|
+
|
|
33
|
+
import { describe, it, expect } from 'vitest';
|
|
34
|
+
import { driver, type OpenWOPResponse } from '../lib/driver.js';
|
|
35
|
+
import { gateFamily } from '../lib/v2.js';
|
|
36
|
+
import { isFixtureAdvertised } from '../lib/fixtures.js';
|
|
37
|
+
import { readErrorCode } from '../lib/error-envelope.js';
|
|
38
|
+
import { softSkip } from '../lib/soft-skip.js';
|
|
39
|
+
import { req } from '../lib/requirement-ids.js';
|
|
40
|
+
import { scaledTimeoutMs } from '../lib/polling.js';
|
|
41
|
+
|
|
42
|
+
const ID = 'openwop.requirement.replay.suppression-execution-ordinal';
|
|
43
|
+
const DOC = 'spec/v2/core/replay.md §Suppression';
|
|
44
|
+
const FIXTURE = 'conformance-replay-ordinal-loop';
|
|
45
|
+
const EFFECT = 'effect';
|
|
46
|
+
const TERMINAL = new Set(['completed', 'failed', 'cancelled']);
|
|
47
|
+
|
|
48
|
+
interface Ev { readonly sequence?: unknown; readonly type?: unknown; readonly nodeId?: unknown; readonly payload?: { readonly nodeId?: unknown; readonly error?: { readonly code?: unknown } } }
|
|
49
|
+
|
|
50
|
+
const enc = (id: string): string => encodeURIComponent(id);
|
|
51
|
+
async function http(fn: () => Promise<OpenWOPResponse>): Promise<OpenWOPResponse | null> { try { return await fn(); } catch { return null; } }
|
|
52
|
+
const nodeOf = (e: Ev): unknown => e.nodeId ?? e.payload?.nodeId;
|
|
53
|
+
|
|
54
|
+
async function logOf(runId: string): Promise<Ev[]> {
|
|
55
|
+
const res = await http(() => driver.get(`/runs/${enc(runId)}/events/poll?timeout=1&streamMode=debug`));
|
|
56
|
+
const events = (res?.json as { events?: unknown } | null)?.events;
|
|
57
|
+
return Array.isArray(events) ? (events as Ev[]).filter((e) => typeof e.sequence === 'number').sort((a, b) => (a.sequence as number) - (b.sequence as number)) : [];
|
|
58
|
+
}
|
|
59
|
+
async function statusOf(runId: string): Promise<string | null> {
|
|
60
|
+
const res = await http(() => driver.get(`/runs/${enc(runId)}`));
|
|
61
|
+
return res?.status === 200 ? String((res.json as { status?: unknown } | null)?.status ?? '') : null;
|
|
62
|
+
}
|
|
63
|
+
async function until(check: () => Promise<boolean>, ms: number): Promise<boolean> {
|
|
64
|
+
const deadline = Date.now() + ms;
|
|
65
|
+
for (;;) {
|
|
66
|
+
if (await check()) return true;
|
|
67
|
+
if (Date.now() > deadline) return false;
|
|
68
|
+
await new Promise((r) => setTimeout(r, 200));
|
|
69
|
+
}
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
describe('v2 replay-suppression-ordinal (replay.md §Suppression rule 2, openwop#1718)', () => {
|
|
73
|
+
it('a replay fork resolves each execution of a looped side-effecting node by its own ordinal, failing closed where none was recorded', async () => {
|
|
74
|
+
if (!(await gateFamily('replay'))) return softSkip('inapplicable', 'replay family not advertised (gate recorded under openwop.family.replay)');
|
|
75
|
+
if (!isFixtureAdvertised(FIXTURE)) return softSkip('inapplicable', `fixture ${FIXTURE} is not advertised — the host does not run cycles with a side-effecting node`);
|
|
76
|
+
|
|
77
|
+
const created = await http(() => driver.post('/runs', { workflowId: FIXTURE, inputs: { delayMs: 3000 } }));
|
|
78
|
+
const runId = (created?.json as { runId?: unknown } | null)?.runId;
|
|
79
|
+
if (created === null || created.status !== 201 || typeof runId !== 'string') return softSkip('blocked', `POST /runs (${FIXTURE}) answered ${created?.status ?? 'nothing'} ${readErrorCode(created?.json) ?? ''}`.trim());
|
|
80
|
+
|
|
81
|
+
// Cancel inside `wait`: effect's first execution completed, its second not yet started.
|
|
82
|
+
const firstDone = await until(async () => (await logOf(runId)).some((e) => e.type === 'node.completed' && nodeOf(e) === EFFECT), scaledTimeoutMs(10_000));
|
|
83
|
+
if (!firstDone) return softSkip('blocked', `the source run never completed ${EFFECT}'s first execution`);
|
|
84
|
+
await http(() => driver.post(`/runs/${enc(runId)}/cancel`, {}));
|
|
85
|
+
if (!(await until(async () => TERMINAL.has((await statusOf(runId)) ?? ''), scaledTimeoutMs(10_000)))) return softSkip('blocked', 'the cancelled source run did not settle');
|
|
86
|
+
const source = await logOf(runId);
|
|
87
|
+
const effectStarts = source.filter((e) => e.type === 'node.started' && nodeOf(e) === EFFECT);
|
|
88
|
+
const effectDone = source.filter((e) => e.type === 'node.completed' && nodeOf(e) === EFFECT);
|
|
89
|
+
if (effectStarts.length !== 1 || effectDone.length !== 1) return softSkip('blocked', `the source was cancelled with ${effectStarts.length} starts and ${effectDone.length} completions of ${EFFECT}, not exactly one of each — the second execution cannot be isolated`);
|
|
90
|
+
|
|
91
|
+
const fromSeq = effectStarts[0]!.sequence as number;
|
|
92
|
+
const fork = await http(() => driver.post(`/runs/${enc(runId)}:fork`, { mode: 'replay', fromSeq }));
|
|
93
|
+
const forkId = (fork?.json as { runId?: unknown } | null)?.runId;
|
|
94
|
+
if (fork === null || fork.status !== 201 || typeof forkId !== 'string') return softSkip('blocked', `POST /runs/{runId}:fork answered ${fork?.status ?? 'nothing'} ${readErrorCode(fork?.json) ?? ''} — forkRun owns that contract`.trim());
|
|
95
|
+
await until(async () => TERMINAL.has((await statusOf(forkId)) ?? ''), scaledTimeoutMs(20_000));
|
|
96
|
+
const forked = await logOf(forkId);
|
|
97
|
+
const forkEffect = forked.filter((e) => nodeOf(e) === EFFECT);
|
|
98
|
+
|
|
99
|
+
expect(
|
|
100
|
+
forkEffect.filter((e) => e.type === 'node.completed').length,
|
|
101
|
+
req(ID, DOC, 'only the execution with a recorded outcome (n = 1) may complete in the fork; execution 2 has none and MUST NOT complete'),
|
|
102
|
+
).toBeLessThanOrEqual(1);
|
|
103
|
+
expect(
|
|
104
|
+
forkEffect.some((e) => e.type === 'node.failed' && e.payload?.error?.code === 'replay_source_missing'),
|
|
105
|
+
req(ID, `${DOC} rules 2–3`, `the second execution of ${EFFECT} (n = 2) has no recorded source outcome and MUST fail closed with replay_source_missing — a host keyed on (nodeId, attempt), or on the latest outcome, resolves it from the first execution instead`),
|
|
106
|
+
).toBe(true);
|
|
107
|
+
}, 90_000);
|
|
108
|
+
});
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* `spec/v2/core/runs.md` §Diff and ancestry — a fork has no ancestry parent
|
|
3
|
+
* (suite 2.43.2, target major 2; gated on `replay` and on
|
|
4
|
+
* `multiAgent.executionModel.crossHostCausation.ancestryEndpointSupported`;
|
|
5
|
+
* one run created plus one branch fork).
|
|
6
|
+
*
|
|
7
|
+
* `getRunAncestry`'s `parent` is the run that dispatched this one, and `cause`
|
|
8
|
+
* names the composition mechanism it used (`core.subWorkflow`,
|
|
9
|
+
* `core.dispatch`, `mcp-tool-call`, `a2a-message`). A fork is not dispatched:
|
|
10
|
+
* its lineage is `sourceRunId` on the fork `201` and `parentRunId` on the
|
|
11
|
+
* snapshot, and its ancestry `parent` MUST be `null`. A host that reuses its
|
|
12
|
+
* fork link as an ancestry parent reports `cause: core.subWorkflow` for a run
|
|
13
|
+
* no parent dispatched, and fails here.
|
|
14
|
+
*
|
|
15
|
+
* The source run's own ancestry is the control: it was created by `POST /runs`,
|
|
16
|
+
* so its `parent` is `null` too. Without it a host answering `null` for every
|
|
17
|
+
* run would pass the fork leg without being measured on it, and one answering
|
|
18
|
+
* non-null for every run would fail it for the wrong reason.
|
|
19
|
+
*
|
|
20
|
+
* Its own file, not a leg of `v2-run-fork-refusals`: the ancestry gate is
|
|
21
|
+
* absent on most v2 hosts, and an `inapplicable` leg would turn that file's
|
|
22
|
+
* row into a `partial-witness` pass on every one of them.
|
|
23
|
+
*
|
|
24
|
+
* @see spec/v2/core/runs.md §Diff and ancestry
|
|
25
|
+
* @see spec/v2/core/runs.md §Fork
|
|
26
|
+
* @see schemas/v2/run-ancestry-response.schema.json
|
|
27
|
+
* @see RFCS/0040-multi-agent-cross-host-causation.md §C
|
|
28
|
+
*/
|
|
29
|
+
|
|
30
|
+
import { describe, it, expect } from 'vitest';
|
|
31
|
+
import { driver, type OpenWOPResponse } from '../lib/driver.js';
|
|
32
|
+
import { v2Discovery, gateFamily, familyAdvertised } from '../lib/v2.js';
|
|
33
|
+
import { readErrorCode } from '../lib/error-envelope.js';
|
|
34
|
+
import { softSkip } from '../lib/soft-skip.js';
|
|
35
|
+
import { req } from '../lib/requirement-ids.js';
|
|
36
|
+
|
|
37
|
+
const ID = 'openwop.requirement.0040.fork-has-no-ancestry-parent';
|
|
38
|
+
const DOC = 'spec/v2/core/runs.md §Diff and ancestry';
|
|
39
|
+
const NOOP = 'conformance-noop';
|
|
40
|
+
const TERMINAL = new Set(['completed', 'failed', 'cancelled']);
|
|
41
|
+
|
|
42
|
+
async function discovery(): Promise<Record<string, unknown> | null> { try { return await v2Discovery(); } catch { return null; } }
|
|
43
|
+
async function http(fn: () => Promise<OpenWOPResponse>): Promise<OpenWOPResponse | null> { try { return await fn(); } catch { return null; } }
|
|
44
|
+
const enc = (id: string): string => encodeURIComponent(id);
|
|
45
|
+
async function status(runId: string): Promise<string | null> {
|
|
46
|
+
const res = await http(() => driver.get(`/runs/${enc(runId)}`));
|
|
47
|
+
return res?.status === 200 ? String((res.json as { status?: unknown } | null)?.status ?? '') : null;
|
|
48
|
+
}
|
|
49
|
+
|
|
50
|
+
describe('v2 run-fork-ancestry (runs.md §Diff and ancestry)', () => {
|
|
51
|
+
it('a branch fork is not a composition child: its ancestry parent is null', async () => {
|
|
52
|
+
if (!(await discovery())) return softSkip('blocked', 'v2 discovery unreachable');
|
|
53
|
+
if (!(await gateFamily('replay'))) return softSkip('inapplicable', 'replay family not advertised (gate recorded under openwop.family.replay) — forkRun is gated on replay (runs.md §Surface)');
|
|
54
|
+
const chc = ((await familyAdvertised('multiAgent'))?.['executionModel'] as { crossHostCausation?: { ancestryEndpointSupported?: unknown } } | undefined)?.crossHostCausation;
|
|
55
|
+
if (chc?.ancestryEndpointSupported !== true) return softSkip('inapplicable', 'multiAgent.executionModel.crossHostCausation.ancestryEndpointSupported not advertised — getRunAncestry is gated on it (runs.md §Surface)');
|
|
56
|
+
|
|
57
|
+
const created = await http(() => driver.post('/runs', { workflowId: NOOP }));
|
|
58
|
+
if (created === null) return softSkip('blocked', 'POST /runs unreachable (fetch failed)');
|
|
59
|
+
const runId = (created.json as { runId?: unknown } | null)?.runId;
|
|
60
|
+
if (created.status !== 201 || typeof runId !== 'string') return softSkip('blocked', `POST /runs answered ${created.status} ${readErrorCode(created.json) ?? ''}`.trim());
|
|
61
|
+
const t0 = Date.now(); let settled: string | null = null;
|
|
62
|
+
while (Date.now() - t0 < 10_000) { settled = await status(runId); if (settled !== null && TERMINAL.has(settled)) break; await new Promise((r) => setTimeout(r, 250)); }
|
|
63
|
+
if (settled === null || !TERMINAL.has(settled)) return softSkip('blocked', 'the noop run did not settle within 10 s');
|
|
64
|
+
|
|
65
|
+
const own = await http(() => driver.get(`/runs/${enc(runId)}/ancestry`));
|
|
66
|
+
if (own === null) return softSkip('blocked', 'GET /runs/{runId}/ancestry unreachable (fetch failed)');
|
|
67
|
+
expect(own.status, req(ID, DOC, `a host advertising ancestryEndpointSupported MUST serve getRunAncestry — got ${own.status} ${readErrorCode(own.json) ?? ''}`.trim())).toBe(200);
|
|
68
|
+
expect((own.json as { parent?: unknown } | null)?.parent, req(ID, DOC, 'the control: a run created by POST /runs was dispatched by no parent, so its ancestry parent MUST be null')).toBeNull();
|
|
69
|
+
|
|
70
|
+
const fork = await http(() => driver.post(`/runs/${enc(runId)}:fork`, { mode: 'branch', fromSeq: 0 }));
|
|
71
|
+
if (fork === null) return softSkip('blocked', 'POST /runs/{runId}:fork unreachable (fetch failed)');
|
|
72
|
+
if (fork.status === 404) return softSkip('blocked', 'POST /runs/{runId}:fork answered 404 with replay advertised — forkRun is not mounted');
|
|
73
|
+
const forkId = (fork.json as { runId?: unknown } | null)?.runId;
|
|
74
|
+
if (fork.status !== 201 || typeof forkId !== 'string') return softSkip('blocked', `a branch fork at fromSeq 0 answered ${fork.status} ${readErrorCode(fork.json) ?? ''} — no fork to read the ancestry of`.trim());
|
|
75
|
+
expect((fork.json as { sourceRunId?: unknown }).sourceRunId, req(ID, 'spec/v2/core/runs.md §Fork', 'the fork\'s lineage is carried by the 201: sourceRunId MUST name the source')).toBe(runId);
|
|
76
|
+
|
|
77
|
+
const anc = await http(() => driver.get(`/runs/${enc(forkId)}/ancestry`));
|
|
78
|
+
if (anc === null) return softSkip('blocked', 'GET /runs/{fork}/ancestry unreachable (fetch failed)');
|
|
79
|
+
expect(anc.status, req(ID, DOC, `getRunAncestry on a fork MUST answer 200 — got ${anc.status} ${readErrorCode(anc.json) ?? ''}`.trim())).toBe(200);
|
|
80
|
+
const parent = (anc.json as { parent?: unknown } | null)?.parent;
|
|
81
|
+
expect(parent, req(ID, DOC, `a fork is not dispatched, so its ancestry parent MUST be null — got ${JSON.stringify(parent)}. Its lineage is sourceRunId and the snapshot's parentRunId; ancestry's cause names a composition mechanism no fork used`)).toBeNull();
|
|
82
|
+
}, 30_000);
|
|
83
|
+
});
|
package/src/setup.ts
CHANGED
|
@@ -36,7 +36,7 @@ import { basename, join } from 'node:path';
|
|
|
36
36
|
import { existsSync, readFileSync } from 'node:fs';
|
|
37
37
|
import { PKG_ROOT_PATH } from './lib/paths.js';
|
|
38
38
|
import { recordRequirement, hasRequirement, journalLength, journalSince } from './lib/requirement-ledger.js';
|
|
39
|
-
import { requirementIdForFile, resolveFileRecord, resolveItRecord, type FileTestState, type TestFailure } from './lib/scenario-disposition.js';
|
|
39
|
+
import { requirementIdForFile, resolveFileRecord, resolveItRecord, unrecordedTests, type FileTestState, type FinishedTest, type TestFailure } from './lib/scenario-disposition.js';
|
|
40
40
|
import { softSkipDisposition, softSkipDispositionSince, softSkipMark } from './lib/soft-skip.js';
|
|
41
41
|
import { ItIdAllocator, takeExplicitRequirementId } from './lib/requirement-ids.js';
|
|
42
42
|
import { SPEC_COHERENCE_SCENARIOS, SPEC_COHERENCE_DETAIL } from './lib/spec-coherence.js';
|
|
@@ -250,6 +250,8 @@ const _fileStates = new Map<string, FileTestState[]>();
|
|
|
250
250
|
const _fileAssertions = new Map<string, number>();
|
|
251
251
|
/** Per file: the failing cases, so the file row can NAME them (2.37.0). */
|
|
252
252
|
const _fileFailures = new Map<string, TestFailure[]>();
|
|
253
|
+
/** Task ids the per-test recorder saw, so `afterAll` can recover the ones a throwing scenario hook hid (`unrecordedTests`). */
|
|
254
|
+
const _fileRecorded = new Map<string, Set<string>>();
|
|
253
255
|
const _ledgerMarks = new Map<string, number>();
|
|
254
256
|
// Per-`it` recording (v2 charter Phase 1, suite 1.153.0 — the durable G8 fix
|
|
255
257
|
// named in scenario-disposition.ts). Each test gets its own ledger row under
|
|
@@ -296,6 +298,20 @@ function registeredExplicitId(file: string, title: string): string | null {
|
|
|
296
298
|
return _registry.get(`${file}\u0000${title}`) ?? null;
|
|
297
299
|
}
|
|
298
300
|
|
|
301
|
+
/** Every `it` under a suite or file task, with the state vitest settled on. */
|
|
302
|
+
function _finishedTests(node: unknown): FinishedTest[] {
|
|
303
|
+
const out: FinishedTest[] = [];
|
|
304
|
+
const walk = (n: unknown): void => {
|
|
305
|
+
const t = n as { id?: string; name?: string; type?: string; tasks?: unknown[]; result?: { state?: string; errors?: Array<{ message?: string }> } };
|
|
306
|
+
if (t.type === 'test' && typeof t.id === 'string') {
|
|
307
|
+
const message = t.result?.errors?.[0]?.message;
|
|
308
|
+
out.push({ id: t.id, name: t.name ?? '', state: t.result?.state, ...(message === undefined ? {} : { message }) });
|
|
309
|
+
}
|
|
310
|
+
for (const c of t.tasks ?? []) walk(c);
|
|
311
|
+
};
|
|
312
|
+
walk(node);
|
|
313
|
+
return out;
|
|
314
|
+
}
|
|
299
315
|
function _fileOf(task: { file?: { filepath?: string; name?: string } } | undefined): string | null {
|
|
300
316
|
const f = task?.file?.filepath ?? task?.file?.name;
|
|
301
317
|
return typeof f === 'string' && f.length > 0 ? basename(f) : null;
|
|
@@ -414,6 +430,9 @@ afterEach(({ task }) => {
|
|
|
414
430
|
if (file === null) return;
|
|
415
431
|
if (!_ledgerMarks.has(file)) _ledgerMarks.set(file, 0);
|
|
416
432
|
const state = task.result?.state;
|
|
433
|
+
const seen = _fileRecorded.get(file) ?? new Set<string>();
|
|
434
|
+
seen.add(task.id);
|
|
435
|
+
_fileRecorded.set(file, seen);
|
|
417
436
|
const arr = _fileStates.get(file) ?? [];
|
|
418
437
|
arr.push(state === 'pass' ? 'pass' : state === 'fail' ? 'fail' : 'skip');
|
|
419
438
|
_fileStates.set(file, arr);
|
|
@@ -506,7 +525,10 @@ afterAll(({}, suite) => {
|
|
|
506
525
|
const s = suite as unknown as { filepath?: string; name?: string; file?: { filepath?: string; name?: string } };
|
|
507
526
|
const file = _fileOf({ file: s.file ?? s });
|
|
508
527
|
if (file === null || !file.endsWith('.test.ts')) return;
|
|
509
|
-
|
|
528
|
+
// A scenario hook that threw before the recorder ran hid its test; recover it from vitest's own results.
|
|
529
|
+
const hidden = unrecordedTests(_finishedTests(s.file ?? s), _fileRecorded.get(file) ?? new Set());
|
|
530
|
+
const states = [...(_fileStates.get(file) ?? []), ...hidden.states];
|
|
531
|
+
if (hidden.failures.length > 0) _fileFailures.set(file, [...(_fileFailures.get(file) ?? []), ...hidden.failures]);
|
|
510
532
|
// Did a behaviorGate in this file record inapplicable/skipped for its profile?
|
|
511
533
|
const mark = _ledgerMarks.get(file) ?? 0;
|
|
512
534
|
const since = journalSince(mark);
|
|
@@ -550,6 +572,7 @@ afterAll(({}, suite) => {
|
|
|
550
572
|
_fileStates.delete(file);
|
|
551
573
|
_fileAssertions.delete(file);
|
|
552
574
|
_fileFailures.delete(file);
|
|
575
|
+
_fileRecorded.delete(file);
|
|
553
576
|
_ledgerMarks.delete(file);
|
|
554
577
|
_itAllocators.delete(file);
|
|
555
578
|
_itMarks.delete(file);
|