@openwop/openwop-conformance 2.42.0 → 2.42.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +1 -1
- package/coverage.md +1 -0
- package/dist/lib/certification-bundle-verify.js +13 -0
- package/dist/spec-artifacts.lock.json +2 -2
- package/package.json +2 -2
- package/requirements.json +25 -19
- package/schemas/CORPUS-STAMP.json +16 -16
- package/src/lib/carries-payload.ts +42 -0
- package/src/lib/certification-bundle-verify.ts +13 -0
- package/src/lib/effect-refire.ts +34 -0
- package/src/lib/polling.ts +21 -0
- package/src/scenarios/certification-bundle-redaction.test.ts +1 -1
- package/src/scenarios/context-budget-transcript-bound.test.ts +3 -3
- package/src/scenarios/context-summarization-replay.test.ts +4 -4
- package/src/scenarios/v2-a2ui-v09-surface.test.ts +8 -12
- package/src/scenarios/v2-effect-identity-business-key.test.ts +17 -3
- package/src/scenarios/v2-effect-seam-no-refire.test.ts +18 -5
- package/src/scenarios/v2-mcp-mount-map.test.ts +20 -4
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,19 @@
|
|
|
1
1
|
# `@openwop/openwop-conformance` Changelog
|
|
2
2
|
|
|
3
|
+
## [2.42.2] — 2026-09-26 — four conformance legs now measure what their spec says: effect re-fire, keying fallback, A2UI JSON values, digest-safe scrubbing
|
|
4
|
+
|
|
5
|
+
- **`0173.effect-seam-no-refire` can detect a re-fire.** The leg asserted that the replay fork's ledger count was at most the parent's. The seam fires one attempt on the source run, so a host that re-fires on replay read 1 ≤ 1 and passed, over exactly the defect the leg exists for. ws3 measured it: openwop-app with `sideEffecting` dropped still passed. The leg now fails if the fork records any attempt the parent never made, using the new `lib/effect-refire.ts`. An inherited row repeats a parent attempt's `(nodeId, attempt, at)`, while a re-fire is a new attempt, and one parent row covers at most one fork row. A host whose fork projection is empty passes, as does one that projects the parent's attempts as history. The `effect-ledger-projection` schema description, which said the scenario "compares row COUNTS", is corrected. New suite self-test `src/lib/effect-refire.test.ts` uses a re-firing host stub as its first leg. It is sabotage-proved: with the old count rule restored, that leg fails.
|
|
6
|
+
- **RFC 0209 §C.11 compares JSON values, not JSON text, so a JSONB host can pass.** In `v2-a2ui-v09-surface`, `carries()` compared `JSON.stringify(subtree) === JSON.stringify(sent)`, which is order-sensitive. RFC 8259 §4 gives object member order no meaning, and a Postgres JSONB host re-sorts keys. ws5 measured `{version,catalogId,surfaceId,messages}` coming back as `{version,messages,catalogId,surfaceId}` on pg16. So `recorded-as-recorded` and `0209.legacy-readable` could never pass on a conforming JSONB host. The comparison now uses RFC 8785 (JCS) forms, via the new `lib/carries-payload.ts` built on `lib/jcs.ts`. A changed value, an added or dropped member, and reordered messages still fail. New suite self-test `src/lib/carries-payload.test.ts` covers the measured reordering as the positive leg and four negative controls. It is sabotage-proved: with the text comparison restored, the reordering leg fails. RFC 0209 carries a dated correction note. A sweep of every scenario for `JSON.stringify` equality found no other comparison of a suite-sent payload with a host-returned one. The rest compare two answers from the same host and route (`tools/list` twice, source vs fork, `v1` vs `v2` discovery), or arrays of numbers.
|
|
7
|
+
- **The 0173 retry leg accepts the keying the spec allows.** `v2-effect-identity-business-key` demanded `keying: business-identity` on every retried attempt. `idempotency.md` §Layer 2 makes the activity recipe "the fallback for a provider with no business key", and the seam's effect (a POST to a suite-chosen `providerUrl`) has none. So a host honestly declaring `activity-recipe` failed a rule the spec does not state. The leg now requires every attempt to declare a documented keying, the same one across attempts, and the same provider key. The last check was already there and is unchanged. Raised by openwop-app, whose seam began writing one ledger row per attempt.
|
|
8
|
+
- **`scrubEvidence` never rewrites inside a hex digest.** 2.40.3 stopped the environment sweep from collecting short and bare-integer values, after openwop-app's `OPENWOP_WEBHOOK_SECRET_ROTATION_OVERLAP_S=60` rewrote the "60" inside `discovery.sha256`. That failed the host's certify at random, roughly one bundle in five. But a credential the emitter hands over explicitly is still scrubbed whatever its shape, so the scrub itself could still corrupt a digest. A string that is wholly a hex digest (32+ hex characters, bare or `sha256:`-prefixed) is now left untouched unless it IS a secret. A secret inside a digest is a coincidence, not a leak, and rewriting it only breaks the evidence. A real leak in any other string is still redacted. Name classification needs no separate rule: a `*_S` / `*_MS` / `*_SECONDS` setting is a bare integer, which 2.40.3's floor already excludes. New suite self-test `src/lib/scrub-evidence-digest.test.ts` hands the scrub secrets that are substrings of a digest and asserts the digest survives while a neighbouring leak is redacted. It is sabotage-proved: without the guard, two of its three legs fail. **The 2.40.3 redaction leg was vacuous:** its fixture digest contained no "60", so it passed whether or not the value was scrubbed. The fixture now contains one.
|
|
9
|
+
- **Version moved ahead of publication.** `@openwop/openwop-conformance` and its exact-pinned peer `@openwop/spec-artifacts` move to `2.42.2` because `2.42.1` is tagged, and this cycle changes a shipped file: `spec/v1/gaps.json` records RFC 0111 gap G5 `closed` (RFC 0111 `Accepted`, witnessed by MyndHyve on 2.42.1). No scenario, schema or `MUST` change. Not tagged, not published.
|
|
10
|
+
|
|
11
|
+
## [2.42.1] — 2026-09-26 — RFC 0111's live witness scenarios outlast a real host, and the MCP run-transport leg identifies its own run
|
|
12
|
+
|
|
13
|
+
- **RFC 0111's live scenarios no longer die at vitest's 30 s default.** `context-budget-transcript-bound` and `context-summarization-replay` drive the live-model fixture `conformance-context-budget-live` (six real model turns and six child runs; 30–40 s per run on MyndHyve production, 2026-09-26), and the global `testTimeout: 30_000` killed both before any assertion — which `OPENWOP_POLL_TIMEOUT_SCALE` could not reach. Each now carries a per-test timeout from `liveScenarioTimeoutMs(runs)` (`src/lib/polling.ts`: the scaled sum of `LIVE_RUN_POLL_MS` = 180 s per run plus 60 s for seam reads — 240 s / 420 s at scale 1) and polls with `{ timeoutMs: LIVE_RUN_POLL_MS }`, so the named poll deadline always fires before the test deadline, at every scale. The global timeout is unchanged. Self-test: `src/lib/polling.test.ts` (3 new); an unscaled helper fails 2 of them and a 30 s helper 3.
|
|
14
|
+
- **`v2-mcp-mount-map`'s run-transport leg identifies its own run.** It diffed `listRuns` around one `tools/call` and required exactly one new `conformance-noop` run, but the suite runs files concurrently and that fixture is every file's smallest run: the v2 reference host's CI (4 workers) failed "got 2 new run(s)" on a host whose `tools/call` started exactly one, and `fresh[0]` could have been a sibling's run read for the wrong transport. The leg now requires at least one new run and reads the transport of the run the result names (when `listRuns` shows it), else the only new run, else passes if any new run in the window started `mcp`.
|
|
15
|
+
- **`spec/v2/core/webhooks.md`'s `Stable` banner cites RFC 0217**, now `Accepted` on the v2 reference host's certified 2.42.0 cut. No scenario change.
|
|
16
|
+
|
|
3
17
|
## [2.42.0] — 2026-09-26 — `replay_context_summary_unavailable` is a registered code, and an unregistered subscription has no sink
|
|
4
18
|
|
|
5
19
|
- **RFC 0217: after unregister, the dead-letter read answers `404 not_found`, as for a subscription that never existed.** `v2-webhook-durable-delivery` gains `openwop.requirement.0217.dead-letter-read-after-unregister`, gated on `webhooks.deadLetter`: read the sink (`200`, the control), unregister (`204`), read again, and read a never-minted same-tenant id; both later reads must be `404 not_found`. Closes RFC 0215's gap G5. Sabotage: a v2 reference host whose read answers `200 { deliveries: [] }` for an unknown id fails the row.
|
package/README.md
CHANGED
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
# --legacy-peer-deps is REQUIRED, not optional: the exact peer pin is what npm's
|
|
12
12
|
# default resolver refuses. npm 10.9 fails outright with
|
|
13
13
|
# "Cannot read properties of null (reading 'edgesOut')" — use npm >= 11.
|
|
14
|
-
npm install --legacy-peer-deps @openwop/openwop-conformance@2.42.
|
|
14
|
+
npm install --legacy-peer-deps @openwop/openwop-conformance@2.42.2 @openwop/spec-artifacts@2.42.2
|
|
15
15
|
# or run without install:
|
|
16
16
|
npx @openwop/openwop-conformance --base-url https://api.example.com --api-key hk_test_...
|
|
17
17
|
```
|
package/coverage.md
CHANGED
|
@@ -269,6 +269,7 @@ Every OpenAPI operation should have:
|
|
|
269
269
|
| `putTestPackTarball`, `getTestPackTarball`, `deleteTestPackVersion`, `getTestPackSignature` | `pack-registry-publish.test.ts` covers the 19-code publish error catalog through the RFC 0025 `/v1/packs-test/*` mirror namespace, gated on `capabilities.packs.testMode.supported: true` (RFC 0025 §A). 26 scenarios soft-skip when the advertisement is absent; when present, the suite exercises URL/scope, body-shape, tarball-extraction, manifest-contents, integrity, auth/conflict, unpublish-window, and signature-endpoint pairing. | Soft-skip on advertisement absence; behavioral on advertisement presence | Add real-tarball-builder fixtures so the manifest_mismatch / pack_integrity_failure / unsupported_runtime branches assert against a meaningful gzip+tar payload (currently soft-skipped with explanatory comments). |
|
|
270
270
|
| `getContentPage` | `localized-content-delivery.test.ts` (RFC 0103 §D) — always-on `localized-content-page-response.schema.json` shape + the §C `resolveSection` merge; gated live legs assert negotiated delivery + `Content-Language`. Public (anonymous-capable) endpoint, `security: []`. | Gated behavioral legs: published-only delivery, tenant isolation, no cross-tenant enumeration (`content-no-cross-tenant-enumeration`); cross-tenant slug → same `404`. | Behavioral legs soft-skip without a live `GET /v1/content/pages/{slug}`; non-vacuous against openwop-app + MyndHyve. |
|
|
271
271
|
| `listContentPages`, `createContentPage` | `localized-content-delivery.test.ts` (RFC 0103 §D) — `localized-content-page.schema.json` shape + §A coherence. | Protected admin ops document `401`/`403`; create documents `400`. Tenant-scoped (`content-response-tenant-scoped`). | Behavioral list/create round-trip pending a host admin seam. |
|
|
272
|
+
| `deleteContentPage` | No leg asserts it. `v2-content-locale-keys.test.ts` calls it only as best-effort cleanup, and its result is not asserted (RFC 0103 §D; added to the contract 2026-09-26). | `401`/`403`/`404`; `204` with no body. The same `404` for a `pageId` outside the caller's tenant (§F). | Uncovered. A behavioral delete (the page, its sections and its overlays are gone, and delivery answers `404`) is pending a host admin seam. The other four §D admin operations have no contract yet (RFC 0103 gap G12). |
|
|
272
273
|
| `putContentSection` | `localized-content-delivery.test.ts` (RFC 0103 §D) — `localized-content-section.schema.json` shape + the locale-targeted write body. | `401`/`403`/`404`/`400`; base-locale key + bad-key-case negatives. | Behavioral locale-targeted upsert pending a host admin seam. |
|
|
273
274
|
| `getContentSettings`, `putContentSettings` | `localized-content-delivery.test.ts` (RFC 0103 §B) — `localized-content-language-settings.schema.json` shape (base ∉ supported). | `401`/`403`; put documents `400` for the base ∈ supported violation. | Behavioral settings round-trip pending a host admin seam. |
|
|
274
275
|
|
|
@@ -208,6 +208,8 @@ export function verifyBundleV2(bundle) {
|
|
|
208
208
|
export function redactionMarker(secret) {
|
|
209
209
|
return `«redacted:${createHash('sha256').update(secret).digest('hex').slice(0, 12)}»`;
|
|
210
210
|
}
|
|
211
|
+
/** A whole-string hex digest: `sha256:`-prefixed or bare, 32+ hex characters. */
|
|
212
|
+
const HEX_DIGEST = /^(sha256:)?[0-9a-f]{32,}$/i;
|
|
211
213
|
/**
|
|
212
214
|
* Replace every occurrence of every secret in every string of `value` (keys
|
|
213
215
|
* included) with `redactionMarker(secret)`. Empty / whitespace-only secrets
|
|
@@ -223,6 +225,17 @@ export function scrubEvidence(value, secrets) {
|
|
|
223
225
|
if (live.length === 0)
|
|
224
226
|
return { value, redactedAt };
|
|
225
227
|
const scrubString = (s, path) => {
|
|
228
|
+
// A string that is wholly a hex digest (a sha256, a key fingerprint) is
|
|
229
|
+
// scrubbed only when it IS a secret, never for a secret found inside it.
|
|
230
|
+
// A secret inside a digest is a coincidence, not a leak: the digest does
|
|
231
|
+
// not disclose it. Rewriting it only corrupts the evidence. Suite 2.40.3
|
|
232
|
+
// floored the env sweep after openwop-app's `…_ROTATION_OVERLAP_S=60`
|
|
233
|
+
// rewrote the "60" inside `discovery.sha256`, but an explicitly handed
|
|
234
|
+
// credential is still scrubbed whatever its shape, so a short one (or any
|
|
235
|
+
// hex one) could still hit a digest at random. This closes that class at
|
|
236
|
+
// the scrub instead of at each source of secrets.
|
|
237
|
+
if (HEX_DIGEST.test(s) && !live.includes(s))
|
|
238
|
+
return s;
|
|
226
239
|
let out = s;
|
|
227
240
|
let hit = false;
|
|
228
241
|
for (const secret of live) {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@openwop/openwop-conformance",
|
|
3
|
-
"version": "2.42.
|
|
3
|
+
"version": "2.42.2",
|
|
4
4
|
"description": "Production-ready black-box conformance suite for OpenWOP v1.0 compliant servers.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -56,6 +56,6 @@
|
|
|
56
56
|
"@openwop/spec-artifacts": "file:../spec-artifacts"
|
|
57
57
|
},
|
|
58
58
|
"peerDependencies": {
|
|
59
|
-
"@openwop/spec-artifacts": "2.42.
|
|
59
|
+
"@openwop/spec-artifacts": "2.42.2"
|
|
60
60
|
}
|
|
61
61
|
}
|
package/requirements.json
CHANGED
|
@@ -27246,7 +27246,7 @@
|
|
|
27246
27246
|
{
|
|
27247
27247
|
"id": "openwop.it.v2-a2ui-v09-surface.a-version-2-envelope-carrying-a-version-1-body-is-refused-the-positive-is-admitt",
|
|
27248
27248
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27249
|
-
"line":
|
|
27249
|
+
"line": 131,
|
|
27250
27250
|
"title": "a version-2 envelope carrying a version-1 body is refused; the positive is admitted; above the floor is refused; below it follows envelopeStrictness",
|
|
27251
27251
|
"explicitId": "openwop.requirement.0209.version-selects-branch",
|
|
27252
27252
|
"citations": [
|
|
@@ -27290,7 +27290,7 @@
|
|
|
27290
27290
|
{
|
|
27291
27291
|
"id": "openwop.it.v2-a2ui-v09-surface.no-first-createsurface-a-message-after-deletesurface-and-a-second-createsurface",
|
|
27292
27292
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27293
|
-
"line":
|
|
27293
|
+
"line": 157,
|
|
27294
27294
|
"title": "no first createSurface, a message after deleteSurface, and a second createSurface are each refused; the valid steps are admitted",
|
|
27295
27295
|
"explicitId": "openwop.requirement.0209.fold-guarded",
|
|
27296
27296
|
"citations": [
|
|
@@ -27304,7 +27304,7 @@
|
|
|
27304
27304
|
{
|
|
27305
27305
|
"id": "openwop.it.v2-a2ui-v09-surface.a-createsurface-catalogid-differing-from-the-payload-catalogid-is-refused",
|
|
27306
27306
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27307
|
-
"line":
|
|
27307
|
+
"line": 182,
|
|
27308
27308
|
"title": "a createSurface.catalogId differing from the payload catalogId is refused",
|
|
27309
27309
|
"explicitId": "openwop.requirement.0209.catalog-equality",
|
|
27310
27310
|
"citations": [
|
|
@@ -27322,7 +27322,7 @@
|
|
|
27322
27322
|
{
|
|
27323
27323
|
"id": "openwop.it.v2-a2ui-v09-surface.a-message-whose-surfaceid-differs-from-the-payload-surfaceid-is-refused",
|
|
27324
27324
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27325
|
-
"line":
|
|
27325
|
+
"line": 196,
|
|
27326
27326
|
"title": "a message whose surfaceId differs from the payload surfaceId is refused",
|
|
27327
27327
|
"explicitId": "openwop.it.v2-a2ui-v09-surface.surface-id-equality",
|
|
27328
27328
|
"citations": [
|
|
@@ -27340,7 +27340,7 @@
|
|
|
27340
27340
|
{
|
|
27341
27341
|
"id": "openwop.it.v2-a2ui-v09-surface.a-recorded-version-2-surface-is-returned-byte-equal-on-poll-and-carried-by-a-for",
|
|
27342
27342
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27343
|
-
"line":
|
|
27343
|
+
"line": 214,
|
|
27344
27344
|
"title": "a recorded version-2 surface is returned byte-equal on poll and carried by a fork at a later fromSeq",
|
|
27345
27345
|
"explicitId": "openwop.it.v2-a2ui-v09-surface.recorded-as-recorded",
|
|
27346
27346
|
"citations": [
|
|
@@ -27351,7 +27351,7 @@
|
|
|
27351
27351
|
},
|
|
27352
27352
|
{
|
|
27353
27353
|
"section": "RFC 0209 §C.11",
|
|
27354
|
-
"requirement": "the run's poll MUST return the recorded surface payload
|
|
27354
|
+
"requirement": "the run's poll MUST return the recorded surface payload unchanged as JSON (JCS-equal; member order carries no meaning) — no surface is regenerated"
|
|
27355
27355
|
},
|
|
27356
27356
|
{
|
|
27357
27357
|
"section": "RFC 0209 §C.9",
|
|
@@ -27377,7 +27377,7 @@
|
|
|
27377
27377
|
{
|
|
27378
27378
|
"id": "openwop.it.v2-a2ui-v09-surface.a-version-1-surface-recorded-at-floor-1-is-returned-byte-equal-on-poll-and-carri",
|
|
27379
27379
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27380
|
-
"line":
|
|
27380
|
+
"line": 241,
|
|
27381
27381
|
"title": "a version-1 surface recorded at floor 1 is returned byte-equal on poll and carried by a fork",
|
|
27382
27382
|
"explicitId": "openwop.requirement.0209.legacy-readable",
|
|
27383
27383
|
"citations": [
|
|
@@ -27388,7 +27388,7 @@
|
|
|
27388
27388
|
},
|
|
27389
27389
|
{
|
|
27390
27390
|
"section": "RFC 0209 §C.11",
|
|
27391
|
-
"requirement": "the poll MUST return the version-1 surface
|
|
27391
|
+
"requirement": "the poll MUST return the version-1 surface unchanged as JSON (JCS-equal)"
|
|
27392
27392
|
},
|
|
27393
27393
|
{
|
|
27394
27394
|
"section": "RFC 0209 §C.11",
|
|
@@ -27414,7 +27414,7 @@
|
|
|
27414
27414
|
{
|
|
27415
27415
|
"id": "openwop.it.v2-a2ui-v09-surface.one-untrusted-update-taints-a-trusted-surface-a-later-trusted-update-does-not-la",
|
|
27416
27416
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27417
|
-
"line":
|
|
27417
|
+
"line": 265,
|
|
27418
27418
|
"title": "one untrusted update taints a trusted surface, a later trusted update does not launder it, and the approval does not advance; an all-trusted surface does",
|
|
27419
27419
|
"explicitId": "openwop.requirement.0209.taint-sticky",
|
|
27420
27420
|
"citations": [
|
|
@@ -27428,7 +27428,7 @@
|
|
|
27428
27428
|
{
|
|
27429
27429
|
"id": "openwop.it.v2-a2ui-v09-surface.records-why-render-needs-root-has-no-server-side-witness-and-where-its-witness-l",
|
|
27430
27430
|
"file": "v2-a2ui-v09-surface.test.ts",
|
|
27431
|
-
"line":
|
|
27431
|
+
"line": 304,
|
|
27432
27432
|
"title": "records why render-needs-root has no server-side witness, and where its witness lives",
|
|
27433
27433
|
"explicitId": "openwop.requirement.0209.render-needs-root",
|
|
27434
27434
|
"citations": [
|
|
@@ -28814,6 +28814,11 @@
|
|
|
28814
28814
|
"requirement": null,
|
|
28815
28815
|
"interpolated": true
|
|
28816
28816
|
},
|
|
28817
|
+
{
|
|
28818
|
+
"section": "spec/v2/core/idempotency.md §Layer 2: effect identity",
|
|
28819
|
+
"requirement": null,
|
|
28820
|
+
"interpolated": true
|
|
28821
|
+
},
|
|
28817
28822
|
{
|
|
28818
28823
|
"section": "spec/v2/core/idempotency.md §Layer 2: effect identity",
|
|
28819
28824
|
"requirement": null,
|
|
@@ -28889,7 +28894,7 @@
|
|
|
28889
28894
|
{
|
|
28890
28895
|
"id": "openwop.it.v2-effect-seam-no-refire.driving-one-seam-of-each-kind-observes-no-re-fire-on-replay",
|
|
28891
28896
|
"file": "v2-effect-seam-no-refire.test.ts",
|
|
28892
|
-
"line":
|
|
28897
|
+
"line": 44,
|
|
28893
28898
|
"title": "driving one seam of each kind observes no re-fire on replay",
|
|
28894
28899
|
"explicitId": "openwop.requirement.0173.effect-seam-no-refire",
|
|
28895
28900
|
"citations": [
|
|
@@ -30273,14 +30278,15 @@
|
|
|
30273
30278
|
},
|
|
30274
30279
|
{
|
|
30275
30280
|
"section": "interop-map.json mcp.methods tools/call; runs.md run.started",
|
|
30276
|
-
"requirement":
|
|
30281
|
+
"requirement": null,
|
|
30282
|
+
"interpolated": true
|
|
30277
30283
|
}
|
|
30278
30284
|
]
|
|
30279
30285
|
},
|
|
30280
30286
|
{
|
|
30281
30287
|
"id": "openwop.it.v2-mcp-mount-map.a-suspending-tool-answers-inputrequiredresult-the-retry-resolves-it-requeststate",
|
|
30282
30288
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30283
|
-
"line":
|
|
30289
|
+
"line": 197,
|
|
30284
30290
|
"title": "a suspending tool answers InputRequiredResult; the retry resolves it; requestState is single use and forgery-proof",
|
|
30285
30291
|
"explicitId": null,
|
|
30286
30292
|
"citations": [
|
|
@@ -30322,7 +30328,7 @@
|
|
|
30322
30328
|
{
|
|
30323
30329
|
"id": "openwop.it.v2-mcp-mount-map.the-mrtr-input-request-key-is-the-open-interrupt-s-interruptid-never-the-node-it",
|
|
30324
30330
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30325
|
-
"line":
|
|
30331
|
+
"line": 222,
|
|
30326
30332
|
"title": "the MRTR input-request key is the open interrupt’s interruptId, never the node it suspended on",
|
|
30327
30333
|
"explicitId": null,
|
|
30328
30334
|
"citations": [
|
|
@@ -30361,7 +30367,7 @@
|
|
|
30361
30367
|
{
|
|
30362
30368
|
"id": "openwop.it.v2-mcp-mount-map.a-list-that-differs-per-caller-is-cachescope-private",
|
|
30363
30369
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30364
|
-
"line":
|
|
30370
|
+
"line": 260,
|
|
30365
30371
|
"title": "a list that differs per caller is cacheScope private",
|
|
30366
30372
|
"explicitId": null,
|
|
30367
30373
|
"citations": [
|
|
@@ -30379,7 +30385,7 @@
|
|
|
30379
30385
|
{
|
|
30380
30386
|
"id": "openwop.it.v2-mcp-mount-map.unknown-meta-extension-keys-and-capabilities-extensions-are-opaque-processed-nor",
|
|
30381
30387
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30382
|
-
"line":
|
|
30388
|
+
"line": 277,
|
|
30383
30389
|
"title": "unknown _meta extension keys and capabilities.extensions are opaque: processed normally, granting nothing",
|
|
30384
30390
|
"explicitId": null,
|
|
30385
30391
|
"citations": [
|
|
@@ -30397,7 +30403,7 @@
|
|
|
30397
30403
|
{
|
|
30398
30404
|
"id": "openwop.it.v2-mcp-mount-map.an-unauthenticated-request-is-refused-at-the-boundary-unless-anonymousactor-is-a",
|
|
30399
30405
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30400
|
-
"line":
|
|
30406
|
+
"line": 286,
|
|
30401
30407
|
"title": "an unauthenticated request is refused at the boundary unless anonymousActor is advertised",
|
|
30402
30408
|
"explicitId": null,
|
|
30403
30409
|
"citations": [
|
|
@@ -30420,7 +30426,7 @@
|
|
|
30420
30426
|
{
|
|
30421
30427
|
"id": "openwop.it.v2-mcp-mount-map.a-credential-interrupt-is-answered-in-url-mode-url-connecturl-iserror-without-ur",
|
|
30422
30428
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30423
|
-
"line":
|
|
30429
|
+
"line": 303,
|
|
30424
30430
|
"title": "a credential interrupt is answered in URL mode (url = connectUrl), isError without URL support, and input_required again on an accept retry with no credential",
|
|
30425
30431
|
"explicitId": "openwop.requirement.0199.mcp-url-mode",
|
|
30426
30432
|
"citations": [
|
|
@@ -30458,7 +30464,7 @@
|
|
|
30458
30464
|
{
|
|
30459
30465
|
"id": "openwop.it.v2-mcp-mount-map.form-mode-is-never-emitted-for-a-nested-schema-or-a-sensitive-format-password-fi",
|
|
30460
30466
|
"file": "v2-mcp-mount-map.test.ts",
|
|
30461
|
-
"line":
|
|
30467
|
+
"line": 344,
|
|
30462
30468
|
"title": "form mode is never emitted for a nested schema or a sensitive (format password) field",
|
|
30463
30469
|
"explicitId": "openwop.requirement.0199.form-mode-no-secret",
|
|
30464
30470
|
"citations": [
|
|
@@ -1,17 +1,17 @@
|
|
|
1
1
|
{
|
|
2
2
|
"_comment": "Provenance of @openwop/spec-artifacts (RFC 0168 §D.2). files: SHA-256 per file; the conformance suite compares the installed peer against dist/spec-artifacts.lock.json at start.",
|
|
3
3
|
"package": "@openwop/spec-artifacts",
|
|
4
|
-
"version": "2.42.
|
|
5
|
-
"corpusTag": "v2.42.
|
|
4
|
+
"version": "2.42.2",
|
|
5
|
+
"corpusTag": "v2.42.2",
|
|
6
6
|
"files": {
|
|
7
7
|
"api/.redocly.lint-ignore.yaml": "bf5a8350b88a72fa43f59605ed8d903ed24b6cfccda5e45509c9f6ed9ee4e712",
|
|
8
8
|
"api/asyncapi.yaml": "d5ecb9ee6114582be3b1f662c84bfac9ae96dae7bacb853e461168f70a8e1c7d",
|
|
9
9
|
"api/grpc/openwop.proto": "c3e72bb17cba514ee98feb6434e6c9b6ea6795bfd086489ec69fd882dd1ad977",
|
|
10
|
-
"api/openapi.yaml": "
|
|
10
|
+
"api/openapi.yaml": "69464bc0e1e71ef3b543a73d4fa4f67202343d2603c80af8033f6c3619ef1b30",
|
|
11
11
|
"api/redocly.yaml": "b0604c89b2ca6d5076ec25725c539dad44a741a811fe524439ee6daef8baa09f",
|
|
12
|
-
"api/seams-v2.yaml": "
|
|
13
|
-
"api/v2/asyncapi.yaml": "
|
|
14
|
-
"api/v2/openapi.yaml": "
|
|
12
|
+
"api/seams-v2.yaml": "0777cffd47bdaad6482bc910064a42f793d74ae8e8921c84f498b1670f254f21",
|
|
13
|
+
"api/v2/asyncapi.yaml": "b87e6d39591f6967c6a6154657863c4573dd9833e7b6b8f4ca3472e9d92c9352",
|
|
14
|
+
"api/v2/openapi.yaml": "12ee6cf90544a49863f9238ba8ddd2d1333f913f26e122819e399ec65e6b682b",
|
|
15
15
|
"api/v2/redocly.yaml": "1e66b60e6118ad11a823bb620678be464d99dfe50a40e3e6f93ec9429b88b34c",
|
|
16
16
|
"schemas/README.md": "0c0b737ffcf8f30e7d2809cec8a498232de710f41443212922ad8337cdde0b51",
|
|
17
17
|
"schemas/a2a-task-state.schema.json": "75d5049dea8bd873ff0e7546f1c60c8d36c219bec7084264360be54a8be30ae1",
|
|
@@ -132,7 +132,7 @@
|
|
|
132
132
|
"schemas/v2/credential-reference.schema.json": "ea699ce30f58fb2c32d4279158b3f730f70be7b2dc8c3037693a76e66eb3e001",
|
|
133
133
|
"schemas/v2/debug-bundle.schema.json": "0bfa4b1d91fefbc12c95670dba9d5eacb72b879bb9a5c074e821ca9ed0fee6fd",
|
|
134
134
|
"schemas/v2/dispatch-config.schema.json": "5718c1ed7d95672e618bcb1f1da8e38180e144b12151307afc3aba1314dfcef0",
|
|
135
|
-
"schemas/v2/effect-ledger-projection.schema.json": "
|
|
135
|
+
"schemas/v2/effect-ledger-projection.schema.json": "0d118d385a43921118a76888449c5c3c5618eab1b80914b84a0a1b48d18822c1",
|
|
136
136
|
"schemas/v2/effect-seam-manifest.schema.json": "197ef9944ab027a50d1b9352f317cbbc12c8ad955fdb14650fb90477052ffef0",
|
|
137
137
|
"schemas/v2/envelopes/clarification.request.schema.json": "9e39d7b09f0b3ce7902d6aa19ca0eb52867b634efc78a881469c4d44fb3333e7",
|
|
138
138
|
"schemas/v2/envelopes/error.schema.json": "6a11fa3b9d61736fcb5675e6ac9825919aa609dceba13ef5615993d54647802b",
|
|
@@ -204,17 +204,17 @@
|
|
|
204
204
|
"schemas/workspace-file.schema.json": "464de85c2a068243084ee9c1d969bc7cd5d8f7948574e58450d6493c38a0e1e4",
|
|
205
205
|
"spec/v1/alias-detectors.json": "2401fcb1c18cdd688c018b3d716ae6bca85c5e356220ee2b793bd9d81872412d",
|
|
206
206
|
"spec/v1/capability-declaration-classes.json": "e7729aed5c4b4e1dd02abab0530f14cc95f5d4070fe51fb139e7f5cccefa00c6",
|
|
207
|
-
"spec/v1/core-standard-manifest.json": "
|
|
207
|
+
"spec/v1/core-standard-manifest.json": "c38981a02d063f68cffd5d22cf8e10390a3b12af4271e2819dc753711fc9616f",
|
|
208
208
|
"spec/v1/deprecations.json": "2d03f4729810280147ea08c630ea65434f2dc5370567d0aac41c3595f8337d40",
|
|
209
209
|
"spec/v1/deprecations.schema.json": "3e393c405d2a41b467d8c5e3c468549078df1ce6a6d2588fc488097b95d9b55b",
|
|
210
210
|
"spec/v1/event-codemap.json": "3da60d884157793a360da532a9fcbbfb5285636db325a74cec94b34622186d97",
|
|
211
211
|
"spec/v1/event-codemap.schema.json": "d05933b2e88103aff51a2f774df97b9bda065e0fc88dbd9b114cc47a32189174",
|
|
212
212
|
"spec/v1/extensions.json": "79a60754aa16cbdbdb604f8e5c00038af535a1690a4d07bd5ff7c38f0256a7bc",
|
|
213
|
-
"spec/v1/gaps.json": "
|
|
213
|
+
"spec/v1/gaps.json": "e99603a9c0e14f69fc3dd5ff6269f75a7b25aa2eb28051f9e2e1758fd08265b9",
|
|
214
214
|
"spec/v1/gaps.schema.json": "8fd83259f556553c9df0f53e7a82ca8c2a4771d471197f69b4e2ae0ceeacfd66",
|
|
215
215
|
"spec/v1/migrations.json": "8839497d6a4830b53fa7867014a8336f0d725ef2346352e54fb580b7947897af",
|
|
216
216
|
"spec/v1/migrations.schema.json": "886779aa6c22e646db097f5df210adb018a4dd14a7b815465a18c8a7056c8f72",
|
|
217
|
-
"spec/v1/operation-path-manifest.json": "
|
|
217
|
+
"spec/v1/operation-path-manifest.json": "fada0342815fb129c1ae61f1829bd6e34417717408ce86fcc54280f373c19e74",
|
|
218
218
|
"spec/v1/spec-gaps.json": "ea9612af55bea9dae3e82223d4f714243f8f590b0193f1188257b13491a38afc",
|
|
219
219
|
"spec/v2/README.md": "ccfed7d978ce7293a6d8f3c0f6045835dfb7e488f2b67a68f6b04d67c4b0cfc9",
|
|
220
220
|
"spec/v2/core/capabilities.md": "e22381fed3a93376ad478fa7f07571fdad9653aefd7c05f189a2ef45a49f6c63",
|
|
@@ -224,7 +224,7 @@
|
|
|
224
224
|
"spec/v2/core/errors.md": "870f1a9c9481fb49c8bd39170f0656c2ae46124afd4472fb7a3c18c67be0e561",
|
|
225
225
|
"spec/v2/core/events.md": "3df1866a89fb765c78adf79b8d51f96c59584dee9aa88d59179b88a681d827da",
|
|
226
226
|
"spec/v2/core/form-content-packs.md": "55ea021b5f676f5851f420b3597262fbf6642108bddeb726922f8c5dd96c968b",
|
|
227
|
-
"spec/v2/core/headers.md": "
|
|
227
|
+
"spec/v2/core/headers.md": "10ec6a6c3e0437b1501cb224273e86c2f6aede92fd469605f3353904d3ebd230",
|
|
228
228
|
"spec/v2/core/host-services.md": "c3ff93201d642edf835e9ae84f06a17edd207f9d351f8bfd70572be05bbfd894",
|
|
229
229
|
"spec/v2/core/idempotency.md": "b61d4a200526485937897dd7ebca092e68424a80624e0be7a0b7db8cdbd54403",
|
|
230
230
|
"spec/v2/core/identity.md": "d15667552019f64316be9b65c6e5696c10766581405f05da3907fb47f36e6459",
|
|
@@ -240,7 +240,7 @@
|
|
|
240
240
|
"spec/v2/core/security-defaults.md": "500471a8db7af9b776ac40b4a9a278c9f3d45e1c1dead880114e7fadc7bee37d",
|
|
241
241
|
"spec/v2/core/tool-catalog.md": "35f1a3fd509db0fc0490ddf400686fe94dc4b20c98064d5ebe776da516153e60",
|
|
242
242
|
"spec/v2/core/versioning.md": "0aded71d090c8358c3ce17763c120cfc1cc42531a5eaa1b8a2a7016ed9594069",
|
|
243
|
-
"spec/v2/core/webhooks.md": "
|
|
243
|
+
"spec/v2/core/webhooks.md": "78d7e237866bf1ee95bbfa9d5b62b619ca2f8ec677d1b752ba96ed5efcfc77ef",
|
|
244
244
|
"spec/v2/core/workflow-chain-packs.md": "ee45d3fede3bb6f0cc6abcf0d6c929c7cd213a0fbe862b738a998f929c4d1102",
|
|
245
245
|
"spec/v2/corrections.json": "cc74ffde74384b0f261a4f625871d922e8661d654d647a8534e21036595cb280",
|
|
246
246
|
"spec/v2/corrections.schema.json": "4ac595b9a6d7f66d53de03fe0edbe0dc58de46426932387bfdc0ceb3cf949a3e",
|
|
@@ -291,12 +291,12 @@
|
|
|
291
291
|
"spec/v2/interop-map.schema.json": "0017c882df34ee215e2e0729de051490e09d43aac3d2c949285698ed51832f12",
|
|
292
292
|
"spec/v2/migrations.json": "651954938a1e154c0b6f1e3dd6e105889d30de3bb65789f21de7f77a12869e9e",
|
|
293
293
|
"spec/v2/migrations.schema.json": "b91ade6320ad1b1185c9495b9424f3aa63438ed8e412c45171c227d82c0afd86",
|
|
294
|
-
"spec/v2/path-manifest.json": "
|
|
294
|
+
"spec/v2/path-manifest.json": "17298a245c5da9a70aa7ab0b7f2db4277e06209d6fcdd7187a8b4cea179c08f3",
|
|
295
295
|
"spec/v2/peer-dependency-aliases.json": "d10299280abee08258502925bc327293ee413e0108cd6e6ec75ff6110653308d",
|
|
296
296
|
"spec/v2/profiles.json": "0636f19fceae625390003a347e70ef4797d84766b5c24ce8a02cea52aadebca4",
|
|
297
|
-
"spec/v2/release.json": "
|
|
297
|
+
"spec/v2/release.json": "f8aeec2581bb0b934b2667ac72680f2caf6259b2ca7312e716598e9302573966",
|
|
298
298
|
"spec/v2/retention-floors.json": "eaf3722d95c79947af1d4269ef85117e126518c588cfcf1a2b21b97269f51624",
|
|
299
|
-
"spec/v2/surface-baseline.json": "
|
|
299
|
+
"spec/v2/surface-baseline.json": "e2ce7ec18f217f08dd13d6ff1895ea4030b0a76965b774689fb6b52bb20b745a"
|
|
300
300
|
},
|
|
301
|
-
"corpusCommit": "
|
|
301
|
+
"corpusCommit": "10e78e04bc0dac6cb93a0ef00eda5dc7547b4ee8"
|
|
302
302
|
}
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Does any event's payload carry `sent`, as JSON — member order ignored?
|
|
3
|
+
*
|
|
4
|
+
* `v2-a2ui-v09-surface` asserted RFC 0209 §C.11 ("recorded surface envelopes
|
|
5
|
+
* are returned as recorded") by comparing `JSON.stringify(subtree)` with
|
|
6
|
+
* `JSON.stringify(sent)`. That comparison is ORDER-sensitive, and RFC 8259 §4
|
|
7
|
+
* gives object member order no meaning. A Postgres JSONB host returns every
|
|
8
|
+
* member and every value but re-sorts keys (shorter first, then bytewise):
|
|
9
|
+
* measured on pg16, `{version,catalogId,surfaceId,messages}` came back as
|
|
10
|
+
* `{version,messages,catalogId,surfaceId}`. So a conforming host could never
|
|
11
|
+
* pass the leg (suite 2.42.2 correction).
|
|
12
|
+
*
|
|
13
|
+
* The comparison is now between RFC 8785 (JCS) forms, the corpus's canonical
|
|
14
|
+
* JSON (`lib/jcs.ts`; conformance.md §"Canonical JSON"). JCS sorts members and
|
|
15
|
+
* fixes number and string serialization, so two values are JCS-equal exactly
|
|
16
|
+
* when they are the same JSON value. What §C.11 forbids, a regenerated or
|
|
17
|
+
* altered surface, still fails: a changed value, a dropped or added member, or
|
|
18
|
+
* a reordered ARRAY (message order is meaningful) all change the JCS form.
|
|
19
|
+
*/
|
|
20
|
+
|
|
21
|
+
import { canonicalJSON } from './jcs.js';
|
|
22
|
+
|
|
23
|
+
/** Every subtree of `v`, depth-first. */
|
|
24
|
+
function* subtrees(v: unknown): Generator<unknown> {
|
|
25
|
+
yield v;
|
|
26
|
+
if (Array.isArray(v)) for (const x of v) yield* subtrees(x);
|
|
27
|
+
else if (v && typeof v === 'object') for (const x of Object.values(v)) yield* subtrees(x);
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
/** JCS form, or null for a value JCS refuses (not I-JSON), which then matches nothing. */
|
|
31
|
+
function jcs(v: unknown): string | null {
|
|
32
|
+
try { return canonicalJSON(v); } catch { return null; }
|
|
33
|
+
}
|
|
34
|
+
|
|
35
|
+
export function carriesPayload(events: readonly unknown[], sent: unknown): boolean {
|
|
36
|
+
const want = jcs(sent);
|
|
37
|
+
if (want === null) return false;
|
|
38
|
+
return events.some((e) => {
|
|
39
|
+
const payload = e !== null && typeof e === 'object' ? (e as Record<string, unknown>)['payload'] : undefined;
|
|
40
|
+
return [...subtrees(payload)].some((s) => jcs(s) === want);
|
|
41
|
+
});
|
|
42
|
+
}
|
|
@@ -298,6 +298,9 @@ export interface ScrubResult<T> {
|
|
|
298
298
|
readonly redactedAt: readonly string[];
|
|
299
299
|
}
|
|
300
300
|
|
|
301
|
+
/** A whole-string hex digest: `sha256:`-prefixed or bare, 32+ hex characters. */
|
|
302
|
+
const HEX_DIGEST = /^(sha256:)?[0-9a-f]{32,}$/i;
|
|
303
|
+
|
|
301
304
|
/**
|
|
302
305
|
* Replace every occurrence of every secret in every string of `value` (keys
|
|
303
306
|
* included) with `redactionMarker(secret)`. Empty / whitespace-only secrets
|
|
@@ -312,6 +315,16 @@ export function scrubEvidence<T>(value: T, secrets: readonly string[]): ScrubRes
|
|
|
312
315
|
const redactedAt: string[] = [];
|
|
313
316
|
if (live.length === 0) return { value, redactedAt };
|
|
314
317
|
const scrubString = (s: string, path: string): string => {
|
|
318
|
+
// A string that is wholly a hex digest (a sha256, a key fingerprint) is
|
|
319
|
+
// scrubbed only when it IS a secret, never for a secret found inside it.
|
|
320
|
+
// A secret inside a digest is a coincidence, not a leak: the digest does
|
|
321
|
+
// not disclose it. Rewriting it only corrupts the evidence. Suite 2.40.3
|
|
322
|
+
// floored the env sweep after openwop-app's `…_ROTATION_OVERLAP_S=60`
|
|
323
|
+
// rewrote the "60" inside `discovery.sha256`, but an explicitly handed
|
|
324
|
+
// credential is still scrubbed whatever its shape, so a short one (or any
|
|
325
|
+
// hex one) could still hit a digest at random. This closes that class at
|
|
326
|
+
// the scrub instead of at each source of secrets.
|
|
327
|
+
if (HEX_DIGEST.test(s) && !live.includes(s)) return s;
|
|
315
328
|
let out = s;
|
|
316
329
|
let hit = false;
|
|
317
330
|
for (const secret of live) {
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Which rows of a replay fork's effect ledger record an attempt the PARENT
|
|
3
|
+
* never made, i.e. a re-fire (RFC 0173 §C.2; replay.md §Suppression rule 1)?
|
|
4
|
+
*
|
|
5
|
+
* `v2-effect-seam-no-refire` used to assert `fork count ≤ parent count`. The
|
|
6
|
+
* seam fires ONE attempt on the source run, so a host that re-fires on replay
|
|
7
|
+
* records one row on the fork and reads 1 ≤ 1: the leg passed over exactly the
|
|
8
|
+
* defect it exists for. Measured: openwop-app with `sideEffecting` dropped
|
|
9
|
+
* still passed (suite 2.42.2 correction).
|
|
10
|
+
*
|
|
11
|
+
* A fork's projection may legitimately carry nothing (a host that keys the
|
|
12
|
+
* ledger per run: its fork's own ledger is empty when nothing re-fired) or the
|
|
13
|
+
* parent's rows as fixed history (a host that projects inherited attempts). So
|
|
14
|
+
* the witness is not a count. It is whether any fork row records an attempt
|
|
15
|
+
* that is not one of the parent's: an inherited row repeats a parent attempt's
|
|
16
|
+
* `(nodeId, attempt, at)`, while a re-fire is a new attempt with a new `at`.
|
|
17
|
+
* The multiset match means N inherited rows cover at most N parent rows.
|
|
18
|
+
*/
|
|
19
|
+
|
|
20
|
+
export interface EffectRow { readonly nodeId?: unknown; readonly attempt?: unknown; readonly at?: unknown }
|
|
21
|
+
|
|
22
|
+
const key = (r: EffectRow): string => `${String(r.nodeId)}|${String(r.attempt)}|${String(r.at)}`;
|
|
23
|
+
|
|
24
|
+
export function refiredAttempts(parent: readonly EffectRow[], fork: readonly EffectRow[]): EffectRow[] {
|
|
25
|
+
const available = new Map<string, number>();
|
|
26
|
+
for (const r of parent) available.set(key(r), (available.get(key(r)) ?? 0) + 1);
|
|
27
|
+
const out: EffectRow[] = [];
|
|
28
|
+
for (const r of fork) {
|
|
29
|
+
const n = available.get(key(r)) ?? 0;
|
|
30
|
+
if (n > 0) available.set(key(r), n - 1);
|
|
31
|
+
else out.push(r);
|
|
32
|
+
}
|
|
33
|
+
return out;
|
|
34
|
+
}
|
package/src/lib/polling.ts
CHANGED
|
@@ -71,6 +71,27 @@ export function scaledTimeoutMs(timeoutMs: number): number {
|
|
|
71
71
|
return scale === 1 ? timeoutMs : Math.ceil(timeoutMs * scale);
|
|
72
72
|
}
|
|
73
73
|
|
|
74
|
+
/**
|
|
75
|
+
* Live-model scenarios (RFC 0111's `conformance-context-budget-live`) drive real
|
|
76
|
+
* model turns and child runs; one run takes 30–40 s on a production host
|
|
77
|
+
* (MyndHyve, 2026-09-26), so vitest's global 30 s `testTimeout` killed them
|
|
78
|
+
* before any assertion — a suite defect that looks like a host failure.
|
|
79
|
+
*
|
|
80
|
+
* `LIVE_RUN_POLL_MS` is the base bound for ONE live run to reach a terminal
|
|
81
|
+
* status (below the fixture's own `settings.timeout` of 300 s); pass it as
|
|
82
|
+
* `pollUntilTerminal(runId, { timeoutMs: LIVE_RUN_POLL_MS })` — `pollUntil`
|
|
83
|
+
* scales it. `liveScenarioTimeoutMs(runs)` is the per-test vitest timeout: the
|
|
84
|
+
* scaled sum of the scenario's poll bounds plus 60 s for its seam reads, so the
|
|
85
|
+
* poll deadline always fires first and a hung host fails with a named poll
|
|
86
|
+
* message, never a bare vitest timeout. Both scale with
|
|
87
|
+
* `OPENWOP_POLL_TIMEOUT_SCALE`, so their order holds at every scale.
|
|
88
|
+
*/
|
|
89
|
+
export const LIVE_RUN_POLL_MS = 180_000;
|
|
90
|
+
|
|
91
|
+
export function liveScenarioTimeoutMs(runs: number): number {
|
|
92
|
+
return scaledTimeoutMs(runs * LIVE_RUN_POLL_MS + 60_000);
|
|
93
|
+
}
|
|
94
|
+
|
|
74
95
|
const TERMINAL = new Set(['completed', 'failed', 'cancelled']);
|
|
75
96
|
|
|
76
97
|
export async function getRun(runId: string): Promise<RunSnapshot> {
|
|
@@ -158,7 +158,7 @@ describe('RFC 0148 §C — certification-bundle-redaction: secret canaries never
|
|
|
158
158
|
// Suite 2.40.2: openwop-app set OPENWOP_WEBHOOK_SECRET_ROTATION_OVERLAP_S=60,
|
|
159
159
|
// the name matched SECRET, and scrubEvidence rewrote the "60" inside
|
|
160
160
|
// discovery.sha256 — the bundle failed ^[0-9a-f]{64}$ and main could not deploy.
|
|
161
|
-
const digest = '
|
|
161
|
+
const digest = '4b8d99c002a061a0255e8224ee683160007febadd251b2d171b47c0849695315';
|
|
162
162
|
const env = {
|
|
163
163
|
OPENWOP_WEBHOOK_SECRET_ROTATION_OVERLAP_S: '60',
|
|
164
164
|
OPENWOP_TOKEN_TTL_SECONDS_KEY: '86400000',
|
|
@@ -55,7 +55,7 @@
|
|
|
55
55
|
|
|
56
56
|
import { describe, it, expect } from 'vitest';
|
|
57
57
|
import { driver } from '../lib/driver.js';
|
|
58
|
-
import { pollUntilTerminal } from '../lib/polling.js';
|
|
58
|
+
import { LIVE_RUN_POLL_MS, liveScenarioTimeoutMs, pollUntilTerminal } from '../lib/polling.js';
|
|
59
59
|
import { behaviorGate } from '../lib/behavior-gate.js';
|
|
60
60
|
import { isFixtureAdvertised } from '../lib/fixtures.js';
|
|
61
61
|
import { readCapabilityFamily } from '../lib/discovery-capabilities.js';
|
|
@@ -111,7 +111,7 @@ describe('context-budget-transcript-bound (RFC 0111 §"Context economy")', () =>
|
|
|
111
111
|
const runId = runIdOf(create.json);
|
|
112
112
|
expect(runId, req(ID, 'RFC 0111', 'the create response MUST carry a runId')).toBeDefined();
|
|
113
113
|
if (runId === undefined) return softSkip('blocked', 'no runId');
|
|
114
|
-
await pollUntilTerminal(runId);
|
|
114
|
+
await pollUntilTerminal(runId, { timeoutMs: LIVE_RUN_POLL_MS });
|
|
115
115
|
|
|
116
116
|
const windows: Array<{ iteration: number; window: TranscriptWindow }> = [];
|
|
117
117
|
for (let iteration = 1; iteration <= MAX_ITERATIONS_PROBED; iteration += 1) {
|
|
@@ -153,5 +153,5 @@ describe('context-budget-transcript-bound (RFC 0111 §"Context economy")', () =>
|
|
|
153
153
|
|
|
154
154
|
if (log === null) return softSkip('blocked', 'the run event-log seam is unavailable, so the real-event, recent-tail and pressure rules were not measured');
|
|
155
155
|
if (!pressure) softSkip('inapplicable', `no iteration shows budget pressure — every eligible event fit under transcriptTokenBudget ${budget}, so the bound was never exercised (a budget the run never reaches is not a witness)`);
|
|
156
|
-
});
|
|
156
|
+
}, liveScenarioTimeoutMs(1));
|
|
157
157
|
});
|
|
@@ -38,7 +38,7 @@
|
|
|
38
38
|
|
|
39
39
|
import { describe, it, expect } from 'vitest';
|
|
40
40
|
import { driver } from '../lib/driver.js';
|
|
41
|
-
import { pollUntilTerminal } from '../lib/polling.js';
|
|
41
|
+
import { LIVE_RUN_POLL_MS, liveScenarioTimeoutMs, pollUntilTerminal } from '../lib/polling.js';
|
|
42
42
|
import { behaviorGate } from '../lib/behavior-gate.js';
|
|
43
43
|
import { isFixtureAdvertised } from '../lib/fixtures.js';
|
|
44
44
|
import { readCapabilityFamily } from '../lib/discovery-capabilities.js';
|
|
@@ -125,7 +125,7 @@ describe('context-summarization-replay (RFC 0111 §"Replay determinism")', () =>
|
|
|
125
125
|
const sourceRunId = runIdOf(create.json);
|
|
126
126
|
expect(sourceRunId, req(ID, 'rest-endpoints.md POST /v1/runs', 'the create response MUST carry a runId')).toBeDefined();
|
|
127
127
|
if (sourceRunId === undefined) return softSkip('blocked', 'no runId');
|
|
128
|
-
await pollUntilTerminal(sourceRunId);
|
|
128
|
+
await pollUntilTerminal(sourceRunId, { timeoutMs: LIVE_RUN_POLL_MS });
|
|
129
129
|
|
|
130
130
|
const sourceQ = await queryTestEvents(sourceRunId);
|
|
131
131
|
if (!sourceQ.ok) return softSkip('blocked', 'the run event-log seam is unavailable');
|
|
@@ -141,7 +141,7 @@ describe('context-summarization-replay (RFC 0111 §"Replay determinism")', () =>
|
|
|
141
141
|
const forkRunId = runIdOf(fork.json);
|
|
142
142
|
expect(forkRunId, req(ID, 'rest-endpoints.md POST /v1/runs/{runId}:fork', 'replay fork MUST return a runId')).toBeDefined();
|
|
143
143
|
if (forkRunId === undefined) return softSkip('blocked', 'no fork runId');
|
|
144
|
-
await pollUntilTerminal(forkRunId);
|
|
144
|
+
await pollUntilTerminal(forkRunId, { timeoutMs: LIVE_RUN_POLL_MS });
|
|
145
145
|
|
|
146
146
|
const forkQ = await queryTestEvents(forkRunId);
|
|
147
147
|
if (!forkQ.ok) return softSkip('blocked', 'the event-log seam is unavailable for the fork');
|
|
@@ -153,5 +153,5 @@ describe('context-summarization-replay (RFC 0111 §"Replay determinism")', () =>
|
|
|
153
153
|
if (sourceTexts === null || forkTexts === null) return softSkip('inapplicable', 'the transcript-window seam serves no entries[] for these runs, so the model-facing summary text was not compared (summaryRef reuse was)');
|
|
154
154
|
expect(sourceTexts.length, req(ID, 'RFC 0111 §"Replay determinism"', 'the source run summarized, so its transcript windows MUST carry the summary text it fed')).toBeGreaterThan(0);
|
|
155
155
|
expect(forkTexts, req(ID, 'RFC 0111 §"Replay determinism"', 'the replay MUST feed the model the recorded summary text, byte for byte — never a re-summarization')).toEqual(sourceTexts);
|
|
156
|
-
});
|
|
156
|
+
}, liveScenarioTimeoutMs(2));
|
|
157
157
|
});
|
|
@@ -42,6 +42,7 @@ import { readErrorCode } from '../lib/error-envelope.js';
|
|
|
42
42
|
import { FIXTURES_DIR } from '../lib/paths.js';
|
|
43
43
|
import { softSkip } from '../lib/soft-skip.js';
|
|
44
44
|
import { req } from '../lib/requirement-ids.js';
|
|
45
|
+
import { carriesPayload } from '../lib/carries-payload.js';
|
|
45
46
|
|
|
46
47
|
type Json = Record<string, unknown>;
|
|
47
48
|
const KIND = 'ui.a2ui-surface';
|
|
@@ -117,16 +118,9 @@ async function gate(want: (floor: number) => boolean, wantText: string): Promise
|
|
|
117
118
|
}
|
|
118
119
|
async function done(runId: string): Promise<void> { await driver.post(`/runs/${encodeURIComponent(runId)}/cancel`, {}); }
|
|
119
120
|
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
if (Array.isArray(v)) for (const x of v) yield* subtrees(x);
|
|
124
|
-
else if (v && typeof v === 'object') for (const x of Object.values(v)) yield* subtrees(x);
|
|
125
|
-
}
|
|
126
|
-
const carries = (events: unknown[], payload: unknown): boolean => {
|
|
127
|
-
const want = JSON.stringify(payload);
|
|
128
|
-
return events.some((e) => [...subtrees((e as Json)['payload'])].some((s) => JSON.stringify(s) === want));
|
|
129
|
-
};
|
|
121
|
+
// RFC 0209 §C.11 compares as JSON, not as text: member order carries no meaning
|
|
122
|
+
// (RFC 8259 §4), and a JSONB host re-sorts it. See lib/carries-payload.ts.
|
|
123
|
+
const carries = carriesPayload;
|
|
130
124
|
async function eventsOf(runId: string): Promise<unknown[]> {
|
|
131
125
|
const r = await driver.get(`/runs/${encodeURIComponent(runId)}/events/poll?timeout=1`);
|
|
132
126
|
const ev = (r.json as { events?: unknown } | null)?.events;
|
|
@@ -215,6 +209,8 @@ describe('RFC 0209 §A.2 — the cross-field rules are checked at admission (sea
|
|
|
215
209
|
});
|
|
216
210
|
|
|
217
211
|
describe('RFC 0209 §C.11 — recorded surfaces are returned as recorded (seam-gated)', () => {
|
|
212
|
+
// "byte-equal" in these two titles means JCS-equal (lib/carries-payload.ts). The titles
|
|
213
|
+
// are kept because they mint the rows' title ids.
|
|
218
214
|
it('a recorded version-2 surface is returned byte-equal on poll and carried by a fork at a later fromSeq', async () => {
|
|
219
215
|
const g = await gate((f) => f >= 2, 'a floor of at least 2');
|
|
220
216
|
if (!g.ok) return softSkip(g.kind, g.reason);
|
|
@@ -225,7 +221,7 @@ describe('RFC 0209 §C.11 — recorded surfaces are returned as recorded (seam-g
|
|
|
225
221
|
expect(r.status, req('openwop.it.v2-a2ui-v09-surface.recorded-as-recorded', 'RFC 0209 §C.11', `the positive surface MUST be admitted, got ${code(r)}`)).toBe(201);
|
|
226
222
|
const sequence = (r.json as { sequence: number }).sequence;
|
|
227
223
|
const events = await eventsOf(runId);
|
|
228
|
-
expect(carries(events, payload), req('openwop.it.v2-a2ui-v09-surface.recorded-as-recorded', 'RFC 0209 §C.11', 'the run\'s poll MUST return the recorded surface payload
|
|
224
|
+
expect(carries(events, payload), req('openwop.it.v2-a2ui-v09-surface.recorded-as-recorded', 'RFC 0209 §C.11', 'the run\'s poll MUST return the recorded surface payload unchanged as JSON (JCS-equal; member order carries no meaning) — no surface is regenerated')).toBe(true);
|
|
229
225
|
// The surface is the last event of the suspended run, so `sequence + 1`
|
|
230
226
|
// names no event and runs.md §Fork REQUIRES 422 fork_point_invalid for it.
|
|
231
227
|
// Record one more envelope on the same surface and fork AT that event:
|
|
@@ -250,7 +246,7 @@ describe('RFC 0209 §C.11 — recorded surfaces are returned as recorded (seam-g
|
|
|
250
246
|
const r = await emit(runId, envelope(1, V1_SURFACE));
|
|
251
247
|
expect(r.status, req('openwop.requirement.0209.legacy-readable', 'RFC 0209 §C.11', `a version-1 surface at floor 1 MUST be admitted, got ${code(r)}`)).toBe(201);
|
|
252
248
|
const sequence = (r.json as { sequence: number }).sequence;
|
|
253
|
-
expect(carries(await eventsOf(runId), V1_SURFACE), req('openwop.requirement.0209.legacy-readable', 'RFC 0209 §C.11', 'the poll MUST return the version-1 surface
|
|
249
|
+
expect(carries(await eventsOf(runId), V1_SURFACE), req('openwop.requirement.0209.legacy-readable', 'RFC 0209 §C.11', 'the poll MUST return the version-1 surface unchanged as JSON (JCS-equal)')).toBe(true);
|
|
254
250
|
// As in the version-2 leg: fork at a recorded event after the surface, never at `sequence + 1` (no such event; 422 fork_point_invalid).
|
|
255
251
|
const later = await emit(runId, envelope(1, V1_LATER));
|
|
256
252
|
expect(later.status, req('openwop.requirement.0209.legacy-readable', 'RFC 0209 §C.11', `a second version-1 surface at floor 1 MUST be admitted (it supplies the fork point), got ${code(later)}`)).toBe(201);
|
|
@@ -122,11 +122,25 @@ describe('RFC 0173 §B — effect-identity-business-key (gated on idempotency)',
|
|
|
122
122
|
keys.size,
|
|
123
123
|
req('openwop.requirement.0173.effect-identity-business-key.retry', 'spec/v2/core/idempotency.md §Layer 2: effect identity', `every attempt of one effect MUST present the same provider key across a transport retry — ${attempts.length} attempt(s) presented ${keys.size} distinct key(s)`),
|
|
124
124
|
).toBe(1);
|
|
125
|
+
// idempotency.md §Layer 2: business identity is the rule, and the activity
|
|
126
|
+
// recipe is "the fallback for a provider with no business key". The seam's
|
|
127
|
+
// effect is a POST to a suite-chosen providerUrl, which has no business key,
|
|
128
|
+
// so `activity-recipe` is the CONFORMANT keying here. Until 2.42.2 this leg
|
|
129
|
+
// demanded `business-identity` on every attempt, which is stricter than the
|
|
130
|
+
// spec, and a host honestly declaring the fallback failed. What the spec does
|
|
131
|
+
// require of a retry is one effect, one key: every attempt carries a documented
|
|
132
|
+
// keying, and the SAME one. A host that switches modes between attempts has
|
|
133
|
+
// re-derived the effect's identity mid-flight.
|
|
134
|
+
const keyings = new Set(attempts.map((a) => String(a['keying'])));
|
|
125
135
|
for (const a of attempts) {
|
|
126
136
|
expect(
|
|
127
|
-
|
|
128
|
-
req('openwop.requirement.0173.effect-identity-business-key.retry', 'spec/v2/core/idempotency.md §Layer 2: effect identity', `
|
|
129
|
-
).
|
|
137
|
+
KEYING,
|
|
138
|
+
req('openwop.requirement.0173.effect-identity-business-key.retry', 'spec/v2/core/idempotency.md §Layer 2: effect identity', `every attempt MUST declare a documented keying — business-identity, or activity-recipe for a provider with no business key (attempt ${String(a['attempt'])} declares ${String(a['keying'])})`),
|
|
139
|
+
).toContain(a['keying']);
|
|
130
140
|
}
|
|
141
|
+
expect(
|
|
142
|
+
keyings.size,
|
|
143
|
+
req('openwop.requirement.0173.effect-identity-business-key.retry', 'spec/v2/core/idempotency.md §Layer 2: effect identity', `every attempt of one effect MUST declare the same keying — ${attempts.length} attempt(s) declared ${[...keyings].join(', ')}`),
|
|
144
|
+
).toBe(1);
|
|
131
145
|
});
|
|
132
146
|
});
|
|
@@ -3,8 +3,9 @@
|
|
|
3
3
|
* on `replay`, driven through the seams profile).
|
|
4
4
|
*
|
|
5
5
|
* A guarded manifest row that states `branchReFires: false` MUST NOT be fired
|
|
6
|
-
* again by a replay fork: the fork's Layer-2 effect ledger
|
|
7
|
-
* parent
|
|
6
|
+
* again by a replay fork: the fork's Layer-2 effect ledger records no attempt
|
|
7
|
+
* the parent did not make (suite 2.42.2; until then "cannot grow past the
|
|
8
|
+
* parent's", which a one-attempt source let a re-firing host pass at 1 ≤ 1). The witness is the host's own ledger, reached by
|
|
8
9
|
* firing a named manifest row inside a run through the catalogued seam
|
|
9
10
|
* (`api/seams-v2.yaml` `fireEffectSeam`), forking in `replay` mode, and reading
|
|
10
11
|
* `GET /runs/{runId}/effects` on both.
|
|
@@ -28,6 +29,7 @@ import { v2Discovery, gateFamily } from '../lib/v2.js';
|
|
|
28
29
|
import { seamsProfileAdvertised, SEAMS_PREFIX } from '../lib/seams.js';
|
|
29
30
|
import { softSkip } from '../lib/soft-skip.js';
|
|
30
31
|
import { req } from '../lib/requirement-ids.js';
|
|
32
|
+
import { refiredAttempts, type EffectRow } from '../lib/effect-refire.js';
|
|
31
33
|
|
|
32
34
|
const MANIFEST_PATH = '/host/effect-seams';
|
|
33
35
|
|
|
@@ -115,9 +117,20 @@ describe('RFC 0173 §C.2 — effect-seam-no-refire (gated on replay, seam-driven
|
|
|
115
117
|
}
|
|
116
118
|
const forkEffects = await driver.get(`/runs/${encodeURIComponent(String(forkBody.runId))}/effects`);
|
|
117
119
|
if (forkEffects.status !== 200) return softSkip('blocked', `GET /runs/{runId}/effects answered ${forkEffects.status} on the replay fork`);
|
|
120
|
+
// Not a count comparison. The seam fires ONE attempt, so a host that
|
|
121
|
+
// re-fires on replay reads 1 on the fork and `fork <= parent` passed it:
|
|
122
|
+
// measured, openwop-app with `sideEffecting` dropped still passed (suite
|
|
123
|
+
// 2.42.2). The witness is whether the fork records any attempt the parent
|
|
124
|
+
// never made (lib/effect-refire.ts): inherited history repeats a parent
|
|
125
|
+
// attempt, a re-fire is a new one.
|
|
126
|
+
const rowsOf = (r: OpenWOPResponse): EffectRow[] => {
|
|
127
|
+
const b = r.json as { effects?: unknown } | null;
|
|
128
|
+
return Array.isArray(b?.effects) ? (b.effects as EffectRow[]) : [];
|
|
129
|
+
};
|
|
130
|
+
const refired = refiredAttempts(rowsOf(parentEffects), rowsOf(forkEffects));
|
|
118
131
|
expect(
|
|
119
|
-
|
|
120
|
-
req('openwop.requirement.0173.effect-seam-no-refire', 'spec/v2/core/replay.md §Suppression', `seam ${String(target.seam)} is guarded, so a replay fork MUST NOT issue a further attempt through it — suppression is unconditional for mode: replay (§Suppression rule 1) and does not depend on branchReFires, which states only what a BRANCH may re-fire (§Branch: a host "MUST NOT report that as replay suppression"). The fork's
|
|
121
|
-
).
|
|
132
|
+
refired.length,
|
|
133
|
+
req('openwop.requirement.0173.effect-seam-no-refire', 'spec/v2/core/replay.md §Suppression', `seam ${String(target.seam)} is guarded, so a replay fork MUST NOT issue a further attempt through it — suppression is unconditional for mode: replay (§Suppression rule 1) and does not depend on branchReFires, which states only what a BRANCH may re-fire (§Branch: a host "MUST NOT report that as replay suppression"). The fork's ledger records ${refired.length} attempt(s) the parent never made: ${JSON.stringify(refired).slice(0, 300)}`),
|
|
134
|
+
).toBe(0);
|
|
122
135
|
});
|
|
123
136
|
});
|
|
@@ -172,10 +172,26 @@ describe('RFC 0208 — v2-mcp-mount-map (host as MCP 2026-07-28 server, gated on
|
|
|
172
172
|
const before = new Set(await list());
|
|
173
173
|
const ok = await toolCall(m.url, 'conformance-noop');
|
|
174
174
|
const fresh = (await list()).filter((x) => !before.has(x));
|
|
175
|
-
expect(fresh.length, req(R('mcp-run-transport'), 'interop-map.json mcp.methods tools/call', `tools/call MUST start a run (v2Operation createRun) that listRuns shows the same Subject (got ${fresh.length} new run(s); result ${JSON.stringify(ok.error ?? ok.result)})`)).
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
175
|
+
expect(fresh.length, req(R('mcp-run-transport'), 'interop-map.json mcp.methods tools/call', `tools/call MUST start a run (v2Operation createRun) that listRuns shows the same Subject (got ${fresh.length} new run(s); result ${JSON.stringify(ok.error ?? ok.result)})`)).toBeGreaterThanOrEqual(1);
|
|
176
|
+
// WHICH new run is ours (2.42.1). The leg counted `fresh.length === 1`, but
|
|
177
|
+
// the suite runs files concurrently and conformance-noop is every file's
|
|
178
|
+
// smallest run, so a sibling's run landed in the same window: the v2
|
|
179
|
+
// reference host's CI (4 workers) failed "got 2 new run(s)" on a host whose
|
|
180
|
+
// own tools/call started exactly one, and `fresh[0]` could have been the
|
|
181
|
+
// sibling's run, read for the wrong transport. Ours is the runId the result
|
|
182
|
+
// names when it names one listRuns shows, else the only new run; when the
|
|
183
|
+
// window stays ambiguous, the requirement holds if any new run started mcp.
|
|
184
|
+
const named = ((): string | null => {
|
|
185
|
+
const m = /"runId"\s*:\s*"([^"]+)"/.exec(JSON.stringify(ok.result ?? ''));
|
|
186
|
+
return m && fresh.includes(m[1]!) ? m[1]! : null;
|
|
187
|
+
})();
|
|
188
|
+
const transportOf = async (runId: string): Promise<string | undefined> => {
|
|
189
|
+
const poll = await driver.get(`/runs/${encodeURIComponent(runId)}/events/poll?timeout=1`);
|
|
190
|
+
return ((poll.json as { events?: Array<{ type?: string; payload?: { transport?: string } }> } | undefined)?.events ?? []).find((e) => e.type === 'run.started')?.payload?.transport;
|
|
191
|
+
};
|
|
192
|
+
const ours = named ?? (fresh.length === 1 ? fresh[0]! : null);
|
|
193
|
+
const transports = ours !== null ? [await transportOf(ours)] : await Promise.all(fresh.map(transportOf));
|
|
194
|
+
expect(transports.includes('mcp') ? 'mcp' : transports[0], req(R('mcp-run-transport'), 'interop-map.json mcp.methods tools/call; runs.md run.started', `the run starts with run.started.transport mcp (${ours !== null ? `run ${ours}` : `${fresh.length} runs started in the window, none named by the result`}; transports ${JSON.stringify(transports)})`)).toBe('mcp');
|
|
179
195
|
});
|
|
180
196
|
|
|
181
197
|
it('a suspending tool answers InputRequiredResult; the retry resolves it; requestState is single use and forgery-proof', async () => {
|