@openwop/openwop-conformance 2.45.2 → 2.45.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/README.md +2 -2
- package/dist/cli.js +1 -1
- package/dist/lib/certification-bundle-verify.js +22 -1
- package/dist/lib/profiles.js +17 -2
- package/dist/lib/requirement-registry.js +15 -0
- package/dist/lib/scenario-disposition.js +60 -9
- package/dist/spec-artifacts.lock.json +2 -2
- package/fixtures/conformance-budget-tool-calls.json +81 -0
- package/fixtures/conformance-fs-probe.json +25 -0
- package/fixtures/conformance-queue-consume.json +27 -0
- package/fixtures/conformance-queue-publish.json +29 -0
- package/fixtures/conformance-safefetch-probe.json +23 -0
- package/fixtures/conformance-secret-resolve-then-fail.json +33 -0
- package/fixtures/conformance-storage-probe.json +27 -0
- package/fixtures/conformance-tool-scope-probe.json +27 -0
- package/fixtures/openwop-secrets-run-witness.json +29 -0
- package/fixtures.md +113 -1
- package/package.json +2 -2
- package/requirement-aliases.json +72 -71
- package/requirements.json +1081 -43
- package/scenario-majors.json +45 -3
- package/schemas/CORPUS-STAMP.json +22 -22
- package/src/cli.ts +1 -1
- package/src/lib/backpressure-witness.ts +163 -0
- package/src/lib/budget-witness.ts +165 -0
- package/src/lib/certification-bundle-verify.ts +23 -2
- package/src/lib/driver.ts +61 -2
- package/src/lib/major-profile.ts +93 -0
- package/src/lib/memoryAttribution.ts +41 -6
- package/src/lib/polling.ts +2 -18
- package/src/lib/profiles.ts +26 -2
- package/src/lib/requirement-registry.ts +12 -0
- package/src/lib/run-secrets-witness.ts +240 -0
- package/src/lib/scenario-disposition.ts +54 -7
- package/src/lib/scratch-host.ts +130 -0
- package/src/lib/secret-scan.ts +141 -0
- package/src/lib/timeout-scale.ts +25 -0
- package/src/lib/triggerBridge.ts +69 -1
- package/src/scenarios/byok-roundtrip.test.ts +20 -4
- package/src/scenarios/runner-ledger.test.ts +2 -0
- package/src/scenarios/secrets-run-witness.test.ts +154 -0
- package/src/scenarios/trigger-bridge-delivery.test.ts +183 -126
- package/src/scenarios/trigger-refused-event-keeps-subscription.test.ts +141 -0
- package/src/scenarios/v2-budget-enforcement.test.ts +94 -0
- package/src/scenarios/v2-eval-mode-unadvertised-refused.test.ts +55 -0
- package/src/scenarios/v2-fs-sandbox-escape-refused.test.ts +161 -0
- package/src/scenarios/v2-memory-cross-tenant-isolation.test.ts +113 -0
- package/src/scenarios/v2-production-backpressure.test.ts +74 -0
- package/src/scenarios/v2-queue-cross-tenant-isolation.test.ts +118 -0
- package/src/scenarios/v2-safefetch-ssrf-refused.test.ts +198 -0
- package/src/scenarios/v2-secret-canary-absent.test.ts +211 -0
- package/src/scenarios/v2-secrets-run-witness.test.ts +167 -0
- package/src/scenarios/v2-storage-cross-tenant-isolation.test.ts +213 -0
- package/src/scenarios/v2-tool-authorization-fail-closed.test.ts +197 -0
- package/src/scenarios/v2-workspace-scope-from-identity.test.ts +146 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,82 @@
|
|
|
1
1
|
# `@openwop/openwop-conformance` Changelog
|
|
2
2
|
|
|
3
|
+
## [2.45.4] — 2026-10-01 — v2 ports of production-backpressure and budget-enforcement, on a per-major profile
|
|
4
|
+
|
|
5
|
+
- **The gate runs every host-free scenario** (not packed; no scenario changes). New `scripts/list-host-free-scenarios.mjs` lists the scenarios whose import closure never reaches `lib/driver`, `lib/sse` or `lib/env` and never calls `fetch(`: 70 files today. `scripts/openwop-check.sh` runs them with `OPENWOP_BASE_URL` and `OPENWOP_API_KEY` unset, after the nine hand-listed files.
|
|
6
|
+
- **Why.** `host-callback-declaration` and `runner-ledger` failed on every host in the 2.45.3 candidate and nothing in CI said so (#1835).
|
|
7
|
+
- **Proof.** Clean tree: 66 files pass and 4 skip, no failure. With `REQUIRES_HOST_CALLBACK` removed from `v2-tool-authorization-fail-closed`, the step fails on `host-callback-declaration`.
|
|
8
|
+
- **Limits.** The walk reads static relative imports. `--check` fails if the set is empty or loses a known member. `ai-envelope-shape` is on the hand list and outside the derived set, so the hand list stays.
|
|
9
|
+
- **The 2.45.4 cycle opens.** Two new scenarios and one new fixture change the packed content.
|
|
10
|
+
- **Fix: an unclaimed any-of floor no longer blocks a major-1 bundle** (a 2.45.3 defect). The runner wrote the `openwop.floor.anyof.byok-roundtrip+secrets-run-witness` summary row for every floor in the v1 table, claimed or not. A host that does not advertise secrets records both members `inapplicable`, so it got one `blocked` row for `openwop-secrets`, a profile it never claimed, and a v3 bundle with any `blocked` row certifies nothing (RFC 0168 §E.1). The row is now written only when the profile is claimed. A claimed profile is unchanged: unwitnessed members still record `blocked` and deny it.
|
|
11
|
+
- **Who was affected:** a major-1 host cutting a v3 bundle on 2.45.3 without claiming `openwop-secrets`. Major-2 bundles never carried the row.
|
|
12
|
+
- **Proof.** `floor-any-of.test.ts`: the same evidence yields no row unclaimed and a `blocked` row claimed. Found by running the derivation, after reading the code had suggested the row was harmless.
|
|
13
|
+
- **What differs between majors is now data** (`src/lib/major-profile.ts`). One table row per major holds the run path, the version headers, the event names (from the codemap), where retry timing lives, how a family is advertised, and how a run budget rides on `createRun`. A shared witness takes a profile and never names a major, so a port to a later major is one row plus a thin scenario file. A major with no row throws; it never falls back to an older major's rules.
|
|
14
|
+
- **New scenario `v2-production-backpressure`** (major 2, unaided; `conformance.md` §Production profile). Gated on `production.backpressure.inflightCap`. The suite holds that many event streams open, waits for each to open, and sends one more request.
|
|
15
|
+
- `openwop.requirement.production.backpressure-refusal`: the extra request answers `503 service_unavailable` with `Retry-After`, equal to `retryAfterSeconds` where advertised.
|
|
16
|
+
- `openwop.requirement.0171.error-registry.no-retry-details`: the refusal carries no `details.retryAfter*`. v1 required `details.retryAfter`; v2 forbids it.
|
|
17
|
+
- **Dispositions.** No `production`, `backpressure` or `inflightCap`, a cap above 64, or an unadvertised hold or probe fixture: `inapplicable`. A slot that cannot be held, or a probe with no response: `blocked`.
|
|
18
|
+
- **Not ported:** v1's "discovery is exempt from the cap" leg. No v2 document states it.
|
|
19
|
+
- **New scenario `v2-budget-enforcement`** (major 2, unaided; `runs.md` §`budget` section). At major 2 a budget rides on `createRun` (`configurable.budget`), so no seam is needed. The suite runs the new fixture with `{ maxToolCalls: 2, thresholdPercent: 50, onExhaustion: "fail" }` and reads the log through the poll.
|
|
20
|
+
- `openwop.requirement.runs.budget-lifecycle`: `budget.reserved`, `budget.threshold-crossed` and `budget.exhausted`, in order.
|
|
21
|
+
- `openwop.requirement.runs.budget-enforcement`: a `hard` host emits `cap.breached` (`budget-tool-calls`) and fails the run `budget_exhausted`; an `advisory` host does not stop it.
|
|
22
|
+
- `openwop.requirement.runs.budget-content-free`: no `budget.*` or `cap.breached` payload carries pricing or a credential.
|
|
23
|
+
- **Dispositions.** No `budget`, no `toolCalls` dimension, or the fixture unadvertised: `inapplicable`. The fixture is the opt-in, so a host that advertises `budget` without seeding it is unwitnessed, not blocked. A refused valid create: `executed-fail`.
|
|
24
|
+
- **Not ported:** v1's `budget_model_denied` leg. It needs a fixture that resolves a model without a seam.
|
|
25
|
+
- **New fixture `conformance-budget-tool-calls`**: one `core.conformance.mock-agent` with three scripted tool calls.
|
|
26
|
+
- **Proof.** No host serves either family unaided at major 2, so both witnesses are proven against `src/lib/scratch-host.ts`, a profile-driven test double: 29 self-tests, one defect each (`backpressure-witness.test.ts`, `budget-witness.test.ts`, `major-profile.test.ts`). The scenario files were also run end to end against a conforming and a defective scratch host.
|
|
27
|
+
|
|
28
|
+
## [2.45.3] — 2026-10-01 — RFC 0229 (Active): a production host can witness secret resolution without an oracle; v2 tenant-isolation witnesses for storage, fs, memory, workspace, queues and secrets
|
|
29
|
+
|
|
30
|
+
- **Two self-checks that failed on every host are fixed before the cut.** Found by running the candidate on openwop-app at major 1 against 2.45.2: four rows went `executed-pass` → `executed-fail`, and both files failed the same way with no host.
|
|
31
|
+
- `host-callback-declaration`: `v2-tool-authorization-fail-closed` imports `lib/effect-receiver` and now exports `REQUIRES_HOST_CALLBACK`.
|
|
32
|
+
- `runner-ledger`: the derivation emits the any-of summary row for `openwop-secrets` whatever profile is claimed, as it does the prefix rows. The test's synthetic report holds neither member, so the row is `blocked`; the test now asserts that and excludes it from the claimed floor's pass count. The derivation is unchanged.
|
|
33
|
+
- The rest of that diff was as intended: no row newly `blocked`, the four `0083.trigger-delivery.*` rows and the any-of row `executed-pass`, `secrets-run-witness` and `trigger-refused-event-keeps-subscription` `inapplicable`.
|
|
34
|
+
- **The 2.45.3 cycle opens.** RFC 0229's gap register changes the packed `spec/v1/gaps.json`.
|
|
35
|
+
- **A lost response is `blocked`, not `executed-fail`** (#1829).
|
|
36
|
+
- **Bound.** `driver.request` gives each request 20 s (`REQUEST_TIMEOUT_MS`, scaled by `OPENWOP_POLL_TIMEOUT_SCALE`), below vitest's 30 s `testTimeout`. The body read is under the same bound. A caller's own `signal` replaces the bound, and its abort is rethrown as-is.
|
|
37
|
+
- **Named.** A timeout or a failed connection throws `TransportError` (`transport-loss: <METHOD> <origin+path> got no response …`). The query string is dropped from the message.
|
|
38
|
+
- **Recorded.** A test that fails on a transport loss before its first assertion records `blocked`; the file row follows when every failure in it is one and nothing was asserted. A loss after an assertion stays `executed-fail`, and an assertion failure is never excused.
|
|
39
|
+
- **Why.** On MyndHyve's 2.45.2 cut the runner's network dropped for about 9 s. `v2-mcp-client-results` hung to the harness timeout and recorded `executed-fail` with 0 assertions against a host that had answered in 2.6 ms. `blocked` still denies certification, so a flaky runner cannot certify; it can no longer convict the host.
|
|
40
|
+
- **Proof.** `src/lib/driver-transport-loss.test.ts` (10 cases): a local server that never answers yields a timeout `TransportError` well inside the bound; without the classification the same failure is `executed-fail`.
|
|
41
|
+
- **New scenario `v2-eval-mode-unadvertised-refused`** (major 2, unaided): a host that does not advertise `agents.evalSuite` answers `POST /runs {mode: "eval"}` with `422 capability_not_provided` (`runs.md` §Refusals; `openwop.requirement.runs.eval-mode-unadvertised-refused`). The eval scenarios ran at major 1 only, so nothing measured this at major 2. A host that advertises the facet records `inapplicable`. If a host wrongly accepts the create, the leg cancels the run before asserting.
|
|
42
|
+
- **Proof.** Passes on the v2 reference host with openwop-examples #143; fails on the host without it, which answered `capability_required`.
|
|
43
|
+
- **`trigger-bridge-delivery`: a normative-surface witness path (RFC 0230), and one requirement per leg.**
|
|
44
|
+
- **Split.** The single combined `it` becomes four, each with its own requirement id: `openwop.requirement.0083.trigger-delivery.{dedup,dead-letter,causation,runless-content-free}`. The old per-`it` id aliases to the dedup leg.
|
|
45
|
+
- **Seam path unchanged.** It stays the primary witness whenever the host serves the seams.
|
|
46
|
+
- **Normative-surface path.** When the seams are absent and the host advertises `triggerBridge.ingestion.inboundSigning: ["standard-webhooks-1"]`, legs 1–3 run it. The suite registers a webhook subscription, then POSTs Standard-Webhooks-signed bodies to its `ingestUrl` with no OpenWOP credential:
|
|
47
|
+
- dedup: 202, then `200 duplicate` with the same `runId`;
|
|
48
|
+
- dead-letter: a bad signature gets 401 with no run, the subscription stays `active` (RFC 0230 §C.1), and a signed post then gets 202;
|
|
49
|
+
- causation: `run.started.causationId` via the run's own event poll.
|
|
50
|
+
- **Run-less content-freeness** stays seam-only, since those events are on no run's log. It **stays in the floor**, which never narrows: a seam-free host records one honest red row instead of three.
|
|
51
|
+
- **Proof.** Measured on openwop-app with the RFC 0230 route on and the seams off: legs 1–3 pass in strict mode on the new path. Sabotaged host builds each fail only their own leg:
|
|
52
|
+
- a duplicate `webhook-id` starting a new run fails dedup;
|
|
53
|
+
- a bad signature accepted fails dead-letter;
|
|
54
|
+
- a refused post dead-lettering the subscription fails dead-letter;
|
|
55
|
+
- a missing `causationId` fails causation.
|
|
56
|
+
- **RFC 0229 `Active`** (the window is waived by steward override of RFC 0147 §A.6). Two new scenarios, `secrets-run-witness` (major 1) and `v2-secrets-run-witness` (major 2), share `lib/run-secrets-witness.ts`. Each draws a fresh 64-character value `C` per run, supplies it as `createRun.runSecrets` under `run:openwop-witness`, and runs the fixture `openwop-secrets-run-witness`. Four rows, the same ids at both majors:
|
|
57
|
+
- `openwop.requirement.secrets.run-witness-resolves`: `matched: true` for `C`, and `false` when `expectedSha256` names another value;
|
|
58
|
+
- `openwop.requirement.secrets.run-witness-redacted`: `C` is absent from the create answer, snapshot, polled and `debug`-mode events, the run list and the debug bundle where served. The scan reads raw text and covers every encoding RFC 0229 §F lists. It also covers the digest, wherever the suite did not send that digest itself;
|
|
59
|
+
- `openwop.requirement.secrets.run-witness-scope-bound`: a non-`run:` ref fails `credential_forbidden`; an unsupplied or earlier run's `run:` ref fails `credential_not_found`; a branch fork, where served, fails `credential_not_found`;
|
|
60
|
+
- `openwop.requirement.secrets.run-secrets-outside-request-digest`: a same-key retry differing only in `runSecrets` replays with the original `runId` and `OpenWOP-Idempotent-Replay: true`, is never `409 idempotency_key_mismatch`, and its value appears nowhere.
|
|
61
|
+
|
|
62
|
+
**Dispositions.** Facet absent: `inapplicable`. Facet without the fixture: `blocked`. A refused well-formed `runSecrets`: `executed-fail`.
|
|
63
|
+
|
|
64
|
+
**Sabotage.** Each row failed against a scratch host with its defect: a witness that resolves a stored secret by a non-`run:` name, an echoed value, an emitted digest, a digest-covering idempotency record, a fork that inherits, and a cross-run fallback. The clean host passed all four rows at both majors.
|
|
65
|
+
- **The `openwop-secrets` floor is an any-of group** (`requiredAnyOf`, RFC 0229 §E). It is satisfied by a witnessed pass of `byok-roundtrip` or `secrets-run-witness` with no member failing, and never by an `inapplicable` one. The runner writes one summary row, `openwop.floor.anyof.byok-roundtrip+secrets-run-witness`, and both bundle verifiers and the `--certify` witness count read it. `src/lib/floor-any-of.test.ts` pins this.
|
|
66
|
+
- `byok-roundtrip` records `inapplicable` instead of `blocked` when the canary fixture is withheld and the host advertises `secrets.runSecrets` and the witness fixture. A host that takes the other path is not denied the bundle over a canary it does not serve.
|
|
67
|
+
- Otherwise `byok-roundtrip` is unchanged.
|
|
68
|
+
- **Host impact: none today.** No committed bundle and neither live production host advertises `secrets.runSecrets`, so both new files record `inapplicable` everywhere, and `byok-roundtrip` keeps its current disposition.
|
|
69
|
+
- **New v2 tenant-isolation witnesses** from `docs/V2-WITNESS-COVERAGE.md`, each sabotage-proved against a host with the exact defect: `v2-storage-cross-tenant-isolation` (eight families, one id each), `v2-fs-sandbox-escape-refused` (absolute, `..`, symlink), `v2-memory-cross-tenant-isolation`, `v2-workspace-scope-from-identity`, `v2-queue-cross-tenant-isolation`, and `v2-secret-canary-absent` (a host-resolved secret on no readable surface, including after a failure and in a fork). Each is `inapplicable` where its family is not advertised; MyndHyve runs the memory and secret-canary legs.
|
|
70
|
+
- **`memory-attribution-replay-stable` runs on v2 hosts.** `emitsWriteEvents()` (`lib/memoryAttribution.ts`) required `attribution.supported === true`, a field a v2 family record does not have (presence is the claim, RFC 0169 §A.2), so the file recorded `inapplicable` on every v2 host, including the two that advertise `memory.attribution.emitsWriteEvents: true`. At major 2 it now reads `emitsWriteEvents` on the present record; major 1 still requires `supported`. `memoryWrittenEvents()` read `GET /runs/{runId}/events` at major 2, which is the SSE stream, so once the gate opened every leg would have recorded `blocked` ("run wrote no memory"). It now reads `GET /runs/{runId}/events/poll` and pages on `afterSequence`. No other `conformance/src/lib` helper that a major-2 scenario imports has the same bug (`toolCatalog`, `context-budget` and `run-secrets-witness` were already major-aware). Sabotage: against a loopback v2 stub, the clean host passes, and a replay that re-mints ids, one that suppresses them, and one that adds a fresh id each fail. A host without the facet records `inapplicable`. Self-test: `lib/memoryAttribution.test.ts`.
|
|
71
|
+
- **New scenario `trigger-refused-event-keeps-subscription` (major 1), for the §F.2 Class 3 correction (RFC 0230 risk R3).** Row `openwop.requirement.trigger-bridge.refused-event-keeps-subscription`. The row is gated on the `openwop-trigger-bridge` profile and on `webhook` in `triggerBridge.ingestion.externalSources`. It drives the delivery seam's new OPTIONAL `scenario: "refused"`, then `deliver` on the same `subscriptionId`, and asserts three things:
|
|
72
|
+
- the refused event is a dead-lettered `trigger.delivery.attempted` and starts no run;
|
|
73
|
+
- no `trigger.subscription.state.changed` is emitted, and the subscription reads back `active`;
|
|
74
|
+
- the next verified event on that subscription is delivered.
|
|
75
|
+
|
|
76
|
+
**Dispositions.** Seam without `refused` (`400`): `blocked` (fails under strict mode). No `webhook` ingestion: `inapplicable`.
|
|
77
|
+
|
|
78
|
+
**Sabotage.** Proved against a loopback stub. The corrected host passes (default and strict). Each defective variant fails: the old subscription-level transition with its `state.changed`; the same transition with no event (the read shows `dead-lettered`); a refused event that starts a run. **Host impact:** trigger scenarios are major-1 only and no committed bundle carries one; a major-1 host that derives the profile with `webhook` ingestion records `blocked` until its seam serves `refused`.
|
|
79
|
+
|
|
3
80
|
## [2.45.2] — 2026-09-30 — a run-failure code must be registered or a vendor code (no longer advisory)
|
|
4
81
|
|
|
5
82
|
- **`errors.event-code-registered` is required** (#1698). The leg in `v2-error-registry` shipped ADVISORY in 2.43.1: an unregistered, non-vendor code on `run.failed`, `node.failed` or the snapshot's `error` recorded a partial witness naming the codes, while hosts remapped. That branch is removed, so such a code now fails the row, which names each `where=code`.
|
package/README.md
CHANGED
|
@@ -135,7 +135,7 @@ Exit code is non-zero on any failed assertion. `--certify` distinguishes: `0`
|
|
|
135
135
|
|
|
136
136
|
## What's Covered
|
|
137
137
|
|
|
138
|
-
The current suite has
|
|
138
|
+
The current suite has 589 scenario files under `src/scenarios/`.
|
|
139
139
|
- 2026-09-29 (suite 2.45.0, openwop#1763): NEW `v2-cors-preflight.test.ts` — `headers.md` §Cross-origin preflight. For every operation it preflights with `Origin`, the method, and the operation's declared request headers (plus `authorization` / `content-type` where they apply, from `spec/v2/path-manifest.json`), and reads the answer as the Fetch standard's CORS-preflight check does. Self-gated: `inapplicable` when no operation's preflight grants `OPENWOP_CORS_ORIGIN` (default `https://conformance.invalid`).
|
|
140
140
|
- 2026-09-23 (suite 2.37.0 cycle, RFC 0213): NEW `v2-sse-last-event-id-cursor.test.ts` (a `Last-Event-ID` past the log is an exclusive cursor; a malformed id, when refused, is `400 validation_error`; the cursor never changes the answer for an unknown or foreign-tenant run — public test of `event-cursor-after-authorization`), `v2-idempotency-in-flight.test.ts` (five concurrent same-key creates yield one run; each loser is a marked replay or `409 idempotency_in_flight` with no retry timing in `details`; `partial-witness` when no loser was refused in flight) and `v2-interrupt-resolve-terminal.test.ts` (a run-scoped resolve after cancel or completion is `409 interrupt_already_resolved`, never `interrupt_cancelled`). All three sit off the core-standard floor until measured on the three bundle hosts.
|
|
141
141
|
- 2026-09-03 (suite `1.157.0 -> 1.158.0`, gap G17): NEW `idempotency-concurrent-claim.test.ts` — drives the new `host-sample-test-seams.md` §25 concurrent duplicate-delivery seam for the RFC 0150 §B / `idempotency.md` §"Concurrent duplicates (Layer 2)" atomic-claim MUST, which is unconditional and had no witness of any kind. Asserts every executor mints the SAME `logicalInvocationId` **before** asserting `delivered === 1` — without the identity check a host passes by minting different ids and never colliding, one effect because nothing raced. Not profile-gated and so not opt-out-able (the obligation is unconditional); an unmounted seam records `blocked`, which is not certifiable. Graduates `layer2-invocation-claim-atomic` reference-impl -> protocol.
|
|
@@ -482,7 +482,7 @@ Server-required (added in 1.7.0):
|
|
|
482
482
|
| ------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
483
483
|
| **Redaction** | [`capabilities.md`](../spec/v1/capabilities.md) §"Secrets" + NFR-7 + §"aiProviders" | Vendor-neutral assertions that the server doesn't leak secret material. Three scenario groups: (a) discovery shape contract — `secrets` + `aiProviders` advertisements are well-formed regardless of `secrets.supported`; when `supported === true`, scopes MUST be non-empty + `resolution === 'host-managed'`; `byok ⊆ supported`. (b) bearer-token redaction — invalid Bearer canary in `Authorization` header is not echoed in the 401 response body. (c) credentialRef echo control — gated on `secrets.supported === true`; canary planted in `configurable.ai.credentialRef` MUST NOT appear in any RunEvent payload (poll-based capture; transport-agnostic). Uses runtime-built canary fixtures (`lib/canaries.ts`) that defeat static secret scanners. 6 scenarios. |
|
|
484
484
|
|
|
485
|
-
Current source tree:
|
|
485
|
+
Current source tree: 589 scenario files. Use [`coverage.md`](./coverage.md) for current grade/gap tracking.
|
|
486
486
|
|
|
487
487
|
## Remaining Gaps
|
|
488
488
|
|
package/dist/cli.js
CHANGED
|
@@ -646,7 +646,7 @@ async function runCertify(args, baseUrl, apiKey) {
|
|
|
646
646
|
const floor = PROFILE_FLOOR_SCENARIOS[profile];
|
|
647
647
|
if (!floor)
|
|
648
648
|
return 0;
|
|
649
|
-
const onFloor = (scenario) => floor.required.includes(scenario) || (floor.requiredAnyPrefix ?? []).some((pre) => scenario.startsWith(pre));
|
|
649
|
+
const onFloor = (scenario) => floor.required.includes(scenario) || (floor.requiredAnyPrefix ?? []).some((pre) => scenario.startsWith(pre)) || (floor.requiredAnyOf ?? []).some((g) => g.includes(scenario));
|
|
650
650
|
return derived.requirements.filter((r) => r.disposition === 'executed-pass' && onFloor(r.scenarioId)).length;
|
|
651
651
|
};
|
|
652
652
|
// `certified` IS the verdict (RFC 0148 §A; RFC 0168 §E.1 adds the bundle-wide
|
|
@@ -26,7 +26,7 @@
|
|
|
26
26
|
import { createHash } from 'node:crypto';
|
|
27
27
|
import { PROFILE_FLOOR_SCENARIOS, profileDerivable, } from './profiles.js';
|
|
28
28
|
import { CERTIFIABLE, DISPOSITIONS } from './requirement-ledger.js';
|
|
29
|
-
import { floorFilesFor, requirementIdForPrefix, requirementIdForScenario } from './requirement-registry.js';
|
|
29
|
+
import { floorFilesFor, requirementIdForAnyOf, requirementIdForPrefix, requirementIdForScenario } from './requirement-registry.js';
|
|
30
30
|
import { UNCLASSIFIED_RETURN_DETAIL } from './soft-skip.js';
|
|
31
31
|
/**
|
|
32
32
|
* The BYOK conformance canary (`fixtures.md` §conformance-secrets-roundtrip;
|
|
@@ -84,6 +84,7 @@ function requiredFor(profile, document) {
|
|
|
84
84
|
return {
|
|
85
85
|
files: [...files],
|
|
86
86
|
prefixes: [...(floor.requiredAnyPrefix ?? [])],
|
|
87
|
+
anyOf: [...(floor.requiredAnyOf ?? [])],
|
|
87
88
|
discoveryOnly: floor.discoveryOnly === true,
|
|
88
89
|
};
|
|
89
90
|
}
|
|
@@ -192,6 +193,26 @@ export function verifyBundleV2(bundle) {
|
|
|
192
193
|
if (!considered.some((r) => isWitnessedPass(r)))
|
|
193
194
|
notCertifiable.push(id);
|
|
194
195
|
}
|
|
196
|
+
for (const group of req.anyOf) {
|
|
197
|
+
// RFC 0229 §E: one requirement, satisfied by a WITNESSED pass of any
|
|
198
|
+
// member with no member failing. The emitter writes a summary row under
|
|
199
|
+
// the group id; member rows are read when it is absent.
|
|
200
|
+
const id = requirementIdForAnyOf(group);
|
|
201
|
+
required.push(id);
|
|
202
|
+
const summary = byId.get(id)?.[0];
|
|
203
|
+
const members = group.map((f) => byId.get(requirementIdForScenario(scenarioBasename(f)))?.[0]).filter((r) => r !== undefined);
|
|
204
|
+
if (summary === undefined && members.length === 0) {
|
|
205
|
+
profileRejections.push({ kind: 'unwitnessed-requirement', profile, requirementId: id, detail: `${id}: no alternative (${group.join(', ')}) recorded a row` });
|
|
206
|
+
continue;
|
|
207
|
+
}
|
|
208
|
+
const considered = summary !== undefined ? [summary] : members;
|
|
209
|
+
if (considered.some((r) => r.disposition === 'executed-pass' && !isWitnessedPass(r))) {
|
|
210
|
+
profileRejections.push({ kind: 'vacuous-pass', profile, requirementId: id, detail: `${id}: executed-pass with no witnessed assertion` });
|
|
211
|
+
continue;
|
|
212
|
+
}
|
|
213
|
+
if (considered.some((r) => r.disposition === 'executed-fail') || !considered.some((r) => isWitnessedPass(r)))
|
|
214
|
+
notCertifiable.push(id);
|
|
215
|
+
}
|
|
195
216
|
const evidenceValid = profileRejections.length === 0;
|
|
196
217
|
const certified = derivable && evidenceValid && notCertifiable.length === 0 && (req.discoveryOnly || required.length > 0);
|
|
197
218
|
return { profile, derivable, floorUnspecified: false, required, rejections: profileRejections, notCertifiable, evidenceValid, certified };
|
package/dist/lib/profiles.js
CHANGED
|
@@ -462,6 +462,10 @@ export function hasProfile(c, profile) {
|
|
|
462
462
|
return isTriggerBridge(c);
|
|
463
463
|
}
|
|
464
464
|
}
|
|
465
|
+
/** Every scenario FILE a floor names: `required`, every `conditional` branch, and every `requiredAnyOf` member. */
|
|
466
|
+
export function floorMemberFiles(floor) {
|
|
467
|
+
return [...floor.required, ...(floor.conditional ?? []).flatMap((c) => c.required), ...(floor.requiredAnyOf ?? []).flat()];
|
|
468
|
+
}
|
|
465
469
|
export const PROFILE_FLOOR_SCENARIOS = {
|
|
466
470
|
'openwop-core-standard': {
|
|
467
471
|
required: [
|
|
@@ -531,7 +535,11 @@ export const PROFILE_FLOOR_SCENARIOS = {
|
|
|
531
535
|
// `profiles.md` §`openwop-secrets`: credential resolution per `run-options.md`
|
|
532
536
|
// §"Credential references" — the BYOK canary round-trip (`fixtures.md`
|
|
533
537
|
// §conformance-secrets-roundtrip, SR-1) is the profile's proof.
|
|
534
|
-
|
|
538
|
+
// RFC 0229 §E (2026-09-30): a SECOND floor path, as an any-of group. The
|
|
539
|
+
// canary fixture's node is an oracle over whatever its name reaches, so a
|
|
540
|
+
// production host rightly withholds it; the run-witness reads only a value
|
|
541
|
+
// the caller supplied with the same run and outputs only `matched`.
|
|
542
|
+
'openwop-secrets': { required: [], requiredAnyOf: [['byok-roundtrip.test.ts', 'secrets-run-witness.test.ts']] },
|
|
535
543
|
// `profiles.md` §`openwop-provider-policy`: the four-mode taxonomy shape and
|
|
536
544
|
// its enforcement on the wire.
|
|
537
545
|
'openwop-provider-policy': { required: ['policies.test.ts', 'providerPolicyEnforcement.test.ts'] },
|
|
@@ -650,9 +658,16 @@ export function verifyBundleProfile(bundle, profile) {
|
|
|
650
658
|
}
|
|
651
659
|
const missingFloor = requiredFiles.filter((r) => !passed.has(scenarioBasename(r)));
|
|
652
660
|
const prefixOk = (floor.requiredAnyPrefix ?? []).every((p) => [...passed].some((s) => s.startsWith(p)));
|
|
661
|
+
// RFC 0229 §E any-of groups: a v1 bundle lists only what passed, so a group
|
|
662
|
+
// holds when one member is in `results.passed`; an unmet group is reported
|
|
663
|
+
// as missing, spelled as its alternatives.
|
|
664
|
+
const groups = floor.requiredAnyOf ?? [];
|
|
665
|
+
for (const g of groups)
|
|
666
|
+
if (!g.some((f) => passed.has(scenarioBasename(f))))
|
|
667
|
+
missingFloor.push(g.join(' | '));
|
|
653
668
|
// A conditional floor none of whose branches matched requires nothing — that
|
|
654
669
|
// is unprovable for a non-discovery-only profile, not proven.
|
|
655
|
-
const evaluable = floor.discoveryOnly === true || requiredFiles.length > 0 || (floor.requiredAnyPrefix ?? []).length > 0;
|
|
670
|
+
const evaluable = floor.discoveryOnly === true || requiredFiles.length > 0 || (floor.requiredAnyPrefix ?? []).length > 0 || groups.length > 0;
|
|
656
671
|
const floorProven = evaluable && missingFloor.length === 0 && prefixOk;
|
|
657
672
|
return {
|
|
658
673
|
profile,
|
|
@@ -32,6 +32,14 @@ export function requirementIdForScenario(scenarioFile) {
|
|
|
32
32
|
export function requirementIdForPrefix(prefix) {
|
|
33
33
|
return `openwop.floor.any.${prefix}`;
|
|
34
34
|
}
|
|
35
|
+
/**
|
|
36
|
+
* Any-of groups (RFC 0229 §E) become one requirement:
|
|
37
|
+
* `['byok-roundtrip.test.ts', 'secrets-run-witness.test.ts']` →
|
|
38
|
+
* `openwop.floor.anyof.byok-roundtrip+secrets-run-witness`.
|
|
39
|
+
*/
|
|
40
|
+
export function requirementIdForAnyOf(files) {
|
|
41
|
+
return `openwop.floor.anyof.${files.map((f) => f.replace(/\.test\.ts$/, '')).join('+')}`;
|
|
42
|
+
}
|
|
35
43
|
/**
|
|
36
44
|
* The requirement IDs a profile's certification rests on.
|
|
37
45
|
*
|
|
@@ -121,6 +129,7 @@ export function requirementsFor(profile, document) {
|
|
|
121
129
|
return [
|
|
122
130
|
...files.map(requirementIdForScenario),
|
|
123
131
|
...(floor.requiredAnyPrefix ?? []).map(requirementIdForPrefix),
|
|
132
|
+
...(floor.requiredAnyOf ?? []).map(requirementIdForAnyOf),
|
|
124
133
|
];
|
|
125
134
|
}
|
|
126
135
|
/** Every registered requirement ID across every profile with a runtime floor. */
|
|
@@ -137,6 +146,12 @@ export function allRequirements() {
|
|
|
137
146
|
ids.add(requirementIdForScenario(f));
|
|
138
147
|
for (const p of floor.requiredAnyPrefix ?? [])
|
|
139
148
|
ids.add(requirementIdForPrefix(p));
|
|
149
|
+
// an any-of group is one requirement, and each member file records under its own floor id
|
|
150
|
+
for (const g of floor.requiredAnyOf ?? []) {
|
|
151
|
+
ids.add(requirementIdForAnyOf(g));
|
|
152
|
+
for (const f of g)
|
|
153
|
+
ids.add(requirementIdForScenario(f));
|
|
154
|
+
}
|
|
140
155
|
void profile;
|
|
141
156
|
}
|
|
142
157
|
return [...ids].sort();
|
|
@@ -26,24 +26,20 @@
|
|
|
26
26
|
* host or a vitest subprocess.
|
|
27
27
|
*/
|
|
28
28
|
import { scenarioFileOfItId } from './requirement-ids.js';
|
|
29
|
-
import { PROFILE_FLOOR_SCENARIOS } from './profiles.js';
|
|
29
|
+
import { DEPRECATED_PROFILE_ALIASES, PROFILE_FLOOR_SCENARIOS, floorMemberFiles } from './profiles.js';
|
|
30
30
|
import { targetMajor } from './seams.js';
|
|
31
31
|
import { PKG_ROOT_PATH } from './paths.js';
|
|
32
32
|
import { v2ProfileFloorFiles } from './requirement-registry.js';
|
|
33
|
-
import { requirementIdForScenario, requirementIdForPrefix, requirementsFor, v2FloorsActive } from './requirement-registry.js';
|
|
33
|
+
import { requirementIdForScenario, requirementIdForPrefix, requirementIdForAnyOf, requirementsFor, v2FloorsActive } from './requirement-registry.js';
|
|
34
34
|
import { UNCLASSIFIED_RETURN_DETAIL } from './soft-skip.js';
|
|
35
35
|
import { SPEC_COHERENCE_SCENARIOS, SPEC_COHERENCE_DETAIL } from './spec-coherence.js';
|
|
36
36
|
import { CERTIFIABLE } from './requirement-ledger.js';
|
|
37
37
|
/** All scenario basenames that appear in some profile's runtime floor. */
|
|
38
38
|
export function floorScenarioFiles() {
|
|
39
39
|
const out = new Set();
|
|
40
|
-
for (const floor of Object.values(PROFILE_FLOOR_SCENARIOS))
|
|
41
|
-
for (const f of floor
|
|
40
|
+
for (const floor of Object.values(PROFILE_FLOOR_SCENARIOS))
|
|
41
|
+
for (const f of floorMemberFiles(floor))
|
|
42
42
|
out.add(f);
|
|
43
|
-
for (const c of floor.conditional ?? [])
|
|
44
|
-
for (const f of c.required)
|
|
45
|
-
out.add(f);
|
|
46
|
-
}
|
|
47
43
|
// At target major 2 the floors come from the declaration, not the v1 hand
|
|
48
44
|
// table. The ledger and --certify MUST agree on this set, or a floor file is
|
|
49
45
|
// minted `openwop.scenario.*` and looked up as `openwop.floor.*` — which is
|
|
@@ -77,10 +73,13 @@ export function requirementIdForFile(basename) {
|
|
|
77
73
|
* the thing that passed.
|
|
78
74
|
*/
|
|
79
75
|
export const PARTIAL_WITNESS_PREFIX = 'partial-witness: ';
|
|
76
|
+
/** `driver.ts`'s `TRANSPORT_LOSS_PREFIX`, restated so this module stays free of the driver (a test pins the two equal). */
|
|
77
|
+
export const TRANSPORT_LOSS_MARK = 'transport-loss: ';
|
|
80
78
|
/**
|
|
81
79
|
* The per-`it` record (RFC 0148 §A at test granularity), as `setup.ts`
|
|
82
80
|
* computes it in `afterEach`. Pure so a lib test can pin it:
|
|
83
81
|
* - fail ⇒ executed-fail (detail = the first error message)
|
|
82
|
+
* - fail on a driver transport loss, 0 assertions ⇒ blocked (#1829: nothing was observed)
|
|
84
83
|
* - pass with ≥ 1 assertion ⇒ executed-pass
|
|
85
84
|
* - a behaviorGate entry journaled during the test ⇒ that gate's disposition
|
|
86
85
|
* - pass with 0 assertions ⇒ the softSkip note written DURING THIS TEST
|
|
@@ -114,6 +113,15 @@ blockedStands = false,
|
|
|
114
113
|
/** The `it` title, so a failed row says WHICH leg of the requirement failed. */
|
|
115
114
|
testName) {
|
|
116
115
|
if (state === 'fail') {
|
|
116
|
+
// #1829: the driver got no response (`TransportError`) before the test
|
|
117
|
+
// asserted anything. Nothing about the host was observed, so this is not a
|
|
118
|
+
// verdict on it. `blocked` still denies certification, so a flaky runner
|
|
119
|
+
// cannot certify; it just cannot convict the host either. A loss AFTER an
|
|
120
|
+
// assertion stays `executed-fail`: the test was mid-measurement, and a host
|
|
121
|
+
// that stops answering part-way is a finding.
|
|
122
|
+
if (assertionCalls === 0 && firstError !== undefined && firstError.includes(TRANSPORT_LOSS_MARK)) {
|
|
123
|
+
return { disposition: 'blocked', detail: firstError.slice(firstError.indexOf(TRANSPORT_LOSS_MARK), firstError.indexOf(TRANSPORT_LOSS_MARK) + 300) };
|
|
124
|
+
}
|
|
117
125
|
const where = testName === undefined || testName.trim() === '' ? '' : ` in "${testName.slice(0, 120)}"`;
|
|
118
126
|
return { disposition: 'executed-fail', detail: `the test executed and failed${where}: ${(firstError ?? 'no message').slice(0, 300)}` };
|
|
119
127
|
}
|
|
@@ -198,6 +206,11 @@ export function failureDetail(failures, failedCount) {
|
|
|
198
206
|
* the ONE disposition the file records. */
|
|
199
207
|
export function fileDisposition(states, gateReason, assertionCount, failures = []) {
|
|
200
208
|
const failed = states.filter((s) => s === 'fail').length;
|
|
209
|
+
// #1829, the file-level half of `resolveItRecord`'s rule: every failure was a
|
|
210
|
+
// driver transport loss and the file asserted nothing, so nothing was observed.
|
|
211
|
+
if (failed > 0 && assertionCount === 0 && failures.length === failed && failures.every((f) => f.message?.includes(TRANSPORT_LOSS_MARK) === true)) {
|
|
212
|
+
return { disposition: 'blocked', detail: `no response reached the suite in ${failed} test(s); first: ${failures[0].message.slice(failures[0].message.indexOf(TRANSPORT_LOSS_MARK)).slice(0, 300)}` };
|
|
213
|
+
}
|
|
201
214
|
if (failed > 0)
|
|
202
215
|
return { disposition: 'executed-fail', detail: failureDetail(failures, failed) };
|
|
203
216
|
if (states.some((s) => s === 'pass')) {
|
|
@@ -381,6 +394,44 @@ document) {
|
|
|
381
394
|
}
|
|
382
395
|
rows.push(row);
|
|
383
396
|
}
|
|
397
|
+
// Any-of requirements (RFC 0229 §E): one summary row per group, derived from
|
|
398
|
+
// its member files. Satisfied only by a WITNESSED pass (assertions > 0) with
|
|
399
|
+
// no member failing; `inapplicable`/`skipped` members never satisfy it. v1
|
|
400
|
+
// hand table only, like the prefix groups (the major-2 floors have none).
|
|
401
|
+
const groups = new Map();
|
|
402
|
+
//
|
|
403
|
+
// Only for a CLAIMED profile's floor (2.45.4). The row is `blocked` when no
|
|
404
|
+
// member was witnessed, and a v1 floor reads `inapplicable` as certifiable, so
|
|
405
|
+
// the row cannot simply say `inapplicable`. But 2.45.3 wrote it for every
|
|
406
|
+
// floor in the table: a host that does not advertise secrets records both
|
|
407
|
+
// members `inapplicable`, got one `blocked` row for a profile it never
|
|
408
|
+
// claimed, and a bundle with any `blocked` row certifies nothing (RFC 0168
|
|
409
|
+
// §E.1). A group nobody claims is not a requirement on this host.
|
|
410
|
+
const claimed = new Set(claimedProfiles.map((p) => DEPRECATED_PROFILE_ALIASES[p] ?? p));
|
|
411
|
+
if (!v2FloorsActive())
|
|
412
|
+
for (const [profile, floor] of Object.entries(PROFILE_FLOOR_SCENARIOS))
|
|
413
|
+
if (claimed.has(profile))
|
|
414
|
+
for (const g of floor.requiredAnyOf ?? [])
|
|
415
|
+
groups.set(requirementIdForAnyOf(g), g);
|
|
416
|
+
for (const [id, members] of [...groups.entries()].sort((a, b) => a[0].localeCompare(b[0]))) {
|
|
417
|
+
const scenarioId = `anyof:${members.join('|')}`;
|
|
418
|
+
const matching = members.map((f) => perFile.get(f)).filter((r) => r !== undefined);
|
|
419
|
+
const witnessed = matching.filter((r) => r.disposition === 'executed-pass' && (r.assertionCount ?? 0) > 0);
|
|
420
|
+
let row;
|
|
421
|
+
if (matching.some((r) => r.disposition === 'executed-fail')) {
|
|
422
|
+
row = { requirementId: id, scenarioId, disposition: 'executed-fail', detail: `an alternative failed: ${matching.filter((r) => r.disposition === 'executed-fail').map((r) => r.scenarioId).join(', ')}` };
|
|
423
|
+
}
|
|
424
|
+
else if (witnessed.length > 0) {
|
|
425
|
+
row = { requirementId: id, scenarioId, disposition: 'executed-pass', assertionCount: witnessed.reduce((n, r) => n + (r.assertionCount ?? 0), 0) };
|
|
426
|
+
}
|
|
427
|
+
else if (matching.length === 0) {
|
|
428
|
+
row = { requirementId: id, scenarioId, disposition: 'blocked', detail: `none of ${members.join(', ')} ran — unclassified return` };
|
|
429
|
+
}
|
|
430
|
+
else {
|
|
431
|
+
row = { requirementId: id, scenarioId, disposition: 'blocked', detail: `no alternative recorded a witnessed pass (${matching.map((r) => `${r.scenarioId}: ${r.disposition}`).join(', ')})` };
|
|
432
|
+
}
|
|
433
|
+
rows.push(row);
|
|
434
|
+
}
|
|
384
435
|
// Per-`it` rows (suite 1.153.0): every ledger entry keyed `openwop.it.<file>.<slug>`
|
|
385
436
|
// becomes its own bundle row, attributed to its scenario file. Additive — the
|
|
386
437
|
// file-level and prefix rows above are unchanged, and the floors still key on
|
|
@@ -440,7 +491,7 @@ document) {
|
|
|
440
491
|
const atMajor2 = v2FloorsActive();
|
|
441
492
|
for (const id of ids) {
|
|
442
493
|
const r = rowById.get(id);
|
|
443
|
-
const fromLedger = byId.has(id) || (r !== undefined && r.scenarioId.endsWith('*'));
|
|
494
|
+
const fromLedger = byId.has(id) || (r !== undefined && (r.scenarioId.endsWith('*') || r.scenarioId.startsWith('anyof:')));
|
|
444
495
|
if (r !== undefined && r.disposition === 'executed-pass' && (r.assertionCount ?? 0) > 0)
|
|
445
496
|
witnessedPasses += 1;
|
|
446
497
|
// Unclassified: no row, or a report-derived blocked (nothing recorded), or a
|
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-budget-tool-calls",
|
|
3
|
+
"name": "Conformance: Budget (tool calls)",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "Spends a run budget without a model or a seam. One `core.conformance.mock-agent` node makes three scripted tool calls, so a run created with `budget.maxToolCalls: 2` crosses its threshold on the first call and is exhausted before the third. See v2-budget-enforcement.test.ts.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "spender",
|
|
9
|
+
"typeId": "core.conformance.mock-agent",
|
|
10
|
+
"name": "Budget Spender",
|
|
11
|
+
"position": {
|
|
12
|
+
"x": 0,
|
|
13
|
+
"y": 0
|
|
14
|
+
},
|
|
15
|
+
"config": {
|
|
16
|
+
"mockToolCalls": [
|
|
17
|
+
{
|
|
18
|
+
"toolId": "openwop.echo:echo",
|
|
19
|
+
"arguments": {
|
|
20
|
+
"x": 1
|
|
21
|
+
},
|
|
22
|
+
"result": {
|
|
23
|
+
"x": 1
|
|
24
|
+
},
|
|
25
|
+
"durationMs": 1
|
|
26
|
+
},
|
|
27
|
+
{
|
|
28
|
+
"toolId": "openwop.echo:echo",
|
|
29
|
+
"arguments": {
|
|
30
|
+
"x": 2
|
|
31
|
+
},
|
|
32
|
+
"result": {
|
|
33
|
+
"x": 2
|
|
34
|
+
},
|
|
35
|
+
"durationMs": 1
|
|
36
|
+
},
|
|
37
|
+
{
|
|
38
|
+
"toolId": "openwop.echo:echo",
|
|
39
|
+
"arguments": {
|
|
40
|
+
"x": 3
|
|
41
|
+
},
|
|
42
|
+
"result": {
|
|
43
|
+
"x": 3
|
|
44
|
+
},
|
|
45
|
+
"durationMs": 1
|
|
46
|
+
}
|
|
47
|
+
],
|
|
48
|
+
"mockDecision": {
|
|
49
|
+
"decision": {
|
|
50
|
+
"next": "done"
|
|
51
|
+
},
|
|
52
|
+
"confidence": 1
|
|
53
|
+
}
|
|
54
|
+
},
|
|
55
|
+
"inputs": {},
|
|
56
|
+
"agent": {
|
|
57
|
+
"agentId": "core.conformance.budget-spender",
|
|
58
|
+
"modelClass": "reasoning"
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
],
|
|
62
|
+
"edges": [],
|
|
63
|
+
"triggers": [
|
|
64
|
+
{
|
|
65
|
+
"id": "manual",
|
|
66
|
+
"type": "manual",
|
|
67
|
+
"enabled": true
|
|
68
|
+
}
|
|
69
|
+
],
|
|
70
|
+
"variables": [],
|
|
71
|
+
"metadata": {
|
|
72
|
+
"tags": [
|
|
73
|
+
"conformance",
|
|
74
|
+
"budget",
|
|
75
|
+
"rfc-0084"
|
|
76
|
+
]
|
|
77
|
+
},
|
|
78
|
+
"settings": {
|
|
79
|
+
"timeout": 15000
|
|
80
|
+
}
|
|
81
|
+
}
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-fs-probe",
|
|
3
|
+
"name": "Conformance: v2 fs sandbox probe",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "Fixture for `v2-fs-sandbox-escape-refused.test.ts` (spec/v2/core/storage.md §`fs`). One node, `core.conformance.fs-probe` (conformance-RESERVED): a host that advertises `fs` maps it to a node that calls its OWN `ctx.fs` with the run's inputs and nothing else. `op: write` calls `write(path, utf8(content), \"text/plain\")` and completes with `outputs.result` = the resolved `{ path, sizeBytes }`; `op: read` calls `read(path)` and completes with `outputs.result = { content }`, the bytes decoded as UTF-8. The node passes `path` through UNCHANGED (no normalisation of its own, so the host's sandbox check is what is exercised). A rejection fails the node with the rejection's code and details, unchanged. Operator contract: `<sandboxRoot>/conformance/escape-link` is a symbolic link to a readable file OUTSIDE the root. A host that does not advertise `fs` MUST NOT advertise this fixture.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "fs-probe",
|
|
9
|
+
"typeId": "core.conformance.fs-probe",
|
|
10
|
+
"name": "Probe ctx.fs",
|
|
11
|
+
"position": { "x": 0, "y": 0 },
|
|
12
|
+
"config": {},
|
|
13
|
+
"inputs": {}
|
|
14
|
+
}
|
|
15
|
+
],
|
|
16
|
+
"edges": [],
|
|
17
|
+
"triggers": [{ "id": "manual", "type": "manual", "enabled": true }],
|
|
18
|
+
"variables": [
|
|
19
|
+
{ "name": "op", "type": "string", "defaultValue": "read" },
|
|
20
|
+
{ "name": "path", "type": "string" },
|
|
21
|
+
{ "name": "content", "type": "string" }
|
|
22
|
+
],
|
|
23
|
+
"metadata": { "tags": ["conformance", "storage", "fs", "security"] },
|
|
24
|
+
"settings": { "timeout": 0 }
|
|
25
|
+
}
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-queue-consume",
|
|
3
|
+
"name": "Conformance: Queue Consume",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "host-services.md §queueBus — node 'consume' (conformance.queue.probe, action 'consume') calls ctx.queueBus.consume on `topic` under the run's own tenant for up to config.waitMs, acks every message it receives, and outputs { consumed: [<payload.message of each, in order>] } — [] when none arrived, never unset. Paired with conformance-queue-publish. See v2-queue-cross-tenant-isolation.test.ts.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "consume",
|
|
9
|
+
"typeId": "conformance.queue.probe",
|
|
10
|
+
"name": "Consume",
|
|
11
|
+
"position": { "x": 0, "y": 0 },
|
|
12
|
+
"config": { "action": "consume", "waitMs": 1000 },
|
|
13
|
+
"inputs": {
|
|
14
|
+
"topic": { "type": "variable", "variableName": "topic" }
|
|
15
|
+
}
|
|
16
|
+
}
|
|
17
|
+
],
|
|
18
|
+
"edges": [],
|
|
19
|
+
"triggers": [
|
|
20
|
+
{ "id": "manual", "type": "manual", "enabled": true }
|
|
21
|
+
],
|
|
22
|
+
"variables": [
|
|
23
|
+
{ "name": "topic", "type": "string", "description": "Topic to consume from.", "required": true }
|
|
24
|
+
],
|
|
25
|
+
"metadata": { "tags": ["conformance", "queue", "security"] },
|
|
26
|
+
"settings": { "timeout": 10000, "maxRetries": 0 }
|
|
27
|
+
}
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-queue-publish",
|
|
3
|
+
"name": "Conformance: Queue Publish",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "host-services.md §queueBus — node 'publish' (conformance.queue.probe, action 'publish') calls ctx.queueBus.publish({ topic, payload: { message } }) under the run's own tenant and outputs { published: true }. Paired with conformance-queue-consume. See v2-queue-cross-tenant-isolation.test.ts.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "publish",
|
|
9
|
+
"typeId": "conformance.queue.probe",
|
|
10
|
+
"name": "Publish",
|
|
11
|
+
"position": { "x": 0, "y": 0 },
|
|
12
|
+
"config": { "action": "publish" },
|
|
13
|
+
"inputs": {
|
|
14
|
+
"topic": { "type": "variable", "variableName": "topic" },
|
|
15
|
+
"message": { "type": "variable", "variableName": "message" }
|
|
16
|
+
}
|
|
17
|
+
}
|
|
18
|
+
],
|
|
19
|
+
"edges": [],
|
|
20
|
+
"triggers": [
|
|
21
|
+
{ "id": "manual", "type": "manual", "enabled": true }
|
|
22
|
+
],
|
|
23
|
+
"variables": [
|
|
24
|
+
{ "name": "topic", "type": "string", "description": "Topic to publish on.", "required": true },
|
|
25
|
+
{ "name": "message", "type": "string", "description": "Opaque message string carried as payload.message.", "required": true }
|
|
26
|
+
],
|
|
27
|
+
"metadata": { "tags": ["conformance", "queue", "security"] },
|
|
28
|
+
"settings": { "timeout": 10000, "maxRetries": 0 }
|
|
29
|
+
}
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-safefetch-probe",
|
|
3
|
+
"name": "Conformance: v2 safeFetch SSRF probe",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "Fixture for `v2-safefetch-ssrf-refused.test.ts` (spec/v2/core/host-services.md §`httpClient`). One node, `core.conformance.safefetch-probe` (conformance-RESERVED): a host that advertises `httpClient.safeFetch` maps it to a node that calls its OWN `ctx.http.safeFetch(url)` with the run's `url` input, unchanged, and nothing else (method GET, no headers, no body). On success it completes with `outputs.result = { status }`, the response status only (never the body, so the fixture is not a read proxy). A rejection fails the node with the rejection's code and details, unchanged (a refused target is `egress_denied`, `details.reason: ssrf-blocked`; an unreachable one `upstream_unavailable`). A host that does not advertise `httpClient.safeFetch` MUST NOT advertise this fixture.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "safefetch-probe",
|
|
9
|
+
"typeId": "core.conformance.safefetch-probe",
|
|
10
|
+
"name": "Probe ctx.http.safeFetch",
|
|
11
|
+
"position": { "x": 0, "y": 0 },
|
|
12
|
+
"config": {},
|
|
13
|
+
"inputs": {}
|
|
14
|
+
}
|
|
15
|
+
],
|
|
16
|
+
"edges": [],
|
|
17
|
+
"triggers": [{ "id": "manual", "type": "manual", "enabled": true }],
|
|
18
|
+
"variables": [
|
|
19
|
+
{ "name": "url", "type": "string" }
|
|
20
|
+
],
|
|
21
|
+
"metadata": { "tags": ["conformance", "httpClient", "safeFetch", "ssrf", "security"] },
|
|
22
|
+
"settings": { "timeout": 0 }
|
|
23
|
+
}
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
{
|
|
2
|
+
"id": "conformance-secret-resolve-then-fail",
|
|
3
|
+
"name": "Conformance: Secret Resolve Then Fail",
|
|
4
|
+
"version": "1.0",
|
|
5
|
+
"description": "host-services.md §secrets — node 'resolve-secret' (conformance.secret.echo) resolves the host-provisioned canary 'openwop-conformance-canary-secret' exactly as openwop-smoke-byok-roundtrip does and outputs only {secretSha256, secretLength}; node 'fail' (core.fail) then fails the run. The canary MUST appear on no readable surface of the failed run: snapshot, error, events, or any error envelope. See v2-secret-canary-absent.test.ts.",
|
|
6
|
+
"nodes": [
|
|
7
|
+
{
|
|
8
|
+
"id": "resolve-secret",
|
|
9
|
+
"typeId": "conformance.secret.echo",
|
|
10
|
+
"name": "Resolve Canary Secret",
|
|
11
|
+
"position": { "x": 0, "y": 0 },
|
|
12
|
+
"config": { "secretId": "openwop-conformance-canary-secret" },
|
|
13
|
+
"inputs": {}
|
|
14
|
+
},
|
|
15
|
+
{
|
|
16
|
+
"id": "fail",
|
|
17
|
+
"typeId": "core.fail",
|
|
18
|
+
"name": "Fail After Resolve",
|
|
19
|
+
"position": { "x": 200, "y": 0 },
|
|
20
|
+
"config": { "message": "Intentional failure after resolving the conformance canary" },
|
|
21
|
+
"inputs": {}
|
|
22
|
+
}
|
|
23
|
+
],
|
|
24
|
+
"edges": [
|
|
25
|
+
{ "id": "resolve-then-fail", "sourceNodeId": "resolve-secret", "targetNodeId": "fail" }
|
|
26
|
+
],
|
|
27
|
+
"triggers": [
|
|
28
|
+
{ "id": "manual", "type": "manual", "enabled": true }
|
|
29
|
+
],
|
|
30
|
+
"variables": [],
|
|
31
|
+
"metadata": { "tags": ["conformance", "byok", "secrets", "security"] },
|
|
32
|
+
"settings": { "timeout": 10000, "maxRetries": 0 }
|
|
33
|
+
}
|