assertledger 1.0.0 → 1.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/README.fr.md +3 -3
  2. package/README.md +3 -3
  3. package/SECURITY.md +11 -6
  4. package/conformance/schema-extensions.json +36 -0
  5. package/dist/build-info.d.ts +16 -0
  6. package/dist/build-info.d.ts.map +1 -0
  7. package/dist/build-info.js +18 -0
  8. package/dist/build-info.js.map +1 -0
  9. package/dist/build-info.json +7 -0
  10. package/dist/cli.d.ts.map +1 -1
  11. package/dist/cli.js +117 -19
  12. package/dist/cli.js.map +1 -1
  13. package/dist/contracts/index.d.ts +943 -0
  14. package/dist/contracts/index.d.ts.map +1 -1
  15. package/dist/contracts/index.js +561 -21
  16. package/dist/contracts/index.js.map +1 -1
  17. package/dist/core/index.d.ts +10 -1
  18. package/dist/core/index.d.ts.map +1 -1
  19. package/dist/core/index.js +484 -8
  20. package/dist/core/index.js.map +1 -1
  21. package/dist/diagnostics.d.ts.map +1 -1
  22. package/dist/diagnostics.js +42 -2
  23. package/dist/diagnostics.js.map +1 -1
  24. package/dist/engine/adapters/node-test-runtime.d.ts +17 -0
  25. package/dist/engine/adapters/node-test-runtime.d.ts.map +1 -1
  26. package/dist/engine/adapters/node-test-runtime.js +45 -23
  27. package/dist/engine/adapters/node-test-runtime.js.map +1 -1
  28. package/dist/engine/container.d.ts +54 -0
  29. package/dist/engine/container.d.ts.map +1 -0
  30. package/dist/engine/container.js +464 -0
  31. package/dist/engine/container.js.map +1 -0
  32. package/dist/engine/git-regression.d.ts +12 -2
  33. package/dist/engine/git-regression.d.ts.map +1 -1
  34. package/dist/engine/git-regression.js +63 -16
  35. package/dist/engine/git-regression.js.map +1 -1
  36. package/dist/engine/index.d.ts +7 -1
  37. package/dist/engine/index.d.ts.map +1 -1
  38. package/dist/engine/index.js +300 -74
  39. package/dist/engine/index.js.map +1 -1
  40. package/dist/mcp/index.d.ts.map +1 -1
  41. package/dist/mcp/index.js +53 -1
  42. package/dist/mcp/index.js.map +1 -1
  43. package/dist/proof-planner/carry-over.d.ts +21 -0
  44. package/dist/proof-planner/carry-over.d.ts.map +1 -0
  45. package/dist/proof-planner/carry-over.js +212 -0
  46. package/dist/proof-planner/carry-over.js.map +1 -0
  47. package/dist/proof-planner/index.d.ts +11 -0
  48. package/dist/proof-planner/index.d.ts.map +1 -0
  49. package/dist/proof-planner/index.js +11 -0
  50. package/dist/proof-planner/index.js.map +1 -0
  51. package/dist/proof-planner/model.d.ts +272 -0
  52. package/dist/proof-planner/model.d.ts.map +1 -0
  53. package/dist/proof-planner/model.js +215 -0
  54. package/dist/proof-planner/model.js.map +1 -0
  55. package/dist/proof-planner/plan.d.ts +140 -0
  56. package/dist/proof-planner/plan.d.ts.map +1 -0
  57. package/dist/proof-planner/plan.js +1124 -0
  58. package/dist/proof-planner/plan.js.map +1 -0
  59. package/dist/proof-planner/policy.d.ts +100 -0
  60. package/dist/proof-planner/policy.d.ts.map +1 -0
  61. package/dist/proof-planner/policy.js +407 -0
  62. package/dist/proof-planner/policy.js.map +1 -0
  63. package/dist/proof-planner/render.d.ts +4 -0
  64. package/dist/proof-planner/render.d.ts.map +1 -0
  65. package/dist/proof-planner/render.js +73 -0
  66. package/dist/proof-planner/render.js.map +1 -0
  67. package/dist/sdk/index.d.ts +16 -6
  68. package/dist/sdk/index.d.ts.map +1 -1
  69. package/dist/sdk/index.js +63 -5
  70. package/dist/sdk/index.js.map +1 -1
  71. package/docs/agentic-test-profile.md +5 -0
  72. package/docs/architecture.md +15 -3
  73. package/docs/ci.md +7 -1
  74. package/docs/conformance-v1.md +11 -2
  75. package/docs/container-isolation.md +156 -0
  76. package/docs/evidence-export.md +146 -0
  77. package/docs/git-regression.md +7 -3
  78. package/docs/migration-timeout-discovery-inconclusive.md +126 -0
  79. package/docs/migration-verification-v2.md +55 -0
  80. package/docs/project-intent.md +3 -2
  81. package/docs/proof-model.md +28 -6
  82. package/docs/proof-planner.md +368 -0
  83. package/docs/reference.md +59 -12
  84. package/docs/roadmap.md +12 -4
  85. package/examples/agentic-profile/profile-manifest.mjs +5 -1
  86. package/examples/evidence-export/consumer.mjs +35 -0
  87. package/package.json +5 -5
  88. package/schemas/evidence-export-replay-result.v1.json +49 -0
  89. package/schemas/evidence-export-request.v1.json +681 -0
  90. package/schemas/evidence-export.v1.json +1479 -0
  91. package/schemas/evidence-manifest.v2.json +886 -0
  92. package/schemas/evidence-provider-manifest.v1.json +269 -0
  93. package/schemas/verification-request.v2.json +457 -0
@@ -0,0 +1,156 @@
1
+ # Container isolation
2
+
3
+ [Back to the reference](reference.md) · [Security policy](../SECURITY.md)
4
+
5
+ Verification request v2 adds a `container` isolation backend. Every control and candidate
6
+ execution then runs in a fresh Linux container created from a digest-pinned image that is already
7
+ present on an operator-administered Docker-compatible daemon. `trusted-local` stays available in v1
8
+ and v2, remains explicitly `UNSANDBOXED`, and still requires external authorization.
9
+
10
+ Container isolation is a containment layer around an execution, not a proof that a campaign is
11
+ correct. The daemon, its host kernel, the image and the AssertLedger engine remain trusted.
12
+
13
+ ## Select the backend
14
+
15
+ The Git workflow selects the backend explicitly. Without `--container-image` or
16
+ `--allow-unsafe-execution`, `check` stops before reading the repository and names both options:
17
+
18
+ ```sh
19
+ assertledger check . --before BEFORE --after AFTER --neutral NEUTRAL --neutral-reason "explicit control reason" --test tests/regression.test.js --base-test tests/base.test.js --out .assertledger/evidence --container-image node@sha256:DIGEST
20
+ ```
21
+
22
+ `verify` runs a v2 request whose `isolation.kind` is `container` without
23
+ `--allow-unsafe-execution`:
24
+
25
+ ```json
26
+ {
27
+ "schemaVersion": "2.0.0",
28
+ "isolation": {
29
+ "kind": "container",
30
+ "image": "node@sha256:DIGEST",
31
+ "environment": [{ "name": "TZ", "value": "UTC" }],
32
+ "limits": {
33
+ "memoryBytes": 1073741824,
34
+ "cpuMillicores": 2000,
35
+ "pids": 256,
36
+ "temporaryDirectoryBytes": 67108864
37
+ }
38
+ }
39
+ }
40
+ ```
41
+
42
+ The other request fields keep their v1 meaning; `assertledger schema verification-request-v2 --json`
43
+ prints the complete contract. `check` uses the limits shown above and a 30-second timeout per
44
+ execution.
45
+
46
+ Combining a container image or request with `--allow-unsafe-execution` fails with
47
+ `ISOLATION_MODE_CONFLICT` before any runtime call. Combining `--container-runtime` with a
48
+ trusted-local request fails the same way.
49
+
50
+ The runtime command belongs to the operator, never to the request. It defaults to `["docker"]` and
51
+ is passed as a JSON argument array; it is never interpreted by a shell:
52
+
53
+ ```sh
54
+ assertledger verify request.json --container-runtime '["docker"]' --json
55
+ ```
56
+
57
+ On Windows with Docker Engine inside WSL, use
58
+ `--container-runtime '["wsl.exe","-d","Ubuntu","--exec","docker"]'` from a POSIX shell. Shells that
59
+ strip inner double quotes from native arguments, such as Windows PowerShell 5.1, need escaped
60
+ quotes. The SDK takes the same values through its v2 entry points, which return v2 manifests:
61
+
62
+ ```js
63
+ const ledger = new AssertLedger();
64
+ await ledger.verifyV2(request, { containerRuntime: { command: ["docker"] } });
65
+ await ledger.checkGitRegressionV2({ ...options, container: { image: "node@sha256:DIGEST" } });
66
+ ```
67
+
68
+ `verify()` and `checkGitRegression()` keep their v1 contracts and result types.
69
+
70
+ The MCP `check` and `verify` tools still execute only trusted-local v1 requests.
71
+
72
+ ## Backend checks before execution
73
+
74
+ AssertLedger queries the runtime before creating a workspace container. Each refusal happens before
75
+ any repository code runs and has an [`explain`](diagnostics.md) entry:
76
+
77
+ | Reason code | Meaning |
78
+ | --- | --- |
79
+ | `CONTAINER_RUNTIME_COMMAND_INVALID` | The runtime argv is not a JSON array of 1 to 16 non-empty strings. |
80
+ | `CONTAINER_RUNTIME_NOT_FOUND` | The runtime executable cannot be started. |
81
+ | `CONTAINER_RUNTIME_UNAVAILABLE` | The runtime reports no reachable daemon. |
82
+ | `CONTAINER_RUNTIME_PLATFORM_UNSUPPORTED` | The daemon does not run Linux containers. |
83
+ | `CONTAINER_IMAGE_REFERENCE_INVALID` | The `check` image is not pinned by `@sha256:`; a v2 request with such an image fails request validation. |
84
+ | `CONTAINER_IMAGE_NOT_PRESENT` | The pinned image is absent locally. AssertLedger never pulls. |
85
+ | `CONTAINER_IMAGE_DIGEST_MISMATCH` | The local image does not carry the pinned digest. |
86
+ | `CONTAINER_IMAGE_PLATFORM_UNSUPPORTED` | The image is not a Linux image. |
87
+ | `CONTAINER_CLEANUP_FAILED` | A finished container could not be removed; no result is produced. |
88
+
89
+ Review an image, then pull it yourself by the same digest, for example
90
+ `docker pull node@sha256:DIGEST`. For `node:test`, the engine also probes `node` inside the image
91
+ twice and records its path, version and SHA-256 digest, as it does locally.
92
+
93
+ ## Controls applied to every execution
94
+
95
+ | Boundary | Control |
96
+ | --- | --- |
97
+ | Files | No host path is mounted. The prepared workspace is streamed as an archive into an anonymous volume; the root file system is read-only; `/tmp` is a `tmpfs` bounded by `temporaryDirectoryBytes`. |
98
+ | Result | Only a regular file named `result.json`, no larger than the report bound, is read back as an archive in memory. A link, a renamed entry or a malformed archive yields `INFRA_ERROR`. |
99
+ | Identity | User and group `65534`, all capabilities dropped, `no-new-privileges`. |
100
+ | Network | `--network=none`: only the loopback interface exists. |
101
+ | Environment | The image's own variables, the variables the runtime sets itself such as `HOSTNAME` and `HOME`, the declared `environment` entries and the adapter protocol variables. The host environment is never forwarded. |
102
+ | Resources | `pids`, `memoryBytes` with swap disabled, and `cpuMillicores` limits. |
103
+ | Time | On timeout the container is killed, which ends every process in its PID namespace, including detached descendants. |
104
+ | Output | Standard output and error are bounded by `maximumOutputBytes`; the runtime log driver is disabled. |
105
+ | Lifetime | One container per execution, labeled `assertledger.execution`, removed with its volume afterwards. |
106
+
107
+ A timeout is recorded as `TIMEOUT`. A process stopped by the memory or process limit exits without a
108
+ valid report and is recorded as an infrastructure or process failure. Neither can count as target
109
+ detection: only an attributed `ASSERTION_FAILURE` kills a target.
110
+
111
+ ## Evidence
112
+
113
+ A v2 manifest binds the backend facts to the decision digest in
114
+ `evidenceContext.execution.backend`: image reference, identifier, OS and architecture; runtime
115
+ client and server versions, server OS and architecture, cgroup version and reported security
116
+ options; and the controls, declared environment and limits above. The top-level
117
+ `isolation` summary records `kind: "container"`, `level: "CONTAINER"` and the runtime argv; it is
118
+ covered by the artifact digest. The manifest's `limitations` state the main limits of this backend.
119
+
120
+ Declared environment values are recorded verbatim. Never place a secret in `isolation.environment`.
121
+
122
+ Replay validates v1 and v2 manifests. Evidence export, Agentic Test Profiles and benchmarks still
123
+ accept only v1 manifests. The v1 request and manifest contracts are unchanged.
124
+
125
+ ## Limits and non-claims
126
+
127
+ - Containers share the daemon host's kernel; this is not a virtual machine boundary.
128
+ - Access to the Docker daemon is equivalent to administrative access on its host. Run it on a host
129
+ without secrets, as you would a CI runner.
130
+ - AssertLedger does not bound the writable workspace volume. Storage quotas remain an
131
+ operator-administered daemon setting.
132
+ - `/dev/shm` keeps the runtime default size (64 MiB on Docker Engine), charged to the memory limit.
133
+ - The image and declared environment values are operator inputs. AssertLedger verifies the image
134
+ digest, not its contents. With no network, the image must already contain every runtime the
135
+ adapter needs.
136
+ - An interrupted AssertLedger process can leave a labeled container. Remove it with
137
+ `docker ps --all --filter label=assertledger.execution` followed by `docker rm --force --volumes`.
138
+ - Only Linux daemons are supported. The host side is platform-neutral Node.js: CI runs the backend
139
+ selection, diagnostic, archive and cleanup tests against a fake runtime on Linux, Windows and
140
+ macOS. The hostile scenario suite runs against real Docker Engine on Linux in CI and was run
141
+ locally against Docker Engine in WSL on Windows. GitHub-hosted Windows and macOS runners provide
142
+ no Linux daemon, so no real daemon is exercised from a macOS host.
143
+
144
+ ## Real-daemon test suite
145
+
146
+ `tests/container-isolation-docker.test.ts` runs hostile scenarios against a real daemon: network
147
+ egress, root file system writes, host paths, environment leakage, privileges, temporary storage
148
+ and process count, a detached descendant after timeout, a result link, a memory limit, an absent
149
+ image and an end-to-end `check`. It is skipped unless both variables below are set, and
150
+ `ASSERTLEDGER_REQUIRE_CONTAINER_TESTS=1` turns a missing configuration into a failure:
151
+
152
+ ```sh
153
+ ASSERTLEDGER_CONTAINER_RUNTIME='["docker"]' ASSERTLEDGER_CONTAINER_IMAGE='node@sha256:DIGEST' pnpm test
154
+ ```
155
+
156
+ The suite never pulls; provide the image first.
@@ -0,0 +1,146 @@
1
+ # Interoperable evidence export
2
+
3
+ AssertLedger can hand its results to an external engine, such as a change-governance tool, as
4
+ typed evidence. The export is a thin, deterministic projection of a replay-valid
5
+ [evidence manifest](proof-model.md). It adds no gate, no authority, and no runtime dependency on
6
+ any consumer: AssertLedger behaves identically whether a consumer exists, accepts the evidence,
7
+ degrades it, or ignores it.
8
+
9
+ Three public contracts are involved, each with a versioned JSON Schema:
10
+
11
+ | Contract | Schema | Produced by |
12
+ | --- | --- | --- |
13
+ | Provider manifest | [`evidence-provider-manifest.v1.json`](../schemas/evidence-provider-manifest.v1.json) | `assertledger provider`, `AssertLedger.providerManifest()`, `assertledger_provider` |
14
+ | Export request | [`evidence-export-request.v1.json`](../schemas/evidence-export-request.v1.json) | The consumer or operator |
15
+ | Evidence export | [`evidence-export.v1.json`](../schemas/evidence-export.v1.json) | `assertledger export`, `AssertLedger.exportEvidence()`, `assertledger_export` |
16
+ | Export replay result | [`evidence-export-replay-result.v1.json`](../schemas/evidence-export-replay-result.v1.json) | `assertledger export-replay`, `AssertLedger.replayEvidenceExport()`, `assertledger_export_replay` |
17
+
18
+ ## Provider manifest: what is announced
19
+
20
+ The provider manifest describes the installed provider: its version, the source revision recorded
21
+ at build time, its scope, the formats it accepts and emits, its capabilities, its supported
22
+ adapters, its cost model, and its limits. It carries a `manifestDigest` over its canonical
23
+ projection.
24
+
25
+ The source revision is `RECORDED` with the commit and a `CLEAN` or `DIRTY` worktree only when the
26
+ package build captured it (`dist/build-info.json`). A source checkout or a build without Git reports
27
+ `UNKNOWN`; the current checkout never fills in a missing value.
28
+
29
+ Capabilities are announcements. `CONTROL_WITHOUT_CANDIDATE`, `REFERENCE_PASS`,
30
+ `REGRESSION_DETECTION`, `NEUTRAL_PASS`, and `STABILITY_REPETITION` are `SUPPORTED` test-observed
31
+ controls. `GIT_REVISION_PROVENANCE` is `SUPPORTED_WHEN_RECORDED`. `EXECUTION_FRESHNESS`,
32
+ `PRODUCER_AUTHENTICATION`, and `SANDBOXED_EXECUTION` are `UNSUPPORTED`. A capability never proves
33
+ that a control ran: only the recorded observations of an exported manifest do.
34
+
35
+ ## Export request
36
+
37
+ ```json
38
+ {
39
+ "schemaVersion": "1.0.0",
40
+ "manifest": { "...": "a replay-valid evidence manifest v1" },
41
+ "consumerRequest": {
42
+ "reference": "change-42",
43
+ "profileId": null,
44
+ "obligations": [
45
+ { "id": "detect", "control": "REGRESSION_DETECTION" },
46
+ { "id": "sandbox", "control": "SANDBOXED_EXECUTION" }
47
+ ]
48
+ }
49
+ }
50
+ ```
51
+
52
+ `consumerRequest` may be `null`. Obligation identifiers must be unique. The request is not a
53
+ policy for AssertLedger: it only lets the export state which requested controls were executed,
54
+ not executed, or unsupported. No particular plan format is required.
55
+
56
+ The export refuses a manifest that fails replay with `EVIDENCE_EXPORT_SOURCE_INVALID` (CLI exit
57
+ code `4`, no output). A malformed request fails with `EVIDENCE_EXPORT_REQUEST_INVALID`, as does a
58
+ request embedding a v2 manifest from [container isolation](container-isolation.md): this export
59
+ version accepts only v1 manifests.
60
+
61
+ ## Evidence export: what was observed
62
+
63
+ The export embeds its `sourceManifest` and is self-contained. Its sections are deliberately
64
+ separate:
65
+
66
+ - `result`: the verbatim campaign decision, a per-candidate and per-world view, and a single
67
+ conservative `detection` with a stable `reasonCode`.
68
+ - `integrity`: the replay rails that were verified before export and the bound digests (artifact,
69
+ decision, repository, policy, and world digests).
70
+ - `authenticity`: always `UNAUTHENTICATED` with attestation `NONE`; the producer name and version
71
+ are only declared by the manifest.
72
+ - `environment`: `trusted-local`/`UNSANDBOXED` isolation, the environment allowlist names, and the
73
+ adapter. The framework is `RECORDED` only for a `node:test` manifest that recorded its official
74
+ adapter profile; otherwise it stays `UNKNOWN`.
75
+ - `confidence`: the level `REPLAY_CONSISTENT_UNAUTHENTICATED`, with the properties that replay
76
+ established and those it did not (producer authenticity, observation truthfulness, execution
77
+ isolation, execution freshness, and the semantic relevance of the declared worlds).
78
+ - `scope`: attempts, candidate and observation counts, and each world with its digest, declared
79
+ provenance, and Git revision. Git commits and trees are exported only when the provenance is the
80
+ exact canonical record written by [Git regression qualification](git-regression.md);
81
+ `gitRevisions` is `RECORDED`, `PARTIAL`, or `NOT_RECORDED`.
82
+ - `controls`: executed controls with their observation counts, requested obligations with
83
+ `EXECUTED`, `NOT_EXECUTED`, or `UNSUPPORTED` coverage, and gates that did not run.
84
+ - `profile`: `NOT_REQUESTED`, or `UNKNOWN_PROFILE` for any requested profile identifier. No usage
85
+ profile is defined yet, so none is ever applied or inferred.
86
+ - `policy`: the effective AssertLedger policy version and digest.
87
+ - `cost`: the estimated process executions derived from the recorded campaign shape, the observed
88
+ executions and recorded wall time with its coverage, and execution freshness `UNKNOWN` with cache
89
+ provenance `NOT_RECORDED`, because evidence manifest v1 does not record them.
90
+
91
+ ### Detection mapping
92
+
93
+ The first matching row applies.
94
+
95
+ | Situation | `detection` | `modality` | `reasonCode` |
96
+ | --- | --- | --- | --- |
97
+ | Campaign `VERIFIED` | `OBSERVED` | `TEST_OBSERVED` | `REGRESSION_ASSERTION_OBSERVED` |
98
+ | Campaign `ENGINE_ERROR` | `NOT_ESTABLISHED` | `NONE` | `ENGINE_ERROR` |
99
+ | Invalid controls | `NOT_ESTABLISHED` | `NONE` | `CONTROL_EVIDENCE_INVALID` |
100
+ | An `UNSTABLE` or `INCONCLUSIVE` candidate | `NOT_ESTABLISHED` | `NONE` | `CANDIDATE_EVIDENCE_INCONCLUSIVE` |
101
+ | An `INVALID` candidate | `NOT_ESTABLISHED` | `NONE` | `CANDIDATE_EVIDENCE_INVALID` |
102
+ | A target outcome that is neither a stable attributed `PASS` nor a stable attributed assertion | `NOT_ESTABLISHED` | `NONE` | `OPERATIONAL_OUTCOME_NOT_DETECTION` |
103
+ | At least one target observed passing | `NOT_OBSERVED` | `TEST_OBSERVED` | `TARGET_PASSED_WITHOUT_DETECTION` |
104
+ | Otherwise, such as every target killed without meeting the policy | `NOT_ESTABLISHED` | `NONE` | `TARGET_STRENGTH_INSUFFICIENT` |
105
+
106
+ The targeted test is the candidate: its `id` and content `digest`, the SHA-256 of the canonical JSON
107
+ of its `[{ "path", "content" }]` overlay files. Evidence manifest v1 does not record the file paths
108
+ themselves; a consumer holding the verification request (for example `executed-request.json`
109
+ written by `assertledger check`) can recompute the digest to bind paths to the exported evidence.
110
+
111
+ Per world, `signal` is `RED` only for complete, stable, attributed `ASSERTION_FAILURE` runs and
112
+ `GREEN` only for complete, stable, attributed `PASS` runs. A target world's `detection` is
113
+ `OBSERVED` or `NOT_OBSERVED` only when controls are valid and the candidate is `ELIGIBLE` or
114
+ `WEAK_ORACLE`. Compilation, collection, crash, timeout, infrastructure, and no-test outcomes are
115
+ never promoted to detection evidence in either direction.
116
+
117
+ ## Determinism and replay
118
+
119
+ The same manifest and consumer request always produce the same export bytes after canonical
120
+ serialization, independent of JSON key order; obligations are ordered by identifier. The
121
+ `exportDigest` covers the whole export except itself.
122
+
123
+ `assertledger export-replay` validates the export schema, replays the embedded source manifest,
124
+ recomputes `exportDigest`, and rebuilds the export from the embedded manifest and consumer request.
125
+ `valid` requires all four rails. It exits with `0` when valid and `4` otherwise. A re-digested
126
+ forgery keeps `exportDigestValid` true but fails `semanticsValid`.
127
+
128
+ ## Consumer responsibilities
129
+
130
+ A consumer decides admissibility with its own obligations. The
131
+ [`consumer example`](../examples/evidence-export/consumer.mjs) rejects an export that fails replay,
132
+ ignores exports without observed detection, and degrades observed detection to advisory when
133
+ controls or Git revisions are missing or the evidence is unauthenticated, which is always the case
134
+ for this version. Rejecting, degrading, or ignoring an export changes nothing in AssertLedger.
135
+
136
+ ## Non-claims
137
+
138
+ The export does not rerun tests, authenticate the producer, attest isolation, prove that recorded
139
+ observations were truthful, prove freshness, or make declared worlds semantically relevant. It
140
+ covers only the recorded worlds, candidates, and attempts.
141
+
142
+ ## Versioning
143
+
144
+ These four schemas are additive to the frozen [conformance v1](conformance-v1.md) set and are
145
+ locked by `conformance/schema-extensions.json`. Any change to their bytes, their mapping, or their
146
+ digest projections requires a new schema version, a migration note, and compatibility tests.
@@ -11,14 +11,18 @@ assertledger check . --before BEFORE --after AFTER --neutral NEUTRAL --neutral-r
11
11
 
12
12
  `--after` utilise `HEAD` par défaut. `--base-test` peut être répété. Le dossier donné à `--out` doit être un nouveau chemin relatif au dépôt. AssertLedger le réserve de manière exclusive, écrit `executed-request.json` et `summary.md`, puis publie `manifest.json` en dernier par renommage atomique. La présence de `manifest.json` est le marqueur de complétion ; un lecteur doit ignorer un dossier qui ne le contient pas.
13
13
 
14
- L’exécution `trusted-local` est volontairement **UNSANDBOXED**. Le drapeau `--allow-unsafe-execution` constitue l’autorisation distincte de l’opérateur. L’API équivalente est `new AssertLedger().checkGitRegression(options)` et exige `allowUnsafeExecution: true`.
14
+ Le mode d’exécution est toujours choisi explicitement. Sans `--container-image` ni `--allow-unsafe-execution`, `check` refuse de s’exécuter ; les deux ensemble sont refusés avec `ISOLATION_MODE_CONFLICT`.
15
15
 
16
- MCP expose le même parcours avec `assertledger_check` (alias `testforge_check`) seulement si
16
+ Pour isoler l’exécution, remplacez `--allow-unsafe-execution` par `--container-image NOM@sha256:DIGEST`. Chaque contrôle et chaque candidat s’exécute alors dans un conteneur Linux neuf, sans réseau ni montage de l’hôte, avec les limites par défaut et un délai de 30 secondes par exécution. L’image doit déjà être présente sur le démon et contenir `node` : AssertLedger ne la télécharge jamais. `--container-runtime` accepte la commande du runtime sous forme de tableau JSON, `["docker"]` par défaut. Le manifeste produit est alors en version `2.0.0`. Voir [l’isolation par conteneur](container-isolation.md).
17
+
18
+ L’exécution `trusted-local` est volontairement **UNSANDBOXED**. Le drapeau `--allow-unsafe-execution` constitue l’autorisation distincte de l’opérateur. Côté SDK, `new AssertLedger().checkGitRegression(options)` exige `allowUnsafeExecution: true` et renvoie un manifeste v1 ; `checkGitRegressionV2(options)` exige `container: { image }` et renvoie un manifeste v2.
19
+
20
+ MCP expose le parcours `trusted-local` avec `assertledger_check` (alias `testforge_check`) seulement si
17
21
  l’opérateur a démarré le serveur avec `--allow-unsafe-execution`. L’entrée reprend les options
18
22
  du SDK, sans le champ de permission : `repository`, `before`, `after` facultatif, `neutral`,
19
23
  `neutralReason`, `test`, `baseTests` et `out`. La racine est confinée aux dépôts autorisés ; le
20
24
  dossier de sortie suit les mêmes contrôles que la CLI. Un client ne peut pas s’accorder cette
21
- permission dans son message.
25
+ permission dans son message. Le mode conteneur n’est pas encore exposé par MCP.
22
26
 
23
27
  ## Limites de cette première tranche
24
28
 
@@ -0,0 +1,126 @@
1
+ # Timeouts and infrastructure errors no longer fail discovery
2
+
3
+ A candidate run that ends in `TIMEOUT` or `INFRA_ERROR` never reached a verdict. It can neither
4
+ prove nor disprove that the candidate test was discovered and attributed. The core now classifies
5
+ such a candidate as `INCONCLUSIVE` instead of `INVALID`. This resolves an ambiguity of the
6
+ [proof model](proof-model.md), which listed the case both under `INVALID` (discovery failed) and
7
+ under `INCONCLUSIVE` (`TIMEOUT` or `INFRA_ERROR`); the core applied the first.
8
+
9
+ ## What changed
10
+
11
+ The engine reports every `TIMEOUT` and `INFRA_ERROR` observation with `attributed: false` and,
12
+ normally, `candidateTestsDiscovered: 0`. The `DISCOVERY` gate requires every candidate run to report
13
+ an attributed candidate test, so those runs made it fail, and the candidate became `INVALID` before
14
+ the core looked at execution outcomes. A candidate that hangs on the faulty code was reported as a
15
+ broken test, and the campaign as `REJECTED`.
16
+
17
+ Once completeness and stability hold, `DISCOVERY` now distinguishes the runs that fail it:
18
+
19
+ | Runs that fail discovery | Before | After |
20
+ | --- | --- | --- |
21
+ | None | `DISCOVERY` passed | unchanged |
22
+ | At least one completed run (`PASS`, `ASSERTION_FAILURE`, compile, collection, crash or no-test outcome) | `INVALID`, `CANDIDATE_DISCOVERY_INVALID` | unchanged |
23
+ | Only `TIMEOUT` or `INFRA_ERROR` runs | `INVALID`, `CANDIDATE_DISCOVERY_INVALID` | `INCONCLUSIVE`, `CANDIDATE_EXECUTION_INCONCLUSIVE` |
24
+
25
+ In the last case the `DISCOVERY` gate is `FAILED` with reason `CANDIDATE_EXECUTION_INCONCLUSIVE`,
26
+ and `REFERENCE`, `NEUTRAL` and `TARGET_STRENGTH` stay `NOT_RUN` with `PREREQUISITE_GATE_FAILED`, as
27
+ after any failed prerequisite. The candidate kills nothing: a red reference or neutral world, or a
28
+ target its other runs killed, is not recorded as a failed gate or a kill, whereas a timeout reported
29
+ with an attributed candidate test, a core input the engine does not produce, still lets those gates
30
+ run. Once its attempts are complete and agree, a completed run that disproves discovery still makes
31
+ the candidate `INVALID`, whatever else timed out; missing attempts still make it `INCONCLUSIVE` and
32
+ attempts that disagree `UNSTABLE` first. Inconclusive execution still takes precedence over a red
33
+ reference or neutral world, as it already did for timeouts that passed discovery.
34
+
35
+ ## Observable effects
36
+
37
+ For a campaign where a candidate falls in the last row:
38
+
39
+ - candidate status `INVALID` becomes `INCONCLUSIVE`, with the reason codes above, whatever the
40
+ campaign decision, including in the candidate list of evidence exports built from the manifest
41
+ and in the evidence status profile v1 reports for that candidate;
42
+ - when no other candidate is selected, the campaign decision `REJECTED` / `NO_ELIGIBLE_CANDIDATE`
43
+ becomes `INCONCLUSIVE` / `CANDIDATE_EVIDENCE_INCONCLUSIVE`, and the CLI exit code of `verify` and
44
+ `check` becomes `3` instead of `2`; in that case the evidence export reason code becomes
45
+ `CANDIDATE_EVIDENCE_INCONCLUSIVE` instead of `CANDIDATE_EVIDENCE_INVALID`. A campaign with a
46
+ selected candidate stays `VERIFIED`,
47
+ invalid controls keep their precedence (`CONTROL_EVIDENCE_INVALID`), and a campaign that another
48
+ unstable or inconclusive candidate already made `INCONCLUSIVE` keeps its decision;
49
+ - the `decisionDigest` and `artifactDigest` values of the manifest change. The digest projections
50
+ do not.
51
+
52
+ The core does not look at where those outcomes come from. They include a candidate that hangs on a
53
+ target world and a structured-command adapter that dies without a report, writes a malformed report
54
+ or reports `INFRA_ERROR` itself (each witnessed below), as well as other engine paths such as a run
55
+ that ends without an exit code and without a valid report (for the structured-command adapter, with
56
+ any report but `PASS`), a `node:test` run without a valid report, a contradictory report, a process
57
+ that fails to start or a container execution failure. A candidate that makes its own run end in
58
+ `INFRA_ERROR`, for example by forcing a zero exit code despite failing tests, is therefore
59
+ inconclusive rather than invalid. None of them can be selected: only an `ELIGIBLE` candidate is, so
60
+ no `VERIFIED` decision changes.
61
+
62
+ ## Replay of existing manifests
63
+
64
+ Replay recomputes the decision from the recorded observations, and no field distinguishes the
65
+ decision semantics before and after this change (`policyVersion` stays `1.0.0`). Therefore:
66
+
67
+ - a manifest sealed before this change, with at least one candidate in the last row above, no
68
+ longer replays: `decisionDigestValid` and `artifactDigestValid` stay `true`,
69
+ `decisionSemanticsValid` and `valid` become `false`. This holds per candidate, not per decision:
70
+ a `VERIFIED` manifest with such a neighbouring candidate is affected too, and so are the evidence
71
+ exports, profiles and benchmarks built from it;
72
+ - a manifest sealed after this change with such a candidate fails replay the same way under a
73
+ verifier that predates it (1.1.0 and earlier).
74
+
75
+ Re-decide affected manifests from their observations with the release that includes this change,
76
+ and replay them with a verifier of that release or a later one. Manifests without a candidate in
77
+ the last row replay as before.
78
+
79
+ ## Compatibility decision
80
+
81
+ `CONTRIBUTING.md` asks for a schema-version discussion before a contract-breaking change. This
82
+ change keeps `schemaVersion` and `policyVersion` unchanged:
83
+
84
+ - the previous decision contradicted the proof model's own rule that `TIMEOUT` and `INFRA_ERROR`
85
+ are inconclusive. Campaign requests declare `policyVersion`, and the core accepts only `1.0.0`;
86
+ versioning would make the core accept a second version that requests opt into, while every
87
+ request that does not opt in would keep the misclassification;
88
+ - replay fails closed: an affected manifest never replays as valid under the other semantics, so no
89
+ verdict is silently reinterpreted, and every other manifest replays unchanged.
90
+
91
+ The cost of this choice: `policyVersion` `1.0.0` no longer names a single decision function, so
92
+ archived manifests of the affected case, including `VERIFIED` ones with an affected neighbour, stop
93
+ replaying; and a candidate can turn its own `INVALID` status into `INCONCLUSIVE` by ending its run
94
+ in `INFRA_ERROR`, which can turn the campaign decision from `REJECTED` into `INCONCLUSIVE` (exit
95
+ code 3 instead of 2), never into a selection.
96
+
97
+ The alternative is a `policyVersion` bump under which `1.0.0` manifests keep replaying with the old
98
+ rule. Choosing it would replace this section and the replay section above.
99
+
100
+ ## What does not change
101
+
102
+ No schema version, manifest field, gate order, digest projection, adapter protocol, diagnostic
103
+ catalogue entry or unsafe-execution permission changes, and no reason code is added or removed;
104
+ `CANDIDATE_EXECUTION_INCONCLUSIVE` can now appear on the `DISCOVERY` gate. Its explanation already
105
+ asks to fix the timeout or infrastructure failure and rerun; the campaign explanation of
106
+ `CANDIDATE_EVIDENCE_INCONCLUSIVE` still speaks of unstable or incomplete evidence, as it did for
107
+ timeouts after discovery. The conformance v1 bundle and its lock are unchanged: every candidate run
108
+ in its fixtures reports an attributed candidate test, so none reaches the new case. The self-hosted
109
+ core campaign keeps its mutation anchors and its existing-test snapshot.
110
+
111
+ ## Compatibility witnesses
112
+
113
+ - `tests/core-inconclusive-discovery.test.ts` fails against the previous core for `TIMEOUT` and
114
+ `INFRA_ERROR` runs as the engine reports them, a candidate that hangs everywhere, a hung neighbour
115
+ of an eligible candidate, and a red reference with a timed-out target. It also pins, to the values
116
+ the previous core produced for the same host-independent fixtures, the decision digests of the
117
+ inputs whose verdict must not change (a completed run that disproves discovery beside a `TIMEOUT`
118
+ or an `INFRA_ERROR`, unattributed assertion, compile, collection and crash outcomes, a completed
119
+ crash reported without an exit code, no test discovered, diverging timeouts, a timed-out control,
120
+ an attributed timeout), and shows that a manifest sealed before this change, rebuilt with both of
121
+ its digests, fails replay only on decision semantics. The reverse direction was observed by
122
+ replaying a manifest of this change with the previous core, which the repository tests cannot
123
+ import.
124
+ - `tests/engine.test.ts` runs a real `node:test` candidate that hangs on the target world for every
125
+ attempt, and a structured-command adapter that dies without a report, writes a malformed report
126
+ or reports `INFRA_ERROR` on the target. All four fail against the previous core and pass now.
@@ -0,0 +1,55 @@
1
+ # Verification request and evidence manifest v2
2
+
3
+ Version `2.0.0` of the verification request and evidence manifest adds
4
+ [container isolation](container-isolation.md). It is a new major contract beside v1, not a change
5
+ to v1.
6
+
7
+ ## What stays the same
8
+
9
+ - The 34 v1 schema bytes, the conformance-v1 lock root, v1 decision and artifact digest
10
+ projections, outcomes, gates and reason-code semantics are unchanged.
11
+ - A v1 request still produces a v1 manifest, and v1 manifests replay exactly as before.
12
+ - `parseVerificationRequest()` and `parseEvidenceManifest()` still accept only v1 and reject v2.
13
+ - The SDK methods `verify()` and `checkGitRegression()` keep their v1 inputs, results and TypeScript
14
+ types; `verify()` still refuses a v2 request with `SCHEMA_VERSION_UNSUPPORTED`.
15
+ - `trusted-local` keeps its fields, stays `UNSANDBOXED` and still requires explicit authorization.
16
+ - MCP tool schemas, evidence export, Agentic Test Profiles and benchmarks still accept only v1.
17
+
18
+ ## What v2 adds
19
+
20
+ - `schemas/verification-request.v2.json`: the v1 request with `schemaVersion: "2.0.0"` and an
21
+ `isolation` union of the unchanged `trusted-local` object and a `container` object with a
22
+ digest-pinned image, declared environment and limits. A request cannot name the runtime command
23
+ or forward host variables.
24
+ - `schemas/evidence-manifest.v2.json`: the v1 manifest with `evidenceContext.execution.backend`,
25
+ which the decision digest covers, and a top-level `isolation` union whose container form records
26
+ the runtime argv. A v2 `trusted-local` manifest records `{ "kind": "trusted-local", "level":
27
+ "UNSANDBOXED" }` as its backend.
28
+ - Both schemas join the additive schema-extension lock, whose published digest changes
29
+ accordingly.
30
+ - `parseVerificationRequestV2()` and `parseEvidenceManifestV2()` accept only v2;
31
+ `parseVersionedVerificationRequest()` and `parseVersionedEvidenceManifest()` accept either version.
32
+ - The SDK adds `verifyV2(request, options)` and `checkGitRegressionV2(options)`, which return
33
+ `EvidenceManifestV2Contract`, and the `GitRegressionV2Options` type.
34
+
35
+ ## Changes visible to existing callers
36
+
37
+ - A request declaring `schemaVersion: "2.0.0"` previously failed with
38
+ `SCHEMA_VERSION_UNSUPPORTED` everywhere. The CLI `verify` command, `verifyV2()` and the exported
39
+ engine function `verifyCampaign()` now validate it as v2. Other unknown versions still fail with
40
+ that code.
41
+ - CLI and SDK container failures use new `CONTAINER_*` and `ISOLATION_MODE_CONFLICT` reason codes,
42
+ with CLI exit code `4`, before any repository code runs.
43
+ - `assertledger schema` also prints `verification-request-v2` and `evidence-manifest-v2`.
44
+
45
+ ## Compatibility witnesses
46
+
47
+ - `tests/verification-v2.test.ts`: v2 acceptance and rejection, digest binding of the backend,
48
+ replay tampering, v1 parsers and export closed to v2, and published schema identities.
49
+ - `tests/conformance-v1.test.ts` and `tests/schema-registry.test.ts`: unchanged v1 bytes and lock,
50
+ registered v2 schemas.
51
+ - `tests/container-backend.test.ts` and `tests/container-facades.test.ts`: backend selection,
52
+ hardened runtime arguments, archive handling and CLI/SDK behavior against a fake runtime.
53
+ - `tests/container-isolation-docker.test.ts`: hostile scenarios against a real Docker Engine.
54
+ - `tests/engine.test.ts` and `tests/core.test.ts` now use `3.0.0` as their unsupported-version
55
+ witness, because `2.0.0` became a supported version.
@@ -50,8 +50,9 @@ référence, un monde neutre rouge, une observation instable ou une attribution
50
50
 
51
51
  Le socle local comprend les contrats versionnés, le noyau déterministe, les digests, le replay
52
52
  sémantique, le moteur d’exécution, les façades CLI/SDK/MCP, l’audit statique et l’initialisation.
53
- Le registre contient 34 schémas JSON, tous présents dans le verrou de conformance. L’ancien
54
- décompte de 30 dans la roadmap était périmé.
53
+ Le registre contient 38 schémas JSON : les 34 figés par le verrou de conformance v1 et les 4
54
+ schémas d’export d’évidence, verrouillés par une extension additive qui ne modifie pas le bundle v1.
55
+ L’ancien décompte de 30 dans la roadmap était périmé.
55
56
 
56
57
  L’adaptateur officiel intégré est `node:test`. Le protocole `testforge-command` permet d’intégrer
57
58
  un autre framework avec un adaptateur fourni par l’opérateur. Détecter Vitest, Jest, Bun ou pytest
@@ -55,15 +55,33 @@ Candidate statuses are:
55
55
 
56
56
  - `ELIGIBLE`: every gate needed by policy passed;
57
57
  - `WEAK_ORACLE`: reference and neutral evidence passed, but target strength did not;
58
- - `INVALID`: discovery, reference, or neutral evidence failed;
58
+ - `INVALID`: a completed run disproved discovery, or reference or neutral evidence failed;
59
59
  - `UNSTABLE`: complete attempts produced different normalized outcomes;
60
60
  - `INCONCLUSIVE`: required observations were missing or execution produced `TIMEOUT` or
61
61
  `INFRA_ERROR`.
62
62
 
63
+ When several apply, the first gate that fails decides, and inconclusive execution takes precedence
64
+ over later failures:
65
+
66
+ 1. missing observations make the candidate `INCONCLUSIVE` (`COMPLETENESS`);
67
+ 2. diverging attempts make it `UNSTABLE` (`STABILITY`);
68
+ 3. a completed run without an attributed candidate test makes it `INVALID`
69
+ (`CANDIDATE_DISCOVERY_INVALID`);
70
+ 4. when only `TIMEOUT` or `INFRA_ERROR` runs lack an attributed candidate test, `DISCOVERY` is
71
+ not established and the candidate is `INCONCLUSIVE` (`CANDIDATE_EXECUTION_INCONCLUSIVE`): such a
72
+ run never reached a verdict, so it can neither prove nor disprove discovery;
73
+ 5. after discovery, any `TIMEOUT` or `INFRA_ERROR` run makes the candidate `INCONCLUSIVE` before a
74
+ red reference or neutral world makes it `INVALID`.
75
+
76
+ See [the migration note](migration-timeout-discovery-inconclusive.md) for the effect on existing
77
+ manifests.
78
+
63
79
  ## Campaign statuses
64
80
 
65
81
  - `VERIFIED`: at least one eligible candidate was selected.
66
- - `REJECTED`: evidence was complete and conclusive, but no candidate was eligible.
82
+ - `REJECTED`: evidence was complete and conclusive, but no candidate was eligible. Once attempts are
83
+ complete and agree, a completed run that disproves discovery is conclusive, even if other runs
84
+ timed out.
67
85
  - `INCONCLUSIVE`: controls were invalid, or at least one candidate was unstable or inconclusive.
68
86
  - `ENGINE_ERROR`: the core could not normalize the supplied evidence safely.
69
87
 
@@ -98,10 +116,12 @@ observations were truthful, authenticate the producer, or replace a signed exter
98
116
 
99
117
  The manifest alone is not a self-contained reproduction bundle. It stores candidate and world
100
118
  digests rather than their file bodies, stdout/stderr digests rather than raw logs, and environment
101
- allowlist names rather than effective values. The built-in `node:test` adapter records its resolved
102
- executable real path, version, and SHA-256 digest; the structured-command adapter records only its
103
- configured command. Preserve the original request, repository snapshot, dependencies, missing
104
- executable identities, and raw logs separately when independent audit matters.
119
+ allowlist names rather than effective values. A v2 container manifest also binds the image identity,
120
+ runtime facts, declared environment values, and limits to the decision. The built-in `node:test`
121
+ adapter records its resolved executable real path, version, and SHA-256 digest; the
122
+ structured-command adapter records only its configured command. Preserve the original request,
123
+ repository snapshot, dependencies, missing executable identities, and raw logs separately when
124
+ independent audit matters.
105
125
 
106
126
  ## Explicit non-claims
107
127
 
@@ -113,4 +133,6 @@ AssertLedger does not prove:
113
133
  - semantic relevance or correctness of operator-supplied worlds;
114
134
  - resistance to a candidate designed to recognize the worlds;
115
135
  - containment of hostile code under `trusted-local`;
136
+ - containment beyond the recorded [container controls](container-isolation.md), which share the
137
+ daemon host's kernel and trust the daemon, its host, and the image;
116
138
  - provenance authenticity without a separate signed attestation.