assertledger 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (203) hide show
  1. package/CONTRIBUTING.md +31 -0
  2. package/LICENSE +21 -0
  3. package/README.fr.md +236 -0
  4. package/README.md +224 -0
  5. package/SECURITY.md +51 -0
  6. package/benchmarks/agentic-profile/README.md +15 -0
  7. package/benchmarks/agentic-profile/public/README.md +5 -0
  8. package/benchmarks/self-hosted-core/README.md +113 -0
  9. package/benchmarks/self-hosted-core/adapter.mjs +293 -0
  10. package/benchmarks/self-hosted-core/builder.ts +193 -0
  11. package/benchmarks/self-hosted-core/campaign.ts +233 -0
  12. package/benchmarks/self-hosted-core/existing-tests-builder.ts +217 -0
  13. package/benchmarks/self-hosted-core/existing-tests.ts +146 -0
  14. package/benchmarks/self-hosted-core/liveness.test.mjs +8 -0
  15. package/conformance/v1/bundle.json +104 -0
  16. package/conformance/v1/expected/canonical-order-a.json +4 -0
  17. package/conformance/v1/expected/canonical-order-b.json +4 -0
  18. package/conformance/v1/expected/create-benchmark-v1-measured.json +575 -0
  19. package/conformance/v1/expected/create-profile-v1-qualified.json +280 -0
  20. package/conformance/v1/expected/decide-collection-failure-non-kill.json +192 -0
  21. package/conformance/v1/expected/decide-compile-failure-non-kill.json +192 -0
  22. package/conformance/v1/expected/decide-infra-error-non-kill.json +192 -0
  23. package/conformance/v1/expected/decide-no-test-discovered-non-kill.json +192 -0
  24. package/conformance/v1/expected/decide-process-crash-non-kill.json +192 -0
  25. package/conformance/v1/expected/decide-timeout-non-kill.json +192 -0
  26. package/conformance/v1/expected/decide-verified.json +192 -0
  27. package/conformance/v1/expected/replay-benchmark-v1-resealed-summary-forgery.json +12 -0
  28. package/conformance/v1/expected/replay-evidence-raw-tamper.json +6 -0
  29. package/conformance/v1/expected/replay-evidence-resealed-semantic-forgery.json +6 -0
  30. package/conformance/v1/inputs/canonical-order-a.json +8 -0
  31. package/conformance/v1/inputs/canonical-order-b.json +8 -0
  32. package/conformance/v1/inputs/create-benchmark-v1-measured.json +459 -0
  33. package/conformance/v1/inputs/create-profile-v1-qualified.json +228 -0
  34. package/conformance/v1/inputs/decide-collection-failure-non-kill.json +143 -0
  35. package/conformance/v1/inputs/decide-compile-failure-non-kill.json +143 -0
  36. package/conformance/v1/inputs/decide-infra-error-non-kill.json +143 -0
  37. package/conformance/v1/inputs/decide-no-test-discovered-non-kill.json +143 -0
  38. package/conformance/v1/inputs/decide-process-crash-non-kill.json +143 -0
  39. package/conformance/v1/inputs/decide-timeout-non-kill.json +143 -0
  40. package/conformance/v1/inputs/decide-verified.json +143 -0
  41. package/conformance/v1/inputs/replay-benchmark-v1-resealed-summary-forgery.json +575 -0
  42. package/conformance/v1/inputs/replay-evidence-raw-tamper.json +201 -0
  43. package/conformance/v1/inputs/replay-evidence-resealed-semantic-forgery.json +192 -0
  44. package/conformance/v1/schemas/expected-digests.json +175 -0
  45. package/dist/cli.d.ts +9 -0
  46. package/dist/cli.d.ts.map +1 -0
  47. package/dist/cli.js +951 -0
  48. package/dist/cli.js.map +1 -0
  49. package/dist/contracts/diagnostics.d.ts +18 -0
  50. package/dist/contracts/diagnostics.d.ts.map +1 -0
  51. package/dist/contracts/diagnostics.js +13 -0
  52. package/dist/contracts/diagnostics.js.map +1 -0
  53. package/dist/contracts/index.d.ts +3908 -0
  54. package/dist/contracts/index.d.ts.map +1 -0
  55. package/dist/contracts/index.js +2569 -0
  56. package/dist/contracts/index.js.map +1 -0
  57. package/dist/contracts/runtime-doctor.d.ts +107 -0
  58. package/dist/contracts/runtime-doctor.d.ts.map +1 -0
  59. package/dist/contracts/runtime-doctor.js +91 -0
  60. package/dist/contracts/runtime-doctor.js.map +1 -0
  61. package/dist/core/index.d.ts +200 -0
  62. package/dist/core/index.d.ts.map +1 -0
  63. package/dist/core/index.js +2587 -0
  64. package/dist/core/index.js.map +1 -0
  65. package/dist/diagnostics.d.ts +7 -0
  66. package/dist/diagnostics.d.ts.map +1 -0
  67. package/dist/diagnostics.js +252 -0
  68. package/dist/diagnostics.js.map +1 -0
  69. package/dist/engine/adapters/node-test-profile.d.ts +14 -0
  70. package/dist/engine/adapters/node-test-profile.d.ts.map +1 -0
  71. package/dist/engine/adapters/node-test-profile.js +14 -0
  72. package/dist/engine/adapters/node-test-profile.js.map +1 -0
  73. package/dist/engine/adapters/node-test-runtime.d.ts +39 -0
  74. package/dist/engine/adapters/node-test-runtime.d.ts.map +1 -0
  75. package/dist/engine/adapters/node-test-runtime.js +173 -0
  76. package/dist/engine/adapters/node-test-runtime.js.map +1 -0
  77. package/dist/engine/adapters/runtime-facts.d.ts +26 -0
  78. package/dist/engine/adapters/runtime-facts.d.ts.map +1 -0
  79. package/dist/engine/adapters/runtime-facts.js +73 -0
  80. package/dist/engine/adapters/runtime-facts.js.map +1 -0
  81. package/dist/engine/connection.d.ts +22 -0
  82. package/dist/engine/connection.d.ts.map +1 -0
  83. package/dist/engine/connection.js +343 -0
  84. package/dist/engine/connection.js.map +1 -0
  85. package/dist/engine/git-regression.d.ts +25 -0
  86. package/dist/engine/git-regression.d.ts.map +1 -0
  87. package/dist/engine/git-regression.js +803 -0
  88. package/dist/engine/git-regression.js.map +1 -0
  89. package/dist/engine/index.d.ts +55 -0
  90. package/dist/engine/index.d.ts.map +1 -0
  91. package/dist/engine/index.js +2782 -0
  92. package/dist/engine/index.js.map +1 -0
  93. package/dist/engine/node-test-reporter.d.ts +2 -0
  94. package/dist/engine/node-test-reporter.d.ts.map +1 -0
  95. package/dist/engine/node-test-reporter.js +70 -0
  96. package/dist/engine/node-test-reporter.js.map +1 -0
  97. package/dist/engine/runtime-doctor.d.ts +16 -0
  98. package/dist/engine/runtime-doctor.d.ts.map +1 -0
  99. package/dist/engine/runtime-doctor.js +100 -0
  100. package/dist/engine/runtime-doctor.js.map +1 -0
  101. package/dist/evaluation/agentic-corpus.d.ts +161 -0
  102. package/dist/evaluation/agentic-corpus.d.ts.map +1 -0
  103. package/dist/evaluation/agentic-corpus.js +710 -0
  104. package/dist/evaluation/agentic-corpus.js.map +1 -0
  105. package/dist/index.d.ts +8 -0
  106. package/dist/index.d.ts.map +1 -0
  107. package/dist/index.js +8 -0
  108. package/dist/index.js.map +1 -0
  109. package/dist/mcp/index.d.ts +13 -0
  110. package/dist/mcp/index.d.ts.map +1 -0
  111. package/dist/mcp/index.js +391 -0
  112. package/dist/mcp/index.js.map +1 -0
  113. package/dist/mcp/stdio.d.ts +3 -0
  114. package/dist/mcp/stdio.d.ts.map +1 -0
  115. package/dist/mcp/stdio.js +13 -0
  116. package/dist/mcp/stdio.js.map +1 -0
  117. package/dist/sdk/index.d.ts +52 -0
  118. package/dist/sdk/index.d.ts.map +1 -0
  119. package/dist/sdk/index.js +224 -0
  120. package/dist/sdk/index.js.map +1 -0
  121. package/dist/version.d.ts +2 -0
  122. package/dist/version.d.ts.map +1 -0
  123. package/dist/version.js +10 -0
  124. package/dist/version.js.map +1 -0
  125. package/docs/adapter-protocol.md +196 -0
  126. package/docs/agentic-benchmark.md +118 -0
  127. package/docs/agentic-corpus-experiment-h3.md +89 -0
  128. package/docs/agentic-corpus-plan.md +105 -0
  129. package/docs/agentic-corpus-provenance.md +59 -0
  130. package/docs/agentic-test-profile-pilot.md +57 -0
  131. package/docs/agentic-test-profile-v2.md +116 -0
  132. package/docs/agentic-test-profile.md +274 -0
  133. package/docs/architecture.md +157 -0
  134. package/docs/ci.md +37 -0
  135. package/docs/client-connections.md +61 -0
  136. package/docs/conformance-v1.md +72 -0
  137. package/docs/decisions/0001-typescript-runtime.md +24 -0
  138. package/docs/developer-experience.md +55 -0
  139. package/docs/diagnostics.md +35 -0
  140. package/docs/distribution.md +40 -0
  141. package/docs/git-regression.md +39 -0
  142. package/docs/migration-repository-validation-order.md +35 -0
  143. package/docs/migration-testforge-to-assertledger.md +64 -0
  144. package/docs/project-intent.md +173 -0
  145. package/docs/proof-model.md +116 -0
  146. package/docs/reference.md +336 -0
  147. package/docs/release-1.0.md +63 -0
  148. package/docs/repository-audit.md +52 -0
  149. package/docs/repository-init.md +60 -0
  150. package/docs/research-basis.md +27 -0
  151. package/docs/roadmap.md +74 -0
  152. package/docs/runtime-doctor.md +65 -0
  153. package/docs/testexplora-calibration.md +71 -0
  154. package/examples/agentic-benchmark/benchmark-request.mjs +19 -0
  155. package/examples/agentic-benchmark/structured-phase-adapter-fixture.mjs +35 -0
  156. package/examples/agentic-profile/profile-benchmark.mjs +34 -0
  157. package/examples/agentic-profile/profile-manifest.mjs +28 -0
  158. package/examples/git-history/README.md +44 -0
  159. package/examples/git-history/create-demo.mjs +128 -0
  160. package/examples/git-history/escape-string-regexp/LICENSE +9 -0
  161. package/examples/git-history/escape-string-regexp/before.cjs.txt +11 -0
  162. package/examples/git-history/escape-string-regexp/fixed.cjs.txt +13 -0
  163. package/examples/git-history/escape-string-regexp/provenance.json +28 -0
  164. package/examples/node-test/repository/package.json +5 -0
  165. package/examples/node-test/repository/src/is-even.js +3 -0
  166. package/examples/node-test/repository/tests/base.test.js +6 -0
  167. package/examples/node-test/request.json +93 -0
  168. package/integrations/skill/SKILL.md +51 -0
  169. package/package.json +88 -0
  170. package/schemas/agentic-benchmark-acquisition-replay-result.v1.json +70 -0
  171. package/schemas/agentic-benchmark-acquisition-request.v1.json +564 -0
  172. package/schemas/agentic-benchmark-acquisition-result.v1.json +1409 -0
  173. package/schemas/agentic-benchmark-artifact.v1.json +1251 -0
  174. package/schemas/agentic-benchmark-replay-result.v1.json +84 -0
  175. package/schemas/agentic-benchmark-request.v1.json +1034 -0
  176. package/schemas/agentic-corpus-allocation-commitment-replay-result.v1.json +58 -0
  177. package/schemas/agentic-corpus-allocation-commitment.v1.json +141 -0
  178. package/schemas/agentic-corpus-allocation-replay-result.v1.json +34 -0
  179. package/schemas/agentic-corpus-allocation-request.v1.json +65 -0
  180. package/schemas/agentic-corpus-allocation-reveal.v1.json +66 -0
  181. package/schemas/agentic-corpus-allocation.v1.json +167 -0
  182. package/schemas/agentic-corpus-experiment-artifact.v1.json +329 -0
  183. package/schemas/agentic-corpus-experiment-plan-replay-result.v1.json +50 -0
  184. package/schemas/agentic-corpus-experiment-plan.v1.json +424 -0
  185. package/schemas/agentic-corpus-experiment-replay-request.v1.json +336 -0
  186. package/schemas/agentic-corpus-experiment-replay-result.v1.json +106 -0
  187. package/schemas/agentic-corpus-experiment-request.v1.json +204 -0
  188. package/schemas/agentic-corpus-provenance.v1.json +143 -0
  189. package/schemas/agentic-corpus-trust-policy.v1.json +133 -0
  190. package/schemas/agentic-profile-replay-result.v1.json +56 -0
  191. package/schemas/agentic-profile-replay-result.v2.json +63 -0
  192. package/schemas/agentic-profile-report.v1.json +961 -0
  193. package/schemas/agentic-profile-report.v2.json +1674 -0
  194. package/schemas/agentic-profile-request.v1.json +671 -0
  195. package/schemas/agentic-profile-request.v2.json +1338 -0
  196. package/schemas/evidence-manifest.v1.json +636 -0
  197. package/schemas/replay-result.v1.json +49 -0
  198. package/schemas/repository-analysis.v1.json +119 -0
  199. package/schemas/repository-audit.v1.json +811 -0
  200. package/schemas/repository-init-config.v1.json +183 -0
  201. package/schemas/repository-init-lock.v1.json +162 -0
  202. package/schemas/repository-init-result.v1.json +212 -0
  203. package/schemas/verification-request.v1.json +389 -0
@@ -0,0 +1,336 @@
1
+ # CLI, SDK and MCP reference
2
+
3
+ [Back to the introduction](../README.md) · [Introduction en français](../README.fr.md)
4
+
5
+ ## JSON CLI
6
+
7
+ This reference uses the installed or locally linked `assertledger` command. When working from
8
+ the source branch, build first and replace `assertledger` with `node dist/cli.js`.
9
+
10
+ For a first run, use `assertledger doctor .` and the [Git qualification guide](git-regression.md).
11
+ `assertledger connect . --client codex` prints a read-only MCP configuration; see
12
+ [developer entry points](developer-experience.md) for creation, conflicts and trust requirements.
13
+
14
+ ```sh
15
+ assertledger init . --dry-run --json
16
+ assertledger init . --json
17
+ assertledger audit . --json
18
+ assertledger analyze . --json
19
+ assertledger schema verification-request --json
20
+ assertledger verify assertledger.request.json --allow-unsafe-execution --json
21
+ assertledger replay assertledger.manifest.json --json
22
+ assertledger profile assertledger.profile-request.json --json
23
+ assertledger profile-replay assertledger.profile-report.json --json
24
+ assertledger profile-v2 assertledger.profile-v2-request.json --json
25
+ assertledger profile-v2-replay assertledger.profile-v2-report.json --json
26
+ assertledger benchmark assertledger.benchmark-request.json --json
27
+ assertledger benchmark-replay assertledger.benchmark-artifact.json --json
28
+ assertledger benchmark-acquire assertledger.benchmark-acquisition-request.json --allow-unsafe-execution --json
29
+ assertledger benchmark-acquire-replay assertledger.benchmark-acquisition-result.json --json
30
+ assertledger mcp
31
+ # Operator-only opt-in: assertledger mcp --allow-unsafe-execution
32
+ ```
33
+
34
+ Every command above also runs identically under the legacy `testforge` binary name.
35
+
36
+ `init` is static: it never executes detected commands or adapters and manages only
37
+ `assertledger.config.json` and `assertledger.lock.json`. It requires operator-owned worlds and
38
+ candidates rather than inventing a verification request. See
39
+ [`docs/repository-init.md`](repository-init.md) for detection conflicts, recovery behavior,
40
+ adapter availability, and the trust boundary.
41
+
42
+ Repository analysis reports a test framework only from framework-specific evidence such as a
43
+ declared dependency, a dedicated configuration file, or a real `node:test` import below a test
44
+ directory or in a colocated `*.test.*`/`*.spec.*` source file. Comments, string examples, and
45
+ documentation-like config filenames do not count; no evidence produces an empty list. Detection is
46
+ intentionally incomplete: an unrecognized manifest layout or test-file convention yields no claim.
47
+
48
+ `verify`, `replay`, `profile`, `profile-replay`, `profile-v2`, `profile-v2-replay`, `benchmark`, and
49
+ `benchmark-replay`, `benchmark-acquire`, and `benchmark-acquire-replay` also accept `-`
50
+ or an omitted file argument and
51
+ then read JSON from stdin. The CLI rejects file and stdin JSON inputs larger than 16 MiB. JSON
52
+ results go to stdout. Diagnostics go to stderr. `assertledger mcp` reserves stdout for JSON-RPC.
53
+
54
+ `--allow-unsafe-execution` is an external authorization signal. The CLI requires it for every
55
+ campaign and sets the request's local acknowledgement before validation. The flag does not create a
56
+ sandbox.
57
+
58
+ Versioned JSON Schemas are published for the
59
+ [`verification request`](../schemas/verification-request.v1.json),
60
+ [`repository analysis`](../schemas/repository-analysis.v1.json),
61
+ [`repository init config`](../schemas/repository-init-config.v1.json),
62
+ [`repository init lock`](../schemas/repository-init-lock.v1.json),
63
+ [`repository init result`](../schemas/repository-init-result.v1.json),
64
+ [`evidence manifest`](../schemas/evidence-manifest.v1.json),
65
+ [`replay result`](../schemas/replay-result.v1.json),
66
+ [`Agentic Test Profile request`](../schemas/agentic-profile-request.v1.json),
67
+ [`Agentic Test Profile report`](../schemas/agentic-profile-report.v1.json),
68
+ [`Agentic Test Profile replay result`](../schemas/agentic-profile-replay-result.v1.json),
69
+ [`Agentic Benchmark request`](../schemas/agentic-benchmark-request.v1.json),
70
+ [`Agentic Benchmark artifact`](../schemas/agentic-benchmark-artifact.v1.json), and
71
+ [`Agentic Benchmark replay result`](../schemas/agentic-benchmark-replay-result.v1.json),
72
+ [`Agentic Benchmark acquisition request`](../schemas/agentic-benchmark-acquisition-request.v1.json),
73
+ [`Agentic Benchmark acquisition result`](../schemas/agentic-benchmark-acquisition-result.v1.json),
74
+ [`Agentic Benchmark acquisition replay result`](../schemas/agentic-benchmark-acquisition-replay-result.v1.json),
75
+ [`Agentic Test Profile v2 request`](../schemas/agentic-profile-request.v2.json),
76
+ [`Agentic Test Profile v2 report`](../schemas/agentic-profile-report.v2.json), and
77
+ [`Agentic Test Profile v2 replay result`](../schemas/agentic-profile-replay-result.v2.json), plus the
78
+ [`corpus trust policy`](../schemas/agentic-corpus-trust-policy.v1.json) and paired
79
+ [`corpus provenance`](../schemas/agentic-corpus-provenance.v1.json), six corpus allocation and
80
+ two-party [`commitment/reveal`](../schemas/agentic-corpus-allocation-commitment.v1.json) contracts,
81
+ plus six pre-declared [`H3 experiment`](../schemas/agentic-corpus-experiment-plan.v1.json) contracts.
82
+ The main CLI `schema` command prints the thirty-one facade schemas by name; the corpus evaluator consumes
83
+ the two trust/provenance schemas directly.
84
+ A complete runnable verification request
85
+ is available at
86
+ [`examples/node-test/request.json`](../examples/node-test/request.json).
87
+ After building the package, the
88
+ [`Agentic Test Profile example`](../examples/agentic-profile/profile-manifest.mjs) derives a report
89
+ directly from any replay-valid manifest. The first two real-campaign results and their limitations
90
+ are recorded in [`docs/agentic-test-profile-pilot.md`](agentic-test-profile-pilot.md).
91
+ The [`Agentic Benchmark example`](../examples/agentic-benchmark/benchmark-request.mjs) consumes a
92
+ strict benchmark request. See [`docs/agentic-benchmark.md`](agentic-benchmark.md) for the fixed
93
+ protocol, comparison scope, replay rails, and current acquisition limitation.
94
+ The benchmark-backed [`Agentic Test Profile v2`](agentic-test-profile-v2.md) replaces v1's
95
+ campaign wall-time proxy with an exact scoped warm-total-wall p95 cost basis. It is additive: v1
96
+ artifacts and commands remain supported.
97
+
98
+ The checked-in [`conformance v1 bundle`](conformance-v1.md) locks autonomous inputs, complete
99
+ expected outputs, negative replay witnesses, all published schema bytes, and selected public
100
+ digests. `pnpm check` validates this static oracle without regenerating it.
101
+
102
+ ## TypeScript SDK
103
+
104
+ ```ts
105
+ import { readFile } from "node:fs/promises";
106
+ import { AssertLedger } from "assertledger";
107
+
108
+ const assertLedger = new AssertLedger();
109
+ const initialized = await assertLedger.init("/absolute/path/to/repository", { dryRun: true });
110
+ const context = await assertLedger.analyze("/absolute/path/to/repository");
111
+ const request = JSON.parse(await readFile("assertledger.request.json", "utf8"));
112
+ const manifest = await assertLedger.verify(request);
113
+ const integrity = assertLedger.replay(manifest);
114
+
115
+ const profile = assertLedger.profile({
116
+ schemaVersion: "1.0.0",
117
+ manifest,
118
+ policy: {
119
+ profileVersion: "1.0.0",
120
+ profileId: "my-repository/default",
121
+ mode: "HARDENING",
122
+ minimumTimingSamples: 3,
123
+ lanes: [
124
+ { id: "instant", maximumReferenceP95Ms: 2_000 },
125
+ { id: "loop", maximumReferenceP95Ms: 10_000 },
126
+ ],
127
+ },
128
+ });
129
+ const profileIntegrity = assertLedger.replayProfile(profile);
130
+
131
+ const benchmarkRequest = JSON.parse(await readFile("assertledger.benchmark-request.json", "utf8"));
132
+ const benchmark = assertLedger.benchmark(benchmarkRequest);
133
+ const benchmarkIntegrity = assertLedger.replayBenchmark(benchmark);
134
+
135
+ const profileV2 = assertLedger.profileV2({
136
+ schemaVersion: "2.0.0",
137
+ benchmarkArtifact: benchmark,
138
+ policy: {
139
+ profileVersion: "2.0.0",
140
+ profileId: "my-repository/hardening-v2",
141
+ mode: "HARDENING",
142
+ requiredComparisonScopeDigest: benchmark.comparisonScopeDigest,
143
+ costBasis: {
144
+ regime: "WARM",
145
+ measure: "WALL",
146
+ aggregation: "TOTAL",
147
+ statistic: "P95",
148
+ unit: "MICROSECOND",
149
+ portfolioAggregation: "SUM_OF_INDIVIDUAL_P95",
150
+ },
151
+ lanes: [{ id: "loop", maximumWarmTotalWallP95Us: 10_000_000 }],
152
+ },
153
+ });
154
+ const profileV2Integrity = assertLedger.replayProfileV2(profileV2);
155
+ ```
156
+
157
+ `AssertLedger` is exported alongside a deprecated `TestForge` subclass alias with an identical
158
+ surface, for consumers migrating from the prior name.
159
+
160
+ The SDK accepts plain JSON-compatible values and validates them against the same contracts as the
161
+ CLI. Unlike the CLI and MCP tool, `AssertLedger.verify()` has no separate authorization parameter: the
162
+ caller must set `isolation.acknowledgedUnsafeExecution` to `true` after applying its own policy.
163
+
164
+ `AssertLedger.replay()` reports schema validity, both digest checks, and deterministic
165
+ decision-semantic validity. Its aggregate `valid` field is true only when all four checks pass. Replay
166
+ does not rerun the campaign, authenticate the producer, or establish that the observations were
167
+ truthful.
168
+
169
+ `AssertLedger.profile()` accepts only a replay-valid manifest. It qualifies selected `ELIGIBLE`
170
+ hardening candidates against repository-scoped latency lanes and publishes evidence strength,
171
+ observed consistency, nearest-rank p50/p95, marginal target weight, and Pareto status. Latency never
172
+ compensates for a failed evidence gate. For each declared lane, the report also selects an
173
+ explainable greedy portfolio that maximizes new target weight per unit of recorded reference p95
174
+ within the lane's total budget. See
175
+ [`docs/agentic-test-profile.md`](agentic-test-profile.md) for the contract and its non-claims.
176
+
177
+ `AssertLedger.benchmark()` accepts only a replay-valid `VERIFIED` source manifest plus a complete,
178
+ fingerprinted cold/warm run plan. It summarizes declared microsecond phase timings and never changes
179
+ the source decision. `await assertLedger.acquireBenchmark()` runs the fresh campaign and strict
180
+ framework-neutral phase acquisition directly. The built-in `node:test` adapter remains unsupported
181
+ because it cannot faithfully expose all four phase boundaries.
182
+ `AssertLedger.replayBenchmarkAcquisition()` independently checks the source, artifact, context, result
183
+ digest, and exact status/reason semantics.
184
+
185
+ H3 replay requires three externally pinned digests and separate SDK options for the trust policy,
186
+ two-party commitment/reveal, allocation, pre-declared plan, subject provenance and evidence bytes.
187
+ Allocation is source-stratified; arms use content-addressed candidate references and derived suite
188
+ digests. Replay resolves and re-hashes every candidate's supplied bytes; this proves byte identity,
189
+ not that those bytes were executed. Scheduled process counts and planned timeout ceilings must be
190
+ exactly equal. H3
191
+ non-inferiority must hold per source and in aggregate. Receipts attest the applied timeout limit but
192
+ do not independently prove enforcement or equal observed CPU/wall time.
193
+ The CLI exposes the same boundary through mandatory `--trust-policy-digest`,
194
+ `--allocation-commitment-digest`, and `--experiment-plan-digest` flags plus separate files. See
195
+ [`docs/agentic-corpus-experiment-h3.md`](agentic-corpus-experiment-h3.md).
196
+
197
+ `AssertLedger.profileV2()` accepts only a replay-valid benchmark artifact. Its comparison scope must
198
+ match the policy, all warm measurements must be usable, and every cost is the declared warm total
199
+ wall p95 in microseconds. It publishes a Pareto frontier and deterministic greedy portfolios only
200
+ over source-selected eligible candidates. The portfolio cost is a sum of individual p95 values,
201
+ not a measured portfolio p95 and not a universal optimum.
202
+
203
+ ## Calibration corpus
204
+
205
+ The versioned scaffold under [`benchmarks/agentic-profile`](../benchmarks/agentic-profile) evaluates
206
+ the falsifiable H1-H4 hypotheses without an LLM judge. It remains `NOT_READY` until at least 20
207
+ strict cases from three sources cover every hypothesis and include physically separate public and
208
+ private splits.
209
+
210
+ ```sh
211
+ tsx scripts/evaluate-agentic-profile-corpus.ts status benchmarks/agentic-profile \
212
+ --trust-policy operator-policy.json --trust-policy-digest sha256:...
213
+ tsx scripts/evaluate-agentic-profile-corpus.ts evaluate-public benchmarks/agentic-profile \
214
+ --trust-policy operator-policy.json --trust-policy-digest sha256:...
215
+ tsx scripts/evaluate-agentic-profile-corpus.ts evaluate-holdout benchmarks/agentic-profile \
216
+ --trust-policy operator-policy.json --trust-policy-digest sha256:...
217
+ ```
218
+
219
+ Public evaluation returns per-case reason codes. Holdout evaluation returns aggregate hypothesis
220
+ counts only and suppresses case identifiers, source identifiers, raw values, and readiness counts.
221
+ The repository intentionally contains no fabricated cases; private `*.case.json` files are ignored.
222
+ The first empirical calibration is frozen in the
223
+ [24-case corpus plan](agentic-corpus-plan.md). The plan fixes sources and admission rules; it
224
+ is not evidence that every tranche has already been executed. The receipt-linked
225
+ [TestExplora calibration](testexplora-calibration.md) contributes eight real, stable cases but
226
+ is explicitly curated calibration rather than holdout or `READY_H3` evidence.
227
+
228
+ ## MCP v2 over stdio
229
+
230
+ Start the read-only server with `assertledger mcp`, or configure an MCP client with `assertledger` as
231
+ the command and `["mcp"]` as its argument array (`testforge mcp` runs the identical server). The
232
+ server name reported to clients is `assertledger`. Every tool is registered twice, under a preferred
233
+ `assertledger_*` name and a legacy `testforge_*` name bound to the same handler and the same tool
234
+ configuration. The schema lookup pair has no single fixed output schema because its selected JSON
235
+ Schema document varies; every other pair shares the same output-schema object. The default server
236
+ exposes twenty-eight read-only tools:
237
+
238
+ | Preferred tool | Legacy alias | Purpose |
239
+ | --- | --- | --- |
240
+ | `assertledger_analyze` | `testforge_analyze` | Produce repository context for test generation |
241
+ | `assertledger_doctor` | `testforge_doctor` | Return a static repository initialization plan without writing files or executing repository code |
242
+ | `assertledger_explain` | `testforge_explain` | Explain bounded reason codes with versioned guidance and safe next actions |
243
+ | `assertledger_benchmark` | `testforge_benchmark` | Derive scoped cold/warm phase summaries from declared raw runs |
244
+ | `assertledger_benchmark_replay` | `testforge_benchmark_replay` | Replay a self-contained benchmark artifact |
245
+ | `assertledger_profile` | `testforge_profile` | Derive an Agentic Test Profile from replay-valid evidence |
246
+ | `assertledger_profile_replay` | `testforge_profile_replay` | Replay a self-contained profile report |
247
+ | `assertledger_profile_v2` | `testforge_profile_v2` | Derive a benchmark-backed strength and warm-cost profile |
248
+ | `assertledger_profile_v2_replay` | `testforge_profile_v2_replay` | Replay a self-contained Profile v2 report |
249
+ | `assertledger_schema` | `testforge_schema` | Return any of the published JSON Schemas |
250
+ | `assertledger_replay` | `testforge_replay` | Validate and replay a manifest's schema, digests, and decision semantics |
251
+ | `assertledger_benchmark_acquire_replay` | `testforge_benchmark_acquire_replay` | Replay acquisition source, artifact, context, digest, and status bindings |
252
+ | `assertledger_corpus_allocate` | `testforge_corpus_allocate` | Create a deterministic calibration/holdout allocation |
253
+ | `assertledger_corpus_allocation_replay` | `testforge_corpus_allocation_replay` | Replay allocation scores, partition, and digest semantics |
254
+
255
+ H3 creation and replay are deliberately absent from the default agent-facing MCP server. Its
256
+ operator-owned policy, commitment and plan pins cannot be supplied safely as self-attested tool
257
+ input.
258
+
259
+ `assertledger_verify`/`testforge_verify` and `assertledger_benchmark_acquire`/`testforge_benchmark_acquire`
260
+ are absent by default, as are `assertledger_check`/`testforge_check` and
261
+ `assertledger_doctor_runtime`/`testforge_doctor_runtime`. A server operator may register these pairs by starting
262
+ `assertledger mcp --allow-unsafe-execution`, or by calling `createAssertLedgerServer({
263
+ allowUnsafeExecution: true })` (the deprecated `createTestForgeServer` alias calls the identical
264
+ factory). The tool then accepts a verification request and executes it without a second per-call
265
+ authorization field. Run that server only inside the intended isolation boundary; do not let an MCP
266
+ caller decide whether the capability exists. The server resolves repository roots to real paths and
267
+ confines them to the server process's current working directory by default. Programmatic operators
268
+ may supply a different `allowedRepositoryRoots` allowlist.
269
+
270
+ The doctor pair accepts a strict `{ "root": "..." }` input and returns the existing
271
+ `repository-init-result` contract. It is read-only in both the default and operator-enabled server;
272
+ enabling unsafe execution does not change doctor behavior. Dynamic runtime and client diagnostics
273
+ remain outside this static readiness result.
274
+
275
+ `doctor_runtime` accepts the same strict root input and returns the separate
276
+ [runtime diagnostic contract](runtime-doctor.md). `check` accepts the
277
+ [high-level Git options](git-regression.md), without a permission field, and returns the existing
278
+ evidence manifest. The operator's capability is required for both tools.
279
+
280
+ ## Continuous integration
281
+
282
+ Run `pnpm check` on every change. The included GitHub Actions workflow runs this gate on Node.js 22
283
+ and 24 on Ubuntu and Windows. A separate matrix installs and exercises the packed artifact on both
284
+ operating systems with Node.js 22.15.0 and 24. A CI job that executes campaigns must
285
+ also treat `trusted-local` as `UNSANDBOXED`: use an isolated runner without secrets or host
286
+ credentials, and pass `--allow-unsafe-execution` only from reviewed CI configuration.
287
+
288
+ ## Decision semantics
289
+
290
+ - `VERIFIED`: at least one candidate completed all required evidence and was selected.
291
+ - `REJECTED`: the campaign completed, but no candidate satisfied the policy.
292
+ - `INCONCLUSIVE`: controls or candidate evidence were incomplete, unstable, timed out, or affected
293
+ by infrastructure failure.
294
+ - `ENGINE_ERROR`: the deterministic core could not normalize the supplied evidence safely.
295
+
296
+ Only an attributed `ASSERTION_FAILURE` can kill a target in protocol v1. Compilation errors,
297
+ collection failures, crashes, timeouts, and infrastructure errors never count as target evidence.
298
+
299
+ Campaign budgets cover aggregate candidate, world, repository, overlay, and execution counts or
300
+ bytes. `timeoutMsPerExecution` and `maximumOutputBytes` apply to each process execution; the timeout
301
+ and process-tree termination are best effort on the local host.
302
+
303
+ Before a built-in `node:test` campaign runs, AssertLedger probes the requested executable, resolves its
304
+ real path, probes that resolved file again, and requires matching Node.js versions of at least
305
+ 22.15. The manifest records the requested executable, resolved path, Node.js version, and executable
306
+ SHA-256 digest.
307
+
308
+ Each candidate records the ordered gates `COMPLETENESS`, `STABILITY`, `DISCOVERY`, `REFERENCE`,
309
+ `NEUTRAL`, and `TARGET_STRENGTH`, including evidence run IDs and reason codes. See
310
+ [docs/proof-model.md](proof-model.md) for candidate statuses, controls, selection, and digest
311
+ scope.
312
+
313
+ The manifest is sufficient to replay AssertLedger's deterministic decision, but it is not a complete
314
+ audit archive. It stores content and process-output digests, not candidate or world bodies, a
315
+ repository archive, or raw logs. Preserve the original request, repository snapshot or trusted
316
+ source reference, raw logs, dependencies, and any structured-command executable identity separately
317
+ when independent audit or reproduction matters. Manifest digests detect changes; they do not
318
+ authenticate the producer.
319
+
320
+ ## Project status and roadmap
321
+
322
+ The [intent and restart assessment](project-intent.md) records the current implementation,
323
+ adoption gaps, and proposed delivery order. It distinguishes local evidence from release and
324
+ empirical claims.
325
+
326
+ Version `1.0.0` qualifies the supported node:test workflow. The
327
+ deterministic core and protocols are framework-independent; `node:test` is the first built-in
328
+ framework adapter. Other frameworks integrate through the structured-command protocol described in
329
+ [docs/adapter-protocol.md](adapter-protocol.md).
330
+
331
+ Planned work is not shipped behavior. Priorities include a real sandbox backend, additional
332
+ framework reporters with runtime attribution, signed provenance, cross-runtime conformance
333
+ fixtures, and more built-in adapters. See [docs/roadmap.md](roadmap.md).
334
+
335
+ Contributions are welcome under the [MIT license](../LICENSE). Read [CONTRIBUTING.md](../CONTRIBUTING.md)
336
+ before changing a public contract.
@@ -0,0 +1,63 @@
1
+ # AssertLedger 1.0
2
+
3
+ Cap décidé par le propriétaire le 8 septembre 2026. Ce document fixe les critères de la version 1.0 ;
4
+ la publication est attestée séparément par le tag, le registre npm et la release GitHub.
5
+ Le nom du package reste `assertledger` et le dépôt public est
6
+ [hoklims/assertledger](https://github.com/hoklims/assertledger).
7
+
8
+ ## Résultat utilisateur
9
+
10
+ Un développeur ou son agent apporte le test d’une correction. AssertLedger vérifie que ce test
11
+ passe sur la correction et les contrôles déclarés, puis détecte la faute connue par assertion.
12
+ Le résultat identifie les révisions, les octets du test, les essais, les raisons et les limites.
13
+ Le même artefact se rejoue sans modèle dans le terminal et en CI.
14
+
15
+ La première cible est une équipe Node.js/TypeScript utilisant des agents. La première exécution
16
+ officielle repose sur `node:test` ; les variantes de TypeScript ou de dépendances non qualifiées
17
+ doivent être annoncées comme telles. Détecter un framework ne signifie pas savoir le qualifier.
18
+
19
+ ## Parcours et critères de livraison
20
+
21
+ | Étape | Critère observable | Issue |
22
+ | --- | --- | --- |
23
+ | Première preuve | Version fautive, correction, test et contrôle neutre déclarés ; test utile et inutile distingués ; manifeste rejouable | [HOK-658](https://linear.app/hoklims/issue/HOK-658) |
24
+ | Adaptateur de référence | Matrice adversariale, trois mondes, deux essais, candidat fort seul sélectionné | [HOK-570](https://linear.app/hoklims/issue/HOK-570) |
25
+ | Diagnostic | Prérequis, permissions et limites expliqués avant exécution ; prochaine action sûre | [HOK-420](https://linear.app/hoklims/issue/HOK-420) |
26
+ | Agents | CLI/SDK/MCP compatibles ; configuration idempotente et confinée aux racines autorisées | [HOK-423](https://linear.app/hoklims/issue/HOK-423) |
27
+ | Restitution | Verdict humain et JSON cohérents, raisons stables, aucune suggestion qui contourne un gate | [HOK-428](https://linear.app/hoklims/issue/HOK-428) |
28
+ | Distribution | Installation du tarball dans un consommateur neuf ; binaires, SDK, fichiers runtime et replay vérifiés | [HOK-659](https://linear.app/hoklims/issue/HOK-659) |
29
+ | Publication | CI et revue du candidat exact, tag/package/release publiés, réinstallation de la version publiée | [HOK-399](https://linear.app/hoklims/issue/HOK-399) |
30
+
31
+ La documentation doit être copiée et exécutée sans correction implicite. Mesurer le temps jusqu’à
32
+ la première preuve comprise ; cinq minutes reste un objectif, sans résultat annoncé d’avance.
33
+ Une démo locale n’est pas une preuve d’adoption. [Trois pilotes](https://linear.app/hoklims/issue/HOK-660)
34
+ doivent ensuite documenter réutilisation, effort, abandon et volonté de conserver l’intégration.
35
+
36
+ ## Frontières de la version
37
+
38
+ - Les gates, schémas, projections de digest et aliases TestForge conservent leur compatibilité.
39
+ - Seule une assertion attribuée peut compter comme détection ; les erreurs opérationnelles restent
40
+ des non-kills. La vitesse ne compense aucun défaut de preuve.
41
+ - Les mondes attendus relèvent de l’opérateur. Un contrôle identique à la référence ne devient pas
42
+ une preuve indépendante de robustesse.
43
+ - `trusted-local` reste `UNSANDBOXED`, opt-in, réservé au code de confiance. Les exemples CI refusent
44
+ les contributions non fiables et l’exécution sur un runner privilégié.
45
+ - Le replay vérifie intégrité et cohérence ; il n’authentifie pas le producteur des observations.
46
+ - Profils statistiques et holdout gardent leurs critères avant toute affirmation empirique.
47
+
48
+ ## Extensions conservées
49
+
50
+ Les adaptateurs Vitest, Jest, Bun et pytest suivent les besoins des pilotes. Cache incrémental,
51
+ landing automatique, profils de performance et isolation des contributions hostiles sont des
52
+ travaux distincts. Leurs critères restent dans Linear ; ils ne sont pas annoncés comme livrés par
53
+ la seule publication d’un package 1.0.
54
+
55
+ ## État de départ vérifié
56
+
57
+ Le candidat local du 8 septembre passe `pnpm check` avec 307 tests. La version du package est
58
+ encore `0.1.0`. Le dépôt public dédié est créé et `origin` configuré ; aucune release n’est publiée.
59
+ Le contrôle npm initial retournait `E401`. L’authentification du compte personnel `hoklims` a
60
+ ensuite été vérifiée ; elle ne vaut pas preuve de publication.
61
+ Le diff hérité est conservé ; sa présence ne constitue pas une acceptation de tous ses travaux.
62
+ La preuve externe [HOK-406](https://linear.app/hoklims/issue/HOK-406) reste ouverte et ne donne pas
63
+ autorité pour modifier la politique active.
@@ -0,0 +1,52 @@
1
+ # Repository audit v1
2
+
3
+ `assertledger audit` inventories a repository without executing tests, adapters, candidates, or
4
+ campaigns. Its output is descriptive: it contains no score, note, quality label, coverage claim, or
5
+ verification verdict.
6
+
7
+ ```sh
8
+ assertledger audit . --json
9
+ assertledger audit . --verification-request request.json --json
10
+ assertledger audit . --emit-verification-request --no-git | assertledger verify - --allow-unsafe-execution
11
+ ```
12
+
13
+ The command discovers `assertledger.request.json` at the audited root unless
14
+ `--verification-request` is supplied. A request is emitted only after v1 validation and exact
15
+ repository-root binding. Emitted requests always have
16
+ `acknowledgedUnsafeExecution: false`; the receiving `verify` command owns the explicit unsafe
17
+ authorization. Exit code `3` means no request was available for `--emit-verification-request`, and
18
+ exit code `4` means the supplied or discovered request was invalid.
19
+
20
+ ## Signals and limits
21
+
22
+ JavaScript and TypeScript files use the TypeScript lexical scanner. Decision points count `if`,
23
+ loop, `catch`, `case`, conditional, `&&`, `||`, and `??` tokens. Declared exports and test-root
24
+ identifier occurrences produce `noObservedTestReference`; this name deliberately does not claim
25
+ coverage. Test association is conservative: source and `.test`/`.spec` files must share a basename.
26
+ Weak assertion shapes are syntactic occurrences of truthiness-only, constant-boolean equality,
27
+ definedness-only (`toBeDefined`), type-only (`typeof` asserted against a type name), and
28
+ non-throw-only forms. Unsupported languages report `supported: false` and nullable measurements.
29
+
30
+ Git history is optional. The 90-day window is anchored to the commit timestamp of `HEAD`, never the
31
+ wall clock, and is capped at 1,000 commits. Correction commits are messages containing the explicit
32
+ words `fix`, `bug`, `repair`, `hotfix`, or `regression` (case-insensitive). Missing, shallow, failed,
33
+ or truncated history is represented by nulls and reason codes, not invented zeroes.
34
+
35
+ Modules use `PATH_PREFIX_V1` (the first two path segments when available). Their contiguous order is
36
+ the descending factual tuple: decision points, exports with no observed test reference, correction
37
+ commits, commits in the window, sources newer than associated tests, weak assertion occurrences;
38
+ the POSIX path prefix breaks ties. The tuple is not a score.
39
+
40
+ Campaign bounds use decimal BigInt strings. For candidates `C`, worlds `W`, attempts `A`, repository
41
+ bytes `R`, total world-overlay bytes `OW`, total candidate bytes `OC`, timeout `T`, and output limit
42
+ `O`:
43
+
44
+ - executions = `(C + 1) * W * A`
45
+ - controls = `W * A`; candidate executions = `C * W * A`
46
+ - overlay bytes = `A * ((C + 1) * OW + W * OC)`
47
+ - materialization bytes = `R * (executions + 1) + overlay bytes`
48
+ - timeout limit = `executions * T`
49
+ - captured stream limit = `executions * O * 2`
50
+
51
+ Two complete inventories bind the audit. A digest change aborts with
52
+ `REPOSITORY_CHANGED_DURING_AUDIT`.
@@ -0,0 +1,60 @@
1
+ # Repository initialization
2
+
3
+ `assertledger init` performs a static, provider-neutral repository detection pass and manages exactly
4
+ two root files: `assertledger.config.json` and `assertledger.lock.json`. It never runs a test command,
5
+ adapter, candidate, or campaign, and it never edits package manifests, lockfiles, CI configuration,
6
+ or tests.
7
+
8
+ ```sh
9
+ assertledger init . --dry-run --json
10
+ assertledger init . --json
11
+ assertledger audit . --json
12
+ ```
13
+
14
+ The dry run emits the exact canonical bytes and SHA-256 digests that a subsequent write plans to
15
+ use. Writes use a same-directory temporary file followed by an atomic rename. A second run returns
16
+ `UNCHANGED` without rewriting matching files. A missing lock or a stale, structurally valid lock
17
+ bound to the same config can be recovered; invalid files and config-binding conflicts fail closed
18
+ without a `--force` mode.
19
+
20
+ Detection is conservative. Package-manager declarations and root lockfiles must converge. Multiple
21
+ plausible test frameworks, composite shell scripts, contradictory overrides, and invalid explicit
22
+ adapter configurations return `CONFLICT`. Multiple CI providers are only sorted evidence and do not
23
+ block initialization.
24
+
25
+ Files at or below the managed `candidateRoots` are deliberately excluded from framework inference,
26
+ evidence, built-in control tests, and repository-change comparison. Candidate generation therefore
27
+ cannot silently redefine initialization facts or invalidate an otherwise unchanged lock.
28
+
29
+ The built-in ready adapter is currently `node-test`. Bun, pytest, Vitest, and Jest can be detected,
30
+ but initialization returns `BLOCKED` with `OFFICIAL_ADAPTER_UNAVAILABLE` unless the operator supplies
31
+ an existing structured adapter configuration:
32
+
33
+ ```sh
34
+ assertledger init . --adapter-config integrations/my-adapter.json --json
35
+ ```
36
+
37
+ The adapter document is parsed through the public adapter contract and is operator-owned. Its
38
+ executable is recorded as argv but is not resolved or executed by `init`. This is not an official
39
+ adapter endorsement and does not reduce the later `trusted-local` execution boundary.
40
+ `node-test` adapters are accepted only for the `node:test` framework; every other framework requires
41
+ an operator-owned `testforge-command` adapter.
42
+
43
+ All detections and evidence digests come from one byte snapshot. Immediately before returning or
44
+ writing managed files, initialization rechecks the in-scope inventory and every evidence byte. A
45
+ change returns `CONFLICT` with `REPOSITORY_CHANGED_DURING_INIT` and writes nothing.
46
+
47
+ Initialization deliberately does not generate a verification request. Both `worlds` and
48
+ `candidates` remain required operator or harness inputs because repository detection cannot infer
49
+ valid semantic faults, neutral transformations, or oracle expectations. The next command is
50
+ `assertledger audit . --json`; that audit may report `VERIFICATION_REQUEST_UNAVAILABLE` until the
51
+ harness supplies those inputs.
52
+
53
+ Exit codes are `0` for `CREATED`, `UNCHANGED`, or `WOULD_CREATE`; `3` for `BLOCKED`; `4` for
54
+ ambiguity, invalid overrides, conflicts, or contract validation; `5` for unexpected I/O; and `64`
55
+ for CLI usage errors.
56
+
57
+ The public `repository-init-config.v1`, `repository-init-lock.v1`, and
58
+ `repository-init-result.v1` contracts contain no timestamps, absolute repository roots, environment
59
+ values, worlds, or candidates. The lock binds normalized detections and sorted evidence digests; it
60
+ does not authenticate the detector, repository, adapter, or later execution evidence.
@@ -0,0 +1,27 @@
1
+ # Research basis
2
+
3
+ AssertLedger's evidence model follows several established testing and provenance principles.
4
+
5
+ - Mutation testing distinguishes killed mutants from live mutants and reveals oracle weakness that
6
+ coverage alone cannot show: [PIT](https://pitest.org/) and
7
+ [Just et al., FSE 2014](https://homes.cs.washington.edu/~mernst/pubs/mutation-effectiveness-fse2014-abstract.html).
8
+ - Repetition measures observed flakiness; a successful retry does not erase a contradictory run:
9
+ [Gruber et al., 2021](https://arxiv.org/abs/2101.09077) and
10
+ [Lam et al., ISSTA 2019](https://www.microsoft.com/en-us/research/publication/root-causing-flaky-tests-in-a-large-scale-industrial-setting/).
11
+ - Selection is stated only over executed evidence; “safe” regression-test selection requires stronger
12
+ assumptions: [Rothermel and Harrold, IEEE TSE 1998](https://digitalcommons.unl.edu/csearticles/11/).
13
+ - The manifest separates external parameters, resolved inputs, builder identity, and byproducts,
14
+ inspired by [SLSA build provenance v1.2](https://slsa.dev/spec/v1.2/build-provenance). AssertLedger does
15
+ not claim a SLSA level.
16
+ - The proposed Agentic Test Profile keeps evidence strength and observed execution cost as separate
17
+ dimensions. This follows cost-aware regression-test prioritization rather than treating line
18
+ coverage as an oracle-quality score: [Elbaum, Rothermel, and Penix, FSE
19
+ 2014](https://research.google/pubs/techniques-for-improving-regression-testing-in-continuous-integration-development-environments/)
20
+ and [Meta predictive test
21
+ selection](https://engineering.fb.com/2018/11/21/developer-tools/predictive-test-selection/).
22
+ - Mutation-guided agentic testing now has industrial hardening and catching-test evidence. The two
23
+ modes have different reference directions and must remain distinct: [Mutation-Guided LLM-based
24
+ Test Generation at Meta](https://arxiv.org/abs/2501.12862) and [Just-in-Time Catching Test
25
+ Generation at Meta](https://arxiv.org/abs/2601.22832).
26
+
27
+ These sources motivate the design; they do not turn mutant detection into a proof of correctness.
@@ -0,0 +1,74 @@
1
+ # Roadmap
2
+
3
+ This roadmap separates locally implemented v0.1 behavior from the agreed 1.0 delivery scope.
4
+ See the [intent assessment](project-intent.md) and [1.0 release contract](release-1.0.md) for the
5
+ regression-test qualification workflow, acceptance criteria, and evidence boundaries.
6
+
7
+ ## Implemented locally in v0.1
8
+
9
+ - deterministic repository analysis;
10
+ - thirty-four versioned public JSON Schemas, all frozen by conformance v1;
11
+ - reference, target, and neutral overlay worlds;
12
+ - one campaign snapshot and disposable execution workspaces;
13
+ - built-in `node:test` runtime reporter and structured-command adapter;
14
+ - deterministic gates, candidate selection, reason codes, and evidence run links;
15
+ - decision and artifact integrity digests;
16
+ - derived Agentic Test Profiles with replay, repository-scoped latency lanes, marginal evidence,
17
+ and a multi-dimensional Pareto frontier;
18
+ - replayable Agentic Benchmark Artifacts with fingerprinted comparison scopes, cold/warm regimes,
19
+ deterministic phase quantiles, and explicit sample/failure states;
20
+ - fresh phase-aware benchmark acquisition for `testforge-command`, gated by replay-valid
21
+ `VERIFIED` evidence and bound to adapter and dependency identities;
22
+ - benchmark-backed Agentic Test Profile v2 reports with an exact warm-total-wall p95 cost basis,
23
+ selected-only cohort disclosure, deterministic Pareto comparison, and summed-cost portfolios;
24
+ - externally pinned Ed25519 corpus trust policies and canonical dual-signed provenance sidecars;
25
+ - deterministic source-stratified corpus allocations with dual-principal commitment/reveal and
26
+ private holdout custody;
27
+ - externally pinned H3 experiment plans and artifacts with byte-resolved run receipts, exact
28
+ planned-process and planned-timeout equality, content-addressed suites, per-source
29
+ non-inferiority, and a shared frozen candidate universe;
30
+ - a static conformance-v1 oracle locking canonicalization, decisions, replay witnesses, Profile v1,
31
+ Benchmark v1, and all 34 published schema bytes;
32
+ - a pinned 24-case, three-source empirical corpus plan with signed admission and holdout rules,
33
+ including eight receipt-linked TestExplora cases admitted only for curated calibration;
34
+ - JSON CLI, TypeScript SDK, MCP v2 stdio server, and integration skill;
35
+ - `trusted-local`, explicitly recorded as `UNSANDBOXED`.
36
+
37
+ ## Release 1.0 priority
38
+
39
+ 1. Qualify a regression test against declared buggy, fixed, and neutral revisions using existing
40
+ deterministic gates, with useful diagnostics and preserved artifacts.
41
+ 2. Verify installation and the documented workflow from the packed package in a fresh consumer.
42
+ 3. Add repository diagnostics and agent integration, then a PR/CI report and replay path.
43
+ 4. Publish only after exact-candidate checks and required independent review; install and verify
44
+ the published version before closing delivery.
45
+
46
+ The first built-in adapter is `node:test`. Broader adapters and optimizations follow pilot needs.
47
+ The empirical work below remains required for scientific claims, rather than for releasing the
48
+ scoped deterministic qualification workflow.
49
+
50
+ ## Further evidence and extensions
51
+
52
+ 1. Complete the [24-case corpus plan](agentic-corpus-plan.md), including provenance sidecars, a
53
+ separately selected and pre-committed holdout, reviewed historical faults, mutants, and neutral
54
+ rewrites, before claiming an empirically established sweet spot. The current TestExplora 8/8
55
+ tranche is curated calibration evidence and cannot satisfy the holdout requirement.
56
+ 2. Add faithful framework-specific phase adapters. The framework-neutral acquisition path is
57
+ shipped, but the built-in `node:test` reporter cannot attribute all four phases.
58
+ 3. Add an isolation backend backed by an independently administered container or VM boundary.
59
+ 4. Add framework reporters beyond `node:test` that derive discovery and attribution from runtime
60
+ events.
61
+ 5. Publish cross-runtime conformance fixtures for canonicalization, decisions, and both digests.
62
+ 6. Add mature replayable experiment artifacts for H1, H2, and H4 without allowing them to
63
+ compensate for an H3 failure.
64
+ 7. Add adapters only when their outcome mapping has adversarial contract tests.
65
+ 8. Extend provenance to cover executed dependencies, structured-command binary identities, and
66
+ effective environment inputs without leaking secrets.
67
+
68
+ ## Explicitly out of scope for v0.1
69
+
70
+ - hostile-code containment;
71
+ - automatic trust in agent-supplied worlds or policy;
72
+ - proof of program correctness or complete fault detection;
73
+ - proof that a test will never become flaky;
74
+ - SLSA conformance or authenticated provenance.