assertledger 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CONTRIBUTING.md +31 -0
- package/LICENSE +21 -0
- package/README.fr.md +236 -0
- package/README.md +224 -0
- package/SECURITY.md +51 -0
- package/benchmarks/agentic-profile/README.md +15 -0
- package/benchmarks/agentic-profile/public/README.md +5 -0
- package/benchmarks/self-hosted-core/README.md +113 -0
- package/benchmarks/self-hosted-core/adapter.mjs +293 -0
- package/benchmarks/self-hosted-core/builder.ts +193 -0
- package/benchmarks/self-hosted-core/campaign.ts +233 -0
- package/benchmarks/self-hosted-core/existing-tests-builder.ts +217 -0
- package/benchmarks/self-hosted-core/existing-tests.ts +146 -0
- package/benchmarks/self-hosted-core/liveness.test.mjs +8 -0
- package/conformance/v1/bundle.json +104 -0
- package/conformance/v1/expected/canonical-order-a.json +4 -0
- package/conformance/v1/expected/canonical-order-b.json +4 -0
- package/conformance/v1/expected/create-benchmark-v1-measured.json +575 -0
- package/conformance/v1/expected/create-profile-v1-qualified.json +280 -0
- package/conformance/v1/expected/decide-collection-failure-non-kill.json +192 -0
- package/conformance/v1/expected/decide-compile-failure-non-kill.json +192 -0
- package/conformance/v1/expected/decide-infra-error-non-kill.json +192 -0
- package/conformance/v1/expected/decide-no-test-discovered-non-kill.json +192 -0
- package/conformance/v1/expected/decide-process-crash-non-kill.json +192 -0
- package/conformance/v1/expected/decide-timeout-non-kill.json +192 -0
- package/conformance/v1/expected/decide-verified.json +192 -0
- package/conformance/v1/expected/replay-benchmark-v1-resealed-summary-forgery.json +12 -0
- package/conformance/v1/expected/replay-evidence-raw-tamper.json +6 -0
- package/conformance/v1/expected/replay-evidence-resealed-semantic-forgery.json +6 -0
- package/conformance/v1/inputs/canonical-order-a.json +8 -0
- package/conformance/v1/inputs/canonical-order-b.json +8 -0
- package/conformance/v1/inputs/create-benchmark-v1-measured.json +459 -0
- package/conformance/v1/inputs/create-profile-v1-qualified.json +228 -0
- package/conformance/v1/inputs/decide-collection-failure-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-compile-failure-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-infra-error-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-no-test-discovered-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-process-crash-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-timeout-non-kill.json +143 -0
- package/conformance/v1/inputs/decide-verified.json +143 -0
- package/conformance/v1/inputs/replay-benchmark-v1-resealed-summary-forgery.json +575 -0
- package/conformance/v1/inputs/replay-evidence-raw-tamper.json +201 -0
- package/conformance/v1/inputs/replay-evidence-resealed-semantic-forgery.json +192 -0
- package/conformance/v1/schemas/expected-digests.json +175 -0
- package/dist/cli.d.ts +9 -0
- package/dist/cli.d.ts.map +1 -0
- package/dist/cli.js +951 -0
- package/dist/cli.js.map +1 -0
- package/dist/contracts/diagnostics.d.ts +18 -0
- package/dist/contracts/diagnostics.d.ts.map +1 -0
- package/dist/contracts/diagnostics.js +13 -0
- package/dist/contracts/diagnostics.js.map +1 -0
- package/dist/contracts/index.d.ts +3908 -0
- package/dist/contracts/index.d.ts.map +1 -0
- package/dist/contracts/index.js +2569 -0
- package/dist/contracts/index.js.map +1 -0
- package/dist/contracts/runtime-doctor.d.ts +107 -0
- package/dist/contracts/runtime-doctor.d.ts.map +1 -0
- package/dist/contracts/runtime-doctor.js +91 -0
- package/dist/contracts/runtime-doctor.js.map +1 -0
- package/dist/core/index.d.ts +200 -0
- package/dist/core/index.d.ts.map +1 -0
- package/dist/core/index.js +2587 -0
- package/dist/core/index.js.map +1 -0
- package/dist/diagnostics.d.ts +7 -0
- package/dist/diagnostics.d.ts.map +1 -0
- package/dist/diagnostics.js +252 -0
- package/dist/diagnostics.js.map +1 -0
- package/dist/engine/adapters/node-test-profile.d.ts +14 -0
- package/dist/engine/adapters/node-test-profile.d.ts.map +1 -0
- package/dist/engine/adapters/node-test-profile.js +14 -0
- package/dist/engine/adapters/node-test-profile.js.map +1 -0
- package/dist/engine/adapters/node-test-runtime.d.ts +39 -0
- package/dist/engine/adapters/node-test-runtime.d.ts.map +1 -0
- package/dist/engine/adapters/node-test-runtime.js +173 -0
- package/dist/engine/adapters/node-test-runtime.js.map +1 -0
- package/dist/engine/adapters/runtime-facts.d.ts +26 -0
- package/dist/engine/adapters/runtime-facts.d.ts.map +1 -0
- package/dist/engine/adapters/runtime-facts.js +73 -0
- package/dist/engine/adapters/runtime-facts.js.map +1 -0
- package/dist/engine/connection.d.ts +22 -0
- package/dist/engine/connection.d.ts.map +1 -0
- package/dist/engine/connection.js +343 -0
- package/dist/engine/connection.js.map +1 -0
- package/dist/engine/git-regression.d.ts +25 -0
- package/dist/engine/git-regression.d.ts.map +1 -0
- package/dist/engine/git-regression.js +803 -0
- package/dist/engine/git-regression.js.map +1 -0
- package/dist/engine/index.d.ts +55 -0
- package/dist/engine/index.d.ts.map +1 -0
- package/dist/engine/index.js +2782 -0
- package/dist/engine/index.js.map +1 -0
- package/dist/engine/node-test-reporter.d.ts +2 -0
- package/dist/engine/node-test-reporter.d.ts.map +1 -0
- package/dist/engine/node-test-reporter.js +70 -0
- package/dist/engine/node-test-reporter.js.map +1 -0
- package/dist/engine/runtime-doctor.d.ts +16 -0
- package/dist/engine/runtime-doctor.d.ts.map +1 -0
- package/dist/engine/runtime-doctor.js +100 -0
- package/dist/engine/runtime-doctor.js.map +1 -0
- package/dist/evaluation/agentic-corpus.d.ts +161 -0
- package/dist/evaluation/agentic-corpus.d.ts.map +1 -0
- package/dist/evaluation/agentic-corpus.js +710 -0
- package/dist/evaluation/agentic-corpus.js.map +1 -0
- package/dist/index.d.ts +8 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +8 -0
- package/dist/index.js.map +1 -0
- package/dist/mcp/index.d.ts +13 -0
- package/dist/mcp/index.d.ts.map +1 -0
- package/dist/mcp/index.js +391 -0
- package/dist/mcp/index.js.map +1 -0
- package/dist/mcp/stdio.d.ts +3 -0
- package/dist/mcp/stdio.d.ts.map +1 -0
- package/dist/mcp/stdio.js +13 -0
- package/dist/mcp/stdio.js.map +1 -0
- package/dist/sdk/index.d.ts +52 -0
- package/dist/sdk/index.d.ts.map +1 -0
- package/dist/sdk/index.js +224 -0
- package/dist/sdk/index.js.map +1 -0
- package/dist/version.d.ts +2 -0
- package/dist/version.d.ts.map +1 -0
- package/dist/version.js +10 -0
- package/dist/version.js.map +1 -0
- package/docs/adapter-protocol.md +196 -0
- package/docs/agentic-benchmark.md +118 -0
- package/docs/agentic-corpus-experiment-h3.md +89 -0
- package/docs/agentic-corpus-plan.md +105 -0
- package/docs/agentic-corpus-provenance.md +59 -0
- package/docs/agentic-test-profile-pilot.md +57 -0
- package/docs/agentic-test-profile-v2.md +116 -0
- package/docs/agentic-test-profile.md +274 -0
- package/docs/architecture.md +157 -0
- package/docs/ci.md +37 -0
- package/docs/client-connections.md +61 -0
- package/docs/conformance-v1.md +72 -0
- package/docs/decisions/0001-typescript-runtime.md +24 -0
- package/docs/developer-experience.md +55 -0
- package/docs/diagnostics.md +35 -0
- package/docs/distribution.md +40 -0
- package/docs/git-regression.md +39 -0
- package/docs/migration-repository-validation-order.md +35 -0
- package/docs/migration-testforge-to-assertledger.md +64 -0
- package/docs/project-intent.md +173 -0
- package/docs/proof-model.md +116 -0
- package/docs/reference.md +336 -0
- package/docs/release-1.0.md +63 -0
- package/docs/repository-audit.md +52 -0
- package/docs/repository-init.md +60 -0
- package/docs/research-basis.md +27 -0
- package/docs/roadmap.md +74 -0
- package/docs/runtime-doctor.md +65 -0
- package/docs/testexplora-calibration.md +71 -0
- package/examples/agentic-benchmark/benchmark-request.mjs +19 -0
- package/examples/agentic-benchmark/structured-phase-adapter-fixture.mjs +35 -0
- package/examples/agentic-profile/profile-benchmark.mjs +34 -0
- package/examples/agentic-profile/profile-manifest.mjs +28 -0
- package/examples/git-history/README.md +44 -0
- package/examples/git-history/create-demo.mjs +128 -0
- package/examples/git-history/escape-string-regexp/LICENSE +9 -0
- package/examples/git-history/escape-string-regexp/before.cjs.txt +11 -0
- package/examples/git-history/escape-string-regexp/fixed.cjs.txt +13 -0
- package/examples/git-history/escape-string-regexp/provenance.json +28 -0
- package/examples/node-test/repository/package.json +5 -0
- package/examples/node-test/repository/src/is-even.js +3 -0
- package/examples/node-test/repository/tests/base.test.js +6 -0
- package/examples/node-test/request.json +93 -0
- package/integrations/skill/SKILL.md +51 -0
- package/package.json +88 -0
- package/schemas/agentic-benchmark-acquisition-replay-result.v1.json +70 -0
- package/schemas/agentic-benchmark-acquisition-request.v1.json +564 -0
- package/schemas/agentic-benchmark-acquisition-result.v1.json +1409 -0
- package/schemas/agentic-benchmark-artifact.v1.json +1251 -0
- package/schemas/agentic-benchmark-replay-result.v1.json +84 -0
- package/schemas/agentic-benchmark-request.v1.json +1034 -0
- package/schemas/agentic-corpus-allocation-commitment-replay-result.v1.json +58 -0
- package/schemas/agentic-corpus-allocation-commitment.v1.json +141 -0
- package/schemas/agentic-corpus-allocation-replay-result.v1.json +34 -0
- package/schemas/agentic-corpus-allocation-request.v1.json +65 -0
- package/schemas/agentic-corpus-allocation-reveal.v1.json +66 -0
- package/schemas/agentic-corpus-allocation.v1.json +167 -0
- package/schemas/agentic-corpus-experiment-artifact.v1.json +329 -0
- package/schemas/agentic-corpus-experiment-plan-replay-result.v1.json +50 -0
- package/schemas/agentic-corpus-experiment-plan.v1.json +424 -0
- package/schemas/agentic-corpus-experiment-replay-request.v1.json +336 -0
- package/schemas/agentic-corpus-experiment-replay-result.v1.json +106 -0
- package/schemas/agentic-corpus-experiment-request.v1.json +204 -0
- package/schemas/agentic-corpus-provenance.v1.json +143 -0
- package/schemas/agentic-corpus-trust-policy.v1.json +133 -0
- package/schemas/agentic-profile-replay-result.v1.json +56 -0
- package/schemas/agentic-profile-replay-result.v2.json +63 -0
- package/schemas/agentic-profile-report.v1.json +961 -0
- package/schemas/agentic-profile-report.v2.json +1674 -0
- package/schemas/agentic-profile-request.v1.json +671 -0
- package/schemas/agentic-profile-request.v2.json +1338 -0
- package/schemas/evidence-manifest.v1.json +636 -0
- package/schemas/replay-result.v1.json +49 -0
- package/schemas/repository-analysis.v1.json +119 -0
- package/schemas/repository-audit.v1.json +811 -0
- package/schemas/repository-init-config.v1.json +183 -0
- package/schemas/repository-init-lock.v1.json +162 -0
- package/schemas/repository-init-result.v1.json +212 -0
- package/schemas/verification-request.v1.json +389 -0
|
@@ -0,0 +1,336 @@
|
|
|
1
|
+
# CLI, SDK and MCP reference
|
|
2
|
+
|
|
3
|
+
[Back to the introduction](../README.md) · [Introduction en français](../README.fr.md)
|
|
4
|
+
|
|
5
|
+
## JSON CLI
|
|
6
|
+
|
|
7
|
+
This reference uses the installed or locally linked `assertledger` command. When working from
|
|
8
|
+
the source branch, build first and replace `assertledger` with `node dist/cli.js`.
|
|
9
|
+
|
|
10
|
+
For a first run, use `assertledger doctor .` and the [Git qualification guide](git-regression.md).
|
|
11
|
+
`assertledger connect . --client codex` prints a read-only MCP configuration; see
|
|
12
|
+
[developer entry points](developer-experience.md) for creation, conflicts and trust requirements.
|
|
13
|
+
|
|
14
|
+
```sh
|
|
15
|
+
assertledger init . --dry-run --json
|
|
16
|
+
assertledger init . --json
|
|
17
|
+
assertledger audit . --json
|
|
18
|
+
assertledger analyze . --json
|
|
19
|
+
assertledger schema verification-request --json
|
|
20
|
+
assertledger verify assertledger.request.json --allow-unsafe-execution --json
|
|
21
|
+
assertledger replay assertledger.manifest.json --json
|
|
22
|
+
assertledger profile assertledger.profile-request.json --json
|
|
23
|
+
assertledger profile-replay assertledger.profile-report.json --json
|
|
24
|
+
assertledger profile-v2 assertledger.profile-v2-request.json --json
|
|
25
|
+
assertledger profile-v2-replay assertledger.profile-v2-report.json --json
|
|
26
|
+
assertledger benchmark assertledger.benchmark-request.json --json
|
|
27
|
+
assertledger benchmark-replay assertledger.benchmark-artifact.json --json
|
|
28
|
+
assertledger benchmark-acquire assertledger.benchmark-acquisition-request.json --allow-unsafe-execution --json
|
|
29
|
+
assertledger benchmark-acquire-replay assertledger.benchmark-acquisition-result.json --json
|
|
30
|
+
assertledger mcp
|
|
31
|
+
# Operator-only opt-in: assertledger mcp --allow-unsafe-execution
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Every command above also runs identically under the legacy `testforge` binary name.
|
|
35
|
+
|
|
36
|
+
`init` is static: it never executes detected commands or adapters and manages only
|
|
37
|
+
`assertledger.config.json` and `assertledger.lock.json`. It requires operator-owned worlds and
|
|
38
|
+
candidates rather than inventing a verification request. See
|
|
39
|
+
[`docs/repository-init.md`](repository-init.md) for detection conflicts, recovery behavior,
|
|
40
|
+
adapter availability, and the trust boundary.
|
|
41
|
+
|
|
42
|
+
Repository analysis reports a test framework only from framework-specific evidence such as a
|
|
43
|
+
declared dependency, a dedicated configuration file, or a real `node:test` import below a test
|
|
44
|
+
directory or in a colocated `*.test.*`/`*.spec.*` source file. Comments, string examples, and
|
|
45
|
+
documentation-like config filenames do not count; no evidence produces an empty list. Detection is
|
|
46
|
+
intentionally incomplete: an unrecognized manifest layout or test-file convention yields no claim.
|
|
47
|
+
|
|
48
|
+
`verify`, `replay`, `profile`, `profile-replay`, `profile-v2`, `profile-v2-replay`, `benchmark`, and
|
|
49
|
+
`benchmark-replay`, `benchmark-acquire`, and `benchmark-acquire-replay` also accept `-`
|
|
50
|
+
or an omitted file argument and
|
|
51
|
+
then read JSON from stdin. The CLI rejects file and stdin JSON inputs larger than 16 MiB. JSON
|
|
52
|
+
results go to stdout. Diagnostics go to stderr. `assertledger mcp` reserves stdout for JSON-RPC.
|
|
53
|
+
|
|
54
|
+
`--allow-unsafe-execution` is an external authorization signal. The CLI requires it for every
|
|
55
|
+
campaign and sets the request's local acknowledgement before validation. The flag does not create a
|
|
56
|
+
sandbox.
|
|
57
|
+
|
|
58
|
+
Versioned JSON Schemas are published for the
|
|
59
|
+
[`verification request`](../schemas/verification-request.v1.json),
|
|
60
|
+
[`repository analysis`](../schemas/repository-analysis.v1.json),
|
|
61
|
+
[`repository init config`](../schemas/repository-init-config.v1.json),
|
|
62
|
+
[`repository init lock`](../schemas/repository-init-lock.v1.json),
|
|
63
|
+
[`repository init result`](../schemas/repository-init-result.v1.json),
|
|
64
|
+
[`evidence manifest`](../schemas/evidence-manifest.v1.json),
|
|
65
|
+
[`replay result`](../schemas/replay-result.v1.json),
|
|
66
|
+
[`Agentic Test Profile request`](../schemas/agentic-profile-request.v1.json),
|
|
67
|
+
[`Agentic Test Profile report`](../schemas/agentic-profile-report.v1.json),
|
|
68
|
+
[`Agentic Test Profile replay result`](../schemas/agentic-profile-replay-result.v1.json),
|
|
69
|
+
[`Agentic Benchmark request`](../schemas/agentic-benchmark-request.v1.json),
|
|
70
|
+
[`Agentic Benchmark artifact`](../schemas/agentic-benchmark-artifact.v1.json), and
|
|
71
|
+
[`Agentic Benchmark replay result`](../schemas/agentic-benchmark-replay-result.v1.json),
|
|
72
|
+
[`Agentic Benchmark acquisition request`](../schemas/agentic-benchmark-acquisition-request.v1.json),
|
|
73
|
+
[`Agentic Benchmark acquisition result`](../schemas/agentic-benchmark-acquisition-result.v1.json),
|
|
74
|
+
[`Agentic Benchmark acquisition replay result`](../schemas/agentic-benchmark-acquisition-replay-result.v1.json),
|
|
75
|
+
[`Agentic Test Profile v2 request`](../schemas/agentic-profile-request.v2.json),
|
|
76
|
+
[`Agentic Test Profile v2 report`](../schemas/agentic-profile-report.v2.json), and
|
|
77
|
+
[`Agentic Test Profile v2 replay result`](../schemas/agentic-profile-replay-result.v2.json), plus the
|
|
78
|
+
[`corpus trust policy`](../schemas/agentic-corpus-trust-policy.v1.json) and paired
|
|
79
|
+
[`corpus provenance`](../schemas/agentic-corpus-provenance.v1.json), six corpus allocation and
|
|
80
|
+
two-party [`commitment/reveal`](../schemas/agentic-corpus-allocation-commitment.v1.json) contracts,
|
|
81
|
+
plus six pre-declared [`H3 experiment`](../schemas/agentic-corpus-experiment-plan.v1.json) contracts.
|
|
82
|
+
The main CLI `schema` command prints the thirty-one facade schemas by name; the corpus evaluator consumes
|
|
83
|
+
the two trust/provenance schemas directly.
|
|
84
|
+
A complete runnable verification request
|
|
85
|
+
is available at
|
|
86
|
+
[`examples/node-test/request.json`](../examples/node-test/request.json).
|
|
87
|
+
After building the package, the
|
|
88
|
+
[`Agentic Test Profile example`](../examples/agentic-profile/profile-manifest.mjs) derives a report
|
|
89
|
+
directly from any replay-valid manifest. The first two real-campaign results and their limitations
|
|
90
|
+
are recorded in [`docs/agentic-test-profile-pilot.md`](agentic-test-profile-pilot.md).
|
|
91
|
+
The [`Agentic Benchmark example`](../examples/agentic-benchmark/benchmark-request.mjs) consumes a
|
|
92
|
+
strict benchmark request. See [`docs/agentic-benchmark.md`](agentic-benchmark.md) for the fixed
|
|
93
|
+
protocol, comparison scope, replay rails, and current acquisition limitation.
|
|
94
|
+
The benchmark-backed [`Agentic Test Profile v2`](agentic-test-profile-v2.md) replaces v1's
|
|
95
|
+
campaign wall-time proxy with an exact scoped warm-total-wall p95 cost basis. It is additive: v1
|
|
96
|
+
artifacts and commands remain supported.
|
|
97
|
+
|
|
98
|
+
The checked-in [`conformance v1 bundle`](conformance-v1.md) locks autonomous inputs, complete
|
|
99
|
+
expected outputs, negative replay witnesses, all published schema bytes, and selected public
|
|
100
|
+
digests. `pnpm check` validates this static oracle without regenerating it.
|
|
101
|
+
|
|
102
|
+
## TypeScript SDK
|
|
103
|
+
|
|
104
|
+
```ts
|
|
105
|
+
import { readFile } from "node:fs/promises";
|
|
106
|
+
import { AssertLedger } from "assertledger";
|
|
107
|
+
|
|
108
|
+
const assertLedger = new AssertLedger();
|
|
109
|
+
const initialized = await assertLedger.init("/absolute/path/to/repository", { dryRun: true });
|
|
110
|
+
const context = await assertLedger.analyze("/absolute/path/to/repository");
|
|
111
|
+
const request = JSON.parse(await readFile("assertledger.request.json", "utf8"));
|
|
112
|
+
const manifest = await assertLedger.verify(request);
|
|
113
|
+
const integrity = assertLedger.replay(manifest);
|
|
114
|
+
|
|
115
|
+
const profile = assertLedger.profile({
|
|
116
|
+
schemaVersion: "1.0.0",
|
|
117
|
+
manifest,
|
|
118
|
+
policy: {
|
|
119
|
+
profileVersion: "1.0.0",
|
|
120
|
+
profileId: "my-repository/default",
|
|
121
|
+
mode: "HARDENING",
|
|
122
|
+
minimumTimingSamples: 3,
|
|
123
|
+
lanes: [
|
|
124
|
+
{ id: "instant", maximumReferenceP95Ms: 2_000 },
|
|
125
|
+
{ id: "loop", maximumReferenceP95Ms: 10_000 },
|
|
126
|
+
],
|
|
127
|
+
},
|
|
128
|
+
});
|
|
129
|
+
const profileIntegrity = assertLedger.replayProfile(profile);
|
|
130
|
+
|
|
131
|
+
const benchmarkRequest = JSON.parse(await readFile("assertledger.benchmark-request.json", "utf8"));
|
|
132
|
+
const benchmark = assertLedger.benchmark(benchmarkRequest);
|
|
133
|
+
const benchmarkIntegrity = assertLedger.replayBenchmark(benchmark);
|
|
134
|
+
|
|
135
|
+
const profileV2 = assertLedger.profileV2({
|
|
136
|
+
schemaVersion: "2.0.0",
|
|
137
|
+
benchmarkArtifact: benchmark,
|
|
138
|
+
policy: {
|
|
139
|
+
profileVersion: "2.0.0",
|
|
140
|
+
profileId: "my-repository/hardening-v2",
|
|
141
|
+
mode: "HARDENING",
|
|
142
|
+
requiredComparisonScopeDigest: benchmark.comparisonScopeDigest,
|
|
143
|
+
costBasis: {
|
|
144
|
+
regime: "WARM",
|
|
145
|
+
measure: "WALL",
|
|
146
|
+
aggregation: "TOTAL",
|
|
147
|
+
statistic: "P95",
|
|
148
|
+
unit: "MICROSECOND",
|
|
149
|
+
portfolioAggregation: "SUM_OF_INDIVIDUAL_P95",
|
|
150
|
+
},
|
|
151
|
+
lanes: [{ id: "loop", maximumWarmTotalWallP95Us: 10_000_000 }],
|
|
152
|
+
},
|
|
153
|
+
});
|
|
154
|
+
const profileV2Integrity = assertLedger.replayProfileV2(profileV2);
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
`AssertLedger` is exported alongside a deprecated `TestForge` subclass alias with an identical
|
|
158
|
+
surface, for consumers migrating from the prior name.
|
|
159
|
+
|
|
160
|
+
The SDK accepts plain JSON-compatible values and validates them against the same contracts as the
|
|
161
|
+
CLI. Unlike the CLI and MCP tool, `AssertLedger.verify()` has no separate authorization parameter: the
|
|
162
|
+
caller must set `isolation.acknowledgedUnsafeExecution` to `true` after applying its own policy.
|
|
163
|
+
|
|
164
|
+
`AssertLedger.replay()` reports schema validity, both digest checks, and deterministic
|
|
165
|
+
decision-semantic validity. Its aggregate `valid` field is true only when all four checks pass. Replay
|
|
166
|
+
does not rerun the campaign, authenticate the producer, or establish that the observations were
|
|
167
|
+
truthful.
|
|
168
|
+
|
|
169
|
+
`AssertLedger.profile()` accepts only a replay-valid manifest. It qualifies selected `ELIGIBLE`
|
|
170
|
+
hardening candidates against repository-scoped latency lanes and publishes evidence strength,
|
|
171
|
+
observed consistency, nearest-rank p50/p95, marginal target weight, and Pareto status. Latency never
|
|
172
|
+
compensates for a failed evidence gate. For each declared lane, the report also selects an
|
|
173
|
+
explainable greedy portfolio that maximizes new target weight per unit of recorded reference p95
|
|
174
|
+
within the lane's total budget. See
|
|
175
|
+
[`docs/agentic-test-profile.md`](agentic-test-profile.md) for the contract and its non-claims.
|
|
176
|
+
|
|
177
|
+
`AssertLedger.benchmark()` accepts only a replay-valid `VERIFIED` source manifest plus a complete,
|
|
178
|
+
fingerprinted cold/warm run plan. It summarizes declared microsecond phase timings and never changes
|
|
179
|
+
the source decision. `await assertLedger.acquireBenchmark()` runs the fresh campaign and strict
|
|
180
|
+
framework-neutral phase acquisition directly. The built-in `node:test` adapter remains unsupported
|
|
181
|
+
because it cannot faithfully expose all four phase boundaries.
|
|
182
|
+
`AssertLedger.replayBenchmarkAcquisition()` independently checks the source, artifact, context, result
|
|
183
|
+
digest, and exact status/reason semantics.
|
|
184
|
+
|
|
185
|
+
H3 replay requires three externally pinned digests and separate SDK options for the trust policy,
|
|
186
|
+
two-party commitment/reveal, allocation, pre-declared plan, subject provenance and evidence bytes.
|
|
187
|
+
Allocation is source-stratified; arms use content-addressed candidate references and derived suite
|
|
188
|
+
digests. Replay resolves and re-hashes every candidate's supplied bytes; this proves byte identity,
|
|
189
|
+
not that those bytes were executed. Scheduled process counts and planned timeout ceilings must be
|
|
190
|
+
exactly equal. H3
|
|
191
|
+
non-inferiority must hold per source and in aggregate. Receipts attest the applied timeout limit but
|
|
192
|
+
do not independently prove enforcement or equal observed CPU/wall time.
|
|
193
|
+
The CLI exposes the same boundary through mandatory `--trust-policy-digest`,
|
|
194
|
+
`--allocation-commitment-digest`, and `--experiment-plan-digest` flags plus separate files. See
|
|
195
|
+
[`docs/agentic-corpus-experiment-h3.md`](agentic-corpus-experiment-h3.md).
|
|
196
|
+
|
|
197
|
+
`AssertLedger.profileV2()` accepts only a replay-valid benchmark artifact. Its comparison scope must
|
|
198
|
+
match the policy, all warm measurements must be usable, and every cost is the declared warm total
|
|
199
|
+
wall p95 in microseconds. It publishes a Pareto frontier and deterministic greedy portfolios only
|
|
200
|
+
over source-selected eligible candidates. The portfolio cost is a sum of individual p95 values,
|
|
201
|
+
not a measured portfolio p95 and not a universal optimum.
|
|
202
|
+
|
|
203
|
+
## Calibration corpus
|
|
204
|
+
|
|
205
|
+
The versioned scaffold under [`benchmarks/agentic-profile`](../benchmarks/agentic-profile) evaluates
|
|
206
|
+
the falsifiable H1-H4 hypotheses without an LLM judge. It remains `NOT_READY` until at least 20
|
|
207
|
+
strict cases from three sources cover every hypothesis and include physically separate public and
|
|
208
|
+
private splits.
|
|
209
|
+
|
|
210
|
+
```sh
|
|
211
|
+
tsx scripts/evaluate-agentic-profile-corpus.ts status benchmarks/agentic-profile \
|
|
212
|
+
--trust-policy operator-policy.json --trust-policy-digest sha256:...
|
|
213
|
+
tsx scripts/evaluate-agentic-profile-corpus.ts evaluate-public benchmarks/agentic-profile \
|
|
214
|
+
--trust-policy operator-policy.json --trust-policy-digest sha256:...
|
|
215
|
+
tsx scripts/evaluate-agentic-profile-corpus.ts evaluate-holdout benchmarks/agentic-profile \
|
|
216
|
+
--trust-policy operator-policy.json --trust-policy-digest sha256:...
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
Public evaluation returns per-case reason codes. Holdout evaluation returns aggregate hypothesis
|
|
220
|
+
counts only and suppresses case identifiers, source identifiers, raw values, and readiness counts.
|
|
221
|
+
The repository intentionally contains no fabricated cases; private `*.case.json` files are ignored.
|
|
222
|
+
The first empirical calibration is frozen in the
|
|
223
|
+
[24-case corpus plan](agentic-corpus-plan.md). The plan fixes sources and admission rules; it
|
|
224
|
+
is not evidence that every tranche has already been executed. The receipt-linked
|
|
225
|
+
[TestExplora calibration](testexplora-calibration.md) contributes eight real, stable cases but
|
|
226
|
+
is explicitly curated calibration rather than holdout or `READY_H3` evidence.
|
|
227
|
+
|
|
228
|
+
## MCP v2 over stdio
|
|
229
|
+
|
|
230
|
+
Start the read-only server with `assertledger mcp`, or configure an MCP client with `assertledger` as
|
|
231
|
+
the command and `["mcp"]` as its argument array (`testforge mcp` runs the identical server). The
|
|
232
|
+
server name reported to clients is `assertledger`. Every tool is registered twice, under a preferred
|
|
233
|
+
`assertledger_*` name and a legacy `testforge_*` name bound to the same handler and the same tool
|
|
234
|
+
configuration. The schema lookup pair has no single fixed output schema because its selected JSON
|
|
235
|
+
Schema document varies; every other pair shares the same output-schema object. The default server
|
|
236
|
+
exposes twenty-eight read-only tools:
|
|
237
|
+
|
|
238
|
+
| Preferred tool | Legacy alias | Purpose |
|
|
239
|
+
| --- | --- | --- |
|
|
240
|
+
| `assertledger_analyze` | `testforge_analyze` | Produce repository context for test generation |
|
|
241
|
+
| `assertledger_doctor` | `testforge_doctor` | Return a static repository initialization plan without writing files or executing repository code |
|
|
242
|
+
| `assertledger_explain` | `testforge_explain` | Explain bounded reason codes with versioned guidance and safe next actions |
|
|
243
|
+
| `assertledger_benchmark` | `testforge_benchmark` | Derive scoped cold/warm phase summaries from declared raw runs |
|
|
244
|
+
| `assertledger_benchmark_replay` | `testforge_benchmark_replay` | Replay a self-contained benchmark artifact |
|
|
245
|
+
| `assertledger_profile` | `testforge_profile` | Derive an Agentic Test Profile from replay-valid evidence |
|
|
246
|
+
| `assertledger_profile_replay` | `testforge_profile_replay` | Replay a self-contained profile report |
|
|
247
|
+
| `assertledger_profile_v2` | `testforge_profile_v2` | Derive a benchmark-backed strength and warm-cost profile |
|
|
248
|
+
| `assertledger_profile_v2_replay` | `testforge_profile_v2_replay` | Replay a self-contained Profile v2 report |
|
|
249
|
+
| `assertledger_schema` | `testforge_schema` | Return any of the published JSON Schemas |
|
|
250
|
+
| `assertledger_replay` | `testforge_replay` | Validate and replay a manifest's schema, digests, and decision semantics |
|
|
251
|
+
| `assertledger_benchmark_acquire_replay` | `testforge_benchmark_acquire_replay` | Replay acquisition source, artifact, context, digest, and status bindings |
|
|
252
|
+
| `assertledger_corpus_allocate` | `testforge_corpus_allocate` | Create a deterministic calibration/holdout allocation |
|
|
253
|
+
| `assertledger_corpus_allocation_replay` | `testforge_corpus_allocation_replay` | Replay allocation scores, partition, and digest semantics |
|
|
254
|
+
|
|
255
|
+
H3 creation and replay are deliberately absent from the default agent-facing MCP server. Its
|
|
256
|
+
operator-owned policy, commitment and plan pins cannot be supplied safely as self-attested tool
|
|
257
|
+
input.
|
|
258
|
+
|
|
259
|
+
`assertledger_verify`/`testforge_verify` and `assertledger_benchmark_acquire`/`testforge_benchmark_acquire`
|
|
260
|
+
are absent by default, as are `assertledger_check`/`testforge_check` and
|
|
261
|
+
`assertledger_doctor_runtime`/`testforge_doctor_runtime`. A server operator may register these pairs by starting
|
|
262
|
+
`assertledger mcp --allow-unsafe-execution`, or by calling `createAssertLedgerServer({
|
|
263
|
+
allowUnsafeExecution: true })` (the deprecated `createTestForgeServer` alias calls the identical
|
|
264
|
+
factory). The tool then accepts a verification request and executes it without a second per-call
|
|
265
|
+
authorization field. Run that server only inside the intended isolation boundary; do not let an MCP
|
|
266
|
+
caller decide whether the capability exists. The server resolves repository roots to real paths and
|
|
267
|
+
confines them to the server process's current working directory by default. Programmatic operators
|
|
268
|
+
may supply a different `allowedRepositoryRoots` allowlist.
|
|
269
|
+
|
|
270
|
+
The doctor pair accepts a strict `{ "root": "..." }` input and returns the existing
|
|
271
|
+
`repository-init-result` contract. It is read-only in both the default and operator-enabled server;
|
|
272
|
+
enabling unsafe execution does not change doctor behavior. Dynamic runtime and client diagnostics
|
|
273
|
+
remain outside this static readiness result.
|
|
274
|
+
|
|
275
|
+
`doctor_runtime` accepts the same strict root input and returns the separate
|
|
276
|
+
[runtime diagnostic contract](runtime-doctor.md). `check` accepts the
|
|
277
|
+
[high-level Git options](git-regression.md), without a permission field, and returns the existing
|
|
278
|
+
evidence manifest. The operator's capability is required for both tools.
|
|
279
|
+
|
|
280
|
+
## Continuous integration
|
|
281
|
+
|
|
282
|
+
Run `pnpm check` on every change. The included GitHub Actions workflow runs this gate on Node.js 22
|
|
283
|
+
and 24 on Ubuntu and Windows. A separate matrix installs and exercises the packed artifact on both
|
|
284
|
+
operating systems with Node.js 22.15.0 and 24. A CI job that executes campaigns must
|
|
285
|
+
also treat `trusted-local` as `UNSANDBOXED`: use an isolated runner without secrets or host
|
|
286
|
+
credentials, and pass `--allow-unsafe-execution` only from reviewed CI configuration.
|
|
287
|
+
|
|
288
|
+
## Decision semantics
|
|
289
|
+
|
|
290
|
+
- `VERIFIED`: at least one candidate completed all required evidence and was selected.
|
|
291
|
+
- `REJECTED`: the campaign completed, but no candidate satisfied the policy.
|
|
292
|
+
- `INCONCLUSIVE`: controls or candidate evidence were incomplete, unstable, timed out, or affected
|
|
293
|
+
by infrastructure failure.
|
|
294
|
+
- `ENGINE_ERROR`: the deterministic core could not normalize the supplied evidence safely.
|
|
295
|
+
|
|
296
|
+
Only an attributed `ASSERTION_FAILURE` can kill a target in protocol v1. Compilation errors,
|
|
297
|
+
collection failures, crashes, timeouts, and infrastructure errors never count as target evidence.
|
|
298
|
+
|
|
299
|
+
Campaign budgets cover aggregate candidate, world, repository, overlay, and execution counts or
|
|
300
|
+
bytes. `timeoutMsPerExecution` and `maximumOutputBytes` apply to each process execution; the timeout
|
|
301
|
+
and process-tree termination are best effort on the local host.
|
|
302
|
+
|
|
303
|
+
Before a built-in `node:test` campaign runs, AssertLedger probes the requested executable, resolves its
|
|
304
|
+
real path, probes that resolved file again, and requires matching Node.js versions of at least
|
|
305
|
+
22.15. The manifest records the requested executable, resolved path, Node.js version, and executable
|
|
306
|
+
SHA-256 digest.
|
|
307
|
+
|
|
308
|
+
Each candidate records the ordered gates `COMPLETENESS`, `STABILITY`, `DISCOVERY`, `REFERENCE`,
|
|
309
|
+
`NEUTRAL`, and `TARGET_STRENGTH`, including evidence run IDs and reason codes. See
|
|
310
|
+
[docs/proof-model.md](proof-model.md) for candidate statuses, controls, selection, and digest
|
|
311
|
+
scope.
|
|
312
|
+
|
|
313
|
+
The manifest is sufficient to replay AssertLedger's deterministic decision, but it is not a complete
|
|
314
|
+
audit archive. It stores content and process-output digests, not candidate or world bodies, a
|
|
315
|
+
repository archive, or raw logs. Preserve the original request, repository snapshot or trusted
|
|
316
|
+
source reference, raw logs, dependencies, and any structured-command executable identity separately
|
|
317
|
+
when independent audit or reproduction matters. Manifest digests detect changes; they do not
|
|
318
|
+
authenticate the producer.
|
|
319
|
+
|
|
320
|
+
## Project status and roadmap
|
|
321
|
+
|
|
322
|
+
The [intent and restart assessment](project-intent.md) records the current implementation,
|
|
323
|
+
adoption gaps, and proposed delivery order. It distinguishes local evidence from release and
|
|
324
|
+
empirical claims.
|
|
325
|
+
|
|
326
|
+
Version `1.0.0` qualifies the supported node:test workflow. The
|
|
327
|
+
deterministic core and protocols are framework-independent; `node:test` is the first built-in
|
|
328
|
+
framework adapter. Other frameworks integrate through the structured-command protocol described in
|
|
329
|
+
[docs/adapter-protocol.md](adapter-protocol.md).
|
|
330
|
+
|
|
331
|
+
Planned work is not shipped behavior. Priorities include a real sandbox backend, additional
|
|
332
|
+
framework reporters with runtime attribution, signed provenance, cross-runtime conformance
|
|
333
|
+
fixtures, and more built-in adapters. See [docs/roadmap.md](roadmap.md).
|
|
334
|
+
|
|
335
|
+
Contributions are welcome under the [MIT license](../LICENSE). Read [CONTRIBUTING.md](../CONTRIBUTING.md)
|
|
336
|
+
before changing a public contract.
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# AssertLedger 1.0
|
|
2
|
+
|
|
3
|
+
Cap décidé par le propriétaire le 8 septembre 2026. Ce document fixe les critères de la version 1.0 ;
|
|
4
|
+
la publication est attestée séparément par le tag, le registre npm et la release GitHub.
|
|
5
|
+
Le nom du package reste `assertledger` et le dépôt public est
|
|
6
|
+
[hoklims/assertledger](https://github.com/hoklims/assertledger).
|
|
7
|
+
|
|
8
|
+
## Résultat utilisateur
|
|
9
|
+
|
|
10
|
+
Un développeur ou son agent apporte le test d’une correction. AssertLedger vérifie que ce test
|
|
11
|
+
passe sur la correction et les contrôles déclarés, puis détecte la faute connue par assertion.
|
|
12
|
+
Le résultat identifie les révisions, les octets du test, les essais, les raisons et les limites.
|
|
13
|
+
Le même artefact se rejoue sans modèle dans le terminal et en CI.
|
|
14
|
+
|
|
15
|
+
La première cible est une équipe Node.js/TypeScript utilisant des agents. La première exécution
|
|
16
|
+
officielle repose sur `node:test` ; les variantes de TypeScript ou de dépendances non qualifiées
|
|
17
|
+
doivent être annoncées comme telles. Détecter un framework ne signifie pas savoir le qualifier.
|
|
18
|
+
|
|
19
|
+
## Parcours et critères de livraison
|
|
20
|
+
|
|
21
|
+
| Étape | Critère observable | Issue |
|
|
22
|
+
| --- | --- | --- |
|
|
23
|
+
| Première preuve | Version fautive, correction, test et contrôle neutre déclarés ; test utile et inutile distingués ; manifeste rejouable | [HOK-658](https://linear.app/hoklims/issue/HOK-658) |
|
|
24
|
+
| Adaptateur de référence | Matrice adversariale, trois mondes, deux essais, candidat fort seul sélectionné | [HOK-570](https://linear.app/hoklims/issue/HOK-570) |
|
|
25
|
+
| Diagnostic | Prérequis, permissions et limites expliqués avant exécution ; prochaine action sûre | [HOK-420](https://linear.app/hoklims/issue/HOK-420) |
|
|
26
|
+
| Agents | CLI/SDK/MCP compatibles ; configuration idempotente et confinée aux racines autorisées | [HOK-423](https://linear.app/hoklims/issue/HOK-423) |
|
|
27
|
+
| Restitution | Verdict humain et JSON cohérents, raisons stables, aucune suggestion qui contourne un gate | [HOK-428](https://linear.app/hoklims/issue/HOK-428) |
|
|
28
|
+
| Distribution | Installation du tarball dans un consommateur neuf ; binaires, SDK, fichiers runtime et replay vérifiés | [HOK-659](https://linear.app/hoklims/issue/HOK-659) |
|
|
29
|
+
| Publication | CI et revue du candidat exact, tag/package/release publiés, réinstallation de la version publiée | [HOK-399](https://linear.app/hoklims/issue/HOK-399) |
|
|
30
|
+
|
|
31
|
+
La documentation doit être copiée et exécutée sans correction implicite. Mesurer le temps jusqu’à
|
|
32
|
+
la première preuve comprise ; cinq minutes reste un objectif, sans résultat annoncé d’avance.
|
|
33
|
+
Une démo locale n’est pas une preuve d’adoption. [Trois pilotes](https://linear.app/hoklims/issue/HOK-660)
|
|
34
|
+
doivent ensuite documenter réutilisation, effort, abandon et volonté de conserver l’intégration.
|
|
35
|
+
|
|
36
|
+
## Frontières de la version
|
|
37
|
+
|
|
38
|
+
- Les gates, schémas, projections de digest et aliases TestForge conservent leur compatibilité.
|
|
39
|
+
- Seule une assertion attribuée peut compter comme détection ; les erreurs opérationnelles restent
|
|
40
|
+
des non-kills. La vitesse ne compense aucun défaut de preuve.
|
|
41
|
+
- Les mondes attendus relèvent de l’opérateur. Un contrôle identique à la référence ne devient pas
|
|
42
|
+
une preuve indépendante de robustesse.
|
|
43
|
+
- `trusted-local` reste `UNSANDBOXED`, opt-in, réservé au code de confiance. Les exemples CI refusent
|
|
44
|
+
les contributions non fiables et l’exécution sur un runner privilégié.
|
|
45
|
+
- Le replay vérifie intégrité et cohérence ; il n’authentifie pas le producteur des observations.
|
|
46
|
+
- Profils statistiques et holdout gardent leurs critères avant toute affirmation empirique.
|
|
47
|
+
|
|
48
|
+
## Extensions conservées
|
|
49
|
+
|
|
50
|
+
Les adaptateurs Vitest, Jest, Bun et pytest suivent les besoins des pilotes. Cache incrémental,
|
|
51
|
+
landing automatique, profils de performance et isolation des contributions hostiles sont des
|
|
52
|
+
travaux distincts. Leurs critères restent dans Linear ; ils ne sont pas annoncés comme livrés par
|
|
53
|
+
la seule publication d’un package 1.0.
|
|
54
|
+
|
|
55
|
+
## État de départ vérifié
|
|
56
|
+
|
|
57
|
+
Le candidat local du 8 septembre passe `pnpm check` avec 307 tests. La version du package est
|
|
58
|
+
encore `0.1.0`. Le dépôt public dédié est créé et `origin` configuré ; aucune release n’est publiée.
|
|
59
|
+
Le contrôle npm initial retournait `E401`. L’authentification du compte personnel `hoklims` a
|
|
60
|
+
ensuite été vérifiée ; elle ne vaut pas preuve de publication.
|
|
61
|
+
Le diff hérité est conservé ; sa présence ne constitue pas une acceptation de tous ses travaux.
|
|
62
|
+
La preuve externe [HOK-406](https://linear.app/hoklims/issue/HOK-406) reste ouverte et ne donne pas
|
|
63
|
+
autorité pour modifier la politique active.
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Repository audit v1
|
|
2
|
+
|
|
3
|
+
`assertledger audit` inventories a repository without executing tests, adapters, candidates, or
|
|
4
|
+
campaigns. Its output is descriptive: it contains no score, note, quality label, coverage claim, or
|
|
5
|
+
verification verdict.
|
|
6
|
+
|
|
7
|
+
```sh
|
|
8
|
+
assertledger audit . --json
|
|
9
|
+
assertledger audit . --verification-request request.json --json
|
|
10
|
+
assertledger audit . --emit-verification-request --no-git | assertledger verify - --allow-unsafe-execution
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
The command discovers `assertledger.request.json` at the audited root unless
|
|
14
|
+
`--verification-request` is supplied. A request is emitted only after v1 validation and exact
|
|
15
|
+
repository-root binding. Emitted requests always have
|
|
16
|
+
`acknowledgedUnsafeExecution: false`; the receiving `verify` command owns the explicit unsafe
|
|
17
|
+
authorization. Exit code `3` means no request was available for `--emit-verification-request`, and
|
|
18
|
+
exit code `4` means the supplied or discovered request was invalid.
|
|
19
|
+
|
|
20
|
+
## Signals and limits
|
|
21
|
+
|
|
22
|
+
JavaScript and TypeScript files use the TypeScript lexical scanner. Decision points count `if`,
|
|
23
|
+
loop, `catch`, `case`, conditional, `&&`, `||`, and `??` tokens. Declared exports and test-root
|
|
24
|
+
identifier occurrences produce `noObservedTestReference`; this name deliberately does not claim
|
|
25
|
+
coverage. Test association is conservative: source and `.test`/`.spec` files must share a basename.
|
|
26
|
+
Weak assertion shapes are syntactic occurrences of truthiness-only, constant-boolean equality,
|
|
27
|
+
definedness-only (`toBeDefined`), type-only (`typeof` asserted against a type name), and
|
|
28
|
+
non-throw-only forms. Unsupported languages report `supported: false` and nullable measurements.
|
|
29
|
+
|
|
30
|
+
Git history is optional. The 90-day window is anchored to the commit timestamp of `HEAD`, never the
|
|
31
|
+
wall clock, and is capped at 1,000 commits. Correction commits are messages containing the explicit
|
|
32
|
+
words `fix`, `bug`, `repair`, `hotfix`, or `regression` (case-insensitive). Missing, shallow, failed,
|
|
33
|
+
or truncated history is represented by nulls and reason codes, not invented zeroes.
|
|
34
|
+
|
|
35
|
+
Modules use `PATH_PREFIX_V1` (the first two path segments when available). Their contiguous order is
|
|
36
|
+
the descending factual tuple: decision points, exports with no observed test reference, correction
|
|
37
|
+
commits, commits in the window, sources newer than associated tests, weak assertion occurrences;
|
|
38
|
+
the POSIX path prefix breaks ties. The tuple is not a score.
|
|
39
|
+
|
|
40
|
+
Campaign bounds use decimal BigInt strings. For candidates `C`, worlds `W`, attempts `A`, repository
|
|
41
|
+
bytes `R`, total world-overlay bytes `OW`, total candidate bytes `OC`, timeout `T`, and output limit
|
|
42
|
+
`O`:
|
|
43
|
+
|
|
44
|
+
- executions = `(C + 1) * W * A`
|
|
45
|
+
- controls = `W * A`; candidate executions = `C * W * A`
|
|
46
|
+
- overlay bytes = `A * ((C + 1) * OW + W * OC)`
|
|
47
|
+
- materialization bytes = `R * (executions + 1) + overlay bytes`
|
|
48
|
+
- timeout limit = `executions * T`
|
|
49
|
+
- captured stream limit = `executions * O * 2`
|
|
50
|
+
|
|
51
|
+
Two complete inventories bind the audit. A digest change aborts with
|
|
52
|
+
`REPOSITORY_CHANGED_DURING_AUDIT`.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Repository initialization
|
|
2
|
+
|
|
3
|
+
`assertledger init` performs a static, provider-neutral repository detection pass and manages exactly
|
|
4
|
+
two root files: `assertledger.config.json` and `assertledger.lock.json`. It never runs a test command,
|
|
5
|
+
adapter, candidate, or campaign, and it never edits package manifests, lockfiles, CI configuration,
|
|
6
|
+
or tests.
|
|
7
|
+
|
|
8
|
+
```sh
|
|
9
|
+
assertledger init . --dry-run --json
|
|
10
|
+
assertledger init . --json
|
|
11
|
+
assertledger audit . --json
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
The dry run emits the exact canonical bytes and SHA-256 digests that a subsequent write plans to
|
|
15
|
+
use. Writes use a same-directory temporary file followed by an atomic rename. A second run returns
|
|
16
|
+
`UNCHANGED` without rewriting matching files. A missing lock or a stale, structurally valid lock
|
|
17
|
+
bound to the same config can be recovered; invalid files and config-binding conflicts fail closed
|
|
18
|
+
without a `--force` mode.
|
|
19
|
+
|
|
20
|
+
Detection is conservative. Package-manager declarations and root lockfiles must converge. Multiple
|
|
21
|
+
plausible test frameworks, composite shell scripts, contradictory overrides, and invalid explicit
|
|
22
|
+
adapter configurations return `CONFLICT`. Multiple CI providers are only sorted evidence and do not
|
|
23
|
+
block initialization.
|
|
24
|
+
|
|
25
|
+
Files at or below the managed `candidateRoots` are deliberately excluded from framework inference,
|
|
26
|
+
evidence, built-in control tests, and repository-change comparison. Candidate generation therefore
|
|
27
|
+
cannot silently redefine initialization facts or invalidate an otherwise unchanged lock.
|
|
28
|
+
|
|
29
|
+
The built-in ready adapter is currently `node-test`. Bun, pytest, Vitest, and Jest can be detected,
|
|
30
|
+
but initialization returns `BLOCKED` with `OFFICIAL_ADAPTER_UNAVAILABLE` unless the operator supplies
|
|
31
|
+
an existing structured adapter configuration:
|
|
32
|
+
|
|
33
|
+
```sh
|
|
34
|
+
assertledger init . --adapter-config integrations/my-adapter.json --json
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
The adapter document is parsed through the public adapter contract and is operator-owned. Its
|
|
38
|
+
executable is recorded as argv but is not resolved or executed by `init`. This is not an official
|
|
39
|
+
adapter endorsement and does not reduce the later `trusted-local` execution boundary.
|
|
40
|
+
`node-test` adapters are accepted only for the `node:test` framework; every other framework requires
|
|
41
|
+
an operator-owned `testforge-command` adapter.
|
|
42
|
+
|
|
43
|
+
All detections and evidence digests come from one byte snapshot. Immediately before returning or
|
|
44
|
+
writing managed files, initialization rechecks the in-scope inventory and every evidence byte. A
|
|
45
|
+
change returns `CONFLICT` with `REPOSITORY_CHANGED_DURING_INIT` and writes nothing.
|
|
46
|
+
|
|
47
|
+
Initialization deliberately does not generate a verification request. Both `worlds` and
|
|
48
|
+
`candidates` remain required operator or harness inputs because repository detection cannot infer
|
|
49
|
+
valid semantic faults, neutral transformations, or oracle expectations. The next command is
|
|
50
|
+
`assertledger audit . --json`; that audit may report `VERIFICATION_REQUEST_UNAVAILABLE` until the
|
|
51
|
+
harness supplies those inputs.
|
|
52
|
+
|
|
53
|
+
Exit codes are `0` for `CREATED`, `UNCHANGED`, or `WOULD_CREATE`; `3` for `BLOCKED`; `4` for
|
|
54
|
+
ambiguity, invalid overrides, conflicts, or contract validation; `5` for unexpected I/O; and `64`
|
|
55
|
+
for CLI usage errors.
|
|
56
|
+
|
|
57
|
+
The public `repository-init-config.v1`, `repository-init-lock.v1`, and
|
|
58
|
+
`repository-init-result.v1` contracts contain no timestamps, absolute repository roots, environment
|
|
59
|
+
values, worlds, or candidates. The lock binds normalized detections and sorted evidence digests; it
|
|
60
|
+
does not authenticate the detector, repository, adapter, or later execution evidence.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Research basis
|
|
2
|
+
|
|
3
|
+
AssertLedger's evidence model follows several established testing and provenance principles.
|
|
4
|
+
|
|
5
|
+
- Mutation testing distinguishes killed mutants from live mutants and reveals oracle weakness that
|
|
6
|
+
coverage alone cannot show: [PIT](https://pitest.org/) and
|
|
7
|
+
[Just et al., FSE 2014](https://homes.cs.washington.edu/~mernst/pubs/mutation-effectiveness-fse2014-abstract.html).
|
|
8
|
+
- Repetition measures observed flakiness; a successful retry does not erase a contradictory run:
|
|
9
|
+
[Gruber et al., 2021](https://arxiv.org/abs/2101.09077) and
|
|
10
|
+
[Lam et al., ISSTA 2019](https://www.microsoft.com/en-us/research/publication/root-causing-flaky-tests-in-a-large-scale-industrial-setting/).
|
|
11
|
+
- Selection is stated only over executed evidence; “safe” regression-test selection requires stronger
|
|
12
|
+
assumptions: [Rothermel and Harrold, IEEE TSE 1998](https://digitalcommons.unl.edu/csearticles/11/).
|
|
13
|
+
- The manifest separates external parameters, resolved inputs, builder identity, and byproducts,
|
|
14
|
+
inspired by [SLSA build provenance v1.2](https://slsa.dev/spec/v1.2/build-provenance). AssertLedger does
|
|
15
|
+
not claim a SLSA level.
|
|
16
|
+
- The proposed Agentic Test Profile keeps evidence strength and observed execution cost as separate
|
|
17
|
+
dimensions. This follows cost-aware regression-test prioritization rather than treating line
|
|
18
|
+
coverage as an oracle-quality score: [Elbaum, Rothermel, and Penix, FSE
|
|
19
|
+
2014](https://research.google/pubs/techniques-for-improving-regression-testing-in-continuous-integration-development-environments/)
|
|
20
|
+
and [Meta predictive test
|
|
21
|
+
selection](https://engineering.fb.com/2018/11/21/developer-tools/predictive-test-selection/).
|
|
22
|
+
- Mutation-guided agentic testing now has industrial hardening and catching-test evidence. The two
|
|
23
|
+
modes have different reference directions and must remain distinct: [Mutation-Guided LLM-based
|
|
24
|
+
Test Generation at Meta](https://arxiv.org/abs/2501.12862) and [Just-in-Time Catching Test
|
|
25
|
+
Generation at Meta](https://arxiv.org/abs/2601.22832).
|
|
26
|
+
|
|
27
|
+
These sources motivate the design; they do not turn mutant detection into a proof of correctness.
|
package/docs/roadmap.md
ADDED
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Roadmap
|
|
2
|
+
|
|
3
|
+
This roadmap separates locally implemented v0.1 behavior from the agreed 1.0 delivery scope.
|
|
4
|
+
See the [intent assessment](project-intent.md) and [1.0 release contract](release-1.0.md) for the
|
|
5
|
+
regression-test qualification workflow, acceptance criteria, and evidence boundaries.
|
|
6
|
+
|
|
7
|
+
## Implemented locally in v0.1
|
|
8
|
+
|
|
9
|
+
- deterministic repository analysis;
|
|
10
|
+
- thirty-four versioned public JSON Schemas, all frozen by conformance v1;
|
|
11
|
+
- reference, target, and neutral overlay worlds;
|
|
12
|
+
- one campaign snapshot and disposable execution workspaces;
|
|
13
|
+
- built-in `node:test` runtime reporter and structured-command adapter;
|
|
14
|
+
- deterministic gates, candidate selection, reason codes, and evidence run links;
|
|
15
|
+
- decision and artifact integrity digests;
|
|
16
|
+
- derived Agentic Test Profiles with replay, repository-scoped latency lanes, marginal evidence,
|
|
17
|
+
and a multi-dimensional Pareto frontier;
|
|
18
|
+
- replayable Agentic Benchmark Artifacts with fingerprinted comparison scopes, cold/warm regimes,
|
|
19
|
+
deterministic phase quantiles, and explicit sample/failure states;
|
|
20
|
+
- fresh phase-aware benchmark acquisition for `testforge-command`, gated by replay-valid
|
|
21
|
+
`VERIFIED` evidence and bound to adapter and dependency identities;
|
|
22
|
+
- benchmark-backed Agentic Test Profile v2 reports with an exact warm-total-wall p95 cost basis,
|
|
23
|
+
selected-only cohort disclosure, deterministic Pareto comparison, and summed-cost portfolios;
|
|
24
|
+
- externally pinned Ed25519 corpus trust policies and canonical dual-signed provenance sidecars;
|
|
25
|
+
- deterministic source-stratified corpus allocations with dual-principal commitment/reveal and
|
|
26
|
+
private holdout custody;
|
|
27
|
+
- externally pinned H3 experiment plans and artifacts with byte-resolved run receipts, exact
|
|
28
|
+
planned-process and planned-timeout equality, content-addressed suites, per-source
|
|
29
|
+
non-inferiority, and a shared frozen candidate universe;
|
|
30
|
+
- a static conformance-v1 oracle locking canonicalization, decisions, replay witnesses, Profile v1,
|
|
31
|
+
Benchmark v1, and all 34 published schema bytes;
|
|
32
|
+
- a pinned 24-case, three-source empirical corpus plan with signed admission and holdout rules,
|
|
33
|
+
including eight receipt-linked TestExplora cases admitted only for curated calibration;
|
|
34
|
+
- JSON CLI, TypeScript SDK, MCP v2 stdio server, and integration skill;
|
|
35
|
+
- `trusted-local`, explicitly recorded as `UNSANDBOXED`.
|
|
36
|
+
|
|
37
|
+
## Release 1.0 priority
|
|
38
|
+
|
|
39
|
+
1. Qualify a regression test against declared buggy, fixed, and neutral revisions using existing
|
|
40
|
+
deterministic gates, with useful diagnostics and preserved artifacts.
|
|
41
|
+
2. Verify installation and the documented workflow from the packed package in a fresh consumer.
|
|
42
|
+
3. Add repository diagnostics and agent integration, then a PR/CI report and replay path.
|
|
43
|
+
4. Publish only after exact-candidate checks and required independent review; install and verify
|
|
44
|
+
the published version before closing delivery.
|
|
45
|
+
|
|
46
|
+
The first built-in adapter is `node:test`. Broader adapters and optimizations follow pilot needs.
|
|
47
|
+
The empirical work below remains required for scientific claims, rather than for releasing the
|
|
48
|
+
scoped deterministic qualification workflow.
|
|
49
|
+
|
|
50
|
+
## Further evidence and extensions
|
|
51
|
+
|
|
52
|
+
1. Complete the [24-case corpus plan](agentic-corpus-plan.md), including provenance sidecars, a
|
|
53
|
+
separately selected and pre-committed holdout, reviewed historical faults, mutants, and neutral
|
|
54
|
+
rewrites, before claiming an empirically established sweet spot. The current TestExplora 8/8
|
|
55
|
+
tranche is curated calibration evidence and cannot satisfy the holdout requirement.
|
|
56
|
+
2. Add faithful framework-specific phase adapters. The framework-neutral acquisition path is
|
|
57
|
+
shipped, but the built-in `node:test` reporter cannot attribute all four phases.
|
|
58
|
+
3. Add an isolation backend backed by an independently administered container or VM boundary.
|
|
59
|
+
4. Add framework reporters beyond `node:test` that derive discovery and attribution from runtime
|
|
60
|
+
events.
|
|
61
|
+
5. Publish cross-runtime conformance fixtures for canonicalization, decisions, and both digests.
|
|
62
|
+
6. Add mature replayable experiment artifacts for H1, H2, and H4 without allowing them to
|
|
63
|
+
compensate for an H3 failure.
|
|
64
|
+
7. Add adapters only when their outcome mapping has adversarial contract tests.
|
|
65
|
+
8. Extend provenance to cover executed dependencies, structured-command binary identities, and
|
|
66
|
+
effective environment inputs without leaking secrets.
|
|
67
|
+
|
|
68
|
+
## Explicitly out of scope for v0.1
|
|
69
|
+
|
|
70
|
+
- hostile-code containment;
|
|
71
|
+
- automatic trust in agent-supplied worlds or policy;
|
|
72
|
+
- proof of program correctness or complete fault detection;
|
|
73
|
+
- proof that a test will never become flaky;
|
|
74
|
+
- SLSA conformance or authenticated provenance.
|