assertledger 1.1.0 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.fr.md +14 -4
- package/README.md +14 -4
- package/conformance/schema-extensions.json +25 -0
- package/dist/build-info.json +1 -1
- package/dist/cli.d.ts.map +1 -1
- package/dist/cli.js +68 -20
- package/dist/cli.js.map +1 -1
- package/dist/contracts/index.d.ts +510 -4
- package/dist/contracts/index.d.ts.map +1 -1
- package/dist/contracts/index.js +230 -16
- package/dist/contracts/index.js.map +1 -1
- package/dist/contracts/runtime-doctor.d.ts +123 -0
- package/dist/contracts/runtime-doctor.d.ts.map +1 -1
- package/dist/contracts/runtime-doctor.js +64 -0
- package/dist/contracts/runtime-doctor.js.map +1 -1
- package/dist/core/index.d.ts.map +1 -1
- package/dist/core/index.js +16 -6
- package/dist/core/index.js.map +1 -1
- package/dist/diagnostics.d.ts.map +1 -1
- package/dist/diagnostics.js +26 -2
- package/dist/diagnostics.js.map +1 -1
- package/dist/engine/adapters/bun-test-profile.d.ts +17 -0
- package/dist/engine/adapters/bun-test-profile.d.ts.map +1 -0
- package/dist/engine/adapters/bun-test-profile.js +17 -0
- package/dist/engine/adapters/bun-test-profile.js.map +1 -0
- package/dist/engine/git-regression.d.ts +2 -2
- package/dist/engine/git-regression.d.ts.map +1 -1
- package/dist/engine/git-regression.js.map +1 -1
- package/dist/engine/index.d.ts +11 -5
- package/dist/engine/index.d.ts.map +1 -1
- package/dist/engine/index.js +454 -49
- package/dist/engine/index.js.map +1 -1
- package/dist/engine/runtime-doctor.d.ts +13 -4
- package/dist/engine/runtime-doctor.d.ts.map +1 -1
- package/dist/engine/runtime-doctor.js +80 -2
- package/dist/engine/runtime-doctor.js.map +1 -1
- package/dist/mcp/index.d.ts.map +1 -1
- package/dist/mcp/index.js +36 -6
- package/dist/mcp/index.js.map +1 -1
- package/dist/proof-planner/carry-over.d.ts +21 -0
- package/dist/proof-planner/carry-over.d.ts.map +1 -0
- package/dist/proof-planner/carry-over.js +212 -0
- package/dist/proof-planner/carry-over.js.map +1 -0
- package/dist/proof-planner/index.d.ts +11 -0
- package/dist/proof-planner/index.d.ts.map +1 -0
- package/dist/proof-planner/index.js +11 -0
- package/dist/proof-planner/index.js.map +1 -0
- package/dist/proof-planner/model.d.ts +272 -0
- package/dist/proof-planner/model.d.ts.map +1 -0
- package/dist/proof-planner/model.js +215 -0
- package/dist/proof-planner/model.js.map +1 -0
- package/dist/proof-planner/plan.d.ts +140 -0
- package/dist/proof-planner/plan.d.ts.map +1 -0
- package/dist/proof-planner/plan.js +1124 -0
- package/dist/proof-planner/plan.js.map +1 -0
- package/dist/proof-planner/policy.d.ts +100 -0
- package/dist/proof-planner/policy.d.ts.map +1 -0
- package/dist/proof-planner/policy.js +407 -0
- package/dist/proof-planner/policy.js.map +1 -0
- package/dist/proof-planner/render.d.ts +4 -0
- package/dist/proof-planner/render.d.ts.map +1 -0
- package/dist/proof-planner/render.js +73 -0
- package/dist/proof-planner/render.js.map +1 -0
- package/dist/sdk/index.d.ts +7 -6
- package/dist/sdk/index.d.ts.map +1 -1
- package/dist/sdk/index.js +20 -7
- package/dist/sdk/index.js.map +1 -1
- package/docs/adapter-protocol.md +61 -0
- package/docs/architecture.md +11 -1
- package/docs/developer-experience.md +7 -5
- package/docs/distribution.md +8 -3
- package/docs/migration-timeout-discovery-inconclusive.md +126 -0
- package/docs/migration-verification-v3.md +33 -0
- package/docs/project-intent.md +7 -5
- package/docs/proof-model.md +20 -2
- package/docs/proof-planner.md +368 -0
- package/docs/reference.md +9 -6
- package/docs/repository-init.md +38 -7
- package/docs/runtime-doctor.md +12 -9
- package/integrations/bun/assertions.d.mts +4 -0
- package/integrations/bun/assertions.mjs +25 -0
- package/integrations/bun/driver.d.mts +27 -0
- package/integrations/bun/driver.mjs +383 -0
- package/integrations/bun/preload.mjs +165 -0
- package/integrations/skill/SKILL.md +4 -1
- package/package.json +8 -4
- package/schemas/evidence-manifest.v3.json +897 -0
- package/schemas/repository-init-config.v2.json +210 -0
- package/schemas/repository-init-lock.v2.json +162 -0
- package/schemas/repository-init-result.v2.json +212 -0
- package/schemas/verification-request.v3.json +484 -0
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# Verification v3 and Bun test migration
|
|
2
|
+
|
|
3
|
+
Verification request and evidence manifest v3 add the built-in `bun-test` adapter. V1 and v2
|
|
4
|
+
schema bytes, adapter unions, and historical manifests remain unchanged and replayable. V3 keeps
|
|
5
|
+
the v2 execution-backend record and accepts `node-test` and `testforge-command` as before.
|
|
6
|
+
`bun-test` requires `trusted-local`; a v3 Bun request with container isolation is refused before
|
|
7
|
+
execution.
|
|
8
|
+
|
|
9
|
+
`assertledger init` writes repository initialization config, lock, and result v2 for a detected
|
|
10
|
+
`bun:test` repository with one or more base test files. Existing Node initialization remains v1.
|
|
11
|
+
If a repository contains multiple test frameworks, select `--framework bun:test` explicitly.
|
|
12
|
+
Local-only linked entries still need explicit `--exclude NAME` declarations. Static initialization
|
|
13
|
+
does not execute tests or qualify the installed Bun binary. Run the unsafe runtime doctor, then
|
|
14
|
+
a v3 campaign with operator-supplied worlds and candidates.
|
|
15
|
+
|
|
16
|
+
New Bun candidate tests use `assertSame` from `assertledger/bun`. Existing `bun:test` tests can
|
|
17
|
+
remain as base controls. Their native `expect` failures block a campaign as operational failures;
|
|
18
|
+
they are not assertion evidence. The helper compares only with `Object.is`, so use an explicit
|
|
19
|
+
scalar or boolean observation when structural equality is needed. Calculation errors occur before
|
|
20
|
+
the helper and remain ordinary errors. The helper is installed from the package into each
|
|
21
|
+
disposable workspace; no adapter implementation is required in the consumer repository.
|
|
22
|
+
|
|
23
|
+
The Bun runtime profile pins version `1.4.2` and revision
|
|
24
|
+
`744846f844374847c902b5e7fd59b4342a51ef99`. New manifests bind the resolved executable
|
|
25
|
+
digest, Bun revision, driver digest, preload digest, helper digest, command shape,
|
|
26
|
+
and fresh runtime preflight. They may have new decision and artifact digests because v3 evidence
|
|
27
|
+
has a new schema identity. No old manifest is resealed. The v3 schemas are added to the
|
|
28
|
+
post-conformance schema lock; the frozen v1 bundle and its root digest are unchanged.
|
|
29
|
+
|
|
30
|
+
The v1 provider/export contracts still describe v1 manifests only. Consumers of v3 Bun evidence
|
|
31
|
+
should parse and replay the v3 manifest directly; they must not project it through the v1 export
|
|
32
|
+
contract. Windows, Linux, and macOS support is claimed only for the exact Bun version on which
|
|
33
|
+
the corresponding CI campaign gate passes.
|
package/docs/project-intent.md
CHANGED
|
@@ -50,13 +50,15 @@ référence, un monde neutre rouge, une observation instable ou une attribution
|
|
|
50
50
|
|
|
51
51
|
Le socle local comprend les contrats versionnés, le noyau déterministe, les digests, le replay
|
|
52
52
|
sémantique, le moteur d’exécution, les façades CLI/SDK/MCP, l’audit statique et l’initialisation.
|
|
53
|
-
Le registre contient
|
|
54
|
-
|
|
53
|
+
Le registre contient 45 schémas JSON : les 34 figés par le verrou de conformance v1 et 11
|
|
54
|
+
extensions, verrouillées sans modifier le bundle v1.
|
|
55
55
|
L’ancien décompte de 30 dans la roadmap était périmé.
|
|
56
56
|
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
57
|
+
Les adaptateurs officiels intégrés sont `node:test` et `bun:test`. Ce dernier est limité à Bun 1.4.2
|
|
58
|
+
et aux assertions explicites du helper `assertSame` ; les échecs natifs de `expect` ne deviennent pas
|
|
59
|
+
des preuves d’assertion. Le protocole `testforge-command` permet d’intégrer un autre framework avec
|
|
60
|
+
un adaptateur fourni par l’opérateur. Détecter Vitest, Jest ou pytest pendant `init` ne fournit pas
|
|
61
|
+
leur adaptateur officiel.
|
|
60
62
|
|
|
61
63
|
Les profils v1 et v2, les artefacts de benchmark, les contrats de provenance et le protocole
|
|
62
64
|
expérimental H3 sont présents. Le benchmark par phases refuse encore l’adaptateur `node:test` :
|
package/docs/proof-model.md
CHANGED
|
@@ -55,15 +55,33 @@ Candidate statuses are:
|
|
|
55
55
|
|
|
56
56
|
- `ELIGIBLE`: every gate needed by policy passed;
|
|
57
57
|
- `WEAK_ORACLE`: reference and neutral evidence passed, but target strength did not;
|
|
58
|
-
- `INVALID`: discovery, reference
|
|
58
|
+
- `INVALID`: a completed run disproved discovery, or reference or neutral evidence failed;
|
|
59
59
|
- `UNSTABLE`: complete attempts produced different normalized outcomes;
|
|
60
60
|
- `INCONCLUSIVE`: required observations were missing or execution produced `TIMEOUT` or
|
|
61
61
|
`INFRA_ERROR`.
|
|
62
62
|
|
|
63
|
+
When several apply, the first gate that fails decides, and inconclusive execution takes precedence
|
|
64
|
+
over later failures:
|
|
65
|
+
|
|
66
|
+
1. missing observations make the candidate `INCONCLUSIVE` (`COMPLETENESS`);
|
|
67
|
+
2. diverging attempts make it `UNSTABLE` (`STABILITY`);
|
|
68
|
+
3. a completed run without an attributed candidate test makes it `INVALID`
|
|
69
|
+
(`CANDIDATE_DISCOVERY_INVALID`);
|
|
70
|
+
4. when only `TIMEOUT` or `INFRA_ERROR` runs lack an attributed candidate test, `DISCOVERY` is
|
|
71
|
+
not established and the candidate is `INCONCLUSIVE` (`CANDIDATE_EXECUTION_INCONCLUSIVE`): such a
|
|
72
|
+
run never reached a verdict, so it can neither prove nor disprove discovery;
|
|
73
|
+
5. after discovery, any `TIMEOUT` or `INFRA_ERROR` run makes the candidate `INCONCLUSIVE` before a
|
|
74
|
+
red reference or neutral world makes it `INVALID`.
|
|
75
|
+
|
|
76
|
+
See [the migration note](migration-timeout-discovery-inconclusive.md) for the effect on existing
|
|
77
|
+
manifests.
|
|
78
|
+
|
|
63
79
|
## Campaign statuses
|
|
64
80
|
|
|
65
81
|
- `VERIFIED`: at least one eligible candidate was selected.
|
|
66
|
-
- `REJECTED`: evidence was complete and conclusive, but no candidate was eligible.
|
|
82
|
+
- `REJECTED`: evidence was complete and conclusive, but no candidate was eligible. Once attempts are
|
|
83
|
+
complete and agree, a completed run that disproves discovery is conclusive, even if other runs
|
|
84
|
+
timed out.
|
|
67
85
|
- `INCONCLUSIVE`: controls were invalid, or at least one candidate was unstable or inconclusive.
|
|
68
86
|
- `ENGINE_ERROR`: the core could not normalize the supplied evidence safely.
|
|
69
87
|
|
|
@@ -0,0 +1,368 @@
|
|
|
1
|
+
# Proof planner (internal module)
|
|
2
|
+
|
|
3
|
+
The proof planner answers one question: **given a change, the claims it reaches, their criticality
|
|
4
|
+
and an explicit assurance policy, which evidence is proportionate?** It lives in
|
|
5
|
+
`src/proof-planner/`, is pure and deterministic, and is not exported from the package entry points.
|
|
6
|
+
|
|
7
|
+
```text
|
|
8
|
+
impact provider (e.g. a semantic context tool) describes what MAY be affected
|
|
9
|
+
│ ChangeImpact (provider-neutral)
|
|
10
|
+
▼
|
|
11
|
+
proof planner decides what MUST be proved
|
|
12
|
+
│ AssurancePlan
|
|
13
|
+
▼
|
|
14
|
+
AssertLedger decides whether required evidence exists and is
|
|
15
|
+
valid: sufficiency, provenance, freshness, identity
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
The planner never claims that evidence exists. AssertLedger never decides what is proportionate.
|
|
19
|
+
V1 implements the planner only; the AssertLedger-side satisfaction check is future work.
|
|
20
|
+
|
|
21
|
+
## 1. Audit of the existing model
|
|
22
|
+
|
|
23
|
+
| Concept | Present | Where | What exists / what is missing |
|
|
24
|
+
| --- | --- | --- | --- |
|
|
25
|
+
| Claims | No | `src/contracts/index.ts` (profile `consistency.claim`) | The only `claim` field is a stability label. No claim id, criticality or claim-to-evidence link. The implicit claim of a campaign is "this candidate test detects the declared fault". |
|
|
26
|
+
| Evidence | Yes | `EvidenceObservation`, gates, `EvidenceManifest` v1/v2, evidence export | One kind only: repeated test executions across REFERENCE/TARGET/NEUTRAL worlds plus candidate-free controls, scoped to one campaign. Typecheck, CI jobs, reviews or corpora are not AssertLedger evidence. |
|
|
27
|
+
| Provenance | Partial | world `provenance` string, `assertledger-git-regression/1`, provider `sourceRevision` | Declared, unauthenticated (`authenticity: UNAUTHENTICATED`). Git commit/tree only inside a world provenance string. |
|
|
28
|
+
| Freshness | No | export `cost.execution.freshness: "UNKNOWN"`, capability `EXECUTION_FRESHNESS: UNSUPPORTED` | No timestamp, expiry or validity window; the core forbids the clock. The only usable proxy is digest identity. |
|
|
29
|
+
| Candidate identity | Partial | `repositoryDigest` (content digest of the snapshot), git-regression revisions | "Candidate" means a candidate **test**, not a candidate change. No first-class commit or tree of the change under proof. |
|
|
30
|
+
| Invalidation | No | replay checks integrity only | Nothing marks evidence stale when the repository changes. Outputs are append-only. |
|
|
31
|
+
| Confidence | Partial | export `confidence.level: "REPLAY_CONSISTENT_UNAUTHENTICATED"` | Constant, qualitative; `established` / `notEstablished` token lists are a good residual-uncertainty vocabulary. |
|
|
32
|
+
| Verdicts | Yes | `VERIFIED/REJECTED/INCONCLUSIVE/ENGINE_ERROR`, candidate statuses, gate `PASSED/FAILED/NOT_RUN` | All verdicts are about test-evidence quality for one campaign; none says "this change is sufficiently proved". |
|
|
33
|
+
| Infra vs product | Partial | outcome taxonomy; `TIMEOUT`/`INFRA_ERROR` are inconclusive | Observation-level only. Nothing distinguishes a change to the proof infrastructure from a change to the product. |
|
|
34
|
+
|
|
35
|
+
### Where sufficiency is implicit or global
|
|
36
|
+
|
|
37
|
+
These rules are correct for their own purpose and are **not** weakened by the planner:
|
|
38
|
+
|
|
39
|
+
- `pnpm check` is the single completion gate for every change, and CI runs it on the full matrix
|
|
40
|
+
without path filters.
|
|
41
|
+
- One invalid control makes a whole campaign `INCONCLUSIVE`; one timeout attempt among passes makes a
|
|
42
|
+
candidate `UNSTABLE`; every required target must be killed.
|
|
43
|
+
- Consumer `obligations` are `COVERED` only when all are `EXECUTED`, and `EXECUTED` means "observed at
|
|
44
|
+
least once", not "satisfied".
|
|
45
|
+
- The init lock digests every test, lockfile and CI file; one changed byte makes the runtime doctor
|
|
46
|
+
report `RUNTIME_CONFIGURATION_STALE`, whatever the change touches.
|
|
47
|
+
- The decision digest folds product identity (repository, worlds) and proof-infrastructure identity
|
|
48
|
+
(engine, adapter, budgets, backend) into one value: a budget change re-identifies the evidence.
|
|
49
|
+
|
|
50
|
+
What was missing is upstream of these gates: a planner that decides **whether** a given gate is
|
|
51
|
+
proportionate for a change, and **which** earlier evidence stays admissible after a follow-up commit.
|
|
52
|
+
|
|
53
|
+
## 2. Model
|
|
54
|
+
|
|
55
|
+
### Inputs
|
|
56
|
+
|
|
57
|
+
- `ChangeImpact`: the revision (git object id or `sha256:` digest), its merge-base `baseline`, an
|
|
58
|
+
optional `runtimeTreeDigest`, and the surfaces changed since the baseline (cumulative). Each
|
|
59
|
+
surface has a role (`product`, `execution-context`, `evaluation`, `test`, `proof-infrastructure`,
|
|
60
|
+
`documentation`), a reach (`direct`, `transitive`), a runtime (`none`, `build`,
|
|
61
|
+
`offline-analysis`, `live`), boundaries (protocol, persistence, replay, analysis, decision,
|
|
62
|
+
admission, security, public-api, benchmark, holdout), a declared behavior change (`none`,
|
|
63
|
+
`suspected`, `fix`, `feature`), coverage, environment sensitivity, an optional typed content
|
|
64
|
+
digest, and for test and proof-infrastructure surfaces the surfaces they exercise. The analysis
|
|
65
|
+
states its method (`static-graph` or `declared`), completeness, uncertainty, localized unknowns
|
|
66
|
+
and omitted surfaces.
|
|
67
|
+
- `AssuranceClaim[]`: the project's claim registry or the claims a change makes. Pass the registry,
|
|
68
|
+
not a hand-picked subset: untouched claims are filtered by the planner, not by the caller.
|
|
69
|
+
- `ProofSignal[]`: what proof runs revealed, for one revision each. A signal may also name the
|
|
70
|
+
baseline and runtime tree it was observed on: an unresolved signal (a product signal, or a failure
|
|
71
|
+
that may exercise the change) then still counts on a later revision whose product is
|
|
72
|
+
byte-identical, so a proof-only commit cannot erase a regression.
|
|
73
|
+
- `AssurancePolicy`: versioned data, digested into the plan (default `DEFAULT_ASSURANCE_POLICY`).
|
|
74
|
+
|
|
75
|
+
### Facts, not a level table
|
|
76
|
+
|
|
77
|
+
Code derives a closed vocabulary of **facts** (`product:fix`, `boundary:replay:behavior`,
|
|
78
|
+
`runtime:live:behavior`, `claim:critical:direct`, `impact:unbounded`, `observed:regression`, …).
|
|
79
|
+
The policy is data that maps facts to:
|
|
80
|
+
|
|
81
|
+
- **floors**: the minimal level a fact implies;
|
|
82
|
+
- **raises**: one step up, once per subject, capped at P4 (a raise never creates P5);
|
|
83
|
+
- **triggers**: evidence a fact requires or recommends;
|
|
84
|
+
- **status**: facts that put the plan on `HOLD` or `BLOCKED`.
|
|
85
|
+
|
|
86
|
+
Levels are computed **per subject** (product, evaluation, proof infrastructure, documentation) and
|
|
87
|
+
the plan level is their maximum. A fact may concern several subjects. A level baseline (P1 static
|
|
88
|
+
checks and affected tests, P2 targeted regression and revision identity, P3 targeted integration and
|
|
89
|
+
production-path test, P4 independent review, P5 preregistration and provenance) only applies to the
|
|
90
|
+
subjects that reached that level, and only to the impact surfaces of that subject: a holdout change
|
|
91
|
+
at P5 does not drag product ceremony in, and criticality deepens proof without widening it. Breadth
|
|
92
|
+
(full suite, full corpus, multi-environment) comes only from impact facts.
|
|
93
|
+
|
|
94
|
+
Subjects follow what the evidence is about, not only the role of the surface:
|
|
95
|
+
|
|
96
|
+
- benchmark and holdout boundaries always concern the evaluation subject, whatever role carries them;
|
|
97
|
+
- security, admission and decision boundaries keep their weight on a changed gate or oracle
|
|
98
|
+
(proof-infrastructure surfaces, and test surfaces whose oracle changed), under the
|
|
99
|
+
proof-infrastructure subject;
|
|
100
|
+
- within the impact, a claim is reached through the product only by product, execution-context or
|
|
101
|
+
evaluation surfaces (outside an unbounded impact, its surfaces are treated as reached product).
|
|
102
|
+
A high or critical claim whose test or gate changed (the test lists a claim surface in
|
|
103
|
+
`exercises`, or the claim lists the test) gets `claim:<criticality>:oracle` under the
|
|
104
|
+
proof-infrastructure subject: an `ORACLE_WITNESS` (the changed oracle still fails where the
|
|
105
|
+
violation is present) and, for a critical claim, an independent review, but no product
|
|
106
|
+
requalification;
|
|
107
|
+
- a documentation claim whose documented surfaces change behavior requires the documentation check;
|
|
108
|
+
otherwise it is unaffected.
|
|
109
|
+
|
|
110
|
+
| Level | Typical source facts |
|
|
111
|
+
| --- | --- |
|
|
112
|
+
| P0 | documentation only |
|
|
113
|
+
| P1 | refactor; changed tests that no product evidence runs; proof infrastructure; evaluation touched |
|
|
114
|
+
| P2 | product behavior change, new behavior, execution context, compatibility boundary touched; oracle of a high claim changed |
|
|
115
|
+
| P3 | behavior across protocol, persistence, replay, analysis or public API; live surface touched; security boundary touched; high claim reached directly; oracle of a critical claim changed; unbounded impact; costly reversal |
|
|
116
|
+
| P4 | live behavior; decision, admission or security behavior; critical claim reached directly; system-scope claim; irreversible change; observed corpus divergence or live runtime |
|
|
117
|
+
| P5 | empirical claim; benchmark or holdout behavior |
|
|
118
|
+
|
|
119
|
+
### Impact bound and exemptions
|
|
120
|
+
|
|
121
|
+
| Bound | When | Effect on heavy evidence |
|
|
122
|
+
| --- | --- | --- |
|
|
123
|
+
| `no-product-runtime` | a complete, static, low-uncertainty analysis finds no product, execution-context or evaluation surface | may be `NOT REQUIRED` |
|
|
124
|
+
| `bounded-confident` | the same analysis, with runtime surfaces | may be `NOT REQUIRED`; untouched claims are unaffected |
|
|
125
|
+
| `bounded-uncertain` | declared analysis, partial with named unknowns, non-low uncertainty or localized unknowns | computed against the **worst case** of the enumerated surfaces; what the worst case needs becomes `RECOMMENDED`, the rest `NOT REQUIRED`; a high or critical claim outside the impact earns a recommended invariant check, never a floor |
|
|
126
|
+
| `unbounded` | unknown completeness, unnamed unknowns, omitted surfaces, execution context changed, unknowns touching live or high-critical surfaces, observed unknown dependency or unplanned impact, whatever the enumerated surfaces are | full test suite `REQUIRED`; claims outside the impact are treated as reached transitively; everything not selected is `UNDETERMINED`, never `NOT REQUIRED` |
|
|
127
|
+
|
|
128
|
+
The bound is checked before the roles: an incomplete analysis of what looks like a proof-only change
|
|
129
|
+
is unbounded, because the omitted part may be product. An unbounded impact is treated as reaching
|
|
130
|
+
the product. The worst case keeps the declared analysis, so it is never more trusting than the plan;
|
|
131
|
+
if the worst case is itself unbounded, nothing can be exempted.
|
|
132
|
+
|
|
133
|
+
Size is not uncertainty: a large, enumerated transitive set bounds the targeted evidence's scope,
|
|
134
|
+
it does not trigger a global suite. A declared impact recommends the full suite, because surfaces
|
|
135
|
+
it omits are not bounded.
|
|
136
|
+
|
|
137
|
+
Every evidence kind of the catalog ends in exactly one status: required, recommended, not required
|
|
138
|
+
or undetermined. Each `NOT REQUIRED` entry carries its activation condition (the triggers and level
|
|
139
|
+
baselines that would require it, with current values) and states whether it was evaluated against
|
|
140
|
+
the actual change or the worst case.
|
|
141
|
+
|
|
142
|
+
### Product evidence versus proof-infrastructure evidence
|
|
143
|
+
|
|
144
|
+
Product signals (`UNEXPECTED_BEHAVIOR`, `UNPLANNED_IMPACT`, `UNKNOWN_DEPENDENCY`,
|
|
145
|
+
`LIVE_RUNTIME_TOUCHED`, `CORPUS_DIVERGENCE`, `WITNESS_NOT_CAUSAL`, `REGRESSION`) become facts.
|
|
146
|
+
A regression blocks; unexpected behavior and missing causality hold the plan and raise it. A
|
|
147
|
+
regression, an unexpected behavior or a corpus divergence that reproduces on the baseline is a
|
|
148
|
+
pre-existing defect: it does not block or escalate, but the reproduction becomes required evidence
|
|
149
|
+
(`FAILURE_ATTRIBUTION`). An unplanned impact reopens the bound unless the impact already names the
|
|
150
|
+
observed surfaces; a live-runtime observation escalates unless every observed surface is already
|
|
151
|
+
declared live.
|
|
152
|
+
|
|
153
|
+
Infrastructure signals (`TIMEOUT`, `ENVIRONMENT_FAILURE`, `TOOLING_FAILURE`) are attributed in a
|
|
154
|
+
fixed order, after the bound is known:
|
|
155
|
+
|
|
156
|
+
1. `exercises` is recomputed: if the failing job's surfaces are changed runtime surfaces, or changed
|
|
157
|
+
proof surfaces that exercise one, it is `yes`, whatever was declared. A changed test that
|
|
158
|
+
exercises nothing changed stays proof evidence, so a repaired ceiling can still be attributed
|
|
159
|
+
when it flakes. A declared `no` is only trusted under a confident bound and when the failure
|
|
160
|
+
names its surfaces.
|
|
161
|
+
2. A job that does not exercise the change is attributed to the proof infrastructure by any
|
|
162
|
+
admissible basis (baseline reproduction, pass on the same revision, outside impact, reported
|
|
163
|
+
infrastructure error, declared environment factor).
|
|
164
|
+
3. A job that may exercise the change is attributed only by `REPRODUCES_ON_BASELINE`. A retry that
|
|
165
|
+
passes proves nondeterminism, not innocence: a slower code path can time out once and pass once.
|
|
166
|
+
An attribution that rests on the reproduction alone requires `FAILURE_ATTRIBUTION`.
|
|
167
|
+
4. Otherwise the failure is unattributed: the plan holds and requires `FAILURE_ATTRIBUTION`, but the
|
|
168
|
+
level does **not** rise automatically.
|
|
169
|
+
|
|
170
|
+
An attributed infrastructure failure never changes the level, the status or the evidence required
|
|
171
|
+
on the changed surfaces; it adds `AFFECTED_JOB_RERUN` on the failing job, and `FAILURE_ATTRIBUTION`
|
|
172
|
+
there too when a reproduction on the baseline is its only basis. An unattributed one adds
|
|
173
|
+
`FAILURE_ATTRIBUTION` on the failing job; it widens the product evidence only to a failing job that
|
|
174
|
+
is itself a product surface of the impact, never to a test, a harness or a job outside the impact.
|
|
175
|
+
A proof-infrastructure surface without an `exercises` list is taken not to exercise the change when
|
|
176
|
+
its failure is classified, so an outside-impact attribution declared for it stands under a
|
|
177
|
+
confident bound.
|
|
178
|
+
|
|
179
|
+
### Evidence bindings and carry-over
|
|
180
|
+
|
|
181
|
+
Each kind declares what it stays valid for:
|
|
182
|
+
|
|
183
|
+
- `exact-revision`: static checks and revision identity (cheap, re-run on every revision);
|
|
184
|
+
- `surface-content`: valid per scoped surface while that surface, the proof surfaces that exercise
|
|
185
|
+
it (a test without an `exercises` list exercises everything; a proof-infrastructure surface
|
|
186
|
+
counts where it names the surface, and a directly changed one whose behavior change is not
|
|
187
|
+
declared `none` and that names nothing is reported as `proof.exercises-unknown`), the baseline and the runtime tree
|
|
188
|
+
keep their digests. Only documentation-only evidence about a surface both plans call
|
|
189
|
+
documentation ignores the runtime tree: a gate delta
|
|
190
|
+
review is redone when the code its budget measures changes, and a documentation claim scoped to a
|
|
191
|
+
runtime surface is checked again when the behavior it describes may have changed. Without a
|
|
192
|
+
scope, it is revision-wide;
|
|
193
|
+
- `runtime-tree`: full suite, full corpus, multi-environment, live shadow, system requalification,
|
|
194
|
+
benchmark and holdout runs: valid while the baseline and runtime tree digest are unchanged.
|
|
195
|
+
|
|
196
|
+
`carryOverEvidence(previous, next)` lists what may be reused and what must be produced again. Reuse
|
|
197
|
+
is keyed on the kind's semantic digest (id, weight, binding, subjects, verification,
|
|
198
|
+
`semanticsVersion`), never on its wording and never on the whole policy digest. Missing digests, a
|
|
199
|
+
changed baseline (rebase) or a changed runtime tree force re-production of every kind that depends
|
|
200
|
+
on them, and evidence never crosses to another revision on the surfaces that a signal unresolved in
|
|
201
|
+
either plan may concern, read in both plans: the surfaces it names and what a named test or proof
|
|
202
|
+
surface lists as exercised, or everything when it names no surface, a surface either impact does not
|
|
203
|
+
know, an execution context, or a proof surface that does not list what it exercises. A signal the
|
|
204
|
+
next plan observes counts too, because it contradicts evidence produced before it. Within one
|
|
205
|
+
revision, only an observation the previous plan did not already have counts, compared by what was
|
|
206
|
+
observed (id, revision, signal, classification, exercises, attribution basis and surfaces), not by
|
|
207
|
+
id: evidence produced beside an observation answers it. Evidence that answers observations, such as
|
|
208
|
+
a failure attribution or a job rerun, is reused only for the observations it was produced for, each
|
|
209
|
+
scoped by the surfaces it names; exact-revision evidence is reused within its revision whatever was
|
|
210
|
+
observed. Otherwise, a fact that names a signal its plan does not list, which only a stored or
|
|
211
|
+
foreign plan can carry, blocks the reuse of the evidence behind it. A block lifts when the previous
|
|
212
|
+
plan is planned again; the next plan does not revise its classification. A failure attributed to the
|
|
213
|
+
proof infrastructure, or unattributed on a job that a confidently bounded impact places outside the
|
|
214
|
+
change, concerns no product evidence. AssertLedger must still verify that reused evidence exists and
|
|
215
|
+
carries the listed digests.
|
|
216
|
+
|
|
217
|
+
### Escalations
|
|
218
|
+
|
|
219
|
+
Each plan precomputes, by replanning with a canonical hypothetical signal, what every product
|
|
220
|
+
signal and an attributed or unattributed infrastructure failure would do to its level, status and
|
|
221
|
+
required evidence. The tests pin every row of the bounded replay fix with values derived from the
|
|
222
|
+
policy by hand, and check the product rows for consistency with a replan through the public API
|
|
223
|
+
(the same planner, so that check alone would not catch a planner error).
|
|
224
|
+
|
|
225
|
+
## 3. Examples
|
|
226
|
+
|
|
227
|
+
Produced by the V1 default policy; the tests in `tests/proof-planner.test.ts` assert these
|
|
228
|
+
decisions.
|
|
229
|
+
|
|
230
|
+
| Scenario | Level | Required | Heavy evidence |
|
|
231
|
+
| --- | --- | --- | --- |
|
|
232
|
+
| A documentation only | P0 | documentation check | all 11 broad or ceremonial kinds not required |
|
|
233
|
+
| B local refactor | P1 | static checks, affected tests (+ characterization tests if untested) | all not required |
|
|
234
|
+
| C local bugfix | P2 | causal witness, affected tests, targeted regression, static checks, revision identity | all not required |
|
|
235
|
+
| D bounded replay fix | P3 | C + targeted integration, production-path test, targeted corpus, documentation check | all not required, including full corpus and system requalification |
|
|
236
|
+
| D, declared impact | P3 | D | full test suite, characterization tests and an invariant check for the untouched critical claim recommended; the rest not required against the worst case |
|
|
237
|
+
| D, partial analysis | P3 | D + full test suite, invariant check (the untouched critical claim is treated as reached) | full corpus and independent review recommended; everything else undetermined |
|
|
238
|
+
| E live decoder fix | P4 | C + boundary compatibility, targeted integration, production-path test, live shadow, rollback plan, full corpus, independent review | benchmark, holdout, full suite, multi-environment, system requalification not required |
|
|
239
|
+
| E tactical decision | P4 | acceptance test, affected tests, targeted regression and integration, production-path test, live shadow, rollback plan, full corpus, system requalification, independent review, static checks, revision identity | benchmark, holdout, full suite, multi-environment not required |
|
|
240
|
+
| F holdout evaluator + empirical claim | P5 | preregistration, provenance, benchmark protocol, holdout evaluation, contamination check, independent review, affected tests, static checks, revision identity | live shadow, full suite, full corpus, system requalification not required |
|
|
241
|
+
| G timeout ceiling repair | P1 | gate delta review, affected job rerun, static checks | all not required; no functional requalification |
|
|
242
|
+
| G, the repaired test checks a critical claim | P3 (proof infrastructure) | G + oracle witness, independent review of the test, revision identity | no product requalification: every requirement is scoped to the changed test or revision-wide |
|
|
243
|
+
|
|
244
|
+
Rendered plan for the bounded replay fix (abridged):
|
|
245
|
+
|
|
246
|
+
```text
|
|
247
|
+
AssurancePlan P3 PROVE - revision 1111… (baseline 0000…)
|
|
248
|
+
policy assertledger.default-assurance@1.0.0; impact bounded-confident; product P3, documentation P0
|
|
249
|
+
|
|
250
|
+
WHY
|
|
251
|
+
- impact: bounded-confident: static analysis, complete, low uncertainty
|
|
252
|
+
- claim: live-decoder-pins-unchanged: none of its runtime surfaces is in an impact bounded with confidence
|
|
253
|
+
- floor P3 boundary:replay:behavior: channel/replay-fold, channel/replay-safe-keys: replay behavior may change
|
|
254
|
+
- floor P2 product:behavior-change: channel/replay-safe-keys: product behavior may change
|
|
255
|
+
|
|
256
|
+
REQUIRED
|
|
257
|
+
- CAUSAL_WITNESS (surface-content) [channel/replay-safe-keys] <- product:fix
|
|
258
|
+
- TARGETED_CORPUS (surface-content) [channel/replay-fold, channel/replay-safe-keys] <- boundary:analysis:behavior, boundary:replay:behavior
|
|
259
|
+
- PRODUCTION_PATH_TEST, TARGETED_INTEGRATION, TARGETED_REGRESSION, AFFECTED_TESTS, REVISION_IDENTITY, STATIC_CHECKS, DOCUMENTATION_CHECK
|
|
260
|
+
|
|
261
|
+
NOT REQUIRED
|
|
262
|
+
- FULL_CORPUS: Not required: the policy asks for FULL_CORPUS when one of [boundary:decision:behavior,
|
|
263
|
+
impact:unbounded, observed:corpus-divergence, observed:live-runtime, runtime:live:behavior] holds;
|
|
264
|
+
none holds for this change, whose impact is bounded with confidence.
|
|
265
|
+
- SYSTEM_REQUALIFICATION, INDEPENDENT_REVIEW, FULL_TEST_SUITE, BENCHMARK_PROTOCOL, HOLDOUT_EVALUATION, …
|
|
266
|
+
|
|
267
|
+
ESCALATE IF
|
|
268
|
+
- CORPUS_DIVERGENCE (product): Evidence about the product: level moves P3 -> P4, status PROVE, adds FAILURE_ATTRIBUTION, FULL_CORPUS, INDEPENDENT_REVIEW.
|
|
269
|
+
- REGRESSION (product): Evidence about the product: level stays P3, status BLOCKED, adds nothing.
|
|
270
|
+
- TIMEOUT (infrastructure-attributed): Evidence about the proof infrastructure: level stays P3, status PROVE, adds AFFECTED_JOB_RERUN, FAILURE_ATTRIBUTION.
|
|
271
|
+
- TIMEOUT (infrastructure-unattributed): Unattributed failure: level stays P3, status HOLD, adds FAILURE_ATTRIBUTION.
|
|
272
|
+
- …
|
|
273
|
+
```
|
|
274
|
+
|
|
275
|
+
## 4. Before and after: a localized fix with two peripheral timeouts
|
|
276
|
+
|
|
277
|
+
Abstract reproduction of an observed consumer pull request: a three-line fix in a replay-only
|
|
278
|
+
adapter, legitimate targeted proof, then two wall-clock ceilings in untouched packages that timed
|
|
279
|
+
out one after the other (a property at 30.4 s for 30 s that had passed in 12.2 s on the same tree,
|
|
280
|
+
then a performance guard at 2.78 s for 1.5 s whose literal ignored the declared CI latency factor).
|
|
281
|
+
Observed cost before: 3 candidate commits, 5 CI runs, about 6 fresh audits; each ceiling repair
|
|
282
|
+
created a new commit that invalidated the whole aggregate proof although the product claims never
|
|
283
|
+
changed.
|
|
284
|
+
|
|
285
|
+
The test suite replays the sequence with the default policy and with a modeled revision-bound
|
|
286
|
+
global policy (every kind bound to the exact revision, full suite and independent audit at P3):
|
|
287
|
+
|
|
288
|
+
| Step | Before (revision-bound global model) | After (default policy) |
|
|
289
|
+
| --- | --- | --- |
|
|
290
|
+
| r1, replay fix | P3, 12 required kinds | P3, 10 required kinds (full suite and audit not required) |
|
|
291
|
+
| timeout 1 (outside impact, passed on the same tree) | proof infrastructure | proof infrastructure; level, status and product evidence unchanged |
|
|
292
|
+
| r1 → r2, first ceiling routed | 13 kinds re-produced | 4 kinds: static checks, revision identity, gate delta review on the ceiling, job reruns on the ceiling and on the job of timeout 2 |
|
|
293
|
+
| timeout 2 (outside impact, environment factor) | proof infrastructure | proof infrastructure |
|
|
294
|
+
| r2 → r3, second ceiling routed | 13 kinds re-produced | 4 kinds, on the new ceiling only; the first ceiling's review is reused |
|
|
295
|
+
|
|
296
|
+
The planner stays conservative on the same path: a corpus divergence moves the plan to P4 with the
|
|
297
|
+
full corpus; a timeout on a job that exercises the replay surface holds the plan unless it
|
|
298
|
+
reproduces on the baseline; a rebase, a changed runtime tree, a missing digest, a test without an
|
|
299
|
+
`exercises` list, a harness that names the replay surface, or a product digest changed by a
|
|
300
|
+
so-called infrastructure commit forces the product evidence to be produced again; and a regression
|
|
301
|
+
left unresolved on r1 blocks the reuse of the evidence on its surfaces. Each of these is a test.
|
|
302
|
+
|
|
303
|
+
## 5. Mapping to neighbours
|
|
304
|
+
|
|
305
|
+
**Impact provider.** Any provider can fill `ChangeImpact`. From semctx, for example:
|
|
306
|
+
`changedFiles`/`changedSymbols` give direct surfaces; `impactedConsumers` and `control_impact`
|
|
307
|
+
paths give transitive ones; `impactedInvariants`/`impactedContracts` and change-contract
|
|
308
|
+
`preserves` give claims (criticality from `criticalInvariantTags`); `unknowns` and the
|
|
309
|
+
`analysis_scope_incomplete` / `index_binding_stale` findings give completeness and unknowns;
|
|
310
|
+
`recommendedTests` only feed test surfaces, never required evidence. semctx reports neither a
|
|
311
|
+
candidate commit nor content digests through MCP: the caller supplies `revision`, `baseline` and
|
|
312
|
+
digests. Such an adapter belongs outside `src/` (a test forbids that name in source files).
|
|
313
|
+
|
|
314
|
+
**AssertLedger.** Today AssertLedger can verify `CAUSAL_WITNESS` and `ACCEPTANCE_TEST` as a
|
|
315
|
+
qualification campaign (TARGET = baseline, REFERENCE = revision, NEUTRAL = justified variation) with
|
|
316
|
+
replay-valid manifests. Other kinds (static checks, CI jobs, reviews, corpora, preregistration) have
|
|
317
|
+
no AssertLedger evidence type yet; kinds marked `attested` can only be recorded, not executed.
|
|
318
|
+
|
|
319
|
+
## 6. Limits and risks
|
|
320
|
+
|
|
321
|
+
- Inputs are trusted declarations. A caller can mislabel a live surface as offline or omit a surface;
|
|
322
|
+
the planner limits the damage (declared analysis never yields confident exemptions, contradicted
|
|
323
|
+
`exercises` values are overridden, incomplete analyses are unbounded whatever they enumerate,
|
|
324
|
+
untouched critical claims come back as recommendations under uncertainty and as requirements when
|
|
325
|
+
unbounded) but cannot detect a consistent lie. A policy-owned surface classifier is not in V1.
|
|
326
|
+
- Claims come from the caller. Pass the full registry; V1 has no registry of its own. A changed proof
|
|
327
|
+
surface without an `exercises` list only reaches the claims that list it; the plan reports it as
|
|
328
|
+
residual uncertainty (`proof.exercises-unknown`).
|
|
329
|
+
- The default policy is a first calibration, not a measured optimum. A caller-supplied policy must be
|
|
330
|
+
at least as strict as the default: every kind (verification, binding, subjects), baseline, floor,
|
|
331
|
+
trigger, raise and status rule of the default must still hold with at least the same strength, or
|
|
332
|
+
parsing fails with `PROOF_PLANNER_POLICY_BELOW_MINIMUM`. The tests pin the default by digest and
|
|
333
|
+
check its non-negotiable rules against an independent list. A project that wants a looser policy
|
|
334
|
+
must fork the default; that is deliberate in V1.
|
|
335
|
+
- Signals carried across revisions need the caller to report the baseline and runtime tree they were
|
|
336
|
+
observed on. Without them, the carry-over still refuses reuse on the surfaces an unresolved signal
|
|
337
|
+
may concern, but the next plan does not hold or block by itself. An unattributed failure on a job
|
|
338
|
+
that a confidently bounded impact places outside the change holds its own revision only: it is
|
|
339
|
+
never carried and blocks no reuse, so the next revision must observe that job again.
|
|
340
|
+
- Freshness is identity-based (digests, baseline, runtime tree), not time-based: the core forbids a
|
|
341
|
+
clock. Time-bound validity windows remain an AssertLedger-side concern.
|
|
342
|
+
- The satisfaction check (does evidence exist for each requirement, bound to the listed digests) is
|
|
343
|
+
not implemented; the plan is advisory until it is.
|
|
344
|
+
- Attribution bases are declared by whoever reports the signal. A pre-existing product defect, and
|
|
345
|
+
an infrastructure attribution that rests on `REPRODUCES_ON_BASELINE` alone, require
|
|
346
|
+
`FAILURE_ATTRIBUTION`, which AssertLedger should eventually verify; the other bases are trusted
|
|
347
|
+
as declared, within the limits above.
|
|
348
|
+
- The module is compiled into `dist/proof-planner/` but not reachable through the package exports;
|
|
349
|
+
its types and digests are not a public contract yet.
|
|
350
|
+
|
|
351
|
+
## 7. When should it become a separate tool?
|
|
352
|
+
|
|
353
|
+
Not now. Extract it only when several of these hold, with evidence:
|
|
354
|
+
|
|
355
|
+
1. **Used without AssertLedger**: at least one consumer plans assurance without producing or
|
|
356
|
+
verifying AssertLedger evidence (for example a CI router or a review bot).
|
|
357
|
+
2. **Own policy and configuration lifecycle**: projects version their assurance policy independently
|
|
358
|
+
of AssertLedger releases, with their own compatibility promises.
|
|
359
|
+
3. **Several consumers**: two or more independent harnesses or repositories call it, so its release
|
|
360
|
+
cadence conflicts with AssertLedger's.
|
|
361
|
+
4. **Autonomous data model**: `ChangeImpact`, `AssurancePlan` and the policy need published JSON
|
|
362
|
+
schemas and conformance fixtures of their own.
|
|
363
|
+
5. **Significant algorithms**: work beyond V1's fact derivation, such as learned calibration,
|
|
364
|
+
cross-revision planning or cost models, would bloat AssertLedger's evidence core.
|
|
365
|
+
6. **Stable API**: the input and output shapes survive a few real projects without breaking changes.
|
|
366
|
+
|
|
367
|
+
The module is built for that move: it imports only `zod`, its own files and `sha256Canonical` from
|
|
368
|
+
the public `./core` API, and a test enforces that boundary.
|
package/docs/reference.md
CHANGED
|
@@ -310,20 +310,21 @@ caller decide whether the capability exists. The server resolves repository root
|
|
|
310
310
|
confines them to the server process's current working directory by default. Programmatic operators
|
|
311
311
|
may supply a different `allowedRepositoryRoots` allowlist.
|
|
312
312
|
|
|
313
|
-
The doctor pair accepts a strict `{ "root": "..." }` input
|
|
314
|
-
`
|
|
313
|
+
The doctor pair accepts a strict `{ "root": "...", "exclude": ["..."] }` input, where the optional
|
|
314
|
+
`exclude` entry names follow the [repository initialization](repository-init.md) exclusion rules,
|
|
315
|
+
and returns the existing `repository-init-result` contract. It is read-only in both the default and operator-enabled server;
|
|
315
316
|
enabling unsafe execution does not change doctor behavior. Dynamic runtime and client diagnostics
|
|
316
317
|
remain outside this static readiness result.
|
|
317
318
|
|
|
318
|
-
`doctor_runtime` accepts the
|
|
319
|
-
[runtime diagnostic contract](runtime-doctor.md). `check` accepts the
|
|
319
|
+
`doctor_runtime` accepts only the strict root input and returns the separate
|
|
320
|
+
[runtime diagnostic contract](runtime-doctor.md), v2 for a generated Bun configuration. `check` accepts the
|
|
320
321
|
[high-level Git options](git-regression.md), without a permission field, and returns the existing
|
|
321
322
|
evidence manifest. The operator's capability is required for both tools.
|
|
322
323
|
|
|
323
324
|
## Continuous integration
|
|
324
325
|
|
|
325
326
|
Run `pnpm check` on every change. The included GitHub Actions workflow runs this gate on Node.js 22
|
|
326
|
-
and 24 on Ubuntu, Windows and macOS. A separate matrix installs and exercises the packed artifact
|
|
327
|
+
and 24 with Bun 1.4.2 on Ubuntu, Windows and macOS. A separate matrix installs and exercises the packed artifact
|
|
327
328
|
on Ubuntu and Windows with Node.js 22.15.0 and 24. On Ubuntu, the gate also runs the real-daemon
|
|
328
329
|
[container isolation](container-isolation.md) suite. A CI job that executes campaigns must
|
|
329
330
|
also treat `trusted-local` as `UNSANDBOXED`: use an isolated runner without secrets or host
|
|
@@ -334,7 +335,9 @@ credentials, and pass `--allow-unsafe-execution` only from reviewed CI configura
|
|
|
334
335
|
- `VERIFIED`: at least one candidate completed all required evidence and was selected.
|
|
335
336
|
- `REJECTED`: the campaign completed, but no candidate satisfied the policy.
|
|
336
337
|
- `INCONCLUSIVE`: controls or candidate evidence were incomplete, unstable, timed out, or affected
|
|
337
|
-
by infrastructure failure.
|
|
338
|
+
by infrastructure failure. Once its attempts are complete and agree, a completed candidate run
|
|
339
|
+
that reports no attributed candidate test still makes that candidate invalid, whatever else timed
|
|
340
|
+
out; see the [proof model](proof-model.md#candidate-gates).
|
|
338
341
|
- `ENGINE_ERROR`: the deterministic core could not normalize the supplied evidence safely.
|
|
339
342
|
|
|
340
343
|
Only an attributed `ASSERTION_FAILURE` can kill a target in protocol v1. Compilation errors,
|
package/docs/repository-init.md
CHANGED
|
@@ -22,13 +22,45 @@ plausible test frameworks, composite shell scripts, contradictory overrides, and
|
|
|
22
22
|
adapter configurations return `CONFLICT`. Multiple CI providers are only sorted evidence and do not
|
|
23
23
|
block initialization.
|
|
24
24
|
|
|
25
|
+
The static inventory walks the filesystem, not the Git index, and never follows or copies a
|
|
26
|
+
symbolic link: a link inside it returns `CONFLICT` with `UNSUPPORTED_REPOSITORY_SYMLINK` and writes
|
|
27
|
+
nothing. It always skips entries named `.git`, `.testforge`, and `node_modules`, at any depth.
|
|
28
|
+
When a local-only entry holds a link that no campaign needs, the operator can declare its name:
|
|
29
|
+
|
|
30
|
+
```sh
|
|
31
|
+
assertledger doctor . --exclude .claude --exclude .omx --json
|
|
32
|
+
assertledger init . --exclude .claude --exclude .omx --json
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Each `--exclude` value is one portable entry name without separators; every file or directory with
|
|
36
|
+
that name is skipped at any depth. Anything else returns `CONFLICT` with
|
|
37
|
+
`INVALID_REPOSITORY_EXCLUDE`. The declared names are written to `repository.exclude` beside the
|
|
38
|
+
defaults. `init` and `doctor` without `--exclude`, `analyze` and runtime doctor's configuration check
|
|
39
|
+
then reuse that list from a valid `assertledger.config.json`; an unreadable or invalid file leaves
|
|
40
|
+
only the defaults, which widens the inventory. A configured entry that contains a separator matches
|
|
41
|
+
nothing, as in a verification request. An explicit list, including an empty MCP `exclude` array,
|
|
42
|
+
replaces the configured one, so a different declaration never silently widens or narrows the
|
|
43
|
+
inventory: it fails closed, for example with `CONFIG_CONFLICT`, or with
|
|
44
|
+
`UNSUPPORTED_REPOSITORY_SYMLINK` when a narrower list exposes a link again. Links outside the declared names stay fail-closed.
|
|
45
|
+
|
|
46
|
+
The configured list governs only these static diagnostics and the evidence digests of
|
|
47
|
+
`assertledger.lock.json`; it produces no campaign evidence. `audit`, a campaign's repository copy and
|
|
48
|
+
its manifest `repositoryDigest` keep their own exclusions: the defaults, plus the verification
|
|
49
|
+
request's `repository.exclude` for a campaign. `audit` therefore still refuses a linked local-only
|
|
50
|
+
entry, and a request must declare the same names to leave it out of its copy; a committed
|
|
51
|
+
configuration can never remove files from campaign evidence. Declaring a name is an operator decision
|
|
52
|
+
recorded in a reviewable file, not a sandbox.
|
|
53
|
+
|
|
25
54
|
Files at or below the managed `candidateRoots` are deliberately excluded from framework inference,
|
|
26
55
|
evidence, built-in control tests, and repository-change comparison. Candidate generation therefore
|
|
27
56
|
cannot silently redefine initialization facts or invalidate an otherwise unchanged lock.
|
|
28
57
|
|
|
29
|
-
The built-in ready
|
|
30
|
-
|
|
31
|
-
|
|
58
|
+
The built-in ready adapters are `node-test` and `bun-test`. Bun initialization writes v2 config,
|
|
59
|
+
lock, and result contracts; Node initialization stays v1. Bun's static plan does not qualify the
|
|
60
|
+
installed runtime. Run runtime doctor and use a v3 verification request before claiming campaign
|
|
61
|
+
evidence. Pytest, Vitest, and Jest can be detected, but initialization returns `BLOCKED` with
|
|
62
|
+
`OFFICIAL_ADAPTER_UNAVAILABLE` unless the operator supplies an existing structured adapter
|
|
63
|
+
configuration:
|
|
32
64
|
|
|
33
65
|
```sh
|
|
34
66
|
assertledger init . --adapter-config integrations/my-adapter.json --json
|
|
@@ -37,8 +69,8 @@ assertledger init . --adapter-config integrations/my-adapter.json --json
|
|
|
37
69
|
The adapter document is parsed through the public adapter contract and is operator-owned. Its
|
|
38
70
|
executable is recorded as argv but is not resolved or executed by `init`. This is not an official
|
|
39
71
|
adapter endorsement and does not reduce the later `trusted-local` execution boundary.
|
|
40
|
-
`node-test` adapters are accepted only for
|
|
41
|
-
an operator-owned `testforge-command` adapter.
|
|
72
|
+
`node-test` adapters are accepted only for `node:test`; `bun-test` adapters are accepted only for
|
|
73
|
+
`bun:test`. Pytest, Vitest, and Jest require an operator-owned `testforge-command` adapter.
|
|
42
74
|
|
|
43
75
|
All detections and evidence digests come from one byte snapshot. Immediately before returning or
|
|
44
76
|
writing managed files, initialization rechecks the in-scope inventory and every evidence byte. A
|
|
@@ -54,7 +86,6 @@ Exit codes are `0` for `CREATED`, `UNCHANGED`, or `WOULD_CREATE`; `3` for `BLOCK
|
|
|
54
86
|
ambiguity, invalid overrides, conflicts, or contract validation; `5` for unexpected I/O; and `64`
|
|
55
87
|
for CLI usage errors.
|
|
56
88
|
|
|
57
|
-
The public
|
|
58
|
-
`repository-init-result.v1` contracts contain no timestamps, absolute repository roots, environment
|
|
89
|
+
The public v1 and v2 initialization contracts contain no timestamps, absolute repository roots, environment
|
|
59
90
|
values, worlds, or candidates. The lock binds normalized detections and sorted evidence digests; it
|
|
60
91
|
does not authenticate the detector, repository, adapter, or later execution evidence.
|
package/docs/runtime-doctor.md
CHANGED
|
@@ -23,20 +23,23 @@ configuration inspection or executable probes. Unknown flags, repeated flags, an
|
|
|
23
23
|
|
|
24
24
|
## What it checks
|
|
25
25
|
|
|
26
|
-
Version 1.0 supports the generated official `node:test` adapter.
|
|
27
|
-
|
|
26
|
+
Version 1.0 supports the generated official `node:test` adapter. Version 2.0 supports the
|
|
27
|
+
generated official `bun:test` adapter on the qualified Bun 1.4.2 revision. Both fail closed for
|
|
28
|
+
other frameworks and operator-supplied adapters. In order, they check:
|
|
28
29
|
|
|
29
30
|
1. explicit trusted-local authorization;
|
|
30
31
|
2. a current AssertLedger configuration and evidence lock;
|
|
31
|
-
3. the supported generated
|
|
32
|
-
4. a stable Node.js executable at version 22.15 or newer;
|
|
33
|
-
5. availability of
|
|
32
|
+
3. the supported generated adapter;
|
|
33
|
+
4. a stable Node.js executable at version 22.15 or newer, or the exact Bun 1.4.2 revision;
|
|
34
|
+
5. availability of the runner's built-in test module;
|
|
34
35
|
6. permission to create, write, and clean up an operating-system temporary workspace;
|
|
35
36
|
7. controlled reporter discovery, liveness, and attribution probes.
|
|
36
37
|
|
|
37
38
|
The final probes use disposable synthetic tests. One assertion failure must be attributed to the
|
|
38
39
|
candidate; a generic throw with a nested assertion cause must remain a process crash and must not
|
|
39
|
-
be attributed.
|
|
40
|
+
be attributed. The Bun probes also check its owned `assertSame` helper against a plain throw,
|
|
41
|
+
a caught assertion, an operand error, and a native `expect` failure. Malformed reporter output,
|
|
42
|
+
missing discovery, cleanup failure, timeout, or process
|
|
40
43
|
failure blocks readiness. AssertLedger does not write to the repository during runtime doctor.
|
|
41
44
|
|
|
42
45
|
## Result contract
|
|
@@ -44,15 +47,15 @@ failure blocks readiness. AssertLedger does not write to the repository during r
|
|
|
44
47
|
The SDK exposes the same strict additive contract:
|
|
45
48
|
|
|
46
49
|
```ts
|
|
47
|
-
import { AssertLedger,
|
|
50
|
+
import { AssertLedger, parseVersionedRuntimeDoctorResult } from "assertledger";
|
|
48
51
|
|
|
49
52
|
const result = await new AssertLedger().doctorRuntime(repositoryRoot, {
|
|
50
53
|
allowUnsafeExecution: true,
|
|
51
54
|
});
|
|
52
|
-
|
|
55
|
+
parseVersionedRuntimeDoctorResult(result);
|
|
53
56
|
```
|
|
54
57
|
|
|
55
|
-
`schemaVersion` is `1.0.0
|
|
58
|
+
`schemaVersion` is `1.0.0` for Node and `2.0.0` for Bun. Each check has `PASS`, `BLOCKED`, or `LIMITATION`, a stable reason code
|
|
56
59
|
when relevant, a concise summary, and a safe next action. The top-level status is `READY` only when
|
|
57
60
|
every supported runtime boundary passes. JSON results contain no subprocess output, environment
|
|
58
61
|
values, credentials, or repository source.
|