@sabaiway/agent-workflow-kit 2.1.0 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +45 -0
- package/README.md +6 -5
- package/SKILL.md +12 -8
- package/bin/install.mjs +1 -1
- package/bridges/antigravity-cli-bridge/SKILL.md +8 -1
- package/bridges/antigravity-cli-bridge/bin/agy-review-honesty.test.mjs +150 -0
- package/bridges/antigravity-cli-bridge/bin/agy-review-model-screen.test.mjs +76 -0
- package/bridges/antigravity-cli-bridge/bin/agy-review.sh +80 -15
- package/bridges/antigravity-cli-bridge/bin/agy-review.test.mjs +95 -4
- package/bridges/antigravity-cli-bridge/capability.json +3 -2
- package/bridges/codex-cli-bridge/SKILL.md +8 -1
- package/bridges/codex-cli-bridge/bin/codex-review-honesty.test.mjs +143 -0
- package/bridges/codex-cli-bridge/bin/codex-review.sh +71 -10
- package/bridges/codex-cli-bridge/bin/codex-review.test.mjs +93 -10
- package/bridges/codex-cli-bridge/capability.json +3 -2
- package/capability.json +1 -1
- package/migrations/3.0.0-hardened-core-loop.md +44 -0
- package/migrations/README.md +1 -1
- package/package.json +2 -2
- package/references/hooks/gate-approve.mjs +1 -1
- package/references/modes/bootstrap.md +3 -3
- package/references/modes/commit-guard.md +19 -0
- package/references/modes/core-evidence.md +15 -0
- package/references/modes/coverage-check.md +12 -0
- package/references/modes/doc-parity.md +7 -6
- package/references/modes/gates.md +7 -6
- package/references/modes/grounding.md +1 -2
- package/references/modes/hook.md +1 -1
- package/references/modes/review-state.md +3 -3
- package/references/modes/upgrade.md +6 -6
- package/references/modes/velocity.md +2 -2
- package/references/scripts/archive-decisions.mjs +1 -1
- package/references/scripts/install-git-hooks-repo-exec.test.mjs +82 -0
- package/references/scripts/install-git-hooks.mjs +90 -18
- package/references/scripts/install-git-hooks.test.mjs +102 -0
- package/references/scripts/migrate-gates-branches.test.mjs +157 -0
- package/references/scripts/migrate-gates.mjs +395 -0
- package/references/scripts/migrate-gates.test.mjs +284 -0
- package/references/templates/agent_rules.md +2 -2
- package/tools/ack-write.mjs +1 -1
- package/tools/atomic-write.mjs +2 -2
- package/tools/autonomy-config.mjs +1 -1
- package/tools/autonomy-doctor.mjs +1 -1
- package/tools/autonomy-write.mjs +1 -1
- package/tools/bridge-settings-read.mjs +1 -1
- package/tools/bridge-settings.mjs +1 -1
- package/tools/changed-surface.mjs +8 -69
- package/tools/cheap-agents.mjs +2 -2
- package/tools/commands.mjs +17 -10
- package/tools/commit-guard.mjs +167 -0
- package/tools/core-evidence.mjs +914 -0
- package/tools/coverage-check.mjs +260 -0
- package/tools/delegation.mjs +1 -1
- package/tools/detect-backends.mjs +3 -3
- package/tools/doc-parity.mjs +11 -27
- package/tools/engine-source.mjs +1 -1
- package/tools/family-members.mjs +1 -1
- package/tools/family-registry.mjs +1 -1
- package/tools/fs-safe.mjs +1 -1
- package/tools/gate-hook.mjs +1 -1
- package/tools/{seed-gates.mjs → gates-init.mjs} +58 -52
- package/tools/grounding.mjs +6 -83
- package/tools/hide-footprint.mjs +1 -1
- package/tools/inject-methodology.mjs +1 -1
- package/tools/known-footprint.mjs +1 -1
- package/tools/labels.mjs +1 -1
- package/tools/lcov.mjs +6 -10
- package/tools/lens-region.mjs +1 -1
- package/tools/manifest/validate.mjs +19 -0
- package/tools/migrate-adr-store.mjs +2 -2
- package/tools/orchestration-config.mjs +1 -1
- package/tools/orchestration-write.mjs +1 -1
- package/tools/presentation.mjs +1 -1
- package/tools/procedures.mjs +2 -2
- package/tools/recipes.mjs +60 -9
- package/tools/recommendations.mjs +63 -7
- package/tools/release-scan.mjs +1 -1
- package/tools/renderers.mjs +1 -1
- package/tools/review-state.mjs +288 -340
- package/tools/run-gates.mjs +216 -92
- package/tools/sandbox-masks.mjs +1 -1
- package/tools/semver-lite.mjs +1 -1
- package/tools/set-autonomy.mjs +1 -1
- package/tools/set-recipe.mjs +1 -1
- package/tools/setup-backends.mjs +1 -1
- package/tools/surface.mjs +1 -1
- package/tools/uninstall.mjs +1 -1
- package/tools/velocity-profile.mjs +2 -2
- package/tools/view-model.mjs +1 -1
- package/references/modes/fold-completeness.md +0 -30
- package/references/modes/review-ledger.md +0 -34
- package/references/templates/verification-profile.json +0 -10
- package/tools/fold-completeness-run.mjs +0 -1120
- package/tools/fold-completeness.mjs +0 -672
- package/tools/review-ledger-core.mjs +0 -428
- package/tools/review-ledger-write.mjs +0 -647
- package/tools/review-ledger.mjs +0 -630
- package/tools/sarif.mjs +0 -52
- package/tools/verification-profile.mjs +0 -219
package/tools/run-gates.mjs
CHANGED
|
@@ -18,38 +18,31 @@
|
|
|
18
18
|
// Honest outcomes — each distinct, never a silent green (the exit-code table + the summary-line
|
|
19
19
|
// schema are pinned by run-gates.test.mjs):
|
|
20
20
|
// 0 ok · 1 gate failure · 2 usage · 3 missing declaration · 4 empty gates list ·
|
|
21
|
-
// 5 malformed/invalid declaration · 6 bash unavailable ·
|
|
22
|
-
//
|
|
23
|
-
//
|
|
24
|
-
// under another shell.
|
|
21
|
+
// 5 malformed/invalid declaration · 6 bash unavailable · 8 --final receipt not written.
|
|
22
|
+
// Gate `cmd` lines are BASH command lines (brace/glob expansion); a host without bash gets the
|
|
23
|
+
// loud exit-6 preflight error, never a silent reinterpretation under another shell.
|
|
25
24
|
//
|
|
26
|
-
// The runner itself WRITES NOTHING
|
|
27
|
-
//
|
|
28
|
-
//
|
|
29
|
-
// runner never opens the ledger file itself (an import/structure pin holds the boundary). The
|
|
30
|
-
// record carries the FULL declaration + exactly what ran (a --only subset records honestly as a
|
|
31
|
-
// subset — it never satisfies quality-green) + the tree fingerprint BEFORE and AFTER the run (a
|
|
32
|
-
// mutating gate attests no particular tree). Dependency-free, Node >= 18. No side effects on import.
|
|
25
|
+
// The runner itself WRITES NOTHING on a plain run. `--final` (D3(a)) is the ONE writing mode:
|
|
26
|
+
// every attempt lands in the core-evidence store via its sole writer. Dependency-free.
|
|
27
|
+
// No side effects on import.
|
|
33
28
|
|
|
34
|
-
import { readFileSync, lstatSync } from 'node:fs';
|
|
35
|
-
import { join } from 'node:path';
|
|
29
|
+
import { readFileSync, lstatSync, unlinkSync, realpathSync } from 'node:fs';
|
|
30
|
+
import { join, isAbsolute } from 'node:path';
|
|
36
31
|
import { spawnSync } from 'node:child_process';
|
|
37
|
-
import { pathToFileURL } from 'node:url';
|
|
32
|
+
import { pathToFileURL, fileURLToPath } from 'node:url';
|
|
33
|
+
import { createHash, randomUUID } from 'node:crypto';
|
|
38
34
|
import { computeTreeFingerprint } from './review-state.mjs';
|
|
39
|
-
|
|
40
|
-
//
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
import { foldSuiteCredit } from './fold-completeness.mjs';
|
|
44
|
-
|
|
45
|
-
// The gate id the (a) credit applies to — the SAME command the fold runner resolves as its suite.
|
|
46
|
-
export const UNIT_TESTS_GATE_ID = 'unit-tests';
|
|
35
|
+
// The D3(a) final receipt rides the core-evidence SOLE WRITER (the sole-writer boundary — this
|
|
36
|
+
// runner never opens the store itself) + the canonical per-kind serialization its hashes bind.
|
|
37
|
+
import { appendEvidenceRecord, resolveEvidencePath, readEvidence, canonicalKindSerialization, EVIDENCE_SCHEMA_VERSION } from './core-evidence.mjs';
|
|
38
|
+
import { LCOV_BASENAME } from './coverage-check.mjs';
|
|
47
39
|
|
|
48
40
|
// The per-project declaration (strict JSON, hand-editable). cwd-relative — errors show a path the
|
|
49
41
|
// user can open (the orchestration-config CONFIG_REL idiom).
|
|
50
42
|
export const GATES_REL = 'docs/ai/gates.json';
|
|
51
43
|
|
|
52
44
|
// The full exit-code table — one distinct code per honest outcome (never a silent green).
|
|
45
|
+
// 7 is RETIRED (the deleted --record arm's outcome) — never reused for a new meaning.
|
|
53
46
|
export const EXIT = Object.freeze({
|
|
54
47
|
ok: 0,
|
|
55
48
|
fail: 1,
|
|
@@ -58,10 +51,9 @@ export const EXIT = Object.freeze({
|
|
|
58
51
|
empty: 4,
|
|
59
52
|
malformed: 5,
|
|
60
53
|
noBash: 6,
|
|
61
|
-
// --
|
|
62
|
-
//
|
|
63
|
-
|
|
64
|
-
recordFailed: 7,
|
|
54
|
+
// --final could not write its receipt (a corrupt store, an fs refusal): green gates WITHOUT a
|
|
55
|
+
// written receipt never read as success (D3(a)).
|
|
56
|
+
finalFailed: 8,
|
|
65
57
|
});
|
|
66
58
|
|
|
67
59
|
// A tagged failure carrying its process exit code (the shared orchestration-config idiom).
|
|
@@ -74,21 +66,23 @@ const SPAWN_FAILED_CODE = -1;
|
|
|
74
66
|
const MAX_GATE_OUTPUT_BYTES = 64 * 1024 * 1024;
|
|
75
67
|
|
|
76
68
|
const USAGE = [
|
|
77
|
-
'usage: run-gates.mjs [--cwd <dir>] [--only <id>]... [--
|
|
69
|
+
'usage: run-gates.mjs [--cwd <dir>] [--only <id>]... [--final] [--help]',
|
|
70
|
+
'',
|
|
71
|
+
'--final runs the FULL declared matrix as the D3(a) final verification run: it refuses a',
|
|
72
|
+
'declaration lacking the canonical core checks (the review-state + coverage-check gates),',
|
|
73
|
+
'deletes the stale git-dir lcov first, exports AW_GIT_DIR to every gate cmd, records EVERY',
|
|
74
|
+
'attempt (start + completed green/red) in the core-evidence store, and binds the receipt to',
|
|
75
|
+
'{ fingerprint before/after, the full declaration, per-gate results, the canonical red-proof +',
|
|
76
|
+
'degrade evidence hashes, the lcov sha }. --final refuses --only (a subset never attests).',
|
|
78
77
|
'',
|
|
79
78
|
`Runs the gates declared in <cwd>/${GATES_REL} (one bash command line each, project root as cwd).`,
|
|
80
79
|
'Prints a per-gate PASS/FAIL table + one machine-readable summary line; exit 0 iff all green.',
|
|
81
|
-
'--record additionally mints ONE v4 gate-run record into the review ledger via its sole writer',
|
|
82
|
-
'(the D5 green-baseline receipt; needs a single in-flight plan): the full declaration + what ran',
|
|
83
|
-
'+ the pre/post tree fingerprints. A red run records honestly; a --only subset records as a subset.',
|
|
84
|
-
'In --record mode the unit-tests gate is CREDITED (not re-spawned) when it is the FIRST declared gate',
|
|
85
|
-
'AND the fold-completeness runner already ran that EXACT command green at the current tree',
|
|
86
|
-
'(fingerprint-bound + exit-0 + cmd-identity); positioned after another gate, it always re-spawns.',
|
|
87
80
|
'Sandbox-safe: the runner itself needs no network and writes only repo-local state — the D4 sandbox',
|
|
88
81
|
'lane; each DECLARED gate command is the project\'s own, so ITS sandbox-safety is command-shape',
|
|
89
82
|
'dependent (first try the sandbox-safe shape — cache under $TMPDIR, offline/notifier off).',
|
|
90
83
|
`Exit codes: 0 ok · 1 gate failure · 2 usage · 3 missing declaration · 4 empty gates list ·`,
|
|
91
|
-
'5 malformed/invalid declaration · 6 bash unavailable ·
|
|
84
|
+
'5 malformed/invalid declaration · 6 bash unavailable ·',
|
|
85
|
+
'8 --final asked but its receipt could not be written (green gates never read as success without it).',
|
|
92
86
|
].join('\n');
|
|
93
87
|
|
|
94
88
|
// ── declaration validation (malformed → exit 5, loud `path: reason`) ─────────────────
|
|
@@ -187,13 +181,11 @@ export const selectGates = (gates, onlyIds) => {
|
|
|
187
181
|
// Spawn one gate cmd via bash from the project root. `cmd` is a BASH command line by contract
|
|
188
182
|
// (the declaration's _README states it): this repo's own gate matrix needs brace+glob expansion,
|
|
189
183
|
// which /bin/sh does not perform — hence bash explicitly, never the platform default shell.
|
|
190
|
-
// NODE_TEST_CONTEXT is stripped
|
|
191
|
-
//
|
|
192
|
-
//
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
export const spawnGateViaBash = (cmd, cwd) => {
|
|
196
|
-
const env = { ...process.env };
|
|
184
|
+
// NODE_TEST_CONTEXT is stripped: a `node --test` gate spawned while run-gates is itself running
|
|
185
|
+
// under a parent test context would otherwise inherit it, hit Node's recursive-run guard, silently
|
|
186
|
+
// skip every file, and exit 0 — a vacuous false green.
|
|
187
|
+
export const spawnGateViaBash = (cmd, cwd, extraEnv = {}) => {
|
|
188
|
+
const env = { ...process.env, ...extraEnv };
|
|
197
189
|
delete env.NODE_TEST_CONTEXT;
|
|
198
190
|
return spawnSync('bash', ['-c', cmd], { cwd, env, encoding: 'utf8', maxBuffer: MAX_GATE_OUTPUT_BYTES });
|
|
199
191
|
};
|
|
@@ -209,26 +201,17 @@ const trimTrailingNewline = (text) => text.replace(/\n$/, '');
|
|
|
209
201
|
|
|
210
202
|
// Run the selected gates sequentially (declaration order). A green gate logs one PASS line; a
|
|
211
203
|
// failing gate logs FAIL + its captured stdout/stderr VERBATIM (triage without re-running).
|
|
212
|
-
export const runGates = (gates, { cwd, spawn = spawnGateViaBash, now = Date.now, log = console.log
|
|
204
|
+
export const runGates = (gates, { cwd, spawn = spawnGateViaBash, now = Date.now, log = console.log }) => {
|
|
213
205
|
const results = [];
|
|
214
206
|
for (const gate of gates) {
|
|
215
207
|
log(`── ${gate.id} — ${gate.title}`);
|
|
216
|
-
// (a) the one-suite-run credit: for the unit-tests gate, if the fold runner already ran this exact
|
|
217
|
-
// command green at the current tree, CREDIT it instead of re-spawning (no quality loss — same
|
|
218
|
-
// execution, same tree, recorded once). Any mismatch → credit is null → the normal spawn below.
|
|
219
|
-
const credited = credit ? credit(gate) : null;
|
|
220
|
-
if (credited) {
|
|
221
|
-
log(' PASS (credited from the fold-completeness suite run — no re-spawn)');
|
|
222
|
-
results.push({ id: gate.id, title: gate.title, ok: true, code: 0, elapsedMs: 0, credited: true });
|
|
223
|
-
continue;
|
|
224
|
-
}
|
|
225
208
|
const startedAt = now();
|
|
226
209
|
const res = spawn(gate.cmd, cwd);
|
|
227
210
|
const elapsedMs = now() - startedAt;
|
|
228
211
|
const spawnError = res.error ? `spawn error: ${res.error.code || res.error.message}` : null;
|
|
229
212
|
const ok = spawnError === null && res.status === 0;
|
|
230
213
|
const code = spawnError === null ? res.status : SPAWN_FAILED_CODE;
|
|
231
|
-
results.push({ id: gate.id, title: gate.title, ok, code, elapsedMs });
|
|
214
|
+
results.push({ id: gate.id, title: gate.title, ok, code, elapsedMs, stdout: res.stdout ? String(res.stdout) : '' });
|
|
232
215
|
if (ok) {
|
|
233
216
|
log(` PASS (${formatSeconds(elapsedMs)})`);
|
|
234
217
|
} else {
|
|
@@ -264,7 +247,7 @@ export const composeSummaryLine = ({ status, results = [] }) => {
|
|
|
264
247
|
// ── CLI ───────────────────────────────────────────────────────────────────────────────
|
|
265
248
|
|
|
266
249
|
const parseArgs = (argv) => {
|
|
267
|
-
const opts = { cwd: null, only: [],
|
|
250
|
+
const opts = { cwd: null, only: [], final: false, help: false };
|
|
268
251
|
for (let i = 0; i < argv.length; i += 1) {
|
|
269
252
|
const arg = argv[i];
|
|
270
253
|
if (arg === '--help' || arg === '-h') {
|
|
@@ -277,15 +260,53 @@ const parseArgs = (argv) => {
|
|
|
277
260
|
i += 1;
|
|
278
261
|
if (argv[i] === undefined) throw fail(EXIT.usage, '--only requires a gate id argument');
|
|
279
262
|
opts.only.push(argv[i]);
|
|
280
|
-
} else if (arg === '--
|
|
281
|
-
opts.
|
|
263
|
+
} else if (arg === '--final') {
|
|
264
|
+
opts.final = true;
|
|
282
265
|
} else {
|
|
283
266
|
throw fail(EXIT.usage, `unknown argument "${arg}"\n${USAGE}`);
|
|
284
267
|
}
|
|
285
268
|
}
|
|
269
|
+
if (opts.final && opts.only.length > 0) {
|
|
270
|
+
throw fail(EXIT.usage, '--final refuses --only — a subset never attests (the D3(a) receipt binds the FULL declaration)');
|
|
271
|
+
}
|
|
286
272
|
return opts;
|
|
287
273
|
};
|
|
288
274
|
|
|
275
|
+
// The canonical core checks a --final declaration must carry (D3(a)), matched as STRICT FULL
|
|
276
|
+
// commands: `node` + ONE (quoted or bare) path token + the exact tool basename + ` --check` +
|
|
277
|
+
// END — and the path token must REALPATH-RESOLVE to the kit's OWN tool (the canonical sibling of
|
|
278
|
+
// this runner). Masked forms (`--check --help`, `--check || true`, prefix commands) never match
|
|
279
|
+
// the shape; a lookalike file that merely carries the basename — whatever it prints — never
|
|
280
|
+
// resolves to the canonical tool. Any form that DOES resolve (bare, relative, absolute, quoted)
|
|
281
|
+
// is accepted, so the anchor adds no false refusals.
|
|
282
|
+
const coreCheckRe = (basename) => new RegExp(`^node\\s+(?:"((?:[^"]*[/\\\\])?${basename})"|((?:[^\\s"]*[/\\\\])?${basename}))\\s+--check$`);
|
|
283
|
+
const FINAL_CORE_CHECKS = [
|
|
284
|
+
{ name: 'review-state', re: coreCheckRe('review-state\\.mjs'), canonical: fileURLToPath(new URL('./review-state.mjs', import.meta.url)) },
|
|
285
|
+
{ name: 'coverage-check', re: coreCheckRe('coverage-check\\.mjs'), canonical: fileURLToPath(new URL('./coverage-check.mjs', import.meta.url)) },
|
|
286
|
+
];
|
|
287
|
+
const matchesCanonicalCheck = (check, cmd, projectDir) => {
|
|
288
|
+
const m = check.re.exec(cmd.trim());
|
|
289
|
+
if (!m) return false;
|
|
290
|
+
const token = m[1] ?? m[2];
|
|
291
|
+
const abs = isAbsolute(token) ? token : join(projectDir, token);
|
|
292
|
+
try {
|
|
293
|
+
return realpathSync(abs) === realpathSync(check.canonical);
|
|
294
|
+
} catch {
|
|
295
|
+
return false; // unresolvable → never canonical (fail closed)
|
|
296
|
+
}
|
|
297
|
+
};
|
|
298
|
+
|
|
299
|
+
// isFinalCapableDeclaration(gates, projectDir) → whether --final would accept this declaration
|
|
300
|
+
// (every canonical core check present + the checker LAST) — the ONE home consumers (the
|
|
301
|
+
// recommendations guard-install probe) read instead of re-deriving the rule.
|
|
302
|
+
export const isFinalCapableDeclaration = (gates, projectDir) => {
|
|
303
|
+
if (!Array.isArray(gates) || gates.length === 0) return false;
|
|
304
|
+
const missing = FINAL_CORE_CHECKS.filter((c) => !gates.some((g) => matchesCanonicalCheck(c, g.cmd, projectDir)));
|
|
305
|
+
if (missing.length > 0) return false;
|
|
306
|
+
return matchesCanonicalCheck(FINAL_CORE_CHECKS[1], gates[gates.length - 1].cmd, projectDir);
|
|
307
|
+
};
|
|
308
|
+
const sha256Hex = (data) => createHash('sha256').update(data).digest('hex');
|
|
309
|
+
|
|
289
310
|
// The full CLI, dependency-injected for hermetic tests. Returns the process exit code; the two
|
|
290
311
|
// output sinks split human-facing report (log) from error channel (logError). The summary line is
|
|
291
312
|
// emitted via `log` as the final line of every non-usage outcome.
|
|
@@ -299,9 +320,7 @@ export const runCli = (argv, deps = {}) => {
|
|
|
299
320
|
readFile,
|
|
300
321
|
lstat,
|
|
301
322
|
now,
|
|
302
|
-
record = recordGateRun,
|
|
303
323
|
fingerprint = computeTreeFingerprint,
|
|
304
|
-
foldCredit = foldSuiteCredit,
|
|
305
324
|
} = deps;
|
|
306
325
|
try {
|
|
307
326
|
const opts = parseArgs(argv);
|
|
@@ -326,6 +345,39 @@ export const runCli = (argv, deps = {}) => {
|
|
|
326
345
|
return EXIT.empty;
|
|
327
346
|
}
|
|
328
347
|
const selected = selectGates(declaration.gates, opts.only);
|
|
348
|
+
// AW_GIT_DIR rides EVERY gate child inside a git tree (plain and --only alike): declared
|
|
349
|
+
// cmds reference fixed git-dir artifacts (the unit-tests lcov destination) — a plain red-run
|
|
350
|
+
// must exercise the SAME cmd line --final will, never a broken-only-outside---final variant.
|
|
351
|
+
const dirRes = spawnSync('git', ['rev-parse', '--absolute-git-dir'], { cwd: projectDir, windowsHide: true });
|
|
352
|
+
const gitDir = dirRes.error || dirRes.status !== 0 ? null : dirRes.stdout.toString('utf8').replace(/\r?\n$/, '');
|
|
353
|
+
let gateSpawn = gitDir === null ? spawn : (cmd, cwd2) => spawn(cmd, cwd2, { AW_GIT_DIR: gitDir });
|
|
354
|
+
// ── --final preflight (D3(a)): the declaration must carry the canonical core checks; the TRUE
|
|
355
|
+
// git dir resolves AW_GIT_DIR; the stale lcov dies BEFORE the suite so it is never consumed.
|
|
356
|
+
if (opts.final) {
|
|
357
|
+
const missing = FINAL_CORE_CHECKS.filter((c) => !declaration.gates.some((g) => matchesCanonicalCheck(c, g.cmd, projectDir)));
|
|
358
|
+
if (missing.length > 0) {
|
|
359
|
+
throw fail(EXIT.malformed, `--final refuses a weakened declaration — missing the canonical core check(s): ${missing.map((c) => c.name).join(', ')} (each must be ONE plain --check invocation of the kit's OWN tool in ${GATES_REL} — a masked form, a compound, or a lookalike path never counts)`);
|
|
360
|
+
}
|
|
361
|
+
const lastGate = declaration.gates[declaration.gates.length - 1];
|
|
362
|
+
if (!matchesCanonicalCheck(FINAL_CORE_CHECKS[1], lastGate.cmd, projectDir)) {
|
|
363
|
+
throw fail(EXIT.malformed, `--final refuses the declaration — the CANONICAL coverage-check gate must be the LAST declared gate (nothing may run after the checker consumed the lcov; "${lastGate.id}" is declared last)`);
|
|
364
|
+
}
|
|
365
|
+
if (gitDir === null) throw fail(EXIT.fail, '--final needs a git work tree (cannot resolve the git dir)');
|
|
366
|
+
try {
|
|
367
|
+
unlinkSync(join(gitDir, LCOV_BASENAME)); // a stale lcov is never consumed
|
|
368
|
+
} catch (err) {
|
|
369
|
+
if (err.code !== 'ENOENT') {
|
|
370
|
+
// Anything but "absent" leaves a readable stale artifact behind — fail closed BEFORE
|
|
371
|
+
// the attempt starts (a swallowed delete error could fake coverage).
|
|
372
|
+
logError(`[run-gates] --final could not delete the stale lcov at ${join(gitDir, LCOV_BASENAME)}: ${err.message} (fail closed)`);
|
|
373
|
+
log(composeSummaryLine({ status: 'fail' }));
|
|
374
|
+
return EXIT.finalFailed;
|
|
375
|
+
}
|
|
376
|
+
}
|
|
377
|
+
// ONE lcov path for delete/check/hash/guard: the fixed git-dir file is FORCED into every
|
|
378
|
+
// gate child env — a stray host AW_LCOV_FILE can never desync the checker from the receipt.
|
|
379
|
+
gateSpawn = (cmd, cwd2) => spawn(cmd, cwd2, { AW_GIT_DIR: gitDir, AW_LCOV_FILE: join(gitDir, LCOV_BASENAME) });
|
|
380
|
+
}
|
|
329
381
|
const probe = spawn(BASH_PROBE_CMD, projectDir);
|
|
330
382
|
if (probe.error && probe.error.code === 'ENOENT') {
|
|
331
383
|
logError(
|
|
@@ -335,51 +387,123 @@ export const runCli = (argv, deps = {}) => {
|
|
|
335
387
|
log(composeSummaryLine({ status: 'no-bash' }));
|
|
336
388
|
return EXIT.noBash;
|
|
337
389
|
}
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
390
|
+
// --final needs the pre-run fingerprint (the receipt binds before == after == current).
|
|
391
|
+
const finalFingerprintBefore = opts.final ? fingerprint(projectDir) : null;
|
|
392
|
+
const finalAttempt = opts.final ? randomUUID() : null;
|
|
393
|
+
let finalError = null;
|
|
394
|
+
let startEvidenceHashes = null;
|
|
395
|
+
if (opts.final) {
|
|
396
|
+
try {
|
|
397
|
+
appendEvidenceRecord({
|
|
398
|
+
path: resolveEvidencePath(projectDir, env),
|
|
399
|
+
record: { schema: EVIDENCE_SCHEMA_VERSION, kind: 'final-start', fingerprint: finalFingerprintBefore, attempt: finalAttempt, timestamp: new Date().toISOString() },
|
|
400
|
+
});
|
|
401
|
+
// The drift tooth's anchor: the canonical red-proof/degrade serializations BEFORE any
|
|
402
|
+
// gate runs — no legitimate writer appends those kinds DURING a final run, so any change
|
|
403
|
+
// by receipt time is an integrity failure, never a green.
|
|
404
|
+
const { records } = readEvidence(resolveEvidencePath(projectDir, env));
|
|
405
|
+
startEvidenceHashes = {
|
|
406
|
+
redProof: sha256Hex(canonicalKindSerialization(records, 'red-proof')),
|
|
407
|
+
degrade: sha256Hex(canonicalKindSerialization(records, 'degrade')),
|
|
408
|
+
};
|
|
409
|
+
} catch (err) {
|
|
410
|
+
// EVERY attempt is recorded — an unwritable store refuses the whole attempt up front
|
|
411
|
+
// (green gates without a written receipt never read as success).
|
|
412
|
+
logError(`[run-gates] --final could not record the attempt start: ${err.message}`);
|
|
413
|
+
log(composeSummaryLine({ status: 'fail' }));
|
|
414
|
+
return EXIT.finalFailed;
|
|
415
|
+
}
|
|
349
416
|
}
|
|
350
|
-
const results = runGates(selected, { cwd: projectDir, spawn, log, now
|
|
417
|
+
const results = runGates(selected, { cwd: projectDir, spawn: gateSpawn, log, now });
|
|
351
418
|
for (const line of formatTable(results)) log(line);
|
|
352
419
|
const allGreen = results.every((result) => result.ok);
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
357
|
-
|
|
358
|
-
|
|
420
|
+
if (opts.final) {
|
|
421
|
+
// The checker's verbatim diagnostics surface even on green — skipped-no-lcov and the
|
|
422
|
+
// out-of-domain/unsupported lists must never vanish into a suppressed green stdout.
|
|
423
|
+
const checkerRow = results[results.length - 1];
|
|
424
|
+
if (checkerRow?.ok && checkerRow.stdout) {
|
|
425
|
+
log('── coverage-check diagnostics (verbatim)');
|
|
426
|
+
log(trimTrailingNewline(checkerRow.stdout));
|
|
427
|
+
}
|
|
428
|
+
// The completed attempt (D3(a)): status green/red DERIVED from results + integrity, the
|
|
429
|
+
// FULL declaration, the per-gate results, the evidence hashes over the CANONICAL
|
|
430
|
+
// authoritative serializations, and the lcov sha the CHECKER printed for the bytes it
|
|
431
|
+
// actually read — with an end re-hash agreement check (the checker's own children are the
|
|
432
|
+
// one write window that survives "coverage-check runs last").
|
|
359
433
|
try {
|
|
360
|
-
const
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
434
|
+
const storePath = resolveEvidencePath(projectDir, env);
|
|
435
|
+
const endRead = readEvidence(storePath);
|
|
436
|
+
const endBroken = Boolean(endRead.readError) || (endRead.malformed ?? 0) > 0;
|
|
437
|
+
let integrityFailure = null;
|
|
438
|
+
if (endBroken) {
|
|
439
|
+
integrityFailure = 'the evidence store became unreadable under the final run';
|
|
440
|
+
} else {
|
|
441
|
+
const endHashes = {
|
|
442
|
+
redProof: sha256Hex(canonicalKindSerialization(endRead.records, 'red-proof')),
|
|
443
|
+
degrade: sha256Hex(canonicalKindSerialization(endRead.records, 'degrade')),
|
|
444
|
+
};
|
|
445
|
+
if (endHashes.redProof !== startEvidenceHashes.redProof || endHashes.degrade !== startEvidenceHashes.degrade) {
|
|
446
|
+
integrityFailure = 'the evidence store moved under the final run (the canonical red-proof/degrade serialization changed)';
|
|
447
|
+
}
|
|
448
|
+
}
|
|
449
|
+
// Exactly ONE full machine line binds the receipt — an unanchored first-match would let
|
|
450
|
+
// an injected/duplicated line shadow the real one and skip the end re-hash.
|
|
451
|
+
const shaLineRe = /^coverage-check: lcov-sha256=([0-9a-f]{64}|none)$/;
|
|
452
|
+
const shaLines = String(checkerRow?.stdout ?? '').split(/\r?\n/).filter((l) => shaLineRe.test(l));
|
|
453
|
+
const shaValue = shaLines.length === 1 ? shaLineRe.exec(shaLines[0])[1] : null;
|
|
454
|
+
const lcovSha256 = shaValue !== null && shaValue !== 'none' ? shaValue : null;
|
|
455
|
+
if (allGreen && integrityFailure === null) {
|
|
456
|
+
if (shaLines.length !== 1) {
|
|
457
|
+
integrityFailure = shaLines.length === 0
|
|
458
|
+
? 'the coverage-check gate printed no lcov-sha256 line — the consumed lcov is unknowable (fail closed)'
|
|
459
|
+
: `the coverage-check gate printed ${shaLines.length} lcov-sha256 lines — exactly ONE full machine line binds the receipt`;
|
|
460
|
+
} else if (lcovSha256 !== null) {
|
|
461
|
+
let endSha = null;
|
|
462
|
+
try {
|
|
463
|
+
const st = lstatSync(join(gitDir, LCOV_BASENAME));
|
|
464
|
+
if (st.isFile()) endSha = sha256Hex(readFileSync(join(gitDir, LCOV_BASENAME)));
|
|
465
|
+
} catch { /* vanished → disagreement below */ }
|
|
466
|
+
if (endSha !== lcovSha256) integrityFailure = 'the lcov moved under the checker (the end re-hash differs from the checker-read bytes)';
|
|
467
|
+
}
|
|
468
|
+
}
|
|
469
|
+
const status = allGreen && integrityFailure === null ? 'green' : 'red';
|
|
470
|
+
appendEvidenceRecord({
|
|
471
|
+
path: storePath,
|
|
472
|
+
record: {
|
|
473
|
+
schema: EVIDENCE_SCHEMA_VERSION,
|
|
474
|
+
kind: 'final',
|
|
475
|
+
status,
|
|
476
|
+
attempt: finalAttempt,
|
|
477
|
+
fingerprintBefore: finalFingerprintBefore,
|
|
478
|
+
fingerprintAfter: fingerprint(projectDir),
|
|
479
|
+
declared: declaration.gates.map(({ id, cmd }) => ({ id, cmd })),
|
|
480
|
+
results: results.map(({ id, ok, code }) => ({ id, ok, code })),
|
|
481
|
+
evidenceHashes: endBroken
|
|
482
|
+
? startEvidenceHashes
|
|
483
|
+
: {
|
|
484
|
+
redProof: sha256Hex(canonicalKindSerialization(endRead.records, 'red-proof')),
|
|
485
|
+
degrade: sha256Hex(canonicalKindSerialization(endRead.records, 'degrade')),
|
|
486
|
+
},
|
|
487
|
+
lcovSha256,
|
|
488
|
+
integrityFailure,
|
|
489
|
+
timestamp: new Date().toISOString(),
|
|
371
490
|
},
|
|
372
|
-
fingerprintBefore,
|
|
373
|
-
fingerprintAfter: fingerprint(projectDir),
|
|
374
491
|
});
|
|
375
|
-
|
|
492
|
+
if (integrityFailure) {
|
|
493
|
+
finalError = new Error(integrityFailure);
|
|
494
|
+
logError(`[run-gates] --final integrity failure: ${integrityFailure} — the receipt is RED`);
|
|
495
|
+
}
|
|
496
|
+
log(`[run-gates] final receipt recorded (${status}) → ${storePath}`);
|
|
497
|
+
if (status === 'green' && lcovSha256 === null) {
|
|
498
|
+
logError(`[run-gates] --final consumed NO lcov — the coverage arm ran skipped-no-lcov; the receipt records lcovSha256:null (the declared unit-tests gate must produce <git-dir>/${LCOV_BASENAME})`);
|
|
499
|
+
}
|
|
376
500
|
} catch (err) {
|
|
377
|
-
|
|
378
|
-
logError(`[run-gates] --
|
|
501
|
+
finalError = err;
|
|
502
|
+
logError(`[run-gates] --final could not write its receipt: ${err.message}`);
|
|
379
503
|
}
|
|
380
504
|
}
|
|
381
505
|
log(composeSummaryLine({ status: allGreen ? 'ok' : 'fail', results }));
|
|
382
|
-
if (
|
|
506
|
+
if (finalError) return EXIT.finalFailed;
|
|
383
507
|
return allGreen ? EXIT.ok : EXIT.fail;
|
|
384
508
|
} catch (err) {
|
|
385
509
|
logError(`[run-gates] ${err.message}`);
|
package/tools/sandbox-masks.mjs
CHANGED
|
@@ -31,7 +31,7 @@
|
|
|
31
31
|
// from the hidden-mode reconcile's block (hide-footprint.mjs — that lane noops on visible
|
|
32
32
|
// deployments; this one serves them too). A malformed fence (start without end / duplicated
|
|
33
33
|
// markers) fails CLOSED: loud report, file untouched. Patterns are root-anchored (`/rel`),
|
|
34
|
-
// glob-metacharacters escaped, NUL-safe walk. Dependency-free, Node >=
|
|
34
|
+
// glob-metacharacters escaped, NUL-safe walk. Dependency-free, Node >= 22; no side effects on
|
|
35
35
|
// import (the isDirectRun idiom).
|
|
36
36
|
|
|
37
37
|
import { lstatSync, readFileSync, writeFileSync, mkdirSync } from 'node:fs';
|
package/tools/semver-lite.mjs
CHANGED
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
// freshness probe that cannot parse a version must degrade to "unknown" — never to a false ordering
|
|
8
8
|
// claim in either direction (INV-B). No `let`: a small functional comparison (AGENTS.md §2.3).
|
|
9
9
|
//
|
|
10
|
-
// Pure, no imports, no side effects, Node >=
|
|
10
|
+
// Pure, no imports, no side effects, Node >= 22.
|
|
11
11
|
|
|
12
12
|
export const parseSemver = (str) => {
|
|
13
13
|
const m = typeof str === 'string' ? str.trim().match(/^(\d+)\.(\d+)\.(\d+)/) : null;
|
package/tools/set-autonomy.mjs
CHANGED
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
// (malformed/unreadable policy) or a write STOP (no deployment / symlinked leaf). main(argv, ctx) →
|
|
20
20
|
// { code, stdout, stderr }; cwd / fs are injectable for host-independent tests.
|
|
21
21
|
//
|
|
22
|
-
// Dependency-free, Node >=
|
|
22
|
+
// Dependency-free, Node >= 22. No side effects on import (the isDirectRun idiom).
|
|
23
23
|
|
|
24
24
|
import { readFileSync, lstatSync } from 'node:fs';
|
|
25
25
|
import { pathToFileURL } from 'node:url';
|
package/tools/set-recipe.mjs
CHANGED
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
// or a write STOP (no deployment / symlinked leaf). main(argv, ctx) → { code, stdout, stderr }; cwd /
|
|
20
20
|
// env / home / detect / fs are injectable for host-independent tests.
|
|
21
21
|
//
|
|
22
|
-
// Dependency-free, Node >=
|
|
22
|
+
// Dependency-free, Node >= 22. No side effects on import (the isDirectRun idiom).
|
|
23
23
|
|
|
24
24
|
import { readFileSync, lstatSync } from 'node:fs';
|
|
25
25
|
import { homedir } from 'node:os';
|
package/tools/setup-backends.mjs
CHANGED
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
// wrappers are POSIX .sh — report unsupported / use WSL and mutate nothing. Every fs primitive is
|
|
22
22
|
// injectable (deps.*) so the whole module is unit-testable without touching the real filesystem.
|
|
23
23
|
//
|
|
24
|
-
// Dependency-free, Node >=
|
|
24
|
+
// Dependency-free, Node >= 22.
|
|
25
25
|
|
|
26
26
|
import {
|
|
27
27
|
existsSync, lstatSync, statSync, readdirSync, readlinkSync, symlinkSync, mkdirSync, copyFileSync,
|
package/tools/surface.mjs
CHANGED
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
// `/agent-workflow-kit status` surface always consumes `--json` and localizes itself.
|
|
7
7
|
//
|
|
8
8
|
// Pure + fully injectable (argv / env / isTTY / columns / platform are inputs), no side effects on
|
|
9
|
-
// import, Node >=
|
|
9
|
+
// import, Node >= 22. Tested in isolation against the §4.5 table.
|
|
10
10
|
|
|
11
11
|
export const FORMATS = Object.freeze(['auto', 'plain', 'ansi', 'json']);
|
|
12
12
|
export const FORMAT_ENV = 'AGENT_WORKFLOW_FORMAT';
|
package/tools/uninstall.mjs
CHANGED
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
// untouched and reported, never removed.
|
|
20
20
|
//
|
|
21
21
|
// Pure planner (buildPlan) + guarded executor (executePlan), both dependency-injectable so the whole
|
|
22
|
-
// module is unit-testable without the real filesystem. Dependency-free, Node >=
|
|
22
|
+
// module is unit-testable without the real filesystem. Dependency-free, Node >= 22. No side effects on
|
|
23
23
|
// import (the isDirectRun idiom).
|
|
24
24
|
|
|
25
25
|
import { existsSync, statSync, lstatSync, readlinkSync, readFileSync, readdirSync, realpathSync } from 'node:fs';
|
|
@@ -18,7 +18,7 @@ import { findOnPath } from './detect-backends.mjs';
|
|
|
18
18
|
|
|
19
19
|
// Deployment-lineage head this velocity build targets; bump together with agent-workflow-memory
|
|
20
20
|
// LINEAGE_HEAD when the deployed docs/ai structure changes.
|
|
21
|
-
export const EXPECTED_WORKFLOW_VERSION = '
|
|
21
|
+
export const EXPECTED_WORKFLOW_VERSION = '3.0.0';
|
|
22
22
|
export const SETTINGS_FILE = '.claude/settings.json';
|
|
23
23
|
export const SETTINGS_LOCAL_FILE = '.claude/settings.local.json';
|
|
24
24
|
export const CLAUDE_DIR = '.claude';
|
|
@@ -490,7 +490,7 @@ const isScriptMap = (scripts) => Boolean(scripts) && typeof scripts === 'object'
|
|
|
490
490
|
const isMutatingScriptName = (name) =>
|
|
491
491
|
MUTATING_SCRIPT_NAME_PATTERN.test(name) || MUTATING_SCRIPT_HOOK_PATTERN.test(name);
|
|
492
492
|
|
|
493
|
-
// `scriptName` is ADDITIVE (AD-042): the
|
|
493
|
+
// `scriptName` is ADDITIVE (AD-042): the gates-init offer layer maps a candidate to a
|
|
494
494
|
// package-manager-aware `{ id, title, cmd }` and needs the raw script name for that derivation;
|
|
495
495
|
// `command` stays the advisory's own npm-run spelling (this fn is otherwise unchanged).
|
|
496
496
|
const makeGateCandidate = (name) => ({
|
package/tools/view-model.mjs
CHANGED
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
// filenames never reach them). This module resolves public tokens → English phrases (presentation.mjs)
|
|
7
7
|
// and computes the headline counts; the renderers do layout + glyphs only.
|
|
8
8
|
//
|
|
9
|
-
// Pure, no side effects, Node >=
|
|
9
|
+
// Pure, no side effects, Node >= 22.
|
|
10
10
|
|
|
11
11
|
import { STATE_PHRASING, VISIBILITY_PHRASING } from './presentation.mjs';
|
|
12
12
|
|
|
@@ -1,30 +0,0 @@
|
|
|
1
|
-
### Mode: fold-completeness
|
|
2
|
-
|
|
3
|
-
The FOLD-COMPLETENESS gate (M3a / AD-046, honest red→green BUGFREE-1 / AD-047) — the sibling of the review-round ledger: the ledger computes WHEN the plan-execution review loop stops; this tool pair verifies that the tests the loop's folds bind to ACTUALLY pin the changed code, and that each bound test was SEEN FAILING before its fix (no fix theater). One recorded run proves (a) every changed executable line is executed by the suite and (b) every `fixable-bug` bound testId is N/N green with an observed-red receipt behind it; a read-only checker turns the recorded evidence into a gate exit code. **Since schema v3 (BUGFREE-2 / AD-048, D7) every record carries `base` and the whole proof is SEGMENT-scoped** — bound testIds, receipts, custody chains, and tamper obligations all filter to (loop, `base` = the current HEAD): a committed phase's custody obligations **close with its commit** (editing a custody file in a LATER segment no longer re-demands the closed segment's proof — the cross-phase churn class is dead), while a fold whose red was observed in a PREVIOUS segment still needs the `red-proof` lane (a receipt never crosses a commit boundary — the stated residual). Pre-v3 records never enter a segment; a dirty loop holding only pre-v3 records fails with a reason naming the schema upgrade. **Since schema v4 (BUGFREE-3 / AD-049)** a run record additionally carries the **suite-execution evidence** ((a): `cmd` + exit + pre/post fingerprints — so `run-gates --record` can CREDIT the `unit-tests` gate instead of re-spawning it), and a new record kind **`reattest`** joins `run`/`red-probe` ((c): an operator-minted, recorded custody re-attestation — see the runner's `--reattest`). A placed pre-v4 checker rejects a v4 record outright (the version-skew guard); v1/v2/v3 rules are untouched.
|
|
4
|
-
|
|
5
|
-
**The fold-time order (the operational contract — red is observed BEFORE the fix):** classify the fixable-bug with its testId → write the failing test → `--red "<testId>"` observes it FAIL on the real pre-fold tree and mints the receipt → fold the fix → the normal run observes green → `--check`. Folding first is not recoverable without reverting the fix to re-observe red — observe red when it actually happens.
|
|
6
|
-
|
|
7
|
-
**Two tools — the read/run split:**
|
|
8
|
-
|
|
9
|
-
1. **Runner** — `node ${CLAUDE_SKILL_DIR}/tools/fold-completeness-run.mjs [--suite "<cmd>"] [--cwd <dir>]`, plus the red verb `--red "<test-file>#<test-name-pattern>"` — the SOLE tree-toucher and the SOLE result writer. The default run, over the in-flight plan-execution loop's dirty tree: resolve the loop (the single in-flight plan stem — 0 or >1 plans is a typed refusal), classify the changed surface by a CLOSED extension rule (assessable JS / unsupported TS-JSX / out-of-domain), run the suite ONCE under `NODE_V8_COVERAGE` (the coverage dir lives OUTSIDE the work tree — no repo artifacts), map every changed executable line to covered/uncovered from the V8 report itself, probe each bound testId N times (`AW_FOLD_RERUNS`, default 3; a shell-free spawn of the resolver's canonical path; each run under its own per-run `AW_FOLD_PROBE_TIMEOUT_S`, default 120 — probes only, never the suite) recording rerun counts + the test file's content hash (custody), and append ONE machine-only v3 run record — segment-framed (`base`) and bound to BOTH the tree fingerprint AND the segment's sorted `fixable-bug` testId set — to `<git dir>/agent-workflow-fold-completeness.jsonl` (never committable by construction; `AW_FOLD_RESULTS` overrides). **`--red` observes a testId on the CURRENT (pre-fold) tree**: failing on N/N runs → it appends the observed-red receipt (testId, counts, content hash, fingerprint); observed-green / unresolvable / mixed / timed-out are DISTINGUISHED typed refusals and nothing is written. The verdict algebra is strict N/N: RED = N/N fail, GREEN = N/N pass, anything mixed or timed out is **QUARANTINE** — it never converts and has NO override lane (a flaky pin proves nothing — replace the test). A test that cannot even LOAD pre-fold (it imports an export the fix introduces) is authored with a dynamic `import()` so it loads and fails pre-fold.
|
|
10
|
-
2. **Read-only checker** — `node ${CLAUDE_SKILL_DIR}/tools/fold-completeness.mjs [--check | --status | --json]` — never runs tests, never touches the tree, never imports the runner (an import-split test pins the structural split). Per bound testId (of the CURRENT segment) the gate requires: an N/N-green probe, an observed-red receipt in this segment, that receipt PRECEDING the loop's latest run in ledger order (a post-hoc red proves nothing — a late receipt forces one fresh run), and content **custody** — the run's recorded test-file hash equals the latest custody-eligible receipt hash on that file (eligible = the receipt's own testId is bound AND it precedes the run), so the test that is green now is byte-identical to a test that provably failed; any edit after the receipt requires re-observing red, and appending the NEXT fold's red test to the same file re-attests it without ceremony. `--status` (default) is the human report; **`--check` is the gate exit code — the normative exit contract lives in the CHECKER's tool header (the single home — do not re-enumerate it elsewhere)**; `--json` is the structured state + decision.
|
|
11
|
-
|
|
12
|
-
**The verification PROFILE — language independence (BUGFREE-3 / AD-049).** An OPTIONAL per-project `docs/ai/verification-profile.json` (seeded by the memory bootstrap; a memory-canon template with a kit mirror) generalizes the fold signal beyond JS/V8 so a consumer on another language/runner drives the SAME gate. It declares three inputs — **`coverage.kind`** `"v8"` (default) or `"lcov"` (+ `lcovPath`: your suite leaves an LCOV file there, parsed by a dependency-free reader — the diff-cover pattern); **`singleTest`** an argv template (`{file}`/`{pattern}`, plus `{resultPath}` for a file-based format) + `resultFormat` `"tap-stdout"` (default) / `"tap-file"` / `"junit-xml"` (the "0 tests matched" → unresolvable guard holds for EVERY format — a zero-match is never read green); and an optional **`findings.sarifPath`** for ADVISORY findings (`--findings` prints them — never recorded, never gate-blocking; a malformed SARIF fails the advisory read loudly but leaves `--check` unaffected). An ABSENT profile reproduces today's exact behaviour (V8 + node:test TAP on stdout); env knobs (`AW_FOLD_SUITE_CMD`/`AW_FOLD_BOUND_CMD`) still override. Every DECLARED path (`lcovPath`, `sarifPath`) MUST be gitignored or outside the repo (a symlink is refused) — else the suite writing there would move the review fingerprint. The suite COMMAND is NOT a profile field: it stays the `gates.json` `unit-tests` cmd (command-identity — the (a) credit).
|
|
13
|
-
|
|
14
|
-
**The operator verbs (recorded, self-discipline — never auto-detection):** besides `--red`, the runner offers **`--reattest "<testId>"`** ((c)): re-anchor a bound test FILE's custody at its CURRENT bytes after a GREEN-ONLY append — the honest replacement for a `red-proof` waiver (there is no red to observe); it never converts a red baseline or waives the observed-red proof, and the guard still fails closed on any un-reattested change. **`--preflight`** ((f)): the CHEAP half — from the git diff + ledgers (seconds, no suite run, no probe, no coverage) it prints the overrides/re-attests to RECORD BEFORE the expensive pass, routed by kind (`oracle-change` for a tampered file, `--reattest` for a green-only custody delta, `--red`/`red-proof` for a bound testId with no observed-red receipt). **`--findings`**: the SARIF advisory above.
|
|
15
|
-
|
|
16
|
-
**The oracle-tamper guard + the recorded override lanes (review-ledger schema v3).** The runner also records a **tamper surface** over the tracked working-vs-HEAD diff — the union of test-classified paths and the bound-testId file halves, restricted to files existing at HEAD — classified by hunk line polarity: any removed/modified line (or a deleted file) is TAMPER; pure additions and brand-new test files never trip it. Note two edges: a **rename** reads as delete+add under the pinned `--no-renames` diff, so renaming a test file trips the guard (record the override or split the rename from the fix); and a run record written by a **pre-tamper runner** (no recorded surface) fails the gate with a re-run reason — after an upgrade, one fresh run refreshes the evidence. The gate fails closed on any tampered file not named by a recorded **`oracle-change`** override (`review-ledger-write override --json '{"loop":…,"round":…,"scope":"oracle-change","files":[…],"reason":…}'` — a durable, auditable ledger entry, never a silent waiver; one override covers the named files for the rest of the loop). The second lane, scope **`red-proof`**, names ONE testId whose red is genuinely unestablishable pre-fold (or whose custody file was legitimately edited post-fix) and waives the receipt + custody proof for exactly that testId — never the N/N-green requirement, never QUARANTINE. Overrides are loop + payload scoped, minted only inside the single in-flight loop.
|
|
17
|
-
|
|
18
|
-
**Wire the gate BY HAND, once, BEFORE the round's runner run.** The candidate line for your own `docs/ai/gates.json`: `{ "id": "fold-completeness", "title": "Fold-completeness — every changed executable line executed by tests (M3a)", "cmd": "node \"<path-to-this-skill>/tools/fold-completeness.mjs\" --check" }` — the path your project reaches the kit by, kept inside the escaped quotes so a path with spaces survives. Wiring comes FIRST because where `docs/ai` is tracked, the gate-file edit itself moves the tree fingerprint — wiring after a run stales the record you just wrote. The consent-gated seeder does NOT offer this entry yet — a deliberate hold: the language-aware verification profile now EXISTS (BUGFREE-3 / AD-049 — coverage via LCOV, single-test via TAP-file/JUnit-XML), so the JS/V8-only blocker is lifted, but consumer-seeding of this gate is a SEPARATE follow-up the profile merely unblocks.
|
|
19
|
-
|
|
20
|
-
**Run-then-check usage (each plan-execution round):** run the runner AFTER the round's LAST edit — any later tree edit (a source fix, a doc touch, a gate-file tweak) moves the fingerprint and requires a re-run — then the wired `--check` attests it. Staleness is by BOTH bindings: a tree edit moves the fingerprint, and a new `fixable-bug` triage moves the bound-testId set — either makes the previous record STALE, so a green check always attests the CURRENT tree against the CURRENT fold set.
|
|
21
|
-
|
|
22
|
-
**Suite + bound-test commands:** the suite command defaults to the `unit-tests` gate `cmd` in the project's `docs/ai/gates.json`; absent that, `--suite "<cmd>"` (or `AW_FOLD_SUITE_CMD`) is REQUIRED — refused loudly otherwise. Bound-test probes default to the `node --test --test-name-pattern` shape over the testId halves (`<test-file>#<test-name-pattern>`); a consumer on another runner sets the profile's `singleTest.argv` (or the `AW_FOLD_BOUND_CMD` env override, which wins) — a JSON ARRAY of argv strings with `{file}` / `{pattern}` substitution (plus `{resultPath}` for a file-based `resultFormat`), spawned WITHOUT a shell (testId content never reaches a shell). Coverage comes from V8 by default, or from an LCOV file the suite writes (`coverage.kind:"lcov"` + `lcovPath`) — either source feeds the same uncovered-changed decision. Note: the suite's own child processes must inherit `NODE_V8_COVERAGE` for their lines to count (node child spawns inherit it by default; a test runner that scrubs its child env needs consumer configuration).
|
|
23
|
-
|
|
24
|
-
**v1 scope (stated):** assessable = `.mjs` / `.cjs` / `.js`, excluding test files (`*.test.*` / `*.spec.*`); changed `.ts` / `.tsx` / `.jsx` / `.mts` / `.cts` files are recorded as unsupported and FAIL the gate CLOSED — the signal never vouches for JS-family source it cannot assess; everything else (docs/config the suite does not execute) is recorded out-of-domain and LOUDLY listed, never gate-blocking. Executability is coverage-derived (V8 by default, or LCOV via the profile): a changed assessable file entirely absent from the coverage report is RED at file level ("never imported" must not read as "nothing executable"). The LANGUAGE-INDEPENDENCE contract (BUGFREE-3) generalizes the coverage source + single-test result format without changing this closed JS assessable rule — a consumer whose suite emits LCOV + JUnit-XML drives the same gate.
|
|
25
|
-
|
|
26
|
-
**Mutation is NOT shipped (v1).** The result record carries a reserved, always-empty `mutation` shape; the shipped runner performs NO mutation and never modifies a source file, and the checker FAILS CLOSED on a record carrying ANY mutation data (such a record was not produced by this runner version). The researched mutation half was shelved: a bounded operator set tests local expression boundaries, not the cross-module interactions that produce real fold bugs, and it is not language-independent. The `AW_FOLD_MUTANTS_MAX` / `AW_FOLD_HUNK_MUTANTS_MAX` / `AW_FOLD_TIME_BUDGET_S` budgets are parsed and recorded but inert in v1 (reserved).
|
|
27
|
-
|
|
28
|
-
**Honest residuals (stated, accepted — exactly like review-state's / review-ledger's):** coverage proves EXECUTION, not assertion — an executed-but-unasserted line still reads covered; the per-fold proof remains the M2 discipline (a `fixable-bug` testId written red→green, probed green here), the coverage run is the whole-surface prefilter; the result records, testIds, and ledgers are forgeable (`git commit --no-verify`, file editing) — a self-discipline mechanism against silent process drift, NOT a security boundary.
|
|
29
|
-
|
|
30
|
-
**Invariants:** the checker is read-only · never commits · never runs a subscription CLI · spawns read-only `git` queries only · the SOLE tree-toucher + result writer is `fold-completeness-run.mjs`, which the checker NEVER imports (read/run split, import-split test) · the runner's only writes land OUTSIDE the work tree (the coverage temp dir + the git-dir result ledger) — the tree fingerprint is byte-identical before and after a run.
|