@christang/keel 5.67.0 → 5.69.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +35 -0
- package/assets/bootstrap/AGENTS.md +1 -1
- package/package.json +1 -1
- package/plugins/keel/.claude-plugin/plugin.json +1 -1
- package/plugins/keel/.codex-plugin/plugin.json +1 -1
- package/plugins/keel/skills/keel-tdd-or-test-first/SKILL.md +4 -1
- package/scripts/validate_plugin.py +661 -2
- package/src/core/gates.js +185 -1
- package/src/core/task-contract.js +122 -1
package/README.md
CHANGED
|
@@ -344,6 +344,41 @@ wiring, and `keel --init` whenever it tells you the repository is behind its ins
|
|
|
344
344
|
protocol version lives in your `AGENTS.md`, and updating the package does not move it.
|
|
345
345
|
Everything below is the vocabulary the agent uses on your behalf.
|
|
346
346
|
|
|
347
|
+
### When the right answer is "nothing changed"
|
|
348
|
+
|
|
349
|
+
A refactor, a move, a flow upgrade that claims the numbers hold — the correct evidence for these is
|
|
350
|
+
*zero difference*, and red-green has no shape for it. The repository that reported this had a task
|
|
351
|
+
whose `tasks.md` said "this one has no honest red" several times, and re-recorded its contract twice
|
|
352
|
+
trying to fit. The criterion was right; it had nowhere to live.
|
|
353
|
+
|
|
354
|
+
```
|
|
355
|
+
- Verify:
|
|
356
|
+
- Strategy: equivalence
|
|
357
|
+
- Base: origin/main
|
|
358
|
+
- Fields: wns, tns, cell_count
|
|
359
|
+
- M1: node compare.js --base --head reports every field equal
|
|
360
|
+
```
|
|
361
|
+
|
|
362
|
+
`equivalence` owes no red. Its criterion is that base and head agree on the fields you named, which is
|
|
363
|
+
*stronger* than red-green — it also catches the change that incidentally moved a result. What the gate
|
|
364
|
+
checks is every way that shape can look complete and compare nothing: a missing `Base:` or `Fields:`,
|
|
365
|
+
an empty field set, a ref that resolves to nothing, and a `Base:` that resolves to HEAD.
|
|
366
|
+
|
|
367
|
+
It is not a way out of red-green. A task declaring `equivalence` while covering a scenario its own
|
|
368
|
+
change *adds* is refused unless a sibling task covers that same entry under a red-green strategy —
|
|
369
|
+
behavior that is new is not behavior that is unchanged, and a task cannot prove both.
|
|
370
|
+
|
|
371
|
+
And Evidence no longer has to retell the output:
|
|
372
|
+
|
|
373
|
+
```
|
|
374
|
+
- M1: artifact openspec/changes/<change>/evidence/compare.json sha256:9f2c…
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
The gate checks the file is there and the digest matches, and refuses a path outside the change's own
|
|
378
|
+
directory, because archiving moves that directory and the pointer would break. Keel hashes the bytes
|
|
379
|
+
and reads nothing inside them — the claim stays yours, and what the digest buys is that the file your
|
|
380
|
+
reviewer opens is the file you meant.
|
|
381
|
+
|
|
347
382
|
## Verification layering
|
|
348
383
|
|
|
349
384
|
Keel splits verification into two layers so a slow suite never blocks your push:
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "keel",
|
|
3
|
-
"version": "5.
|
|
3
|
+
"version": "5.69.0",
|
|
4
4
|
"description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "TanglmChris",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "keel",
|
|
3
|
-
"version": "5.
|
|
3
|
+
"version": "5.69.0",
|
|
4
4
|
"description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "TanglmChris",
|
|
@@ -20,10 +20,13 @@ Read the selected task's compiled capsule: resolved Acceptance, Verify strategy
|
|
|
20
20
|
- `regression-first`: an observable defect; reproduce it through the public interface first, then prove the fix with the same check.
|
|
21
21
|
- `characterization` / `snapshot-characterization`: deterministic or generated outputs kept stable by byte or snapshot comparison; not downgraded to build success.
|
|
22
22
|
- `rendered-behavior`: interactive surfaces exercised through the real rendered interface; strict red-green optional by cost.
|
|
23
|
-
- `
|
|
23
|
+
- `equivalence`: work whose criterion is that a measured result does not change — a refactor, a move, a flow upgrade claiming the numbers hold. It declares `Base:` (a resolvable git ref) and `Fields:` (the compared field set) beside `Strategy:` and owes no red; the A/B command is an ordinary `M<n>` check. `keel gate task-start` refuses a missing `Base:` or `Fields:`, an empty field set, an unresolvable ref, and a `Base:` that is HEAD, because an A/B against itself always agrees. A task covering a scenario its own change *adds* is refused unless a sibling task covers that same entry under a red-green strategy: behavior that is new is not behavior that is unchanged.
|
|
24
|
+
- `evidence-first`: docs, configuration, or diagnosis work whose checks state the observable artifact or evidence instead of a red-green loop. It is scoped by an absence — nothing here can fail first — which is why it is not the home for `equivalence` work, whose criterion is stronger than red-green rather than missing.
|
|
24
25
|
|
|
25
26
|
Red-green strategies (`vertical-tdd`, `regression-first`) must record concrete per-label `.red` and `.green` Evidence entries for the same check; `keel gate task-complete` rejects absent or pending entries.
|
|
26
27
|
|
|
28
|
+
A check's Evidence may read `artifact <path> sha256:<digest>` instead of retelling the output. `keel gate task-complete` checks the file exists and the digest matches, and refuses a path outside the change's own directory — archiving moves that directory, so a pointer outside it breaks. Keel hashes the bytes and reads nothing inside them: the claim the artifact supports stays yours.
|
|
29
|
+
|
|
27
30
|
## Domain lenses
|
|
28
31
|
|
|
29
32
|
When the change's proposal/design/specs or the task's Touch extensions signal a domain, consult the matching lens's `Execution and review checks` section from `keel/lenses/` — the lens whose `Applies when:` header matches — before finalizing the strategy and the first check, and load only that one. When no lens matches, load nothing.
|
|
@@ -38,8 +38,8 @@ REQUIRED_SCRIPTS = [
|
|
|
38
38
|
"scripts/validate_plugin.py",
|
|
39
39
|
]
|
|
40
40
|
|
|
41
|
-
PACKAGE_VERSION = "5.
|
|
42
|
-
PROTOCOL_VERSION = "5.
|
|
41
|
+
PACKAGE_VERSION = "5.69.0"
|
|
42
|
+
PROTOCOL_VERSION = "5.69.0"
|
|
43
43
|
LEGACY_MANAGED_START = "<!-- keel:start version=2.1 -->"
|
|
44
44
|
OPENSPEC_SCHEMA_NAME = "keel-spec-driven"
|
|
45
45
|
# Mirrors KEEL_PACKAGE_NAME in scripts/install_to_repo.py, one of the two
|
|
@@ -2302,6 +2302,61 @@ def validate_domain_lenses_scenario() -> int:
|
|
|
2302
2302
|
return 0
|
|
2303
2303
|
|
|
2304
2304
|
|
|
2305
|
+
# `bump_version.js` writes a stub section for the new release, and the convention
|
|
2306
|
+
# is to fill it in. Writing the entry *above* it instead leaves an orphan, and
|
|
2307
|
+
# the stub carries the `Version alignment:` line this scenario reads — so the
|
|
2308
|
+
# section describing nothing was the section that passed (#151). It happened in
|
|
2309
|
+
# 5.67.0 and 5.68.0, with the suite green both times.
|
|
2310
|
+
#
|
|
2311
|
+
# Takes the text rather than reading the file, so the rule can be exercised on a
|
|
2312
|
+
# planted copy: a rule that could only ever see the changelog already known to be
|
|
2313
|
+
# correct would pass forever without anyone learning whether it fires.
|
|
2314
|
+
def release_description_problem(changelog: str, version: str) -> str | None:
|
|
2315
|
+
"""Return the problem the released version's section has, or None.
|
|
2316
|
+
|
|
2317
|
+
Scoped to the version being released. An older section is history a current
|
|
2318
|
+
author cannot act on, and a rule that swept the whole file would fail on the
|
|
2319
|
+
archive — which is how a check gets disabled rather than fixed. The file also
|
|
2320
|
+
legitimately quotes `TODO` when describing the rules that refuse it.
|
|
2321
|
+
"""
|
|
2322
|
+
pattern = re.compile(
|
|
2323
|
+
r"^## " + re.escape(version) + r"\b[^\n]*$", re.M
|
|
2324
|
+
)
|
|
2325
|
+
headings = pattern.findall(changelog)
|
|
2326
|
+
if not headings:
|
|
2327
|
+
return (
|
|
2328
|
+
f"the changelog carries no `## {version}` section, so the release "
|
|
2329
|
+
"describes nothing at all."
|
|
2330
|
+
)
|
|
2331
|
+
if len(headings) > 1:
|
|
2332
|
+
return (
|
|
2333
|
+
f"the changelog describes {version} twice — {len(headings)} `## "
|
|
2334
|
+
f"{version}` headings: {', '.join(h.strip() for h in headings)}. A "
|
|
2335
|
+
"release entry written above the stub instead of into it leaves both."
|
|
2336
|
+
)
|
|
2337
|
+
for heading in headings:
|
|
2338
|
+
if "TODO" in heading:
|
|
2339
|
+
return (
|
|
2340
|
+
f"the changelog still carries the stub heading `{heading.strip()}`"
|
|
2341
|
+
f", so the {version} release was never described. Fill that "
|
|
2342
|
+
"section in rather than writing the entry above it."
|
|
2343
|
+
)
|
|
2344
|
+
# The section body, bounded by the next `## ` heading, so a `TODO` belonging
|
|
2345
|
+
# to an older release is not read as this one's.
|
|
2346
|
+
start = changelog.index(headings[0])
|
|
2347
|
+
rest = changelog[start + len(headings[0]):]
|
|
2348
|
+
end = rest.find("\n## ")
|
|
2349
|
+
body = rest if end < 0 else rest[:end]
|
|
2350
|
+
for line in body.splitlines():
|
|
2351
|
+
if re.match(r"^\s*-\s*TODO\b", line):
|
|
2352
|
+
return (
|
|
2353
|
+
f"the {version} section still carries the stub bullet "
|
|
2354
|
+
f"`{line.strip().lstrip('- ')}`, so the release is titled but "
|
|
2355
|
+
"not described."
|
|
2356
|
+
)
|
|
2357
|
+
return None
|
|
2358
|
+
|
|
2359
|
+
|
|
2305
2360
|
def validate_version_alignment_scenario() -> int:
|
|
2306
2361
|
package = json.loads((ROOT / "package.json").read_text(encoding="utf-8"))
|
|
2307
2362
|
if package.get("version") != PACKAGE_VERSION:
|
|
@@ -2366,6 +2421,103 @@ def validate_version_alignment_scenario() -> int:
|
|
|
2366
2421
|
)
|
|
2367
2422
|
return 1
|
|
2368
2423
|
|
|
2424
|
+
# The released version's section has to say what the release did. Asserted on
|
|
2425
|
+
# planted copies first, so the rule is known to fire, and then on the
|
|
2426
|
+
# repository's own changelog.
|
|
2427
|
+
stub = (
|
|
2428
|
+
f"# Keel Changelog\n\n## {PACKAGE_VERSION} - TODO: summarize this "
|
|
2429
|
+
f"release\n\n- TODO: describe the change.\n- Version alignment: "
|
|
2430
|
+
f"shares Keel {PACKAGE_VERSION}.\n\n## 5.0.0 - older\n\n- Done.\n"
|
|
2431
|
+
)
|
|
2432
|
+
planted = release_description_problem(stub, PACKAGE_VERSION)
|
|
2433
|
+
if not planted:
|
|
2434
|
+
report(
|
|
2435
|
+
"version-alignment scenario: an unfilled stub was accepted — a "
|
|
2436
|
+
f"changelog whose {PACKAGE_VERSION} section is the stub "
|
|
2437
|
+
"`bump_version.js` writes produced no problem, so a release can "
|
|
2438
|
+
"ship describing nothing while this scenario passes."
|
|
2439
|
+
)
|
|
2440
|
+
return 1
|
|
2441
|
+
if "TODO: summarize this release" not in planted:
|
|
2442
|
+
report(
|
|
2443
|
+
"version-alignment scenario: the refusal does not name the offending "
|
|
2444
|
+
f"line, so an author is sent to read the whole changelog; got "
|
|
2445
|
+
f"{planted!r}."
|
|
2446
|
+
)
|
|
2447
|
+
return 1
|
|
2448
|
+
|
|
2449
|
+
# Still M1 — the stub survives renaming its heading. The bullet is the other
|
|
2450
|
+
# half of what `bump_version.js` writes, and a section titled correctly while
|
|
2451
|
+
# its only content is `- TODO: describe the change.` describes nothing.
|
|
2452
|
+
renamed = (
|
|
2453
|
+
f"# Keel Changelog\n\n## {PACKAGE_VERSION} - a real sounding title\n\n"
|
|
2454
|
+
f"- TODO: describe the change.\n- Version alignment: shares Keel "
|
|
2455
|
+
f"{PACKAGE_VERSION}.\n"
|
|
2456
|
+
)
|
|
2457
|
+
bullet = release_description_problem(renamed, PACKAGE_VERSION)
|
|
2458
|
+
if not bullet:
|
|
2459
|
+
report(
|
|
2460
|
+
"version-alignment scenario: an unfilled stub was accepted — a "
|
|
2461
|
+
"section with a real heading whose content is still `- TODO: describe "
|
|
2462
|
+
"the change.` produced no problem, so renaming the stub is enough to "
|
|
2463
|
+
"ship an undescribed release."
|
|
2464
|
+
)
|
|
2465
|
+
return 1
|
|
2466
|
+
if "TODO: describe the change." not in bullet:
|
|
2467
|
+
report(
|
|
2468
|
+
"version-alignment scenario: the bullet refusal does not name the "
|
|
2469
|
+
f"line it found; got {bullet!r}."
|
|
2470
|
+
)
|
|
2471
|
+
return 1
|
|
2472
|
+
|
|
2473
|
+
# M2 — two sections for one version. This is the shape the real defect had,
|
|
2474
|
+
# and neither TODO rule catches it once the stub's heading has been renamed:
|
|
2475
|
+
# the reader then sees two identical version headings, one of which describes
|
|
2476
|
+
# nothing.
|
|
2477
|
+
doubled = (
|
|
2478
|
+
f"# Keel Changelog\n\n## {PACKAGE_VERSION} - the real entry\n\n"
|
|
2479
|
+
f"- Did the thing.\n- Version alignment: shares Keel {PACKAGE_VERSION}."
|
|
2480
|
+
f"\n\n## {PACKAGE_VERSION} - summarized\n\n- Also this.\n"
|
|
2481
|
+
)
|
|
2482
|
+
problem_found = release_description_problem(doubled, PACKAGE_VERSION)
|
|
2483
|
+
if not problem_found:
|
|
2484
|
+
report(
|
|
2485
|
+
"version-alignment scenario: a version described twice was accepted "
|
|
2486
|
+
f"— two `## {PACKAGE_VERSION}` sections produced no problem, which is "
|
|
2487
|
+
"the shape the defect this rule exists for actually had."
|
|
2488
|
+
)
|
|
2489
|
+
return 1
|
|
2490
|
+
if "twice" not in problem_found:
|
|
2491
|
+
report(
|
|
2492
|
+
"version-alignment scenario: the duplicate-section refusal does not "
|
|
2493
|
+
f"say the version is described twice; got {problem_found!r}."
|
|
2494
|
+
)
|
|
2495
|
+
return 1
|
|
2496
|
+
|
|
2497
|
+
# M3 — only the version being released is judged. An older section is history,
|
|
2498
|
+
# and this file legitimately quotes `TODO` when describing the rules that
|
|
2499
|
+
# refuse it, so a rule sweeping the whole file would fire on its own
|
|
2500
|
+
# documentation.
|
|
2501
|
+
older = (
|
|
2502
|
+
f"# Keel Changelog\n\n## {PACKAGE_VERSION} - the real entry\n\n"
|
|
2503
|
+
f"- Did the thing.\n- Version alignment: shares Keel {PACKAGE_VERSION}."
|
|
2504
|
+
"\n\n## 5.0.0 - TODO: summarize this release\n\n- TODO: describe it.\n"
|
|
2505
|
+
)
|
|
2506
|
+
if release_description_problem(older, PACKAGE_VERSION):
|
|
2507
|
+
report(
|
|
2508
|
+
"version-alignment scenario: an older section was refused — a `TODO` "
|
|
2509
|
+
"in a section for another version failed the current release, which "
|
|
2510
|
+
"is a check the archive can only be edited to satisfy."
|
|
2511
|
+
)
|
|
2512
|
+
return 1
|
|
2513
|
+
if release_description_problem(changelog, PACKAGE_VERSION):
|
|
2514
|
+
report(
|
|
2515
|
+
"version-alignment scenario: an older section was refused — the "
|
|
2516
|
+
"repository's own changelog does not satisfy the rule: "
|
|
2517
|
+
f"{release_description_problem(changelog, PACKAGE_VERSION)}"
|
|
2518
|
+
)
|
|
2519
|
+
return 1
|
|
2520
|
+
|
|
2369
2521
|
with tempfile.TemporaryDirectory(prefix="keel-version-") as raw_tmp:
|
|
2370
2522
|
repo = Path(raw_tmp) / "repo"
|
|
2371
2523
|
repo.mkdir()
|
|
@@ -29042,6 +29194,509 @@ def validate_guidance_is_referenced_and_carries_no_criterion_scenario() -> int:
|
|
|
29042
29194
|
return 0
|
|
29043
29195
|
|
|
29044
29196
|
|
|
29197
|
+
def equivalence_task(
|
|
29198
|
+
*,
|
|
29199
|
+
base: str | None = "HEAD~1",
|
|
29200
|
+
fields: str | None = "wns, tns, cell_count",
|
|
29201
|
+
covers: tuple[str, ...] = ("E1: the measured result does not move",),
|
|
29202
|
+
label: str = "1.1",
|
|
29203
|
+
title: str = "Move the attribute without moving the numbers",
|
|
29204
|
+
) -> str:
|
|
29205
|
+
"""One `equivalence` task, with either declaration omittable."""
|
|
29206
|
+
lines = [
|
|
29207
|
+
f"- [ ] {label} {title}",
|
|
29208
|
+
" - Covers:",
|
|
29209
|
+
]
|
|
29210
|
+
lines.extend(f" - {entry}" for entry in covers)
|
|
29211
|
+
lines.extend(
|
|
29212
|
+
[
|
|
29213
|
+
" - Read:",
|
|
29214
|
+
" - README.md",
|
|
29215
|
+
" - Touch:",
|
|
29216
|
+
" - src/example.js",
|
|
29217
|
+
" - Verify:",
|
|
29218
|
+
" - Strategy: equivalence",
|
|
29219
|
+
]
|
|
29220
|
+
)
|
|
29221
|
+
if base is not None:
|
|
29222
|
+
lines.append(f" - Base: {base}")
|
|
29223
|
+
if fields is not None:
|
|
29224
|
+
lines.append(f" - Fields: {fields}")
|
|
29225
|
+
lines.extend(
|
|
29226
|
+
[
|
|
29227
|
+
" - M1: node compare.js --base --head reports every field equal",
|
|
29228
|
+
" - Autonomy boundary:",
|
|
29229
|
+
" - Default: hard-stop",
|
|
29230
|
+
" - Pre-authorized fallback: none",
|
|
29231
|
+
" - Stop Rules:",
|
|
29232
|
+
" - Stop on any field that differs.",
|
|
29233
|
+
" - Evidence:",
|
|
29234
|
+
" - Contract: pending",
|
|
29235
|
+
" - M1: pending",
|
|
29236
|
+
" - Review:",
|
|
29237
|
+
" - Status: pass",
|
|
29238
|
+
" - Acceptance check: every declared field agreed.",
|
|
29239
|
+
" - Scope check: writes stayed inside Touch.",
|
|
29240
|
+
" - Findings: none",
|
|
29241
|
+
" - Blocker: none",
|
|
29242
|
+
]
|
|
29243
|
+
)
|
|
29244
|
+
return "\n".join(lines) + "\n"
|
|
29245
|
+
|
|
29246
|
+
|
|
29247
|
+
def validate_an_equivalence_claim_names_its_base_scenario() -> int:
|
|
29248
|
+
"""Issue #142: the correct evidence for a refactor is zero difference.
|
|
29249
|
+
|
|
29250
|
+
Red-green has no shape for it. The reporting repository re-recorded one
|
|
29251
|
+
task's contract twice to get past the shape — not because a criterion was
|
|
29252
|
+
wrong, but because the criterion had nowhere to live. An A/B against a base
|
|
29253
|
+
is *stronger* than red-green: it catches the change that also, incidentally,
|
|
29254
|
+
moved a result. What it needs is a place to say what it compares against and
|
|
29255
|
+
on which fields, plus a refusal for every way that shape can be complete and
|
|
29256
|
+
still compare nothing.
|
|
29257
|
+
"""
|
|
29258
|
+
label = "an-equivalence-claim-names-its-base"
|
|
29259
|
+
|
|
29260
|
+
def git(repo: Path, *args: str) -> subprocess.CompletedProcess[str]:
|
|
29261
|
+
return subprocess.run(
|
|
29262
|
+
["git", "-C", str(repo), *args], capture_output=True, text=True
|
|
29263
|
+
)
|
|
29264
|
+
|
|
29265
|
+
with tempfile.TemporaryDirectory(prefix="keel-equivalence-") as raw:
|
|
29266
|
+
root = Path(raw)
|
|
29267
|
+
|
|
29268
|
+
def fixture(name: str, task: str) -> Path:
|
|
29269
|
+
repo = (root / name).resolve()
|
|
29270
|
+
repo.mkdir()
|
|
29271
|
+
git(repo, "init", "-q")
|
|
29272
|
+
git(repo, "config", "user.email", "t@example.com")
|
|
29273
|
+
git(repo, "config", "user.name", "keel-test")
|
|
29274
|
+
write_gate_fixture(repo, tasks=task)
|
|
29275
|
+
write_text(repo / "src/example.js", "// product\n")
|
|
29276
|
+
git(repo, "add", "-A")
|
|
29277
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
|
|
29278
|
+
write_text(repo / "src/example.js", "// product, moved\n")
|
|
29279
|
+
git(repo, "add", "-A")
|
|
29280
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
|
|
29281
|
+
return repo
|
|
29282
|
+
|
|
29283
|
+
def start(name: str, **kwargs) -> dict:
|
|
29284
|
+
repo = fixture(name, equivalence_task(**kwargs))
|
|
29285
|
+
result = run_keel(
|
|
29286
|
+
repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
|
|
29287
|
+
"--no-guard", "--json",
|
|
29288
|
+
)
|
|
29289
|
+
try:
|
|
29290
|
+
return json.loads(result.stdout)
|
|
29291
|
+
except json.JSONDecodeError:
|
|
29292
|
+
return {
|
|
29293
|
+
"status": "unparsed",
|
|
29294
|
+
"problems": [{"message": result.stdout[:400]}],
|
|
29295
|
+
}
|
|
29296
|
+
|
|
29297
|
+
# M1 — the strategy is accepted, and it owes no red.
|
|
29298
|
+
payload = start("accepted")
|
|
29299
|
+
if payload.get("status") != "pass":
|
|
29300
|
+
report(
|
|
29301
|
+
f"{label}: unsupported verification strategy — a complete "
|
|
29302
|
+
"equivalence task was refused, so the one task class with the "
|
|
29303
|
+
"strongest criterion still has nowhere to declare it. "
|
|
29304
|
+
f"{problem_text(payload)}"
|
|
29305
|
+
)
|
|
29306
|
+
return 1
|
|
29307
|
+
capsule = (payload.get("contract") or {}).get("capsule") or {}
|
|
29308
|
+
verification = capsule.get("verification") or {}
|
|
29309
|
+
if verification.get("strategy", "").lower() != "equivalence":
|
|
29310
|
+
report(
|
|
29311
|
+
f"{label}: the compiled capsule does not carry the strategy; got "
|
|
29312
|
+
f"{verification.get('strategy')!r}."
|
|
29313
|
+
)
|
|
29314
|
+
return 1
|
|
29315
|
+
if not verification.get("base") or not verification.get("fields"):
|
|
29316
|
+
report(
|
|
29317
|
+
f"{label}: the capsule drops the declarations — base "
|
|
29318
|
+
f"{verification.get('base')!r}, fields "
|
|
29319
|
+
f"{verification.get('fields')!r}. A declaration absent from the "
|
|
29320
|
+
"capsule is absent from the fingerprint, so it could be edited "
|
|
29321
|
+
"after the run without moving the contract."
|
|
29322
|
+
)
|
|
29323
|
+
return 1
|
|
29324
|
+
warnings = " ".join(str(w) for w in (payload.get("warnings") or []))
|
|
29325
|
+
if ".red" in warnings or ".green" in warnings:
|
|
29326
|
+
report(
|
|
29327
|
+
f"{label}: equivalence was given a red-green obligation; the "
|
|
29328
|
+
f"criterion is agreement with a base, not a failing first run. "
|
|
29329
|
+
f"{warnings}"
|
|
29330
|
+
)
|
|
29331
|
+
return 1
|
|
29332
|
+
|
|
29333
|
+
# M2 — every way the shape can be complete and still compare nothing,
|
|
29334
|
+
# each named for the declaration it is about. A diagnostic that named
|
|
29335
|
+
# the strategy would send the author to the line that is correct.
|
|
29336
|
+
for name, kwargs, code, expected in (
|
|
29337
|
+
("no-base", {"base": None}, "missing-equivalence-base", "Base:"),
|
|
29338
|
+
("no-fields", {"fields": None}, "missing-equivalence-fields", "Fields:"),
|
|
29339
|
+
("empty-fields", {"fields": " , "}, "missing-equivalence-fields", "Fields:"),
|
|
29340
|
+
(
|
|
29341
|
+
"bad-base",
|
|
29342
|
+
{"base": "no-such-ref"},
|
|
29343
|
+
"unresolvable-equivalence-base",
|
|
29344
|
+
"no-such-ref",
|
|
29345
|
+
),
|
|
29346
|
+
):
|
|
29347
|
+
payload = start(name, **kwargs)
|
|
29348
|
+
if payload.get("status") == "pass":
|
|
29349
|
+
report(
|
|
29350
|
+
f"{label}: accepted an equivalence task that compares "
|
|
29351
|
+
f"nothing — the {name} fixture passed, so the declaration "
|
|
29352
|
+
"is optional in practice."
|
|
29353
|
+
)
|
|
29354
|
+
return 1
|
|
29355
|
+
if code not in problem_codes(payload):
|
|
29356
|
+
report(
|
|
29357
|
+
f"{label}: accepted an equivalence task that compares "
|
|
29358
|
+
f"nothing — the {name} fixture was refused for another "
|
|
29359
|
+
f"reason; expected {code}, got {problem_codes(payload)!r}."
|
|
29360
|
+
)
|
|
29361
|
+
return 1
|
|
29362
|
+
message = problem_text(payload)
|
|
29363
|
+
if expected not in message:
|
|
29364
|
+
report(
|
|
29365
|
+
f"{label}: the {name} refusal does not name {expected!r}, so "
|
|
29366
|
+
f"the author is sent to find which line is wrong; got "
|
|
29367
|
+
f"{message!r}."
|
|
29368
|
+
)
|
|
29369
|
+
return 1
|
|
29370
|
+
# Absent and empty are the same state to the comparison and different
|
|
29371
|
+
# states to the author, so the two are asserted to read differently.
|
|
29372
|
+
absent = problem_text(start("no-fields-message", fields=None))
|
|
29373
|
+
empty = problem_text(start("empty-fields-message", fields=" , "))
|
|
29374
|
+
if absent == empty:
|
|
29375
|
+
report(
|
|
29376
|
+
f"{label}: a missing `Fields:` and one that resolves to an empty "
|
|
29377
|
+
"set produce the same sentence, so an author who wrote the line "
|
|
29378
|
+
"is told they did not."
|
|
29379
|
+
)
|
|
29380
|
+
return 1
|
|
29381
|
+
|
|
29382
|
+
# M1 of 1.2 — `equivalence` owes no red, which makes it the first thing
|
|
29383
|
+
# reached for by a task that should have one. A task covering a scenario
|
|
29384
|
+
# the change *adds* is claiming new behavior and unchanged behavior at
|
|
29385
|
+
# once, and one of the two claims has no proof anywhere.
|
|
29386
|
+
added_spec = (
|
|
29387
|
+
"## ADDED Requirements\n\n"
|
|
29388
|
+
"### Requirement: The moved attribute keeps its effect\n\n"
|
|
29389
|
+
"The attribute SHALL keep its effect after the move.\n\n"
|
|
29390
|
+
"#### Scenario: The effect survives the move\n\n"
|
|
29391
|
+
"- **WHEN** the attribute moves into the flow\n"
|
|
29392
|
+
"- **THEN** the effect is unchanged\n\n"
|
|
29393
|
+
# A second real scenario, so the mismatched-sibling control fails
|
|
29394
|
+
# because the guard refused it and not because its Covers entry
|
|
29395
|
+
# resolves to nothing. A first attempt pointed the sibling at an
|
|
29396
|
+
# invented scenario and passed for that unrelated reason.
|
|
29397
|
+
"#### Scenario: The flow reports the attribute\n\n"
|
|
29398
|
+
"- **WHEN** the flow runs\n"
|
|
29399
|
+
"- **THEN** it reports the attribute\n"
|
|
29400
|
+
)
|
|
29401
|
+
covered = (
|
|
29402
|
+
"demo / The moved attribute keeps its effect / The effect survives "
|
|
29403
|
+
"the move"
|
|
29404
|
+
)
|
|
29405
|
+
|
|
29406
|
+
def escape_fixture(name: str, *, sibling: str | None) -> dict:
|
|
29407
|
+
repo = (root / name).resolve()
|
|
29408
|
+
repo.mkdir()
|
|
29409
|
+
git(repo, "init", "-q")
|
|
29410
|
+
git(repo, "config", "user.email", "t@example.com")
|
|
29411
|
+
git(repo, "config", "user.name", "keel-test")
|
|
29412
|
+
tasks = equivalence_task(covers=(covered,))
|
|
29413
|
+
if sibling is not None:
|
|
29414
|
+
tasks += (
|
|
29415
|
+
"- [ ] 1.2 Add the behavior\n"
|
|
29416
|
+
" - Covers:\n"
|
|
29417
|
+
f" - {sibling}\n"
|
|
29418
|
+
" - Read:\n - README.md\n"
|
|
29419
|
+
" - Touch:\n - src/other.js\n"
|
|
29420
|
+
" - Verify:\n"
|
|
29421
|
+
" - Strategy: vertical-tdd\n"
|
|
29422
|
+
" - M1: node test.js asserts the new behavior\n"
|
|
29423
|
+
" - Autonomy boundary:\n"
|
|
29424
|
+
" - Default: hard-stop\n"
|
|
29425
|
+
" - Pre-authorized fallback: none\n"
|
|
29426
|
+
" - Stop Rules:\n - Stop on failure.\n"
|
|
29427
|
+
" - Evidence:\n - Contract: pending\n - M1: pending\n"
|
|
29428
|
+
" - Review:\n - Status: pass\n"
|
|
29429
|
+
" - Acceptance check: behavior asserted.\n"
|
|
29430
|
+
" - Scope check: inside Touch.\n"
|
|
29431
|
+
" - Findings: none\n"
|
|
29432
|
+
" - Blocker: none\n"
|
|
29433
|
+
)
|
|
29434
|
+
write_gate_fixture(repo, tasks=tasks)
|
|
29435
|
+
write_text(repo / "openspec/changes/demo/specs/demo/spec.md", added_spec)
|
|
29436
|
+
write_text(repo / "src/example.js", "// product\n")
|
|
29437
|
+
git(repo, "add", "-A")
|
|
29438
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
|
|
29439
|
+
write_text(repo / "src/example.js", "// moved\n")
|
|
29440
|
+
git(repo, "add", "-A")
|
|
29441
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
|
|
29442
|
+
result = run_keel(
|
|
29443
|
+
repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
|
|
29444
|
+
"--no-guard", "--json",
|
|
29445
|
+
)
|
|
29446
|
+
try:
|
|
29447
|
+
return json.loads(result.stdout)
|
|
29448
|
+
except json.JSONDecodeError:
|
|
29449
|
+
return {
|
|
29450
|
+
"status": "unparsed",
|
|
29451
|
+
"problems": [{"message": result.stdout[:400]}],
|
|
29452
|
+
}
|
|
29453
|
+
|
|
29454
|
+
payload = escape_fixture("escape-alone", sibling=None)
|
|
29455
|
+
if payload.get("status") == "pass":
|
|
29456
|
+
report(
|
|
29457
|
+
f"{label}: accepted new behavior with no red anywhere — an "
|
|
29458
|
+
"equivalence task covering a scenario the change adds passed, so "
|
|
29459
|
+
"the strategy is a way to author a feature with no red in the "
|
|
29460
|
+
"whole change."
|
|
29461
|
+
)
|
|
29462
|
+
return 1
|
|
29463
|
+
if "equivalence-covers-added-behavior" not in problem_codes(payload):
|
|
29464
|
+
report(
|
|
29465
|
+
f"{label}: accepted new behavior with no red anywhere — refused "
|
|
29466
|
+
f"for another reason; got {problem_codes(payload)!r}."
|
|
29467
|
+
)
|
|
29468
|
+
return 1
|
|
29469
|
+
message = problem_text(payload)
|
|
29470
|
+
if "The effect survives the move" not in message:
|
|
29471
|
+
report(
|
|
29472
|
+
f"{label}: the refusal does not name the covered scenario, so "
|
|
29473
|
+
f"the author cannot tell which Covers entry is the problem; got "
|
|
29474
|
+
f"{message!r}."
|
|
29475
|
+
)
|
|
29476
|
+
return 1
|
|
29477
|
+
if "unchanged" not in message:
|
|
29478
|
+
report(
|
|
29479
|
+
f"{label}: the refusal does not say why the two claims conflict; "
|
|
29480
|
+
f"got {message!r}."
|
|
29481
|
+
)
|
|
29482
|
+
return 1
|
|
29483
|
+
|
|
29484
|
+
# M2 of 1.2 — the guard is satisfied by coverage of the entry it
|
|
29485
|
+
# objected to, and not by a red-green task merely existing in the change.
|
|
29486
|
+
payload = escape_fixture("escape-sibling", sibling=covered)
|
|
29487
|
+
if payload.get("status") != "pass":
|
|
29488
|
+
report(
|
|
29489
|
+
f"{label}: a sibling task covering the same scenario under "
|
|
29490
|
+
f"vertical-tdd did not satisfy the guard. {problem_text(payload)}"
|
|
29491
|
+
)
|
|
29492
|
+
return 1
|
|
29493
|
+
payload = escape_fixture(
|
|
29494
|
+
"escape-other-sibling",
|
|
29495
|
+
sibling=(
|
|
29496
|
+
"demo / The moved attribute keeps its effect / The flow reports "
|
|
29497
|
+
"the attribute"
|
|
29498
|
+
),
|
|
29499
|
+
)
|
|
29500
|
+
if payload.get("status") == "pass":
|
|
29501
|
+
report(
|
|
29502
|
+
f"{label}: any sibling satisfied the guard — a red-green task "
|
|
29503
|
+
"covering a different scenario was accepted as proof of this "
|
|
29504
|
+
"one, which makes the guard a check that a change contains at "
|
|
29505
|
+
"least one red-green task."
|
|
29506
|
+
)
|
|
29507
|
+
return 1
|
|
29508
|
+
if "equivalence-covers-added-behavior" not in problem_codes(payload):
|
|
29509
|
+
report(
|
|
29510
|
+
f"{label}: any sibling satisfied the guard — refused for another "
|
|
29511
|
+
f"reason; got {problem_codes(payload)!r}."
|
|
29512
|
+
)
|
|
29513
|
+
return 1
|
|
29514
|
+
|
|
29515
|
+
# M3 of 1.2 — the guard fires on new behavior, not on the strategy. A
|
|
29516
|
+
# task covering only identifiers, or a requirement the change does not
|
|
29517
|
+
# add, is an ordinary equivalence task.
|
|
29518
|
+
plain = start("escape-plain")
|
|
29519
|
+
if plain.get("status") != "pass":
|
|
29520
|
+
report(
|
|
29521
|
+
f"{label}: an equivalence task covering no added scenario was "
|
|
29522
|
+
f"refused. {problem_text(plain)}"
|
|
29523
|
+
)
|
|
29524
|
+
return 1
|
|
29525
|
+
|
|
29526
|
+
# M3 — the shape that is complete, resolvable, and still empty. This is
|
|
29527
|
+
# the refusal worth having: nothing about the task looks wrong, and the
|
|
29528
|
+
# check passes having compared a thing against itself.
|
|
29529
|
+
payload = start("base-is-head", base="HEAD")
|
|
29530
|
+
if payload.get("status") == "pass":
|
|
29531
|
+
report(
|
|
29532
|
+
f"{label}: accepted a base that is head — an A/B against itself "
|
|
29533
|
+
"always agrees, so the check proves nothing and looks complete "
|
|
29534
|
+
"doing it."
|
|
29535
|
+
)
|
|
29536
|
+
return 1
|
|
29537
|
+
if "equivalence-base-is-head" not in problem_codes(payload):
|
|
29538
|
+
report(
|
|
29539
|
+
f"{label}: accepted a base that is head — refused for another "
|
|
29540
|
+
f"reason; got {problem_codes(payload)!r}."
|
|
29541
|
+
)
|
|
29542
|
+
return 1
|
|
29543
|
+
message = problem_text(payload)
|
|
29544
|
+
head = subprocess.run(
|
|
29545
|
+
["git", "-C", str((root / "base-is-head").resolve()), "rev-parse",
|
|
29546
|
+
"HEAD"],
|
|
29547
|
+
capture_output=True, text=True,
|
|
29548
|
+
).stdout.strip()
|
|
29549
|
+
if head and head not in message:
|
|
29550
|
+
report(
|
|
29551
|
+
f"{label}: the refusal does not name the resolved commit, so the "
|
|
29552
|
+
"author cannot tell which ref collapsed onto HEAD; got "
|
|
29553
|
+
f"{message!r}."
|
|
29554
|
+
)
|
|
29555
|
+
return 1
|
|
29556
|
+
if "always agrees" not in message:
|
|
29557
|
+
report(
|
|
29558
|
+
f"{label}: the refusal does not say why a base that is HEAD is "
|
|
29559
|
+
f"empty rather than merely redundant; got {message!r}."
|
|
29560
|
+
)
|
|
29561
|
+
return 1
|
|
29562
|
+
|
|
29563
|
+
# 1.3 — Evidence may point at the machine output instead of retelling
|
|
29564
|
+
# it. A 244-line tasks.md that is mostly transcribed test output is a
|
|
29565
|
+
# transcription that can be wrong and that nobody can re-check.
|
|
29566
|
+
def artifact_fixture(
|
|
29567
|
+
name: str,
|
|
29568
|
+
*,
|
|
29569
|
+
artifact_path: str = "openspec/changes/demo/evidence/compare.json",
|
|
29570
|
+
body: str = '{"wns": 0.0, "tns": 0.0}\n',
|
|
29571
|
+
recorded: str | None = None,
|
|
29572
|
+
write_at: str | None = None,
|
|
29573
|
+
) -> dict:
|
|
29574
|
+
repo = (root / name).resolve()
|
|
29575
|
+
repo.mkdir()
|
|
29576
|
+
git(repo, "init", "-q")
|
|
29577
|
+
git(repo, "config", "user.email", "t@example.com")
|
|
29578
|
+
git(repo, "config", "user.name", "keel-test")
|
|
29579
|
+
digest = hashlib.sha256(body.encode("utf-8")).hexdigest()
|
|
29580
|
+
evidence = recorded or f"artifact {artifact_path} sha256:{digest}"
|
|
29581
|
+
tasks = equivalence_task().replace(
|
|
29582
|
+
" - M1: pending", f" - M1: {evidence}"
|
|
29583
|
+
)
|
|
29584
|
+
write_gate_fixture(repo, tasks=tasks)
|
|
29585
|
+
if write_at is not None:
|
|
29586
|
+
write_text(repo / write_at, body)
|
|
29587
|
+
write_text(repo / "src/example.js", "// product\n")
|
|
29588
|
+
git(repo, "add", "-A")
|
|
29589
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
|
|
29590
|
+
write_text(repo / "src/example.js", "// moved\n")
|
|
29591
|
+
git(repo, "add", "-A")
|
|
29592
|
+
git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
|
|
29593
|
+
run_keel(
|
|
29594
|
+
repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
|
|
29595
|
+
"--record", "--no-guard",
|
|
29596
|
+
)
|
|
29597
|
+
result = run_keel(
|
|
29598
|
+
repo, "gate", "task-complete", "--change", "demo", "--task",
|
|
29599
|
+
"1.1", "--json",
|
|
29600
|
+
)
|
|
29601
|
+
try:
|
|
29602
|
+
return json.loads(result.stdout)
|
|
29603
|
+
except json.JSONDecodeError:
|
|
29604
|
+
return {
|
|
29605
|
+
"status": "unparsed",
|
|
29606
|
+
"problems": [{"message": result.stdout[:400]}],
|
|
29607
|
+
}
|
|
29608
|
+
|
|
29609
|
+
inside = "openspec/changes/demo/evidence/compare.json"
|
|
29610
|
+
payload = artifact_fixture("artifact-ok", write_at=inside)
|
|
29611
|
+
if payload.get("status") != "pass":
|
|
29612
|
+
report(
|
|
29613
|
+
f"{label}: an artifact reference was not verified — a reference "
|
|
29614
|
+
"to a file that is there, with a digest that matches, was "
|
|
29615
|
+
f"refused. {problem_text(payload)}"
|
|
29616
|
+
)
|
|
29617
|
+
return 1
|
|
29618
|
+
payload = artifact_fixture("artifact-absent", write_at=None)
|
|
29619
|
+
if payload.get("status") == "pass":
|
|
29620
|
+
report(
|
|
29621
|
+
f"{label}: an artifact reference was not verified — a reference "
|
|
29622
|
+
"to a file that does not exist was accepted, so the form is "
|
|
29623
|
+
"tolerated as prose rather than checked. Any sentence would "
|
|
29624
|
+
"have passed the same way."
|
|
29625
|
+
)
|
|
29626
|
+
return 1
|
|
29627
|
+
if "artifact-missing" not in problem_codes(payload):
|
|
29628
|
+
report(
|
|
29629
|
+
f"{label}: an artifact reference was not verified — refused for "
|
|
29630
|
+
f"another reason; got {problem_codes(payload)!r}."
|
|
29631
|
+
)
|
|
29632
|
+
return 1
|
|
29633
|
+
|
|
29634
|
+
# M2 of 1.3 — the digest is what makes the pointer worth more than a
|
|
29635
|
+
# path. A file that moved after the digest was recorded is the case a
|
|
29636
|
+
# bare path cannot see, and it is the common one: the command gets
|
|
29637
|
+
# re-run.
|
|
29638
|
+
stale = hashlib.sha256(b"different\n").hexdigest()
|
|
29639
|
+
payload = artifact_fixture(
|
|
29640
|
+
"artifact-stale",
|
|
29641
|
+
recorded=f"artifact {inside} sha256:{stale}",
|
|
29642
|
+
write_at=inside,
|
|
29643
|
+
)
|
|
29644
|
+
if payload.get("status") == "pass":
|
|
29645
|
+
report(
|
|
29646
|
+
f"{label}: accepted a stale digest — the artifact's content does "
|
|
29647
|
+
"not hash to the recorded digest and the reference was accepted, "
|
|
29648
|
+
"so Review reads whatever the file says now."
|
|
29649
|
+
)
|
|
29650
|
+
return 1
|
|
29651
|
+
if "artifact-digest-mismatch" not in problem_codes(payload):
|
|
29652
|
+
report(
|
|
29653
|
+
f"{label}: accepted a stale digest — refused for another reason; "
|
|
29654
|
+
f"got {problem_codes(payload)!r}."
|
|
29655
|
+
)
|
|
29656
|
+
return 1
|
|
29657
|
+
message = problem_text(payload)
|
|
29658
|
+
for expected in (inside, stale[:12]):
|
|
29659
|
+
if expected not in message:
|
|
29660
|
+
report(
|
|
29661
|
+
f"{label}: the refusal does not name {expected!r}, so a "
|
|
29662
|
+
"reader cannot tell a stale record from a wrong path; got "
|
|
29663
|
+
f"{message!r}."
|
|
29664
|
+
)
|
|
29665
|
+
return 1
|
|
29666
|
+
|
|
29667
|
+
# M3 of 1.3 — the inverse of the `Durable owner:` rule, for the opposite
|
|
29668
|
+
# reason: a follow-up pointer must outlive the change, and an evidence
|
|
29669
|
+
# artifact must travel with it. `openspec archive` moves the change
|
|
29670
|
+
# directory, so a path outside it is one the archive leaves behind.
|
|
29671
|
+
outside = "evidence/compare.json"
|
|
29672
|
+
payload = artifact_fixture("artifact-outside", artifact_path=outside, write_at=outside)
|
|
29673
|
+
if payload.get("status") == "pass":
|
|
29674
|
+
report(
|
|
29675
|
+
f"{label}: accepted a path archiving would leave behind — an "
|
|
29676
|
+
f"artifact at {outside!r} was accepted although the archive "
|
|
29677
|
+
"moves only the change directory."
|
|
29678
|
+
)
|
|
29679
|
+
return 1
|
|
29680
|
+
if "artifact-outside-change" not in problem_codes(payload):
|
|
29681
|
+
report(
|
|
29682
|
+
f"{label}: accepted a path archiving would leave behind — "
|
|
29683
|
+
f"refused for another reason; got {problem_codes(payload)!r}."
|
|
29684
|
+
)
|
|
29685
|
+
return 1
|
|
29686
|
+
if "archiv" not in problem_text(payload):
|
|
29687
|
+
report(
|
|
29688
|
+
f"{label}: the refusal does not say that archiving is what "
|
|
29689
|
+
f"breaks the pointer; got {problem_text(payload)!r}."
|
|
29690
|
+
)
|
|
29691
|
+
return 1
|
|
29692
|
+
|
|
29693
|
+
if label not in {name for name, _ in SCENARIOS}:
|
|
29694
|
+
report(f"{label}: the scenario registry does not include it.")
|
|
29695
|
+
return 1
|
|
29696
|
+
report(f"{label} scenario passed.")
|
|
29697
|
+
return 0
|
|
29698
|
+
|
|
29699
|
+
|
|
29045
29700
|
SCENARIOS: tuple = (
|
|
29046
29701
|
("stateless-continuity", validate_stateless_continuity_scenario),
|
|
29047
29702
|
("core-gates", validate_core_gates_scenario),
|
|
@@ -29404,6 +30059,10 @@ SCENARIOS: tuple = (
|
|
|
29404
30059
|
"a-negation-is-not-a-marker",
|
|
29405
30060
|
validate_a_negation_is_not_a_marker_scenario,
|
|
29406
30061
|
),
|
|
30062
|
+
(
|
|
30063
|
+
"an-equivalence-claim-names-its-base",
|
|
30064
|
+
validate_an_equivalence_claim_names_its_base_scenario,
|
|
30065
|
+
),
|
|
29407
30066
|
(
|
|
29408
30067
|
"guidance-is-referenced-and-carries-no-criterion",
|
|
29409
30068
|
validate_guidance_is_referenced_and_carries_no_criterion_scenario,
|
package/src/core/gates.js
CHANGED
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
// Keel 4.1.0 deterministic gate contract.
|
|
4
4
|
|
|
5
|
+
const crypto = require("crypto");
|
|
5
6
|
const fs = require("fs");
|
|
6
7
|
const path = require("path");
|
|
7
8
|
const { spawnSync } = require("child_process");
|
|
@@ -293,6 +294,112 @@ function taskShapeWarnings(repo, selection, task, compiled) {
|
|
|
293
294
|
];
|
|
294
295
|
}
|
|
295
296
|
|
|
297
|
+
// The requirement names a change declares as new, read from the `## ADDED
|
|
298
|
+
// Requirements` sections of its delta specs. A `## MODIFIED` requirement is
|
|
299
|
+
// deliberately not here: behavior that changed is still behavior an equivalence
|
|
300
|
+
// task can legitimately claim is measurement-stable.
|
|
301
|
+
function addedRequirementNames(repo, change) {
|
|
302
|
+
const specsRoot = path.join(repo, "openspec", "changes", change, "specs");
|
|
303
|
+
const names = new Set();
|
|
304
|
+
let capabilities = [];
|
|
305
|
+
try {
|
|
306
|
+
capabilities = fs.readdirSync(specsRoot);
|
|
307
|
+
} catch {
|
|
308
|
+
return names;
|
|
309
|
+
}
|
|
310
|
+
for (const capability of capabilities) {
|
|
311
|
+
const specPath = path.join(specsRoot, capability, "spec.md");
|
|
312
|
+
let content = "";
|
|
313
|
+
try {
|
|
314
|
+
content = fs.readFileSync(specPath, "utf8");
|
|
315
|
+
} catch {
|
|
316
|
+
continue;
|
|
317
|
+
}
|
|
318
|
+
// Only the ADDED section, bounded by the next `## ` heading, so a
|
|
319
|
+
// requirement listed under MODIFIED or REMOVED is not read as new.
|
|
320
|
+
for (const section of content.split(/^##\s+/m).slice(1)) {
|
|
321
|
+
if (!/^ADDED Requirements\s*$/m.test(section.split(/\r?\n/)[0])) continue;
|
|
322
|
+
for (const match of section.matchAll(/^###\s+Requirement:\s*(.+?)\s*$/gm)) {
|
|
323
|
+
names.add(match[1]);
|
|
324
|
+
}
|
|
325
|
+
}
|
|
326
|
+
}
|
|
327
|
+
return names;
|
|
328
|
+
}
|
|
329
|
+
|
|
330
|
+
// A task's declared strategy, from the text it wrote. Both forms: the compact
|
|
331
|
+
// `Strategy:` entry under `Verify`, and the expanded `Verification Strategy`
|
|
332
|
+
// field beside `Commands`.
|
|
333
|
+
function declaredStrategy(task) {
|
|
334
|
+
const compact = String(field(task, "Verify") || "").match(
|
|
335
|
+
/^\s*-?\s*Strategy:\s*(.+?)\s*$/m
|
|
336
|
+
);
|
|
337
|
+
const value = compact
|
|
338
|
+
? compact[1]
|
|
339
|
+
: String(field(task, "Verification Strategy") || "");
|
|
340
|
+
return value.trim().toLowerCase();
|
|
341
|
+
}
|
|
342
|
+
|
|
343
|
+
// A `Covers` entry naming a spec scenario: `capability / Requirement / Scenario`.
|
|
344
|
+
// Identifier entries (`D4`, `F1`, `E2`) have no slashes and are not spec claims.
|
|
345
|
+
function specCoverEntries(task) {
|
|
346
|
+
return String(field(task, "Covers") || "")
|
|
347
|
+
.split(/\r?\n/)
|
|
348
|
+
.map((line) => line.replace(/^\s*-\s*/, "").trim())
|
|
349
|
+
.filter((entry) => entry.split("/").length >= 3)
|
|
350
|
+
.map((entry) => entry.split("/").map((part) => part.trim()));
|
|
351
|
+
}
|
|
352
|
+
|
|
353
|
+
// `equivalence` owes no red, which makes it the first strategy reached for by a
|
|
354
|
+
// task that should have one. A task covering a scenario its own change *adds* is
|
|
355
|
+
// claiming that behavior is new and that behavior is unchanged at the same time,
|
|
356
|
+
// and without this guard one of the two claims has no proof anywhere in the
|
|
357
|
+
// change. Satisfied only by a sibling covering the same entry, never by the mere
|
|
358
|
+
// presence of a red-green task somewhere in the change.
|
|
359
|
+
function equivalenceEscapeProblems(repo, selection, task, compiled) {
|
|
360
|
+
const strategy = String(
|
|
361
|
+
((compiled.capsule || {}).verification || {}).strategy || ""
|
|
362
|
+
).toLowerCase();
|
|
363
|
+
if (strategy !== "equivalence") return [];
|
|
364
|
+
const added = addedRequirementNames(repo, selection.change);
|
|
365
|
+
if (added.size === 0) return [];
|
|
366
|
+
const problems = [];
|
|
367
|
+
for (const parts of specCoverEntries(task)) {
|
|
368
|
+
const requirement = parts[1];
|
|
369
|
+
if (!added.has(requirement)) continue;
|
|
370
|
+
const entry = parts.join(" / ");
|
|
371
|
+
const scenario = parts.slice(2).join(" / ");
|
|
372
|
+
const covered = selection.tasks.some((sibling) => {
|
|
373
|
+
if (sibling.id === task.id) return false;
|
|
374
|
+
// Read from the sibling's own `Verify` text rather than by compiling it.
|
|
375
|
+
// Compiling made the guard depend on the sibling being otherwise valid: a
|
|
376
|
+
// sibling with any unrelated contract error produced no capsule, so its
|
|
377
|
+
// strategy read as absent and it silently stopped satisfying the guard.
|
|
378
|
+
const siblingStrategy = declaredStrategy(sibling);
|
|
379
|
+
if (!RED_GREEN_VERIFICATION_STRATEGIES.has(siblingStrategy)) return false;
|
|
380
|
+
// The same entry, not any entry. A guard satisfied by the presence of a
|
|
381
|
+
// red-green task would be a check that the change contains one, which
|
|
382
|
+
// every change with more than one task passes.
|
|
383
|
+
return specCoverEntries(sibling).some(
|
|
384
|
+
(other) => other.join(" / ") === entry
|
|
385
|
+
);
|
|
386
|
+
});
|
|
387
|
+
if (!covered) {
|
|
388
|
+
problems.push(
|
|
389
|
+
problem(
|
|
390
|
+
"equivalence-covers-added-behavior",
|
|
391
|
+
`This task declares \`Strategy: equivalence\` and covers `
|
|
392
|
+
+ `"${scenario}", a scenario this change adds. Behavior that is new `
|
|
393
|
+
+ "is not behavior that is unchanged, and no task of this change "
|
|
394
|
+
+ "proves the new half: name a red-green task that covers the same "
|
|
395
|
+
+ "entry, or move this Covers entry to the task that implements it."
|
|
396
|
+
)
|
|
397
|
+
);
|
|
398
|
+
}
|
|
399
|
+
}
|
|
400
|
+
return problems;
|
|
401
|
+
}
|
|
402
|
+
|
|
296
403
|
function taskStart(repo, options) {
|
|
297
404
|
const selection = loadSelection(repo, options);
|
|
298
405
|
const task = selection.selected[0];
|
|
@@ -300,6 +407,7 @@ function taskStart(repo, options) {
|
|
|
300
407
|
const problems = [
|
|
301
408
|
...compiled.diagnostics,
|
|
302
409
|
...invalidationProblems(repo, selection.content, selection.tasks, selection.change),
|
|
410
|
+
...equivalenceEscapeProblems(repo, selection, task, compiled),
|
|
303
411
|
];
|
|
304
412
|
// Recording the current fingerprint is idempotent: --record replaces the
|
|
305
413
|
// selected task's Contract anchor whatever it holds, so reauthorizing a task
|
|
@@ -454,6 +562,72 @@ function commandLabels(task) {
|
|
|
454
562
|
].map((match) => match[1]);
|
|
455
563
|
}
|
|
456
564
|
|
|
565
|
+
// `artifact <path> sha256:<digest>`, or null for prose. The path is
|
|
566
|
+
// repo-relative; the digest is what Review is entitled to assume it is reading.
|
|
567
|
+
const ARTIFACT_EVIDENCE =
|
|
568
|
+
/^artifact\s+(\S+)\s+sha256:([0-9a-f]{64})\s*$/i;
|
|
569
|
+
|
|
570
|
+
function artifactReference(value) {
|
|
571
|
+
const match = String(value || "").trim().match(ARTIFACT_EVIDENCE);
|
|
572
|
+
return match ? { path: match[1], digest: match[2].toLowerCase() } : null;
|
|
573
|
+
}
|
|
574
|
+
|
|
575
|
+
// Keel checks identity and reads nothing else: it does not parse the artifact,
|
|
576
|
+
// does not know what a field is, and compares nothing in it. The claim that the
|
|
577
|
+
// numbers agree stays the author's, recorded before Review exactly as
|
|
578
|
+
// `Fails with:` and `Detects:` are. What the digest buys is that the file Review
|
|
579
|
+
// opens is the file the author meant, which a retelling cannot offer.
|
|
580
|
+
function artifactProblems(repo, change, label, reference) {
|
|
581
|
+
if (!reference) return [];
|
|
582
|
+
// The inverse of the `Durable owner:` rule, and for the opposite reason: a
|
|
583
|
+
// follow-up pointer has to outlive the change, while an evidence artifact has
|
|
584
|
+
// to travel with it. `openspec archive` moves the change directory and
|
|
585
|
+
// nothing else, so a path outside it is one the archive is guaranteed to
|
|
586
|
+
// leave behind — and an evidence pointer that breaks on archive is worse than
|
|
587
|
+
// a retelling, which at least survives.
|
|
588
|
+
const changeDir = path.join("openspec", "changes", change);
|
|
589
|
+
const normalized = reference.path.split(path.sep).join("/");
|
|
590
|
+
if (!normalized.startsWith(`${changeDir.split(path.sep).join("/")}/`)) {
|
|
591
|
+
return [
|
|
592
|
+
problem(
|
|
593
|
+
"artifact-outside-change",
|
|
594
|
+
`${label} Evidence references \`${reference.path}\`, which is outside `
|
|
595
|
+
+ `\`${changeDir}\`. Archiving moves the change directory and nothing `
|
|
596
|
+
+ "else, so this pointer breaks the moment the change is archived. "
|
|
597
|
+
+ "Put the artifact inside the change's own directory."
|
|
598
|
+
),
|
|
599
|
+
];
|
|
600
|
+
}
|
|
601
|
+
const absolute = path.join(repo, reference.path);
|
|
602
|
+
if (!fs.existsSync(absolute) || !fs.statSync(absolute).isFile()) {
|
|
603
|
+
return [
|
|
604
|
+
problem(
|
|
605
|
+
"artifact-missing",
|
|
606
|
+
`${label} Evidence references the artifact \`${reference.path}\`, and `
|
|
607
|
+
+ "no file is there. An unresolvable reference is worse than a "
|
|
608
|
+
+ "retelling: the retelling at least carries the result."
|
|
609
|
+
),
|
|
610
|
+
];
|
|
611
|
+
}
|
|
612
|
+
const actual = crypto
|
|
613
|
+
.createHash("sha256")
|
|
614
|
+
.update(fs.readFileSync(absolute))
|
|
615
|
+
.digest("hex");
|
|
616
|
+
if (actual !== reference.digest) {
|
|
617
|
+
return [
|
|
618
|
+
problem(
|
|
619
|
+
"artifact-digest-mismatch",
|
|
620
|
+
`${label} Evidence records \`${reference.path}\` at `
|
|
621
|
+
+ `sha256:${reference.digest}, and the file there hashes to `
|
|
622
|
+
+ `sha256:${actual}. Re-record the digest if the command was re-run, `
|
|
623
|
+
+ "or correct the path — naming both is what lets a reader tell a "
|
|
624
|
+
+ "stale record from a pointer at the wrong file."
|
|
625
|
+
),
|
|
626
|
+
];
|
|
627
|
+
}
|
|
628
|
+
return [];
|
|
629
|
+
}
|
|
630
|
+
|
|
457
631
|
function evidenceValue(task, label) {
|
|
458
632
|
const escaped = label.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
459
633
|
const match = field(task, "Evidence").match(
|
|
@@ -1049,11 +1223,21 @@ function completionChecks(repo, task, contract = null, changeVerify = null, chan
|
|
|
1049
1223
|
problems.push(problem("missing-commands", "Commands must define at least one M<n>."));
|
|
1050
1224
|
}
|
|
1051
1225
|
for (const label of commands) {
|
|
1052
|
-
|
|
1226
|
+
const recorded = evidenceValue(task, label);
|
|
1227
|
+
if (!isConcrete(recorded)) {
|
|
1053
1228
|
problems.push(
|
|
1054
1229
|
problem("missing-evidence", `Missing concrete Evidence for ${label}.`)
|
|
1055
1230
|
);
|
|
1231
|
+
continue;
|
|
1056
1232
|
}
|
|
1233
|
+
// An A/B pairing or a sweep summary *is* the output of one command, and
|
|
1234
|
+
// retelling it into tasks.md is a transcription that can be wrong and that
|
|
1235
|
+
// nobody can re-check (issue #142). A reference is checked rather than
|
|
1236
|
+
// tolerated: as prose it would already have passed as concrete, so without
|
|
1237
|
+
// this the form would buy nothing at all.
|
|
1238
|
+
problems.push(
|
|
1239
|
+
...artifactProblems(repo, change, label, artifactReference(recorded))
|
|
1240
|
+
);
|
|
1057
1241
|
}
|
|
1058
1242
|
// A declared measurement is held against the check's own bare `M<n>` Evidence
|
|
1059
1243
|
// — the entry where the command and its output are recorded. A literal that
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
"use strict";
|
|
2
2
|
|
|
3
3
|
const crypto = require("crypto");
|
|
4
|
+
const { execFileSync } = require("child_process");
|
|
4
5
|
const fs = require("fs");
|
|
5
6
|
const path = require("path");
|
|
6
7
|
|
|
@@ -143,6 +144,13 @@ const SUPPORTED_VERIFICATION_STRATEGIES = [
|
|
|
143
144
|
"snapshot-characterization",
|
|
144
145
|
"rendered-behavior",
|
|
145
146
|
"evidence-first",
|
|
147
|
+
// Zero difference, not "nothing could fail first". `evidence-first` is the
|
|
148
|
+
// only other strategy without a red-green obligation, and it is scoped by an
|
|
149
|
+
// absence; an A/B against a base is the opposite — a criterion stronger than
|
|
150
|
+
// red-green, because it also catches the change that incidentally moved a
|
|
151
|
+
// result. Issue #142 measured a task re-recording its contract twice to get
|
|
152
|
+
// past the shape rather than the criterion.
|
|
153
|
+
"equivalence",
|
|
146
154
|
];
|
|
147
155
|
|
|
148
156
|
const RED_GREEN_VERIFICATION_STRATEGIES = new Set([
|
|
@@ -257,6 +265,77 @@ function isPassingReviewStatus(value) {
|
|
|
257
265
|
);
|
|
258
266
|
}
|
|
259
267
|
|
|
268
|
+
// A git ref resolved locally, or null. Reads the repository the gate is already
|
|
269
|
+
// reading and reaches nothing else: a gate that fetched would stop being local
|
|
270
|
+
// and offline, which is the property its verdict rests on.
|
|
271
|
+
function resolveCommit(repo, ref) {
|
|
272
|
+
try {
|
|
273
|
+
return execFileSync(
|
|
274
|
+
"git",
|
|
275
|
+
["-C", repo, "rev-parse", "--verify", "--quiet", `${ref}^{commit}`],
|
|
276
|
+
{ encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }
|
|
277
|
+
).trim() || null;
|
|
278
|
+
} catch {
|
|
279
|
+
return null;
|
|
280
|
+
}
|
|
281
|
+
}
|
|
282
|
+
|
|
283
|
+
// Four ways an `equivalence` task's shape can be complete and still compare
|
|
284
|
+
// nothing. Each names the declaration it is about: a diagnostic naming the
|
|
285
|
+
// strategy would send the author to the one line that is correct.
|
|
286
|
+
function equivalenceProblems(repo, taskVerification) {
|
|
287
|
+
const problems = [];
|
|
288
|
+
const declaredFields = taskVerification.fieldsDeclared;
|
|
289
|
+
if (!isConcrete(taskVerification.base)) {
|
|
290
|
+
problems.push({
|
|
291
|
+
code: "missing-equivalence-base",
|
|
292
|
+
message:
|
|
293
|
+
"equivalence compares one code path at two commits and declares which "
|
|
294
|
+
+ "one it is compared against. Add a `Base:` entry beside `Strategy:` "
|
|
295
|
+
+ "naming a git ref — without it the criterion is `the numbers are the "
|
|
296
|
+
+ "same as some other numbers`.",
|
|
297
|
+
});
|
|
298
|
+
}
|
|
299
|
+
if (taskVerification.fields.length === 0) {
|
|
300
|
+
problems.push({
|
|
301
|
+
code: "missing-equivalence-fields",
|
|
302
|
+
message: declaredFields
|
|
303
|
+
? "`Fields:` is declared and resolves to an empty set, so the "
|
|
304
|
+
+ "comparison has nothing to compare. It reads as a declaration to "
|
|
305
|
+
+ "every reader except the comparison; name the fields, separated by "
|
|
306
|
+
+ "commas."
|
|
307
|
+
: "equivalence declares which fields are compared. Add a `Fields:` "
|
|
308
|
+
+ "entry beside `Strategy:` listing them, separated by commas — a "
|
|
309
|
+
+ "comparison with no field set agrees with everything.",
|
|
310
|
+
});
|
|
311
|
+
}
|
|
312
|
+
if (isConcrete(taskVerification.base)) {
|
|
313
|
+
const base = resolveCommit(repo, taskVerification.base);
|
|
314
|
+
if (!base) {
|
|
315
|
+
problems.push({
|
|
316
|
+
code: "unresolvable-equivalence-base",
|
|
317
|
+
message:
|
|
318
|
+
`\`Base: ${taskVerification.base}\` resolves to no commit in this `
|
|
319
|
+
+ "repository. The base is read locally and never fetched, so a ref "
|
|
320
|
+
+ "that exists only on a remote is not one this gate can see.",
|
|
321
|
+
});
|
|
322
|
+
} else {
|
|
323
|
+
const head = resolveCommit(repo, "HEAD");
|
|
324
|
+
if (head && head === base) {
|
|
325
|
+
problems.push({
|
|
326
|
+
code: "equivalence-base-is-head",
|
|
327
|
+
message:
|
|
328
|
+
`\`Base: ${taskVerification.base}\` resolves to ${base}, which is `
|
|
329
|
+
+ "HEAD. An A/B against itself always agrees, so the check would "
|
|
330
|
+
+ "pass having compared nothing — the one shape here that is "
|
|
331
|
+
+ "complete, resolvable, and still empty.",
|
|
332
|
+
});
|
|
333
|
+
}
|
|
334
|
+
}
|
|
335
|
+
}
|
|
336
|
+
return problems;
|
|
337
|
+
}
|
|
338
|
+
|
|
260
339
|
function verification(task) {
|
|
261
340
|
const compact = fieldValues(task, "Verify");
|
|
262
341
|
const strategyEntry = compact.find((entry) => /^Strategy:\s*/i.test(entry));
|
|
@@ -266,8 +345,19 @@ function verification(task) {
|
|
|
266
345
|
// command, and a reason that took an `M<n>` label would be a check the author
|
|
267
346
|
// never wrote and evidence nobody can record.
|
|
268
347
|
const reasonEntry = compact.find((entry) => /^Reason:\s*/i.test(entry));
|
|
348
|
+
// `equivalence` compares one code path at two commits. Neither half of that
|
|
349
|
+
// fits in a check: the check is the command, and what it cannot say by itself
|
|
350
|
+
// is which commit it is compared against and which fields are compared. #142
|
|
351
|
+
// proposed a third field for the command too; commands already have exactly
|
|
352
|
+
// one home here, and a second would put half of them outside the labelled
|
|
353
|
+
// evidence `task-complete` enforces.
|
|
354
|
+
const baseEntry = compact.find((entry) => /^Base:\s*/i.test(entry));
|
|
355
|
+
const fieldsEntry = compact.find((entry) => /^Fields:\s*/i.test(entry));
|
|
269
356
|
const isVerificationField = (entry) =>
|
|
270
|
-
/^Strategy:\s*/i.test(entry)
|
|
357
|
+
/^Strategy:\s*/i.test(entry)
|
|
358
|
+
|| /^Reason:\s*/i.test(entry)
|
|
359
|
+
|| /^Base:\s*/i.test(entry)
|
|
360
|
+
|| /^Fields:\s*/i.test(entry);
|
|
271
361
|
const commandSource = compact.length > 0
|
|
272
362
|
? compact.filter((entry) => !isVerificationField(entry))
|
|
273
363
|
: fieldValues(task, "Commands");
|
|
@@ -342,6 +432,24 @@ function verification(task) {
|
|
|
342
432
|
? reasonEntry.replace(/^Reason:\s*/i, "")
|
|
343
433
|
: field(task, "Verification Reason")
|
|
344
434
|
),
|
|
435
|
+
base: normalizeText(
|
|
436
|
+
baseEntry ? baseEntry.replace(/^Base:\s*/i, "") : field(task, "Verification Base")
|
|
437
|
+
),
|
|
438
|
+
// A set, so "declared but empty" is a state the gate can see. `Fields:` with
|
|
439
|
+
// nothing behind it is the shape that passes while comparing nothing, and it
|
|
440
|
+
// reads as a declaration to everyone except the comparison.
|
|
441
|
+
// Whether the line was written at all, kept beside the parsed set so a
|
|
442
|
+
// refusal can tell an author who wrote nothing from one who wrote an empty
|
|
443
|
+
// set. To the comparison they are the same state; to the author they are
|
|
444
|
+
// opposite mistakes.
|
|
445
|
+
fieldsDeclared: Boolean(fieldsEntry || field(task, "Verification Fields")),
|
|
446
|
+
fields: (fieldsEntry
|
|
447
|
+
? fieldsEntry.replace(/^Fields:\s*/i, "")
|
|
448
|
+
: field(task, "Verification Fields") || ""
|
|
449
|
+
)
|
|
450
|
+
.split(",")
|
|
451
|
+
.map((entry) => normalizeText(entry))
|
|
452
|
+
.filter(Boolean),
|
|
345
453
|
commands,
|
|
346
454
|
};
|
|
347
455
|
}
|
|
@@ -1191,6 +1299,11 @@ function compileTaskContract(repo, change, task) {
|
|
|
1191
1299
|
+ "needs no reason.",
|
|
1192
1300
|
});
|
|
1193
1301
|
}
|
|
1302
|
+
if (taskVerification.strategy.toLowerCase() === "equivalence") {
|
|
1303
|
+
resolved.diagnostics.push(
|
|
1304
|
+
...equivalenceProblems(repo, taskVerification)
|
|
1305
|
+
);
|
|
1306
|
+
}
|
|
1194
1307
|
const couplingMode = normalizeText(field(task, "Coupling")).toLowerCase()
|
|
1195
1308
|
|| "none";
|
|
1196
1309
|
const candidateBoundary = normalizedValues(task, "Candidate Boundary", {
|
|
@@ -1311,6 +1424,14 @@ function compileTaskContract(repo, change, task) {
|
|
|
1311
1424
|
// other task keeps the capsule shape and fingerprint it had before the
|
|
1312
1425
|
// field existed.
|
|
1313
1426
|
...(taskVerification.reason ? { reason: taskVerification.reason } : {}),
|
|
1427
|
+
// Same rule: emitted only by the strategy that declares them, so every
|
|
1428
|
+
// existing task's capsule shape and fingerprint are untouched. They belong
|
|
1429
|
+
// in the capsule rather than only in the file because a declaration
|
|
1430
|
+
// outside the fingerprint could be edited after the run it describes.
|
|
1431
|
+
...(taskVerification.base ? { base: taskVerification.base } : {}),
|
|
1432
|
+
...(taskVerification.fields.length > 0
|
|
1433
|
+
? { fields: taskVerification.fields }
|
|
1434
|
+
: {}),
|
|
1314
1435
|
// Emit a tag only when the check opts out of a default, so an untagged
|
|
1315
1436
|
// check keeps the capsule shape and fingerprint it had before either tag
|
|
1316
1437
|
// existed. `layer` appears only for `fast`, `regression` only when true.
|