@christang/keel 5.67.0 → 5.69.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -344,6 +344,41 @@ wiring, and `keel --init` whenever it tells you the repository is behind its ins
344
344
  protocol version lives in your `AGENTS.md`, and updating the package does not move it.
345
345
  Everything below is the vocabulary the agent uses on your behalf.
346
346
 
347
+ ### When the right answer is "nothing changed"
348
+
349
+ A refactor, a move, a flow upgrade that claims the numbers hold — the correct evidence for these is
350
+ *zero difference*, and red-green has no shape for it. The repository that reported this had a task
351
+ whose `tasks.md` said "this one has no honest red" several times, and re-recorded its contract twice
352
+ trying to fit. The criterion was right; it had nowhere to live.
353
+
354
+ ```
355
+ - Verify:
356
+ - Strategy: equivalence
357
+ - Base: origin/main
358
+ - Fields: wns, tns, cell_count
359
+ - M1: node compare.js --base --head reports every field equal
360
+ ```
361
+
362
+ `equivalence` owes no red. Its criterion is that base and head agree on the fields you named, which is
363
+ *stronger* than red-green — it also catches the change that incidentally moved a result. What the gate
364
+ checks is every way that shape can look complete and compare nothing: a missing `Base:` or `Fields:`,
365
+ an empty field set, a ref that resolves to nothing, and a `Base:` that resolves to HEAD.
366
+
367
+ It is not a way out of red-green. A task declaring `equivalence` while covering a scenario its own
368
+ change *adds* is refused unless a sibling task covers that same entry under a red-green strategy —
369
+ behavior that is new is not behavior that is unchanged, and a task cannot prove both.
370
+
371
+ And Evidence no longer has to retell the output:
372
+
373
+ ```
374
+ - M1: artifact openspec/changes/<change>/evidence/compare.json sha256:9f2c…
375
+ ```
376
+
377
+ The gate checks the file is there and the digest matches, and refuses a path outside the change's own
378
+ directory, because archiving moves that directory and the pointer would break. Keel hashes the bytes
379
+ and reads nothing inside them — the claim stays yours, and what the digest buys is that the file your
380
+ reviewer opens is the file you meant.
381
+
347
382
  ## Verification layering
348
383
 
349
384
  Keel splits verification into two layers so a slow suite never blocks your push:
@@ -1,4 +1,4 @@
1
- <!-- keel:start version=5.67.0 -->
1
+ <!-- keel:start version=5.69.0 -->
2
2
  ## Keel Bootstrap
3
3
 
4
4
  - Start every session with `keel context`; OpenSpec artifacts and Git are the only durable authority — never native memory, goals, or transcripts.
package/package.json CHANGED
@@ -2,7 +2,7 @@
2
2
  "name": "@christang/keel",
3
3
  "displayName": "Keel",
4
4
  "description": "Keel OpenSpec execution discipline CLI for Claude Code, Codex, and OpenCode.",
5
- "version": "5.67.0",
5
+ "version": "5.69.0",
6
6
  "license": "MIT",
7
7
  "repository": {
8
8
  "type": "git",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "keel",
3
- "version": "5.67.0",
3
+ "version": "5.69.0",
4
4
  "description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
5
5
  "author": {
6
6
  "name": "TanglmChris",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "keel",
3
- "version": "5.67.0",
3
+ "version": "5.69.0",
4
4
  "description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
5
5
  "author": {
6
6
  "name": "TanglmChris",
@@ -20,10 +20,13 @@ Read the selected task's compiled capsule: resolved Acceptance, Verify strategy
20
20
  - `regression-first`: an observable defect; reproduce it through the public interface first, then prove the fix with the same check.
21
21
  - `characterization` / `snapshot-characterization`: deterministic or generated outputs kept stable by byte or snapshot comparison; not downgraded to build success.
22
22
  - `rendered-behavior`: interactive surfaces exercised through the real rendered interface; strict red-green optional by cost.
23
- - `evidence-first`: docs, configuration, or diagnosis work whose checks state the observable artifact or evidence instead of a red-green loop.
23
+ - `equivalence`: work whose criterion is that a measured result does not change — a refactor, a move, a flow upgrade claiming the numbers hold. It declares `Base:` (a resolvable git ref) and `Fields:` (the compared field set) beside `Strategy:` and owes no red; the A/B command is an ordinary `M<n>` check. `keel gate task-start` refuses a missing `Base:` or `Fields:`, an empty field set, an unresolvable ref, and a `Base:` that is HEAD, because an A/B against itself always agrees. A task covering a scenario its own change *adds* is refused unless a sibling task covers that same entry under a red-green strategy: behavior that is new is not behavior that is unchanged.
24
+ - `evidence-first`: docs, configuration, or diagnosis work whose checks state the observable artifact or evidence instead of a red-green loop. It is scoped by an absence — nothing here can fail first — which is why it is not the home for `equivalence` work, whose criterion is stronger than red-green rather than missing.
24
25
 
25
26
  Red-green strategies (`vertical-tdd`, `regression-first`) must record concrete per-label `.red` and `.green` Evidence entries for the same check; `keel gate task-complete` rejects absent or pending entries.
26
27
 
28
+ A check's Evidence may read `artifact <path> sha256:<digest>` instead of retelling the output. `keel gate task-complete` checks the file exists and the digest matches, and refuses a path outside the change's own directory — archiving moves that directory, so a pointer outside it breaks. Keel hashes the bytes and reads nothing inside them: the claim the artifact supports stays yours.
29
+
27
30
  ## Domain lenses
28
31
 
29
32
  When the change's proposal/design/specs or the task's Touch extensions signal a domain, consult the matching lens's `Execution and review checks` section from `keel/lenses/` — the lens whose `Applies when:` header matches — before finalizing the strategy and the first check, and load only that one. When no lens matches, load nothing.
@@ -38,8 +38,8 @@ REQUIRED_SCRIPTS = [
38
38
  "scripts/validate_plugin.py",
39
39
  ]
40
40
 
41
- PACKAGE_VERSION = "5.67.0"
42
- PROTOCOL_VERSION = "5.67.0"
41
+ PACKAGE_VERSION = "5.69.0"
42
+ PROTOCOL_VERSION = "5.69.0"
43
43
  LEGACY_MANAGED_START = "<!-- keel:start version=2.1 -->"
44
44
  OPENSPEC_SCHEMA_NAME = "keel-spec-driven"
45
45
  # Mirrors KEEL_PACKAGE_NAME in scripts/install_to_repo.py, one of the two
@@ -2302,6 +2302,61 @@ def validate_domain_lenses_scenario() -> int:
2302
2302
  return 0
2303
2303
 
2304
2304
 
2305
+ # `bump_version.js` writes a stub section for the new release, and the convention
2306
+ # is to fill it in. Writing the entry *above* it instead leaves an orphan, and
2307
+ # the stub carries the `Version alignment:` line this scenario reads — so the
2308
+ # section describing nothing was the section that passed (#151). It happened in
2309
+ # 5.67.0 and 5.68.0, with the suite green both times.
2310
+ #
2311
+ # Takes the text rather than reading the file, so the rule can be exercised on a
2312
+ # planted copy: a rule that could only ever see the changelog already known to be
2313
+ # correct would pass forever without anyone learning whether it fires.
2314
+ def release_description_problem(changelog: str, version: str) -> str | None:
2315
+ """Return the problem the released version's section has, or None.
2316
+
2317
+ Scoped to the version being released. An older section is history a current
2318
+ author cannot act on, and a rule that swept the whole file would fail on the
2319
+ archive — which is how a check gets disabled rather than fixed. The file also
2320
+ legitimately quotes `TODO` when describing the rules that refuse it.
2321
+ """
2322
+ pattern = re.compile(
2323
+ r"^## " + re.escape(version) + r"\b[^\n]*$", re.M
2324
+ )
2325
+ headings = pattern.findall(changelog)
2326
+ if not headings:
2327
+ return (
2328
+ f"the changelog carries no `## {version}` section, so the release "
2329
+ "describes nothing at all."
2330
+ )
2331
+ if len(headings) > 1:
2332
+ return (
2333
+ f"the changelog describes {version} twice — {len(headings)} `## "
2334
+ f"{version}` headings: {', '.join(h.strip() for h in headings)}. A "
2335
+ "release entry written above the stub instead of into it leaves both."
2336
+ )
2337
+ for heading in headings:
2338
+ if "TODO" in heading:
2339
+ return (
2340
+ f"the changelog still carries the stub heading `{heading.strip()}`"
2341
+ f", so the {version} release was never described. Fill that "
2342
+ "section in rather than writing the entry above it."
2343
+ )
2344
+ # The section body, bounded by the next `## ` heading, so a `TODO` belonging
2345
+ # to an older release is not read as this one's.
2346
+ start = changelog.index(headings[0])
2347
+ rest = changelog[start + len(headings[0]):]
2348
+ end = rest.find("\n## ")
2349
+ body = rest if end < 0 else rest[:end]
2350
+ for line in body.splitlines():
2351
+ if re.match(r"^\s*-\s*TODO\b", line):
2352
+ return (
2353
+ f"the {version} section still carries the stub bullet "
2354
+ f"`{line.strip().lstrip('- ')}`, so the release is titled but "
2355
+ "not described."
2356
+ )
2357
+ return None
2358
+
2359
+
2305
2360
  def validate_version_alignment_scenario() -> int:
2306
2361
  package = json.loads((ROOT / "package.json").read_text(encoding="utf-8"))
2307
2362
  if package.get("version") != PACKAGE_VERSION:
@@ -2366,6 +2421,103 @@ def validate_version_alignment_scenario() -> int:
2366
2421
  )
2367
2422
  return 1
2368
2423
 
2424
+ # The released version's section has to say what the release did. Asserted on
2425
+ # planted copies first, so the rule is known to fire, and then on the
2426
+ # repository's own changelog.
2427
+ stub = (
2428
+ f"# Keel Changelog\n\n## {PACKAGE_VERSION} - TODO: summarize this "
2429
+ f"release\n\n- TODO: describe the change.\n- Version alignment: "
2430
+ f"shares Keel {PACKAGE_VERSION}.\n\n## 5.0.0 - older\n\n- Done.\n"
2431
+ )
2432
+ planted = release_description_problem(stub, PACKAGE_VERSION)
2433
+ if not planted:
2434
+ report(
2435
+ "version-alignment scenario: an unfilled stub was accepted — a "
2436
+ f"changelog whose {PACKAGE_VERSION} section is the stub "
2437
+ "`bump_version.js` writes produced no problem, so a release can "
2438
+ "ship describing nothing while this scenario passes."
2439
+ )
2440
+ return 1
2441
+ if "TODO: summarize this release" not in planted:
2442
+ report(
2443
+ "version-alignment scenario: the refusal does not name the offending "
2444
+ f"line, so an author is sent to read the whole changelog; got "
2445
+ f"{planted!r}."
2446
+ )
2447
+ return 1
2448
+
2449
+ # Still M1 — the stub survives renaming its heading. The bullet is the other
2450
+ # half of what `bump_version.js` writes, and a section titled correctly while
2451
+ # its only content is `- TODO: describe the change.` describes nothing.
2452
+ renamed = (
2453
+ f"# Keel Changelog\n\n## {PACKAGE_VERSION} - a real sounding title\n\n"
2454
+ f"- TODO: describe the change.\n- Version alignment: shares Keel "
2455
+ f"{PACKAGE_VERSION}.\n"
2456
+ )
2457
+ bullet = release_description_problem(renamed, PACKAGE_VERSION)
2458
+ if not bullet:
2459
+ report(
2460
+ "version-alignment scenario: an unfilled stub was accepted — a "
2461
+ "section with a real heading whose content is still `- TODO: describe "
2462
+ "the change.` produced no problem, so renaming the stub is enough to "
2463
+ "ship an undescribed release."
2464
+ )
2465
+ return 1
2466
+ if "TODO: describe the change." not in bullet:
2467
+ report(
2468
+ "version-alignment scenario: the bullet refusal does not name the "
2469
+ f"line it found; got {bullet!r}."
2470
+ )
2471
+ return 1
2472
+
2473
+ # M2 — two sections for one version. This is the shape the real defect had,
2474
+ # and neither TODO rule catches it once the stub's heading has been renamed:
2475
+ # the reader then sees two identical version headings, one of which describes
2476
+ # nothing.
2477
+ doubled = (
2478
+ f"# Keel Changelog\n\n## {PACKAGE_VERSION} - the real entry\n\n"
2479
+ f"- Did the thing.\n- Version alignment: shares Keel {PACKAGE_VERSION}."
2480
+ f"\n\n## {PACKAGE_VERSION} - summarized\n\n- Also this.\n"
2481
+ )
2482
+ problem_found = release_description_problem(doubled, PACKAGE_VERSION)
2483
+ if not problem_found:
2484
+ report(
2485
+ "version-alignment scenario: a version described twice was accepted "
2486
+ f"— two `## {PACKAGE_VERSION}` sections produced no problem, which is "
2487
+ "the shape the defect this rule exists for actually had."
2488
+ )
2489
+ return 1
2490
+ if "twice" not in problem_found:
2491
+ report(
2492
+ "version-alignment scenario: the duplicate-section refusal does not "
2493
+ f"say the version is described twice; got {problem_found!r}."
2494
+ )
2495
+ return 1
2496
+
2497
+ # M3 — only the version being released is judged. An older section is history,
2498
+ # and this file legitimately quotes `TODO` when describing the rules that
2499
+ # refuse it, so a rule sweeping the whole file would fire on its own
2500
+ # documentation.
2501
+ older = (
2502
+ f"# Keel Changelog\n\n## {PACKAGE_VERSION} - the real entry\n\n"
2503
+ f"- Did the thing.\n- Version alignment: shares Keel {PACKAGE_VERSION}."
2504
+ "\n\n## 5.0.0 - TODO: summarize this release\n\n- TODO: describe it.\n"
2505
+ )
2506
+ if release_description_problem(older, PACKAGE_VERSION):
2507
+ report(
2508
+ "version-alignment scenario: an older section was refused — a `TODO` "
2509
+ "in a section for another version failed the current release, which "
2510
+ "is a check the archive can only be edited to satisfy."
2511
+ )
2512
+ return 1
2513
+ if release_description_problem(changelog, PACKAGE_VERSION):
2514
+ report(
2515
+ "version-alignment scenario: an older section was refused — the "
2516
+ "repository's own changelog does not satisfy the rule: "
2517
+ f"{release_description_problem(changelog, PACKAGE_VERSION)}"
2518
+ )
2519
+ return 1
2520
+
2369
2521
  with tempfile.TemporaryDirectory(prefix="keel-version-") as raw_tmp:
2370
2522
  repo = Path(raw_tmp) / "repo"
2371
2523
  repo.mkdir()
@@ -29042,6 +29194,509 @@ def validate_guidance_is_referenced_and_carries_no_criterion_scenario() -> int:
29042
29194
  return 0
29043
29195
 
29044
29196
 
29197
+ def equivalence_task(
29198
+ *,
29199
+ base: str | None = "HEAD~1",
29200
+ fields: str | None = "wns, tns, cell_count",
29201
+ covers: tuple[str, ...] = ("E1: the measured result does not move",),
29202
+ label: str = "1.1",
29203
+ title: str = "Move the attribute without moving the numbers",
29204
+ ) -> str:
29205
+ """One `equivalence` task, with either declaration omittable."""
29206
+ lines = [
29207
+ f"- [ ] {label} {title}",
29208
+ " - Covers:",
29209
+ ]
29210
+ lines.extend(f" - {entry}" for entry in covers)
29211
+ lines.extend(
29212
+ [
29213
+ " - Read:",
29214
+ " - README.md",
29215
+ " - Touch:",
29216
+ " - src/example.js",
29217
+ " - Verify:",
29218
+ " - Strategy: equivalence",
29219
+ ]
29220
+ )
29221
+ if base is not None:
29222
+ lines.append(f" - Base: {base}")
29223
+ if fields is not None:
29224
+ lines.append(f" - Fields: {fields}")
29225
+ lines.extend(
29226
+ [
29227
+ " - M1: node compare.js --base --head reports every field equal",
29228
+ " - Autonomy boundary:",
29229
+ " - Default: hard-stop",
29230
+ " - Pre-authorized fallback: none",
29231
+ " - Stop Rules:",
29232
+ " - Stop on any field that differs.",
29233
+ " - Evidence:",
29234
+ " - Contract: pending",
29235
+ " - M1: pending",
29236
+ " - Review:",
29237
+ " - Status: pass",
29238
+ " - Acceptance check: every declared field agreed.",
29239
+ " - Scope check: writes stayed inside Touch.",
29240
+ " - Findings: none",
29241
+ " - Blocker: none",
29242
+ ]
29243
+ )
29244
+ return "\n".join(lines) + "\n"
29245
+
29246
+
29247
+ def validate_an_equivalence_claim_names_its_base_scenario() -> int:
29248
+ """Issue #142: the correct evidence for a refactor is zero difference.
29249
+
29250
+ Red-green has no shape for it. The reporting repository re-recorded one
29251
+ task's contract twice to get past the shape — not because a criterion was
29252
+ wrong, but because the criterion had nowhere to live. An A/B against a base
29253
+ is *stronger* than red-green: it catches the change that also, incidentally,
29254
+ moved a result. What it needs is a place to say what it compares against and
29255
+ on which fields, plus a refusal for every way that shape can be complete and
29256
+ still compare nothing.
29257
+ """
29258
+ label = "an-equivalence-claim-names-its-base"
29259
+
29260
+ def git(repo: Path, *args: str) -> subprocess.CompletedProcess[str]:
29261
+ return subprocess.run(
29262
+ ["git", "-C", str(repo), *args], capture_output=True, text=True
29263
+ )
29264
+
29265
+ with tempfile.TemporaryDirectory(prefix="keel-equivalence-") as raw:
29266
+ root = Path(raw)
29267
+
29268
+ def fixture(name: str, task: str) -> Path:
29269
+ repo = (root / name).resolve()
29270
+ repo.mkdir()
29271
+ git(repo, "init", "-q")
29272
+ git(repo, "config", "user.email", "t@example.com")
29273
+ git(repo, "config", "user.name", "keel-test")
29274
+ write_gate_fixture(repo, tasks=task)
29275
+ write_text(repo / "src/example.js", "// product\n")
29276
+ git(repo, "add", "-A")
29277
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
29278
+ write_text(repo / "src/example.js", "// product, moved\n")
29279
+ git(repo, "add", "-A")
29280
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
29281
+ return repo
29282
+
29283
+ def start(name: str, **kwargs) -> dict:
29284
+ repo = fixture(name, equivalence_task(**kwargs))
29285
+ result = run_keel(
29286
+ repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
29287
+ "--no-guard", "--json",
29288
+ )
29289
+ try:
29290
+ return json.loads(result.stdout)
29291
+ except json.JSONDecodeError:
29292
+ return {
29293
+ "status": "unparsed",
29294
+ "problems": [{"message": result.stdout[:400]}],
29295
+ }
29296
+
29297
+ # M1 — the strategy is accepted, and it owes no red.
29298
+ payload = start("accepted")
29299
+ if payload.get("status") != "pass":
29300
+ report(
29301
+ f"{label}: unsupported verification strategy — a complete "
29302
+ "equivalence task was refused, so the one task class with the "
29303
+ "strongest criterion still has nowhere to declare it. "
29304
+ f"{problem_text(payload)}"
29305
+ )
29306
+ return 1
29307
+ capsule = (payload.get("contract") or {}).get("capsule") or {}
29308
+ verification = capsule.get("verification") or {}
29309
+ if verification.get("strategy", "").lower() != "equivalence":
29310
+ report(
29311
+ f"{label}: the compiled capsule does not carry the strategy; got "
29312
+ f"{verification.get('strategy')!r}."
29313
+ )
29314
+ return 1
29315
+ if not verification.get("base") or not verification.get("fields"):
29316
+ report(
29317
+ f"{label}: the capsule drops the declarations — base "
29318
+ f"{verification.get('base')!r}, fields "
29319
+ f"{verification.get('fields')!r}. A declaration absent from the "
29320
+ "capsule is absent from the fingerprint, so it could be edited "
29321
+ "after the run without moving the contract."
29322
+ )
29323
+ return 1
29324
+ warnings = " ".join(str(w) for w in (payload.get("warnings") or []))
29325
+ if ".red" in warnings or ".green" in warnings:
29326
+ report(
29327
+ f"{label}: equivalence was given a red-green obligation; the "
29328
+ f"criterion is agreement with a base, not a failing first run. "
29329
+ f"{warnings}"
29330
+ )
29331
+ return 1
29332
+
29333
+ # M2 — every way the shape can be complete and still compare nothing,
29334
+ # each named for the declaration it is about. A diagnostic that named
29335
+ # the strategy would send the author to the line that is correct.
29336
+ for name, kwargs, code, expected in (
29337
+ ("no-base", {"base": None}, "missing-equivalence-base", "Base:"),
29338
+ ("no-fields", {"fields": None}, "missing-equivalence-fields", "Fields:"),
29339
+ ("empty-fields", {"fields": " , "}, "missing-equivalence-fields", "Fields:"),
29340
+ (
29341
+ "bad-base",
29342
+ {"base": "no-such-ref"},
29343
+ "unresolvable-equivalence-base",
29344
+ "no-such-ref",
29345
+ ),
29346
+ ):
29347
+ payload = start(name, **kwargs)
29348
+ if payload.get("status") == "pass":
29349
+ report(
29350
+ f"{label}: accepted an equivalence task that compares "
29351
+ f"nothing — the {name} fixture passed, so the declaration "
29352
+ "is optional in practice."
29353
+ )
29354
+ return 1
29355
+ if code not in problem_codes(payload):
29356
+ report(
29357
+ f"{label}: accepted an equivalence task that compares "
29358
+ f"nothing — the {name} fixture was refused for another "
29359
+ f"reason; expected {code}, got {problem_codes(payload)!r}."
29360
+ )
29361
+ return 1
29362
+ message = problem_text(payload)
29363
+ if expected not in message:
29364
+ report(
29365
+ f"{label}: the {name} refusal does not name {expected!r}, so "
29366
+ f"the author is sent to find which line is wrong; got "
29367
+ f"{message!r}."
29368
+ )
29369
+ return 1
29370
+ # Absent and empty are the same state to the comparison and different
29371
+ # states to the author, so the two are asserted to read differently.
29372
+ absent = problem_text(start("no-fields-message", fields=None))
29373
+ empty = problem_text(start("empty-fields-message", fields=" , "))
29374
+ if absent == empty:
29375
+ report(
29376
+ f"{label}: a missing `Fields:` and one that resolves to an empty "
29377
+ "set produce the same sentence, so an author who wrote the line "
29378
+ "is told they did not."
29379
+ )
29380
+ return 1
29381
+
29382
+ # M1 of 1.2 — `equivalence` owes no red, which makes it the first thing
29383
+ # reached for by a task that should have one. A task covering a scenario
29384
+ # the change *adds* is claiming new behavior and unchanged behavior at
29385
+ # once, and one of the two claims has no proof anywhere.
29386
+ added_spec = (
29387
+ "## ADDED Requirements\n\n"
29388
+ "### Requirement: The moved attribute keeps its effect\n\n"
29389
+ "The attribute SHALL keep its effect after the move.\n\n"
29390
+ "#### Scenario: The effect survives the move\n\n"
29391
+ "- **WHEN** the attribute moves into the flow\n"
29392
+ "- **THEN** the effect is unchanged\n\n"
29393
+ # A second real scenario, so the mismatched-sibling control fails
29394
+ # because the guard refused it and not because its Covers entry
29395
+ # resolves to nothing. A first attempt pointed the sibling at an
29396
+ # invented scenario and passed for that unrelated reason.
29397
+ "#### Scenario: The flow reports the attribute\n\n"
29398
+ "- **WHEN** the flow runs\n"
29399
+ "- **THEN** it reports the attribute\n"
29400
+ )
29401
+ covered = (
29402
+ "demo / The moved attribute keeps its effect / The effect survives "
29403
+ "the move"
29404
+ )
29405
+
29406
+ def escape_fixture(name: str, *, sibling: str | None) -> dict:
29407
+ repo = (root / name).resolve()
29408
+ repo.mkdir()
29409
+ git(repo, "init", "-q")
29410
+ git(repo, "config", "user.email", "t@example.com")
29411
+ git(repo, "config", "user.name", "keel-test")
29412
+ tasks = equivalence_task(covers=(covered,))
29413
+ if sibling is not None:
29414
+ tasks += (
29415
+ "- [ ] 1.2 Add the behavior\n"
29416
+ " - Covers:\n"
29417
+ f" - {sibling}\n"
29418
+ " - Read:\n - README.md\n"
29419
+ " - Touch:\n - src/other.js\n"
29420
+ " - Verify:\n"
29421
+ " - Strategy: vertical-tdd\n"
29422
+ " - M1: node test.js asserts the new behavior\n"
29423
+ " - Autonomy boundary:\n"
29424
+ " - Default: hard-stop\n"
29425
+ " - Pre-authorized fallback: none\n"
29426
+ " - Stop Rules:\n - Stop on failure.\n"
29427
+ " - Evidence:\n - Contract: pending\n - M1: pending\n"
29428
+ " - Review:\n - Status: pass\n"
29429
+ " - Acceptance check: behavior asserted.\n"
29430
+ " - Scope check: inside Touch.\n"
29431
+ " - Findings: none\n"
29432
+ " - Blocker: none\n"
29433
+ )
29434
+ write_gate_fixture(repo, tasks=tasks)
29435
+ write_text(repo / "openspec/changes/demo/specs/demo/spec.md", added_spec)
29436
+ write_text(repo / "src/example.js", "// product\n")
29437
+ git(repo, "add", "-A")
29438
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
29439
+ write_text(repo / "src/example.js", "// moved\n")
29440
+ git(repo, "add", "-A")
29441
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
29442
+ result = run_keel(
29443
+ repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
29444
+ "--no-guard", "--json",
29445
+ )
29446
+ try:
29447
+ return json.loads(result.stdout)
29448
+ except json.JSONDecodeError:
29449
+ return {
29450
+ "status": "unparsed",
29451
+ "problems": [{"message": result.stdout[:400]}],
29452
+ }
29453
+
29454
+ payload = escape_fixture("escape-alone", sibling=None)
29455
+ if payload.get("status") == "pass":
29456
+ report(
29457
+ f"{label}: accepted new behavior with no red anywhere — an "
29458
+ "equivalence task covering a scenario the change adds passed, so "
29459
+ "the strategy is a way to author a feature with no red in the "
29460
+ "whole change."
29461
+ )
29462
+ return 1
29463
+ if "equivalence-covers-added-behavior" not in problem_codes(payload):
29464
+ report(
29465
+ f"{label}: accepted new behavior with no red anywhere — refused "
29466
+ f"for another reason; got {problem_codes(payload)!r}."
29467
+ )
29468
+ return 1
29469
+ message = problem_text(payload)
29470
+ if "The effect survives the move" not in message:
29471
+ report(
29472
+ f"{label}: the refusal does not name the covered scenario, so "
29473
+ f"the author cannot tell which Covers entry is the problem; got "
29474
+ f"{message!r}."
29475
+ )
29476
+ return 1
29477
+ if "unchanged" not in message:
29478
+ report(
29479
+ f"{label}: the refusal does not say why the two claims conflict; "
29480
+ f"got {message!r}."
29481
+ )
29482
+ return 1
29483
+
29484
+ # M2 of 1.2 — the guard is satisfied by coverage of the entry it
29485
+ # objected to, and not by a red-green task merely existing in the change.
29486
+ payload = escape_fixture("escape-sibling", sibling=covered)
29487
+ if payload.get("status") != "pass":
29488
+ report(
29489
+ f"{label}: a sibling task covering the same scenario under "
29490
+ f"vertical-tdd did not satisfy the guard. {problem_text(payload)}"
29491
+ )
29492
+ return 1
29493
+ payload = escape_fixture(
29494
+ "escape-other-sibling",
29495
+ sibling=(
29496
+ "demo / The moved attribute keeps its effect / The flow reports "
29497
+ "the attribute"
29498
+ ),
29499
+ )
29500
+ if payload.get("status") == "pass":
29501
+ report(
29502
+ f"{label}: any sibling satisfied the guard — a red-green task "
29503
+ "covering a different scenario was accepted as proof of this "
29504
+ "one, which makes the guard a check that a change contains at "
29505
+ "least one red-green task."
29506
+ )
29507
+ return 1
29508
+ if "equivalence-covers-added-behavior" not in problem_codes(payload):
29509
+ report(
29510
+ f"{label}: any sibling satisfied the guard — refused for another "
29511
+ f"reason; got {problem_codes(payload)!r}."
29512
+ )
29513
+ return 1
29514
+
29515
+ # M3 of 1.2 — the guard fires on new behavior, not on the strategy. A
29516
+ # task covering only identifiers, or a requirement the change does not
29517
+ # add, is an ordinary equivalence task.
29518
+ plain = start("escape-plain")
29519
+ if plain.get("status") != "pass":
29520
+ report(
29521
+ f"{label}: an equivalence task covering no added scenario was "
29522
+ f"refused. {problem_text(plain)}"
29523
+ )
29524
+ return 1
29525
+
29526
+ # M3 — the shape that is complete, resolvable, and still empty. This is
29527
+ # the refusal worth having: nothing about the task looks wrong, and the
29528
+ # check passes having compared a thing against itself.
29529
+ payload = start("base-is-head", base="HEAD")
29530
+ if payload.get("status") == "pass":
29531
+ report(
29532
+ f"{label}: accepted a base that is head — an A/B against itself "
29533
+ "always agrees, so the check proves nothing and looks complete "
29534
+ "doing it."
29535
+ )
29536
+ return 1
29537
+ if "equivalence-base-is-head" not in problem_codes(payload):
29538
+ report(
29539
+ f"{label}: accepted a base that is head — refused for another "
29540
+ f"reason; got {problem_codes(payload)!r}."
29541
+ )
29542
+ return 1
29543
+ message = problem_text(payload)
29544
+ head = subprocess.run(
29545
+ ["git", "-C", str((root / "base-is-head").resolve()), "rev-parse",
29546
+ "HEAD"],
29547
+ capture_output=True, text=True,
29548
+ ).stdout.strip()
29549
+ if head and head not in message:
29550
+ report(
29551
+ f"{label}: the refusal does not name the resolved commit, so the "
29552
+ "author cannot tell which ref collapsed onto HEAD; got "
29553
+ f"{message!r}."
29554
+ )
29555
+ return 1
29556
+ if "always agrees" not in message:
29557
+ report(
29558
+ f"{label}: the refusal does not say why a base that is HEAD is "
29559
+ f"empty rather than merely redundant; got {message!r}."
29560
+ )
29561
+ return 1
29562
+
29563
+ # 1.3 — Evidence may point at the machine output instead of retelling
29564
+ # it. A 244-line tasks.md that is mostly transcribed test output is a
29565
+ # transcription that can be wrong and that nobody can re-check.
29566
+ def artifact_fixture(
29567
+ name: str,
29568
+ *,
29569
+ artifact_path: str = "openspec/changes/demo/evidence/compare.json",
29570
+ body: str = '{"wns": 0.0, "tns": 0.0}\n',
29571
+ recorded: str | None = None,
29572
+ write_at: str | None = None,
29573
+ ) -> dict:
29574
+ repo = (root / name).resolve()
29575
+ repo.mkdir()
29576
+ git(repo, "init", "-q")
29577
+ git(repo, "config", "user.email", "t@example.com")
29578
+ git(repo, "config", "user.name", "keel-test")
29579
+ digest = hashlib.sha256(body.encode("utf-8")).hexdigest()
29580
+ evidence = recorded or f"artifact {artifact_path} sha256:{digest}"
29581
+ tasks = equivalence_task().replace(
29582
+ " - M1: pending", f" - M1: {evidence}"
29583
+ )
29584
+ write_gate_fixture(repo, tasks=tasks)
29585
+ if write_at is not None:
29586
+ write_text(repo / write_at, body)
29587
+ write_text(repo / "src/example.js", "// product\n")
29588
+ git(repo, "add", "-A")
29589
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "first")
29590
+ write_text(repo / "src/example.js", "// moved\n")
29591
+ git(repo, "add", "-A")
29592
+ git(repo, "-c", "commit.gpgsign=false", "commit", "-q", "-m", "second")
29593
+ run_keel(
29594
+ repo, "gate", "task-start", "--change", "demo", "--task", "1.1",
29595
+ "--record", "--no-guard",
29596
+ )
29597
+ result = run_keel(
29598
+ repo, "gate", "task-complete", "--change", "demo", "--task",
29599
+ "1.1", "--json",
29600
+ )
29601
+ try:
29602
+ return json.loads(result.stdout)
29603
+ except json.JSONDecodeError:
29604
+ return {
29605
+ "status": "unparsed",
29606
+ "problems": [{"message": result.stdout[:400]}],
29607
+ }
29608
+
29609
+ inside = "openspec/changes/demo/evidence/compare.json"
29610
+ payload = artifact_fixture("artifact-ok", write_at=inside)
29611
+ if payload.get("status") != "pass":
29612
+ report(
29613
+ f"{label}: an artifact reference was not verified — a reference "
29614
+ "to a file that is there, with a digest that matches, was "
29615
+ f"refused. {problem_text(payload)}"
29616
+ )
29617
+ return 1
29618
+ payload = artifact_fixture("artifact-absent", write_at=None)
29619
+ if payload.get("status") == "pass":
29620
+ report(
29621
+ f"{label}: an artifact reference was not verified — a reference "
29622
+ "to a file that does not exist was accepted, so the form is "
29623
+ "tolerated as prose rather than checked. Any sentence would "
29624
+ "have passed the same way."
29625
+ )
29626
+ return 1
29627
+ if "artifact-missing" not in problem_codes(payload):
29628
+ report(
29629
+ f"{label}: an artifact reference was not verified — refused for "
29630
+ f"another reason; got {problem_codes(payload)!r}."
29631
+ )
29632
+ return 1
29633
+
29634
+ # M2 of 1.3 — the digest is what makes the pointer worth more than a
29635
+ # path. A file that moved after the digest was recorded is the case a
29636
+ # bare path cannot see, and it is the common one: the command gets
29637
+ # re-run.
29638
+ stale = hashlib.sha256(b"different\n").hexdigest()
29639
+ payload = artifact_fixture(
29640
+ "artifact-stale",
29641
+ recorded=f"artifact {inside} sha256:{stale}",
29642
+ write_at=inside,
29643
+ )
29644
+ if payload.get("status") == "pass":
29645
+ report(
29646
+ f"{label}: accepted a stale digest — the artifact's content does "
29647
+ "not hash to the recorded digest and the reference was accepted, "
29648
+ "so Review reads whatever the file says now."
29649
+ )
29650
+ return 1
29651
+ if "artifact-digest-mismatch" not in problem_codes(payload):
29652
+ report(
29653
+ f"{label}: accepted a stale digest — refused for another reason; "
29654
+ f"got {problem_codes(payload)!r}."
29655
+ )
29656
+ return 1
29657
+ message = problem_text(payload)
29658
+ for expected in (inside, stale[:12]):
29659
+ if expected not in message:
29660
+ report(
29661
+ f"{label}: the refusal does not name {expected!r}, so a "
29662
+ "reader cannot tell a stale record from a wrong path; got "
29663
+ f"{message!r}."
29664
+ )
29665
+ return 1
29666
+
29667
+ # M3 of 1.3 — the inverse of the `Durable owner:` rule, for the opposite
29668
+ # reason: a follow-up pointer must outlive the change, and an evidence
29669
+ # artifact must travel with it. `openspec archive` moves the change
29670
+ # directory, so a path outside it is one the archive leaves behind.
29671
+ outside = "evidence/compare.json"
29672
+ payload = artifact_fixture("artifact-outside", artifact_path=outside, write_at=outside)
29673
+ if payload.get("status") == "pass":
29674
+ report(
29675
+ f"{label}: accepted a path archiving would leave behind — an "
29676
+ f"artifact at {outside!r} was accepted although the archive "
29677
+ "moves only the change directory."
29678
+ )
29679
+ return 1
29680
+ if "artifact-outside-change" not in problem_codes(payload):
29681
+ report(
29682
+ f"{label}: accepted a path archiving would leave behind — "
29683
+ f"refused for another reason; got {problem_codes(payload)!r}."
29684
+ )
29685
+ return 1
29686
+ if "archiv" not in problem_text(payload):
29687
+ report(
29688
+ f"{label}: the refusal does not say that archiving is what "
29689
+ f"breaks the pointer; got {problem_text(payload)!r}."
29690
+ )
29691
+ return 1
29692
+
29693
+ if label not in {name for name, _ in SCENARIOS}:
29694
+ report(f"{label}: the scenario registry does not include it.")
29695
+ return 1
29696
+ report(f"{label} scenario passed.")
29697
+ return 0
29698
+
29699
+
29045
29700
  SCENARIOS: tuple = (
29046
29701
  ("stateless-continuity", validate_stateless_continuity_scenario),
29047
29702
  ("core-gates", validate_core_gates_scenario),
@@ -29404,6 +30059,10 @@ SCENARIOS: tuple = (
29404
30059
  "a-negation-is-not-a-marker",
29405
30060
  validate_a_negation_is_not_a_marker_scenario,
29406
30061
  ),
30062
+ (
30063
+ "an-equivalence-claim-names-its-base",
30064
+ validate_an_equivalence_claim_names_its_base_scenario,
30065
+ ),
29407
30066
  (
29408
30067
  "guidance-is-referenced-and-carries-no-criterion",
29409
30068
  validate_guidance_is_referenced_and_carries_no_criterion_scenario,
package/src/core/gates.js CHANGED
@@ -2,6 +2,7 @@
2
2
 
3
3
  // Keel 4.1.0 deterministic gate contract.
4
4
 
5
+ const crypto = require("crypto");
5
6
  const fs = require("fs");
6
7
  const path = require("path");
7
8
  const { spawnSync } = require("child_process");
@@ -293,6 +294,112 @@ function taskShapeWarnings(repo, selection, task, compiled) {
293
294
  ];
294
295
  }
295
296
 
297
+ // The requirement names a change declares as new, read from the `## ADDED
298
+ // Requirements` sections of its delta specs. A `## MODIFIED` requirement is
299
+ // deliberately not here: behavior that changed is still behavior an equivalence
300
+ // task can legitimately claim is measurement-stable.
301
+ function addedRequirementNames(repo, change) {
302
+ const specsRoot = path.join(repo, "openspec", "changes", change, "specs");
303
+ const names = new Set();
304
+ let capabilities = [];
305
+ try {
306
+ capabilities = fs.readdirSync(specsRoot);
307
+ } catch {
308
+ return names;
309
+ }
310
+ for (const capability of capabilities) {
311
+ const specPath = path.join(specsRoot, capability, "spec.md");
312
+ let content = "";
313
+ try {
314
+ content = fs.readFileSync(specPath, "utf8");
315
+ } catch {
316
+ continue;
317
+ }
318
+ // Only the ADDED section, bounded by the next `## ` heading, so a
319
+ // requirement listed under MODIFIED or REMOVED is not read as new.
320
+ for (const section of content.split(/^##\s+/m).slice(1)) {
321
+ if (!/^ADDED Requirements\s*$/m.test(section.split(/\r?\n/)[0])) continue;
322
+ for (const match of section.matchAll(/^###\s+Requirement:\s*(.+?)\s*$/gm)) {
323
+ names.add(match[1]);
324
+ }
325
+ }
326
+ }
327
+ return names;
328
+ }
329
+
330
+ // A task's declared strategy, from the text it wrote. Both forms: the compact
331
+ // `Strategy:` entry under `Verify`, and the expanded `Verification Strategy`
332
+ // field beside `Commands`.
333
+ function declaredStrategy(task) {
334
+ const compact = String(field(task, "Verify") || "").match(
335
+ /^\s*-?\s*Strategy:\s*(.+?)\s*$/m
336
+ );
337
+ const value = compact
338
+ ? compact[1]
339
+ : String(field(task, "Verification Strategy") || "");
340
+ return value.trim().toLowerCase();
341
+ }
342
+
343
+ // A `Covers` entry naming a spec scenario: `capability / Requirement / Scenario`.
344
+ // Identifier entries (`D4`, `F1`, `E2`) have no slashes and are not spec claims.
345
+ function specCoverEntries(task) {
346
+ return String(field(task, "Covers") || "")
347
+ .split(/\r?\n/)
348
+ .map((line) => line.replace(/^\s*-\s*/, "").trim())
349
+ .filter((entry) => entry.split("/").length >= 3)
350
+ .map((entry) => entry.split("/").map((part) => part.trim()));
351
+ }
352
+
353
+ // `equivalence` owes no red, which makes it the first strategy reached for by a
354
+ // task that should have one. A task covering a scenario its own change *adds* is
355
+ // claiming that behavior is new and that behavior is unchanged at the same time,
356
+ // and without this guard one of the two claims has no proof anywhere in the
357
+ // change. Satisfied only by a sibling covering the same entry, never by the mere
358
+ // presence of a red-green task somewhere in the change.
359
+ function equivalenceEscapeProblems(repo, selection, task, compiled) {
360
+ const strategy = String(
361
+ ((compiled.capsule || {}).verification || {}).strategy || ""
362
+ ).toLowerCase();
363
+ if (strategy !== "equivalence") return [];
364
+ const added = addedRequirementNames(repo, selection.change);
365
+ if (added.size === 0) return [];
366
+ const problems = [];
367
+ for (const parts of specCoverEntries(task)) {
368
+ const requirement = parts[1];
369
+ if (!added.has(requirement)) continue;
370
+ const entry = parts.join(" / ");
371
+ const scenario = parts.slice(2).join(" / ");
372
+ const covered = selection.tasks.some((sibling) => {
373
+ if (sibling.id === task.id) return false;
374
+ // Read from the sibling's own `Verify` text rather than by compiling it.
375
+ // Compiling made the guard depend on the sibling being otherwise valid: a
376
+ // sibling with any unrelated contract error produced no capsule, so its
377
+ // strategy read as absent and it silently stopped satisfying the guard.
378
+ const siblingStrategy = declaredStrategy(sibling);
379
+ if (!RED_GREEN_VERIFICATION_STRATEGIES.has(siblingStrategy)) return false;
380
+ // The same entry, not any entry. A guard satisfied by the presence of a
381
+ // red-green task would be a check that the change contains one, which
382
+ // every change with more than one task passes.
383
+ return specCoverEntries(sibling).some(
384
+ (other) => other.join(" / ") === entry
385
+ );
386
+ });
387
+ if (!covered) {
388
+ problems.push(
389
+ problem(
390
+ "equivalence-covers-added-behavior",
391
+ `This task declares \`Strategy: equivalence\` and covers `
392
+ + `"${scenario}", a scenario this change adds. Behavior that is new `
393
+ + "is not behavior that is unchanged, and no task of this change "
394
+ + "proves the new half: name a red-green task that covers the same "
395
+ + "entry, or move this Covers entry to the task that implements it."
396
+ )
397
+ );
398
+ }
399
+ }
400
+ return problems;
401
+ }
402
+
296
403
  function taskStart(repo, options) {
297
404
  const selection = loadSelection(repo, options);
298
405
  const task = selection.selected[0];
@@ -300,6 +407,7 @@ function taskStart(repo, options) {
300
407
  const problems = [
301
408
  ...compiled.diagnostics,
302
409
  ...invalidationProblems(repo, selection.content, selection.tasks, selection.change),
410
+ ...equivalenceEscapeProblems(repo, selection, task, compiled),
303
411
  ];
304
412
  // Recording the current fingerprint is idempotent: --record replaces the
305
413
  // selected task's Contract anchor whatever it holds, so reauthorizing a task
@@ -454,6 +562,72 @@ function commandLabels(task) {
454
562
  ].map((match) => match[1]);
455
563
  }
456
564
 
565
+ // `artifact <path> sha256:<digest>`, or null for prose. The path is
566
+ // repo-relative; the digest is what Review is entitled to assume it is reading.
567
+ const ARTIFACT_EVIDENCE =
568
+ /^artifact\s+(\S+)\s+sha256:([0-9a-f]{64})\s*$/i;
569
+
570
+ function artifactReference(value) {
571
+ const match = String(value || "").trim().match(ARTIFACT_EVIDENCE);
572
+ return match ? { path: match[1], digest: match[2].toLowerCase() } : null;
573
+ }
574
+
575
+ // Keel checks identity and reads nothing else: it does not parse the artifact,
576
+ // does not know what a field is, and compares nothing in it. The claim that the
577
+ // numbers agree stays the author's, recorded before Review exactly as
578
+ // `Fails with:` and `Detects:` are. What the digest buys is that the file Review
579
+ // opens is the file the author meant, which a retelling cannot offer.
580
+ function artifactProblems(repo, change, label, reference) {
581
+ if (!reference) return [];
582
+ // The inverse of the `Durable owner:` rule, and for the opposite reason: a
583
+ // follow-up pointer has to outlive the change, while an evidence artifact has
584
+ // to travel with it. `openspec archive` moves the change directory and
585
+ // nothing else, so a path outside it is one the archive is guaranteed to
586
+ // leave behind — and an evidence pointer that breaks on archive is worse than
587
+ // a retelling, which at least survives.
588
+ const changeDir = path.join("openspec", "changes", change);
589
+ const normalized = reference.path.split(path.sep).join("/");
590
+ if (!normalized.startsWith(`${changeDir.split(path.sep).join("/")}/`)) {
591
+ return [
592
+ problem(
593
+ "artifact-outside-change",
594
+ `${label} Evidence references \`${reference.path}\`, which is outside `
595
+ + `\`${changeDir}\`. Archiving moves the change directory and nothing `
596
+ + "else, so this pointer breaks the moment the change is archived. "
597
+ + "Put the artifact inside the change's own directory."
598
+ ),
599
+ ];
600
+ }
601
+ const absolute = path.join(repo, reference.path);
602
+ if (!fs.existsSync(absolute) || !fs.statSync(absolute).isFile()) {
603
+ return [
604
+ problem(
605
+ "artifact-missing",
606
+ `${label} Evidence references the artifact \`${reference.path}\`, and `
607
+ + "no file is there. An unresolvable reference is worse than a "
608
+ + "retelling: the retelling at least carries the result."
609
+ ),
610
+ ];
611
+ }
612
+ const actual = crypto
613
+ .createHash("sha256")
614
+ .update(fs.readFileSync(absolute))
615
+ .digest("hex");
616
+ if (actual !== reference.digest) {
617
+ return [
618
+ problem(
619
+ "artifact-digest-mismatch",
620
+ `${label} Evidence records \`${reference.path}\` at `
621
+ + `sha256:${reference.digest}, and the file there hashes to `
622
+ + `sha256:${actual}. Re-record the digest if the command was re-run, `
623
+ + "or correct the path — naming both is what lets a reader tell a "
624
+ + "stale record from a pointer at the wrong file."
625
+ ),
626
+ ];
627
+ }
628
+ return [];
629
+ }
630
+
457
631
  function evidenceValue(task, label) {
458
632
  const escaped = label.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
459
633
  const match = field(task, "Evidence").match(
@@ -1049,11 +1223,21 @@ function completionChecks(repo, task, contract = null, changeVerify = null, chan
1049
1223
  problems.push(problem("missing-commands", "Commands must define at least one M<n>."));
1050
1224
  }
1051
1225
  for (const label of commands) {
1052
- if (!isConcrete(evidenceValue(task, label))) {
1226
+ const recorded = evidenceValue(task, label);
1227
+ if (!isConcrete(recorded)) {
1053
1228
  problems.push(
1054
1229
  problem("missing-evidence", `Missing concrete Evidence for ${label}.`)
1055
1230
  );
1231
+ continue;
1056
1232
  }
1233
+ // An A/B pairing or a sweep summary *is* the output of one command, and
1234
+ // retelling it into tasks.md is a transcription that can be wrong and that
1235
+ // nobody can re-check (issue #142). A reference is checked rather than
1236
+ // tolerated: as prose it would already have passed as concrete, so without
1237
+ // this the form would buy nothing at all.
1238
+ problems.push(
1239
+ ...artifactProblems(repo, change, label, artifactReference(recorded))
1240
+ );
1057
1241
  }
1058
1242
  // A declared measurement is held against the check's own bare `M<n>` Evidence
1059
1243
  // — the entry where the command and its output are recorded. A literal that
@@ -1,6 +1,7 @@
1
1
  "use strict";
2
2
 
3
3
  const crypto = require("crypto");
4
+ const { execFileSync } = require("child_process");
4
5
  const fs = require("fs");
5
6
  const path = require("path");
6
7
 
@@ -143,6 +144,13 @@ const SUPPORTED_VERIFICATION_STRATEGIES = [
143
144
  "snapshot-characterization",
144
145
  "rendered-behavior",
145
146
  "evidence-first",
147
+ // Zero difference, not "nothing could fail first". `evidence-first` is the
148
+ // only other strategy without a red-green obligation, and it is scoped by an
149
+ // absence; an A/B against a base is the opposite — a criterion stronger than
150
+ // red-green, because it also catches the change that incidentally moved a
151
+ // result. Issue #142 measured a task re-recording its contract twice to get
152
+ // past the shape rather than the criterion.
153
+ "equivalence",
146
154
  ];
147
155
 
148
156
  const RED_GREEN_VERIFICATION_STRATEGIES = new Set([
@@ -257,6 +265,77 @@ function isPassingReviewStatus(value) {
257
265
  );
258
266
  }
259
267
 
268
+ // A git ref resolved locally, or null. Reads the repository the gate is already
269
+ // reading and reaches nothing else: a gate that fetched would stop being local
270
+ // and offline, which is the property its verdict rests on.
271
+ function resolveCommit(repo, ref) {
272
+ try {
273
+ return execFileSync(
274
+ "git",
275
+ ["-C", repo, "rev-parse", "--verify", "--quiet", `${ref}^{commit}`],
276
+ { encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }
277
+ ).trim() || null;
278
+ } catch {
279
+ return null;
280
+ }
281
+ }
282
+
283
+ // Four ways an `equivalence` task's shape can be complete and still compare
284
+ // nothing. Each names the declaration it is about: a diagnostic naming the
285
+ // strategy would send the author to the one line that is correct.
286
+ function equivalenceProblems(repo, taskVerification) {
287
+ const problems = [];
288
+ const declaredFields = taskVerification.fieldsDeclared;
289
+ if (!isConcrete(taskVerification.base)) {
290
+ problems.push({
291
+ code: "missing-equivalence-base",
292
+ message:
293
+ "equivalence compares one code path at two commits and declares which "
294
+ + "one it is compared against. Add a `Base:` entry beside `Strategy:` "
295
+ + "naming a git ref — without it the criterion is `the numbers are the "
296
+ + "same as some other numbers`.",
297
+ });
298
+ }
299
+ if (taskVerification.fields.length === 0) {
300
+ problems.push({
301
+ code: "missing-equivalence-fields",
302
+ message: declaredFields
303
+ ? "`Fields:` is declared and resolves to an empty set, so the "
304
+ + "comparison has nothing to compare. It reads as a declaration to "
305
+ + "every reader except the comparison; name the fields, separated by "
306
+ + "commas."
307
+ : "equivalence declares which fields are compared. Add a `Fields:` "
308
+ + "entry beside `Strategy:` listing them, separated by commas — a "
309
+ + "comparison with no field set agrees with everything.",
310
+ });
311
+ }
312
+ if (isConcrete(taskVerification.base)) {
313
+ const base = resolveCommit(repo, taskVerification.base);
314
+ if (!base) {
315
+ problems.push({
316
+ code: "unresolvable-equivalence-base",
317
+ message:
318
+ `\`Base: ${taskVerification.base}\` resolves to no commit in this `
319
+ + "repository. The base is read locally and never fetched, so a ref "
320
+ + "that exists only on a remote is not one this gate can see.",
321
+ });
322
+ } else {
323
+ const head = resolveCommit(repo, "HEAD");
324
+ if (head && head === base) {
325
+ problems.push({
326
+ code: "equivalence-base-is-head",
327
+ message:
328
+ `\`Base: ${taskVerification.base}\` resolves to ${base}, which is `
329
+ + "HEAD. An A/B against itself always agrees, so the check would "
330
+ + "pass having compared nothing — the one shape here that is "
331
+ + "complete, resolvable, and still empty.",
332
+ });
333
+ }
334
+ }
335
+ }
336
+ return problems;
337
+ }
338
+
260
339
  function verification(task) {
261
340
  const compact = fieldValues(task, "Verify");
262
341
  const strategyEntry = compact.find((entry) => /^Strategy:\s*/i.test(entry));
@@ -266,8 +345,19 @@ function verification(task) {
266
345
  // command, and a reason that took an `M<n>` label would be a check the author
267
346
  // never wrote and evidence nobody can record.
268
347
  const reasonEntry = compact.find((entry) => /^Reason:\s*/i.test(entry));
348
+ // `equivalence` compares one code path at two commits. Neither half of that
349
+ // fits in a check: the check is the command, and what it cannot say by itself
350
+ // is which commit it is compared against and which fields are compared. #142
351
+ // proposed a third field for the command too; commands already have exactly
352
+ // one home here, and a second would put half of them outside the labelled
353
+ // evidence `task-complete` enforces.
354
+ const baseEntry = compact.find((entry) => /^Base:\s*/i.test(entry));
355
+ const fieldsEntry = compact.find((entry) => /^Fields:\s*/i.test(entry));
269
356
  const isVerificationField = (entry) =>
270
- /^Strategy:\s*/i.test(entry) || /^Reason:\s*/i.test(entry);
357
+ /^Strategy:\s*/i.test(entry)
358
+ || /^Reason:\s*/i.test(entry)
359
+ || /^Base:\s*/i.test(entry)
360
+ || /^Fields:\s*/i.test(entry);
271
361
  const commandSource = compact.length > 0
272
362
  ? compact.filter((entry) => !isVerificationField(entry))
273
363
  : fieldValues(task, "Commands");
@@ -342,6 +432,24 @@ function verification(task) {
342
432
  ? reasonEntry.replace(/^Reason:\s*/i, "")
343
433
  : field(task, "Verification Reason")
344
434
  ),
435
+ base: normalizeText(
436
+ baseEntry ? baseEntry.replace(/^Base:\s*/i, "") : field(task, "Verification Base")
437
+ ),
438
+ // A set, so "declared but empty" is a state the gate can see. `Fields:` with
439
+ // nothing behind it is the shape that passes while comparing nothing, and it
440
+ // reads as a declaration to everyone except the comparison.
441
+ // Whether the line was written at all, kept beside the parsed set so a
442
+ // refusal can tell an author who wrote nothing from one who wrote an empty
443
+ // set. To the comparison they are the same state; to the author they are
444
+ // opposite mistakes.
445
+ fieldsDeclared: Boolean(fieldsEntry || field(task, "Verification Fields")),
446
+ fields: (fieldsEntry
447
+ ? fieldsEntry.replace(/^Fields:\s*/i, "")
448
+ : field(task, "Verification Fields") || ""
449
+ )
450
+ .split(",")
451
+ .map((entry) => normalizeText(entry))
452
+ .filter(Boolean),
345
453
  commands,
346
454
  };
347
455
  }
@@ -1191,6 +1299,11 @@ function compileTaskContract(repo, change, task) {
1191
1299
  + "needs no reason.",
1192
1300
  });
1193
1301
  }
1302
+ if (taskVerification.strategy.toLowerCase() === "equivalence") {
1303
+ resolved.diagnostics.push(
1304
+ ...equivalenceProblems(repo, taskVerification)
1305
+ );
1306
+ }
1194
1307
  const couplingMode = normalizeText(field(task, "Coupling")).toLowerCase()
1195
1308
  || "none";
1196
1309
  const candidateBoundary = normalizedValues(task, "Candidate Boundary", {
@@ -1311,6 +1424,14 @@ function compileTaskContract(repo, change, task) {
1311
1424
  // other task keeps the capsule shape and fingerprint it had before the
1312
1425
  // field existed.
1313
1426
  ...(taskVerification.reason ? { reason: taskVerification.reason } : {}),
1427
+ // Same rule: emitted only by the strategy that declares them, so every
1428
+ // existing task's capsule shape and fingerprint are untouched. They belong
1429
+ // in the capsule rather than only in the file because a declaration
1430
+ // outside the fingerprint could be edited after the run it describes.
1431
+ ...(taskVerification.base ? { base: taskVerification.base } : {}),
1432
+ ...(taskVerification.fields.length > 0
1433
+ ? { fields: taskVerification.fields }
1434
+ : {}),
1314
1435
  // Emit a tag only when the check opts out of a default, so an untagged
1315
1436
  // check keeps the capsule shape and fingerprint it had before either tag
1316
1437
  // existed. `layer` appears only for `fast`, `regression` only when true.