@ccoalm/ccl-skills 0.16.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +8 -7
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +22 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-review-covers-head.sh +128 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-untracked-background.sh +53 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_review_covers_head.sh +104 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_untracked_background.sh +75 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +8 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -4
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +13 -1
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +78 -126
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/wording-only-review.md +136 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +521 -32
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +25 -2
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +30 -15
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +439 -19
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +14 -7
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +1 -1
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +6 -6
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +1 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +48 -119
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +9 -9
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +22 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +18 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +2 -2
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check_review_evidence_present.py +122 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +4 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +37 -8
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +102 -20
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +16 -27
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +3 -5
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_review_evidence_present.sh +81 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +137 -279
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +41 -1
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +49 -3
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +3 -2
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/closeout-reread.md +40 -0
  38. package/dist/assets/release.json +72 -52
  39. package/package.json +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +0 -1242
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +0 -978
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +0 -1477
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +0 -1183
@@ -30,15 +30,22 @@ CORE="$TMP/core.txt"
30
30
  LATEST="$TMP/latest.txt"
31
31
  OLD_INTENT="$TMP/old-intent.txt"
32
32
 
33
- python3 - "$PLAN" "$APPEND" "$CORE" "$LATEST" "$OLD_INTENT" <<'PY'
33
+ python3 - "$PLAN" "$APPEND" "$CORE" "$LATEST" "$OLD_INTENT" "$SCRIPT_DIR" <<'PY'
34
34
  import json
35
35
  import hashlib
36
+ import subprocess
36
37
  import sys
37
38
  from pathlib import Path
38
39
 
39
40
  plan_path, append_path, core_path, latest_path, old_intent_path = map(
40
- Path, sys.argv[1:]
41
+ Path, sys.argv[1:6]
41
42
  )
43
+ required_concerns = subprocess.run(
44
+ [sys.executable, str(Path(sys.argv[6]) / "review_gate.py"),
45
+ "--print-required-concerns", "--stage", "build"],
46
+ capture_output=True, text=True, check=True,
47
+ ).stdout.split()
48
+ assert required_concerns, "the controller printed no required concerns"
42
49
  old = "scope:" + ("x" * (3995 - len("scope:") - len("c27"))) + "c27"
43
50
  latest = "latest-round:c28"
44
51
  core = old[: 3892 - len("\n\n") - len(latest)]
@@ -55,12 +62,12 @@ plan = {
55
62
  "conclusion": conclusion,
56
63
  "evidence_refs": ["focused-test"],
57
64
  }
65
+ # Derived, not copied: a fixture holding its own copy of the required set
66
+ # stops satisfying the gate the moment that set changes, and the runner
67
+ # aborts at its first failing target so the drift surfaces rounds later.
58
68
  for concern, conclusion in (
59
- ("correctness", "The focused checks cover the updater's accepted state transitions."),
60
- ("safety", "The focused checks cover no-write failures and file integrity boundaries."),
61
- ("failure_paths", "The focused checks cover overflow, stale input, and malformed text paths."),
62
- ("tests_evidence", "The focused regression fails when bounded update guarantees are removed."),
63
- ("compatibility", "The focused checks preserve the plan schema and original file permissions."),
69
+ (concern, f"The focused updater checks cover {concern} on this fixture.")
70
+ for concern in required_concerns
64
71
  )
65
72
  ],
66
73
  "evidence": [
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: multi-perspective-research
3
- description: Skip 优先判定:一查便知的单点事实问题 → 直接回答不套流程;请求点名了候选让选、或用裁决措辞要结论(选哪个 / 要不要上 X / 该不该做 / 可不可行)→ 先 `product-rd-workflow`(它可回调本技能产调研底稿);纯主题/领域/写作准备调研留在本技能;拷问已有方案 → `grill-me`;bug 根因 → `defect-diagnosis`;复盘沉淀 → `skill-extraction-workflow`。触发:调研一个主题 / 深度调研 / 多视角研究 / 帮我研究一下 X / 写作前调研 / 领域·组织研究(不喂决策)/ research a topic → run a structured research loop(视角枚举 → 矛盾图 → 合成简报 → 自评)with evidence-grounding discipline, instead of one-question surface answers.
3
+ description: 调研一个主题 / 深度调研 / deep research / 调研某个产品·公司·赛道 / 多视角研究 / 帮我研究一下 X / 写作前调研 / 领域·组织研究(不喂决策)/ research a topic → run a structured research loop(视角枚举 → 矛盾图 → 合成简报 → 自评)with evidence-grounding discipline, instead of one-question surface answers; another installed deep-research skill may serve as a search tool inside this loop, not replace it. Skip:一查便知的单点事实问题 → 直接回答不套流程;请求点名了候选让选、或用裁决措辞要结论(选哪个 / 要不要上 X / 该不该做 / 可不可行)→ 先 `product-rd-workflow`(它可回调本技能产调研底稿);纯主题/领域/产品/写作准备调研留在本技能;拷问已有方案 → `grill-me`;bug 根因 → `defect-diagnosis`;复盘沉淀 → `skill-extraction-workflow`。
4
4
  ---
5
5
 
6
6
  # Multi-Perspective Research(多视角主题调研)
@@ -122,7 +122,7 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
122
122
  - **Bounded mechanical-move exception** — scoped review of a move-dominated diff only under the reference's infeasibility precondition and four conditions.
123
123
  - **Cross-cutting shared-runtime primitives** — explicit adversarial pass over edge paths plus a regression test per path found.
124
124
  - **Implementer self-review row** — persisted before the independent review in a record whose history proves ordering; a late row invalidates the review.
125
- - **Changed candidate = refreshed row + fresh full-scope rerun** — any material post-review change mechanically triggers a full-scope rerun; the author cannot narrow the rerun's scope or self-classify the change as whitespace-only.
125
+ - **Changed candidate = refreshed row + fresh full-scope rerun** — any post-review change, tests/docs included, mechanically triggers a full-scope rerun; the author cannot narrow its scope or call the change immaterial.
126
126
  - **Independent gate surfacing basics = process defect** — repair the self-review/deterministic-gate loop before rerunning, and findings still require disposition, not waiver.
127
127
  - **Green-tests-alone merge is the same defect** as reaching implementation with only spec plus plan.
128
128
  - **Human/team sign-off** — for high-risk money/permission/data paths and any contract/API change with an external consumer, before implementation, merge, or launch.
@@ -26,7 +26,7 @@ Turn observed experience into reusable skills without business-specific details.
26
26
 
27
27
  - **R0 (mandatory clean-landing gate)**: Before marking any skill or reference change `R0-clean`, `landing-clean`, merge-ready, or cleanly landed, run the leakage audit (`audit_cmd` from the maintainer's private alias YAML) and require **zero hits** across all leakage categories — design-source file keys/URLs/node-ids, project/team identifiers, real subproject paths / repo / branch names, contributor names/emails, ticket ids, internal domains/hosts, and any non-distilled business/product noun pointing to one specific organization.
28
28
  - **Probe, don't infer**: a local environment without `ALIAS_AUDIT_CMD` is normal, but you must *observe* that — before recording `interim` / `r0_status`, run `check-ccl-skills.sh` and cite the actual `r0_status=` / final-token line it produced; never infer the branch from "unset is common" or a bare `printenv ALIAS_AUDIT_CMD` pre-check. Skipping straight to the interim fallback is a blocked-verification miss (a set var means the private audit is *configured*; it is *available* only once it runs clean to `alias_audit_ok`, and a set-but-broken var is itself a blocked-verification failure to fix or waive, not a fallback licence).
29
- - **Interim commit semantics**: when `ALIAS_AUDIT_CMD` is genuinely unset, the contributor may commit, push, or open a Draft/WIP MR only as `interim` with `R0 pending maintainer audit` recorded, but must not merge or claim clean landing until R0 passes or an explicit risk-owner waiver is recorded. (Label semantics: `interim` alone does not decide commit permission — this R0-pending `interim` allows commit/push/Draft-MR but blocks merge/clean-landing, while the review/challenge gate's `interim` is an *uncommitted checkpoint* that blocks commit itself; each gate's failure carries its own commit permission, and when both fail the stricter applies.)
29
+ - **Interim commit semantics**: when `ALIAS_AUDIT_CMD` is genuinely unset, the contributor may commit, push, or open a Draft/WIP MR only as `interim` with `R0 pending maintainer audit` recorded, but must not merge or claim clean landing until R0 passes or an explicit risk-owner waiver is recorded. (Label semantics: `interim` alone does not decide commit permission — this R0-pending `interim` allows commit/push/Draft-MR but blocks merge/clean-landing, while the review/challenge gate's `interim` is an *uncommitted checkpoint* that blocks commit itself, except the dual-track local checkpoint commit; each gate's failure carries its own commit permission, and when both fail the stricter applies.)
30
30
  - **Edit-time gates that stay inline**: every new sanitized label MUST already exist in the alias YAML before clean landing (fail-closed); pre-existing leakage may be `known_debt` but new/modified content MUST stay zero-hit; grep cannot catch source-shaped example identifiers (variable/function/class/file/package names lifted verbatim) so adversarial review (codex challenge or equivalent) is the practical safety net.
31
31
  - **False-green guard**: the public fallback is public interim evidence only — the private alias audit (`alias_audit_ok`) has NOT run when those tokens print, so record `interim / R0 pending` and never treat `ccl_skill_check_ok`, `ccl_skill_check_interim_ok`, `generic_r0_leak_scan_ok`, or `alias_audit_unavailable` as clean-landing R0 evidence; the clean-landing signal is `ccl_skill_check_clean_ok` (`r0_status=private-ok`); no project alias is not a waiver (use the generic process-retro profile when no product corpus applies; neither the generic fallback nor an ad-hoc `grep` is the clean-landing gate).
32
32
  - For category definitions, alias-YAML structure, the false-green guard and fallback/token semantics, `known_debt` semantics, example-identifier substitution, the generic process-retro profile, and `ALIAS_AUDIT_CMD` enforcement, read `references/r0-leakage-audit.md`.
@@ -154,20 +154,20 @@ Turn observed experience into reusable skills without business-specific details.
154
154
  - **"The owner skill already states the rule" / "avoid monotonic growth" does NOT license a memory-only or no-op landing when the gate demonstrably failed to fire.** If a rule exists yet the failure still happened (and would recur for another agent or project), adequate *content* is not adequate *enforcement*: land the firing mechanism — the trigger, closeout step, validator, or merged clause that makes the existing rule actually catch this case, in the owning shared skill — or prove it now fires. Retreating to a personal memory or "no change, content is fine" while the gate stays un-fired is the dodge this prevents (memory is supplement only, per the memory-only-insufficient rule).
155
155
  - For changed upstream owners, `check-ccl-skills.sh` (via `scripts/impact-chain-gate.rb`) is the mechanical closeout over every added source-register row; the declaration format (behavioral-evidence / observed-failure / firing-path fragments), anchor rules, wording-only classification, the `RED-baseline` floor, and the author-declaration trust model: `references/external-practice-controls.md#behavioral-evidence-and-attestation`.
156
156
  - For shared-skill changes, classify the diff before finalization: `shared-skill change` is defined in `references/dual-track-review-gate.md` (it also covers the plugin-shipped command/behavior surfaces outside `skills/`); `wording-only` means punctuation, grammar, typo, formatting, or synonym substitution with no change to trigger, scope, routing, validation, condition, example, owner, or acceptance meaning, and touching no frontmatter (any `description`/frontmatter edit, even a pure typo fix, is a routing-surface change, never wording-only); all other changes are non-wording.
157
- - Every shared-skill change, including wording-only edits, must include a recorded independent review row before commit.
157
+ - Every shared-skill change, including wording-only edits, must include a recorded independent review row before commit. A local unpushed checkpoint commit on an isolated worktree branch is allowed so each pass can name the commit it reviewed; push, pull request, readiness claim and landing still wait for the owed passes.
158
158
  - Non-wording shared-skill changes must also include a recorded challenge row AND a recorded behavioral-evidence row before commit (actual behavior/routing deltas require a true `RED-baseline`; unchanged controls use paired `semantic-control`; status rules in `references/dual-track-review-gate.md`). A non-wording owner package must include at least one `RED-baseline` row, so stable-control labels cannot self-clear the package. **For a DESTRUCTIVE/irreversible change, a `RED-baseline` row must show the protected predicates were mutated, not merely that negative cases ran.** Executing the must-NOT-touch cases is necessary but not sufficient: a negative probe that would still pass with its protecting predicate removed is evidence of nothing (recurring shape: a degenerate fixture short-circuits every probe on an unrelated conservative branch, so the safety predicate is never reached and the green suite certifies the hole).
159
159
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
160
160
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
161
161
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
162
- - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=1` (2 rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled new rounds. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
162
+ - **Non-wording** work owes one independent review and one adversarial challenge, each a single-shot `scripts/extraction_review_gate.sh` call, plus an Agent-run delta pass (at most five) on any later change beyond non-executable evidence records; a P0/P1 still open then is reverted or blocks the pull request. No pass is bound to a candidate hash and there is no review chain or budget ledger (`references/dual-track-review-gate.md`). Proven wording-only work keeps the single-review path. Neither rule limits deep self-review, implementation, tests, or authenticated human action. Candidate input cannot assert human authority.
163
163
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
164
- - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At budget end validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
164
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
165
165
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
166
166
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
167
167
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
168
168
  - Independent review is **dual-track** for any non-wording shared-skill change, deep extraction, multi-skill landing, new shared skill, or skill change that ships operational/architectural rules: run BOTH (a) a fact/consistency review (`codex review` or equivalent — catches inaccuracies, contradictions across references, sanitization gaps, over-prescription) AND (b) an adversarial challenge (`codex exec` with adversarial prompt — hunts race conditions, data-loss paths, security holes, algorithm flaws, operational footguns). Skipping the challenge mode is how P0/P1 production-safety issues survive into shared skills. The two are not interchangeable: review catches what's wrong; challenge catches what would break under chaos. See `references/dual-track-review-gate.md` for the runbook.
169
169
  - **Draft-time corollary — pre-cover the recurring first-draft blind-spot AXES before challenge (not security alone).** Sweep EACH applicable axis explicitly before handing off (record ≥1 negative case per axis, or a reasoned `not-applicable`): **(1) security / privacy / authority / data-loss**, **(2) concurrency & lifecycle**, **(3) resource bounds**, **(4) rollout / migration ordering**, **(5) over-broad absolute**, **(6) enumeration-completeness (the mirror of (5) — under-listing, not over-listing)**. A draft that reads clean on the happy path almost always misses ≥1 of these; the challenge is the safety net, not the first line. When drafting any rule/code that touches execution, deletion, optimization, deployment, capability/permission, concurrency, resource lifetime, or a cross-service rollout, self-cover the applicable axes first. "Pre-cover" means recording at least one relevant negative case — or, for axis (6), the set-diff against the complete set — (or a reasoned `not-applicable`) per applicable axis, NOT a "covered" badge — and it never narrows or downgrades the mandatory adversarial challenge. The per-axis instance lists: `references/dual-track-review-gate.md` (pre-cover axis detail).
170
- - **Repeated same-class adversarial-challenge findings across rounds are a design/scope smell, not just more patches** — and the resolution depends on WHAT recurs. When the challenge surfaces NEW instances of the SAME risk class in two or more rounds, evaluate whether that capability should EXIST, not only how to patch this instance: convergence-by-deletion is often cleaner and more complete than convergence-by-patching, and is the right call when the capability serves no real need — judged from the convention's primary source and/or product/process evidence, not from the patch count alone. Record the decision explicitly as `keep / delete / narrow / replace` with the same-class evidence (the rounds and findings), the real user need it serves, a safer alternative, and blast-radius/migration; this complements the challenge-round convergence standard in step 6 (which counts remaining P0/P1 fixes) — same-class recurrence is the cue to question the design, not to open another patch round (observed: an auto-rewrite-of-a-human-file capability spawned a new data-loss footgun every round until it was deleted, after which the gate converged immediately). Two siblings share that decision record: **cross-landing** — the same class returning in production after a previous landing already "fixed" it means the fix shape itself is wrong, the tell being that each prior fix **re-instantiated the same predicate on new inputs**; when the predicate is the **vocabulary of an artifact the control does not own** (a field/value/version list from an upstream tool, format, or API) the second occurrence is already the design signal — re-express it over an **invariant the control does own** (shape, arity, type, or the independently-verified property the check actually needs) or accept the maintenance and say so, recording which invariant replaced the vocabulary and the residual risk the looser predicate accepts; **scope-direction** — a finding whose remediation would implement concerns the design explicitly defers to a later phase is a classification checkpoint, not a cut or another implement round: establish current-phase impact, split a compound finding, and let only a distinct risk owner (never the proposing controller) ratify a scope cut, so a cut defers rather than discards — the failure it prevents is the reviewer silently becoming a scope-expansion engine while the artifact gold-plates (observed: a first-cut state-machine contract ballooned across rounds until the controller reset to the in-scope diff). Failure shapes and mechanics: `references/dual-track-review-gate.md` (`Findings, autonomous budget, and human authority`; design-time operability check) and `references/external-practice-controls.md` (predicates over a vocabulary the control does not own).
170
+ - **Repeated same-class adversarial-challenge findings across rounds are a design/scope smell, not just more patches** — and the resolution depends on WHAT recurs. When the challenge surfaces NEW instances of the SAME risk class in two or more rounds, evaluate whether that capability should EXIST, not only how to patch this instance: convergence-by-deletion is often cleaner and more complete than convergence-by-patching, and is the right call when the capability serves no real need — judged from the convention's primary source and/or product/process evidence, not from the patch count alone. Record the decision explicitly as `keep / delete / narrow / replace` with the same-class evidence (the rounds and findings), the real user need it serves, a safer alternative, and blast-radius/migration; this complements the challenge-round convergence standard in step 6 (which counts remaining P0/P1 fixes) — same-class recurrence is the cue to question the design, not to open another patch round (observed: an auto-rewrite-of-a-human-file capability spawned a new data-loss footgun every round until it was deleted, after which the gate converged immediately). Two siblings share that decision record: **cross-landing** — the same class returning in production after a previous landing already "fixed" it means the fix shape itself is wrong, the tell being that each prior fix **re-instantiated the same predicate on new inputs**; when the predicate is the **vocabulary of an artifact the control does not own** (a field/value/version list from an upstream tool, format, or API) the second occurrence is already the design signal — re-express it over an **invariant the control does own** (shape, arity, type, or the independently-verified property the check actually needs) or accept the maintenance and say so, recording which invariant replaced the vocabulary and the residual risk the looser predicate accepts; **scope-direction** — a finding whose remediation would implement concerns the design explicitly defers to a later phase is a classification checkpoint, not a cut or another implement round: establish current-phase impact, split a compound finding, and let only a distinct risk owner (never the proposing controller) ratify a scope cut, so a cut defers rather than discards — the failure it prevents is the reviewer silently becoming a scope-expansion engine while the artifact gold-plates (observed: a first-cut state-machine contract ballooned across rounds until the controller reset to the in-scope diff). Failure shapes and mechanics: `references/dual-track-review-gate.md` (`Findings and dispositions`; design-time operability check) and `references/external-practice-controls.md` (predicates over a vocabulary the control does not own).
171
171
  - **Global-iteration boundary:** any `STOP`/`revert`/terminal wording in this gate applies to the reviewer lane, readiness claim, or defective dependent slice, never automatically to the overall task. Repeated root cause or no-progress rounds trigger method/design/validation changes or a parked human-decision item while independent runnable work continues. Only an explicit authenticated human stop ends the overall iteration.
172
172
  - **The should-it-exist question also applies at DESIGN TIME for any new mechanical gate, validator, or evidence apparatus, AND for any change to its verdict in EITHER direction — a loosening must be checked hardest, since it removes evidence instead of adding a red — run the check before landing AND before offering a human design options.** Legs: **author-dogfood** (the authoring workflow passes it end-to-end under every base resolution CI uses), **marginal-cost**, **trust-model fit** (for a loosening, which class stops owing evidence), **premise check** when tightening (a clean run on the current corpus is not evidence). Per-leg method and failure shapes: `references/dual-track-review-gate.md` (design-time operability check).
173
173
  - Reference files containing CLI commands, external API calls, or runnable code examples ship operational behavior and are NOT exempt from the dual-track gate; treat them the same as skill core-rule sections. Specifically, any reference file with recommended runnable commands or code that encodes operational behavior, external-state mutation, security-sensitive behavior, deployment/release/audit workflows, or nontrivial reusable recipes — including dangerous anti-patterns a teammate could plausibly copy — must pass both the fact/consistency review and the adversarial challenge before landing. Harmless validation commands and truly non-runnable static illustrative snippets (e.g., pseudocode, redacted/placeholder examples) remain eligible for a trivial-scope challenge skip only when the changed artifact is not a shared-skill change; any non-wording shared-skill change still requires challenge.
@@ -281,7 +281,7 @@ Turn observed experience into reusable skills without business-specific details.
281
281
  - A recorded behavioral-evidence row (`RED-baseline` / `semantic-control` / `not-applicable: docs-only` — see `references/dual-track-review-gate.md`) shows the change moves agent behavior as intended; independent review alone is not behavioral evidence.
282
282
  - If independent review was attempted through Claude Code and the first prompt hung or produced no output, validation must show the wrapper remediation path per the primary-reviewer remediation rule (Core Rules `Validation & the dual-track gate` group): the `code-review` wrapper invocation used (`claude_review.sh`, host/direct recovery), its structured-output validity, and whether the result was `findings applied`, `no blocking findings`, or `review unavailable after remediation`. The legacy bounded review packet is debugging/advisory context only — it does not satisfy this row.
283
283
  - Dual-track closeout record: the gate definition stays in the Core Rules `Validation & the dual-track gate` group; Step 6 records separate rows for fact/consistency review, adversarial challenge or table-permitted `challenge: not-required`, and the behavioral-evidence row required above. Use `references/dual-track-review-gate.md` and `references/description-authoring.md` for trigger/skip/convergence mechanics.
284
- - Challenge convergence means no **undispositioned** P0/P1 and a fresh unbiased full challenge of the exact landing candidate, including the final round; never use a scoped "confirm my prior fixes" pass as the final challenge. Record per-round log and final pass/abort decision per `references/dual-track-review-gate.md`.
284
+ - Challenge convergence means no **undispositioned** P0/P1 after one unprimed full challenge and the delta passes the post-review changes owe; never replace the full challenge with a "confirm my prior fixes" pass. Record the passes and the post-review delta per `references/dual-track-review-gate.md`.
285
285
  - Recurring-anti-patterns checklist: before committing operational/architectural skill changes, run the grep panel in `references/recurring-anti-patterns-checklist.md` against the changed file set; hits must be fixed or recorded as `known_debt` in the project alias map (pre-existing hits only — hits introduced by new/modified content must be fixed before landing, per the R0 rule's zero-hit requirement). The checklist grows over time; add a new pattern when a class appears in 2+ skills.
286
286
  - For fresh extraction wiring, start from `references/extraction-quickstart.md`. For specialized sources, read `references/two-source-extraction-pattern.md` for design+code extraction, `references/incident-postmortem-extraction.md` before landing incident/postmortem-derived rules, and `references/parallel-stack-references-pattern.md` for stack-agnostic + stack-specific sibling references.
287
287
  - Match validation depth to risk: light checks for small updates, strict review/test/install checks for new, shared, high-impact, or generator-backed skills.
@@ -25,6 +25,7 @@ The read side already defends against oversized files (chunked reads under ~200
25
25
  - An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
26
26
  - Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
27
27
  - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
+ - **Funding an addition by trimming prose means editing text that may be pinned — resolve the pins before rewriting, not after.** The ratchet's per-file freeze makes every addition to a legacy surface a rewrite of something else in the same file, and load-bearing sentences are pinned in two places: declaratively in `../../skill-extraction-workflow/scripts/contract-anchors.tsv`, which the fast repo gate checks, and as `grep -Fq` assertions inside owner suites, whose break a full lane run reports half an hour later. Read BOTH for the file you are about to trim — `awk -F'\t' '$2 == "<path>"' skills/skill-extraction-workflow/scripts/contract-anchors.tsv` lists the registry rows pinning it, and `grep -rn 'grep -Fq' skills/*/scripts/*.sh` finds the suite assertions — because a funded trim that silently retired two pinned wait-contract obligations was reported by the slow lane only, long after the edit. A pinned sentence may be reworded only together with whatever pins it, in the same landing.
28
29
  - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
29
30
 
30
31
  ## Retirement and relocation signal (usage census)
@@ -19,7 +19,7 @@ table. At closeout:
19
19
 
20
20
  - **Count lane names, never rounds**: enumerate the lanes the slice's own gate section requires, and confirm each one has a recorded outcome.
21
21
  - **A rounds table naming one lane while the gate requires two is `interim`, never landed** — no round count and no green deterministic gate substitutes for a lane that never ran.
22
- - **Where lanes leave evidence in a local store, a chain must exist for THIS slice** — check what the stored packets actually covered rather than assuming one is there.
22
+ - **Where lanes leave evidence in a local store, a stored result must exist for THIS slice** — check what the stored packets actually covered rather than assuming one is there.
23
23
 
24
24
  Observed: a mandatory, fail-closed repository gate landed with nine review-mode
25
25
  rounds recorded, no challenge lane run, and a landing state still reading
@@ -47,7 +47,7 @@ Before invoking either pass, self-audit the candidate to the point where *you* e
47
47
 
48
48
  **A design pivot resets the self-audit obligation.** When the candidate is substantially redesigned mid-gate (a mechanism replaced, a capability torn down, a rewrite beyond the findings being fixed), the earlier self-audit covered the OLD candidate: redo the closure self-audit on the NEW candidate before re-entering the challenge. Sliding from a pivot straight back into challenge → fix → re-challenge is the exact loop this section forbids, and it recurs precisely at pivots because the prior audit feels "already done."
49
49
 
50
- **Partition findings before fixing: mechanical fixes vs design decisions.** When a round returns findings, classify each before starting fix work: a *mechanical* finding (bug, missing check, wrong value) goes on the fix list; a *design-level* finding — one that questions a mechanism's cost, operability, trust-model fit, or existence (the remove-the-capability signal in `SKILL.md`), OR one whose remediation would expand scope into an explicitly-deferred concern (the scope-direction / controller-cut-scope signal in `SKILL.md`), recognized on its FIRST appearance (recurrence across rounds is only the reviewer-lane stop/reframe escalation, never the point at which you first classify) — is a **risk-owner decision item**: present keep / delete / narrow / replace — or, only after the current-phase-impact test and compound split (see `Findings, autonomous budget, and human authority`) leave a residual with no current-phase impact, cut-scope-to-phase-boundary — to the user/maintainer BEFORE investing hardening rounds in the questioned mechanism. Executing "fix the findings" by hardening a mechanism whose design finding was never decided pays the hardening cost twice — once to build, once to tear down.
50
+ **Partition findings before fixing: mechanical fixes vs design decisions.** When a round returns findings, classify each before starting fix work: a *mechanical* finding (bug, missing check, wrong value) goes on the fix list; a *design-level* finding — one that questions a mechanism's cost, operability, trust-model fit, or existence (the remove-the-capability signal in `SKILL.md`), OR one whose remediation would expand scope into an explicitly-deferred concern (the scope-direction / controller-cut-scope signal in `SKILL.md`), recognized on its FIRST appearance (recurrence across rounds is only the reviewer-lane stop/reframe escalation, never the point at which you first classify) — is a **risk-owner decision item**: present keep / delete / narrow / replace — or, only after the current-phase-impact test and compound split (see `Findings and dispositions`) leave a residual with no current-phase impact, cut-scope-to-phase-boundary — to the user/maintainer BEFORE investing hardening rounds in the questioned mechanism. Executing "fix the findings" by hardening a mechanism whose design finding was never decided pays the hardening cost twice — once to build, once to tear down.
51
51
 
52
52
  **A finding-fix that widens or hardens a validator needs a false-positive sweep against an explicit accept-set oracle.** Before landing a fix that strengthens a check in response to a finding (advisory → blocking, a narrower accept-set, a new rejection class), set-diff the strengthened check against an authoritative universe of legitimate content — a finite schema, the surface's own inventory of shapes, or the existing corpus the check will scan; when no finite oracle exists, name the representative legitimate classes checked, the sampling boundary, and the uncovered residual as explicit risk. Listing the salient examples that came to mind is not a sweep — the fix must state its precision exposure, not only close the recall gap. Failure shape: a "block malformed ledger rows" fix that would have false-positived on the register's other legitimate table shapes, reverted one round later.
53
53
 
@@ -59,7 +59,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
59
59
  - **Independent oracle.** Where the property has no executable test — rule text, a register row, a doc — the enumeration is discharged only by an **independent oracle**: name the concrete observation that would contradict the property, say where that observation lives (the owner file and line, the primary source, the command whose output would differ), and go look. **Validate the oracle before trusting its verdict**: a check that returns "clean" because it looked in the wrong place, matched case-sensitively, used too narrow a pattern, or swallowed an error is indistinguishable from a passing property, and it fails in the dangerous direction. Before accepting a clean result you must PROVE THE CHECK CAN FAIL — point it at something you know is broken and watch it report that. Confirming it enumerated the inputs you meant is a necessary extra step, never a substitute: correct inputs say nothing about whether the predicate detects a mismatch or whether a non-zero exit was swallowed, so a check that can only ever say clean passes that weaker test. An unvalidated oracle is not weaker evidence than an imagined mutation; it is the same thing wearing a command prompt.
60
60
  - **Dimension walk.** Adding cases inside an axis you already had buys nothing against one you did not: the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering) before values — `testing-strategy` owns that list and the precision-row obligation that goes with it. A walk whose rows are all imagined mutations is exhortation wearing a checklist's clothes.
61
61
  - **Re-owe after fixes.** Whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration — that newly-added mechanism is the most dangerous line in the diff, because it has no test yet and you wrote it with your attention on the defect it repairs. "The whole enumeration" includes the **pre-cover axes sweep** (concurrency & lifecycle above all): remediation text written mid-round re-owes the draft-time axes BEFORE the candidate goes back to the reviewer, because a fix written with attention on one defect systematically re-opens the same blind-spot axes the original draft missed. And the loop has an escalation point: when the same blind-spot axis or finding class supplies findings in a **third** round, stop the per-finding loop and run one full-matrix implementer self-enumeration (the artifact's own states × failure points × orderings × residues × cross-references) on the current candidate before any further external round — letting the reviewer surface one hole per round is the reviewer-as-defect-finder failure at its most expensive (observed shape: a multi-round program burned twenty-plus single-finding rounds on one axis family; the one full-lifecycle enumeration, run at the maintainer's correction, found the remaining holes in a single batch).
62
- - **Classify before fixing.** Persist one transition row per finding class: a stable semantic class key, its root-cause predicate, affected surface, and one disposition per occurrence. Each occurrence names the SHA-256 of its controller receipt plus the canonical JSON SHA-256 of a finding that actually appears in that receipt; the closeout must classify every controller finding exactly once. New wording or a new file is not a new class when the predicate is the same. Root-cause predicates must be unique across classes after case-folding and whitespace collapse; **only that exact normalization is mechanical**. It catches cosmetic case/whitespace key splits but does not decide whether differently worded predicates are semantically equivalent, which remains reviewer-contestable. Resolution is per occurrence, not "the last disposition wins": every `fixed`, `accepted_tradeoff`, `pre_existing_out_of_scope`, or `source_refuted` occurrence names one same-directory structured disposition-evidence file plus its SHA-256. That file binds schema version, exact current candidate, its controller receipt and finding, the disposition, a non-empty evidence list, and the ordered current/prior class occurrences it resolves. It must resolve its own occurrence; a later closing disposition that lists only itself leaves an earlier `open` occurrence unresolved until later evidence explicitly includes that exact receipt/finding pair. A `needs_human_decision` occurrence cannot be resolved by any candidate-local disposition evidence and stays unresolved until an authenticated external human/platform decision takes the separately authorized path. A controller finding disproved by first-hand source or failure-path evidence records `source_refuted`; this is a classification of an invalid finding, not a fourth disposition for a valid P0/P1. The validator binds every JSON input using duplicate-key rejection, plus the evidence file, digest, candidate, occurrence, disposition, and transition links; it does not judge the evidence text or authenticate tradeoff/scope acceptance. The reviewer or human decision-maker still owns those semantic and authority verdicts, and `ready_for_human_decision` is not approval or merge authority. On the class's third appearance, stop patching individual instances and enumerate the complete authoritative surface against that predicate. The sweep names a same-directory manifest plus its SHA-256; its candidate and exact ordered `searched_set` must match the row, and its unmatched list supplies the recorded count. `ready_for_human_decision` requires zero unresolved occurrences and zero unmatched instances; `continuation_authorization_required` and `baseline_race` retain non-zero unmatched evidence instead of lying about closure. Classification is reviewer-contestable evidence, not an author-controlled escape hatch; splitting one predicate into cosmetic sub-classes does not reset the count. `scripts/validate_extraction_review_state.py` enforces these bindings within the referenced receipt/evidence set.
62
+ - **Classify before fixing.** Group findings by root-cause predicate, not by wording or file: new wording or a new file is not a new class when the predicate is the same, and splitting one predicate into cosmetic sub-classes does not reset the count. Record one disposition per finding (fixed, accepted with who/when/why, pre-existing with evidence, or source-refuted with the first-hand evidence that disproves it — a classification of an invalid finding, not a fourth disposition for a valid P0/P1). A finding that needs a human decision stays open until that human decides. On a class's third appearance, stop patching individual instances and enumerate the complete surface against that predicate; record the searched set and what it found next to the disposition. Classification is reviewer-contestable evidence, not an author-controlled escape hatch.
63
63
  - **Graded verdict shape.** When the assessed reality is multi-dimensional or partial (capability, coverage, feasibility, quality, completion), collapsing it into one binary verdict — "done/not-done", "possible/impossible", "all correct/all wrong" — is the over-broad-absolute axis applied to your own claim layer: the swing to whichever pole feels safest to assert misrepresents a distribution, and the opposite-pole absolute ("structurally impossible", "nothing works") is the SAME defect as an unearned "done", not a humbler one. Report per-dimension status — what's strong, what's weak, what wasn't checked — with the confidence each part actually earned; and where a binary gate genuinely applies (a pass/fail check, a blocked/allowed decision), still give the clear top-line verdict after the per-dimension basis — calibration is not hedged mush. A user correcting your answers as too absolute ("每次都很绝对") is this defect's recurrence signal, same escalation as the `SKILL.md` rule states.
64
64
  - **Honesty (descriptive, not permissive).** This is recognition-dependent salience, not a mechanical gate — an agent that doesn't notice it is done-claiming cannot self-fire it; the mechanical backstops remain the closeout `interim` gates + user-signal escalation. "I didn't notice I was claiming done" does NOT waive the rule — any non-trivial completion/coverage/convergence wording must carry clean-pass evidence or an explicit interim/downscope disposition *before* you emit it. The rule targets completion/coverage/convergence assertions on work whose failure a check could catch, and never narrows the mandatory dual-track challenge (it is the always-on generalization of *self-audit to convergence*, not a replacement for the gate).
65
65
 
@@ -198,7 +198,7 @@ This is the largest class in the round-059 corpus by a wide margin. The counts b
198
198
  | Generator-owned shared skill regenerated through its tool | required independent review; generator validation is additional, not a substitute | required for any non-wording shared-skill change |
199
199
  | Generator-owned non-shared skill regenerated through its tool | required (the generator's own) | required if rules changed |
200
200
 
201
- Only the challenge pass may be skipped, and only when this table marks challenge not required; record an explicit `challenge: not-required, reason: ...` row in the validation log. Independent review for shared-skill changes has no skip row. A required review or challenge that is missing, inconclusive, or skipped blocks commit and landing; the work may only be reported as an uncommitted interim checkpoint until the required pass succeeds.
201
+ Only the challenge pass may be skipped, and only when this table marks challenge not required; record an explicit `challenge: not-required, reason: ...` row in the validation log. Independent review for shared-skill changes has no skip row. A required review or challenge that is missing, inconclusive, or skipped blocks commit and landing (a local unpushed checkpoint commit on an isolated worktree branch aside, per `SKILL.md`); the work may only be reported as an interim checkpoint until the required pass succeeds.
202
202
 
203
203
  **Canonical wording-only criterion + the deterministic scope check (challenge-skip gate).** This is the single source of truth the L0/L1/L2 risk view (`l0-l1-l2-routing.md`) defers to; do not redefine it elsewhere. An edit qualifies as `wording-only` — and so may take the challenge-not-required row above — **only when BOTH** of the following hold, never on the author's say-so:
204
204
 
@@ -256,32 +256,11 @@ A stale base puts the upstream's newer fixes into the packet **reversed**, so th
256
256
  4. **Independent of the candidate.** The base comes from the landing-target authority, outside candidate-controlled inputs. Immutability is not independence — a candidate can commit a tuned fixture and cite its perfectly immutable SHA.
257
257
  5. **Contained in HEAD.** `git rev-list --count HEAD..<pinned SHA>` must be `0`, and nothing above implies it: the count passes while a wrapper builds from a stale local `dev` or an old SHA you supplied (a property of HEAD, not of the packet), and a correctly pinned current tip still shows the target's newer commits as *reversals* once the candidate branch has diverged. Count against the **branch tip** — not the derived base commit or a `merge-base HEAD <upstream>` (ancestors of HEAD by construction, so `0` for free), not the *local* branch (the fetch did not advance it), not a *different* branch than the real base. A non-zero count is not a note-and-continue: integrate the target first (`worktree-isolation` owns that sequence and its stale-overwrite hazard), then re-pin and rebuild. **Scoping the packet to the merge base instead does not discharge this** — a three-dot range hides the target's newer commits rather than reversing them, which fixes the reviewer-facing artifact and leaves the real gap: the candidate was never exercised against the code it will land on top of. That is the right bound for reviewing what an author wrote; it is the wrong bound for a landing candidate, whose artifact is the merge.
258
258
 
259
- Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, the review is not landing evidence until the base is re-pinned and the round re-run. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
260
-
261
- When that check detects drift, rebuild instead of improvising: stop the reviewer
262
- lane; preserve the named candidate manifest/patch and the caller-owned ordered evidence rows;
263
- integrate the newly attested target in an isolated worktree; reapply only the
264
- named candidate paths/rows; regenerate derived artifacts; compare the resulting
265
- path set with the manifest; rerun the selected tests; then re-pin and rebuild the
266
- packet. Never use a broad reset plus `add -A`, which can silently absorb ambient
267
- work. Keep every base attestation in one ledger scoped from the first packet
268
- until landing or a human scheduling decision; repinning, retrying, or rebuilding
269
- does not reset it. Each row uses one remote/ref, a contiguous sequence, a strictly
270
- increasing RFC3339 confirmation time, the corresponding controller-receipt hash
271
- when a round consumed it, and a same-directory hash-bound file containing the
272
- canonical raw `ls-remote` line. A second ordered SHA change (A→B→C or A→B→A) is
273
- the second drift and must terminate the lane as `baseline_race`, with the
274
- unreviewed delta; the drift row and every later row must not map another
275
- controller receipt. For every non-race closeout — `ready_for_human_decision`
276
- and `continuation_authorization_required` alike — the final controller receipt
277
- must consume the latest attested SHA; later same-SHA live rechecks are allowed,
278
- but an unconsumed newer SHA is not reviewed evidence (a post-final-round drift
279
- belongs in the next round's ledger, not appended unconsumed to this one). Do not open another
280
- automatic rebuild after the second drift. The state validator counts these
281
- transitions and receipt/base associations inside the complete referenced row
282
- set. Keep independent work moving while a human chooses a landing window.
283
-
284
- **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks the five live-Git properties (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section can still pass a stale base. The v3 closeout validator makes referenced evidence tampering and broken round association detectable, but it cannot prove that a caller supplied every historical attestation or that the recorded `ls-remote` output is still current; candidate-local receipts are consistency evidence, not remote authority or an append-only log. Therefore **each review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is `base-unattested`, not landing evidence**, the caller retains the complete history across rebuilds, and item 2's live authority query is still repeated immediately before landing. Mechanising the live checks and history retention into a trusted wrapper/platform is follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
259
+ Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, handle the drift as the next paragraph says before landing. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
260
+
261
+ When that check detects drift, integrate the new target in the worktree (`worktree-isolation` owns the sequence), rerun the affected tests, and continue. A rebase that does not conflict with the candidate's own paths does not re-open the review; a conflicting one goes into the post-review delta (see the extraction review lane below). Never use a broad reset plus `add -A`, which can silently absorb ambient work.
262
+
263
+ **The caller owns these checks.** No wrapper checks the five live-Git properties (`review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so record the base next to each pass — remote, ref, SHA, confirmation moment — and repeat item 2's live authority query immediately before landing. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
285
264
 
286
265
  > **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
287
266
  >
@@ -295,13 +274,13 @@ set. Keep independent work moving while a human chooses a landing window.
295
274
 
296
275
  Primary reviewer failure is a remediation branch only when the owning gate classifies it as candidate-local. Use this ladder separately for the review lane and challenge lane:
297
276
 
298
- 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` from round 1; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
277
+ 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` once per pass; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
299
278
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
300
279
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
301
280
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
302
281
  5. A fallback result satisfies only the exact lane it ran, and only when it meets the same evidence bar — do not call a manual ad-hoc run "fallback review". Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence. If no candidate returns a conclusive verdict, keep the work `interim`.
303
282
 
304
- - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
283
+ - **Run the two passes single-shot.** A non-wording extraction's review and challenge are two separate `scripts/extraction_review_gate.sh` calls with no review-chain id; the wrapper refuses chain flags. A strictly proven wording-only change instead uses the proof-bound single-review exception. Controller options and the plan schema are owned by `code-review/SKILL.md` and `code-review/references/staged-review-contract.md`.
305
284
 
306
285
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
307
286
 
@@ -471,7 +450,7 @@ If the pass returned no findings within seconds, treat it as failed (likely a pr
471
450
 
472
451
  ## Iterating: challenge → fix → re-challenge
473
452
 
474
- A single challenge pass is not always enough. Fix-ups can introduce new bugs, and challenge passes have stochastic depth — what one pass missed, the next may surface. Plan for multiple rounds when the change is non-trivial.
453
+ A later challenge pass is not automatic (see the extraction review lane below). When a human requests one, or a recorded redesign owes one, the rules in this section apply to it.
475
454
 
476
455
  ### The pattern
477
456
 
@@ -481,7 +460,7 @@ Each round produces three classes of finding:
481
460
  2. **New issues introduced by the previous round's fixes** — e.g. a fix narrows a regex but the narrowed version misses a real case; a fix moves a routing pointer but the new location creates a different collision.
482
461
  3. **Pre-existing issues codex notices on second look** — often P2/P3, often deferrable, but worth recording.
483
462
 
484
- After applying fixes from round N, re-run the challenge. The next round should produce strictly fewer findings AND no new P0/P1 from the round-N fixes themselves. If round N+1 surfaces a P0 the round-N fix introduced, the fix was wrong — revert or redesign before continuing.
463
+ If a later pass surfaces a P0 that an earlier fix introduced, the fix was wrong — revert or redesign before continuing.
485
464
 
486
465
  ### Do not bias the re-challenge (gate integrity)
487
466
 
@@ -506,7 +485,7 @@ Neutral, non-leading context is encouraged, not withheld (starving the reviewer
506
485
 
507
486
  (Self-extracted: an agent-runtime reference in this tree was reported CONVERGED off a fix-list-primed re-challenge; an unbiased re-run surfaced real remaining P1s — partial-stream double-dispatch, cancellation-vs-error path split, finalization-vs-idle-wake race, abort-cleanup deadlock.)
508
487
 
509
- ### Findings, autonomous budget, and human authority
488
+ ### Findings and dispositions
510
489
 
511
490
  A candidate may be claimed review-ready only when **every** remaining P0/P1 finding has a disposition — there are exactly three, each evidence-backed, not the author's word. Missing disposition blocks the readiness claim and the next external review; it does not stop implementation or unrelated work:
512
491
 
@@ -520,107 +499,57 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
520
499
 
521
500
  "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*. Do NOT iterate to zero *findings* either — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing, and forcing the count to zero either over-corrects or scope-creeps; the bar is zero *undispositioned* P0/P1, which differs from both *zero findings* and *no new P0/P1 appeared*.
522
501
 
523
- **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
524
-
525
- The initial independent review plus Agent-initiated challenges share a **bounded external-review sequence of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. An authorized task includes its necessary fixes, tests and review by default; a sequence limit triggers the checkpoint below, not a new permission request. Candidate edits, commits, rebases, amended plans or renamed slices never erase cumulative spending or broaden that task authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
526
-
527
- Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **Each extraction receipt sequence spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There that receipt sequence ends. Necessary further review follows the checkpoint and original-task authority rule below in another bounded sequence, with complete cumulative history retained. Record a later round as human-requested only when a human actually requested that round. Unused generic capacity alone never justifies another call. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). An integration branch that accumulated several reviewed rounds is promoted as one pull request without a new ledger: when neither a single ledger nor a manifest binds the promotion, the gate walks HEAD's first-parent chain down to the first commit already on the target and rebinds each round merge in a detached checkout of its second parent against its first parent, judged with the landing tree's own controller and validator rather than the round's (a round could carry a hollowed validator that a later round restores); a merge whose second parent is already on the target is a sync merge and owes nothing; every step must be exactly the automatic merge of its parents (a hand resolution or an extra file in the merge commit is refused as unreviewed), a non-merge commit on the chain is refused, and the chain is consulted only for the default path set. Rounds that appended to the same register therefore no longer force the promotion to be split by round. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
528
-
529
- **Self-hosted chains break on owner edits; sum rounds across chains within each sequence and retain cumulative spending across sequences.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break still consumes its sequence's budget; a later sequence requires the checkpoint below and retains every earlier round. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
530
-
531
- 1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and requires a recorded recovery checkpoint before a fresh bounded sequence. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
532
- 2. **Sum spent rounds across all chains before opening one more; each sequence's cap is three rounds across two chains.** Record every prior external round — every sequence and chain, finished or broken — and the cumulative count in the caller-owned task artifact. Within one sequence, the only restart is the single succession challenge opened with `--predecessor-chain-result-file`; never append an over-budget receipt or clear earlier spending. At the cap, or when remaining rounds cannot fund that sequence's closeout floor, run the checkpoint before starting another bounded sequence under existing task authority.
533
- 3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
534
- 4. **At the cap — or at effective exhaustion — finish disposition and reassess the method before another sequence.** On the unchanged candidate, when review and challenge are conclusive and every finding occurrence is source-refuted by first-hand evidence, run the local `complete --finding-dispositions-file` path in the [staged contract](../../code-review/references/staged-review-contract.md#mechanical-self-review-gate). Preserve original findings and the full history; an empty model verdict is not required after evidenced refutation. Otherwise, apply or disposition the final batch, name every post-review fix and unreviewed delta in the MR/PR description, and retain `continuation_authorization_required` or an honest interim record. Unreviewed changes and unresolved risks still need their applicable review or human decision. First trace findings to source, run targeted tests and deep self-review, then fix the evidenced cause, improve missing packet context or change the failed review/diagnostic method. Name the specific remaining verification before starting another necessary bounded sequence under the rule below. A reviewer cap alone never asks the user to renew task authority. Never repeat calls solely to obtain zero findings, report an unreviewed batch as reviewed, or reset cumulative spending.
535
-
536
- A strictly proven wording-only change has no convergence loop: it uses one
537
- generic `code-review` pass, records the independent-review row and the
538
- challenge-not-required proof, and does not create a schema-v3 multi-round
539
- terminal ledger. This exception does not apply to frontmatter, routing,
540
- validation, acceptance, example, owner or behavior changes.
541
-
542
- This budget bounds each reviewer sequence. It does **not** stop implementation, tests, debugging or deep self-review, and reaching it requires a method checkpoint before necessary review continues under the existing task scope. Explicit user cost, round-count and stop limits still govern:
543
-
544
- - A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
545
- - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
546
- - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
547
- - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
548
- - **`continuation_authorization`** must first be checked against the original task authorization: necessary in-scope fixes, tests and review are already authorized by default. At a sequence checkpoint, record `continuation_basis=existing-task-scope` in the caller-owned task artifact, with the original authorization reference and scope, the reason another bounded sequence is needed, changed method or added evidence, cumulative rounds, and links between the old sequence's terminal evidence and the new sequence. Each new sequence uses fresh current-candidate bindings and preserves every prior receipt, focus, finding and disposition; no CLI flag or runtime receipt field is added. The existing per-sequence format, timeout and validation bounds remain unchanged. This is inherited task authority, not a new human request for each round: never relabel these calls as newly human-requested or erase earlier spending. Ask only for scope or authority the original task lacks, an explicit user limit that prevents the next action, or a genuine unresolved product/design/risk decision; continue independent authorized work. Continuation waives no review, test or evidence obligation and grants no merge, publication or risk-acceptance authority. Never infer a lane waiver from silence or from authorization to continue.
502
+ **A convergence or closure declaration must be written falsifiably.** Name the commit it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
549
503
 
550
- When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
504
+ ### The extraction review lane: one review, one challenge
551
505
 
552
- The mechanical reminder is `self_review_gate`, not prose alone. It records outstanding and satisfied triggers, the narrow blocked actions, and the productive actions that remain allowed. It fires before external review, after findings, after a tracked candidate change, on risk/scope escalation, at the post-budget checkpoint, and before a completion claim. A final passed review stays `completion_gated=true` until the exact-candidate local `complete` checkpoint validates the new deep-self-review plan; this checkpoint invokes no reviewer and grants no human authority.
506
+ A non-wording shared-skill change owes exactly two external passes: one independent review and one adversarial challenge, each a single-shot `scripts/extraction_review_gate.sh` call (`--mode review`, then `--mode challenge`). There is no tracked review chain, no round budget, no succession round and no closeout ledger, and the two passes are not bound to one candidate hash: the challenge normally runs after the review's fixes are applied.
553
507
 
554
- In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
508
+ 1. Run the full test lane once on the candidate before the review. After the review, a fix confined to one suite reruns that suite plus `check-ccl-skills.sh`; a fix to a controller, a contract or a shared gate reruns the full lane. CI reruns everything on the pull request either way.
509
+ 2. Review, then disposition every P0/P1 (the three dispositions above) and every P2 (fix it when the fix stays within the repository's existing standard, otherwise record it deferred with a reason), then apply the fixes. Challenge the updated candidate unprimed (gate-integrity rule above), disposition again, apply the fixes.
510
+ 3. Record both passes in the round's `evidence/` directory: each pass's controller result JSON, the commit it reviewed, and one disposition line per P0/P1 (format below). CI refuses a pull request that changes `skills/` or `hooks/` without at least one conclusive review result there (`scripts/check_review_evidence_present.py`); it checks presence only, never which candidate a result reviewed, and it does not check the challenge — that obligation stays with this lane.
511
+ 4. **Every post-review delta gets a delta pass, run by the Agent, never left to a human reader.** Everything committed after the last pass's reviewed commit is the post-review delta. When it changes anything other than non-executable record files in the round's own `evidence/` directory (controller results, disposition notes) — a P0/P1 fix, a P2 fix, a late edit, a register row, an executable probe, a rebase that is not path-disjoint — run a delta pass on it before claiming the round ready. The pull-request description lists each pass and the commit it reviewed, for traceability; nobody is expected to re-review the delta by hand.
512
+ 5. **A delta pass reviews only the delta.** Its packet is the delta from the reviewed commit — pass `--base <reviewed commit>`, which binds it to the worktree and records the local receipt the pull-request hook reads — plus, for a fix, the original finding verbatim as an open item, asking for any P0/P1 in that delta — never a fix-claim (gate-integrity rule above). A new P0/P1 in the delta is fixed and gets one more delta pass. After five delta passes a still-open P0/P1 is not a human decision: revert the change that introduced it, or mark the pull request blocked and do not report it ready. Only an exact rollback of that change to a previously accepted state — the base or a version a pass reviewed — with its dependent changes owes no further pass; any other deletion leaves the pull request blocked. A delta pass never re-reviews unchanged content and never voids an earlier pass. Any pass uses the same adversarial framing; a softer prompt after fixes defeats it. P2/P3 findings owe a disposition (step 2), not a pass of their own.
555
513
 
556
- At the final round of a bounded sequence, perform the checkpoint before another necessary sequence. If findings remain:
514
+ A rebase owes nothing only when it is path-disjoint: `git diff --name-only <old base> <new base>` shares no path with the candidate's changed files. When the target's new commits touched a file the candidate also touches — with or without a textual conflict — the combination was never reviewed, so the delta pass covers those files.
557
515
 
558
- - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`; these legacy fields describe the spent sequence, so check existing task authority before asking for permission;
559
- - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
560
- - mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
561
- - enter `awaiting_human` only for an actual missing decision or authority after available authorized work and checkpoint recovery are exhausted. A sequence cap alone is not that blocker.
516
+ No step of this lane waits on a human reading the diff. Human authority is unchanged where it is exercised: a human may request another pass, stop a live pass or the whole iteration, or merge. A `review_waiver` clears only the review lane for the named change; a `merge_authorization` is the human's final merge decision, with every CI failure still reported and never rewritten as passed. The Agent cannot grant either to itself.
562
517
 
563
- The terminal checkpoint is an extraction closeout record, not a state emitted by
564
- `review_gate.py`, and its schema-v3 state is derived from evidence rather than
565
- trusted as an author assertion. Schema-v2 closeout ledgers are rejected rather
566
- than silently reinterpreted under the breaking occurrence/evidence shape. The
567
- ledger and every referenced controller, completion, base, and sweep file live
568
- in one directory and carry SHA-256s. The validator walks the ordered schema-v3
569
- controller chain (same chain and scope,
570
- review then contiguous challenges, packet=candidate, complete prior-result hash
571
- prefix, the wrapper-fixed `challenge_budget`) and binds every closeout candidate to its
572
- last receipt. Ready requires at least review + challenge; a second base drift may
573
- stop as race immediately after round 1 rather than spending an illegal challenge
574
- after the terminal predicate already fired.
575
- It ends in exactly one state:
518
+ **Task authority.** An authorized task includes its necessary in-scope fixes, tests and review passes by default, so the Agent never asks permission per pass. It never relabels an Agent-run pass as newly human-requested, and never infers a lane waiver from silence or from authorization to continue. Ask only for scope or authority the original task lacks, or when an explicit user limit prevents the next action. Running passes waives no review, test or evidence obligation and grants no merge, publication or risk-acceptance authority.
576
519
 
577
- - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance. For `completion_basis=source_refuted_findings`, add the same-directory `finding_dispositions: {file, sha256}` reference. Its digest must match the completion receipt; every original occurrence must also retain its separately bound `source_refuted` class evidence with exactly the same evidence array. This proves consistency and coverage, not the truth of the reasoning or permission to accept risk.
578
- - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation. Preserve this legacy schema value and first check the original task scope; it does not unconditionally require another user grant.
579
- - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
520
+ When a pass returns findings, hand them to the implementer before any further external call: verify each failure path, classify it (local fix, false positive, deferred risk, human decision), and record targeted self-review plus tests. Do not use the reviewer as the primary defect finder, and never repeat a pass solely to reach zero findings.
580
521
 
581
- Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
582
- the state. This proves internal consistency and coverage of the files the ledger
583
- references. It does **not** authenticate that no earlier receipt/attestation was
584
- omitted and does not replace the live remote recheck above; the caller still owns
585
- complete-history retention until a trusted platform owns it. An exhausted budget,
586
- stale review, omitted evidence, or unknown lane state is never represented as
587
- convergence.
522
+ In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause or recurring findings trigger a method change, redesign or parked decision item; they never auto-stop unrelated runnable work.
588
523
 
589
- ### Concrete cadence
524
+ A strictly proven wording-only change has no challenge: it uses one generic `code-review` pass and records the challenge-not-required proof (table above). A release version bump that changes only version fields is not a shared-skill change and owes no external pass; the release-version gate in CI is its check.
590
525
 
591
- For a focused single-skill change:
592
- - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
593
- - **Round 2 — challenge**: after implementer triage — fixes stay HELD: applying any fix before this round breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt.
594
- - **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the sequence's final external round; its findings feed the method/authority checkpoint before any necessary next bounded sequence, and the batch lands MR/PR-listed.
526
+ **Why there is no identity binding.** An earlier design bound every pass to one packet hash and refused a landing whose tree differed from the last reviewed one. In a skill repository nearly every fix edits the reviewed owner, so nearly every fix voided the recorded passes and restarted the sequence; the binding then accumulated budget accounting, succession rounds, a partition manifest for large candidates and a first-parent rebinding walk for promotions, and single rounds landed only their last two of many external passes. The binding was introduced to close a consistency gap in its own cadence, not after an unreviewed change shipped a defect. Its protection — nothing lands that no reviewer read — is kept by the delta pass over every post-review change, run by the Agent; what is dropped is the hash identity and the restart it forced. If an unreviewed post-review change ever ships a defect, that incident is the evidence for a narrower mechanical check sized to it.
595
527
 
596
- Broad extractions use the same per-sequence bounds, round 3 included on the same condition. Continue necessary implementation and checkpoint-qualified review within the original task scope; retain cumulative history rather than resetting the task.
528
+ ### Recording the passes
597
529
 
598
- ### Anti-patterns
530
+ Add these rows to the validation log:
599
531
 
600
- - **Landing a fix batch no round ever saw**. The round-1 fix-up itself may introduce bugs, so a batch that moved the candidate owes the succession challenge of round 3 above — the earlier absolute ("always re-challenge after a non-trivial fix-up") was unreachable while the budget was two rounds, and an unreachable obligation reads as satisfied. A candidate the batch did not move owes nothing: the condition is the candidate hash, not the author's sense of how big the fix was.
601
- - **Iterating external review until zero findings**. Stop blind repetition at the configured sequence limit and run the checkpoint. Stabilized or repeated findings require source disposition and a method/design or evidence change before necessary review continues under existing task authority; a genuine decision blocks only its dependent slice.
602
- - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
603
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another review round only when recorded verification needs and current risk call for it, or when a human explicitly requests one. A fresh bounded sequence never resets task history or cumulative spending — retain both and apply the self-hosted-chain checkpoint above.
604
- - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
532
+ ```
533
+ ## Review pass
534
+ - Reviewed commit: <sha>
535
+ - Findings: N total (a P0 / b P1 / c P2)
536
+ - R0 evidence: <same value menu as the single-pass rows above>
537
+ - Dispositions: <one line per P0/P1: fixed | accepted (who/when/why) | pre-existing (evidence)>
605
538
 
606
- ### Recording the loop
539
+ ## Challenge pass
540
+ - Reviewed commit: <sha>
541
+ - Findings: N total (a P0 / b P1 / c P2)
542
+ - R0 evidence: <same value menu>
543
+ - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator | n/a>
544
+ - Dispositions: <one line per P0/P1>
607
545
 
608
- Add one row per round to the validation log:
546
+ ## Delta pass N (after any post-review change)
547
+ - Reviewed commit: <sha>; packet: git diff <previous reviewed commit>..<sha>
548
+ - Findings / Dispositions: <as above>
609
549
 
550
+ ## Post-review delta
551
+ - <git diff --stat last-reviewed-commit..HEAD>, one line per change
610
552
  ```
611
- ## Challenge pass — round N (codex exec adversarial)
612
- - Diff scope: <files / commit range / sha>
613
- - Findings: N total (a P0 / b P1 / c P2)
614
- - R0 evidence: <same value menu as the review-pass row above>
615
- - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
616
- - New since prior round: <count> (subset of above; flag round-introduced bugs)
617
- - Stabilized: <list of findings carried over without change>
618
- - Applied: M fixes (commit: <sha>)
619
- - Deferred: <list with reason>
620
- - Decision: continue implementation / park dependent slice / await human / human stop, because <reason>
621
- ```
622
-
623
- A complete dual-track-validated change names every round explicitly. Skipping rounds without recording the decision is the same as not running them.
624
553
 
625
554
  ## What does NOT count as dual-track
626
555