@ccoalm/ccl-skills 0.17.0 → 0.18.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +8 -7
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +22 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-review-covers-head.sh +128 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-untracked-background.sh +53 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_review_covers_head.sh +104 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_untracked_background.sh +75 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +9 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +13 -1
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -3
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +401 -21
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +8 -1
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +230 -7
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +1 -1
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +1 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +6 -6
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +48 -119
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +9 -9
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +8 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +2 -2
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check_review_evidence_present.py +122 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +0 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +37 -8
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +102 -20
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +16 -27
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +3 -5
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_review_evidence_present.sh +81 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +130 -310
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +49 -3
  29. package/dist/assets/release.json +55 -45
  30. package/package.json +1 -1
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +0 -1242
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +0 -978
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +0 -1477
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +0 -1183
@@ -19,7 +19,7 @@ table. At closeout:
19
19
 
20
20
  - **Count lane names, never rounds**: enumerate the lanes the slice's own gate section requires, and confirm each one has a recorded outcome.
21
21
  - **A rounds table naming one lane while the gate requires two is `interim`, never landed** — no round count and no green deterministic gate substitutes for a lane that never ran.
22
- - **Where lanes leave evidence in a local store, a chain must exist for THIS slice** — check what the stored packets actually covered rather than assuming one is there.
22
+ - **Where lanes leave evidence in a local store, a stored result must exist for THIS slice** — check what the stored packets actually covered rather than assuming one is there.
23
23
 
24
24
  Observed: a mandatory, fail-closed repository gate landed with nine review-mode
25
25
  rounds recorded, no challenge lane run, and a landing state still reading
@@ -47,7 +47,7 @@ Before invoking either pass, self-audit the candidate to the point where *you* e
47
47
 
48
48
  **A design pivot resets the self-audit obligation.** When the candidate is substantially redesigned mid-gate (a mechanism replaced, a capability torn down, a rewrite beyond the findings being fixed), the earlier self-audit covered the OLD candidate: redo the closure self-audit on the NEW candidate before re-entering the challenge. Sliding from a pivot straight back into challenge → fix → re-challenge is the exact loop this section forbids, and it recurs precisely at pivots because the prior audit feels "already done."
49
49
 
50
- **Partition findings before fixing: mechanical fixes vs design decisions.** When a round returns findings, classify each before starting fix work: a *mechanical* finding (bug, missing check, wrong value) goes on the fix list; a *design-level* finding — one that questions a mechanism's cost, operability, trust-model fit, or existence (the remove-the-capability signal in `SKILL.md`), OR one whose remediation would expand scope into an explicitly-deferred concern (the scope-direction / controller-cut-scope signal in `SKILL.md`), recognized on its FIRST appearance (recurrence across rounds is only the reviewer-lane stop/reframe escalation, never the point at which you first classify) — is a **risk-owner decision item**: present keep / delete / narrow / replace — or, only after the current-phase-impact test and compound split (see `Findings, autonomous budget, and human authority`) leave a residual with no current-phase impact, cut-scope-to-phase-boundary — to the user/maintainer BEFORE investing hardening rounds in the questioned mechanism. Executing "fix the findings" by hardening a mechanism whose design finding was never decided pays the hardening cost twice — once to build, once to tear down.
50
+ **Partition findings before fixing: mechanical fixes vs design decisions.** When a round returns findings, classify each before starting fix work: a *mechanical* finding (bug, missing check, wrong value) goes on the fix list; a *design-level* finding — one that questions a mechanism's cost, operability, trust-model fit, or existence (the remove-the-capability signal in `SKILL.md`), OR one whose remediation would expand scope into an explicitly-deferred concern (the scope-direction / controller-cut-scope signal in `SKILL.md`), recognized on its FIRST appearance (recurrence across rounds is only the reviewer-lane stop/reframe escalation, never the point at which you first classify) — is a **risk-owner decision item**: present keep / delete / narrow / replace — or, only after the current-phase-impact test and compound split (see `Findings and dispositions`) leave a residual with no current-phase impact, cut-scope-to-phase-boundary — to the user/maintainer BEFORE investing hardening rounds in the questioned mechanism. Executing "fix the findings" by hardening a mechanism whose design finding was never decided pays the hardening cost twice — once to build, once to tear down.
51
51
 
52
52
  **A finding-fix that widens or hardens a validator needs a false-positive sweep against an explicit accept-set oracle.** Before landing a fix that strengthens a check in response to a finding (advisory → blocking, a narrower accept-set, a new rejection class), set-diff the strengthened check against an authoritative universe of legitimate content — a finite schema, the surface's own inventory of shapes, or the existing corpus the check will scan; when no finite oracle exists, name the representative legitimate classes checked, the sampling boundary, and the uncovered residual as explicit risk. Listing the salient examples that came to mind is not a sweep — the fix must state its precision exposure, not only close the recall gap. Failure shape: a "block malformed ledger rows" fix that would have false-positived on the register's other legitimate table shapes, reverted one round later.
53
53
 
@@ -59,7 +59,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
59
59
  - **Independent oracle.** Where the property has no executable test — rule text, a register row, a doc — the enumeration is discharged only by an **independent oracle**: name the concrete observation that would contradict the property, say where that observation lives (the owner file and line, the primary source, the command whose output would differ), and go look. **Validate the oracle before trusting its verdict**: a check that returns "clean" because it looked in the wrong place, matched case-sensitively, used too narrow a pattern, or swallowed an error is indistinguishable from a passing property, and it fails in the dangerous direction. Before accepting a clean result you must PROVE THE CHECK CAN FAIL — point it at something you know is broken and watch it report that. Confirming it enumerated the inputs you meant is a necessary extra step, never a substitute: correct inputs say nothing about whether the predicate detects a mismatch or whether a non-zero exit was swallowed, so a check that can only ever say clean passes that weaker test. An unvalidated oracle is not weaker evidence than an imagined mutation; it is the same thing wearing a command prompt.
60
60
  - **Dimension walk.** Adding cases inside an axis you already had buys nothing against one you did not: the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering) before values — `testing-strategy` owns that list and the precision-row obligation that goes with it. A walk whose rows are all imagined mutations is exhortation wearing a checklist's clothes.
61
61
  - **Re-owe after fixes.** Whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration — that newly-added mechanism is the most dangerous line in the diff, because it has no test yet and you wrote it with your attention on the defect it repairs. "The whole enumeration" includes the **pre-cover axes sweep** (concurrency & lifecycle above all): remediation text written mid-round re-owes the draft-time axes BEFORE the candidate goes back to the reviewer, because a fix written with attention on one defect systematically re-opens the same blind-spot axes the original draft missed. And the loop has an escalation point: when the same blind-spot axis or finding class supplies findings in a **third** round, stop the per-finding loop and run one full-matrix implementer self-enumeration (the artifact's own states × failure points × orderings × residues × cross-references) on the current candidate before any further external round — letting the reviewer surface one hole per round is the reviewer-as-defect-finder failure at its most expensive (observed shape: a multi-round program burned twenty-plus single-finding rounds on one axis family; the one full-lifecycle enumeration, run at the maintainer's correction, found the remaining holes in a single batch).
62
- - **Classify before fixing.** Persist one transition row per finding class: a stable semantic class key, its root-cause predicate, affected surface, and one disposition per occurrence. Each occurrence names the SHA-256 of its controller receipt plus the canonical JSON SHA-256 of a finding that actually appears in that receipt; the closeout must classify every controller finding exactly once. New wording or a new file is not a new class when the predicate is the same. Root-cause predicates must be unique across classes after case-folding and whitespace collapse; **only that exact normalization is mechanical**. It catches cosmetic case/whitespace key splits but does not decide whether differently worded predicates are semantically equivalent, which remains reviewer-contestable. Resolution is per occurrence, not "the last disposition wins": every `fixed`, `accepted_tradeoff`, `pre_existing_out_of_scope`, or `source_refuted` occurrence names one same-directory structured disposition-evidence file plus its SHA-256. That file binds schema version, exact current candidate, its controller receipt and finding, the disposition, a non-empty evidence list, and the ordered current/prior class occurrences it resolves. It must resolve its own occurrence; a later closing disposition that lists only itself leaves an earlier `open` occurrence unresolved until later evidence explicitly includes that exact receipt/finding pair. A `needs_human_decision` occurrence cannot be resolved by any candidate-local disposition evidence and stays unresolved until an authenticated external human/platform decision takes the separately authorized path. A controller finding disproved by first-hand source or failure-path evidence records `source_refuted`; this is a classification of an invalid finding, not a fourth disposition for a valid P0/P1. The validator binds every JSON input using duplicate-key rejection, plus the evidence file, digest, candidate, occurrence, disposition, and transition links; it does not judge the evidence text or authenticate tradeoff/scope acceptance. The reviewer or human decision-maker still owns those semantic and authority verdicts, and `ready_for_human_decision` is not approval or merge authority. On the class's third appearance, stop patching individual instances and enumerate the complete authoritative surface against that predicate. The sweep names a same-directory manifest plus its SHA-256; its candidate and exact ordered `searched_set` must match the row, and its unmatched list supplies the recorded count. `ready_for_human_decision` requires zero unresolved occurrences and zero unmatched instances; `continuation_authorization_required` and `baseline_race` retain non-zero unmatched evidence instead of lying about closure. Classification is reviewer-contestable evidence, not an author-controlled escape hatch; splitting one predicate into cosmetic sub-classes does not reset the count. `scripts/validate_extraction_review_state.py` enforces these bindings within the referenced receipt/evidence set.
62
+ - **Classify before fixing.** Group findings by root-cause predicate, not by wording or file: new wording or a new file is not a new class when the predicate is the same, and splitting one predicate into cosmetic sub-classes does not reset the count. Record one disposition per finding (fixed, accepted with who/when/why, pre-existing with evidence, or source-refuted with the first-hand evidence that disproves it — a classification of an invalid finding, not a fourth disposition for a valid P0/P1). A finding that needs a human decision stays open until that human decides. On a class's third appearance, stop patching individual instances and enumerate the complete surface against that predicate; record the searched set and what it found next to the disposition. Classification is reviewer-contestable evidence, not an author-controlled escape hatch.
63
63
  - **Graded verdict shape.** When the assessed reality is multi-dimensional or partial (capability, coverage, feasibility, quality, completion), collapsing it into one binary verdict — "done/not-done", "possible/impossible", "all correct/all wrong" — is the over-broad-absolute axis applied to your own claim layer: the swing to whichever pole feels safest to assert misrepresents a distribution, and the opposite-pole absolute ("structurally impossible", "nothing works") is the SAME defect as an unearned "done", not a humbler one. Report per-dimension status — what's strong, what's weak, what wasn't checked — with the confidence each part actually earned; and where a binary gate genuinely applies (a pass/fail check, a blocked/allowed decision), still give the clear top-line verdict after the per-dimension basis — calibration is not hedged mush. A user correcting your answers as too absolute ("每次都很绝对") is this defect's recurrence signal, same escalation as the `SKILL.md` rule states.
64
64
  - **Honesty (descriptive, not permissive).** This is recognition-dependent salience, not a mechanical gate — an agent that doesn't notice it is done-claiming cannot self-fire it; the mechanical backstops remain the closeout `interim` gates + user-signal escalation. "I didn't notice I was claiming done" does NOT waive the rule — any non-trivial completion/coverage/convergence wording must carry clean-pass evidence or an explicit interim/downscope disposition *before* you emit it. The rule targets completion/coverage/convergence assertions on work whose failure a check could catch, and never narrows the mandatory dual-track challenge (it is the always-on generalization of *self-audit to convergence*, not a replacement for the gate).
65
65
 
@@ -198,7 +198,7 @@ This is the largest class in the round-059 corpus by a wide margin. The counts b
198
198
  | Generator-owned shared skill regenerated through its tool | required independent review; generator validation is additional, not a substitute | required for any non-wording shared-skill change |
199
199
  | Generator-owned non-shared skill regenerated through its tool | required (the generator's own) | required if rules changed |
200
200
 
201
- Only the challenge pass may be skipped, and only when this table marks challenge not required; record an explicit `challenge: not-required, reason: ...` row in the validation log. Independent review for shared-skill changes has no skip row. A required review or challenge that is missing, inconclusive, or skipped blocks commit and landing; the work may only be reported as an uncommitted interim checkpoint until the required pass succeeds.
201
+ Only the challenge pass may be skipped, and only when this table marks challenge not required; record an explicit `challenge: not-required, reason: ...` row in the validation log. Independent review for shared-skill changes has no skip row. A required review or challenge that is missing, inconclusive, or skipped blocks commit and landing (a local unpushed checkpoint commit on an isolated worktree branch aside, per `SKILL.md`); the work may only be reported as an interim checkpoint until the required pass succeeds.
202
202
 
203
203
  **Canonical wording-only criterion + the deterministic scope check (challenge-skip gate).** This is the single source of truth the L0/L1/L2 risk view (`l0-l1-l2-routing.md`) defers to; do not redefine it elsewhere. An edit qualifies as `wording-only` — and so may take the challenge-not-required row above — **only when BOTH** of the following hold, never on the author's say-so:
204
204
 
@@ -256,32 +256,11 @@ A stale base puts the upstream's newer fixes into the packet **reversed**, so th
256
256
  4. **Independent of the candidate.** The base comes from the landing-target authority, outside candidate-controlled inputs. Immutability is not independence — a candidate can commit a tuned fixture and cite its perfectly immutable SHA.
257
257
  5. **Contained in HEAD.** `git rev-list --count HEAD..<pinned SHA>` must be `0`, and nothing above implies it: the count passes while a wrapper builds from a stale local `dev` or an old SHA you supplied (a property of HEAD, not of the packet), and a correctly pinned current tip still shows the target's newer commits as *reversals* once the candidate branch has diverged. Count against the **branch tip** — not the derived base commit or a `merge-base HEAD <upstream>` (ancestors of HEAD by construction, so `0` for free), not the *local* branch (the fetch did not advance it), not a *different* branch than the real base. A non-zero count is not a note-and-continue: integrate the target first (`worktree-isolation` owns that sequence and its stale-overwrite hazard), then re-pin and rebuild. **Scoping the packet to the merge base instead does not discharge this** — a three-dot range hides the target's newer commits rather than reversing them, which fixes the reviewer-facing artifact and leaves the real gap: the candidate was never exercised against the code it will land on top of. That is the right bound for reviewing what an author wrote; it is the wrong bound for a landing candidate, whose artifact is the merge.
258
258
 
259
- Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, the review is not landing evidence until the base is re-pinned and the round re-run. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
260
-
261
- When that check detects drift, rebuild instead of improvising: stop the reviewer
262
- lane; preserve the named candidate manifest/patch and the caller-owned ordered evidence rows;
263
- integrate the newly attested target in an isolated worktree; reapply only the
264
- named candidate paths/rows; regenerate derived artifacts; compare the resulting
265
- path set with the manifest; rerun the selected tests; then re-pin and rebuild the
266
- packet. Never use a broad reset plus `add -A`, which can silently absorb ambient
267
- work. Keep every base attestation in one ledger scoped from the first packet
268
- until landing or a human scheduling decision; repinning, retrying, or rebuilding
269
- does not reset it. Each row uses one remote/ref, a contiguous sequence, a strictly
270
- increasing RFC3339 confirmation time, the corresponding controller-receipt hash
271
- when a round consumed it, and a same-directory hash-bound file containing the
272
- canonical raw `ls-remote` line. A second ordered SHA change (A→B→C or A→B→A) is
273
- the second drift and must terminate the lane as `baseline_race`, with the
274
- unreviewed delta; the drift row and every later row must not map another
275
- controller receipt. For every non-race closeout — `ready_for_human_decision`
276
- and `continuation_authorization_required` alike — the final controller receipt
277
- must consume the latest attested SHA; later same-SHA live rechecks are allowed,
278
- but an unconsumed newer SHA is not reviewed evidence (a post-final-round drift
279
- belongs in the next round's ledger, not appended unconsumed to this one). Do not open another
280
- automatic rebuild after the second drift. The state validator counts these
281
- transitions and receipt/base associations inside the complete referenced row
282
- set. Keep independent work moving while a human chooses a landing window.
283
-
284
- **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks the five live-Git properties (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section can still pass a stale base. The v3 closeout validator makes referenced evidence tampering and broken round association detectable, but it cannot prove that a caller supplied every historical attestation or that the recorded `ls-remote` output is still current; candidate-local receipts are consistency evidence, not remote authority or an append-only log. Therefore **each review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is `base-unattested`, not landing evidence**, the caller retains the complete history across rebuilds, and item 2's live authority query is still repeated immediately before landing. Mechanising the live checks and history retention into a trusted wrapper/platform is follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
259
+ Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, handle the drift as the next paragraph says before landing. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
260
+
261
+ When that check detects drift, integrate the new target in the worktree (`worktree-isolation` owns the sequence), rerun the affected tests, and continue. A rebase that does not conflict with the candidate's own paths does not re-open the review; a conflicting one goes into the post-review delta (see the extraction review lane below). Never use a broad reset plus `add -A`, which can silently absorb ambient work.
262
+
263
+ **The caller owns these checks.** No wrapper checks the five live-Git properties (`review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so record the base next to each pass — remote, ref, SHA, confirmation moment — and repeat item 2's live authority query immediately before landing. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
285
264
 
286
265
  > **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
287
266
  >
@@ -295,13 +274,13 @@ set. Keep independent work moving while a human chooses a landing window.
295
274
 
296
275
  Primary reviewer failure is a remediation branch only when the owning gate classifies it as candidate-local. Use this ladder separately for the review lane and challenge lane:
297
276
 
298
- 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` from round 1; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
277
+ 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` once per pass; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
299
278
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
300
279
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
301
280
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
302
281
  5. A fallback result satisfies only the exact lane it ran, and only when it meets the same evidence bar — do not call a manual ad-hoc run "fallback review". Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence. If no candidate returns a conclusive verdict, keep the work `interim`.
303
282
 
304
- - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
283
+ - **Run the two passes single-shot.** A non-wording extraction's review and challenge are two separate `scripts/extraction_review_gate.sh` calls with no review-chain id; the wrapper refuses chain flags. A strictly proven wording-only change instead uses the proof-bound single-review exception. Controller options and the plan schema are owned by `code-review/SKILL.md` and `code-review/references/staged-review-contract.md`.
305
284
 
306
285
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
307
286
 
@@ -471,7 +450,7 @@ If the pass returned no findings within seconds, treat it as failed (likely a pr
471
450
 
472
451
  ## Iterating: challenge → fix → re-challenge
473
452
 
474
- A single challenge pass is not always enough. Fix-ups can introduce new bugs, and challenge passes have stochastic depth — what one pass missed, the next may surface. Plan for multiple rounds when the change is non-trivial.
453
+ A later challenge pass is not automatic (see the extraction review lane below). When a human requests one, or a recorded redesign owes one, the rules in this section apply to it.
475
454
 
476
455
  ### The pattern
477
456
 
@@ -481,7 +460,7 @@ Each round produces three classes of finding:
481
460
  2. **New issues introduced by the previous round's fixes** — e.g. a fix narrows a regex but the narrowed version misses a real case; a fix moves a routing pointer but the new location creates a different collision.
482
461
  3. **Pre-existing issues codex notices on second look** — often P2/P3, often deferrable, but worth recording.
483
462
 
484
- After applying fixes from round N, re-run the challenge. The next round should produce strictly fewer findings AND no new P0/P1 from the round-N fixes themselves. If round N+1 surfaces a P0 the round-N fix introduced, the fix was wrong — revert or redesign before continuing.
463
+ If a later pass surfaces a P0 that an earlier fix introduced, the fix was wrong — revert or redesign before continuing.
485
464
 
486
465
  ### Do not bias the re-challenge (gate integrity)
487
466
 
@@ -506,7 +485,7 @@ Neutral, non-leading context is encouraged, not withheld (starving the reviewer
506
485
 
507
486
  (Self-extracted: an agent-runtime reference in this tree was reported CONVERGED off a fix-list-primed re-challenge; an unbiased re-run surfaced real remaining P1s — partial-stream double-dispatch, cancellation-vs-error path split, finalization-vs-idle-wake race, abort-cleanup deadlock.)
508
487
 
509
- ### Findings, autonomous budget, and human authority
488
+ ### Findings and dispositions
510
489
 
511
490
  A candidate may be claimed review-ready only when **every** remaining P0/P1 finding has a disposition — there are exactly three, each evidence-backed, not the author's word. Missing disposition blocks the readiness claim and the next external review; it does not stop implementation or unrelated work:
512
491
 
@@ -520,107 +499,57 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
520
499
 
521
500
  "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*. Do NOT iterate to zero *findings* either — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing, and forcing the count to zero either over-corrects or scope-creeps; the bar is zero *undispositioned* P0/P1, which differs from both *zero findings* and *no new P0/P1 appeared*.
522
501
 
523
- **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
524
-
525
- The initial independent review plus Agent-initiated challenges share a **bounded external-review sequence of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. An authorized task includes its necessary fixes, tests and review by default; a sequence limit triggers the checkpoint below, not a new permission request. Candidate edits, commits, rebases, amended plans or renamed slices never erase cumulative spending or broaden that task authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
526
-
527
- Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **Each extraction receipt sequence spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There that receipt sequence ends. Necessary further review follows the checkpoint and original-task authority rule below in another bounded sequence, with complete cumulative history retained. Record a later round as human-requested only when a human actually requested that round. Unused generic capacity alone never justifies another call. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). An integration branch that accumulated several reviewed rounds is promoted as one pull request without a new ledger: when neither a single ledger nor a manifest binds the promotion, the gate walks HEAD's first-parent chain down to the first commit already on the target and rebinds each round merge in a detached checkout of its second parent against its first parent, judged with the landing tree's own controller and validator rather than the round's (a round could carry a hollowed validator that a later round restores); a merge whose second parent is already on the target is a sync merge and owes nothing; every step must be exactly the automatic merge of its parents (a hand resolution or an extra file in the merge commit is refused as unreviewed), a non-merge commit on the chain is refused, and the chain is consulted only for the default path set. Rounds that appended to the same register therefore no longer force the promotion to be split by round. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
528
-
529
- **Self-hosted chains break on owner edits; sum rounds across chains within each sequence and retain cumulative spending across sequences.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break still consumes its sequence's budget; a later sequence requires the checkpoint below and retains every earlier round. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
530
-
531
- 1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and requires a recorded recovery checkpoint before a fresh bounded sequence. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
532
- 2. **Sum spent rounds across all chains before opening one more; each sequence's cap is three rounds across two chains.** Record every prior external round — every sequence and chain, finished or broken — and the cumulative count in the caller-owned task artifact. Within one sequence, the only restart is the single succession challenge opened with `--predecessor-chain-result-file`; never append an over-budget receipt or clear earlier spending. At the cap, or when remaining rounds cannot fund that sequence's closeout floor, run the checkpoint before starting another bounded sequence under existing task authority.
533
- 3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
534
- 4. **At the cap — or at effective exhaustion — finish disposition and reassess the method before another sequence.** On the unchanged candidate, when review and challenge are conclusive and every finding occurrence is source-refuted by first-hand evidence, run the local `complete --finding-dispositions-file` path in the [staged contract](../../code-review/references/staged-review-contract.md#mechanical-self-review-gate). Preserve original findings and the full history; an empty model verdict is not required after evidenced refutation. Otherwise, apply or disposition the final batch, name every post-review fix and unreviewed delta in the MR/PR description, and retain `continuation_authorization_required` or an honest interim record. Unreviewed changes and unresolved risks still need their applicable review or human decision. First trace findings to source, run targeted tests and deep self-review, then fix the evidenced cause, improve missing packet context or change the failed review/diagnostic method. Name the specific remaining verification before starting another necessary bounded sequence under the rule below. A reviewer cap alone never asks the user to renew task authority. Never repeat calls solely to obtain zero findings, report an unreviewed batch as reviewed, or reset cumulative spending.
535
-
536
- A strictly proven wording-only change has no convergence loop: it uses one
537
- generic `code-review` pass, records the independent-review row and the
538
- challenge-not-required proof, and does not create a schema-v3 multi-round
539
- terminal ledger. This exception does not apply to frontmatter, routing,
540
- validation, acceptance, example, owner or behavior changes.
541
-
542
- This budget bounds each reviewer sequence. It does **not** stop implementation, tests, debugging or deep self-review, and reaching it requires a method checkpoint before necessary review continues under the existing task scope. Explicit user cost, round-count and stop limits still govern:
543
-
544
- - A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
545
- - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
546
- - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
547
- - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
548
- - **`continuation_authorization`** must first be checked against the original task authorization: necessary in-scope fixes, tests and review are already authorized by default. At a sequence checkpoint, record `continuation_basis=existing-task-scope` in the caller-owned task artifact, with the original authorization reference and scope, the reason another bounded sequence is needed, changed method or added evidence, cumulative rounds, and links between the old sequence's terminal evidence and the new sequence. Each new sequence uses fresh current-candidate bindings and preserves every prior receipt, focus, finding and disposition; no CLI flag or runtime receipt field is added. The existing per-sequence format, timeout and validation bounds remain unchanged. This is inherited task authority, not a new human request for each round: never relabel these calls as newly human-requested or erase earlier spending. Ask only for scope or authority the original task lacks, an explicit user limit that prevents the next action, or a genuine unresolved product/design/risk decision; continue independent authorized work. Continuation waives no review, test or evidence obligation and grants no merge, publication or risk-acceptance authority. Never infer a lane waiver from silence or from authorization to continue.
502
+ **A convergence or closure declaration must be written falsifiably.** Name the commit it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
549
503
 
550
- When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
504
+ ### The extraction review lane: one review, one challenge
551
505
 
552
- The mechanical reminder is `self_review_gate`, not prose alone. It records outstanding and satisfied triggers, the narrow blocked actions, and the productive actions that remain allowed. It fires before external review, after findings, after a tracked candidate change, on risk/scope escalation, at the post-budget checkpoint, and before a completion claim. A final passed review stays `completion_gated=true` until the exact-candidate local `complete` checkpoint validates the new deep-self-review plan; this checkpoint invokes no reviewer and grants no human authority.
506
+ A non-wording shared-skill change owes exactly two external passes: one independent review and one adversarial challenge, each a single-shot `scripts/extraction_review_gate.sh` call (`--mode review`, then `--mode challenge`). There is no tracked review chain, no round budget, no succession round and no closeout ledger, and the two passes are not bound to one candidate hash: the challenge normally runs after the review's fixes are applied.
553
507
 
554
- In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
508
+ 1. Run the full test lane once on the candidate before the review. After the review, a fix confined to one suite reruns that suite plus `check-ccl-skills.sh`; a fix to a controller, a contract or a shared gate reruns the full lane. CI reruns everything on the pull request either way.
509
+ 2. Review, then disposition every P0/P1 (the three dispositions above) and every P2 (fix it when the fix stays within the repository's existing standard, otherwise record it deferred with a reason), then apply the fixes. Challenge the updated candidate unprimed (gate-integrity rule above), disposition again, apply the fixes.
510
+ 3. Record both passes in the round's `evidence/` directory: each pass's controller result JSON, the commit it reviewed, and one disposition line per P0/P1 (format below). CI refuses a pull request that changes `skills/` or `hooks/` without at least one conclusive review result there (`scripts/check_review_evidence_present.py`); it checks presence only, never which candidate a result reviewed, and it does not check the challenge — that obligation stays with this lane.
511
+ 4. **Every post-review delta gets a delta pass, run by the Agent, never left to a human reader.** Everything committed after the last pass's reviewed commit is the post-review delta. When it changes anything other than non-executable record files in the round's own `evidence/` directory (controller results, disposition notes) — a P0/P1 fix, a P2 fix, a late edit, a register row, an executable probe, a rebase that is not path-disjoint — run a delta pass on it before claiming the round ready. The pull-request description lists each pass and the commit it reviewed, for traceability; nobody is expected to re-review the delta by hand.
512
+ 5. **A delta pass reviews only the delta.** Its packet is the delta from the reviewed commit — pass `--base <reviewed commit>`, which binds it to the worktree and records the local receipt the pull-request hook reads — plus, for a fix, the original finding verbatim as an open item, asking for any P0/P1 in that delta — never a fix-claim (gate-integrity rule above). A new P0/P1 in the delta is fixed and gets one more delta pass. After five delta passes a still-open P0/P1 is not a human decision: revert the change that introduced it, or mark the pull request blocked and do not report it ready. Only an exact rollback of that change to a previously accepted state — the base or a version a pass reviewed — with its dependent changes owes no further pass; any other deletion leaves the pull request blocked. A delta pass never re-reviews unchanged content and never voids an earlier pass. Any pass uses the same adversarial framing; a softer prompt after fixes defeats it. P2/P3 findings owe a disposition (step 2), not a pass of their own.
555
513
 
556
- At the final round of a bounded sequence, perform the checkpoint before another necessary sequence. If findings remain:
514
+ A rebase owes nothing only when it is path-disjoint: `git diff --name-only <old base> <new base>` shares no path with the candidate's changed files. When the target's new commits touched a file the candidate also touches — with or without a textual conflict — the combination was never reviewed, so the delta pass covers those files.
557
515
 
558
- - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`; these legacy fields describe the spent sequence, so check existing task authority before asking for permission;
559
- - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
560
- - mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
561
- - enter `awaiting_human` only for an actual missing decision or authority after available authorized work and checkpoint recovery are exhausted. A sequence cap alone is not that blocker.
516
+ No step of this lane waits on a human reading the diff. Human authority is unchanged where it is exercised: a human may request another pass, stop a live pass or the whole iteration, or merge. A `review_waiver` clears only the review lane for the named change; a `merge_authorization` is the human's final merge decision, with every CI failure still reported and never rewritten as passed. The Agent cannot grant either to itself.
562
517
 
563
- The terminal checkpoint is an extraction closeout record, not a state emitted by
564
- `review_gate.py`, and its schema-v3 state is derived from evidence rather than
565
- trusted as an author assertion. Schema-v2 closeout ledgers are rejected rather
566
- than silently reinterpreted under the breaking occurrence/evidence shape. The
567
- ledger and every referenced controller, completion, base, and sweep file live
568
- in one directory and carry SHA-256s. The validator walks the ordered schema-v3
569
- controller chain (same chain and scope,
570
- review then contiguous challenges, packet=candidate, complete prior-result hash
571
- prefix, the wrapper-fixed `challenge_budget`) and binds every closeout candidate to its
572
- last receipt. Ready requires at least review + challenge; a second base drift may
573
- stop as race immediately after round 1 rather than spending an illegal challenge
574
- after the terminal predicate already fired.
575
- It ends in exactly one state:
518
+ **Task authority.** An authorized task includes its necessary in-scope fixes, tests and review passes by default, so the Agent never asks permission per pass. It never relabels an Agent-run pass as newly human-requested, and never infers a lane waiver from silence or from authorization to continue. Ask only for scope or authority the original task lacks, or when an explicit user limit prevents the next action. Running passes waives no review, test or evidence obligation and grants no merge, publication or risk-acceptance authority.
576
519
 
577
- - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance. For `completion_basis=source_refuted_findings`, add the same-directory `finding_dispositions: {file, sha256}` reference. Its digest must match the completion receipt; every original occurrence must also retain its separately bound `source_refuted` class evidence with exactly the same evidence array. This proves consistency and coverage, not the truth of the reasoning or permission to accept risk.
578
- - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation. Preserve this legacy schema value and first check the original task scope; it does not unconditionally require another user grant.
579
- - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
520
+ When a pass returns findings, hand them to the implementer before any further external call: verify each failure path, classify it (local fix, false positive, deferred risk, human decision), and record targeted self-review plus tests. Do not use the reviewer as the primary defect finder, and never repeat a pass solely to reach zero findings.
580
521
 
581
- Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
582
- the state. This proves internal consistency and coverage of the files the ledger
583
- references. It does **not** authenticate that no earlier receipt/attestation was
584
- omitted and does not replace the live remote recheck above; the caller still owns
585
- complete-history retention until a trusted platform owns it. An exhausted budget,
586
- stale review, omitted evidence, or unknown lane state is never represented as
587
- convergence.
522
+ In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause or recurring findings trigger a method change, redesign or parked decision item; they never auto-stop unrelated runnable work.
588
523
 
589
- ### Concrete cadence
524
+ A strictly proven wording-only change has no challenge: it uses one generic `code-review` pass and records the challenge-not-required proof (table above). A release version bump that changes only version fields is not a shared-skill change and owes no external pass; the release-version gate in CI is its check.
590
525
 
591
- For a focused single-skill change:
592
- - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
593
- - **Round 2 — challenge**: after implementer triage — fixes stay HELD: applying any fix before this round breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt.
594
- - **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the sequence's final external round; its findings feed the method/authority checkpoint before any necessary next bounded sequence, and the batch lands MR/PR-listed.
526
+ **Why there is no identity binding.** An earlier design bound every pass to one packet hash and refused a landing whose tree differed from the last reviewed one. In a skill repository nearly every fix edits the reviewed owner, so nearly every fix voided the recorded passes and restarted the sequence; the binding then accumulated budget accounting, succession rounds, a partition manifest for large candidates and a first-parent rebinding walk for promotions, and single rounds landed only their last two of many external passes. The binding was introduced to close a consistency gap in its own cadence, not after an unreviewed change shipped a defect. Its protection — nothing lands that no reviewer read — is kept by the delta pass over every post-review change, run by the Agent; what is dropped is the hash identity and the restart it forced. If an unreviewed post-review change ever ships a defect, that incident is the evidence for a narrower mechanical check sized to it.
595
527
 
596
- Broad extractions use the same per-sequence bounds, round 3 included on the same condition. Continue necessary implementation and checkpoint-qualified review within the original task scope; retain cumulative history rather than resetting the task.
528
+ ### Recording the passes
597
529
 
598
- ### Anti-patterns
530
+ Add these rows to the validation log:
599
531
 
600
- - **Landing a fix batch no round ever saw**. The round-1 fix-up itself may introduce bugs, so a batch that moved the candidate owes the succession challenge of round 3 above — the earlier absolute ("always re-challenge after a non-trivial fix-up") was unreachable while the budget was two rounds, and an unreachable obligation reads as satisfied. A candidate the batch did not move owes nothing: the condition is the candidate hash, not the author's sense of how big the fix was.
601
- - **Iterating external review until zero findings**. Stop blind repetition at the configured sequence limit and run the checkpoint. Stabilized or repeated findings require source disposition and a method/design or evidence change before necessary review continues under existing task authority; a genuine decision blocks only its dependent slice.
602
- - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
603
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another review round only when recorded verification needs and current risk call for it, or when a human explicitly requests one. A fresh bounded sequence never resets task history or cumulative spending — retain both and apply the self-hosted-chain checkpoint above.
604
- - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
532
+ ```
533
+ ## Review pass
534
+ - Reviewed commit: <sha>
535
+ - Findings: N total (a P0 / b P1 / c P2)
536
+ - R0 evidence: <same value menu as the single-pass rows above>
537
+ - Dispositions: <one line per P0/P1: fixed | accepted (who/when/why) | pre-existing (evidence)>
605
538
 
606
- ### Recording the loop
539
+ ## Challenge pass
540
+ - Reviewed commit: <sha>
541
+ - Findings: N total (a P0 / b P1 / c P2)
542
+ - R0 evidence: <same value menu>
543
+ - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator | n/a>
544
+ - Dispositions: <one line per P0/P1>
607
545
 
608
- Add one row per round to the validation log:
546
+ ## Delta pass N (after any post-review change)
547
+ - Reviewed commit: <sha>; packet: git diff <previous reviewed commit>..<sha>
548
+ - Findings / Dispositions: <as above>
609
549
 
550
+ ## Post-review delta
551
+ - <git diff --stat last-reviewed-commit..HEAD>, one line per change
610
552
  ```
611
- ## Challenge pass — round N (codex exec adversarial)
612
- - Diff scope: <files / commit range / sha>
613
- - Findings: N total (a P0 / b P1 / c P2)
614
- - R0 evidence: <same value menu as the review-pass row above>
615
- - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
616
- - New since prior round: <count> (subset of above; flag round-introduced bugs)
617
- - Stabilized: <list of findings carried over without change>
618
- - Applied: M fixes (commit: <sha>)
619
- - Deferred: <list with reason>
620
- - Decision: continue implementation / park dependent slice / await human / human stop, because <reason>
621
- ```
622
-
623
- A complete dual-track-validated change names every round explicitly. Skipping rounds without recording the decision is the same as not running them.
624
553
 
625
554
  ## What does NOT count as dual-track
626
555
 
@@ -20,13 +20,13 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
20
20
  ├─ d. Sanitization pass with checklist (cheap, seconds)
21
21
  ├─ e. Owner review gate per mandatory table (deep, minutes)
22
22
  │ ├─ Strict wording-only → one independent code-review pass
23
- │ └─ Non-wording → extraction_review_gate review + challenge (wrapper-fixed budget)
23
+ │ └─ Non-wording → extraction_review_gate review + challenge (single-shot each)
24
24
  ├─ f. Apply fixes, re-sanitize
25
25
  ├─ g. Commit per batch on a feature branch → MR pending review (never push to main)
26
26
  └─ h. Update charter completion log
27
27
  │
28
28
  4. Closeout → ~/.<host>/skills/.extraction-work/<project>-completion.md
29
- Non-wording terminal ledger validator + final state + deferred backlog
29
+ Recorded passes + post-review delta + final state + deferred backlog
30
30
  │
31
31
  5. Provenance migration → ~/.<host>/.private-aliases/<project>.yaml
32
32
  Move file keys / paths / counts / dates out of working files
@@ -91,10 +91,11 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
91
91
 
92
92
  - When required: see `references/dual-track-review-gate.md` table.
93
93
  - Choose the review tier from that table, not from intuition. Do not restate the rows locally; record the exact `dual-track-review-gate.md` table row used. Record `challenge: not-required` only when that row classifies the actual diff as challenge-not-required (for shared skills, this means strict wording-only with deterministic scope proof + independent review confirmation). Non-wording shared-skill changes cannot skip challenge.
94
- - Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row; any rerun of review/challenge against the new candidate draws on the remaining cross-chain Agent budget (the self-hosted-chain rule in `references/dual-track-review-gate.md`), and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path governs instead of a rerun. Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
95
- - Review pass: persist the complete self-review row and encode it in the review plan. For a **non-wording** lane, resolve the repository-owned `scripts/extraction_review_gate.sh` and use it from round 1; never substitute the generic controller, scan writable plugin roots, or supply a caller-selected budget. For a strictly proven **wording-only** lane, use the generic `code-review` proof-bound single-review recipe in `code-review/references/staged-review-contract.md` and record `challenge: not-required`; require its controller-derived wording scope plus the independent `wording_only_boundary` confirmation. This is the only extraction path that stays outside the multi-round wrapper and terminal ledger; the gate, not this page, decides whether a chainless review is legal, and it may still demand the tracked pair. Take all controller options from that runnable recipe, supplying the actual stage and exact candidate rather than an example default. The non-wording chain cannot be retrofitted, so a run started outside its owner wrapper is thrown away and restarted. Read the chain-opening and packet-composition rules in `references/dual-track-review-gate.md` first. Require conclusive JSON, selected-client attribution, packet/profile binding, family exclusion, and wrapper runtime evidence. When the host returns a live execution handle (`session_id`, `cell_id`, or equivalent), keep polling that exact handle until terminal exit; empty current output is progress, not a verdict, and no replacement/fallback reviewer may start while the original process is live. The result row records handle type, an opaque host transcript/tool-call reference and terminal exit status. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle. Never copy a credential-like raw handle into shared evidence. `findings` is not pass; inconclusive, malformed, or free-form output stays interim. Do not add a separate behavior probe.
96
- - Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh` separately with the same plan, stage, candidate, family and tracked chain. From the second round on, the packet's `--paths` exclusion list must exclude every evidence JSON the round has already added (receipts, dispositions, closeout files) — the merge-side binder excludes exactly those, so a packet that excludes only receipts binds a different candidate than the one that lands and the round is thrown away (observed twice in consecutive rounds); bound evidence such as base attestations and excerpt files is committed before the round, never after. Pass the next one-based index; later rounds include a distinct focus and all prior focuses. Preserve a separate result row with the same binding, exclusion, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
97
- - Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a round; when both lenses are required and cross-chain Agent budget remains, re-run both on the updated candidate before landing — every re-run sums into the same wrapper-fixed budget, and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path in `references/dual-track-review-gate.md` replaces further re-runs.
94
+ - Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row; a changed candidate does not by itself owe another external pass (the extraction review lane in `references/dual-track-review-gate.md` decides which passes are owed). Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
95
+ - Review pass: persist the complete self-review row and encode it in the review plan. For a **non-wording** lane, resolve the repository-owned `scripts/extraction_review_gate.sh` once per pass; never substitute the generic controller, scan writable plugin roots, or pass chain or budget options (the wrapper refuses them). For a strictly proven **wording-only** lane, use the generic `code-review` proof-bound single-review recipe in `code-review/references/staged-review-contract.md` and record `challenge: not-required`; require its controller-derived wording scope plus the independent `wording_only_boundary` confirmation. The gate, not this page, decides whether the wording-only single review is legal, and it may still demand the review-plus-challenge pair. Take all controller options from that runnable recipe, supplying the actual stage and exact candidate rather than an example default. Read the packet-composition rules in `references/dual-track-review-gate.md` first. Require conclusive JSON, selected-client attribution, packet/profile binding, family exclusion, and wrapper runtime evidence. When the host returns a live execution handle (`session_id`, `cell_id`, or equivalent), keep polling that exact handle until terminal exit; empty current output is progress, not a verdict, and no replacement/fallback reviewer may start while the original process is live. The result row records handle type, an opaque host transcript/tool-call reference and terminal exit status. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle. Never copy a credential-like raw handle into shared evidence. `findings` is not pass; inconclusive, malformed, or free-form output stays interim. Do not add a separate behavior probe.
96
+ - Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh --mode challenge` separately, with a focus, on the candidate after the review's fixes are applied (a local checkpoint commit is allowed, see SKILL.md). It is not bound to the review's candidate. Preserve a separate result row with the same binding, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
97
+ - Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a pass. Every commit after the last pass that changes more than non-executable record files in the round's evidence directory (register rows and scripts included) owes an Agent-run delta pass on that delta only (at most five; a P0/P1 still open after that is reverted or blocks the pull request). The pull-request description lists each pass and the commit it reviewed (`references/dual-track-review-gate.md`, extraction review lane).
98
+ - Test lane: run the full lane once before the review; after it, a fix confined to one suite reruns that suite plus `check-ccl-skills.sh`, and a controller, contract or shared-gate fix reruns the full lane.
98
99
  - Skipping a required challenge = work can only land as interim, not complete.
99
100
 
100
101
  #### 3f. Apply fixes, re-sanitize
@@ -120,7 +121,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
120
121
  - Final state: which batches done, which deferred, which sources unavailable.
121
122
  - Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
122
123
  - Cost row (process toil is measured, not felt): review/challenge rounds run, findings fixed / accepted / deferred, wall-clock from charter to PR, and net body-word delta per touched entrypoint (`scripts/check-size-budget.sh` prints it). A round that grew a `severe_debt` entrypoint's references without retiring anything records that as the outcome; the next round must read this row before deciding its batch shape.
123
- - For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. When the whole candidate exceeds one packet, split the review rather than the pull request: `--print-manifest --partition <paths> ...` renders a landing partition manifest, and one validated ledger per partition plus the committed manifest is what the gate binds. When an integration branch that accumulated several already-bound rounds is promoted as one pull request, no new ledger is owed: the gate walks the branch's first-parent chain and rebinds each round merge at its own base, provided every merge is the automatic merge of its parents and nothing was pushed to the branch outside a round. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
124
+ - Record the passes and the post-review delta in the round's `evidence/` directory and the pull-request description per `references/dual-track-review-gate.md` (Recording the passes). There is no closeout ledger and no merge-side candidate binding.
124
125
 
125
126
  ### 5. Provenance migration
126
127
 
@@ -150,8 +151,7 @@ Skip this step when nothing transferable surfaced.
150
151
  | Anti-pattern grep panel | `references/recurring-anti-patterns-checklist.md` | Every commit; ~30s |
151
152
  | `check-ccl-skills.sh` | `scripts/check-ccl-skills.sh` | Every commit; ~10s |
152
153
  | Generic `code-review` gate | repository-owned skill | Strict wording-only independent review; ~5-10 min |
153
- | `scripts/extraction_review_gate.sh` | this skill package | Non-wording review plus the wrapper-fixed challenge budget; ~5-15 min each |
154
- | `scripts/validate_extraction_review_state.py <closeout.json>` | this skill package | Every non-wording terminal checkpoint |
154
+ | `scripts/extraction_review_gate.sh` | this skill package | Non-wording review and challenge, one single-shot call each; ~5-15 min each |
155
155
  | Source-read fallback ladder | `SKILL.md` Source-read remediation | When a source read fails or times out |
156
156
  | Sibling mini-map | `SKILL.md` Step 4 stack-specific updates | Every stack-specific change |
157
157
  | Private alias map `audit_cmd` | `~/.<host>/.private-aliases/<project>.yaml` or process-retro profile | Every commit's R0 audit |
@@ -689,3 +689,11 @@ The pending classification above is superseded by the executed source comparison
689
689
  | 语域漂移的查法要先隔离本轮新句:必须把新写的句子单独拎出来单独看,混在整篇里读就看不出用词是谁的 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#必须把本轮**新写的句子**单独拎出来单独看 | `updated` | Owner key `tighten-doc/SKILL.md`。与上三行同属本轮那条收尾规则,分行是归属选择——守卫支持一条 firing-path 里放多个 locator 并各自解析,分行只是为了让红时能指到具体那一步。RED-baseline(applied,differential):删掉或改写该规范列表行,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。这一行来自人工授权轮的评审:三个 locator 都不覆盖「隔离新句」,删掉它仍能过闸。 |
690
690
  | 语域漂移的候选判定必须以改动之前的文本为对照:不得拿改完的文档做对照,否则新词已经在里面,永远查不出候选 | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/closeout-reread.md#逐个术语查它**在改动之前的文本里**出现过没有;不得拿改完的文档做对照,否则新词自己就在里面,永远查不出候选。没出现过的就是候选 | `updated` | Owner key `tighten-doc/SKILL.md`。与上三行同属本轮那条收尾规则,分行是归属选择——守卫支持一条 firing-path 里放多个 locator 并各自解析,分行只是为了让红时能指到具体那一步。RED-baseline(applied,differential):删掉或改写该规范列表行,`register-firing-path-resolution.rb` rc=1 并点名该 locator;控制组与恢复后 rc=0。这一行同样来自人工授权轮。锚点从「逐个术语」起,覆盖遍历范围、对照对象、禁令与候选定义四段连写:上一轮 challenge 实测出,只钉禁令时把「在改动之前的文本里」换成「在术语表里」,五条 locator 全部照旧匹配而查法已废;扩锚后该替换直接失配。同类 finding 已连出三轮(只护一半规则 / 锚点截断 / 换掉正面对照对象),据此在账本里把结论写死:substring 锚钉的是字面不是语义,它保证的只是**被锚定的那段字面**被删或被改写时会红;它不保证规则被改写时会红——实测:把查法的引导词改成「以下查法仅在用户明确要求时执行」,五条 locator 全部照常匹配而整条查法已成可选。适用条件与语义完整性由评审与人读负责,本行不作此声称。 |
691
691
  | 钉住散文规则的锚点闸必须自带能失败的测试:删掉被锚定的规则行必须让闸变红并点名该 locator,而锚点之外的文字被掏空时闸不得报红——后一条把「锚钉字面不钉语义」这条边界写成被执行的事实,而不是账本里的一句声明 | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`(本轮未改,改动落在同包的该测试脚本)。Observed failure:本轮连续三轮 challenge 都指出同一件事——变异记录依赖的守卫不在按 diff 装的评审包里,包内无法核验;先后用「路径+blob 哈希」和「候选内的风险责任人接受书」作答都被驳回,后者尤其是错的:被评审的东西不能自己给自己发授权。RED-baseline(applied,differential,隔离路径归因):完整套件 fail-fast,任何守卫变异都先撞红最早受影响的既有用例,新增两条根本跑不到——本轮 challenge 正是据此推翻了先前那条「改诊断串」的证据,那次变异撞的是既有的 reworded-anchor 用例,什么也没归因到。改用**每条用例各一份单例副本**:把 `next if body.include?(anchor)` 换成整行相等时,边界用例红、删除用例绿;换成从不报缺失锚点时,删除用例红、边界用例绿;控制组与恢复四格全绿。两次不翻转的格子都是实跑观测;探针已提交为 `specs/125-doc-closeout-and-register-drift/evidence/attribution-probe.sh`,一条命令重跑整张矩阵;它在被测提交的一次性 detached worktree 里执行,调用者的检出只读不写(首版写进活动检出,被本轮评审判为 P1 并已重写)。三次被推翻的归因尝试(改诊断串撞到既有用例/两条用例同副本致后一条不执行/副本临时不可复核)连同原因一并留在走查里。两次变异与原始输出见 `specs/125-doc-closeout-and-register-drift/evidence/mutation-walk.txt` 的第二段。 |
692
+ | 非 wording 的提炼改动只欠一次独立 review 加一次对抗 challenge,各为单次调用、不挂评审链;最后一次评审之后凡改了证据与台账以外内容的提交,都由 agent 自动跑一次只看该增量的复查,最多两次,仍有 P0/P1 就撤回该改动或把 PR 标为阻塞,不交人逐行把关;不再要求落地树等于某次评审的候选哈希,CI 只检查改了 skills/ 或 hooks/ 的 PR 至少带一份结论性的 review 结果,challenge 仍由流程规则约束 | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#Every post-review delta gets a delta pass | `updated` | Owner key `skill-extraction-workflow/SKILL.md`;配套改动 `references/dual-track-review-gate.md`、`references/extraction-quickstart.md`、`references/validation-and-landing.md`、`scripts/extraction_review_gate.sh`;删除 `scripts/review_ledger_binding.py`、`scripts/validate_extraction_review_state.py` 及两套测试,CI 的合并侧绑定步骤换成 `scripts/check_review_evidence_present.py` 存在性检查(只查有没有,不比对候选)。Observed failure:候选哈希绑定让技能仓里几乎每次修复都作废已有评审结果并重开序列,此前连续几轮的证据目录各自只留下最后两轮收据,其间又为它补了分区清单、首父链逐轮重跑、候选与评审包拆分、成本收窄等轮次;引入绑定的提交补的是自身节奏里的一致性缺口,不是有未审改动出过事。RED-baseline(applied,differential):新测试放进 base 的一次性 detached worktree 对旧 wrapper 跑,第一条断言即红(review 调用带着 `--challenge-budget 1`),当前 wrapper 全绿。 |
693
+ | 被正式废止的台账锚点可以由一条豁免绑定一组固定的历史行:每行按 SHA-256 摘要逐一核对,删掉、改写或多出一行引用都必须报红;豁免的 command 锚点所指脚本不存在时视为已记录的退休 | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh | `updated` | Owner key `skill-extraction-workflow/SKILL.md`;改动在 `scripts/register-firing-path-resolution.rb` 与其接线测试。本行也是退休锚点的 supersede 行:`test_review_ledger_binding.sh`(19 行引用)、`test_validate_extraction_review_state.sh`(5 行)、`review_ledger_binding.py`(1 行),以及 `dual-track-review-gate.md` 与 `extraction-quickstart.md` 各两句被本轮删除或改写的锚点,逐条原因见该脚本的 EXEMPT 表。Observed failure:删除脚本后整本台账校验 rc=1、30 个 locator 不可解析,而原豁免每个锚点只容许 1 行,表达不了一个脚本被 19 行引用的退休。RED-baseline(applied):加豁免前 rc=1 并逐条点名,加后 rc=0;接线测试新增删一行、改一行、多一行三条回归,各自以准确诊断报红。 |
694
+ | code-review 的分阶段契约不得再许诺一个在合并时重算候选的脚本:该脚本已退休,契约只说明收据同时记录评审包与候选两个哈希,需要比对落地树的调用方自行比对 | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md`(本轮未改);改动在本包分阶段评审契约的一句、控制器两处注释、控制器测试新增一条断言。通用控制器的评审链模式保留给其他调用方,未改。RED-baseline(applied,differential):新断言对 base 的契约文字为红(含 1 处已退休脚本名),对当前文字为绿。 |
695
+ | 调研类请求(含深度调研、deep research、调研某个产品)必须先进 multi-perspective-research:宿主另装的通用 deep-research 技能只能作为本技能循环里的检索工具,不替代它;常驻路由层同步加这一条 | `multi-perspective-research` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/multi-perspective-research/SKILL.md#description; bank-evidence: file:eval/routing-tasks.jsonl#route-product-deep-research | `updated` | Owner key `multi-perspective-research/SKILL.md`;包内只改 frontmatter description(触发词前置,Skip 条款原样后移),另改 `agent-context/session-start.md` 常驻路由加一行(17051→17042 字节,删三处冗余措辞抵消)与 `docs/SKILLS.md` 目录标记由 leaf 改为 entry。Observed failure:codex 宿主上一条「深度调研这个产品」的请求选中了宿主自带的通用 deep-research 技能(其描述为 Use only when the user asks for deep research),CCL 调研 owner 直到用户追问才加载;当时常驻路由层没有调研条目,本技能描述开头是近 200 字的 Skip 条款。RED-baseline:失败侧是记录在维护者私有宿主会话里的实际事件;改后侧是题库新增用例与两条相邻用例各跑 3 次、9/9 判定正确(筛查精度,未达 10 次观测下限);codex 端同句复测要等宿主插件更新后才能跑,记为待办。 |
696
+ | 评审之后候选的任何改动(任意严重度的修复、补的测试、changelog)都要在开 / 就绪 / 合并 PR 或报完成之前由 agent 按所属闸自己重审:通用流程完整重跑、最多 5 次,仍有 P0/P1 就撤回或报阻塞,不交人工 review,最后一次结论性评审覆盖 HEAD 时才能说 HEAD 已评审;控制器在 worktree 的 git 目录里记下每次结论性评审时的 HEAD,插件的 PR hook 在开 / 就绪 / 合并命令前比对并把漏审提交注入给 agent | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/references/development-completion.md#Pushing it for a human to review does not discharge it | `updated` | Owner key `code-review/SKILL.md`(入口未改);改动在 `skills/code-review/references/development-completion.md`(新增「Before a pull request or a ready report」五步必走检查)、`skills/code-review/references/staged-review-contract.md`、`skills/code-review/scripts/review_gate.py`(本机回执:只在候选由 `--base` 从整个工作区派生、冻结前后 HEAD 与干净度一致、结果结论性时写,锚点记下解析后的 git 目录路径与身份,写入时从根逐级不跟随链接打开它、核对身份,再经 no-follow 目录描述符写入)、`hooks/remind-review-covers-head.sh`(新 PreToolUse 提醒,只提醒不拦截,OpenCode 插件同步镜像)、常驻层 `agent-context/session-start.md` 一句(净 -4 字节)。Observed failure:另一仓库的生产会话在最后一次评审之后提交了为一条 P2 补的测试和 changelog,随即推分支开 MR,报告写「没有再单独 review、等人工 review」,合并后又把未评审的 HEAD 说成评审过的。当时规则已写「候选变了要重审」(本 reference 与 product-rd 验证门都有),失效在触发点:从修复到开 MR 的动作序列里没有一步比对评审覆盖的提交与 HEAD。RED-baseline:红的一半只有这次生产失效;两种隔离探针(直接问下一步、流程中只给命令,均 n=3/臂,判分标准先于运行冻结)在改前文本上都没复现(问答 3/3 先重审再开 MR;流程中 0/3 推送或开 MR),所以文本改动本身的效果未被证明,归为假设,摘要见 `specs/128-renewed-review-and-repo-contract/evidence/probe-summary.md`。hook 的机械部分有 applied 差分:在副本上分别去掉内容树相等判断、祖先判断、`--ready` 匹配、引号屏蔽,各自只让对应的那一条用例变红;评审后补的「同一命令里先 commit 再开 PR」「`--help ;` 后接真正的开 PR」「带引号带空格的 `cd`」「先开草稿、再提交、再就绪」四类用例在修复前的 hook 上为红(现共 25 条);控制器的去重、冻结前后锚点一致两条用例,以及「原 git 目录被挪走、旧路径换成指向它的链接」场景,在修复前的控制器上为红;回执目录被换成链接的竞态按构造封住、没有确定性测试。提醒是否改变 agent 行为未测量。 |
697
+ | 评审模式把被评审仓库自己入库的约定文件引在候选之后交给评审员:从仓库根到每个改动路径的每层目录取 `AGENTS.override.md`(否则 `AGENTS.md`)、`CLAUDE.md`、`.claude/CLAUDE.md`,根在前,32 KiB 以内;未入库和 `CLAUDE.local.md` 一律不读(不得外发给别家评审模型),链接、非 UTF-8、超预算的整份省略并记原因,永不因约定文件让评审失败;有引用时加 `repository_contract` concern,要求评审员报改动违反的规则和规则本身的缺陷(`Contract defect:`),规则只作数据、不能为缺陷开脱 | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_review_gate.sh | `updated` | Owner key `code-review/SKILL.md`(入口未改);改动在 `skills/code-review/scripts/review_gate.py`(`repository_contract_section`、profile 与结果的 `repository_contract` 字段、条件 concern、trust boundary 一句)、`skills/code-review/references/staged-review-contract.md`、`skills/code-review/references/development-completion.md`(偏离约定要在 plan intent 里声明、`Contract defect:` 的处置、不许悄悄放松约定文件),`skills/code-review/scripts/test_review_client_compat.py` 的 provider mock 改为放行控制器自己的 git 读取;共享读取函数 `read_bounded_regular_file` 逐级打开目录中途失败时不再泄漏已持有的目录描述符(修复前 100 次失败读取后打开的描述符由 4 个涨到 254 个,约定文件被逐个省略时会反复触发)。发现规则按两家宿主一手文档核对(Codex:每层目录 override 优先、根到深拼接、默认 32 KiB;Claude Code:`CLAUDE.md` 或 `.claude/CLAUDE.md`,`CLAUDE.local.md` 是个人文件)。`candidate_sha256` 在追加前算定,引用内容不能增加候选路径;challenge、complete、wording-only 不附加。不读 `@path` 导入、`.claude/rules`、宿主配置的备用文件名(结果里 `sources_not_read` 列明)。RED-baseline(applied,differential):在副本上分别去掉「只收已入库文件」、去掉 override 优先、不加 concern,三条约定用例里恰好两条变红、其余用例全绿;base 版控制器配新套件时三条约定用例全红;中间目录被换成链接、带第二个硬链接的约定文件都被省略且内容不进评审包。 |
698
+ | 产品研发验证门里「评审后有实质改动才全量重跑」的「实质」一词删除:评审后的任何改动(含测试 / 文档)都触发全量重跑,作者不能自行判定改动无关紧要 | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#any post-review change, tests/docs included | `updated` | Owner key `product-rd-workflow/SKILL.md`(改一句,字数与字节都不增)。Observed failure 同上一条生产会话:它把评审后的提交归为「只有测试和文档」而没有重审,原句的「实质」正好给了这个归类余地。RED-baseline 同上:红的一半是生产失效,隔离探针在改前文本上未复现,改动效果未被证明。 |
699
+ | 提炼流程的增量复查上限由两次改为五次(用户裁决),同步 `skills/skill-extraction-workflow/SKILL.md`、`dual-track-review-gate.md`、`extraction-quickstart.md` 与被测试钉住的原文;本行按指针取代 127 轮评审线那一行里的「最多两次」 | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#After five delta passes a still-open P0/P1 | `updated` | Owner key `skill-extraction-workflow/SKILL.md`(同一行内省 6 字节,入口不增长)。delta 复查改为 `--base <已评审提交>` 取增量,使它绑定工作区并留下 PR hook 读取的本机回执。RED-baseline(applied):`test_extraction_review_gate.sh` 钉住的短语由两次改为五次后,对 base 文本变红、对 head 文本变绿。 |
@@ -141,7 +141,7 @@ Before commit, for any rule that appears in more than one authored file (`SKILL.
141
141
 
142
142
  ## Bounded Independent Review Packet
143
143
 
144
- Use this when an independent review is required but a broad reviewer prompt hangs, returns no output, or starts expanding beyond the intended review scope. For non-wording extraction work, the gate-valid path is `scripts/extraction_review_gate.sh`, which owns the fixed autonomous budget while delegating transport to the provider-neutral `code-review` controller. A strictly proven wording-only change uses the proof-bound generic single-review recipe in `code-review/references/staged-review-contract.md`; its controller-derived wording scope and independent `wording_only_boundary` result replace neither one another nor a failed semantic check. Both paths preserve the frozen packet, family exclusion, structured validation and client-specific recovery; raw provider CLI packets are debugging/advisory only and must not be recorded as passing review evidence.
144
+ Use this when an independent review is required but a broad reviewer prompt hangs, returns no output, or starts expanding beyond the intended review scope. For non-wording extraction work, the gate-valid path is `scripts/extraction_review_gate.sh`, which makes each pass single-shot while delegating transport to the provider-neutral `code-review` controller. A strictly proven wording-only change uses the proof-bound generic single-review recipe in `code-review/references/staged-review-contract.md`; its controller-derived wording scope and independent `wording_only_boundary` result replace neither one another nor a failed semantic check. Both paths preserve the frozen packet, family exclusion, structured validation and client-specific recovery; raw provider CLI packets are debugging/advisory only and must not be recorded as passing review evidence.
145
145
 
146
146
  Required flow:
147
147
 
@@ -164,7 +164,7 @@ Review only stdin. Do not use tools. Do not request more context. Return finding
164
164
  5. Apply or explicitly reject actionable findings.
165
165
  6. Rerun the owning skill validators and `git diff --check`.
166
166
  7. Record the result as `findings applied`, `no blocking findings`, or `review unavailable after remediation` only when the wrapper or approved alternate produced valid structured evidence. Raw packet output is recorded separately as debugging/advisory and cannot close the gate.
167
- 8. At a non-wording terminal checkpoint, build the receipt-bound ledger and run `scripts/validate_extraction_review_state.py <closeout.json>`. A wording-only review records its single independent-review row and does not invent a multi-round ledger.
167
+ 8. Record each pass and the post-review delta per `references/dual-track-review-gate.md` (Recording the passes). A wording-only review records its single independent-review row.
168
168
 
169
169
  Do not count as completed review:
170
170