@ccoalm/ccl-skills 0.9.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +4 -3
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +23 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +178 -4
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +127 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +2 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +1 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +11 -1
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +16 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +1 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +8 -8
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +9 -1
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +2 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +9 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +17 -20
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +9 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +26 -44
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +16 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +59 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +2 -2
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +3 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +35 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +16 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +35 -6
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +476 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +81 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +28 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +336 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +41 -1
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +141 -28
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +139 -38
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +1 -1
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -8
  56. package/dist/assets/release.json +110 -45
  57. package/package.json +1 -1
@@ -67,7 +67,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
67
67
 
68
68
  Across a long operational-rule/code extraction series the challenge supplies *the same handful of axes* as the recurring P0/P1 — first drafts systematically nail the functional / cost / happy-path and omit a predictable set. The per-axis instance lists for the six first-draft blind-spot axes:
69
69
 
70
- - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox.
70
+ - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox; and — because skill/reference text is itself a prompt agents execute — the rollout-safety screen an eval harness runs on skill text before release: wording that induces context/secret exfiltration (prompt leakage), assumes or grants permissions beyond the task (overreach), or automates a destructive or confirmation-skipping step (unsafe automation) — text a draft must not carry except as an explicitly labelled anti-example.
71
71
  - **(2) concurrency & lifecycle** — races, deadlock (e.g. holding a lock through a drain/callback), use-after-close/free, resurrection after delete, double-free/double-close, cleanup ordering, at-most-once/fires-once.
72
72
  - **(3) resource bounds** — an unbounded default/timeout/buffer/retry, a leaked registry/goroutine/task entry, a missing max backstop.
73
73
  - **(4) rollout / migration ordering** — a step that breaks not-yet-upgraded consumers, or abandons a live bug to do the clean refactor first.
@@ -82,7 +82,7 @@ Three per-edit instances of the axes above that recur because their canonical ru
82
82
 
83
83
  For any new mechanical gate, validator, or evidence apparatus — **and for any change that makes an existing one's verdict stricter** — run the four legs at design time, not after challenge rounds force them. These four legs all fire at design time; the check has a second firing point they do not cover, because its actor is not the gate's author: when a landing **withdraws or downgrades the evidentiary claim an existing gate rests on**, that gate is re-based or retired in the same landing — see the claim-liveness rule in `product-rd-workflow/references/design-review-gate-mechanics.md`, which owns it, including the obligation walk a retirement owes.
84
84
 
85
- - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under the SAME base resolution CI uses, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep).
85
+ - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under **every base resolution CI uses — enumerate the landing faces, never assume one**, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep). A base-relative verdict is a quantity *relative to its base*, and CI resolves that base per event, so this leg is discharged by a **recorded manifest, never by a walk you assert**: derive the branch set from the workflow itself — its `push` branch filter plus every branch a pull request is actually landed into — and record one entry per branch carrying the resolved base ref and this gate's verdict on the candidate, so a reviewer can diff the manifest against the workflow and `every` stops being author-adjudicated. Green on the round's own base says nothing about the others: the difference is every commit that landed on the shared branch before the gate existed, and it surfaces as one accumulated violation on the face nobody measured — typically the promotion PR, long after the authors who could have funded it moved on. **A set that cannot be enumerated is a blocking residual, not a pass** — an unrestricted `pull_request` trigger with no fixed target set has no finite manifest, so take leg (d)'s non-blocking or risk-owner-deferral exit and name the faces left unmeasured rather than claiming coverage. Enumerating the faces costs a loop; the exemption you will reach for instead is the loosening leg (e) forbids.
86
86
  - **(b) marginal-cost statement** — record what the cheapest routine change costs under the gate (recompute/regenerate/rerun burden); a gate whose per-iteration cost defeats normal development gets lightened or redesigned at design time.
87
87
  - **(c) trust-model fit** — name what the mechanism defends against under its DECLARED trust model; machinery that only defends against adversaries the trust model already excludes (e.g. content digests where the author can regenerate every hash) buys redundant detection at full complexity cost — prefer the lighter mechanism that keeps the enforceable core.
88
88
  - **(e) loosening check — an exemption must name the class of change that stops owing evidence, and that class must be one with no behaviour to evidence.** Fires whenever a change makes a gate accept what it used to reject: a new exemption class, a waived requirement, a widened accept set. A loosening is easier to get wrong than a tightening and shows up later, because it produces no red for anyone to notice — the gate simply stops asking. Two obligations, both outcomes rather than procedures. **First, state the exempted class in behavioural terms and check it against the repo's own definition of that class**: an exemption named `not-required` asserts *no behaviour*, so if any rule in the tree already classifies that same diff shape as behaviour-changing, the exemption contradicts it and the answer is to fix the *anchor/evidence form* for that shape, never to drop the evidence. **Second, if the class does carry behaviour, the exemption must be replaced by a way to SUPPLY the evidence** — widen where the anchor may land, add an evidence form the shape can satisfy — because the class with the most behaviour is exactly the one an exemption hurts most. **A precedent of the same shape is not a justification**: reaching for an existing exemption class because the root cause rhymes with an earlier one transfers the solution without checking the disanalogy, and the disanalogy is usually the load-bearing part. Failure shape: a gate anchor that structurally cannot bind to a frontmatter-only change was answered with a third `not-required` class by analogy to two existing ones, even though the same repository elsewhere states that any frontmatter edit is a routing-surface change and the author had just measured its routing delta; the adversarial review caught it on the first finding, and the correct fix was to let the anchor bind to the changed description instead.
@@ -230,7 +230,7 @@ Review + challenge check the change for defects; neither checks whether it actua
230
230
  | Status | Use when |
231
231
  |---|---|
232
232
  | `RED-baseline` | the change alters behavior or routing (trigger / scope / routing / validation / acceptance). Run the scenario WITHOUT the change first (baseline failure), then WITH it (compliance) |
233
- | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance, and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
233
+ | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance — an author cannot self-assert it — and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
234
234
  | `not-applicable: docs-only` | the change touches NO file under `skills/**` and no skill-loaded guidance — i.e. `README` / `ARCHITECTURE` / `CONTRIBUTING` / `docs/**` only. Forbidden for `SKILL.md`, `references/**`, validators, templates, examples, and the plugin-shipped command/behavior surfaces (`hooks/*`, `scripts/install.sh`, `bin/`, `.mcp.json`, `.lsp.json`, `monitors/**`, `settings.json`): those are behavioral or executable source even when they read like prose |
235
235
 
236
236
  For a `RED-baseline`, the evidence form can be a before-after task diff, a golden trace, or a pressure scenario — these are *how* you show baseline→compliance, not standalone substitutes for it. For a routing-surface / hub-skill change you can run the golden-trace form as a REAL headless-agent run via the F4 Tier-3 harness (optional, higher-fidelity than a recorded scenario — see `validation-and-landing.md` Behavioral Validation); a recorded scenario is not by itself a failing baseline.
@@ -241,7 +241,6 @@ For a `RED-baseline`, the evidence form can be a before-after task diff, a golde
241
241
 
242
242
  Rules:
243
243
 
244
- - `RED-baseline` is required whenever the change alters behavior or routing. `semantic-control` is valid ONLY with reviewer confirmation that none of trigger/scope/routing/validation/acceptance changed — an author cannot self-assert it.
245
244
  - The row must give a concrete locator + evidence shape, not a bare status: artifact path / commit / transcript / command, the exact prompt or scenario, and expected-vs-actual. For `RED-baseline`, record BOTH the without-change (baseline failure) and with-change (compliance) results, and name the baseline's **provenance type**: a *recorded incident* (cite where the failure is actually recorded — transcript, note, issue; the cited record must describe a failure that occurred, not prescribe a method) or a *constructed scenario* (run against BOTH the unchanged baseline and the changed rule, with an openable artifact for each run — a scenario "run" only mentally, only against the patched text, or only as reviewer discussion does not count). A prescriptive source — a method-bar or best-practice note with no failure recorded — cannot be cited as an occurred failure and is not by itself a valid `RED-baseline`; it may seed the constructed scenario's design or support the rule's rationale, but the RED evidence is the run artifact. Narrating what "would have" failed as if it happened is a fabricated evidence row, the same defect class as fabricated verification output. A status word with no openable artifact is not a valid row, same as a missing review row.
246
245
  - A missing or unreviewable behavioral-evidence row blocks landing the same way a missing review row does; until it exists the work is an uncommitted interim checkpoint.
247
246
 
@@ -300,18 +299,13 @@ Primary reviewer failure is a remediation branch only when the owning gate class
300
299
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
301
300
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
302
301
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
303
- 5. A fallback result satisfies only the exact lane it ran. If no candidate returns a conclusive verdict, keep the work `interim`.
304
-
305
- Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
302
+ 5. A fallback result satisfies only the exact lane it ran, and only when it meets the same evidence bar — do not call a manual ad-hoc run "fallback review". Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence. If no candidate returns a conclusive verdict, keep the work `interim`.
306
303
 
307
304
  - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
308
305
 
309
306
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
310
307
 
311
- The following raw CLI shapes are **debugging diagnostics only**. They may help
312
- isolate an owner-wrapper failure, but neither output is review/challenge evidence
313
- and neither may replace `scripts/extraction_review_gate.sh` for a non-wording
314
- lane.
308
+ The raw CLI shapes below are **debugging diagnostics only** — they can help isolate an owner-wrapper failure, but are never review/challenge evidence and never replace `scripts/extraction_review_gate.sh` on a non-wording lane.
315
309
 
316
310
  Debug a wrapper with raw `codex review` against the target diff:
317
311
 
@@ -423,7 +417,7 @@ The external reviewer judges what the bounded packet can show; the packet struct
423
417
  - The reviewer owns CONTENT SEMANTICS: wording coherence, dropped qualifiers/obligations, source-accuracy, sanitization, scope drift — everything decidable from the packet plus the reviewer's own reasoning.
424
418
  - Deterministic-gate claims (pinned literals present, size ratchet net-zero, obligation audit green, R0 clean, parity) are verified by the repository's CI re-running those gates on the actual branch — never by the reviewer, and never accepted from the implementer's prose alone; a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round.
425
419
  - Historical-process claims (a pre-fix RED, a measurement taken before landing) are session-record-grade unless bound to an immutable revision or a candidate-bound receipt; treat them as the implementer's testimony, and say so in the disposition instead of demanding evidence the packet cannot hold.
426
- - Standing backlog: teaching the gate to embed candidate-SHA-bound receipts of deterministic-gate output into the packet removes the third bullet's limitation mechanically; until that lands, this boundary is the accepted state.
420
+ - The receipt channel for the third bullet is landed: `scripts/gate_receipt.py mint --out <ledger-dir>/gate-receipt-N.json -- <gate command>` binds one deterministic-gate run to the committed candidate — clean tree required, HEAD commit recorded, argv + repository-relative cwd + exit code + output SHA-256 + RFC3339 mint time, with an off-by-default opt-in output tail (a verbatim excerpt would copy a leaked token into the ledger; receipts are created 0600) and a post-run candidate re-read that refuses a result when HEAD moved mid-run; a RED run mints fine (the pre-fix RED is the canonical use). Receipts live OUTSIDE the candidate tree (the chain ledger directory) and ride the packet by reference: name the receipt file plus its own SHA-256 in a review-plan `evidence` entry or a disposition-evidence `evidence` item, so `review_context_sha256` / the v3 ledger's sibling-hash discipline bind it. Anyone re-checks with `gate_receipt.py verify <file>` (structural) or `verify <file> --rerun -- <the gate command you expect>` at the recorded commit — the verifier types the command and the tool compares it against the recorded argv before executing only the verifier's own words, because a receipt is untrusted input and its recorded argv must never be executed as-is (differential exit-code + output-hash comparison; `--exit-only` for a legitimately nondeterministic gate, and say so in the referencing row). Trust model, stated so it is not oversold: a receipt is candidate-bound, falsifiable consistency evidence — it does not authenticate who ran the command; deterministic authority stays with CI re-running the gates on the actual branch (second bullet), and a process claim carrying no receipt remains testimony under the third bullet.
427
421
 
428
422
  ## Recording findings + fixes
429
423
 
@@ -438,7 +432,7 @@ For each pass, record in the extraction's working file (e.g. `<project>-extracti
438
432
 
439
433
  ## Challenge pass (codex exec adversarial)
440
434
  - Findings: N total (a P0 / b P1 / c P2)
441
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
435
+ - R0 evidence: <same value menu as the review-pass row above>
442
436
  - Gate-fireability applicability: <yes — change adds/edits a semantic rule/gate/status/verdict | no — valid ONLY when the diff is wording-only or adds/edits no semantic rule/gate/status/verdict>
443
437
  - Item 9 exercised: <locator to the captured prompt/transcript/JSONL showing the bypass-by-omission probe actually ran (not a pasted self-assertion) | n/a per line above>
444
438
  - Applied: M fixes (commit: <sha>)
@@ -524,32 +518,20 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
524
518
 
525
519
  - When the recurring surface is a **self-adjudication clause** — decidable test: the clause's classification verb has NO named test whose output produces the classification, so the receiver/author judges it — the `keep / delete / narrow / replace` decision must first ask whether an existing mechanical or semi-mechanical test can carry the adjudication: name the test whose output settles the classification, **and the mapping from its output to the classes** — which output means which class — plus the actual result or the evidence contract that will produce it. Naming a test is not routing to it: a test whose output cannot discriminate the classes, or one named with no recorded output-to-class mapping, leaves the adjudication exactly where it was and does not satisfy this rule. A `keep` that retains self-adjudication prose, or a `narrow`/`replace` that adds more prose bindings, is landed only when the decision record (the same-class rule's recorded decision, in the commit body or register row) names the reason no existing test could carry it — a reason left in chat does not count; two challenge rounds attacking the same self-adjudication surface are the signal that prose is the wrong layer — a clause routed to an existing test inherits that test's evidence bar instead of the adjudicator's say-so (worked instance: a review-reception clause that left "is this finding scope-adding" to the receiver survived two rounds of attacks on that self-classification until the classification was routed to the existing structural-minimality test). The same question fires at drafting time for any new reception/discipline-style clause that would grant a self-adjudication.
526
520
 
527
- "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*, not *no new P0/P1 appeared*.
521
+ "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*. Do NOT iterate to zero *findings* either — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing, and forcing the count to zero either over-corrects or scope-creeps; the bar is zero *undispositioned* P0/P1, which differs from both *zero findings* and *no new P0/P1 appeared*.
528
522
 
529
523
  **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
530
524
 
531
- Do NOT iterate to zero *findings* — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing. Forcing the finding count to zero either over-corrects or scope-creeps. The bar is zero *undispositioned P0/P1*, which is different from zero findings.
532
-
533
525
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
534
526
 
535
- Five is the generic `code-review` transport ceiling, not this extraction lane's
536
- spend. Non-wording Agent-autonomous extraction calls go through
537
- `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=2`: the
538
- initial review plus at most two challenges. At round 3 the autonomous lane ends.
539
- An authenticated human may request later review, but that is separately
540
- attributed human-requested evidence outside this chain/budget, not an Agent
541
- round 4 or 5. Unused generic capacity never authorizes automatic continuation.
542
- The v3 closeout validator rejects referenced receipts whose recorded budget is
543
- not 2 and checks budget and ordering consistency within the caller-supplied
544
- set. It cannot authenticate that the wrapper produced those receipts or that
545
- the caller retained every earlier chain or receipt. The wrapper does not mint or
546
- persist
547
- `review_chain_id` or `autonomous_review_index`: the caller still supplies both,
548
- and could start a fresh-looking chain after round 3. The validator detects bad
549
- order inside the referenced set but cannot detect a prior chain the caller
550
- omitted, so complete caller-owned ledger retention—and treating an Agent reset
551
- as a contract violation—remains part of the boundary rather than a property the
552
- local scripts prove.
527
+ Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
528
+
529
+ **Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
530
+
531
+ 1. **Batch dispositions; never re-chain per finding — and hold every fix until the round-2 challenge has run.** Triage the whole batch through the disposition bar and deep-self-review once, then hold, never deciding by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, and round 2 binds the round-1 candidate, so a fix applied between the two forfeits the double-receipt terminal and costs a fresh human-authorized chain to recover — an observed failure, not a hypothetical: a round that landed its review fixes before the challenge had to be closed by a user-granted continuation chain. So the rule through round 2 is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, then apply the held batch, MR/PR-listed, and let round 3's succession challenge — owed exactly when the batch moved the candidate — be what inspects it.
532
+ 2. **Sum spent rounds across all chains before opening one more; the lane's cap is three rounds across two chains.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against that cap. The only restart the lane funds is the single succession challenge, opened with `--predecessor-chain-result-file` so the ended chain is carried rather than laundered into a fresh-looking loop; any OTHER restart opens with a fresh review by contract, trading a challenge round for a review round, and must not be opened autonomously at the cap. Effective exhaustion is reached when the remaining rounds cannot fund the closeout floor for any continuation; treat it exactly like the cap.
533
+ 3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
534
+ 4. **At the cap — or at effective exhaustion — the designed terminal is disposition, never another chain.** Apply or disposition the final batch, name every post-review fix in the MR/PR description, record the honest terminal state (`continuation_authorization_required` when the lane's final round ran and itself returned findings; otherwise — including a chain broken before its challenge could run — an interim record naming the last externally reviewed candidate and every later delta), and hand continuation or merge to the human. The post-review batch sits only on the pending MR/PR branch beside that record — the human's authenticated continuation, waiver, or merge decision is what certifies it, and it is never reported as reviewed. This is the bounded outcome working as designed, so do not report it as convergence and do not launder it through a fresh-looking chain.
553
535
 
554
536
  A strictly proven wording-only change has no convergence loop: it uses one
555
537
  generic `code-review` pass, records the independent-review row and the
@@ -563,7 +545,7 @@ This budget limits only automatic reviewer invocation. It does **not** stop impl
563
545
  - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
564
546
  - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
565
547
  - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
566
- - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. One dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. The recovery is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
548
+ - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. The dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. While cross-chain budget remains, that break is handled autonomously by the self-hosted-chain rule's ledger-counted restart; it becomes this bullet's human-decision dead-end when the remaining budget cannot fund the re-review. The recovery at that point is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
567
549
 
568
550
  When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
569
551
 
@@ -571,7 +553,7 @@ The mechanical reminder is `self_review_gate`, not prose alone. It records outst
571
553
 
572
554
  In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
573
555
 
574
- At the third Agent-autonomous round, do not start a fourth automatically. If findings remain:
556
+ At the final Agent-autonomous round, do not start another automatically. If findings remain:
575
557
 
576
558
  - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`;
577
559
  - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
@@ -586,14 +568,14 @@ ledger and every referenced controller, completion, base, and sweep file live
586
568
  in one directory and carry SHA-256s. The validator walks the ordered schema-v3
587
569
  controller chain (same chain and scope,
588
570
  review then contiguous challenges, packet=candidate, complete prior-result hash
589
- prefix, fixed `challenge_budget=2`) and binds every closeout candidate to its
571
+ prefix, the wrapper-fixed `challenge_budget`) and binds every closeout candidate to its
590
572
  last receipt. Ready requires at least review + challenge; a second base drift may
591
573
  stop as race immediately after round 1 rather than spending an illegal challenge
592
574
  after the terminal predicate already fired.
593
575
  It ends in exactly one state:
594
576
 
595
577
  - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
596
- - `continuation_authorization_required`: round 3 itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
578
+ - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
597
579
  - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
598
580
 
599
581
  Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
@@ -608,17 +590,17 @@ convergence.
608
590
 
609
591
  For a focused single-skill change:
610
592
  - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
611
- - **Round 2 — challenge 1**: after implementer triage, attack the highest-risk unresolved surface with an unprimed prompt.
612
- - **Round 3 — challenge 2**: verify remaining/new attack paths. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic round 4.
593
+ - **Round 2 — challenge**: after implementer triage — fixes stay HELD: applying any fix before this round breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt.
594
+ - **Round 3 — succession challenge, owed only when the fix batch moved the candidate**: apply the held batch, commit it, then ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate now hashes to. Unchanged (every finding accepted, pre-existing, or source-refuted) ⇒ the lane ends at round 2 and owes nothing. Changed ⇒ open ONE succeeding chain with `--predecessor-chain-result-file <round-2 receipt>` and challenge the landing candidate on a focus distinct from round 2's. This is the final Agent-initiated external round; its findings feed the post-budget checkpoint rather than an automatic further round, and the batch lands MR/PR-listed.
613
595
 
614
- Broad extractions use the same three-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
596
+ Broad extractions use the same lane budget, round 3 included on the same condition. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
615
597
 
616
598
  ### Anti-patterns
617
599
 
618
- - **Single-round challenge → done**. The round-1 fix-up itself may introduce bugs. Always do at least one re-challenge after a non-trivial fix-up.
600
+ - **Landing a fix batch no round ever saw**. The round-1 fix-up itself may introduce bugs, so a batch that moved the candidate owes the succession challenge of round 3 above — the earlier absolute ("always re-challenge after a non-trivial fix-up") was unreachable while the budget was two rounds, and an unreachable obligation reads as satisfied. A candidate the batch did not move owes nothing: the condition is the candidate hash, not the author's sense of how big the fix was.
619
601
  - **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
620
602
  - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
621
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one.
603
+ - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one. The observed extreme is chain multiplication: a chain broken by your own fix and restarted at index 1 is the same budget, not a new loop — sum rounds across chains per the self-hosted-chain rule above.
622
604
  - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
623
605
 
624
606
  ### Recording the loop
@@ -629,7 +611,7 @@ Add one row per round to the validation log:
629
611
  ## Challenge pass — round N (codex exec adversarial)
630
612
  - Diff scope: <files / commit range / sha>
631
613
  - Findings: N total (a P0 / b P1 / c P2)
632
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
614
+ - R0 evidence: <same value menu as the review-pass row above>
633
615
  - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
634
616
  - New since prior round: <count> (subset of above; flag round-introduced bugs)
635
617
  - Stabilized: <list of findings carried over without change>
@@ -41,10 +41,11 @@
41
41
 
42
42
  ## Tier-2:路由 task-bank + 廉价 grader(已落地,advisory)
43
43
 
44
- `scripts/eval-routing-bank.rb <repo-root> [--bank p] [--model m] [--limit N] [--dry-run] [--json p] [--baseline p] [--desc-budget-chars N]`。`--desc-budget-chars N` 把每条 description 截到前 N 字符再评(消费端截断臂——如 Codex 在 ~2% 上下文预算/未知窗口 8,000 字符下压缩技能清单);plain 与 budgeted 各跑一遍即可定位"只活在描述尾部"的路由触发词。`make eval-routing-bank`。
44
+ `scripts/eval-routing-bank.rb <repo-root> [--bank p] [--model m] [--limit N] [--dry-run] [--json p] [--baseline p] [--desc-budget-chars N] [--replicas N]`。`--desc-budget-chars N` 把每条 description 截到前 N 字符再评(消费端截断臂——如 Codex 在 ~2% 上下文预算/未知窗口 8,000 字符下压缩技能清单);plain 与 budgeted 各跑一遍即可定位"只活在描述尾部"的路由触发词。`--replicas N` 每 task 评 N 次:task 判定取保守共识(任一有效副本 FAIL 即 FAIL),并报告副本 top1 一致率——不同 (bank, replicas) 配置是不同尺子,不得互相 diff 当回归。`make eval-routing-bank`。
45
45
 
46
- - 冻结 task-bank `eval/routing-tasks.jsonl`:每行 `{id, utterance, expected_skill, must_not_route_to?, source, why_expected, frozen_at_sha}`,种子取自 source-register 历史 miss + bootstrap 路由规则。
47
- - grader = 每 task 一次本机 `claude --print --tools "" --model <haiku>`,喂 utterance + 全部 skill description(agent 真正路由的那份面),要 `{selected_skill, confidence, rationale_short}`。
46
+ - 冻结 task-bank `eval/routing-tasks.jsonl`:每行 `{id, utterance, expected_skill, acceptable?, must_not_route_to?, source, why_expected, frozen_at_sha}`,种子取自 source-register 历史 miss + bootstrap 路由规则。**已修复的路由 miss 必须把其 utterance 冻结成 bank task 落在同一交付里**(修复不冻结=下次同类漂移无回归面)。`expected_skill: "none"` 是否定对照/覆盖空洞哨兵:正确结果是没有技能认领;`acceptable` 列出可辩护替代结果(如空洞探针上 coordinator 接管与拒绝都对);`must_not_route_to` 点名吸入诱饵邻居。结构由 `test_routing_bank_integrity.sh` 确定性把关(sentinel 只准出现在 expected/acceptable,不准进 must_not)。
47
+ - **冻结案例神圣(regressions-are-sacred)**:已冻结案例的删除或判定面改写(bank task 的 expected/acceptable/must_not,golden trace 的 assert 块)是一次回归裁决事项,与邻居回归同权——平均改善不得抵消单条冻结案例的失守,且「曾 yes 现非 yes」**含降级为 unsure/INCONCLUSIVE** 都算回归。删除/改判的每条案例必须在同一轮的 register 追加行里写 `case-retired: <id>` 或 `case-rescoped: <id>` 并给理由,交独立评审裁决;确定性半边由 `test_frozen_case_sanctity.sh` 按 `CCL_SKILL_BASE_REF` 把关——无裁决行即红,无 base ref 时打印显式 skip token(skipped ≠ passed),它只保证交易可见,不裁决交易正当性。
48
+ - grader = 每 task(×replicas)一次本机 `claude --print --tools "" --model <haiku>`,喂 utterance + 全部 skill description(agent 真正路由的那份面),要 `{selected_skill|none, clarify, confidence, rationale_short}`。**clarify 率、低置信率(<0.5)、副本一致率是一等报告字段**,不是旁注——路由质量的残余风险常在"高置信直选却选错、无自纠路径"这类 pass/fail 看不见的分布里。
48
49
  - **路由兼容性信号,不是真值预言机** —— grader 自己可能错;只衡量"当前 description 能否让廉价模型把固定 utterance 路由到 expected"。
49
50
  - **advisory**:不接 `check-ccl-skills.sh`,不挡 merge。退出码:`0` = 跑完;`2` = 用法;`3` = grader 整体不可用(claude CLI 缺失则打印 skipped 后 `0`)。路由 miss 永不非 0。
50
51
  - **防作弊**:runner 校验每 task 的 `frozen_at_sha` 是 HEAD 祖先(非祖先 = drift,排除出回归判定);同一改动若同时动 task-bank 和 SKILL.md description 会显式告警(防"改 skill 顺手改测试让它过")。
@@ -58,6 +59,18 @@
58
59
  2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
59
60
  3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
60
61
  4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
62
+ 5. **单变量归因**:一次改前/改后对照只准动**一个路由变量**(一条 description,或同一 skill 不可分割的一组路由面)。同时动多条 description 的批量改动,其对照差值不可归因到任何一条,只能按整包回归读——要归因就拆成逐条 A/B。(源侧实测形态:仅替换一条 description 的成对子集对照,把命中从约 2/3 提到 95%,且提升可归因到那一条改动——多条同动时这句话说不出口。)
63
+
64
+ ## 路由失效形态词表(跨层;报告与修复讨论用这套名字)
65
+
66
+ pass/fail 之外,路由失败有可命名的形态;每个形态有不同的检测器与不同的修法,混称"路由不准"会修错面:
67
+
68
+ | 形态 | 判据 | 检测器 | 典型修法 |
69
+ |---|---|---|---|
70
+ | **吸入 (absorbed)** | 应被拒绝(expected none)或应远离诱饵邻居(must_not)的 utterance 被某技能认领——覆盖空洞不被承认而被最近邻低置信/clarify 拉走 | bank 否定对照+空洞探针,runner `absorbed` 标签 | 认领空洞(补 owner)或在诱饵邻居 description 加排除句;不要靠 grader 自觉 |
71
+ | **归属分裂 (ownership_split)** | 同一 utterance 多副本给出不同 top1——所有权不稳定,谁都像 owner | `--replicas ≥2` 一致率 + `ownership_split` 标签 | Skip-when 互相消歧或触发词改具体(description-authoring 的 80% 阈值) |
72
+ | **静默跳过 (silent skip)** | 路由"成功"但被选技能的正文从未被读/其硬规则从未被应用——选择层过了,质量层空转。判据阶梯:mounted → invoked → 文件真被读 → 下游行为改变;mounted-only 不证明任何生效 | 非 Tier-1/2 可见;B 面 body-compliance 探针 + Tier-3 真 agent 事件流 | 修正文 firing point/read routing(body 指针补齐四元组,见 description-authoring.md「Body routing pointers」),不是修 description |
73
+ | **高置信错选** | 错选但 confidence 高、无 clarify——事后无自纠路径,比低置信错选更危险 | 报告里 FAIL ∩ 高置信 ∩ clarify=false 的行 | 触发词消歧;必要时在正文加入口自检;残余风险如实记录 |
61
74
 
62
75
  ## Tier-3:hub golden trace 真 agent 回放(已落地,advisory,人工判定)
63
76
 
@@ -71,6 +84,14 @@
71
84
  - **随机性**:agent 非确定;判定先人工、nightly 起步,有稳定史前不自动 gate。防作弊同 T2(`frozen_at_sha` 祖先校验)。
72
85
  - **双用途**:除回归外,Tier-3 还可当**改技能前的 RED-baseline**(改前手动跑触发场景看真 agent 是否真路由错,改后看 compliance)——可选;只有真观察到 miss 才算 RED(PASS/INCONCLUSIVE 不算),小 N + 非确定有噪声,手动跑两次自己留两份报告。落地 + 防作弊注意见 [validation-and-landing.md](validation-and-landing.md) "Optional real-agent RED-baseline"。
73
86
 
87
+ ## B 面:正文合规探针 body-compliance(已落地,advisory)
88
+
89
+ `eval/body-compliance-eval.rb <repo-root> [--arm L] [--json p] [--model m] [--timeout s] [--ids a,b]`。`make eval-body-compliance`。路由三层测「选没选对技能」;B 面测**已激活技能的正文硬规则是否真被应用**——正文即 prompt,逐探针 required/forbidden marker 契约判分,覆盖是 NAMED SUBSET(见 runner 头部声明)。
90
+
91
+ - 含 product-rd-workflow 停机谓词的**成对分类探针**(prd-stop-*/prd-continue-*):每对场景恒定、只变谓词判别特征,按闸自身的字面 `continuing:`/`blocked:` marker 判分——确定性锚钉的是这些谓词的**措辞存在性**,只有这些探针检验**案例被分到哪边**。
92
+ - 触发纪律:改动触及某技能正文硬规则或停机谓词时,must run the affected `--ids` probe subset on this machine before landing,结果按该改动既有门禁处置;advisory 契约与升级路径(提炼确定性不变量,永不阻断 LLM 判定)见 `eval/AGENTS.md` 与 [f4 手册](../../../docs/f4-skill-effectiveness-harness.md)「双轨承载面与运行分层」。
93
+ - 实测的双轨互补边界(锚管措辞、探针管行为漂移、不是埋句绊线)与小样本告诫,单一落点在 f4 手册「双轨承载面与运行分层」,此处不复述。
94
+
74
95
  ## Health roll-up:描述性仪表盘(已落地,advisory)
75
96
 
76
97
  把上面各信号卷成**一个加权 0–10 显示值 + 同尺子变化**,用于定位值得继续检查的维度。它借用 OpenSSF Scorecard 的呈现形态,但不把不同性质的 F4 信号变成“仓库整体变好/变差”的总判决。映射与限制见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。
@@ -139,6 +139,22 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
139
139
  - 给列表内的外部 CLI 传长 flag 之前,**必须先在本机读过该 CLI 的 help 输出**;未读过就用,属于把记忆当一手源。
140
140
  - 一条命令因无法识别的 flag 失败时,**不得**先改用法之外的东西——先怀疑该 flag 在本机这个版本不受支持。语料里多次出现「以为是用法错、其实是这版没这个 flag」。
141
141
 
142
+ ### 记录半边:把外部契约钉进控制时,同时钉住它的时效
143
+
144
+ 上两节讲**读法**(读一手源、把谓词挂在自己拥有的东西上);这一节讲**记录**。一旦本仓的某个控制**编码了**对外部契约的一次读取——闸强制它、脚本依赖它、recipe 假设它——控制自身必须带三样东西,且写在控制旁边,不是写在 commit message 或某轮私档里:
145
+
146
+ - **出处到条款级**:读的是哪一份一手源的哪一节,不是「按官方文档」。
147
+ - **核验时点与被核对象的版本**:哪天核的、对着工具/宿主的哪个版本或哪个 pin 核的。
148
+ - **失效条件**:必须写明什么变化会使这条结论不再成立(pin 的大版本被抬、宿主换了解析链、上游把该字段改成会读的)。
149
+
150
+ 缺这三样,下一个读到该控制的 agent 分不清它是**经过判断的决定**还是**没人敢动的化石**,唯一的复核方式是把整轮外部调研重做一遍——这正是「继承来的读法」能一路存活的原因:不是没人想核,是核一次的代价被隐藏了。三样都在时,复核退化成一次针对性的对照。
151
+
152
+ - A control whose annotation lacks any of the three must be treated as carrying an unverified reading: the next change to it re-reads the primary source instead of inheriting the claim, and a reviewer must not accept the inherited reading as established merely because the control has shipped for a long time.
153
+
154
+ 判据窄,别泛用:只对**控制依赖的外部契约**成立(宿主/工具/格式/协议的行为)。纯内部不变量不适用;「谁提出的」那一档归 `attribution-verification.md` 的源质量与 pending 语义,与本条不同 facet。
155
+
156
+ 本轮两个活体实例(都是读代码时先撞上「这为什么在这」才发现的):`verify-packed.mjs` 用机械闸禁止 `plugin.json` 带 `version`——这是一条关于宿主版本解析链的主张,落地时正确、但当时没留出处与时点;以及 `check-release-version.py` 依赖 `actions/checkout` 的 `fetch-depth: 0` 会取到 tag。两处已按上述三样补齐。
157
+
142
158
  ## 只据内部语料的「深度提炼」:失败形态
143
159
 
144
160
  一次仅以内部 launch-SOP 为源的「深度提炼」会落出**看起来完整**的规则集,却从未核过这套技能声称代表的**公开实践**;下一个用户于是问「参考网上优秀实践了么 / did you check industry practice」。
@@ -20,7 +20,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
20
20
  ├─ d. Sanitization pass with checklist (cheap, seconds)
21
21
  ├─ e. Owner review gate per mandatory table (deep, minutes)
22
22
  │ ├─ Strict wording-only → one independent code-review pass
23
- │ └─ Non-wording → extraction_review_gate review + at most two challenges
23
+ │ └─ Non-wording → extraction_review_gate review + challenge (wrapper-fixed budget)
24
24
  ├─ f. Apply fixes, re-sanitize
25
25
  ├─ g. Commit per batch on a feature branch → MR pending review (never push to main)
26
26
  └─ h. Update charter completion log
@@ -91,10 +91,10 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
91
91
 
92
92
  - When required: see `references/dual-track-review-gate.md` table.
93
93
  - Choose the review tier from that table, not from intuition. Do not restate the rows locally; record the exact `dual-track-review-gate.md` table row used. Record `challenge: not-required` only when that row classifies the actual diff as challenge-not-required (for shared skills, this means strict wording-only with deterministic scope proof + independent review confirmation). Non-wording shared-skill changes cannot skip challenge.
94
- - Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row and rerun review/challenge against the new candidate. Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
94
+ - Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row; any rerun of review/challenge against the new candidate draws on the remaining cross-chain Agent budget (the self-hosted-chain rule in `references/dual-track-review-gate.md`), and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path governs instead of a rerun. Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
95
95
  - Review pass: persist the complete self-review row and encode it in the review plan. For a **non-wording** lane, resolve the repository-owned `scripts/extraction_review_gate.sh` and use it from round 1; never substitute the generic controller, scan writable plugin roots, or supply a caller-selected budget. For a strictly proven **wording-only** lane, use the generic `code-review` proof-bound single-review recipe in `code-review/references/staged-review-contract.md` and record `challenge: not-required`; require its controller-derived wording scope plus the independent `wording_only_boundary` confirmation. This is the only extraction path that stays outside the multi-round wrapper and terminal ledger; the gate, not this page, decides whether a chainless review is legal, and it may still demand the tracked pair. Take all controller options from that runnable recipe, supplying the actual stage and exact candidate rather than an example default. The non-wording chain cannot be retrofitted, so a run started outside its owner wrapper is thrown away and restarted. Read the chain-opening and packet-composition rules in `references/dual-track-review-gate.md` first. Require conclusive JSON, selected-client attribution, packet/profile binding, family exclusion, and wrapper runtime evidence. When the host returns a live execution handle (`session_id`, `cell_id`, or equivalent), keep polling that exact handle until terminal exit; empty current output is progress, not a verdict, and no replacement/fallback reviewer may start while the original process is live. The result row records handle type, an opaque host transcript/tool-call reference and terminal exit status. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle. Never copy a credential-like raw handle into shared evidence. `findings` is not pass; inconclusive, malformed, or free-form output stays interim. Do not add a separate behavior probe.
96
96
  - Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh` separately with the same plan, stage, candidate, family and tracked chain. Pass the next one-based index; later rounds include a distinct focus and all prior focuses. Preserve a separate result row with the same binding, exclusion, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
97
- - Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a round; when both lenses are required, re-run both on the updated candidate before landing.
97
+ - Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a round; when both lenses are required and cross-chain Agent budget remains, re-run both on the updated candidate before landing — every re-run sums into the same wrapper-fixed budget, and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path in `references/dual-track-review-gate.md` replaces further re-runs.
98
98
  - Skipping a required challenge = work can only land as interim, not complete.
99
99
 
100
100
  #### 3f. Apply fixes, re-sanitize
@@ -119,7 +119,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
119
119
  - File: `~/.<host>/skills/.extraction-work/<project>-completion.md`
120
120
  - Final state: which batches done, which deferred, which sources unavailable.
121
121
  - Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
122
- - For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. A clean Round 2 plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 3 findings validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row but does not fabricate a multi-round ledger.
122
+ - For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
123
123
 
124
124
  ### 5. Provenance migration
125
125
 
@@ -149,7 +149,7 @@ Skip this step when nothing transferable surfaced.
149
149
  | Anti-pattern grep panel | `references/recurring-anti-patterns-checklist.md` | Every commit; ~30s |
150
150
  | `check-ccl-skills.sh` | `scripts/check-ccl-skills.sh` | Every commit; ~10s |
151
151
  | Generic `code-review` gate | repository-owned skill | Strict wording-only independent review; ~5-10 min |
152
- | `scripts/extraction_review_gate.sh` | this skill package | Non-wording review plus at most two challenges; ~5-15 min each |
152
+ | `scripts/extraction_review_gate.sh` | this skill package | Non-wording review plus the wrapper-fixed challenge budget; ~5-15 min each |
153
153
  | `scripts/validate_extraction_review_state.py <closeout.json>` | this skill package | Every non-wording terminal checkpoint |
154
154
  | Source-read fallback ladder | `SKILL.md` Source-read remediation | When a source read fails or times out |
155
155
  | Sibling mini-map | `SKILL.md` Step 4 stack-specific updates | Every stack-specific change |
@@ -24,10 +24,12 @@ The at-add-time check above decides *where* a rule lands (merge vs new bullet);
24
24
 
25
25
  | Baseline failure | Right form | Wrong form |
26
26
  |---|---|---|
27
- | Agent knows the gate and walks past it under pressure (discipline slip — our merge-authorization / R0 / done-claim class) | Prohibition + rationalization-vs-reality pairs + red-flag self-check | Soft guidance ("prefer…", "consider…") |
27
+ | Agent knows the gate and walks past it under pressure (discipline slip — our merge-authorization / R0 / done-claim class) | Prohibition + rationalization-vs-reality pairs + red-flag self-check — pairs quote excuses actually captured verbatim from baseline/pressure runs (`validation-and-landing.md` eval-first step 2), never invented ones | Soft guidance ("prefer…", "consider…") |
28
28
  | Agent complies but the artifact's SHAPE is wrong (bloated review packet, buried verdict, register row restating the source) | Positive recipe/contract: state what the artifact IS — its parts, in order | A "don't"-list about the shape — the measured backfire above |
29
29
  | Agent omits a required element from an artifact it already produces (missing status/evidence cell, absent map row) | A REQUIRED slot in the template/validator it must fill (our closeout rows and register gate are this form) | Prose reminders near the template |
30
30
  | Behavior should differ by situation | Conditional keyed to an observable predicate ("fan-out → name the tier") | Unconditional rule + exemption clauses |
31
+ | Every step is right and the aggregate is still wrong — a batch, destructive, or otherwise high-stakes operation names a target that does not exist, misses a required item, or applies two conflicting changes, and the damage is done before anything checks | **plan-validate-execute**: the step emits a structured plan artifact, a script validates the plan, and only a valid plan is executed. Errors name the offending item *and* the set it was checked against ("field `signature_date` not found. Available: …"), because a verdict the agent cannot act on sends it back to guessing. The plan is also what makes the operation reversible while it is still cheap — iterating on the plan touches nothing (Anthropic, *Skill authoring best practices*, "Create verifiable intermediate outputs"; verified 2026-09-01) | Prose care ("verify each target before applying"), or a validator whose output is a bare pass/fail |
32
+ | Instruction latitude does not match the operation's fragility — a fixed-sequence fragile operation written as a heuristic (the agent improvises a variant), or a genuinely multi-solution judgement written as one exact script (the agent follows it off a cliff when the context differs) | Pick the tier from the operation, not from the author's confidence: **high** freedom (prose heuristics) when several approaches are valid and context decides; **medium** (parameterised template or pseudocode) when a preferred pattern exists and configuration varies; **low** (one exact invocation, few or no parameters, stated as not-to-be-modified) when the operation is fragile, consistency is critical, or a specific sequence must hold (same source, "Set appropriate degrees of freedom"; verified 2026-09-01) | One uniform specificity across a whole skill |
31
33
 
32
34
  The rows are not mutually exclusive — a real failure often carries several axes (e.g. a required field omitted only under pressure). Give each axis its own form **on its own target**: the discipline axis keeps a prohibition aimed at the *act of skipping/violating*; the shape axis gets the recipe/slot describing *what the output is*. Do NOT re-express the shape requirement as a prohibition merely because a discipline axis is also present — a "don't restate X"-style prohibition riding along with a recipe is exactly the measured backfire. (This combination guidance is an inference beyond the source's separately-tested arms; like any load-bearing form choice it is subject to the behavioral-evidence row below.)
33
35