@ccoalm/ccl-skills 0.13.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +19 -24
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +32 -32
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +16 -14
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +24 -26
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +11 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +60 -209
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +114 -367
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +52 -672
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +10 -2
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +4 -2
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +77 -444
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +33 -98
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +57 -173
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +1 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +1 -1
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +1 -1
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +2 -2
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +2 -2
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +1 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +14 -41
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +11 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/correction-routing-map.md +22 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +7 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +26 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +2 -2
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +8 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +7 -1
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +4 -2
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +8 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +8 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +73 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +11 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +11 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/entrypoint_form_census.py +169 -0
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +62 -3
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +114 -10
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/reference-access-census.sh +157 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +483 -109
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +188 -14
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +10 -0
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_form_census.sh +174 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_resolution.sh +253 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +49 -25
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_reference_access_census.sh +209 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +394 -5
  51. package/dist/assets/release.json +77 -47
  52. package/package.json +1 -1
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: requirement-scope
3
- description: 改动范围 / 影响范围 / scope / 需求拆分 / MVP 边界 / 非目标 / 版本切片 / 变更影响 / appetite / timebox —— 交付物是**变更边界**:in/out scope、受影响对象、依赖、MVP 与后续切片、appetite 与砍项、每个切片的验收范围。前提是方向已定。Skip 方向还没定、要先弄清「到底要什么」→ requirement-intent;要的是现状清单(现在怎么运作、有什么能力)→ requirement-baseline;风险定级与要哪些 gate → feature-risk-router;实现/发布计划 → product-rd-workflow;测试范围 → testing-strategy。
3
+ description: 改动范围 / 影响范围 / scope / 需求拆分 / MVP 边界 / 非目标 / 版本切片 / 兼容·回滚降级边界 / 变更影响 / appetite / timebox —— 交付物是**变更边界**:in/out scope、受影响对象、依赖、MVP 与后续切片、appetite 与砍项、每个切片的验收范围。前提是方向已定。Skip 方向还没定、要先弄清「到底要什么」→ requirement-intent;要的是现状清单(现在怎么运作、有什么能力)→ requirement-baseline;风险定级与要哪些 gate → feature-risk-router;实现/发布计划 → product-rd-workflow;测试范围 → testing-strategy。
4
4
  ---
5
5
 
6
6
  # Requirement Scope
@@ -41,7 +41,7 @@ Use this skill to turn observed experience into durable agent skills without cop
41
41
  ### Evidence, RCA, charter & attribution(证据 / RCA / charter / 出处核实)
42
42
 
43
43
  - **Charter-before-editing red-line**: do not read sources or edit skills until the complete charter in `references/source-to-skill-extraction.md#extraction-charter` is filled cell-by-cell. A remembered field subset is not a charter. Step 0 owns the procedure; this rule owns the hard stop. Core-Rule canonicality governs same-facet drift only and never narrows that field set to this bullet.
44
- - Classify every result — including task/session summaries and lessons-learned requests, not only failures — as **failure/correction**, **stable success**, or **unstable/insufficient evidence**. Failure runs RCA; stable success requires mechanism, non-luck evidence, reuse conditions, firing point, and owner; insufficient evidence stays an observation. Never invent a failure story to justify learning from success. The author classifies, so the class is not self-elective: correction/finding/failure-triggered work defaults to `failure/correction`, and only independent review may accept a relabel into an RCA-skipping class. Missing classification, unaccepted relabel, or missing analysis leaves it `interim`. Full method: `references/source-to-skill-extraction.md#result-learning-baseline-for-every-extraction`.
44
+ - Classify every result — including task/session summaries and lessons-learned requests, not only failures — as **failure/correction**, **stable success**, or **unstable/insufficient evidence**. Failure runs RCA; stable success requires mechanism, non-luck evidence, reuse conditions, firing point, and owner; insufficient evidence stays an observation. Never invent a failure story to justify learning from success. The author classifies, so the class is not self-elective: correction/finding/failure-triggered work defaults to `failure/correction`, and only independent review may accept a relabel into an RCA-skipping class. Missing classification, unaccepted relabel, or missing analysis leaves it `interim`. Full method: `references/source-to-skill-extraction.md#result-learning-baseline-for-every-extraction`; the success half (work-as-done mechanism, re-examination trigger): `references/incident-postmortem-extraction.md` §Success reviews.
45
45
  - **RCA must go wider than one 5-Why chain.** 5 Why is the entry technique to get past a symptom, but a single linear chain to one "root cause" is its documented failure mode: real process/agent failures need multiple concurrent causes, the stop point is arbitrary, and "why" drifts toward "who"/blame.
46
46
  - For any non-trivial extraction, RCA must (a) **widen** — enumerate the multiple contributing factors across categories (trigger/routing, stale process-model, missing mechanical control, missing feedback, latent authored-earlier condition, detection gap) before deepening one — a straight chain with no branches means you stopped early; (b) **counterfactually test** each candidate to rank causal weight — *if removed or changed, would the failure still happen?* — keeping necessary/sufficient factors **and** failed redundant safeguards as secondary controls / defence-in-depth, and dropping only genuine coincidence (a one-trace factor is a hypothesis — mark it probabilistic, don't hard-drop); (c) frame prevention as a **mechanical control on the failure CLASS** — an enforced constraint plus the feedback that confirms it fired on a surface the next agent actually reaches in time (not a clause buried in a deep reference) — not agent diligence, preferring the highest-leverage **practical** control (a named artifact, an owner who can change it today, an observable check — never deleting a useful narrow gate or inflating one miss into an over-broad hook) over the first patchable point.
47
47
  - Reject hindsight causes ("agent careless" / "need more attention" / "should have known"): ask why the action made sense given what the agent could see, not what it should have done. Full method, category prompts, stopping points, and sources: `references/source-to-skill-extraction.md` (Deep RCA For Extraction).
@@ -59,9 +59,7 @@ Use this skill to turn observed experience into durable agent skills without cop
59
59
  - The trigger is a *public-best-practice / state-of-the-art claim* (architecture, testing-strategy, LLM/inference, observability, release, security, design), not every rule: a rule that encodes an internal-only operating constraint or a postmortem-derived guard, stated with explicit internal scope and no state-of-art claim, does not need external grounding. For rules that do claim to represent industry practice, the charter's evidence plan MUST include authoritative external sources (standards, canonical vendor/tool docs, widely-cited literature, or ≥2 independent practitioner sources) used to confirm, refine, or contradict each such rule — or record per-rule why external grounding is not applicable. Verify any named attribution per the attribution rule; generic established terms (e.g. a well-known named problem or method confirmed by ≥2 independent sources) may be used without person-attribution. Failure shape 见 `references/external-practice-controls.md`。
60
60
  - **Implementing or depending on a named external convention/spec/format verifies it against the primary source FIRST — before building, not after.** This fires on *implementing the named thing itself* — a convention/spec/standard/file-format/protocol such as `AGENTS.md`/CODEOWNERS placement, an RFC, a wire or file format, or a tool's config contract — even with no best-practice claim (distinct from the industry-practice trigger above, which fires on a state-of-the-art *claim*). A prior agent's or a prior commit's reading of that convention is **hypothesis-grade**: re-verify against the primary source before extending it, because a fix-forward built on an inherited interpretation propagates the original error. When the convention uses a term with stack-specific meanings (e.g. "package" = a manifest-bearing directory in npm but *every directory* in Go; likewise module/project/workspace), map it to each concrete target stack before encoding scope. Failure shape 见 `references/external-practice-controls.md`。
61
61
  - Match evidence claims to evidence depth. "Full", "complete", "all", "re-read", and "source inventory" claims require named source categories, inspected artifacts, and concrete observations. Use "targeted check" or "no new source read" when that is the real coverage.
62
- - **A curated digest is one source class, not the repo — its exhaustion is not the repo's exhaustion (digest-masks-corpus trap).** A high-quality maintainer digest (`AGENTS.md`/`CLAUDE.md`, README, CONTRIBUTING, architecture/design doc) that *summarizes* a larger code corpus is a **distinct source class** from the code; a strong digest masks how much went unread. Gates: **(1)** an "exhausted / complete / no-gap / fully-extracted" claim requires the **code corpus as its own register row with a terminal status** (deep-read, inventory+owner-mapping, or a downscope citing an actual user instruction — not self-declared); until then scope the claim ("digest-layer covered; code corpus `pending`") — the sweep is usually inventory+owner-mapping+novelty-spot depth and commonly low-yield, so record that outcome, don't skip the row. **(2) Enumerate the source's OWN top-level structure** before any exhausted claim — a doc's `##`/`###` sections, a repo's top-level dirs (or the next unit: TOC/pages/anchors/line-chunks for a doc; package/module/test/script/config for a repo) — and mark each `read`/`skipped`; un-enumerated structure = unsupported claim (the trap recurs even within one artifact).
63
- The invariant under both shapes is **an exhaustion claim must be scoped to a unit you actually enumerated** — for the digest/corpus shape that unit is the artifact's own structure; the rule keeps its `digest-masks-corpus` name for continuity, so do not skip it just because no digest is present. **When the claim is over an ACQUISITION CHANNEL SET rather than one artifact** ("the public sources are mined out", "there is no more data"), walking a seed list of channel classes is the cheap way to catch a class you never considered — but a seed list is not a universe, so **the honest output is which classes you walked and with what search boundary, never an exhaustion claim**; the tell that this is the live shape is that each pushback surfaces a class you had not considered rather than another artifact in a known class. Closure and downgrade stay with whatever coverage gate the owning skill already has — do not introduce a parallel status vocabulary here (variant (c) in `references/coverage-exhaustion-traps.md`).
64
- **(3)** Four variants share the same enumerate-the-real-artifact fix. **(a) self-retrospective** (digest = your OWN summaries/MR-descriptions/project-memory, a lossy digest of the change set) → enumerate the **actual change artifacts** (commit list, diffs, touched files); a project-memory fact ≠ a lesson promoted to the owning shared skill; **(b) benchmark / no-gap** ("we already cover this" is itself a coverage claim) → read the source's **load-bearing decision surface** (decision rules, claim-vs-ground-truth checks, data-loss guards — the code, or the full rule text where the load-bearing artifact is executable prose like a `SKILL.md`/spec), not its advertising overview/feature-docs. **(a0) produced-artifact + next-run-delta** — a source that produced runnable artifacts owes BOTH the produced-artifacts register row family (under the existing statuses `read`/`deep-read`/`excluded`/`routed` — "project-specific only" recorded as the *reason* on an `excluded` row, never a new status) AND a next-run delta row (what this run's cost went to; what the next run must do, in what order), and attaching the artifacts to a deliverable or archiving them in the project is NOT `landed`; **(4)** the verdict owes a per-item disposition row EVEN WHEN THE ROUND LANDS NOTHING — load-bearing surface actually read, owning artifact or route, mechanical firing path for `covered`; a `quick`/`light`/`sweep`/`triage` framing does NOT waive the load-bearing read; the row lands on a concrete surface (the final response OR persistent scratch/source-register), never chat-ephemeral; a repeated 深度分析了么 / did-you-deep-read is the recurrence signal — treat it as a validation-gate defect. Gate detail and worked failure shapes (50KB-digest "exhausted" / multi-repo "fully summarized" / benchmark "already covered"): `references/coverage-exhaustion-traps.md`.
62
+ - **A curated digest is one source class, not the repo — its exhaustion is not the repo's exhaustion (digest-masks-corpus trap).** A maintainer digest (`AGENTS.md`/`CLAUDE.md`, README, CONTRIBUTING, architecture/design doc) that *summarizes* a larger corpus is a **distinct source class** from that corpus, and the invariant under every shape is **an exhaustion claim must be scoped to a unit you actually enumerated** — the rule keeps its `digest-masks-corpus` name for continuity, so do not skip it just because no digest is present. Gates: **(1)** an "exhausted / complete / no-gap / fully-extracted" claim requires the code corpus as its own register row with a terminal status (deep-read, inventory+owner-mapping, or a downscope citing an actual user instruction — not self-declared); until then scope the claim ("digest-layer covered; code corpus `pending`"). **(2)** Enumerate the source's OWN top-level structure before any exhausted claim and mark each unit `read`/`skipped`; un-enumerated structure = unsupported claim. **(3)** The variants share the same enumerate-the-real-artifact fix — **(a)** self-retrospective (your own summaries, MR descriptions, and project memory are a lossy digest of the change set; enumerate the actual change artifacts); **(b)** benchmark / no-gap ("we already cover this" is itself a coverage claim; read the source's load-bearing decision surface, not its overview); **(c)** an acquisition-channel SET (walk a seed list of channel classes, then report which classes you walked with what search boundary — never an exhaustion claim, and never a parallel status vocabulary); **(a0)** produced-artifact + next-run-delta (a source that produced runnable artifacts owes BOTH rows under the existing status vocabulary; attaching or archiving the artifacts is NOT `landed`). **(4)** The verdict owes a per-item disposition row EVEN WHEN THE ROUND LANDS NOTHING — load-bearing surface actually read, owning artifact or route, mechanical firing path for `covered` — on a concrete surface, never chat-ephemeral; a `quick`/`light`/`sweep`/`triage` framing does NOT waive the load-bearing read; a repeated 深度分析了么 / did-you-deep-read is the recurrence signal — treat it as a validation-gate defect. Gate detail (the per-gate unit lists, the (a0)/(4) obligation text) and the worked failure shapes (50KB-digest "exhausted" / multi-repo "fully summarized" / benchmark "already covered" / channel set "mined out"): `references/coverage-exhaustion-traps.md`.
65
63
  - For repository evidence, do not conclude a source is empty or unavailable from the checked-out default branch alone. If the default branch only contains a template, README, or obvious placeholder, inspect local and remote branches, tags, and `git ls-tree`/`git show` for candidate feature/dev branches before marking the source unavailable or routing around it.
66
64
  - When the user asks for full, complete, deep, final, whole-codebase, all-Figma, all-docs, or broad extraction for a task, workflow, product, source set, or multi-skill suite, representative sampling is not an acceptable substitute. Build a source register first, define the minimum read depth for each source class, and do not land a final/complete claim until the register is closed or explicitly downscoped by the user.
67
65
  - Conflicts require a keep/merge/discard decision. Do not append incompatible rules side by side.
@@ -85,13 +83,7 @@ Use this skill to turn observed experience into durable agent skills without cop
85
83
  - **Advisory** (report, never block) = fuzzy collision / possible-dangling bare redirect.
86
84
  - Body-only, typo, or non-routing reference/doc edits run only the existing structural validation, not the routing gate as a landing blocker.
87
85
  - For the analyzer contract, finding classes, blocking-vs-advisory policy, and the staged Tier-2/Tier-3 plan, read `references/eval-routing.md`.
88
- - **Routing-surface fixes must cover the coordinator-vs-executor axis, not only the executor.** When a routing miss is fixed by editing a skill's `description`/trigger surface so a request auto-surfaces the right skill, the owner-generalization map MUST ask whether a lifecycle-COORDINATOR/router owner (e.g. `product-rd-workflow`) also needs the trigger — not only the stack/EXECUTOR owner (`*-dev` / `*-architecture` / a single domain skill).
89
- - A fix that advertises the request type only on the executor's description is incomplete: *multi-stage deliveries* of that type (spec → plan → test-first → impl → verify) keep auto-routing to the executor and skip the coordinator's lifecycle gates.
90
- - The failure shape that triggers this rule: adding a trigger word (e.g. "重构"/refactor) to stack `*-dev`/`*-architecture` descriptions while the coordinator workflow's description never advertises it, even though the coordinator's BODY already claims ownership — a body↔routing-surface contradiction.
91
- - Required check before landing any description/trigger edit: (a) does the coordinator's description advertise this request type with a scale qualifier (narrow/single-file → executor; multi-stage/cross-cutting → coordinator)? (b) is there a body↔description contradiction where the body claims ownership the routing surface omits? (c) when the description paraphrases a body trigger / skip-condition that has **multiple clauses with different scopes** (e.g. clause 1 fires on *any edit* of a surface, clause 2 only on a *semantics change*), preserve each clause's scope separately — collapsing a multi-scope trigger to its most salient clause silently over- or under-fires; the body trigger is the primary source, so re-read every clause (the first reading is hypothesis-grade) before asserting the description over/under-covers it OR rewording it. Resolve all, or record per-owner why the coordinator is `unchanged`.
92
- - Validation gate: a repeated routing miss of the same request type across two rounds is evidence the prior fix only patched the executor surface — re-run this coordinator check before claiming the routing class is closed.
93
- - **Routing-trigger fixes must cover the utterance-variant axis, not only the canonical phrasing.** When fixing a routing miss by adding a trigger to a `description`, the same delivery type arrives under many utterances — the canonical name PLUS restart/redo/from-scratch/continue variants (e.g. a *refactor delivery* arrives as "重构 X" but also "重新开发", "完全重新开始", "推倒重来", "清除代码重新开发", "redo/rewrite from scratch"); advertising only the canonical phrase leaves the variants unmatched, so the workflow silently fails to auto-trigger on them. "The rule/trigger exists but the utterance class is uncovered" is a validation-gate defect: enumerate the restart/redo/continue variants of the request type and add the high-value ones (within the 800-char cap; route overflow to the body entry-precedence text). A repeated miss of the SAME delivery type arriving via a different utterance across rounds is the signal this axis was skipped.
94
- - **In Claude Code, before fixing a routing miss with a trigger-word edit, check whether the skill's description was even IN the host listing — it may be budget-dropped.** A name-only entry silently voids a keyword-based routing fix (the skill stays reachable by explicit `/skill-name`). **Evidence bar (do not over-apply this as a catch-all):** conclude budget-dropped only from concrete evidence — this turn's listing shows that skill as name-only, `/doctor` output, or other host-listing proof; absent that, treat it as a hypothesis and STILL do the normal trigger / utterance-variant / coordinator fix (this rule does not replace them; a *visible* description misjudged as name-only is the reverse trap). Corollaries: (a) a routing/bootstrap doc assuming "every description is always visible" is wrong under budget pressure and must not be a low-traffic gate skill's sole discovery path; (b) on non-Claude-Code hosts verify the host's own listing behavior from its primary docs — any LLM's "it's a host bug" guess is hypothesis-grade until so checked. The budget mechanism, the cold-start trap, and the remediation levers: `references/skill-listing-budget.md`.
86
+ - **A routing-miss fix must cover three axes, not only the executor's canonical phrase.** Enumerate them before landing any `description`/trigger edit; a repeated miss of the same request type across two rounds is evidence the prior fix skipped an axis — re-run the check before claiming the routing class is closed. **(a) Coordinator-vs-executor:** the owner-generalization map MUST ask whether a lifecycle-COORDINATOR/router owner (e.g. `product-rd-workflow`) also needs the trigger, not only the stack/EXECUTOR owner — a fix that advertises the request type only on the executor leaves *multi-stage deliveries* of that type skipping the coordinator's lifecycle gates; check that the coordinator's description carries it with a scale qualifier (narrow/single-file → executor; multi-stage/cross-cutting → coordinator), that no body↔description contradiction remains where the body claims ownership the routing surface omits, and — when the description paraphrases a body trigger with **multiple clauses of different scopes** — that each clause's scope is preserved separately (the body trigger is the primary source; a first reading of it is hypothesis-grade); resolve all, or record per-owner why the coordinator is `unchanged`. **(b) Utterance-variant:** the same delivery type arrives as the canonical name PLUS restart/redo/from-scratch/continue variants; "the trigger exists but the utterance class is uncovered" is a validation-gate defect — enumerate the variants and add the high-value ones within the 800-char cap (route overflow to the body entry-precedence text). **(c) Listing budget (Claude Code):** check whether the skill's description was even IN the host listing — a name-only entry silently voids a keyword-based fix (the skill stays reachable by explicit `/skill-name`) — but conclude budget-dropped only from concrete evidence (this turn's listing shows it name-only, `/doctor` output, other host-listing proof) and STILL do the (a)/(b) fix regardless; a routing/bootstrap doc must not assume every description is always visible, and on non-Claude-Code hosts verify the host's own listing behavior from its primary docs (an LLM's "host bug" guess is hypothesis-grade). Failure shapes, the variant list, and the axis checklist: `references/description-authoring.md` (routing-miss fix checklist); the budget mechanism, the cold-start trap, and the levers: `references/skill-listing-budget.md`.
95
87
  - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings.
96
88
  - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it becomes extraction once the output is meant to change reusable skill behavior.
97
89
  - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim` — rebuild charter + target-output map and pass the dual-track before claiming landed.
@@ -125,12 +117,7 @@ Use this skill to turn observed experience into durable agent skills without cop
125
117
  - Extract behavior, decision rules, quality gates, evidence patterns, and routing boundaries; do not extract business nouns, repo names, IDs, one-off incidents, or stale implementation details.
126
118
  - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files. Each reference links one level from the entrypoint and stays inside the reference line budget; `references/attention-budget-ratchet.md` owns that budget, the write-side authoring norms, and the design invariants any size/budget gate must satisfy.
127
119
  - A skill must be executable, not only directional. For design, client, testing, debugging, or review skills, include concrete workflow steps, decision points, state/checklist coverage, and verification evidence so future agents do not produce work that is compliant but weak.
128
- - Design/client extraction must cover the judgment layer, not only the engineering layer. For UI/UX, extract aesthetic logic, interaction logic, behavioral logic, and user psychology from source evidence before landing rules about layout, components, breakpoints, or tests.
129
- - UI/UX judgment extraction must use observable proxies, not adjectives. Read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming behavioral or psychology rules. Use `references/uiux-judgment-extraction.md` for the required method.
130
- - UI/UX lessons usually route to multiple owners. Before editing, map each candidate to design, web, app, miniapp, testing, product workflow, or this extraction workflow using `references/uiux-routing-map.md`; do not land only the design rule when implementation or scenario testing is required. For mini-program lessons, `testing-strategy` owns layer/scenario selection, while `miniapp-product-dev` owns host-platform implementation, developer-tool or real-device evidence, review/release mechanics, and miniapp runtime constraints.
131
- - Judgment-layer extraction must name what changed. For UI/UX/client sources, record whether each judgment layer produced a new rule, confirmed an existing rule, narrowed an existing rule, or found no new evidence. If the pass only improves execution/validation, say so instead of implying new aesthetic, behavioral, psychology, or interaction knowledge.
132
- - For UI/UX/client extraction, the judgment-dimension axis enumeration lives in `references/uiux-judgment-extraction.md`. When the adjacency-scan rule fires on a UI/UX source, walk that enumeration — do not re-derive the axis list from memory.
133
- - A UI/UX judgment-delta row is not complete with labels such as `confirmed`, `narrowed`, or `no new evidence` alone. Each visual direction/tokens row must satisfy the field list in `references/uiux-judgment-extraction.md`; if those fields were not inspected, mark the row `pending` or `out of scope` and do not claim design-judgment extraction.
120
+ - UI/UX, Figma, frontend, app, miniapp, or client sources carry six entrypoint-level obligations — cover the four judgment layers (aesthetics / interaction / behavioral / psychology) from source evidence before any layout/component/breakpoint/test rule; use observable proxies, never adjectives — read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming a behavioral or psychology rule; route each lesson across design / web / app / miniapp / testing / product-workflow owners via `references/uiux-routing-map.md` (for mini-program lessons `testing-strategy` owns layer/scenario selection while `miniapp-product-dev` owns host-platform implementation, developer-tool or real-device evidence, review/release mechanics, and runtime constraints); name per-layer what changed (new / confirmed / narrowed / no new evidence) and say so when a pass only improves execution/validation rather than implying new judgment knowledge; walk the judgment-dimension enumeration instead of re-deriving it; and never close a judgment-delta row whose visual direction/tokens fields were not inspected — mark it `pending` or `out of scope` and do not claim design-judgment extraction. Their canonical, verbatim text is `references/uiux-judgment-extraction.md` §Entrypoint obligations (relocated from this entrypoint; low-frequency detail per the placement rule above) — load it before drafting any design/client rule; Step 6's UI/UX validation rows point there.
134
121
 
135
122
  ### Retrospectives, corrections & auto-triggered learning(复盘 / 纠正 / 自触发学习)
136
123
 
@@ -139,24 +126,18 @@ Use this skill to turn observed experience into durable agent skills without cop
139
126
  - **复盘 / 纠正 / retro / postmortem / bug-hunt / review-follow-up is an extraction by default, not a chat-only retro**. Triggers (full duty): any standalone "复盘" / "沉淀" / retro / postmortem invoked as such — **复盘 is an extraction by default, no missed gate required**; a missed testing/design/delivery/review gate followed by such a request or an equivalent user correction; any user correction about extraction quality; a retro/review/postmortem the user then challenges with "why wasn't this workflow used". Treat it as a **process defect**, not a conversation detail, and do **not resume the in-flight extraction until the RCA is recorded**:
140
127
  - (1) run RCA first — failed decision → root cause → missing skill rule/validation gate → durable prevention; (2) build a target-output map deciding for each kept lesson whether it belongs in a target skill, in this workflow, or both, mapping each to an owning skill/reference/process gate (include `skill-extraction-workflow` itself when the miss is routing/extraction/under-trigger behavior); (3) **land** the smallest durable prevention OR explicitly mark every candidate `unchanged`/`routed`/`discarded` with evidence (subject to bullet B path (i)'s always-land-here for failure-class exposures), and verify the **final diff matches the whole map** (not intent — every mapped target actually landed, not merely that one change is real). A chat-only owner recommendation is not enough.
141
128
  - When the user challenges why the workflow wasn't used, the earlier response is a **failed extraction, not a completed retro** — run the same correction RCA → map → land sequence; another chat-only explanation is not an answer. A final-response-only retro is allowed ONLY after a minimal RCA + candidate map where every candidate is explicitly `discarded`/`routed`/`no-durable-output`, with the source/evidence boundary stated. Ordinary bug-hunts / review follow-ups (not corrections) carry only the lighter duty — don't stop at the project outcome: decide target / this-workflow / both and make the diff match the map — without the stop-the-line RCA; routine bug/QA/review nits stay in their owner skill per the auto-trigger boundary below. If the extraction workflow itself allowed the shallow response, add the narrow guard HERE, not only in the downstream skill.
142
- - **Deferred-evidence over-polishing corrections** route to `product-rd-workflow`'s `DFE-CONT` rule: when a delivery keeps hardening tests/verifiers after the real/external evidence was unavailable, blocked, unreachable, deferred, skipped, or mock-substituted, or reports such real evidence as complete, the target-output map must name `DFE-CONT` (paraphrases like "why keep fixing tests when the cluster wasn't reachable" count — exact deferred/skipped wording is not required). A bare mention of runtime/external access is not enough: a test-hardening, coverage, or test-order correction with no unavailable/blocked/deferred real-evidence element and no deferral-as-terminal claim defaults to `testing-strategy`; when both a deferral signal and a test-order signal are present, map to both (`DFE-CONT` + `testing-strategy`, the test-case-first rule still mandatory). Validate by confirming the `DFE-CONT` block and its non-completion rule are present in the installed `product-rd-workflow` skill's `SKILL.md` (resolve via routing/skill discovery; inside the ccl-skills repo: `skills/product-rd-workflow/SKILL.md` — the file is not at that relative path on installed hosts), and keep this routing token in sync if that block is renamed.
143
- - If a retrospective correction says tests were run before test cases, or asks why test cases were not written first, the target-output map must include `testing-strategy` and any coordinating workflow such as `product-rd-workflow`. A final answer without a durable test-case-first prevention rule, validation command, and challenge or explicit no-update reason is only `interim`.
129
+ - **Correction-type routing for test / verifier / real-evidence corrections** is canonical, verbatim, in `references/correction-routing-map.md` (relocated from this entrypoint): a deferred-evidence over-polishing correction — a delivery that keeps hardening tests/verifiers after real/external evidence was unavailable, blocked, unreachable, deferred, skipped, or mock-substituted, or that reports such evidence as complete — must name `product-rd-workflow`'s `DFE-CONT` in the target-output map (a bare mention of runtime or external access is not enough; a test-hardening / coverage / test-order correction with no deferral element defaults to `testing-strategy`; both signals → both owners, test-case-first still mandatory; validate by confirming the `DFE-CONT` block and its non-completion rule are present in the installed `product-rd-workflow` skill's `SKILL.md`, and keep this routing token in sync if that block is renamed); if a retrospective correction says tests were run before test cases, the target-output map must include `testing-strategy` and any coordinating workflow such as `product-rd-workflow`, and a final answer without a durable test-case-first prevention rule, validation command, and challenge or explicit no-update reason is only `interim`. Consult that map whenever a correction names tests, verifiers, coverage, or unreachable real evidence.
144
130
  - **A reusable lesson must land in a SHARED artifact, and `skill-extraction-workflow` itself gets a prevention point**. Two paths with DIFFERENT strength:
145
131
  - (i) **exposure-by-failure-class** — when a reusable lesson is exposed by a routing miss, shallow retro, validation gap, wrong-ownership decision, or proxy/subset/slice/local-bound objective narrowing, **always** land at least one durable prevention point in this workflow itself (naming the **failure class, correction path, and validation gate** that stops recurrence), **even when the concrete lesson also belongs in another skill or project doc** — **no opt-out**; a target-skill-only landing is insufficient.
146
132
  - (ii) **explicit-ask** — when the user asks whether it was "沉淀到提炼技能" / "放进提炼技能" / "固化到 extraction workflow", the target-output map must include this workflow: either land a real prevention rule here OR state why it is `unchanged` and name the owning skill that received the rule (a target-skill-only update is insufficient when the failure was that extraction was skipped, shallow, or only oral).
147
133
  - Either way the landing must be a **shared** artifact — CCL skill, shared reference, validator, checklist, or project template — naming the exact trigger/gate teammates will hit; for a reusable routing/process/team failure, classify a memory-only landing as insufficient — local-only and "I'll remember next time" count the same (local memory supplements user/workspace context only). A candidate that turns out NOT genuinely reusable may be `discarded` with evidence. When neither path fires, ordinary `routed`/`unchanged` disposition to a different owning skill stays available per bullet A step 3.
148
134
  - For any analysis-parse-fix-test-challenge loop, separate five stages explicitly: analysis, parse/decompose, fix, test/verification, and challenge. Add a replay step when validating reusable lessons: rerun the same task shape or a close analog through the proposed workflow and check whether the required outputs and gates still appear in order. Keep four outputs explicit: the project-level fix, the test/verification evidence, the challenge findings, and the reusable workflow lesson. If the same pattern can recur across different domains, lift only the workflow lesson into `skill-extraction-workflow`; keep domain-specific implementation details in the owning project or target skill. See `references/analysis-parse-fix-test-challenge-replay.md` for the replay validation runbook.
149
135
  - **The agent failing to self-invoke this workflow (the user had to point out that `skill-extraction-workflow` should have been used) is a tracked failure class that recurs across the session / different tasks, not only within one extraction thread** (so the "twice in one extraction thread" scope above does not catch it). **Honesty:** an in-the-moment self-trigger is recognition-dependent — the always-on bootstrap layer raises its salience but is NOT a mechanical gate; do not overclaim a passive rule "fixes" the recurrence.
150
- - **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill", OR **enumerates what an external source has that we lack** (a gap list vs another pack; see `references/firing-point-placement.md`), that naming is a trigger to RECOGNISE the owner and load it** — not a licence to widen scope: shared-skill edits still need the authority you already have, so when the user's request covered only a status review or a narrow fix, record the extraction as `pending` with the owner named and ask rather than self-authorising a shared-skill change off your own suggestion.
151
- - The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` (per the closeout gate's evidence forms) — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check (was there a similar miss earlier this session, even on another task?) and, if so, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
152
- - **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — at the SAME transition, NOT another bootstrap/per-case bullet or more prose.
153
- - **Record-field corollary (the forgery surface):** any field that NAMES an owner — a checklist row, a CLI flag, a "decision:" slot — is fillable without invoking that owner, and filling it is what *feels* like discharging the gate, so it carries an explicit invoke bar on its triggered values.
154
- - The self-detect firing point's authority boundary and observed shape, the record-field corollary expansion (the invoke-bar coverage set-diff mechanics, the delegation-dispatch worked case), the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
155
- - **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge (it is the always-on generalization of self-audit-to-convergence, not a replacement for the gate).
156
- That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never buy a RED by disabling a guard against a shared or live dependency; where no isolated path exists record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point it at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A failing anchor is first a question about the ANCHOR, not a verdict on the implementation** (§Self-audit). **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
157
- A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`: `Y not run` blocks a done/complete/landed claim unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / verify / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above).
158
- The full self-adversary method — the mutation enumeration, the applied-mutation discipline, the independent-oracle validation, the re-owe-after-fixes rule, the graded-verdict calibration, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
159
- - Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent, preserved in the visible conversation or quoted exactly in trusted host-owned session/compaction state — never reconstructed, broadened, or substituted — or the user's correction literally naming the paused action and scope) — a semantic compaction paraphrase or a bare "why did you stop" complaint is not path-(b) authority, and a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid — asking again when neither binds, the user intervened, or scope/gates changed, and never copying real conversation text into a shared repository record. Do not let correction RCA or extraction extend a still-authorized delivery, and do not let stale assent bypass a newly pending or inconclusive gate. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. The full binding rules, the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
136
+ - **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill", OR **enumerates what an external source has that we lack** (a gap list vs another pack), that naming is a trigger to RECOGNISE the owner and load it** — not a licence to widen scope: when the user's request covered only a status review or a narrow fix, the extraction is recorded `pending` with the owner named, and the shared-skill edit waits for authority you were actually given.
137
+ - The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check across the whole session (even other tasks) and, on recurrence, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
138
+ - **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* at the SAME transition — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — NOT another bootstrap/per-case bullet or more prose. **Record-field corollary (the forgery surface):** any field that NAMES an owner is fillable without invoking that owner, and filling it is what *feels* like discharging the gate, so it carries an explicit invoke bar on its triggered values. The self-detect firing point's authority boundary and observed shape, the output-shape and option-set siblings, the invoke-bar set-diff mechanics, the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
139
+ - **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge. That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified; **a mutation you did not APPLY is a hypothesis, not evidence** (bound its blast radius; where no isolated path exists record the property `unverified`); **prove the oracle can fail before trusting its clean verdict**; **a failing anchor is first a question about the ANCHOR, not a verdict on the implementation**; a clean run is reported with the dimensions it crossed, walked before values; a property with no contradicting observation is `unverified`, never counted as audited. A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`, unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above). The full method — mutation enumeration, the applied-mutation discipline and its blast-radius bound, independent-oracle validation, the dimension walk (`testing-strategy` owns the axis list), re-owe-after-fixes, graded-verdict calibration, the failing-anchor section, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
140
+ - Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent preserved verbatim, or the user's correction literally naming the paused action and scope — never reconstructed, broadened, or substituted, and never copied into a shared repository record); a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid; asking again is required when neither path binds, the user intervened, or scope/gates changed; and neither correction RCA nor stale assent may extend a still-authorized delivery or slip past a gate that is newly pending or inconclusive. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. What counts as a binding (the two paths, and why a compaction paraphrase or a bare "why did you stop" is not one), the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
160
141
  - The trigger is a correction about *reusable skill/process behavior*, NOT every bug/QA/review nit handled inside its own owner skill. Do not wait for the user to say "沉淀": if the user points out a missed source, missed sibling skill, shallow rule, overclaim, domain leakage, missing trigger, missing verification, repeated correction, **or that this workflow should have been invoked at all (an under-trigger / "should you have used 提炼/复盘" correction, including outside an active extraction)**, run correction RCA, update the smallest owning skill/reference/validator, and verify the prevention point before finalizing the turn.
161
142
  - When the user asks whether a lesson was durably landed after a failed extraction, verify the actual skill diff or file content first. Do not answer from memory or intent. If the prevention rule is not present in the owning skill, add it or state that it has not been durably landed.
162
143
  - **Consolidate and retire rules; a skill's rule set must not grow monotonically.** Every correction adds a guard, but an N-bullet wall on one theme is itself the over-prescription/unreadability failure, and "just append another bullet" is how it regrows.
@@ -186,22 +167,14 @@ Use this skill to turn observed experience into durable agent skills without cop
186
167
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
187
168
  - Independent review is **dual-track** for any non-wording shared-skill change, deep extraction, multi-skill landing, new shared skill, or skill change that ships operational/architectural rules: run BOTH (a) a fact/consistency review (`codex review` or equivalent — catches inaccuracies, contradictions across references, sanitization gaps, over-prescription) AND (b) an adversarial challenge (`codex exec` with adversarial prompt — hunts race conditions, data-loss paths, security holes, algorithm flaws, operational footguns). Skipping the challenge mode is how P0/P1 production-safety issues survive into shared skills. The two are not interchangeable: review catches what's wrong; challenge catches what would break under chaos. See `references/dual-track-review-gate.md` for the runbook.
188
169
  - **Draft-time corollary — pre-cover the recurring first-draft blind-spot AXES before challenge (not security alone).** Sweep EACH applicable axis explicitly before handing off (record ≥1 negative case per axis, or a reasoned `not-applicable`): **(1) security / privacy / authority / data-loss**, **(2) concurrency & lifecycle**, **(3) resource bounds**, **(4) rollout / migration ordering**, **(5) over-broad absolute**, **(6) enumeration-completeness (the mirror of (5) — under-listing, not over-listing)**. A draft that reads clean on the happy path almost always misses ≥1 of these; the challenge is the safety net, not the first line. When drafting any rule/code that touches execution, deletion, optimization, deployment, capability/permission, concurrency, resource lifetime, or a cross-service rollout, self-cover the applicable axes first. "Pre-cover" means recording at least one relevant negative case — or, for axis (6), the set-diff against the complete set — (or a reasoned `not-applicable`) per applicable axis, NOT a "covered" badge — and it never narrows or downgrades the mandatory adversarial challenge. The per-axis instance lists: `references/dual-track-review-gate.md` (pre-cover axis detail).
189
- - **Repeated same-class adversarial-challenge findings across rounds are a design/scope smell, not just more patches** — and the resolution depends on WHAT recurs. When the challenge surfaces NEW instances of the SAME risk class in two or more rounds (e.g. each round finds another way one capability loses or corrupts data), treat it as a design smell: evaluate whether that capability should EXIST, not only how to patch this instance. Convergence-by-deletion (removing the risky capability) is often cleaner and more complete than convergence-by-patching, and is the right call when the capability serves no real need — judged from the convention's primary source and/or product/process evidence, not from the patch count alone. Record the decision explicitly as `keep / delete / narrow / replace` with: the same-class evidence (the rounds and findings), the real user need it serves, a safer alternative, and blast-radius/migration. This complements the challenge-round convergence standard in step 6 (which counts remaining P0/P1 fixes): same-class recurrence across rounds is the cue to question the design, not to open another patch round. Failure shape: an auto-rewrite-of-a-human-file capability spawns a new data-loss footgun every challenge round until it is deleted, after which the gate converges immediately.
190
- - **The cross-landing sibling — the same class recurring across separate LANDINGS (not rounds) means the fix shape itself is wrong.** The parent rule counts recurrence inside one gate's rounds; this fires when the *same* failure class returns in production after a previous landing already "fixed" it, which the within-round counter cannot see. The tell that distinguishes it from ordinary maintenance: each prior fix **re-instantiated the same predicate on new inputs** — registering the newly-seen value/name/version rather than changing what the check is predicated on.
191
- - When a control's predicate is the **vocabulary of an artifact the control does not own** (a field/value/version list from an upstream tool, format, or API), every upstream release re-breaks it, so the second occurrence is already the design signal: re-express the predicate over an **invariant the control does own** — shape, arity, type, or the independently-verified property the check actually needs — or accept the maintenance and say so. Record the same `keep / delete / narrow / replace` decision, plus which invariant replaced the vocabulary, and state the residual risk the looser predicate accepts. Failure shape: a reviewer-isolation gate pinned to an upstream CLI's init field names/enum values took every reviewer lane down on three separate upstream releases; the first two landings each registered the new vocabulary — which guaranteed the third — until the predicate moved to shape.
192
- - **The scope-direction sibling — recurrence that expands SCOPE into explicitly-deferred concerns is a controller-cut-scope signal, not another implement round.** When a challenge finding's remediation would push the artifact to implement concerns the design/spec/architecture explicitly defers to a later phase (a future domain, an unbuilt consumer topology, a not-yet-enabled capability), that is a **classification checkpoint, not a cut** (the checkpoint fires on first appearance; sustained *undispositioned* recurrence across rounds is the reviewer-lane stop/reframe escalation — a legitimately-repeated already-accepted/deferred finding, re-verified candidate-relative against the `pre-existing & out-of-scope` freshness bar, is not): before implementing, establish the finding's current-phase impact and split a compound finding, with any scope-cut ratified by a distinct risk owner (never the proposing controller). The failure it prevents: the reviewer silently becomes a scope-expansion engine and the artifact gold-plates. Failure shape: a first-cut state-machine contract ballooned across rounds as the challenge kept inventing machinery for an explicitly-deferred future domain, until the controller reset to the in-scope diff and it converged. The mechanics (current-phase-impact test, compound split, canonical-disposition routing, and a tracked follow-up bound to the deferred phase's entry gate so a cut defers rather than discards) live in `references/dual-track-review-gate.md`'s `Findings, autonomous budget, and human authority`.
170
+ - **Repeated same-class adversarial-challenge findings across rounds are a design/scope smell, not just more patches** — and the resolution depends on WHAT recurs. When the challenge surfaces NEW instances of the SAME risk class in two or more rounds, evaluate whether that capability should EXIST, not only how to patch this instance: convergence-by-deletion is often cleaner and more complete than convergence-by-patching, and is the right call when the capability serves no real need — judged from the convention's primary source and/or product/process evidence, not from the patch count alone. Record the decision explicitly as `keep / delete / narrow / replace` with the same-class evidence (the rounds and findings), the real user need it serves, a safer alternative, and blast-radius/migration; this complements the challenge-round convergence standard in step 6 (which counts remaining P0/P1 fixes) — same-class recurrence is the cue to question the design, not to open another patch round (observed: an auto-rewrite-of-a-human-file capability spawned a new data-loss footgun every round until it was deleted, after which the gate converged immediately). Two siblings share that decision record: **cross-landing** — the same class returning in production after a previous landing already "fixed" it means the fix shape itself is wrong, the tell being that each prior fix **re-instantiated the same predicate on new inputs**; when the predicate is the **vocabulary of an artifact the control does not own** (a field/value/version list from an upstream tool, format, or API) the second occurrence is already the design signal — re-express it over an **invariant the control does own** (shape, arity, type, or the independently-verified property the check actually needs) or accept the maintenance and say so, recording which invariant replaced the vocabulary and the residual risk the looser predicate accepts; **scope-direction** — a finding whose remediation would implement concerns the design explicitly defers to a later phase is a classification checkpoint, not a cut or another implement round: establish current-phase impact, split a compound finding, and let only a distinct risk owner (never the proposing controller) ratify a scope cut, so a cut defers rather than discards — the failure it prevents is the reviewer silently becoming a scope-expansion engine while the artifact gold-plates (observed: a first-cut state-machine contract ballooned across rounds until the controller reset to the in-scope diff). Failure shapes and mechanics: `references/dual-track-review-gate.md` (`Findings, autonomous budget, and human authority`; design-time operability check) and `references/external-practice-controls.md` (predicates over a vocabulary the control does not own).
193
171
  - **Global-iteration boundary:** any `STOP`/`revert`/terminal wording in this gate applies to the reviewer lane, readiness claim, or defective dependent slice, never automatically to the overall task. Repeated root cause or no-progress rounds trigger method/design/validation changes or a parked human-decision item while independent runnable work continues. Only an explicit authenticated human stop ends the overall iteration.
194
- - **The should-it-exist question also applies at DESIGN TIME for any new mechanical gate, validator, or evidence apparatus, AND for any change to its verdict in EITHER direction — a loosening must be checked hardest, since it removes evidence instead of adding a red — run the check before landing AND before offering a human design options.** Four legs — **author-dogfood** (the authoring workflow passes it end-to-end under CI's base resolution), **marginal-cost** (what the cheapest routine change costs), **trust-model fit** (what it defends against; for a loosening, which class stops owing evidence), **premise check** when tightening (a clean run on the current corpus is not evidence). Per-leg method and failure shapes: `references/dual-track-review-gate.md` (design-time operability check).
172
+ - **The should-it-exist question also applies at DESIGN TIME for any new mechanical gate, validator, or evidence apparatus, AND for any change to its verdict in EITHER direction — a loosening must be checked hardest, since it removes evidence instead of adding a red — run the check before landing AND before offering a human design options.** Legs: **author-dogfood** (the authoring workflow passes it end-to-end under every base resolution CI uses), **marginal-cost**, **trust-model fit** (for a loosening, which class stops owing evidence), **premise check** when tightening (a clean run on the current corpus is not evidence). Per-leg method and failure shapes: `references/dual-track-review-gate.md` (design-time operability check).
195
173
  - Reference files containing CLI commands, external API calls, or runnable code examples ship operational behavior and are NOT exempt from the dual-track gate; treat them the same as skill core-rule sections. Specifically, any reference file with recommended runnable commands or code that encodes operational behavior, external-state mutation, security-sensitive behavior, deployment/release/audit workflows, or nontrivial reusable recipes — including dangerous anti-patterns a teammate could plausibly copy — must pass both the fact/consistency review and the adversarial challenge before landing. Harmless validation commands and truly non-runnable static illustrative snippets (e.g., pseudocode, redacted/placeholder examples) remain eligible for a trivial-scope challenge skip only when the changed artifact is not a shared-skill change; any non-wording shared-skill change still requires challenge.
196
174
  - For reference files or skill examples that include CLI commands or API calls, treat challenge findings as valid until specifically verified or refuted. Use the live tool (`--help` or equivalent), official docs, or version evidence to verify flag existence and semantics. If the live tool is unavailable or the finding is semantic/security-related rather than flag-existence-related, keep the finding `pending` — do not discard it and do not land the example as complete until verification is resolved, unless the finding is explicitly accepted or deferred under the dual-track gate's documented rules.
197
175
  - When adding a new section or materially editing any existing section, rule, or reference pointer in a skill, cross-check the changed content against the host skill's existing Core Rules before landing. A contradiction or unresolved gap with the host skill's own rules is a P0 issue and must be resolved before the change is landed as complete; split the change or downgrade to interim if the conflict cannot be resolved in the same logical landing set.
198
176
  - **A change that TIGHTENS or removes a previously-allowed path (an escape hatch, a permissive default, an "only when X you may Y" carve-out) is an impact-chain edit, not a one-spot edit.** Grep for every surface that asserted the OLD allowance — the skill body, the **always-on `bootstrap`/SessionStart layer**, AND the **`description` routing surface** — and tighten ALL of them in one landing, or a stale permissive bootstrap/description silently re-permits what the skill now forbids (a P0 cross-surface contradiction). Adding a rule cross-checks the host's own Core Rules; subtracting an allowance additionally cross-checks the always-on layer and the routing surface.
199
- - Reference example code that performs read-modify-write on external mutable state (Bitable records, database rows, file contents, API state) must:
200
- - Read failure: raise explicitly; never return `{}`, `""`, `None`, or any empty-success value that silently drops the prior state.
201
- - Write failure: raise or skip, never silently continue.
202
- - Uniqueness invariant: when the example declares, assumes, or depends on one, detect violations such as duplicate unique keys and raise before propagating bad state.
203
- - Lost-update control: use optimistic concurrency controls (ETag, version field, CAS, transaction, or compare-and-swap), append-only API semantics, or an explicitly declared single-writer precondition — RMW examples that silently assume no concurrent writers will produce lost-update bugs under normal conditions.
204
- - Data-loss anti-pattern: log-and-continue after a read failure on an append-only field.
177
+ - Reference example code that performs read-modify-write on external mutable state must satisfy the five RMW obligations — explicit raise on read failure (never an empty-success value), no silent continue on write failure, raise on a violated uniqueness invariant, an explicit lost-update control (optimistic concurrency, append-only semantics, or a declared single-writer precondition), and no log-and-continue on an append-only field — canonical, verbatim, in `references/validation-and-landing.md` §Read-modify-write example code (relocated from this entrypoint).
205
178
 
206
179
  ## Do Not Extract When
207
180
 
@@ -331,4 +304,4 @@ Use this skill to turn observed experience into durable agent skills without cop
331
304
  - Extraction planning and coverage: `references/source-to-skill-extraction.md`, `references/source-register.md` (required before editing target skills or claiming full/complete coverage), `references/coverage-exhaustion-traps.md` (before exhausted/complete/no-gap/fully-extracted claims), `references/evidence-card-template.md` (L0/L1 evidence record), `references/review-feedback-mining.md` (source = human review feedback on merged changes: author filter, diff-fact adoption, idempotent classification, operator decides).
332
305
  - UI/UX and multi-skill routing: `references/uiux-judgment-extraction.md`, `references/uiux-routing-map.md`, `references/firing-point-placement.md`, `references/skill-listing-budget.md`, `references/online-skill-review.md`, `references/harness-patterns-and-eval.md`.
333
306
  - Validation and authoring gates: `references/validation-and-landing.md`, `references/description-authoring.md` (required before routing/workflow/cross-cutting `description` rewrites), `references/rule-consolidation.md` (required before rewriting/retiring/consolidating rules or drafting behavior-shaping rules; its form-by-failure table owns the form), `references/l0-l1-l2-routing.md` (required for `discard`/`route`/L0–L2 or gate-strength decisions), `references/review-rubric.md` (dual-track support; does not replace `references/dual-track-review-gate.md`), `references/external-practice-controls.md` (public-evidence trust boundaries: research, leaf delegation, behavioral evidence).
334
- - Retrospective/correction support: `references/resume-paused-delivery.md` (premature-stop resume requires affirmative continuation).
307
+ - Retrospective/correction support: `references/resume-paused-delivery.md` (premature-stop resume requires affirmative continuation), `references/correction-routing-map.md` (required when a correction names tests, verifiers, coverage, test order, or deferred/unreachable real evidence).
@@ -27,6 +27,17 @@ The read side already defends against oversized files (chunked reads under ~200
27
27
  - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
28
  - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
29
29
 
30
+ ## Retirement and relocation signal (usage census)
31
+
32
+ The ratchet only stops growth; it never says *what* to retire or relocate, and "not pulling its weight" is an author's opinion until something is measured. Three independent lines converge on the same instrument: context-evolution methods keep per-bullet usage counters (helpful/harmful marks) and prune or merge on them rather than on a single monolithic rewrite that collapses detail; trajectory-distillation work places broadly applicable procedure in the root document and *lower-frequency* detail in auxiliary files, and finds joint consolidation over many traces stronger than order-dependent one-lesson-at-a-time edits; and the official skill-authoring guidance tells authors to watch how the agent navigates a skill — a bundled file the agent never accesses is unnecessary or poorly signaled, one it reads on every run belongs in the entrypoint. The repo instantiation:
33
+
34
+ - `scripts/reference-access-census.sh [--skill <name>] [--days <n>]` reads the host's own agent transcripts (Claude Code and Codex session logs; both consume this tree) and prints, per `SKILL.md`/reference file, how many sessions in the window mentioned it and when it was last touched (the last-touched column must be the newest touching transcript's date, never the oldest or an unparsed value). Counts only — no transcript text, prompts, absolute log paths, or session ids — so the output is safe for a private charter; it is still per-host data and never lands in the shared tree.
35
+ - **Placement by firing frequency, not only by kind.** The entrypoint's content-placement rule sorts by kind (trigger / routing / core workflow / non-negotiables stay; detail moves). Add the frequency axis: a rule that fires on a narrow source class or correction type (one client surface, one correction shape, one artifact kind) is *low-frequency detail* even when it is non-negotiable, and must live verbatim in the reference the entrypoint already points at, with a one-bullet summary that keeps the load-bearing obligations inline (`rule-consolidation.md` condensing rule). Ledger `file:` anchors and script pins name a path, so a pinned phrase stays where it is — enumerate them before choosing what moves.
36
+ - **Advisory, never a gate** (Goodhart, same as the health roll-up): a count that becomes a target gets gamed by mentioning files. A zero-session reference is a relocation/merge *candidate* that still owes the zero-loss obligation map; a high-share reference is a promotion candidate, not an automatic move. Pair the census with the closeout cost row in `extraction-quickstart.md` §4 so that a round records what it cost and what it retired, and the next round can tell whether the corpus is shrinking toward the cap or only holding level.
37
+
38
+ - **A mention count is not an open count — classify HOW the file is reached before relocating on a census figure.** The census matches the path anywhere in a transcript line, so a file that a gate names in its own output, that an agent appends to, or that a bounded `sed`/`tail`/`grep` touches for one row scores the same as one an agent loads whole. Before a census figure justifies a split, a promotion, or a retirement, the read shape must be counted in the same window — whole-file reads versus bounded reads versus gate echo versus writes — and let the whole-read count, not the mention share, carry the read-side cost argument. Observed: the append-only ledger scored the package's highest mention share, and the read-shape count showed whole-file loads in a small minority of those mentions, with bounded reads and gate echo making up the rest — the split its share seemed to demand would have bought nothing on the read side.
39
+ - The census is host-format-bound (it matches fixed-string skill paths inside transcript lines). When no transcript is found, or any transcript is unreadable or vanishes mid-scan, it prints `reference_access_census_unevaluated` and withholds the table (exit 2 on input errors) — never zeros, and never a path: a run that could not evaluate must say so with counts only. An explicitly supplied log root that does not exist must be treated as an input error (an absent default root is normal), overlapping or repeated log roots must count a transcript once, a scan batch that fails without writing stderr must still withhold the table, the scan must not misreport an all-no-match batch under an inherited errexit, and no usage error may echo the caller's argument.
40
+
30
41
  ## Official-clause verdicts (provenance)
31
42
 
32
43
  Registered claims were re-verified against the primary source (Anthropic "Skill authoring best practices", docs.claude.com, read 2026-08-31) before landing; per-clause disposition:
@@ -0,0 +1,22 @@
1
+ # Correction Routing Map
2
+
3
+ Use this when a retrospective, review follow-up, or user correction names tests, verifiers, coverage, test order, or real/external evidence that was unavailable, blocked, deferred, skipped, or mock-substituted. It decides which owner(s) the target-output map must name for that correction type; the extraction workflow itself still gets its prevention point per the entrypoint's always-land rule.
4
+
5
+ The two rules below were relocated verbatim from `SKILL.md`'s `Retrospectives, corrections & auto-triggered learning` Core Rules group (low-frequency detail per the entrypoint's content-placement rule); the entrypoint keeps a one-bullet summary that points here. Wording changes here go through the same shared-skill gates as an entrypoint edit.
6
+
7
+ ## Deferred-evidence over-polishing corrections
8
+
9
+ - **Deferred-evidence over-polishing corrections** route to `product-rd-workflow`'s `DFE-CONT` rule: when a delivery keeps hardening tests/verifiers after the real/external evidence was unavailable, blocked, unreachable, deferred, skipped, or mock-substituted, or reports such real evidence as complete, the target-output map must name `DFE-CONT` (paraphrases like "why keep fixing tests when the cluster wasn't reachable" count — exact deferred/skipped wording is not required). A bare mention of runtime/external access is not enough: a test-hardening, coverage, or test-order correction with no unavailable/blocked/deferred real-evidence element and no deferral-as-terminal claim defaults to `testing-strategy`; when both a deferral signal and a test-order signal are present, map to both (`DFE-CONT` + `testing-strategy`, the test-case-first rule still mandatory). Validate by confirming the `DFE-CONT` block and its non-completion rule are present in the installed `product-rd-workflow` skill's `SKILL.md` (resolve via routing/skill discovery; inside the ccl-skills repo: `skills/product-rd-workflow/SKILL.md` — the file is not at that relative path on installed hosts), and keep this routing token in sync if that block is renamed.
10
+
11
+ ## Tests-before-test-cases corrections
12
+
13
+ - If a retrospective correction says tests were run before test cases, or asks why test cases were not written first, the target-output map must include `testing-strategy` and any coordinating workflow such as `product-rd-workflow`. A final answer without a durable test-case-first prevention rule, validation command, and challenge or explicit no-update reason is only `interim`.
14
+
15
+ ## Decision table
16
+
17
+ | Correction signal present | Owners the target-output map must name |
18
+ | --- | --- |
19
+ | Deferred / blocked / unreachable / mock-substituted real evidence, then continued test or verifier hardening, or such evidence reported as complete | `product-rd-workflow` (`DFE-CONT`) + this workflow |
20
+ | Test-hardening, coverage, or test-order correction with no deferral element and no deferral-as-terminal claim | `testing-strategy` + this workflow |
21
+ | Both a deferral signal and a test-order signal | `product-rd-workflow` (`DFE-CONT`) + `testing-strategy` + this workflow (test-case-first still mandatory) |
22
+ | Tests were run before test cases / why were test cases not written first | `testing-strategy` + the coordinating workflow (`product-rd-workflow`) + this workflow |
@@ -88,3 +88,10 @@ When the source is a task/session that PRODUCED runnable artifacts (scripts, com
88
88
  ## Gate (4) — the per-item verdict-disposition obligation (relocated gate detail)
89
89
 
90
90
  The verdict owes a **per-item disposition row EVEN WHEN THE ROUND LANDS NOTHING** (no commit fires to catch an advertising-level verdict): each row records the load-bearing surface actually read (section/line/rule — NOT headers/`When to invoke` blurb/feature-list), the owning organization artifact (`file:rule`) OR `route-to-tool` / `out-of-scope: <reason>`, and for `covered` the mechanical firing path. A `quick`/`light`/`sweep`/`triage` framing does **NOT** waive the load-bearing read. Record the row on a **concrete surface** — the triage turn's final response OR persistent scratch / `source-register`, never chat-ephemeral. (Honesty: absent a host/review hook checking that surface, this stays recognition-dependent — salience, not mechanical enforcement.) A repeated user "深度分析了么 / did you deep-read" across rounds is the recurrence signal that an advertising-level verdict shipped; treat it as a validation-gate defect.
91
+
92
+ ## Gates (1) and (2) — the unit lists (relocated gate detail)
93
+
94
+ The Core Rule states the two gates; this is the detail the entrypoint no longer carries.
95
+
96
+ - **Gate (1), the terminal-status row.** The code-corpus row's terminal status is one of `deep-read`, `inventory+owner-mapping`, or a downscope that cites an actual user instruction — a self-declared downscope is not terminal. The depth note under *Core* above governs what that sweep usually looks like and why its outcome is recorded rather than skipped.
97
+ - **Gate (2), the enumeration unit.** "Top-level structure" means the artifact's own next unit, chosen so every part is either `read` or `skipped`: for a document its `##`/`###` sections, or TOC / pages / anchors / line-chunks when sections are absent; for a repository its top-level directories, or package / module / test / script / config units when the top level is flat. The trap recurs even within one artifact — a doc read to its third section with the remaining sections un-enumerated supports no exhausted claim over the doc.
@@ -173,3 +173,29 @@ Re-run the challenge after each fix-up. Multiple rounds are normal: in practice
173
173
  ## Iteration after deployment
174
174
 
175
175
  Even a carefully written description will miss real-user phrasings. Plan to collect actual miss / over-trigger reports for 1–2 weeks after landing, then increment the trigger list and Skip-when based on data. Do not pre-emptively stuff in every imaginable phrase — that fails the 80% threshold and creates new collisions.
176
+
177
+ ## Routing-miss fix checklist — the three axes (relocated from `SKILL.md`)
178
+
179
+ The Core Rule names the axes; this section carries the checks and the observed failure shapes. Walk all three before any `description`/trigger edit lands; a repeated miss of the same request type across two rounds means the previous fix skipped one.
180
+
181
+ ### (a) Coordinator-vs-executor
182
+
183
+ - A fix that advertises the request type only on the executor's description is incomplete: *multi-stage deliveries* of that type (spec → plan → test-first → impl → verify) keep auto-routing to the executor and skip the coordinator's lifecycle gates.
184
+ - Check 1 — does the coordinator's description advertise this request type with a scale qualifier — narrow/single-file work goes to the executor, multi-stage/cross-cutting work to the coordinator?
185
+ - Check 2 — does the body claim ownership that the routing surface leaves out (a body↔description contradiction)?
186
+ - Check 3 — when the description paraphrases a body trigger or skip-condition that has **multiple clauses with different scopes** (clause 1 fires on *any edit* of a surface, clause 2 only on a *semantics change*), preserve each clause's scope separately; collapsing a multi-scope trigger to its most salient clause silently over- or under-fires. The body trigger is the primary source, so re-read every clause (the first reading is hypothesis-grade) before asserting the description over/under-covers it or rewording it.
187
+ - Resolve all three checks, or record for that owner why the coordinator stays `unchanged`.
188
+ - Failure shape: a trigger word ("重构"/refactor) was added to stack `*-dev`/`*-architecture` descriptions while the coordinator workflow's description never advertised it, even though the coordinator's BODY already claimed ownership — a body↔routing-surface contradiction that kept multi-stage refactor deliveries on the executor.
189
+
190
+ ### (b) Utterance-variant
191
+
192
+ - The same delivery type arrives under many utterances — the canonical name PLUS restart / redo / from-scratch / continue variants. A *refactor delivery* arrives as "重构 X" but also "重新开发", "完全重新开始", "推倒重来", "清除代码重新开发", "redo/rewrite from scratch"; advertising only the canonical phrase leaves the variants unmatched, so the workflow silently fails to auto-trigger on them.
193
+ - Enumerate the restart/redo/continue variants of the request type and add the high-value ones within the 800-char cap; overflow goes to the body's entry-precedence text.
194
+ - Failure shape: a repeated miss of the SAME delivery type arriving via a different utterance across rounds — the axis was skipped the first time.
195
+
196
+ ### (c) Listing budget (Claude Code)
197
+
198
+ - Before fixing a routing miss with a trigger-word edit, check whether the skill's description was even IN the host listing; a name-only entry voids a keyword-based fix while the skill stays reachable by explicit `/skill-name`.
199
+ - Evidence bar (do not over-apply as a catch-all): treat budget-dropped as established only on concrete evidence — this turn's listing shows that skill as name-only, `/doctor` output, or other host-listing proof. Absent that, treat it as a hypothesis and STILL do the (a)/(b) fix; a *visible* description misjudged as name-only is the reverse trap.
200
+ - Corollaries: a routing/bootstrap doc that assumes "every description is always visible" is wrong under budget pressure and must not be a low-traffic gate skill's sole discovery path; on non-Claude-Code hosts check the host's own listing behavior against its primary docs — any LLM's "it's a host bug" guess is hypothesis-grade until so checked.
201
+ - Mechanism, cold-start trap, and levers: `skill-listing-budget.md`.
@@ -100,7 +100,7 @@ The self-audit rule says to prove the oracle can fail before trusting its clean
100
100
  - **A pipeline reports its LAST stage's status.** `make test | tail -20` exits with `tail`'s status, so a failing `make` reads as exit 0; the same holds for `| grep`, `| head`, and any `$?` read after a pipe. Redirect and read the command's own status (`cmd > log 2>&1; echo $?`) or set `pipefail`. Observed: a full-suite failure was reported as passing, and the mistake was caught only because the summary line the suite prints on success was missing from the captured tail.
101
101
  - **A check run against the wrong tree passes silently.** When the work lives in a git worktree, an ambient `cwd` the harness may reset between tool calls sends relative-path commands to the primary checkout, which is clean — so the gate evaluates a tree that does not contain the change. Resolve the target by absolute path (`git -C <abs>`, `bash <abs>/script.sh <abs>`); `worktree-isolation` owns that discipline. Observed: a catalog test's pristine-tree case passed against the primary checkout while the real candidate was blocked.
102
102
 
103
- Both are the same defect as a check that can only ever say clean: if you cannot produce the red path on demand, the green is not evidence.
103
+ Both are the same defect as an oracle that cannot produce red: if you cannot produce the red path on demand, the green is not evidence.
104
104
 
105
105
  ### A new gate owes the negative half (relocated from `SKILL.md` §behavioral-evidence)
106
106
 
@@ -524,7 +524,7 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
524
524
 
525
525
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
526
526
 
527
- Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
527
+ Five is the generic `code-review` transport ceiling, not this extraction lane's spend. Non-wording Agent-autonomous extraction calls go through `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1` per chain: one review plus one challenge. **The lane spans at most two chains and three rounds; the third exists only because a fix batch moved the candidate.** Holding fixes keeps the challenge on the frozen round-1 candidate, so the batch that lands is unreviewed until a succeeding chain challenges it — and a fix touching a selected owner's `SKILL.md` or `references/**.md` moves that owner digest and ends the first chain anyway. The trigger is the candidate, never a disposition label the author writes: **landing hash equal to the challenged hash owes nothing; different owes one succession challenge bound to what lands.** There the autonomous lane ends. An authenticated human may request later review, but that is separately attributed human-requested evidence outside this chain/budget, never an additional Agent round. Unused generic capacity never authorizes automatic continuation. The closeout validator rejects referenced receipts whose recorded budget is not the wrapper-fixed value, rejects any post-chain round that is not a succession, and checks budget and ordering consistency within the caller-supplied set. `scripts/review_ledger_binding.py` is its merge-side half: it recomputes the candidate with the controller's own packet freeze and refuses a landing whose evidence binds a different one. Evidence lives outside the reviewed paths, so committing the ledger cannot move the hash it records. A candidate larger than one packet is not split as a pull request but as a review: `--print-manifest --partition <paths> [--partition <paths> ...]` renders a landing partition manifest whose path partitions cover every changed file exactly once, each partition hashing to what `--print-candidate --paths <partition>` answers; commit the manifest with one validated closeout ledger per partition, and the gate recomputes every partition and refuses a manifest whose parts do not add up to the whole (an uncovered or overlapping file, a partition that no longer reproduces, a base other than the fork point, or an aggregate hash that does not reproduce its partitions). An integration branch that accumulated several reviewed rounds is promoted as one pull request without a new ledger: when neither a single ledger nor a manifest binds the promotion, the gate walks HEAD's first-parent chain down to the first commit already on the target and rebinds each round merge in a detached checkout of its second parent against its first parent, judged with the landing tree's own controller and validator rather than the round's (a round could carry a hollowed validator that a later round restores); a merge whose second parent is already on the target is a sync merge and owes nothing; every step must be exactly the automatic merge of its parents (a hand resolution or an extra file in the merge commit is refused as unreviewed), a non-merge commit on the chain is refused, and the chain is consulted only for the default path set. Rounds that appended to the same register therefore no longer force the promotion to be split by round. It cannot authenticate that the wrapper produced those receipts or that the caller retained every earlier chain or receipt. The wrapper does not mint or persist `review_chain_id` or `autonomous_review_index`: the caller still supplies both, and could start a fresh-looking chain after the final round. The validator detects bad order inside the referenced set but cannot detect a prior chain the caller omitted, so complete caller-owned ledger retention—and treating an Agent reset as a contract violation—remains part of the boundary rather than a property the local scripts prove.
528
528
 
529
529
  **Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
530
530
 
@@ -45,6 +45,8 @@
45
45
 
46
46
  - 冻结 task-bank `eval/routing-tasks.jsonl`:每行 `{id, utterance, expected_skill, acceptable?, must_not_route_to?, source, why_expected, frozen_at_sha}`,种子取自 source-register 历史 miss + bootstrap 路由规则。**已修复的路由 miss 必须把其 utterance 冻结成 bank task 落在同一交付里**(修复不冻结=下次同类漂移无回归面)。`expected_skill: "none"` 是否定对照/覆盖空洞哨兵:正确结果是没有技能认领;`acceptable` 列出可辩护替代结果(如空洞探针上 coordinator 接管与拒绝都对);`must_not_route_to` 点名吸入诱饵邻居。结构由 `test_routing_bank_integrity.sh` 确定性把关(sentinel 只准出现在 expected/acceptable,不准进 must_not)。
47
47
  - **冻结案例神圣(regressions-are-sacred)**:已冻结案例的删除或判定面改写(bank task 的 expected/acceptable/must_not,golden trace 的 assert 块)是一次回归裁决事项,与邻居回归同权——平均改善不得抵消单条冻结案例的失守,且「曾 yes 现非 yes」**含降级为 unsure/INCONCLUSIVE** 都算回归。删除/改判的每条案例必须在同一轮的 register 追加行里写 `case-retired: <id>` 或 `case-rescoped: <id>` 并给理由,交独立评审裁决;确定性半边由 `test_frozen_case_sanctity.sh` 按 `CCL_SKILL_BASE_REF` 把关——无裁决行即红,无 base ref 时打印显式 skip token(skipped ≠ passed),它只保证交易可见,不裁决交易正当性。
48
+ - **描述触发评测的三条纪律**(Anthropic skill-creator 的 description-optimization 形态):① utterance 要真实——带文件名、口语、错字、上下文,不写抽象短句,否定例要是**近似负例**(共享关键词但该走别的技能),不是明显无关句;② 触发率按重复取——each bank utterance must be graded at least three times per description version(`--replicas 3` 起),单次 pass/fail 不是触发率;③ 用 grader 自动改写 description 时,bank 必须按 60/40 切成训练/held-out,只按 held-out 分选优,否则描述会过拟合 bank(本仓尚无自动改写循环;这条是它的前置条件,不是已有能力)。
49
+ - **改 bank 的那一轮必须同轮重建基线**——runner 按 `bank_sha256` 与 `replicas` 判定「不同尺子」并抑制 diff,所以往 bank 增删一行就把此前**每一份**基线永久作废:不是变旧,是不可比。旧基线仍留在 `eval/evidence/` 里读着像基线,没有任何东西提示它已被孤立,于是路由面可以长期无人测量而全绿。改 bank 的交付要么同轮跑出新基线并落 evidence(报告自带 `bank_sha256`/`replicas`,将来才可比),要么在 register 追加行里显式写明旧基线自此孤立、下一轮补测。失败形态:bank 从 136 行长到 158 行,其间六条 description 变更、一个技能新增,而唯一的全量基线停在三周前且与当前 bank 在任何 replicas 设置下都无法比对——这一点直到有人主动传 `--baseline` 才由 runner 自己说出来。
48
50
  - grader = 每 task(×replicas)一次本机 `claude --print --tools "" --model <haiku>`,喂 utterance + 全部 skill description(agent 真正路由的那份面),要 `{selected_skill|none, clarify, confidence, rationale_short}`。**clarify 率、低置信率(<0.5)、副本一致率是一等报告字段**,不是旁注——路由质量的残余风险常在"高置信直选却选错、无自纠路径"这类 pass/fail 看不见的分布里。
49
51
  - **路由兼容性信号,不是真值预言机** —— grader 自己可能错;只衡量"当前 description 能否让廉价模型把固定 utterance 路由到 expected"。
50
52
  - **advisory**:不接 `check-ccl-skills.sh`,不挡 merge。退出码:`0` = 跑完;`2` = 用法;`3` = grader 整体不可用(claude CLI 缺失则打印 skipped 后 `0`)。路由 miss 永不非 0。
@@ -55,8 +57,14 @@
55
57
 
56
58
  以下是生成可比较证据的默认协议。**降级的是「F4 自己充当统一合并门禁」这个声称,不是「必须测、且必须有人裁决」这个义务**——这两件事分开:落地判断交给本轮实际的 owner/风险/评审门禁,但**测量本身不可选**。任何动 routing 面(SKILL.md description、task-bank 判定面)的改动都必须按下列协议产出证据;没跑就是没收敛,不得进入独立评审、也不得声称本轮无回归。采用不同样本量时,须随工件记录理由,且该理由与本轮证据一同进入独立评审——「记了理由」本身不是豁免,自审通过的理由不构成已裁决:
57
59
 
60
+ **筛查分辨率 ≠ 行动分辨率(是两个数,不是一个)**:全量基线默认 `--replicas 3`,它是**筛查器**——把候选捞出来,不是给结论。判定面是保守共识(任一副本 FAIL 即 FAIL),于是三副本下一条真实认领率 80–90% 的用例读起来就是红的,一次孤立偏离读起来就是一条 finding。**任何按用例采取的行动——改它 owner 的 description、判它是回归、判某条冻结期望已过期——都要求那条用例自己有 ≥10 个有效观测**;runner 对任何 `replicas < 10` 的报告打 `screening_resolution_only`、并在 JSON 里落 `action_resolution: false`,两侧的 10 由 `test_eval_routing_bank_resolution.sh` 钉在一起,防止文档与执行体漂移。实测形态(115 轮,同一轮内三次反向出错):3 副本全量基线报 14 条失败,10 副本下其中 5 条是抖动、六条已起草的 description 改动全部建立在它们上面;其中 `p3-spec-then-tc` 的单次偏离被据以论证某冻结期望已过期,10 副本下该对手 3/17、论证撤回;`ab-c5` 被判成「改前 PASS、改后 FAIL」的邻居回归,钉在未改动树上的对照臂显示它改前就是 8/10 的边缘失败。**留在 `eval/evidence/` 里的三副本基线因此是候选清单,不是 findings 清单**——引用它开轮的人要先为自己要动的每条用例补齐观测。
61
+
62
+ **地板管「能不能动手」,不等于「点估计已经准到能和阈值比」**:10 次有效观测在 70–85% 区间的抽样误差约 ±15–20 个点。实测形态(115 轮,同一个候选、同一条用例 `ab-b5`):一次 10 副本得 5/10(50%),紧接着 20 副本得 17/20(85%),合并 22/30(73%)——单看前者会判成「稳定失败」并据以改描述,单看后者会判成「健康」。所以**任何按阈值分档的判断(稳定失败 / 边缘 / 抖动)必须读合并观测,落在约 45–75% 之间的读数在 10 副本下不构成判定**,要么补到 30 次以上,要么如实记成「区间未定」。同理,改前/改后的差值也按合并观测比:本轮那条 8/10→5/10 的「回归」在 24/30 vs 22/30 下相差两次命中,不可分。
63
+
58
64
  1. 动任何 description 之前必须先跑 **≥10 轮有效观测**的稳定性基线,把稳定失败与抖动分开;抖动不得作为修改依据(grader 超时/不可解析轮不算有效观测,须补跑)。
59
65
  2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
66
+ **`newly_failed` 是候选,不是回归判定。** runner 的 `--baseline` diff 在三副本下按保守共识判 status,于是一次孤立偏离就把一条用例记进 `newly_failed`;而这个集合**每跑一次就换一批**。实测(115 轮,同一条分支上四次全量三副本运行):`{route-opencode-project-config, route-nodejs-arch}`、`{skip-pytest-cmd, ctrl-ai-risk, miss-refactor-python-unqualified, route-nodejs-arch}`、`{mem-api-log-redact, route-nodejs-arch}`——除 `route-nodejs-arch` 外每一条只出现过一次、再未复现,逐条做成对 20 副本探针后**无一可归因于该轮改动**(两例两臂分布完全相同,一例两臂都红,一例合并后相差两次命中)。所以:`newly_failed` 的每一条都要按「同一用例、改前/改后两棵树、合并 ≥20 次观测」复测才能称为回归,不得直接写进轮记录当回归清单;同样地,不得因为它每轮都有内容就把整轮判红。
67
+
60
68
  3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
61
69
  4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
62
70
  5. **单变量归因**:一次改前/改后对照只准动**一个路由变量**(一条 description,或同一 skill 不可分割的一组路由面)。同时动多条 description 的批量改动,其对照差值不可归因到任何一条,只能按整包回归读——要归因就拆成逐条 A/B。(源侧实测形态:仅替换一条 description 的成对子集对照,把命中从约 2/3 提到 95%,且提升可归因到那一条改动——多条同动时这句话说不出口。)
@@ -48,7 +48,7 @@ Disposition:
48
48
  - Local trust model: register rows remain honest-but-fallible workflow evidence, not a hostile-author security boundary. The gate proves that a changed, normative, owner-scoped rule line (or changed owner executable) exists for every claimed firing path; it does not prove a model run occurred, that a named executable implements the claimed enforcement (a shebang stub passes the static check), that a mangled or ambiguous ledger row was honest (those are warned, not blocked, to avoid false positives on other table shapes), author identity, or non-tampering by an authorized contributor. Independent review/challenge and the fixed checker remain the assurance case.
49
49
  - Machine format (relocated from the `SKILL.md` firing-mechanism rule; the local evidence policy above carries the rationale): every added source-register row must carry `behavioral-evidence: RED-baseline` (any observed delta — `observed-failure: yes` requires it) or `semantic-control` (only with `observed-failure: no`), an `observed-failure: yes/no` state, and an owner-scoped `firing-path` — each declaration in its own semicolon-delimited fragment of the cell (`…prose; behavioral-evidence: …; observed-failure: …; firing-path: …`), so a key embedded mid-prose never parses as a declaration. The firing-path anchor is at least 16 characters, occurs once in the file and once in its round's added lines, and lands on a numbered/list Markdown rule with a normative action. A row that survives at HEAD must also resolve to an owner this range actually changes — an owner reverted to its base bytes by a rebase or a base-side conflict resolution leaves the changed set while its row stays behind, and the row then vouches for a change the delivered diff does not contain. There is no author-declared escape from this: a corrective rewrite that back-fills a row for a round which merged red produces the same shape, and it is a deliberate, person-adjudicated repair that can adjudicate this refusal too.
50
50
 
51
- - **Round scoping — a row is judged against the round it landed in, never the accumulating range.** A row is authored against one round's diff, so reading the whole `base..HEAD` range to classify it judges the row against work it never described. That mismatch produced both directions of the same defect: an already-gated row turned red once a LATER round touched the same owner (which is what the ledger's superseded-row notes were absorbing), and a description-only round lost its routing-surface locator because an EARLIER round had edited that owner's body. The gate cuts rounds at the commits that touch the ledger, walked first-parent so one merged worktree round is one boundary, and each round spans from the previous boundary so work commits sit in the round whose ledger append describes them. The partition is derived from git alone — an author cannot nominate, widen, or move their own scope.
51
+ - **Round scoping — a row is judged against the round it landed in, never the accumulating range.** A row is authored against one round's diff, so reading the whole `base..HEAD` range to classify it judges the row against work it never described. That mismatch produced both directions of the same defect: an already-gated row turned red once a LATER round touched the same owner (which is what the ledger's superseded-row notes were absorbing), and a description-only round lost its routing-surface locator because an EARLIER round had edited that owner's body. The gate cuts rounds at the commits that touch the ledger along a first-parent line, and each round spans from the previous boundary so work commits sit in the round whose ledger append describes them. A merge that git rebuilds from its two parents is expanded into its branch's own rounds, so a merged worktree round is judged exactly as its pull request was — the same history must not partition differently after it lands; a merge git cannot rebuild (a hand resolution, a conflict) keeps a single boundary at the merge, so content that came from neither parent is never left in no round. The partition is derived from git alone — an author cannot nominate, widen, or move their own scope.
52
52
  - Both obligations move together, in opposite directions. **Classification** narrows to the round: whether a diff is wording-only, an identifier retarget, or description-only is asked of that round's bytes, which is what makes a verdict stable once it lands. **Presence** narrows to the round too: the round that changed an owner is the round that owes the row, so owner work committed after a ledger append can no longer ride on an earlier round's row. Narrowing classification without narrowing presence would have opened exactly that laundering route.
53
53
  - **Known exception, inherited not introduced: a renamed-away owner escapes round presence.** The round's subject set is intersected with the cumulative one, and the cumulative pass drops an owner whose package was renamed away (a row citing it would be refused for naming a SKILL.md that no longer exists). So an owner changed substantively in an earlier round, with an intervening ledger boundary that closed that round without its row, and renamed away in a later round, is demanded by neither round. Differential check on the same fixture: the pre-round-scoping gate passes it too, so this is not a regression of the narrowing — but the per-round presence contract above does not hold in this shape, and saying so is the point of this bullet. Closing it needs the evidence-file existence check to resolve at the row's round head instead of the worktree, which is its own change.
54
54
  - The owner-level `RED-baseline` floor deliberately stays **cumulative**: it asks whether anything in the owner's whole change is left unvouched-for, and per-round would let a package self-clear on the one round that happened to be punctuation-only. Cumulative is the stricter of the two readings, so the round scoping cannot loosen it.
@@ -114,6 +114,8 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
114
114
  | [OpenAI, *GPT-5.1 Prompting Guide*](https://cookbook.openai.com/examples/gpt-5/gpt-5-1_prompting_guide) | check-conflicts-first; a published metaprompt recipe for finding contradictions in your own system prompt | same class |
115
115
  | [RECAST](https://arxiv.org/html/2505.19030) | joint satisfaction degrades sharply as constraint count grows — and it is already low at small counts. The metric for "all of them at once" is the paper's **OSR** (§4.1: "the HSR of all constraints, both rule-based and model-based, that are successfully satisfied simultaneously"). Across all 29 model rows of Table 1 the **ceiling** on OSR is **25.0** at Level 1, falling to **19.0 / 13.0 / 13.5** at Levels 2–4, whose constraint counts are **5 / 10 / 15 / all** (§B.4). So at five constraints no model held the whole set more than about a quarter of the time, and by fifteen none exceeded ~13% | benchmark paper proposing its own dataset and method — a low baseline flatters the contribution; these are one hard benchmark's order of magnitude, not a usage failure rate, and its constraints are generation-task instruction constraints rather than preconditions of a procedure, so transfer is by analogy. **Citation corrected 2026-08 against Table 1:** the widely quoted 39.75% is the **Average column of the single best-by-average row** (Gemini-2.5-Pro) — the arithmetic mean of that row's twelve MSR/RSR/OSR cells (sum 477; 477/12 = 39.75, confirmed) — **not** an all-constraints-satisfied rate, and it must not be cited as one. The two orderings differ: Gemini leads on Average while Qwen3-235B-A22B holds the highest Level-1 OSR, so do not carry "best model" across from one column to the other. Cite the OSR ceilings above |
116
116
 
117
+ | [IFScale](https://arxiv.org/abs/2507.11538) | instruction-following accuracy degrades as instruction density rises (500 keyword-inclusion instructions; best frontier model 68% at max density across 20 models / 7 providers); three degradation shapes correlated with model size and reasoning; a **bias toward earlier instructions** (primacy), and omission as a distinct error category | benchmark paper on a synthetic keyword-inclusion task — density and primacy transfer by analogy only; it measures a list of independent constraints, not a procedure's preconditions; note it reports primacy where the vendor guides report recency, so position is a bias with no single direction |
118
+
117
119
  **Why recency is a hazard, not a rule.** Vendor guidance reports that models *tend to follow* whichever instruction sits later — an observation about behaviour, not a licence to resolve conflicts by position. Two ways position becomes dangerous if read as a rule: a later permissive line beats an earlier stricter one (directly contradicting `Conflict Resolution`, which keeps the stricter data-loss/security/contract guard); and text embedded in **untrusted data** — a diff under review, a retrieved document, tool output — sits later within the same authority level and would win by placement alone, which is prompt injection with extra steps. Treat recency as a bias to design against: put the load-bearing rule where the decision happens, and never let placement confer authority.
118
120
 
119
121
  **How to use them.** Vendor guidance and benchmarks are **hypotheses with good provenance** — they tell you what to test on your own corpus, they do not substitute for testing it. Citing them as settled is the same error as landing an unverified claim; so is overruling one with an underpowered probe.
@@ -134,6 +136,10 @@ Referenced from `SKILL.md`'s "The mechanism underneath" rule. This section holds
134
136
 
135
137
  **谓词选择是这条里唯一承重的设计决定**:hook 认的是**工具身份**,不是「已知坏 flag」的清单。flag 清单是上游项目拥有的词表,上游每发一版就把控制打回原形——这正是本文件上一节所述「控制建立在自己不拥有的东西上」的形态,也是这个类跨多次落地反复回来的原因。工具名集合缓慢、有限、可由本仓拥有;flag 集合不是。代价是明写的:未列入的工具是**不触发**,这是刻意接受的残留,扩列表的判据是该工具已积累出记录在案的 flag 失败。
136
138
 
139
+ ### Cross-landing recurrence: the reviewer-isolation instance and the decision record (relocated from `SKILL.md`)
140
+
141
+ The same predicate-ownership rule has an English-side worked instance and a required record. When a failure class returns in production after a previous landing already "fixed" it, and each prior fix registered the newly-seen value, name, or version on the same predicate rather than changing what the check is predicated on, the second occurrence is already the design signal: re-express the predicate over an **invariant the control owns** (shape, arity, type, or the verified property the check really needs), or accept the maintenance and say so. Either way the landing records a `keep / delete / narrow / replace` decision, **the invariant that took the vocabulary's place**, and **the residual risk accepted by the looser predicate**. Failure shape: a reviewer-isolation gate pinned to an upstream CLI's init field names and enum values took every reviewer lane down on three separate upstream releases; the first two landings each registered the new vocabulary — which guaranteed the third — until the predicate moved to shape.
142
+
137
143
  可执行形态:
138
144
 
139
145
  - 给列表内的外部 CLI 传长 flag 之前,**必须先在本机读过该 CLI 的 help 输出**;未读过就用,属于把记忆当一手源。
@@ -93,7 +93,7 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
93
93
  - Choose the review tier from that table, not from intuition. Do not restate the rows locally; record the exact `dual-track-review-gate.md` table row used. Record `challenge: not-required` only when that row classifies the actual diff as challenge-not-required (for shared skills, this means strict wording-only with deterministic scope proof + independent review confirmation). Non-wording shared-skill changes cannot skip challenge.
94
94
  - Run deterministic checks and implementer self-review first, and record what each proves before invoking review/challenge (this self-review-before-review ordering applies to every non-wording shared-skill change the dual-track table requires review for, not only the rows that look high-risk): `git diff --check` proves whitespace/conflict-marker hygiene only; validators prove schema/link/routing invariants; leakage/sanitization scans prove only their configured patterns; scope checks must name the changed files or expected file set; the self-review row is conclusive only when each required field is non-empty (acceptance criteria, changed-file scope, edge/failure paths, known residual risks) and the changed-file scope equals the candidate diff's changed-file set, or explicitly explains any excluded generated/irrelevant file. Persist it before the review/challenge run in a fresh, non-overwritten task-evidence path outside the candidate diff, pass that exact file as the gate's review plan, and retain the gate result that binds its profile hash; do not edit the candidate merely to record self-review or review outcome, because that creates self-referential candidate churn. A candidate-local row is appropriate only when the row itself is a substantive deliverable under review. A plain in-place-editable MR description or scratch log is not ordering proof unless its edit history is retrievable and checked; a backfilled row is invalid and forces a rerun. If the candidate diff changes after the row is saved — a file added/removed OR the content of any listed file materially changed — refresh the row; any rerun of review/challenge against the new candidate draws on the remaining cross-chain Agent budget (the self-hosted-chain rule in `references/dual-track-review-gate.md`), and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path governs instead of a rerun. Changing only the external self-review record refreshes the profile binding; it does not by itself invalidate implementation tests or the candidate packet. A missing field, "ok" placeholder, mismatched scope, or unprovable ordering makes the row inconclusive. Do not spend LLM review rounds on issues a script or implementer-side checklist can decide. If the independent pass is the first place basic scope, contract, privacy, or test issues surface, apply those findings to the diff, close the self-review gap, and rerun the deterministic gates before rerunning review/challenge; the process-defect repair is in addition to resolving the findings, not a way to discard or downgrade them.
95
95
  - Review pass: persist the complete self-review row and encode it in the review plan. For a **non-wording** lane, resolve the repository-owned `scripts/extraction_review_gate.sh` and use it from round 1; never substitute the generic controller, scan writable plugin roots, or supply a caller-selected budget. For a strictly proven **wording-only** lane, use the generic `code-review` proof-bound single-review recipe in `code-review/references/staged-review-contract.md` and record `challenge: not-required`; require its controller-derived wording scope plus the independent `wording_only_boundary` confirmation. This is the only extraction path that stays outside the multi-round wrapper and terminal ledger; the gate, not this page, decides whether a chainless review is legal, and it may still demand the tracked pair. Take all controller options from that runnable recipe, supplying the actual stage and exact candidate rather than an example default. The non-wording chain cannot be retrofitted, so a run started outside its owner wrapper is thrown away and restarted. Read the chain-opening and packet-composition rules in `references/dual-track-review-gate.md` first. Require conclusive JSON, selected-client attribution, packet/profile binding, family exclusion, and wrapper runtime evidence. When the host returns a live execution handle (`session_id`, `cell_id`, or equivalent), keep polling that exact handle until terminal exit; empty current output is progress, not a verdict, and no replacement/fallback reviewer may start while the original process is live. The result row records handle type, an opaque host transcript/tool-call reference and terminal exit status. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle. Never copy a credential-like raw handle into shared evidence. `findings` is not pass; inconclusive, malformed, or free-form output stays interim. Do not add a separate behavior probe.
96
- - Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh` separately with the same plan, stage, candidate, family and tracked chain. Pass the next one-based index; later rounds include a distinct focus and all prior focuses. Preserve a separate result row with the same binding, exclusion, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
96
+ - Challenge pass: for a non-wording lane, invoke `scripts/extraction_review_gate.sh` separately with the same plan, stage, candidate, family and tracked chain. From the second round on, the packet's `--paths` exclusion list must exclude every evidence JSON the round has already added (receipts, dispositions, closeout files) — the merge-side binder excludes exactly those, so a packet that excludes only receipts binds a different candidate than the one that lands and the round is thrown away (observed twice in consecutive rounds); bound evidence such as base attestations and excerpt files is committed before the round, never after. Pass the next one-based index; later rounds include a distinct focus and all prior focuses. Preserve a separate result row with the same binding, exclusion, egress, attribution and conclusive checks. Review never satisfies challenge; missing or inconclusive required challenge keeps extraction interim. A wording-only lane has no challenge pass.
97
97
  - Treat review/challenge as batch-level gates over the landing candidate, not as a per-bullet or per-line edit loop. Apply all findings from a round; when both lenses are required and cross-chain Agent budget remains, re-run both on the updated candidate before landing — every re-run sums into the same wrapper-fixed budget, and at the cap, or at the effective exhaustion that rule defines, the terminal-disposition path in `references/dual-track-review-gate.md` replaces further re-runs.
98
98
  - Skipping a required challenge = work can only land as interim, not complete.
99
99
 
@@ -119,7 +119,8 @@ For maintainers running a fresh codebase / Figma / doc extraction. Read this fir
119
119
  - File: `~/.<host>/skills/.extraction-work/<project>-completion.md`
120
120
  - Final state: which batches done, which deferred, which sources unavailable.
121
121
  - Lessons: what surprised; what would change in next extraction; what to add to skill-extraction-workflow.
122
- - For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. When the whole candidate exceeds one packet, split the review rather than the pull request: `--print-manifest --partition <paths> ...` renders a landing partition manifest, and one validated ledger per partition plus the committed manifest is what the gate binds. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
122
+ - Cost row (process toil is measured, not felt): review/challenge rounds run, findings fixed / accepted / deferred, wall-clock from charter to PR, and net body-word delta per touched entrypoint (`scripts/check-size-budget.sh` prints it). A round that grew a `severe_debt` entrypoint's references without retiring anything records that as the outcome; the next round must read this row before deciding its batch shape.
123
+ - For every non-wording review chain, build the receipt-bound closeout ledger and run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting a terminal state. Ask `scripts/review_ledger_binding.py --print-candidate` what the landing candidate hashes to first: if the held fix batch moved it, the ledger owes the succession challenge bound to that hash, and the same script is the merge-side gate that refuses a landing whose evidence binds a different candidate. When the whole candidate exceeds one packet, split the review rather than the pull request: `--print-manifest --partition <paths> ...` renders a landing partition manifest, and one validated ledger per partition plus the committed manifest is what the gate binds. When an integration branch that accumulated several already-bound rounds is promoted as one pull request, no new ledger is owed: the gate walks the branch's first-parent chain and rebinds each round merge at its own base, provided every merge is the automatic merge of its parents and nothing was pushed to the branch outside a round. A clean Round 2 challenge plus its exact-candidate completion receipt may validate as `ready_for_human_decision`; Round 2 findings at the exhausted budget validate as `continuation_authorization_required`; a second ordered base drift validates as `baseline_race`. Unknown, stale, omitted, or invalid evidence remains `interim`. The strict wording-only single-review path records its independent review row and does not fabricate a multi-round ledger — but note the cost the merge-side gate imposes on it: that gate accepts only a validator-checked ledger, because it cannot authenticate a hand-writable receipt, so a wording-only change that touches the bound paths still owes the two-round chain before it can land.
123
124
 
124
125
  ### 5. Provenance migration
125
126
 
@@ -154,6 +155,7 @@ Skip this step when nothing transferable surfaced.
154
155
  | Source-read fallback ladder | `SKILL.md` Source-read remediation | When a source read fails or times out |
155
156
  | Sibling mini-map | `SKILL.md` Step 4 stack-specific updates | Every stack-specific change |
156
157
  | Private alias map `audit_cmd` | `~/.<host>/.private-aliases/<project>.yaml` or process-retro profile | Every commit's R0 audit |
158
+ | `scripts/reference-access-census.sh` | this skill package | Before relocating/retiring entrypoint or reference text; advisory usage counts from local transcripts (`references/attention-budget-ratchet.md`) |
157
159
 
158
160
  ## Typical timings
159
161