@ccoalm/ccl-skills 0.9.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +7 -7
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -11
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +9 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +29 -31
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -1
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +32 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +1 -1
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +30 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +15 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +25 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +27 -21
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +25 -15
  33. package/dist/assets/release.json +79 -24
  34. package/package.json +1 -1
@@ -314,6 +314,11 @@ self-review plus explicit task reframing; it does not silently create a new
314
314
  Agent budget. An untracked challenge is one-off advisory evidence; it cannot
315
315
  enter a later Agent round or satisfy the local completion checkpoint.
316
316
 
317
+ Two consequences follow from those stable bindings and must be planned for before round 1:
318
+
319
+ - The selected-owner digest hashes each selected owner package's current working tree, and owners derive from the candidate's own paths — so a candidate edit inside any selected owner package invalidates every prior receipt and the next tracked round fails `review_chain_invalid`. For a self-hosted candidate (a skill-repo diff editing the package that owns it) that is nearly every applied fix — one confined to files outside every selected owner drifts only the candidate hash and may continue in-chain: "do not reset Agent authority" promises no continuation, and the in-chain tolerance for older candidate hashes is reachable only while the fix stays outside its selected owners.
320
+ - A chain restarted after such a break re-enters the same cumulative Agent budget and must never be counted as fresh authority; the bounded restart recipe for the extraction lane (batched dispositions, cross-chain round accounting, full-context first packet, terminal disposition at the cap) is owned by the extraction workflow's dual-track gate reference.
321
+
317
322
  The controller is stateless and prevents accidental/cooperative resets only. A
318
323
  trusted host or platform must retain the ledger when hostile local callers are in
319
324
  scope; a repository-local counter cannot authenticate human authority.
@@ -7,7 +7,7 @@ description: 加功能 / 新需求 / 技术方案 / 方案评估 / 技术选型
7
7
 
8
8
  Use this skill as the top-level workflow for new product development, feature delivery, bug handling, refactoring, release preparation, or repeated process improvement. It is not a replacement for stack-specific skills; it decides which skill should own each stage and what evidence is required before moving on.
9
9
 
10
- **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack or execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
10
+ **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack/execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
11
11
 
12
12
  **Continuation-proposal output contract (session-wide for product delivery).** Every assistant message in a delivery session routed by or through this workflow carries exactly one literal line until the user explicitly ends/pauses the delivery or changes scope: `proposed-next: <action and scope>` when the message has imperative/future/next-step wording or a proposed action, otherwise `proposed-next: none — status only`. At the start of every subsequent user turn, read that line before interpreting the reply: absence, multiplicity, or a marker/wording conflict enters `blocked:`/`interim` by default, never `not-applicable`. Coverage detail and the host-layer caveat: `references/pre-final-continuation-gate.md` (Continuation-proposal output contract).
13
13
 
@@ -21,11 +21,11 @@ Use this skill as the top-level workflow for new product development, feature de
21
21
  - A general-purpose process skill that merely *looks* like the obvious start — brainstorm/scope-shaping, plan-writing, or TDD auto-suggested by ANY channel: a session-start prompt, an optional skill package, or the host platform's native skills listing (including a listed entry skill's own self-invocation mandate, e.g. "must invoke if there is a 1% chance") — does not replace this entry: suggestion-channel wording is channel self-promotion, not routing authority (a host-mandated preflight — mandated by a host-authored system/developer-level or equivalent higher-priority instruction — may run first without thereby becoming the delivery owner; the test is AUTHORSHIP, not rendering position: a host-authored instruction counts even when rendered within the listing surface, while a skill's own description/content claiming preflight status never does); invoke this workflow as the delivery entry (immediately after any genuine host-mandated preflight), then call that skill inside the stage it serves (for example, requirement shaping in Workflow step 1).
22
22
  - For delegated agent execution, `multi-agent-delegation` owns the execution recipe and worker verification while this workflow owns the lifecycle gate and acceptance boundary; delegated agents resuming after a pause must receive or re-verify the current plan/spec artifact set before editing.
23
23
  - For multi-repo delivery, or any delivery that changes remote branch, MR, pipeline, release, or deployable-artifact state, maintain a compact per-changed-unit delivery-status ledger (row schema + persistence rules in `references/status-tracker-sync.md`) before claiming done, recommending MR/merge, or choosing the next slice. A required verification/review gate passes only when its state is `success`, `not-applicable`, or `not-required`; any other state blocks a done/merge recommendation, so remediate, wait to a terminal state, or report the delivery as pending with the unblock action. Residual-risk acceptance by the user only permits the recommendation/handoff label for that concrete action; it does **not** authorize merge, auto-merge, default-branch push, or cleanup, which still require the user's explicit merge instruction for the current MR (per `worktree-isolation`). Small local-only multi-file edits use the normal concise status unless they introduce remote, CI, MR, release, or deployable-artifact state.
24
- - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack or execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
24
+ - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack/execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
25
25
 
26
26
  ## Scope
27
27
 
28
- - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed or stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
28
+ - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed/stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
29
29
  - **Implementation completeness/minimality gate.** For behavior-changing delivery this gate fires — functional completeness and structural minimality are independent gates — and gaps block `complete` — **load `references/implementation-completeness-and-minimality.md` before the mapping**: map every in-scope acceptance point to implementation plus fresh evidence, and speculative future need and omitted required behavior both fail — the independent-gates rule and the acceptance/concept matrices live there, and every retained new concept must map to a current acceptance point or hard constraint.
30
30
  - Product requirement artifacts: route clarification to `requirement-intent`, conditionally-required current-state evidence to `requirement-baseline`, scope/version/appetite to `requirement-scope`, and Ready-only human-readable PRD assembly to `requirement-doc-writer`. All four use the canonical `requirement-doc-writer/references/requirement-closure-contract.md`. In active delivery, each narrow artifact returns here after completion; only a standalone non-PRD narrow artifact may return directly to the user. “只写个 PRD / 不走流程”仍需 lifecycle-issued Ready;WIP/会议材料不得命名为 PRD. This workflow owns the complete lifecycle's `PRD Ready` / `PRD Not Ready` verdict and cross-owner closure coordination.
31
31
  - Design routing: decide when interaction, information architecture, visual system, design-system, or UX acceptance work must happen before implementation.
@@ -35,9 +35,9 @@ Use this skill as the top-level workflow for new product development, feature de
35
35
  - Defect routing: use `defect-diagnosis` for hands-on reproduction, isolation, instrumentation, fix, regression verification, root-cause analysis, and prevention routing.
36
36
  - Research routing: For research groundwork behind a selection or assessment decision (技术选型、方案评估前的主题调研), this workflow must call `multi-perspective-research` for an evidence-grounded brief; the verdict itself stays with this workflow's gates.
37
37
  - Existing-project assessment routing: for requests that ask to analyze a repository, product codebase, or project quality across architecture, implementation, tests, UI/UX, bugs, or risks, run codebase understanding first, then route each assessment dimension to the smallest owning skill instead of treating understanding as the final answer.
38
- - AI/algorithm product launch discipline: for new or iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
- - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness or freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
- - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new or public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
38
+ - AI/algorithm product launch discipline: for new/iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
+ - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness/freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
+ - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new/public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
41
41
  - **Artifact-egress confidentiality gate**: before a delivery artifact (spec/plan/requirement, status/writeback doc, launch/task card, retrospective) — or any text generated from it — **crosses the local trusted boundary**, run a confidentiality pass *before* the write/create/push/send call; it fires **only on cross-boundary egress** and owns only the **semantic confidentiality axis** that secret scanners miss. A block stops the *egress* but preserves the draft locally and never deletes the work; exception authority stays with the user/owner. Egress channels include chat/tracker surfaces (Feishu/Bitable writeback, shared/public docs, MR/issue body or review comment, external trackers), durable VCS/release metadata (commit message, branch/tag name, release note/changelog, CI metadata), external assistants/models, and delegated-worker prompts; the gate's own block reason/finding is itself egress — never quote the raw sensitive span across the boundary; secrets/PII/credentials/raw-logs/customer-data route to their existing owners (`platform-observability` redaction, `defect-diagnosis` evidence sanitization, the `feature-risk-router` security-review gate); per-category actions, severity, and the delegated-worker firing point: `references/artifact-egress-confidentiality.md`.
42
42
  - Learning loop: after bugs, review findings, incidents, repeated friction, or external skill research, use `skill-extraction-workflow` to update the right skill or reference instead of leaving knowledge only in chat.
43
43
  - Skill/process extraction: when asked to summarize delivery experience, preserve a workflow lesson, update a reusable skill, or decide where a lesson belongs, route to `skill-extraction-workflow` first. Update this skill only when the lesson changes product R&D routing, gates, ownership, or lifecycle policy.
@@ -195,7 +195,7 @@ Run this gate before finalizing a product R&D turn after any delivery slice land
195
195
  - **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — **load it before selecting `continuing:` on any assent**, and any concrete next-slice proposal you issue must itself carry the `proposed-next:` marker or a later assent cannot bind — an unmarked referent is ambiguous, never self-cleared; the rule fires only when the immediately preceding assistant message itself states one concrete next action and its scope, and `continuing:` binds to that proposal, never to adjacent status or response-format prose; ambiguous assent, referent, or authority ⇒ `blocked:` with step-4 precedence — restate the proposed action/scope plus the specific ambiguity/authority, cite the step-1 evidence and ask one concise question in the same turn; self-classifying the reply or marker away is never an exit, and the `continuing:` default applies only when assent is unambiguous and no step-4 condition holds (the full rule and its fallback are stated there); the visible `continuing:`/`blocked:` outcome obligation is unchanged.
196
196
  3. Continue automatically only when no step-4 stop condition fires and: the next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned, verifiable with existing commands, and needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision, or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
197
197
  4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a failed/pending/inconclusive required/blocking gate, dirty/conflicting worktree that can't be isolated, required environment unavailable after remediation, high-impact product/architecture/compliance decision, destructive action, external purchase/financial commitment, unclear owner, ambiguous assent, missing stricter authorization, materially differing viable approaches (none dominant-and-reversible), a fix lacking evidenced cause, or no low-risk slice. Exactly one dominant reversible approach and no other stop condition firing: do not stop at a recommendation: deliver a tested reviewable draft.
198
- 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form, turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
198
+ 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
199
199
  6. **Assent-outcome closeout check.** Before every final response in a product-delivery session, walk the literal immediately preceding marker and visible outcome; the agent cannot exclude a status, question, review, or dispatched-owner turn by reclassifying it outside the session. An action marker or a plausibly affirmative user reply requires exactly one already-visible `continuing:`/`blocked:` outcome; until `continuing:` has executed the accepted slice or `blocked:` has named the blocker, the current turn may not use `proposed-next: none — status only` or `not-applicable`. A valid status-only marker permits `not-applicable` only when the preceding prose has no imperative/future/next-step wording and no affirmative reply pending. A missing/multiple/conflicting marker forces a visible `blocked:`/`interim` outcome with one clarifying question; it never produces `not-applicable`. The current assistant message itself must end with exactly one action-form or status-only marker for the next turn. Omitting the marker cannot justify a silent stop.
200
200
 
201
201
  If a user later challenges "why did you stop" or "was the rule too weak", treat it as a product workflow defect: route through `skill-extraction-workflow`, strengthen the smallest owning skill or validation checklist, validate the diff, and only then claim the process issue is solved.
@@ -123,7 +123,7 @@ Use this skill to turn observed experience into durable agent skills without cop
123
123
  ### What to extract, content placement & domain (UI/UX) judgment(抽什么 / 内容放置 / 领域判断)
124
124
 
125
125
  - Extract behavior, decision rules, quality gates, evidence patterns, and routing boundaries; do not extract business nouns, repo names, IDs, one-off incidents, or stale implementation details.
126
- - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files.
126
+ - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files. Each reference links one level from the entrypoint and stays inside the reference line budget; `references/attention-budget-ratchet.md` owns that budget, the write-side authoring norms, and the design invariants any size/budget gate must satisfy.
127
127
  - A skill must be executable, not only directional. For design, client, testing, debugging, or review skills, include concrete workflow steps, decision points, state/checklist coverage, and verification evidence so future agents do not produce work that is compliant but weak.
128
128
  - Design/client extraction must cover the judgment layer, not only the engineering layer. For UI/UX, extract aesthetic logic, interaction logic, behavioral logic, and user psychology from source evidence before landing rules about layout, components, breakpoints, or tests.
129
129
  - UI/UX judgment extraction must use observable proxies, not adjectives. Read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming behavioral or psychology rules. Use `references/uiux-judgment-extraction.md` for the required method.
@@ -178,9 +178,9 @@ Use this skill to turn observed experience into durable agent skills without cop
178
178
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
179
179
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
180
180
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
181
- - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=2` (three rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled round 4/5. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
181
+ - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=1` (2 rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled new rounds. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
182
182
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
183
- - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At round 3 validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
183
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At budget end validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
184
184
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
185
185
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
186
186
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
@@ -240,7 +240,7 @@ Use this skill to turn observed experience into durable agent skills without cop
240
240
  - For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
241
241
  - Trigger situations and users/tasks it should serve.
242
242
  - What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
243
- - For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
243
+ - For NEW skills and subjective/high-impact skills (design, UX, client, product workflow, architecture, review), must define eval/pressure scenarios, baselines, acceptance criteria pre-draft.
244
244
  - For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
245
245
 
246
246
  3. Inventory evidence.
@@ -262,17 +262,15 @@ Use this skill to turn observed experience into durable agent skills without cop
262
262
  4. Extract candidate rules.
263
263
  - Convert observations into reusable rules: "when X, do Y, verify Z".
264
264
  - Separate invariant rules from stack-specific examples.
265
- - Before editing any stack-specific skill, build the sibling-generalization mini-map: source stack, sibling stacks, shared workflow owner, per-sibling decision, and reason. If a candidate is language-agnostic, route it to the shared workflow/testing/architecture skill first, then add stack-specific implementation notes only where needed.
265
+ - Before editing any stack-specific skill, build the sibling-generalization mini-map the Core Rules owner-generalization group defines; route a language-agnostic candidate to the shared owner first, then add stack-specific implementation notes only where needed.
266
266
  - Mark each candidate as keep, merge, discard, or route to another skill.
267
267
  - Map each kept or routed candidate to its owning target in the target-output map before editing. Do not finish a design/client extraction until implementation, testing, product workflow, and sibling-client implications have been checked and either updated or explicitly marked unchanged. For mini-program surfaces, include `miniapp-product-dev` in that owner check.
268
268
  - Keep an extraction ledger during analysis: rule origin (`observed` or `hypothesis`), source IDs inspected before the rule was drafted, evidence grade, candidate rule, conflict, decision, target skill/reference, and reason.
269
269
 
270
270
  5. Generalize and place content.
271
- - Put trigger, routing, core workflow, and non-negotiable rules in the skill entrypoint.
272
- - Put detailed reference material in direct reference files.
271
+ - Place content by the Core Rules content-placement rule (entrypoint owns trigger, routing, core workflow, and non-negotiables; direct reference files own the detail, within their budget).
273
272
  - Generalize from the evidence ledger, not from a polished rule draft. Do not search for examples to justify a rule that has already been written.
274
- - Choose capability names that describe what future users can do with the skill, such as complex workspace patterns, service architecture, test strategy, or document tightening. Do not name durable outputs after the source file, source project, old feature, or temporary extraction task.
275
- - Keep source identity in provenance fields only. If a source name is useful for audit, label it as source evidence; do not make normal users route through that source name to understand or trigger the skill.
273
+ - Name durable outputs per the Core Rules naming rule (complex workspace patterns, service architecture, test strategy, document tightening are capability names; the source file, source project, old feature, and this extraction task are not), and do not make normal users route through a source name to understand or trigger the skill.
276
274
  - Turn source lessons into execution recipes: analyze first, implement with ownership boundaries, debug by layer, test at the right level, and verify on the real rendered/runtime surface.
277
275
  - For design and client skills, place the four judgment layers explicitly (aesthetics / interaction logic / behavioral logic / psychology — per-layer semantics and decision fields: `references/uiux-judgment-extraction.md`); do not hide them inside generic "UI polish" wording.
278
276
  - Write each judgment layer's delta per the Core Rules judgment-delta rule (new / confirmed / narrowed / routed / no new evidence — a restatement of existing principles is not newly extracted knowledge). For visual direction/tokens, use the token provenance fields in `references/uiux-judgment-extraction.md` and state what they mean for design, implementation, and testing; otherwise it is only a static source note.
@@ -290,8 +288,7 @@ Use this skill to turn observed experience into durable agent skills without cop
290
288
  - **Cross-section facet-ownership check** (Core Rules ↔ Step 0–6): for any edit touching `## Core Rules` or a Workflow step, record whether the changed facet is rule/invariant-owned (→ Core Rules) or procedure/checklist-owned (→ the step), and confirm the opposite surface only *points* to it, not restates it. Same-facet text living in both surfaces is a drift defect — converge toward the canonical surface per the Start-here boundary contract before landing.
291
289
  - All referenced files exist and are one level from the skill entrypoint.
292
290
  - No business or source-repo leakage remains in executable guidance.
293
- - Capability naming is source-neutral: old source artifact names, old page names, and old scenario labels are absent from executable guidance or clearly marked as provenance.
294
- - For any rename or generalization, residual searches for the old file name, old source label, old English shorthand, and old capability phrase are clean or justified as provenance.
291
+ - Capability naming is source-neutral, and after any rename or generalization residual searches for the old source artifact name, file name, source label, page name, scenario label, English shorthand, and capability phrase are clean, absent from executable guidance, or clearly marked as provenance.
295
292
  - Trigger boundaries do not collide with sibling skills.
296
293
  - Coverage matrix has no unexplained gaps: each relevant source category is used, routed, discarded, or marked unavailable with a reason.
297
294
  - For broad or multi-skill extraction, validation must include the durable source register and a target-output map showing which target skills or outputs were updated, which sources informed them, which sources were excluded, and which source classes remain pending. If any required row is pending, the work can be landed only as an interim checkpoint, not as complete.
@@ -0,0 +1,37 @@
1
+ # Attention-budget ratchet — design invariants for size/budget gates
2
+
3
+ An attention-budget gate limits how much prose an agent must hold to use a surface: the entrypoint body-word/byte gate, the every-session-injection byte gate, the `description` 800-char cap, and the reference-file line gate (all enforced by `scripts/check-size-budget.sh` or the canonical validator). This file owns the design invariants those gates share and the write-side norms for reference files. Companions: `rule-consolidation.md` owns the prose doctrine (merge-into-canonical, why rule sets must not grow monotonically); `skill-listing-budget.md` owns the host's listing-budget mechanism; `description-authoring.md` owns the description surface.
4
+
5
+ ## The five ratchet invariants
6
+
7
+ Any NEW budget/size gate, and any modification to an existing one, is checked against all five before landing (this is the budget-gate instantiation of the design-time operability check in `dual-track-review-gate.md` — run that check's four legs too). A gate missing one of these fails in a predictable way, named per item:
8
+
9
+ 1. **Stable proxy estimator** — the metric is deterministic and environment-independent: Unicode letter/number word runs with Han counted per ideograph, raw byte size, or physical line count. Never a model/tokenizer-dependent estimate: two environments disagreeing on the measure turns the gate into noise, and a changed estimator silently invalidates every recorded allowance. Deterministic also means encoding-normalized before measuring — a line count taken over raw bytes reads a CR-delimited file as one line, so line endings are folded to LF first.
10
+ 2. **Anti-false-green sentinel** — a run that could not evaluate says so: base-unresolvable prints an `*_unevaluated` token (never the ok token), probe failures fail closed as partials, and on any block the last token is the failure marker. The ok token must be unearnable by losing the base; a consumer grepping for ok must never read an un-run gate as a pass.
11
+ 3. **Zero tolerance for new debt** — a NEW surface over budget blocks outright. There is no exempt marker, no waiver flag, and no way for a candidate to nominate its own baseline; structural exclusions live in the gate, owned by the gate.
12
+ 4. **Legacy may only shrink** — an existing over-budget surface is frozen at its base measure: level or shrinking lands, any growth blocks. Rename credit is path-paired and non-growing (move plus growth blocks as growth). This is what makes a uniform cap deployable over a corpus that already exceeds it, without a rewrite round and without rewarding a rush to pre-shrink.
13
+ 5. **A missing baseline is never a pass** — the comparison base comes from revision history (`CCL_SKILL_BASE_REF`, upstream, or merge-base), so there is no stored manifest to go stale; when no base resolves, the verdict is unevaluated (invariant 2), and reddening that state is the caller's pipeline decision (CI always exports the base ref).
14
+
15
+ Two cross-cutting corollaries:
16
+
17
+ - Average headroom must never fund a single over-budget surface: the ratchet judges each file alone, and corpus-level counters stay visibility-only.
18
+ - Debt counters and advisory bands are never a clean-landing waiver nor authorization to keep growing a surface; only the delta verdict blocks.
19
+
20
+ ## Reference-file write-side norms
21
+
22
+ The read side already defends against oversized files (chunked reads under ~200 lines, references one level deep). These norms are the write side, enforced as a delta ratchet over `skills/*/references/**/*.md`:
23
+
24
+ - A NEW reference file over 500 physical lines must not land — split it by subtopic before landing (the gate blocks new-or-crossing files; 500 exactly passes).
25
+ - An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
26
+ - Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
27
+ - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
+ - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
29
+
30
+ ## Official-clause verdicts (provenance)
31
+
32
+ Registered claims were re-verified against the primary source (Anthropic "Skill authoring best practices", docs.claude.com, read 2026-08-31) before landing; per-clause disposition:
33
+
34
+ - "Keep SKILL.md body under 500 lines" — present verbatim, but it scopes to SKILL.md, not references. Already covered more strictly here by the 5000-body-word delta ratchet. The 500-LINE reference cap above is a repo-internal norm motivated by the read-side chunking evidence, and is labeled as such — never cite it as an official requirement.
35
+ - Table-of-contents mandate for long references — NOT PRESENT in the current official text. The official remedy for `head -100` partial reads is keeping references one level deep (already a Step 6 validation rule). The `##`-section advisory above rests on repo-internal evidence only.
36
+ - "Tested with Haiku, Sonnet, and Opus" — present, conditional on the models you plan to ship to. This repo's skills inherit the session model and the eval layer exercises real sessions, so no multi-model matrix is mechanized; the clause fires only if the repo starts shipping model-pinned skills.
37
+ - Checklist anti-patterns (time-sensitive info, terminology, concrete examples, forward-slash paths, voodoo constants, scripts-solve-not-defer) — present; landed above as authoring norms. The recurring-anti-patterns grep panel is NOT their landing surface: its admission rule requires a class observed in 2+ skills of this repo.
@@ -68,6 +68,15 @@ Bad: `Proactively invoke when the user shares a draft doc / spec / plan and asks
68
68
 
69
69
  Good: `Skip when the ask is already scoped to one stack (e.g. "fix this React render bug" → web-react-dev; "add a GORM index" → go-microservice-dev; "调下这个按钮间距" → product-ui-ux-design), or when the user is reporting a defect / failure / regression → defect-diagnosis owns reproduction and root cause first.`
70
70
 
71
+ ## Body routing pointers: the quadruple
72
+
73
+ The description is not the only routing surface an author writes: skill/reference BODY text routes too, through cross-skill and cross-reference pointers ("route to X", "read Y before Z"). Tier-1 static analysis parses only the description, so body pointers are held to an authoring contract instead:
74
+
75
+ - **Every cross-skill or cross-reference routing pointer in body text must carry the routing quadruple**: trigger (when to go read the other surface), scope (which file or small subset), output (what decision/artifact to extract), and return point (where to resume in the owning workflow). A pointer whose parts are obvious from sentence position may state them compactly ("at step 3, read X's §Y for the Z decision, then continue step 4").
76
+ - A bare "refer to X if useful" / "参见 X" with no trigger and no extraction target is never a landing shape: such pointers rarely fire, and when they do fire they read too much (source-observed: unbounded pointers were the dominant dead-routing shape in an adopting skill pack; the pack that enforced the quadruple had two verified adopters and no dead pointers).
77
+ - The description-side Skip-when `→ skill` idiom already satisfies the quadruple (trigger = the skip condition, scope = the target skill's entry, output = ownership transfer, return = none) — no extra wording needed there.
78
+ - Detector pairing: a pointer that routes but is never read shows up as the **silent skip** failure mode in `eval-routing.md`'s failure-mode vocabulary (B-side probes / Tier-3 traces), not in Tier-1/2 — fix the pointer's trigger and scope, not the description.
79
+
71
80
  ## Precedence against session-injected process skills
72
81
 
73
82
  A routing / workflow skill competes not only with sibling skills but also with process-discipline skills that some hosts INJECT at session start with very forceful language (e.g. a brainstorming skill whose description says "MUST use before creating features / adding functionality"). At initial routing time the router primarily sees skill names, descriptions, and host/session rules — the workflow body has not loaded yet, so an "Entry precedence" paragraph in the body does NOT win the routing decision. If your routing skill should own the entry point for a request class that an injected process skill also claims, the description must:
@@ -67,7 +67,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
67
67
 
68
68
  Across a long operational-rule/code extraction series the challenge supplies *the same handful of axes* as the recurring P0/P1 — first drafts systematically nail the functional / cost / happy-path and omit a predictable set. The per-axis instance lists for the six first-draft blind-spot axes:
69
69
 
70
- - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox.
70
+ - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox; and — because skill/reference text is itself a prompt agents execute — the rollout-safety screen an eval harness runs on skill text before release: wording that induces context/secret exfiltration (prompt leakage), assumes or grants permissions beyond the task (overreach), or automates a destructive or confirmation-skipping step (unsafe automation) — text a draft must not carry except as an explicitly labelled anti-example.
71
71
  - **(2) concurrency & lifecycle** — races, deadlock (e.g. holding a lock through a drain/callback), use-after-close/free, resurrection after delete, double-free/double-close, cleanup ordering, at-most-once/fires-once.
72
72
  - **(3) resource bounds** — an unbounded default/timeout/buffer/retry, a leaked registry/goroutine/task entry, a missing max backstop.
73
73
  - **(4) rollout / migration ordering** — a step that breaks not-yet-upgraded consumers, or abandons a live bug to do the clean refactor first.
@@ -82,7 +82,7 @@ Three per-edit instances of the axes above that recur because their canonical ru
82
82
 
83
83
  For any new mechanical gate, validator, or evidence apparatus — **and for any change that makes an existing one's verdict stricter** — run the four legs at design time, not after challenge rounds force them. These four legs all fire at design time; the check has a second firing point they do not cover, because its actor is not the gate's author: when a landing **withdraws or downgrades the evidentiary claim an existing gate rests on**, that gate is re-based or retired in the same landing — see the claim-liveness rule in `product-rd-workflow/references/design-review-gate-mechanics.md`, which owns it, including the obligation walk a retirement owes.
84
84
 
85
- - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under the SAME base resolution CI uses, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep).
85
+ - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under **every base resolution CI uses — enumerate the landing faces, never assume one**, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep). A base-relative verdict is a quantity *relative to its base*, and CI resolves that base per event, so this leg is discharged by a **recorded manifest, never by a walk you assert**: derive the branch set from the workflow itself — its `push` branch filter plus every branch a pull request is actually landed into — and record one entry per branch carrying the resolved base ref and this gate's verdict on the candidate, so a reviewer can diff the manifest against the workflow and `every` stops being author-adjudicated. Green on the round's own base says nothing about the others: the difference is every commit that landed on the shared branch before the gate existed, and it surfaces as one accumulated violation on the face nobody measured — typically the promotion PR, long after the authors who could have funded it moved on. **A set that cannot be enumerated is a blocking residual, not a pass** — an unrestricted `pull_request` trigger with no fixed target set has no finite manifest, so take leg (d)'s non-blocking or risk-owner-deferral exit and name the faces left unmeasured rather than claiming coverage. Enumerating the faces costs a loop; the exemption you will reach for instead is the loosening leg (e) forbids.
86
86
  - **(b) marginal-cost statement** — record what the cheapest routine change costs under the gate (recompute/regenerate/rerun burden); a gate whose per-iteration cost defeats normal development gets lightened or redesigned at design time.
87
87
  - **(c) trust-model fit** — name what the mechanism defends against under its DECLARED trust model; machinery that only defends against adversaries the trust model already excludes (e.g. content digests where the author can regenerate every hash) buys redundant detection at full complexity cost — prefer the lighter mechanism that keeps the enforceable core.
88
88
  - **(e) loosening check — an exemption must name the class of change that stops owing evidence, and that class must be one with no behaviour to evidence.** Fires whenever a change makes a gate accept what it used to reject: a new exemption class, a waived requirement, a widened accept set. A loosening is easier to get wrong than a tightening and shows up later, because it produces no red for anyone to notice — the gate simply stops asking. Two obligations, both outcomes rather than procedures. **First, state the exempted class in behavioural terms and check it against the repo's own definition of that class**: an exemption named `not-required` asserts *no behaviour*, so if any rule in the tree already classifies that same diff shape as behaviour-changing, the exemption contradicts it and the answer is to fix the *anchor/evidence form* for that shape, never to drop the evidence. **Second, if the class does carry behaviour, the exemption must be replaced by a way to SUPPLY the evidence** — widen where the anchor may land, add an evidence form the shape can satisfy — because the class with the most behaviour is exactly the one an exemption hurts most. **A precedent of the same shape is not a justification**: reaching for an existing exemption class because the root cause rhymes with an earlier one transfers the solution without checking the disanalogy, and the disanalogy is usually the load-bearing part. Failure shape: a gate anchor that structurally cannot bind to a frontmatter-only change was answered with a third `not-required` class by analogy to two existing ones, even though the same repository elsewhere states that any frontmatter edit is a routing-surface change and the author had just measured its routing delta; the adversarial review caught it on the first finding, and the correct fix was to let the anchor bind to the changed description instead.
@@ -230,7 +230,7 @@ Review + challenge check the change for defects; neither checks whether it actua
230
230
  | Status | Use when |
231
231
  |---|---|
232
232
  | `RED-baseline` | the change alters behavior or routing (trigger / scope / routing / validation / acceptance). Run the scenario WITHOUT the change first (baseline failure), then WITH it (compliance) |
233
- | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance, and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
233
+ | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance — an author cannot self-assert it — and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
234
234
  | `not-applicable: docs-only` | the change touches NO file under `skills/**` and no skill-loaded guidance — i.e. `README` / `ARCHITECTURE` / `CONTRIBUTING` / `docs/**` only. Forbidden for `SKILL.md`, `references/**`, validators, templates, examples, and the plugin-shipped command/behavior surfaces (`hooks/*`, `scripts/install.sh`, `bin/`, `.mcp.json`, `.lsp.json`, `monitors/**`, `settings.json`): those are behavioral or executable source even when they read like prose |
235
235
 
236
236
  For a `RED-baseline`, the evidence form can be a before-after task diff, a golden trace, or a pressure scenario — these are *how* you show baseline→compliance, not standalone substitutes for it. For a routing-surface / hub-skill change you can run the golden-trace form as a REAL headless-agent run via the F4 Tier-3 harness (optional, higher-fidelity than a recorded scenario — see `validation-and-landing.md` Behavioral Validation); a recorded scenario is not by itself a failing baseline.
@@ -241,7 +241,6 @@ For a `RED-baseline`, the evidence form can be a before-after task diff, a golde
241
241
 
242
242
  Rules:
243
243
 
244
- - `RED-baseline` is required whenever the change alters behavior or routing. `semantic-control` is valid ONLY with reviewer confirmation that none of trigger/scope/routing/validation/acceptance changed — an author cannot self-assert it.
245
244
  - The row must give a concrete locator + evidence shape, not a bare status: artifact path / commit / transcript / command, the exact prompt or scenario, and expected-vs-actual. For `RED-baseline`, record BOTH the without-change (baseline failure) and with-change (compliance) results, and name the baseline's **provenance type**: a *recorded incident* (cite where the failure is actually recorded — transcript, note, issue; the cited record must describe a failure that occurred, not prescribe a method) or a *constructed scenario* (run against BOTH the unchanged baseline and the changed rule, with an openable artifact for each run — a scenario "run" only mentally, only against the patched text, or only as reviewer discussion does not count). A prescriptive source — a method-bar or best-practice note with no failure recorded — cannot be cited as an occurred failure and is not by itself a valid `RED-baseline`; it may seed the constructed scenario's design or support the rule's rationale, but the RED evidence is the run artifact. Narrating what "would have" failed as if it happened is a fabricated evidence row, the same defect class as fabricated verification output. A status word with no openable artifact is not a valid row, same as a missing review row.
246
245
  - A missing or unreviewable behavioral-evidence row blocks landing the same way a missing review row does; until it exists the work is an uncommitted interim checkpoint.
247
246
 
@@ -300,18 +299,13 @@ Primary reviewer failure is a remediation branch only when the owning gate class
300
299
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
301
300
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
302
301
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
303
- 5. A fallback result satisfies only the exact lane it ran. If no candidate returns a conclusive verdict, keep the work `interim`.
304
-
305
- Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
302
+ 5. A fallback result satisfies only the exact lane it ran, and only when it meets the same evidence bar — do not call a manual ad-hoc run "fallback review". Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence. If no candidate returns a conclusive verdict, keep the work `interim`.
306
303
 
307
304
  - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
308
305
 
309
306
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
310
307
 
311
- The following raw CLI shapes are **debugging diagnostics only**. They may help
312
- isolate an owner-wrapper failure, but neither output is review/challenge evidence
313
- and neither may replace `scripts/extraction_review_gate.sh` for a non-wording
314
- lane.
308
+ The raw CLI shapes below are **debugging diagnostics only** — they can help isolate an owner-wrapper failure, but are never review/challenge evidence and never replace `scripts/extraction_review_gate.sh` on a non-wording lane.
315
309
 
316
310
  Debug a wrapper with raw `codex review` against the target diff:
317
311
 
@@ -423,7 +417,7 @@ The external reviewer judges what the bounded packet can show; the packet struct
423
417
  - The reviewer owns CONTENT SEMANTICS: wording coherence, dropped qualifiers/obligations, source-accuracy, sanitization, scope drift — everything decidable from the packet plus the reviewer's own reasoning.
424
418
  - Deterministic-gate claims (pinned literals present, size ratchet net-zero, obligation audit green, R0 clean, parity) are verified by the repository's CI re-running those gates on the actual branch — never by the reviewer, and never accepted from the implementer's prose alone; a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round.
425
419
  - Historical-process claims (a pre-fix RED, a measurement taken before landing) are session-record-grade unless bound to an immutable revision or a candidate-bound receipt; treat them as the implementer's testimony, and say so in the disposition instead of demanding evidence the packet cannot hold.
426
- - Standing backlog: teaching the gate to embed candidate-SHA-bound receipts of deterministic-gate output into the packet removes the third bullet's limitation mechanically; until that lands, this boundary is the accepted state.
420
+ - The receipt channel for the third bullet is landed: `scripts/gate_receipt.py mint --out <ledger-dir>/gate-receipt-N.json -- <gate command>` binds one deterministic-gate run to the committed candidate — clean tree required, HEAD commit recorded, argv + repository-relative cwd + exit code + output SHA-256 + RFC3339 mint time, with an off-by-default opt-in output tail (a verbatim excerpt would copy a leaked token into the ledger; receipts are created 0600) and a post-run candidate re-read that refuses a result when HEAD moved mid-run; a RED run mints fine (the pre-fix RED is the canonical use). Receipts live OUTSIDE the candidate tree (the chain ledger directory) and ride the packet by reference: name the receipt file plus its own SHA-256 in a review-plan `evidence` entry or a disposition-evidence `evidence` item, so `review_context_sha256` / the v3 ledger's sibling-hash discipline bind it. Anyone re-checks with `gate_receipt.py verify <file>` (structural) or `verify <file> --rerun -- <the gate command you expect>` at the recorded commit — the verifier types the command and the tool compares it against the recorded argv before executing only the verifier's own words, because a receipt is untrusted input and its recorded argv must never be executed as-is (differential exit-code + output-hash comparison; `--exit-only` for a legitimately nondeterministic gate, and say so in the referencing row). Trust model, stated so it is not oversold: a receipt is candidate-bound, falsifiable consistency evidence — it does not authenticate who ran the command; deterministic authority stays with CI re-running the gates on the actual branch (second bullet), and a process claim carrying no receipt remains testimony under the third bullet.
427
421
 
428
422
  ## Recording findings + fixes
429
423
 
@@ -438,7 +432,7 @@ For each pass, record in the extraction's working file (e.g. `<project>-extracti
438
432
 
439
433
  ## Challenge pass (codex exec adversarial)
440
434
  - Findings: N total (a P0 / b P1 / c P2)
441
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
435
+ - R0 evidence: <same value menu as the review-pass row above>
442
436
  - Gate-fireability applicability: <yes — change adds/edits a semantic rule/gate/status/verdict | no — valid ONLY when the diff is wording-only or adds/edits no semantic rule/gate/status/verdict>
443
437
  - Item 9 exercised: <locator to the captured prompt/transcript/JSONL showing the bypass-by-omission probe actually ran (not a pasted self-assertion) | n/a per line above>
444
438
  - Applied: M fixes (commit: <sha>)
@@ -524,33 +518,38 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
524
518
 
525
519
  - When the recurring surface is a **self-adjudication clause** — decidable test: the clause's classification verb has NO named test whose output produces the classification, so the receiver/author judges it — the `keep / delete / narrow / replace` decision must first ask whether an existing mechanical or semi-mechanical test can carry the adjudication: name the test whose output settles the classification, **and the mapping from its output to the classes** — which output means which class — plus the actual result or the evidence contract that will produce it. Naming a test is not routing to it: a test whose output cannot discriminate the classes, or one named with no recorded output-to-class mapping, leaves the adjudication exactly where it was and does not satisfy this rule. A `keep` that retains self-adjudication prose, or a `narrow`/`replace` that adds more prose bindings, is landed only when the decision record (the same-class rule's recorded decision, in the commit body or register row) names the reason no existing test could carry it — a reason left in chat does not count; two challenge rounds attacking the same self-adjudication surface are the signal that prose is the wrong layer — a clause routed to an existing test inherits that test's evidence bar instead of the adjudicator's say-so (worked instance: a review-reception clause that left "is this finding scope-adding" to the receiver survived two rounds of attacks on that self-classification until the classification was routed to the existing structural-minimality test). The same question fires at drafting time for any new reception/discipline-style clause that would grant a self-adjudication.
526
520
 
527
- "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*, not *no new P0/P1 appeared*.
521
+ "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*. Do NOT iterate to zero *findings* either — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing, and forcing the count to zero either over-corrects or scope-creeps; the bar is zero *undispositioned* P0/P1, which differs from both *zero findings* and *no new P0/P1 appeared*.
528
522
 
529
523
  **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
530
524
 
531
- Do NOT iterate to zero *findings* — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing. Forcing the finding count to zero either over-corrects or scope-creeps. The bar is zero *undispositioned P0/P1*, which is different from zero findings.
532
-
533
525
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
534
526
 
535
527
  Five is the generic `code-review` transport ceiling, not this extraction lane's
536
528
  spend. Non-wording Agent-autonomous extraction calls go through
537
- `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=2`: the
538
- initial review plus at most two challenges. At round 3 the autonomous lane ends.
529
+ `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1`: the
530
+ initial review plus at most one challenge. At round 2 the autonomous lane ends.
539
531
  An authenticated human may request later review, but that is separately
540
- attributed human-requested evidence outside this chain/budget, not an Agent
541
- round 4 or 5. Unused generic capacity never authorizes automatic continuation.
532
+ attributed human-requested evidence outside this chain/budget, never an
533
+ additional Agent round. Unused generic capacity never authorizes automatic continuation.
542
534
  The v3 closeout validator rejects referenced receipts whose recorded budget is
543
- not 2 and checks budget and ordering consistency within the caller-supplied
535
+ not the wrapper-fixed value and checks budget and ordering consistency within the caller-supplied
544
536
  set. It cannot authenticate that the wrapper produced those receipts or that
545
537
  the caller retained every earlier chain or receipt. The wrapper does not mint or
546
538
  persist
547
539
  `review_chain_id` or `autonomous_review_index`: the caller still supplies both,
548
- and could start a fresh-looking chain after round 3. The validator detects bad
540
+ and could start a fresh-looking chain after the final round. The validator detects bad
549
541
  order inside the referenced set but cannot detect a prior chain the caller
550
542
  omitted, so complete caller-owned ledger retention—and treating an Agent reset
551
543
  as a contract violation—remains part of the boundary rather than a property the
552
544
  local scripts prove.
553
545
 
546
+ **Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
547
+
548
+ 1. **Batch dispositions; never re-chain per finding — and hold fixes until the ledger can afford their application.** Triage every finding through the disposition bar and deep-self-review once, then decide when the batch lands by budget arithmetic, never by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, so apply-now is legal only when the remaining ledger can still fund a restarted chain's ready floor. **Under the wrapper-fixed 1+1 budget that condition is NEVER true after the review round — the one remaining round cannot fund review plus challenge — so there the rule is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, and land the whole batch only after the full review+challenge chain has run, MR/PR-listed.** A fix applied between review and challenge breaks the chain (the challenge binds to the round-1 candidate), forfeits the double-receipt terminal, and costs a fresh human-authorized chain to recover — an observed failure, not a hypothetical: a round that landed its review fixes before the challenge had to be closed by a user-granted continuation chain.
549
+ 2. **Sum spent rounds across all chains before opening one more.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against the lane's wrapper-fixed budget; at the cap a restarted chain must not be opened autonomously, and below it budget the restart so the final chain can still hold review plus one challenge (the closeout ready floor) — a restarted chain opens with a fresh review by contract, so every restart trades a challenge round for a review round. Effective exhaustion is reached when the remaining rounds cannot fund that floor for any continuation; treat it exactly like the cap.
550
+ 3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
551
+ 4. **At the cap — or at effective exhaustion — the designed terminal is disposition, never another chain.** Apply or disposition the final batch, name every post-review fix in the MR/PR description, record the honest terminal state (`continuation_authorization_required` when the lane's final round ran and itself returned findings; otherwise — including a chain broken before its challenge could run — an interim record naming the last externally reviewed candidate and every later delta), and hand continuation or merge to the human. The post-review batch sits only on the pending MR/PR branch beside that record — the human's authenticated continuation, waiver, or merge decision is what certifies it, and it is never reported as reviewed. This is the bounded outcome working as designed, so do not report it as convergence and do not launder it through a fresh-looking chain.
552
+
554
553
  A strictly proven wording-only change has no convergence loop: it uses one
555
554
  generic `code-review` pass, records the independent-review row and the
556
555
  challenge-not-required proof, and does not create a schema-v3 multi-round
@@ -563,7 +562,7 @@ This budget limits only automatic reviewer invocation. It does **not** stop impl
563
562
  - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
564
563
  - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
565
564
  - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
566
- - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. One dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. The recovery is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
565
+ - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. The dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. While cross-chain budget remains, that break is handled autonomously by the self-hosted-chain rule's ledger-counted restart; it becomes this bullet's human-decision dead-end when the remaining budget cannot fund the re-review. The recovery at that point is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
567
566
 
568
567
  When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
569
568
 
@@ -571,7 +570,7 @@ The mechanical reminder is `self_review_gate`, not prose alone. It records outst
571
570
 
572
571
  In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
573
572
 
574
- At the third Agent-autonomous round, do not start a fourth automatically. If findings remain:
573
+ At the final Agent-autonomous round, do not start another automatically. If findings remain:
575
574
 
576
575
  - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`;
577
576
  - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
@@ -586,14 +585,14 @@ ledger and every referenced controller, completion, base, and sweep file live
586
585
  in one directory and carry SHA-256s. The validator walks the ordered schema-v3
587
586
  controller chain (same chain and scope,
588
587
  review then contiguous challenges, packet=candidate, complete prior-result hash
589
- prefix, fixed `challenge_budget=2`) and binds every closeout candidate to its
588
+ prefix, the wrapper-fixed `challenge_budget`) and binds every closeout candidate to its
590
589
  last receipt. Ready requires at least review + challenge; a second base drift may
591
590
  stop as race immediately after round 1 rather than spending an illegal challenge
592
591
  after the terminal predicate already fired.
593
592
  It ends in exactly one state:
594
593
 
595
594
  - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
596
- - `continuation_authorization_required`: round 3 itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
595
+ - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
597
596
  - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
598
597
 
599
598
  Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
@@ -608,17 +607,16 @@ convergence.
608
607
 
609
608
  For a focused single-skill change:
610
609
  - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
611
- - **Round 2 — challenge 1**: after implementer triage, attack the highest-risk unresolved surface with an unprimed prompt.
612
- - **Round 3 — challenge 2**: verify remaining/new attack paths. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic round 4.
610
+ - **Round 2 — challenge**: after implementer triage — fixes stay HELD: under the 1+1 budget applying any fix before this round always breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic further round, and the post-review fix batch lands MR/PR-listed.
613
611
 
614
- Broad extractions use the same three-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
612
+ Broad extractions use the same fixed two-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
615
613
 
616
614
  ### Anti-patterns
617
615
 
618
616
  - **Single-round challenge → done**. The round-1 fix-up itself may introduce bugs. Always do at least one re-challenge after a non-trivial fix-up.
619
617
  - **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
620
618
  - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
621
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one.
619
+ - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one. The observed extreme is chain multiplication: a chain broken by your own fix and restarted at index 1 is the same budget, not a new loop — sum rounds across chains per the self-hosted-chain rule above.
622
620
  - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
623
621
 
624
622
  ### Recording the loop
@@ -629,7 +627,7 @@ Add one row per round to the validation log:
629
627
  ## Challenge pass — round N (codex exec adversarial)
630
628
  - Diff scope: <files / commit range / sha>
631
629
  - Findings: N total (a P0 / b P1 / c P2)
632
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
630
+ - R0 evidence: <same value menu as the review-pass row above>
633
631
  - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
634
632
  - New since prior round: <count> (subset of above; flag round-introduced bugs)
635
633
  - Stabilized: <list of findings carried over without change>
@@ -41,10 +41,11 @@
41
41
 
42
42
  ## Tier-2:路由 task-bank + 廉价 grader(已落地,advisory)
43
43
 
44
- `scripts/eval-routing-bank.rb <repo-root> [--bank p] [--model m] [--limit N] [--dry-run] [--json p] [--baseline p] [--desc-budget-chars N]`。`--desc-budget-chars N` 把每条 description 截到前 N 字符再评(消费端截断臂——如 Codex 在 ~2% 上下文预算/未知窗口 8,000 字符下压缩技能清单);plain 与 budgeted 各跑一遍即可定位"只活在描述尾部"的路由触发词。`make eval-routing-bank`。
44
+ `scripts/eval-routing-bank.rb <repo-root> [--bank p] [--model m] [--limit N] [--dry-run] [--json p] [--baseline p] [--desc-budget-chars N] [--replicas N]`。`--desc-budget-chars N` 把每条 description 截到前 N 字符再评(消费端截断臂——如 Codex 在 ~2% 上下文预算/未知窗口 8,000 字符下压缩技能清单);plain 与 budgeted 各跑一遍即可定位"只活在描述尾部"的路由触发词。`--replicas N` 每 task 评 N 次:task 判定取保守共识(任一有效副本 FAIL 即 FAIL),并报告副本 top1 一致率——不同 (bank, replicas) 配置是不同尺子,不得互相 diff 当回归。`make eval-routing-bank`。
45
45
 
46
- - 冻结 task-bank `eval/routing-tasks.jsonl`:每行 `{id, utterance, expected_skill, must_not_route_to?, source, why_expected, frozen_at_sha}`,种子取自 source-register 历史 miss + bootstrap 路由规则。
47
- - grader = 每 task 一次本机 `claude --print --tools "" --model <haiku>`,喂 utterance + 全部 skill description(agent 真正路由的那份面),要 `{selected_skill, confidence, rationale_short}`。
46
+ - 冻结 task-bank `eval/routing-tasks.jsonl`:每行 `{id, utterance, expected_skill, acceptable?, must_not_route_to?, source, why_expected, frozen_at_sha}`,种子取自 source-register 历史 miss + bootstrap 路由规则。**已修复的路由 miss 必须把其 utterance 冻结成 bank task 落在同一交付里**(修复不冻结=下次同类漂移无回归面)。`expected_skill: "none"` 是否定对照/覆盖空洞哨兵:正确结果是没有技能认领;`acceptable` 列出可辩护替代结果(如空洞探针上 coordinator 接管与拒绝都对);`must_not_route_to` 点名吸入诱饵邻居。结构由 `test_routing_bank_integrity.sh` 确定性把关(sentinel 只准出现在 expected/acceptable,不准进 must_not)。
47
+ - **冻结案例神圣(regressions-are-sacred)**:已冻结案例的删除或判定面改写(bank task 的 expected/acceptable/must_not,golden trace 的 assert 块)是一次回归裁决事项,与邻居回归同权——平均改善不得抵消单条冻结案例的失守,且「曾 yes 现非 yes」**含降级为 unsure/INCONCLUSIVE** 都算回归。删除/改判的每条案例必须在同一轮的 register 追加行里写 `case-retired: <id>` 或 `case-rescoped: <id>` 并给理由,交独立评审裁决;确定性半边由 `test_frozen_case_sanctity.sh` 按 `CCL_SKILL_BASE_REF` 把关——无裁决行即红,无 base ref 时打印显式 skip token(skipped ≠ passed),它只保证交易可见,不裁决交易正当性。
48
+ - grader = 每 task(×replicas)一次本机 `claude --print --tools "" --model <haiku>`,喂 utterance + 全部 skill description(agent 真正路由的那份面),要 `{selected_skill|none, clarify, confidence, rationale_short}`。**clarify 率、低置信率(<0.5)、副本一致率是一等报告字段**,不是旁注——路由质量的残余风险常在"高置信直选却选错、无自纠路径"这类 pass/fail 看不见的分布里。
48
49
  - **路由兼容性信号,不是真值预言机** —— grader 自己可能错;只衡量"当前 description 能否让廉价模型把固定 utterance 路由到 expected"。
49
50
  - **advisory**:不接 `check-ccl-skills.sh`,不挡 merge。退出码:`0` = 跑完;`2` = 用法;`3` = grader 整体不可用(claude CLI 缺失则打印 skipped 后 `0`)。路由 miss 永不非 0。
50
51
  - **防作弊**:runner 校验每 task 的 `frozen_at_sha` 是 HEAD 祖先(非祖先 = drift,排除出回归判定);同一改动若同时动 task-bank 和 SKILL.md description 会显式告警(防"改 skill 顺手改测试让它过")。
@@ -58,6 +59,18 @@
58
59
  2. 改后通过数必须在**最终措辞**上重测:中间稿的通过数在措辞再变的那一刻作废,不得挪用到最终候选的证据里。
59
60
  3. 受影响邻居用例集默认改前/改后各 **≥3 轮**,集合须含期望 owner 自己的兄弟用例与高词面重叠的他 owner 用例;邻居回归作为独立 finding 交由本轮实际门禁处置——**该 finding 须以 blocking 记入本轮 dual-track 评审记录,且只能由独立评审方豁免,不能由实现者自行判定「本轮没有门禁采用这组证据」而放行**。降级的是「F4 自己充当合并门禁」这一声称,不是「回归必须被人裁决」这一义务;后者若也随之消失,这一条就只剩被裁决方自审。
60
61
  4. 每轮判决必须连同 **runner 调用、grader 模型身份、候选身份**(commit 或描述内容指纹)与**原始逐轮工件的持久定位符**一并记入轮记录;没有定位符的通过数只能标注为 operator-reported,不得据以宣称修复轮已 concluded。
62
+ 5. **单变量归因**:一次改前/改后对照只准动**一个路由变量**(一条 description,或同一 skill 不可分割的一组路由面)。同时动多条 description 的批量改动,其对照差值不可归因到任何一条,只能按整包回归读——要归因就拆成逐条 A/B。(源侧实测形态:仅替换一条 description 的成对子集对照,把命中从约 2/3 提到 95%,且提升可归因到那一条改动——多条同动时这句话说不出口。)
63
+
64
+ ## 路由失效形态词表(跨层;报告与修复讨论用这套名字)
65
+
66
+ pass/fail 之外,路由失败有可命名的形态;每个形态有不同的检测器与不同的修法,混称"路由不准"会修错面:
67
+
68
+ | 形态 | 判据 | 检测器 | 典型修法 |
69
+ |---|---|---|---|
70
+ | **吸入 (absorbed)** | 应被拒绝(expected none)或应远离诱饵邻居(must_not)的 utterance 被某技能认领——覆盖空洞不被承认而被最近邻低置信/clarify 拉走 | bank 否定对照+空洞探针,runner `absorbed` 标签 | 认领空洞(补 owner)或在诱饵邻居 description 加排除句;不要靠 grader 自觉 |
71
+ | **归属分裂 (ownership_split)** | 同一 utterance 多副本给出不同 top1——所有权不稳定,谁都像 owner | `--replicas ≥2` 一致率 + `ownership_split` 标签 | Skip-when 互相消歧或触发词改具体(description-authoring 的 80% 阈值) |
72
+ | **静默跳过 (silent skip)** | 路由"成功"但被选技能的正文从未被读/其硬规则从未被应用——选择层过了,质量层空转。判据阶梯:mounted → invoked → 文件真被读 → 下游行为改变;mounted-only 不证明任何生效 | 非 Tier-1/2 可见;B 面 body-compliance 探针 + Tier-3 真 agent 事件流 | 修正文 firing point/read routing(body 指针补齐四元组,见 description-authoring.md「Body routing pointers」),不是修 description |
73
+ | **高置信错选** | 错选但 confidence 高、无 clarify——事后无自纠路径,比低置信错选更危险 | 报告里 FAIL ∩ 高置信 ∩ clarify=false 的行 | 触发词消歧;必要时在正文加入口自检;残余风险如实记录 |
61
74
 
62
75
  ## Tier-3:hub golden trace 真 agent 回放(已落地,advisory,人工判定)
63
76
 
@@ -71,6 +84,14 @@
71
84
  - **随机性**:agent 非确定;判定先人工、nightly 起步,有稳定史前不自动 gate。防作弊同 T2(`frozen_at_sha` 祖先校验)。
72
85
  - **双用途**:除回归外,Tier-3 还可当**改技能前的 RED-baseline**(改前手动跑触发场景看真 agent 是否真路由错,改后看 compliance)——可选;只有真观察到 miss 才算 RED(PASS/INCONCLUSIVE 不算),小 N + 非确定有噪声,手动跑两次自己留两份报告。落地 + 防作弊注意见 [validation-and-landing.md](validation-and-landing.md) "Optional real-agent RED-baseline"。
73
86
 
87
+ ## B 面:正文合规探针 body-compliance(已落地,advisory)
88
+
89
+ `eval/body-compliance-eval.rb <repo-root> [--arm L] [--json p] [--model m] [--timeout s] [--ids a,b]`。`make eval-body-compliance`。路由三层测「选没选对技能」;B 面测**已激活技能的正文硬规则是否真被应用**——正文即 prompt,逐探针 required/forbidden marker 契约判分,覆盖是 NAMED SUBSET(见 runner 头部声明)。
90
+
91
+ - 含 product-rd-workflow 停机谓词的**成对分类探针**(prd-stop-*/prd-continue-*):每对场景恒定、只变谓词判别特征,按闸自身的字面 `continuing:`/`blocked:` marker 判分——确定性锚钉的是这些谓词的**措辞存在性**,只有这些探针检验**案例被分到哪边**。
92
+ - 触发纪律:改动触及某技能正文硬规则或停机谓词时,must run the affected `--ids` probe subset on this machine before landing,结果按该改动既有门禁处置;advisory 契约与升级路径(提炼确定性不变量,永不阻断 LLM 判定)见 `eval/AGENTS.md` 与 [f4 手册](../../../docs/f4-skill-effectiveness-harness.md)「双轨承载面与运行分层」。
93
+ - 实测的双轨互补边界(锚管措辞、探针管行为漂移、不是埋句绊线)与小样本告诫,单一落点在 f4 手册「双轨承载面与运行分层」,此处不复述。
94
+
74
95
  ## Health roll-up:描述性仪表盘(已落地,advisory)
75
96
 
76
97
  把上面各信号卷成**一个加权 0–10 显示值 + 同尺子变化**,用于定位值得继续检查的维度。它借用 OpenSSF Scorecard 的呈现形态,但不把不同性质的 F4 信号变成“仓库整体变好/变差”的总判决。映射与限制见 [harness-patterns-and-eval.md](harness-patterns-and-eval.md) §3.4。