@ccoalm/ccl-skills 0.9.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +4 -3
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +23 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +178 -4
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +127 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +2 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +1 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +11 -1
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +16 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +1 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +8 -8
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +9 -1
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +2 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +9 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +17 -20
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +9 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +26 -44
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +16 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +3 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +59 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +2 -2
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +3 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +35 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +16 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +35 -6
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +476 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +81 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +28 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +336 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +41 -1
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +141 -28
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +139 -38
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +1 -1
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -8
  56. package/dist/assets/release.json +110 -45
  57. package/package.json +1 -1
@@ -7,7 +7,7 @@ description: 加功能 / 新需求 / 技术方案 / 方案评估 / 技术选型
7
7
 
8
8
  Use this skill as the top-level workflow for new product development, feature delivery, bug handling, refactoring, release preparation, or repeated process improvement. It is not a replacement for stack-specific skills; it decides which skill should own each stage and what evidence is required before moving on.
9
9
 
10
- **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack or execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
10
+ **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack/execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
11
11
 
12
12
  **Continuation-proposal output contract (session-wide for product delivery).** Every assistant message in a delivery session routed by or through this workflow carries exactly one literal line until the user explicitly ends/pauses the delivery or changes scope: `proposed-next: <action and scope>` when the message has imperative/future/next-step wording or a proposed action, otherwise `proposed-next: none — status only`. At the start of every subsequent user turn, read that line before interpreting the reply: absence, multiplicity, or a marker/wording conflict enters `blocked:`/`interim` by default, never `not-applicable`. Coverage detail and the host-layer caveat: `references/pre-final-continuation-gate.md` (Continuation-proposal output contract).
13
13
 
@@ -21,23 +21,23 @@ Use this skill as the top-level workflow for new product development, feature de
21
21
  - A general-purpose process skill that merely *looks* like the obvious start — brainstorm/scope-shaping, plan-writing, or TDD auto-suggested by ANY channel: a session-start prompt, an optional skill package, or the host platform's native skills listing (including a listed entry skill's own self-invocation mandate, e.g. "must invoke if there is a 1% chance") — does not replace this entry: suggestion-channel wording is channel self-promotion, not routing authority (a host-mandated preflight — mandated by a host-authored system/developer-level or equivalent higher-priority instruction — may run first without thereby becoming the delivery owner; the test is AUTHORSHIP, not rendering position: a host-authored instruction counts even when rendered within the listing surface, while a skill's own description/content claiming preflight status never does); invoke this workflow as the delivery entry (immediately after any genuine host-mandated preflight), then call that skill inside the stage it serves (for example, requirement shaping in Workflow step 1).
22
22
  - For delegated agent execution, `multi-agent-delegation` owns the execution recipe and worker verification while this workflow owns the lifecycle gate and acceptance boundary; delegated agents resuming after a pause must receive or re-verify the current plan/spec artifact set before editing.
23
23
  - For multi-repo delivery, or any delivery that changes remote branch, MR, pipeline, release, or deployable-artifact state, maintain a compact per-changed-unit delivery-status ledger (row schema + persistence rules in `references/status-tracker-sync.md`) before claiming done, recommending MR/merge, or choosing the next slice. A required verification/review gate passes only when its state is `success`, `not-applicable`, or `not-required`; any other state blocks a done/merge recommendation, so remediate, wait to a terminal state, or report the delivery as pending with the unblock action. Residual-risk acceptance by the user only permits the recommendation/handoff label for that concrete action; it does **not** authorize merge, auto-merge, default-branch push, or cleanup, which still require the user's explicit merge instruction for the current MR (per `worktree-isolation`). Small local-only multi-file edits use the normal concise status unless they introduce remote, CI, MR, release, or deployable-artifact state.
24
- - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack or execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
24
+ - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack/execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
25
25
 
26
26
  ## Scope
27
27
 
28
- - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed or stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
28
+ - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed/stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
29
29
  - **Implementation completeness/minimality gate.** For behavior-changing delivery this gate fires — functional completeness and structural minimality are independent gates — and gaps block `complete` — **load `references/implementation-completeness-and-minimality.md` before the mapping**: map every in-scope acceptance point to implementation plus fresh evidence, and speculative future need and omitted required behavior both fail — the independent-gates rule and the acceptance/concept matrices live there, and every retained new concept must map to a current acceptance point or hard constraint.
30
30
  - Product requirement artifacts: route clarification to `requirement-intent`, conditionally-required current-state evidence to `requirement-baseline`, scope/version/appetite to `requirement-scope`, and Ready-only human-readable PRD assembly to `requirement-doc-writer`. All four use the canonical `requirement-doc-writer/references/requirement-closure-contract.md`. In active delivery, each narrow artifact returns here after completion; only a standalone non-PRD narrow artifact may return directly to the user. “只写个 PRD / 不走流程”仍需 lifecycle-issued Ready;WIP/会议材料不得命名为 PRD. This workflow owns the complete lifecycle's `PRD Ready` / `PRD Not Ready` verdict and cross-owner closure coordination.
31
31
  - Design routing: decide when interaction, information architecture, visual system, design-system, or UX acceptance work must happen before implementation.
32
32
  - Architecture routing: decide when architecture work is needed before implementation.
33
- - Development routing: invoke stack-specific skills for implementation details, such as `app-cross-platform-dev` for Flutter/Android/iOS apps, `miniapp-product-dev` for WeChat/Alipay/Douyin/Baidu mini-programs, `web-react-dev` for React web, `terminal-cli-dev` for terminal/TUI product surfaces and CLI interface design (the rendered interface — layout, input, ANSI, scrollback — plus the command/subcommand/flag/help contract, which it owns even when nothing is rendered); CLI/tooling implementation without terminal-UI concerns must go to that language's dev skill when one owns it (`go-microservice-dev`, `python-service-dev`) and must not fall back to `terminal-cli-dev` merely because the deliverable is a command; a CLI in a language with no such dev owner stays with `terminal-cli-dev` as the default CLI owner, `go-microservice-architecture`/`go-microservice-dev` for Go services, and `python-service-architecture`/`python-service-dev` for Python services.
33
+ - Development routing: invoke stack-specific skills, such as `app-cross-platform-dev` for Flutter/Android/iOS apps, `miniapp-product-dev` for WeChat/Alipay/Douyin/Baidu mini-programs, `web-react-dev` for React web, `terminal-cli-dev` for terminal/TUI surfaces and CLI interface design (the rendered surface — layout, input, ANSI, scrollback — plus the command/subcommand/flag/help contract, which it owns even when nothing is rendered); CLI/tooling without terminal-UI concerns must go to that language's dev skill when one owns it (`go-microservice-dev`, `python-service-dev`, `nodejs-service-dev`) and must not fall back to `terminal-cli-dev` merely because it is a command; a CLI in a language with no such dev owner stays with `terminal-cli-dev` as the default CLI owner, `go-microservice-architecture`/`go-microservice-dev` for Go services, `python-service-architecture`/`python-service-dev` for Python services, and `nodejs-service-dev` for Node.js services.
34
34
  - Quality discipline: define test scope, review scope, release checks, and bug-fix verification.
35
35
  - Defect routing: use `defect-diagnosis` for hands-on reproduction, isolation, instrumentation, fix, regression verification, root-cause analysis, and prevention routing.
36
36
  - Research routing: For research groundwork behind a selection or assessment decision (技术选型、方案评估前的主题调研), this workflow must call `multi-perspective-research` for an evidence-grounded brief; the verdict itself stays with this workflow's gates.
37
37
  - Existing-project assessment routing: for requests that ask to analyze a repository, product codebase, or project quality across architecture, implementation, tests, UI/UX, bugs, or risks, run codebase understanding first, then route each assessment dimension to the smallest owning skill instead of treating understanding as the final answer.
38
- - AI/algorithm product launch discipline: for new or iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
- - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness or freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
- - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new or public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
38
+ - AI/algorithm product launch discipline: for new/iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
+ - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness/freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
+ - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new/public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
41
41
  - **Artifact-egress confidentiality gate**: before a delivery artifact (spec/plan/requirement, status/writeback doc, launch/task card, retrospective) — or any text generated from it — **crosses the local trusted boundary**, run a confidentiality pass *before* the write/create/push/send call; it fires **only on cross-boundary egress** and owns only the **semantic confidentiality axis** that secret scanners miss. A block stops the *egress* but preserves the draft locally and never deletes the work; exception authority stays with the user/owner. Egress channels include chat/tracker surfaces (Feishu/Bitable writeback, shared/public docs, MR/issue body or review comment, external trackers), durable VCS/release metadata (commit message, branch/tag name, release note/changelog, CI metadata), external assistants/models, and delegated-worker prompts; the gate's own block reason/finding is itself egress — never quote the raw sensitive span across the boundary; secrets/PII/credentials/raw-logs/customer-data route to their existing owners (`platform-observability` redaction, `defect-diagnosis` evidence sanitization, the `feature-risk-router` security-review gate); per-category actions, severity, and the delegated-worker firing point: `references/artifact-egress-confidentiality.md`.
42
42
  - Learning loop: after bugs, review findings, incidents, repeated friction, or external skill research, use `skill-extraction-workflow` to update the right skill or reference instead of leaving knowledge only in chat.
43
43
  - Skill/process extraction: when asked to summarize delivery experience, preserve a workflow lesson, update a reusable skill, or decide where a lesson belongs, route to `skill-extraction-workflow` first. Update this skill only when the lesson changes product R&D routing, gates, ownership, or lifecycle policy.
@@ -195,7 +195,7 @@ Run this gate before finalizing a product R&D turn after any delivery slice land
195
195
  - **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — **load it before selecting `continuing:` on any assent**, and any concrete next-slice proposal you issue must itself carry the `proposed-next:` marker or a later assent cannot bind — an unmarked referent is ambiguous, never self-cleared; the rule fires only when the immediately preceding assistant message itself states one concrete next action and its scope, and `continuing:` binds to that proposal, never to adjacent status or response-format prose; ambiguous assent, referent, or authority ⇒ `blocked:` with step-4 precedence — restate the proposed action/scope plus the specific ambiguity/authority, cite the step-1 evidence and ask one concise question in the same turn; self-classifying the reply or marker away is never an exit, and the `continuing:` default applies only when assent is unambiguous and no step-4 condition holds (the full rule and its fallback are stated there); the visible `continuing:`/`blocked:` outcome obligation is unchanged.
196
196
  3. Continue automatically only when no step-4 stop condition fires and: the next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned, verifiable with existing commands, and needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision, or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
197
197
  4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a failed/pending/inconclusive required/blocking gate, dirty/conflicting worktree that can't be isolated, required environment unavailable after remediation, high-impact product/architecture/compliance decision, destructive action, external purchase/financial commitment, unclear owner, ambiguous assent, missing stricter authorization, materially differing viable approaches (none dominant-and-reversible), a fix lacking evidenced cause, or no low-risk slice. Exactly one dominant reversible approach and no other stop condition firing: do not stop at a recommendation: deliver a tested reviewable draft.
198
- 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form, turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
198
+ 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
199
199
  6. **Assent-outcome closeout check.** Before every final response in a product-delivery session, walk the literal immediately preceding marker and visible outcome; the agent cannot exclude a status, question, review, or dispatched-owner turn by reclassifying it outside the session. An action marker or a plausibly affirmative user reply requires exactly one already-visible `continuing:`/`blocked:` outcome; until `continuing:` has executed the accepted slice or `blocked:` has named the blocker, the current turn may not use `proposed-next: none — status only` or `not-applicable`. A valid status-only marker permits `not-applicable` only when the preceding prose has no imperative/future/next-step wording and no affirmative reply pending. A missing/multiple/conflicting marker forces a visible `blocked:`/`interim` outcome with one clarifying question; it never produces `not-applicable`. The current assistant message itself must end with exactly one action-form or status-only marker for the next turn. Omitting the marker cannot justify a silent stop.
200
200
 
201
201
  If a user later challenges "why did you stop" or "was the rule too weak", treat it as a product workflow defect: route through `skill-extraction-workflow`, strengthen the smallest owning skill or validation checklist, validate the diff, and only then claim the process issue is solved.
@@ -75,7 +75,7 @@ Run an architecture pass before implementation when any of these change:
75
75
  - async workflow, idempotency, retry, lock, queue, or scheduled job behavior.
76
76
  - release runtime, canary, rollback, observability, or operational ownership.
77
77
 
78
- For Go backend work, use `go-microservice-architecture` for this gate. For Python backend, AI-service host, worker, SDK/package, or batch-job architecture, use `python-service-architecture`.
78
+ For Go backend work, use `go-microservice-architecture` for this gate. For Python backend, AI-service host, worker, SDK/package, or batch-job architecture, use `python-service-architecture`. For Node.js backend, worker, or CLI/tooling architecture there is no `*-architecture` sibling by decision: run this gate here, with the relevant `platform-*` owners for connectivity/observability/release surfaces and `nodejs-service-dev` for the Node-side implementation contract (runtime/module/type path, async lifecycle, outbound client shape). Do not substitute the Go or Python architecture skill — a stack without an `*-architecture` owner routes to a named owner, never to the nearest-looking one (`references/dispatch-owner-skills.md`).
79
79
 
80
80
  Architecture evidence should name the runtime surface that proves the design can operate after launch: generated-contract ownership, dependency timeout/secret/discovery contracts, health/readiness, trace/log context, durable state for long-running work, retry/idempotency policy, and rollback or repair path. Do not accept a plan that only names the happy-path API or screen while leaving async tasks, generated code, or runtime dependencies implicit.
81
81
 
@@ -4,7 +4,7 @@ Use this reference for the owner-dispatch discipline behind the technical design
4
4
 
5
5
  ## Why the complete owner set loads at design time
6
6
 
7
- When the technical design gate is triggered and the deliverable's substance spans more than one owning skill — backend/platform/architecture (`go-microservice-architecture` / `python-service-architecture` + the relevant `platform-*` skills), LLM/inference/RAG/eval (`llm-inference-integration`), the storage/data-contract owner, the client stack owner (`web-react-dev` / `app-cross-platform-dev` / `miniapp-product-dev`), `testing-strategy`, `product-ui-ux-design` — those owning skills own BOTH (a) producing the design substance and (b) the review. Invoking the router does not discharge them.
7
+ When the technical design gate is triggered and the deliverable's substance spans more than one owning skill — backend/platform/architecture (`go-microservice-architecture` / `python-service-architecture` for those stacks; for a stack with no `*-architecture` owner, this workflow's architecture gate plus the relevant `platform-*` skills — see the stack owner map below), LLM/inference/RAG/eval (`llm-inference-integration`), the storage/data-contract owner, the client stack owner (`web-react-dev` / `app-cross-platform-dev` / `miniapp-product-dev`), `testing-strategy`, `product-ui-ux-design` — those owning skills own BOTH (a) producing the design substance and (b) the review. Invoking the router does not discharge them.
8
8
 
9
9
  - Before producing or reviewing, enumerate the concerns the deliverable touches and load the COMPLETE owner set for those concerns during design, not only at review. Do not load owners for untouched concerns, and do not discover required owners one at a time mid-review — if a new touched concern appears, add it to the inventory and load its owner before continuing.
10
10
  - A cold dispatched worker is the worst case (see `multi-agent-delegation`).
@@ -33,3 +33,11 @@ This gate is enforceable in code, not only prose:
33
33
  - ccl-skills itself ships no config (its own edits are shared-skill/process edits, exempt from owner-load).
34
34
 
35
35
  The closeout-acquire rule (`status` must read `opted-in: yes`, or install, or record why exempt) stays in the entrypoint because it is a closeout item, not a map mechanic.
36
+
37
+ ## Stack dev owners and the CLI implementation carve-out
38
+
39
+ The per-stack implementation owners are `go-microservice-dev` (Go), `python-service-dev` (Python), and `nodejs-service-dev` (Node.js). A newly added stack dev owner joins this map in the same landing that adds the owner, and the entrypoint's Development routing line carries the same list — the two must not drift.
40
+
41
+ - **CLI/tooling implementation without terminal-UI concerns goes to that language's dev owner** and must not fall back to `terminal-cli-dev` merely because the deliverable is a command. `terminal-cli-dev` keeps the interface contract — the rendered surface plus the command/subcommand/flag/help/exit contract it owns even when nothing is rendered — and each stack dev owner carries the reciprocal skip leg. A CLI in a language with no dev owner stays with `terminal-cli-dev` as the default CLI owner.
42
+ - **Node.js has no `*-architecture` sibling by decision, not by oversight.** Its architecture, service-boundary, data-ownership, and reliability decisions run through this workflow's architecture gate plus the relevant `platform-*` owners, with `nodejs-service-dev` owning the Node-side implementation mechanics. Do not borrow `go-microservice-architecture` or `python-service-architecture` for a Node.js service: their stack-specific RPC, storage, and codegen contracts do not transfer, and treating a missing sibling as "route to the nearest architecture skill" is the failure this rule prevents.
43
+ - A stack whose dev owner exists but whose architecture decisions have no owner is a **routing vacuum, not a valid `not-applicable`**. Name the owner that absorbs them, as Node.js does here; recording the absence itself as the reason restates the symptom and leaves the stack unreachable from this gate.
@@ -54,6 +54,8 @@ Update the smallest correct durable place:
54
54
  - Go backend implementation, testing, codegen, DB, Redis, MQ, or protobuf issue: update `go-microservice-dev` references.
55
55
  - Python backend, AI-service host, worker, SDK/package, or batch-job architecture issue: update `python-service-architecture`.
56
56
  - Python implementation, pytest, packaging, schema, ORM/migration, Redis, queue, async, or service-wiring issue: update `python-service-dev` references.
57
+ - Node.js implementation, runtime/module/type path, package-manager or lockfile, async lifecycle, stream, worker, outbound-client, or `node:test`/runner issue: update `nodejs-service-dev` references.
58
+ - Node.js architecture or service-boundary issue: there is no Node architecture sibling skill by decision — update this workflow's architecture gate text or the relevant `platform-*` owner, and record which one absorbed it rather than filing it under a sibling stack's architecture skill.
57
59
  - UI/product interaction issue: update the relevant design or frontend skill.
58
60
  - One-off business/domain issue: do not turn it into a generic skill rule.
59
61
 
@@ -30,7 +30,7 @@ Use this for implementation of Python backend products, services, microservices,
30
30
  - Convert domain-specific source patterns into reusable mechanics: route shape, schema validation, repository/unit-of-work boundary, transaction scope, cache key strategy, idempotency, job lease, config object, trace context, fake client, or test style.
31
31
  - If an observed pattern only works for one product domain, discard it instead of turning it into a rule.
32
32
  - Resolve conflicts by choosing the safer generic default: explicit schemas over dicts, typed settings over ad hoc environment reads, reviewed migrations over blind autogeneration, bounded async concurrency over unbounded gather, dependency injection over import-time clients, focused pytest tests over live-infra tests by default, and fail-closed for auth/permission/data-integrity paths.
33
- - When adding or revising durable Python implementation guidance, check whether the lesson is generic backend service practice that should also update `go-microservice-dev`, or belongs in a shared workflow skill instead. If the rule depends on Python tooling, FastAPI/Flask/Django, Pydantic, asyncio, pytest, or Python package layout, keep it here and do not force a Go mirror.
33
+ - When adding or revising durable Python implementation guidance, check whether the lesson is generic backend service practice that should also update the sibling stack owners `go-microservice-dev` and `nodejs-service-dev`, or belongs in a shared workflow skill instead. Record each sibling as `update`, `unchanged`, or `route-to-shared` rather than leaving it unexamined. If the rule depends on Python tooling, FastAPI/Flask/Django, Pydantic, asyncio, pytest, or Python package layout, keep it here and do not force a Go or Node mirror.
34
34
 
35
35
  ## Development Workflow
36
36
 
@@ -18,3 +18,12 @@ After pushing a tag, read back:
18
18
  - Produced image/digest/version evidence when available.
19
19
 
20
20
  If the tag target is wrong after push, do not force-move a published production tag. Stop and escalate to the release owner for the corrective version/tag path.
21
+
22
+ ## The version pointer is under the same immutability, one step earlier
23
+
24
+ A published version cannot be changed or reused, so the source tree's version pointer may never sit **below** the highest already-released version. Treat that as a checked invariant, not a convention:
25
+
26
+ - **Check it at merge time, not only at tag time.** A tag-time check blocks the bad release but leaves the wrong pointer on the integration branch until a person happens to notice, and the next bump then lands on a corrupted base.
27
+ - **The commit that lowers it is usually not a release commit.** The version line is a both-sides-changed hunk, so the observed shape is a conflict resolved the wrong way inside a change about something else entirely — the commit subject gives no warning, and reviewers reading it for its stated purpose skip the hunk.
28
+ - **Repair every site.** The version is stated in the manifest and again in the lockfile (twice, in current npm lockfile versions); a partial repair leaves the sites disagreeing, which is its own release defect.
29
+ - **Compare against the released record, not against a base branch.** "Did this branch lower it" is a different, weaker question than "does the tree point under something already published"; derive the record from release tags or the registry.
@@ -49,10 +49,10 @@ Use this skill to turn observed experience into durable agent skills without cop
49
49
  - Task-retrospective extraction must inspect the whole delivery chain, not only the final fix. For incidents, regressions, contract drift, weak UI, missed tests, bad reviews, or repeated user corrections, trace the failure through definition, implementation, verification, review/MR or release readiness, and retrospective quality. If the root cause includes this extraction workflow allowing a shallow summary, update this skill or its references before claiming the lesson is landed.
50
50
  - **A retrospective over a LARGE multi-batch / multi-phase session has a second axis beyond the per-delivery chain: distinct lesson-TYPE axes that must each be covered or explicitly marked `no-new-lesson` — (a) per-artifact CONTENT lessons (the specific bug / contract value / domain rule; for research/writing/design programs this axis is the METHOD/CRAFT — how the work was done well), (b) PROGRAM/PROCESS lessons (how the multi-batch effort was structured and driven), (c) WORKFLOW/META lessons (did the retro or extraction itself recur shallow, under-trigger, or stop at the most salient content lesson), and (d) SUSTAIN lessons (what went RIGHT and how the next run reuses it — counts only with mechanism + non-luck evidence + owner routing; **axes (a)-craft and (d) read from the produced-artifact class — an enumeration driven by correction turns cannot reach them and will come back falsely empty**). Landing only the loudest content lesson and declaring the session "fully summarized / 复盘完成" is incomplete.** "LARGE" is not a vibe — it fires when the session already carries a coverage/program structure: a source register or named batch-progress standard was applied, OR the work spanned multiple explicit phases/batches/verticals. A user re-ask after a "done" claim = same-scope correction signal — classify first; never manufacture a lesson. Per-axis detail, re-ask classification, DO-CONFIRM card, `covered-through` watermark: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
51
51
  - **A long operational delivery session also needs a separate non-lesson delivery-state axis (in addition to the content/program/meta lesson axes above).** When the session changed operational delivery state across multiple repositories, branches, MRs, pipelines, releases, or deployable artifacts, the source register must carry that axis — changed artifact set, branch/worktree state, remote/MR state, CI or local verification state, cancelled/retried pipeline state, unresolved risks, and the next concrete action — before "whole-session retro complete" is claimed; closeout records either the axis rows' locator (sanitized labels in the shared landing, real per-repo evidence in scratch/private archive) or `artifact/status axis: not-applicable` with a reason. If required rows are absent, the retro can be reported only as `interim`, even when the extracted lesson text is correct. The row-family fields and closeout-evidence forms: `references/source-to-skill-extraction.md` (Task Retrospective Extraction).
52
- - A blocked verification item is not closed by naming the blockage. Before marking a test, device, browser, service, credential, or environment layer unavailable, attempt the normal remediation path for that layer, such as starting the emulator/browser/service, waiting for readiness, restarting the client daemon, checking local setup scripts, or running the documented fallback. Only record `unavailable` after remediation fails, with command evidence, residual risk, and the next concrete unblock action.
53
- - **A landed CONCLUSION that a tool / capability / lane is unavailable, impossible, or must permanently fail-closed is itself a blocked-verification claim — "fail-closed is the safe default" does not waive the in-env attempt.** Before landing such a conclusion, exercise the tool/capability in the current environment to try to falsify it — **but only within existing sandbox/permission, non-destructive, synthetic-target, and credential-safety boundaries** (the falsification attempt never licenses unsafe mutation, prod/live-credential use, secret-bearing state, or a permission-boundary bypass; that would just trade this rule for the security/authority/data-loss axis). If a safe falsification attempt is genuinely impossible after normal remediation, that is a real `unavailable`/`pending`-with-remediation+residual-risk record; a capability declared impossible without an available safe attempt is `pending`, not `unavailable`/`fail-closed`. **When REVIEWING a change that asserts impossibility/unavailability, independently run the same safe falsification attempt before accepting it** — an inherited "it can't be done" is hypothesis-grade (see the named-convention primary-source re-verify rule). The avoidance-form analysis and failure shape: `references/validation-and-landing.md` (Behavioral Validation).
54
- - A blocked source read is not closed by naming the blockage. If Figma, code, document, API, or repository reads time out, return partial output, or fail transport, switch to a smaller or different read strategy before extracting rules — and when the source is **missing rather than unreadable**, change WHERE you enumerate instead: a store keyed by something other than the unit you are asking about makes "not in this project's directory" read as absence. Both ladders: `references/source-to-skill-extraction.md#blocked-verification-and-source-read-remediation`. Failed or timed-out reads do not count as coverage.
55
- - **Large reads can lose the middle with no reliable signal — the trigger is read-OUTPUT size, so chunk proactively.** A read whose OUTPUT exceeds ~256 lines / ~10 KiB can be silently head+tail truncated (no marker guaranteed), so a single `cat`/whole-file read does not count as coverage even when it returns no error. Whenever you need a **complete** view — whole-file coverage, a no-findings/absence claim, or a load-bearing section read — chunk it under **both ~200 lines AND ~8 KiB** and confirm a mid-file section was ingested. Thresholds, version drift, and measurement: `references/source-to-skill-extraction.md#read-in-chunks-large-reads-lose-the-middle`.
52
+ - A blocked verification item is not closed by naming the blockage. Before marking a test, device, browser, service, credential, or environment layer unavailable, attempt the normal remediation path for that layer — restart the client daemon, run the documented fallback (full rung list and sandbox-denial triage: the next ladder). Only record `unavailable` after remediation fails, with command evidence, residual risk, and the next concrete unblock action.
53
+ - **A landed CONCLUSION is a hypothesis until an operation that could have falsified it has been run** — both a claim that a tool / capability / lane is unavailable, impossible, or must permanently fail-closed ("fail-closed is the safe default" does not waive the in-env attempt) and a DIAGNOSIS of why an observed failure happened — a search hit proves the text EXISTS, not that it RAN on the path that failed, so run the falsifying operation first — exercise the suspected mechanism on the failing path for an observation only IT predicts, or build a paired control differing in exactly ONE variable — **only within existing sandbox/permission, non-destructive, synthetic-target, and credential-safety boundaries** (never unsafe mutation, prod/live credentials, secret-bearing state, or a permission-boundary bypass — trading this rule for the security/authority/data-loss axis). Where no safe attempt is available after remediation the record is `pending` with remediation and residual risk — never `unavailable`, `fail-closed`, or a stated cause — and an unfalsified cause is `hypothesis`, kept off shared surfaces, because withdrawing a landed cause costs more than testing it. **When REVIEWING a change that asserts impossibility/unavailability or rests on a diagnosis, independently run the same falsification attempt before accepting it** — an inherited "it can't be done" or "this is why it broke" is hypothesis-grade (see the named-convention primary-source re-verify rule). Both forms and failure shapes: `references/validation-and-landing.md` (Behavioral Validation).
54
+ - A blocked source read is not closed by naming the blockage. If Figma, code, document, API, or repository reads time out, return partial output, or fail transport, switch to a smaller or different read strategy before extracting rules — and when the source is **missing rather than unreadable**, change WHERE you enumerate instead. Both ladders: `references/source-to-skill-extraction.md#blocked-verification-and-source-read-remediation`. Failed or timed-out reads do not count as coverage.
55
+ - **Large reads can lose the middle with no reliable signal — the trigger is read-OUTPUT size, so chunk proactively.** A read whose OUTPUT exceeds ~256 lines / ~10 KiB can be silently head+tail truncated (no marker guaranteed), so a single `cat`/whole-file read does not count as coverage even when it returns no error. Whenever you need a **complete** view — whole-file coverage, a no-findings/absence claim, or a load-bearing section read — chunk it under **both ~200 lines AND ~8 KiB** and confirm a mid-file section was ingested. Detail: `references/source-to-skill-extraction.md#read-in-chunks-large-reads-lose-the-middle`.
56
56
  - Think across the full delivery lifecycle before editing: product intent, design/UX, implementation, debugging, test strategy, launch acceptance, iteration feedback, team onboarding, and normal users without source access. A rule that improves only one slice while leaving another slice ambiguous is incomplete or belongs in a narrower skill.
57
57
  - Evidence must come before new rules. Do not add a new conceptual layer, workflow gate, or strong claim first and then backfill supporting sources. If a useful rule appears before source review, keep it as a working hypothesis and do not land it until evidence confirms it, narrows it, or routes it elsewhere. For subjective design, UX, frontend/client, product, architecture, or review rules, unverified external expertise is not enough to land executable guidance.
58
58
  - **Product-agnostic / industry-practice skills require an external authoritative source class in the evidence plan, not internal corpus alone.** An extraction sourced only from one internal corpus (an SOP, one repo, one project doc) shows what *this org* does, not whether the skill matches the public state of the art.
@@ -123,7 +123,7 @@ Use this skill to turn observed experience into durable agent skills without cop
123
123
  ### What to extract, content placement & domain (UI/UX) judgment(抽什么 / 内容放置 / 领域判断)
124
124
 
125
125
  - Extract behavior, decision rules, quality gates, evidence patterns, and routing boundaries; do not extract business nouns, repo names, IDs, one-off incidents, or stale implementation details.
126
- - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files.
126
+ - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files. Each reference links one level from the entrypoint and stays inside the reference line budget; `references/attention-budget-ratchet.md` owns that budget, the write-side authoring norms, and the design invariants any size/budget gate must satisfy.
127
127
  - A skill must be executable, not only directional. For design, client, testing, debugging, or review skills, include concrete workflow steps, decision points, state/checklist coverage, and verification evidence so future agents do not produce work that is compliant but weak.
128
128
  - Design/client extraction must cover the judgment layer, not only the engineering layer. For UI/UX, extract aesthetic logic, interaction logic, behavioral logic, and user psychology from source evidence before landing rules about layout, components, breakpoints, or tests.
129
129
  - UI/UX judgment extraction must use observable proxies, not adjectives. Read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming behavioral or psychology rules. Use `references/uiux-judgment-extraction.md` for the required method.
@@ -150,10 +150,10 @@ Use this skill to turn observed experience into durable agent skills without cop
150
150
  - **One self-detectable firing point does exist and must be used: the moment YOUR OWN output names 沉淀 / 提炼 / 复盘 / "distil this into a skill", OR **enumerates what an external source has that we lack** (a gap list vs another pack; see `references/firing-point-placement.md`), that naming is a trigger to RECOGNISE the owner and load it** — not a licence to widen scope: shared-skill edits still need the authority you already have, so when the user's request covered only a status review or a narrow fix, record the extraction as `pending` with the owner named and ask rather than self-authorising a shared-skill change off your own suggestion.
151
151
  - The mechanical backstops are (a) the closeout gate — a committed skill-change with neither a visible in-session `skill-extraction-workflow` invocation nor the round's durable charter/target-output record is `interim` (per the closeout gate's evidence forms) — and (b) **user-signal escalation**: you generally cannot self-count misses you did not notice, so a user-pointed-out under-trigger is a recurrence check (was there a similar miss earlier this session, even on another task?) and, if so, escalates to tightening the always-on discipline rather than landing another narrow per-case trigger.
152
152
  - **Firing-point-placement corollary:** when the SAME meta-class (a precise gate walked past at the routing → pre-code/design transition) recurs at a *new* lifecycle sub-point despite prior bootstrap-salience + the closeout gate, the durable lever is **moving the owning gate's firing point ONTO the transition itself** (pre-substance-draft AND pre-first-impl-edit) and sharpening *name→invoke* — naming/knowing an owner is NOT invoking/loading it, and a named-but-unloaded owner's mechanical rules never fire — at the SAME transition, NOT another bootstrap/per-case bullet or more prose.
153
- - **Record-field corollary (the forgery surface):** when you land an owner gate as a *field in a record* — a checklist row, a boundary-record line, a CLI flag taking owner names, a "decision:" slot — that field is fillable without invoking the owner, and filling it is what *feels* like discharging the gate; any field naming an owner therefore carries an explicit invoke bar on its triggered values.
153
+ - **Record-field corollary (the forgery surface):** any field that NAMES an owner — a checklist row, a CLI flag, a "decision:" slot — is fillable without invoking that owner, and filling it is what *feels* like discharging the gate, so it carries an explicit invoke bar on its triggered values.
154
154
  - The self-detect firing point's authority boundary and observed shape, the record-field corollary expansion (the invoke-bar coverage set-diff mechanics, the delegation-dispatch worked case), the worked recurrence-chain, and the landed owner-dispatch implementation: `references/firing-point-placement.md`.
155
155
  - **Run your own adversary to convergence BEFORE any "done / fixed / passing / covered / converged / complete" claim — your own such claim is the least-trustworthy thing you emit.** For any non-trivial completion/coverage/convergence claim, you must have already run — **yourself, not deferred to the user** — the verification or adversarial pass that would catch its failure, to a **clean fresh result** (a first clean pass on the current candidate, never a "confirm my fix" pass), OR **downgrade the claim to `interim` and name what you ran vs. didn't**. "Covered / converged / already handled" is a claim, not a status — back it with firing-path or clean-pass evidence or do not emit it; this self-adversary duty never narrows the mandatory dual-track challenge (it is the always-on generalization of self-audit-to-convergence, not a replacement for the gate).
156
- That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never disable an authorization, idempotency, or deletion guard and exercise it against a shared or live dependency; mutate against isolated dependencies or at the lowest layer that avoids them, and where neither is possible record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point the check at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A failing anchor is first a question about the ANCHOR, not a verdict on the implementation** (§Self-audit). **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
156
+ That pass is a **walked enumeration over the properties the candidate asserts, never a re-read**: a property whose killing mutation you cannot name was never verified, and re-reading your own prose can only ever confirm that the prose is self-consistent with itself. **A mutation you did not APPLY is a hypothesis, not evidence** — bound its blast radius (never buy a RED by disabling a guard against a shared or live dependency; where no isolated path exists record the property `unverified`). **Prove the oracle can fail before trusting its clean verdict** — point it at something you know is broken and watch it report that; a check that can only ever say clean is no evidence, and whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration. **A failing anchor is first a question about the ANCHOR, not a verdict on the implementation** (§Self-audit). **A validated oracle is still clean only over the DIMENSIONS it crossed** — proving it can fail says nothing about the axis you never varied, so a clean run is reported with the dimensions it covers, and the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering — `testing-strategy` owns that list) before values. If no contradicting observation exists, the property is `unverified` and must be labelled that way rather than counted as audited.
157
157
  A scoped "X verified; Y not run" is an interim checkpoint, **not** `done`/`complete`/`landed`: `Y not run` blocks a done/complete/landed claim unless a **risk owner — the user/maintainer, never the agent self-accepting — explicitly accepts the gap AND it is tracked to that owner** (agent self-labeling "risk accepted" or "deferred" does not qualify; scoping is a downgrade, never a license to call the narrowed slice done). **Recurrence signal:** a user prompting you to keep digging / verify / disputing a "covered/converged/done" is a premature-completion signal — on the **2nd** such correction in a session (even across different tasks) escalate to tightening this discipline, not just fixing the one case (per the repeated-correction escalation above).
158
158
  The full self-adversary method — the mutation enumeration, the applied-mutation discipline, the independent-oracle validation, the re-owe-after-fixes rule, the graded-verdict calibration, and the recognition-dependent honesty caveat: `references/dual-track-review-gate.md` §Self-audit.
159
159
  - Automatically trigger durable learning when extraction work exposes a reusable failure — **and when ordinary delivery work does, capture it here too, but without extraction taking over the delivery**: let the active owner (`product-rd-workflow` / `defect-diagnosis` / `testing-strategy` / …) handle the immediate work first, then route the durable lesson here. **For a premature-stop correction after affirmative continuation**, immediate recovery means first rerun the active owner's current continuation/blocking gate in full (for product R&D, Pre-Final Continuation Gate steps 1–6) against current state, then follow its observable outcome — proceeding only when a literal binding exists (the original proposed-next action/scope plus literal assent, preserved in the visible conversation or quoted exactly in trusted host-owned session/compaction state — never reconstructed, broadened, or substituted — or the user's correction literally naming the paused action and scope) — a semantic compaction paraphrase or a bare "why did you stop" complaint is not path-(b) authority, and a `blocked:` recovery without the step-1 evidence and a specific missing authority/ambiguity is invalid — asking again when neither binds, the user intervened, or scope/gates changed, and never copying real conversation text into a shared repository record. Do not let correction RCA or extraction extend a still-authorized delivery, and do not let stale assent bypass a newly pending or inconclusive gate. After delivery recovery, correction RCA plus the durable prevention landing and verification are still due before the turn can be reported complete; otherwise report `interim`. The full binding rules, the `continuing:`-line form, and the invalid-`blocked:`-recovery rule: `references/resume-paused-delivery.md`.
@@ -172,15 +172,15 @@ Use this skill to turn observed experience into durable agent skills without cop
172
172
  - For any correction where the missed step was covered by any rule in this workflow that a reasonable reader would apply to the scenario, add or tighten a closeout gate that would have blocked the exact premature final answer.
173
173
  - **"The owner skill already states the rule" / "avoid monotonic growth" does NOT license a memory-only or no-op landing when the gate demonstrably failed to fire.** If a rule exists yet the failure still happened (and would recur for another agent or project), adequate *content* is not adequate *enforcement*: land the firing mechanism — the trigger, closeout step, validator, or merged clause that makes the existing rule actually catch this case, in the owning shared skill — or prove it now fires. Retreating to a personal memory or "no change, content is fine" while the gate stays un-fired is the dodge this prevents (memory is supplement only, per the memory-only-insufficient rule).
174
174
  - For changed upstream owners, `check-ccl-skills.sh` (via `scripts/impact-chain-gate.rb`) is the mechanical closeout over every added source-register row; the declaration format (behavioral-evidence / observed-failure / firing-path fragments), anchor rules, wording-only classification, the `RED-baseline` floor, and the author-declaration trust model: `references/external-practice-controls.md#behavioral-evidence-and-attestation`.
175
- - For shared-skill changes, classify the diff before finalization: `shared-skill change` means any change under a shared skill package, including `SKILL.md`, references, scripts, validators, templates, generated outputs, metadata, and examples; `wording-only` means punctuation, grammar, typo, formatting, or synonym substitution with no change to trigger, scope, routing, validation, condition, example, owner, or acceptance meaning, and touching no frontmatter (any `description`/frontmatter edit, even a pure typo fix, is a routing-surface change, never wording-only); all other changes are non-wording.
175
+ - For shared-skill changes, classify the diff before finalization: `shared-skill change` is defined in `references/dual-track-review-gate.md` (it also covers the plugin-shipped command/behavior surfaces outside `skills/`); `wording-only` means punctuation, grammar, typo, formatting, or synonym substitution with no change to trigger, scope, routing, validation, condition, example, owner, or acceptance meaning, and touching no frontmatter (any `description`/frontmatter edit, even a pure typo fix, is a routing-surface change, never wording-only); all other changes are non-wording.
176
176
  - Every shared-skill change, including wording-only edits, must include a recorded independent review row before commit.
177
177
  - Non-wording shared-skill changes must also include a recorded challenge row AND a recorded behavioral-evidence row before commit (actual behavior/routing deltas require a true `RED-baseline`; unchanged controls use paired `semantic-control`; status rules in `references/dual-track-review-gate.md`). A non-wording owner package must include at least one `RED-baseline` row, so stable-control labels cannot self-clear the package. **For a DESTRUCTIVE/irreversible change, a `RED-baseline` row must show the protected predicates were mutated, not merely that negative cases ran.** Executing the must-NOT-touch cases is necessary but not sufficient: a negative probe that would still pass with its protecting predicate removed is evidence of nothing (recurring shape: a degenerate fixture short-circuits every probe on an unrelated conservative branch, so the safety predicate is never reached and the green suite certifies the hole).
178
178
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
179
179
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
180
180
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
181
- - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=2` (three rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled round 4/5. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
181
+ - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=1` (2 rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled new rounds. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
182
182
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
183
- - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At round 3 validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
183
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At budget end validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
184
184
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
185
185
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
186
186
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
@@ -240,7 +240,7 @@ Use this skill to turn observed experience into durable agent skills without cop
240
240
  - For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
241
241
  - Trigger situations and users/tasks it should serve.
242
242
  - What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
243
- - For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
243
+ - For NEW skills and subjective/high-impact skills (design, UX, client, product workflow, architecture, review), must define eval/pressure scenarios, baselines, acceptance criteria pre-draft.
244
244
  - For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
245
245
 
246
246
  3. Inventory evidence.
@@ -256,23 +256,21 @@ Use this skill to turn observed experience into durable agent skills without cop
256
256
  - Before extracting rules from broad or mixed sources, summarize coverage, contradictions, thin areas, and source-quality gaps.
257
257
  - For tool failures during source reading, keep a retry ledger: failed method, observed error, fallback method, source rows recovered, source rows still missing, and whether the fallback evidence is strong enough for the target rule. Do not report a pressure test, re-read, or extraction as complete when the only successful step was rephrasing existing notes.
258
258
  - When changing a source-derived skill's conceptual layer, refresh the relevant source classes before editing. If the source was not refreshed, label the change as wording/routing cleanup, not source extraction.
259
- - Record coverage using honest depth labels that map to the charter: wording cleanup/no new source read, targeted check, file-level refresh, node/artifact-level inventory, full workflow extraction, or generator/tooling change.
259
+ - Record coverage using one of the depth labels defined under `references/source-to-skill-extraction.md#extraction-charter`; invent none outside that full set.
260
260
  - If evidence is thin, constrain the output: fewer rules, explicit low-confidence notes, wider "do not use when" boundaries, and a clear list of evidence that would improve the skill.
261
261
 
262
262
  4. Extract candidate rules.
263
263
  - Convert observations into reusable rules: "when X, do Y, verify Z".
264
264
  - Separate invariant rules from stack-specific examples.
265
- - Before editing any stack-specific skill, build the sibling-generalization mini-map: source stack, sibling stacks, shared workflow owner, per-sibling decision, and reason. If a candidate is language-agnostic, route it to the shared workflow/testing/architecture skill first, then add stack-specific implementation notes only where needed.
265
+ - Before editing any stack-specific skill, build the sibling-generalization mini-map the Core Rules owner-generalization group defines; route a language-agnostic candidate to the shared owner first, then add stack-specific implementation notes only where needed.
266
266
  - Mark each candidate as keep, merge, discard, or route to another skill.
267
267
  - Map each kept or routed candidate to its owning target in the target-output map before editing. Do not finish a design/client extraction until implementation, testing, product workflow, and sibling-client implications have been checked and either updated or explicitly marked unchanged. For mini-program surfaces, include `miniapp-product-dev` in that owner check.
268
268
  - Keep an extraction ledger during analysis: rule origin (`observed` or `hypothesis`), source IDs inspected before the rule was drafted, evidence grade, candidate rule, conflict, decision, target skill/reference, and reason.
269
269
 
270
270
  5. Generalize and place content.
271
- - Put trigger, routing, core workflow, and non-negotiable rules in the skill entrypoint.
272
- - Put detailed reference material in direct reference files.
271
+ - Place content by the Core Rules content-placement rule (entrypoint owns trigger, routing, core workflow, and non-negotiables; direct reference files own the detail, within their budget).
273
272
  - Generalize from the evidence ledger, not from a polished rule draft. Do not search for examples to justify a rule that has already been written.
274
- - Choose capability names that describe what future users can do with the skill, such as complex workspace patterns, service architecture, test strategy, or document tightening. Do not name durable outputs after the source file, source project, old feature, or temporary extraction task.
275
- - Keep source identity in provenance fields only. If a source name is useful for audit, label it as source evidence; do not make normal users route through that source name to understand or trigger the skill.
273
+ - Name durable outputs per the Core Rules naming rule (complex workspace patterns, service architecture, test strategy, document tightening are capability names; the source file, source project, old feature, and this extraction task are not), and do not make normal users route through a source name to understand or trigger the skill.
276
274
  - Turn source lessons into execution recipes: analyze first, implement with ownership boundaries, debug by layer, test at the right level, and verify on the real rendered/runtime surface.
277
275
  - For design and client skills, place the four judgment layers explicitly (aesthetics / interaction logic / behavioral logic / psychology — per-layer semantics and decision fields: `references/uiux-judgment-extraction.md`); do not hide them inside generic "UI polish" wording.
278
276
  - Write each judgment layer's delta per the Core Rules judgment-delta rule (new / confirmed / narrowed / routed / no new evidence — a restatement of existing principles is not newly extracted knowledge). For visual direction/tokens, use the token provenance fields in `references/uiux-judgment-extraction.md` and state what they mean for design, implementation, and testing; otherwise it is only a static source note.
@@ -290,8 +288,7 @@ Use this skill to turn observed experience into durable agent skills without cop
290
288
  - **Cross-section facet-ownership check** (Core Rules ↔ Step 0–6): for any edit touching `## Core Rules` or a Workflow step, record whether the changed facet is rule/invariant-owned (→ Core Rules) or procedure/checklist-owned (→ the step), and confirm the opposite surface only *points* to it, not restates it. Same-facet text living in both surfaces is a drift defect — converge toward the canonical surface per the Start-here boundary contract before landing.
291
289
  - All referenced files exist and are one level from the skill entrypoint.
292
290
  - No business or source-repo leakage remains in executable guidance.
293
- - Capability naming is source-neutral: old source artifact names, old page names, and old scenario labels are absent from executable guidance or clearly marked as provenance.
294
- - For any rename or generalization, residual searches for the old file name, old source label, old English shorthand, and old capability phrase are clean or justified as provenance.
291
+ - Capability naming is source-neutral, and after any rename or generalization residual searches for the old source artifact name, file name, source label, page name, scenario label, English shorthand, and capability phrase are clean, absent from executable guidance, or clearly marked as provenance.
295
292
  - Trigger boundaries do not collide with sibling skills.
296
293
  - Coverage matrix has no unexplained gaps: each relevant source category is used, routed, discarded, or marked unavailable with a reason.
297
294
  - For broad or multi-skill extraction, validation must include the durable source register and a target-output map showing which target skills or outputs were updated, which sources informed them, which sources were excluded, and which source classes remain pending. If any required row is pending, the work can be landed only as an interim checkpoint, not as complete.
@@ -304,7 +301,7 @@ Use this skill to turn observed experience into durable agent skills without cop
304
301
  - Compare lifecycle impact against the target-output map. If an affected stage has no owning target or no-output reason, validation fails even if all listed targets have correct diffs.
305
302
  - For UI/UX, Figma, frontend, app, miniapp, or client extraction, validation must include the judgment-delta matrix per the Core Rules `What to extract` section (each layer marked new / confirmed / narrowed / routed / no new evidence with source observations; visual direction/tokens rows carry token provenance and the design/dev/test owner decision). If the matrix is absent or lacks the required visual-token fields, the work can be reported only as source inventory or execution hardening, not design-judgment extraction.
306
303
  - UI/UX validation must include the minimum pressure set from `references/uiux-judgment-extraction.md`: long content, empty/no-data, error/retry, slow/weak network, permission/disabled, narrow/responsive, accessibility text scaling, orientation or viewport change where relevant, interruption/return recovery, and repeated-use/cache-hit behavior; mark irrelevant cases explicitly instead of silently skipping them.
307
- - Source-map and final-response claims match actual evidence depth. Broad coverage language is a blocker unless the evidence ledger supports it.
304
+ - Source-map and final-response claims match actual evidence depth. Broad coverage language is a blocker unless the evidence ledger supports it. The same bar governs CAUSAL claims: a stated cause names the falsifying operation run against it or is `hypothesis`.
308
305
  - Pressure-test authenticity check: if the final claim says "tested", "pressure-tested", "re-read", "re-extracted", or equivalent, validation must show the source artifact reopened in that validation pass. If no source artifact was reopened, downgrade the claim to "static review" or rerun the test correctly.
309
306
  - Repeated-correction lessons have a visible prevention point in the owning skill, reference, checklist, validator, or final-response constraint recorded in a skill/reference. A chat-only apology, summary, or ad-hoc wording in the current reply does not count as durable landing.
310
307
  - Every landed source-derived rule has an `observed` ledger entry with concrete observations. Any rule that started as a hypothesis is either converted to observed evidence, narrowed, routed, discarded, or explicitly left out of the landed skill.
@@ -0,0 +1,37 @@
1
+ # Attention-budget ratchet — design invariants for size/budget gates
2
+
3
+ An attention-budget gate limits how much prose an agent must hold to use a surface: the entrypoint body-word/byte gate, the every-session-injection byte gate, the `description` 800-char cap, and the reference-file line gate (all enforced by `scripts/check-size-budget.sh` or the canonical validator). This file owns the design invariants those gates share and the write-side norms for reference files. Companions: `rule-consolidation.md` owns the prose doctrine (merge-into-canonical, why rule sets must not grow monotonically); `skill-listing-budget.md` owns the host's listing-budget mechanism; `description-authoring.md` owns the description surface.
4
+
5
+ ## The five ratchet invariants
6
+
7
+ Any NEW budget/size gate, and any modification to an existing one, is checked against all five before landing (this is the budget-gate instantiation of the design-time operability check in `dual-track-review-gate.md` — run that check's four legs too). A gate missing one of these fails in a predictable way, named per item:
8
+
9
+ 1. **Stable proxy estimator** — the metric is deterministic and environment-independent: Unicode letter/number word runs with Han counted per ideograph, raw byte size, or physical line count. Never a model/tokenizer-dependent estimate: two environments disagreeing on the measure turns the gate into noise, and a changed estimator silently invalidates every recorded allowance. Deterministic also means encoding-normalized before measuring — a line count taken over raw bytes reads a CR-delimited file as one line, so line endings are folded to LF first.
10
+ 2. **Anti-false-green sentinel** — a run that could not evaluate says so: base-unresolvable prints an `*_unevaluated` token (never the ok token), probe failures fail closed as partials, and on any block the last token is the failure marker. The ok token must be unearnable by losing the base; a consumer grepping for ok must never read an un-run gate as a pass.
11
+ 3. **Zero tolerance for new debt** — a NEW surface over budget blocks outright. There is no exempt marker, no waiver flag, and no way for a candidate to nominate its own baseline; structural exclusions live in the gate, owned by the gate.
12
+ 4. **Legacy may only shrink** — an existing over-budget surface is frozen at its base measure: level or shrinking lands, any growth blocks. Rename credit is path-paired and non-growing (move plus growth blocks as growth). This is what makes a uniform cap deployable over a corpus that already exceeds it, without a rewrite round and without rewarding a rush to pre-shrink.
13
+ 5. **A missing baseline is never a pass** — the comparison base comes from revision history (`CCL_SKILL_BASE_REF`, upstream, or merge-base), so there is no stored manifest to go stale; when no base resolves, the verdict is unevaluated (invariant 2), and reddening that state is the caller's pipeline decision (CI always exports the base ref).
14
+
15
+ Two cross-cutting corollaries:
16
+
17
+ - Average headroom must never fund a single over-budget surface: the ratchet judges each file alone, and corpus-level counters stay visibility-only.
18
+ - Debt counters and advisory bands are never a clean-landing waiver nor authorization to keep growing a surface; only the delta verdict blocks.
19
+
20
+ ## Reference-file write-side norms
21
+
22
+ The read side already defends against oversized files (chunked reads under ~200 lines, references one level deep). These norms are the write side, enforced as a delta ratchet over `skills/*/references/**/*.md`:
23
+
24
+ - A NEW reference file over 500 physical lines must not land — split it by subtopic before landing (the gate blocks new-or-crossing files; 500 exactly passes).
25
+ - An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
26
+ - Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
27
+ - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
+ - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
29
+
30
+ ## Official-clause verdicts (provenance)
31
+
32
+ Registered claims were re-verified against the primary source (Anthropic "Skill authoring best practices", docs.claude.com, read 2026-08-31) before landing; per-clause disposition:
33
+
34
+ - "Keep SKILL.md body under 500 lines" — present verbatim, but it scopes to SKILL.md, not references. Already covered more strictly here by the 5000-body-word delta ratchet. The 500-LINE reference cap above is a repo-internal norm motivated by the read-side chunking evidence, and is labeled as such — never cite it as an official requirement.
35
+ - Table-of-contents mandate for long references — NOT PRESENT in the current official text. The official remedy for `head -100` partial reads is keeping references one level deep (already a Step 6 validation rule). The `##`-section advisory above rests on repo-internal evidence only.
36
+ - "Tested with Haiku, Sonnet, and Opus" — present, conditional on the models you plan to ship to. This repo's skills inherit the session model and the eval layer exercises real sessions, so no multi-model matrix is mechanized; the clause fires only if the repo starts shipping model-pinned skills.
37
+ - Checklist anti-patterns (time-sensitive info, terminology, concrete examples, forward-slash paths, voodoo constants, scripts-solve-not-defer) — present; landed above as authoring norms. The recurring-anti-patterns grep panel is NOT their landing surface: its admission rule requires a class observed in 2+ skills of this repo.
@@ -68,6 +68,15 @@ Bad: `Proactively invoke when the user shares a draft doc / spec / plan and asks
68
68
 
69
69
  Good: `Skip when the ask is already scoped to one stack (e.g. "fix this React render bug" → web-react-dev; "add a GORM index" → go-microservice-dev; "调下这个按钮间距" → product-ui-ux-design), or when the user is reporting a defect / failure / regression → defect-diagnosis owns reproduction and root cause first.`
70
70
 
71
+ ## Body routing pointers: the quadruple
72
+
73
+ The description is not the only routing surface an author writes: skill/reference BODY text routes too, through cross-skill and cross-reference pointers ("route to X", "read Y before Z"). Tier-1 static analysis parses only the description, so body pointers are held to an authoring contract instead:
74
+
75
+ - **Every cross-skill or cross-reference routing pointer in body text must carry the routing quadruple**: trigger (when to go read the other surface), scope (which file or small subset), output (what decision/artifact to extract), and return point (where to resume in the owning workflow). A pointer whose parts are obvious from sentence position may state them compactly ("at step 3, read X's §Y for the Z decision, then continue step 4").
76
+ - A bare "refer to X if useful" / "参见 X" with no trigger and no extraction target is never a landing shape: such pointers rarely fire, and when they do fire they read too much (source-observed: unbounded pointers were the dominant dead-routing shape in an adopting skill pack; the pack that enforced the quadruple had two verified adopters and no dead pointers).
77
+ - The description-side Skip-when `→ skill` idiom already satisfies the quadruple (trigger = the skip condition, scope = the target skill's entry, output = ownership transfer, return = none) — no extra wording needed there.
78
+ - Detector pairing: a pointer that routes but is never read shows up as the **silent skip** failure mode in `eval-routing.md`'s failure-mode vocabulary (B-side probes / Tier-3 traces), not in Tier-1/2 — fix the pointer's trigger and scope, not the description.
79
+
71
80
  ## Precedence against session-injected process skills
72
81
 
73
82
  A routing / workflow skill competes not only with sibling skills but also with process-discipline skills that some hosts INJECT at session start with very forceful language (e.g. a brainstorming skill whose description says "MUST use before creating features / adding functionality"). At initial routing time the router primarily sees skill names, descriptions, and host/session rules — the workflow body has not loaded yet, so an "Entry precedence" paragraph in the body does NOT win the routing decision. If your routing skill should own the entry point for a request class that an injected process skill also claims, the description must: