@ccoalm/ccl-skills 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +5 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +1 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +12 -12
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +2 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +2 -2
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +1 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -15
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +13 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +37 -30
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -1
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +81 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +12 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +30 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +15 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +25 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +27 -21
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +25 -15
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -9
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  77. package/dist/assets/release.json +127 -67
  78. package/package.json +1 -1
@@ -311,7 +311,7 @@ Implementation rules:
311
311
  Motion must have a product purpose — comprehension, orientation, feedback, or deliberate brand/emotional expression. Operational, finance, moderation, dense-data, and destructive flows default to calmer motion (this complements, not contradicts, the expressive-defaults note in the next section). Apply across iOS and Android:
312
312
 
313
313
  - **Budget attention-grabbing motion.** Avoid more than roughly two *attention-grabbing or decorative* animations competing at once in one view; essential status indicators (a progress spinner, a skeleton shimmer) and a single choreographed timeline are exempt. Layered competing motion reads as jank, not polish.
314
- - **Let platform and design-system motion tokens own duration and curve; only tune micro-feedback.** Micro-feedback (tap, toggle, small in-place state change) defaults to a short band (about 150–350ms), but system navigation, sheet presentation, predictive-back/gesture, hero choreography, and design-system motion tokens (e.g. M3 Expressive) carry their own longer, tuned durations — do not clamp them to the micro band. Reserve custom playful bounce/overshoot for light surfaces; on serious, destructive, financial, or trust-sensitive flows do not add custom overshoot that makes finality feel reversible or celebratory (standard platform component motion, including system spring settling, is fine).
314
+ - **Let platform and design-system motion tokens own duration and curve; only tune micro-feedback.** Micro-feedback (tap, toggle, small in-place state change) defaults to a short band (about 150–350ms — a team heuristic, not a platform mandate: it sits inside Material 3's official duration tokens, which span 50–400ms across the short/medium steps, while Apple HIG gives no numeric duration guidance), but system navigation, sheet presentation, predictive-back/gesture, hero choreography, and design-system motion tokens (e.g. M3 Expressive) carry their own longer, tuned durations — do not clamp them to the micro band. Reserve custom playful bounce/overshoot for light surfaces; on serious, destructive, financial, or trust-sensitive flows do not add custom overshoot that makes finality feel reversible or celebratory (standard platform component motion, including system spring settling, is fine).
315
315
  - **Respect the OS reduce-motion setting natively, not only via web `prefers-reduced-motion`.** iOS exposes it directly: `UIAccessibility.isReduceMotionEnabled` / SwiftUI `\.accessibilityReduceMotion`. Android has no single reduce-motion boolean — gate custom animation on `ValueAnimator.areAnimatorsEnabled()` (or the framework duration scale), treating the animation scale as a capability signal rather than a reduce-motion *intent* flag, and fall back to `Settings.Global.*_ANIMATION_SCALE` only when needed. Classify each motion as decorative / spatial-orientation / essential: reduced motion drops decorative movement and shortens or simplifies spatial/essential movement, but must still show required feedback and state changes (progress, a status flip, a gesture preview).
316
316
  - **Motion must not shift layout or delay the task.** Confirm active state without reflowing surrounding content (per the bottom-tab rule above), and never hold loading/disabled/error feedback behind an entrance animation.
317
317
 
@@ -320,5 +320,5 @@ Motion must have a product purpose — comprehension, orientation, feedback, or
320
320
  When the target product ships on iOS 26+ / Android 16+ / Material 3 Expressive defaults, the design baseline shifts. Treat these as platform-default changes that affect token tuning, motion budget, and gesture geometry — not as visual style copies.
321
321
 
322
322
  - **iOS 26 Liquid Glass (Apple, WWDC 2025; iOS 26 / iPadOS 26 / macOS Tahoe 26 / watchOS 26 / tvOS 26)** introduces a translucent system material that reflects and refracts surrounding content and dynamically transforms across controls, navigation, app icons, and widgets. Apps built with standard SwiftUI / UIKit / AppKit components inherit the new design automatically when rebuilt against the Xcode 26 / iOS 26 SDK; custom-drawn UI (custom CALayers, manual gradients, hand-rolled tab bars) does NOT inherit it and must adopt explicitly. Apple ships a temporary opt-out in Xcode 26 so teams can ramp on their schedule rather than be forced to ship Liquid Glass the day they upgrade SDK. Design impact: (1) custom translucent / blur / glass material tokens MUST be tuned separately for Light, Dark, AND Increased Contrast appearances — Apple's own system colors were re-tuned across all three; (2) typography baseline became bolder and left-aligned, so if the product design uses centered or thinner type to "feel premium" on iOS, re-validate hierarchy on iOS 26; (3) chrome-on-content (tab bars, sidebars) now refract content beneath, so check that overlay surface tokens still keep on-surface text readable when chrome sits over high-contrast or saturated media. Do not adopt Apple's Liquid Glass as a cross-platform default token — it is a system material with system-tuned color/blur/refraction, and cloning it on web / Android as "default brand glass" produces a hard-to-maintain knock-off. Web / Android may use a deliberate glassmorphism treatment when the product needs it, but it must declare its own contrast budget, performance fallback (opaque mode when GPU / battery / low-end device requires), and an opaque-mode trigger that does NOT rely solely on `prefers-reduced-transparency` (the CSS media feature is real but not Baseline — Chrome desktop / Firefox stable lag — so back it up with an in-app "Reduce transparency" setting, platform-equivalent OS preference where available, or a default-opaque variant for non-supporting browsers); cite the explicit rationale in the design spec rather than treating glass as a free aesthetic upgrade.
323
- - **Material 3 Expressive (Google, 2025)** is an opt-in expansion of Material Design 3 with research-backed motion theming tokens, more expressive shape / color / typography, and an explicit emotional-design dimension (research showed expressive variants outperformed baseline on "energetic / emotive / positive / playful / friendly" perception). When the project uses M3 Expressive defaults (Jetpack Compose with M3 expressive themes, libraries pulling expressive motion tokens — note AndroidX `MotionScheme.expressive()` is alpha at the time of writing, not a stable everywhere-default), expect default animation durations and easings to be more energetic than baseline M3. Review whether *operational* / *finance* / *moderation* / *dense-data* surfaces should override motion tokens to a calmer set (`MotionScheme.standard()`) rather than inheriting expressive defaults, because dense workbench surfaces work against the expressive tone and feel jittery under it.
323
+ - **Material 3 Expressive (Google, 2025)** is an opt-in expansion of Material Design 3 with research-backed motion theming tokens, more expressive shape / color / typography, and an explicit emotional-design dimension (Google's design.google research article reports 46 studies with 18,000+ participants; expressive variants outperformed baseline on "energetic / emotive / positive / playful / friendly" perception). When the project uses M3 Expressive defaults (Jetpack Compose with M3 expressive themes, libraries pulling expressive motion tokens — note AndroidX `MotionScheme.expressive()` is alpha at the time of writing, not a stable everywhere-default), expect default animation durations and easings to be more energetic than baseline M3. Review whether *operational* / *finance* / *moderation* / *dense-data* surfaces should override motion tokens to a calmer set (`MotionScheme.standard()`) rather than inheriting expressive defaults, because dense workbench surfaces work against the expressive tone and feel jittery under it.
324
324
  - **Android Predictive Back is default-enforced for apps targeting API 36 (Android 16, 2025)**. Design implication: the back gesture is no longer a single instant action but a *preview-then-commit* gesture — during the swipe the inner area scales down and the destination peeks behind; on commit-threshold crossing the contents fade-through to the destination (Android recommends `STANDARD_DECELERATE` or `PathInterpolator(0f, 0f, 0f, 1f)` for the progress easing). The system handles the previous-destination snapshot automatically for stock navigation; the design only needs to define preview behavior for custom-managed states: modals, bottom sheets, full-screen overlays, in-screen multi-step wizards, and any flow that owns its own back stack. Avoid placing draggable controls or custom horizontal-edge gestures inside the system gesture inset; they fight the OS back gesture and feel broken. For multi-step in-screen flows (form wizards, multi-pane), the design owns the *semantic back contract* (back pops inner step, not the whole screen) — *implementation* should integrate through the owning navigation stack's predictive-back support (Jetpack Navigation predictive-back APIs, `react-native-screens` predictive-back, Flutter `PopScope` / `NavigatorPopHandler`, native Fragment back-stack handlers) rather than wiring an ad-hoc `OnBackPressedCallback` at the screen level. Compose's lower-level `PredictiveBackHandler` is appropriate when the Compose screen owns its own back stack (no navigation library on top); when a navigation library is present, prefer the library's predictive-back hook so the system snapshot and inner-step pop stay in sync. Ad-hoc handlers on top of a navigation library double-pop, desync the system snapshot animation, or bypass the library's intended back stack and the regression is hard to reproduce because the OS-level animation still looks right.
@@ -20,6 +20,7 @@ Desktop and mobile catalogs below are illustrative component vocabularies, not a
20
20
  - Use product UI files for page composition, state coverage, and interaction-pattern evidence; do not inherit their old domain requirements.
21
21
  - Use third-party UI kits and icon libraries only as reference-only coverage checks for component categories, state variants, icon discipline, and documentation quality. Do not copy their brand, marketing IA, or visual identity into the product skill.
22
22
  - Do not hardcode colors, spacing, radii, or typography when a design-system token exists.
23
+ - Existing non-compliant code is not permission: an old component's hard-coded color or ad-hoc style is recorded debt, never a precedent to copy into new work. When the requested result cannot be achieved within the design system's current rules, stop and route the gap to the design-system owner instead of silently inventing a new visual rule.
23
24
  - **Color tokens 优先 HSL 而非 Hex / RGB**(hand-tuning same-hue 变体):HSL 让"同色不同亮度"(hover、disabled、bg tint、border-on-bg shade)通过只改 L 直接派生;Hex / RGB 改 1 个亮度需要算 3 通道易调不准。toolchain 支持时**优先 OKLCH / LCH** 做感知一致的 color ramp(HSL 在跨 hue 时亮度不感知统一)。Token 源用 HSL/OKLCH 表达 intent;输出层(CSS / iOS / Android)按需 convert;设计工具如 Figma 可能存 RGB,token spec 保留 HSL/OKLCH 语义即可。
24
25
 
25
26
  ## Token Sync Pipeline (Figma → Front End)
@@ -49,6 +49,14 @@ Always include concrete file/line references for code reviews and Figma file/pag
49
49
  9. **Check serious-domain adaptation** when relevant: source, timestamp, partial data, confirmation, audit labels, and no unsafe optimistic UI.
50
50
  10. **Check platform-convention conformance** on the named target: use `external-ui-ux-quality-benchmarks.md` to classify authority and boundary, recheck the current first-party platform source, and map applicable criteria into `delivery-contract.md`. Preserve requirement versus recommendation strength and verify on that platform's rendered runtime; do not reuse a combined HIG/Material checklist as a cross-platform standard.
51
51
 
52
+ ## Diff-Scoped Review
53
+
54
+ When the audit target is a change (a PR/MR diff), scope the verdict to the change while still scanning mechanically:
55
+
56
+ - Run a hard-coded visual-value scan over the changed paths covering **every governed visual category** — color, background, border/stroke, shadow/elevation, gradient, spacing/padding/margin, radius, size/layout (width/height/gap), typography (font-size/weight/family/line-height), opacity, and motion (transition/animation/duration/easing/transform) (starter regex: `rg -n '#[0-9a-fA-F]{3}|rgb\(|rgba\(|hsl\(|color:|background:|border:|box-shadow|gradient|padding:|margin:|border-radius|font-size|font-weight|line-height|opacity:|width:|height:|gap:|transition|animation|transform:' <changed-paths>`; extend per stack: inline style props, CSS-in-JS literals, imperative theme config). The regex is a recall aid, not the boundary: any unmatched style declaration in a changed hunk still gets read and classified — prefer a stack-aware style/token lint where one exists. Full-audit sweep obligations stay in `multi-project-token-consistency.md`.
57
+ - Classify every hit into exactly one of three buckets: **approved design-system usage** (a token/semantic reference, or an exception carrying the design-system owner's recorded approval — approver, scope, and unexpired validity, covering THIS usage; age is not approval: reusing or extending an old undocumented/expired exception in a changed hunk is a new violation), **pre-existing code outside the requested change**, or **new violation**. Only new violations block the change; pre-existing hits are recorded as debt for the token-consistency audit, never reported as caused by this change.
58
+ - Match the fix duty to the bucket: fix new violations in this change; do not silently expand the change to migrate pre-existing debt (route it), and do not let pre-existing debt normalize new violations ("the file already does this" is not approval — the old code is debt, not a license to copy).
59
+
52
60
  ## UI Checks
53
61
 
54
62
  - Typography hierarchy matches the surface: display for brand/product moments, compact headings for dashboards, readable body text for feed/detail/comment/AI output.
@@ -162,7 +162,7 @@ Tenant commitments around where data lives and who can see it shape the isolatio
162
162
  - **Residency** — "tenant X's data stays in region Y" is a region-per-tenant or region-pinned commitment; the data plane (DB, object storage, backup, analytics) all honor it; the routing layer enforces it.
163
163
  - **Sovereignty** — government / regulated tenants may require a separate stack with no cross-border access; this is a deployment-level isolation, not a runtime knob.
164
164
  - **Encryption** — at-rest encryption per tenant (separate keys per tenant) is a stronger model than a shared key.
165
- - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with conditions; some interpretations of GDPR distinguish anonymization (irreversible) from pseudonymization (key-linkable encrypted data may still be personal data). Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
165
+ - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with explicit conditions — do not use CE when data predates encryption enablement or when keys were backed up/escrowed without verified protection (conditions as stated in Rev.1 §2.6; Rev.2, 2025, supersedes Rev.1 and continues the CE-conditions framework — verify the corresponding Rev.2 section when citing it as the authority); and EDPB guidance (e.g. Guidelines 02/2025) holds that encrypted personal data remains personal data "at least until the algorithm is broken", so key destruction is a conditional control, not automatic GDPR erasure. Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
166
166
  - *(a) key hierarchy diagram* — scoped to this store and this tenant's key version, dated within a defined freshness window (e.g., last 30 days), showing every key that wraps or could reconstruct the data.
167
167
  - *(b) key-backup inventory* — for this store, scoped to this tenant's keys, naming every backup location, rotation policy, and the holder of each backup; dated within the freshness window.
168
168
  - *(c) restore-path test result* — run against *this* backup store with *this* tenant's key-version destroyed, confirming the restore fails because the key is gone. A restore test on a different store or a different key-version is not evidence.
@@ -24,6 +24,8 @@ Sibling note: `go-microservice-dev/references/state-machine-task-patterns.md` ca
24
24
  - Start transitions move pending work to processing before expensive work begins.
25
25
  - Failure transitions capture canonical error code, safe message, retryable flag, retry count, and last trace/log id.
26
26
  - Success transitions persist the result pointer or summary before publishing completion events; completion events are idempotent.
27
+ - All timestamps the transition itself stamps come from a single captured `now` (double clock capture inside one transition produces `finished_at < started_at` records or audit/state disagreement under load); domain-provided times — upstream completion time, event time — are recorded as received, never re-stamped with the local `now`.
28
+ - Validate external/dependency response structure before mapping it (pydantic/schema parse with a typed error path, never a blind attribute access) — a malformed upstream payload must become a failure transition with the canonical error, not an exception mid-transition or a silently-defaulted field.
27
29
 
28
30
  ## Async Processing
29
31
 
@@ -61,7 +61,7 @@ This skill coordinates gates; it does **not** itself authorize merge, tag push,
61
61
  | Reset dev/test-like branches | Yes | Target/env refs, before SHAs, dry-run/plan, force-with-lease semantics |
62
62
  | Post-merge cleanup of the merged temp feature branch (worktree/local/remote) | No — covered by the user's merge authorization (`worktree-isolation` 收尾) | The authorized MR/PR read back as merged at the current head SHA and target; the live remote source ref is absent (already cleaned by the platform) or still equals the merged MR source head (moved → preserve and ask, remote path only — eligible local cleanup proceeds per `worktree-isolation`); no other open or plan-declared MR/PR still consumes the source branch; source branch is a temp feature branch (unclear role → preserve and ask); mechanics/safety rails per `worktree-isolation` |
63
63
 
64
- **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work.
64
+ **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work. **Authorization is never inferred**: a generic instruction ("跑下测试" / "run the pipeline"), a prior run's report, or the mere presence of working credentials/config for a mutating lane does not authorize that lane's mutations — the authorization must name the action category in the current task. Destructive cleanup of test/experiment resources is additionally scope-bound to the resources this run observed itself creating (registry/run-id based), never a name-pattern or global sweep.
65
65
 
66
66
  ## Minimal checklist
67
67
 
@@ -92,10 +92,10 @@ Use this skill to turn observed experience into durable agent skills without cop
92
92
  - Validation gate: a repeated routing miss of the same request type across two rounds is evidence the prior fix only patched the executor surface — re-run this coordinator check before claiming the routing class is closed.
93
93
  - **Routing-trigger fixes must cover the utterance-variant axis, not only the canonical phrasing.** When fixing a routing miss by adding a trigger to a `description`, the same delivery type arrives under many utterances — the canonical name PLUS restart/redo/from-scratch/continue variants (e.g. a *refactor delivery* arrives as "重构 X" but also "重新开发", "完全重新开始", "推倒重来", "清除代码重新开发", "redo/rewrite from scratch"); advertising only the canonical phrase leaves the variants unmatched, so the workflow silently fails to auto-trigger on them. "The rule/trigger exists but the utterance class is uncovered" is a validation-gate defect: enumerate the restart/redo/continue variants of the request type and add the high-value ones (within the 800-char cap; route overflow to the body entry-precedence text). A repeated miss of the SAME delivery type arriving via a different utterance across rounds is the signal this axis was skipped.
94
94
  - **In Claude Code, before fixing a routing miss with a trigger-word edit, check whether the skill's description was even IN the host listing — it may be budget-dropped.** A name-only entry silently voids a keyword-based routing fix (the skill stays reachable by explicit `/skill-name`). **Evidence bar (do not over-apply this as a catch-all):** conclude budget-dropped only from concrete evidence — this turn's listing shows that skill as name-only, `/doctor` output, or other host-listing proof; absent that, treat it as a hypothesis and STILL do the normal trigger / utterance-variant / coordinator fix (this rule does not replace them; a *visible* description misjudged as name-only is the reverse trap). Corollaries: (a) a routing/bootstrap doc assuming "every description is always visible" is wrong under budget pressure and must not be a low-traffic gate skill's sole discovery path; (b) on non-Claude-Code hosts verify the host's own listing behavior from its primary docs — any LLM's "it's a host bug" guess is hypothesis-grade until so checked. The budget mechanism, the cold-start trap, and the remediation levers: `references/skill-listing-budget.md`.
95
- - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings across many turns.
96
- - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it crosses into extraction only once the output is meant to change reusable skill behavior.
97
- - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim`, not landed — reconstruct the charter + target-output map and pass the dual-track before claiming it landed.
98
- - The benchmarked external packs are reference-only: route a missing capability that belongs to the method/tool layer to that pack; only land a ccl-layer rule (domain / governance / cross-cutting principle) here, never a copy of the external skill.
95
+ - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings.
96
+ - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it becomes extraction once the output is meant to change reusable skill behavior.
97
+ - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim` — rebuild charter + target-output map and pass the dual-track before claiming landed.
98
+ - The benchmarked external packs are reference-only: route a missing method/tool-layer capability to that pack; only land a ccl-layer rule (domain/governance/cross-cutting) here, never a copy of the external skill. Verdicts: P/I/M/W, grep-anchored; M needs the functional-equivalent check.
99
99
  - Treat remembered tool, script, installed-skill, repo-root, validator, and executable paths as stale until re-resolved in the current workspace. A path from memory, a prior-round plan, a compacted summary, shell history, handoff notes, or another agent's report must first be re-resolved against the current workspace (reopen the owning skill/reference or inspect the current repository), then probed for existence and executability before running it or reporting it missing. Prefer repo-local or skill-relative scripts only when they resolve inside the loaded skill package or trusted ccl-skills repository root, pass containment and no-symlink/hardlink trust checks, and are not merely same-named scripts in an arbitrary product repo or fork; otherwise fall back to a trusted installed-skill path or manual checklist. If the remembered path fails but a current-context trusted path succeeds, record the failed path source, fallback probe, resolved path class, trust check, and whether shared-file content validation was affected; do not classify shared skill content as broken only because a stale tool path failed.
100
100
 
101
101
  ### Owner-generalization, target-output & impact-chain mapping(owner / 目标映射 / impact-chain)
@@ -123,7 +123,7 @@ Use this skill to turn observed experience into durable agent skills without cop
123
123
  ### What to extract, content placement & domain (UI/UX) judgment(抽什么 / 内容放置 / 领域判断)
124
124
 
125
125
  - Extract behavior, decision rules, quality gates, evidence patterns, and routing boundaries; do not extract business nouns, repo names, IDs, one-off incidents, or stale implementation details.
126
- - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files.
126
+ - Keep the skill entrypoint as the trigger and routing surface; move detailed variants, source-derived patterns, and examples into reference files. Each reference links one level from the entrypoint and stays inside the reference line budget; `references/attention-budget-ratchet.md` owns that budget, the write-side authoring norms, and the design invariants any size/budget gate must satisfy.
127
127
  - A skill must be executable, not only directional. For design, client, testing, debugging, or review skills, include concrete workflow steps, decision points, state/checklist coverage, and verification evidence so future agents do not produce work that is compliant but weak.
128
128
  - Design/client extraction must cover the judgment layer, not only the engineering layer. For UI/UX, extract aesthetic logic, interaction logic, behavioral logic, and user psychology from source evidence before landing rules about layout, components, breakpoints, or tests.
129
129
  - UI/UX judgment extraction must use observable proxies, not adjectives. Read state families, navigation/entry/return paths, disabled reasons, recovery controls, timing/feedback, accessibility, responsive/device variants, and code state machines before claiming behavioral or psychology rules. Use `references/uiux-judgment-extraction.md` for the required method.
@@ -178,9 +178,9 @@ Use this skill to turn observed experience into durable agent skills without cop
178
178
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
179
179
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
180
180
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
181
- - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=2` (three rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled round 4/5. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
181
+ - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=1` (2 rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled new rounds. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
182
182
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
183
- - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At round 3 validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
183
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At budget end validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
184
184
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
185
185
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
186
186
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
@@ -240,7 +240,7 @@ Use this skill to turn observed experience into durable agent skills without cop
240
240
  - For any extraction beyond wording-only cleanup, include the provenance-to-target diff shape before editing: source mechanism, provenance row, target file, executable landing, test or acceptance owner, and status.
241
241
  - Trigger situations and users/tasks it should serve.
242
242
  - What future failure or drift it should prevent, or which evidenced success mechanism it should preserve and reuse.
243
- - For subjective or high-impact skills such as design, UX, frontend/client, product workflow, architecture, or review, define pressure scenarios and acceptance criteria before editing the skill.
243
+ - For NEW skills and subjective/high-impact skills (design, UX, client, product workflow, architecture, review), must define eval/pressure scenarios, baselines, acceptance criteria pre-draft.
244
244
  - For UI/UX or client-facing skills, the pressure scenario must ask whether a person without source access can produce a good-looking and behaviorally sound screen: clear visual hierarchy, fitting density, risk-matched feedback, recoverable state transitions, responsive/device adaptation, and rendered acceptance evidence.
245
245
 
246
246
  3. Inventory evidence.
@@ -262,17 +262,15 @@ Use this skill to turn observed experience into durable agent skills without cop
262
262
  4. Extract candidate rules.
263
263
  - Convert observations into reusable rules: "when X, do Y, verify Z".
264
264
  - Separate invariant rules from stack-specific examples.
265
- - Before editing any stack-specific skill, build the sibling-generalization mini-map: source stack, sibling stacks, shared workflow owner, per-sibling decision, and reason. If a candidate is language-agnostic, route it to the shared workflow/testing/architecture skill first, then add stack-specific implementation notes only where needed.
265
+ - Before editing any stack-specific skill, build the sibling-generalization mini-map the Core Rules owner-generalization group defines; route a language-agnostic candidate to the shared owner first, then add stack-specific implementation notes only where needed.
266
266
  - Mark each candidate as keep, merge, discard, or route to another skill.
267
267
  - Map each kept or routed candidate to its owning target in the target-output map before editing. Do not finish a design/client extraction until implementation, testing, product workflow, and sibling-client implications have been checked and either updated or explicitly marked unchanged. For mini-program surfaces, include `miniapp-product-dev` in that owner check.
268
268
  - Keep an extraction ledger during analysis: rule origin (`observed` or `hypothesis`), source IDs inspected before the rule was drafted, evidence grade, candidate rule, conflict, decision, target skill/reference, and reason.
269
269
 
270
270
  5. Generalize and place content.
271
- - Put trigger, routing, core workflow, and non-negotiable rules in the skill entrypoint.
272
- - Put detailed reference material in direct reference files.
271
+ - Place content by the Core Rules content-placement rule (entrypoint owns trigger, routing, core workflow, and non-negotiables; direct reference files own the detail, within their budget).
273
272
  - Generalize from the evidence ledger, not from a polished rule draft. Do not search for examples to justify a rule that has already been written.
274
- - Choose capability names that describe what future users can do with the skill, such as complex workspace patterns, service architecture, test strategy, or document tightening. Do not name durable outputs after the source file, source project, old feature, or temporary extraction task.
275
- - Keep source identity in provenance fields only. If a source name is useful for audit, label it as source evidence; do not make normal users route through that source name to understand or trigger the skill.
273
+ - Name durable outputs per the Core Rules naming rule (complex workspace patterns, service architecture, test strategy, document tightening are capability names; the source file, source project, old feature, and this extraction task are not), and do not make normal users route through a source name to understand or trigger the skill.
276
274
  - Turn source lessons into execution recipes: analyze first, implement with ownership boundaries, debug by layer, test at the right level, and verify on the real rendered/runtime surface.
277
275
  - For design and client skills, place the four judgment layers explicitly (aesthetics / interaction logic / behavioral logic / psychology — per-layer semantics and decision fields: `references/uiux-judgment-extraction.md`); do not hide them inside generic "UI polish" wording.
278
276
  - Write each judgment layer's delta per the Core Rules judgment-delta rule (new / confirmed / narrowed / routed / no new evidence — a restatement of existing principles is not newly extracted knowledge). For visual direction/tokens, use the token provenance fields in `references/uiux-judgment-extraction.md` and state what they mean for design, implementation, and testing; otherwise it is only a static source note.
@@ -290,8 +288,7 @@ Use this skill to turn observed experience into durable agent skills without cop
290
288
  - **Cross-section facet-ownership check** (Core Rules ↔ Step 0–6): for any edit touching `## Core Rules` or a Workflow step, record whether the changed facet is rule/invariant-owned (→ Core Rules) or procedure/checklist-owned (→ the step), and confirm the opposite surface only *points* to it, not restates it. Same-facet text living in both surfaces is a drift defect — converge toward the canonical surface per the Start-here boundary contract before landing.
291
289
  - All referenced files exist and are one level from the skill entrypoint.
292
290
  - No business or source-repo leakage remains in executable guidance.
293
- - Capability naming is source-neutral: old source artifact names, old page names, and old scenario labels are absent from executable guidance or clearly marked as provenance.
294
- - For any rename or generalization, residual searches for the old file name, old source label, old English shorthand, and old capability phrase are clean or justified as provenance.
291
+ - Capability naming is source-neutral, and after any rename or generalization residual searches for the old source artifact name, file name, source label, page name, scenario label, English shorthand, and capability phrase are clean, absent from executable guidance, or clearly marked as provenance.
295
292
  - Trigger boundaries do not collide with sibling skills.
296
293
  - Coverage matrix has no unexplained gaps: each relevant source category is used, routed, discarded, or marked unavailable with a reason.
297
294
  - For broad or multi-skill extraction, validation must include the durable source register and a target-output map showing which target skills or outputs were updated, which sources informed them, which sources were excluded, and which source classes remain pending. If any required row is pending, the work can be landed only as an interim checkpoint, not as complete.
@@ -0,0 +1,37 @@
1
+ # Attention-budget ratchet — design invariants for size/budget gates
2
+
3
+ An attention-budget gate limits how much prose an agent must hold to use a surface: the entrypoint body-word/byte gate, the every-session-injection byte gate, the `description` 800-char cap, and the reference-file line gate (all enforced by `scripts/check-size-budget.sh` or the canonical validator). This file owns the design invariants those gates share and the write-side norms for reference files. Companions: `rule-consolidation.md` owns the prose doctrine (merge-into-canonical, why rule sets must not grow monotonically); `skill-listing-budget.md` owns the host's listing-budget mechanism; `description-authoring.md` owns the description surface.
4
+
5
+ ## The five ratchet invariants
6
+
7
+ Any NEW budget/size gate, and any modification to an existing one, is checked against all five before landing (this is the budget-gate instantiation of the design-time operability check in `dual-track-review-gate.md` — run that check's four legs too). A gate missing one of these fails in a predictable way, named per item:
8
+
9
+ 1. **Stable proxy estimator** — the metric is deterministic and environment-independent: Unicode letter/number word runs with Han counted per ideograph, raw byte size, or physical line count. Never a model/tokenizer-dependent estimate: two environments disagreeing on the measure turns the gate into noise, and a changed estimator silently invalidates every recorded allowance. Deterministic also means encoding-normalized before measuring — a line count taken over raw bytes reads a CR-delimited file as one line, so line endings are folded to LF first.
10
+ 2. **Anti-false-green sentinel** — a run that could not evaluate says so: base-unresolvable prints an `*_unevaluated` token (never the ok token), probe failures fail closed as partials, and on any block the last token is the failure marker. The ok token must be unearnable by losing the base; a consumer grepping for ok must never read an un-run gate as a pass.
11
+ 3. **Zero tolerance for new debt** — a NEW surface over budget blocks outright. There is no exempt marker, no waiver flag, and no way for a candidate to nominate its own baseline; structural exclusions live in the gate, owned by the gate.
12
+ 4. **Legacy may only shrink** — an existing over-budget surface is frozen at its base measure: level or shrinking lands, any growth blocks. Rename credit is path-paired and non-growing (move plus growth blocks as growth). This is what makes a uniform cap deployable over a corpus that already exceeds it, without a rewrite round and without rewarding a rush to pre-shrink.
13
+ 5. **A missing baseline is never a pass** — the comparison base comes from revision history (`CCL_SKILL_BASE_REF`, upstream, or merge-base), so there is no stored manifest to go stale; when no base resolves, the verdict is unevaluated (invariant 2), and reddening that state is the caller's pipeline decision (CI always exports the base ref).
14
+
15
+ Two cross-cutting corollaries:
16
+
17
+ - Average headroom must never fund a single over-budget surface: the ratchet judges each file alone, and corpus-level counters stay visibility-only.
18
+ - Debt counters and advisory bands are never a clean-landing waiver nor authorization to keep growing a surface; only the delta verdict blocks.
19
+
20
+ ## Reference-file write-side norms
21
+
22
+ The read side already defends against oversized files (chunked reads under ~200 lines, references one level deep). These norms are the write side, enforced as a delta ratchet over `skills/*/references/**/*.md`:
23
+
24
+ - A NEW reference file over 500 physical lines must not land — split it by subtopic before landing (the gate blocks new-or-crossing files; 500 exactly passes).
25
+ - An existing over-limit reference is frozen per invariant 4: shrink or stay level; growth blocks. Additions to a frozen reference are funded by consolidating existing text in the same file.
26
+ - Append-only ledgers are structurally excluded: `references/source-register.md` grows by contract (append-only, supersede-by-pointer, rows never edited), so a line cap would block the ledger discipline itself; the gate skips it and prints a visibility token when it is over the figure. Residual risk, accepted under the same trusted-contributor model as the entrypoint gate: a prose file named `source-register.md` would dodge the cap — review owns that shape.
27
+ - A new reference over 100 lines must be structured with `##` sections so chunked reads and greps can navigate it; a heading-less long file draws an advisory token (never a block). A table-of-contents list is optional — section structure is the invariant, not a TOC block.
28
+ - Authoring anti-patterns (verified against the official skill-authoring checklist, see verdicts below): time-sensitive facts outside an explicit old-patterns section; inconsistent terminology for one concept; abstract examples where a concrete input/output pair fits; Windows-style paths; unexplained constants; scripts that defer error handling to the model instead of solving it.
29
+
30
+ ## Official-clause verdicts (provenance)
31
+
32
+ Registered claims were re-verified against the primary source (Anthropic "Skill authoring best practices", docs.claude.com, read 2026-08-31) before landing; per-clause disposition:
33
+
34
+ - "Keep SKILL.md body under 500 lines" — present verbatim, but it scopes to SKILL.md, not references. Already covered more strictly here by the 5000-body-word delta ratchet. The 500-LINE reference cap above is a repo-internal norm motivated by the read-side chunking evidence, and is labeled as such — never cite it as an official requirement.
35
+ - Table-of-contents mandate for long references — NOT PRESENT in the current official text. The official remedy for `head -100` partial reads is keeping references one level deep (already a Step 6 validation rule). The `##`-section advisory above rests on repo-internal evidence only.
36
+ - "Tested with Haiku, Sonnet, and Opus" — present, conditional on the models you plan to ship to. This repo's skills inherit the session model and the eval layer exercises real sessions, so no multi-model matrix is mechanized; the clause fires only if the repo starts shipping model-pinned skills.
37
+ - Checklist anti-patterns (time-sensitive info, terminology, concrete examples, forward-slash paths, voodoo constants, scripts-solve-not-defer) — present; landed above as authoring norms. The recurring-anti-patterns grep panel is NOT their landing surface: its admission rule requires a class observed in 2+ skills of this repo.
@@ -68,6 +68,15 @@ Bad: `Proactively invoke when the user shares a draft doc / spec / plan and asks
68
68
 
69
69
  Good: `Skip when the ask is already scoped to one stack (e.g. "fix this React render bug" → web-react-dev; "add a GORM index" → go-microservice-dev; "调下这个按钮间距" → product-ui-ux-design), or when the user is reporting a defect / failure / regression → defect-diagnosis owns reproduction and root cause first.`
70
70
 
71
+ ## Body routing pointers: the quadruple
72
+
73
+ The description is not the only routing surface an author writes: skill/reference BODY text routes too, through cross-skill and cross-reference pointers ("route to X", "read Y before Z"). Tier-1 static analysis parses only the description, so body pointers are held to an authoring contract instead:
74
+
75
+ - **Every cross-skill or cross-reference routing pointer in body text must carry the routing quadruple**: trigger (when to go read the other surface), scope (which file or small subset), output (what decision/artifact to extract), and return point (where to resume in the owning workflow). A pointer whose parts are obvious from sentence position may state them compactly ("at step 3, read X's §Y for the Z decision, then continue step 4").
76
+ - A bare "refer to X if useful" / "参见 X" with no trigger and no extraction target is never a landing shape: such pointers rarely fire, and when they do fire they read too much (source-observed: unbounded pointers were the dominant dead-routing shape in an adopting skill pack; the pack that enforced the quadruple had two verified adopters and no dead pointers).
77
+ - The description-side Skip-when `→ skill` idiom already satisfies the quadruple (trigger = the skip condition, scope = the target skill's entry, output = ownership transfer, return = none) — no extra wording needed there.
78
+ - Detector pairing: a pointer that routes but is never read shows up as the **silent skip** failure mode in `eval-routing.md`'s failure-mode vocabulary (B-side probes / Tier-3 traces), not in Tier-1/2 — fix the pointer's trigger and scope, not the description.
79
+
71
80
  ## Precedence against session-injected process skills
72
81
 
73
82
  A routing / workflow skill competes not only with sibling skills but also with process-discipline skills that some hosts INJECT at session start with very forceful language (e.g. a brainstorming skill whose description says "MUST use before creating features / adding functionality"). At initial routing time the router primarily sees skill names, descriptions, and host/session rules — the workflow body has not loaded yet, so an "Entry precedence" paragraph in the body does NOT win the routing decision. If your routing skill should own the entry point for a request class that an injected process skill also claims, the description must:
@@ -81,6 +90,10 @@ For teammates using the shared skill on description-based routing hosts, this is
81
90
 
82
91
  A quoted trigger phrase should belong to its skill in at least 80% of real-use contexts. If a phrase would commonly mean something else, it must be DROPPED or ANCHORED.
83
92
 
93
+ **Activation is closer to keyword match than semantic match** — a public sandbox measurement (2025-2026; sources/locators in the specs ledger) found prompts containing a skill's name or a distinctive description token activate near-100% while conceptual paraphrases activate near-0%; a second independent eval corroborates the keyword-dependence and adds that routing accuracy degrades as the installed-skill count nears ~20 similar skills, recovering when consolidated to ~12.
94
+
95
+ - Both findings are host/model/catalog-conditional: treat them as directional and do not rely on the numbers without reproducing against your own catalog. Two consequences for authoring: (a) the description must contain the distinctive tokens users actually type (measure real utterances, don't invent vocabulary — the discovery-vocabulary rule); (b) when routing degrades across the catalog, merging/pruning similar skills beats adding more trigger words to each.
96
+
84
97
  ### Drop (too generic, no rescue possible)
85
98
 
86
99
  - `"改下样式"` — almost always means "change CSS now", which is implementation, not design ownership. Drop.
@@ -67,7 +67,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
67
67
 
68
68
  Across a long operational-rule/code extraction series the challenge supplies *the same handful of axes* as the recurring P0/P1 — first drafts systematically nail the functional / cost / happy-path and omit a predictable set. The per-axis instance lists for the six first-draft blind-spot axes:
69
69
 
70
- - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox.
70
+ - **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox; and — because skill/reference text is itself a prompt agents execute — the rollout-safety screen an eval harness runs on skill text before release: wording that induces context/secret exfiltration (prompt leakage), assumes or grants permissions beyond the task (overreach), or automates a destructive or confirmation-skipping step (unsafe automation) — text a draft must not carry except as an explicitly labelled anti-example.
71
71
  - **(2) concurrency & lifecycle** — races, deadlock (e.g. holding a lock through a drain/callback), use-after-close/free, resurrection after delete, double-free/double-close, cleanup ordering, at-most-once/fires-once.
72
72
  - **(3) resource bounds** — an unbounded default/timeout/buffer/retry, a leaked registry/goroutine/task entry, a missing max backstop.
73
73
  - **(4) rollout / migration ordering** — a step that breaks not-yet-upgraded consumers, or abandons a live bug to do the clean refactor first.
@@ -82,7 +82,7 @@ Three per-edit instances of the axes above that recur because their canonical ru
82
82
 
83
83
  For any new mechanical gate, validator, or evidence apparatus — **and for any change that makes an existing one's verdict stricter** — run the four legs at design time, not after challenge rounds force them. These four legs all fire at design time; the check has a second firing point they do not cover, because its actor is not the gate's author: when a landing **withdraws or downgrades the evidentiary claim an existing gate rests on**, that gate is re-based or retired in the same landing — see the claim-liveness rule in `product-rd-workflow/references/design-review-gate-mechanics.md`, which owns it, including the obligation walk a retirement owes.
84
84
 
85
- - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under the SAME base resolution CI uses, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep).
85
+ - **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under **every base resolution CI uses — enumerate the landing faces, never assume one**, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep). A base-relative verdict is a quantity *relative to its base*, and CI resolves that base per event, so this leg is discharged by a **recorded manifest, never by a walk you assert**: derive the branch set from the workflow itself — its `push` branch filter plus every branch a pull request is actually landed into — and record one entry per branch carrying the resolved base ref and this gate's verdict on the candidate, so a reviewer can diff the manifest against the workflow and `every` stops being author-adjudicated. Green on the round's own base says nothing about the others: the difference is every commit that landed on the shared branch before the gate existed, and it surfaces as one accumulated violation on the face nobody measured — typically the promotion PR, long after the authors who could have funded it moved on. **A set that cannot be enumerated is a blocking residual, not a pass** — an unrestricted `pull_request` trigger with no fixed target set has no finite manifest, so take leg (d)'s non-blocking or risk-owner-deferral exit and name the faces left unmeasured rather than claiming coverage. Enumerating the faces costs a loop; the exemption you will reach for instead is the loosening leg (e) forbids.
86
86
  - **(b) marginal-cost statement** — record what the cheapest routine change costs under the gate (recompute/regenerate/rerun burden); a gate whose per-iteration cost defeats normal development gets lightened or redesigned at design time.
87
87
  - **(c) trust-model fit** — name what the mechanism defends against under its DECLARED trust model; machinery that only defends against adversaries the trust model already excludes (e.g. content digests where the author can regenerate every hash) buys redundant detection at full complexity cost — prefer the lighter mechanism that keeps the enforceable core.
88
88
  - **(e) loosening check — an exemption must name the class of change that stops owing evidence, and that class must be one with no behaviour to evidence.** Fires whenever a change makes a gate accept what it used to reject: a new exemption class, a waived requirement, a widened accept set. A loosening is easier to get wrong than a tightening and shows up later, because it produces no red for anyone to notice — the gate simply stops asking. Two obligations, both outcomes rather than procedures. **First, state the exempted class in behavioural terms and check it against the repo's own definition of that class**: an exemption named `not-required` asserts *no behaviour*, so if any rule in the tree already classifies that same diff shape as behaviour-changing, the exemption contradicts it and the answer is to fix the *anchor/evidence form* for that shape, never to drop the evidence. **Second, if the class does carry behaviour, the exemption must be replaced by a way to SUPPLY the evidence** — widen where the anchor may land, add an evidence form the shape can satisfy — because the class with the most behaviour is exactly the one an exemption hurts most. **A precedent of the same shape is not a justification**: reaching for an existing exemption class because the root cause rhymes with an earlier one transfers the solution without checking the disanalogy, and the disanalogy is usually the load-bearing part. Failure shape: a gate anchor that structurally cannot bind to a frontmatter-only change was answered with a third `not-required` class by analogy to two existing ones, even though the same repository elsewhere states that any frontmatter edit is a routing-surface change and the author had just measured its routing delta; the adversarial review caught it on the first finding, and the correct fix was to let the anchor bind to the changed description instead.
@@ -230,7 +230,7 @@ Review + challenge check the change for defects; neither checks whether it actua
230
230
  | Status | Use when |
231
231
  |---|---|
232
232
  | `RED-baseline` | the change alters behavior or routing (trigger / scope / routing / validation / acceptance). Run the scenario WITHOUT the change first (baseline failure), then WITH it (compliance) |
233
- | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance, and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
233
+ | `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance — an author cannot self-assert it — and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
234
234
  | `not-applicable: docs-only` | the change touches NO file under `skills/**` and no skill-loaded guidance — i.e. `README` / `ARCHITECTURE` / `CONTRIBUTING` / `docs/**` only. Forbidden for `SKILL.md`, `references/**`, validators, templates, examples, and the plugin-shipped command/behavior surfaces (`hooks/*`, `scripts/install.sh`, `bin/`, `.mcp.json`, `.lsp.json`, `monitors/**`, `settings.json`): those are behavioral or executable source even when they read like prose |
235
235
 
236
236
  For a `RED-baseline`, the evidence form can be a before-after task diff, a golden trace, or a pressure scenario — these are *how* you show baseline→compliance, not standalone substitutes for it. For a routing-surface / hub-skill change you can run the golden-trace form as a REAL headless-agent run via the F4 Tier-3 harness (optional, higher-fidelity than a recorded scenario — see `validation-and-landing.md` Behavioral Validation); a recorded scenario is not by itself a failing baseline.
@@ -241,7 +241,6 @@ For a `RED-baseline`, the evidence form can be a before-after task diff, a golde
241
241
 
242
242
  Rules:
243
243
 
244
- - `RED-baseline` is required whenever the change alters behavior or routing. `semantic-control` is valid ONLY with reviewer confirmation that none of trigger/scope/routing/validation/acceptance changed — an author cannot self-assert it.
245
244
  - The row must give a concrete locator + evidence shape, not a bare status: artifact path / commit / transcript / command, the exact prompt or scenario, and expected-vs-actual. For `RED-baseline`, record BOTH the without-change (baseline failure) and with-change (compliance) results, and name the baseline's **provenance type**: a *recorded incident* (cite where the failure is actually recorded — transcript, note, issue; the cited record must describe a failure that occurred, not prescribe a method) or a *constructed scenario* (run against BOTH the unchanged baseline and the changed rule, with an openable artifact for each run — a scenario "run" only mentally, only against the patched text, or only as reviewer discussion does not count). A prescriptive source — a method-bar or best-practice note with no failure recorded — cannot be cited as an occurred failure and is not by itself a valid `RED-baseline`; it may seed the constructed scenario's design or support the rule's rationale, but the RED evidence is the run artifact. Narrating what "would have" failed as if it happened is a fabricated evidence row, the same defect class as fabricated verification output. A status word with no openable artifact is not a valid row, same as a missing review row.
246
245
  - A missing or unreviewable behavioral-evidence row blocks landing the same way a missing review row does; until it exists the work is an uncommitted interim checkpoint.
247
246
 
@@ -300,18 +299,13 @@ Primary reviewer failure is a remediation branch only when the owning gate class
300
299
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
301
300
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
302
301
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
303
- 5. A fallback result satisfies only the exact lane it ran. If no candidate returns a conclusive verdict, keep the work `interim`.
304
-
305
- Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
302
+ 5. A fallback result satisfies only the exact lane it ran, and only when it meets the same evidence bar — do not call a manual ad-hoc run "fallback review". Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence. If no candidate returns a conclusive verdict, keep the work `interim`.
306
303
 
307
304
  - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
308
305
 
309
306
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
310
307
 
311
- The following raw CLI shapes are **debugging diagnostics only**. They may help
312
- isolate an owner-wrapper failure, but neither output is review/challenge evidence
313
- and neither may replace `scripts/extraction_review_gate.sh` for a non-wording
314
- lane.
308
+ The raw CLI shapes below are **debugging diagnostics only** — they can help isolate an owner-wrapper failure, but are never review/challenge evidence and never replace `scripts/extraction_review_gate.sh` on a non-wording lane.
315
309
 
316
310
  Debug a wrapper with raw `codex review` against the target diff:
317
311
 
@@ -416,6 +410,15 @@ item 9 signs off on it.
416
410
 
417
411
  Use `--json` to capture reasoning traces and tool calls cleanly. Parse the JSONL stream with a small Python or jq script as documented in `gstack-codex` skill.
418
412
 
413
+ ## Reviewer verification scope (packet-verifiability boundary)
414
+
415
+ The external reviewer judges what the bounded packet can show; the packet structurally cannot carry every deterministic oracle its acceptance claims depend on (frozen preservation-mapping rows, the checker's complete pinned-literal sets, whole-file postimages, immutable pre-fix revisions). The division of labor is fixed and documented here so it is ruled on once, not re-litigated per round:
416
+
417
+ - The reviewer owns CONTENT SEMANTICS: wording coherence, dropped qualifiers/obligations, source-accuracy, sanitization, scope drift — everything decidable from the packet plus the reviewer's own reasoning.
418
+ - Deterministic-gate claims (pinned literals present, size ratchet net-zero, obligation audit green, R0 clean, parity) are verified by the repository's CI re-running those gates on the actual branch — never by the reviewer, and never accepted from the implementer's prose alone; a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round.
419
+ - Historical-process claims (a pre-fix RED, a measurement taken before landing) are session-record-grade unless bound to an immutable revision or a candidate-bound receipt; treat them as the implementer's testimony, and say so in the disposition instead of demanding evidence the packet cannot hold.
420
+ - The receipt channel for the third bullet is landed: `scripts/gate_receipt.py mint --out <ledger-dir>/gate-receipt-N.json -- <gate command>` binds one deterministic-gate run to the committed candidate — clean tree required, HEAD commit recorded, argv + repository-relative cwd + exit code + output SHA-256 + RFC3339 mint time, with an off-by-default opt-in output tail (a verbatim excerpt would copy a leaked token into the ledger; receipts are created 0600) and a post-run candidate re-read that refuses a result when HEAD moved mid-run; a RED run mints fine (the pre-fix RED is the canonical use). Receipts live OUTSIDE the candidate tree (the chain ledger directory) and ride the packet by reference: name the receipt file plus its own SHA-256 in a review-plan `evidence` entry or a disposition-evidence `evidence` item, so `review_context_sha256` / the v3 ledger's sibling-hash discipline bind it. Anyone re-checks with `gate_receipt.py verify <file>` (structural) or `verify <file> --rerun -- <the gate command you expect>` at the recorded commit — the verifier types the command and the tool compares it against the recorded argv before executing only the verifier's own words, because a receipt is untrusted input and its recorded argv must never be executed as-is (differential exit-code + output-hash comparison; `--exit-only` for a legitimately nondeterministic gate, and say so in the referencing row). Trust model, stated so it is not oversold: a receipt is candidate-bound, falsifiable consistency evidence — it does not authenticate who ran the command; deterministic authority stays with CI re-running the gates on the actual branch (second bullet), and a process claim carrying no receipt remains testimony under the third bullet.
421
+
419
422
  ## Recording findings + fixes
420
423
 
421
424
  For each pass, record in the extraction's working file (e.g. `<project>-extraction-summary.md`):
@@ -429,7 +432,7 @@ For each pass, record in the extraction's working file (e.g. `<project>-extracti
429
432
 
430
433
  ## Challenge pass (codex exec adversarial)
431
434
  - Findings: N total (a P0 / b P1 / c P2)
432
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
435
+ - R0 evidence: <same value menu as the review-pass row above>
433
436
  - Gate-fireability applicability: <yes — change adds/edits a semantic rule/gate/status/verdict | no — valid ONLY when the diff is wording-only or adds/edits no semantic rule/gate/status/verdict>
434
437
  - Item 9 exercised: <locator to the captured prompt/transcript/JSONL showing the bypass-by-omission probe actually ran (not a pasted self-assertion) | n/a per line above>
435
438
  - Applied: M fixes (commit: <sha>)
@@ -515,33 +518,38 @@ A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`
515
518
 
516
519
  - When the recurring surface is a **self-adjudication clause** — decidable test: the clause's classification verb has NO named test whose output produces the classification, so the receiver/author judges it — the `keep / delete / narrow / replace` decision must first ask whether an existing mechanical or semi-mechanical test can carry the adjudication: name the test whose output settles the classification, **and the mapping from its output to the classes** — which output means which class — plus the actual result or the evidence contract that will produce it. Naming a test is not routing to it: a test whose output cannot discriminate the classes, or one named with no recorded output-to-class mapping, leaves the adjudication exactly where it was and does not satisfy this rule. A `keep` that retains self-adjudication prose, or a `narrow`/`replace` that adds more prose bindings, is landed only when the decision record (the same-class rule's recorded decision, in the commit body or register row) names the reason no existing test could carry it — a reason left in chat does not count; two challenge rounds attacking the same self-adjudication surface are the signal that prose is the wrong layer — a clause routed to an existing test inherits that test's evidence bar instead of the adjudicator's say-so (worked instance: a review-reception clause that left "is this finding scope-adding" to the receiver survived two rounds of attacks on that self-classification until the classification was routed to the existing structural-minimality test). The same question fires at drafting time for any new reception/discipline-style clause that would grant a self-adjudication.
517
520
 
518
- "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*, not *no new P0/P1 appeared*.
521
+ "No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*. Do NOT iterate to zero *findings* either — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing, and forcing the count to zero either over-corrects or scope-creeps; the bar is zero *undispositioned* P0/P1, which differs from both *zero findings* and *no new P0/P1 appeared*.
519
522
 
520
523
  **A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
521
524
 
522
- Do NOT iterate to zero *findings* — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing. Forcing the finding count to zero either over-corrects or scope-creeps. The bar is zero *undispositioned P0/P1*, which is different from zero findings.
523
-
524
525
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
525
526
 
526
527
  Five is the generic `code-review` transport ceiling, not this extraction lane's
527
528
  spend. Non-wording Agent-autonomous extraction calls go through
528
- `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=2`: the
529
- initial review plus at most two challenges. At round 3 the autonomous lane ends.
529
+ `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=1`: the
530
+ initial review plus at most one challenge. At round 2 the autonomous lane ends.
530
531
  An authenticated human may request later review, but that is separately
531
- attributed human-requested evidence outside this chain/budget, not an Agent
532
- round 4 or 5. Unused generic capacity never authorizes automatic continuation.
532
+ attributed human-requested evidence outside this chain/budget, never an
533
+ additional Agent round. Unused generic capacity never authorizes automatic continuation.
533
534
  The v3 closeout validator rejects referenced receipts whose recorded budget is
534
- not 2 and checks budget and ordering consistency within the caller-supplied
535
+ not the wrapper-fixed value and checks budget and ordering consistency within the caller-supplied
535
536
  set. It cannot authenticate that the wrapper produced those receipts or that
536
537
  the caller retained every earlier chain or receipt. The wrapper does not mint or
537
538
  persist
538
539
  `review_chain_id` or `autonomous_review_index`: the caller still supplies both,
539
- and could start a fresh-looking chain after round 3. The validator detects bad
540
+ and could start a fresh-looking chain after the final round. The validator detects bad
540
541
  order inside the referenced set but cannot detect a prior chain the caller
541
542
  omitted, so complete caller-owned ledger retention—and treating an Agent reset
542
543
  as a contract violation—remains part of the boundary rather than a property the
543
544
  local scripts prove.
544
545
 
546
+ **Self-hosted chains break on every fix; the budget is summed across chains, never per chain.** In a skill repository the candidate edits its own owner package by construction, so the chain's stable bindings make the dead-end the norm, not an edge case: the selected-owner digest hashes each owner package's current working tree and owners derive from the candidate's own paths, so a fix that touches any selected-owner tree ends the tracked chain (`review_chain_invalid`) — in an extraction round that is nearly every fix, while a fix confined to files outside every selected owner drifts only the candidate hash and continues in-chain — and a plan edit that changes the normalized review scope (intent, acceptance, stage/depth, risk tags, budget) ends it as `review_scope_changed` — a self-review- or evidence-only plan refresh keeps the scope digest and the chain (binding mechanics are owned by the staged review contract in `code-review`). A chain restarted at index 1 after such a break spends the SAME Agent-autonomous budget. Treating each restarted chain as a procedurally required fresh review loop is the observed way the budget hollows out: two consecutive extraction rounds ran 20+ reviewer rounds and then 12 restarted chains — 21 reviewer invocations to land a three-line diff — each restart looking locally mandatory. When a round returns findings, walk this enumeration before any further external call:
547
+
548
+ 1. **Batch dispositions; never re-chain per finding — and hold fixes until the ledger can afford their application.** Triage every finding through the disposition bar and deep-self-review once, then decide when the batch lands by budget arithmetic, never by the urge to fix now: applying any fix to a selected-owner tree ends the tracked chain, so apply-now is legal only when the remaining ledger can still fund a restarted chain's ready floor. **Under the wrapper-fixed 1+1 budget that condition is NEVER true after the review round — the one remaining round cannot fund review plus challenge — so there the rule is unconditional: accumulate every fix unapplied, run the challenge on the frozen, unchanged round-1 candidate, and land the whole batch only after the full review+challenge chain has run, MR/PR-listed.** A fix applied between review and challenge breaks the chain (the challenge binds to the round-1 candidate), forfeits the double-receipt terminal, and costs a fresh human-authorized chain to recover — an observed failure, not a hypothetical: a round that landed its review fixes before the challenge had to be closed by a user-granted continuation chain.
549
+ 2. **Sum spent rounds across all chains before opening one more.** Count every prior external round in the caller-retained ledger — every chain, finished or broken — against the lane's wrapper-fixed budget; at the cap a restarted chain must not be opened autonomously, and below it budget the restart so the final chain can still hold review plus one challenge (the closeout ready floor) — a restarted chain opens with a fresh review by contract, so every restart trades a challenge round for a review round. Effective exhaustion is reached when the remaining rounds cannot fund that floor for any continuation; treat it exactly like the cap.
550
+ 3. **Front-load packet quality in chain 1.** The first chain's packet must already be the full-context diff (`--unified` wide enough to carry whole files, e.g. `-U200`) with the plan frozen alongside the candidate; narrow packets breed packet-boundary pseudo-findings whose fixes break chains and burn rounds on artifacts of the packet itself.
551
+ 4. **At the cap — or at effective exhaustion — the designed terminal is disposition, never another chain.** Apply or disposition the final batch, name every post-review fix in the MR/PR description, record the honest terminal state (`continuation_authorization_required` when the lane's final round ran and itself returned findings; otherwise — including a chain broken before its challenge could run — an interim record naming the last externally reviewed candidate and every later delta), and hand continuation or merge to the human. The post-review batch sits only on the pending MR/PR branch beside that record — the human's authenticated continuation, waiver, or merge decision is what certifies it, and it is never reported as reviewed. This is the bounded outcome working as designed, so do not report it as convergence and do not launder it through a fresh-looking chain.
552
+
545
553
  A strictly proven wording-only change has no convergence loop: it uses one
546
554
  generic `code-review` pass, records the independent-review row and the
547
555
  challenge-not-required proof, and does not create a schema-v3 multi-round
@@ -554,7 +562,7 @@ This budget limits only automatic reviewer invocation. It does **not** stop impl
554
562
  - A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
555
563
  - A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
556
564
  - A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
557
- - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. One dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. The recovery is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
565
+ - A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. The dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. While cross-chain budget remains, that break is handled autonomously by the self-hosted-chain rule's ledger-counted restart; it becomes this bullet's human-decision dead-end when the remaining budget cannot fund the re-review. The recovery at that point is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
558
566
 
559
567
  When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
560
568
 
@@ -562,7 +570,7 @@ The mechanical reminder is `self_review_gate`, not prose alone. It records outst
562
570
 
563
571
  In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
564
572
 
565
- At the third Agent-autonomous round, do not start a fourth automatically. If findings remain:
573
+ At the final Agent-autonomous round, do not start another automatically. If findings remain:
566
574
 
567
575
  - keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`;
568
576
  - record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
@@ -577,14 +585,14 @@ ledger and every referenced controller, completion, base, and sweep file live
577
585
  in one directory and carry SHA-256s. The validator walks the ordered schema-v3
578
586
  controller chain (same chain and scope,
579
587
  review then contiguous challenges, packet=candidate, complete prior-result hash
580
- prefix, fixed `challenge_budget=2`) and binds every closeout candidate to its
588
+ prefix, the wrapper-fixed `challenge_budget`) and binds every closeout candidate to its
581
589
  last receipt. Ready requires at least review + challenge; a second base drift may
582
590
  stop as race immediately after round 1 rather than spending an illegal challenge
583
591
  after the terminal predicate already fired.
584
592
  It ends in exactly one state:
585
593
 
586
594
  - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
587
- - `continuation_authorization_required`: round 3 itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
595
+ - `continuation_authorization_required`: the final round itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
588
596
  - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
589
597
 
590
598
  Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
@@ -599,17 +607,16 @@ convergence.
599
607
 
600
608
  For a focused single-skill change:
601
609
  - **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
602
- - **Round 2 — challenge 1**: after implementer triage, attack the highest-risk unresolved surface with an unprimed prompt.
603
- - **Round 3 — challenge 2**: verify remaining/new attack paths. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic round 4.
610
+ - **Round 2 — challenge**: after implementer triage — fixes stay HELD: under the 1+1 budget applying any fix before this round always breaks the chain, so the challenge runs on the frozen round-1 candidate (self-hosted-chain rule; enumeration item 1 above) — attack the highest-risk unresolved surface with an unprimed prompt. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic further round, and the post-review fix batch lands MR/PR-listed.
604
611
 
605
- Broad extractions use the same three-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
612
+ Broad extractions use the same fixed two-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
606
613
 
607
614
  ### Anti-patterns
608
615
 
609
616
  - **Single-round challenge → done**. The round-1 fix-up itself may introduce bugs. Always do at least one re-challenge after a non-trivial fix-up.
610
617
  - **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
611
618
  - **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
612
- - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one.
619
+ - **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one. The observed extreme is chain multiplication: a chain broken by your own fix and restarted at index 1 is the same budget, not a new loop — sum rounds across chains per the self-hosted-chain rule above.
613
620
  - **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
614
621
 
615
622
  ### Recording the loop
@@ -620,7 +627,7 @@ Add one row per round to the validation log:
620
627
  ## Challenge pass — round N (codex exec adversarial)
621
628
  - Diff scope: <files / commit range / sha>
622
629
  - Findings: N total (a P0 / b P1 / c P2)
623
- - R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
630
+ - R0 evidence: <same value menu as the review-pass row above>
624
631
  - Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
625
632
  - New since prior round: <count> (subset of above; flag round-introduced bugs)
626
633
  - Stabilized: <list of findings carried over without change>