@ccoalm/ccl-skills 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (127) hide show
  1. package/README.md +2 -2
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +8 -7
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +6 -1
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +16 -17
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +1 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +195 -7
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +3 -3
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +13 -5
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +9 -3
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +9 -3
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/normalize_review_timeout.sh +22 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +9 -3
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +1540 -129
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +8 -3
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +76 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +1858 -3
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +789 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/update_review_plan_intent.py +513 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +4 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -1
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +11 -10
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +64 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml +4 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +73 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/runtime-and-project-contract.md +58 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +41 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/verification-diagnostics-and-security.md +63 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +14 -16
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +10 -14
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +1 -1
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +135 -86
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +66 -80
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/delivery-contract.md +275 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +88 -214
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +2 -2
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +10 -8
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +6 -5
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +112 -95
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +30 -21
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +22 -3
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +20 -17
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +7 -9
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +14 -10
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +2 -0
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +3 -3
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +9 -6
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +3 -0
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +37 -10
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +8 -1
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +16 -5
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +16 -5
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +4 -2
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +4 -1
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +4 -0
  77. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +104 -5
  78. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +11 -9
  79. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +102 -0
  80. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +103 -0
  81. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +20 -0
  82. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +6 -6
  83. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +4 -3
  84. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +69 -2
  85. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +22 -0
  86. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +49 -4
  87. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +2748 -0
  88. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +20 -5
  89. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py +1142 -0
  90. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +17 -0
  91. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +41 -4
  92. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +120 -0
  93. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +82 -8
  94. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +336 -0
  95. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +82 -10
  96. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +1416 -0
  97. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh +57 -0
  98. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +141 -4
  99. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +3 -1
  100. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh +1696 -0
  101. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2117 -0
  102. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +316 -0
  103. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +1176 -0
  104. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +31 -1
  105. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +9 -4
  106. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +980 -0
  107. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +8 -6
  108. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  109. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  110. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  111. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +16 -15
  112. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  113. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +10 -2
  114. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  115. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  116. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  117. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  118. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  119. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  120. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  121. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +7 -5
  122. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +1 -1
  123. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  124. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  125. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  126. package/dist/assets/release.json +215 -105
  127. package/package.json +1 -1
@@ -4,6 +4,8 @@ Use this reference when the task is to review, QA, compare, or improve UI/UX in
4
4
 
5
5
  This is not a domain-compliance audit. For finance, healthcare, legal, or regulated products, use this as the UI/UX layer and add the appropriate domain rules separately.
6
6
 
7
+ Audit findings are criteria/evidence inputs to `delivery-contract.md`, not an independent acceptance path. A runtime-visible audit names every affected React/other-Web, H5, native, mini-app, terminal/CLI/TUI, Electron/desktop/TV, other-client, and composite-host layer; ready/complete still requires the complete design/test/producer/client binding set, Test Phase 1, and an allowed design verdict.
8
+
7
9
  ## Evidence Sources
8
10
 
9
11
  Use these sources in this order:
@@ -12,10 +14,10 @@ Use these sources in this order:
12
14
  2. Relevant source Figma file or design-system file.
13
15
  3. `interaction-design-patterns.md` for flow, feedback, state, gesture, and trust-sensitive behavior.
14
16
  4. `visual-craft.md` for brand feel, hierarchy, anti-slop, motion, and polish.
15
- 5. `tokens-and-components.md`, `platform-mobile-patterns.md`, and `platform-web-desktop-patterns.md` for tokens and component semantics.
17
+ 5. `tokens-and-components.md`, the applicable platform lens, and every affected client owner's convention for tokens, component/command semantics, host behavior, and adaptation.
16
18
  6. `frontend-code-evidence-map.md` for local code evidence classification and reusable behavior patterns, never as product-domain requirements.
17
19
 
18
- Reusable capability observations already extracted:
20
+ Reusable Mobile/desktop capability observations already extracted (illustrative, not a closed platform set):
19
21
 
20
22
  - Mobile design system: Button, List, Card, Image, ImageViewer, NoticeBar, FloatingPanel, Dialog, Empty, ErrorBlock, Modal, Progress, Result, Skeleton, SwipeAction, Toast, NavBar, Popup, TabBar, SafeArea, ImageUploader, PasscodeInput, and Example Pages.
21
23
  - Desktop design system: Button, Layout, Splitter, Menu, Dropdown, Steps, Form, Input, Select, Upload, Card, Empty, List, Tag, Tooltip, Alert, Drawer, Message, Modal, Notification, Progress, Result, Skeleton, and Table/Tabs marked `【todo】` as weak guidance only.
@@ -36,16 +38,24 @@ Always include concrete file/line references for code reviews and Figma file/pag
36
38
 
37
39
  ## Audit Procedure
38
40
 
39
- 1. **Classify the surface**: mobile consumer, web consumer, web operational, AI workspace, shared component, onboarding, settings, trust/safety, or analytics.
41
+ 1. **Classify every rendered layer**: React/other Web, H5, native mobile/host, mini-app, ordinary CLI or terminal/TUI, Electron/desktop/TV, other client, or composite host; then name the consumer/operational/AI/shared/onboarding/settings/trust/analytics task shape.
40
42
  2. **Name the primary task**: what the user must be able to do in one sentence.
41
43
  3. **Trace the flow**: entry, context, action, feedback, recovery, return.
42
44
  4. **Map required states**: happy, first-use, empty, loading, partial, error, retry, permission, disabled, success, undo/cancel, long-content, and responsive states.
43
45
  5. **Compare UI primitives**: check whether the implementation uses the closest existing component and token semantics instead of one-off UI.
44
46
  6. **Check UX clarity**: hierarchy, action priority, copy, affordance, consequence, source/provenance, and next action.
45
- 7. **Check accessibility basics**: readable contrast, keyboard/focus where relevant, touch target size, visible labels, reduced-motion risk, safe-area/keyboard behavior on mobile.
47
+ 7. **Check accessibility basics**: readable contrast, labels, keyboard/focus or touch/input semantics where relevant, reduced-motion/capability fallback, and the affected owner's host-specific accessibility/adaptation behavior.
46
48
  8. **Check visual craft**: anti-slop, product-level identity, spacing rhythm, typography scale, consistent iconography, appropriate density.
47
49
  9. **Check serious-domain adaptation** when relevant: source, timestamp, partial data, confirmation, audit labels, and no unsafe optimistic UI.
48
- 10. **Check platform-convention conformance** for iOS/Android app surfaces: run the pass/fail criteria in `external-ui-ux-quality-benchmarks.md` Platform Convention Walkthrough (HIG/Material state completeness, platform accessibility minima, and platform conventions) against the rendered surface.
50
+ 10. **Check platform-convention conformance** on the named target: use `external-ui-ux-quality-benchmarks.md` to classify authority and boundary, recheck the current first-party platform source, and map applicable criteria into `delivery-contract.md`. Preserve requirement versus recommendation strength and verify on that platform's rendered runtime; do not reuse a combined HIG/Material checklist as a cross-platform standard.
51
+
52
+ ## Diff-Scoped Review
53
+
54
+ When the audit target is a change (a PR/MR diff), scope the verdict to the change while still scanning mechanically:
55
+
56
+ - Run a hard-coded visual-value scan over the changed paths covering **every governed visual category** — color, background, border/stroke, shadow/elevation, gradient, spacing/padding/margin, radius, size/layout (width/height/gap), typography (font-size/weight/family/line-height), opacity, and motion (transition/animation/duration/easing/transform) (starter regex: `rg -n '#[0-9a-fA-F]{3}|rgb\(|rgba\(|hsl\(|color:|background:|border:|box-shadow|gradient|padding:|margin:|border-radius|font-size|font-weight|line-height|opacity:|width:|height:|gap:|transition|animation|transform:' <changed-paths>`; extend per stack: inline style props, CSS-in-JS literals, imperative theme config). The regex is a recall aid, not the boundary: any unmatched style declaration in a changed hunk still gets read and classified — prefer a stack-aware style/token lint where one exists. Full-audit sweep obligations stay in `multi-project-token-consistency.md`.
57
+ - Classify every hit into exactly one of three buckets: **approved design-system usage** (a token/semantic reference, or an exception carrying the design-system owner's recorded approval — approver, scope, and unexpired validity, covering THIS usage; age is not approval: reusing or extending an old undocumented/expired exception in a changed hunk is a new violation), **pre-existing code outside the requested change**, or **new violation**. Only new violations block the change; pre-existing hits are recorded as debt for the token-consistency audit, never reported as caused by this change.
58
+ - Match the fix duty to the bucket: fix new violations in this change; do not silently expand the change to migrate pre-existing debt (route it), and do not let pre-existing debt normalize new violations ("the file already does this" is not approval — the old code is debt, not a license to copy).
49
59
 
50
60
  ## UI Checks
51
61
 
@@ -58,6 +68,7 @@ Always include concrete file/line references for code reviews and Figma file/pag
58
68
  - Data-heavy screens keep scan lines stable: sticky headers, aligned controls, consistent row/card heights, and visible active filters.
59
69
  - Mobile screens respect safe area, bottom actions, keyboard visibility, and one-handed reach.
60
70
  - Web screens collapse secondary panels before damaging primary content readability.
71
+ - Mini-app, ordinary CLI or terminal/TUI, other-Web, Electron/desktop/TV, and composite-host screens apply the actual owner-specific host, input, geometry, fallback, bridge, and recovery checks.
61
72
 
62
73
  ## UX Checks
63
74
 
@@ -1,6 +1,6 @@
1
1
  # UI/UX Design And Development
2
2
 
3
- Use this reference when the task is to design a UI/UX surface and implement it in frontend code. It bridges Figma-derived design rules with reusable implementation patterns observed in mobile and web frontend sources.
3
+ Use this reference when the task is to design a UI/UX surface and implement it in client code. It is a primitive catalog and translation guide; `delivery-contract.md` is authoritative for the design brief, testing selection, producer/client returns, evidence semantics, and design verdict.
4
4
 
5
5
  This reference is for new-product UI/UX execution. Do not copy education-domain workflows, wording, assets, product assumptions, or information architecture from source artifacts.
6
6
 
@@ -13,6 +13,8 @@ Before coding, define:
13
13
  - **Density**: consumer relaxed, productive compact, or hybrid.
14
14
  - **Source pattern**: which Figma/code pattern is being reused and what domain details are discarded.
15
15
  - **State set**: use the canonical taxonomy in `product-surface-patterns.md`, then add implementation-specific loading, retry, cancellation, permission, long-content, and responsive behavior.
16
+ - **Client owner set**: follow every affected rendered layer in `delivery-contract.md`. React web → `web-react-dev`; Vue/Svelte/static/vendor/other web → its installed web-content owner or fail-closed project-convention lookup; native mobile/host → `app-cross-platform-dev`; mini-app → `miniapp-product-dev`; terminal/CLI/TUI → `terminal-cli-dev`; Electron/desktop/TV shell → its installed owner or the same project-convention lookup. Composite hosts keep separate content and shell members.
17
+ - **Producer owner**: every changed or claim-bearing backend, config, content, or inference source that supplies a rendered value or behavior; record its exact artifact/version identity before client execution.
16
18
 
17
19
  Do not start from a decorative layout. Start from the user job, interaction loop, and required states.
18
20
 
@@ -25,8 +27,8 @@ Do not start from a decorative layout. Start from the user job, interaction loop
25
27
  5. **Implement state model**: represent loading/error/empty/partial/success/permission explicitly in data and UI.
26
28
  6. **Implement responsive behavior**: mobile safe area and keyboard behavior; web secondary panel collapse and min/max widths.
27
29
  7. **Implement feedback**: inline validation, toast/message, alert/notice, drawer/sheet, modal/dialog, result page.
28
- 8. **Verify by screenshot**: check realistic happy, empty, loading, error, long-content, and narrow-width screenshots before claiming done.
29
- 9. **Verify against UI/UX audit**: run the checks in `ui-ux-audit.md` before claiming done.
30
+ 8. **Return execution evidence**: have every changed or claim-bearing producer and every affected client return its own immutable record, then have the test owner bind its definition and execution records to the exact producer/client versions exercised. Keep commands, artifacts, criterion results, dimensions, coverage boundary, and gaps in their owning records as required by `delivery-contract.md`; a screenshot proves only its captured state.
31
+ 9. **Record the design verdict**: run the relevant checks in `ui-ux-audit.md`, evaluate every criterion, and bind the design record plus `candidate`, `accepted`, `rejected`, or `pending` to the complete design/test/producer/client candidate-binding set.
30
32
 
31
33
  ## Mobile Frontend Patterns
32
34
 
@@ -134,6 +136,8 @@ Avoid silent catches and console-only errors for user-triggered actions.
134
136
 
135
137
  ## Responsive Acceptance
136
138
 
139
+ Treat the Mobile and Web checks below as stack-specific examples. Build the complete affected client-owner set from `delivery-contract.md`; add mini-app host/device, ordinary CLI or terminal/TUI, Electron/desktop/TV shell, other-Web, and composite-host evidence whenever those layers are affected.
140
+
137
141
  Mobile:
138
142
 
139
143
  - Safe top/bottom areas are respected.
@@ -150,6 +154,12 @@ Web:
150
154
  - Sidebar collapsed mode remains discoverable.
151
155
  - Tables/lists preserve row identity and selected/filter state.
152
156
 
157
+ Other affected clients:
158
+
159
+ - Mini-app evidence covers the shipped host/tool, supported device class, safe area, permissions/capabilities, package/platform constraints, and embedded web-view bridge when present.
160
+ - Ordinary CLI or terminal/TUI evidence covers command/help/default/exit/recovery semantics, TTY and non-TTY/plain modes as applicable, width/capability/color fallback, and interactive lifecycle only where used.
161
+ - Electron/desktop/TV and other-Web evidence comes from the actual content owner plus shell owner or fail-closed project convention, including supported sizes/scaling, input/focus, bridge, and content-shell integration.
162
+
153
163
  Screenshot acceptance:
154
164
 
155
165
  - First viewport shows the primary workflow, not a decorative banner or empty dead area.
@@ -159,13 +169,13 @@ Screenshot acceptance:
159
169
 
160
170
  ## Development Review Checklist
161
171
 
162
- Before finishing UI/UX implementation, verify:
172
+ Use this list as client-side criteria before returning the client record. It cannot by itself finish the slice; completion requires bound design and test records, every changed producer and affected client return, Test Phase 1 sufficiency, and an allowed design verdict under `delivery-contract.md`.
163
173
 
164
174
  - The code uses local primitives and tokens before custom markup/styles.
165
175
  - The flow maps to `discover -> inspect -> act -> confirm -> return`.
166
176
  - Canonical states from `product-surface-patterns.md` are implemented, not just documented.
167
177
  - Feedback strength follows `interaction-design-patterns.md`.
168
- - Mobile safe-area/keyboard and web responsive behavior are covered.
178
+ - Every affected client's adaptation contract is covered; Mobile safe-area/keyboard and Web responsive behavior are examples, not the closed set.
169
179
  - Global feedback providers, async wrappers, and route/workspace state are mounted at the shell level when multiple feature pages rely on them.
170
180
  - Long-running uploads, imports, downloads, generation, or review jobs remain visible after route changes and have retry/fail/complete states.
171
181
  - Designed states have a code owner: shell/provider state, route state, feature state, server task state, or local draft state. Do not leave a designed state as static markup with no data transition.
@@ -174,3 +184,4 @@ Before finishing UI/UX implementation, verify:
174
184
  - Visual polish passes `visual-craft.md`.
175
185
  - Screenshot acceptance passes `layout-recipes-and-screenshot-acceptance.md`.
176
186
  - UI/UX review passes `ui-ux-audit.md`.
187
+ - The client return names the exact producer member/version exercised and contributes its immutable member to the complete design/test/producer/client binding set.
@@ -102,10 +102,12 @@ If a pattern appears in a Figma source, preserve it only when it has a clear pro
102
102
 
103
103
  ## Acceptance Questions
104
104
 
105
- Before calling frontend visual work polished, ask:
105
+ These questions contribute visual-craft criteria; they cannot mark a runtime slice ready or complete. Evaluate them on every affected rendered layer and bind the resulting evidence through the complete design/test/producer/client set and design verdict in `delivery-contract.md`.
106
+
107
+ Before calling client visual work polished, ask:
106
108
 
107
109
  - Does the screen have a clear product-level visual point of view?
108
110
  - Does it avoid generic AI frontend patterns?
109
111
  - Does the visual direction support the target product loop instead of distracting from it?
110
112
  - Are typography, color, spacing, radius, motion, and background choices tied to existing tokens or an explicit product reason?
111
- - Does the screen remain readable, accessible, and performant on mobile and desktop?
113
+ - Does the surface remain readable, accessible, and performant across the supported sizes, host modes, input/capability modes, and adaptation matrix of every affected client—not only Mobile and desktop Web?
@@ -162,7 +162,7 @@ Tenant commitments around where data lives and who can see it shape the isolatio
162
162
  - **Residency** — "tenant X's data stays in region Y" is a region-per-tenant or region-pinned commitment; the data plane (DB, object storage, backup, analytics) all honor it; the routing layer enforces it.
163
163
  - **Sovereignty** — government / regulated tenants may require a separate stack with no cross-border access; this is a deployment-level isolation, not a runtime knob.
164
164
  - **Encryption** — at-rest encryption per tenant (separate keys per tenant) is a stronger model than a shared key.
165
- - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with conditions; some interpretations of GDPR distinguish anonymization (irreversible) from pseudonymization (key-linkable encrypted data may still be personal data). Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
165
+ - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with explicit conditions — do not use CE when data predates encryption enablement or when keys were backed up/escrowed without verified protection (conditions as stated in Rev.1 §2.6; Rev.2, 2025, supersedes Rev.1 and continues the CE-conditions framework — verify the corresponding Rev.2 section when citing it as the authority); and EDPB guidance (e.g. Guidelines 02/2025) holds that encrypted personal data remains personal data "at least until the algorithm is broken", so key destruction is a conditional control, not automatic GDPR erasure. Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
166
166
  - *(a) key hierarchy diagram* — scoped to this store and this tenant's key version, dated within a defined freshness window (e.g., last 30 days), showing every key that wraps or could reconstruct the data.
167
167
  - *(b) key-backup inventory* — for this store, scoped to this tenant's keys, naming every backup location, rotation policy, and the holder of each backup; dated within the freshness window.
168
168
  - *(c) restore-path test result* — run against *this* backup store with *this* tenant's key-version destroyed, confirming the restore fails because the key is gone. A restore test on a different store or a different key-version is not evidence.
@@ -19,7 +19,10 @@ Use this for implementation of Python backend products, services, microservices,
19
19
  - Use `go-microservice-dev` for Go services. Do not load Go implementation rules for Python work unless the task is explicitly cross-language contract or generated-client integration.
20
20
  - Use codebase-specific skills only when the task is explicitly about an existing repository.
21
21
  - For money, billing, quota, permission, tenant/user data isolation, high-impact AI, repeated writes, async finality, or incident-explanation risk, apply `product-rd-workflow` high-risk resilience gates and route test-layer design through `testing-strategy`.
22
- - When a change edits strings, templates, or config values that are returned to, persisted for, emitted to, served to, synchronized with, or configured for client consumption (error copy, labels, notification text, localization payloads, content/CMS/seed rows, message or notification templates, flag-delivered content), classify the consumers with a recorded bounded check (client repo / contract / locale search) before closing on API/log evidence; if any client surface renders the value user-facing, or consumers are unknown, load `product-ui-ux-design` and record its implementation-owner checkpoint — including the consuming client stack owner(s) and `testing-strategy` per that checkpoint's field list — with client-side rendered-evidence routing. Backend-only closure without that recorded consumer check is invalid.
22
+ - When a change can alter what a client renders or which state, action, or decision path it offers—including strings/templates/config/flags and API/event/schema fields, enums, status/progress, permission/capability signals, defaults, or result shapes—load `../product-ui-ux-design/references/delivery-contract.md`, create the applicable full or lightweight record in that contract, and follow its canonical consumer-universe classification, design/test/client handoffs, and terminal-status rules.
23
+ - This Python owner returns only its `producer_record` delta: immutable binding, build/schema/config artifact identity, exact command/environment, and API/event/log/output observation.
24
+
25
+ - For a standalone Python CLI, this skill owns Python parser/library implementation mechanics. Any change to a user-facing command tree, subcommand, flag/default/action path, help/output/exit behavior, confirmation, progress, or recovery path also loads `terminal-cli-dev`, which owns the terminal contract and its UI/UX/testing handoff. Only internal parser refactors proven to preserve all user-visible semantics may skip that owner.
23
26
 
24
27
  ## Generalization Discipline
25
28
 
@@ -24,6 +24,8 @@ Sibling note: `go-microservice-dev/references/state-machine-task-patterns.md` ca
24
24
  - Start transitions move pending work to processing before expensive work begins.
25
25
  - Failure transitions capture canonical error code, safe message, retryable flag, retry count, and last trace/log id.
26
26
  - Success transitions persist the result pointer or summary before publishing completion events; completion events are idempotent.
27
+ - All timestamps the transition itself stamps come from a single captured `now` (double clock capture inside one transition produces `finished_at < started_at` records or audit/state disagreement under load); domain-provided times — upstream completion time, event time — are recorded as received, never re-stamped with the local `now`.
28
+ - Validate external/dependency response structure before mapping it (pydantic/schema parse with a typed error path, never a blind attribute access) — a malformed upstream payload must become a failure transition with the canonical error, not an exception mid-transition or a silently-defaulted field.
27
29
 
28
30
  ## Async Processing
29
31
 
@@ -61,7 +61,7 @@ This skill coordinates gates; it does **not** itself authorize merge, tag push,
61
61
  | Reset dev/test-like branches | Yes | Target/env refs, before SHAs, dry-run/plan, force-with-lease semantics |
62
62
  | Post-merge cleanup of the merged temp feature branch (worktree/local/remote) | No — covered by the user's merge authorization (`worktree-isolation` 收尾) | The authorized MR/PR read back as merged at the current head SHA and target; the live remote source ref is absent (already cleaned by the platform) or still equals the merged MR source head (moved → preserve and ask, remote path only — eligible local cleanup proceeds per `worktree-isolation`); no other open or plan-declared MR/PR still consumes the source branch; source branch is a temp feature branch (unclear role → preserve and ask); mechanics/safety rails per `worktree-isolation` |
63
63
 
64
- **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work.
64
+ **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work. **Authorization is never inferred**: a generic instruction ("跑下测试" / "run the pipeline"), a prior run's report, or the mere presence of working credentials/config for a mutating lane does not authorize that lane's mutations — the authorization must name the action category in the current task. Destructive cleanup of test/experiment resources is additionally scope-bound to the resources this run observed itself creating (registry/run-id based), never a name-pattern or global sweep.
65
65
 
66
66
  ## Minimal checklist
67
67
 
@@ -92,10 +92,10 @@ Use this skill to turn observed experience into durable agent skills without cop
92
92
  - Validation gate: a repeated routing miss of the same request type across two rounds is evidence the prior fix only patched the executor surface — re-run this coordinator check before claiming the routing class is closed.
93
93
  - **Routing-trigger fixes must cover the utterance-variant axis, not only the canonical phrasing.** When fixing a routing miss by adding a trigger to a `description`, the same delivery type arrives under many utterances — the canonical name PLUS restart/redo/from-scratch/continue variants (e.g. a *refactor delivery* arrives as "重构 X" but also "重新开发", "完全重新开始", "推倒重来", "清除代码重新开发", "redo/rewrite from scratch"); advertising only the canonical phrase leaves the variants unmatched, so the workflow silently fails to auto-trigger on them. "The rule/trigger exists but the utterance class is uncovered" is a validation-gate defect: enumerate the restart/redo/continue variants of the request type and add the high-value ones (within the 800-char cap; route overflow to the body entry-precedence text). A repeated miss of the SAME delivery type arriving via a different utterance across rounds is the signal this axis was skipped.
94
94
  - **In Claude Code, before fixing a routing miss with a trigger-word edit, check whether the skill's description was even IN the host listing — it may be budget-dropped.** A name-only entry silently voids a keyword-based routing fix (the skill stays reachable by explicit `/skill-name`). **Evidence bar (do not over-apply this as a catch-all):** conclude budget-dropped only from concrete evidence — this turn's listing shows that skill as name-only, `/doctor` output, or other host-listing proof; absent that, treat it as a hypothesis and STILL do the normal trigger / utterance-variant / coordinator fix (this rule does not replace them; a *visible* description misjudged as name-only is the reverse trap). Corollaries: (a) a routing/bootstrap doc assuming "every description is always visible" is wrong under budget pressure and must not be a low-traffic gate skill's sole discovery path; (b) on non-Claude-Code hosts verify the host's own listing behavior from its primary docs — any LLM's "it's a host bug" guess is hypothesis-grade until so checked. The budget mechanism, the cold-start trap, and the remediation levers: `references/skill-listing-budget.md`.
95
- - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings across many turns.
96
- - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it crosses into extraction only once the output is meant to change reusable skill behavior.
97
- - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim`, not landed — reconstruct the charter + target-output map and pass the dual-track before claiming it landed.
98
- - The benchmarked external packs are reference-only: route a missing capability that belongs to the method/tool layer to that pack; only land a ccl-layer rule (domain / governance / cross-cutting principle) here, never a copy of the external skill.
95
+ - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings.
96
+ - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it becomes extraction once the output is meant to change reusable skill behavior.
97
+ - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim` — rebuild charter + target-output map and pass the dual-track before claiming landed.
98
+ - The benchmarked external packs are reference-only: route a missing method/tool-layer capability to that pack; only land a ccl-layer rule (domain/governance/cross-cutting) here, never a copy of the external skill. Verdicts: P/I/M/W, grep-anchored; M needs the functional-equivalent check.
99
99
  - Treat remembered tool, script, installed-skill, repo-root, validator, and executable paths as stale until re-resolved in the current workspace. A path from memory, a prior-round plan, a compacted summary, shell history, handoff notes, or another agent's report must first be re-resolved against the current workspace (reopen the owning skill/reference or inspect the current repository), then probed for existence and executability before running it or reporting it missing. Prefer repo-local or skill-relative scripts only when they resolve inside the loaded skill package or trusted ccl-skills repository root, pass containment and no-symlink/hardlink trust checks, and are not merely same-named scripts in an arbitrary product repo or fork; otherwise fall back to a trusted installed-skill path or manual checklist. If the remembered path fails but a current-context trusted path succeeds, record the failed path source, fallback probe, resolved path class, trust check, and whether shared-file content validation was affected; do not classify shared skill content as broken only because a stale tool path failed.
100
100
 
101
101
  ### Owner-generalization, target-output & impact-chain mapping(owner / 目标映射 / impact-chain)
@@ -111,10 +111,10 @@ Use this skill to turn observed experience into durable agent skills without cop
111
111
  - Method: scan the current session's available-skills list before building the owner-generalization map; for each lifecycle stage, name the external-skill candidate alongside the ccl-skill candidate; route to the external skill when it owns the operational recipe and keep the CCL skill as the gate-keeper / cross-cutting rule layer.
112
112
  - When external packages are absent in a teammate's environment, the CCL skill's principle wording must stand alone (no broken `superpowers:*` / `gstack-*` references in executable guidance) — name them as "if installed, route to X; otherwise apply the principle inline".
113
113
  - When a user correction or self-check exposes one missed extraction dimension, sibling owner, or lifecycle stage, scan the immediate neighbors on the same axis before landing the fix. Reuse existing machinery: target-output map for lifecycle, sibling-generalization mini-map for stack/owner, and the source type's dimension enumeration for judgment axes. Do not invent new axes per task and do not walk beyond immediate neighbors. Land only the smallest needed updates and record one line per neighbor as `update`, `unchanged`, or `routed`. The trigger is a discovered miss, not every extraction.
114
- - **Class-wide routing/trigger changes require COMPLETE-set coverage in one pass — the immediate-neighbor scan above does NOT apply.** The neighbor scan covers a *single discovered miss*; a change whose rule is "every skill of class C should advertise / scope / carry X" (e.g. "every stack `*-dev`/`*-architecture` should advertise a localized-refactor trigger") is out of its scope.
114
+ - **Class-wide routing/trigger changes require COMPLETE-set coverage in one pass — the immediate-neighbor scan above does NOT apply.** The neighbor scan covers a *single discovered miss*; a change whose rule is "every skill of class C should advertise / scope / carry X" is out of its scope.
115
115
  - For such a change, first write an **explicit, narrow, source-backed class predicate** (from how the user/source phrased the class — not an expansive inference from examples; if the predicate is ambiguous or huge, downscope or ask, and record non-member exclusions).
116
116
  - Then the COMPLETE set matching that predicate is the required coverage: enumerate the installed members from the session available-skills list PLUS any referenced repo-present CCL members (absent ones get `install-drift: pending` per the next rule, never silent omission), and update or explicitly mark each `unchanged`/`routed` in ONE landing before claiming the class closed.
117
- - Landing one member-pair (e.g. Python+Go) and waiting for the user to surface each remaining member across rounds is the defect this prevents; a repeated same-class miss across rounds is the signal.
117
+ - Landing one member-pair and waiting for the user to surface each remaining member across rounds is the defect this prevents; a repeated same-class miss across rounds is the signal; member-JOIN duty: `references/source-to-skill-extraction.md#member-join-inheritance`.
118
118
  - Validation gate: the closeout map lists every predicate-matching member with a status, or the class is not closed.
119
119
  - (Ordinary single-skill edits still use owner/sibling checks, not a full-class sweep.)
120
120
  - **A referenced ccl-owned/vendored skill not installed in a host is install-drift, not a valid `not-applicable`.** When enumeration reaches a skill that exists in the canonical CCL repo (or is vendored here) and is referenced by the tree but not installed in a host, do NOT mark it `not-applicable: not installed` and move on — that records a symptom as a reason and the skill silently never fires there. Record `install-drift: pending`, surface the exact remediation, and install it ONLY via an approved user/maintainer instruction or the managed install script — do NOT silently create host symlinks or install unprompted (a host mutation changes future routing globally and can point at the wrong checkout). Until installed, that member's coverage stays interim, not omitted. This applies ONLY to ccl-owned/vendored skills; an absent *external/system* package routes to maintainer/upstream, never local install/edit.
@@ -178,9 +178,9 @@ Use this skill to turn observed experience into durable agent skills without cop
178
178
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
179
179
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
180
180
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
181
- - The Agent-autonomous review budget is the initial review plus at most four challenges, five external rounds total; it never limits deep self-review, implementation, tests, or an authenticated human request. `self_review_gate` mechanically fires before external review, after findings/candidate/scope changes, at the post-budget checkpoint, and before an Agent completion claim. Its `blocks` are narrow: another external review and/or the Agent completion claim, not productive work or an authenticated human merge authorization. Candidate-controlled input cannot assert either human decision.
181
+ - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=2` (three rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled round 4/5. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
182
182
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
183
- - If a required review/challenge row or its required pass is absent, skipped, inconclusive, or unavailable — a recorded `review unavailable after remediation` (or equivalent non-success) row does NOT satisfy the gate — block the Agent's completion/commit claim, run/remediate the review lane when Agent budget remains, or use an approved alternate independent reviewer with the same bounded-scope, attribution, timeout, and output-validity requirements. When Agent budget is exhausted or every lane remains inconclusive, report an `interim` checkpoint, continue self-review/implementation/tests and independent runnable work, and park only decision-dependent work. Only an authenticated human may waive the review-process gate or stop the overall iteration; neither action is inferred from a local file or Agent statement.
183
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At round 3 validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
184
184
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
185
185
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
186
186
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
@@ -81,6 +81,10 @@ For teammates using the shared skill on description-based routing hosts, this is
81
81
 
82
82
  A quoted trigger phrase should belong to its skill in at least 80% of real-use contexts. If a phrase would commonly mean something else, it must be DROPPED or ANCHORED.
83
83
 
84
+ **Activation is closer to keyword match than semantic match** — a public sandbox measurement (2025-2026; sources/locators in the specs ledger) found prompts containing a skill's name or a distinctive description token activate near-100% while conceptual paraphrases activate near-0%; a second independent eval corroborates the keyword-dependence and adds that routing accuracy degrades as the installed-skill count nears ~20 similar skills, recovering when consolidated to ~12.
85
+
86
+ - Both findings are host/model/catalog-conditional: treat them as directional and do not rely on the numbers without reproducing against your own catalog. Two consequences for authoring: (a) the description must contain the distinctive tokens users actually type (measure real utterances, don't invent vocabulary — the discovery-vocabulary rule); (b) when routing degrades across the catalog, merging/pruning similar skills beats adding more trigger words to each.
87
+
84
88
  ### Drop (too generic, no rescue possible)
85
89
 
86
90
  - `"改下样式"` — almost always means "change CSS now", which is implementation, not design ownership. Drop.
@@ -59,6 +59,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
59
59
  - **Independent oracle.** Where the property has no executable test — rule text, a register row, a doc — the enumeration is discharged only by an **independent oracle**: name the concrete observation that would contradict the property, say where that observation lives (the owner file and line, the primary source, the command whose output would differ), and go look. **Validate the oracle before trusting its verdict**: a check that returns "clean" because it looked in the wrong place, matched case-sensitively, used too narrow a pattern, or swallowed an error is indistinguishable from a passing property, and it fails in the dangerous direction. Before accepting a clean result you must PROVE THE CHECK CAN FAIL — point it at something you know is broken and watch it report that. Confirming it enumerated the inputs you meant is a necessary extra step, never a substitute: correct inputs say nothing about whether the predicate detects a mismatch or whether a non-zero exit was swallowed, so a check that can only ever say clean passes that weaker test. An unvalidated oracle is not weaker evidence than an imagined mutation; it is the same thing wearing a command prompt.
60
60
  - **Dimension walk.** Adding cases inside an axis you already had buys nothing against one you did not: the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering) before values — `testing-strategy` owns that list and the precision-row obligation that goes with it. A walk whose rows are all imagined mutations is exhortation wearing a checklist's clothes.
61
61
  - **Re-owe after fixes.** Whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration — that newly-added mechanism is the most dangerous line in the diff, because it has no test yet and you wrote it with your attention on the defect it repairs. "The whole enumeration" includes the **pre-cover axes sweep** (concurrency & lifecycle above all): remediation text written mid-round re-owes the draft-time axes BEFORE the candidate goes back to the reviewer, because a fix written with attention on one defect systematically re-opens the same blind-spot axes the original draft missed. And the loop has an escalation point: when the same blind-spot axis or finding class supplies findings in a **third** round, stop the per-finding loop and run one full-matrix implementer self-enumeration (the artifact's own states × failure points × orderings × residues × cross-references) on the current candidate before any further external round — letting the reviewer surface one hole per round is the reviewer-as-defect-finder failure at its most expensive (observed shape: a multi-round program burned twenty-plus single-finding rounds on one axis family; the one full-lifecycle enumeration, run at the maintainer's correction, found the remaining holes in a single batch).
62
+ - **Classify before fixing.** Persist one transition row per finding class: a stable semantic class key, its root-cause predicate, affected surface, and one disposition per occurrence. Each occurrence names the SHA-256 of its controller receipt plus the canonical JSON SHA-256 of a finding that actually appears in that receipt; the closeout must classify every controller finding exactly once. New wording or a new file is not a new class when the predicate is the same. Root-cause predicates must be unique across classes after case-folding and whitespace collapse; **only that exact normalization is mechanical**. It catches cosmetic case/whitespace key splits but does not decide whether differently worded predicates are semantically equivalent, which remains reviewer-contestable. Resolution is per occurrence, not "the last disposition wins": every `fixed`, `accepted_tradeoff`, `pre_existing_out_of_scope`, or `source_refuted` occurrence names one same-directory structured disposition-evidence file plus its SHA-256. That file binds schema version, exact current candidate, its controller receipt and finding, the disposition, a non-empty evidence list, and the ordered current/prior class occurrences it resolves. It must resolve its own occurrence; a later closing disposition that lists only itself leaves an earlier `open` occurrence unresolved until later evidence explicitly includes that exact receipt/finding pair. A `needs_human_decision` occurrence cannot be resolved by any candidate-local disposition evidence and stays unresolved until an authenticated external human/platform decision takes the separately authorized path. A controller finding disproved by first-hand source or failure-path evidence records `source_refuted`; this is a classification of an invalid finding, not a fourth disposition for a valid P0/P1. The validator binds every JSON input using duplicate-key rejection, plus the evidence file, digest, candidate, occurrence, disposition, and transition links; it does not judge the evidence text or authenticate tradeoff/scope acceptance. The reviewer or human decision-maker still owns those semantic and authority verdicts, and `ready_for_human_decision` is not approval or merge authority. On the class's third appearance, stop patching individual instances and enumerate the complete authoritative surface against that predicate. The sweep names a same-directory manifest plus its SHA-256; its candidate and exact ordered `searched_set` must match the row, and its unmatched list supplies the recorded count. `ready_for_human_decision` requires zero unresolved occurrences and zero unmatched instances; `continuation_authorization_required` and `baseline_race` retain non-zero unmatched evidence instead of lying about closure. Classification is reviewer-contestable evidence, not an author-controlled escape hatch; splitting one predicate into cosmetic sub-classes does not reset the count. `scripts/validate_extraction_review_state.py` enforces these bindings within the referenced receipt/evidence set.
62
63
  - **Graded verdict shape.** When the assessed reality is multi-dimensional or partial (capability, coverage, feasibility, quality, completion), collapsing it into one binary verdict — "done/not-done", "possible/impossible", "all correct/all wrong" — is the over-broad-absolute axis applied to your own claim layer: the swing to whichever pole feels safest to assert misrepresents a distribution, and the opposite-pole absolute ("structurally impossible", "nothing works") is the SAME defect as an unearned "done", not a humbler one. Report per-dimension status — what's strong, what's weak, what wasn't checked — with the confidence each part actually earned; and where a binary gate genuinely applies (a pass/fail check, a blocked/allowed decision), still give the clear top-line verdict after the per-dimension basis — calibration is not hedged mush. A user correcting your answers as too absolute ("每次都很绝对") is this defect's recurrence signal, same escalation as the `SKILL.md` rule states.
63
64
  - **Honesty (descriptive, not permissive).** This is recognition-dependent salience, not a mechanical gate — an agent that doesn't notice it is done-claiming cannot self-fire it; the mechanical backstops remain the closeout `interim` gates + user-signal escalation. "I didn't notice I was claiming done" does NOT waive the rule — any non-trivial completion/coverage/convergence wording must carry clean-pass evidence or an explicit interim/downscope disposition *before* you emit it. The rule targets completion/coverage/convergence assertions on work whose failure a check could catch, and never narrows the mandatory dual-track challenge (it is the always-on generalization of *self-audit to convergence*, not a replacement for the gate).
64
65
 
@@ -204,6 +205,16 @@ Only the challenge pass may be skipped, and only when this table marks challenge
204
205
  - **(a) Bounded change class.** The edit changes ONLY typo, grammar, formatting, or a meaning-preserving synonym, and changes NO trigger, scope, routing, validation, acceptance, rule/threshold/boundary text, or any other meaning. Reference/body prose that states a rule, threshold, boundary, rubric, or applies/does-not-apply line IS a semantic surface — editing it is NOT wording-only unless the change is purely typo/grammar/formatting with no meaning change. A synonym substitution usually cannot satisfy (b)'s deterministic evidence bar unless the proof avoids intent/meaning judgment. Description / frontmatter is never wording-only (see below).
205
206
  - **(b) Deterministic scope check + independent review.** Dropping challenge for the edit requires a **deterministic scope check** — recorded controller-side or diff-based scope evidence that is decidable without judging intent or meaning, such as: touched files are formatting-only by formatter output; hunks are only meaning-inert whitespace / punctuation / markdown table alignment; or token-level changes are limited to a named typo correction while the surrounding rule sentence is byte-identical. If the proof depends on a human or LLM deciding whether revised prose changes a rule, threshold, boundary, applicability, or acceptance meaning, it is NOT deterministic and challenge stays required. The deterministic evidence **AND** an independent review row confirming the same must both be present. Either piece missing, or any reviewer-flagged / unconfirmed meaning / scope / trigger / routing / validation / acceptance change, **re-arms challenge + the behavioral-evidence row** (never demoted to "recommended").
206
207
 
208
+ The executable controller proof is intentionally narrower than every edit a
209
+ human might call wording-only. It accepts only a canonical full-context
210
+ Markdown diff inside one existing skill and recomputes either punctuation-only
211
+ changed lines or one named whole-token replacement with an exact count. Use the
212
+ schema and command in `code-review/references/staged-review-contract.md`; the
213
+ result must carry both `wording_only_scope.status=passed` and the independently
214
+ reviewed `wording_only_boundary` concern. Other grammar/synonym edits, custom or
215
+ context-augmented packets, multiple skills, and any unconfirmed meaning change
216
+ take the normal challenge path.
217
+
207
218
  **If no such deterministic scope check can be formed for the change, challenge stays required.** This is a hard gate, not a default the author may waive: an LLM independent review is hypothesis-grade, so review alone never downgrades a non-wording change. Every shared-skill change still requires the independent review row regardless of class.
208
219
 
209
220
  If either required pass times out, returns empty output, is rate-limited, cannot access auth, exits nonzero, emits malformed or truncated output, fails JSON/shape parsing when structured output was requested, lacks evidence that the pass inspected the target diff/files, shows a prompt/tool-scope mismatch, or returns any `inconclusive` status, the dual-track gate has not passed **for that lane via that reviewer**. A *recoverable-lane* failure — auth, quota, rate limit, timeout, local cache/db failure, or missing capability — is NOT a terminal stop and is NOT "the gate is unrunnable on this host": before you record `blocked`, you MUST walk the **Primary reviewer failure** remediation ladder below and route to an approved independent third-party reviewer (preferably a different model family) — either your runtime's native multi-model subagent (e.g. an OpenCode `Task`/council subagent on a separate model) or a shell wrapper such as `opencode_review.sh --model <provider/model> --implementer-family <author-family> --mode review|challenge`. A primary CLI lacking auth is a routing trigger to that fallback lane, never a license to declare the gate unrunnable. Only after that ladder is exhausted — every approved independent reviewer probed and unavailable — do you record the row as `pending` or `blocked`, include the remediation attempted (which ladder steps were tried) and the next unblock action, and do not describe the skill change as solved, complete, landed, or fully closed. (A large reviewer INPUT can also be silently middle-truncated, not just the reviewer's OUTPUT — see **Read-coverage of large inputs** under *Sanity checks the gate must enforce*.)
@@ -248,7 +259,30 @@ A stale base puts the upstream's newer fixes into the packet **reversed**, so th
248
259
 
249
260
  Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, the review is not landing evidence until the base is re-pinned and the round re-run. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
250
261
 
251
- **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks any of the five (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section passes a stale base and gets a verdict that looks exactly like a good one. Make that difference visible rather than silent: **the review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is recorded `base-unattested`, which is not landing evidence.** A missing attestation is then a detectable state instead of an indistinguishable one; that is the containment available to prose, and it is weaker than a gate. Mechanising the list into the wrapper is the durable form and is registered as follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
262
+ When that check detects drift, rebuild instead of improvising: stop the reviewer
263
+ lane; preserve the named candidate manifest/patch and the caller-owned ordered evidence rows;
264
+ integrate the newly attested target in an isolated worktree; reapply only the
265
+ named candidate paths/rows; regenerate derived artifacts; compare the resulting
266
+ path set with the manifest; rerun the selected tests; then re-pin and rebuild the
267
+ packet. Never use a broad reset plus `add -A`, which can silently absorb ambient
268
+ work. Keep every base attestation in one ledger scoped from the first packet
269
+ until landing or a human scheduling decision; repinning, retrying, or rebuilding
270
+ does not reset it. Each row uses one remote/ref, a contiguous sequence, a strictly
271
+ increasing RFC3339 confirmation time, the corresponding controller-receipt hash
272
+ when a round consumed it, and a same-directory hash-bound file containing the
273
+ canonical raw `ls-remote` line. A second ordered SHA change (A→B→C or A→B→A) is
274
+ the second drift and must terminate the lane as `baseline_race`, with the
275
+ unreviewed delta; the drift row and every later row must not map another
276
+ controller receipt. For every non-race closeout — `ready_for_human_decision`
277
+ and `continuation_authorization_required` alike — the final controller receipt
278
+ must consume the latest attested SHA; later same-SHA live rechecks are allowed,
279
+ but an unconsumed newer SHA is not reviewed evidence (a post-final-round drift
280
+ belongs in the next round's ledger, not appended unconsumed to this one). Do not open another
281
+ automatic rebuild after the second drift. The state validator counts these
282
+ transitions and receipt/base associations inside the complete referenced row
283
+ set. Keep independent work moving while a human chooses a landing window.
284
+
285
+ **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks the five live-Git properties (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section can still pass a stale base. The v3 closeout validator makes referenced evidence tampering and broken round association detectable, but it cannot prove that a caller supplied every historical attestation or that the recorded `ls-remote` output is still current; candidate-local receipts are consistency evidence, not remote authority or an append-only log. Therefore **each review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is `base-unattested`, not landing evidence**, the caller retains the complete history across rebuilds, and item 2's live authority query is still repeated immediately before landing. Mechanising the live checks and history retention into a trusted wrapper/platform is follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
252
286
 
253
287
  > **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
254
288
  >
@@ -262,7 +296,7 @@ Walking all five is a **point-in-time attestation, not a lock**: the target can
262
296
 
263
297
  Primary reviewer failure is a remediation branch only when the owning gate classifies it as candidate-local. Use this ladder separately for the review lane and challenge lane:
264
298
 
265
- 1. Persist the owner-guided self-review, pass it through the required `--review-plan-file`, and run `review_gate.sh` once for the lane. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
299
+ 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` from round 1; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
266
300
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
267
301
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
268
302
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
@@ -270,11 +304,16 @@ Primary reviewer failure is a remediation branch only when the owning gate class
270
304
 
271
305
  Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
272
306
 
273
- - **Open the Agent chain on the FIRST review — it cannot be retrofitted.** An extraction's required review and challenge are one tracked multi-round run, so the round budget is decided before round 1, not after reading the review; a run that starts untracked is thrown away and restarted. Every trigger, default, flag, index, prior-result, and advisory rule behind that obligation is owned by `code-review/references/staged-review-contract.md` (Agent review chain), with the runnable pair in `code-review/SKILL.md` — take the command from there and never reconstruct it from this bullet.
307
+ - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
274
308
 
275
309
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
276
310
 
277
- Standard `codex review` against the target diff:
311
+ The following raw CLI shapes are **debugging diagnostics only**. They may help
312
+ isolate an owner-wrapper failure, but neither output is review/challenge evidence
313
+ and neither may replace `scripts/extraction_review_gate.sh` for a non-wording
314
+ lane.
315
+
316
+ Debug a wrapper with raw `codex review` against the target diff:
278
317
 
279
318
  ```bash
280
319
  cd <skills-repo>
@@ -310,7 +349,7 @@ Output: one finding per line with severity (P0/P1/P2), file:line, scenario, fix.
310
349
 
311
350
  ## Running the challenge pass
312
351
 
313
- `codex exec` with adversarial prompt (read-only):
352
+ Debug a wrapper with raw `codex exec` and an adversarial prompt (read-only):
314
353
 
315
354
  ```bash
316
355
  timeout 540 codex exec "<adversarial-prompt>" </dev/null \
@@ -377,6 +416,15 @@ item 9 signs off on it.
377
416
 
378
417
  Use `--json` to capture reasoning traces and tool calls cleanly. Parse the JSONL stream with a small Python or jq script as documented in `gstack-codex` skill.
379
418
 
419
+ ## Reviewer verification scope (packet-verifiability boundary)
420
+
421
+ The external reviewer judges what the bounded packet can show; the packet structurally cannot carry every deterministic oracle its acceptance claims depend on (frozen preservation-mapping rows, the checker's complete pinned-literal sets, whole-file postimages, immutable pre-fix revisions). The division of labor is fixed and documented here so it is ruled on once, not re-litigated per round:
422
+
423
+ - The reviewer owns CONTENT SEMANTICS: wording coherence, dropped qualifiers/obligations, source-accuracy, sanitization, scope drift — everything decidable from the packet plus the reviewer's own reasoning.
424
+ - Deterministic-gate claims (pinned literals present, size ratchet net-zero, obligation audit green, R0 clean, parity) are verified by the repository's CI re-running those gates on the actual branch — never by the reviewer, and never accepted from the implementer's prose alone; a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round.
425
+ - Historical-process claims (a pre-fix RED, a measurement taken before landing) are session-record-grade unless bound to an immutable revision or a candidate-bound receipt; treat them as the implementer's testimony, and say so in the disposition instead of demanding evidence the packet cannot hold.
426
+ - Standing backlog: teaching the gate to embed candidate-SHA-bound receipts of deterministic-gate output into the packet removes the third bullet's limitation mechanically; until that lands, this boundary is the accepted state.
427
+
380
428
  ## Recording findings + fixes
381
429
 
382
430
  For each pass, record in the extraction's working file (e.g. `<project>-extraction-summary.md`):
@@ -484,6 +532,31 @@ Do NOT iterate to zero *findings* — some are intentional design tradeoffs the
484
532
 
485
533
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
486
534
 
535
+ Five is the generic `code-review` transport ceiling, not this extraction lane's
536
+ spend. Non-wording Agent-autonomous extraction calls go through
537
+ `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=2`: the
538
+ initial review plus at most two challenges. At round 3 the autonomous lane ends.
539
+ An authenticated human may request later review, but that is separately
540
+ attributed human-requested evidence outside this chain/budget, not an Agent
541
+ round 4 or 5. Unused generic capacity never authorizes automatic continuation.
542
+ The v3 closeout validator rejects referenced receipts whose recorded budget is
543
+ not 2 and checks budget and ordering consistency within the caller-supplied
544
+ set. It cannot authenticate that the wrapper produced those receipts or that
545
+ the caller retained every earlier chain or receipt. The wrapper does not mint or
546
+ persist
547
+ `review_chain_id` or `autonomous_review_index`: the caller still supplies both,
548
+ and could start a fresh-looking chain after round 3. The validator detects bad
549
+ order inside the referenced set but cannot detect a prior chain the caller
550
+ omitted, so complete caller-owned ledger retention—and treating an Agent reset
551
+ as a contract violation—remains part of the boundary rather than a property the
552
+ local scripts prove.
553
+
554
+ A strictly proven wording-only change has no convergence loop: it uses one
555
+ generic `code-review` pass, records the independent-review row and the
556
+ challenge-not-required proof, and does not create a schema-v3 multi-round
557
+ terminal ledger. This exception does not apply to frontmatter, routing,
558
+ validation, acceptance, example, owner or behavior changes.
559
+
487
560
  This budget limits only automatic reviewer invocation. It does **not** stop implementation, tests, debugging, or deep self-review, and it does not limit an authenticated human:
488
561
 
489
562
  - A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
@@ -505,6 +578,32 @@ At the third Agent-autonomous round, do not start a fourth automatically. If fin
505
578
  - mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
506
579
  - enter `awaiting_human` only when no independent runnable work remains. This is a scheduling state, not task failure and not a human merge prohibition.
507
580
 
581
+ The terminal checkpoint is an extraction closeout record, not a state emitted by
582
+ `review_gate.py`, and its schema-v3 state is derived from evidence rather than
583
+ trusted as an author assertion. Schema-v2 closeout ledgers are rejected rather
584
+ than silently reinterpreted under the breaking occurrence/evidence shape. The
585
+ ledger and every referenced controller, completion, base, and sweep file live
586
+ in one directory and carry SHA-256s. The validator walks the ordered schema-v3
587
+ controller chain (same chain and scope,
588
+ review then contiguous challenges, packet=candidate, complete prior-result hash
589
+ prefix, fixed `challenge_budget=2`) and binds every closeout candidate to its
590
+ last receipt. Ready requires at least review + challenge; a second base drift may
591
+ stop as race immediately after round 1 rather than spending an illegal challenge
592
+ after the terminal predicate already fired.
593
+ It ends in exactly one state:
594
+
595
+ - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
596
+ - `continuation_authorization_required`: round 3 itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
597
+ - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
598
+
599
+ Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
600
+ the state. This proves internal consistency and coverage of the files the ledger
601
+ references. It does **not** authenticate that no earlier receipt/attestation was
602
+ omitted and does not replace the live remote recheck above; the caller still owns
603
+ complete-history retention until a trusted platform owns it. An exhausted budget,
604
+ stale review, omitted evidence, or unknown lane state is never represented as
605
+ convergence.
606
+
508
607
  ### Concrete cadence
509
608
 
510
609
  For a focused single-skill change: