@ccoalm/ccl-skills 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (127) hide show
  1. package/README.md +2 -2
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +8 -7
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +6 -1
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +16 -17
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +1 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +195 -7
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +3 -3
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +13 -5
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +9 -3
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +9 -3
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/normalize_review_timeout.sh +22 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +9 -3
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +1540 -129
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +8 -3
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +76 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +1858 -3
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +789 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/update_review_plan_intent.py +513 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +4 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -1
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +11 -10
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +64 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml +4 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +73 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/runtime-and-project-contract.md +58 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +41 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/verification-diagnostics-and-security.md +63 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +14 -16
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +10 -14
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +1 -1
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +135 -86
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +66 -80
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/delivery-contract.md +275 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +88 -214
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +2 -2
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +10 -8
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +6 -5
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +112 -95
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +30 -21
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +22 -3
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +20 -17
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +7 -9
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +14 -10
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +2 -0
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +3 -3
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +9 -6
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +3 -0
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +37 -10
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +8 -1
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +16 -5
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +16 -5
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +4 -2
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +4 -1
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +4 -0
  77. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +104 -5
  78. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +11 -9
  79. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +102 -0
  80. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +103 -0
  81. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +20 -0
  82. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +6 -6
  83. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +4 -3
  84. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +69 -2
  85. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +22 -0
  86. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +49 -4
  87. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +2748 -0
  88. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +20 -5
  89. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py +1142 -0
  90. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +17 -0
  91. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +41 -4
  92. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +120 -0
  93. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +82 -8
  94. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +336 -0
  95. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +82 -10
  96. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +1416 -0
  97. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh +57 -0
  98. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +141 -4
  99. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +3 -1
  100. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh +1696 -0
  101. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2117 -0
  102. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +316 -0
  103. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +1176 -0
  104. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +31 -1
  105. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +9 -4
  106. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +980 -0
  107. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +8 -6
  108. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  109. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  110. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  111. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +16 -15
  112. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  113. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +10 -2
  114. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  115. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  116. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  117. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  118. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  119. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  120. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  121. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +7 -5
  122. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +1 -1
  123. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  124. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  125. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  126. package/dist/assets/release.json +215 -105
  127. package/package.json +1 -1
@@ -1,23 +1,24 @@
1
1
  ---
2
2
  name: terminal-cli-dev
3
- description: Use when designing, implementing, reviewing, debugging, testing, or shipping command-line, terminal, text UI, PTY, ANSI-rendered, keyboard-driven, or console product surfaces, including the command/subcommand/flag/help contract (owned here even when nothing is rendered), layout, input, color, wrapping, selection, scrollback, accessibility, performance, and real terminal verification. Triggers also include "命令行/TUI 界面怎么做", "CLI 界面怎么写", "重构这个终端/TUI 命令或界面(局部)", "refactor a terminal command / TUI view". Skip when the ask is CLI/tooling implementation in a language whose dev skill owns it, without terminal-UI concerns ("用 Python 写个命令行工具" → python-service-dev, Go CLI → go-microservice-dev); a CLI in a language with no such owner stays here.
3
+ description: CLI/terminal/console/PTY/ANSI/keyboard/TUI design, implementation, review, debugging, testing, or shipping. Owns the command/subcommand/flag/help contract (owned here even when nothing is rendered), with defaults/output/exit/action/confirmation/progress/recovery, plus layout, input, accessibility, and real-terminal evidence. Triggers include "命令行/TUI 界面怎么做", "CLI 界面怎么写", "refactor a terminal command / TUI view". Skip only parser/library/tooling internals owned by a language skill ("用 Python 写个命令行工具" → python-service-dev, Go CLI → go-microservice-dev, Node.js CLI → nodejs-service-dev) that provably preserve every user-facing command tree, default/action path, help/output/exit behavior, confirmation, progress, and recovery path; compose both owners when user-visible semantics change.
4
4
  ---
5
5
 
6
6
  # Terminal CLI Dev
7
7
 
8
- Use this skill for terminal and command-line product surfaces. It owns implementation mechanics for text UIs, console workflows, ANSI-rendered output, PTY-backed interaction, keyboard input, terminal capability handling, and real terminal verification. It does not own web browsers, mobile apps, mini-program hosts, backend services, or product design judgment.
8
+ Use this skill for terminal and command-line product surfaces. It owns the user-facing command/subcommand/flag/default/help/output/exit/action/confirmation/progress/recovery contract, plus implementation mechanics for text UIs, console workflows, ANSI-rendered output, PTY-backed interaction, keyboard input, terminal capability handling, and real terminal verification. A language skill may own parser or library mechanics, but those mechanics do not displace this user-visible contract. This skill does not own web browsers, mobile apps, mini-program hosts, backend services, or product design judgment.
9
9
 
10
10
  ## Routing
11
11
 
12
12
  - Use `product-rd-workflow` first when the work spans product intent, design, implementation, testing, release, or postmortem follow-up.
13
- - Use `product-ui-ux-design` before or alongside coding when the terminal surface is user-facing: hierarchy, density, interaction model, copy, states, accessibility, and visual acceptance.
13
+ - A user-facing command/subcommand/flag/default/help/output/exit/action/confirmation/progress/recovery path is a terminal surface even when it emits only plain text and never enters an alternate screen.
14
+ - Use `product-ui-ux-design` before or alongside coding for that user-facing terminal surface: hierarchy, density, interaction model, copy, states, accessibility, consequence, recovery, and visual/textual acceptance.
14
15
  - Use `testing-strategy` to choose unit, snapshot, PTY, integration, and real-terminal evidence; return here for terminal-specific implementation mechanics.
15
16
  - Use `defect-diagnosis` first for rendering regressions, input bugs, flicker, selection/copy issues, broken resize behavior, color/readability defects, or flaky terminal tests.
16
17
  - Use `platform-observability` for telemetry/log schema and `platform-release-engineering` for rollout of behavior-changing defaults, persisted settings migrations, or terminal capability fallbacks. Do not treat CLI package distribution, installer, or updater mechanics as covered unless the release skill has explicit terminal distribution guidance.
17
18
 
18
19
  ## Core Workflow
19
20
 
20
- Before editing terminal UI code, output formatting, keyboard handling, PTY integration, layout, color/theme logic, or terminal tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a small copy or formatting change can use a short note; a new interactive surface, renderer, input mode, terminal capability change, high-risk action, release behavior, or bug fix needs explicit scenarios, target terminal environments, verification commands, and stop conditions.
21
+ Before editing a user-facing command tree, flag/default/action path, help/output/exit behavior, confirmation/progress/recovery flow, terminal UI code, output formatting, keyboard handling, PTY integration, layout, color/theme logic, or terminal tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a small copy or formatting change can use a short note; a new command path, changed default, interactive surface, renderer, input mode, terminal capability change, high-risk action, release behavior, or bug fix needs explicit scenarios, target terminal environments, verification commands, and stop conditions.
21
22
 
22
23
  Repo-local agent contracts (`AGENTS.md` at the repo root and in source directories) are part of the delivery contract: when a change moves a stable boundary, generated surface, workflow, or directory-local rule, update the nearest contract in the same MR and keep coverage in sync per `product-rd-workflow`'s spec / repo-contract sync gate.
23
24
 
@@ -33,8 +34,8 @@ When checking a terminal/CLI project against team standards, split conformance i
33
34
  - Whether the feature requires a TTY, raw mode, cursor control, bracketed paste, mouse reporting, focus reporting, hyperlinks, truecolor, or scrollback control.
34
35
  - Behavior when capabilities are missing, disabled by user preference, blocked by a multiplexer, proxied through a remote shell, or unavailable in CI.
35
36
  - State ownership for input focus, modal overlays, selection, scroll position, unseen output, pending operations, and resize recovery.
36
- - For UI/UX redesign slices meeting `product-ui-ux-design`'s page-slice trigger conditions — that gate's trigger list is authoritative and must be checked, not paraphrased, whenever a screen/surface change could be a redesign, restyle, new-style declaration, structural/visual-system change, continuation, or redesigned-surface reference — apply its cross-stack page-slice gate before terminal mechanics: RED-first focused assertion, IA regrouping by user intent/consequence, behavior-contract preservation, state matrix, rendered evidence, and the design verdict (`accepted` / `rejected` / `pending`; missing = `pending`, and `design-rejected` blocks complete/MR-ready/normal/draft MR per the **Rejected-surface rule**). Translate Web/App examples into terminal cell-grid proof instead of copying browser or device commands.
37
- - For any user-facing visible terminal UI change, also record `product-ui-ux-design`'s implementation-owner checkpoint before the first edit — its field list (design/stack/test owners, entry-rule evidence, rendered/device evidence status) and copy-only path are authoritative there; load the named owner skills rather than only naming them, and treat a completion claim without `captured/verified` rendered evidence as incomplete — an explicitly accepted gap closes the slice only as `pre-runtime-test ready` / handoff, never as complete/done.
37
+ - For every user-facing terminal/CLI contract change—including command tree, subcommand, flag/default/action path, help/output/exit behavior, confirmation, progress, or recovery—load `../product-ui-ux-design/references/delivery-contract.md` and consume either its full Design brief + Phase 0 or its valid low-risk copy-only record + lightweight Phase 0 before coding. Only parser/library internals that preserve all of those user-visible semantics may mark UI/UX `not-applicable`, with the preservation evidence recorded. The lightweight path checks semantics, accessible text, localization/width, cell extent, and target-terminal render without inventing unrelated matrices; risk-bearing copy or behavior uses the full path. For full slices, map structure, state/adaptation matrices, behavior and criteria to the screen buffer/lifecycle; record terminal classes, dimensions, input/capability modes, resize/scrollback/selection, and preserved state.
38
+ - Before the first implementation edit, add the canonical `client_entry` defined there: local rule identifier or short quote and implementation decision, target surface/runtime, planned run/capture command, and behavior that must remain unchanged.
38
39
 
39
40
  3. Render by terminal cells, not string length.
40
41
  - Measure display width with ANSI-stripped, Unicode-aware logic. Cover combining marks, emoji, East Asian width, zero-width code points, variation selectors, and control characters.
@@ -78,6 +79,7 @@ When checking a terminal/CLI project against team standards, split conformance i
78
79
  - Use a PTY or equivalent integration test for raw mode, resize, key/mouse/paste sequences, terminal responses, and process lifecycle.
79
80
  - Run at least one real terminal smoke for visible interactive changes when lower layers cannot prove color, cursor, scrollback, focus, selection, or resize behavior.
80
81
  - For UI/UX redesign evidence, include target terminal class, size and narrow/short stress size, color mode or no-color fallback, keyboard-only path, empty/loading/error/final states, scrollback behavior, selection/copy boundary, resize behavior, and screenshot/transcript/PTY artifact; mark each dimension covered or `N/A` with a one-line reason. `N/A` is valid only when the reason names a verifiable structural fact, explains why that fact makes the dimension unreachable or unchanged for this slice, and includes a checkable pointer such as a file path, config key, or commit that resolves at review time. Persist evidence artifacts where reviewers can access them after redacting tokens, PII, credentials, private paths, command secrets, and raw personal data; remove temporary smoke files or PTY capture helpers before commit unless the repo intentionally owns them.
82
+ - Return the complete canonical client-record member defined in `../product-ui-ux-design/references/delivery-contract.md` for testing Phase 1 and the design verdict. The member includes its applied rule/decision, affected files/components, preserved behavior, exact command, immutable candidate binding, producer member/version actually exercised, artifacts, tested terminal classes/dimensions/states/input/capability modes, criterion-mapped observations, coverage boundary, and gaps. A terminal capture proves only the captured states; it cannot close an unbound producer member. `testing-strategy` records aggregate sufficiency before the design owner records the candidate-bound verdict.
81
83
  - Capture evidence: command, terminal class, size, color mode, before/after screenshot or transcript, and any unavailable capability with attempted remediation.
82
84
 
83
85
  ## Reference Loading
@@ -13,7 +13,7 @@
13
13
  - **James Bach & Michael Bolton**, *Rapid Software Testing* — HTSM / SFDPOT
14
14
  - **Elisabeth Hendrickson**, *Explore It!* — test heuristics cheatsheet
15
15
  - **James Whittaker**, *Exploratory Software Testing* — tours
16
- - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(多被测系统 2-way 捕到 50–90% 的缺陷,差异大);工具:Microsoft PICT、NIST ACTS、Hexawise
16
+ - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(NIST SP 800-142 Table 1 复现其数据:各被测域 2-way 累计触发 53–97%,多数域 70–97%;NIST 同文提醒 pairwise 仍可能漏掉 10–40% 或更多缺陷,mission-critical 不足恃);工具:Microsoft PICT、NIST ACTS、Hexawise
17
17
  - **Hans Buwalda 2004** — soap opera testing
18
18
  - **Lisa Crispin & Janet Gregory**, *Agile Testing*
19
19
  - **Glenford Myers**, *The Art of Software Testing* (1979) — error guessing 起源
@@ -61,7 +61,7 @@ SKILL.md 现状 P0 / P1 / P2 是经验判断("blocking / important / nice")
61
61
 
62
62
  ### 2.1 风险公式(最通用)
63
63
 
64
- **Risk = Probability × Impact**(业内 30 年共识,ISTQB / ISO 29119 同源)
64
+ **Risk = Probability × Impact**(出处:ISTQB CTFL v4.0.1 §5.2——风险级别由 likelihood 与 impact 决定,**定量法**为二者相乘,**定性法**用风险矩阵,二者皆合规;ISO/IEC/IEEE 29119-1 采用类似 likelihood/impact 框架,原文付费墙未逐字核)
65
65
 
66
66
  **Probability**(发生概率,1-5 分):
67
67
  - 1 = 罕见(新代码 + 简单逻辑 + 测试覆盖好)
@@ -49,6 +49,8 @@ python test/scripts/gen_report.py \
49
49
  - md 同步匹配必须以 用例ID 为唯一键;模块名或功能点相同不代表是同一条记录
50
50
  - 若 md 先于 Bitable 被修改(如直接编辑文件),需将 md 变更反向同步到 Bitable,再按 update 工作流补信息流转
51
51
  - **漂移检查**:定期运行 `python gen_report.py --config test/.report-config.json --diff-md test/cases/all.md` 显示 md 与 Bitable 的 added/removed/changed;废弃记录自动排除(不在 md 里属正常);多人编辑后必跑一次再 commit
52
+ - **版本钉扎**:当仓内测试代码/脚本按某条 TC 实现时,在实现侧记录该 TC 的修订标识——用**定义字段的规范化快照哈希**(或 Bitable 的不可变 record 修订号,若可得);`[姓名 日期]` 不够(同人同日二次修订会撞标识,钉住旧版仍比对相等)。实现前先比对,**回写/提测前再验一次,且回写本身走与认领相同的修订条件更新**(先验后写仍是两步——验证通过与写入之间的并发修改只有条件写能挡)——哈希只有相等性判断:**任何一次不一致都视为分歧,停下、先做源对账**(更新钉扎侧并使差异可评审后再继续);「更新/更旧」的说法仅当平台提供不可变单调修订号时才可用;发现源记录 stale/自相矛盾时**停下先修源记录**,不得按旧版实现后事后补
53
+ - **认领防冲突**:多人/多 agent 并发按 TC 实现时,认领必须是**原子的修订条件更新**(仅当记录修订仍等于读取时的修订才写入——平台支持的 CAS/乐观锁语义)或走**串行化认领协调者**(单写入口)。「读空→写」是 TOCTOU;「写后回读」也不等价——A 回读成功后仍可被 B 顶掉、双方各自都验证通过。两种安全机制都不可得时,**并发认领在该表上不受支持:停下改走人工/单线分配,不得按回读结果继续**。字段已有他人活跃认领时不得覆盖,改为联系认领人或换条目
52
54
  - `+record-upsert --record-id` 中的 record_id 是 Bitable 内部 ID,必须先用 `get_record_index()` 从 用例ID 查出,不能直接用 用例ID 代替
53
55
 
54
56
  ## update 工作流
@@ -71,26 +71,25 @@ Use this skill to decide whether a specialized or non-functional test belongs in
71
71
  - Conditional skips (missing-optional-dependency guards such as module-level `importorskip`, platform/env markers) combined with per-job test selection can leave an entire test file executed in NO CI job while every pipeline stays green: the job that selects the file lacks the optional dependency (the skip fires for the whole module), and the job that has the dependency does not select the file. When a suite mixes conditional skips with job-scoped test selection, the job that owns those tests must carry an executed-count guard — the per-file invariant and the floor fallback live in `references/ci-fixtures-and-flake-control.md`. Any change to job-level selection re-verifies which files each job actually executes (run with skip reporting and read the executed/skipped counts per file). "The tests exist and CI is green" is not evidence they ran anywhere.
72
72
  - A new, ported, or mirrored enforcement mechanism (pre-edit hook, permission guard, write-blocking plugin, validator) is not verified by loading, parsing, or config inspection — those prove installation, not enforcement. Require a behavioral matrix before a completion claim — blocked case per deny-condition, allowed/no-collateral cases, the fail-open/swallowed-exception bypass set, and the canonicalization/symlink/worktree edge cases; the matrix cells, the fail-open bypass set, the safe-unavailable-gap disposition for a case that cannot be exercised safely, and the port/mirror parity procedure live in `references/verify-enforcement-mechanisms.md`. Run the matrix only against scratch/synthetic targets (a throwaway checkout/worktree, fixture repo, or dry-run mode) — never a live workspace, real user data, or live credentials. Parity claimed from code reading alone is hypothesis-grade, not evidence. When a test is **ported/mirrored to a sibling stack**, input-fixture fidelity is part of that parity — see `references/test-code-authoring-patterns.md` (跨栈移植:移植对抗输入本身).
73
73
  - A "skip CI" / "no runner" instruction does not by itself lower verification rigor, only ceremony. When the blocking CI gate is skipped or unavailable, substitute a same-risk independent check before treating the change as verified — for a tiny/doc/test-only change a local command or `diff --check` is enough; for a change that can break its own gate (it edits the test/tripwire/CI config it is guarded by), an adversarial review/challenge of the diff is what catches the self-break CI would have. Note where a local run is not equivalent to CI (secrets, OS matrix, merge-result pipeline) rather than treating it as full proof. Separately, a project-enforced merge gate (pipeline-must-pass, required review) is not waived by a "skip CI" instruction for convenience: require green status, or an explicit authorized break-glass/override with recorded reason + residual risk — surface the conflict and stop rather than silently bypassing.
74
- - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture or inspect console errors and failed network requests when the browser tool supports it.
74
+ - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture console errors and failed network requests when tools support it.
75
75
  - UI tests and screenshots must prove design quality layers, not only DOM existence. Assert or visually inspect aesthetic hierarchy/density, interaction path, behavioral recovery states, and psychology-critical cues such as disabled reasons, progress certainty, retry safety, confirmation consequences, and return context.
76
- - For UI/UX redesign slices meeting `product-ui-ux-design`'s page-slice trigger conditions — that gate's trigger list is authoritative and must be checked, not paraphrased, whenever a screen/surface change could be a redesign, restyle, new-style declaration, structural/visual-system change, continuation, or redesigned-surface reference — derive scenario selection from that gate's recorded state matrix and RED baseline instead of re-inventing scenarios, and route rendered-evidence layer choice here as usual. A test plan that ignores a triggered page-slice record, or that lets component/DOM layers stand in for the gate's stack-routed rendered evidence, is incomplete.
77
- - For any runtime visible UI/UX slice, consume `product-ui-ux-design`'s implementation-owner checkpoint (its field list and copy-only path are authoritative there): this skill is the named test owner for assertion and rendered-evidence layer selection, and a test plan or completion claim whose rendered/device evidence status is not `captured/verified` is incomplete. If no checkpoint exists for the slice, that absence blocks implementation-facing test plans and completion claims — load `product-ui-ux-design` and remediate the checkpoint first; producing this skill's assertion-layer and rendered-evidence choice as an input to creating that checkpoint is the expected first step, not a blocked action. An explicitly accepted gap recorded per that checkpoint closes the slice only as `pre-runtime-test ready` / handoff with the gap stated, never as complete/done.
76
+ - For every runtime-visible UI/UX slice, load the canonical sequence in `../product-ui-ux-design/references/delivery-contract.md` and `references/client-runtime-test-matrices.md` §UI/UX Delivery Contract before Phase 0 and after producer/client execution. Testing owns layer selection and sufficiency, binds and cites the complete design/test/producer/client record and candidate-binding sets, confirms every affected client wrote its canonical pre-edit `client_entry` and complete client-record member naming the producer version it exercised, fails closed on a missing/incomplete/mismatched/stale/changed-after-run/unexercised member, and never issues the holistic design verdict.
78
77
  - Authentication and account surfaces need an explicit scenario matrix before they can be called complete. Cover identity-input validation across relevant entries, available sign-in methods, registration, account recovery or password reset/change, logout/account switching, sensitive storage/log cleanup, permission or host-authorization denial, and UI/UX acceptance for error copy, disabled reasons, keyboard/safe-area/touch behavior, and visual evidence. If a capability such as recovery, host authorization, real message delivery, or live account verification is absent or external, record it as `product gap`, `blocked`, or `live-only` instead of silently excluding it from the test claim.
79
- - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the reproducible command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
80
- - Do not answer "tests are complete" from command output alone. The claim must map each important scenario to a written case or to an explicit `blocked`, `live-only`, `product gap`, or `not applicable` row.
78
+ - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, `infra-error`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the repro command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
79
+ - Do not answer "tests are complete" from command output alone. Map each important scenario to a written case or an explicit `blocked`, `live-only`, `product gap`, `infra-error`, or `not applicable` row. (verdict definitions and their mutual exclusivity: `references/ci-fixtures-and-flake-control.md`).
81
80
  - For report-only QA, baseline comparison, or "testing only" branches with no product-code changes, failing tests can be the intended deliverable.
82
81
  - This exception applies only when a human reviewer, PR owner, or user explicitly states in the current work item, PR description, or current-turn context that the deliverable is test coverage, evidence, or a QA report rather than a product fix; an agent or automated process cannot infer or self-apply this exception from prior-session memory or summarized context.
83
82
  - Do not weaken the test or patch product code just to go green.
84
83
  - For disputed or high-stakes defects, prefer splitting verification and fix into two deliverables: a test-only verification slice first pins the confirm/deny verdict and root-cause attribution, and the fix is a separate change that references it — keeping the verification verdict uncontaminated by fix intent. Its failing regression test may land skip-marked as the trace only as a bounded state, not an escape: the skip carries the reason, an owner, and the linked fix item, and accepting the fix requires un-skipping it into the blocking regression set (or an explicitly owner-signed quarantine lane) — a RED test that stays skipped after its fix merges is the bypass this rule exists to prevent.
85
84
  - The QA report is not complete until pass evidence, red-light evidence, blocked/live-only gaps, and baseline comparison are each present and non-empty or explicitly marked `not applicable`; red-light evidence and baseline comparison cannot both be `not applicable`.
86
- - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner and do not resolve it by trusting either side.
85
+ - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner; do not resolve it by trusting either side.
87
86
  - A QA report with an open live-contradiction gap is `blocked` until the discrepancy is escalated and an owner assigns a resolution path.
88
- - Generated starter tests are not regression evidence: replace scaffold placeholders (e.g. a counter widget test) in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
89
- - Verification warnings are not automatically follow-up work: classify build/bundle-size/lint/flaky/deprecation/security/performance warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
90
- - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix has been written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are a release risk, not an afterthought.
87
+ - Generated starter tests are not regression evidence: replace scaffold placeholders in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
88
+ - Verification warnings are not automatically follow-up work: classify build/bundle/lint/flaky/deprecation/security/perf warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
89
+ - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix is written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are release risk, not afterthought.
91
90
  - A multi-stack development-standard family is incomplete without a testing standard. The testing standard must define test deliverables, layer policy, harness expectations, CI gates, high-risk coverage, evidence format, and stack handoff rules; stack docs may specialize commands but must not redefine the layer policy.
92
- - Do not mark a browser/device/E2E layer unavailable just because discovery returns no device, browser, server, or dependency. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
93
- - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; this is a handoff-only label, not merge-ready or release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, ready to merge, or ready to release.
91
+ - Do not mark a browser/device/E2E layer unavailable just because discovery returns nothing. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
92
+ - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; a handoff-only label, not merge-ready/release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, merge-ready, or release-ready.
94
93
 
95
94
  ## Entry Decision: TC Source and Scope
96
95
 
@@ -112,16 +111,18 @@ TCs cover functional and interaction scope (QA perspective: user journeys, accep
112
111
  - A TC entry without automated test code is also valid: manual test, deferred automation, or blocked environment.
113
112
  - The automated report (`gen_report.py`) tracks TC-mapped results (tests that register TC IDs via the `tc(...)` helper — see `test-artifact-management/references/tc-marker-conventions.md`) plus a separate "未链接 TC 的测试" section for tests without TC links. Broader code coverage is a code quality concern tracked separately (e.g. coverage reports, CI pass/fail).
114
113
 
115
- ### Step B — Determine scope (if not specified by user)
114
+ ### Step B — Determine scope (if not already bounded)
116
115
 
117
- If the user has not specified scope, **ask before proceeding**:
116
+ A current acceptance source or reviewable task artifact may already confirm scope. In particular, a UI/UX Design brief with a stable slice/surface, authoritative consumer inventory, affected owner(s), and criterion IDs is the confirmed bounded scope for its Phase 0/Phase 1 work; consume it instead of asking the user to restate global/file/function scope.
117
+
118
+ If neither the request nor a current authoritative artifact bounds the work, inspect the current task, repository contract, diff/target, and available acceptance sources first. Ask only when two or more plausible scopes remain and choosing among them would materially change the test plan. Then ask the smallest concrete question, for example:
118
119
 
119
120
  > 请确认测试范围:
120
121
  > 1. 全局 — 整个仓库 / 当前 feature 所有文件
121
122
  > 2. 文件 — 指定文件(请提供路径)
122
123
  > 3. 函数 / 接口 — 指定函数或 API endpoint(请提供名称)
123
124
 
124
- Do not infer scope from context alone — an incorrect scope wastes implementation work. Wait for the user's answer before moving to layer assignment.
125
+ Do not infer scope from stale conversation or an unverified guess. A current, resolvable Design brief or accepted scope artifact is evidence, not inference. If the inspection leaves one material scope, proceed and record its source; if ambiguity remains, wait for the user's answer before layer assignment.
125
126
 
126
127
  When scope is confirmed:
127
128
  - **全局**: run the full scenario matrix from `testing-strategy` workflow; consult TC list for all active TCs.
@@ -171,7 +172,7 @@ Before editing tests, CI gates, mocks/fakes, fixtures, test scripts, verificatio
171
172
  - High-risk resilience matrix: for each triggered class, choose the lowest layer that can prove the invariant, then add one release drill or real-flow smoke for the most expensive failure. Examples: idempotency unit/contract plus duplicate callback integration; auth timeout unit plus cross-tenant API integration; AI fallback eval/replay plus visible refusal E2E; client double-submit component test plus server idempotency integration.
172
173
  - High-risk backend tests should include missing tenant/actor/subject/resource-scope rejection, durable idempotency beyond cache TTL, mutation-plus-audit/outbox atomicity or repair visibility, stale/pending worker status, and operator/request/trace evidence for admin repair paths when those risks exist.
173
174
 
174
- For client API-backed surfaces, mini-program/mobile/device runtime smoke, runtime-client mechanisms (route guards, permission trees, request interceptors, generated clients, upload wrappers, long-task polling, safe-area/keyboard/orientation handling, foreground/background restore, native bridges, app-hosted H5), terminal/CLI/TUI runtime tests, and streaming/async-finality changes (model streams, queued jobs, MQ consumers, scheduled prompt tasks, cron tasks, persisted tool-output artifacts, long-running exports, long-lived connections), load `references/client-runtime-test-matrices.md` before assigning layers — and again before declaring any runtime/device/browser evidence unavailable, not only at layer assignment. Non-negotiable anchors kept in view here: the default three-boundary client split (unit/component + API client/contract + browser/device smoke) applies unless the repository has a stronger convention; runtime-dependent smoke is **blocking** when lower layers cannot prove the changed behavior, and a missing runner after normal remediation stops at `pre-runtime-test ready` or `blocked` with owner, commands attempted, residual risk, and next unblock action — never converts into a code correctness claim; developer-tool compile/preview is structural evidence only; dangerous or irreversible operations require operator/confirmation/audit/duplicate-submit/final-status assertions before the UI or API is called ready. Build scenario matrices from the reference's reusable dimensions (host/container, identity/permission, data/state, async/finality, visual/interaction, high-consequence); do not paste product-specific matrices into this skill — product-specific lists belong in the project checklist or the owning product/domain skill.
175
+ For client API-backed surfaces, mini-program/mobile/device runtime smoke, runtime-client mechanisms (route guards, permission trees, request interceptors, generated clients, upload wrappers, long-task polling, safe-area/keyboard/orientation handling, foreground/background restore, native bridges, app-hosted H5), terminal/CLI/TUI runtime tests, and streaming/async-finality changes (model streams, queued jobs, MQ consumers, scheduled prompt tasks, cron tasks, persisted tool-output artifacts, long-running exports, long-lived connections), load `references/client-runtime-test-matrices.md` before assigning layers — and again before declaring any runtime/device/browser evidence unavailable, not only at layer assignment. Non-negotiable anchors kept in view here: the default three-boundary client split (unit/component + API client/contract + browser/device smoke) applies unless the repository has a stronger convention; runtime-dependent smoke is **blocking** when lower layers cannot prove the changed behavior, and a missing runner after normal remediation stops at `pre-runtime-test-ready` or `blocked` with owner, commands attempted, residual risk, and next unblock action — never converts into a code correctness claim; developer-tool compile/preview is structural evidence only; dangerous or irreversible operations require operator/confirmation/audit/duplicate-submit/final-status assertions before the UI or API is called ready. Build scenario matrices from the reference's reusable dimensions (host/container, identity/permission, data/state, async/finality, visual/interaction, high-consequence); do not paste product-specific matrices into this skill — product-specific lists belong in the project checklist or the owning product/domain skill.
175
176
 
176
177
  4. Define data and dependency strategy.
177
178
  - Use small fixtures named by scenario. For repeated/complex fixture construction, pick §4 (factory/builder) from the decision table in `references/test-code-authoring-patterns.md`.
@@ -16,7 +16,7 @@ Duplication and dead-code gates run with explicit configuration, not defaults: t
16
16
 
17
17
  ## Frozen Regression Set And Adversarial Passes
18
18
 
19
- Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
19
+ Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests. The tiering criterion is cost and side effects, not speed: a case that consumes paid resources (model/API spend, sandbox creation, render jobs) or mutates external state never belongs in the default always-on lane even when it happens to be fast — the default lane is reserved for cases that are free and side-effect-free to run on every change.
20
20
 
21
21
  The proactive complement — the adversarial pass over code already considered "done": run it as an active defect-discovery step, deliberately hunting coverage blind spots — error-mapping boundaries, concurrency-protection bypass, double-release/double-close paths — instead of waiting for review or production to surface them. Each confirmed gap lands as a failing test first. Candidate blind-spot classes: the risk-matrix failure classes in `scenario-testing.md`, plus the dependency fault-injection and concurrency/cache cases in `integration-contract-testing.md`.
22
22
 
@@ -24,6 +24,10 @@ The proactive complement — the adversarial pass over code already considered "
24
24
 
25
25
  Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
26
26
 
27
+ - `infra-error` verdict semantics (the entrypoint's status family). One discriminating predicate decides the verdict — **fault origin**, not symptom: a fault in the **evidence infrastructure** (collector, fixture cache/manifest, state-preparation or controlled-fault harness — anything outside the system under test) = `infra-error`, which is never a pass, never a business fail, and never a silent skip; a fault **in the system under test** (crash, malformed product response, product-path timeout) or an assertion that evaluates to false on collected evidence = business `fail`; an **external prerequisite missing before any attempt** = `blocked`. Exactly one verdict per case; ambiguous origin is resolved by investigation, never by defaulting to whichever verdict looks better; page text, a success toast, or another weaker surface must not substitute for the missing evidence. A required case standing at `infra-error` keeps the aggregate claim incomplete — it counts against merge/release readiness exactly like `blocked`, and a report that excludes `infra-error` cases to present a clean total is a false-green report. Before coding a case, each acceptance criterion names its collector and assertion.
28
+
29
+ External-asset fixtures (media files, documents, large binaries fetched from an external system) form a supply chain that gets pinned end to end: test execution reads only a local read-only cache — never downloads from the external system at run time; a committed manifest pins each asset's identity/hash and CI verifies the manifest plus every blob before the suite runs; cache/manifest verification is part of each affected case's attempted preparation, so a missing or changed cached asset maps to `infra-error` for exactly the cases that need it (a manifest failure aborting before any case attempt marks those cases `infra-error` too, not `blocked` — the infrastructure was configured and failed), never a skip and never a fallback download; seeding/refreshing the cache is a separate offline step on a trusted host, not part of the test run. This composes the network-isolation default and manifest regenerate-and-diff rules in `test-data-and-determinism.md` with the missing-dependency-is-failure rule below into one chain.
30
+
27
31
  ### Fault-Injection Layers For External-Provider Recovery Paths
28
32
 
29
33
  Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
@@ -16,10 +16,18 @@ For client API-backed surfaces, apply this default split unless the repository h
16
16
  - Browser/device/E2E smoke: open the real page or app screen, perform the primary action with controlled data or a stable test backend, verify the visible success path, then verify one realistic failure path is readable and non-crashing.
17
17
  - Mini-program smoke: compile or preview in the relevant platform developer tool, open the real page or preview build, verify route/scene params, loading/success/failure states, auth or permission behavior when relevant, and one host-platform capability path such as share, payment, subscribe message, camera, scan, or webview bridge when touched. Developer-tool compile or preview is structural evidence only; it does not replace assertion-based behavior tests or rendered flow checks.
18
18
 
19
+ ## UI/UX Delivery Contract
20
+
21
+ For every runtime-visible UI/UX slice, load the canonical sequence in `../../product-ui-ux-design/references/delivery-contract.md`.
22
+
23
+ - For a full UI/UX slice, Test selection Phase 0 derives its case set from the applicable Design brief's state/adaptation matrix, criterion IDs/outcomes, and RED baseline, while testing owns verifier type, assertion/rendered-evidence layers, commands/targets, independent oracles, and gaps; add risk-driven cases when needed, but do not silently replace or narrow the recorded design obligations. A valid low-risk copy-only record uses the contract's lightweight Phase 0—semantic, accessible-name, localization, extent, and target-render checks—instead of unrelated state/adaptation matrices.
24
+ - After producer/client execution, Phase 1 cites the complete shared `design_record_ids`, `test_record_ids`, `producer_record_ids`, `client_record_ids`, and `candidate_binding_set`, adds only bound test-owned executions/artifacts, maps results to criteria using producer observations, client observations, and test artifacts, and records evidence sufficiency, combined coverage boundary, and unresolved gaps. Every changed or claim-bearing brief/criteria/source/review artifact, harness/oracle/fixture/config/protocol, backend/config/prompt/model producer, and every affected rendered layer has a keyed member; each client record names the producer member/version it exercised. A missing canonical `client_entry`, incomplete client-record member, or missing, mismatched, stale, changed-after-run, or unexercised member blocks sufficiency. Do not copy role-owned raw fields into a second record. A truly single-member `candidate_binding` is shorthand only; a branch name, abbreviated SHA, empty digest, mutable external path, or `commit:HEAD` after dirty execution does not bind evidence. Dirty bundles include result-affecting ignored inputs; inputs outside the checkout use their own exact-bytes `artifact-sha256` member. Testing does not issue the holistic design verdict or its status combination, and `accepted + complete` is invalid unless bound Phase 1 is `sufficient` with no required evidence gap.
25
+ - Phase 0 is incomplete when the applicable full/lightweight design inputs are missing, or when DOM/component existence, a build pass, or an unreviewed screenshot is used in place of required rendered, interaction, accessibility, recovery, or design evidence. A user-accepted evidence gap remains at the contract's scoped handoff state, never `complete`.
26
+
19
27
  ## Runtime Smoke: Blocking Rules And Evidence Surface
20
28
 
21
- - Runtime-dependent client smoke is blocking when lower layers cannot prove the changed behavior, including platform request/chunking, streaming finality, foreground/background restore, host navigation, permission/capability prompts, native or mini-program bridge callbacks, storage/session restore, and host-rendered error or recovery states. If the runner is missing, first attempt normal setup/remediation; if still unavailable, stop at `pre-runtime-test ready` or `blocked` with owner, commands attempted, residual risk, and next unblock action.
22
- - For mini-program and mobile runtime smoke, treat app identity, plugin authorization, dev-tool login, service-port availability, and host permissions as part of the evidence surface. If those inputs are wrong or missing after remediation, classify the slice as `blocked` or `pre-runtime-test ready` rather than converting the gap into a code correctness claim.
29
+ - Runtime-dependent client smoke is blocking when lower layers cannot prove the changed behavior, including platform request/chunking, streaming finality, foreground/background restore, host navigation, permission/capability prompts, native or mini-program bridge callbacks, storage/session restore, and host-rendered error or recovery states. If the runner is missing, first attempt normal setup/remediation; if still unavailable, stop at `pre-runtime-test-ready` or `blocked` with owner, commands attempted, residual risk, and next unblock action.
30
+ - For mini-program and mobile runtime smoke, treat app identity, plugin authorization, dev-tool login, service-port availability, and host permissions as part of the evidence surface. If those inputs are wrong or missing after remediation, classify the slice as `blocked` or `pre-runtime-test-ready` rather than converting the gap into a code correctness claim.
23
31
  - For mini-program flows whose correctness depends on real-host completion state, classify the strategy as "real WeChat / real device required" and point the project checklist at a reusable host-flow template: target app entry, minimal real input, first-chunk observation, final closure, and evidence capture. Keep the concrete app name and search path in the project checklist; keep the skill-level rule generic.
24
32
  - Device/environment readiness: when a mobile, browser, or service runner is required, verify readiness with the runner's own discovery command and one direct health/state command. For Android, for example, do not trust a wrapper script alone; confirm `adb devices -l`, `adb -s <serial> get-state`, and boot readiness before treating the device E2E layer as available or unavailable.
25
33
 
@@ -23,7 +23,7 @@ Before adding a new E2E test, create or update the scenario matrix in `scenario-
23
23
  - Save authenticated state only when the test is not about login.
24
24
  - Capture console errors, failed network requests, screenshots/traces/video when useful.
25
25
  - Use network interception only to control nondeterminism or assert payloads; do not mock away the contract that the E2E test is meant to prove.
26
- - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing. Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing** (`@playwright/experimental-ct-{react,vue,svelte}`) is still labeled experimental — use Vitest browser mode for component-level tests until Playwright component-test API stabilizes. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
26
+ - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing (community-adoption evidence, reproducible: `curl -s https://api.npmjs.org/downloads/point/2026-07-31:2026-08-29/<pkg>` for `playwright` / `@playwright/test` / `cypress` returned 339.9M / 216.2M / 30.3M (fixed range, re-verified 2026-08-30), ≈11×; State of JS 2024 testing section, `2024.stateofjs.com/en-US/libraries/testing/`, shows Playwright leading E2E usage/retention — re-run the query before citing as current). Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing**: Playwright's current component-testing guide replaces the former `@playwright/experimental-ct-{react,vue}` packages (`playwright.dev/docs/test-components`) — that replacement statement is all the source establishes; draw no stability or package-layout inference from it, and follow the pinned Playwright version's own installation instructions before adding or removing any component-testing package. Vitest browser mode remains a valid component-level alternative when the portfolio already standardizes on Vitest. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
27
27
 
28
28
  ## Runtime QA Sweep
29
29
 
@@ -58,7 +58,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
58
58
  - Do not reproduce every unit branch through E2E.
59
59
  - Keep E2E flows few, stable, and tied to user/business risk.
60
60
  - Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
61
- - A click-through without assertions is not E2E evidence.
61
+ - A click-through without assertions is not E2E evidence — and weak proxy signals are not business assertions: page loaded, URL changed, non-empty body text, a generic button/canvas/heading visible, or a success toast alone do not prove the business outcome. Anchor the pass condition on objective effects — API response fields, persisted records, balance/count deltas, generated artifact URLs (sufficient alone only when URL issuance is the claimed contract; a generation-success case dereferences the URL and validates artifact status/metadata/content, since a request can issue a valid URL and fail before storing the artifact), or a stable user-visible terminal state (alone only when the visible terminal presentation IS the claimed contract — a rendered "completed" can outrun persistence/billing/artifact creation, so business-outcome cases pair it with the durable effect) — and match assertion strength to what the test title and scenario row claim; a shallow signal is acceptable only when the case explicitly tests just that shallow signal.
62
62
  - If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
63
63
  - External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
64
64
 
@@ -113,6 +113,16 @@ the implementation.
113
113
  - For cross-RPC typed error envelopes, test a roundtrip: server raises a typed error, the wire-format payload is captured, the client reconstructs a typed error of the same class with the same code/message. Include the unknown-shape path: a wire payload that does not match the canonical envelope returns a transport/unknown error without silent loss of the original cause.
114
114
  - For Code-range allocation, test that a service trying to register a code outside its allocated range fails at build/test time, not at runtime.
115
115
 
116
+ ### Cross-Repo Field Change — End-to-End Checklist
117
+
118
+ Adding, renaming, or retyping a field that crosses a repo/service boundary is one end-to-end contract change, not N independent edits. Before calling it covered, walk all five steps (each is a distinct failure site with its own evidence):
119
+
120
+ 1. **Producer fallback / bridge** — for an additive field the producer emits a safe default/absent form for consumers that have not upgraded; for a rename/retype a default is NOT enough — keep a compatibility bridge (dual-write old+new representation, or a versioned mapping) until every active consumer is confirmed reading the new form, then remove it in the cleanup stage. Asserted, not assumed.
121
+ 2. **Every transport mapper preserves the field explicitly** — do not assume an object spread/copy crosses a mapper or DTO boundary; each mapper in the chain gets an assertion that the field survives it.
122
+ 3. **Consumer coverage spans all active consumer variants** — enumerate them from the delivery record's consumer inventory (`../../product-ui-ux-design/references/delivery-contract.md` consumer_inventory for UI variants); testing one variant of a multi-variant consumer is the classic escape.
123
+ 4. **One real inbound frame through the mapper, plus one unchanged generic path as control** — the real-frame test proves the new field flows; the untouched-path test proves the change did not perturb everything else (the control catches over-broad mapping edits).
124
+ 5. **Paired changes are cross-linked and land in a compatibility-safe order** — not by mutually blocking merges (that deadlocks): consumer tolerance for the field's absence/new form lands first, then the producer emission (its fallback from step 1 keeps not-yet-upgraded consumers safe), then cleanup removes the fallback once all consumers are confirmed upgraded. Each stage gates on the previous stage's **deployed** compatibility state, and the MRs cross-reference per `../../product-rd-workflow/references/cross-repo-coordination.md` for visibility.
125
+
116
126
  ## Platform Contract / Protobuf / RPC Test Obligations
117
127
 
118
128
  Platform-service-connectivity owns policy and proof mechanics for protobuf-backed HTTP, response envelopes, RPC/base fields, and boundary exposure. `testing-strategy` owns assertion coverage, verdict shape, and CI placement.
@@ -78,7 +78,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
78
78
 
79
79
  ## 3. Test smells(测试异味)
80
80
 
81
- **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书列约 18 项分 code/behavior/project 三类;下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
81
+ **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书顶层列 15 项 smell,分 code/behavior/project 三类(5/6/4;xunitpatterns.com "All Test Smells" 目录,另有类下变体/别名未计入);下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
82
82
 
83
83
  | 异味 | 含义 | 后果 |
84
84
  |---|---|---|
@@ -229,7 +229,7 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
229
229
  - **MC/DC**(Modified Condition/Decision Coverage):每个 boolean 子条件独立影响过决策。**DO-178C 航空 / 医疗 / 汽车 functional safety 才用**;普通业务代码无监管要求时不必上。
230
230
  - **Mutation coverage**(见 source-to-case-workflows §C.1):才是真"测得好"的 proxy
231
231
 
232
- **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。
232
+ **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。数值出处:60%/90% 对齐 Google Testing Blog "Code Coverage Best Practices"(2020)的 60% acceptable / 75% commendable / 90% exemplary 分档——注意该文同时反对自上而下的强制统一阈值,floor 应按仓现状起步再棘轮;branch 50% 为团队启发值,无外部权威出处。
233
233
 
234
234
  **不用**:
235
235
  - **不用 100% 作 KPI** — 强行凑 100% 会产生 lazy assertions(`assert result is not None` 这种)
@@ -64,7 +64,7 @@ Classify commands before running them:
64
64
  - E2E/release smoke: browser/API/device real flows in an isolated environment.
65
65
  - Long gate: compatibility matrix, migration dry-run, load/replay, visual regression, or full-suite release checks.
66
66
 
67
- If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it.
67
+ If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it. Expensive includes billable: tests that consume metered external resources — paid AI/model inference, per-call third-party APIs, real payment/checkout flows, cloud sandboxes or device farms — get their own explicit marker/lane, stay out of default PR and scheduled-frequent lanes by default, and run only with a named budget owner and an authorized test account (credential/provisioning discipline per `ci-fixtures-and-flake-control.md`). Payment/checkout tests default to the provider's sandbox/test mode — a budget owner bounds spend but does not make live payment mutation safe; an unavoidable live-money canary needs its own explicit approval, a hard spending cap, and refund/cleanup handling. A case that silently creates paid resources under a generic `e2e` tag is a lane-classification finding.
68
68
 
69
69
  When existing commands use richer labels such as `contract_fake`, `contract_mysql`, `api`, `e2e_smoke`, `failure_mode`, `drill`, `shadow`, `smoke`, or `replay`, preserve the local meaning instead of flattening everything into unit/integration/E2E. Map them to the scenario matrix and CI gate they actually serve.
70
70
 
@@ -99,7 +99,7 @@ For a 域卡/执行卡 (a card that sets WHAT a domain must achieve + who owns i
99
99
  - Required flow: STOP line edits → confirm the corrected core with the user (one short question, don't guess again — this failure class recurs precisely from re-guessing) → re-derive 负责人/红线/里程碑/验收/依赖兜底 from the corrected core as a **draft for review** (not a blind blast-write) → publish on approval.
100
100
  - Repeat signal: repeated user "这是什么/什么玩意儿" on the same card = the premise is wrong; escalate to re-derive, do not keep tightening.
101
101
 
102
- ## 句子层(吸收 Strunk《风格的要素》与 Google Technical Writing 课程,仅取适合中文交付文档的;英文语法/标点规则不适用,已剔除)
102
+ ## 句子层(吸收 Strunk 与 Google Tech Writing,仅取中文交付文档适用项)
103
103
 
104
104
  > 英文文档:用完整 Strunk 规则(含被本节剔除的语法/标点条),本节只是中文交付子集。
105
105
 
@@ -140,6 +140,7 @@ Never destroy collaborative comments. Before editing a collaborative doc, fetch
140
140
 
141
141
  ## WORKFLOW
142
142
 
143
+ 0. 读者批注:判根因类、全文修同类(`references/annotation-driven-revision.md`)。
143
144
  1. Extract the decided-points checklist from the current text.
144
145
  2. Apply the DELETE list; keep everything in KEEP.
145
146
  3. Rewrite to FORM; confirm every decided point still present.
@@ -0,0 +1,9 @@
1
+ # 批注驱动修订
2
+
3
+ 读者批注/评审意见不是孤立改句请求,而是**阅读断裂的证据**。处理协议:
4
+
5
+ - 逐条判根因类:背景缺失 / 概念未定义 / 逻辑跳跃 / 措辞 / **事实・引用错误**(日期、数字、出处错——此类不走措辞同类扫,改走源核验:对一手源改正并按同源扫其余引用处)。
6
+ - 按根因类**必须全文扫同类位置一起修,不得只改被标记的那一句**——点修复会把同类断裂留给下一位读者复发。**扫描全文、编辑限权**:当授权只覆盖某条批注/某节时,全文扫描产出同类候选清单,但自动编辑只落在授权范围内;范围外的同类位置先报告、经批准再修(不得以「修同类」为名越权改动已定内容)。
7
+ - 改完以首次读者身份通读被改段落(standalone-paste 逐行读,同 closeout 判法)。
8
+ - 批注本体的保全走 SKILL.md 的 COMMENT-SAFE 硬规则(先取真实评论数、定向编辑、改后复核锚点/条数)。
9
+ - 出处:读者差集原则(好文档=读者需要的知识−已有的知识,Google Technical Writing audience 章)——批注正是「差集没算对」的实测信号。
@@ -16,8 +16,14 @@
16
16
 
17
17
  **`[禁]` 档(已核,勿再立)**:加粗/高亮密度;每 N 字一图的图表密度(唯一数字是学术期刊的
18
18
  **印刷页数配额**,成因是版面成本不是可读性);"一行不超过 40 汉字"引 WCAG(该条的 CJK 40 是从
19
- 拉丁文 80 折半推导,且原文要求是「提供**机制**让用户改」不是「正文必须排这么宽」);"句子不超过
20
- 25 词";"留白提升理解约 20%"(错误引用链)。
19
+ 拉丁文 80 折半推导,且原文要求是「提供**机制**让用户改」不是「正文必须排这么宽」;作为**社区规范**
20
+ 另有可追溯来源——阮一峰《中文技术文档的写作规范》"多于40个字的句子不能接受"——个人规范可引用,
21
+ 不可当标准/厂商级权威);"留白提升理解约 20%"(错误引用链;2026-08 复核仍无任何标准/厂商/学术
22
+ 一手源给出该比例)。
23
+
24
+ **降档更正(2026-08 复核)**:"句子不超过 25 词"从 `[禁]` 移出——GOV.UK 官方写作指引有一手源
25
+ ("Try to split up sentences that are over 25 words long" + GDS 博客专文),属**单一机构 house
26
+ style**:可引用(点名 GOV.UK),不可当行业标准立硬规则,归 `[工]` 档强度。
21
27
 
22
28
  **同体裁实测分布 ≠ 规范阈值**:可以说"本稿在同体裁公开样本分布的哪个位置",不能由此推出"写得好"。
23
29
  分布定位的合法输出是描述不是裁决——这条与 `tighten-doc` SKILL.md「外部基线」条同源,按那条执行。
@@ -33,7 +33,7 @@ Use this skill for React web client engineering. It covers browser-rendered Reac
33
33
 
34
34
  ## Core Workflow
35
35
 
36
- Before editing components, routes, state, API clients, styles, configs, or tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a simple low-risk single-component change can use a short inline plan; multi-file, user-visible, API-visible, accessibility-sensitive, release, bug-fix, branch/MR, unclear-risk, or high-risk work needs explicit task split, design checkpoint, acceptance checks, verification commands, rollback or stop conditions, and named handoffs to design, testing, miniapp/app, backend, or diagnosis skills before edits.
36
+ Before editing components, routes, state, API clients, styles, configs, or tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a simple low-risk single-component change can use a short inline plan; multi-file, API-visible, accessibility-sensitive, release, bug-fix, branch/MR, unclear-risk, or high-risk work needs explicit task split, acceptance checks, verification commands, rollback or stop conditions, and named handoffs to testing, miniapp/app, backend, or diagnosis skills before edits. Runtime-visible work additionally consumes the canonical UI/UX delivery contract's Design brief and Test selection Phase 0 before the first implementation edit.
37
37
 
38
38
  Repo-local agent contracts (`AGENTS.md` at the repo root and in source directories) are part of the delivery contract: when a change moves a stable boundary, generated surface, workflow, or directory-local rule, update the nearest contract in the same MR and keep coverage in sync per `product-rd-workflow`'s spec / repo-contract sync gate.
39
39
 
@@ -50,9 +50,10 @@ When checking a React project against team standards, split findings into determ
50
50
  - Locate the owning route/page, component tree, state owner, API client, data-fetching layer, styling system, tests, and build scripts before editing.
51
51
  - Identify whether state belongs in URL/query params, cache/server state, form state, local component state, browser storage, or global app state.
52
52
  - Read repo wrappers first: package manager, dev/build scripts, lint/typecheck/test runners, browser/E2E tools, environment variables, and generated clients.
53
+ - Before generating component-library code: the **workspace lockfile resolution is the version authority** (`npm ls <pkg>` / `pnpm why` / yarn equivalent — a library config file or global CLI can resolve a different release than the workspace); take configuration ground truth (framework, aliases, installed components) from the library's own introspection surface (config file such as `components.json`, official info CLI/MCP, or the installed package's exports/types); write APIs against the resolved version, never from memory of "current" APIs (prop names and defaults shift across majors). After editing, close with the library's own linter/codemod check on the changed files when one exists (deprecated-usage and a11y rules the generic lint config does not know); for a library major-version migration, follow the official migration checklist + changelog for the exact from→to pair, apply, then re-run the library lint to prove no deprecated usage remains.
53
54
  - If a design exists, map visible states and interactions to component ownership before implementing.
54
- - For any visible UI change, map the design checkpoint to implementation ownership before coding: visual hierarchy/density, interaction flow, behavioral feedback, user psychology, responsive collapse, and screenshot acceptance. Do not reduce the design to component names. Also record `product-ui-ux-design`'s implementation-owner checkpoint before the first edit — its field list (design/stack/test owners, entry-rule evidence, rendered/device evidence status) and copy-only path are authoritative there; load the named owner skills rather than only naming them, and treat a completion claim without `captured/verified` rendered evidence as incomplete — an explicitly accepted gap closes the slice only as `pre-runtime-test ready` / handoff, never as complete/done.
55
- - For UI/UX redesign slices meeting `product-ui-ux-design`'s page-slice trigger conditions — that gate's trigger list is authoritative and must be checked, not paraphrased, whenever a screen/surface change could be a redesign, restyle, new-style declaration, structural/visual-system change, continuation, or redesigned-surface reference — apply its cross-stack page-slice gate before React mechanics: RED-first focused assertion, IA regrouping by user intent/consequence, behavior-contract preservation, state matrix, rendered evidence, and the design verdict (`accepted` / `rejected` / `pending`; missing = `pending`, and `design-rejected` blocks complete/MR-ready/normal/draft MR per the **Rejected-surface rule**). Web is not a lower-evidence surface than app; component tests or DOM snapshots must be paired with browser-rendered evidence for changed layout, interaction, or visual states.
55
+ - For every visible UI change, load `../product-ui-ux-design/references/delivery-contract.md` and consume either its full Design brief + Phase 0 or its valid low-risk copy-only record + lightweight Phase 0 before coding. The lightweight path checks semantics, accessible name, localization, rendered extent, and target render without inventing unrelated matrices; risk-bearing copy uses the full path. For full slices, map structure, state/adaptation matrices, behavior and criteria to React ownership; record route/server, component/state owners, viewports/themes/input modes, and preserved behavior. When React is embedded in a native WebView, mini-program `web-view`, or Electron shell, this skill owns the content-layer member; the native/mini/desktop host owner must add its separate entry, binding and runtime record, even when host code is unchanged.
56
+ - Before the first implementation edit, add the canonical `client_entry` defined there: local rule identifier or short quote and implementation decision, target surface/runtime, planned run/capture command, and behavior that must remain unchanged.
56
57
 
57
58
  3. Structure React code by ownership.
58
59
  - Decompose UI by responsibility, not by arbitrary visual fragments.
@@ -98,7 +99,8 @@ When checking a React project against team standards, split findings into determ
98
99
  - For API-backed UI, test component states, API client parsing/error translation, and at least one browser/E2E smoke path when feasible.
99
100
  - Inspect the rendered page in a browser for any visible UI change, responsive behavior, empty/error states, and console/network errors.
100
101
  - For UI/UX redesign evidence, include the declared stress viewport, or when none exists use the minimum supported width plus one narrow stress width such as 320px; text wrapping/overflow; loading/empty/error/final states; keyboard/focus path; and a browser screenshot or equivalent visual artifact. Mark each dimension covered or `N/A` with a one-line reason; `N/A` is valid only when the reason names a verifiable structural fact, explains why that fact makes the dimension unreachable or unchanged for this slice, and includes a checkable pointer such as a file path, config key, or commit that resolves at review time. Persist evidence artifacts where reviewers can access them using sanitized/test accounts and redacting tokens, PII, credentials, private paths, and raw personal data; delete temporary smoke pages or helper scripts before commit unless the repo intentionally owns them.
101
- - For browser-runtime changes, browser smoke is a completion gate when lower layers cannot prove the behavior. This includes changes to routing, browser storage/session restore, streaming/fetch finality, visibility or foreground/background behavior, permission/capability prompts, WebView bridge callbacks, upload/media flows, and rendered loading/error/final states. If the browser or app server is missing, first attempt normal setup; if still unavailable, stop at `pre-runtime-test ready` or `blocked` and name the owner, attempted commands, residual risk, and next unblock action. `pre-runtime-test ready` is handoff-only, not merge-ready, release-ready, or complete.
102
+ - Return the complete canonical client-record member defined in `../product-ui-ux-design/references/delivery-contract.md` for testing Phase 1 and the design verdict. The member includes its applied rule/decision, affected files/components, preserved behavior, exact command, immutable candidate binding, producer member/version actually exercised, artifacts, tested route/server, viewport/container sizes, themes/input modes/states, criterion-mapped observations, console/network checks, coverage boundary, and gaps. The browser render proves only the captured content layer; it cannot close an embedded host or unbound producer member. `testing-strategy` records aggregate sufficiency before the design owner records the candidate-bound verdict.
103
+ - For browser-runtime changes, browser smoke is a completion gate when lower layers cannot prove the behavior. This includes changes to routing, browser storage/session restore, streaming/fetch finality, visibility or foreground/background behavior, permission/capability prompts, WebView bridge callbacks, upload/media flows, and rendered loading/error/final states. If the browser or app server is missing, first attempt normal setup; if still unavailable, stop at `pre-runtime-test-ready` or `blocked` and name the owner, attempted commands, residual risk, and next unblock action. `pre-runtime-test-ready` is handoff-only, not merge-ready, release-ready, or complete.
102
104
  - Check accessibility names, labels, focus order, keyboard navigation, aria only when semantic HTML is insufficient, contrast, and text wrapping.
103
105
  - Check performance when relevant: bundle impact, unnecessary renders, long lists, image loading, code splitting, hydration/runtime errors, and Core Web Vitals risk.
104
106
 
@@ -114,7 +116,7 @@ When checking a React project against team standards, split findings into determ
114
116
  - Do not treat a frontend API client as done until empty response, invalid JSON, non-2xx envelope, auth expiry, network failure, cancellation, and backend error message extraction are covered at the client or component boundary when relevant.
115
117
  - Do not scatter backend enum/string literals through React components, URL/query handling, analytics, or tests. Centralize finite-value parsing, display labels, defaults, and unknown-value behavior at the API/client-domain boundary, and keep raw literals only in clearly named boundary conversion tests that cover every known external value plus unknown/default behavior. Migrate existing non-boundary test raw literals for that value in the same pull request or mark each remaining use with `finite-value-debt: <task-ref> <owner> <deadline> <reason>`, even when the current slice does not introduce a new mapper.
116
118
  - Do not debug React/browser failures from code inspection alone when a browser reproduction, console output, network trace, screenshot, or focused test can be collected.
117
- - Do not claim a web client fix is complete without naming the browser/rendered verification that was run. If required browser/runtime verification is unavailable after remediation, the status is `pre-runtime-test ready` or `blocked`, not complete.
119
+ - Do not claim a web client fix is complete without naming the browser/rendered verification that was run. If required browser/runtime verification is unavailable after remediation, the status is `pre-runtime-test-ready` or `blocked`, not complete.
118
120
 
119
121
  ## Reference Loading
120
122
 
@@ -44,4 +44,4 @@ Use this reference after `web-react-dev/SKILL.md` identifies a React surface as
44
44
 
45
45
  - State owner map exists for every selected pattern family.
46
46
  - Long content, empty/no-data, error/retry, slow/weak network, permission/disabled, narrow/responsive, accessibility text scaling, interruption/return recovery, and repeated-use/cache-hit behavior are either tested or explicitly out of scope.
47
- - Browser or host-container evidence captures the declared stress widths and the primary pending/final/error states. If rendered evidence cannot run after normal remediation, status is `pre-runtime-test ready` or `blocked`, not complete.
47
+ - Browser or host-container evidence captures the declared stress widths and the primary pending/final/error states. If rendered evidence cannot run after normal remediation, status is `pre-runtime-test-ready` or `blocked`, not complete.
@@ -50,6 +50,9 @@ Three reuse patterns recur in React libraries; pick by what the consumer needs t
50
50
  - **Custom hook** — when reuse is logic only, no rendering shape required. Default choice for state machines, subscriptions, side-effect orchestration. See discipline above.
51
51
  - **Headless / unstyled primitives** (Radix UI, Headless UI, Ariakit, downshift, react-aria) — when reuse is a11y + behavior (focus trap, keyboard navigation, ARIA contract) but every consumer needs different visual treatment. Default for design-system primitives: the primitive owns keyboard / focus / ARIA / portal / dismiss, the consumer owns styling. **Do not invent a custom focus-trap or ARIA implementation when a maintained headless primitive exists** — the bug surface is well-known and the maintained library has fixed bugs you have not heard of yet.
52
52
  - **Compound components** (e.g., `<Select><Select.Trigger/><Select.Content/><Select.Item/></Select>`) — when consumers need to arrange children but share an implicit parent context. Cleaner than render props for this case. Default for `Tabs` / `Accordion` / `Select` / `Menu` / `Tooltip` shells.
53
+ - **Type the shared context as a grouped contract — `{ state, actions, meta }`** (data / state-changing functions / refs & config; a component with no refs or config declares `meta` as an optional key — `meta?:` — and consumers handle its absence) rather than a flat grab-bag. Subcomponents consume the contract, never a specific state hook, so any provider implementing it can inject the state — local `useState` for an ephemeral form, a global or server-synced store for a live surface — and the same composed UI works under either provider. One packed context value means a consumer that only calls `actions` still re-renders whenever `state` changes; when that measurably matters, split state and actions into two contexts (the React-docs reducer + context split) instead of one value. The split pays off only when the provided actions value is itself referentially stable — pass `dispatch` directly or memoize the actions object with its complete dependency list, never dropping a changing dependency (an `onSave` prop, a scoped API client) to force identity — that ships stale calls; an inline `{ dispatch }` wrapper takes a new identity on every provider render, and the optimization holds only while the declared dependencies are actually stable; actions must not close over current state (a reducer's `dispatch` qualifies; state-closing memoized callbacks do not), and reducers/updaters themselves stay pure — a side-effecting command that needs the current state snapshot, like a submit posting the form, accepts it as an argument from a state-reading consumer; when actions inherently depend on state, keep the single combined context.
54
+ - **The provider boundary, not the visual shell, decides who can share state.** When consumers outside the visual shell need the state, lift it into a dedicated provider component — the lowest-common-owner rule above still decides how high, and lift only the shareable model `state`/`actions`: each compound root keeps focus refs, generated IDs, and item registration instance-local — open/highlight state stays local by default too, lifting only when outside consumers genuinely coordinate it and then scoped by an explicit compound-instance identity — or two shells under one provider cross-wire focus and ARIA targets; a preview panel or submit button rendered outside the visual shell but inside the provider then reads `state` and calls `actions.submit`. Two smells that say state should have been lifted into a provider: syncing a child's state upward with an on-change `useEffect` callback, and imperatively reading child-owned React state through a ref because another consumer must coordinate with it — uncontrolled DOM values read at submit (`FormData`, file inputs) and imperative third-party-widget integrations are not this smell and stay valid per the uncontrolled-forms rule above.
55
+ - **Mode booleans multiplying on one component are a composition finding.** When a new component API — or an existing API already undergoing an intentional redesign — accumulates mode props (`isEditing` / `isCompact` / `isInline`-style) whose combinations multiply and some are impossible, replace the modes with explicit variant components — each variant composes the shared compound parts it needs and composes under the appropriate shared provider, or declares its own when the variant is the lowest common state owner. The `never`-union prop typing above makes illegal combinations uncallable at the type level; explicit variants remove the combinations altogether. Prefer variants when the modes are stable, meaningfully different compositions or ownership boundaries and conditional render branches, not just prop types, fork on the booleans; modes sharing one structural contract stay one component with a discriminated-union variant prop and exhaustive branching. Do not refactor a stable existing API solely for pattern conformance.
53
56
 
54
57
  **Legacy patterns to recognize, not to reach for**:
55
58
  - **Higher-Order Components (HOCs)** — `withAuth(Component)` / `connect(mapStateToProps)(Component)`. For the same concern in NEW code, prefer a custom hook (`useAuth()` / `useSelector()`). Existing HOC APIs are acceptable when a library / framework contract requires them (React-Redux `connect` is still a valid public API) — do not refactor working HOC integrations on cosmetic grounds. Do not introduce a new HOC when a hook does the same job.