@ccoalm/ccl-skills 0.8.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +5 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +1 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +6 -6
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +2 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +2 -2
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +1 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +4 -4
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +4 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +9 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +49 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +12 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -9
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  50. package/dist/assets/release.json +56 -51
  51. package/package.json +1 -1
@@ -61,7 +61,7 @@ This skill coordinates gates; it does **not** itself authorize merge, tag push,
61
61
  | Reset dev/test-like branches | Yes | Target/env refs, before SHAs, dry-run/plan, force-with-lease semantics |
62
62
  | Post-merge cleanup of the merged temp feature branch (worktree/local/remote) | No — covered by the user's merge authorization (`worktree-isolation` 收尾) | The authorized MR/PR read back as merged at the current head SHA and target; the live remote source ref is absent (already cleaned by the platform) or still equals the merged MR source head (moved → preserve and ask, remote path only — eligible local cleanup proceeds per `worktree-isolation`); no other open or plan-declared MR/PR still consumes the source branch; source branch is a temp feature branch (unclear role → preserve and ask); mechanics/safety rails per `worktree-isolation` |
63
63
 
64
- **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work.
64
+ **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work. **Authorization is never inferred**: a generic instruction ("跑下测试" / "run the pipeline"), a prior run's report, or the mere presence of working credentials/config for a mutating lane does not authorize that lane's mutations — the authorization must name the action category in the current task. Destructive cleanup of test/experiment resources is additionally scope-bound to the resources this run observed itself creating (registry/run-id based), never a name-pattern or global sweep.
65
65
 
66
66
  ## Minimal checklist
67
67
 
@@ -92,10 +92,10 @@ Use this skill to turn observed experience into durable agent skills without cop
92
92
  - Validation gate: a repeated routing miss of the same request type across two rounds is evidence the prior fix only patched the executor surface — re-run this coordinator check before claiming the routing class is closed.
93
93
  - **Routing-trigger fixes must cover the utterance-variant axis, not only the canonical phrasing.** When fixing a routing miss by adding a trigger to a `description`, the same delivery type arrives under many utterances — the canonical name PLUS restart/redo/from-scratch/continue variants (e.g. a *refactor delivery* arrives as "重构 X" but also "重新开发", "完全重新开始", "推倒重来", "清除代码重新开发", "redo/rewrite from scratch"); advertising only the canonical phrase leaves the variants unmatched, so the workflow silently fails to auto-trigger on them. "The rule/trigger exists but the utterance class is uncovered" is a validation-gate defect: enumerate the restart/redo/continue variants of the request type and add the high-value ones (within the 800-char cap; route overflow to the body entry-precedence text). A repeated miss of the SAME delivery type arriving via a different utterance across rounds is the signal this axis was skipped.
94
94
  - **In Claude Code, before fixing a routing miss with a trigger-word edit, check whether the skill's description was even IN the host listing — it may be budget-dropped.** A name-only entry silently voids a keyword-based routing fix (the skill stays reachable by explicit `/skill-name`). **Evidence bar (do not over-apply this as a catch-all):** conclude budget-dropped only from concrete evidence — this turn's listing shows that skill as name-only, `/doctor` output, or other host-listing proof; absent that, treat it as a hypothesis and STILL do the normal trigger / utterance-variant / coordinator fix (this rule does not replace them; a *visible* description misjudged as name-only is the reverse trap). Corollaries: (a) a routing/bootstrap doc assuming "every description is always visible" is wrong under budget pressure and must not be a low-traffic gate skill's sole discovery path; (b) on non-Claude-Code hosts verify the host's own listing behavior from its primary docs — any LLM's "it's a host bug" guess is hypothesis-grade until so checked. The budget mechanism, the cold-start trap, and the remediation levers: `references/skill-listing-budget.md`.
95
- - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings across many turns.
96
- - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it crosses into extraction only once the output is meant to change reusable skill behavior.
97
- - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim`, not landed — reconstruct the charter + target-output map and pass the dual-track before claiming it landed.
98
- - The benchmarked external packs are reference-only: route a missing capability that belongs to the method/tool layer to that pack; only land a ccl-layer rule (domain / governance / cross-cutting principle) here, never a copy of the external skill.
95
+ - **A deep review / benchmark of the skill repo is an extraction once it produces durable skill-change findings.** Reviewing or auditing the ccl-skills repo — especially benchmarking it against external/reference skill packs (`superpowers` / `gstack` / etc) — becomes an extraction the moment the output is meant to change reusable skill behavior (a target-output map, a gap list, or a remediation plan). Invoke `skill-extraction-workflow` and set the charter BEFORE the first findings/gap-report turn, not after accumulating findings.
96
+ - Boundary (avoid over-fire): a casual "看一下 / 这写得怎么样", a one-off opinion, or a pure code/doc PR review stays ordinary review; it becomes extraction once the output is meant to change reusable skill behavior.
97
+ - A multi-turn review that lands a change-plan without an upfront extraction charter is `interim` — rebuild charter + target-output map and pass the dual-track before claiming landed.
98
+ - The benchmarked external packs are reference-only: route a missing method/tool-layer capability to that pack; only land a ccl-layer rule (domain/governance/cross-cutting) here, never a copy of the external skill. Verdicts: P/I/M/W, grep-anchored; M needs the functional-equivalent check.
99
99
  - Treat remembered tool, script, installed-skill, repo-root, validator, and executable paths as stale until re-resolved in the current workspace. A path from memory, a prior-round plan, a compacted summary, shell history, handoff notes, or another agent's report must first be re-resolved against the current workspace (reopen the owning skill/reference or inspect the current repository), then probed for existence and executability before running it or reporting it missing. Prefer repo-local or skill-relative scripts only when they resolve inside the loaded skill package or trusted ccl-skills repository root, pass containment and no-symlink/hardlink trust checks, and are not merely same-named scripts in an arbitrary product repo or fork; otherwise fall back to a trusted installed-skill path or manual checklist. If the remembered path fails but a current-context trusted path succeeds, record the failed path source, fallback probe, resolved path class, trust check, and whether shared-file content validation was affected; do not classify shared skill content as broken only because a stale tool path failed.
100
100
 
101
101
  ### Owner-generalization, target-output & impact-chain mapping(owner / 目标映射 / impact-chain)
@@ -81,6 +81,10 @@ For teammates using the shared skill on description-based routing hosts, this is
81
81
 
82
82
  A quoted trigger phrase should belong to its skill in at least 80% of real-use contexts. If a phrase would commonly mean something else, it must be DROPPED or ANCHORED.
83
83
 
84
+ **Activation is closer to keyword match than semantic match** — a public sandbox measurement (2025-2026; sources/locators in the specs ledger) found prompts containing a skill's name or a distinctive description token activate near-100% while conceptual paraphrases activate near-0%; a second independent eval corroborates the keyword-dependence and adds that routing accuracy degrades as the installed-skill count nears ~20 similar skills, recovering when consolidated to ~12.
85
+
86
+ - Both findings are host/model/catalog-conditional: treat them as directional and do not rely on the numbers without reproducing against your own catalog. Two consequences for authoring: (a) the description must contain the distinctive tokens users actually type (measure real utterances, don't invent vocabulary — the discovery-vocabulary rule); (b) when routing degrades across the catalog, merging/pruning similar skills beats adding more trigger words to each.
87
+
84
88
  ### Drop (too generic, no rescue possible)
85
89
 
86
90
  - `"改下样式"` — almost always means "change CSS now", which is implementation, not design ownership. Drop.
@@ -416,6 +416,15 @@ item 9 signs off on it.
416
416
 
417
417
  Use `--json` to capture reasoning traces and tool calls cleanly. Parse the JSONL stream with a small Python or jq script as documented in `gstack-codex` skill.
418
418
 
419
+ ## Reviewer verification scope (packet-verifiability boundary)
420
+
421
+ The external reviewer judges what the bounded packet can show; the packet structurally cannot carry every deterministic oracle its acceptance claims depend on (frozen preservation-mapping rows, the checker's complete pinned-literal sets, whole-file postimages, immutable pre-fix revisions). The division of labor is fixed and documented here so it is ruled on once, not re-litigated per round:
422
+
423
+ - The reviewer owns CONTENT SEMANTICS: wording coherence, dropped qualifiers/obligations, source-accuracy, sanitization, scope drift — everything decidable from the packet plus the reviewer's own reasoning.
424
+ - Deterministic-gate claims (pinned literals present, size ratchet net-zero, obligation audit green, R0 clean, parity) are verified by the repository's CI re-running those gates on the actual branch — never by the reviewer, and never accepted from the implementer's prose alone; a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round.
425
+ - Historical-process claims (a pre-fix RED, a measurement taken before landing) are session-record-grade unless bound to an immutable revision or a candidate-bound receipt; treat them as the implementer's testimony, and say so in the disposition instead of demanding evidence the packet cannot hold.
426
+ - Standing backlog: teaching the gate to embed candidate-SHA-bound receipts of deterministic-gate output into the packet removes the third bullet's limitation mechanically; until that lands, this boundary is the accepted state.
427
+
419
428
  ## Recording findings + fixes
420
429
 
421
430
  For each pass, record in the extraction's working file (e.g. `<project>-extraction-summary.md`):
@@ -412,3 +412,52 @@ and inverted the sense (production, not product), and the coordinator now shares
412
412
  | Supersedes the identity-grammar and tag-layer clauses of the shared-scan supersede row: the metadata scan recognizes current cross-provider model words, session-task URLs, and only canonically-shaped tag objects | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh; firing-path: command:skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. The same external round surfaced: bare `GPT-*` identities and `Gemini * Flash/Ultra` model forms escaped every surface (now in the registry and unambiguous grammar with trailer regressions); Codex task URLs on session origins (`.../codex/tasks/<id>`) matched no session-path grammar (now matched, with a product-page near-miss control); a literal tag object could duplicate its `object` header so the scanner followed a decoy chain while Git peels the first target (canonical header shape now required, duplicates fail closed); the closeout validator's completion-receipt `schema_version` guard lacked the exact-integer type check applied everywhere else (a float `3.0` validated; now rejected with a regression), and `bounded_text` accepted C1 controls and U+2028/U+2029 line separators into diagnostics (now rejected with per-codepoint regressions). The second pass additionally rejected default-ignorable format characters that could split one recurrence class into visually identical keys, and bound every counted chain round to the exact ledger candidate rather than only the final receipt. |
413
413
  | Supersedes the fixture-portability posture of the plan-intent suite row: a mode assertion probes GNU stat before BSD stat | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_update_review_plan_intent.sh; result-class: failure | updated | `code-review/SKILL.md` is the owner key. GNU stat echoes unknown BSD directives (`%Lp`) verbatim with exit 0, so the suite's BSD-first probe never reached its GNU fallback and CI shard 2 failed `plan mode was not preserved` on every Linux run while macOS stayed green. The assertion now probes `stat -c` first and falls back to `stat -f '%Lp'`, mirroring the pattern `opencode_review.sh` already runs on both platforms. |
414
414
  | Supersedes the comparison-domain clause of the obligation-preservation audit: a repository-frozen ledger pins BOTH ends of its domain, so unrelated later changes owe it nothing | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh; firing-path: command:skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh; result-class: failure | updated | `skill-extraction-workflow/SKILL.md` is the owner key. The 065 obligation audit derived its row set from pinned-base..WORKING-TREE, so the first post-landing PR that rewrote any obligation line in any `skills/**/*.md` went red on preservation rows it never owed (observed: 12 phantom rows for this candidate's own code-review contract-paragraph rewrite). `obligation-ledger.py` now accepts `--head`, the ledger header pins `Head revision` beside the base, and the repo audit reads and requires it; carrier-drift detection still reads current files, so rewriting a BOUND carrier stays red (probed: the pinned audit still fails `CARRIER_COMPOSITE_NOT_UNIQUE` on a carrier rewrite). The synthetic suite adds the differential: a post-head non-carrier rewrite passes the pinned audit and fails the unpinned one with `ROW_SET_MISMATCH`. |
415
+ | Browser E2E business-success assertions must anchor on objective effects — weak proxy signals never suffice as the sole pass condition; billable metered-resource tests take an explicit lane with a named budget owner; the Playwright component-testing maturity claim is corrected against its primary source | `testing-strategy` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/references/e2e-real-flow-testing.md#weak proxy signals are not business assertions | updated | Owner key `testing-strategy/SKILL.md` (unchanged this round; the changes are merges into two references, no new bullets). (1) Weak-proxy blacklist (page loaded / URL changed / non-empty text / generic element visible / success toast alone) merged into the existing click-through clause in `e2e-real-flow-testing.md` §E2E Scope Control, pairing the already-present objective-effect positive list in §Backend/API Real Flows with its negative list; (2) billable-resource lane semantics (paid model inference, per-call third-party APIs, real payment flows, cloud sandboxes/device farms → explicit marker/lane, out of default PR and scheduled-frequent lanes, named budget owner plus authorized test account) merged into the expensive-test clause in `test-topology-and-commands.md` §Command Tiers, with credential/provisioning discipline routed to `ci-fixtures-and-flake-control.md`; (3) factual correction in `e2e-real-flow-testing.md` §Browser E2E Rules item (d), scoped to exactly what the excerpt establishes: Playwright's current component-testing guide supersedes the experimental add-on packages — primary source `https://playwright.dev/docs/test-components` (fetched 2026-08-30): "This guide replaces the experimental `@playwright/experimental-ct-react` and `@playwright/experimental-ct-vue` packages"; the sibling Biome correction rests on `https://biomejs.dev/linter/` ("a total of 526 rules", "many of them inspired from other linters" — the latter covers the skill text's rule-origin parenthetical) plus `https://biomejs.dev/blog/biome-v1-9/` ("CSS formatter and linter are now considered stable"; the prior GraphQL-timeline clause was removed from the edited skill text rather than carried beyond its excerpt), and the S3 note on `https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html` plus its upload subpage `.../checking-object-integrity-upload.html` (both fetched 2026-08-30: `CRC64NVME` "is the default checksum algorithm"; "A composite checksum is calculated based on the individual checksums of each part in a multipart upload" — establishing the composite branch the skill text describes; the per-algorithm support matrix lives on that subpage and is not restated here). RED-baseline (applied, differential; evidence scope: gate wiring and package integrity only, not per-clause semantics): on the committed candidate, `CCL_SKILL_BASE_REF=origin/dev check-ccl-skills.sh` ran green (`ccl_skill_check_clean_ok`, control); a probe commit deleting exactly this row turned the same run RED printing `impact_chain_gate_missing: upstream-owner skill changed without a matching source-register impact-chain row / missing evidence path: testing-strategy/SKILL.md` (attributable to the owning gate); restoring the row returned it to green — the oracle command is reproducible in-repo on any checkout of this candidate. Corpus-identifier scan over the five changed files (grep for the private source-vocabulary set) returned zero hits, and the repository R0 audit printed the clean private token in the same run. Clause semantics rest on this round's independent review and adversarial challenge rows. Scope of what THIS ROW asserts is deliberately narrow: the three inlined public excerpts above, the reproducible in-repo gate/scan runs, and nothing further — verification of other claims in the same round's diff is recorded in the maintainer's private round archive per the extraction-lifecycle handoff policy and is NOT certified by this row; a reviewer should judge those clauses against their own named public sources directly. Dual-track terminal record: multiple independent codex review rounds plus three full review+challenge×2 chains ran against successive candidates; every P0/P1 with a demonstrable failure path was applied (streaming retry/latch/resume/cancel semantics, canonical/hreflang interplay, submit latch, billable/payment-sandbox lanes, token-scan hardening, excerpt-scoping narrowings); the terminal chain's last three finding-fixes (cancel-vs-buffered-end, idle/heartbeat timeout, latch release on terminal state) were merged AFTER that chain per the repository's fixed-budget review-loop rule, so the exact landing candidate carries them un-rechallenged — listed for the merge decision-maker; the recurring bounded-packet self-certification findings (a packet reviewer cannot execute the in-repo oracle or see the private vocabulary) are dispositioned as verify-by-rerun: running the repository check script (`check-ccl-skills.sh` with `CCL_SKILL_BASE_REF=origin/dev`) on this candidate is the reviewer-executable oracle. Source class: local external engineering-skill corpus (sanitized to capability labels) plus public primary docs; sibling map: `web-react-dev` updated in the same round (its own package, not machine-gated), `product-ui-ux-design` unchanged (aesthetic-direction and anti-generic coverage already equal or stronger than the external source), `design-closed-contract-oracles.md` unchanged (criterion-to-proof already covered), `ci-fixtures-and-flake-control.md` unchanged (credential/provisioning face already covered). |
416
+ | Evidence-pipeline failure is a per-case third verdict: infra-error is never pass, never business-fail, never silent skip, and weaker surfaces cannot substitute for missing evidence | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#never a pass, never a business fail, and never a silent skip | `updated` | Owner key `testing-strategy/SKILL.md`. Benchmark round (private provenance alias: mediagen-platform) plus primary-source verification. RED baseline (replayed): `git show origin/dev:skills/testing-strategy/references/test-code-authoring-patterns.md` asserts the xUnit smell corpus as ~18 items while xunitpatterns.com lists 15 top-level (5/6/4), and the coverage floors carried no provenance — head corrects both and adds the infra-error verdict rule; baseline grep for the anchor line is zero-hit, head exactly one. |
417
+ | A continuation gate distinguishes multi-viable-approach and no-evidence-cause stops from a single dominant reversible path, which is carried through to a reviewable draft instead of stopping at a recommendation | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#do not stop at a recommendation | `updated` | Owner key `product-rd-workflow/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/product-rd-workflow/references/delivery-lifecycle.md` attributes DORA metrics' origin to the 2018 book while dora.dev/insights/dora-metrics-history dates the research line to 2014 — head corrects it; baseline zero-hit for the new stop/carry-through predicate, head exactly one. Review checklist gains guarantee grading and disabled-path walk (same package). |
418
+ | Query/lookup evidence reports 0, 1, or N matches distinctly; silently taking the first row of N is forbidden and an empty result over a named scope is itself evidence | `defect-diagnosis` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/defect-diagnosis/SKILL.md#never silently take the first row of N | `updated` | Owner key `defect-diagnosis/SKILL.md`. Source: external production skill-pack doctrine (alias mediagen-platform), mechanism verified as stack-agnostic. RED baseline (replayed): baseline grep for the cardinality predicate is zero-hit across the package, head exactly one — the baseline tree gave no instruction against first-row-of-N, so an agent following it could silently mis-resolve identity lookups. |
419
+ | Risk classification runs on the change's objective shape: reporter tone or executive pressure never escalates tags and a small diff never de-escalates them | `feature-risk-router` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/feature-risk-router/SKILL.md#must never escalate a change's tags | `updated` | Owner key `feature-risk-router/SKILL.md`. Baseline covered only the small-diff de-escalation side (low-risk intuition clause); the tone/pressure escalation side was absent — RED baseline (replayed): baseline grep zero-hit for the anti-escalation predicate, head exactly one. Six objective evaluation axes recorded inline. |
420
+ | Reader annotations are evidence of reading breakdown: classify the root-cause class first, then sweep the whole document for the same class instead of patching only the flagged sentence | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/annotation-driven-revision.md#必须全文扫同类位置一起修,不得只改被标记的那一句 | `updated` | Owner key `tighten-doc/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/tighten-doc/references/figure-and-table-craft.md` lists the 25-word sentence limit as no-reliable-source while GOV.UK's writing guideline states it as its house style — head regrades it to a sourced single-institution style and records the 40-char claim's traceable community source; baseline zero-hit for the annotation-sweep predicate, head exactly one. |
421
+ | Review reuse depth is a deterministic function of the candidate delta with no manual downgrade, the reviewed-identity record is a freshness guard not cryptographic proof, and relayed blocking findings pass a false-positive check with auditable dispositions | `code-review` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/code-review/references/manual-invocation-and-prompts.md#deterministic function of the delta, never a manual downgrade | `updated` | Owner key `code-review/SKILL.md`. Benchmark adjudication recorded: the source pack's receipt survives rebase/amend; our stricter voids-on-any-edit line is deliberately KEPT and only the tiering-determinism and honest-trust-model clauses are absorbed (keep-stricter per Conflict Resolution). RED baseline (replayed): baseline grep zero-hit for the determinism predicate, head exactly one. |
422
+ | Observation code never intrudes on the observed path, cross-layer conclusions require aligned identifiers or time, existing signals are discovered before new ones are created, and explicit environment targets always win | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/references/metrics-conventions.md#must not add retries, blocking waits, or business-logic branches | `updated` | Owner key `platform-observability/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/platform-observability/SKILL.md` attributes the good/valid SLI formula to Google SRE generically while the SRE Workbook's own text is good/total (the valid-events refinement is the Art of SLOs / GCP-blog line) — head corrects the attribution and self-labels the SLO ladder as team heuristic; baseline zero-hit for the non-intrusion predicate, head exactly one. |
423
+ | Benchmark rounds issue per-mechanism P/I/M/W verdicts with grep-anchored evidence, and an M verdict passes the functional-equivalent check before any borrow lands | `skill-extraction-workflow` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#requires the zero-hit grep recorded as replayable evidence | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Method absorbed from an external theory-audit corpus (alias mediagen-platform) whose own run found all 8 W-verdicts were internal drift rather than external knowledge gaps — the W-type internal-consistency sweep, disposition middle states (待实验/条件化采纳), and three conflict-synthesis shapes land in source-to-skill-extraction.md; keyword-activation evidence lands in description-authoring.md. RED baseline (replayed): baseline grep zero-hit for the functional-equivalent predicate, head exactly one. |
424
+ | A diff-scoped design review classifies every hard-coded visual-value hit into approved usage, pre-existing outside the change, or new violation — only new violations block, and pre-existing debt is routed, never blamed on the change | `product-ui-ux-design` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-ui-ux-design/references/ui-ux-audit.md#never reported as caused by this change | `updated` | Owner key `product-ui-ux-design/SKILL.md`. Changed refs: ui-ux-audit.md (diff-scoped three-bucket review), design-system-source-of-truth.md (definition-matrix completeness check), tokens-and-components.md (old code is not permission), platform-mobile-patterns.md (motion band self-labeled team heuristic per M3 tokens/Apple HIG verification; M3 Expressive research figures cited). RED baseline (replayed): baseline grep zero-hit for the three-bucket predicate, head exactly one. |
425
+ | Repo-pinned TC implementations pin the catalog revision, claim before implementing, and stop to fix a stale/contradictory source record before implementing against it | `test-artifact-management` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/references/update-lifecycle.md#停下先修源记录**,不得按旧版实现后事后补 | `updated` | Owner key `test-artifact-management/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/test-artifact-management/references/classical-test-design-techniques.md` states 2-way捕到 50–90% while NIST SP 800-142 Table 1 gives 53–97% with an explicit 10–40%+ miss warning, and tc-review-and-prioritization.md claimed a 30-year P×I consensus while ISTQB CTFL v4.0.1 §5.2 defines multiplication as the quantitative approach beside a qualitative matrix — head corrects both; baseline zero-hit for the revision-pin predicate, head exactly one. |
426
+ | Evaluation records keep human-review and machine fields with distinct writers, normalize cross-provider token accounting before comparison, and reconcile usage against billing where available | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#must never overwrite a human verdict | `updated` | Owner key `llm-inference-integration/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/llm-inference-integration/references/model-prompt-evaluation.md` states extended thinking is off by default per docs while the current platform docs deprecate the manual mode on 4.6, reject it on 4.7+, and default thinking on for the Claude 5 family — head rewrites the claim per model generation; prompt-caching modes and the provider-evaluation evidence ladder (spend contract, seven layers, open-loop load per NSDI'06) added in the same package. |
427
+ | Durable-state transitions capture the clock once and validate external-response structure with a typed error path before mapping | `nodejs-service-dev` | result-class: stable-success; behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/nodejs-service-dev/references/async-lifecycle-and-performance.md#never a crash or a silently-defaulted field | `updated` | Owner key `nodejs-service-dev/SKILL.md`. Sibling-parity landing with the go/python state-machine references (same two predicates, stack-idiomatic wording). RED baseline (replayed): baseline grep zero-hit for the predicate, head exactly one. |
428
+ | Vendor/ecosystem status claims carry their verified level and date — the RPC-alternative entry records CNCF sandbox status and the accurate interop-test wording instead of an inflated maturity tier | `go-microservice-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/go-microservice-architecture/references/architecture-playbook.md#CNCF **sandbox** project (per connectrpc.com | `updated` | Owner key `go-microservice-architecture/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/go-microservice-architecture/references/architecture-playbook.md` asserts CNCF-incubated and Google-validated interop while connectrpc.com states sandbox level and self-run extended interop tests — head corrects both; crypto-erase citation upgraded to SP 800-88 r2 + EDPB 02/2025 in the same package. |
429
+ | Crypto-erase evidence conditions cite the current sanitization standard revision and regulator guidance rather than a withdrawn revision | `python-service-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/python-service-architecture/references/multi-tenant-isolation.md#do not use CE when data predates encryption enablement | `updated` | Owner key `python-service-architecture/SKILL.md`. RED baseline (replayed): baseline cites NIST SP 800-88 generically (Rev.1 withdrawn 2025) with no CE-condition specifics; head cites Rev.2's explicit do-not-use conditions and EDPB 02/2025's encrypted-data-is-still-personal-data holding — mirrored with the go sibling. |
430
+ | Gitflow-style release trains use direction-sensitive merges (squash only feature→develop), prompt dual back-merge, human-confirmed tags, and full-ladder hotfixes; canary template thresholds are labeled as tool example values, and user-bucketed ramps follow exposure/funnel/salt data-literacy rules | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#must not squash** — prefer fast-forward/plain merge | `updated` | Owner key `platform-release-engineering/SKILL.md`. RED baseline (replayed): `git show origin/dev:skills/platform-release-engineering/references/canary-and-rollout-strategy.md` presents the 1%/500ms thresholds as bare template defaults with no provenance, while Flagger's builtin checks carry exactly those example values and Argo Rollouts' official example uses 95% — head attributes and scopes them; the merge-topology section is grounded in AWS Prescriptive Guidance's Gitflow pattern (verified 2026-08). |
431
+ | Verdict assignment is decided by fault origin with evaluated-false as business fail, and required infra-error cases block aggregate readiness | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/references/ci-fixtures-and-flake-control.md#never by defaulting to whichever verdict looks better | `updated` | Owner key `testing-strategy/SKILL.md`. Fix-round rows for the review-chain findings: the pre-fix span (replayed via `git show` at the round base) carried the enumerated verdict mapping whose edges four review rounds broke in turn; the fault-origin predicate replaced the enumeration and the aggregate-readiness clause closed the clean-total false green. |
432
+ | Design-review exceptions are approved only by the design-system owner's recorded, unexpired, usage-covering approval; age is not approval; scan categories cover size/layout and motion | `product-ui-ux-design` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-ui-ux-design/references/ui-ux-audit.md#age is not approval: reusing or extending an old undocumented/expired exception | `updated` | Owner key `product-ui-ux-design/SKILL.md`. Fix-round: the pre-fix predates-arm allowed laundering an old undocumented/expired exception through a changed hunk (challenge-confirmed bypass, replayed at the round base); the arm was deleted (convergence-by-deletion) and the definition matrix aligned with the governed-category list. |
433
+ | TC claiming and write-back both ride atomic revision-conditioned updates; hash-pin mismatch always stops for source reconciliation | `test-artifact-management` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/references/update-lifecycle.md#停下改走人工/单线分配,不得按回读结果继续 | `updated` | Owner key `test-artifact-management/SKILL.md`. Fix-round: read-back-after-write was challenge-proven non-CAS (later writer silently replaces a verified claim, replayed at the round base); the landed rule requires platform CAS/optimistic-lock or a serialized coordinator, else concurrent claiming is unsupported and stops. |
434
+ | Relayed findings pass an existence-and-severity check, unverifiable stays blocking, and reuse voids on base/profile/lens change | `code-review` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/references/manual-invocation-and-prompts.md#a confirmed-but-advisory or overstated-severity issue is relayed at its true severity | `updated` | Owner key `code-review/SKILL.md`. Fix-round: challenge rounds showed the fp-check validated existence only (severity inflation passed) and the unverifiable disposition could demote real defects; both closed, with the reuse boundary extended to base/profile/lens changes. |
435
+ | Auto-continue is subordinate to every stop condition, the metered-account carve-out keeps its existing/configured/self-use qualifiers, and stop wording preserves the pinned anchors | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#report interim or blocked with the next unblock step | `updated` | Owner key `product-rd-workflow/SKILL.md`. Fix-round: size-offset compression had silently widened the metered-account carve-out and detached carry-through from the stop list (review-confirmed, replayed at the round base); qualifiers restored, precedence bound, and checker-pinned anchor phrases restored after the pinned-phrase gate fired. |
436
+ | Vendor-standard citations are layered honestly: CE conditions cite Rev.1 §2.6 with Rev.2 superseding and continuing the framework | `go-microservice-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/go-microservice-architecture/references/multi-tenant-isolation.md#verify the corresponding Rev.2 section when citing it as the authority | `updated` | Owner key `go-microservice-architecture/SKILL.md`. Fix-round: the earlier candidate attributed CE do-not-use conditions to Rev.2 while the verification ledger located them in Rev.1 §2.6 (review-caught attribution mismatch); the landed text layers the citation and instructs Rev.2 section verification before citing it as authority. |
437
+ | The python sibling mirrors the layered CE citation byte-for-byte per the parity discipline | `python-service-architecture` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/python-service-architecture/references/multi-tenant-isolation.md#verify the corresponding Rev.2 section when citing it as the authority | `updated` | Owner key `python-service-architecture/SKILL.md`. Fix-round mirror of the go row above; parity gate keeps the mirrored region byte-identical. |
438
+ | Node's malformed-payload rule mirrors the go/python persisted-failure-transition semantics and single-now scopes to transition-stamped timestamps | `nodejs-service-dev` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/nodejs-service-dev/references/async-lifecycle-and-performance.md#mirroring the go/python state-machine rendering | `updated` | Owner key `nodejs-service-dev/SKILL.md`. Fix-round: review caught the node wording drifting weaker than the go/python rendering (policy-handled vs persisted failure transition) and the single-now literalism re-stamping domain-provided times; both aligned. |
439
+ | Best-effort observation is scoped to diagnostic telemetry; audit/billing/deletion/release-gate records are business writes that never fail open | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/references/metrics-conventions.md#their loss is a failure, never shrugged off as telemetry | `updated` | Owner key `platform-observability/SKILL.md`. Fix-round: review showed the unbounded best-effort license could excuse a mandatory record failing open; the boundary now names the mandatory-record classes and their durable-delivery semantics. |
440
+ | Annotation-driven revision scans the whole document but edits only within authorization, and factual/citation errors route to source verification | `tighten-doc` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/references/annotation-driven-revision.md#不得以「修同类」为名越权改动已定内容 | `updated` | Owner key `tighten-doc/SKILL.md`. Fix-round: review caught the unconditional same-class sweep expanding mutation past the approved scope and the root-cause classes omitting factual/citation errors; both landed with the scan-wide/edit-scoped split. |
441
+ | Keyword-activation guidance carries its sources and conditionality, and benchmark-figure reliance requires local reproduction | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/description-authoring.md#do not rely on the numbers without reproducing against your own catalog | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Fix-round: review flagged the activation claims as unsourced-in-ledger and unconditional; sources landed in the source-verification ledger and the conditionality bullet forbids relying on the numbers without local reproduction. This round also carries the crypto-erase supersede note above. |
442
+ | Hotfix back-merge covers every open release train and canary thresholds carry their tool-example provenance | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#an active train missing the hotfix reverts it at that train's own merge | `updated` | Owner key `platform-release-engineering/SKILL.md`. Challenge-round: the earlier back-merge wording covered main and develop but not an open release branch (replayed at the round base), so a concurrent train could revert a shipped hotfix; the funnel invariant was also scoped as sanity-not-attribution and the error-budget comment renamed to a rolling error-rate threshold. |
443
+ | Benchmark M-verdict evidence pins command, scope, and baseline revision so replays run against the recorded baseline | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#replay runs against the recorded baseline | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Challenge-round: a NO_HITS record without its baseline revision self-hits once the borrow lands (review-caught, replayed at the round base); this round also relocates the ledger's supersede note below the row table so machine and human readers see one uninterrupted row stream. |
444
+ | Provider-evaluation spend is enforced at admission with spent-plus-in-flight headroom and billed attempts missing usage count against the provider | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#its cost enters via billing records, never silent exclusion | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: delayed usage/billing signals let an open-loop driver overshoot the cap and a provider omitting usage on billed failures flattered its ratios (replayed at the round base); the admission budget now subtracts recorded spend plus in-flight worst case, the stop latch halts admissions, and ratio exclusions are reported. |
445
+
446
+ | Spend enforcement is a reservation invariant (cap ≥ reconciled spend + outstanding reservations) and ratio exclusion never touches cost totals | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#total-cost accounting is the separate aggregate that never excludes | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: three successive edge findings on the admission-budget arithmetic converged by replacing the enumeration with the reservation invariant (reserve worst-case at admission, release only on billing reconciliation), and the ratio-exclusion clause was split from total-cost accounting so neither aggregate can be flattered by usage omission. |
447
+
448
+ | The billing-records anchor phrase is preserved inside the split-aggregate wording so prior ledger locators keep resolving | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/model-prompt-evaluation.md#never silent exclusion from totals | `updated` | Owner key `llm-inference-integration/SKILL.md`. Locator-repair round: the split-aggregate rewrite had dropped the substring an earlier row anchors on (register_firing_path_unresolved, replayed at the round base); the phrase is restored within the new semantics so both locators resolve. |
449
+
450
+ | Spend reservations are atomic check-and-reserve on one ledger, closing the concurrent-headroom TOCTOU | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#read-headroom-then-reserve is a TOCTOU that lets concurrent workers jointly overshoot | `updated` | Owner key `llm-inference-integration/SKILL.md`. Challenge-round: two open-loop workers reading the same headroom could each reserve and jointly exceed the cap (replayed at the round base); the reservation is now an atomic check-and-reserve, completing the spend-invariant class alongside admission-halt and billing-release. |
451
+
452
+ | The hotfix full-ladder rule names the emergency-override section as its one sanctioned, loudly-logged exception | `platform-release-engineering` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-release-engineering/references/promotion-gate-and-review.md#the emergency-override section below is the one sanctioned, loudly-logged exception | `updated` | Owner key `platform-release-engineering/SKILL.md`. Review-round: the unconditional never-skips-a-gate wording contradicted the file's own emergency-override section during a production incident (replayed at the round base); the exception is now named inline so the two sections compose instead of conflicting. |
453
+
454
+ | Sub-threshold sample counts are reported as unreliable, never averaged into a verdict | `llm-inference-integration` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/inference-capacity-operations.md#must be reported as unreliable, never averaged into a verdict | `updated` | Owner key `llm-inference-integration/SKILL.md`. Review-round: the ~3-run floor carried no provenance label (replayed at the round base); it is now explicitly a team heuristic with a raise-per-variance instruction, and the rule moved to its own normative bullet. |
455
+
456
+ | The latency-SLI formula divides by the valid set, matching the good/valid definition in the same rule | `platform-observability` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/SKILL.md#the denominator is the availability SLI's valid set, never raw total | `updated` | Owner key `platform-observability/SKILL.md`. Challenge-round: the bullet preferred good/valid while its own latency formula divided by raw total (replayed at the round base); the denominator now names the valid set explicitly. |
457
+ | Verbatim-bound obligation carriers survive wording compression only verbatim: a size-budget compression must first check the sentence against the frozen preservation mapping, and byte offsets come from sentences the same round added | `testing-strategy` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#Do not describe such work as done, fixed, merge-ready, or release-ready | `updated` | Owner key `testing-strategy/SKILL.md`. Observed failure: two wording compressions in this round's size-budget offsets rewrote exact carrier sentences bound by the specs/065 obligation mapping (such-as→like; are-done→pass), and the heavy lane's real-repository obligation audit went red (CARRIER_COMPOSITE_NOT_UNIQUE count=0) while every entrypoint-scope gate stayed green — the preservation mapping is a verification surface the size-budget workflow did not consult. Fix: carrier sentences restored verbatim; equal-byte offsets taken from sentences this branch itself added (which the frozen mapping cannot bind); reader index regenerated at the mapping's pinned base/head. RED baseline (replayed): the repo-audit suite for the obligation ledger (`test_obligation_ledger_repo_audit.sh`) red before the restoration, `audit_ok domain=50 rows=1240 unresolved=0` plus `test_obligation_ledger_repo_audit_ok` after. |
458
+ | A wording compression that deletes a sentence boundary corrupts the hosting rule: the reinserted stop/continuation sentence must keep its full punctuation, and a fresh-eyes review of the reinsertion is what catches the truncation | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#explicit stop/pause needs no reconfirmation. A `continuing:` outcome | `updated` | Owner key `product-rd-workflow/SKILL.md`. Observed failure: the size-budget reinsertion of pinned continuation literals dropped a sentence boundary, leaving "needs no reconfirm A `continuing:` outcome" — two rules joined without punctuation on the entrypoint's stop/continue surface; caught by the supplementary post-delta review round (P1), invisible to every deterministic gate because pinned-phrase gates match their own literals only. Fix: boundary restored ("no reconfirmation. A"). Offset provenance, stated exactly: this entrypoint's continuation-gate block is wholly branch-rewritten (six modified base lines, no pure additions), so offsets necessarily live inside that rewritten block; a word-level diff against the base revision audited every token those compressions dropped — the two load-bearing drops it surfaced (metered model/tool scope; the missing-capability routing qualifier) are restored in their owners' rows, the ambiguity-or join was additionally reverted with its replacement byte taken from a demonstrably branch-added sentence (em-dash tightened to a colon in the dominant-approach clause), and the remaining list joins are recorded as audited-neutral. Where genuinely branch-added sentences exist, offsets come from them first. The truncation shape joins the round's carrier-restoration lesson: compression edits need a substring check against the frozen preservation mapping and plain sentence-boundary integrity. RED baseline (replayed): repo-wide grep for "needs no reconfirm A" one hit before the fix, zero after; the restored phrase greps exactly once. |
459
+ | A compression that drops a scope qualifier widens the rule it hosts: the external-pack routing clause routes only a MISSING method/tool-layer capability, and restoring the dropped word is the fix, not rewording around it | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#M needs the functional-equivalent check | `updated` | Owner key `skill-extraction-workflow/SKILL.md`. Observed failure: a size-budget compression rewrote the reference-only routing clause from routing a missing method/tool-layer capability to routing that layer categorically, which would displace locally covered P-verdict capabilities and contradict the functional-equivalent check landed in the same round; caught by the supplementary post-delta challenge round enumerating compressed sentences (same class as the metered model/tool qualifier drop fixed in the sibling owner). Fix: the missing-capability qualifier restored; byte offsets from this branch's own pointer sentence, whose semantics live in the owning reference. RED baseline (replayed): word-level diff against the base revision showed the dropped qualifier before the fix and shows it restored after; the class sweep over all three owner entrypoints found no further load-bearing drops. |
460
+ | Stop reporting keeps its specificity qualifiers: the entrypoint demands the concrete stop reason and the exact evidence checked, and budget offsets come from relocating clauses whose semantics already live verbatim in the owning reference | `product-rd-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#state the concrete stop reason and the exact evidence checked | `updated` | Owner key `product-rd-workflow/SKILL.md`. Observed failure: a size-budget compression dropped the-concrete/the-exact from the stop-reporting sentence, licensing generic stop reports; the challenger graded it load-bearing, the implementer's word-sweep had graded it neutral, and the maintainer's standing delegation resolves such token disputes by the repository's fail-closed obligation standard, so the qualifiers are restored. Byte offset: the stale-source parenthetical is removed from the entrypoint because its full sentence lives verbatim in references/pre-final-continuation-gate.md (Status-source reconciliation) which the same sentence already cites — relocation, not compression. RED baseline (replayed): grep for the restored phrase zero-hit on the pre-fix entrypoint, exactly one hit after; the removed parenthetical greps once in the owning reference. |
461
+ | The dual-track reviewer's verification scope is a documented boundary: content semantics belong to the reviewer, deterministic-gate claims to CI, historical-process claims are testimony unless receipt-bound — ruled on once so packet-verifiability findings stop recurring per round | `skill-extraction-workflow` | result-class: failure; behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#a finding that only restates this boundary is dispositioned against this rule, never re-litigated per round | `updated` | Owner key `skill-extraction-workflow/SKILL.md`; the change lands in references/dual-track-review-gate.md (new Reviewer verification scope section). Observed failure: the packet-verifiability finding class recurred across four supplementary review rounds and roughly a dozen occurrences in this round's chains — every reviewer independently rediscovered that the packet cannot carry the deterministic oracles, and every round paid the same finding again because the boundary was undocumented. The maintainer confirmed the operating reality (all consumers and reviewers are agents; the human role is authority, not readership), so the boundary is now standing text agents can disposition against, with the receipt-embedding backlog item named in place. RED baseline (replayed): grep for the boundary phrase zero-hit before this change, exactly one hit after, on an added normative list line. |
462
+
463
+ Supersede note (this round, before landing): the two crypto-erase rows above ("go-microservice-architecture" and "python-service-architecture") describe an earlier candidate state; the landed text attributes the CE do-not-use conditions to SP 800-88 Rev.1 §2.6 with Rev.2 (2025) superseding and continuing the framework — per the source-verification ledger row "CE conditions text location". The rows' RED-baseline probes and firing-path anchors are unaffected.
@@ -677,3 +677,15 @@ Use these checks before treating a source-derived skill as ready:
677
677
  - New skill: create only when the future user would naturally ask for a different task type.
678
678
  - Reference file: use for detailed checklists, variants, examples, and source-derived heuristics.
679
679
  - Script: use for repeatable validation, scanning, generation, or formatting that should be deterministic.
680
+
681
+ ## Benchmark Verdict Discipline
682
+
683
+ For an external-pack / sibling-repo benchmark round (the `SKILL.md` benchmark Core Rule owns when this fires), the per-mechanism verdict table replaces impressionistic gap lists. Each source mechanism gets one row:
684
+
685
+ - **Verdict, four values**: `P` (we cover it correctly — requires a decision point, not just the concept named), `I` (covered but incomplete — name the missing branch), `M` (missing — but see the functional-equivalent check), `W` (we state it wrong — quote both sides with file:line and propose the fix).
686
+ - **Grep-anchored evidence**: every row carries the grep terms used against our tree; an `M` verdict requires the zero-hit grep recorded as replayable evidence — terms plus the exact command/mode, path scope, and baseline revision (`NO_HITS: <terms> @ <rev>`), because once the borrow lands, a replay against HEAD hits the landed rule itself; replay runs against the recorded baseline.
687
+ - **Functional-equivalent check before any `M` borrow**: a missing *name* is not a missing *capability* — search for what our tree does in that situation under other vocabulary; when an equivalent exists, the verdict is `P(功能等价)` or `I`, and the borrow shrinks to naming/linking value (usually low). This is the mechanical guard against borrowing-for-the-sake-of-alignment.
688
+ - **Borrowing degree** per `M`/`I`: high / medium / low, scored on one axis only — how much the mechanism changes an agent's actual decisions — never on how impressive the source text reads.
689
+ - **Disposition vocabulary**: beyond keep/merge/discard/route, two middle states are legitimate and prevent binary misjudgment — `待实验` (hypothesis worth testing: record the hypothesis, the experiment, and pass/fail criteria; do not land prose) and `条件化采纳` (adopt with explicit applicability conditions, disable conditions, and a fallback path recorded next to the rule).
690
+ - **Conflict synthesis, three shapes before choosing sides**: when the source contradicts our rule, first test whether the conflict is (a) a missing tier — both are right in different regimes, so merge by adding the regime split; (b) a missing ordering — both rules survive with an explicit precedence; (c) conflated authority — one side owns proposing and the other owns deciding, so split the powers. Only when none fits is it a genuine pick-one decision (keep the stricter per Conflict Resolution).
691
+ - **W-type internal-consistency sweep**: a mature skill tree's highest-yield audit is internal drift, not external comparison. The recurring W classes to grep for: an absolutized single-case rule (one scenario promoted to an unconditional "always/never"), same-name-different-meaning across files, translation/paraphrase drift between sibling-language files, a template violating its own stated hard rule, stale-era residue (rules about tools/versions no longer in the stack), a hard-coded value where a policy/parameter reference should be, and version-lagged claims about fast-moving vendors. Route wording-level dedup to `tighten-doc`; land each confirmed W as a fix with both sides cited.
@@ -13,7 +13,7 @@
13
13
  - **James Bach & Michael Bolton**, *Rapid Software Testing* — HTSM / SFDPOT
14
14
  - **Elisabeth Hendrickson**, *Explore It!* — test heuristics cheatsheet
15
15
  - **James Whittaker**, *Exploratory Software Testing* — tours
16
- - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(多被测系统 2-way 捕到 50–90% 的缺陷,差异大);工具:Microsoft PICT、NIST ACTS、Hexawise
16
+ - **Pairwise / Combinatorial**: **Kuhn / Wallace / Gallo 2004 NIST 实证**(NIST SP 800-142 Table 1 复现其数据:各被测域 2-way 累计触发 53–97%,多数域 70–97%;NIST 同文提醒 pairwise 仍可能漏掉 10–40% 或更多缺陷,mission-critical 不足恃);工具:Microsoft PICT、NIST ACTS、Hexawise
17
17
  - **Hans Buwalda 2004** — soap opera testing
18
18
  - **Lisa Crispin & Janet Gregory**, *Agile Testing*
19
19
  - **Glenford Myers**, *The Art of Software Testing* (1979) — error guessing 起源
@@ -61,7 +61,7 @@ SKILL.md 现状 P0 / P1 / P2 是经验判断("blocking / important / nice")
61
61
 
62
62
  ### 2.1 风险公式(最通用)
63
63
 
64
- **Risk = Probability × Impact**(业内 30 年共识,ISTQB / ISO 29119 同源)
64
+ **Risk = Probability × Impact**(出处:ISTQB CTFL v4.0.1 §5.2——风险级别由 likelihood 与 impact 决定,**定量法**为二者相乘,**定性法**用风险矩阵,二者皆合规;ISO/IEC/IEEE 29119-1 采用类似 likelihood/impact 框架,原文付费墙未逐字核)
65
65
 
66
66
  **Probability**(发生概率,1-5 分):
67
67
  - 1 = 罕见(新代码 + 简单逻辑 + 测试覆盖好)
@@ -49,6 +49,8 @@ python test/scripts/gen_report.py \
49
49
  - md 同步匹配必须以 用例ID 为唯一键;模块名或功能点相同不代表是同一条记录
50
50
  - 若 md 先于 Bitable 被修改(如直接编辑文件),需将 md 变更反向同步到 Bitable,再按 update 工作流补信息流转
51
51
  - **漂移检查**:定期运行 `python gen_report.py --config test/.report-config.json --diff-md test/cases/all.md` 显示 md 与 Bitable 的 added/removed/changed;废弃记录自动排除(不在 md 里属正常);多人编辑后必跑一次再 commit
52
+ - **版本钉扎**:当仓内测试代码/脚本按某条 TC 实现时,在实现侧记录该 TC 的修订标识——用**定义字段的规范化快照哈希**(或 Bitable 的不可变 record 修订号,若可得);`[姓名 日期]` 不够(同人同日二次修订会撞标识,钉住旧版仍比对相等)。实现前先比对,**回写/提测前再验一次,且回写本身走与认领相同的修订条件更新**(先验后写仍是两步——验证通过与写入之间的并发修改只有条件写能挡)——哈希只有相等性判断:**任何一次不一致都视为分歧,停下、先做源对账**(更新钉扎侧并使差异可评审后再继续);「更新/更旧」的说法仅当平台提供不可变单调修订号时才可用;发现源记录 stale/自相矛盾时**停下先修源记录**,不得按旧版实现后事后补
53
+ - **认领防冲突**:多人/多 agent 并发按 TC 实现时,认领必须是**原子的修订条件更新**(仅当记录修订仍等于读取时的修订才写入——平台支持的 CAS/乐观锁语义)或走**串行化认领协调者**(单写入口)。「读空→写」是 TOCTOU;「写后回读」也不等价——A 回读成功后仍可被 B 顶掉、双方各自都验证通过。两种安全机制都不可得时,**并发认领在该表上不受支持:停下改走人工/单线分配,不得按回读结果继续**。字段已有他人活跃认领时不得覆盖,改为联系认领人或换条目
52
54
  - `+record-upsert --record-id` 中的 record_id 是 Bitable 内部 ID,必须先用 `get_record_index()` 从 用例ID 查出,不能直接用 用例ID 代替
53
55
 
54
56
  ## update 工作流
@@ -71,25 +71,25 @@ Use this skill to decide whether a specialized or non-functional test belongs in
71
71
  - Conditional skips (missing-optional-dependency guards such as module-level `importorskip`, platform/env markers) combined with per-job test selection can leave an entire test file executed in NO CI job while every pipeline stays green: the job that selects the file lacks the optional dependency (the skip fires for the whole module), and the job that has the dependency does not select the file. When a suite mixes conditional skips with job-scoped test selection, the job that owns those tests must carry an executed-count guard — the per-file invariant and the floor fallback live in `references/ci-fixtures-and-flake-control.md`. Any change to job-level selection re-verifies which files each job actually executes (run with skip reporting and read the executed/skipped counts per file). "The tests exist and CI is green" is not evidence they ran anywhere.
72
72
  - A new, ported, or mirrored enforcement mechanism (pre-edit hook, permission guard, write-blocking plugin, validator) is not verified by loading, parsing, or config inspection — those prove installation, not enforcement. Require a behavioral matrix before a completion claim — blocked case per deny-condition, allowed/no-collateral cases, the fail-open/swallowed-exception bypass set, and the canonicalization/symlink/worktree edge cases; the matrix cells, the fail-open bypass set, the safe-unavailable-gap disposition for a case that cannot be exercised safely, and the port/mirror parity procedure live in `references/verify-enforcement-mechanisms.md`. Run the matrix only against scratch/synthetic targets (a throwaway checkout/worktree, fixture repo, or dry-run mode) — never a live workspace, real user data, or live credentials. Parity claimed from code reading alone is hypothesis-grade, not evidence. When a test is **ported/mirrored to a sibling stack**, input-fixture fidelity is part of that parity — see `references/test-code-authoring-patterns.md` (跨栈移植:移植对抗输入本身).
73
73
  - A "skip CI" / "no runner" instruction does not by itself lower verification rigor, only ceremony. When the blocking CI gate is skipped or unavailable, substitute a same-risk independent check before treating the change as verified — for a tiny/doc/test-only change a local command or `diff --check` is enough; for a change that can break its own gate (it edits the test/tripwire/CI config it is guarded by), an adversarial review/challenge of the diff is what catches the self-break CI would have. Note where a local run is not equivalent to CI (secrets, OS matrix, merge-result pipeline) rather than treating it as full proof. Separately, a project-enforced merge gate (pipeline-must-pass, required review) is not waived by a "skip CI" instruction for convenience: require green status, or an explicit authorized break-glass/override with recorded reason + residual risk — surface the conflict and stop rather than silently bypassing.
74
- - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture or inspect console errors and failed network requests when the browser tool supports it.
74
+ - Browser/E2E smoke must assert visible outcomes, not just click controls. For frontend API pages, verify loading, success, failure, disabled/retry behavior, and absence of dangerous actions where relevant. Capture console errors and failed network requests when tools support it.
75
75
  - UI tests and screenshots must prove design quality layers, not only DOM existence. Assert or visually inspect aesthetic hierarchy/density, interaction path, behavioral recovery states, and psychology-critical cues such as disabled reasons, progress certainty, retry safety, confirmation consequences, and return context.
76
76
  - For every runtime-visible UI/UX slice, load the canonical sequence in `../product-ui-ux-design/references/delivery-contract.md` and `references/client-runtime-test-matrices.md` §UI/UX Delivery Contract before Phase 0 and after producer/client execution. Testing owns layer selection and sufficiency, binds and cites the complete design/test/producer/client record and candidate-binding sets, confirms every affected client wrote its canonical pre-edit `client_entry` and complete client-record member naming the producer version it exercised, fails closed on a missing/incomplete/mismatched/stale/changed-after-run/unexercised member, and never issues the holistic design verdict.
77
77
  - Authentication and account surfaces need an explicit scenario matrix before they can be called complete. Cover identity-input validation across relevant entries, available sign-in methods, registration, account recovery or password reset/change, logout/account switching, sensitive storage/log cleanup, permission or host-authorization denial, and UI/UX acceptance for error copy, disabled reasons, keyboard/safe-area/touch behavior, and visual evidence. If a capability such as recovery, host authorization, real message delivery, or live account verification is absent or external, record it as `product gap`, `blocked`, or `live-only` instead of silently excluding it from the test claim.
78
- - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the reproducible command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
79
- - Do not answer "tests are complete" from command output alone. The claim must map each important scenario to a written case or to an explicit `blocked`, `live-only`, `product gap`, or `not applicable` row.
78
+ - Test cases come before implementation and broad execution for behavior-changing work. Write a compact test-case register first: scenario, layer, assertion, data/dependency, command, expected current result (`fail`, `pass-existing`, `blocked`, `infra-error`, or `gap`), and owner. For bug fixes and user-visible or contract-visible behavior, at least one relevant case must be added or updated and run RED before implementation unless no harness can support it after normal remediation; then record the evidence gap and strongest alternate check. The same RED-first discipline applies to defect records: a reported defect (issue, QA finding) carries the repro command plus the actual failing output, and the fix change references that failing test — a bug "fixed" from its description alone, without a RED reproduction, is unverified.
79
+ - Do not answer "tests are complete" from command output alone. Map each important scenario to a written case or an explicit `blocked`, `live-only`, `product gap`, `infra-error`, or `not applicable` row. (verdict definitions and their mutual exclusivity: `references/ci-fixtures-and-flake-control.md`).
80
80
  - For report-only QA, baseline comparison, or "testing only" branches with no product-code changes, failing tests can be the intended deliverable.
81
81
  - This exception applies only when a human reviewer, PR owner, or user explicitly states in the current work item, PR description, or current-turn context that the deliverable is test coverage, evidence, or a QA report rather than a product fix; an agent or automated process cannot infer or self-apply this exception from prior-session memory or summarized context.
82
82
  - Do not weaken the test or patch product code just to go green.
83
83
  - For disputed or high-stakes defects, prefer splitting verification and fix into two deliverables: a test-only verification slice first pins the confirm/deny verdict and root-cause attribution, and the fix is a separate change that references it — keeping the verification verdict uncontaminated by fix intent. Its failing regression test may land skip-marked as the trace only as a bounded state, not an escape: the skip carries the reason, an owner, and the linked fix item, and accepting the fix requires un-skipping it into the blocking regression set (or an explicitly owner-signed quarantine lane) — a RED test that stays skipped after its fix merges is the bypass this rule exists to prevent.
84
84
  - The QA report is not complete until pass evidence, red-light evidence, blocked/live-only gaps, and baseline comparison are each present and non-empty or explicitly marked `not applicable`; red-light evidence and baseline comparison cannot both be `not applicable`.
85
- - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner and do not resolve it by trusting either side.
85
+ - Live or production behavior is environment evidence, not a correctness oracle: when live behavior contradicts automated or documented expectations, record the discrepancy as a `live-only gap` with owner; do not resolve it by trusting either side.
86
86
  - A QA report with an open live-contradiction gap is `blocked` until the discrepancy is escalated and an owner assigns a resolution path.
87
- - Generated starter tests are not regression evidence: replace scaffold placeholders (e.g. a counter widget test) in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
88
- - Verification warnings are not automatically follow-up work: classify build/bundle-size/lint/flaky/deprecation/security/performance warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
89
- - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix has been written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are a release risk, not an afterthought.
87
+ - Generated starter tests are not regression evidence: replace scaffold placeholders in the same delivery slice with assertions for the actual app shell, route, state, or user-visible contract.
88
+ - Verification warnings are not automatically follow-up work: classify build/bundle/lint/flaky/deprecation/security/perf warnings from a required gate before reporting success — fix now when caused by the current slice or cheaply local; defer only with reason, residual risk, owner, and follow-up artifact.
89
+ - Do not open, merge, or describe an MR as ready for a contract-visible change until the test matrix is written and executed, or each unavailable layer is explicitly marked unavailable with reason and residual risk. The matrix must include the relevant unit, API/contract, integration, and browser/device/E2E layers; missing layers are release risk, not afterthought.
90
90
  - A multi-stack development-standard family is incomplete without a testing standard. The testing standard must define test deliverables, layer policy, harness expectations, CI gates, high-risk coverage, evidence format, and stack handoff rules; stack docs may specialize commands but must not redefine the layer policy.
91
- - Do not mark a browser/device/E2E layer unavailable just because discovery returns no device, browser, server, or dependency. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
92
- - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; this is a handoff-only label, not merge-ready or release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, ready to merge, or ready to release.
91
+ - Do not mark a browser/device/E2E layer unavailable just because discovery returns nothing. First run the normal remediation path: launch the emulator/browser/server/container, wait for readiness, restart the client daemon if appropriate, run the repo setup script, and re-run discovery. Only after that fails may the layer be reported unavailable, with command evidence, residual risk, and next unblock action.
92
+ - If a browser/device/E2E or host-smoke layer is classified as blocking, unavailable means the delivery is not complete. Use `pre-runtime-test-ready` only when code, lower-layer tests, and build checks are done and a named human/device owner must finish the runtime gate; a handoff-only label, not merge-ready/release-ready. Otherwise use `blocked`. Do not describe such work as done, fixed, merge-ready, or release-ready.
93
93
 
94
94
  ## Entry Decision: TC Source and Scope
95
95
 
@@ -16,7 +16,7 @@ Duplication and dead-code gates run with explicit configuration, not defaults: t
16
16
 
17
17
  ## Frozen Regression Set And Adversarial Passes
18
18
 
19
- Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests.
19
+ Tier the frozen regression set so the gate stays affordable: deterministic frozen cases run in the blocking release gate; cases needing live infra / model calls / real indexes run in the release or pre-ramp gate with an explicit marker, owner, and timeout (do not stuff flaky live cases into the fast gate); human-review-only cases are release evidence, not mislabeled automated tests. The tiering criterion is cost and side effects, not speed: a case that consumes paid resources (model/API spend, sandbox creation, render jobs) or mutates external state never belongs in the default always-on lane even when it happens to be fast — the default lane is reserved for cases that are free and side-effect-free to run on every change.
20
20
 
21
21
  The proactive complement — the adversarial pass over code already considered "done": run it as an active defect-discovery step, deliberately hunting coverage blind spots — error-mapping boundaries, concurrency-protection bypass, double-release/double-close paths — instead of waiting for review or production to surface them. Each confirmed gap lands as a failing test first. Candidate blind-spot classes: the risk-matrix failure classes in `scenario-testing.md`, plus the dependency fault-injection and concurrency/cache cases in `integration-contract-testing.md`.
22
22
 
@@ -24,6 +24,10 @@ The proactive complement — the adversarial pass over code already considered "
24
24
 
25
25
  Use `test-data-and-determinism.md` as the canonical source for fixture shape, anonymization, data builders, golden-file normalization, and deterministic clocks/randomness/ordering.
26
26
 
27
+ - `infra-error` verdict semantics (the entrypoint's status family). One discriminating predicate decides the verdict — **fault origin**, not symptom: a fault in the **evidence infrastructure** (collector, fixture cache/manifest, state-preparation or controlled-fault harness — anything outside the system under test) = `infra-error`, which is never a pass, never a business fail, and never a silent skip; a fault **in the system under test** (crash, malformed product response, product-path timeout) or an assertion that evaluates to false on collected evidence = business `fail`; an **external prerequisite missing before any attempt** = `blocked`. Exactly one verdict per case; ambiguous origin is resolved by investigation, never by defaulting to whichever verdict looks better; page text, a success toast, or another weaker surface must not substitute for the missing evidence. A required case standing at `infra-error` keeps the aggregate claim incomplete — it counts against merge/release readiness exactly like `blocked`, and a report that excludes `infra-error` cases to present a clean total is a false-green report. Before coding a case, each acceptance criterion names its collector and assertion.
28
+
29
+ External-asset fixtures (media files, documents, large binaries fetched from an external system) form a supply chain that gets pinned end to end: test execution reads only a local read-only cache — never downloads from the external system at run time; a committed manifest pins each asset's identity/hash and CI verifies the manifest plus every blob before the suite runs; cache/manifest verification is part of each affected case's attempted preparation, so a missing or changed cached asset maps to `infra-error` for exactly the cases that need it (a manifest failure aborting before any case attempt marks those cases `infra-error` too, not `blocked` — the infrastructure was configured and failed), never a skip and never a fallback download; seeding/refreshing the cache is a separate offline step on a trusted host, not part of the test run. This composes the network-isolation default and manifest regenerate-and-diff rules in `test-data-and-determinism.md` with the missing-dependency-is-failure rule below into one chain.
30
+
27
31
  ### Fault-Injection Layers For External-Provider Recovery Paths
28
32
 
29
33
  Recovery behavior against an external provider (a model API, payment/storage backend, streaming dependency) needs its fault permutations proven below the live layer. Layer the fixtures; prove each fault class at the most protocol-real layer that can still script it deterministically:
@@ -23,7 +23,7 @@ Before adding a new E2E test, create or update the scenario matrix in `scenario-
23
23
  - Save authenticated state only when the test is not about login.
24
24
  - Capture console errors, failed network requests, screenshots/traces/video when useful.
25
25
  - Use network interception only to control nondeterminism or assert payloads; do not mock away the contract that the E2E test is meant to prove.
26
- - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing. Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing** (`@playwright/experimental-ct-{react,vue,svelte}`) is still labeled experimental — use Vitest browser mode for component-level tests until Playwright component-test API stabilizes. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
26
+ - **Playwright 1.5x baseline (Microsoft, ongoing 2024-2026)** is the current default-recommendation browser-E2E framework for new web testing (community-adoption evidence, reproducible: `curl -s https://api.npmjs.org/downloads/point/2026-07-31:2026-08-29/<pkg>` for `playwright` / `@playwright/test` / `cypress` returned 339.9M / 216.2M / 30.3M (fixed range, re-verified 2026-08-30), ≈11×; State of JS 2024 testing section, `2024.stateofjs.com/en-US/libraries/testing/`, shows Playwright leading E2E usage/retention — re-run the query before citing as current). Key features to use deliberately, per `playwright.dev` docs: (a) **Trace Viewer** is the load-bearing debugging surface — every CI failure should produce a `trace.zip` artifact. Set `trace: 'retain-on-failure'` (not `'on-first-retry'`) when CI runs with `retries: 0` for deterministic gating — `'on-first-retry'` produces no trace unless the test actually retries, leaving teams with zero diagnostics on first-failure-then-fix-the-flake debugging cycles. Reviewers open the artifact locally or via `trace.playwright.dev` to step through actions, screenshots, network, and console without re-running the test. (b) **Soft assertions** via `const softExpect = expect.configure({ soft: true })` let one test report multiple failures rather than stopping at the first — useful for state-snapshot assertions where the team wants the whole-page diff in one run, NOT a substitute for the "one test, one behavior" discipline. (c) `toMatchAriaSnapshot()` (Playwright 1.49+) is the structured accessibility-tree assertion — call it on a locator (`expect(page.locator('main')).toMatchAriaSnapshot(...)` or `await page.locator(...).ariaSnapshot()`), preferred over DOM-string snapshots for resilience to non-semantic markup changes. (d) **Component testing**: Playwright's current component-testing guide replaces the former `@playwright/experimental-ct-{react,vue}` packages (`playwright.dev/docs/test-components`) — that replacement statement is all the source establishes; draw no stability or package-layout inference from it, and follow the pinned Playwright version's own installation instructions before adding or removing any component-testing package. Vitest browser mode remains a valid component-level alternative when the portfolio already standardizes on Vitest. (e) **Projects** in `playwright.config.ts` define run matrices (browser × device emulation × baseURL × config variant) and produce one merged HTML report. Note: Projects ≠ sharding — sharding is a separate `--shard=k/n` mechanism that splits a single project's tests across multiple workers/machines; teams often combine both (projects for the matrix, sharding for parallelism per project). Pin worker count and shard count for CI determinism, do not let auto-detect choose. Routing the per-stack Playwright config implementation goes to `web-react-dev/references/web-quality-release.md`; this skill owns the test-layer policy.
27
27
 
28
28
  ## Runtime QA Sweep
29
29
 
@@ -58,7 +58,7 @@ When the deliverable ships as a built or installed artifact — a package `bin`,
58
58
  - Do not reproduce every unit branch through E2E.
59
59
  - Keep E2E flows few, stable, and tied to user/business risk.
60
60
  - Prefer one happy path plus high-risk negative paths over many shallow click-throughs.
61
- - A click-through without assertions is not E2E evidence.
61
+ - A click-through without assertions is not E2E evidence — and weak proxy signals are not business assertions: page loaded, URL changed, non-empty body text, a generic button/canvas/heading visible, or a success toast alone do not prove the business outcome. Anchor the pass condition on objective effects — API response fields, persisted records, balance/count deltas, generated artifact URLs (sufficient alone only when URL issuance is the claimed contract; a generation-success case dereferences the URL and validates artifact status/metadata/content, since a request can issue a valid URL and fail before storing the artifact), or a stable user-visible terminal state (alone only when the visible terminal presentation IS the claimed contract — a rendered "completed" can outrun persistence/billing/artifact creation, so business-outcome cases pair it with the durable effect) — and match assertion strength to what the test title and scenario row claim; a shallow signal is acceptable only when the case explicitly tests just that shallow signal.
62
62
  - If a scenario can be proven with a stable API/contract/integration test and only needs one browser smoke for confidence, do not duplicate all permutations in the browser.
63
63
  - External-provider fault/recovery permutations (disconnects, malformed streams, rate limits) belong at the protocol-real fault-server and recorded-replay layers (`ci-fixtures-and-flake-control.md`, Fault-Injection Layers); the live credentialed e2e keeps one wiring sanity path, not the fault matrix.
64
64
 
@@ -113,6 +113,16 @@ the implementation.
113
113
  - For cross-RPC typed error envelopes, test a roundtrip: server raises a typed error, the wire-format payload is captured, the client reconstructs a typed error of the same class with the same code/message. Include the unknown-shape path: a wire payload that does not match the canonical envelope returns a transport/unknown error without silent loss of the original cause.
114
114
  - For Code-range allocation, test that a service trying to register a code outside its allocated range fails at build/test time, not at runtime.
115
115
 
116
+ ### Cross-Repo Field Change — End-to-End Checklist
117
+
118
+ Adding, renaming, or retyping a field that crosses a repo/service boundary is one end-to-end contract change, not N independent edits. Before calling it covered, walk all five steps (each is a distinct failure site with its own evidence):
119
+
120
+ 1. **Producer fallback / bridge** — for an additive field the producer emits a safe default/absent form for consumers that have not upgraded; for a rename/retype a default is NOT enough — keep a compatibility bridge (dual-write old+new representation, or a versioned mapping) until every active consumer is confirmed reading the new form, then remove it in the cleanup stage. Asserted, not assumed.
121
+ 2. **Every transport mapper preserves the field explicitly** — do not assume an object spread/copy crosses a mapper or DTO boundary; each mapper in the chain gets an assertion that the field survives it.
122
+ 3. **Consumer coverage spans all active consumer variants** — enumerate them from the delivery record's consumer inventory (`../../product-ui-ux-design/references/delivery-contract.md` consumer_inventory for UI variants); testing one variant of a multi-variant consumer is the classic escape.
123
+ 4. **One real inbound frame through the mapper, plus one unchanged generic path as control** — the real-frame test proves the new field flows; the untouched-path test proves the change did not perturb everything else (the control catches over-broad mapping edits).
124
+ 5. **Paired changes are cross-linked and land in a compatibility-safe order** — not by mutually blocking merges (that deadlocks): consumer tolerance for the field's absence/new form lands first, then the producer emission (its fallback from step 1 keeps not-yet-upgraded consumers safe), then cleanup removes the fallback once all consumers are confirmed upgraded. Each stage gates on the previous stage's **deployed** compatibility state, and the MRs cross-reference per `../../product-rd-workflow/references/cross-repo-coordination.md` for visibility.
125
+
116
126
  ## Platform Contract / Protobuf / RPC Test Obligations
117
127
 
118
128
  Platform-service-connectivity owns policy and proof mechanics for protobuf-backed HTTP, response envelopes, RPC/base fields, and boundary exposure. `testing-strategy` owns assertion coverage, verdict shape, and CI placement.
@@ -78,7 +78,7 @@ def test_export_request_returns_signed_url_when_user_has_quota():
78
78
 
79
79
  ## 3. Test smells(测试异味)
80
80
 
81
- **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书列约 18 项分 code/behavior/project 三类;下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
81
+ **定义**(精选自 Meszaros 2007 及其衍生分类):测试代码的反模式。Meszaros 原书顶层列 15 项 smell,分 code/behavior/project 三类(5/6/4;xunitpatterns.com "All Test Smells" 目录,另有类下变体/别名未计入);下面是日常 review 最常碰到的子集(部分名字 / 阈值是团队启发,非原书字面):
82
82
 
83
83
  | 异味 | 含义 | 后果 |
84
84
  |---|---|---|
@@ -229,7 +229,7 @@ internal_helper_mock.parse.assert_called_once() # 重构改 parse 就挂
229
229
  - **MC/DC**(Modified Condition/Decision Coverage):每个 boolean 子条件独立影响过决策。**DO-178C 航空 / 医疗 / 汽车 functional safety 才用**;普通业务代码无监管要求时不必上。
230
230
  - **Mutation coverage**(见 source-to-case-workflows §C.1):才是真"测得好"的 proxy
231
231
 
232
- **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。
232
+ **用**:CI 设 floor(如 line 60% / branch 50%)防覆盖崩塌;critical-path 模块定专项目标(line 90%)。数值出处:60%/90% 对齐 Google Testing Blog "Code Coverage Best Practices"(2020)的 60% acceptable / 75% commendable / 90% exemplary 分档——注意该文同时反对自上而下的强制统一阈值,floor 应按仓现状起步再棘轮;branch 50% 为团队启发值,无外部权威出处。
233
233
 
234
234
  **不用**:
235
235
  - **不用 100% 作 KPI** — 强行凑 100% 会产生 lazy assertions(`assert result is not None` 这种)
@@ -64,7 +64,7 @@ Classify commands before running them:
64
64
  - E2E/release smoke: browser/API/device real flows in an isolated environment.
65
65
  - Long gate: compatibility matrix, migration dry-run, load/replay, visual regression, or full-suite release checks.
66
66
 
67
- If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it.
67
+ If the repo separates markers such as `unit`, `integration`, `contract`, `slow`, or `e2e`, preserve that split. Do not move expensive tests into the default PR path unless the local CI contract already expects it. Expensive includes billable: tests that consume metered external resources — paid AI/model inference, per-call third-party APIs, real payment/checkout flows, cloud sandboxes or device farms — get their own explicit marker/lane, stay out of default PR and scheduled-frequent lanes by default, and run only with a named budget owner and an authorized test account (credential/provisioning discipline per `ci-fixtures-and-flake-control.md`). Payment/checkout tests default to the provider's sandbox/test mode — a budget owner bounds spend but does not make live payment mutation safe; an unavoidable live-money canary needs its own explicit approval, a hard spending cap, and refund/cleanup handling. A case that silently creates paid resources under a generic `e2e` tag is a lane-classification finding.
68
68
 
69
69
  When existing commands use richer labels such as `contract_fake`, `contract_mysql`, `api`, `e2e_smoke`, `failure_mode`, `drill`, `shadow`, `smoke`, or `replay`, preserve the local meaning instead of flattening everything into unit/integration/E2E. Map them to the scenario matrix and CI gate they actually serve.
70
70
 
@@ -99,7 +99,7 @@ For a 域卡/执行卡 (a card that sets WHAT a domain must achieve + who owns i
99
99
  - Required flow: STOP line edits → confirm the corrected core with the user (one short question, don't guess again — this failure class recurs precisely from re-guessing) → re-derive 负责人/红线/里程碑/验收/依赖兜底 from the corrected core as a **draft for review** (not a blind blast-write) → publish on approval.
100
100
  - Repeat signal: repeated user "这是什么/什么玩意儿" on the same card = the premise is wrong; escalate to re-derive, do not keep tightening.
101
101
 
102
- ## 句子层(吸收 Strunk《风格的要素》与 Google Technical Writing 课程,仅取适合中文交付文档的;英文语法/标点规则不适用,已剔除)
102
+ ## 句子层(吸收 Strunk 与 Google Tech Writing,仅取中文交付文档适用项)
103
103
 
104
104
  > 英文文档:用完整 Strunk 规则(含被本节剔除的语法/标点条),本节只是中文交付子集。
105
105
 
@@ -140,6 +140,7 @@ Never destroy collaborative comments. Before editing a collaborative doc, fetch
140
140
 
141
141
  ## WORKFLOW
142
142
 
143
+ 0. 读者批注:判根因类、全文修同类(`references/annotation-driven-revision.md`)。
143
144
  1. Extract the decided-points checklist from the current text.
144
145
  2. Apply the DELETE list; keep everything in KEEP.
145
146
  3. Rewrite to FORM; confirm every decided point still present.
@@ -0,0 +1,9 @@
1
+ # 批注驱动修订
2
+
3
+ 读者批注/评审意见不是孤立改句请求,而是**阅读断裂的证据**。处理协议:
4
+
5
+ - 逐条判根因类:背景缺失 / 概念未定义 / 逻辑跳跃 / 措辞 / **事实・引用错误**(日期、数字、出处错——此类不走措辞同类扫,改走源核验:对一手源改正并按同源扫其余引用处)。
6
+ - 按根因类**必须全文扫同类位置一起修,不得只改被标记的那一句**——点修复会把同类断裂留给下一位读者复发。**扫描全文、编辑限权**:当授权只覆盖某条批注/某节时,全文扫描产出同类候选清单,但自动编辑只落在授权范围内;范围外的同类位置先报告、经批准再修(不得以「修同类」为名越权改动已定内容)。
7
+ - 改完以首次读者身份通读被改段落(standalone-paste 逐行读,同 closeout 判法)。
8
+ - 批注本体的保全走 SKILL.md 的 COMMENT-SAFE 硬规则(先取真实评论数、定向编辑、改后复核锚点/条数)。
9
+ - 出处:读者差集原则(好文档=读者需要的知识−已有的知识,Google Technical Writing audience 章)——批注正是「差集没算对」的实测信号。
@@ -16,8 +16,14 @@
16
16
 
17
17
  **`[禁]` 档(已核,勿再立)**:加粗/高亮密度;每 N 字一图的图表密度(唯一数字是学术期刊的
18
18
  **印刷页数配额**,成因是版面成本不是可读性);"一行不超过 40 汉字"引 WCAG(该条的 CJK 40 是从
19
- 拉丁文 80 折半推导,且原文要求是「提供**机制**让用户改」不是「正文必须排这么宽」);"句子不超过
20
- 25 词";"留白提升理解约 20%"(错误引用链)。
19
+ 拉丁文 80 折半推导,且原文要求是「提供**机制**让用户改」不是「正文必须排这么宽」;作为**社区规范**
20
+ 另有可追溯来源——阮一峰《中文技术文档的写作规范》"多于40个字的句子不能接受"——个人规范可引用,
21
+ 不可当标准/厂商级权威);"留白提升理解约 20%"(错误引用链;2026-08 复核仍无任何标准/厂商/学术
22
+ 一手源给出该比例)。
23
+
24
+ **降档更正(2026-08 复核)**:"句子不超过 25 词"从 `[禁]` 移出——GOV.UK 官方写作指引有一手源
25
+ ("Try to split up sentences that are over 25 words long" + GDS 博客专文),属**单一机构 house
26
+ style**:可引用(点名 GOV.UK),不可当行业标准立硬规则,归 `[工]` 档强度。
21
27
 
22
28
  **同体裁实测分布 ≠ 规范阈值**:可以说"本稿在同体裁公开样本分布的哪个位置",不能由此推出"写得好"。
23
29
  分布定位的合法输出是描述不是裁决——这条与 `tighten-doc` SKILL.md「外部基线」条同源,按那条执行。
@@ -50,6 +50,7 @@ When checking a React project against team standards, split findings into determ
50
50
  - Locate the owning route/page, component tree, state owner, API client, data-fetching layer, styling system, tests, and build scripts before editing.
51
51
  - Identify whether state belongs in URL/query params, cache/server state, form state, local component state, browser storage, or global app state.
52
52
  - Read repo wrappers first: package manager, dev/build scripts, lint/typecheck/test runners, browser/E2E tools, environment variables, and generated clients.
53
+ - Before generating component-library code: the **workspace lockfile resolution is the version authority** (`npm ls <pkg>` / `pnpm why` / yarn equivalent — a library config file or global CLI can resolve a different release than the workspace); take configuration ground truth (framework, aliases, installed components) from the library's own introspection surface (config file such as `components.json`, official info CLI/MCP, or the installed package's exports/types); write APIs against the resolved version, never from memory of "current" APIs (prop names and defaults shift across majors). After editing, close with the library's own linter/codemod check on the changed files when one exists (deprecated-usage and a11y rules the generic lint config does not know); for a library major-version migration, follow the official migration checklist + changelog for the exact from→to pair, apply, then re-run the library lint to prove no deprecated usage remains.
53
54
  - If a design exists, map visible states and interactions to component ownership before implementing.
54
55
  - For every visible UI change, load `../product-ui-ux-design/references/delivery-contract.md` and consume either its full Design brief + Phase 0 or its valid low-risk copy-only record + lightweight Phase 0 before coding. The lightweight path checks semantics, accessible name, localization, rendered extent, and target render without inventing unrelated matrices; risk-bearing copy uses the full path. For full slices, map structure, state/adaptation matrices, behavior and criteria to React ownership; record route/server, component/state owners, viewports/themes/input modes, and preserved behavior. When React is embedded in a native WebView, mini-program `web-view`, or Electron shell, this skill owns the content-layer member; the native/mini/desktop host owner must add its separate entry, binding and runtime record, even when host code is unchanged.
55
56
  - Before the first implementation edit, add the canonical `client_entry` defined there: local rule identifier or short quote and implementation decision, target surface/runtime, planned run/capture command, and behavior that must remain unchanged.
@@ -50,6 +50,9 @@ Three reuse patterns recur in React libraries; pick by what the consumer needs t
50
50
  - **Custom hook** — when reuse is logic only, no rendering shape required. Default choice for state machines, subscriptions, side-effect orchestration. See discipline above.
51
51
  - **Headless / unstyled primitives** (Radix UI, Headless UI, Ariakit, downshift, react-aria) — when reuse is a11y + behavior (focus trap, keyboard navigation, ARIA contract) but every consumer needs different visual treatment. Default for design-system primitives: the primitive owns keyboard / focus / ARIA / portal / dismiss, the consumer owns styling. **Do not invent a custom focus-trap or ARIA implementation when a maintained headless primitive exists** — the bug surface is well-known and the maintained library has fixed bugs you have not heard of yet.
52
52
  - **Compound components** (e.g., `<Select><Select.Trigger/><Select.Content/><Select.Item/></Select>`) — when consumers need to arrange children but share an implicit parent context. Cleaner than render props for this case. Default for `Tabs` / `Accordion` / `Select` / `Menu` / `Tooltip` shells.
53
+ - **Type the shared context as a grouped contract — `{ state, actions, meta }`** (data / state-changing functions / refs & config; a component with no refs or config declares `meta` as an optional key — `meta?:` — and consumers handle its absence) rather than a flat grab-bag. Subcomponents consume the contract, never a specific state hook, so any provider implementing it can inject the state — local `useState` for an ephemeral form, a global or server-synced store for a live surface — and the same composed UI works under either provider. One packed context value means a consumer that only calls `actions` still re-renders whenever `state` changes; when that measurably matters, split state and actions into two contexts (the React-docs reducer + context split) instead of one value. The split pays off only when the provided actions value is itself referentially stable — pass `dispatch` directly or memoize the actions object with its complete dependency list, never dropping a changing dependency (an `onSave` prop, a scoped API client) to force identity — that ships stale calls; an inline `{ dispatch }` wrapper takes a new identity on every provider render, and the optimization holds only while the declared dependencies are actually stable; actions must not close over current state (a reducer's `dispatch` qualifies; state-closing memoized callbacks do not), and reducers/updaters themselves stay pure — a side-effecting command that needs the current state snapshot, like a submit posting the form, accepts it as an argument from a state-reading consumer; when actions inherently depend on state, keep the single combined context.
54
+ - **The provider boundary, not the visual shell, decides who can share state.** When consumers outside the visual shell need the state, lift it into a dedicated provider component — the lowest-common-owner rule above still decides how high, and lift only the shareable model `state`/`actions`: each compound root keeps focus refs, generated IDs, and item registration instance-local — open/highlight state stays local by default too, lifting only when outside consumers genuinely coordinate it and then scoped by an explicit compound-instance identity — or two shells under one provider cross-wire focus and ARIA targets; a preview panel or submit button rendered outside the visual shell but inside the provider then reads `state` and calls `actions.submit`. Two smells that say state should have been lifted into a provider: syncing a child's state upward with an on-change `useEffect` callback, and imperatively reading child-owned React state through a ref because another consumer must coordinate with it — uncontrolled DOM values read at submit (`FormData`, file inputs) and imperative third-party-widget integrations are not this smell and stay valid per the uncontrolled-forms rule above.
55
+ - **Mode booleans multiplying on one component are a composition finding.** When a new component API — or an existing API already undergoing an intentional redesign — accumulates mode props (`isEditing` / `isCompact` / `isInline`-style) whose combinations multiply and some are impossible, replace the modes with explicit variant components — each variant composes the shared compound parts it needs and composes under the appropriate shared provider, or declares its own when the variant is the lowest common state owner. The `never`-union prop typing above makes illegal combinations uncallable at the type level; explicit variants remove the combinations altogether. Prefer variants when the modes are stable, meaningfully different compositions or ownership boundaries and conditional render branches, not just prop types, fork on the booleans; modes sharing one structural contract stay one component with a discriminated-union variant prop and exhaustive branching. Do not refactor a stable existing API solely for pattern conformance.
53
56
 
54
57
  **Legacy patterns to recognize, not to reach for**:
55
58
  - **Higher-Order Components (HOCs)** — `withAuth(Component)` / `connect(mapStateToProps)(Component)`. For the same concern in NEW code, prefer a custom hook (`useAuth()` / `useSelector()`). Existing HOC APIs are acceptable when a library / framework contract requires them (React-Redux `connect` is still a valid public API) — do not refactor working HOC integrations on cosmetic grounds. Do not introduce a new HOC when a hook does the same job.