@ccoalm/ccl-skills 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +5 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +5 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +1 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +12 -12
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +2 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +2 -2
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +1 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +12 -15
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attention-budget-ratchet.md +37 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +13 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +37 -30
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +24 -3
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +5 -5
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +1 -1
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +81 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +12 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +1 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +30 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-contract-anchors.sh +126 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +197 -1
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/contract-anchors.tsv +15 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +210 -36
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +3 -3
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/gate_receipt.py +576 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_antipattern_grep_panel.sh +80 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +99 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +25 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +251 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_contract_anchors.sh +196 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +222 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +16 -10
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity.sh +178 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_frozen_case_sanctity_selfproof.sh +108 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_gate_receipt.sh +431 -0
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_pinned_phrase_mutation_walk.sh +151 -0
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +86 -5
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +27 -21
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +25 -15
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +9 -9
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
  77. package/dist/assets/release.json +127 -67
  78. package/package.json +1 -1
@@ -75,3 +75,8 @@ When the mobile app must talk to internal-CA / enterprise-CA / self-signed hosts
75
75
  - **Exclude at the dependency/target level, per real-user variant — not just a runtime flag, and not just the variant named `release`.** An `if (isDebug)` check still compiles the code — and any bound port/server — into the binary; even `kDebugMode` / `__DEV__` only strip the Dart/JS call path while a native plugin/module/platform-channel stays linked. The primary guarantee is that the debug dependency/target is absent from the release classpath / target membership / build flavor of *every variant that reaches real users*, verified per variant: iOS `#if DEBUG` plus Debug-only target membership — CocoaPods can scope a pod to Debug configurations (`:configurations => ['Debug']`), but SwiftPM has no build-configuration-scoped *package dependency*, so a debug package must be confined to a debug-only target/scheme and left unlinked by every real-user target rather than gated by a `Package.swift` condition; also confirm no real-user scheme (TestFlight / enterprise / staging) inherits `DEBUG` in `SWIFT_ACTIVE_COMPILATION_CONDITIONS` / `GCC_PREPROCESSOR_DEFINITIONS`, or `#if DEBUG` code compiles straight into it; Android a `debugImplementation` / debug source-set dependency, NOT `implementation` plus a runtime check (a manifest `ContentProvider`/initializer or reflection keep-rule runs it anyway) — and note a `debug` / `qaDebug` buildType shipped to dogfood/internal testing IS a real-user build that includes those deps, so real-user channels must use release-class variants with debug deps excluded; Flutter/RN remove the package from the release flavor at the platform level (a native plugin autolinks/registers into the binary via podspec/Gradle/manifest unless excluded from the release target, even when the Dart/JS guard is stripped).
76
76
  - **Verify on the exact final signed/packaged upload artifact per channel and variant, and treat a symbol/string grep as a backstop, not proof.** A single `nm -j` / `strings` / dexdump / grep is defeated by R8 / Hermes / LTO renaming (`DevInspector` → `a.b.C`), dynamic loading (`dlopen`, Android dynamic feature, RN/Flutter OTA bundle, remote WebView JS), and non-code surfaces (debug assets, URL schemes, manifest providers/services, plist keys, debug entitlements). *Discover* the inventory from the dependency graph, manifest/plist, entitlements, and route/deep-link/bridge registries — not a remembered list, since an unlisted transitive SDK initializer or debug bridge handler is exactly the gap — then for each assert absence across binary, dex/native libs, assets/resources, manifest/plist, entitlements (`get-task-allow` is debug-only; Local Network / ATS / cleartext exceptions are security-sensitive plist/capability config that need production justification, not necessarily debug), and the OTA/dynamic-load channels reachable by that release (a passing artifact digest does NOT bind a debug bundle pushed later via CodePush/Expo/remote WebView JS, so real-user OTA/remote channels must not serve debug payloads). Run it on the post-signing artifact you actually upload (record its digest), across every shipped platform / scheme / flavor / extension, never on an intermediate `assembleRelease`. (`__DEV__` / `kDebugMode` elimination is also minifier-dependent and has silently regressed historically, so the source guard alone is never proof.)
77
77
  - **Make it a release-blocking evidence row, not MR prose** (otherwise the gate is decorative — a team ships from local Xcode/Fastlane, self-marks "checked", or never adds the CI job). The row is machine-generated, tied to the uploaded artifact digest + variant, and verified by someone other than the author; it records the digest, distribution channel, variant matrix, and check-command output, and the release/upload lane fails closed on a missing row. **Risk acceptance cannot waive structural exclusion** of debug-only instrumentation in any real-user / real-auth / real-PII build: an exception may only reclassify the channel as non-real-user / no-real-data, or apply to a production-intended diagnostic that ships its own auth + redaction controls (at which point it is no longer debug-only instrumentation) — otherwise the release stays blocked, an unbounded "accepted" is not a valid disposition. Guidance cannot force a given repo's pipeline to fail closed — that mechanical enforcement is the release-pipeline owner's job (route to `platform-release-engineering`); absent the row the status is not release-ready, the same way missing device smoke leaves a slice `pre-runtime-test-ready`. A manual removal/teardown flow is a convenience path for a one-off exit, never the mechanism that keeps debug code out of real-user builds.
78
+
79
+ ## Component-Library Version Discipline
80
+
81
+ - Write third-party UI/component-library APIs against the exact **resolved installed version** — the package-manager lockfile (pubspec.lock / Podfile.lock / package-lock.json / yarn.lock / pnpm-lock.yaml / Gradle dependency-locking or `dependencies` report) or the package manager's installed-package query — never a manifest's semver range (package.json / a catalog alias with dynamic constraints declares a range, not what is installed) and never memory of "current" APIs — prop names and defaults shift across majors.
82
+ - When the library ships its own lint/codemod or migration checklist, close changes and major-version migrations with it on the changed files (deprecated-usage and a11y rules a generic lint config does not know).
@@ -124,3 +124,9 @@ Review the changes on this branch against the base branch. Focus specifically on
124
124
  ```
125
125
 
126
126
  For product, architecture, economics, compliance, IA, or risk decisions, ask Claude to challenge the decision and return concrete blockers, missing gates, or required follow-ups. Keep the scope narrow enough that Claude can finish. When the decision depends on trusted cross-repository evidence already collected by Codex or codegraph and the source repositories cannot be safely exposed through repository consult, use prompt-only consult with `--allow-prompt-only-advisory` instead of broadening Claude's filesystem scope; paste the relevant excerpts with provenance and state that Claude must answer only from that evidence. Prompt-only consult cannot be the deciding evidence for landing, risk acceptance, or gate closure; pair it with repository verification or explicit human risk-owner acceptance. If a repository can be safely scoped with `--cwd`, prefer repository consult for repository-answerable questions.
127
+
128
+ ## Reviewed-Identity Honesty And Finding Relay Integrity
129
+
130
+ - The reviewed-identity record is a freshness guard, not cryptographic proof: it detects staleness but does not authenticate who reviewed or prevent forgery, so it never substitutes for the gate's attribution and output-validity checks.
131
+ - Review depth for a changed candidate is a deterministic function of the delta, never a manual downgrade: an unchanged candidate may reuse its recorded pass only while base, review profile, and lens set are also unchanged — a base advance or profile/lens change voids reuse exactly like an edit; any changed candidate takes a fresh full run — "the change is small, the old pass still counts" is not an agent's call, and any future bounded incremental mode must use mechanical criteria (delta size, files touched, high-risk surface entered) recorded with the run.
132
+ - Reviewer findings pass a false-positive check before being relayed as blocking: verify each claimed defect against the cited code (open the lines; confirm a "missing" guard is absent on the whole path) — the check validates existence AND blocking severity: a confirmed-but-advisory or overstated-severity issue is relayed at its true severity, not as blocking; demote only what is established false-positive-with-reason; an unverifiable load-bearing finding stays blocking until verified or explicitly accepted by the risk owner (unverifiable is a status, never a demotion). Keep every raw finding with its disposition (confirmed / false-positive-with-reason / unverifiable) — the check filters the relay, never the record, so the filtering itself stays auditable.
@@ -314,6 +314,11 @@ self-review plus explicit task reframing; it does not silently create a new
314
314
  Agent budget. An untracked challenge is one-off advisory evidence; it cannot
315
315
  enter a later Agent round or satisfy the local completion checkpoint.
316
316
 
317
+ Two consequences follow from those stable bindings and must be planned for before round 1:
318
+
319
+ - The selected-owner digest hashes each selected owner package's current working tree, and owners derive from the candidate's own paths — so a candidate edit inside any selected owner package invalidates every prior receipt and the next tracked round fails `review_chain_invalid`. For a self-hosted candidate (a skill-repo diff editing the package that owns it) that is nearly every applied fix — one confined to files outside every selected owner drifts only the candidate hash and may continue in-chain: "do not reset Agent authority" promises no continuation, and the in-chain tolerance for older candidate hashes is reachable only while the fix stays outside its selected owners.
320
+ - A chain restarted after such a break re-enters the same cumulative Agent budget and must never be counted as fresh authority; the bounded restart recipe for the extraction lane (batched dispositions, cross-chain round accounting, full-context first packet, terminal disposition at the cap) is owned by the extraction workflow's dual-track gate reference.
321
+
317
322
  The controller is stateless and prevents accidental/cooperative resets only. A
318
323
  trusted host or platform must retain the ledger when hostile local callers are in
319
324
  scope; a repository-local counter cannot authenticate human authority.
@@ -65,6 +65,7 @@ Use this skill for the full defect discipline: diagnose the immediate failure, f
65
65
 
66
66
  5. Verify cause.
67
67
  - Prove the cause with evidence.
68
+ - Report query/lookup evidence by cardinality: a data query, log search, or identity resolution that returns 0, 1, or N matches reports each of those outcomes distinctly — never silently take the first row of N, and never treat 0 rows as "no evidence collected" (an empty result over a named scope IS evidence: record which scopes matched and which were empty).
68
69
  - When the cause is environment/toolchain state, prove it from the tool that owns that state, not only from the high-level wrapper. A wrapper failure is a symptom until the underlying compiler, generator, runtime registry, dependency resolver, or platform destination evidence explains it.
69
70
  - If disproven, return to hypotheses instead of guessing.
70
71
  - Separate symptom, immediate cause, contributing factors, and prevention.
@@ -7,7 +7,9 @@ description: 风险定级 / 要不要灰度 / 需要哪些 gate / 双人 review
7
7
 
8
8
  Use this lightweight router before or during product delivery when the required rigor is unclear. It classifies risk and names the gates that must run; it does not replace `product-rd-workflow`, `testing-strategy`, `product-ui-ux-design`, stack-specific development skills, review skills, or release workflows.
9
9
 
10
- **Objective high-risk change-shapes force classification even when the change feels trivial — a low-risk intuition is not a substitute for running the tags.** The "when rigor is unclear" entry is not a licence to skip the router because you judged a change low-risk: the tags below classify well *once you are here*; the failure mode is never arriving because you felt sure. So regardless of that intuition, if the change deletes / backfills / migrates data or is otherwise irreversible, touches authentication / authorization / roles / tenant isolation, handles secrets / credentials / tokens, moves money or billing / quota, alters an external-facing or service-client API contract, changes a trust boundary or an untrusted-input sink, drives a production rollout / flag, or changes a shared deterministic gate / validator / CI harness / merge-readiness or completion-status rule (which can look docs/scripts-only yet decide future merge and completion semantics), run the classification FIRST and only then conclude low-risk if the relevant tags (`write-finality`, `data-migration`, `permission-access`, `security-review`, `api-contract`, `money-quota`, `release-ops`, `shared-gate`, `ai-action`) actually clear. Self-de-escalating one of those change-shapes to "not needed" without running its tag is the miss this prevents. (The always-on routing layer and `product-rd-workflow` owner-dispatch are the backstops for the case where the router is never invoked at all — this clause governs the case where you are deciding whether it applies.)
10
+ **Objective high-risk change-shapes force classification even when the change feels trivial — a low-risk intuition is not a substitute for running the tags.** The "when rigor is unclear" entry is not a licence to skip the router because you judged a change low-risk: the tags below classify well *once you are here*; the failure mode is never arriving because you felt sure. So regardless of that intuition, if the change deletes / backfills / migrates data or is otherwise irreversible, touches authentication / authorization / roles / tenant isolation, handles secrets / credentials / tokens, moves money or billing / quota, alters an external-facing or service-client API contract, changes a trust boundary or an untrusted-input sink, drives a production rollout / flag, or changes a shared deterministic gate / validator / CI harness / merge-readiness or completion-status rule (which can look docs/scripts-only yet decide future merge and completion semantics), run the classification FIRST and only then conclude low-risk if the relevant tags (`write-finality`, `data-migration`, `permission-access`, `security-review`, `api-contract`, `money-quota`, `release-ops`, `shared-gate`, `ai-action`) actually clear. Self-de-escalating one of those change-shapes to "not needed" without running its tag is the miss this prevents.
11
+
12
+ - Tone and diff size are not classification inputs: a reporter's urgent tone or executive pressure must never escalate a change's tags, and a small diff must never de-escalate them — tags run on the change's objective shape (production impact, frequency, money/data finality, workaround availability, rollback difficulty, verification sufficiency). (The always-on routing layer and `product-rd-workflow` owner-dispatch are the backstops for the case where the router is never invoked at all — this clause governs the case where you are deciding whether it applies.)
11
13
 
12
14
  ## Output Contract
13
15
 
@@ -89,7 +89,7 @@ Not appropriate for:
89
89
  - Keep transport DTOs, application parameters/results, domain objects, and storage models as separate boundaries when behavior or compatibility is non-trivial.
90
90
  - Patch/update contracts need presence semantics, not only zero values.
91
91
  - For finite values used across contracts, domain logic, persistence, clients, analytics, or events, architecture owns the semantic source of truth. Decide the canonical owner, shared contract/package location, allowed representations per boundary, parser/canonicalization owner, unknown/default behavior, and rollout order. If a shared location does not yet exist, the architecture decision must approve a local fallback and a consolidation task with owner and deadline; unowned `finite-value-debt` markers are architecture findings.
92
- - **RPC framework choice: Kitex remains the default; ConnectRPC is the credible 2025-2026 alternative for specific contexts**. Per `connectrpc.com` docs, Connect ships as a single small Go package (single-digit-thousand LOC), built on `net/http` with handlers implementing `http.Handler` and clients wrapping `http.Client` — works with any third-party router, middleware, or server. Supports three protocols (gRPC, gRPC-Web, Connect's own protocol) over both HTTP/1.1 and HTTP/2; any gRPC client in any language can call a Connect server, and Connect clients can call any gRPC server (validated by Google's interop tests). CNCF-incubated as of 2025; production-adopted at CrowdStrike, PlanetScale, Bluesky, Dropbox per `buf.build/blog`. Architecture choice: **choose Connect** when (a) the service hosts a browser-facing API and gRPC-Web is the main use case (Connect collapses gRPC + gRPC-Web + Connect into one server), (b) the service is small and Kitex's framework surface is over-spec'd, (c) the service must interop bidirectionally with multi-language gRPC clients but the team wants to avoid grpc-go's 130k-LOC dependency footprint. **Stay on Kitex** when (a) the team already operates Kitex across many services and the platform observability/middleware/registry contracts are Kitex-native, (b) Thrift IDL (TTHeader / TTHeader Streaming) is in active use alongside protobuf — Connect is protobuf-only, (c) Kitex-specific features (StreamX, FastCodec, generic call) are load-bearing. The two are NOT mutually exclusive: a Connect-based public-edge gateway can fan out to internal Kitex services, but **wire compatibility alone is not interop** — Connect speaks gRPC on the wire when configured for gRPC, and Kitex servers configured with the gRPC meta handler can accept those calls; however, Kitex services in production typically depend on Kitex-specific framework metadata (lane / tenant / caller-identity ctx values propagated via TTHeader or Kitex middleware), and Connect clients do not emit those by default. The fan-out works ONLY when (a) the target Kitex services are configured for the gRPC meta handler (not TTHeader), (b) the Connect-side Go client explicitly attaches the gRPC metadata (`metadata.MD`) that the Kitex service's middleware reads as caller-identity / lane / tenant — typically through a small adapter layer at the Connect-side that pulls fields from the Connect call context and writes them as gRPC headers, (c) any Kitex-specific TTHeader Streaming / generic call / Thrift-only paths stay routed through a Kitex-native edge instead. Plan the adapter layer as a first-class architecture component, not a one-line bridge; otherwise tenant / lane context drops silently at the protocol boundary and downstream services run without isolation.
92
+ - **RPC framework choice: Kitex remains the default; ConnectRPC is the credible 2025-2026 alternative for specific contexts**. Per `connectrpc.com` docs, Connect ships as a single small Go package (single-digit-thousand LOC), built on `net/http` with handlers implementing `http.Handler` and clients wrapping `http.Client` — works with any third-party router, middleware, or server. Supports three protocols (gRPC, gRPC-Web, Connect's own protocol) over both HTTP/1.1 and HTTP/2; any gRPC client in any language can call a Connect server, and Connect clients can call any gRPC server (the project validates gRPC compatibility with an extended version of Google's own gRPC interoperability test suite). CNCF **sandbox** project (per connectrpc.com, 2026-08 — not incubating/graduated); production-adopted at CrowdStrike, PlanetScale, Bluesky, Dropbox per `buf.build/blog`. Architecture choice: **choose Connect** when (a) the service hosts a browser-facing API and gRPC-Web is the main use case (Connect collapses gRPC + gRPC-Web + Connect into one server), (b) the service is small and Kitex's framework surface is over-spec'd, (c) the service must interop bidirectionally with multi-language gRPC clients but the team wants to avoid grpc-go's 130k-LOC dependency footprint. **Stay on Kitex** when (a) the team already operates Kitex across many services and the platform observability/middleware/registry contracts are Kitex-native, (b) Thrift IDL (TTHeader / TTHeader Streaming) is in active use alongside protobuf — Connect is protobuf-only, (c) Kitex-specific features (StreamX, FastCodec, generic call) are load-bearing. The two are NOT mutually exclusive: a Connect-based public-edge gateway can fan out to internal Kitex services, but **wire compatibility alone is not interop** — Connect speaks gRPC on the wire when configured for gRPC, and Kitex servers configured with the gRPC meta handler can accept those calls; however, Kitex services in production typically depend on Kitex-specific framework metadata (lane / tenant / caller-identity ctx values propagated via TTHeader or Kitex middleware), and Connect clients do not emit those by default. The fan-out works ONLY when (a) the target Kitex services are configured for the gRPC meta handler (not TTHeader), (b) the Connect-side Go client explicitly attaches the gRPC metadata (`metadata.MD`) that the Kitex service's middleware reads as caller-identity / lane / tenant — typically through a small adapter layer at the Connect-side that pulls fields from the Connect call context and writes them as gRPC headers, (c) any Kitex-specific TTHeader Streaming / generic call / Thrift-only paths stay routed through a Kitex-native edge instead. Plan the adapter layer as a first-class architecture component, not a one-line bridge; otherwise tenant / lane context drops silently at the protocol boundary and downstream services run without isolation.
93
93
 
94
94
  ## Data Design
95
95
 
@@ -162,7 +162,7 @@ Tenant commitments around where data lives and who can see it shape the isolatio
162
162
  - **Residency** — "tenant X's data stays in region Y" is a region-per-tenant or region-pinned commitment; the data plane (DB, object storage, backup, analytics) all honor it; the routing layer enforces it.
163
163
  - **Sovereignty** — government / regulated tenants may require a separate stack with no cross-border access; this is a deployment-level isolation, not a runtime knob.
164
164
  - **Encryption** — at-rest encryption per tenant (separate keys per tenant) is a stronger model than a shared key.
165
- - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with conditions; some interpretations of GDPR distinguish anonymization (irreversible) from pseudonymization (key-linkable encrypted data may still be personal data). Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
165
+ - **Crypto-deletion is conditional, not universal** — destroying the per-tenant key is acceptable proof of erasure **only when** the key hierarchy, key backups, envelope keys, restore paths, and the relevant regulator's interpretation all support it. NIST SP 800-88 treats cryptographic erase as a sanitization technique with explicit conditions — do not use CE when data predates encryption enablement or when keys were backed up/escrowed without verified protection (conditions as stated in Rev.1 §2.6; Rev.2, 2025, supersedes Rev.1 and continues the CE-conditions framework — verify the corresponding Rev.2 section when citing it as the authority); and EDPB guidance (e.g. Guidelines 02/2025) holds that encrypted personal data remains personal data "at least until the algorithm is broken", so key destruction is a conditional control, not automatic GDPR erasure. Where any condition fails, crypto-deletion is a *beyond-use / suppression* control (one of the per-store deletion modes above), not proof of erasure; the deletion workflow records the actual mode achieved per tenant per store, and the tenant is told what was achieved. **Required evidence before recording "erasure" via crypto-deletion**, per store and per tenant; each artifact must be authentic to *this* deletion, not theatrical paperwork:
166
166
  - *(a) key hierarchy diagram* — scoped to this store and this tenant's key version, dated within a defined freshness window (e.g., last 30 days), showing every key that wraps or could reconstruct the data.
167
167
  - *(b) key-backup inventory* — for this store, scoped to this tenant's keys, naming every backup location, rotation policy, and the holder of each backup; dated within the freshness window.
168
168
  - *(c) restore-path test result* — run against *this* backup store with *this* tenant's key-version destroyed, confirming the restore fails because the key is gone. A restore test on a different store or a different key-version is not evidence.
@@ -22,6 +22,8 @@ For generic async, consumer, scheduled-job, panic recovery, context timeout, and
22
22
  - Start transitions should move pending work to processing before expensive work.
23
23
  - Failure transitions should capture canonical error code, safe message, retryable flag, retry count, and last trace or log id.
24
24
  - Success transitions should persist result pointer or summary before publishing completion events.
25
+ - All timestamps the transition itself stamps come from a single captured `now` (capturing the clock twice inside one transition produces `finished_at < started_at` records or audit rows disagreeing with the state row under load); domain-provided times — an upstream completion time, an event time — are recorded as received, never re-stamped with the local `now`.
26
+ - Validate external/dependency response structure before casting or mapping it: an assertion/parse step with a typed error path, never a blind cast — a malformed upstream payload must become a failure transition with the canonical error, not a panic or silently-zeroed field.
25
27
 
26
28
  ## Async Processing
27
29
 
@@ -43,6 +43,30 @@ Before rollout, run a bounded capacity check for:
43
43
 
44
44
  Use dry-run or report-only modes for migration/backfill/batch jobs whenever possible.
45
45
 
46
+ ## Provider Evaluation Evidence Ladder
47
+
48
+ When the decision is "adopt / switch to / gray-ramp provider X or model Y for a real workload" (procurement or migration, not routine regression), the two questions are: is it cheaper on the real workload, and is it stable at the required load — and neither is answerable from a model list, a public price page, or one successful request.
49
+
50
+ **Contract before the first paid call.** Write down: exact environment/model IDs/protocol/credential scope; the production request-shape distribution being simulated; the capability-parity checklist; **spend cap, request cap, wall-time cap, automatic stop conditions, and who approved the budget** — enforced by reservation at admission (mechanics in the bullet below), so the invariant is cap ≥ reconciled spend + outstanding reservations at all times; the stop latch halts new admissions, not merely accounting; and the acceptance thresholds plus which decision this run feeds. Thresholds precede data (measurement-design discipline); a run whose stop conditions were invented after the spend is not an evaluation. Before any run that costs money or creates external resources, show the (redacted) config plus the exact commands and estimated cost/duration and get explicit confirmation — never proceed on inferred consent.
51
+
52
+ - Reservation mechanics: admission must be an atomic check-and-reserve against a single reservation ledger — read-headroom-then-reserve is a TOCTOU that lets concurrent workers jointly overshoot the cap — and a reservation releases only on billing reconciliation, never on completion alone (usage/billing signals lag).
53
+
54
+ **Seven evidence layers — a stronger layer may use lower layers as context, never substitute for them:**
55
+ 1. published claim (price page / model card) →
56
+ 2. authenticated control-plane fact (the model is actually on this account/region) →
57
+ 3. single-call data-plane success →
58
+ 4. capability parity on the checklist (structured output, tool use, streaming, context length — against the exact model string; aliases/version suffixes/casing variants are non-equivalent until proven) →
59
+ 5. capacity/stability under sustained load →
60
+ 6. billing reconciliation (billed vs calculated within tolerance, e.g. ≤2%; failure/cancel/refund behavior checked against the contract) →
61
+ 7. end-to-end on the real workload path.
62
+ A PASS verdict binds to the **exact** account/model/region/protocol combination that produced layers 3-7; anything less is CONDITIONAL/FAIL/BLOCKED naming the missing layer.
63
+
64
+ **Load-methodology minima:** drive load open-loop at a fixed arrival rate — closed-loop concurrency hides throughput loss when requests slow down (coordinated-omission family: Schroeder et al., "Open Versus Closed", NSDI'06; Gil Tene's coordinated-omission analysis for the latency-distortion half). Report first-attempt results separately from retry-assisted results. Make warmup semantics explicit — including cold start measures the real user experience, excluding it measures steady-state capability; they are two different experiments, name which one you ran. Step the load (fractional → 1× → burst) with pre-set stop thresholds per stage, include a recovery segment (idle then back to 1×), and never continue to a higher stage merely to fill a report.
65
+
66
+ - A sample count below ~3 runs per cell must be reported as unreliable, never averaged into a verdict (team heuristic, no external source — raise the floor per your observed variance).
67
+
68
+ **Cost accounting:** separate the five cost frames — internal charge-back price, incumbent's actually-paid price, candidate's nominal price, candidate's measured price on the real workload, and cash cost (tax/discount/FX) — and state which frames are assumptions rather than quotes. Task submission ≠ success: success requires the expected terminal state, a usable artifact, normalized usage, and billing evidence; an HTTP 200 carrying an error envelope is a failure. The attempt ledger is append-only — a retry never overwrites a failed attempt's record.
69
+
46
70
  ## Fine-Tuning And Local Models
47
71
 
48
72
  - Treat fine-tuned models as registry versions with parent base model, training data lineage, training job id, parameter recipe, eval gate, and rollback path.
@@ -38,7 +38,7 @@
38
38
  - **When evolving the stream protocol, answer identity/lifecycle questions before designing event fields.** "What fields does the new event type carry" and "which existing identity/idempotency/billing model do these events bind to" are two different questions, and the second must be answered first: which answer/revision object the events belong to, what key their usage is accounted under, and whether a retry hides a provider-side billing fork. Event designs that skip this get reworked at review or implementation time.
39
39
  - **Answer-replacing stream events need explicit commit-point semantics — per plane, not one global commit.** For *answer visibility*, bind the revision's commit point to its first visible content delta: when cancel or degrade lands before commit, the previous answer must be preserved while the failure still surfaces to the caller, and caller-visible mutations (staged references/metadata) plus caller-side effects (tool dispatch) stay buffered until commit so a never-committed revision leaves no executed side effects to duplicate on replay. *Accounting is a separate plane*: provider-attempt usage/ledger records book per attempt under their own keys (per the usage rules below) regardless of whether the revision ever commits — an uncommitted revision still consumed provider tokens, and deferring or dropping its usage record is a billing hole, not tidiness. Post-commit failure semantics must also be an explicit spec decision, not an accident: when a revision commits (content already visible) and the stream later fails, either keep the partial replacement visible with the failure surfaced (progressive replacement) or roll back to the prior answer at terminal — name the choice and test it; leaving it implicit is how a partial revision silently clobbers a good answer. Define the semantics explicitly for streams with no content plane (tool-only, refusal, error-terminal): name which event commits them and how their outcome surfaces. This boundary is what chain-level replay tests force out — for stateful stream protocols, write chain-level (multi-event-sequence) tests at spec time, not after implementation.
40
40
  - **OpenAI Responses API supersedes Chat Completions as the recommended surface for new integrations** per `platform.openai.com/docs/guides/migrate-to-responses` and `developers.openai.com/blog/responses-api`. Two load-bearing differences: (a) **reasoning state preservation across turns** — Responses keeps the model's reasoning context alive between API calls, whereas Chat Completions drops it; OpenAI's published evals show ~3% SWE-bench Verified improvement vs Chat Completions with identical prompts on reasoning models, and ~40-80% better cache utilization. (b) **agentic-by-default**: a single API request can call multiple built-in tools (`web_search`, `image_generation`, `file_search`, `code_interpreter`, custom functions) in one round-trip — the orchestration that previously required client-side tool-loop scaffolding moves provider-side. **CRITICAL P0**: provider-side built-in tools bypass the team's own auth / audit / rate-limit / data-egress boundaries — calls to `web_search` happen at OpenAI's edge with OpenAI's network, calls to `file_search` index data the team uploaded to OpenAI's storage, calls to `code_interpreter` execute code in OpenAI's sandbox. Required: explicitly DISABLE the built-in tools in the Responses request unless a specific tool is approved for the use case (default-on is the wrong stance for any team with compliance / data-residency / audit obligations); when enabled, require post-response reconciliation that emits one audit event per built-in tool invocation (which tool, with what arguments, returning what footprint), enforce per-tool rate limits the team owns rather than relying on provider defaults, and document the data-egress policy (what user content can the team's prompts contain when web_search / file_search is on). Migration mechanics: structured-output config moved from `response_format` to `text.format`; `reasoning_effort` default is model-specific and drifts across model versions (see `model-prompt-evaluation.md`) — do not hardcode a flat default; read the current model's documented default per route. **When to migrate**: new integrations on GPT-5.x reasoning models default to Responses; existing Chat Completions integrations stay until reasoning-state preservation OR multi-tool orchestration is load-bearing — do not migrate as drive-by during feature work. Cross-provider abstraction layers (LangChain, LlamaIndex, an in-house adapter) need to surface the API choice explicitly because SDK shapes differ; verify the adapter speaks both forms before flipping callers.
41
- - **Prompt caching is a 2024-2025 industry-standard cost lever, not a niche optimization** — all three majors support it with different ergonomics. **Anthropic**: explicit `cache_control` blocks on prompt segments per `docs.anthropic.com/en/docs/build-with-claude/prompt-caching`; cache-read tokens are billed at ~10% of standard input price (NOT free — the 90% savings is the discount on those read tokens, not zero cost); cache-write is 1.25x base for 5-minute default TTL, 2x base for 1-hour TTL; the 1-hour TTL pays back vs uncached at roughly the third cache read once the write premium is amortized. Anthropic claims up to 90% cost savings on cache hits, but a workload that writes once and reads zero is a net loss. **OpenAI**: automatic caching on prompts ≥1024 tokens with no API change; Responses API improves cache utilization 40-80% vs Chat Completions per OpenAI's own announcement. **Google Gemini**: **implicit caching enabled by default for all Gemini 2.5+ models** per `developers.googleblog.com/en/gemini-2-5-models-now-support-implicit-caching/` — no API change needed; minimum cache-eligible request 1024 tokens (2.5 Flash) / 2048 tokens (2.5 Pro). **Architecture impact**: structure prompts so the cacheable prefix is stable (system prompt + tool/function defs + few-shot examples on top; per-request user content at the bottom). **Tenant-isolation P0**: cache keys / cacheable prefixes MUST be tenant-isolated — NEVER place tenant id, tenant-private context, tenant-specific tool definitions, or tenant-scoped few-shot examples inside the shared stable prefix unless the provider's account-isolation contract is verified AND tested (Anthropic and OpenAI account-scope caches per organization, but the moment two tenants share one API account / organization, a shared prefix that includes one tenant's context is a cross-tenant leak vector). Default position: tenant-scoped content goes BELOW the cache boundary; the stable prefix carries only tenant-neutral system prompts / tool defs / generic examples. Verify with a per-tenant cache-hit-rate audit — if tenant A's content hash ever produces a cache hit on tenant B's first request, the isolation is broken. **Footgun (Anthropic-specific)**: per Anthropic's docs, extended-thinking blocks interact with prompt caching in ways that invalidate cache entries more aggressively when the thinking state changes between turns; pin per-route extended-thinking decisions to avoid silent cost drift. For OpenAI / Google equivalents, this drop is plausible but not documented as a generalized rule at writing — measure cache-hit-rate per route before assuming the same dynamic. Treat cache-hit-rate as a per-route SLI emitted into the observability stack.
41
+ - **Prompt caching is a 2024-2025 industry-standard cost lever, not a niche optimization** — all three majors support it with different ergonomics. **Anthropic**: `cache_control` with two usage modes per `docs.anthropic.com/en/docs/build-with-claude/prompt-caching` — a top-level automatic mode (one request-level `cache_control` field; the system places the breakpoint at the last cacheable block) and explicit per-block breakpoints for fine control; cache-read tokens are billed at ~10% of standard input price (NOT free — the 90% savings is the discount on those read tokens, not zero cost); cache-write is 1.25x base for 5-minute default TTL, 2x base for 1-hour TTL; the 1-hour TTL pays back vs uncached at roughly the third cache read once the write premium is amortized. Anthropic claims up to 90% cost savings on cache hits, but a workload that writes once and reads zero is a net loss. **OpenAI**: automatic caching on prompts ≥1024 tokens with no API change; Responses API improves cache utilization 40-80% vs Chat Completions per OpenAI's own announcement. **Google Gemini**: **implicit caching enabled by default for all Gemini 2.5+ models** per `developers.googleblog.com/en/gemini-2-5-models-now-support-implicit-caching/` — no API change needed; minimum cache-eligible request 1024 tokens (2.5 Flash) / 2048 tokens (2.5 Pro). **Architecture impact**: structure prompts so the cacheable prefix is stable (system prompt + tool/function defs + few-shot examples on top; per-request user content at the bottom). **Tenant-isolation P0**: cache keys / cacheable prefixes MUST be tenant-isolated — NEVER place tenant id, tenant-private context, tenant-specific tool definitions, or tenant-scoped few-shot examples inside the shared stable prefix unless the provider's account-isolation contract is verified AND tested (Anthropic and OpenAI account-scope caches per organization, but the moment two tenants share one API account / organization, a shared prefix that includes one tenant's context is a cross-tenant leak vector). Default position: tenant-scoped content goes BELOW the cache boundary; the stable prefix carries only tenant-neutral system prompts / tool defs / generic examples. Verify with a per-tenant cache-hit-rate audit — if tenant A's content hash ever produces a cache hit on tenant B's first request, the isolation is broken. **Footgun (Anthropic-specific)**: per Anthropic's docs, extended-thinking blocks interact with prompt caching in ways that invalidate cache entries more aggressively when the thinking state changes between turns; pin per-route extended-thinking decisions to avoid silent cost drift. For OpenAI / Google equivalents, this drop is plausible but not documented as a generalized rule at writing — measure cache-hit-rate per route before assuming the same dynamic. Treat cache-hit-rate as a per-route SLI emitted into the observability stack.
42
42
  - **Keeping the prefix stable means making cost-only prefix churn cache-aware — and never deferring an authority/correctness change to save a cache write.** Any mid-conversation change to the cached prefix invalidates the cache from that point and pays the full uncached prefill on the next turn (plus, on providers with explicit cache-write pricing, a fresh cache-write charge), silently every turn after if it keeps mutating. Split such changes into two classes and handle them differently:
43
43
  - **Authority-neutral, cost-only churn** (re-ordering stable prefix content, refreshing an unchanged or non-authority system-prompt section, adding authority-neutral formatting/few-shot guidance that grants no new capability) — here the recommended pattern is **deferred invalidation**: don't rebuild the prefix as a side effect of an unrelated command; instead record the change as a durable `pending` mutation bound to principal/session/config-generation, apply it **before the next model request** (not at some far-off boundary, and never silently dropped — if applying it fails, fail closed to a degraded/visible state rather than continuing on stale-but-uncommitted intent), and offer an explicit opt-in (e.g. a `--now` / "apply immediately" path) to rebuild the prefix this turn.
44
44
  - **Authority/correctness/capability-surface changes** — ANY change to the callable/visible tool or capability surface (add, remove, revoke, schema/trust/exposure change), plus policy/authorization changes, memory deletion, model/behavior changes, and any privacy/principal/tenant/auth change — must invalidate / fence / re-authorize **NOW, never deferred for cost**. These follow the fail-closed snapshot/re-verify/re-authorize rules above and in `references/retrieval-agent-safety.md` / `references/agent-tool-dispatch.md`, which take precedence over cost-driven deferral. Deferring a revocation that leaves a revoked tool callable for the rest of a long turn — or deferring a newly-added tool so the model plans around a surface that isn't really there — is a correctness/security bug, not a saving. The only other routinely-acceptable mid-conversation prefix change is the deliberate context compaction already governed above (which invalidates and re-writes by design).
@@ -15,7 +15,7 @@ Avoid hard-coded model strings scattered through product code. Business logic sh
15
15
  ## Model Version Baseline 2025-2026
16
16
 
17
17
  - **Three major-provider model families are the credible production options at 2026-Q2**; the registry/router MUST verify exact current model strings against the live docs page before pinning, because provider naming cadence has accelerated (Anthropic ships ~quarterly minor revs, OpenAI ships sub-version reasoning-effort variants, Google ships Pro/Flash/Flash-Lite tiers separately). Authoritative-doc sources for live verification: `docs.anthropic.com/en/docs/about-claude/models` (Anthropic) / `developers.openai.com/api/docs/changelog` + OpenAI Help Center model release notes / `ai.google.dev/gemini-api/docs/changelog` (Google). General shape: Anthropic ships Opus (frontier reasoning) + Sonnet (workhorse, 200K default + 1M beta context per `anthropic.com/news/1m-context`) + Haiku (cost/speed); OpenAI ships GPT-5.x with a `reasoning_effort` parameter (default value and available effort tiers are model-version-specific and drift — do not hardcode; verify per target model per the next bullet) plus reasoning summaries; Google ships Gemini 2.5 Pro + Flash + Flash-Lite with explicit thinking-budget control. Per-route choice is a registry decision recorded in the project's model registry, not hardcoded. **Pinning + eval-baseline discipline (P0)**: pin the exact model version string in the registry config (NOT just the family name like `claude-sonnet`); block ENV-var-only flips between model versions in production (flipping `MODEL=...` to a new minor rev silently invalidates every prompt eval result against the old version); require evaluation replay or shadow comparison for ANY model-string or default-param change before flipping production traffic; emit a metric on per-route model-string drift so an unintended autoupgrade is observable, not just discoverable at the next eval. The recurring failure is: team upgrades minor rev for cost/quality reason, eval suite still passes BECAUSE eval prompts didn't exercise the regressed code path, production sees the regression a week later.
18
- - **Extended thinking / reasoning effort is a per-request product decision, not an account-level default**. Anthropic's extended thinking (off by default) trades latency + token cost for accuracy on multi-step reasoning per `docs.anthropic.com/en/docs/build-with-claude/extended-thinking`; impacts prompt-caching efficiency (cache invalidation more aggressive with thinking on). OpenAI's `reasoning_effort` tunes the same axis per `developers.openai.com/api/docs/guides/reasoning`, but the **default differs by model version and is not monotonic** — defaults have varied across GPT-5.x revs (different revs ship different defaults; do not assume a trend) and new effort tiers (e.g. `xhigh`) appear in some revs. Do NOT hardcode a version-specific default here; confirm the default for the exact model you target against its live docs page before relying on "default behavior" — drift in this default is the most common silent cost/latency change across an OpenAI minor-version bump. `revalidate-when: OpenAI ships a new GPT-5.x reasoning rev`. Gemini exposes a thinking budget knob per `developers.googleblog.com/en/gemini-2-5-thinking-model-updates/`. **Pattern**: per-route declare reasoning intent (off / low / medium / high) based on the task's complexity AND record the decision in the prompt/route config so latency and cost regressions are traceable to an explicit knob, not provider-default drift. **Production-safety contract for reasoning enablement (P0)**: declaring reasoning intent is not enough — the per-route config MUST also carry (a) p95/p99 **latency budget** with downstream timeout propagation (reasoning can add seconds-to-tens-of-seconds to single-turn latency; downstream HTTP timeouts that worked at non-reasoning latency will fail), (b) **reasoning-token cap** to prevent runaway thinking burning 10-50× the non-reasoning token budget on a hard prompt, (c) **cost guardrail** as a per-request token + per-user-session aggregate ceiling, (d) **rollback path** to flip reasoning off at the route level under incident. Mid-route enablement of reasoning (turning thinking on after launch) is a load-bearing change that requires fresh load/eval evidence and SRE sign-off, NOT a config-only flip. Reasoning content is a separate channel from user-visible content — never persist reasoning into product history or display it as final output (see streaming rules in `llm-client-gateway.md`).
18
+ - **Extended thinking / reasoning effort is a per-request product decision, not an account-level default**. Anthropic's thinking model is **generation-dependent — verify against the live docs for the exact target model**: on Claude 4.5-and-earlier thinking models, extended thinking is off by default and enabled per request (`thinking: {type:"enabled", budget_tokens}`); that manual mode is deprecated on 4.6 and rejected with a 400 on 4.7+; on the current generation (Claude 5 family) thinking is on by default (adaptive) with `effort` controlling depth (`platform.claude.com/docs/en/build-with-claude/thinking`). Thinking trades latency + token cost for accuracy on multi-step reasoning; it also impacts prompt-caching efficiency (cache invalidation more aggressive when thinking state changes). OpenAI's `reasoning_effort` tunes the same axis per `developers.openai.com/api/docs/guides/reasoning`, but the **default differs by model version and is not monotonic** — defaults have varied across GPT-5.x revs (different revs ship different defaults; do not assume a trend) and new effort tiers (e.g. `xhigh`) appear in some revs. Do NOT hardcode a version-specific default here; confirm the default for the exact model you target against its live docs page before relying on "default behavior" — drift in this default is the most common silent cost/latency change across an OpenAI minor-version bump. `revalidate-when: OpenAI ships a new GPT-5.x reasoning rev`. Gemini exposes a thinking budget knob per `developers.googleblog.com/en/gemini-2-5-thinking-model-updates/`. **Pattern**: per-route declare reasoning intent (off / low / medium / high) based on the task's complexity AND record the decision in the prompt/route config so latency and cost regressions are traceable to an explicit knob, not provider-default drift. **Production-safety contract for reasoning enablement (P0)**: declaring reasoning intent is not enough — the per-route config MUST also carry (a) p95/p99 **latency budget** with downstream timeout propagation (reasoning can add seconds-to-tens-of-seconds to single-turn latency; downstream HTTP timeouts that worked at non-reasoning latency will fail), (b) **reasoning-token cap** to prevent runaway thinking burning 10-50× the non-reasoning token budget on a hard prompt, (c) **cost guardrail** as a per-request token + per-user-session aggregate ceiling, (d) **rollback path** at the route level under incident — on request-gated generations, flip thinking off; on the current always-on generation, drop effort to the minimum tier or route to a non-thinking model, and record which lever the route uses. Mid-route enablement of reasoning (turning thinking on after launch) is a load-bearing change that requires fresh load/eval evidence and SRE sign-off, NOT a config-only flip. Reasoning content is a separate channel from user-visible content — never persist reasoning into product history or display it as final output (see streaming rules in `llm-client-gateway.md`).
19
19
 
20
20
  ## Prompt Registry
21
21
 
@@ -80,7 +80,10 @@ Every material model or prompt change should define:
80
80
  - for compared runs, hold the conditions fixed — everything except the declared variable under test, with both values of that variable recorded (model string, prompt version, retrieval index, tool set, and decoding parameters are the usual fixed set);
81
81
  - quality metrics such as accuracy, recall, consistency, parse success, or groundedness;
82
82
  - runtime metrics such as latency, token usage, success rate, fallback rate, and cost;
83
+ - cross-provider token/usage comparison only after normalization — providers account cache/reasoning tokens differently (some list cache reads beside input, some fold them into prompt tokens), so raw `totalTokens` is comparable only between runs of the same case input after the runner normalizes accounting; per-case ratio aggregates (cache-hit rate, cost/case) include only cases with usage on all compared arms, with the excluded-case count reported next to every ratio; total-cost accounting is the separate aggregate that never excludes — a billed attempt missing usage is a data defect counted against that provider and its cost enters via billing records, never silent exclusion from totals — so omission cannot flatter an arm on either aggregate;
84
+ - where billing records are available, reconcile reported usage against them for at least a sample — usage self-reported by the pipeline is one layer weaker than the invoice;
83
85
  - regression examples and human review notes when judgment is subjective.
86
+ - keep human-review fields (reviewer, rating, notes) and machine-produced fields in separate columns/records with distinct writers: an automated re-run must never overwrite a human verdict, and a retry archives a new attempt record instead of replacing the failed one.
84
87
 
85
88
  Store evaluation reports with the model/prompt versions compared so future changes can reproduce the decision.
86
89
 
@@ -60,3 +60,8 @@ Use this checklist for payment, quota, order, publishing, generated content, acc
60
60
  - Callback/reconciliation path for payment or async work, with order/request id matched on both sides.
61
61
  - Support-visible order/request id surfaced to the user; without it, support cannot disambiguate "I paid but app says canceled".
62
62
  - Safe retry and user explanation for uncertain final state — never present "succeeded" or "failed" when reconciliation has not run.
63
+
64
+ ## 组件库版本纪律
65
+
66
+ - 组件库 API(taroify / NutUI-Taro / tdesign 等)按 lockfile 实装版本写,勿凭训练记忆——prop 名与默认值跨大版本会变。
67
+ - 库自带 lint/codemod 或迁移清单时,变更与大版本迁移以它对 changed files 收口(deprecated 用法与 a11y 规则通用 lint 不认识)。
@@ -40,6 +40,7 @@ Node.js uses a small number of threads to serve many clients. A long callback re
40
40
  ## Errors and process lifecycle
41
41
 
42
42
  - Catch errors at boundaries that can make a valid decision: translate, retry under policy, compensate, or fail the operation. Otherwise preserve `cause` and propagate.
43
+ - For durable/task state written by one operation, capture the clock once per transition (all timestamps the transition itself stamps come from a single `now`; domain-provided times are recorded as received, never re-stamped), and validate external-response structure (schema parse with a typed error path) before mapping — a malformed upstream payload must become a persisted failure transition carrying the canonical error (mirroring the go/python state-machine rendering), never a crash or a silently-defaulted field.
43
44
  - Treat unknown `uncaughtException` and default-throw unhandled rejection paths as fatal. A handler is for synchronous cleanup/diagnostics before termination, not resuming normal operation from an undefined state.
44
45
  - On `SIGTERM`/the platform's shutdown signal:
45
46
  1. mark readiness false or otherwise stop new routing;
@@ -46,6 +46,7 @@ A new service must satisfy all five before it is allowed in production. A releas
46
46
  - Logger MUST extract trace context from `context.Context` and attach `_trace_id`/`_span_id`/`_trace_flags` to the log record automatically. If a log call requires the developer to manually pass trace fields, the framework is broken.
47
47
  - Async work that detaches from the request lifecycle (audit, reporting, ledger side-paths) keeps correlation — but only correlation: capture trace/span/log-id fields before the handoff and carry them (plus durable work-item identity) into the detached context; never a bare `context.Background()` that drops linkage, and never a wholesale request-context copy — `context.WithoutCancel` preserves every request value, so a naive derive smuggles request auth/session/secrets/PII into work that outlives the request and can run under stale, revoked authority. Whitelist correlation fields, strip the rest; the worker re-authorizes or runs under service identity (lifecycle/loss-policy detail owned by the stack dev skills' async side-path rules — this bullet owns only the correlation contract).
48
48
  - Verification: correlation is an acceptance item, not a default — wiring the unified logging component is NOT evidence it works. Assert over the real transport: drive one request through the actual RPC/HTTP server in a test, assert the handler ctx carries a non-empty trace-id/log-id, and that business + access log lines actually contain those fields. Then live: pick any production user action → find one log line → click trace-id → see the full multi-hop trace → all spans share the same log-id.
49
+ - Query/dashboard discipline (evidence-boundary, non-intrusion, discover-before-create, env-resolution): `references/metrics-conventions.md`.
49
50
 
50
51
  ### R2 — Framework default observability, not opt-in
51
52
 
@@ -117,8 +118,8 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
117
118
  ### R8 — SLI/SLO discipline
118
119
 
119
120
  - Define SLIs from the **user's perspective** (success rate of a request type, latency of a user-visible action), not from internal counters.
120
- - Prefer expressing an SLI as **good events / valid events** (or good windows / valid windows), per Google SRE — define valid events/windows first, then the good ratio. A latency SLI is the **proportion of requests faster than a threshold** (`count(latency ≤ T) / total`, from histogram buckets), NOT a percentile value — P95/P99 are dashboard aids, not the SLI. An error-log counter is a diagnostic signal, not an availability-SLI input (it is skewed by log sampling, dedup, and async/non-request errors).
121
- - Signals you cannot compute are blind spots, not near-coverage: record each one explicitly ("can't measure X because Y" — e.g. a failure counter with no attempt total yields no error *rate*; content not collected means input-semantic drift is unmeasurable) in a blind-spot register instead of pretending coverage, so on-call never leans on a signal that does not exist.
121
+ - Prefer expressing an SLI as **good events / valid events** — valid first, then good — per Google's Art of SLOs (the Workbook itself says good/total). A latency SLI is the **share of valid requests faster than a threshold** (`count(latency≤T)/valid` — the denominator is the availability SLI's valid set, never raw total), NOT a percentile — P95/P99 are dashboard aids, not the SLI. An error-log counter is a diagnostic signal, not an availability-SLI input.
122
+ - Uncomputable signals are blind spots, not near-coverage: record each explicitly ("can't measure X because Y" — e.g. a failure counter with no attempt total yields no error *rate*; uncollected content means unmeasurable input-semantic drift) in a blind-spot register instead of pretending coverage, so on-call never leans on a nonexistent signal.
122
123
  - Each user-visible journey gets at least one availability SLI + one latency SLI.
123
124
  - Set SLO targets, error budgets, and burn-rate alerts (SRE Workbook multiwindow tiers, derived for a 30d budget window — recompute for other periods: page 14.4× 1h/5m, page 6× 6h/30m, ticket 1× 3d/6h). Each tier MUST evaluate its long AND short window together, firing only when both burn above threshold; the short (~1/12) window makes paging stop soon after the burn stops.
124
125
  - An SLO without an error-budget-driven release decision is decoration. See `references/sli-slo-design.md`.
@@ -50,7 +50,7 @@ Async observable gauges via OTel SDK reduce noise vs synchronous sets. Pattern:
50
50
  | 外部框架 | 关注什么 | 适用对象 | 映射到本 ref instrument |
51
51
  |---|---|---|---|
52
52
  | **Google 4 Golden Signals**(SRE Book)| Latency / Traffic / Errors / Saturation | service-level(user-facing service)| Latency → Histogram;Traffic → Counter (request rate);Errors → Counter (error count);Saturation → Gauge (queue / pool / cpu / mem) |
53
- | **RED**(Tom Wilkie / Weaveworks)| Rate / Errors / Duration | request-driven service / RPC endpoint | Rate → Counter;Errors → Counter;Duration → Histogram(近似等于 request-side Golden Signals 去 Saturation;非正式 derivation)|
53
+ | **RED**(Tom Wilkie;原始 Weaveworks 博客已随公司关停下线,现存最佳出处为 Grafana 官方博文 "The RED Method")| Rate / Errors / Duration | request-driven service / RPC endpoint | Rate → Counter;Errors → Counter;Duration → Histogram(近似等于 request-side Golden Signals 去 Saturation;非正式 derivation)|
54
54
  | **USE**(Brendan Gregg)| Utilization / Saturation / Errors | resource(CPU / memory / disk / NIC / connection pool)| Utilization → Gauge (%) ;Saturation → Gauge (queue depth);Errors → Counter(资源驱动而非请求驱动)|
55
55
 
56
56
  何时用哪个:
@@ -103,3 +103,10 @@ The 15s reader interval matches Prometheus scrape conventions and keeps OTLP pus
103
103
  - Using gauge for monotonic counters (loses rate semantics on restart).
104
104
  - Histogram with too few buckets (loses percentile precision) or too many (storage cost).
105
105
  - Recording metrics inside a tight loop without rate-limiting (collector receives bursts).
106
+
107
+ ## Observation Discipline(查询/看板/新增信号的通用纪律)
108
+
109
+ - Cross-layer evidence boundary: a client-side event does not prove backend success, and a backend metric does not prove the user saw success — any conclusion crossing the client/backend (or service/service) boundary requires identifiers or time windows aligned across the layers, never a same-shape count on each side.
110
+ - Observation code never intrudes on the observed path: **diagnostic telemetry** (metrics, traces, debug logs) is best-effort and must not add retries, blocking waits, or business-logic branches to the monitored path — observability that changes the behavior it measures is its own defect class. This best-effort license covers diagnostics only: mandatory records — security audit, billing/ledger, deletion/erasure evidence, release-gate records — are business writes with their own durable-delivery and failure semantics (their loss is a failure, never shrugged off as telemetry).
111
+ - Discover before creating: before proposing a new event, metric, label, panel, or query, enumerate what already exists for that surface and extend/reuse it — parallel near-duplicate signals fragment dashboards and split history.
112
+ - Environment-name resolution: a user's explicit component/branch/URL/datasource always wins; never silently substitute a different physical environment for a colloquial environment word — resolve an unqualified name to the recorded default and say which one was used.
@@ -45,11 +45,11 @@ If you cannot write the query against existing metrics, the metric set is incomp
45
45
 
46
46
  ## SLO target
47
47
 
48
- Pick a number that reflects user expectation and product maturity, not aspiration:
48
+ Pick a number that reflects user expectation and product maturity, not aspiration. The ladder below is a **team-heuristic starting point, not an industry standard** — Google SRE literature deliberately gives no numeric ladder; its direction is that each extra nine costs sharply more for marginal utility approaching zero (SRE Workbook Ch.2), and that targets should come from user expectation, not current performance:
49
49
  - New service: 99.0% — generous.
50
50
  - Mature service: 99.9% — three nines.
51
51
  - Critical path (payment, login): 99.95% — push.
52
- - Never start with 99.99% without 24/7 staffing.
52
+ - Never start with 99.99% without 24/7 staffing (team heuristic: at four nines the monthly budget is minutes, which no unstaffed rotation can defend).
53
53
 
54
54
  Lower the target if every release burns the budget. Raise it only after sustained achievement.
55
55
 
@@ -40,10 +40,14 @@ canary_check_task:
40
40
  # Filter by baseline label `lane` (canonical per metrics-conventions
41
41
  # Baseline labels), not raw mesh subset name.
42
42
  query: 'sum(rate(http_server_request_error_total{service="<svc>", lane="prod-canary-<v>"}[5m])) / sum(rate(http_server_request_total{service="<svc>", lane="prod-canary-<v>"}[5m]))'
43
- threshold: < 0.01 # 1% error budget
43
+ threshold: < 0.01 # 1% rolling error-rate threshold (5m window; not an SLO error budget) — Flagger's builtin-check example value
44
44
  - name: latency_p99
45
45
  query: 'histogram_quantile(0.99, sum(rate(http_server_request_duration_seconds_bucket{service="<svc>", lane="prod-canary-<v>"}[5m])) by (le))'
46
- threshold: < 0.5 # seconds (500 ms)
46
+ threshold: < 0.5 # seconds (500 ms) — Flagger's builtin-check example value
47
+ # Threshold provenance: <1% error rate / p99<500ms match Flagger's official
48
+ # builtin metric checks (request-success-rate min 99, request-duration max 500ms).
49
+ # They are NOT a cross-tool default (Argo Rollouts' official example uses 95%
50
+ # success); tune per service SLO — the SLO, not the template, is the authority.
47
51
  abort_thresholds: # immediate abort if these hit
48
52
  - name: error_spike
49
53
  query: 'sum(rate(http_server_request_error_total{service="<svc>", lane="prod-canary-<v>"}[1m]))'
@@ -155,6 +159,16 @@ Rollback path post-promotion:
155
159
  2. Walk the same rollout flow.
156
160
  3. Or, in extreme cases, immediately set the new bad version's weight to 0 and old version's to 100.
157
161
 
162
+ ## Ramp/Rollback Data Literacy (user-level experiments and flags)
163
+
164
+ When the rollout is user-bucketed (A/B, percentage flag, allowlist) rather than pod-weighted, the decision data has its own failure modes; a ramp/rollback/graduate decision that ignores them reads a lying funnel:
165
+
166
+ - **Exposure without a bake point lies.** Count a user as "on treatment" only from an exposure event emitted at the moment the treatment actually took effect for them — an assignment record without exposure means the funnel includes users who never saw the change.
167
+ - **Funnel invariant:** for any feature, outcome unique-users ≤ exposure unique-users — a necessary sanity check, never attribution proof (disjoint sets can satisfy it): attribution requires joining outcome events to the same users' exposure with outcome.time > exposure.time. A dashboard violating the invariant has a data-pipeline defect, not a product effect.
168
+ - **Analytics profile attributes are current state, not post-assignment behavior.** A profile-filtered dashboard cannot express `event.time > assigned_at`; ramp/rollback/graduate decisions must not rely on profile-filtered boards alone — use exposure/outcome event joins, and pair them with failure logs and cost data.
169
+ - **Control-plane query ≠ execution fact.** Reading the flag service proves only what policy is *stored* at query time. Policy, assignment, exposure, actual execution path, and outcome are five separate evidence layers — never infer a user's actual arm from the configured weights.
170
+ - **Bucketing salt is immutable mid-flight.** Changing the salt re-buckets every user (treatment users silently swap arms), destroying the experiment; weight and target changes have defined semantics for existing vs new users — write down which one the platform gives before ramping.
171
+
158
172
  ## Verification
159
173
 
160
174
  - Trigger a synthetic SLI breach during canary → control plane logs detection within bake window → automatic abort fires.
@@ -133,6 +133,15 @@ Rules:
133
133
  - If a metric cannot be computed from the audit log, that is an audit-log gap: fix the event capture, don't estimate the metric.
134
134
  - Before publishing any of the five, verify the audit-log schema actually carries the correlation keys that metric joins on — at minimum: commit timestamp and production-exposure (traffic-reached) timestamp per deployment (lead time); a stable logical-deployment ID plus a distinct per-attempt ID (deployment frequency, and the failed-attempt/retry separation above); a causal link from each remediation event to the deployment it remediates (change fail rate); an incident link and a recovery-confirmed event (rework rate, recovery time). A metric whose keys are absent is unavailable — an audit-log gap per the rule above — not a license to join on wall-clock proximity or event-name heuristics; proximity joins are exactly how a failed attempt gets deduped into its later retry or a recovery gets inferred from an unrelated healthy reading.
135
135
 
136
+ ## Merge Topology For Gitflow-Style Release Trains
137
+
138
+ Applies when a project runs Gitflow-style release branches (the default preference stays short-lived branches / trunk-based per `product-rd-workflow/references/delivery-lifecycle.md`; adopt this section only where release trains are genuinely required). Grounded in AWS Prescriptive Guidance's Gitflow pattern:
139
+
140
+ - **Squash is direction-sensitive.** feature/bugfix → develop merges use squash (clean linear history). release → main and every back-merge **must not squash** — prefer fast-forward/plain merge. Why (git topology): squashing on higher branches rewrites commit identity, so the back-merge can no longer recognize already-merged changes and produces repeated conflicts or silently dropped work (AWS: "Only use a squash merge when you are merging from a feature branch to a develop branch").
141
+ - **Back-merge promptly, both targets.** At release close, merge the release into main AND back into develop as soon as possible ("as soon as possible to consolidate work back into the primary branches") — a skipped back-merge means the next release train overwrites what only main has (the classic lost-hotfix incident).
142
+ - **Tag only after the human confirms the main merge.** The production tag is a deploy trigger, not bookkeeping — never tag ahead of the confirmed merge, and stop after MR creation unless asked to proceed (composes with `release-coordination`'s authorization matrix, which owns who may confirm).
143
+ - **Hotfix keeps the full environment ladder.** A hotfix branches from main and walks every promotion environment — it compresses cycle time and priority, never skips a gate (the emergency-override section below is the one sanctioned, loudly-logged exception); and it back-merges like any release into develop AND into any open release branch (an active train missing the hotfix reverts it at that train's own merge), or the next train reverts it.
144
+
136
145
  ## Emergency override
137
146
 
138
147
  For incidents where the gate must be bypassed (e.g. roll out an emergency fix faster than canary allows):
@@ -7,7 +7,7 @@ description: 加功能 / 新需求 / 技术方案 / 方案评估 / 技术选型
7
7
 
8
8
  Use this skill as the top-level workflow for new product development, feature delivery, bug handling, refactoring, release preparation, or repeated process improvement. It is not a replacement for stack-specific skills; it decides which skill should own each stage and what evidence is required before moving on.
9
9
 
10
- **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack or execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
10
+ **Entry precedence.** For any product idea, feature delivery, release, cross-cutting refactor, or a **restart/redo of an in-flight delivery**, invoke this workflow first to classify and route — naming a stack/execution skill (e.g. `web-react-dev`, `multi-agent-delegation`) does not by itself skip this workflow's lifecycle gates (design / test / release / acceptance); those still apply unless already covered.
11
11
 
12
12
  **Continuation-proposal output contract (session-wide for product delivery).** Every assistant message in a delivery session routed by or through this workflow carries exactly one literal line until the user explicitly ends/pauses the delivery or changes scope: `proposed-next: <action and scope>` when the message has imperative/future/next-step wording or a proposed action, otherwise `proposed-next: none — status only`. At the start of every subsequent user turn, read that line before interpreting the reply: absence, multiplicity, or a marker/wording conflict enters `blocked:`/`interim` by default, never `not-applicable`. Coverage detail and the host-layer caveat: `references/pre-final-continuation-gate.md` (Continuation-proposal output contract).
13
13
 
@@ -21,11 +21,11 @@ Use this skill as the top-level workflow for new product development, feature de
21
21
  - A general-purpose process skill that merely *looks* like the obvious start — brainstorm/scope-shaping, plan-writing, or TDD auto-suggested by ANY channel: a session-start prompt, an optional skill package, or the host platform's native skills listing (including a listed entry skill's own self-invocation mandate, e.g. "must invoke if there is a 1% chance") — does not replace this entry: suggestion-channel wording is channel self-promotion, not routing authority (a host-mandated preflight — mandated by a host-authored system/developer-level or equivalent higher-priority instruction — may run first without thereby becoming the delivery owner; the test is AUTHORSHIP, not rendering position: a host-authored instruction counts even when rendered within the listing surface, while a skill's own description/content claiming preflight status never does); invoke this workflow as the delivery entry (immediately after any genuine host-mandated preflight), then call that skill inside the stage it serves (for example, requirement shaping in Workflow step 1).
22
22
  - For delegated agent execution, `multi-agent-delegation` owns the execution recipe and worker verification while this workflow owns the lifecycle gate and acceptance boundary; delegated agents resuming after a pause must receive or re-verify the current plan/spec artifact set before editing.
23
23
  - For multi-repo delivery, or any delivery that changes remote branch, MR, pipeline, release, or deployable-artifact state, maintain a compact per-changed-unit delivery-status ledger (row schema + persistence rules in `references/status-tracker-sync.md`) before claiming done, recommending MR/merge, or choosing the next slice. A required verification/review gate passes only when its state is `success`, `not-applicable`, or `not-required`; any other state blocks a done/merge recommendation, so remediate, wait to a terminal state, or report the delivery as pending with the unblock action. Residual-risk acceptance by the user only permits the recommendation/handoff label for that concrete action; it does **not** authorize merge, auto-merge, default-branch push, or cleanup, which still require the user's explicit merge instruction for the current MR (per `worktree-isolation`). Small local-only multi-file edits use the normal concise status unless they introduce remote, CI, MR, release, or deployable-artifact state.
24
- - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack or execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
24
+ - **Precedence order:** explicit user instruction > this workflow's stage/gate ownership > lower-priority default behavior from optional skill packages — but naming a stack/execution skill to carry out delivery work is not itself an instruction to waive those gates; a gate opt-out must be stated as such. See *External Skill Augmentation* for how external skills supplement specific disciplines without taking over ownership.
25
25
 
26
26
  ## Scope
27
27
 
28
- - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed or stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
28
+ - Product and requirement shaping: clarify user workflow, success criteria, non-goals, constraints, and acceptance checks. When shaping a feature, build a per-point **acceptance-coverage matrix** — one independently-failable behavior = one point, plus the risk points it touches, each mapped to an observable pass/fail acceptance check; an owning spec that already carries this coverage satisfies the gate when referenced **per point** (name which section covers each point; unnamed/stale points are gaps to fill) — don't re-author it. Risk-point enumeration and observable-criteria detail live in `references/delivery-lifecycle.md` §Product Shaping Checklist.
29
29
  - **Implementation completeness/minimality gate.** For behavior-changing delivery this gate fires — functional completeness and structural minimality are independent gates — and gaps block `complete` — **load `references/implementation-completeness-and-minimality.md` before the mapping**: map every in-scope acceptance point to implementation plus fresh evidence, and speculative future need and omitted required behavior both fail — the independent-gates rule and the acceptance/concept matrices live there, and every retained new concept must map to a current acceptance point or hard constraint.
30
30
  - Product requirement artifacts: route clarification to `requirement-intent`, conditionally-required current-state evidence to `requirement-baseline`, scope/version/appetite to `requirement-scope`, and Ready-only human-readable PRD assembly to `requirement-doc-writer`. All four use the canonical `requirement-doc-writer/references/requirement-closure-contract.md`. In active delivery, each narrow artifact returns here after completion; only a standalone non-PRD narrow artifact may return directly to the user. “只写个 PRD / 不走流程”仍需 lifecycle-issued Ready;WIP/会议材料不得命名为 PRD. This workflow owns the complete lifecycle's `PRD Ready` / `PRD Not Ready` verdict and cross-owner closure coordination.
31
31
  - Design routing: decide when interaction, information architecture, visual system, design-system, or UX acceptance work must happen before implementation.
@@ -35,9 +35,9 @@ Use this skill as the top-level workflow for new product development, feature de
35
35
  - Defect routing: use `defect-diagnosis` for hands-on reproduction, isolation, instrumentation, fix, regression verification, root-cause analysis, and prevention routing.
36
36
  - Research routing: For research groundwork behind a selection or assessment decision (技术选型、方案评估前的主题调研), this workflow must call `multi-perspective-research` for an evidence-grounded brief; the verdict itself stays with this workflow's gates.
37
37
  - Existing-project assessment routing: for requests that ask to analyze a repository, product codebase, or project quality across architecture, implementation, tests, UI/UX, bugs, or risks, run codebase understanding first, then route each assessment dimension to the smallest owning skill instead of treating understanding as the final answer.
38
- - AI/algorithm product launch discipline: for new or iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
- - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness or freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
- - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new or public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
38
+ - AI/algorithm product launch discipline: for new/iterative algorithm capabilities, require product goal, business acceptance baseline, offline evaluation baseline, engineering serving baseline, rollout/rollback plan, and risk owner before launch. The workflow owns the gate; `llm-inference-integration`, testing, release, and stack skills own their narrower execution details.
39
+ - User-visible tips, nudges, release notes, and update notices: treat these as product surfaces, not harmless copy — define audience, eligibility, suppression rules, user-disable path, freshness label, repeat/cooldown policy, maximum interruption level, and success/abuse metrics before launch. Context-derived tips need a privacy review of what local behavior, files, tools, account state, or capability signals may influence eligibility; incomplete, stale, or cached update data must not imply completeness/freshness. Route terminal rendering to `terminal-cli-dev`, visual hierarchy/accessibility to `product-ui-ux-design`, behavior-changing defaults/migrations to `platform-release-engineering`, analytics/diagnostic redaction to `platform-observability`, scenario coverage to `testing-strategy`.
40
+ - **Developer-facing surfaces** (CLI / SDK / library / public API / developer docs): the user is a developer, so developer experience is an acceptance dimension. **For a new/public developer surface, or a change touching onboarding, install/setup, first-success, defaults, error surfaces, or a breaking migration**, prove DX by the **measured onboarding journey** — run the real discover→install→first-success path as a new user; do not infer DX from README / feature-list quality. Error messages are a first-class acceptance item; a non-safety default needs a safe override or a documented no-escape rationale; a **safety / security default stays fail-closed** (widening needs risk-owner approval; DX never licenses an `--insecure` bypass); breaking changes need a migration path. Journey metrics, segment proof, blocked handling, and per-surface executor routing: `references/verify-developer-experience.md`.
41
41
  - **Artifact-egress confidentiality gate**: before a delivery artifact (spec/plan/requirement, status/writeback doc, launch/task card, retrospective) — or any text generated from it — **crosses the local trusted boundary**, run a confidentiality pass *before* the write/create/push/send call; it fires **only on cross-boundary egress** and owns only the **semantic confidentiality axis** that secret scanners miss. A block stops the *egress* but preserves the draft locally and never deletes the work; exception authority stays with the user/owner. Egress channels include chat/tracker surfaces (Feishu/Bitable writeback, shared/public docs, MR/issue body or review comment, external trackers), durable VCS/release metadata (commit message, branch/tag name, release note/changelog, CI metadata), external assistants/models, and delegated-worker prompts; the gate's own block reason/finding is itself egress — never quote the raw sensitive span across the boundary; secrets/PII/credentials/raw-logs/customer-data route to their existing owners (`platform-observability` redaction, `defect-diagnosis` evidence sanitization, the `feature-risk-router` security-review gate); per-category actions, severity, and the delegated-worker firing point: `references/artifact-egress-confidentiality.md`.
42
42
  - Learning loop: after bugs, review findings, incidents, repeated friction, or external skill research, use `skill-extraction-workflow` to update the right skill or reference instead of leaving knowledge only in chat.
43
43
  - Skill/process extraction: when asked to summarize delivery experience, preserve a workflow lesson, update a reusable skill, or decide where a lesson belongs, route to `skill-extraction-workflow` first. Update this skill only when the lesson changes product R&D routing, gates, ownership, or lifecycle policy.
@@ -190,13 +190,13 @@ At each stage boundary, walk the per-stage entry-state enumeration in [Stage-Ent
190
190
  Run this gate before finalizing a product R&D turn after any delivery slice lands. The session-wide continuation-proposal output contract above creates a second, independent trigger at the start of every subsequent user turn in that delivery session when either (a) the immediately preceding assistant message carries an action-form `proposed-next:` and the user replies, (b) its marker is absent/multiple or `none` conflicts with imperative/future/next-step wording, or (c) a user reply reads as affirmative/permissive toward an explicit assistant-proposed next action. Paths (a) and (b) are unconditional literal/fail-closed checks. Before any further action or final response, visibly emit exactly one of `continuing: <action and scope>` or `blocked: <proposed action and scope> — <specific stop, missing authority, or ambiguity>`; emitting neither or both is invalid. Path-(c) examples and the full outcome contract: `references/pre-final-continuation-gate.md` (Gate triggers and outcome contract).
191
191
 
192
192
  1. Confirm the landing state from real evidence (local branch, remote sync, MR/review artifact, CI/pipeline when applicable, review/challenge status when required, status-doc sync, dirty worktree), proving the landing before reading any document (for the remote-backed default, fetch/update the target ref from its remote immediately before classifying the slice landed) per `references/pre-final-continuation-gate.md` (Landing-state proof); content/tree/patch equivalence never by itself proves a slice landed.
193
- 2. Inspect the current product/status source of truth, issue list, repo-local next-step artifact, unresolved acceptance item, or direct user continuation instruction for the next implied slice. **Reconcile it against the current branch/MR/merge/CI/tag state from step 1 before deriving: if it contradicts reality it is stale — stop, repair the status source first, and do NOT derive from the stale source or a git-log/grep scan; needing to grep history to guess the next slice is itself a stale-source signal** (`references/pre-final-continuation-gate.md` §Status-source reconciliation).
194
- - **Deferred-evidence continuation check (`DFE-CONT`).** When real/runtime evidence is due (named by an acceptance item, status source, landing-evidence row, required gate, user correction, or because it is the behavior's only meaningful proof) yet deferred, blocked after remediation, skipped at finalization, or replaced by local/mock verification. Report deferred real evidence as `interim` / outstanding and do NOT report the turn complete while it is outstanding. A local/mock substitution is terminal only when a cited **non-agent** anchor — **agent-authored or agent-co-edited status/router/gate/handoff text never satisfies this** — names the same evidence, declares the deferral terminal, and carries the outstanding command/source forward for the active slice/ref. Never add verifier/config/test hardening motivated only by missing deferred evidence, and never auto-continue past the pending gate. **Load `references/pre-final-continuation-gate.md` before treating any deferral as terminal** — it owns the full valid/invalid-anchor list and hardening boundary.
193
+ 2. Inspect the current product/status source of truth, issue list, repo-local next-step artifact, unresolved acceptance item, or direct user continuation instruction for the next implied slice. **Reconcile it against the current branch/MR/merge/CI/tag state from step 1 before deriving: contradicting reality means stale — stop, repair the status source first, and do NOT derive from the stale source or a git-log/grep scan** (`references/pre-final-continuation-gate.md` §Status-source reconciliation).
194
+ - **Deferred-evidence continuation check (`DFE-CONT`).** When real/runtime evidence is due (named by an acceptance item, status source, landing-evidence row, required gate, user correction, or because it is the behavior's only meaningful proof) yet deferred, blocked after remediation, skipped at finalization, or replaced by local/mock verification. Report deferred real evidence as `interim`/outstanding; do NOT report the turn complete while it is outstanding. A local/mock substitution is terminal only when a cited **non-agent** anchor — **agent-authored or agent-co-edited status/router/gate/handoff text never satisfies this** — names the same evidence, declares the deferral terminal, and carries the outstanding command/source forward for the active slice/ref. Never add verifier/config/test hardening motivated only by missing deferred evidence; never auto-continue past the pending gate. **Load `references/pre-final-continuation-gate.md` before treating any deferral as terminal** — it owns the valid/invalid-anchor list and hardening boundary.
195
195
  - **Affirmative-assent binding rule** lives in `references/pre-final-continuation-gate.md` §Assent binding — **load it before selecting `continuing:` on any assent**, and any concrete next-slice proposal you issue must itself carry the `proposed-next:` marker or a later assent cannot bind — an unmarked referent is ambiguous, never self-cleared; the rule fires only when the immediately preceding assistant message itself states one concrete next action and its scope, and `continuing:` binds to that proposal, never to adjacent status or response-format prose; ambiguous assent, referent, or authority ⇒ `blocked:` with step-4 precedence — restate the proposed action/scope plus the specific ambiguity/authority, cite the step-1 evidence and ask one concise question in the same turn; self-classifying the reply or marker away is never an exit, and the `continuing:` default applies only when assent is unambiguous and no step-4 condition holds (the full rule and its fallback are stated there); the visible `continuing:`/`blocked:` outcome obligation is unchanged.
196
- 3. Continue automatically when all of these are true: the next slice comes from an explicit status/task/acceptance source or active user continuation instruction, is low-risk, local-only or already-authenticated, within the current accepted scope, has clear ownership, can be verified with existing commands, and does not require destructive actions, an external purchase/financial commitment, production access, legal/compliance/product strategy decisions, or high-impact architecture choices. Existing internal developer self-use through configured metered model/tool accounts is not an external purchase for this gate.
197
- 4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a failed, pending, or inconclusive required or blocking gate, dirty/conflicting worktree that cannot be isolated, unavailable required environment after remediation, high-impact product/architecture/compliance decision, destructive action, an external purchase/financial commitment, unclear owner, ambiguous assent, missing stricter authorization, or no low-risk implied next slice.
198
- 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome must use the required action/scope-plus-blocker form and classify the turn as `interim`. Ask one concise question in the same turn when ambiguity or missing authority is the blocker; an explicit stop/pause needs no reconfirmation. A `continuing:` outcome must proceed with the named slice before finalizing. A silent/completion-style stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync when a required review/challenge is pending or inconclusive; report the status as interim or blocked with the next unblock action.
199
- 6. **Assent-outcome closeout check.** Before every final response in any turn of a product-delivery session, walk the literal immediately preceding marker and visible outcome; the agent cannot exclude a status, question, review, or dispatched-owner turn by reclassifying it outside the session. An action marker or a plausibly affirmative user reply requires exactly one already-visible `continuing:`/`blocked:` outcome; until `continuing:` has executed the accepted slice or `blocked:` has named the blocker, the current turn may not use `proposed-next: none — status only` or `not-applicable`. A valid status-only marker permits `not-applicable` only when the preceding prose contains no imperative/future/next-step wording and no affirmative reply is pending. A missing/multiple/conflicting marker forces a visible `blocked:`/`interim` outcome with one clarifying question; it can never produce `not-applicable`. Finally, the current assistant message itself must end with exactly one action-form or status-only marker for the next turn. Omitting the marker cannot justify a silent stop.
196
+ 3. Continue automatically only when no step-4 stop condition fires and: the next slice comes from an explicit status/task/acceptance source or active user continuation, is low-risk, local-only/already-authenticated, in accepted scope, clearly owned, verifiable with existing commands, and needs no destructive action, external purchase/financial commitment, production access, legal/compliance/product-strategy decision, or high-impact architecture choice. Existing configured internal developer-self-use metered model/tool accounts aren't an external purchase here.
197
+ 4. Stop only for an explicit stop/pause instruction, a user-requested status-only answer, a failed/pending/inconclusive required/blocking gate, dirty/conflicting worktree that can't be isolated, required environment unavailable after remediation, high-impact product/architecture/compliance decision, destructive action, external purchase/financial commitment, unclear owner, ambiguous assent, missing stricter authorization, materially differing viable approaches (none dominant-and-reversible), a fix lacking evidenced cause, or no low-risk slice. Exactly one dominant reversible approach and no other stop condition firing: do not stop at a recommendation: deliver a tested reviewable draft.
198
+ 5. If stopping, state the concrete stop reason and the exact evidence checked; an assent-triggered `blocked:` outcome uses the action/scope-plus-blocker form and classifies the turn `interim`. Ask one concise in-turn question when ambiguity or missing authority blocks; explicit stop/pause needs no reconfirmation. A `continuing:` outcome proceeds with the named slice before finalizing. A silent/completion stop is invalid. Do not send a completion-only, solved, fixed, or fully-closed final response after a merge/sync while a required review/challenge is pending or inconclusive; report interim or blocked with the next unblock step.
199
+ 6. **Assent-outcome closeout check.** Before every final response in a product-delivery session, walk the literal immediately preceding marker and visible outcome; the agent cannot exclude a status, question, review, or dispatched-owner turn by reclassifying it outside the session. An action marker or a plausibly affirmative user reply requires exactly one already-visible `continuing:`/`blocked:` outcome; until `continuing:` has executed the accepted slice or `blocked:` has named the blocker, the current turn may not use `proposed-next: none — status only` or `not-applicable`. A valid status-only marker permits `not-applicable` only when the preceding prose has no imperative/future/next-step wording and no affirmative reply pending. A missing/multiple/conflicting marker forces a visible `blocked:`/`interim` outcome with one clarifying question; it never produces `not-applicable`. The current assistant message itself must end with exactly one action-form or status-only marker for the next turn. Omitting the marker cannot justify a silent stop.
200
200
 
201
201
  If a user later challenges "why did you stop" or "was the rule too weak", treat it as a product workflow defect: route through `skill-extraction-workflow`, strengthen the smallest owning skill or validation checklist, validate the diff, and only then claim the process issue is solved.
202
202
 
@@ -66,6 +66,8 @@ Use this for a general product R&D review pass before merge or release. Stack-sp
66
66
  - **Review both negative space and concept delta.** Review from the original acceptance source, the acceptance-to-evidence closure table, the concept-to-current-need table, and the final diff; the diff alone cannot show required behavior that is absent.
67
67
  - For negative space, confirm every in-scope acceptance point maps to an implementation surface and fresh evidence. Reject silent omissions, implementer-created downscoping, and tests derived only from the code that happens to exist.
68
68
  - For concept delta, require every new abstraction, indirection, module/service, state/entity, dependency, config/flag, generalized path, or extension point to identify its current acceptance point or observed hard constraint and the simpler alternative considered. Reject hypothetical future reuse; allow refactoring and testability work tied to a current observed constraint.
69
+ - Grade every stated guarantee at one of four levels — **strong / eventual / best-effort / unsupported** — and reject prose that presents a desired property as an implemented guarantee: "X is guaranteed" in a doc/comment/MR while the code delivers best-effort is a finding, not wording polish.
70
+ - A "no impact when disabled / flag off" claim is verified by walking the disable paths, not by trusting the flag: check imports/side-effectful module init, startup/registration hooks, API surface exposure, shared state/schema touched, UI/entry visibility, and the flag-off runtime path itself (new parsing/computation/IO executed before returning the old result) — any of the six can leak behavior or cost while "disabled".
69
71
  - Reconcile the active acceptance-source stable ID set against the closeout table and verify every out/deferred row cites a real product/human decision. Accept two-axis `not-applicable` only for a reviewed documentation-only diff. For a pure refactor or mechanical maintenance with no contract/behavior delta, accept functional-axis `not-applicable` only with the concept-delta table still present. The implementer's label alone is never evidence, and diffs touching the behavior-bearing surface classes enumerated in `implementation-completeness-and-minimality.md` (contracts, schemas/migrations, permissions, quotas/pricing, config defaults, user-visible copy — that list is canonical) qualify for no exemption. A rejected classification returns the delivery to the implementer to build the required tables; do not reconstruct them in review.
70
72
 
71
73
  ## Testing And Evidence
@@ -84,3 +86,5 @@ Use this for a general product R&D review pass before merge or release. Stack-sp
84
86
  ## Review Output
85
87
 
86
88
  Lead with findings. For each issue include severity, file/line, impact, and suggested correction. If there are no findings, state residual risk and any verification gaps.
89
+
90
+ Tag each load-bearing statement in the review by its evidence class — **Observed** (you observed the fact itself — ran the command, exercised the path, or read the code text — and the tag covers only what that observation proves: a runtime/concurrency/authorization property merely inferred from static code is Hypothesis until exercised; reading someone's assertion, including an implementer-authored CI or runtime claim, is Reported), **Reported** (a doc, comment, or another party asserts it), or **Hypothesis** (inferred, unverified) — and attach a suggested verification step to every non-Observed tag. An Observed tag carries its auditable receipt — the exact command or exercised path, scope/ref, and result — a bare "I exercised it" without the receipt downgrades to Reported. A review whose key claims are all Reported/Hypothesis is a reading summary, not review evidence.
@@ -172,7 +172,7 @@ Before release, confirm:
172
172
 
173
173
  ## Delivery Health Metrics(DORA software delivery metrics)
174
174
 
175
- 评估 engineering delivery health 时用 DORA 系列 metrics(起源于 Forsgren / Humble / Kim *Accelerate* 2018,**模型随年度 State of DevOps Report 演进**)作可量化共同词汇。当前 DORA 模型为 5 metric:
175
+ 评估 engineering delivery health 时用 DORA 系列 metrics(起源于 2014 年 State of DevOps 研究;Forsgren / Humble / Kim *Accelerate* 2018 成书普及,**模型随年度 State of DevOps Report 演进**——演进史见 dora.dev/insights/dora-metrics-history)作可量化共同词汇。当前 DORA 模型为 5 metric:
176
176
 
177
177
  | Metric | 定义 | 备注 |
178
178
  |---|---|---|
@@ -13,6 +13,7 @@ Use this checklist for R&D standards, specs, guidelines, or Feishu/wiki doc fami
13
13
  7. If the family will drive project health checks, add a conformance appendix or sibling checklist that maps each rule to `deterministic`, `agent_review`, `manual`, or `not_automatable_yet`, with severity, evidence, command/prompt owner, and CI behavior. Route architecture fitness functions to `testing-strategy`; route directory-contract coverage to `agents-file-coverage-gate`; route stack mechanics to the owning stack skills.
14
14
  8. Before invoking `tighten-doc`, record `owner-ready` with the authority statement, doc-family layer, enumeration evidence, and conformance mapping evidence when applicable, or `blocked: family enumeration unverified`.
15
15
  9. Re-confirm the doc-set enumeration at completion/sign-off, or add the sync gate if the family has grown.
16
+ 10. Layer authority and write-back: when an execution-layer doc (runbook, repo-local execution Spec, checklist) conflicts with its governing Spec/guideline, the Spec wins the ruling AND the execution doc is corrected in the same change; when the execution doc is outside the change's write scope (other repo/owner), route the correction as a trackable item the receiving owner has accepted, and keep the cross-repo sync explicitly outstanding — the ruling is landed only when the execution doc is corrected or the accepted item is in the owner's queue with the drift marked open; a note naming an owner is not routing. Superseded rules inside a living doc are marked deprecated **with a date** and a replacement pointer, not silently deleted, so a reader can tell "current" from "was once true" (append-only ledgers keep their own stricter rules).
16
17
 
17
18
  ## Architecture / System-Overview Honesty Pass
18
19
 
@@ -69,6 +69,8 @@ When starting work on tokens/theme/design-system:
69
69
  5. For each `A2` business file, confirm `styles_count == 0` and `components_count == 0` via Figma API. Non-zero = file-local forked tokens; route to the design-system maintainer to merge.
70
70
  6. Cross-check deprecation: every file the team treats as deprecated must appear in the private deprecation list. Files-only-marked-by-prefix get a follow-up to add them explicitly.
71
71
 
72
+ 7. Check the canonical source's completeness against a definition matrix before treating it as the acceptance baseline: for each module — color, background, typography, spacing, stroke/divider, shadow/elevation, gradient, opacity, radius, size/layout, motion (aligned with the governed-category list in `ui-ux-audit.md` §Diff-Scoped Review) — the source defines both the token **values** and each token's **role/usage semantics** (a value list without roles cannot answer "which token do I use here", which is what implementation and review actually ask). Record missing modules or missing role columns as design-system gaps routed to the maintainer; do not fill them ad hoc per feature.
73
+
72
74
  ## Anti-Patterns
73
75
 
74
76
  - **Mistaking a third-party mirror for a team design system**: leads to "we already have a complete design system" claims that ignore the missing brand customisation layer.