@ccoalm/ccl-skills 0.7.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +8 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +6 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +16 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +195 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +13 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/normalize_review_timeout.sh +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +9 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +1540 -129
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +8 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +76 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +1858 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +789 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/update_review_plan_intent.py +513 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +11 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/runtime-and-project-contract.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/verification-diagnostics-and-security.md +63 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +8 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +16 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +14 -16
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +10 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +135 -86
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +66 -80
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/delivery-contract.md +275 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +88 -214
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +10 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +6 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +112 -95
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +30 -21
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +22 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +20 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +7 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +14 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +9 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +37 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +8 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +16 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +16 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +4 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +4 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +8 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +104 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +11 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +102 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +103 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +6 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +4 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +69 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +49 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +2748 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +20 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py +1142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +41 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +82 -8
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +336 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +82 -10
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +1416 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +141 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh +1696 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2117 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +316 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +1176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +31 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +9 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +980 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +8 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +16 -15
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +5 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +10 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/annotation-driven-revision.md +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/figure-and-table-craft.md +8 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +7 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +37 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +10 -1
- package/dist/assets/release.json +215 -105
- package/package.json +1 -1
|
@@ -38,7 +38,7 @@
|
|
|
38
38
|
- **When evolving the stream protocol, answer identity/lifecycle questions before designing event fields.** "What fields does the new event type carry" and "which existing identity/idempotency/billing model do these events bind to" are two different questions, and the second must be answered first: which answer/revision object the events belong to, what key their usage is accounted under, and whether a retry hides a provider-side billing fork. Event designs that skip this get reworked at review or implementation time.
|
|
39
39
|
- **Answer-replacing stream events need explicit commit-point semantics — per plane, not one global commit.** For *answer visibility*, bind the revision's commit point to its first visible content delta: when cancel or degrade lands before commit, the previous answer must be preserved while the failure still surfaces to the caller, and caller-visible mutations (staged references/metadata) plus caller-side effects (tool dispatch) stay buffered until commit so a never-committed revision leaves no executed side effects to duplicate on replay. *Accounting is a separate plane*: provider-attempt usage/ledger records book per attempt under their own keys (per the usage rules below) regardless of whether the revision ever commits — an uncommitted revision still consumed provider tokens, and deferring or dropping its usage record is a billing hole, not tidiness. Post-commit failure semantics must also be an explicit spec decision, not an accident: when a revision commits (content already visible) and the stream later fails, either keep the partial replacement visible with the failure surfaced (progressive replacement) or roll back to the prior answer at terminal — name the choice and test it; leaving it implicit is how a partial revision silently clobbers a good answer. Define the semantics explicitly for streams with no content plane (tool-only, refusal, error-terminal): name which event commits them and how their outcome surfaces. This boundary is what chain-level replay tests force out — for stateful stream protocols, write chain-level (multi-event-sequence) tests at spec time, not after implementation.
|
|
40
40
|
- **OpenAI Responses API supersedes Chat Completions as the recommended surface for new integrations** per `platform.openai.com/docs/guides/migrate-to-responses` and `developers.openai.com/blog/responses-api`. Two load-bearing differences: (a) **reasoning state preservation across turns** — Responses keeps the model's reasoning context alive between API calls, whereas Chat Completions drops it; OpenAI's published evals show ~3% SWE-bench Verified improvement vs Chat Completions with identical prompts on reasoning models, and ~40-80% better cache utilization. (b) **agentic-by-default**: a single API request can call multiple built-in tools (`web_search`, `image_generation`, `file_search`, `code_interpreter`, custom functions) in one round-trip — the orchestration that previously required client-side tool-loop scaffolding moves provider-side. **CRITICAL P0**: provider-side built-in tools bypass the team's own auth / audit / rate-limit / data-egress boundaries — calls to `web_search` happen at OpenAI's edge with OpenAI's network, calls to `file_search` index data the team uploaded to OpenAI's storage, calls to `code_interpreter` execute code in OpenAI's sandbox. Required: explicitly DISABLE the built-in tools in the Responses request unless a specific tool is approved for the use case (default-on is the wrong stance for any team with compliance / data-residency / audit obligations); when enabled, require post-response reconciliation that emits one audit event per built-in tool invocation (which tool, with what arguments, returning what footprint), enforce per-tool rate limits the team owns rather than relying on provider defaults, and document the data-egress policy (what user content can the team's prompts contain when web_search / file_search is on). Migration mechanics: structured-output config moved from `response_format` to `text.format`; `reasoning_effort` default is model-specific and drifts across model versions (see `model-prompt-evaluation.md`) — do not hardcode a flat default; read the current model's documented default per route. **When to migrate**: new integrations on GPT-5.x reasoning models default to Responses; existing Chat Completions integrations stay until reasoning-state preservation OR multi-tool orchestration is load-bearing — do not migrate as drive-by during feature work. Cross-provider abstraction layers (LangChain, LlamaIndex, an in-house adapter) need to surface the API choice explicitly because SDK shapes differ; verify the adapter speaks both forms before flipping callers.
|
|
41
|
-
- **Prompt caching is a 2024-2025 industry-standard cost lever, not a niche optimization** — all three majors support it with different ergonomics. **Anthropic**:
|
|
41
|
+
- **Prompt caching is a 2024-2025 industry-standard cost lever, not a niche optimization** — all three majors support it with different ergonomics. **Anthropic**: `cache_control` with two usage modes per `docs.anthropic.com/en/docs/build-with-claude/prompt-caching` — a top-level automatic mode (one request-level `cache_control` field; the system places the breakpoint at the last cacheable block) and explicit per-block breakpoints for fine control; cache-read tokens are billed at ~10% of standard input price (NOT free — the 90% savings is the discount on those read tokens, not zero cost); cache-write is 1.25x base for 5-minute default TTL, 2x base for 1-hour TTL; the 1-hour TTL pays back vs uncached at roughly the third cache read once the write premium is amortized. Anthropic claims up to 90% cost savings on cache hits, but a workload that writes once and reads zero is a net loss. **OpenAI**: automatic caching on prompts ≥1024 tokens with no API change; Responses API improves cache utilization 40-80% vs Chat Completions per OpenAI's own announcement. **Google Gemini**: **implicit caching enabled by default for all Gemini 2.5+ models** per `developers.googleblog.com/en/gemini-2-5-models-now-support-implicit-caching/` — no API change needed; minimum cache-eligible request 1024 tokens (2.5 Flash) / 2048 tokens (2.5 Pro). **Architecture impact**: structure prompts so the cacheable prefix is stable (system prompt + tool/function defs + few-shot examples on top; per-request user content at the bottom). **Tenant-isolation P0**: cache keys / cacheable prefixes MUST be tenant-isolated — NEVER place tenant id, tenant-private context, tenant-specific tool definitions, or tenant-scoped few-shot examples inside the shared stable prefix unless the provider's account-isolation contract is verified AND tested (Anthropic and OpenAI account-scope caches per organization, but the moment two tenants share one API account / organization, a shared prefix that includes one tenant's context is a cross-tenant leak vector). Default position: tenant-scoped content goes BELOW the cache boundary; the stable prefix carries only tenant-neutral system prompts / tool defs / generic examples. Verify with a per-tenant cache-hit-rate audit — if tenant A's content hash ever produces a cache hit on tenant B's first request, the isolation is broken. **Footgun (Anthropic-specific)**: per Anthropic's docs, extended-thinking blocks interact with prompt caching in ways that invalidate cache entries more aggressively when the thinking state changes between turns; pin per-route extended-thinking decisions to avoid silent cost drift. For OpenAI / Google equivalents, this drop is plausible but not documented as a generalized rule at writing — measure cache-hit-rate per route before assuming the same dynamic. Treat cache-hit-rate as a per-route SLI emitted into the observability stack.
|
|
42
42
|
- **Keeping the prefix stable means making cost-only prefix churn cache-aware — and never deferring an authority/correctness change to save a cache write.** Any mid-conversation change to the cached prefix invalidates the cache from that point and pays the full uncached prefill on the next turn (plus, on providers with explicit cache-write pricing, a fresh cache-write charge), silently every turn after if it keeps mutating. Split such changes into two classes and handle them differently:
|
|
43
43
|
- **Authority-neutral, cost-only churn** (re-ordering stable prefix content, refreshing an unchanged or non-authority system-prompt section, adding authority-neutral formatting/few-shot guidance that grants no new capability) — here the recommended pattern is **deferred invalidation**: don't rebuild the prefix as a side effect of an unrelated command; instead record the change as a durable `pending` mutation bound to principal/session/config-generation, apply it **before the next model request** (not at some far-off boundary, and never silently dropped — if applying it fails, fail closed to a degraded/visible state rather than continuing on stale-but-uncommitted intent), and offer an explicit opt-in (e.g. a `--now` / "apply immediately" path) to rebuild the prefix this turn.
|
|
44
44
|
- **Authority/correctness/capability-surface changes** — ANY change to the callable/visible tool or capability surface (add, remove, revoke, schema/trust/exposure change), plus policy/authorization changes, memory deletion, model/behavior changes, and any privacy/principal/tenant/auth change — must invalidate / fence / re-authorize **NOW, never deferred for cost**. These follow the fail-closed snapshot/re-verify/re-authorize rules above and in `references/retrieval-agent-safety.md` / `references/agent-tool-dispatch.md`, which take precedence over cost-driven deferral. Deferring a revocation that leaves a revoked tool callable for the rest of a long turn — or deferring a newly-added tool so the model plans around a surface that isn't really there — is a correctness/security bug, not a saving. The only other routinely-acceptable mid-conversation prefix change is the deliberate context compaction already governed above (which invalidates and re-writes by design).
|
|
@@ -15,7 +15,7 @@ Avoid hard-coded model strings scattered through product code. Business logic sh
|
|
|
15
15
|
## Model Version Baseline 2025-2026
|
|
16
16
|
|
|
17
17
|
- **Three major-provider model families are the credible production options at 2026-Q2**; the registry/router MUST verify exact current model strings against the live docs page before pinning, because provider naming cadence has accelerated (Anthropic ships ~quarterly minor revs, OpenAI ships sub-version reasoning-effort variants, Google ships Pro/Flash/Flash-Lite tiers separately). Authoritative-doc sources for live verification: `docs.anthropic.com/en/docs/about-claude/models` (Anthropic) / `developers.openai.com/api/docs/changelog` + OpenAI Help Center model release notes / `ai.google.dev/gemini-api/docs/changelog` (Google). General shape: Anthropic ships Opus (frontier reasoning) + Sonnet (workhorse, 200K default + 1M beta context per `anthropic.com/news/1m-context`) + Haiku (cost/speed); OpenAI ships GPT-5.x with a `reasoning_effort` parameter (default value and available effort tiers are model-version-specific and drift — do not hardcode; verify per target model per the next bullet) plus reasoning summaries; Google ships Gemini 2.5 Pro + Flash + Flash-Lite with explicit thinking-budget control. Per-route choice is a registry decision recorded in the project's model registry, not hardcoded. **Pinning + eval-baseline discipline (P0)**: pin the exact model version string in the registry config (NOT just the family name like `claude-sonnet`); block ENV-var-only flips between model versions in production (flipping `MODEL=...` to a new minor rev silently invalidates every prompt eval result against the old version); require evaluation replay or shadow comparison for ANY model-string or default-param change before flipping production traffic; emit a metric on per-route model-string drift so an unintended autoupgrade is observable, not just discoverable at the next eval. The recurring failure is: team upgrades minor rev for cost/quality reason, eval suite still passes BECAUSE eval prompts didn't exercise the regressed code path, production sees the regression a week later.
|
|
18
|
-
- **Extended thinking / reasoning effort is a per-request product decision, not an account-level default**. Anthropic's extended thinking
|
|
18
|
+
- **Extended thinking / reasoning effort is a per-request product decision, not an account-level default**. Anthropic's thinking model is **generation-dependent — verify against the live docs for the exact target model**: on Claude 4.5-and-earlier thinking models, extended thinking is off by default and enabled per request (`thinking: {type:"enabled", budget_tokens}`); that manual mode is deprecated on 4.6 and rejected with a 400 on 4.7+; on the current generation (Claude 5 family) thinking is on by default (adaptive) with `effort` controlling depth (`platform.claude.com/docs/en/build-with-claude/thinking`). Thinking trades latency + token cost for accuracy on multi-step reasoning; it also impacts prompt-caching efficiency (cache invalidation more aggressive when thinking state changes). OpenAI's `reasoning_effort` tunes the same axis per `developers.openai.com/api/docs/guides/reasoning`, but the **default differs by model version and is not monotonic** — defaults have varied across GPT-5.x revs (different revs ship different defaults; do not assume a trend) and new effort tiers (e.g. `xhigh`) appear in some revs. Do NOT hardcode a version-specific default here; confirm the default for the exact model you target against its live docs page before relying on "default behavior" — drift in this default is the most common silent cost/latency change across an OpenAI minor-version bump. `revalidate-when: OpenAI ships a new GPT-5.x reasoning rev`. Gemini exposes a thinking budget knob per `developers.googleblog.com/en/gemini-2-5-thinking-model-updates/`. **Pattern**: per-route declare reasoning intent (off / low / medium / high) based on the task's complexity AND record the decision in the prompt/route config so latency and cost regressions are traceable to an explicit knob, not provider-default drift. **Production-safety contract for reasoning enablement (P0)**: declaring reasoning intent is not enough — the per-route config MUST also carry (a) p95/p99 **latency budget** with downstream timeout propagation (reasoning can add seconds-to-tens-of-seconds to single-turn latency; downstream HTTP timeouts that worked at non-reasoning latency will fail), (b) **reasoning-token cap** to prevent runaway thinking burning 10-50× the non-reasoning token budget on a hard prompt, (c) **cost guardrail** as a per-request token + per-user-session aggregate ceiling, (d) **rollback path** at the route level under incident — on request-gated generations, flip thinking off; on the current always-on generation, drop effort to the minimum tier or route to a non-thinking model, and record which lever the route uses. Mid-route enablement of reasoning (turning thinking on after launch) is a load-bearing change that requires fresh load/eval evidence and SRE sign-off, NOT a config-only flip. Reasoning content is a separate channel from user-visible content — never persist reasoning into product history or display it as final output (see streaming rules in `llm-client-gateway.md`).
|
|
19
19
|
|
|
20
20
|
## Prompt Registry
|
|
21
21
|
|
|
@@ -80,7 +80,10 @@ Every material model or prompt change should define:
|
|
|
80
80
|
- for compared runs, hold the conditions fixed — everything except the declared variable under test, with both values of that variable recorded (model string, prompt version, retrieval index, tool set, and decoding parameters are the usual fixed set);
|
|
81
81
|
- quality metrics such as accuracy, recall, consistency, parse success, or groundedness;
|
|
82
82
|
- runtime metrics such as latency, token usage, success rate, fallback rate, and cost;
|
|
83
|
+
- cross-provider token/usage comparison only after normalization — providers account cache/reasoning tokens differently (some list cache reads beside input, some fold them into prompt tokens), so raw `totalTokens` is comparable only between runs of the same case input after the runner normalizes accounting; per-case ratio aggregates (cache-hit rate, cost/case) include only cases with usage on all compared arms, with the excluded-case count reported next to every ratio; total-cost accounting is the separate aggregate that never excludes — a billed attempt missing usage is a data defect counted against that provider and its cost enters via billing records, never silent exclusion from totals — so omission cannot flatter an arm on either aggregate;
|
|
84
|
+
- where billing records are available, reconcile reported usage against them for at least a sample — usage self-reported by the pipeline is one layer weaker than the invoice;
|
|
83
85
|
- regression examples and human review notes when judgment is subjective.
|
|
86
|
+
- keep human-review fields (reviewer, rating, notes) and machine-produced fields in separate columns/records with distinct writers: an automated re-run must never overwrite a human verdict, and a retry archives a new attempt record instead of replacing the failed one.
|
|
84
87
|
|
|
85
88
|
Store evaluation reports with the model/prompt versions compared so future changes can reproduce the decision.
|
|
86
89
|
|
|
@@ -23,9 +23,9 @@ For evaluating whether Taro is the right choice for a given project (vs native,
|
|
|
23
23
|
|
|
24
24
|
## Maturity Baseline
|
|
25
25
|
|
|
26
|
-
The current
|
|
26
|
+
The current baseline is **vendor-spec + framework-canonical**, not `mature confirmed`: host-platform guidelines, Taro documentation/examples, and canonical Taro UI libraries (taroify, NutUI-Taro, tdesign React mapping). No production-quality miniapp portfolio has been observed end-to-end. Apply positive rules as defaults and anti-patterns as guardrails, and mark them `confirmed` only after a correction-free real feature delivery.
|
|
27
27
|
|
|
28
|
-
**
|
|
28
|
+
**Existing team codebases** do not automatically supply positive rules: production use proves distribution, not quality. Until audited end-to-end against this skill or piloted through one real feature with a `skill-extraction-workflow` retrospective, label them `quality-unverified` per `references/source-evidence-map.md`. Use this skill from the start and feed lessons back through extraction, not silently copy patterns.
|
|
29
29
|
|
|
30
30
|
## Runtime Compatibility
|
|
31
31
|
|
|
@@ -101,7 +101,7 @@ Cross-checking rule: when editing code shared with a React web project, also che
|
|
|
101
101
|
|
|
102
102
|
## Core Workflow
|
|
103
103
|
|
|
104
|
-
Before editing Taro/native mini-program code, page config, host capability adapters, platform project files, styles, assets, or tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a simple low-risk single-page change can use a short inline plan; multi-target,
|
|
104
|
+
Before editing Taro/native mini-program code, page config, host capability adapters, platform project files, styles, assets, or tests, complete enough analysis and planning for the change to be reviewable. Scale the plan to risk: a simple low-risk single-page change can use a short inline plan; multi-target, API-visible, host-capability, platform-review/release, bug-fix, branch/MR, unclear-risk, or high-risk work needs explicit task split, host/target verification matrix, acceptance checks, verification commands, rollback or stop conditions, and named handoffs to testing, web/app, backend, release, or diagnosis skills before edits. Runtime-visible work consumes the canonical Design brief and Test Phase 0 before implementation.
|
|
105
105
|
|
|
106
106
|
1. Resolve the miniapp platform and delivery shape.
|
|
107
107
|
- Host platform target(s): WeChat, Alipay, Douyin/TikTok, Baidu, or several at once. Multi-target = a separate verification matrix; one target compiling is not proof another target passes.
|
|
@@ -116,8 +116,8 @@ Before editing Taro/native mini-program code, page config, host capability adapt
|
|
|
116
116
|
- For native miniapp: WeChat uses `app.json` + `project.config.json` + `sitemap.json` + `ext.json` (plugin/extension); Alipay uses `app.json` + `mini.project.json`; Douyin/Baidu use their own platform project files. `manifest.json` + `pages.json` is uni-app shape, not native; only include it when the repo is uni-app. Also inspect package config, build scripts, and CI jobs as applicable.
|
|
117
117
|
- Identify platform-branching code paths: `process.env.TARO_ENV` checks in Taro, conditional compilation blocks, or platform-specific files (`*.weapp.tsx`, `*.alipay.tsx`). Confirm branching lives at the adapter/wrapper layer, not in render code.
|
|
118
118
|
- Identify whether the change must be shared, forked, or guarded by capability detection.
|
|
119
|
-
- For visible
|
|
120
|
-
-
|
|
119
|
+
- For every visible UI change, load `../product-ui-ux-design/references/delivery-contract.md` and consume either its full Design brief + Phase 0 or its valid low-risk copy-only record + lightweight Phase 0 before coding. The lightweight path checks semantics, accessible name, localization, rendered extent, and target-host render without inventing unrelated matrices; risk-bearing copy uses the full path. For full slices, map structure, state/adaptation matrices, behavior and criteria to pages/host adapters; record route/back/share entry, hosts, capabilities, recovery geometry, and preserved behavior. A mini-program `web-view` has a host member here and a separate web-content owner member; browser/H5-only preview satisfies neither the shipped-host bridge nor the complete owner set.
|
|
120
|
+
- Before the first implementation edit, add the canonical `client_entry` defined there: local rule identifier or short quote and implementation decision, target surface/runtime, planned run/capture command, and behavior that must remain unchanged.
|
|
121
121
|
|
|
122
122
|
3. Define the miniapp contract before coding.
|
|
123
123
|
- Pages, route params, tab ownership, back behavior, deep links, scene/query entry, and share/open-from-chat behavior. Treat every scene/share/QR param as untrusted input: schema-parse it, server-authorize the referenced target against the current identity, and require backend-issued, TTL-bounded, replay-protected share tokens for attribution or unlock flows. Client-side attribution is never the final source of truth.
|
|
@@ -155,15 +155,16 @@ Before editing Taro/native mini-program code, page config, host capability adapt
|
|
|
155
155
|
4. 静态资源:图片/字体被 wxml/wxss 引用 → `grep -rn "<filename>" src/`;分包资源 → 查每个分包 `pages` 列表
|
|
156
156
|
5. 平台条件编译(Taro 多端):`#ifdef WEAPP / ALIPAY` 内的 import 在另一端不存在;按目标平台跑 build 看 warning
|
|
157
157
|
- For visible changes, inspect the rendered page in the relevant developer tool (WeChat DevTools, Alipay IDE, Douyin DevTools, Baidu DevTools), simulator, preview build, or real device and capture evidence where feasible.
|
|
158
|
-
- For systemic UI/UX redesign slices, diff the declared host target list against repo-configured build targets; every configured target must be classified as shipped (needs rendered evidence), product-level permanently excluded (can complete) — valid only when the authoritative build/release target source already stopped shipping that target before this slice; removing or disabling a target within the slice is a separate product/release scope change that routes through its product/risk/release owners and cannot satisfy this gate's evidence for the same slice; explanatory docs or an MR comment alone are temporary-skip authority, never permanent exclusion — or temporary slice-skip for a shipped target (leaves that host `pre-runtime-test
|
|
158
|
+
- For systemic UI/UX redesign slices, diff the declared host target list against repo-configured build targets; every configured target must be classified as shipped (needs rendered evidence), product-level permanently excluded (can complete) — valid only when the authoritative build/release target source already stopped shipping that target before this slice; removing or disabling a target within the slice is a separate product/release scope change that routes through its product/risk/release owners and cannot satisfy this gate's evidence for the same slice; explanatory docs or an MR comment alone are temporary-skip authority, never permanent exclusion — or temporary slice-skip for a shipped target (leaves that host `pre-runtime-test-ready` / `blocked`), and any unclassified target blocks completion.
|
|
159
159
|
- For UI/UX redesign evidence, include declared host targets, host developer-tool or real-device channel, loading/empty/error/final states, long text or text-scale behavior where supported, permission/capability prompts, route/share/scene entry when relevant, and screenshot or equivalent host-rendered artifact. Mark each dimension covered or `N/A` with a one-line reason; `N/A` is valid only when the reason names a verifiable structural fact, explains why that fact makes the dimension unreachable or unchanged for this slice, and includes a checkable pointer such as a file path, config key, or commit that resolves at review time. Persist evidence artifacts where reviewers can access them using sanitized/test accounts and redacting tokens, PII, credentials, private paths, and raw personal data; remove temporary smoke files or generated preview helpers before commit unless the repo intentionally owns them.
|
|
160
|
-
-
|
|
160
|
+
- Return the complete canonical client-record member defined in `../product-ui-ux-design/references/delivery-contract.md` for testing Phase 1 and the design verdict. The member includes its applied rule/decision, affected files/components, preserved behavior, exact command, immutable candidate binding, producer member/version actually exercised, artifacts, tested host/tool/device targets, states/dimensions/capabilities, criterion-mapped observations, coverage boundary, and gaps. A host render proves only the captured host/member states; it cannot close an unbound producer member. `testing-strategy` records aggregate sufficiency before the design owner records the candidate-bound verdict.
|
|
161
|
+
- For mini-program runtime changes, developer-tool or real-device smoke is a completion gate, not optional evidence. This includes changes to `Taro.*` or host APIs, `wx.*`/`my.*` calls, chunked/streaming transport, foreground/background recovery, route/share/scene behavior, storage/session restore, permissions/capabilities, and host-rendered loading/error/final states. If the tool or device is missing, first attempt discovery and normal setup; if still unavailable, stop at `pre-runtime-test-ready` or `blocked` and name the owner, attempted commands, residual risk, and next unblock action. `pre-runtime-test-ready` is a handoff-only label; it is not merge-ready, release-ready, or complete.
|
|
161
162
|
- If an automation, remote-control, or screenshot channel reports a blank or stale mini-program surface while a human operator can see the real host page rendering, treat it as an observation-channel conflict before treating it as an app defect. Re-check focus/window/permission state, capture the human-visible state through another channel when possible, label which evidence came from the automation channel versus the human-visible host, and only mark "blank screen" as a product defect after at least one host-visible channel reproduces it.
|
|
162
163
|
- When a human operator's already-authenticated host client or device is used as the runtime test surface, treat it as a human-assisted host test: record observer role/source class, sanitized account class, host client/device, entry path, actions performed, redacted artifacts, state changes such as login/logout or permission prompts, and restoration outcome in the project evidence. Label it as manual, scenario-scoped evidence; it does not replace required automated assertions or other host checks. If requested logout/account-switch/storage/permission restoration is not confirmed, mark the host test blocked or incomplete until restored or handed off to a named owner. Do not record personal phone numbers, personal operator names, chat/contact handles, tokens, private account names, or reviewer credentials in shared artifacts.
|
|
163
164
|
- Treat appid, dev-tool login, plugin authorization, service-port availability, and host identity/configuration as part of the runtime verification surface, not as background noise. A generated preview QR or a backend login success does not prove the mini-program runtime path until the correct host/app identity and permissions are exercised in the tool or on device.
|
|
164
165
|
- Verify route entry, share/deep-link scene params, auth state, storage restore, network error, permission denial, and primary recovery path for affected flows.
|
|
165
|
-
- For detail pages and deep-link/share/QR entry points, test missing or stale route params, missing local storage/cache payloads, expired auth, and direct cold entry. These states must resolve to an explicit empty/error/recovery state or safe redirect; a permanent loading spinner or blank screen under cold entry or missing-param entry is a blocking defect that must be fixed or explicitly marked `blocked` with a real owner, resolution path, and target follow-up point before the flow can be called complete. For stale storage/cache, include at least one real-device or emulator state test with a prior-version or manually seeded cache payload. Developer-tool-only stale-cache evidence is fallback evidence and must be labeled as a `device-state gap`; a flow with an open `device-state gap` is `blocked` or `pre-runtime-test
|
|
166
|
-
- For payment, subscription, login, phone, camera/media, or write-finality changes, verify sandbox/mock plus one platform-specific happy path. If platform evidence is unavailable after remediation, stop at `pre-runtime-test
|
|
166
|
+
- For detail pages and deep-link/share/QR entry points, test missing or stale route params, missing local storage/cache payloads, expired auth, and direct cold entry. These states must resolve to an explicit empty/error/recovery state or safe redirect; a permanent loading spinner or blank screen under cold entry or missing-param entry is a blocking defect that must be fixed or explicitly marked `blocked` with a real owner, resolution path, and target follow-up point before the flow can be called complete. For stale storage/cache, include at least one real-device or emulator state test with a prior-version or manually seeded cache payload. Developer-tool-only stale-cache evidence is fallback evidence and must be labeled as a `device-state gap`; a flow with an open `device-state gap` is `blocked` or `pre-runtime-test-ready`, not complete.
|
|
167
|
+
- For payment, subscription, login, phone, camera/media, or write-finality changes, verify sandbox/mock plus one platform-specific happy path. If platform evidence is unavailable after remediation, stop at `pre-runtime-test-ready` or `blocked`; do not complete the work by only recording the gap.
|
|
167
168
|
- For release work, verify app id/env, version, build output, platform review checklist, gray release/rollback path, analytics version tag, and owner handoff. For every **shipped** host platform, compile + developer-tool/real-device evidence is blocking — "recorded as unverified" is only acceptable for targets the release is not actually shipping. Mini-program rollback through host-platform re-review is slow; risky flows must therefore have a **server-side feature flag with safe default + tested kill-switch runbook** in place before submission. Capture the current official platform-policy doc URL + date for every review-sensitive area touched (payment, privacy, AI/generated content, minors, financial/medical/legal copy) — policy text changes faster than skill rules.
|
|
168
169
|
|
|
169
170
|
## Non-Negotiable Rules
|
|
@@ -184,7 +185,7 @@ Before editing Taro/native mini-program code, page config, host capability adapt
|
|
|
184
185
|
- Do not scatter backend enum/string literals through mini-program pages, scene/share/QR parsing, host bridge payload handling, storage, analytics, or tests. Centralize finite-value parsing, display labels, defaults, and unknown-value behavior at the API/client-domain boundary, and keep raw literals only in clearly named boundary conversion tests that cover every known external value plus unknown/default behavior. Migrate existing non-boundary test raw literals for that value in the same pull request or mark each remaining use with `finite-value-debt: <task-ref> <owner> <deadline> <reason>`, even when the current slice does not introduce a new mapper.
|
|
185
186
|
- Do not ship auth, payment, phone, location, camera, share, subscription, or generated-content flows without explicit denial/error/retry states.
|
|
186
187
|
- Do not ship a host platform without compile + developer-tool/real-device evidence for that platform — "recorded as unverified" is only acceptable for non-shipped targets.
|
|
187
|
-
- Do not claim a mini-program runtime fix is complete when developer-tool or real-device smoke did not run. Build output, unit tests, source-regex checks, and independent code review can make the branch `pre-runtime-test
|
|
188
|
+
- Do not claim a mini-program runtime fix is complete when developer-tool or real-device smoke did not run. Build output, unit tests, source-regex checks, and independent code review can make the branch `pre-runtime-test-ready`; they cannot make host-runtime behavior complete.
|
|
188
189
|
- Do not submit a risky flow for platform review without a server-side feature flag (safe default + kill-switch runbook); platform-review rollback is too slow to be the only lever. The flag does not stop client-only effects (permission prompts triggered at startup, SDK auto-collection on load, host-platform config already submitted) — the flow's **client-side code path itself must no-op when the flag is off**, the SDK must not load until the flag is on, and any host-config change submitted at review time must be reviewed for "what if we need to disable this without a new submission" before approval.
|
|
189
190
|
- Do not ship review-sensitive surfaces (payment, privacy disclosure, AI/generated content, minors, financial/medical/legal copy, account deletion, SDK data collection) without naming the current platform-policy doc URL + date you read.
|
|
190
191
|
- Do not claim platform review or real-device readiness without current evidence.
|
|
@@ -60,3 +60,8 @@ Use this checklist for payment, quota, order, publishing, generated content, acc
|
|
|
60
60
|
- Callback/reconciliation path for payment or async work, with order/request id matched on both sides.
|
|
61
61
|
- Support-visible order/request id surfaced to the user; without it, support cannot disambiguate "I paid but app says canceled".
|
|
62
62
|
- Safe retry and user explanation for uncertain final state — never present "succeeded" or "failed" when reconciliation has not run.
|
|
63
|
+
|
|
64
|
+
## 组件库版本纪律
|
|
65
|
+
|
|
66
|
+
- 组件库 API(taroify / NutUI-Taro / tdesign 等)按 lockfile 实装版本写,勿凭训练记忆——prop 名与默认值跨大版本会变。
|
|
67
|
+
- 库自带 lint/codemod 或迁移清单时,变更与大版本迁移以它对 changed files 收口(deprecated 用法与 a11y 规则通用 lint 不认识)。
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: nodejs-service-dev
|
|
3
|
+
description: Use when implementing, modifying, scaffolding, or testing Node.js backend and service code, including HTTP/RPC handlers, workers, jobs, standalone Node.js CLI/tooling, runtime configuration, TypeScript/JavaScript module setup, async cancellation, streams, graceful shutdown, and Node-specific test mechanics. Triggers include "用 Node.js 写接口", "Node 后端实现", "用 Node.js 写个命令行工具", "重构 Node 服务里的某文件/某类(局部)", "refactor a file/class within a Node.js service", "Fastify/Express/NestJS 服务", "node:test 怎么写", and "event loop / worker_threads 怎么改". Route active failures to defect-diagnosis, test-layer choices to testing-strategy, terminal contracts to terminal-cli-dev, and cross-module architecture or multi-stage delivery to product-rd-workflow.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Node.js Service Development
|
|
7
|
+
|
|
8
|
+
## Skill Routing
|
|
9
|
+
|
|
10
|
+
- Use this skill for Node.js service implementation: handlers, middleware, adapters, workers, jobs, clients, runtime/toolchain mechanics, and focused tests.
|
|
11
|
+
- New capabilities, cross-module architecture, service redesigns, and multi-stage refactors enter `product-rd-workflow`; verified repairs return here after `defect-diagnosis`.
|
|
12
|
+
- `testing-strategy` chooses test layers, coverage policy, contract/E2E scope, and CI gates. This skill owns Node runner, mock, fixture, and command mechanics after that choice.
|
|
13
|
+
- Standalone CLI/tooling stays here; language-stack CLI implementation must not be routed back to `terminal-cli-dev`, owner of the interface contract; logs/metrics/traces to `platform-observability`; cross-service timeout/retry/mTLS to `platform-service-connectivity`; rollout/rollback to `platform-release-engineering`.
|
|
14
|
+
- Browser UI goes to `web-react-dev`; LLM/RAG/agent-runtime behavior to `llm-inference-integration`. Keep language-neutral rules in their existing owner.
|
|
15
|
+
|
|
16
|
+
## Workflow
|
|
17
|
+
|
|
18
|
+
### 1. Recover the contract
|
|
19
|
+
|
|
20
|
+
Read the nearest `AGENTS.md`, then inspect `package.json`, lockfile/package-manager metadata, runtime-version files, `tsconfig*.json`/`jsconfig.json`, build/test scripts, deployment manifests, and the smallest relevant source path. Record:
|
|
21
|
+
|
|
22
|
+
- supported and deployed Node.js line;
|
|
23
|
+
- package manager and authoritative lockfile;
|
|
24
|
+
- ESM/CommonJS and JS/TS execution plus type-check path;
|
|
25
|
+
- framework lifecycle, verification commands, changed external contract, and non-goals.
|
|
26
|
+
|
|
27
|
+
Preserve those choices unless migration is explicit. Read [runtime-and-project-contract.md](references/runtime-and-project-contract.md) when changing one.
|
|
28
|
+
|
|
29
|
+
### 2. Define and implement the narrow boundary
|
|
30
|
+
|
|
31
|
+
State input, output, error, cancellation, timeout, idempotency, and ownership before code. Preserve established contracts unless explicitly changed; validate untrusted data at the owning boundary.
|
|
32
|
+
|
|
33
|
+
- Keep event-loop callbacks and worker-pool tasks short; bound fan-out, queues, retries, payloads, and buffering.
|
|
34
|
+
- Propagate cancellation/deadlines to underlying work. A wrapper timeout that leaves work running is not cancellation.
|
|
35
|
+
- Use streams with backpressure for large/unbounded data. Use a bounded `worker_threads` pool only for measured CPU-intensive JavaScript, not ordinary async I/O.
|
|
36
|
+
- Preserve error causes and map once at the boundary. Do not swallow rejections or resume normal operation after an unknown fatal process error.
|
|
37
|
+
|
|
38
|
+
Read [async-lifecycle-and-performance.md](references/async-lifecycle-and-performance.md) when touching concurrency, streams, CPU work, shutdown, or performance.
|
|
39
|
+
|
|
40
|
+
### 3. Make lifecycle and exposure explicit
|
|
41
|
+
|
|
42
|
+
Validate configuration before traffic and never log secrets. On shutdown, stop intake, drain bounded work, abort owned background work, close resources, and honor one documented deadline; handle upgraded connections separately. Treat readiness and liveness as different contracts.
|
|
43
|
+
|
|
44
|
+
Keep dependency changes minimal, update the lockfile, use frozen/immutable install verification, and review scripts/transitive impact. Bound input work and treat the Node.js Permission Model as optional defense in depth, never a complete sandbox. Read [verification-diagnostics-and-security.md](references/verification-diagnostics-and-security.md) for concrete test, diagnostic, dependency, and security checks.
|
|
45
|
+
|
|
46
|
+
### 4. Verify narrow to broad
|
|
47
|
+
|
|
48
|
+
Use repository commands: focused behavior and failure-path test → touched-package lint/type/check → package/service suite → build/package/start check → required repo gates. Exercise cancellation, malformed input, cleanup, and shutdown when relevant.
|
|
49
|
+
|
|
50
|
+
Performance/reliability claims require a representative workload plus outcome and causal evidence. Re-run the reproducer; inspection alone cannot prove a leak, stall, or regression fixed.
|
|
51
|
+
|
|
52
|
+
## Hard Rules
|
|
53
|
+
|
|
54
|
+
- Preserve module convention and make new package intent explicit; do not rely on ambiguous `.js` syntax detection.
|
|
55
|
+
- Built-in TypeScript stripping is execution support, not type checking or general transpilation. Keep a real type-check gate and verify unsupported syntax/`tsconfig` dependencies.
|
|
56
|
+
- Prefer explicit dependencies and startup wiring over mutable process-wide singletons so tests need not bind ports or mutate globals.
|
|
57
|
+
- Avoid unbounded `Promise.all`, ownerless fire-and-forget promises, synchronous hot-path APIs, per-request workers/processes, and whole-stream buffering by default.
|
|
58
|
+
- Do not add blanket retries, speculative caches, generic base layers, or process-level exception recovery without observed need and an owning contract.
|
|
59
|
+
|
|
60
|
+
## Output Contract
|
|
61
|
+
|
|
62
|
+
Report the runtime/package/module contract, changed behavior, preserved boundaries, exact checks, measurements, skips, version assumptions, and risks.
|
|
63
|
+
|
|
64
|
+
Technical claims and comparison limits are recorded in [source-map.md](references/source-map.md); ordinary implementation should load only the task reference whose decision surface is reached.
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Node.js Service Dev"
|
|
3
|
+
short_description: "Build reliable Node.js backend and service features"
|
|
4
|
+
default_prompt: "Use $nodejs-service-dev to implement a Node.js backend or service change from the repository's runtime, module, lifecycle, and verification contracts."
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Async, Lifecycle, and Performance
|
|
2
|
+
|
|
3
|
+
Use this reference when a change touches request concurrency, cancellation, streams, CPU work, background jobs, shutdown, or a performance claim.
|
|
4
|
+
|
|
5
|
+
## Classify the work first
|
|
6
|
+
|
|
7
|
+
| Work | Default execution | Main risk |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| Network and async file/database I/O | Native async API on the event loop | unbounded concurrency, missing timeout/cancellation, retained buffers |
|
|
10
|
+
| Short JavaScript transformation | Event-loop callback | input-dependent long task, excessive allocation |
|
|
11
|
+
| CPU-intensive JavaScript | Bounded `worker_threads` pool or external worker | worker creation/serialization overhead, queue growth, memory sharing |
|
|
12
|
+
| Blocking/native task using libuv pool | Async API, with measured pool pressure | worker-pool starvation across unrelated requests |
|
|
13
|
+
| Large or unbounded payload | Stream/pipeline with backpressure | full buffering, missing cleanup, partial output |
|
|
14
|
+
|
|
15
|
+
Node.js uses a small number of threads to serve many clients. A long callback reduces event-loop throughput; a long libuv task can starve the worker pool. Both can become denial-of-service paths when complexity or input size is attacker-controlled.
|
|
16
|
+
|
|
17
|
+
## Cancellation and deadlines
|
|
18
|
+
|
|
19
|
+
- Accept an owning cancellation/deadline signal at service boundaries and pass it through every supported downstream API.
|
|
20
|
+
- Prefer an existing `AbortSignal`; combine caller cancellation and timeout without losing the original reason when the supported runtime provides the needed API.
|
|
21
|
+
- A raced timeout that rejects while the database call, fetch, stream, worker, or child process continues is not cancellation. Close/destroy/abort the underlying resource or document why it cannot be stopped and bound the orphaned work.
|
|
22
|
+
- Remove listeners and timers during cleanup. Use `unref()` only when it matches lifecycle ownership; it is not a substitute for cancelling work.
|
|
23
|
+
- Retries must fit inside one overall deadline, use the connectivity owner's policy, and remain bounded. Never retry non-idempotent effects without an idempotency contract.
|
|
24
|
+
|
|
25
|
+
## Bounded concurrency
|
|
26
|
+
|
|
27
|
+
- Replace unbounded `Promise.all(items.map(...))` on variable-size input with a repository-standard limiter, queue, or batch window.
|
|
28
|
+
- Bound queue length as well as worker count. Define overload behavior: reject, shed, defer durably, or backpressure the producer.
|
|
29
|
+
- Track in-flight ownership so shutdown can await or abort it. A detached promise must have an explicit supervisor and error sink.
|
|
30
|
+
- Avoid per-request child processes or workers. If CPU offload is justified, measure task duration and transfer cost, then reuse a bounded pool.
|
|
31
|
+
|
|
32
|
+
## Streams and backpressure
|
|
33
|
+
|
|
34
|
+
- Prefer `node:stream/promises` `pipeline()` or an established equivalent so errors and teardown propagate across the chain.
|
|
35
|
+
- Respect `write()` backpressure/drain semantics and configure object/buffer high-water marks from measurement, not folklore.
|
|
36
|
+
- Set payload/record limits even when streaming. Streaming bounds memory growth; it does not bound total work.
|
|
37
|
+
- Propagate abort signals and verify cleanup on source error, transform error, destination error, client disconnect, and timeout.
|
|
38
|
+
- Do not mix flowing-mode event handlers and async iteration on the same readable unless the lifecycle is deliberately controlled.
|
|
39
|
+
|
|
40
|
+
## Errors and process lifecycle
|
|
41
|
+
|
|
42
|
+
- Catch errors at boundaries that can make a valid decision: translate, retry under policy, compensate, or fail the operation. Otherwise preserve `cause` and propagate.
|
|
43
|
+
- For durable/task state written by one operation, capture the clock once per transition (all timestamps the transition itself stamps come from a single `now`; domain-provided times are recorded as received, never re-stamped), and validate external-response structure (schema parse with a typed error path) before mapping — a malformed upstream payload must become a persisted failure transition carrying the canonical error (mirroring the go/python state-machine rendering), never a crash or a silently-defaulted field.
|
|
44
|
+
- Treat unknown `uncaughtException` and default-throw unhandled rejection paths as fatal. A handler is for synchronous cleanup/diagnostics before termination, not resuming normal operation from an undefined state.
|
|
45
|
+
- On `SIGTERM`/the platform's shutdown signal:
|
|
46
|
+
1. mark readiness false or otherwise stop new routing;
|
|
47
|
+
2. stop accepting new work;
|
|
48
|
+
3. drain bounded in-flight work;
|
|
49
|
+
4. abort/stop owned background loops and consumers;
|
|
50
|
+
5. close database, queue, cache, HTTP, worker, and telemetry resources;
|
|
51
|
+
6. force termination only after the documented grace deadline.
|
|
52
|
+
- `server.close()` and force-closing connections have version-specific semantics. Verify the deployed Node line, long-lived/upgraded connections, keep-alive behavior, and orchestrator grace period.
|
|
53
|
+
- Prefer setting `process.exitCode` and allowing owned flushes to finish. Use immediate `process.exit()` only when the deliberate loss of pending asynchronous work is acceptable.
|
|
54
|
+
|
|
55
|
+
## Evidence-led performance
|
|
56
|
+
|
|
57
|
+
1. Reproduce with a representative payload, concurrency, runtime flags, dependency state, and warm-up period.
|
|
58
|
+
2. Capture an application outcome (latency distribution, throughput, timeout/error rate) and at least one causal signal (CPU profile, event-loop delay/utilization, worker-pool/queue depth, heap/GC, active resources).
|
|
59
|
+
3. Change one mechanism, rerun the same workload, and compare distributions rather than one fastest sample.
|
|
60
|
+
4. Check correctness and resource cleanup under load; faster wrong or leaking code is a regression.
|
|
61
|
+
|
|
62
|
+
Use CPU profiles/flame graphs for CPU attribution and heap/retainer evidence for memory claims. A heap snapshot stops the main thread and can approximately double heap use while being produced; do not take one from a sole production instance or expose an unauthenticated snapshot endpoint.
|
|
63
|
+
|
|
64
|
+
## Focused adversarial cases
|
|
65
|
+
|
|
66
|
+
- large but valid input;
|
|
67
|
+
- malformed input that exercises worst-case parsing/regex behavior;
|
|
68
|
+
- downstream never responds or ignores cancellation;
|
|
69
|
+
- client disconnects mid-stream;
|
|
70
|
+
- queue reaches its bound;
|
|
71
|
+
- shutdown arrives during startup and during in-flight work;
|
|
72
|
+
- worker crashes or returns an unserializable/oversized result;
|
|
73
|
+
- retry budget/deadline is exhausted.
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Runtime and Project Contract
|
|
2
|
+
|
|
3
|
+
Use this reference when a Node.js change touches runtime selection, package-manager state, ESM/CommonJS, TypeScript execution, dependencies, or configuration.
|
|
4
|
+
|
|
5
|
+
## Runtime selection
|
|
6
|
+
|
|
7
|
+
1. Prefer the repository and deployment contract over a generic recommendation. Reconcile `.nvmrc`, `.node-version`, `package.json#engines`, package-manager metadata, container base image, CI matrix, and production runtime; do not silently choose one when they disagree.
|
|
8
|
+
2. For a new production target with no existing contract, select a currently supported Active LTS or Maintenance LTS line from the live [Node.js release page](https://nodejs.org/en/about/previous-releases). Do not hardcode “latest” or infer support from odd/even numbering: the project announced an annual schedule beginning with Node.js 27, and the live [release schedule](https://github.com/nodejs/Release/blob/main/schedule.json) is authoritative when dates drift.
|
|
9
|
+
3. Use a Current or Alpha line only for an explicit compatibility/experimentation goal. Record the fallback and do not widen the production support claim from a development smoke test.
|
|
10
|
+
4. Test the lowest and highest supported runtime when a library or shared package promises a range. An application normally pins one deployment line and tests the upgrade candidate separately.
|
|
11
|
+
|
|
12
|
+
## Package-manager and dependency state
|
|
13
|
+
|
|
14
|
+
- Treat the committed lockfile plus package-manager metadata as the reproducibility contract. Do not switch npm/pnpm/yarn/bun or regenerate a foreign lockfile without an explicit migration.
|
|
15
|
+
- Use the manager's frozen install in CI and verification. For npm, `npm ci` requires an existing lockfile, removes an existing `node_modules`, fails when `package.json` and lock state disagree, and does not rewrite either file.
|
|
16
|
+
- Keep dependency additions proportional to the contract. Check maintenance, runtime support, transitive size, native build/install scripts, license constraints, and whether a built-in API already meets the need.
|
|
17
|
+
- Treat provenance/signatures as origin evidence, not a safety verdict. A vulnerability scan also cannot prove that a dependency is non-malicious or that a vulnerable path is reachable.
|
|
18
|
+
- Never accept automatic dependency updates solely because CI is green. Review behavior, changelog/security impact, lockfile delta, and rollback path.
|
|
19
|
+
|
|
20
|
+
## ESM and CommonJS
|
|
21
|
+
|
|
22
|
+
- Preserve the established module system. For a new package, set `package.json#type` explicitly and choose extensions/exports that match the actual consumers.
|
|
23
|
+
- Do not rely on Node.js syntax detection for ambiguous `.js` files. Explicit intent prevents behavior changes across runtimes and tooling.
|
|
24
|
+
- Treat package `exports` as a public compatibility contract. Test every promised import/require path from a packed artifact when publishing a library; service-internal path aliases still need runtime support, not only editor/type-check support.
|
|
25
|
+
- Keep dynamic import, top-level await, JSON/native modules, test runner, bundler, and deployment loader behavior in the compatibility matrix when the change uses them.
|
|
26
|
+
|
|
27
|
+
## TypeScript paths
|
|
28
|
+
|
|
29
|
+
Choose one explicit path:
|
|
30
|
+
|
|
31
|
+
| Path | Use when | Required proof |
|
|
32
|
+
|---|---|---|
|
|
33
|
+
| Compile/transpile before run | The service uses emitted JavaScript, transforms, decorators, path rewriting, or an older runtime | type-check, emitted artifact, source-map/error behavior, production start command |
|
|
34
|
+
| Runtime loader/tool | The repository already standardizes on one | loader version/runtime matrix, type-check remains separate, production parity |
|
|
35
|
+
| Node.js type stripping | The deployed runtime supports it and source uses erasable syntax only | no unsupported transform syntax, no `tsconfig`-dependent runtime behavior, explicit type-check gate |
|
|
36
|
+
| Plain JavaScript | The repository does not require TypeScript | runtime syntax target, lint/check path, public type contract if shipped as a library |
|
|
37
|
+
|
|
38
|
+
Node.js type stripping ignores `tsconfig.json`, performs no type checking, and does not transform syntax such as enums, runtime namespaces, parameter properties, or import aliases. Do not present it as a drop-in replacement for an existing compiler pipeline.
|
|
39
|
+
|
|
40
|
+
## Configuration contract
|
|
41
|
+
|
|
42
|
+
- Parse and validate configuration once during startup; pass typed/validated values inward rather than reading `process.env` throughout the codebase.
|
|
43
|
+
- Separate presence, format, range, and cross-field validation. Error messages may name a key but must not echo secret values.
|
|
44
|
+
- Define precedence among defaults, env files, environment variables, flags, secret mounts, and remote configuration. A newly available built-in flag is not permission to change repository precedence.
|
|
45
|
+
- Keep build-time and runtime configuration distinct. Verify container/orchestrator injection and local development paths separately.
|
|
46
|
+
|
|
47
|
+
## Contract checkpoint
|
|
48
|
+
|
|
49
|
+
Before implementation, be able to state:
|
|
50
|
+
|
|
51
|
+
```text
|
|
52
|
+
Runtime: <repo/deploy evidence and supported line>
|
|
53
|
+
Package manager: <manager + lockfile>
|
|
54
|
+
Modules: <ESM/CJS boundary>
|
|
55
|
+
Type path: <execution + type-check>
|
|
56
|
+
Public contract changed: <yes/no + exact surface>
|
|
57
|
+
Compatibility matrix: <minimum necessary rows>
|
|
58
|
+
```
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Maintainer Source Map
|
|
2
|
+
|
|
3
|
+
Inspected 2026-08-30. This file records provenance and extraction limits; it is not required reading for ordinary Node.js implementation work.
|
|
4
|
+
|
|
5
|
+
## Primary Node.js and npm sources
|
|
6
|
+
|
|
7
|
+
| Decision surface | Inspected source | Extracted constraint |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| production runtime | [Node.js releases](https://nodejs.org/en/about/previous-releases), [release schedule](https://github.com/nodejs/Release/blob/main/schedule.json), [2026 schedule announcement](https://nodejs.org/en/blog/announcements/evolving-the-nodejs-release-schedule) | production uses supported LTS; live schedule beats memorized odd/even rules; Node.js 27 begins the announced annual model |
|
|
10
|
+
| packages/modules | [Packages API](https://nodejs.org/api/packages.html) | make module intent explicit; ambiguous `.js` syntax detection is not a project contract |
|
|
11
|
+
| TypeScript | [TypeScript API](https://nodejs.org/api/typescript.html) | built-in stripping is stable on documented lines but performs no type checking, ignores `tsconfig`, and supports erasable syntax only |
|
|
12
|
+
| tests | [Test runner API](https://nodejs.org/api/test.html), [CLI API](https://nodejs.org/api/cli.html) | `node:test` is capable but individual coverage/CLI features remain version-gated; preserve the repo runner and verify the supported matrix |
|
|
13
|
+
| concurrency/context | [Worker threads](https://nodejs.org/api/worker_threads.html), [async context](https://nodejs.org/api/async_context.html) | workers suit CPU-intensive JavaScript, not ordinary I/O; prefer optimized `AsyncLocalStorage` over custom `async_hooks` context machinery |
|
|
14
|
+
| event loop / DoS | [Don't block the event loop](https://nodejs.org/en/learn/asynchronous-work/dont-block-the-event-loop) | long event-loop or worker-pool work reduces throughput and can create denial-of-service paths |
|
|
15
|
+
| cancellation/streams | [Global Abort APIs](https://nodejs.org/api/globals.html), [Streams API](https://nodejs.org/api/stream.html) | propagate cancellation to underlying work; pipeline/backpressure owns teardown for large data |
|
|
16
|
+
| fatal errors / shutdown | [Process API](https://nodejs.org/api/process.html), [HTTP API](https://nodejs.org/api/http.html) | unknown uncaught failures are not safe recovery points; connection-closing behavior is version- and protocol-sensitive |
|
|
17
|
+
| diagnostics | [Heap snapshots](https://nodejs.org/en/learn/diagnostics/memory/using-heap-snapshot), [flame graphs](https://nodejs.org/en/learn/diagnostics/flame-graphs) | performance claims need profiles/measurements; heap snapshots can stop the main thread and exhaust memory |
|
|
18
|
+
| security | [Node.js security best practices](https://nodejs.org/en/learn/getting-started/security-best-practices) | bound input work, harden dependencies, and use runtime permissions only as defense in depth |
|
|
19
|
+
| reproducible install / provenance | [npm ci](https://docs.npmjs.com/cli/v11/commands/npm-ci/), [npm provenance](https://docs.npmjs.com/generating-provenance-statements) | frozen lockfile install is the verification contract; provenance proves origin/build linkage, not code safety |
|
|
20
|
+
|
|
21
|
+
## Independent industry controls
|
|
22
|
+
|
|
23
|
+
- [OWASP NodeJS Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Nodejs_Security_Cheat_Sheet.html): used to challenge missing web/runtime security axes. Framework- or version-specific prescriptions were not copied without Node.js primary-source support.
|
|
24
|
+
- [OpenSSF npm supply-chain guidance](https://openssf.org/blog/2022/09/01/npm-best-practices-for-the-supply-chain/): used to challenge lockfile, install-script, dependency, CI, and provenance handling. Supply-chain release governance remains routed to existing CCL owners.
|
|
25
|
+
- [Node.js Best Practices](https://github.com/goldbergyoni/nodebestpractices): broad practitioner checklist consulted as coverage input only; its July 2024 edition and library preferences are not treated as current runtime authority.
|
|
26
|
+
|
|
27
|
+
## Evaluation-method sources
|
|
28
|
+
|
|
29
|
+
- [Anthropic, Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents): grounds the split between a task's inputs and success criteria, multiple trials for variable outputs, combined code/model/human graders, and review of both outcomes and transcripts. This supports separate routing and body-effect measurements; it does not make either one a merge gate.
|
|
30
|
+
- [OpenAI, How evals drive the next chapter in AI for businesses](https://openai.com/index/evals-drive-next-chapter-of-ai/): grounds the specify → measure → improve loop and contextual evals tied to the actual workflow rather than generic benchmark scores.
|
|
31
|
+
- [OpenAI, A shared playbook for trustworthy third party evaluations](https://openai.com/index/trustworthy-third-party-evaluations-foundations/): grounds binding claims to the tested system, harness, task distribution, budget, elicitation method, and validity checks. A skill-content result must therefore identify the exact skill snapshot and host conditions it tested.
|
|
32
|
+
- [On Randomness in Agentic Evals](https://arxiv.org/abs/2602.07150): large-sample evidence that agent trajectories vary even at temperature zero; a future effectiveness claim needs repeated independent trials and uncertainty, not one favorable answer.
|
|
33
|
+
- [Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge](https://aclanthology.org/2025.ijcnlp-long.18/): supports balanced answer order and human inspection for pairwise model judgments. An LLM preference is advisory evidence, not an oracle.
|
|
34
|
+
|
|
35
|
+
## Extraction limits
|
|
36
|
+
|
|
37
|
+
- This is a source-backed design, not proof that every rule improves agent behavior in production. The deterministic RED/GREEN baseline proves discovery/routing registration and repository conformance only.
|
|
38
|
+
- The Node.js body fixtures freeze tasks and rubrics for later paired evaluation. Until repeated with/without runs bind the same model, host, tools, budget, skill snapshot, and blind grading procedure, their result class remains `insufficient-evidence`.
|
|
39
|
+
- Runtime and CLI stability can change. The skill deliberately tells the agent to resolve the live release/support matrix instead of hardcoding Node.js 24/26.
|
|
40
|
+
- Framework-specific internals were not generalized. Express, Fastify, NestJS, Hono, and other frameworks retain their repository-local lifecycle and security contracts.
|
|
41
|
+
- No peer-skill text was copied as authoritative runtime guidance; public peer skills were used for structure and collision analysis, while technical claims were checked against primary Node.js/npm sources. Snapshot details stay in the extraction register rather than the distributed runtime skill.
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Verification, Diagnostics, and Security
|
|
2
|
+
|
|
3
|
+
Use this reference when selecting concrete Node.js test mechanics, proving runtime behavior, changing dependencies, or reviewing Node-specific security exposure.
|
|
4
|
+
|
|
5
|
+
## Test mechanics after strategy is chosen
|
|
6
|
+
|
|
7
|
+
Preserve the repository's runner. `node:test` is a first-class built-in option, not a mandatory migration target. Choose a new runner only from actual needs such as ecosystem integration, transform support, watch/UI workflow, mocking behavior, coverage maturity, or multi-project support.
|
|
8
|
+
|
|
9
|
+
| Behavior | Focused proof |
|
|
10
|
+
|---|---|
|
|
11
|
+
| Handler/domain logic | direct unit test without port/global mutation |
|
|
12
|
+
| HTTP/RPC adapter | in-process integration test for status/schema/error mapping |
|
|
13
|
+
| Database/queue/cache adapter | real protocol dependency or contract-faithful test double at the adapter boundary |
|
|
14
|
+
| Timeout/cancellation | fake/controlled time where sound, plus assertion that underlying work stopped |
|
|
15
|
+
| Stream | backpressure, partial failure, disconnect, cleanup, size bound |
|
|
16
|
+
| Worker/background job | ownership, retry/idempotency, poison input, shutdown/drain |
|
|
17
|
+
| Process lifecycle | child-process test for signal, exit code, readiness/drain deadline |
|
|
18
|
+
| Package/module boundary | start/import/require the built or packed artifact on the supported runtime matrix |
|
|
19
|
+
|
|
20
|
+
Mock the narrow external boundary, not the implementation under test. Reset mocks/timers and avoid cross-test process-global mutation. When concurrency is meaningful, test ordering independence and run the suspected flaky case repeatedly before calling it stable.
|
|
21
|
+
|
|
22
|
+
Coverage is a gap detector, not the acceptance oracle. Keep an existing threshold; change policy through `testing-strategy`. Node's built-in coverage and threshold flags have version/stability differences, so verify them against every supported runtime before making them a required gate.
|
|
23
|
+
|
|
24
|
+
## Diagnostic decision table
|
|
25
|
+
|
|
26
|
+
| Symptom | Start with | Avoid claiming from |
|
|
27
|
+
|---|---|---|
|
|
28
|
+
| high CPU / latency | reproducer, CPU profile/flame graph, event-loop and queue evidence | one stack sample or code inspection |
|
|
29
|
+
| event-loop stall | event-loop delay/utilization plus long-callback attribution | total CPU alone |
|
|
30
|
+
| memory growth | heap/GC trend, retained-object comparison, active resources | RSS snapshot alone |
|
|
31
|
+
| process will not exit | active handles/resources, owned timers/listeners/workers, lifecycle trace | adding forced `process.exit()` |
|
|
32
|
+
| worker-pool starvation | workload class, async-resource duration/concurrency, pool queue symptoms | increasing pool size first |
|
|
33
|
+
| flaky async test | repeated isolated run, seed/time/concurrency capture, leaked-resource check | blanket timeout increase |
|
|
34
|
+
|
|
35
|
+
Bind evidence to the candidate runtime and commit. A diagnostic command that failed, timed out, or could not attach is missing evidence, not a clean result.
|
|
36
|
+
|
|
37
|
+
## Security review axes
|
|
38
|
+
|
|
39
|
+
- **Input and complexity:** validate type/shape/range; limit headers, bodies, decompression, records, regex complexity, recursion, and parsing work. An input-dependent long callback is both performance and DoS risk.
|
|
40
|
+
- **Injection and paths:** use parameterized database/process APIs, avoid shell construction, normalize and constrain filesystem paths, and define archive/symlink behavior.
|
|
41
|
+
- **HTTP boundary:** keep secure parser defaults, schema-validate untrusted network input, set explicit timeouts/limits appropriate to the framework/runtime, and do not expose framework diagnostics or raw errors.
|
|
42
|
+
- **Outbound access:** validate destinations and redirects where user input influences network access; apply platform egress controls for service-level guarantees.
|
|
43
|
+
- **Secrets and logs:** never log credentials/tokens/raw sensitive payloads; redact at structured logging boundaries and test representative failure paths.
|
|
44
|
+
- **Dependencies:** review direct and transitive changes, install scripts, lockfile delta, known advisories, maintainer/package identity, and rollback. `npm audit` or equivalent is one signal, not a pass/fail security proof.
|
|
45
|
+
- **Runtime containment:** the Node.js Permission Model can reduce filesystem/network/process/addon/worker capabilities on supported runtimes. Verify flags and framework needs in the deployment environment. It is defense in depth and does not make malicious in-process code trustworthy.
|
|
46
|
+
- **Prototype/object hazards:** accept only expected keys, use schema validation, avoid unsafe recursive merge of untrusted objects, and keep framework/runtime patched.
|
|
47
|
+
|
|
48
|
+
Route threat-model and required-review gate decisions to `feature-risk-router`. Route service-wide identity, authorization, tenant isolation, data ownership, or network-policy architecture through `product-rd-workflow` and the appropriate platform/security owner before implementation.
|
|
49
|
+
|
|
50
|
+
## Verification transcript
|
|
51
|
+
|
|
52
|
+
Capture a compact table:
|
|
53
|
+
|
|
54
|
+
| Check | Command/probe | Result | Candidate/runtime |
|
|
55
|
+
|---|---|---|---|
|
|
56
|
+
| focused behavior | repository command | pass/fail | SHA + Node line |
|
|
57
|
+
| failure/cancellation | test or reproducer | pass/fail | same |
|
|
58
|
+
| lint/type/check | repository command | pass/fail | same |
|
|
59
|
+
| package/service suite | repository command | pass/fail/skipped | same |
|
|
60
|
+
| build/start/package | repository command | pass/fail/skipped | same |
|
|
61
|
+
| performance/diagnostic | frozen workload/probe | measured/inconclusive | same |
|
|
62
|
+
|
|
63
|
+
Do not collapse skipped, unavailable, inconclusive, and passed into one “green” status.
|
|
@@ -46,6 +46,7 @@ A new service must satisfy all five before it is allowed in production. A releas
|
|
|
46
46
|
- Logger MUST extract trace context from `context.Context` and attach `_trace_id`/`_span_id`/`_trace_flags` to the log record automatically. If a log call requires the developer to manually pass trace fields, the framework is broken.
|
|
47
47
|
- Async work that detaches from the request lifecycle (audit, reporting, ledger side-paths) keeps correlation — but only correlation: capture trace/span/log-id fields before the handoff and carry them (plus durable work-item identity) into the detached context; never a bare `context.Background()` that drops linkage, and never a wholesale request-context copy — `context.WithoutCancel` preserves every request value, so a naive derive smuggles request auth/session/secrets/PII into work that outlives the request and can run under stale, revoked authority. Whitelist correlation fields, strip the rest; the worker re-authorizes or runs under service identity (lifecycle/loss-policy detail owned by the stack dev skills' async side-path rules — this bullet owns only the correlation contract).
|
|
48
48
|
- Verification: correlation is an acceptance item, not a default — wiring the unified logging component is NOT evidence it works. Assert over the real transport: drive one request through the actual RPC/HTTP server in a test, assert the handler ctx carries a non-empty trace-id/log-id, and that business + access log lines actually contain those fields. Then live: pick any production user action → find one log line → click trace-id → see the full multi-hop trace → all spans share the same log-id.
|
|
49
|
+
- Query/dashboard discipline (evidence-boundary, non-intrusion, discover-before-create, env-resolution): `references/metrics-conventions.md`.
|
|
49
50
|
|
|
50
51
|
### R2 — Framework default observability, not opt-in
|
|
51
52
|
|
|
@@ -117,8 +118,8 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
|
|
|
117
118
|
### R8 — SLI/SLO discipline
|
|
118
119
|
|
|
119
120
|
- Define SLIs from the **user's perspective** (success rate of a request type, latency of a user-visible action), not from internal counters.
|
|
120
|
-
- Prefer expressing an SLI as **good events / valid events**
|
|
121
|
-
-
|
|
121
|
+
- Prefer expressing an SLI as **good events / valid events** — valid first, then good — per Google's Art of SLOs (the Workbook itself says good/total). A latency SLI is the **share of valid requests faster than a threshold** (`count(latency≤T)/valid` — the denominator is the availability SLI's valid set, never raw total), NOT a percentile — P95/P99 are dashboard aids, not the SLI. An error-log counter is a diagnostic signal, not an availability-SLI input.
|
|
122
|
+
- Uncomputable signals are blind spots, not near-coverage: record each explicitly ("can't measure X because Y" — e.g. a failure counter with no attempt total yields no error *rate*; uncollected content means unmeasurable input-semantic drift) in a blind-spot register instead of pretending coverage, so on-call never leans on a nonexistent signal.
|
|
122
123
|
- Each user-visible journey gets at least one availability SLI + one latency SLI.
|
|
123
124
|
- Set SLO targets, error budgets, and burn-rate alerts (SRE Workbook multiwindow tiers, derived for a 30d budget window — recompute for other periods: page 14.4× 1h/5m, page 6× 6h/30m, ticket 1× 3d/6h). Each tier MUST evaluate its long AND short window together, firing only when both burn above threshold; the short (~1/12) window makes paging stop soon after the burn stops.
|
|
124
125
|
- An SLO without an error-budget-driven release decision is decoration. See `references/sli-slo-design.md`.
|
|
@@ -50,7 +50,7 @@ Async observable gauges via OTel SDK reduce noise vs synchronous sets. Pattern:
|
|
|
50
50
|
| 外部框架 | 关注什么 | 适用对象 | 映射到本 ref instrument |
|
|
51
51
|
|---|---|---|---|
|
|
52
52
|
| **Google 4 Golden Signals**(SRE Book)| Latency / Traffic / Errors / Saturation | service-level(user-facing service)| Latency → Histogram;Traffic → Counter (request rate);Errors → Counter (error count);Saturation → Gauge (queue / pool / cpu / mem) |
|
|
53
|
-
| **RED**(Tom Wilkie
|
|
53
|
+
| **RED**(Tom Wilkie;原始 Weaveworks 博客已随公司关停下线,现存最佳出处为 Grafana 官方博文 "The RED Method")| Rate / Errors / Duration | request-driven service / RPC endpoint | Rate → Counter;Errors → Counter;Duration → Histogram(近似等于 request-side Golden Signals 去 Saturation;非正式 derivation)|
|
|
54
54
|
| **USE**(Brendan Gregg)| Utilization / Saturation / Errors | resource(CPU / memory / disk / NIC / connection pool)| Utilization → Gauge (%) ;Saturation → Gauge (queue depth);Errors → Counter(资源驱动而非请求驱动)|
|
|
55
55
|
|
|
56
56
|
何时用哪个:
|
|
@@ -103,3 +103,10 @@ The 15s reader interval matches Prometheus scrape conventions and keeps OTLP pus
|
|
|
103
103
|
- Using gauge for monotonic counters (loses rate semantics on restart).
|
|
104
104
|
- Histogram with too few buckets (loses percentile precision) or too many (storage cost).
|
|
105
105
|
- Recording metrics inside a tight loop without rate-limiting (collector receives bursts).
|
|
106
|
+
|
|
107
|
+
## Observation Discipline(查询/看板/新增信号的通用纪律)
|
|
108
|
+
|
|
109
|
+
- Cross-layer evidence boundary: a client-side event does not prove backend success, and a backend metric does not prove the user saw success — any conclusion crossing the client/backend (or service/service) boundary requires identifiers or time windows aligned across the layers, never a same-shape count on each side.
|
|
110
|
+
- Observation code never intrudes on the observed path: **diagnostic telemetry** (metrics, traces, debug logs) is best-effort and must not add retries, blocking waits, or business-logic branches to the monitored path — observability that changes the behavior it measures is its own defect class. This best-effort license covers diagnostics only: mandatory records — security audit, billing/ledger, deletion/erasure evidence, release-gate records — are business writes with their own durable-delivery and failure semantics (their loss is a failure, never shrugged off as telemetry).
|
|
111
|
+
- Discover before creating: before proposing a new event, metric, label, panel, or query, enumerate what already exists for that surface and extend/reuse it — parallel near-duplicate signals fragment dashboards and split history.
|
|
112
|
+
- Environment-name resolution: a user's explicit component/branch/URL/datasource always wins; never silently substitute a different physical environment for a colloquial environment word — resolve an unqualified name to the recorded default and say which one was used.
|
|
@@ -45,11 +45,11 @@ If you cannot write the query against existing metrics, the metric set is incomp
|
|
|
45
45
|
|
|
46
46
|
## SLO target
|
|
47
47
|
|
|
48
|
-
Pick a number that reflects user expectation and product maturity, not aspiration:
|
|
48
|
+
Pick a number that reflects user expectation and product maturity, not aspiration. The ladder below is a **team-heuristic starting point, not an industry standard** — Google SRE literature deliberately gives no numeric ladder; its direction is that each extra nine costs sharply more for marginal utility approaching zero (SRE Workbook Ch.2), and that targets should come from user expectation, not current performance:
|
|
49
49
|
- New service: 99.0% — generous.
|
|
50
50
|
- Mature service: 99.9% — three nines.
|
|
51
51
|
- Critical path (payment, login): 99.95% — push.
|
|
52
|
-
- Never start with 99.99% without 24/7 staffing.
|
|
52
|
+
- Never start with 99.99% without 24/7 staffing (team heuristic: at four nines the monthly budget is minutes, which no unstaffed rotation can defend).
|
|
53
53
|
|
|
54
54
|
Lower the target if every release burns the budget. Raise it only after sustained achievement.
|
|
55
55
|
|