@ccoalm/ccl-skills 0.15.0 → 0.15.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (71) hide show
  1. package/README.md +3 -1
  2. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +1 -1
  3. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +80 -4
  4. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +16 -4
  5. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +13 -2
  6. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +53 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -2
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -5
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +26 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +37 -13
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +5 -2
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +77 -5
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +98 -4
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +48 -1
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +237 -17
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +165 -11
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +143 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +572 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +65 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +3 -1
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +1 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +2 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +7 -5
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +8 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +3 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +16 -14
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +1 -1
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +40 -83
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +2 -2
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +44 -37
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +1 -1
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +7 -7
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +1 -1
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +20 -11
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +7 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +8 -5
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +2 -2
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +17 -17
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +3 -3
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +4 -4
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +3 -3
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +24 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +31 -7
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +74 -3
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +103 -21
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +83 -48
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +80 -2
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +74 -1
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +3 -2
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_runtime.py +428 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +190 -2
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +106 -4
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -1
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -1
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -1
  61. package/dist/assets/release.json +87 -67
  62. package/dist/claude-adapter.js +14 -7
  63. package/dist/codex-host.d.ts +2 -4
  64. package/dist/codex-host.js +40 -19
  65. package/dist/host-probe.d.ts +27 -0
  66. package/dist/host-probe.js +51 -0
  67. package/dist/opencode-adapter.js +24 -19
  68. package/dist/operations.js +34 -8
  69. package/dist/unified.d.ts +1 -1
  70. package/dist/unified.js +18 -9
  71. package/package.json +1 -1
@@ -5,7 +5,9 @@ description: bug / 报错 / test 挂了 / 线上问题 / 接口变慢·性能退
5
5
 
6
6
  # Defect Diagnosis
7
7
 
8
- Use this skill for the full defect discipline: diagnose the immediate failure, fix it with evidence, then decide whether root-cause prevention should update a product, architecture, development, testing, release, or tooling skill.
8
+ Diagnose and fix from evidence; route prevention to product, architecture, development, testing, release or tooling.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  ## Non-Negotiable Rules
11
13
 
@@ -5,7 +5,7 @@ description: Use when implementing, modifying, scaffolding, generating, or testi
5
5
 
6
6
  # Go Microservice Dev
7
7
 
8
- Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase.
8
+ Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase. After code/test edits, self-check and invoke `code-review` automatically before completion.
9
9
 
10
10
  ## Skill Routing
11
11
 
@@ -7,6 +7,8 @@ description: Use when designing, implementing, reviewing, debugging, or operatin
7
7
 
8
8
  Use this for product backend work that calls, hosts, evaluates, or operates LLM and inference systems. Keep the skill generic: extract reusable mechanics only, not business-specific prompts, datasets, provider names, repository paths, or domain nouns.
9
9
 
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
11
+
10
12
  ## Skill Routing
11
13
 
12
14
  - Use this skill for LLM gateway/client design, model registry, prompt versioning, agent/tool orchestration, streaming APIs, fallback, token/cost accounting, evals, replay, shadow comparison, batch inference, and inference observability.
@@ -5,7 +5,7 @@ description: "小程序 / Taro / 微信小程序 / 支付宝小程序 / 抖音
5
5
 
6
6
  # Miniapp Product Dev
7
7
 
8
- Use this skill for mini-program client engineering and platform delivery. It covers product-facing miniapp work across WeChat, Alipay, Douyin/TikTok, Baidu, and similar host platforms. It does not own general product strategy, backend service architecture, or visual design rules.
8
+ Mini-program client engineering and product-facing delivery across WeChat, Alipay, Douyin/TikTok, Baidu and similar hosts; excludes product strategy, backend architecture and visual design. After code/test edits, self-check and invoke `code-review` automatically before completion.
9
9
 
10
10
  ## Framework Scope
11
11
 
@@ -57,7 +57,7 @@ This skill is about **using AI agents / subagents to execute work** — delegati
57
57
  5. **Value**: the task is worth that premium.
58
58
  If independence, value, or breadth is unclear, start with one focused agent or sequential delegation; use parallel multi-agent dispatch when those checks are explicitly satisfied.
59
59
  - Model tier per dispatch is an explicit decision, not an inherited accident. On hosts that inherit the session model for unnamed dispatches (a common default — Claude-family harnesses behave this way; verify yours), an unnamed model is often the most capable and most expensive tier, so a high-volume fan-out silently puts every worker and reviewer on the top tier. The observable triggers are a fan-out — multiple dispatches (workers/reviewers) in one delivery — and any high-risk dispatch: there, name the tier per dispatch and choose by judgment complexity and risk, not token price alone — the cheapest tier routinely takes 2–3× the turns on multi-step work and costs more overall, so use a mid-tier floor for reviewers and for implementers working from prose descriptions; reserve the cheapest tier for transcription-plus-tests tasks (the plan text already contains the code to write) and single-file mechanical fixes; put architecture/design judgment, security/authority/tenant-isolation/data-loss review, and the final whole-scope review on the most capable tier — review tier scales with the diff's size, complexity, and risk, and a high-risk review never silently inherits a cheap session default even as a single dispatch (tier principle: `../skill-extraction-workflow/references/harness-patterns-and-eval.md`). Record the decision as a brief field alongside `required_skills`: `model_tier: <tier>` or `model_tier: host-default (<reason: single low-risk dispatch | no host model selection>)` — an absent field is an unmade decision, not a default, and the field is bookkeeping only until the dispatch call actually passes the matching model selector (verify the effective model where the host exposes it; a field/selector mismatch is a defect, not a recorded decision).
60
- - Stop and escalate when the plan is unclear, a dependency is missing, verification fails repeatedly (same error ~3 times — identical retries, not new findings from successive review rounds), or an agent returns unsupported claims. An escalation message must carry five fields, or it is just "stuck": the specific blocker, the attempts made and their results, the current state (diff / commits / workspace), the safest next action for the human to align on, and whether a lower-risk part can continue meanwhile.
60
+ - Stop and escalate when unclear direction, unavailable dependencies, unknown completion state or missing authority remains after bounded remediation. Repeated identical verification failures (~3 times) require a status report and method/evidence checkpoint, not renewed task permission; stop identical retries and continue necessary work within existing scope, respecting explicit user limits. Escalation must name the blocker, attempts and results, current diff/commits/workspace, safest next action, and lower-risk work that can continue.
61
61
 
62
62
  ## Execution Flow
63
63
 
@@ -5,6 +5,8 @@ description: Use when implementing, modifying, scaffolding, or testing Node.js b
5
5
 
6
6
  # Node.js Service Development
7
7
 
8
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
+
8
10
  ## Skill Routing
9
11
 
10
12
  - Use this skill for Node.js service implementation: handlers, middleware, adapters, workers, jobs, clients, runtime/toolchain mechanics, and focused tests.
@@ -5,7 +5,9 @@ description: Use when designing, reviewing, debugging, or shipping observability
5
5
 
6
6
  # Platform Observability
7
7
 
8
- This skill owns the **evidence layer**: logs, metrics, traces, log/trace correlation, dashboards, alerts, on-call routing, SLI/SLO/error-budget design, and the framework-level wiring that guarantees a new service is observable on day one.
8
+ Owns logs/metrics/traces, correlation, dashboards, alerts, on-call routing, SLI/SLO/error-budget design and framework wiring for new services.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  You do not own:
11
13
  - Traffic routing, retries, timeouts, mTLS, or service mesh policy — go to `platform-service-connectivity`.
@@ -30,7 +32,7 @@ Observability is the chain that turns one user action into searchable, joinable,
30
32
  2. **Instrumentation** — every service emits structured logs, metrics, and spans through the **framework default**, not ad-hoc code. A service whose middleware chain does not include `ctx_inject + metrics + recovery + tracing` is unobservable by design.
31
33
  3. **Transport** — logs go stdout → file collector (DaemonSet) → log pipeline → search index; metrics + traces go SDK → OTLP collector → metric store + trace store. Both transports must survive collector restarts and back-pressure.
32
34
  4. **Storage + display** — logs in a searchable index keyed by log-id and trace-id; metrics in a long-term store separate from scraper; traces queryable by trace-id; one dashboard tool joins all three.
33
- 5. **Evidence consumption** — SLIs are queries against (3) + (4); alerts evaluate SLIs and route to on-call; runbooks live in a wiki that on-call can reach in under one minute.
35
+ 5. **Evidence consumption** — SLIs query (3) + (4); actionable alerts route to on-call; runbooks live in a wiki that on-call can reach in under one minute.
34
36
 
35
37
  A new service must satisfy all five before it is allowed in production. A release must produce evidence at all five before it is promoted.
36
38
 
@@ -106,9 +108,9 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
106
108
 
107
109
  **Retrofit / migration contract** (the rules above are otherwise greenfield-framed): when a schema is introduced over EXISTING services, field names, metric label names, and identity env-var names are a **migration contract** — a blind rename breaks every deployed dashboard, alert, and saved query that keys on the old name. Retrofitting MUST alias or dual-write old→new and migrate consumers before retiring the old name; never rename in place. The cheap time to fix a name is before services adopt it.
108
110
 
109
- ### R7 — Alerts are SLI-driven and route to a human
111
+ ### R7 — Alerts are actionable and route to an owner
110
112
 
111
- - Alerts evaluate against SLIs (query the long-term metric store), not raw metrics.
113
+ - Use SLIs for SLO paging; actionable capacity or impending-failure alerts may use internal measurements. See `references/alerting-and-on-call.md`.
112
114
  - Severity levels: P0 (page on-call now), P1 (notify channel, ack within work hours), P2 (digest).
113
115
  - Every alert MUST link to a runbook entry. If no runbook exists, the alert is not allowed to be P0.
114
116
  - An alert-backed metric is a coverage contract, and the trigger is mechanical, not prose: any new or changed alert rule, SLO, dashboard alert annotation, or metric referenced by an alert policy fires this check (a written "alert on any increase" note also counts, but its absence is not an exemption). Enumerate every site that should feed the metric and verify each is actually instrumented — derive the site list from a static registry or lint where possible; the easiest site to miss is often the riskiest (e.g. the panic counter in a stream reader parsing untrusted bytes). Ship the site checklist with the alert.
@@ -208,7 +210,7 @@ If any phase has missing evidence, the work is not done.
208
210
  - **"Sample everything vs head-sample 1%"** → head-sample low (1–10%) for cost; retaining errors/slow requests is **tail** sampling at the Collector and requires near-full SDK export to it (head-dropped spans never arrive) — pick one model per service, do not claim both. Sampling-by-route is acceptable for known noisy paths.
209
211
  - **"Metrics egress: scrape vs push"** → each platform picks ONE canonical egress and every process uses it. A Prometheus-native platform may default to **scrape** (pull, with `up`/target-health + service discovery); an OTel-first platform, or workers/jobs/runtimes that can't be scraped, may default to **OTLP push** to a collector. Neither is universally better — don't overturn a working pull setup just to push. One egress per process for a given metric; a migration window may dual-write only with isolated pipelines / distinct metric names / a dedup plan. Long-term store stays separate from the short-term scrape/collect layer (R5).
210
212
  - **"Add a new label to a metric"** → answer the cardinality question first. If max distinct values × series count > 1e6, refuse and use an exemplar trace instead.
211
- - **"Alert on this symptom or that cause"** → alert on user-visible symptom; cause-based alerts produce paging spam.
213
+ - **"Alert on this symptom or that cause"** → prefer user-visible symptoms; capacity or impending-failure warnings follow R7.
212
214
 
213
215
  ## Sanitization and Provenance
214
216
 
@@ -1,5 +1,13 @@
1
1
  # Alerting and On-Call
2
2
 
3
+ ## Choose an actionable signal
4
+
5
+ - **User symptoms and SLOs:** use service-level signals for user-symptom paging and the corresponding SLIs for SLO burn-rate alerts. Query the declared metric store; diagnostic counters do not become availability or latency SLIs merely because an alert uses them.
6
+ - **Capacity and impending failures:** white-box measurements may warn before user SLIs degrade, such as a credible forecast of disk exhaustion. Record the expected failure, the evidence for the threshold or forecast, the action, and the responsible owner. Choose urgency from impact and time left to act; a high internal metric alone is not a reason to page.
7
+ - **Diagnostic signals:** keep measurements with no actionable condition in dashboards or queries. Do not alert on every possible cause.
8
+
9
+ Verify new or changed alert rules against relevant failure, healthy, and recovery cases. Preserve the severity, ownership, runbook, and delivery requirements below for both symptom and impending-failure alerts. [Google SRE's Monitoring Distributed Systems](https://sre.google/sre-book/monitoring-distributed-systems/) describes both user symptoms and imminent saturation as alerting inputs.
10
+
3
11
  ## Two patterns
4
12
 
5
13
  ### Pattern A — Prometheus AlertManager
@@ -5,7 +5,9 @@ description: 发布 / 灰度 / canary / rollback / rollout / 环境泳道 / prom
5
5
 
6
6
  # Platform Release Engineering
7
7
 
8
- This skill owns **how a change reaches users and how it gets recalled**: lane / environment topology, build-to-deploy pipeline, traffic-shifting strategies (canary, blue-green, mirror), promotion gates that consume observability evidence, secret + dynamic-config distribution, and the rollback contract.
8
+ Owns lane/environment topology, build/deploy pipelines, traffic shifting (canary, blue-green, mirror), promotion gates using observability evidence, secret/dynamic-config distribution and rollback contracts.
9
+
10
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
11
 
10
12
  You do not own:
11
13
  - What signals exist to judge a release — see `platform-observability`.
@@ -5,6 +5,8 @@ description: 服务互通 / service mesh / service discovery / mTLS / retry / ti
5
5
 
6
6
  # Platform Service Connectivity
7
7
 
8
+ - Code/test changes require self-checks; invoke `code-review` automatically before completion.
9
+
8
10
  This skill owns **how requests move between services**: the transport layer (mesh), service discovery, multi-environment routing, retry/timeout/circuit-breaker policy, and the framework middleware that propagates app-level context across hops.
9
11
 
10
12
  You do not own:
@@ -100,10 +102,11 @@ Confusing these is the source of most "why doesn't this work" connectivity bugs.
100
102
 
101
103
  ### R4 — Default retry/timeout/circuit-breaker live at mesh, app overrides for business reasons
102
104
 
103
- - Mesh DestinationRule provides default per-callee policy: retries (1-2 attempts on 5xx/connect-fail), connect timeout, request timeout ceiling, outlier detection (5xx-percent → eject).
104
- - Framework client SDK provides per-call override: business timeout (always ≤ mesh ceiling), retry policy for idempotent calls, hedging.
105
+ - On Istio HTTP/gRPC paths, `VirtualService` HTTP routes own request `timeout` and `retries`; `DestinationRule` owns connection-pool settings and `outlierDetection`. Other transports use their platform-owned equivalents.
106
+ - Before configuring retries or timers, read `references/retry-timeout-circuit-breaker.md` for field paths, single-owner retry selection, and idempotency checks. Both mesh and SDK follow them; a 5xx or missing response alone never proves replay safe.
107
+ - Framework client SDK may provide method-specific retry/hedging; its total call budget must fit the caller's remaining duration and any platform cap. A longer mesh timeout is a backstop, not an extended caller deadline.
105
108
  - App code MAY override per-RPC. App code MUST NOT silently disable mesh-level outlier detection.
106
- - Budget rule: total upstream timeout = caller deadline minus a safety margin (e.g. 100ms). Cascading timeouts must shrink down the call chain.
109
+ - Budget rule: downstream work, attempts, and backoff fit the remaining duration minus a safety margin (e.g. 100ms). Propagate deadline and cancellation; never reset the full budget at each hop.
107
110
 
108
111
  ### R5 — mTLS is mesh-default, app cannot disable
109
112
 
@@ -151,14 +154,13 @@ The client middleware fills this from ctx automatically; the server middleware e
151
154
 
152
155
  For non-protobuf or metadata-only transports, an equivalent header set is compliant only when it cites a resolvable platform owner record, such as a repo/path, document URL, registry id, gateway policy, or owner-suite id. The owner record must enumerate the concrete header names. In diff-only review without resolver tooling, the diff must cite a stable owner-record locator and list the concrete header names inline; with resolver tooling, the reviewer or owner-suite may resolve the locator to those names instead. The enumerated header names must then be checked against the actual propagated and exposed headers: caller-supplied identity headers are absent unless positive authenticated-caller evidence exists. An unresolvable, non-enumerating, uncited, service-local, or unchecked header set is an open gap, not a permitted alternative. For pure HTTP (no RPC base), the equivalent is a stable owner-recorded header set, also filled by middleware.
153
156
 
154
- ### R8 — gRPC `:authority` and DNS-label hyphenation
157
+ ### R8 — gRPC `:authority` compatibility follows the actual request path
155
158
 
156
- If the platform allows service names with underscores (`<owner>.<class>.<env>` containing `_`), gRPC will reject them in the HTTP/2 `:authority` pseudo-header because it must be a valid DNS label (no `_`). Two acceptable fixes:
159
+ HTTP/2 `:authority` carries URI authority, not a single DNS label. An underscore alone does not prove a gRPC protocol violation. DNS hostnames, certificate identity checks, SDK validation, and proxy routing can impose different constraints; preserve the constraints of the deployed path.
157
160
 
158
- 1. **Disallow underscores in new service names**; enforce in registry registration and CI.
159
- 2. **Mesh-level rewrite**: an EnvoyFilter Lua snippet replaces `_` with `-` in `:authority` for gRPC requests, before routing.
161
+ Before renaming a service or adding a rewrite, capture the failing request's authority, the rejecting layer and version, and the relevant error or trace. A platform may require DNS-compatible service names and enforce that policy at registration and CI, but a registry name need not be the wire authority.
160
162
 
161
- Option 1 is cleaner long-term; option 2 is the live-system workaround. Document which the platform uses; new services should follow option 1.
163
+ A proxy rewrite is an option only when the request reaches that filter before the rejecting layer. It cannot fix a client rejection before transmission or an inbound parser rejection before the filter runs. Scope any verified rewrite to the affected traffic and explicit authority mappings; retain route, TLS identity, and authorization checks. Diagnosis, migration, and verification: `references/grpc-authority-workaround.md`.
162
164
 
163
165
  ### R9 — Ingress and egress are explicit, not implicit
164
166
 
@@ -200,8 +202,8 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
200
202
  ### Phase B — Mesh policy
201
203
 
202
204
  1. PeerAuthentication = STRICT (mTLS namespace-wide).
203
- 2. DestinationRule per critical callee: retry policy, timeouts, outlier detection.
204
- 3. VirtualService rules express lane-based routing: lane header match → lane subset.
205
+ 2. DestinationRule per critical callee: connection-pool limits, connect timeout, outlier detection.
206
+ 3. VirtualService HTTP routes: lane header match → lane subset, request timeout, and the R4 retry policy.
205
207
  4. AuthorizationPolicy expresses which services may call which (zero-trust at network level).
206
208
 
207
209
  ### Phase C — Service discovery
@@ -216,7 +218,7 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
216
218
 
217
219
  1. Read the framework default client/server options module. Confirm middleware chain matches R6.
218
220
  2. Confirm dev cannot build a client/server without inheriting these.
219
- 3. Test: kill a downstream pod → verify mesh outlier-detection ejects, framework retry kicks in for idempotent calls, error propagates up with stable error-code.
221
+ 3. Test: induce the configured outlier threshold on one callee → verify ejection and recovery, the configured retry layer retries only eligible calls within its budget, and exhausted calls propagate a stable error-code.
220
222
 
221
223
  ### Phase E — Failure modes
222
224
 
@@ -234,7 +236,7 @@ Before marking work done:
234
236
 
235
237
  ## Decision Points
236
238
 
237
- - **"Add a new retry policy for service X"** → start at mesh DestinationRule. Move to framework client only if the retry depends on business idempotency knowledge.
239
+ - **"Add a new retry policy for service X"** → apply R4's replay-safety and single-owner checks. Use the matching VirtualService HTTP route for Istio retries; disable and verify route retries when the SDK owns them. DestinationRule limits concurrent retries, not per-request attempts.
238
240
  - **"Service A times out calling Service B"** → check three layers in order: app deadline (ctx timeout) → framework client timeout → mesh request timeout. Whichever is smaller wins; align them.
239
241
  - **"Switch service discovery mode"** → use `references/service-discovery-choice.md`; do not mandate registry unless the platform needs registry-specific capabilities such as per-instance drain, out-of-cluster lookup, or existing registry federation.
240
242
  - **"Need mTLS to a non-mesh external service"** → egress gateway with terminating TLS, not app-managed certs.
@@ -257,7 +259,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
257
259
  - `references/framework-middleware.md` — Server and client middleware chain (HTTP + RPC); the canonical RPC base struct shape; verification commands.
258
260
  - `references/multi-env-routing.md` — Lane label end-to-end recipe; VirtualService patterns; queue-boundary propagation; stress/shadow tags.
259
261
  - `references/retry-timeout-circuit-breaker.md` — Mesh defaults vs SDK overrides; cascading timeout budgets; idempotency awareness; outlier detection tuning.
260
- - `references/grpc-authority-workaround.md` — The `_` → `-` rewrite quirk; when it's needed; the cleaner long-term fix.
262
+ - `references/grpc-authority-workaround.md` — Locate authority rejection; distinguish naming policy from protocol syntax; verify scoped compatibility changes.
261
263
  - `references/dual-sidecar-and-traffic-config-center.md` — Pod-level dual sidecar (mesh + platform), per-protocol mesh injection policy, per-caller-callee traffic config via config center (separate from mesh routing).
262
264
  - `references/rpc-framework-recipe.md` — Concrete kitex/hertz default suite: shared RPC base field contract, RPC base.Request full schema (8 fields incl From/To), dual-channel ctx propagation (metainfo + grpc metadata), 9-code error enum + framework error mapping table, three resolver strategies (registry / FQDN fallback / proxy), platform latency histogram buckets, CORS defaults exposing log-id header, server boot sequence with graceful shutdown.
263
265
  - `references/service-discovery-choice.md` — Decision framework: registry-based vs k8s-native SD; both support lane routing + canary + multi-env; pick by per-instance drain need, laptop access pattern, federation preference, operational burden; mixed mode (k8s east-west + thin registry for laptop) is workable; migration paths in both directions.
@@ -269,7 +271,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
269
271
 
270
272
  1. **Static**: framework default options module includes all R6 middleware; mesh PeerAuthentication is STRICT; DestinationRule and VirtualService exist for every callee that participates in lane routing.
271
273
  2. **Live, identity**: cross-service trace shows log-id + lane consistent across all hops; mesh access log lines include both.
272
- 3. **Live, mesh policy**: kill a callee pod → outlier detection ejects within outlier-detection interval; metrics show client-side error count rise then fall.
274
+ 3. **Live, mesh policy**: trigger the configured outlier threshold with an eligible pool size; verify ejection and recovery against the effective policy and metrics. Consecutive-error checks are inline, not delayed until the periodic analysis interval.
273
275
  4. **Live, lane routing**: send request with non-default lane → confirm only matching-lane instances serve it.
274
276
  5. **Static, no leakage**: grep this skill's content — zero internal hostnames, repo names, or business terms.
275
277
 
@@ -36,7 +36,7 @@ Mesh injection is not all-or-nothing. Real platforms apply per-protocol policy:
36
36
  | App protocol | Inject Istio sidecar? | Reason |
37
37
  |---|---|---|
38
38
  | HTTP (REST) | YES | Envoy handles HTTP/1.1, HTTP/2; full feature support |
39
- | gRPC | YES | HTTP/2 + per-call routing; needs `:authority` quirk handling |
39
+ | gRPC | YES | HTTP/2 + per-call routing; verify `:authority` routing against the actual SDK/proxy path |
40
40
  | TCP (raw) | NO (often) | Envoy TCP proxy is feature-poor; routing/auth less useful at L4 |
41
41
  | Thrift (TTHeader) | NO (often) | Same — L4-ish; tooling limited |
42
42
 
@@ -1,90 +1,47 @@
1
- # gRPC `:authority` and DNS-Label Hyphenation
2
-
3
- ## The quirk
4
-
5
- HTTP/2's `:authority` pseudo-header (which carries the host of a gRPC request) is parsed strictly. RFC 1035 says DNS labels can contain letters, digits, and hyphens — but NOT underscores.
6
-
7
- Many platforms use PSM-style service names like `<owner>.<class>.<env>`. When any segment contains underscores (e.g. an env segment named `prod_v2` or a feature-flagged variant `payments.ledger.prod_2024`), the SDK puts that name into `:authority`. Envoy / strict gRPC implementations reject the request:
8
-
9
- ```
10
- RST_STREAM with INTERNAL_ERROR
11
- ```
12
-
13
- You lose the connection on every call. Frustrating to debug because it looks like network failure.
14
-
15
- ## Two fixes
16
-
17
- ### Fix 1 (clean, long-term): forbid underscores in service names
18
-
19
- - Naming policy: service name = lowercase, dot-separated segments, hyphens within segments allowed, underscores forbidden.
20
- - Enforce at registry registration (reject the registration).
21
- - Enforce at CI / lint when defining new services.
22
- - Migrate existing names by renaming + parallel registration during a deprecation window.
23
-
24
- ### Fix 2 (live-system workaround): mesh-level rewrite
25
-
26
- When you can't break existing names, an Envoy Lua filter rewrites `:authority` on the fly:
27
-
28
- ```yaml
29
- apiVersion: networking.istio.io/v1alpha3
30
- kind: EnvoyFilter
31
- metadata:
32
- name: modify-grpc-authority
33
- namespace: istio-system
34
- spec:
35
- configPatches:
36
- - applyTo: HTTP_FILTER
37
- match:
38
- context: ANY
39
- listener:
40
- filterChain:
41
- filter:
42
- name: "envoy.filters.network.http_connection_manager"
43
- patch:
44
- operation: INSERT_BEFORE
45
- value:
46
- name: envoy.filters.http.lua
47
- typed_config:
48
- "@type": type.googleapis.com/envoy.extensions.filters.http.lua.v3.Lua
49
- inlineCode: |
50
- function envoy_on_request(request_handle)
51
- local authority = request_handle:headers():get(":authority")
52
- local content_type = request_handle:headers():get("content-type")
53
- if authority and content_type and string.find(content_type, "application/grpc") then
54
- local modified_authority = string.gsub(authority, "_", "-")
55
- request_handle:headers():replace(":authority", modified_authority)
56
- end
57
- end
58
- ```
59
-
60
- Effects:
61
- - Applies to gRPC traffic only (content-type check).
62
- - Rewrites `_` to `-` in `:authority`.
63
- - Callee's registry instance name must also use the `-` form so routing matches.
64
-
65
- ## When to use which
66
-
67
- | Situation | Fix |
1
+ # gRPC Authority Compatibility
2
+
3
+ ## Separate the names and constraints
4
+
5
+ HTTP/2 `:authority` conveys the target URI's authority, as defined by [RFC 9113 §8.3.1](https://www.rfc-editor.org/rfc/rfc9113.html#section-8.3.1). [RFC 3986 §3.2.2](https://www.rfc-editor.org/rfc/rfc3986.html#section-3.2.2) permits a registered name containing unreserved characters, including `_`. This syntax does not guarantee DNS resolution, certificate identity matching, or acceptance by every SDK and proxy version.
6
+
7
+ Keep these values distinct when diagnosing a request:
8
+
9
+ - Service-registry identifier and resolved network endpoint.
10
+ - HTTP/2 authority used for virtual-host routing.
11
+ - TLS server name and the certificate identity expected on each TLS hop.
12
+ - Platform naming, routing, and authorization policies.
13
+
14
+ A platform can require DNS-compatible service names and enforce that choice at registration and CI. Existing identifiers do not require migration merely because they contain an underscore; first establish which constraint the actual path violates.
15
+
16
+ ## Locate the rejection before choosing a fix
17
+
18
+ Record the runtime/SDK, proxy versions and relevant configuration, exact authority, and the failing run's error or trace. Use a synthetic payload and redact credentials from captured evidence.
19
+
20
+ | Observed boundary | Next action |
68
21
  |---|---|
69
- | Greenfield platform | Fix 1 — naming policy from day one |
70
- | Mature platform, many existing names | Fix 2 — buy time, then schedule Fix 1 migration |
71
- | Mixed HTTP + gRPC for same service | Fix 2 — HTTP tolerates `_`, gRPC doesn't; rewrite only at gRPC layer |
72
- | Service-mesh-less (direct gRPC) | Fix 1 only — no Envoy to rewrite |
22
+ | Authority syntax is malformed | Validate URI authority syntax, including brackets around an IPv6 literal, before changing service registration or routing. |
23
+ | Client rejects before transmitting HTTP/2 headers | Check that client's authority validation and supported configuration. Fix the client-side mapping or naming contract; a downstream proxy cannot rewrite a request it never receives. |
24
+ | DNS resolution fails | Check the resolved hostname and resolver's naming rules. Changing a later HTTP header does not repair failed resolution. |
25
+ | TLS handshake or certificate identity check fails | Check that hop's endpoint, server name, trust chain, and certificate identities. Retain verification; changing authority is not evidence that TLS is fixed. |
26
+ | Proxy/parser rejects before the HTTP filter runs | Fix the supported parser/input contract or an earlier owned mapping. A Lua filter after the rejection cannot intervene. |
27
+ | Request reaches HTTP filters, then the intended virtual-host route does not match | Compare the received authority with the generated route configuration. A supported authority mapping may be appropriate after proving the mismatch. |
28
+ | Existing path accepts the name and reaches the intended service | Preserve it unless a separate platform naming-policy migration is required. |
73
29
 
74
- ## Verification
30
+ `RST_STREAM` or `INTERNAL_ERROR` alone does not identify an underscore problem. Confirm the first rejecting layer rather than treating every transport failure as the same naming defect.
75
31
 
76
- Fix 1:
77
- - Registry rejects registration with `_` in name.
78
- - Linter / CI rejects PR adding such a name.
32
+ ## Choose a bounded compatibility change
79
33
 
80
- Fix 2:
81
- - Synthetic gRPC call to `<svc>_<env>` → inspect Envoy access log `:authority` field → expect `<svc>-<env>`.
82
- - Callee's access log `host` shows the rewritten form.
34
+ **Naming policy or client mapping.** Where a DNS-compatible name is required, define the allowed form and the mapping from registry identity to endpoint/authority. Check uniqueness before migration: replacing `_` with `-` can collapse two distinct names. Migrate registrations, routes, and callers together with a compatibility window and rollback path. Use supported client options; do not bypass certificate or authorization checks to make a name work.
35
+
36
+ **Proxy mapping.** Use only if the request reaches the chosen filter and the mapping addresses a reproduced failure. Prefer the platform's supported routing mechanism. If an EnvoyFilter is necessary, verify the installed Istio/Envoy API and generated configuration, and scope it to the affected workloads, listener/direction, route, and explicit old-to-new authority mapping. Do not install an all-workload, all-direction underscore replacement. The mapping must preserve the intended destination, tenant/lane routing, authorization, and TLS identity on each hop; rewriting a header does not itself update those contracts.
37
+
38
+ ## Verification
83
39
 
84
- ## Common variants of the same bug
40
+ Retain the original failing case and expected rejecting layer. After the change, verify:
85
41
 
86
- - HTTP/2 SETTINGS frame rejection on cert SAN mismatch (separate issue, but presents similarly — RST_STREAM with INTERNAL_ERROR).
87
- - gRPC `:scheme` set to `http` while upstream expects `https` (mTLS mismatch).
88
- - IPv6 literal in `:authority` not bracketed.
42
+ 1. The same request succeeds through the intended path and reaches the intended service; observe authority before/after any mapping and the selected route.
43
+ 2. Unrelated authorities and non-target traffic retain their behavior. Include potentially colliding names and unknown authorities as negative controls.
44
+ 3. Certificate identity and authorization failures still reject requests; the compatibility change has not disabled those checks.
45
+ 4. Registration/client/route changes can be rolled back together without sending traffic to another service.
89
46
 
90
- The Lua rewrite filter is the right tool for the underscore case; do not generalize it to all of the above. Each bug has its own fix.
47
+ Static configuration validation proves only configuration properties. Claims about a deployed SDK/proxy path require execution evidence from that path.
@@ -33,7 +33,7 @@ Reference topology and component rules. Istio is the recurring example; rules ge
33
33
 
34
34
  ## Istio Ambient Mode (GA Nov 2024, v1.24+)
35
35
 
36
- - **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block: AuthorizationPolicy enforcement loss when flipping a namespace from sidecar to ambient without waypoint** — sidecar mode enforces full L7 `AuthorizationPolicy` rules (HTTP method, path, header, JWT claims) at the per-pod proxy. In ambient mode, ztunnel alone enforces only L4-level policy (workload identity, port); any L7-condition rule (`request.method`, `request.headers[...]`, `request.path`, `request.auth.claims[...]`) silently DOES NOT MATCH on ambient-mode pods until a waypoint proxy is deployed for the namespace/service and the policy is targeted at the waypoint. The failure shape: team flips `istio.io/dataplane-mode=ambient` label, sidecar comes down, ztunnel takes over L4, and every L7 authz policy stops gating traffic — without a visible error, because L4 policy still works. **Pre-flip gate**: inventory every `AuthorizationPolicy` targeting workloads in the namespace, identify the ones using L7 conditions, ensure a waypoint is deployed and the policy `targetRefs` is set to the waypoint (not the workload) BEFORE flipping the namespace label. Validate by re-running the authz integration tests against the ambient-mode namespace before completing the migration. **Observability shift**: ztunnel emits L4 metrics; waypoint emits L7 metrics; existing dashboards keyed on sidecar `istio_requests_total` need a ambient-mode equivalent for the L7-routed traffic — route this update through `platform-observability` rather than redesigning here.
36
+ - **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block: L7 policy enforcement changes when moving from sidecar to ambient.** Ztunnel enforces only L4 policy. A workload-selector policy containing L7 conditions that is picked up by ztunnel [fails safe as a DENY policy](https://istio.io/latest/docs/ambient/usage/l4-policy/#policies-with-layer-7-conditions), potentially blocking legitimate traffic. L7 enforcement requires a waypoint and `targetRefs` bound to the intended Service or waypoint Gateway; deploying a waypoint alone does not migrate selector-based policies. **Pre-flip gate**: inventory affected authorization policies, prepare the waypoint and correctly scoped bindings, and plan the old-policy/pod-restart transition using the deployed version's [migration guide](https://istio.io/latest/docs/ambient/migrate/migrate-policies/). Preserve mTLS and L4 authorization; do not remove a rejecting policy merely to restore connectivity. Verify allowed, denied, and waypoint-bypass paths before completing migration. If the transition cannot maintain required L7 protection, hold migration or use an approved maintenance window. **Observability shift**: ztunnel emits L4 metrics; waypoint emits L7 metrics; existing dashboards keyed on sidecar `istio_requests_total` need a ambient-mode equivalent for the L7-routed traffic — route this update through `platform-observability` rather than redesigning here.
37
37
 
38
38
  ## Sidecar injection rules
39
39
 
@@ -94,7 +94,7 @@ Egress gateway is OFF by default. Turn ON only when a workload needs controlled
94
94
  EnvoyFilter is the escape hatch. Use sparingly; each filter is hard to test and easy to break across Envoy upgrades.
95
95
 
96
96
  Acceptable uses:
97
- - Header rewrite for protocol quirks (e.g. gRPC `:authority` hyphenation — see `grpc-authority-workaround.md`).
97
+ - Header mapping for a reproduced compatibility failure, scoped to the affected workloads and explicit names while preserving routing, TLS, and authorization (see `grpc-authority-workaround.md`).
98
98
  - Adding a Lua filter for one-off business handling that doesn't yet have a first-class WASM filter.
99
99
  - Custom rate-limit before the canonical RLS is rolled out.
100
100