@ccoalm/ccl-skills 0.14.0 → 0.15.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +80 -4
  2. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +16 -4
  3. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +13 -2
  4. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +53 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +19 -24
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +32 -32
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +16 -14
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +24 -26
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +11 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +60 -209
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +114 -367
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +52 -672
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +17 -3
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +4 -2
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +77 -444
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +33 -98
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +57 -173
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +65 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +1 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +1 -1
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +5 -5
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +8 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +1 -1
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +14 -14
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +1 -1
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +40 -83
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +2 -2
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +44 -37
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +1 -1
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +2 -2
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +2 -2
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +1 -1
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +1 -1
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +9 -6
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +6 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +4 -4
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +52 -0
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +31 -7
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +62 -3
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +188 -13
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/review_ledger_binding.py +36 -13
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +103 -21
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +261 -14
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +2 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_resolution.sh +253 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_runtime.py +428 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_verdict_differential.sh +49 -25
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_review_ledger_binding.sh +63 -5
  51. package/dist/assets/release.json +76 -56
  52. package/dist/claude-adapter.js +14 -7
  53. package/dist/codex-host.d.ts +1 -3
  54. package/dist/codex-host.js +6 -9
  55. package/dist/host-probe.d.ts +11 -0
  56. package/dist/host-probe.js +27 -0
  57. package/dist/opencode-adapter.js +24 -19
  58. package/dist/unified.js +11 -9
  59. package/package.json +1 -1
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: platform-observability
3
- description: Use when designing, reviewing, debugging, or shipping observability — service logs, metrics, distributed tracing, log/trace correlation, dashboards, alerts, on-call routing(值班/排班 SOP、P0/P1 打断;发布值班/回滚除外), SLI/SLO, error budgets — for a backend product. Owns the cross-cutting evidence layer — what signals must exist, what fields must propagate, what middleware must auto-wire, what verification proves a change is observable in production. Hand off mesh/routing/mTLS to `platform-service-connectivity`, release gates and rollback evidence to `platform-release-engineering`, service-internal architecture (HTTP/RPC/DB/queue) to `python-service-architecture` / `go-microservice-architecture`. Product-agnostic; do not depend on specific repository names, service names, hostnames, or business domains.
3
+ description: Use when designing, reviewing, debugging, or shipping observability — service logs, metrics, distributed tracing, log/trace correlation, 给接口·服务加结构化日志与 trace 透传, dashboards, alerts, on-call routing(值班/排班 SOP、P0/P1 打断;发布值班/回滚除外), SLI/SLO, error budgets, 线上是否已启用(开关·配置在生产的实际生效状态)— for a backend product. Owns the cross-cutting evidence layer — what signals must exist, what fields must propagate, what middleware must auto-wire, what verification proves a change is observable in production. Hand off mesh/routing/mTLS to `platform-service-connectivity`, release gates and rollback evidence to `platform-release-engineering`, service-internal architecture (HTTP/RPC/DB/queue) to `python-service-architecture` / `go-microservice-architecture`. Product-agnostic.
4
4
  ---
5
5
 
6
6
  # Platform Observability
@@ -30,7 +30,7 @@ Observability is the chain that turns one user action into searchable, joinable,
30
30
  2. **Instrumentation** — every service emits structured logs, metrics, and spans through the **framework default**, not ad-hoc code. A service whose middleware chain does not include `ctx_inject + metrics + recovery + tracing` is unobservable by design.
31
31
  3. **Transport** — logs go stdout → file collector (DaemonSet) → log pipeline → search index; metrics + traces go SDK → OTLP collector → metric store + trace store. Both transports must survive collector restarts and back-pressure.
32
32
  4. **Storage + display** — logs in a searchable index keyed by log-id and trace-id; metrics in a long-term store separate from scraper; traces queryable by trace-id; one dashboard tool joins all three.
33
- 5. **Evidence consumption** — SLIs are queries against (3) + (4); alerts evaluate SLIs and route to on-call; runbooks live in a wiki that on-call can reach in under one minute.
33
+ 5. **Evidence consumption** — SLIs query (3) + (4); actionable alerts route to on-call; runbooks live in a wiki that on-call can reach in under one minute.
34
34
 
35
35
  A new service must satisfy all five before it is allowed in production. A release must produce evidence at all five before it is promoted.
36
36
 
@@ -106,9 +106,9 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
106
106
 
107
107
  **Retrofit / migration contract** (the rules above are otherwise greenfield-framed): when a schema is introduced over EXISTING services, field names, metric label names, and identity env-var names are a **migration contract** — a blind rename breaks every deployed dashboard, alert, and saved query that keys on the old name. Retrofitting MUST alias or dual-write old→new and migrate consumers before retiring the old name; never rename in place. The cheap time to fix a name is before services adopt it.
108
108
 
109
- ### R7 — Alerts are SLI-driven and route to a human
109
+ ### R7 — Alerts are actionable and route to an owner
110
110
 
111
- - Alerts evaluate against SLIs (query the long-term metric store), not raw metrics.
111
+ - Use SLIs for SLO paging; actionable capacity or impending-failure alerts may use internal measurements. See `references/alerting-and-on-call.md`.
112
112
  - Severity levels: P0 (page on-call now), P1 (notify channel, ack within work hours), P2 (digest).
113
113
  - Every alert MUST link to a runbook entry. If no runbook exists, the alert is not allowed to be P0.
114
114
  - An alert-backed metric is a coverage contract, and the trigger is mechanical, not prose: any new or changed alert rule, SLO, dashboard alert annotation, or metric referenced by an alert policy fires this check (a written "alert on any increase" note also counts, but its absence is not an exemption). Enumerate every site that should feed the metric and verify each is actually instrumented — derive the site list from a static registry or lint where possible; the easiest site to miss is often the riskiest (e.g. the panic counter in a stream reader parsing untrusted bytes). Ship the site checklist with the alert.
@@ -208,7 +208,7 @@ If any phase has missing evidence, the work is not done.
208
208
  - **"Sample everything vs head-sample 1%"** → head-sample low (1–10%) for cost; retaining errors/slow requests is **tail** sampling at the Collector and requires near-full SDK export to it (head-dropped spans never arrive) — pick one model per service, do not claim both. Sampling-by-route is acceptable for known noisy paths.
209
209
  - **"Metrics egress: scrape vs push"** → each platform picks ONE canonical egress and every process uses it. A Prometheus-native platform may default to **scrape** (pull, with `up`/target-health + service discovery); an OTel-first platform, or workers/jobs/runtimes that can't be scraped, may default to **OTLP push** to a collector. Neither is universally better — don't overturn a working pull setup just to push. One egress per process for a given metric; a migration window may dual-write only with isolated pipelines / distinct metric names / a dedup plan. Long-term store stays separate from the short-term scrape/collect layer (R5).
210
210
  - **"Add a new label to a metric"** → answer the cardinality question first. If max distinct values × series count > 1e6, refuse and use an exemplar trace instead.
211
- - **"Alert on this symptom or that cause"** → alert on user-visible symptom; cause-based alerts produce paging spam.
211
+ - **"Alert on this symptom or that cause"** → prefer user-visible symptoms; capacity or impending-failure warnings follow R7.
212
212
 
213
213
  ## Sanitization and Provenance
214
214
 
@@ -1,5 +1,13 @@
1
1
  # Alerting and On-Call
2
2
 
3
+ ## Choose an actionable signal
4
+
5
+ - **User symptoms and SLOs:** use service-level signals for user-symptom paging and the corresponding SLIs for SLO burn-rate alerts. Query the declared metric store; diagnostic counters do not become availability or latency SLIs merely because an alert uses them.
6
+ - **Capacity and impending failures:** white-box measurements may warn before user SLIs degrade, such as a credible forecast of disk exhaustion. Record the expected failure, the evidence for the threshold or forecast, the action, and the responsible owner. Choose urgency from impact and time left to act; a high internal metric alone is not a reason to page.
7
+ - **Diagnostic signals:** keep measurements with no actionable condition in dashboards or queries. Do not alert on every possible cause.
8
+
9
+ Verify new or changed alert rules against relevant failure, healthy, and recovery cases. Preserve the severity, ownership, runbook, and delivery requirements below for both symptom and impending-failure alerts. [Google SRE's Monitoring Distributed Systems](https://sre.google/sre-book/monitoring-distributed-systems/) describes both user symptoms and imminent saturation as alerting inputs.
10
+
3
11
  ## Two patterns
4
12
 
5
13
  ### Pattern A — Prometheus AlertManager
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: platform-release-engineering
3
- description: 发布 / 灰度 / canary / rollback / rollout / 环境泳道 / promotion gate → design or review how a change moves from build to traffic and back safely, including rollout strategy, approval, rollback, secrets, config, and deploy control planes. Skip when the ask is the production release lifecycle — 上线范围确认 / 合并 main / 打 tag / 生产构建 / 发布后 reset → release-coordination; release document substance → release-doc-writer.
3
+ description: 发布 / 灰度 / canary / rollback / rollout / 环境泳道 / promotion gate / 发布值班 SOP(P0·P1 打断排班、是否回滚)→ design or review how a change moves from build to traffic and back safely, including rollout strategy, approval, rollback, secrets, config, and deploy control planes. Skip when the ask is the production release lifecycle — 上线范围确认 / 合并 main / 打 tag / 生产构建 / 发布后 reset → release-coordination; release document substance → release-doc-writer.
4
4
  ---
5
5
 
6
6
  # Platform Release Engineering
@@ -100,10 +100,11 @@ Confusing these is the source of most "why doesn't this work" connectivity bugs.
100
100
 
101
101
  ### R4 — Default retry/timeout/circuit-breaker live at mesh, app overrides for business reasons
102
102
 
103
- - Mesh DestinationRule provides default per-callee policy: retries (1-2 attempts on 5xx/connect-fail), connect timeout, request timeout ceiling, outlier detection (5xx-percent → eject).
104
- - Framework client SDK provides per-call override: business timeout (always ≤ mesh ceiling), retry policy for idempotent calls, hedging.
103
+ - On Istio HTTP/gRPC paths, `VirtualService` HTTP routes own request `timeout` and `retries`; `DestinationRule` owns connection-pool settings and `outlierDetection`. Other transports use their platform-owned equivalents.
104
+ - Before configuring retries or timers, read `references/retry-timeout-circuit-breaker.md` for field paths, single-owner retry selection, and idempotency checks. Both mesh and SDK follow them; a 5xx or missing response alone never proves replay safe.
105
+ - Framework client SDK may provide method-specific retry/hedging; its total call budget must fit the caller's remaining duration and any platform cap. A longer mesh timeout is a backstop, not an extended caller deadline.
105
106
  - App code MAY override per-RPC. App code MUST NOT silently disable mesh-level outlier detection.
106
- - Budget rule: total upstream timeout = caller deadline minus a safety margin (e.g. 100ms). Cascading timeouts must shrink down the call chain.
107
+ - Budget rule: downstream work, attempts, and backoff fit the remaining duration minus a safety margin (e.g. 100ms). Propagate deadline and cancellation; never reset the full budget at each hop.
107
108
 
108
109
  ### R5 — mTLS is mesh-default, app cannot disable
109
110
 
@@ -151,14 +152,13 @@ The client middleware fills this from ctx automatically; the server middleware e
151
152
 
152
153
  For non-protobuf or metadata-only transports, an equivalent header set is compliant only when it cites a resolvable platform owner record, such as a repo/path, document URL, registry id, gateway policy, or owner-suite id. The owner record must enumerate the concrete header names. In diff-only review without resolver tooling, the diff must cite a stable owner-record locator and list the concrete header names inline; with resolver tooling, the reviewer or owner-suite may resolve the locator to those names instead. The enumerated header names must then be checked against the actual propagated and exposed headers: caller-supplied identity headers are absent unless positive authenticated-caller evidence exists. An unresolvable, non-enumerating, uncited, service-local, or unchecked header set is an open gap, not a permitted alternative. For pure HTTP (no RPC base), the equivalent is a stable owner-recorded header set, also filled by middleware.
153
154
 
154
- ### R8 — gRPC `:authority` and DNS-label hyphenation
155
+ ### R8 — gRPC `:authority` compatibility follows the actual request path
155
156
 
156
- If the platform allows service names with underscores (`<owner>.<class>.<env>` containing `_`), gRPC will reject them in the HTTP/2 `:authority` pseudo-header because it must be a valid DNS label (no `_`). Two acceptable fixes:
157
+ HTTP/2 `:authority` carries URI authority, not a single DNS label. An underscore alone does not prove a gRPC protocol violation. DNS hostnames, certificate identity checks, SDK validation, and proxy routing can impose different constraints; preserve the constraints of the deployed path.
157
158
 
158
- 1. **Disallow underscores in new service names**; enforce in registry registration and CI.
159
- 2. **Mesh-level rewrite**: an EnvoyFilter Lua snippet replaces `_` with `-` in `:authority` for gRPC requests, before routing.
159
+ Before renaming a service or adding a rewrite, capture the failing request's authority, the rejecting layer and version, and the relevant error or trace. A platform may require DNS-compatible service names and enforce that policy at registration and CI, but a registry name need not be the wire authority.
160
160
 
161
- Option 1 is cleaner long-term; option 2 is the live-system workaround. Document which the platform uses; new services should follow option 1.
161
+ A proxy rewrite is an option only when the request reaches that filter before the rejecting layer. It cannot fix a client rejection before transmission or an inbound parser rejection before the filter runs. Scope any verified rewrite to the affected traffic and explicit authority mappings; retain route, TLS identity, and authorization checks. Diagnosis, migration, and verification: `references/grpc-authority-workaround.md`.
162
162
 
163
163
  ### R9 — Ingress and egress are explicit, not implicit
164
164
 
@@ -200,8 +200,8 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
200
200
  ### Phase B — Mesh policy
201
201
 
202
202
  1. PeerAuthentication = STRICT (mTLS namespace-wide).
203
- 2. DestinationRule per critical callee: retry policy, timeouts, outlier detection.
204
- 3. VirtualService rules express lane-based routing: lane header match → lane subset.
203
+ 2. DestinationRule per critical callee: connection-pool limits, connect timeout, outlier detection.
204
+ 3. VirtualService HTTP routes: lane header match → lane subset, request timeout, and the R4 retry policy.
205
205
  4. AuthorizationPolicy expresses which services may call which (zero-trust at network level).
206
206
 
207
207
  ### Phase C — Service discovery
@@ -216,7 +216,7 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
216
216
 
217
217
  1. Read the framework default client/server options module. Confirm middleware chain matches R6.
218
218
  2. Confirm dev cannot build a client/server without inheriting these.
219
- 3. Test: kill a downstream pod → verify mesh outlier-detection ejects, framework retry kicks in for idempotent calls, error propagates up with stable error-code.
219
+ 3. Test: induce the configured outlier threshold on one callee → verify ejection and recovery, the configured retry layer retries only eligible calls within its budget, and exhausted calls propagate a stable error-code.
220
220
 
221
221
  ### Phase E — Failure modes
222
222
 
@@ -234,7 +234,7 @@ Before marking work done:
234
234
 
235
235
  ## Decision Points
236
236
 
237
- - **"Add a new retry policy for service X"** → start at mesh DestinationRule. Move to framework client only if the retry depends on business idempotency knowledge.
237
+ - **"Add a new retry policy for service X"** → apply R4's replay-safety and single-owner checks. Use the matching VirtualService HTTP route for Istio retries; disable and verify route retries when the SDK owns them. DestinationRule limits concurrent retries, not per-request attempts.
238
238
  - **"Service A times out calling Service B"** → check three layers in order: app deadline (ctx timeout) → framework client timeout → mesh request timeout. Whichever is smaller wins; align them.
239
239
  - **"Switch service discovery mode"** → use `references/service-discovery-choice.md`; do not mandate registry unless the platform needs registry-specific capabilities such as per-instance drain, out-of-cluster lookup, or existing registry federation.
240
240
  - **"Need mTLS to a non-mesh external service"** → egress gateway with terminating TLS, not app-managed certs.
@@ -257,7 +257,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
257
257
  - `references/framework-middleware.md` — Server and client middleware chain (HTTP + RPC); the canonical RPC base struct shape; verification commands.
258
258
  - `references/multi-env-routing.md` — Lane label end-to-end recipe; VirtualService patterns; queue-boundary propagation; stress/shadow tags.
259
259
  - `references/retry-timeout-circuit-breaker.md` — Mesh defaults vs SDK overrides; cascading timeout budgets; idempotency awareness; outlier detection tuning.
260
- - `references/grpc-authority-workaround.md` — The `_` → `-` rewrite quirk; when it's needed; the cleaner long-term fix.
260
+ - `references/grpc-authority-workaround.md` — Locate authority rejection; distinguish naming policy from protocol syntax; verify scoped compatibility changes.
261
261
  - `references/dual-sidecar-and-traffic-config-center.md` — Pod-level dual sidecar (mesh + platform), per-protocol mesh injection policy, per-caller-callee traffic config via config center (separate from mesh routing).
262
262
  - `references/rpc-framework-recipe.md` — Concrete kitex/hertz default suite: shared RPC base field contract, RPC base.Request full schema (8 fields incl From/To), dual-channel ctx propagation (metainfo + grpc metadata), 9-code error enum + framework error mapping table, three resolver strategies (registry / FQDN fallback / proxy), platform latency histogram buckets, CORS defaults exposing log-id header, server boot sequence with graceful shutdown.
263
263
  - `references/service-discovery-choice.md` — Decision framework: registry-based vs k8s-native SD; both support lane routing + canary + multi-env; pick by per-instance drain need, laptop access pattern, federation preference, operational burden; mixed mode (k8s east-west + thin registry for laptop) is workable; migration paths in both directions.
@@ -269,7 +269,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
269
269
 
270
270
  1. **Static**: framework default options module includes all R6 middleware; mesh PeerAuthentication is STRICT; DestinationRule and VirtualService exist for every callee that participates in lane routing.
271
271
  2. **Live, identity**: cross-service trace shows log-id + lane consistent across all hops; mesh access log lines include both.
272
- 3. **Live, mesh policy**: kill a callee pod → outlier detection ejects within outlier-detection interval; metrics show client-side error count rise then fall.
272
+ 3. **Live, mesh policy**: trigger the configured outlier threshold with an eligible pool size; verify ejection and recovery against the effective policy and metrics. Consecutive-error checks are inline, not delayed until the periodic analysis interval.
273
273
  4. **Live, lane routing**: send request with non-default lane → confirm only matching-lane instances serve it.
274
274
  5. **Static, no leakage**: grep this skill's content — zero internal hostnames, repo names, or business terms.
275
275
 
@@ -36,7 +36,7 @@ Mesh injection is not all-or-nothing. Real platforms apply per-protocol policy:
36
36
  | App protocol | Inject Istio sidecar? | Reason |
37
37
  |---|---|---|
38
38
  | HTTP (REST) | YES | Envoy handles HTTP/1.1, HTTP/2; full feature support |
39
- | gRPC | YES | HTTP/2 + per-call routing; needs `:authority` quirk handling |
39
+ | gRPC | YES | HTTP/2 + per-call routing; verify `:authority` routing against the actual SDK/proxy path |
40
40
  | TCP (raw) | NO (often) | Envoy TCP proxy is feature-poor; routing/auth less useful at L4 |
41
41
  | Thrift (TTHeader) | NO (often) | Same — L4-ish; tooling limited |
42
42
 
@@ -1,90 +1,47 @@
1
- # gRPC `:authority` and DNS-Label Hyphenation
2
-
3
- ## The quirk
4
-
5
- HTTP/2's `:authority` pseudo-header (which carries the host of a gRPC request) is parsed strictly. RFC 1035 says DNS labels can contain letters, digits, and hyphens — but NOT underscores.
6
-
7
- Many platforms use PSM-style service names like `<owner>.<class>.<env>`. When any segment contains underscores (e.g. an env segment named `prod_v2` or a feature-flagged variant `payments.ledger.prod_2024`), the SDK puts that name into `:authority`. Envoy / strict gRPC implementations reject the request:
8
-
9
- ```
10
- RST_STREAM with INTERNAL_ERROR
11
- ```
12
-
13
- You lose the connection on every call. Frustrating to debug because it looks like network failure.
14
-
15
- ## Two fixes
16
-
17
- ### Fix 1 (clean, long-term): forbid underscores in service names
18
-
19
- - Naming policy: service name = lowercase, dot-separated segments, hyphens within segments allowed, underscores forbidden.
20
- - Enforce at registry registration (reject the registration).
21
- - Enforce at CI / lint when defining new services.
22
- - Migrate existing names by renaming + parallel registration during a deprecation window.
23
-
24
- ### Fix 2 (live-system workaround): mesh-level rewrite
25
-
26
- When you can't break existing names, an Envoy Lua filter rewrites `:authority` on the fly:
27
-
28
- ```yaml
29
- apiVersion: networking.istio.io/v1alpha3
30
- kind: EnvoyFilter
31
- metadata:
32
- name: modify-grpc-authority
33
- namespace: istio-system
34
- spec:
35
- configPatches:
36
- - applyTo: HTTP_FILTER
37
- match:
38
- context: ANY
39
- listener:
40
- filterChain:
41
- filter:
42
- name: "envoy.filters.network.http_connection_manager"
43
- patch:
44
- operation: INSERT_BEFORE
45
- value:
46
- name: envoy.filters.http.lua
47
- typed_config:
48
- "@type": type.googleapis.com/envoy.extensions.filters.http.lua.v3.Lua
49
- inlineCode: |
50
- function envoy_on_request(request_handle)
51
- local authority = request_handle:headers():get(":authority")
52
- local content_type = request_handle:headers():get("content-type")
53
- if authority and content_type and string.find(content_type, "application/grpc") then
54
- local modified_authority = string.gsub(authority, "_", "-")
55
- request_handle:headers():replace(":authority", modified_authority)
56
- end
57
- end
58
- ```
59
-
60
- Effects:
61
- - Applies to gRPC traffic only (content-type check).
62
- - Rewrites `_` to `-` in `:authority`.
63
- - Callee's registry instance name must also use the `-` form so routing matches.
64
-
65
- ## When to use which
66
-
67
- | Situation | Fix |
1
+ # gRPC Authority Compatibility
2
+
3
+ ## Separate the names and constraints
4
+
5
+ HTTP/2 `:authority` conveys the target URI's authority, as defined by [RFC 9113 §8.3.1](https://www.rfc-editor.org/rfc/rfc9113.html#section-8.3.1). [RFC 3986 §3.2.2](https://www.rfc-editor.org/rfc/rfc3986.html#section-3.2.2) permits a registered name containing unreserved characters, including `_`. This syntax does not guarantee DNS resolution, certificate identity matching, or acceptance by every SDK and proxy version.
6
+
7
+ Keep these values distinct when diagnosing a request:
8
+
9
+ - Service-registry identifier and resolved network endpoint.
10
+ - HTTP/2 authority used for virtual-host routing.
11
+ - TLS server name and the certificate identity expected on each TLS hop.
12
+ - Platform naming, routing, and authorization policies.
13
+
14
+ A platform can require DNS-compatible service names and enforce that choice at registration and CI. Existing identifiers do not require migration merely because they contain an underscore; first establish which constraint the actual path violates.
15
+
16
+ ## Locate the rejection before choosing a fix
17
+
18
+ Record the runtime/SDK, proxy versions and relevant configuration, exact authority, and the failing run's error or trace. Use a synthetic payload and redact credentials from captured evidence.
19
+
20
+ | Observed boundary | Next action |
68
21
  |---|---|
69
- | Greenfield platform | Fix 1 — naming policy from day one |
70
- | Mature platform, many existing names | Fix 2 — buy time, then schedule Fix 1 migration |
71
- | Mixed HTTP + gRPC for same service | Fix 2 — HTTP tolerates `_`, gRPC doesn't; rewrite only at gRPC layer |
72
- | Service-mesh-less (direct gRPC) | Fix 1 only — no Envoy to rewrite |
22
+ | Authority syntax is malformed | Validate URI authority syntax, including brackets around an IPv6 literal, before changing service registration or routing. |
23
+ | Client rejects before transmitting HTTP/2 headers | Check that client's authority validation and supported configuration. Fix the client-side mapping or naming contract; a downstream proxy cannot rewrite a request it never receives. |
24
+ | DNS resolution fails | Check the resolved hostname and resolver's naming rules. Changing a later HTTP header does not repair failed resolution. |
25
+ | TLS handshake or certificate identity check fails | Check that hop's endpoint, server name, trust chain, and certificate identities. Retain verification; changing authority is not evidence that TLS is fixed. |
26
+ | Proxy/parser rejects before the HTTP filter runs | Fix the supported parser/input contract or an earlier owned mapping. A Lua filter after the rejection cannot intervene. |
27
+ | Request reaches HTTP filters, then the intended virtual-host route does not match | Compare the received authority with the generated route configuration. A supported authority mapping may be appropriate after proving the mismatch. |
28
+ | Existing path accepts the name and reaches the intended service | Preserve it unless a separate platform naming-policy migration is required. |
73
29
 
74
- ## Verification
30
+ `RST_STREAM` or `INTERNAL_ERROR` alone does not identify an underscore problem. Confirm the first rejecting layer rather than treating every transport failure as the same naming defect.
75
31
 
76
- Fix 1:
77
- - Registry rejects registration with `_` in name.
78
- - Linter / CI rejects PR adding such a name.
32
+ ## Choose a bounded compatibility change
79
33
 
80
- Fix 2:
81
- - Synthetic gRPC call to `<svc>_<env>` → inspect Envoy access log `:authority` field → expect `<svc>-<env>`.
82
- - Callee's access log `host` shows the rewritten form.
34
+ **Naming policy or client mapping.** Where a DNS-compatible name is required, define the allowed form and the mapping from registry identity to endpoint/authority. Check uniqueness before migration: replacing `_` with `-` can collapse two distinct names. Migrate registrations, routes, and callers together with a compatibility window and rollback path. Use supported client options; do not bypass certificate or authorization checks to make a name work.
35
+
36
+ **Proxy mapping.** Use only if the request reaches the chosen filter and the mapping addresses a reproduced failure. Prefer the platform's supported routing mechanism. If an EnvoyFilter is necessary, verify the installed Istio/Envoy API and generated configuration, and scope it to the affected workloads, listener/direction, route, and explicit old-to-new authority mapping. Do not install an all-workload, all-direction underscore replacement. The mapping must preserve the intended destination, tenant/lane routing, authorization, and TLS identity on each hop; rewriting a header does not itself update those contracts.
37
+
38
+ ## Verification
83
39
 
84
- ## Common variants of the same bug
40
+ Retain the original failing case and expected rejecting layer. After the change, verify:
85
41
 
86
- - HTTP/2 SETTINGS frame rejection on cert SAN mismatch (separate issue, but presents similarly — RST_STREAM with INTERNAL_ERROR).
87
- - gRPC `:scheme` set to `http` while upstream expects `https` (mTLS mismatch).
88
- - IPv6 literal in `:authority` not bracketed.
42
+ 1. The same request succeeds through the intended path and reaches the intended service; observe authority before/after any mapping and the selected route.
43
+ 2. Unrelated authorities and non-target traffic retain their behavior. Include potentially colliding names and unknown authorities as negative controls.
44
+ 3. Certificate identity and authorization failures still reject requests; the compatibility change has not disabled those checks.
45
+ 4. Registration/client/route changes can be rolled back together without sending traffic to another service.
89
46
 
90
- The Lua rewrite filter is the right tool for the underscore case; do not generalize it to all of the above. Each bug has its own fix.
47
+ Static configuration validation proves only configuration properties. Claims about a deployed SDK/proxy path require execution evidence from that path.
@@ -33,7 +33,7 @@ Reference topology and component rules. Istio is the recurring example; rules ge
33
33
 
34
34
  ## Istio Ambient Mode (GA Nov 2024, v1.24+)
35
35
 
36
- - **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block: AuthorizationPolicy enforcement loss when flipping a namespace from sidecar to ambient without waypoint** — sidecar mode enforces full L7 `AuthorizationPolicy` rules (HTTP method, path, header, JWT claims) at the per-pod proxy. In ambient mode, ztunnel alone enforces only L4-level policy (workload identity, port); any L7-condition rule (`request.method`, `request.headers[...]`, `request.path`, `request.auth.claims[...]`) silently DOES NOT MATCH on ambient-mode pods until a waypoint proxy is deployed for the namespace/service and the policy is targeted at the waypoint. The failure shape: team flips `istio.io/dataplane-mode=ambient` label, sidecar comes down, ztunnel takes over L4, and every L7 authz policy stops gating traffic — without a visible error, because L4 policy still works. **Pre-flip gate**: inventory every `AuthorizationPolicy` targeting workloads in the namespace, identify the ones using L7 conditions, ensure a waypoint is deployed and the policy `targetRefs` is set to the waypoint (not the workload) BEFORE flipping the namespace label. Validate by re-running the authz integration tests against the ambient-mode namespace before completing the migration. **Observability shift**: ztunnel emits L4 metrics; waypoint emits L7 metrics; existing dashboards keyed on sidecar `istio_requests_total` need a ambient-mode equivalent for the L7-routed traffic — route this update through `platform-observability` rather than redesigning here.
36
+ - **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block: L7 policy enforcement changes when moving from sidecar to ambient.** Ztunnel enforces only L4 policy. A workload-selector policy containing L7 conditions that is picked up by ztunnel [fails safe as a DENY policy](https://istio.io/latest/docs/ambient/usage/l4-policy/#policies-with-layer-7-conditions), potentially blocking legitimate traffic. L7 enforcement requires a waypoint and `targetRefs` bound to the intended Service or waypoint Gateway; deploying a waypoint alone does not migrate selector-based policies. **Pre-flip gate**: inventory affected authorization policies, prepare the waypoint and correctly scoped bindings, and plan the old-policy/pod-restart transition using the deployed version's [migration guide](https://istio.io/latest/docs/ambient/migrate/migrate-policies/). Preserve mTLS and L4 authorization; do not remove a rejecting policy merely to restore connectivity. Verify allowed, denied, and waypoint-bypass paths before completing migration. If the transition cannot maintain required L7 protection, hold migration or use an approved maintenance window. **Observability shift**: ztunnel emits L4 metrics; waypoint emits L7 metrics; existing dashboards keyed on sidecar `istio_requests_total` need a ambient-mode equivalent for the L7-routed traffic — route this update through `platform-observability` rather than redesigning here.
37
37
 
38
38
  ## Sidecar injection rules
39
39
 
@@ -94,7 +94,7 @@ Egress gateway is OFF by default. Turn ON only when a workload needs controlled
94
94
  EnvoyFilter is the escape hatch. Use sparingly; each filter is hard to test and easy to break across Envoy upgrades.
95
95
 
96
96
  Acceptable uses:
97
- - Header rewrite for protocol quirks (e.g. gRPC `:authority` hyphenation — see `grpc-authority-workaround.md`).
97
+ - Header mapping for a reproduced compatibility failure, scoped to the affected workloads and explicit names while preserving routing, TLS, and authorization (see `grpc-authority-workaround.md`).
98
98
  - Adding a Lua filter for one-off business handling that doesn't yet have a first-class WASM filter.
99
99
  - Custom rate-limit before the canonical RLS is rolled out.
100
100
 
@@ -1,48 +1,48 @@
1
1
  # Retry, Timeout, Circuit Breaker
2
2
 
3
- Where each policy lives, how they compose, and how to keep them from killing each other.
3
+ Where each policy lives, how timeouts compose, and how to bound retry amplification. Istio field names below apply to Istio HTTP/gRPC paths; other transports keep their platform-owned equivalents.
4
4
 
5
5
  ## Layered policy
6
6
 
7
- For a single call A → B, four policies can fire:
7
+ For a single call A → B, several independent timers can end the call:
8
8
 
9
- ```
10
- App business deadline ← ctx.Deadline; "this user can wait at most 800ms"
11
- ⊇
12
- Framework client per-call budget ← SDK option; "give this call 500ms"
13
- ⊇
14
- Mesh transport timeout (ceiling) ← DestinationRule; "any call to B caps at 2s"
15
- ⊇
16
- TCP/HTTP2 connect + idle timeouts ← Envoy defaults; rarely tuned
17
- ```
9
+ | Timer | Owner | Example |
10
+ |---|---|---|
11
+ | Caller deadline | App context | 800ms remaining |
12
+ | Total client-call budget | Framework client | 500ms, including retries and backoff |
13
+ | Mesh request timeout | VirtualService HTTP route `timeout` | 2s platform cap |
14
+ | Per-attempt timeout | VirtualService HTTP route `retries.perTryTimeout`, or the retry-owning SDK | Must fit the remaining call budget |
15
+ | Connection / idle timeout | DestinationRule connection pool / transport | Bounds connection setup or inactivity, not total business work |
18
16
 
19
- Each outer policy must be ≥ the next inner one. Inversions cause "request succeeded internally but caller already returned timeout" — a hard-to-debug class.
17
+ The earliest applicable expiry wins. With the example values measured from call start, the client ends the call at 500ms; the 2s mesh cap can remain a platform backstop. These are not nested durations that must decrease from client to mesh. Propagate cancellation so downstream work stops when the caller's budget expires; a longer transport cap does not extend the caller's deadline.
18
+
19
+ Istio's [HTTPRoute and HTTPRetry fields](https://istio.io/latest/docs/reference/config/networking/virtual-service/) own request/attempt timeout and request retries. [DestinationRule connection-pool settings](https://istio.io/latest/docs/reference/config/networking/destination-rule/) own connection timeouts and concurrent-retry limits; its traffic policy owns outlier detection.
20
20
 
21
21
  ## Timeout budget rule
22
22
 
23
23
  For a chain A → B → C:
24
24
 
25
25
  ```
26
- A.ctx.deadline = D
27
- A's client-to-B budget ≤ D - margin (margin: 50-100ms for serialization)
28
- B's internal work ≤ (B-budget - C-budget)
29
- B's client-to-C budget ≤ B-budget - margin
26
+ A's remaining duration = A.ctx.deadline - now
27
+ A's client-to-B budget ≤ remaining duration - margin (e.g. 50-100ms for serialization)
28
+ B's client-to-C budget ≤ B's remaining duration - margin
29
+ All sequential work, attempts, and backoff fit within that remaining budget
30
30
  ```
31
31
 
32
- A framework helper should compute "remaining ctx deadline" and pass `min(ctx.Deadline, my-budget)` to outbound calls.
32
+ A framework helper computes `min(remaining_duration - margin, configured_call_budget, applicable_platform_cap)`. If no usable duration remains, fail without dispatching another attempt. Pass the resulting deadline downstream and recompute the remainder after work or backoff; do not compare an absolute deadline with a duration or restart the full budget at each hop.
33
33
 
34
34
  ## Retry placement
35
35
 
36
36
  | Layer | Retries WHAT | When |
37
37
  |---|---|---|
38
- | Mesh (DestinationRule) | network errors, gRPC `UNAVAILABLE`, certain 5xx | Always-safe-to-retry transports |
38
+ | Mesh (VirtualService HTTP route) | Explicitly selected network errors or response statuses | Declared idempotent calls, or failures proven to precede server receipt |
39
39
  | Framework client | idempotent business RPCs | Per-method opt-in |
40
40
  | App handler | nothing | App level retry usually wrong |
41
41
  | App business logic | high-level workflows | Saga / orchestration patterns, not "I'll retry the call once" |
42
42
 
43
- Double retry = real bad. If mesh retries 2x and framework retries 2x, you get 4x amplification on every flapping backend.
43
+ Retry layers multiply total attempts. If mesh and framework each make at most two total attempts, one logical call can reach the backend four times. Istio `retries.attempts` counts retries after the initial request: a value of 2 permits up to 3 total attempts, subject to time and retry budgets.
44
44
 
45
- **Rule**: when enabling framework-level retry, disable mesh retry for that callee (DestinationRule `retries.attempts: 0`).
45
+ **Rule**: when enabling framework-level retry, set `retries.attempts: 0` on every matching VirtualService HTTP route used by that call and verify the effective generated route configuration. Keep DestinationRule connection limits and outlier detection; they do not replace the route-level retry switch.
46
46
 
47
47
  Mesh retries back off automatically (Istio/Envoy: jittered exponential backoff with a default 25ms *base* interval — fully jittered, so an actual delay can be shorter than the base; it is not a guaranteed minimum gap); framework-level retry gets no such freebie — it must implement its own jittered backoff that fits inside the caller's remaining deadline.
48
48
 
@@ -51,24 +51,24 @@ Mesh retries back off automatically (Istio/Envoy: jittered exponential backoff w
51
51
  Per-call retry counts bound retries *per request*; they do not bound a caller's total retry share during a partial outage — at high QPS, "2 retries each" is up to a 3× load multiplier at the exact moment the upstream is sickest. Envoy's cluster circuit breakers cap this per proxy:
52
52
 
53
53
  - `max_retries` — max **concurrent** retries to the cluster, per priority. Retries beyond it overflow (fail fast, counted in `upstream_rq_retry_overflow`). The raw Envoy default is 3, but the control plane above Envoy may override it: Istio's `connectionPool.http.maxRetries` defaults to **2^32-1 — effectively unlimited** — so in an Istio mesh "leave it unset and rely on the default cap" is a trap. Set the limit explicitly and verify the *generated* Envoy cluster config, not the assumption.
54
- - `retry_budget` — replaces the fixed cap with a load-proportional one: concurrent retries ≤ `budget_percent` (default 20%) of active + pending requests, with a `min_retry_concurrency` floor so low-traffic clusters can still retry. When set, it overrides `max_retries`. Reachability caveat: Istio's DestinationRule API exposes only `connectionPool.http.maxRetries`, NOT `retry_budget` — on plain Istio, set a finite `maxRetries` first; adopting `retry_budget` there means an EnvoyFilter, acceptable only with the *generated* cluster config verified.
54
+ - `retry_budget` — replaces the fixed cap with a load-proportional one: in the default instantaneous mode, concurrent retries ≤ `budget_percent` (default 20%) of active + pending requests, with a `min_retry_concurrency` floor so low-traffic clusters can still retry. Versions exposing a non-zero `budget_interval` can count requests over that interval instead; verify the configured version and mode. When set, the budget overrides `max_retries`. Reachability caveat: Istio's DestinationRule API exposes only `connectionPool.http.maxRetries`, NOT `retry_budget` — on plain Istio, set a finite `maxRetries` first; adopting `retry_budget` there means an EnvoyFilter, acceptable only with the *generated* cluster config verified.
55
55
  - Know exactly what the budget bounds — and what it doesn't. It bounds **Envoy-originated, concurrent** retries, per proxy. It does NOT bound: retry attempt *rate*; **framework-level retries** (each framework attempt arrives at Envoy as a fresh request and bypasses `max_retries`/`retry_budget` entirely — a platform running framework retries needs a framework-side budget or strict per-call caps); or the **fleet aggregate** (circuit breaking is distributed, not coordinated — each sidecar enforces its own budget and floor, so aggregate retry load still scales with caller replica count). A true service-wide load bound requires callee-side protection (admission control / load shedding, owned by the service-architecture skills) on top.
56
56
  - When tuning for a flaky dependency, set a retry budget rather than raising per-call retry counts — but pick `budget_percent` AND `min_retry_concurrency` deliberately against the callee's capacity: on a very high-QPS caller, an unexamined 20% of active requests is far looser than `max_retries: 3`, and with many low-traffic sidecars the aggregate floor (≈ replicas × `min_retry_concurrency`) dominates instead. Alert on the overflow counter: a growing overflow stat means callers are shedding retries, which is the budget doing its job; do not "fix" it by raising the cap.
57
57
 
58
58
  ## Idempotency awareness
59
59
 
60
- Framework client retries MUST consider idempotency:
60
+ Both mesh and framework client retries MUST consider idempotency:
61
61
 
62
62
  ```
63
63
  RPC method declares "idempotent: true" in IDL or annotation
64
64
  ↓
65
- SDK retry middleware reads this; only retries idempotent methods
65
+ The retry-owning layer applies a policy only to methods covered by that declaration
66
66
  ↓
67
67
  Non-idempotent retry happens only if the network error proves the request didn't reach the server
68
68
  (connect refused, TLS handshake failure — yes; "request sent, no response" — no)
69
69
  ```
70
70
 
71
- Don't trust HTTP method (POST can be idempotent; GET can have side effects). Trust the declaration.
71
+ Don't trust HTTP method (POST can be idempotent; GET can have side effects). Trust the declaration. A 5xx, gRPC `UNAVAILABLE`, timeout, or missing response alone does not prove that a write was not applied. If the mesh cannot distinguish safe methods, keep its request retries disabled and let the method-aware SDK own the policy.
72
72
 
73
73
  ## Circuit breaker / outlier detection
74
74
 
@@ -78,12 +78,12 @@ Mesh outlier-detection (Envoy):
78
78
  trafficPolicy:
79
79
  outlierDetection:
80
80
  consecutive5xxErrors: 5 # 5 consecutive 5xx → eject this pod
81
- interval: 10s # check every 10s
81
+ interval: 10s # periodic ejection analysis / recovery sweep
82
82
  baseEjectionTime: 30s # eject for 30s minimum
83
83
  maxEjectionPercent: 50 # never eject more than 50% of pool
84
84
  ```
85
85
 
86
- This is the right place for circuit breaking. Per-host, automatic, observable via Envoy stats.
86
+ For a meshed path, this provides per-host ejection observable through Envoy stats. [Envoy's ejection algorithm](https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/outlier) checks consecutive-error thresholds inline; periodic analyses use `interval`. An ejection also depends on the configured threshold/enforcement and pool limits. Killing a pod once does not prove those conditions fired.
87
87
 
88
88
  Framework SDK circuit breakers (e.g. hystrix-style) are a fallback for environments without mesh, or for business-logic-driven breaking (e.g. "this dependency's error rate hit 10%, switch to degraded mode").
89
89
 
@@ -91,36 +91,40 @@ Don't run mesh outlier-detection AND SDK circuit breaker simultaneously without
91
91
 
92
92
  ## Cascading cancel
93
93
 
94
- When a caller's ctx is cancelled (deadline, client disconnect, user back-button), the cancel MUST propagate to in-flight downstream calls. Frameworks should support this natively; verify by spawning a long-running downstream call and cancelling the parent — both should terminate.
94
+ When a caller's ctx is cancelled (deadline, client disconnect, user back-button), propagate it to in-flight downstream calls. [gRPC cancellation](https://grpc.io/docs/guides/cancellation/) requires application handlers to cooperate; outgoing-call propagation also depends on the language/runtime. Verify cancellation reaches the handler, stops its work and child calls, and releases resources within the service's documented cancellation bound. A transport cancellation signal alone does not prove application work stopped.
95
95
 
96
96
  If a service swallows ctx cancel, downstream load amplifies during user disconnects (every abandoned tab continues hammering the DB).
97
97
 
98
98
  ## Hedging
99
99
 
100
- Hedging = send a second request after a timeout T, return whichever responds first. Useful for latency-sensitive read APIs.
100
+ Hedging sends an additional request while the first is still in flight and returns the first acceptable response. It can reduce tail latency for idempotent reads, at the cost of concurrent upstream work.
101
101
 
102
102
  Risks:
103
103
  - Doubles load if T is too short.
104
104
  - Not safe for non-idempotent calls.
105
- - Mesh-level hedging is preferred (e.g. Envoy `retry_priority` patterns); SDK hedging is workable but harder to tune.
105
+ - Bound concurrent attempts, total deadline and admitted load; cancel losing attempts after choosing a response and propagate caller cancellation.
106
+ - Envoy's [HedgePolicy](https://www.envoyproxy.io/docs/envoy/latest/api-v3/config/route/v3/route_components.proto#config-route-v3-hedgepolicy) uses `hedge_policy.hedge_on_per_try_timeout` with a finite per-try timeout and a retry policy containing retry conditions and a positive retry limit. `retry_priority` selects upstream priorities; it does not enable hedging.
107
+ - Istio's HTTPRetry API does not expose a hedging field. Use an explicitly supported platform mechanism (such as an EnvoyFilter with generated-route and runtime verification), or a method-aware SDK; do not infer hedging from `perTryTimeout` alone. Keep a single retry/hedge owner.
106
108
 
107
109
  Default: off. Enable per-method after measuring p99 latency distribution.
108
110
 
109
111
  ## Common mistakes
110
112
 
111
- - Setting framework client timeout > mesh timeout → client never sees mesh-side errors, retries the same call, mesh times out anyway. Align them.
112
- - Retrying on 4xx → client errors don't fix themselves; you just amplify load.
113
+ - Treating a shorter mesh timeout as invisible to the client → the client can receive a mesh timeout response before its own deadline. Verify error mapping and the retry-owning layer; do not turn that response into an unsafe replay.
114
+ - Retrying every 4xx → permanent client errors do not fix themselves. Only a documented transient condition, such as a rate-limit response with a bounded retry delay, may be eligible under the idempotency and remaining-budget rules.
113
115
  - Retrying with exponential backoff while caller deadline is 200ms → backoff exceeds deadline, retry never fires, you wasted code.
114
- - Mesh retry + SDK retry both on, both default 3 → 9 actual attempts per logical call.
116
+ - Mesh and SDK each allow 3 retries after the initial request → up to 16 backend attempts per logical call, before deadline/budget limits.
115
117
  - Circuit breaker tripped but no metric → debugging blind.
116
118
 
117
119
  ## Tuning starting points
118
120
 
119
- | Policy | Default |
121
+ These are example starting points for eligible traffic, not vendor defaults or a mandate to enable retries. Apply the idempotency and single-owner rules first.
122
+
123
+ | Policy | Starting point |
120
124
  |---|---|
121
125
  | Mesh request timeout | 5s (HTTP), 10s (RPC); per-callee override |
122
- | Mesh retry attempts | 2 |
123
- | Mesh retry per-try timeout | 1s |
126
+ | Mesh retry attempts | 0 unless mesh owns safe retries; then up to 2 retries within the call budget |
127
+ | Mesh retry per-try timeout | Fit within the remaining call budget; reserve time for backoff and any later attempt |
124
128
  | Mesh outlier detection | 5 consecutive 5xx, 30s eject |
125
129
  | Framework client per-call budget | 500ms or shrunk from ctx deadline |
126
130
  | Framework client retry | OFF by default; opt-in per idempotent method |
@@ -130,6 +134,9 @@ Tune from observed p99 + error rate, not vibes.
130
134
  ## Verification
131
135
 
132
136
  - Trigger downstream 5xx storm → mesh outlier-detection metrics show ejection events; client-side error rate spikes then recovers.
133
- - Force a connect refusal → mesh retry attempts visible in mesh metric; final caller sees one error.
137
+ - Force a connect refusal → only the configured retry owner produces attempts, counts fit its budget, and the final caller sees one error. With SDK-owned retries, confirm effective route retries are disabled.
134
138
  - Set per-call timeout below mesh ceiling → confirm caller sees its own deadline, not mesh's.
135
- - Cancel a request mid-flight → confirm downstream stops within one network round-trip.
139
+ - Set mesh request timeout below the client budget → confirm the mapped mesh error is visible and does not trigger a second, unintended retry layer.
140
+ - For a non-idempotent write with a lost response, confirm neither layer blindly replays it.
141
+ - With hedging enabled, delay an eligible call past the per-try timeout → confirm bounded concurrent attempts, first acceptable response selection, loser cancellation, and no second retry/hedge owner.
142
+ - Cancel a request mid-flight → confirm handler work and downstream calls stop and resources release within the declared cancellation bound.
@@ -79,7 +79,7 @@ The framework client resolver:
79
79
  ## Cross-language registration
80
80
 
81
81
  If services span Go, Python, Java, Node: every language SDK must agree on:
82
- - Service name shape (lowercase, dot-separated, no underscores if gRPC is in scope — see `grpc-authority-workaround.md`).
82
+ - Service name shape and its mapping to endpoint, authority, and TLS identity. Apply the agreed platform naming policy and actual DNS/SDK/proxy constraints; gRPC alone does not require renaming an existing working identifier (see `grpc-authority-workaround.md`).
83
83
  - Tag key names (`lane`, not `env`; pick one).
84
84
  - Heartbeat interval and TTL.
85
85
  - Health state semantics.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: product-rd-workflow
3
- description: 加功能 / 新需求 / 技术方案 / 方案评估 / 技术选型 / 可行性评估 / 工作量评估 / 多阶段重构 / 推倒重来 / 重新开发 / 完全重新开始 / 清除代码重新开发 / redo-from-scratch / 项目分析 / spec·PRD·需求文档实质内容写错要改对(substance 修正) / 方案评审通过开始实现 / 进入实现阶段 → end-to-end product R&D router for requirement shaping, spec/plan, implementation gates, assessment, redo/refactor, and multi-stack standards.
3
+ description: 加功能/新需求/技术方案/方案评估/技术选型/可行性评估/工作量评估/多阶段重构/推倒重来/重新开发/完全重新开始/清除代码重新开发/redo-from-scratch/继续之前的重构·开发/resume-in-flight-delivery/项目分析/无架构兄弟栈的服务边界·数据归属/service-boundary-data-ownership/spec·PRD·需求文档实质内容写错要改对(substance 修正)/方案评审通过·进入实现阶段 → end-to-end product R&D router for requirement shaping, spec/plan, implementation gates, assessment, redo/refactor, and multi-stack standards.
4
4
  ---
5
5
 
6
6
  # Product R&D Workflow
@@ -43,7 +43,7 @@ Use this skill as the top-level workflow for new product development, feature de
43
43
  - Skill/process extraction: when asked to summarize delivery experience, preserve a workflow lesson, update a reusable skill, or decide where a lesson belongs, route to `skill-extraction-workflow` first. Update this skill only when the lesson changes product R&D routing, gates, ownership, or lifecycle policy.
44
44
  - Upstream decision propagation: when architecture, design, testing strategy, release, observability, security, or product workflow guidance changes, treat the next R&D slice as incomplete until the downstream execution owners are named. Confirm which implementation skill, test/review skill, and release or docs owner must apply the decision, or route the gap through `skill-extraction-workflow`.
45
45
  - Text artifact quality: when this workflow creates or updates a SOP, template, checklist, report, Feishu/Lark doc, task card, launch material, or other deliverable text, run `tighten-doc` before sharing, syncing, committing, or publishing. Do not ask the user for a separate optimization confirmation unless substantive decisions may change or collaborative-comment safety is at risk.
46
- - Product R&D standards docs: when creating or updating team development standards, stack guidelines, testing standards, engineering norms, or a multi-doc handbook that belongs to the R&D lifecycle, use the R&D standards checklist in Workflow. Cross-stack or cross-service product Specs live in one authority surface; stack/service repositories keep execution slices and links back to that authority instead of redefining product goals. A multi-stack standards family must include a testing standard owned by `testing-strategy`; stack docs specialize commands and harness mechanics but do not replace the shared test-layer and CI-gate policy. For standalone wording, editing, or polishing requests with no R&D routing decision, use `tighten-doc` directly. Authority/sync-gate and testing-standard detail: `references/rd-standards-doc-family-checklist.md`.
46
+ - Product R&D standards docs: when creating or updating team development standards, stack guidelines, testing standards, engineering norms, or a multi-doc handbook that belongs to the R&D lifecycle, you must use the R&D standards checklist in Workflow. For standalone wording, editing, or polishing requests with no R&D routing decision, use `tighten-doc` directly. Authority/sync-gate and testing-standard detail: `references/rd-standards-doc-family-checklist.md`.
47
47
  - Standards-to-health-gate propagation: when a standards family is intended to check project compliance later, split each norm into deterministic checks and agent review checks, route the executable invariant model to `testing-strategy` fitness functions and the stack-specific mechanics to the owning dev/architecture skills, and do not leave conformance as a human-only checklist (mapping detail in `references/rd-standards-doc-family-checklist.md`).
48
48
  - High-risk resilience gating: when a feature touches money, billing, quota, permissions, tenant/user data isolation, privacy, high-impact AI answers, write-finality risk, repeated submission, async job finality, or incident explanation/compensation, require an explicit resilience gate before implementation or launch; do not escalate low-risk local edits only because they write files. Write-finality classification and gate detail: `references/high-risk-resilience-gates.md`.
49
49
 
@@ -180,7 +180,7 @@ Before release, confirm:
180
180
  | **Change Lead Time** | commit 到 prod 的时长 | — |
181
181
  | **Change Failure Rate (CFR)** | release 中需要 hotfix / rollback / fail forward 的比例 | — |
182
182
  | **Failed Deployment Recovery Time** | failed deployment 恢复时长 | 2023 年 DORA 重命名(原 MTTR)|
183
- | **Deployment Rework Rate** | release 后需要 rework 的比例 | 2024 年新增 |
183
+ | **Deployment Rework Rate** | 由生产事故引发的非计划部署占全部部署的比例 | 2024 年新增;[当前 DORA 口径](https://dora.dev/guides/dora-metrics/) |
184
184
 
185
185
  Elite / High / Medium / Low 具体阈值**按当年 DORA Annual State of DevOps Report 取**(不同年份数字略有变化,本 ref 不固定数字以免过时)。
186
186