@ccoalm/ccl-skills 0.15.0 → 0.15.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +80 -4
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +16 -4
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +13 -2
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +6 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/development-completion.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +37 -13
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +5 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +77 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +98 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +48 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +237 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +165 -11
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +143 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +572 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +7 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +16 -14
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +40 -83
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +44 -37
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +7 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +20 -11
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +7 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +8 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +17 -17
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +4 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +3 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +31 -7
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +74 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +103 -21
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +83 -48
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_body_compliance_grading.sh +80 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +74 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +3 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_runtime.py +428 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +190 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +106 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -1
- package/dist/assets/release.json +87 -67
- package/dist/claude-adapter.js +14 -7
- package/dist/codex-host.d.ts +2 -4
- package/dist/codex-host.js +40 -19
- package/dist/host-probe.d.ts +27 -0
- package/dist/host-probe.js +51 -0
- package/dist/opencode-adapter.js +24 -19
- package/dist/operations.js +34 -8
- package/dist/unified.d.ts +1 -1
- package/dist/unified.js +18 -9
- package/package.json +1 -1
|
@@ -5,7 +5,9 @@ description: bug / 报错 / test 挂了 / 线上问题 / 接口变慢·性能退
|
|
|
5
5
|
|
|
6
6
|
# Defect Diagnosis
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Diagnose and fix from evidence; route prevention to product, architecture, development, testing, release or tooling.
|
|
9
|
+
|
|
10
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
9
11
|
|
|
10
12
|
## Non-Negotiable Rules
|
|
11
13
|
|
|
@@ -5,7 +5,7 @@ description: Use when implementing, modifying, scaffolding, generating, or testi
|
|
|
5
5
|
|
|
6
6
|
# Go Microservice Dev
|
|
7
7
|
|
|
8
|
-
Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase.
|
|
8
|
+
Use this for implementation of new backend products and services. It should adapt to the repo in front of you, but the workflow is independent of any prior codebase. After code/test edits, self-check and invoke `code-review` automatically before completion.
|
|
9
9
|
|
|
10
10
|
## Skill Routing
|
|
11
11
|
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md
CHANGED
|
@@ -7,6 +7,8 @@ description: Use when designing, implementing, reviewing, debugging, or operatin
|
|
|
7
7
|
|
|
8
8
|
Use this for product backend work that calls, hosts, evaluates, or operates LLM and inference systems. Keep the skill generic: extract reusable mechanics only, not business-specific prompts, datasets, provider names, repository paths, or domain nouns.
|
|
9
9
|
|
|
10
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
11
|
+
|
|
10
12
|
## Skill Routing
|
|
11
13
|
|
|
12
14
|
- Use this skill for LLM gateway/client design, model registry, prompt versioning, agent/tool orchestration, streaming APIs, fallback, token/cost accounting, evals, replay, shadow comparison, batch inference, and inference observability.
|
|
@@ -5,7 +5,7 @@ description: "小程序 / Taro / 微信小程序 / 支付宝小程序 / 抖音
|
|
|
5
5
|
|
|
6
6
|
# Miniapp Product Dev
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Mini-program client engineering and product-facing delivery across WeChat, Alipay, Douyin/TikTok, Baidu and similar hosts; excludes product strategy, backend architecture and visual design. After code/test edits, self-check and invoke `code-review` automatically before completion.
|
|
9
9
|
|
|
10
10
|
## Framework Scope
|
|
11
11
|
|
|
@@ -57,7 +57,7 @@ This skill is about **using AI agents / subagents to execute work** — delegati
|
|
|
57
57
|
5. **Value**: the task is worth that premium.
|
|
58
58
|
If independence, value, or breadth is unclear, start with one focused agent or sequential delegation; use parallel multi-agent dispatch when those checks are explicitly satisfied.
|
|
59
59
|
- Model tier per dispatch is an explicit decision, not an inherited accident. On hosts that inherit the session model for unnamed dispatches (a common default — Claude-family harnesses behave this way; verify yours), an unnamed model is often the most capable and most expensive tier, so a high-volume fan-out silently puts every worker and reviewer on the top tier. The observable triggers are a fan-out — multiple dispatches (workers/reviewers) in one delivery — and any high-risk dispatch: there, name the tier per dispatch and choose by judgment complexity and risk, not token price alone — the cheapest tier routinely takes 2–3× the turns on multi-step work and costs more overall, so use a mid-tier floor for reviewers and for implementers working from prose descriptions; reserve the cheapest tier for transcription-plus-tests tasks (the plan text already contains the code to write) and single-file mechanical fixes; put architecture/design judgment, security/authority/tenant-isolation/data-loss review, and the final whole-scope review on the most capable tier — review tier scales with the diff's size, complexity, and risk, and a high-risk review never silently inherits a cheap session default even as a single dispatch (tier principle: `../skill-extraction-workflow/references/harness-patterns-and-eval.md`). Record the decision as a brief field alongside `required_skills`: `model_tier: <tier>` or `model_tier: host-default (<reason: single low-risk dispatch | no host model selection>)` — an absent field is an unmade decision, not a default, and the field is bookkeeping only until the dispatch call actually passes the matching model selector (verify the effective model where the host exposes it; a field/selector mismatch is a defect, not a recorded decision).
|
|
60
|
-
- Stop and escalate when
|
|
60
|
+
- Stop and escalate when unclear direction, unavailable dependencies, unknown completion state or missing authority remains after bounded remediation. Repeated identical verification failures (~3 times) require a status report and method/evidence checkpoint, not renewed task permission; stop identical retries and continue necessary work within existing scope, respecting explicit user limits. Escalation must name the blocker, attempts and results, current diff/commits/workspace, safest next action, and lower-risk work that can continue.
|
|
61
61
|
|
|
62
62
|
## Execution Flow
|
|
63
63
|
|
|
@@ -5,6 +5,8 @@ description: Use when implementing, modifying, scaffolding, or testing Node.js b
|
|
|
5
5
|
|
|
6
6
|
# Node.js Service Development
|
|
7
7
|
|
|
8
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
9
|
+
|
|
8
10
|
## Skill Routing
|
|
9
11
|
|
|
10
12
|
- Use this skill for Node.js service implementation: handlers, middleware, adapters, workers, jobs, clients, runtime/toolchain mechanics, and focused tests.
|
|
@@ -5,7 +5,9 @@ description: Use when designing, reviewing, debugging, or shipping observability
|
|
|
5
5
|
|
|
6
6
|
# Platform Observability
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Owns logs/metrics/traces, correlation, dashboards, alerts, on-call routing, SLI/SLO/error-budget design and framework wiring for new services.
|
|
9
|
+
|
|
10
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
9
11
|
|
|
10
12
|
You do not own:
|
|
11
13
|
- Traffic routing, retries, timeouts, mTLS, or service mesh policy — go to `platform-service-connectivity`.
|
|
@@ -30,7 +32,7 @@ Observability is the chain that turns one user action into searchable, joinable,
|
|
|
30
32
|
2. **Instrumentation** — every service emits structured logs, metrics, and spans through the **framework default**, not ad-hoc code. A service whose middleware chain does not include `ctx_inject + metrics + recovery + tracing` is unobservable by design.
|
|
31
33
|
3. **Transport** — logs go stdout → file collector (DaemonSet) → log pipeline → search index; metrics + traces go SDK → OTLP collector → metric store + trace store. Both transports must survive collector restarts and back-pressure.
|
|
32
34
|
4. **Storage + display** — logs in a searchable index keyed by log-id and trace-id; metrics in a long-term store separate from scraper; traces queryable by trace-id; one dashboard tool joins all three.
|
|
33
|
-
5. **Evidence consumption** — SLIs
|
|
35
|
+
5. **Evidence consumption** — SLIs query (3) + (4); actionable alerts route to on-call; runbooks live in a wiki that on-call can reach in under one minute.
|
|
34
36
|
|
|
35
37
|
A new service must satisfy all five before it is allowed in production. A release must produce evidence at all five before it is promoted.
|
|
36
38
|
|
|
@@ -106,9 +108,9 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
|
|
|
106
108
|
|
|
107
109
|
**Retrofit / migration contract** (the rules above are otherwise greenfield-framed): when a schema is introduced over EXISTING services, field names, metric label names, and identity env-var names are a **migration contract** — a blind rename breaks every deployed dashboard, alert, and saved query that keys on the old name. Retrofitting MUST alias or dual-write old→new and migrate consumers before retiring the old name; never rename in place. The cheap time to fix a name is before services adopt it.
|
|
108
110
|
|
|
109
|
-
### R7 — Alerts are
|
|
111
|
+
### R7 — Alerts are actionable and route to an owner
|
|
110
112
|
|
|
111
|
-
-
|
|
113
|
+
- Use SLIs for SLO paging; actionable capacity or impending-failure alerts may use internal measurements. See `references/alerting-and-on-call.md`.
|
|
112
114
|
- Severity levels: P0 (page on-call now), P1 (notify channel, ack within work hours), P2 (digest).
|
|
113
115
|
- Every alert MUST link to a runbook entry. If no runbook exists, the alert is not allowed to be P0.
|
|
114
116
|
- An alert-backed metric is a coverage contract, and the trigger is mechanical, not prose: any new or changed alert rule, SLO, dashboard alert annotation, or metric referenced by an alert policy fires this check (a written "alert on any increase" note also counts, but its absence is not an exemption). Enumerate every site that should feed the metric and verify each is actually instrumented — derive the site list from a static registry or lint where possible; the easiest site to miss is often the riskiest (e.g. the panic counter in a stream reader parsing untrusted bytes). Ship the site checklist with the alert.
|
|
@@ -208,7 +210,7 @@ If any phase has missing evidence, the work is not done.
|
|
|
208
210
|
- **"Sample everything vs head-sample 1%"** → head-sample low (1–10%) for cost; retaining errors/slow requests is **tail** sampling at the Collector and requires near-full SDK export to it (head-dropped spans never arrive) — pick one model per service, do not claim both. Sampling-by-route is acceptable for known noisy paths.
|
|
209
211
|
- **"Metrics egress: scrape vs push"** → each platform picks ONE canonical egress and every process uses it. A Prometheus-native platform may default to **scrape** (pull, with `up`/target-health + service discovery); an OTel-first platform, or workers/jobs/runtimes that can't be scraped, may default to **OTLP push** to a collector. Neither is universally better — don't overturn a working pull setup just to push. One egress per process for a given metric; a migration window may dual-write only with isolated pipelines / distinct metric names / a dedup plan. Long-term store stays separate from the short-term scrape/collect layer (R5).
|
|
210
212
|
- **"Add a new label to a metric"** → answer the cardinality question first. If max distinct values × series count > 1e6, refuse and use an exemplar trace instead.
|
|
211
|
-
- **"Alert on this symptom or that cause"** →
|
|
213
|
+
- **"Alert on this symptom or that cause"** → prefer user-visible symptoms; capacity or impending-failure warnings follow R7.
|
|
212
214
|
|
|
213
215
|
## Sanitization and Provenance
|
|
214
216
|
|
|
@@ -1,5 +1,13 @@
|
|
|
1
1
|
# Alerting and On-Call
|
|
2
2
|
|
|
3
|
+
## Choose an actionable signal
|
|
4
|
+
|
|
5
|
+
- **User symptoms and SLOs:** use service-level signals for user-symptom paging and the corresponding SLIs for SLO burn-rate alerts. Query the declared metric store; diagnostic counters do not become availability or latency SLIs merely because an alert uses them.
|
|
6
|
+
- **Capacity and impending failures:** white-box measurements may warn before user SLIs degrade, such as a credible forecast of disk exhaustion. Record the expected failure, the evidence for the threshold or forecast, the action, and the responsible owner. Choose urgency from impact and time left to act; a high internal metric alone is not a reason to page.
|
|
7
|
+
- **Diagnostic signals:** keep measurements with no actionable condition in dashboards or queries. Do not alert on every possible cause.
|
|
8
|
+
|
|
9
|
+
Verify new or changed alert rules against relevant failure, healthy, and recovery cases. Preserve the severity, ownership, runbook, and delivery requirements below for both symptom and impending-failure alerts. [Google SRE's Monitoring Distributed Systems](https://sre.google/sre-book/monitoring-distributed-systems/) describes both user symptoms and imminent saturation as alerting inputs.
|
|
10
|
+
|
|
3
11
|
## Two patterns
|
|
4
12
|
|
|
5
13
|
### Pattern A — Prometheus AlertManager
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md
CHANGED
|
@@ -5,7 +5,9 @@ description: 发布 / 灰度 / canary / rollback / rollout / 环境泳道 / prom
|
|
|
5
5
|
|
|
6
6
|
# Platform Release Engineering
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Owns lane/environment topology, build/deploy pipelines, traffic shifting (canary, blue-green, mirror), promotion gates using observability evidence, secret/dynamic-config distribution and rollback contracts.
|
|
9
|
+
|
|
10
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
9
11
|
|
|
10
12
|
You do not own:
|
|
11
13
|
- What signals exist to judge a release — see `platform-observability`.
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md
CHANGED
|
@@ -5,6 +5,8 @@ description: 服务互通 / service mesh / service discovery / mTLS / retry / ti
|
|
|
5
5
|
|
|
6
6
|
# Platform Service Connectivity
|
|
7
7
|
|
|
8
|
+
- Code/test changes require self-checks; invoke `code-review` automatically before completion.
|
|
9
|
+
|
|
8
10
|
This skill owns **how requests move between services**: the transport layer (mesh), service discovery, multi-environment routing, retry/timeout/circuit-breaker policy, and the framework middleware that propagates app-level context across hops.
|
|
9
11
|
|
|
10
12
|
You do not own:
|
|
@@ -100,10 +102,11 @@ Confusing these is the source of most "why doesn't this work" connectivity bugs.
|
|
|
100
102
|
|
|
101
103
|
### R4 — Default retry/timeout/circuit-breaker live at mesh, app overrides for business reasons
|
|
102
104
|
|
|
103
|
-
-
|
|
104
|
-
-
|
|
105
|
+
- On Istio HTTP/gRPC paths, `VirtualService` HTTP routes own request `timeout` and `retries`; `DestinationRule` owns connection-pool settings and `outlierDetection`. Other transports use their platform-owned equivalents.
|
|
106
|
+
- Before configuring retries or timers, read `references/retry-timeout-circuit-breaker.md` for field paths, single-owner retry selection, and idempotency checks. Both mesh and SDK follow them; a 5xx or missing response alone never proves replay safe.
|
|
107
|
+
- Framework client SDK may provide method-specific retry/hedging; its total call budget must fit the caller's remaining duration and any platform cap. A longer mesh timeout is a backstop, not an extended caller deadline.
|
|
105
108
|
- App code MAY override per-RPC. App code MUST NOT silently disable mesh-level outlier detection.
|
|
106
|
-
- Budget rule:
|
|
109
|
+
- Budget rule: downstream work, attempts, and backoff fit the remaining duration minus a safety margin (e.g. 100ms). Propagate deadline and cancellation; never reset the full budget at each hop.
|
|
107
110
|
|
|
108
111
|
### R5 — mTLS is mesh-default, app cannot disable
|
|
109
112
|
|
|
@@ -151,14 +154,13 @@ The client middleware fills this from ctx automatically; the server middleware e
|
|
|
151
154
|
|
|
152
155
|
For non-protobuf or metadata-only transports, an equivalent header set is compliant only when it cites a resolvable platform owner record, such as a repo/path, document URL, registry id, gateway policy, or owner-suite id. The owner record must enumerate the concrete header names. In diff-only review without resolver tooling, the diff must cite a stable owner-record locator and list the concrete header names inline; with resolver tooling, the reviewer or owner-suite may resolve the locator to those names instead. The enumerated header names must then be checked against the actual propagated and exposed headers: caller-supplied identity headers are absent unless positive authenticated-caller evidence exists. An unresolvable, non-enumerating, uncited, service-local, or unchecked header set is an open gap, not a permitted alternative. For pure HTTP (no RPC base), the equivalent is a stable owner-recorded header set, also filled by middleware.
|
|
153
156
|
|
|
154
|
-
### R8 — gRPC `:authority`
|
|
157
|
+
### R8 — gRPC `:authority` compatibility follows the actual request path
|
|
155
158
|
|
|
156
|
-
|
|
159
|
+
HTTP/2 `:authority` carries URI authority, not a single DNS label. An underscore alone does not prove a gRPC protocol violation. DNS hostnames, certificate identity checks, SDK validation, and proxy routing can impose different constraints; preserve the constraints of the deployed path.
|
|
157
160
|
|
|
158
|
-
|
|
159
|
-
2. **Mesh-level rewrite**: an EnvoyFilter Lua snippet replaces `_` with `-` in `:authority` for gRPC requests, before routing.
|
|
161
|
+
Before renaming a service or adding a rewrite, capture the failing request's authority, the rejecting layer and version, and the relevant error or trace. A platform may require DNS-compatible service names and enforce that policy at registration and CI, but a registry name need not be the wire authority.
|
|
160
162
|
|
|
161
|
-
|
|
163
|
+
A proxy rewrite is an option only when the request reaches that filter before the rejecting layer. It cannot fix a client rejection before transmission or an inbound parser rejection before the filter runs. Scope any verified rewrite to the affected traffic and explicit authority mappings; retain route, TLS identity, and authorization checks. Diagnosis, migration, and verification: `references/grpc-authority-workaround.md`.
|
|
162
164
|
|
|
163
165
|
### R9 — Ingress and egress are explicit, not implicit
|
|
164
166
|
|
|
@@ -200,8 +202,8 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
|
|
|
200
202
|
### Phase B — Mesh policy
|
|
201
203
|
|
|
202
204
|
1. PeerAuthentication = STRICT (mTLS namespace-wide).
|
|
203
|
-
2. DestinationRule per critical callee:
|
|
204
|
-
3. VirtualService
|
|
205
|
+
2. DestinationRule per critical callee: connection-pool limits, connect timeout, outlier detection.
|
|
206
|
+
3. VirtualService HTTP routes: lane header match → lane subset, request timeout, and the R4 retry policy.
|
|
205
207
|
4. AuthorizationPolicy expresses which services may call which (zero-trust at network level).
|
|
206
208
|
|
|
207
209
|
### Phase C — Service discovery
|
|
@@ -216,7 +218,7 @@ Option 1 is cleaner long-term; option 2 is the live-system workaround. Document
|
|
|
216
218
|
|
|
217
219
|
1. Read the framework default client/server options module. Confirm middleware chain matches R6.
|
|
218
220
|
2. Confirm dev cannot build a client/server without inheriting these.
|
|
219
|
-
3. Test:
|
|
221
|
+
3. Test: induce the configured outlier threshold on one callee → verify ejection and recovery, the configured retry layer retries only eligible calls within its budget, and exhausted calls propagate a stable error-code.
|
|
220
222
|
|
|
221
223
|
### Phase E — Failure modes
|
|
222
224
|
|
|
@@ -234,7 +236,7 @@ Before marking work done:
|
|
|
234
236
|
|
|
235
237
|
## Decision Points
|
|
236
238
|
|
|
237
|
-
- **"Add a new retry policy for service X"** →
|
|
239
|
+
- **"Add a new retry policy for service X"** → apply R4's replay-safety and single-owner checks. Use the matching VirtualService HTTP route for Istio retries; disable and verify route retries when the SDK owns them. DestinationRule limits concurrent retries, not per-request attempts.
|
|
238
240
|
- **"Service A times out calling Service B"** → check three layers in order: app deadline (ctx timeout) → framework client timeout → mesh request timeout. Whichever is smaller wins; align them.
|
|
239
241
|
- **"Switch service discovery mode"** → use `references/service-discovery-choice.md`; do not mandate registry unless the platform needs registry-specific capabilities such as per-instance drain, out-of-cluster lookup, or existing registry federation.
|
|
240
242
|
- **"Need mTLS to a non-mesh external service"** → egress gateway with terminating TLS, not app-managed certs.
|
|
@@ -257,7 +259,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
|
|
|
257
259
|
- `references/framework-middleware.md` — Server and client middleware chain (HTTP + RPC); the canonical RPC base struct shape; verification commands.
|
|
258
260
|
- `references/multi-env-routing.md` — Lane label end-to-end recipe; VirtualService patterns; queue-boundary propagation; stress/shadow tags.
|
|
259
261
|
- `references/retry-timeout-circuit-breaker.md` — Mesh defaults vs SDK overrides; cascading timeout budgets; idempotency awareness; outlier detection tuning.
|
|
260
|
-
- `references/grpc-authority-workaround.md` —
|
|
262
|
+
- `references/grpc-authority-workaround.md` — Locate authority rejection; distinguish naming policy from protocol syntax; verify scoped compatibility changes.
|
|
261
263
|
- `references/dual-sidecar-and-traffic-config-center.md` — Pod-level dual sidecar (mesh + platform), per-protocol mesh injection policy, per-caller-callee traffic config via config center (separate from mesh routing).
|
|
262
264
|
- `references/rpc-framework-recipe.md` — Concrete kitex/hertz default suite: shared RPC base field contract, RPC base.Request full schema (8 fields incl From/To), dual-channel ctx propagation (metainfo + grpc metadata), 9-code error enum + framework error mapping table, three resolver strategies (registry / FQDN fallback / proxy), platform latency histogram buckets, CORS defaults exposing log-id header, server boot sequence with graceful shutdown.
|
|
263
265
|
- `references/service-discovery-choice.md` — Decision framework: registry-based vs k8s-native SD; both support lane routing + canary + multi-env; pick by per-instance drain need, laptop access pattern, federation preference, operational burden; mixed mode (k8s east-west + thin registry for laptop) is workable; migration paths in both directions.
|
|
@@ -269,7 +271,7 @@ Reused industry patterns (PSM-style identity, `<owner>.<class>.<env>` shape, `tr
|
|
|
269
271
|
|
|
270
272
|
1. **Static**: framework default options module includes all R6 middleware; mesh PeerAuthentication is STRICT; DestinationRule and VirtualService exist for every callee that participates in lane routing.
|
|
271
273
|
2. **Live, identity**: cross-service trace shows log-id + lane consistent across all hops; mesh access log lines include both.
|
|
272
|
-
3. **Live, mesh policy**:
|
|
274
|
+
3. **Live, mesh policy**: trigger the configured outlier threshold with an eligible pool size; verify ejection and recovery against the effective policy and metrics. Consecutive-error checks are inline, not delayed until the periodic analysis interval.
|
|
273
275
|
4. **Live, lane routing**: send request with non-default lane → confirm only matching-lane instances serve it.
|
|
274
276
|
5. **Static, no leakage**: grep this skill's content — zero internal hostnames, repo names, or business terms.
|
|
275
277
|
|
|
@@ -36,7 +36,7 @@ Mesh injection is not all-or-nothing. Real platforms apply per-protocol policy:
|
|
|
36
36
|
| App protocol | Inject Istio sidecar? | Reason |
|
|
37
37
|
|---|---|---|
|
|
38
38
|
| HTTP (REST) | YES | Envoy handles HTTP/1.1, HTTP/2; full feature support |
|
|
39
|
-
| gRPC | YES | HTTP/2 + per-call routing;
|
|
39
|
+
| gRPC | YES | HTTP/2 + per-call routing; verify `:authority` routing against the actual SDK/proxy path |
|
|
40
40
|
| TCP (raw) | NO (often) | Envoy TCP proxy is feature-poor; routing/auth less useful at L4 |
|
|
41
41
|
| Thrift (TTHeader) | NO (often) | Same — L4-ish; tooling limited |
|
|
42
42
|
|
|
@@ -1,90 +1,47 @@
|
|
|
1
|
-
# gRPC
|
|
2
|
-
|
|
3
|
-
##
|
|
4
|
-
|
|
5
|
-
HTTP/2
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
- Enforce at CI / lint when defining new services.
|
|
22
|
-
- Migrate existing names by renaming + parallel registration during a deprecation window.
|
|
23
|
-
|
|
24
|
-
### Fix 2 (live-system workaround): mesh-level rewrite
|
|
25
|
-
|
|
26
|
-
When you can't break existing names, an Envoy Lua filter rewrites `:authority` on the fly:
|
|
27
|
-
|
|
28
|
-
```yaml
|
|
29
|
-
apiVersion: networking.istio.io/v1alpha3
|
|
30
|
-
kind: EnvoyFilter
|
|
31
|
-
metadata:
|
|
32
|
-
name: modify-grpc-authority
|
|
33
|
-
namespace: istio-system
|
|
34
|
-
spec:
|
|
35
|
-
configPatches:
|
|
36
|
-
- applyTo: HTTP_FILTER
|
|
37
|
-
match:
|
|
38
|
-
context: ANY
|
|
39
|
-
listener:
|
|
40
|
-
filterChain:
|
|
41
|
-
filter:
|
|
42
|
-
name: "envoy.filters.network.http_connection_manager"
|
|
43
|
-
patch:
|
|
44
|
-
operation: INSERT_BEFORE
|
|
45
|
-
value:
|
|
46
|
-
name: envoy.filters.http.lua
|
|
47
|
-
typed_config:
|
|
48
|
-
"@type": type.googleapis.com/envoy.extensions.filters.http.lua.v3.Lua
|
|
49
|
-
inlineCode: |
|
|
50
|
-
function envoy_on_request(request_handle)
|
|
51
|
-
local authority = request_handle:headers():get(":authority")
|
|
52
|
-
local content_type = request_handle:headers():get("content-type")
|
|
53
|
-
if authority and content_type and string.find(content_type, "application/grpc") then
|
|
54
|
-
local modified_authority = string.gsub(authority, "_", "-")
|
|
55
|
-
request_handle:headers():replace(":authority", modified_authority)
|
|
56
|
-
end
|
|
57
|
-
end
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
Effects:
|
|
61
|
-
- Applies to gRPC traffic only (content-type check).
|
|
62
|
-
- Rewrites `_` to `-` in `:authority`.
|
|
63
|
-
- Callee's registry instance name must also use the `-` form so routing matches.
|
|
64
|
-
|
|
65
|
-
## When to use which
|
|
66
|
-
|
|
67
|
-
| Situation | Fix |
|
|
1
|
+
# gRPC Authority Compatibility
|
|
2
|
+
|
|
3
|
+
## Separate the names and constraints
|
|
4
|
+
|
|
5
|
+
HTTP/2 `:authority` conveys the target URI's authority, as defined by [RFC 9113 §8.3.1](https://www.rfc-editor.org/rfc/rfc9113.html#section-8.3.1). [RFC 3986 §3.2.2](https://www.rfc-editor.org/rfc/rfc3986.html#section-3.2.2) permits a registered name containing unreserved characters, including `_`. This syntax does not guarantee DNS resolution, certificate identity matching, or acceptance by every SDK and proxy version.
|
|
6
|
+
|
|
7
|
+
Keep these values distinct when diagnosing a request:
|
|
8
|
+
|
|
9
|
+
- Service-registry identifier and resolved network endpoint.
|
|
10
|
+
- HTTP/2 authority used for virtual-host routing.
|
|
11
|
+
- TLS server name and the certificate identity expected on each TLS hop.
|
|
12
|
+
- Platform naming, routing, and authorization policies.
|
|
13
|
+
|
|
14
|
+
A platform can require DNS-compatible service names and enforce that choice at registration and CI. Existing identifiers do not require migration merely because they contain an underscore; first establish which constraint the actual path violates.
|
|
15
|
+
|
|
16
|
+
## Locate the rejection before choosing a fix
|
|
17
|
+
|
|
18
|
+
Record the runtime/SDK, proxy versions and relevant configuration, exact authority, and the failing run's error or trace. Use a synthetic payload and redact credentials from captured evidence.
|
|
19
|
+
|
|
20
|
+
| Observed boundary | Next action |
|
|
68
21
|
|---|---|
|
|
69
|
-
|
|
|
70
|
-
|
|
|
71
|
-
|
|
|
72
|
-
|
|
|
22
|
+
| Authority syntax is malformed | Validate URI authority syntax, including brackets around an IPv6 literal, before changing service registration or routing. |
|
|
23
|
+
| Client rejects before transmitting HTTP/2 headers | Check that client's authority validation and supported configuration. Fix the client-side mapping or naming contract; a downstream proxy cannot rewrite a request it never receives. |
|
|
24
|
+
| DNS resolution fails | Check the resolved hostname and resolver's naming rules. Changing a later HTTP header does not repair failed resolution. |
|
|
25
|
+
| TLS handshake or certificate identity check fails | Check that hop's endpoint, server name, trust chain, and certificate identities. Retain verification; changing authority is not evidence that TLS is fixed. |
|
|
26
|
+
| Proxy/parser rejects before the HTTP filter runs | Fix the supported parser/input contract or an earlier owned mapping. A Lua filter after the rejection cannot intervene. |
|
|
27
|
+
| Request reaches HTTP filters, then the intended virtual-host route does not match | Compare the received authority with the generated route configuration. A supported authority mapping may be appropriate after proving the mismatch. |
|
|
28
|
+
| Existing path accepts the name and reaches the intended service | Preserve it unless a separate platform naming-policy migration is required. |
|
|
73
29
|
|
|
74
|
-
|
|
30
|
+
`RST_STREAM` or `INTERNAL_ERROR` alone does not identify an underscore problem. Confirm the first rejecting layer rather than treating every transport failure as the same naming defect.
|
|
75
31
|
|
|
76
|
-
|
|
77
|
-
- Registry rejects registration with `_` in name.
|
|
78
|
-
- Linter / CI rejects PR adding such a name.
|
|
32
|
+
## Choose a bounded compatibility change
|
|
79
33
|
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
34
|
+
**Naming policy or client mapping.** Where a DNS-compatible name is required, define the allowed form and the mapping from registry identity to endpoint/authority. Check uniqueness before migration: replacing `_` with `-` can collapse two distinct names. Migrate registrations, routes, and callers together with a compatibility window and rollback path. Use supported client options; do not bypass certificate or authorization checks to make a name work.
|
|
35
|
+
|
|
36
|
+
**Proxy mapping.** Use only if the request reaches the chosen filter and the mapping addresses a reproduced failure. Prefer the platform's supported routing mechanism. If an EnvoyFilter is necessary, verify the installed Istio/Envoy API and generated configuration, and scope it to the affected workloads, listener/direction, route, and explicit old-to-new authority mapping. Do not install an all-workload, all-direction underscore replacement. The mapping must preserve the intended destination, tenant/lane routing, authorization, and TLS identity on each hop; rewriting a header does not itself update those contracts.
|
|
37
|
+
|
|
38
|
+
## Verification
|
|
83
39
|
|
|
84
|
-
|
|
40
|
+
Retain the original failing case and expected rejecting layer. After the change, verify:
|
|
85
41
|
|
|
86
|
-
|
|
87
|
-
-
|
|
88
|
-
|
|
42
|
+
1. The same request succeeds through the intended path and reaches the intended service; observe authority before/after any mapping and the selected route.
|
|
43
|
+
2. Unrelated authorities and non-target traffic retain their behavior. Include potentially colliding names and unknown authorities as negative controls.
|
|
44
|
+
3. Certificate identity and authorization failures still reject requests; the compatibility change has not disabled those checks.
|
|
45
|
+
4. Registration/client/route changes can be rolled back together without sending traffic to another service.
|
|
89
46
|
|
|
90
|
-
|
|
47
|
+
Static configuration validation proves only configuration properties. Claims about a deployed SDK/proxy path require execution evidence from that path.
|
|
@@ -33,7 +33,7 @@ Reference topology and component rules. Istio is the recurring example; rules ge
|
|
|
33
33
|
|
|
34
34
|
## Istio Ambient Mode (GA Nov 2024, v1.24+)
|
|
35
35
|
|
|
36
|
-
- **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block:
|
|
36
|
+
- **Istio Ambient Mode reached General Availability in Istio v1.24 (announced November 2024)** per `istio.io/latest/blog/2024/ambient-reaches-ga/` — `ztunnel` + `waypoint` architecture is the stable sidecar alternative for new production mesh deployments. **Two-layer split**: `ztunnel` (Rust-based DaemonSet, one per node, L4-only — mTLS + simple L4 authz + telemetry) handles every pod's transport-layer mesh participation without per-pod sidecar injection; `waypoint` proxies (Envoy-based, scaled independently from app workloads) handle L7 features when needed (rich authz, traffic routing, resilience). Per Istio's own reported numbers, the architecture can save 90%+ memory/CPU vs the sidecar model in dense-pod-per-node workloads (Istio's claim — re-measure on the team's actual workload before quoting savings). **Architecture choice for new mesh**: (a) **choose ambient** when memory/CPU per pod is the binding constraint, sidecar injection causes restart cycles the team wants to avoid, or a subset of namespaces don't need L7 features at all; (b) **stay on sidecar mode** when the team has deep sidecar-specific tooling (custom `EnvoyFilter`s injected per pod, sidecar-aware debugging recipes, app code that expects `localhost:15001` proxy conventions), or when ambient's narrower L7-feature coverage hits a gap the team relies on. **Mixed-mode in one cluster is officially supported** per the Istio docs: sidecar-mode namespaces and ambient-mode namespaces coexist; migration can be incremental namespace-by-namespace. **CRITICAL migration block: L7 policy enforcement changes when moving from sidecar to ambient.** Ztunnel enforces only L4 policy. A workload-selector policy containing L7 conditions that is picked up by ztunnel [fails safe as a DENY policy](https://istio.io/latest/docs/ambient/usage/l4-policy/#policies-with-layer-7-conditions), potentially blocking legitimate traffic. L7 enforcement requires a waypoint and `targetRefs` bound to the intended Service or waypoint Gateway; deploying a waypoint alone does not migrate selector-based policies. **Pre-flip gate**: inventory affected authorization policies, prepare the waypoint and correctly scoped bindings, and plan the old-policy/pod-restart transition using the deployed version's [migration guide](https://istio.io/latest/docs/ambient/migrate/migrate-policies/). Preserve mTLS and L4 authorization; do not remove a rejecting policy merely to restore connectivity. Verify allowed, denied, and waypoint-bypass paths before completing migration. If the transition cannot maintain required L7 protection, hold migration or use an approved maintenance window. **Observability shift**: ztunnel emits L4 metrics; waypoint emits L7 metrics; existing dashboards keyed on sidecar `istio_requests_total` need a ambient-mode equivalent for the L7-routed traffic — route this update through `platform-observability` rather than redesigning here.
|
|
37
37
|
|
|
38
38
|
## Sidecar injection rules
|
|
39
39
|
|
|
@@ -94,7 +94,7 @@ Egress gateway is OFF by default. Turn ON only when a workload needs controlled
|
|
|
94
94
|
EnvoyFilter is the escape hatch. Use sparingly; each filter is hard to test and easy to break across Envoy upgrades.
|
|
95
95
|
|
|
96
96
|
Acceptable uses:
|
|
97
|
-
- Header
|
|
97
|
+
- Header mapping for a reproduced compatibility failure, scoped to the affected workloads and explicit names while preserving routing, TLS, and authorization (see `grpc-authority-workaround.md`).
|
|
98
98
|
- Adding a Lua filter for one-off business handling that doesn't yet have a first-class WASM filter.
|
|
99
99
|
- Custom rate-limit before the canonical RLS is rolled out.
|
|
100
100
|
|