@ccoalm/ccl-skills 0.6.2 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-unverified-cli-flag.sh +309 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_unverified_cli_flag.sh +483 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +14 -11
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +2 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +25 -9
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +7 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +5 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/audit-history-architecture.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +7 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +2 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/notification-architecture.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/replay-comparison-architecture.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/workflow-state-architecture.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +6 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/audit-history-patterns.md +29 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +25 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/notification-patterns.md +40 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/replay-comparison-patterns.md +30 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +48 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +10 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +48 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +21 -2
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +8 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +5 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-parallel-stack-parity.sh +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_parallel_stack_parity.sh +183 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +2 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -4
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +1 -1
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +16 -5
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +5 -3
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +16 -6
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/self-benchmark-baseline.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
- package/dist/assets/release.json +115 -50
- package/package.json +1 -1
|
@@ -6,7 +6,7 @@ Scope split with the sibling `mq-consumer-architecture.md` (which covers consume
|
|
|
6
6
|
|
|
7
7
|
> **Conforms to the parallel-stack references pattern.** This file follows the layout documented in `skill-extraction-workflow/references/parallel-stack-references-pattern.md`: mirrored stack-agnostic core (when-applies through operations checklist), stack-specific implementation patterns section, and the embedded `### Mirrored-section grep gate` at the end of the stack-glue. The sibling `python-service-architecture/references/event-driven-architecture.md` mirrors the same structure. Either this file or `multi-tenant-isolation.md` may be used as a template for new parallel-stack extractions; multi-tenant additionally demonstrates the `## Topic-extension backlog` H2 for topic-wider-than-loop cases.
|
|
8
8
|
|
|
9
|
-
> **Sibling sync.** A parallel `python-service-architecture/references/event-driven-architecture.md` mirrors **all non-stack-specific sections** of this file (when-applies/not-applies, delivery semantics, event vs command vs query, idempotency, outbox, ordering, schema evolution, retry/DLQ/replay, backpressure, fanout, saga, end-to-end exactly-once, anti-patterns, operations checklist). Only the *Go-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change to prevent drift.
|
|
9
|
+
> **Sibling sync.** A parallel `python-service-architecture/references/event-driven-architecture.md` mirrors **all non-stack-specific sections** of this file (when-applies/not-applies, delivery semantics, event vs command vs query, idempotency, outbox, ordering, schema evolution, retry/DLQ/replay, backpressure, fanout, saga, end-to-end exactly-once, anti-patterns, operations checklist). Only the *Go-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change to prevent drift. Tree-specific routing references are written inline for both trees so the mirrored bytes stay identical; cross-file parity is machine-checked by `skill-extraction-workflow/scripts/check-parallel-stack-parity.sh` (wired into `check-ccl-skills.sh`), which diffs the mirrored regions byte-for-byte (no normalization) and blocks on any divergence.
|
|
10
10
|
|
|
11
11
|
> **Sanitization boundary.** The named brokers (Kafka, Pulsar, RabbitMQ, NATS JetStream, Redis Streams) and libraries below are concrete examples for **internal** implementation guidance, scoped to the implementation team's approved audience. Before this file (or excerpts) is copied into any document leaving that audience — external / client-facing materials, customer-specific deliverables, regulator or auditor evidence packages, SOC / compliance reports, procurement responses, or partner architecture appendices — replace the named choices with generic categories (`the broker`, `a partitioned log`, `a confirm-mode AMQP queue`) unless the vendor selection is already approved for disclosure to that specific audience.
|
|
12
12
|
|
|
@@ -20,8 +20,8 @@ Apply when the service:
|
|
|
20
20
|
|
|
21
21
|
Skip when the service:
|
|
22
22
|
- only does in-process pub-sub or fire-and-forget logging,
|
|
23
|
-
- uses synchronous RPC with no durable async boundary (use `protobuf-contract-architecture.md` instead),
|
|
24
|
-
- uses a job queue purely for in-tenant background work where loss is acceptable (use `notification-architecture.md` or `bulk-workflow-architecture.md`).
|
|
23
|
+
- uses synchronous HTTP/RPC with no durable async boundary (use `api-contract-and-schema.md` on the Python tree / `protobuf-contract-architecture.md` on the Go tree instead),
|
|
24
|
+
- uses a job queue purely for in-tenant background work where loss is acceptable (use `background-jobs-and-scheduling.md` or `batch-and-pipeline-architecture.md` on the Python tree / `notification-architecture.md` or `bulk-workflow-architecture.md` on the Go tree).
|
|
25
25
|
|
|
26
26
|
## Delivery semantics taxonomy
|
|
27
27
|
|
|
@@ -39,7 +39,7 @@ The semantic shape determines ownership, schema, and retry posture:
|
|
|
39
39
|
|
|
40
40
|
- **Event** — a fact about something that happened. Past tense. Owned by the producer. Many consumers may subscribe. Schema evolves with backward-compatible additions; semantic meaning is fixed once published.
|
|
41
41
|
- **Command** — a request to do something. Imperative. Owned by the recipient's contract. Typically one consumer (a command handler). May fail validation and be rejected; the sender is told.
|
|
42
|
-
- **Query** — a request for state. Synchronous RPC or query API, not a durable message. If you find yourself sending a query as a message, you probably want a query API plus an event subscription for change notifications.
|
|
42
|
+
- **Query** — a request for state. Synchronous HTTP/RPC or query API, not a durable message. If you find yourself sending a query as a message, you probably want a query API plus an event subscription for change notifications.
|
|
43
43
|
|
|
44
44
|
Mixing these confuses ownership: a command consumer that drops the message because "it's an event, consumers are best-effort" is a bug; an event producer that retries indefinitely because "it's a command, must deliver" creates head-of-line blocking.
|
|
45
45
|
|
|
@@ -77,7 +77,7 @@ The outbox pattern makes "write to DB and publish event" atomic without distribu
|
|
|
77
77
|
|
|
78
78
|
Without one of these (and its prerequisites), `SKIP LOCKED` and per-key ordering are mutually exclusive — pick which property the outbox actually delivers and document it on the event contract.
|
|
79
79
|
|
|
80
|
-
Inbox (the dual) is rarer
|
|
80
|
+
Inbox (the dual) is rarer: when a consumer must read a message and produce a side effect in a different system atomically, the consumer writes the inbox record + side effect in one DB tx, then acks the broker. On crash before ack, redelivery hits the inbox row and skips the side effect.
|
|
81
81
|
|
|
82
82
|
## Delivery ordering
|
|
83
83
|
|
|
@@ -100,13 +100,13 @@ Events outlive the producer's current code. Schema discipline is non-negotiable:
|
|
|
100
100
|
- **Security / compliance exception** — when continuing to emit a field is itself the problem (leaky PII field, secret accidentally embedded, regulator-mandated retraction), additive-only is overruled. The path is: create a new event version that omits the field, inventory active consumers, deploy a coordinated emergency migration plan (consumer flips first, producer flips second, old version retired), and define historical-data handling (purge, redaction in archives, controlled access). Document the security trigger; this path is not the default deprecation flow.
|
|
101
101
|
- **Rolling deprecation (the default)** — create a new event type or version, dual-publish or dual-read across a compatibility window, migrate consumers with monitoring and a rollback path, then retire the old version after explicit approval. Do not stop-the-world.
|
|
102
102
|
- **Schema discovery model** — choose by audience:
|
|
103
|
-
- *Single-team, single-broker boundary*: in-payload schema or schema
|
|
103
|
+
- *Single-team, single-broker boundary*: in-payload schema or a model-derived schema (from the service's typed model layer) in the envelope header is acceptable.
|
|
104
104
|
- *Multi-team, single-broker fanout*: a schema registry (Confluent-style or self-hosted) validates compatibility at build/deploy time.
|
|
105
105
|
- *Multi-broker, multi-tenant, or external consumers* (webhooks, partner integrations, tenant-private contracts): hybrid — versioned envelope (`event_type`, `event_version`) in payload + per-tenant contract catalog/registry where available + signed schema references for external consumers + broker-specific validation gates. Neither a single registry nor in-payload alone is sufficient.
|
|
106
106
|
- **Metadata leakage in cross-trust boundaries** — for external or multi-tenant contracts, the schema discovery layer itself can leak. `event_type` strings, version names, registry paths, enum labels, and catalog visibility can reveal unreleased products, regulated workflows, internal team structure, or tenant-specific capabilities even when payload fields are field-level protected. Before publishing to an external boundary: classify each of (`event_type`, `event_version`, schema-reference path, registry namespace, enum values, error codes, catalog listings) by audience; use opaque public aliases (`event_type=ext.<opaque-name>`, namespace-by-tenant-without-tenant-name) where the internal name is sensitive; never let an internal event type cross the boundary as-is.
|
|
107
107
|
- **Breaking change discipline** — a semantic-level breaking change (a field's meaning changes, an enum value is repurposed) is not caught by schema compatibility checks. Document semantic changes; coordinate consumer rollout before publishing the new shape.
|
|
108
108
|
|
|
109
|
-
For
|
|
109
|
+
For event payloads modelled with the service's typed-model layer, freeze the model at publish time and version the envelope; the producer's evolving model must not reach the wire without a registered new version. The stack-glue section names the specific model library and freezing pattern.
|
|
110
110
|
|
|
111
111
|
## Retry, dead-letter, replay
|
|
112
112
|
|
|
@@ -135,11 +135,11 @@ Document which fixture class is used and which classes are deliberately not cove
|
|
|
135
135
|
Consumer lag, broker queue depth, and producer rate are the three backpressure signals. Wire them:
|
|
136
136
|
|
|
137
137
|
- **Consumer-side** — bound the work-in-flight, but **the shape depends on ordering**:
|
|
138
|
-
- *Unordered consumer*: a bounded work-queue sized to the worker pool, N workers reading from the queue
|
|
139
|
-
- *Ordered partitioned log* (Kafka, Pulsar key-shared, NATS JetStream ordered consumer): **one serial work lane per assigned partition/key**. Feed each partition into its own bounded
|
|
138
|
+
- *Unordered consumer*: a bounded work-queue sized to the worker pool, N workers reading from the queue. A shared shutdown signal stops the workers (the stack-glue section names the specific primitive). Order of completion is undefined.
|
|
139
|
+
- *Ordered partitioned log* (Kafka, Pulsar key-shared, NATS JetStream ordered consumer): **one serial work lane per assigned partition/key**. Feed each partition into its own bounded work-queue + single worker, or process messages serially within the partition's poll loop. Feeding multiple ordered partitions into a single shared work-queue + worker pool loses per-partition ordering, and a slow message on one partition can starve cold partitions or let later offsets overtake earlier ones. The bounded-queue-plus-pool shape is correct for unordered work queues; not for ordered partitioned logs.
|
|
140
140
|
- *Rebalance handling for partition-assigned consumers* — when a Kafka / Pulsar key-shared / similar consumer group rebalances and a partition is revoked, "one lane per partition" is unsafe without explicit rebalance discipline. The revoked owner must (1) stop fetching from the partition immediately, (2) drain or cancel its in-flight lane (await handler completion to a bounded deadline, or cancel with an explicit `partial-failure` disposition), (3) commit or abort offsets according to the handler outcome (commit only completed offsets; do not commit `last poll` blindly), and (4) be fenced so it cannot still publish a side effect after the new owner has started — typically by tagging each in-flight message with the assignment epoch and refusing side effects whose epoch is stale. Without fencing, the new owner and the old owner can process the same business key concurrently; per-partition ordering at steady state is meaningless if the rebalance window allows concurrent processing.
|
|
141
141
|
- Avoid unbounded worker-per-message fanout in all cases (the stack-glue section names the specific anti-pattern API).
|
|
142
|
-
- **Producer-side** — when the broker buffer fills (Kafka producer queue, NATS slow-consumer warning), block the producer's caller with a bounded wait or shed load at the producer entry point. Never block forever; surface a typed error after a bounded wait so upstream can backpressure further.
|
|
142
|
+
- **Producer-side** — when the broker buffer fills (Kafka producer queue, RabbitMQ unconfirmed-publishes limit, NATS slow-consumer warning), block the producer's caller with a bounded wait or shed load at the producer entry point. Never block forever; surface a typed error after a bounded wait so upstream can backpressure further.
|
|
143
143
|
- **Cross-service** — a slow consumer is an upstream producer's problem to know about. Consumer lag must be exposed as a metric and alerted; producers cannot fix what they cannot see.
|
|
144
144
|
|
|
145
145
|
## Fanout patterns
|
|
@@ -187,6 +187,8 @@ Stack-agnostic recipe; document each clause for every event-driven boundary that
|
|
|
187
187
|
|
|
188
188
|
If any clause is missing, the boundary is at-least-once with duplicates. Tell consumers honestly.
|
|
189
189
|
|
|
190
|
+
External grounding (adopted in part): this recipe is an instance of the end-to-end argument — [Saltzer, Reed & Clark, *End-to-End Arguments in System Design*, ACM TOCS 2(4), 1984](https://web.mit.edu/Saltzer/www/publications/endtoend/endtoend.pdf) — a function that "can completely and correctly be implemented only with the knowledge and help of the application standing at the endpoints of the communication system" cannot be delegated to the communication layer, and broker-level transactional features are that paper's "incomplete version … useful as a performance enhancement", never the end-to-end guarantee. Borrowed scope: the placement argument only; the five-clause recipe and the atomic-domain boundary are this skill's own operational criteria.
|
|
191
|
+
|
|
190
192
|
## Anti-patterns
|
|
191
193
|
|
|
192
194
|
- **Post-commit publish (durable cross-process)** — publishing the event after the DB transaction commits, without an outbox, when consumers are in another process. A crash between commit and publish silently drops the event. In-process, same-instance, rebuildable post-commit hooks are not this anti-pattern.
|
|
@@ -198,7 +200,7 @@ If any clause is missing, the boundary is at-least-once with duplicates. Tell co
|
|
|
198
200
|
- **DLQ as graveyard** — messages land in DLQ, nobody looks, no replay tooling. The DLQ becomes a silent data-loss channel.
|
|
199
201
|
- **Exactly-once claimed by broker badge** — broker config has an "exactly-once" mode, but the consumer is not idempotent and the producer is not transactional, OR the side effect is outside the atomic domain. The claim is wrong; record correct end-to-end semantics.
|
|
200
202
|
- **Producer-held subscriber list as code** — subscriber list hardcoded at the producer when broker-side subscription is possible. (CDC bridges and webhook dispatchers are exceptions; their lists must live as config with audit and rotation.)
|
|
201
|
-
- **Sync RPC as command** — a "command" sent via blocking RPC with retry and no replay path. If the call needs the durability of a queue, use a queue; if it needs the latency of RPC, accept best-effort.
|
|
203
|
+
- **Sync HTTP/RPC as command** — a "command" sent via blocking HTTP/RPC with retry and no replay path. If the call needs the durability of a queue, use a queue; if it needs the latency of HTTP/RPC, accept best-effort.
|
|
202
204
|
|
|
203
205
|
## Operations checklist (event-driven boundary launch)
|
|
204
206
|
|
|
@@ -231,6 +233,7 @@ These are stack-localized recipes that implement the stack-agnostic patterns abo
|
|
|
231
233
|
3. **Mark-sent tx (short)**: `BEGIN; UPDATE outbox SET sent_at = NOW(), processing_until = NULL WHERE id = $id AND owner = $owner; COMMIT;` — guarded by `owner` so a re-leased row (after the original lease expired) is not double-marked.
|
|
232
234
|
4. **Abandon tx**: after K publish failures, mark row `abandoned` and alert; do not block the partition key forever on a poison row (see the *Dispatcher that refuses to publish `id=N+1`* strategy).
|
|
233
235
|
For per-key ordering across HA pollers, layer one of the strategies in *SKIP LOCKED and per-key ordering* (hash-routed publisher by `hash(partition_key) MOD N`, per-key advisory lease via the DB's advisory-lock primitive — PostgreSQL `pg_try_advisory_xact_lock(hashtext(partition_key))` for tx-scoped locks bound to the connection's current transaction; MySQL `GET_LOCK(name, timeout)` with explicit session-scoped semantics (release explicitly on success or stall, or rely on session close), and lock names server-wide-scoped so use a fully-bounded namespaced name `lk:<env8>:<svc8>:<purpose8>:<hash16>` (total length = 46 chars including separators, fits inside MySQL's 64-char limit; `<env8>`, `<svc8>`, `<purpose8>` are generated from the canonical environment / service / purpose identities by a **documented deterministic function** (e.g., first-8-of-base32(SHA256(canonical_identity))) or allocated from a **collision-checked registry** — human-readable abbreviations are NOT acceptable unless the registry proves uniqueness in the MySQL server-wide lock namespace; `<hash16>` is the first 16 chars of base32(SHA256(length-prefixed-encoding(canonical_partition_key, versioned_namespace))) — e.g., `SHA256(len(pk) + ':' + pk + len(ns) + ':' + ns)` or canonical JSON/CBOR over `[canonical_partition_key, versioned_namespace]`; raw `pk + ':' + ns` concatenation is **not** acceptable because real partition keys (`acme:prod`, `order:123`, user-supplied ids) can contain `:` and the resulting hash input is not injective. **Namespace-migration safety**: changing `versioned_namespace` requires either a drain / stop-the-world for the poller lane, or a dual-lock period (acquire old + new lock names in canonical order) so that mixed namespace versions across a rolling deploy / rollback cannot acquire different locks for the same partition key and publish concurrently. Without this, a version bump silently splits the per-key serialization lane.), and avoid MySQL NDB / multi-mysqld setups where `GET_LOCK` is not cluster-wide; TiDB supports MySQL-style user-level locks (`GET_LOCK`) cluster-wide in supported versions — verify timeout / deadlock semantics for the deployed TiDB version against a scenario-specific compatibility source: **pinned cluster** → checked-in version pin; **managed channel (TiDB Cloud)** → provider channel/SLA *and* current cluster version (channel alone is insufficient — the version still varies inside the channel); **rolling-upgrade fleet** → min/max active versions across the fleet plus the documented rolling-upgrade policy; **ad-hoc verification** → recorded `tidb_version()` output that includes cluster identity and timestamp. Otherwise fall back to an external coordinator (etcd lease, Redis `SET NX` with TTL) with fencing — with bounded TTL and a fencing token written with each publish; or refuse-newer-id dispatcher with abandoned-row state).
|
|
236
|
+
- **Event payload freezing** — the typed-model layer for Go event payloads is protobuf: freezing means serializing the payload to immutable bytes (marshal, or deep-copy then marshal) at the enqueue/publish boundary and binding those bytes to the envelope (`event_type`, `event_version`) — pinning the generated artifact version fixes the schema, not the instance: a queued mutable message object mutated before serialization changes the wire payload despite the pinned schema. See `protobuf-contract-architecture.md` for `message`/`enum`/`oneof` evolution rules.
|
|
234
237
|
- **Idempotency storage** — pick by impact (see *Idempotency design*):
|
|
235
238
|
- Lossy/rebuildable: Redis `SET dedup:<key> 1 EX <window> NX`, branch on success/skip.
|
|
236
239
|
- Source-of-truth: insert into a `processed_events` table with `event_id` as primary key inside the side-effect transaction; on duplicate-key error, skip. Optionally cache the recent N keys in Redis as a hot-path filter, but Redis is not the authority.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Use when designing the tenant-isolation architecture of a SaaS service: how tenants are kept apart at the data, compute, network, identity, observability, lifecycle, and compliance layers; how queries, jobs, caches, and external calls carry tenant context safely; how a tenant's data can be exported or deleted on demand; and how shared services keep cross-tenant aggregation auditable.
|
|
4
4
|
|
|
5
|
-
> **Sibling sync.** A parallel `python-service-architecture/references/multi-tenant-isolation.md` mirrors **all non-stack-specific sections** of this file. Only the *Go-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library / framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue section. Before commit, run that grep against the file's mirrored sections; zero hits required. The same gate lives in the Python sibling. Routing
|
|
5
|
+
> **Sibling sync.** A parallel `python-service-architecture/references/multi-tenant-isolation.md` mirrors **all non-stack-specific sections** of this file. Only the *Go-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library / framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue section. Before commit, run that grep against the file's mirrored sections; zero hits required. The same gate lives in the Python sibling. Routing text that differs per tree is written inline for both trees (`x.md` on the Python tree / `y.md` on the Go tree), so the mirrored bytes stay identical. Cross-file parity is machine-checked by `skill-extraction-workflow/scripts/check-parallel-stack-parity.sh` (wired into `check-ccl-skills.sh`): it diffs the mirrored regions as a byte-identical region (no normalization; tree-specific routing references are written inline for both trees), and blocks on any divergence.
|
|
6
6
|
|
|
7
7
|
> **Sanitization boundary.** Tenant identifiers, customer names, lane / region names, regulator labels, and quota numbers below are illustrative; concrete values live only in the private alias map. The list of audiences that require sanitization is **positive** (these audiences require it unless explicitly approved otherwise): external / client-facing materials, customer-specific deliverables, regulator or auditor evidence, SOC / compliance reports, procurement responses, internal compliance reviews, sales-engineering or security-questionnaire appendices, partner architecture drafts, and any document that could be forwarded to any of those. "Internal" by itself is not safety; internal documents are routinely forwarded.
|
|
8
8
|
|
|
@@ -16,7 +16,7 @@ Apply when:
|
|
|
16
16
|
|
|
17
17
|
Skip when:
|
|
18
18
|
- the service is single-tenant by deployment (per-customer dedicated stack with no shared layer); route to `platform-release-engineering/SKILL.md` and to this file's *Compliance, residency, sovereignty* section for residency commitments,
|
|
19
|
-
- the service is internal-only with a single owning team (employees of one org are not "tenants" for this purpose); identity/permission boundaries still apply but route to `api-security-boundaries.md
|
|
19
|
+
- the service is internal-only with a single owning team (employees of one org are not "tenants" for this purpose); identity/permission boundaries still apply but route to `web-framework-boundaries.md` or `api-contract-and-schema.md` on the Python tree / `api-security-boundaries.md` on the Go tree,
|
|
20
20
|
- tenant-equivalent isolation is owned entirely by a platform layer above the service (e.g., per-tenant namespace owned by the platform); route to `platform-service-connectivity/SKILL.md` for the platform contract.
|
|
21
21
|
|
|
22
22
|
## Tenant isolation tiers (decision tree)
|
|
@@ -75,6 +75,7 @@ Repo-local agent contracts (`AGENTS.md` at the repo root and in source directori
|
|
|
75
75
|
|
|
76
76
|
5. Verify at the right scope.
|
|
77
77
|
- Run focused unit tests for changed packages.
|
|
78
|
+
- When writing the test code itself (structure, naming, smells, fixtures, behavior-vs-state, coverage, isolation, table-driven parameterization), pick the matching § from the decision table in `testing-strategy/references/test-code-authoring-patterns.md`; enable the per-stack lint executors for its machine-decidable smells (conditional logic / sleep / assertion-free tests) per `testing-strategy/references/fitness-functions.md` §4.1.4 (e.g. `forbidigo` for `time.Sleep` in tests).
|
|
78
79
|
- Run integration-ish tests for DB/Redis/MQ wrappers only when environment is available.
|
|
79
80
|
- **TC traceability**: link tests via `tc.Mark(t, "TC-XX-NNN")` (helper from `test-artifact-management/references/tc_helpers/tc.go`, installed under `internal/testkit/tc/`).
|
|
80
81
|
- **`tc.Mark` MUST be the first non-comment line in the test body, BEFORE any `t.Skip` / `t.Skipf` / setup that may call `t.Fatal`** — otherwise the sidecar entry won't be written for skipped tests.
|
|
@@ -74,7 +74,7 @@ Treat the loop as: **risk-route → product/design → develop (this skill + web
|
|
|
74
74
|
|
|
75
75
|
Repo-local agent contracts (`AGENTS.md` at the repo root and in source directories) are part of the delivery contract: when a change moves a stable boundary, generated surface, workflow, or directory-local rule, update the nearest contract in the same MR and keep coverage in sync per `product-rd-workflow`'s spec / repo-contract sync gate.
|
|
76
76
|
|
|
77
|
-
When checking a mini-program project against team standards, split conformance into deterministic and agent review evidence. Deterministic checks cover project config, target/platform build commands, subpackage config, generated API/IDL client usage, environment/lane config, TC traceability, CI gates, and required request/trace identifiers in central request wrappers. Agent review checks cover page/component boundaries, host capability contracts, shared adapter safety, finite-value mapping, multi-target release risk, and whether tests/device evidence cover the shipped targets. For the deterministic executor list
|
|
77
|
+
When checking a mini-program project against team standards, split conformance into deterministic and agent review evidence. Deterministic checks cover project config, target/platform build commands, subpackage config, generated API/IDL client usage, environment/lane config, TC traceability, CI gates, and required request/trace identifiers in central request wrappers. Agent review checks cover page/component boundaries, host capability contracts, shared adapter safety, finite-value mapping, multi-target release risk, and whether tests/device evidence cover the shipped targets. For the deterministic executor list and the mini-program ESLint config enforcing the two host-boundary invariants (the `react-dom`/DOM-global ban and `TARO_ENV` adapter-layer confinement — ESLint config, not a regex source-scan), see `testing-strategy/references/fitness-functions.md` §4.1.3 (spec 006).
|
|
78
78
|
|
|
79
79
|
## Sibling Boundary With web-react-dev
|
|
80
80
|
|
|
@@ -146,6 +146,7 @@ Before editing Taro/native mini-program code, page config, host capability adapt
|
|
|
146
146
|
|
|
147
147
|
5. Verify in the right environment.
|
|
148
148
|
- Run repo formatter, typecheck/build, focused tests, and platform compile commands. For Taro, run `taro build --type <target>` for every target the change touches; one target's success is not the others' success.
|
|
149
|
+
- Test-code authoring: pick the matching § from the decision table in `testing-strategy/references/test-code-authoring-patterns.md`; smell lint: `testing-strategy/references/fitness-functions.md` §4.1.4.
|
|
149
150
|
- **TC traceability**: link tests via the `createTcSuite(test, describe)` factory wrapper. Registers at collection time so `.skip` / `.skipIf` / `.todo` still map to Bitable status. Full overloads supported: `.concurrent` / `.each` / `(name, options, fn)`. Helper from `test-artifact-management/references/tc_helpers/tc.ts`, installed under `test/tc.ts`. See `test-artifact-management/references/tc-marker-conventions.md`. Before adding tests, `grep -rn 'tcTest\|tcDescribe' src/ __tests__/` plus the sidecar `test/results/tc-map.jsonl` to check for existing coverage — extend rather than duplicate. When a TC is marked 废弃, grep both source and sidecar; follow deprecation cascade in `testing-strategy`. Tests without any TC link: prompt user only when the underlying code is also removed.
|
|
150
151
|
- **废弃级联:业务代码是否仍在用** — 小程序栈混合多种引用机制,单一 grep 不够:
|
|
151
152
|
1. TS/JS 模块:`npx madge --dependents src/<path>` 或 `grep -rEn "from ['\"][./]*<path>"`
|
|
@@ -120,7 +120,7 @@ Add domain fields with a prefix (e.g. `app_*`, `biz_*`) to avoid colliding with
|
|
|
120
120
|
- Prefer expressing an SLI as **good events / valid events** (or good windows / valid windows), per Google SRE — define valid events/windows first, then the good ratio. A latency SLI is the **proportion of requests faster than a threshold** (`count(latency ≤ T) / total`, from histogram buckets), NOT a percentile value — P95/P99 are dashboard aids, not the SLI. An error-log counter is a diagnostic signal, not an availability-SLI input (it is skewed by log sampling, dedup, and async/non-request errors).
|
|
121
121
|
- Signals you cannot compute are blind spots, not near-coverage: record each one explicitly ("can't measure X because Y" — e.g. a failure counter with no attempt total yields no error *rate*; content not collected means input-semantic drift is unmeasurable) in a blind-spot register instead of pretending coverage, so on-call never leans on a signal that does not exist.
|
|
122
122
|
- Each user-visible journey gets at least one availability SLI + one latency SLI.
|
|
123
|
-
- Set SLO targets, error budgets, and burn-rate alerts
|
|
123
|
+
- Set SLO targets, error budgets, and burn-rate alerts (SRE Workbook multiwindow tiers, derived for a 30d budget window — recompute for other periods: page 14.4× 1h/5m, page 6× 6h/30m, ticket 1× 3d/6h). Each tier MUST evaluate its long AND short window together, firing only when both burn above threshold; the short (~1/12) window makes paging stop soon after the burn stops.
|
|
124
124
|
- An SLO without an error-budget-driven release decision is decoration. See `references/sli-slo-design.md`.
|
|
125
125
|
|
|
126
126
|
### R9 — Local dev parity
|
|
@@ -61,19 +61,35 @@ Operational rule: if error budget is exhausted, the service stops shipping non-c
|
|
|
61
61
|
|
|
62
62
|
## Burn-rate alerts
|
|
63
63
|
|
|
64
|
-
Multi-window, multi-burn-rate per Google SRE Workbook:
|
|
64
|
+
Multi-window, multi-burn-rate per Google SRE Workbook ch. 5, Table 5-6 (2% of a 30d budget in 1h, 5% in 6h, 10% in 3d):
|
|
65
65
|
|
|
66
|
-
| Severity |
|
|
67
|
-
|
|
68
|
-
| Page (P0) | 1h |
|
|
69
|
-
| Page (P0) |
|
|
70
|
-
|
|
|
71
|
-
| Digest (P2) | 24h | 1x | Sustained at SLO threshold |
|
|
66
|
+
| Severity | Long window | Short window (~1/12 of long) | Burn rate | Meaning |
|
|
67
|
+
|---|---|---|---|---|
|
|
68
|
+
| Page (P0) | 1h | 5m | 14.4 | Exhausts the 30d budget in ~2 days at this rate |
|
|
69
|
+
| Page (P0) | 6h | 30m | 6 | Exhausts in 5 days |
|
|
70
|
+
| Ticket (P2) | 3d | 6h | 1 | Sustained exactly at the SLO threshold |
|
|
72
71
|
|
|
73
|
-
|
|
72
|
+
A tier fires only when **both** its long AND short window burn above the threshold. The long window gives detection over meaningful budget consumption; the short window confirms the burn is *still happening now*, so the alert resets quickly once the incident ends — with a long window alone, a resolved 1h page keeps firing for up to an hour, and a 3d ticket for days.
|
|
73
|
+
|
|
74
|
+
The Workbook's notification types are page and ticket; P0/P2 above are this platform's local severity mapping. The Workbook maps both the 1h and the 6h tier to **page** — that stays the default (at 6×, 5% of the monthly budget is already gone and the 30m window says it is still burning). Downgrading the 6h tier to a non-paging channel is allowed only under a documented, staffed response-time policy showing that tier is acted on before material further budget loss — record it as a deliberate local deviation, never as the Workbook default.
|
|
75
|
+
|
|
76
|
+
Alert query template (per tier — the recorded metric is the raw error *ratio* per window, as in the Workbook's `slo_errors_per_request:ratio_rate1h`; the burn rate is the multiplier on `(1 - SLO)`, not a pre-divided metric):
|
|
74
77
|
|
|
75
78
|
```
|
|
76
|
-
|
|
79
|
+
slo_error_ratio_1h{...} > 14.4 * (1 - <slo_target>)
|
|
80
|
+
and
|
|
81
|
+
slo_error_ratio_5m{...} > 14.4 * (1 - <slo_target>)
|
|
82
|
+
# <slo_target> is a literal scalar strictly between 0 and 1, e.g.
|
|
83
|
+
# 14.4 * (1 - 0.999) — at 1.0 there is no budget to burn and at 0 the alert
|
|
84
|
+
# can never fire (see the 100%-SLO anti-pattern below) —
|
|
85
|
+
# a bare `SLO` token would parse as a metric selector and silently match nothing.
|
|
86
|
+
# PromQL `and` intersects on the full label set. Invariant: record BOTH window
|
|
87
|
+
# rules aggregated to the SAME alert-identity label set — every label that
|
|
88
|
+
# distinguishes one alert instance from another (service, route, and region/
|
|
89
|
+
# lane if they exist), the window living in the metric NAME, never as a label.
|
|
90
|
+
# Identity labels differing between the rules → empty intersection, the page
|
|
91
|
+
# silently never fires; an `on(...)` that omits an identity label → cross-match
|
|
92
|
+
# (1h burn in one region and-ed with a 5m burn in another) → false page.
|
|
77
93
|
```
|
|
78
94
|
|
|
79
95
|
Each burn rate alert MUST link to a runbook entry.
|
|
@@ -9,3 +9,4 @@ Upstream-owner decision-surface changes record their downstream owners here
|
|
|
9
9
|
| Phase C: non-mesh collector + STRICT mTLS conflict → exception required | `platform-service-connectivity` | owns the mTLS exception YAML (sidecar/ambient PeerAuthentication/DestinationRule); obs only requires it be declared | routed | Istio PeerAuthentication docs |
|
|
10
10
|
| R8 SLI = good/valid events ratio; latency = threshold-bucket ratio | `platform-release-engineering` | error-budget release gate consumes these SLIs | unchanged (already references R8) | Google SRE Workbook |
|
|
11
11
|
| R3 instance-identity not a metric label; instrument discipline | `go-microservice-architecture` / `python-service-architecture` | language-agnostic rule; no stack-specific glue needed | unchanged | Prometheus / OTel Metrics docs |
|
|
12
|
+
| R8 burn-rate tiers aligned to Workbook Table 5-6 (long+short AND windows; 3d slow tier replaces 24h) | `platform-release-engineering` | promotion gate still consumes SLI burn-rate queries unchanged — window/threshold values are obs-owned inputs, gate shape untouched | unchanged | sre.google/workbook/alerting-on-slos (Table 5-6; short window ≈ 1/12 long) |
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md
CHANGED
|
@@ -281,7 +281,7 @@ Reused industry patterns (canary, blue-green, GitOps, control plane, lane, dcc,
|
|
|
281
281
|
- `references/env-and-lane-matrix.md` — Lane as first-class entity; long-lived envs, ephemeral lanes, canary slices, shadow traffic; cluster/region topology.
|
|
282
282
|
- `references/deploy-pipeline.md` — Build → image → manifest → apply; three control-plane patterns; CLI / UI / API surface.
|
|
283
283
|
- `references/canary-and-rollout-strategy.md` — Canary check task shape; bake window; abort thresholds; blue-green and mirror alternatives.
|
|
284
|
-
- `references/promotion-gate-and-review.md` — Evidence-driven gate; SLI
|
|
284
|
+
- `references/promotion-gate-and-review.md` — Evidence-driven gate; SLI wiring; approval workflow; audit-log; DORA release-process metrics.
|
|
285
285
|
- `references/rollback-playbook.md` — Rollback by strategy; data-migration rollback discipline; forward-compatible migration patterns; drill schedule.
|
|
286
286
|
- `references/secret-and-config-management.md` — Static / dynamic / secret tier; rotation; injection paths; scanning.
|
|
287
287
|
- `references/multi-region-and-cluster.md` — Cluster pairing; failover trigger; cross-cluster discovery/federation; DR drill.
|
|
@@ -117,6 +117,22 @@ audit_event:
|
|
|
117
117
|
|
|
118
118
|
Retention: ≥ 1 year for high/critical changes, ≥ 90 days otherwise. Audit log is queryable.
|
|
119
119
|
|
|
120
|
+
## Release-process metrics (DORA)
|
|
121
|
+
|
|
122
|
+
The audit log above is also the source for measuring the release process itself. DORA's current five-factor model (`dora.dev`):
|
|
123
|
+
|
|
124
|
+
- Throughput: **change lead time** (commit → running in prod), **deployment frequency**, **failed deployment recovery time** (successor to "MTTR"; the clock stops when user impact ends — the *current* SLI healthy again over a short confirmation window, NOT the cumulative error budget replenished — rather than when the rollback or hotfix command completed, so the audit log needs a recovery-confirmed event, not only the remediation event).
|
|
125
|
+
- Instability: **change fail rate** (deploys requiring immediate intervention — typically a rollback or hotfix, but any remediation form counts), **deployment rework rate** (unplanned deploys resulting from a production incident — "incident" meaning the production problem itself, not the existence of a formal ticket; missing paperwork doesn't exempt the deploy). Count remediation by what it *is*, not what it is named — a "roll-forward" that exists only to fix a failed deploy is a failed deploy's remediation.
|
|
126
|
+
|
|
127
|
+
Rules:
|
|
128
|
+
|
|
129
|
+
- Compute them from the control plane's own audit log (deploy / rollback / abort events with incident links) — never from a hand-maintained spreadsheet.
|
|
130
|
+
- Define the counting unit before computing, and keep that definition stable over time (changing artifact topology changes the counts, not the process — trends are only meaningful against a constant unit): for the DORA-comparable metrics, the unit is one *logical artifact deployment* to production — dedupe canary steps, per-region waves, and mechanical retries belonging to the same deployment attempt, but never collapse a *failed* attempt into its later successful retry (an attempt that reached production traffic and needed intervention stays a failed deployment even if the same release ID succeeded afterwards; otherwise change-fail rate is gameable by release-ID reuse). A purely mechanical pre-traffic failure — a control-plane error before any exposure — is pipeline noise, tracked separately, not a change failure.
|
|
131
|
+
- Behavior-exposing flag flips, config releases, and traffic shifts stay in the same audit log as release events: they count as the *remediation* (or the cause) of a deployment's failure where causally linked, but they are not deployments — folding them into the deployment-frequency denominator produces a broader custom release metric, which is fine to track but must be labeled as such, not reported as DORA.
|
|
132
|
+
- Use them as feedback on the release *process* — gate friction, batch size, rollback health — never as individual or team performance scores; scoring people on them corrupts the signal (deploys get relabeled, rollbacks get renamed "roll-forwards").
|
|
133
|
+
- If a metric cannot be computed from the audit log, that is an audit-log gap: fix the event capture, don't estimate the metric.
|
|
134
|
+
- Before publishing any of the five, verify the audit-log schema actually carries the correlation keys that metric joins on — at minimum: commit timestamp and production-exposure (traffic-reached) timestamp per deployment (lead time); a stable logical-deployment ID plus a distinct per-attempt ID (deployment frequency, and the failed-attempt/retry separation above); a causal link from each remediation event to the deployment it remediates (change fail rate); an incident link and a recovery-confirmed event (rework rate, recovery time). A metric whose keys are absent is unavailable — an audit-log gap per the rule above — not a license to join on wall-clock proximity or event-name heuristics; proximity joins are exactly how a failed attempt gets deduped into its later retry or a recovery gets inferred from an unrelated healthy reading.
|
|
135
|
+
|
|
120
136
|
## Emergency override
|
|
121
137
|
|
|
122
138
|
For incidents where the gate must be bypassed (e.g. roll out an emergency fix faster than canary allows):
|
|
@@ -31,6 +31,13 @@ Examples:
|
|
|
31
31
|
- `experiment_split = { control: 50, variant_a: 25, variant_b: 25 }`.
|
|
32
32
|
- Kill switch for a degraded-mode path.
|
|
33
33
|
|
|
34
|
+
Feature-flag lifecycle (flags are release levers, not just config values — they decouple *deploy* from *release*: code ships dark, the flag turns it on; per Fowler/Hodgson "Feature Toggles", flags are inventory with a carrying cost):
|
|
35
|
+
|
|
36
|
+
- Classify at creation: **release toggle** (transient, days–weeks), **experiment toggle** (weeks), **ops toggle** (usually short-lived — retire once operational confidence is gained; only a small, deliberate subset become long-lived **kill switches**, each with an owner and periodic review), **permission toggle** (long-lived). The categories age differently; manage them differently.
|
|
37
|
+
- Transient toggles get an owner and an expiry/cleanup task when created (an expiration date on the flag itself is a workable enforcement). A flag past its expiry is debt: surface it (report, lint, or CI warning) — every stale flag is an untested code path and a config surface someone can flip by accident.
|
|
38
|
+
- A production flag flip that exposes new behavior IS a release event, not "just config": deploy-decoupled does not mean gate-decoupled. Risk-class it like a deploy (high-blast-radius flips take the same approval path as R8), keep it in the same audit trail, stage the exposure where blast radius warrants (cohort/percentage ramp with SLI checks, per the promotion gate), and have the kill-switch/rollback path tested before the flip.
|
|
39
|
+
- Test both sides of every mutable flag deployed to production — "we never plan to flip it" is not a waiver, because the untested branch stays one operator click / stale automation run away from live traffic. The full flag combination space is untestable, so test the combinations that will actually run (current production config, the config about to go live, and the fallback/off state — per Fowler), and for cohort/percentage ramps also the *mixed* state the ramp itself creates — old and new behavior running concurrently against shared state (reader/writer compatibility, caches, queues) — before enabling the ramp. Declare dependencies between flags where one implies another, keep the concurrently-active flag count low, and retire release toggles as part of the feature's definition of done.
|
|
40
|
+
|
|
34
41
|
### Secrets
|
|
35
42
|
|
|
36
43
|
What it is: credentials, tokens, certs. Any value that, if leaked, causes a security incident.
|
|
@@ -44,6 +44,17 @@ Double retry = real bad. If mesh retries 2x and framework retries 2x, you get 4x
|
|
|
44
44
|
|
|
45
45
|
**Rule**: when enabling framework-level retry, disable mesh retry for that callee (DestinationRule `retries.attempts: 0`).
|
|
46
46
|
|
|
47
|
+
Mesh retries back off automatically (Istio/Envoy: jittered exponential backoff with a default 25ms *base* interval — fully jittered, so an actual delay can be shorter than the base; it is not a guaranteed minimum gap); framework-level retry gets no such freebie — it must implement its own jittered backoff that fits inside the caller's remaining deadline.
|
|
48
|
+
|
|
49
|
+
## Retry budget (load-proportional guard, per proxy)
|
|
50
|
+
|
|
51
|
+
Per-call retry counts bound retries *per request*; they do not bound a caller's total retry share during a partial outage — at high QPS, "2 retries each" is up to a 3× load multiplier at the exact moment the upstream is sickest. Envoy's cluster circuit breakers cap this per proxy:
|
|
52
|
+
|
|
53
|
+
- `max_retries` — max **concurrent** retries to the cluster, per priority. Retries beyond it overflow (fail fast, counted in `upstream_rq_retry_overflow`). The raw Envoy default is 3, but the control plane above Envoy may override it: Istio's `connectionPool.http.maxRetries` defaults to **2^32-1 — effectively unlimited** — so in an Istio mesh "leave it unset and rely on the default cap" is a trap. Set the limit explicitly and verify the *generated* Envoy cluster config, not the assumption.
|
|
54
|
+
- `retry_budget` — replaces the fixed cap with a load-proportional one: concurrent retries ≤ `budget_percent` (default 20%) of active + pending requests, with a `min_retry_concurrency` floor so low-traffic clusters can still retry. When set, it overrides `max_retries`. Reachability caveat: Istio's DestinationRule API exposes only `connectionPool.http.maxRetries`, NOT `retry_budget` — on plain Istio, set a finite `maxRetries` first; adopting `retry_budget` there means an EnvoyFilter, acceptable only with the *generated* cluster config verified.
|
|
55
|
+
- Know exactly what the budget bounds — and what it doesn't. It bounds **Envoy-originated, concurrent** retries, per proxy. It does NOT bound: retry attempt *rate*; **framework-level retries** (each framework attempt arrives at Envoy as a fresh request and bypasses `max_retries`/`retry_budget` entirely — a platform running framework retries needs a framework-side budget or strict per-call caps); or the **fleet aggregate** (circuit breaking is distributed, not coordinated — each sidecar enforces its own budget and floor, so aggregate retry load still scales with caller replica count). A true service-wide load bound requires callee-side protection (admission control / load shedding, owned by the service-architecture skills) on top.
|
|
56
|
+
- When tuning for a flaky dependency, set a retry budget rather than raising per-call retry counts — but pick `budget_percent` AND `min_retry_concurrency` deliberately against the callee's capacity: on a very high-QPS caller, an unexamined 20% of active requests is far looser than `max_retries: 3`, and with many low-traffic sidecars the aggregate floor (≈ replicas × `min_retry_concurrency`) dominates instead. Alert on the overflow counter: a growing overflow stat means callers are shedding retries, which is the budget doing its job; do not "fix" it by raising the cap.
|
|
57
|
+
|
|
47
58
|
## Idempotency awareness
|
|
48
59
|
|
|
49
60
|
Framework client retries MUST consider idempotency:
|
package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: python-service-architecture
|
|
3
|
-
description: Python 后端架构 / FastAPI 项目结构 / Celery worker 拆分 / Python 微服务边界 / 服务分层重构 / Python 服务重构 → design or review Python backend, service, worker, package, API contract, data ownership, reliability, async/job, and runtime boundaries. Prefer this for architecture/boundary decisions; use python-service-dev for implementation work and localized refactor (某文件/某类); multi-stage / cross-module refactor delivery re-enters product-rd-workflow.
|
|
3
|
+
description: Python 后端架构 / FastAPI 项目结构 / Celery worker 拆分 / Python 微服务边界 / 多租户隔离怎么设计 / 消费 Kafka·消息队列与事件驱动架构 / 数据平台(分库分表·读写分离·备份恢复) / 服务分层重构 / Python 服务重构 → design or review Python backend, service, worker, package, API contract, data ownership, reliability, async/job, and runtime boundaries. Prefer this for architecture/boundary decisions; use python-service-dev for implementation work and localized refactor (某文件/某类); multi-stage / cross-module refactor delivery re-enters product-rd-workflow.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Python Service Architecture
|
|
@@ -145,6 +145,10 @@ Before changing architecture guidance, contracts, service boundaries, diagrams,
|
|
|
145
145
|
- For SQLAlchemy/Django ORM, migrations, transactions, and data ownership, read `references/data-modeling-and-migrations.md`.
|
|
146
146
|
- For asyncio, blocking work, GIL, and concurrency design, read `references/async-execution-model.md`.
|
|
147
147
|
- For Celery/RQ/arq, scheduled jobs, worker leases, and batch/restart policy, read `references/background-jobs-and-scheduling.md`.
|
|
148
|
+
- For durable workflow/task state contracts — state enums, terminal states, transition ownership, duplicate handling, retry/recovery policy — read `references/workflow-state-architecture.md`.
|
|
149
|
+
- For audit logs, operation records, and change-tracking boundaries, read `references/audit-history-architecture.md`.
|
|
150
|
+
- For notification delivery, operator alerting, outbound webhooks, and realtime channel boundaries, read `references/notification-architecture.md`.
|
|
151
|
+
- For replay, shadow-traffic, and response-comparison system design, read `references/replay-comparison-architecture.md`.
|
|
148
152
|
- For event-driven architecture concerns — delivery semantics taxonomy, producer-side patterns, transactional outbox/inbox, idempotency design, partition-key ordering, schema evolution, retry/DLQ/replay strategy, fanout patterns, saga vs choreography, end-to-end "exactly-once" illusion, and Python-specific implementation glue (asyncio consumer + bounded queue, SQLAlchemy outbox poller with `SKIP LOCKED`, exception hierarchy, sync-vs-async consumer choice, graceful-shutdown order) — read `references/event-driven-architecture.md`. Stack-agnostic core sections mirror the sibling `go-microservice-architecture/references/event-driven-architecture.md`; maintainers updating those sections must update both files in the same change.
|
|
149
153
|
- For multi-tenant SaaS isolation concerns — isolation-tier decision tree (RLS / schema-per-tenant / DB-per-tenant / region-per-tenant), tenant context as a first-class value, tenant-aware data access with DB-engine enforcement, per-tenant quota and rate limit at every layer, tenant-aware observability with cardinality management, per-tenant lifecycle (provision / suspend / export / delete / retention / archive), per-tenant rollout and feature flags, cross-tenant capability gating, compliance / residency / sovereignty, migration between tiers, and Python-specific implementation glue (tenant on `contextvars.ContextVar`, FastAPI dependency + Starlette middleware, SQLAlchemy RLS session variables, connection pool reset via pool reset event, cache key helper, async task spawning with `copy_context`, async message consumer pattern, outbox tenant propagation) — read `references/multi-tenant-isolation.md`. Stack-agnostic core sections mirror the sibling `go-microservice-architecture/references/multi-tenant-isolation.md`; maintainers updating those sections must update both files in the same change.
|
|
150
154
|
- For data-platform architecture concerns — DB engine choice axis (single-instance OLTP / sharding middleware like Vitess / distributed SQL like TiDB / managed cloud DB), HA topology and failover model, read scaling and replica routing with staleness budget, sharding and resharding strategy, cross-region replication and data residency, backup with tested recovery (RPO/RTO + restore drill), cluster lifecycle (provision / scale / decommission), capacity planning (storage / IOPS / connections / latency / replica lag), fleet-wide schema-migration coordination, connection-pool and proxy topology (PgBouncer / ProxySQL / Vitess gateway), cost and efficiency, and Python-specific implementation glue (sync-vs-async driver choice, SQLAlchemy 2.x async with asyncpg/asyncmy, Alembic migrations, PgBouncer prepared-statement caveat, health-check FastAPI dependency, connection-storm mitigation) — read `references/data-platform-architecture.md`. Stack-agnostic core sections mirror the sibling `go-microservice-architecture/references/data-platform-architecture.md`; maintainers updating those sections must update both files in the same change.
|
|
@@ -26,7 +26,7 @@ Use this as the first reference for Python backend, Python microservice, AI-serv
|
|
|
26
26
|
|
|
27
27
|
## Layering Depth (Apply In Moderation)
|
|
28
28
|
|
|
29
|
-
- The transport / application / domain / infrastructure split above is the **Ports-and-Adapters / Hexagonal** idea in moderation — `infrastructure` modules are the adapters, the application-service layer hides them behind plain function / class boundaries. Useful when: an external dependency has multiple real implementations (S3 + GCS, OpenAI + local-inference + vendor-N), a domain rule is stable enough that the test fake is reusable across years, or a regulated boundary requires a single audit point. Counter-indicated when: there is one real implementation that will not change, the "adapter" is a thin pass-through with no behavior, or the layering would force every DTO through 3 mappings. Do not introduce an interface per class out of habit — the Java-style "every service has an interface, DTO mirrors entity, entity mirrors row" pattern is over-engineering in Python and produces churn without testability or substitution gains.
|
|
29
|
+
- The transport / application / domain / infrastructure split above is the **[Ports-and-Adapters / Hexagonal](https://alistair.cockburn.us/hexagonal-architecture/)** idea (Alistair Cockburn; borrowed scope: the ports/adapters placement idea only, not the full pattern vocabulary) in moderation — `infrastructure` modules are the adapters, the application-service layer hides them behind plain function / class boundaries. Useful when: an external dependency has multiple real implementations (S3 + GCS, OpenAI + local-inference + vendor-N), a domain rule is stable enough that the test fake is reusable across years, or a regulated boundary requires a single audit point. Counter-indicated when: there is one real implementation that will not change, the "adapter" is a thin pass-through with no behavior, or the layering would force every DTO through 3 mappings. Do not introduce an interface per class out of habit — the Java-style "every service has an interface, DTO mirrors entity, entity mirrors row" pattern is over-engineering in Python and produces churn without testability or substitution gains.
|
|
30
30
|
- **Functional core, imperative shell**: prefer pure functions for calculation / rule / state-transition logic; let FastAPI handlers, SQLAlchemy sessions, Celery tasks, and external clients be the I/O shell that calls them. The pure core is easy to unit-test without fixtures; the shell is small enough to integration-test directly. Conflating them — domain rules sprinkled inside ORM event listeners, or business calculation inside a Celery task — is the recurring source of "we cannot test this without a real database / queue / network." Distinct and **not** optional: domain invariants must not be hidden inside route/FastAPI handlers, repositories, external-client wrappers, SQLAlchemy/ORM event listeners or hooks, or Celery task/worker callbacks — anywhere outside the domain/service layer. That is the layering rule, independent of whether you adopt the pure-core style. Sibling: `go-microservice-architecture/references/architecture-playbook.md` ("Functional Core, Imperative Shell") carries the same principle for Go; keep the two in sync.
|
|
31
31
|
- **Shared foundation/utility packages need stricter tiering than a single service**, scaled to blast radius — a widely-imported cross-service `common` library earns it; a tiny single-consumer helper does not. Such a module inherits one import cycle or one heavyweight coupling into every consumer, and a leaf utility cannot be reused once it transitively drags in unrelated packages. Document an explicit linear package-tier order in the module itself (illustrative: `generated value-types -> constants -> generic pure utils -> logging/metrics -> framework/business adapters`); forbid earlier tiers importing later tiers (one-directional, acyclic). Default to sibling independence so each leaf stays independently importable; when one same-tier package genuinely needs another, extract the shared piece down a tier rather than copy-pasting or adding a micro-tier. Generated pure contracts may be imported anywhere; generated clients carry transport deps and belong in the adapter tier only; all generated code is regenerate-only. Enforce with `import-linter` `layers` + `independence` contracts (a `src/` layout aids packaging isolation but does not enforce tiering) rather than review memory; a new cross-tier or sibling import is an architecture-review item.
|
|
32
32
|
- **How a layer boundary is enforced is itself an architecture decision**, with a strength ladder: physical package boundary (a separate distribution package whose violation is an import or packaging error), `import-linter` contracts in CI (above), then review convention — in decreasing strength; prefer mechanisms where a violation is a CI error, not a review comment. For a core where a frozen contract must coexist with continuous evolution (a gateway data plane, a billing domain), prefer the physical boundary, and add a deterministic digest/conformance anchor as machine proof that evolution has not touched the frozen surface. Record why the chosen strength is enough (cost versus strength); a weaker tier is a documented tradeoff, not a default. Sibling: `go-microservice-architecture/references/architecture-playbook.md` ("Dependency Direction") carries the same rule for Go; keep the two in sync.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# Audit And History Architecture
|
|
2
|
+
|
|
3
|
+
Use this when designing audit logs, operation records, resource history, or change tracking. Implementation mechanics live in `python-service-dev/references/audit-history-patterns.md`.
|
|
4
|
+
|
|
5
|
+
Sibling note: `go-microservice-architecture/references/audit-history-architecture.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
|
|
6
|
+
|
|
7
|
+
## Audit Boundary
|
|
8
|
+
|
|
9
|
+
- Decide which operations require durable audit and which can use best-effort activity logs.
|
|
10
|
+
- Define actor identity, service identity, resource scope, resource id, operation name, result, timestamp, trace id, and canonical error fields.
|
|
11
|
+
- Keep end user, client application, service caller, and operator identity separate.
|
|
12
|
+
- Use stable operation names rather than function names that change during refactors.
|
|
13
|
+
- Define retention, privacy, redaction, query authorization, export, and deletion policy up front (multi-tenant deletion modes: `multi-tenant-isolation.md`).
|
|
14
|
+
|
|
15
|
+
## Data Shape
|
|
16
|
+
|
|
17
|
+
- Store selective before/after summaries or field-level diffs when needed; never raw secrets, tokens, passwords, signatures, or unrelated payloads.
|
|
18
|
+
- Serialize request parameters and result data through redaction helpers.
|
|
19
|
+
- Prefer append-only audit records; corrections are new records unless a legal deletion policy requires removal.
|
|
20
|
+
- For event-derived history, define lag, rebuild, reconciliation, and source event retention.
|
|
21
|
+
|
|
22
|
+
## Write Path
|
|
23
|
+
|
|
24
|
+
- Critical audit belongs in the same transaction or outbox as the state change when correctness depends on it.
|
|
25
|
+
- Best-effort audit is acceptable only when explicitly non-critical and observable through logs or metrics on write failure.
|
|
26
|
+
- Async audit pipelines need bounded queues, retry policy, backpressure or an explicit drop policy, and shutdown flush; the full detached side-path posture is canonical in `python-service-dev/references/async-and-worker-patterns.md`.
|
|
27
|
+
|
|
28
|
+
## Query Path
|
|
29
|
+
|
|
30
|
+
- Query APIs need resource-scope authorization, pagination, time-range filters, and redaction on read.
|
|
31
|
+
- High-volume audit stores need partitioning or retention windows before launch.
|
|
@@ -4,7 +4,7 @@ Use when designing the data-platform substrate of a service or service-fleet: DB
|
|
|
4
4
|
|
|
5
5
|
This complements `data-modeling-and-migrations.md` (which owns schema, index, transaction, outbox, and per-service migration concerns): this file owns the **substrate** that schema and queries sit on. Load both when designing a new data-bound service or auditing an existing one.
|
|
6
6
|
|
|
7
|
-
> **Sibling sync.** A parallel `go-microservice-architecture/references/data-platform-architecture.md` mirrors **all non-stack-specific sections** of this file. Only the *Python-specific implementation patterns* section diverges by stack. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library/framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue.
|
|
7
|
+
> **Sibling sync.** A parallel `go-microservice-architecture/references/data-platform-architecture.md` mirrors **all non-stack-specific sections** of this file. Only the *Python-specific implementation patterns* section diverges by stack. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library/framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue. Tree-specific routing references are written inline for both trees so the mirrored bytes stay identical; cross-file parity is machine-checked by `skill-extraction-workflow/scripts/check-parallel-stack-parity.sh` (wired into `check-ccl-skills.sh`), which diffs the mirrored regions byte-for-byte (no normalization) and blocks on any divergence.
|
|
8
8
|
|
|
9
9
|
> **Sanitization boundary.** Vendor names (PostgreSQL, MySQL, Vitess, TiDB, CockroachDB, Aurora, Cloud Spanner, AlloyDB, Cloud SQL, DynamoDB, RDS Proxy, PgBouncer, ProxySQL, S3, Glacier, gp3, io2, etc.) below are illustrative; concrete topology choices, region names, cluster identifiers, and capacity numbers live only in the maintainer's private alias map. The sanitization audience list is positive (external / client / regulator / SOC / procurement / internal-compliance / sales-engineering / partner draft / forwardable-internal); sanitize before any document leaves the implementation team's approved audience.
|
|
10
10
|
>
|
|
@@ -6,7 +6,7 @@ This complements `async-execution-model.md` (concurrency model choice), `backgro
|
|
|
6
6
|
|
|
7
7
|
> **Conforms to the parallel-stack references pattern.** This file follows the layout documented in `skill-extraction-workflow/references/parallel-stack-references-pattern.md`: mirrored stack-agnostic core (when-applies through operations checklist), stack-specific implementation patterns section, and the embedded `### Mirrored-section grep gate` at the end of the stack-glue. The sibling `go-microservice-architecture/references/event-driven-architecture.md` mirrors the same structure. Either this file or `multi-tenant-isolation.md` may be used as a template for new parallel-stack extractions; multi-tenant additionally demonstrates the `## Topic-extension backlog` H2 for topic-wider-than-loop cases.
|
|
8
8
|
|
|
9
|
-
> **Sibling sync.** A parallel `go-microservice-architecture/references/event-driven-architecture.md` mirrors **all non-stack-specific sections** of this file (when-applies/not-applies, delivery semantics, event vs command vs query, idempotency, outbox, ordering, schema evolution, retry/DLQ/replay, backpressure, fanout, saga, end-to-end exactly-once, anti-patterns, operations checklist). Only the *Python-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change to prevent drift.
|
|
9
|
+
> **Sibling sync.** A parallel `go-microservice-architecture/references/event-driven-architecture.md` mirrors **all non-stack-specific sections** of this file (when-applies/not-applies, delivery semantics, event vs command vs query, idempotency, outbox, ordering, schema evolution, retry/DLQ/replay, backpressure, fanout, saga, end-to-end exactly-once, anti-patterns, operations checklist). Only the *Python-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change to prevent drift. Tree-specific routing references are written inline for both trees so the mirrored bytes stay identical; cross-file parity is machine-checked by `skill-extraction-workflow/scripts/check-parallel-stack-parity.sh` (wired into `check-ccl-skills.sh`), which diffs the mirrored regions byte-for-byte (no normalization) and blocks on any divergence.
|
|
10
10
|
|
|
11
11
|
> **Sanitization boundary.** The named brokers (Kafka, Pulsar, RabbitMQ, NATS JetStream, Redis Streams) and libraries below are concrete examples for **internal** implementation guidance, scoped to the implementation team's approved audience. Before this file (or excerpts) is copied into any document leaving that audience — external / client-facing materials, customer-specific deliverables, regulator or auditor evidence packages, SOC / compliance reports, procurement responses, or partner architecture appendices — replace the named choices with generic categories (`the broker`, `a partitioned log`, `a confirm-mode AMQP queue`) unless the vendor selection is already approved for disclosure to that specific audience.
|
|
12
12
|
|
|
@@ -20,8 +20,8 @@ Apply when the service:
|
|
|
20
20
|
|
|
21
21
|
Skip when the service:
|
|
22
22
|
- only does in-process pub-sub or fire-and-forget logging,
|
|
23
|
-
- uses synchronous HTTP/RPC with no durable async boundary (use `api-contract-and-schema.md` instead),
|
|
24
|
-
- uses a job queue purely for in-tenant background work where loss is acceptable (use `background-jobs-and-scheduling.md` or `batch-and-pipeline-architecture.md`).
|
|
23
|
+
- uses synchronous HTTP/RPC with no durable async boundary (use `api-contract-and-schema.md` on the Python tree / `protobuf-contract-architecture.md` on the Go tree instead),
|
|
24
|
+
- uses a job queue purely for in-tenant background work where loss is acceptable (use `background-jobs-and-scheduling.md` or `batch-and-pipeline-architecture.md` on the Python tree / `notification-architecture.md` or `bulk-workflow-architecture.md` on the Go tree).
|
|
25
25
|
|
|
26
26
|
## Delivery semantics taxonomy
|
|
27
27
|
|
|
@@ -139,7 +139,7 @@ Consumer lag, broker queue depth, and producer rate are the three backpressure s
|
|
|
139
139
|
- *Ordered partitioned log* (Kafka, Pulsar key-shared, NATS JetStream ordered consumer): **one serial work lane per assigned partition/key**. Feed each partition into its own bounded work-queue + single worker, or process messages serially within the partition's poll loop. Feeding multiple ordered partitions into a single shared work-queue + worker pool loses per-partition ordering, and a slow message on one partition can starve cold partitions or let later offsets overtake earlier ones. The bounded-queue-plus-pool shape is correct for unordered work queues; not for ordered partitioned logs.
|
|
140
140
|
- *Rebalance handling for partition-assigned consumers* — when a Kafka / Pulsar key-shared / similar consumer group rebalances and a partition is revoked, "one lane per partition" is unsafe without explicit rebalance discipline. The revoked owner must (1) stop fetching from the partition immediately, (2) drain or cancel its in-flight lane (await handler completion to a bounded deadline, or cancel with an explicit `partial-failure` disposition), (3) commit or abort offsets according to the handler outcome (commit only completed offsets; do not commit `last poll` blindly), and (4) be fenced so it cannot still publish a side effect after the new owner has started — typically by tagging each in-flight message with the assignment epoch and refusing side effects whose epoch is stale. Without fencing, the new owner and the old owner can process the same business key concurrently; per-partition ordering at steady state is meaningless if the rebalance window allows concurrent processing.
|
|
141
141
|
- Avoid unbounded worker-per-message fanout in all cases (the stack-glue section names the specific anti-pattern API).
|
|
142
|
-
- **Producer-side** — when the broker buffer fills (Kafka producer queue, RabbitMQ unconfirmed-publishes limit, NATS slow-consumer warning), block the producer's caller with a bounded wait or shed load at the producer entry point. Never block forever; surface a typed
|
|
142
|
+
- **Producer-side** — when the broker buffer fills (Kafka producer queue, RabbitMQ unconfirmed-publishes limit, NATS slow-consumer warning), block the producer's caller with a bounded wait or shed load at the producer entry point. Never block forever; surface a typed error after a bounded wait so upstream can backpressure further.
|
|
143
143
|
- **Cross-service** — a slow consumer is an upstream producer's problem to know about. Consumer lag must be exposed as a metric and alerted; producers cannot fix what they cannot see.
|
|
144
144
|
|
|
145
145
|
## Fanout patterns
|
|
@@ -187,6 +187,8 @@ Stack-agnostic recipe; document each clause for every event-driven boundary that
|
|
|
187
187
|
|
|
188
188
|
If any clause is missing, the boundary is at-least-once with duplicates. Tell consumers honestly.
|
|
189
189
|
|
|
190
|
+
External grounding (adopted in part): this recipe is an instance of the end-to-end argument — [Saltzer, Reed & Clark, *End-to-End Arguments in System Design*, ACM TOCS 2(4), 1984](https://web.mit.edu/Saltzer/www/publications/endtoend/endtoend.pdf) — a function that "can completely and correctly be implemented only with the knowledge and help of the application standing at the endpoints of the communication system" cannot be delegated to the communication layer, and broker-level transactional features are that paper's "incomplete version … useful as a performance enhancement", never the end-to-end guarantee. Borrowed scope: the placement argument only; the five-clause recipe and the atomic-domain boundary are this skill's own operational criteria.
|
|
191
|
+
|
|
190
192
|
## Anti-patterns
|
|
191
193
|
|
|
192
194
|
- **Post-commit publish (durable cross-process)** — publishing the event after the DB transaction commits, without an outbox, when consumers are in another process. A crash between commit and publish silently drops the event. In-process, same-instance, rebuildable post-commit hooks are not this anti-pattern.
|
|
@@ -226,6 +228,7 @@ These are stack-localized recipes that implement the stack-agnostic patterns abo
|
|
|
226
228
|
- *Ordered partitioned log*: one `asyncio.Task` per assigned partition that processes serially, or a per-partition bounded `asyncio.Queue` with a single worker task. A shared `asyncio.Event` triggers shutdown; never feed multiple ordered partitions into a shared worker pool.
|
|
227
229
|
- Avoid `asyncio.create_task(handle(msg))` inside a `for msg in consumer` loop — unbounded fanout will OOM under load and discards ordering.
|
|
228
230
|
- **Sync vs async consumers** — if the handler is CPU-bound or calls a sync DB driver (psycopg2, sync SQLAlchemy), use a thread pool or process pool; do not block the event loop. For mixed workloads, route async-friendly handlers to the asyncio worker pool and CPU/sync handlers to a `ProcessPoolExecutor` or to Celery/RQ.
|
|
231
|
+
- **Event payload freezing** — the typed-model layer for Python event payloads is the service's Pydantic/typed model: freezing means serializing to immutable bytes at the enqueue/publish boundary (`model_dump_json()` or an explicit serializer writing into the outbox row) and binding those bytes to the envelope (`event_type`, `event_version`) — never enqueue a mutable model instance that later code can mutate before the poller publishes; the schema version alone does not freeze the instance.
|
|
229
232
|
- **Context propagation** — extract correlation id, trace context, lane/env from message headers into `contextvars` at the consumer boundary. OpenTelemetry's `aiokafka`/`pika` instrumentations restore the trace context automatically; verify they are wired in `observability-and-ops.md`'s OTel/startup section, not in `async-execution-model.md`.
|
|
230
233
|
- **Outbox poller with SQLAlchemy** — an asyncio task that runs a **claim → commit → publish → mark-sent** loop, not "publish inside the DB tx" (a broker call inside a SQLAlchemy session held open for the broker round-trip violates the short-transaction rule in `data-modeling-and-migrations.md`):
|
|
231
234
|
1. **Claim tx (short)**: `BEGIN; SELECT … FROM outbox WHERE sent_at IS NULL AND (processing_until IS NULL OR processing_until < NOW()) ORDER BY id LIMIT N FOR UPDATE SKIP LOCKED; UPDATE outbox SET processing_until = NOW() + lease, owner = :owner WHERE id IN (…); COMMIT;` — the row is now leased to this poller; session closes immediately.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Use when designing the tenant-isolation architecture of a SaaS service: how tenants are kept apart at the data, compute, network, identity, observability, lifecycle, and compliance layers; how queries, jobs, caches, and external calls carry tenant context safely; how a tenant's data can be exported or deleted on demand; and how shared services keep cross-tenant aggregation auditable.
|
|
4
4
|
|
|
5
|
-
> **Sibling sync.** A parallel `go-microservice-architecture/references/multi-tenant-isolation.md` mirrors **all non-stack-specific sections** of this file. Only the *Python-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library / framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue section. Before commit, run that grep against the file's mirrored sections; zero hits required. The same gate lives in the Go sibling. Routing
|
|
5
|
+
> **Sibling sync.** A parallel `go-microservice-architecture/references/multi-tenant-isolation.md` mirrors **all non-stack-specific sections** of this file. Only the *Python-specific implementation patterns* section diverges by stack. Maintainers updating any mirrored section here must update the sibling in the same change. The mirrored sections stay free of three categories of stack-specific token: DB-engine-specific syntax, runtime/concurrency-mechanic names, and library / framework API names. The concrete token list and grep command live in the *Mirrored-section grep gate* subsection at the end of this file's stack-glue section. Before commit, run that grep against the file's mirrored sections; zero hits required. The same gate lives in the Go sibling. Routing text that differs per tree is written inline for both trees (`x.md` on the Python tree / `y.md` on the Go tree), so the mirrored bytes stay identical. Cross-file parity is machine-checked by `skill-extraction-workflow/scripts/check-parallel-stack-parity.sh` (wired into `check-ccl-skills.sh`): it diffs the mirrored regions as a byte-identical region (no normalization; tree-specific routing references are written inline for both trees), and blocks on any divergence.
|
|
6
6
|
|
|
7
7
|
> **Sanitization boundary.** Tenant identifiers, customer names, lane / region names, regulator labels, and quota numbers below are illustrative; concrete values live only in the private alias map. The list of audiences that require sanitization is **positive** (these audiences require it unless explicitly approved otherwise): external / client-facing materials, customer-specific deliverables, regulator or auditor evidence, SOC / compliance reports, procurement responses, internal compliance reviews, sales-engineering or security-questionnaire appendices, partner architecture drafts, and any document that could be forwarded to any of those. "Internal" by itself is not safety; internal documents are routinely forwarded.
|
|
8
8
|
|
|
@@ -16,7 +16,7 @@ Apply when:
|
|
|
16
16
|
|
|
17
17
|
Skip when:
|
|
18
18
|
- the service is single-tenant by deployment (per-customer dedicated stack with no shared layer); route to `platform-release-engineering/SKILL.md` and to this file's *Compliance, residency, sovereignty* section for residency commitments,
|
|
19
|
-
- the service is internal-only with a single owning team (employees of one org are not "tenants" for this purpose); identity/permission boundaries still apply but route to `web-framework-boundaries.md` or `api-contract-and-schema.md
|
|
19
|
+
- the service is internal-only with a single owning team (employees of one org are not "tenants" for this purpose); identity/permission boundaries still apply but route to `web-framework-boundaries.md` or `api-contract-and-schema.md` on the Python tree / `api-security-boundaries.md` on the Go tree,
|
|
20
20
|
- tenant-equivalent isolation is owned entirely by a platform layer above the service (e.g., per-tenant namespace owned by the platform); route to `platform-service-connectivity/SKILL.md` for the platform contract.
|
|
21
21
|
|
|
22
22
|
## Tenant isolation tiers (decision tree)
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Notification Architecture
|
|
2
|
+
|
|
3
|
+
Use this when designing notification delivery, operator alerting, outbound webhooks, or realtime client channels. Implementation mechanics live in `python-service-dev/references/notification-patterns.md`.
|
|
4
|
+
|
|
5
|
+
Sibling note: `go-microservice-architecture/references/notification-architecture.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
|
|
6
|
+
|
|
7
|
+
## Boundary
|
|
8
|
+
|
|
9
|
+
- Classify notifications as user-facing, operator-facing, integration callbacks, or internal alerts.
|
|
10
|
+
- Define whether each notification is critical, retryable, idempotent, and auditable.
|
|
11
|
+
- Keep notification templates and delivery endpoints in config or a template store, not inline code.
|
|
12
|
+
- Recipient and endpoint selection is scoped by resource ownership and environment.
|
|
13
|
+
- Notification payloads use safe summaries, not raw request bodies or secrets.
|
|
14
|
+
|
|
15
|
+
## Delivery Policy
|
|
16
|
+
|
|
17
|
+
- Critical notifications need a durable outbox, retry, idempotency key, and terminal delivery state.
|
|
18
|
+
- Best-effort alerts may use async delivery, but failures must be observable.
|
|
19
|
+
- Delivery clients need timeout, status-code validation, response body size limit, and rate limit.
|
|
20
|
+
- Configurable delivery endpoints are an SSRF trust boundary: architecture names the scheme/port policy, private/link-local/metadata blocking, redirect policy, and whether deliveries route through a constrained egress proxy.
|
|
21
|
+
- Retries use bounded backoff and stop on permanent errors.
|
|
22
|
+
- Duplicate delivery must be acceptable to receivers or prevented with stable dedupe keys.
|
|
23
|
+
|
|
24
|
+
## Observability
|
|
25
|
+
|
|
26
|
+
- Track sent, failed, retried, dropped, and suppressed counts by notification type.
|
|
27
|
+
- Include trace/log id and config/template version in delivery logs.
|
|
28
|
+
- Alert storms need grouping, throttling, and suppression policy.
|
|
@@ -16,7 +16,7 @@ Use this for pyproject, uv/poetry/pip, lockfiles, tooling, containers, and deplo
|
|
|
16
16
|
- Define entrypoint: `uvicorn`/`gunicorn`, framework command, worker command, scheduler command, or CLI.
|
|
17
17
|
- Define process model: worker count, async event loop, thread/process workers, memory limits, graceful shutdown, and health probes.
|
|
18
18
|
- Release readiness includes migrations, startup validation, smoke tests, canary, rollback, and observability checks.
|
|
19
|
-
- **ASGI server choice has expanded beyond `uvicorn` / `gunicorn+uvicorn` / `hypercorn`** — `granian` (emmett-framework, Rust-based) is the current credible high-throughput alternative for ASGI services, supporting ASGI/3, RSGI, WSGI, HTTP/1, HTTP/2, TLS, WebSockets (HTTP/3 planned per the project README's "eventually 3" roadmap note — verified not shipped as of
|
|
19
|
+
- **ASGI server choice has expanded beyond `uvicorn` / `gunicorn+uvicorn` / `hypercorn`** — `granian` (emmett-framework, Rust-based) is the current credible high-throughput alternative for ASGI services, supporting ASGI/3, RSGI, WSGI, HTTP/1, HTTP/2, TLS, WebSockets (HTTP/3 planned per the project README's "eventually 3" roadmap note — re-verified not shipped as of Aug 2026, granian 2.8.x). Per the Granian project's own `benchmarks/vs.md` and third-party load-test repos (e.g., `piccolo-orm/asgi_server_performance`, `synodriver/asgi-server-benchmark`), ASGI echo on 10KB payload typically lands granian > uvicorn-httptools > hypercorn by roughly the ratios 58k / 51k / 8k RPS in April-2026-era runs; file-serving gap is wider (granian ~47k vs uvicorn ~18k via `pathsend`). Treat the absolute numbers as benchmark-snapshot-specific; re-run against your workload before basing a switch on them. Architecture impact: when serving throughput is the binding constraint, granian can buy headroom without rewriting the **plain ASGI path**. **"Without rewriting" caveats**: granian's worker / process model differs from `gunicorn+uvicorn` fork-based workers (Rust runtime + Python interpreters with different lifecycle hooks); ASGI lifespan events, contextvars propagation across worker boundaries, custom signal handlers, prometheus/metrics exporters tied to uvicorn internals, and ASGI middleware that depends on uvicorn-specific behavior all need smoke-testing on granian before a switch. If the team plans to adopt granian's bespoke RSGI protocol for max performance (rather than ASGI), application code that uses ASGI-specific middleware, ASGI scope manipulation, or third-party ASGI libraries WILL need rewriting — RSGI is a different protocol, not a faster ASGI. Trade-offs: smaller operational maturity, fewer community recipes, Rust-runtime-on-the-side observability differs from a pure-Python server. **Choose uvicorn** for ecosystem maturity, broad reference material, and known operational patterns; **choose granian** when (a) profiled benchmarks on your workload show uvicorn saturation, (b) the team has Rust-toolchain debugging capacity, (c) the deployment story can absorb a less-common runtime. Hypercorn remains the choice when HTTP/2 + ASGI under pure-Python ops matters more than peak throughput.
|
|
20
20
|
- **Python runtime version baseline (2025-2026)**: Python 3.13 (released October 2024) ships **experimental** free-threaded build per PEP 703 — GIL-disabled, ~40% single-threaded perf hit per python.org "What's New in 3.13" notes, used at the team's risk for parallel-CPU workloads. Python 3.14 (released 7 October 2025) advances free-threading to **supported (Phase II of PEP 703)** per PEP 779 — meaning the free-threaded build is a first-class supported configuration, NOT that it is the default Python build or the default production choice. Per the python.org free-threading howto, the single-threaded penalty narrowed to ~5-10% (specializing adaptive interpreter re-enabled thread-safely); PEP 803 defines the `abi3t` stable ABI for free-threaded C extensions. Architecture impact: for services where parallel CPU work matters (in-process ML inference fan-out, heavy parsing, parallel compression), 3.14 free-threading is the first version where adoption is **a defensible experiment for a selected service**, not yet a defensible default. Pre-flight burn-in required before any production switch: (a) C-extension readiness — all extensions in the service's dependency tree must declare free-threading support; many popular extensions (numpy, pandas, lxml, psycopg native bits, asyncpg native bits, pillow, cryptography) were still mid-migration at 2026-Q1, verify per-version; mixing GIL-only and free-threading-aware extensions in one process is unsupported; (b) GC and runtime behavior under sustained threading load differs from GIL build — measure tail latency, memory residency, and CPU efficiency on the actual workload; (c) debugger / profiler ergonomics (gdb, pdb, py-spy, scalene, prometheus exporters) may have rough edges on free-threaded builds; (d) library-level thread-safety: code paths that were "implicitly safe because of the GIL" can race in free-threaded mode (singletons built at import time, module-level mutable caches, third-party libraries that rely on GIL-protected dict mutation). For pure I/O-bound services, stay on stock GIL build — the 5-10% overhead is pure cost. For services on 3.13 or older, treat free-threading as opt-in research, not default.
|
|
21
21
|
|
|
22
22
|
## Topic-extension backlog
|