@ccoalm/ccl-skills 0.6.2 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (62) hide show
  1. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +11 -0
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-unverified-cli-flag.sh +309 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_unverified_cli_flag.sh +483 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +5 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +2 -1
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +1 -1
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +1 -1
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +14 -11
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +2 -2
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +1 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +2 -1
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +1 -1
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +25 -9
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +1 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +1 -1
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +16 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +7 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +11 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +5 -1
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -1
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/audit-history-architecture.md +31 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +7 -4
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +2 -2
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/notification-architecture.md +28 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +1 -1
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/replay-comparison-architecture.md +28 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/workflow-state-architecture.md +39 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +6 -6
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +8 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/audit-history-patterns.md +29 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +16 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +25 -1
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/notification-patterns.md +40 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +1 -1
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/replay-comparison-patterns.md +30 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +48 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +10 -1
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +2 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +48 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +21 -2
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +8 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +5 -4
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +15 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +2 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +24 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-parallel-stack-parity.sh +119 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_parallel_stack_parity.sh +183 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +2 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +1 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +3 -4
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +16 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +1 -1
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +16 -5
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +5 -3
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +16 -6
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/self-benchmark-baseline.md +37 -0
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +1 -0
  61. package/dist/assets/release.json +115 -50
  62. package/package.json +1 -1
@@ -0,0 +1,28 @@
1
+ # Replay Comparison Architecture
2
+
3
+ Use this when designing replay, shadow-traffic, response-comparison, or migration-verification systems. Implementation mechanics live in `python-service-dev/references/replay-comparison-patterns.md`.
4
+
5
+ Sibling note: `go-microservice-architecture/references/replay-comparison-architecture.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Execution Model
8
+
9
+ - Separate capture, replay, comparison, storage, and reporting.
10
+ - Replays need durable job state, concurrency limit, timeout, delay policy, rate limit, target environment or lane, and cancellation policy.
11
+ - Captured input is redacted and bounded by size before storage.
12
+ - Replay requests preserve only approved headers and metadata; never replay credentials blindly.
13
+ - Shadow execution must not commit side effects unless the target is explicitly isolated (environment, lane, rollback, or dry-run).
14
+
15
+ ## Comparison Policy
16
+
17
+ - Define comparator selection by method, content type, schema, or route.
18
+ - Generic JSON comparison supports ignored fields, custom field comparators, numeric tolerance, null handling, array handling, and type-mismatch reporting.
19
+ - Diff output includes field path, original value summary, replay value summary, diff type, score, and ignored status.
20
+ - Comparison thresholds are config-driven and versioned.
21
+ - Positive, negative, and neutral diffs are classified only when the product has a defensible definition.
22
+
23
+ ## Retention And Reporting
24
+
25
+ - Store aggregate job counts and paginated diff details separately.
26
+ - Define retention for captured requests, replay responses, and diff artifacts.
27
+ - Expose summary metrics: total, processed, success, failed, diff count, similarity, p95 replay latency, and error categories.
28
+ - Treat replay as a confidence signal, not an automatic release approval, unless acceptance gates are explicit.
@@ -0,0 +1,39 @@
1
+ # Workflow State Architecture
2
+
3
+ Use this when designing durable task, async workflow, import/export, backfill, or scheduled processing state. Implementation mechanics live in `python-service-dev/references/state-machine-task-patterns.md`.
4
+
5
+ Sibling note: `go-microservice-architecture/references/workflow-state-architecture.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## State Contract
8
+
9
+ - Define state enum, terminal states, allowed transitions, retry policy, and ownership before handlers are written.
10
+ - Keep allowed transitions in one table, diagram, or policy function so handlers do not invent their own rules.
11
+ - Separate state, progress, result pointer, error metadata, retry metadata, and audit metadata.
12
+ - Persist the task record before starting expensive work, background execution, or external artifact generation so process failure leaves an inspectable, repairable state.
13
+ - Define duplicate handling for every entrypoint: create, start, process, fail, succeed, cancel, and retry.
14
+ - Transitions that affect durable truth need compare-and-update, a transaction, or a lock.
15
+ - Completion events are published after the durable state update and are idempotent.
16
+
17
+ ## Processing Semantics
18
+
19
+ - Create is idempotent when the caller supplies an idempotency key or natural unique key.
20
+ - Start claims ownership before expensive side effects; process re-checks state after ownership is acquired.
21
+ - Success persists result data before publishing completion.
22
+ - Failure records canonical error code, safe message, retryable flag, retry count, and trace/log id.
23
+ - Cancellation defines whether in-flight work is interrupted, allowed to finish, or marked for later stop.
24
+
25
+ ## Retry And Recovery
26
+
27
+ - Retry threshold and backoff belong in policy/config, not inline literals.
28
+ - Terminal-state recovery is explicit; ordinary retries must not move successful or cancelled tasks.
29
+ - Delayed events and queue messages re-check current state before acting.
30
+ - Process restart leaves enough durable state to resume, retry, or safely skip.
31
+ - Background workers and scheduled cleaners need exception recovery and a bounded cursor/batch model; long-lived loops need a stoppable schedule or an explicit process-lifetime owner (`async-execution-model.md`, `background-jobs-and-scheduling.md`).
32
+
33
+ ## Acceptance Checks
34
+
35
+ - Illegal transitions are rejected or no-op according to a documented policy.
36
+ - Duplicate messages, delayed messages, and concurrent processors produce one durable outcome.
37
+ - Every terminal state includes enough result or error context for support and reconciliation.
38
+ - Expensive background work cannot start without an inspectable durable task record, or carries an explicitly documented non-durable rationale.
39
+ - Worker, scheduler, and cleaner paths have exception recovery, bounded scan/batch behavior, and a visible recovery/repair result.
@@ -75,14 +75,10 @@ Repo-local agent contracts (`AGENTS.md` at the repo root and in source directori
75
75
 
76
76
  5. Verify at the right scope.
77
77
  - Run focused pytest tests for changed packages.
78
+ - When writing the test code itself (structure, naming, smells, fixtures, behavior-vs-state, coverage, isolation, parameterization), pick the matching § from the decision table in `testing-strategy/references/test-code-authoring-patterns.md`; enable the per-stack lint executors for its machine-decidable smells (conditional logic / sleep / assertion-free tests) per `testing-strategy/references/fitness-functions.md` §4.1.4 (e.g. Ruff `TID251` banning `time.sleep`).
78
79
  - Run async tests with the repo's configured `pytest-asyncio` mode.
79
80
  - Run integration tests only when required services and credentials are available.
80
- - **TC traceability**: link tests via `@pytest.mark.tc("TC-XX-NNN")` marker. Registers at collection time so `@skip` / `@skipif` / fixture failures still map to Bitable status. Needs the `tc` plugin from `test-artifact-management/references/tc_helpers/tc.py` loaded via `addopts = -p tc`. See `test-artifact-management/references/tc-marker-conventions.md`. Before adding tests, `grep -rn 'pytest\.mark\.tc' tests/` plus the sidecar `test/results/tc-map.jsonl` to check for existing coverage — extend rather than duplicate. When a TC is marked 废弃, grep both source and sidecar; follow deprecation cascade in `testing-strategy`. Tests without any TC link: prompt user only when the underlying code is also removed.
81
- - **废弃级联:业务代码是否仍在用** — grep 只找出"测试函数引用了什么 import"是第一步;判断"该 import 是否还有其他 caller"才能定生死。Python 顺序:
82
- 1. 看测试体导入的模块:`grep -E "^(from |import )" tests/test_<x>.py`
83
- 2. 对每个产品模块(非 stdlib / 非测试 helper),找全仓库 caller:`grep -rEn "from <pkg>\.<mod>|import <pkg>\.<mod>" --include='*.py' --exclude-dir=tests`
84
- 3. 零产品 caller → 同 commit 删该模块 + 测试;有产品 caller → 测试目标仍在用,不删测试(若 TC 已废弃但代码活,先确认产品决策)
85
- 4. 边界:动态 import(`importlib.import_module("...")`)grep 抓不到;含 reflection 的代码人工确认;DI/插件注册(`@register` 装饰器)的产品代码需查注册表而非 import
81
+ - **TC traceability and the deprecation cascade** are mandatory when the repo tracks test cases in Bitable: before adding a test, check existing TC coverage; before deleting one, run the caller-liveness sequence. Mechanics (marker registration, coverage grep, the four-step 废弃级联, dynamic-import boundaries) live in `references/testing-and-quality-patterns.md` (TC Traceability And Deprecation Cascade).
86
82
  - Run ruff, mypy/pyright, formatting, and codegen/migration checks when the repo uses them.
87
83
  - Keep fast tests deterministic; isolate live infrastructure, long sleeps, generated files, and external credentials behind markers.
88
84
 
@@ -135,6 +131,10 @@ Repo-local agent contracts (`AGENTS.md` at the repo root and in source directori
135
131
  - For asyncio, blocking work isolation, concurrency limits, and cancellation, read `references/async-and-worker-patterns.md`.
136
132
  - For Redis, cache, locks, idempotency, counters, and rate limits, read `references/redis-cache-lock-patterns.md`.
137
133
  - For Celery/RQ/arq, queues, scheduled tasks, and job execution, read `references/background-job-patterns.md`.
134
+ - For durable task state machines, status transitions, leases, terminal states, and scheduled repair jobs, read `references/state-machine-task-patterns.md`.
135
+ - For audit logs, operation records, resource history, and change tracking, read `references/audit-history-patterns.md`.
136
+ - For notification delivery, operator alerts, outbound webhooks, and realtime client channels, read `references/notification-patterns.md`.
137
+ - For replay jobs, shadow execution, response comparison, and migration verification, read `references/replay-comparison-patterns.md`.
138
138
  - For external HTTP clients, SDKs, generated clients, service discovery, and dependency adapters, read `references/dependency-client-patterns.md`.
139
139
  - For errors, exception mapping, response envelopes, and validation errors, read `references/error-handling-patterns.md`.
140
140
  - For logs, metrics, traces, health checks, and instrumentation, read `references/observability-implementation-patterns.md`.
@@ -11,6 +11,14 @@ Use this for Python implementation around LLM/RAG/inference calls after `llm-inf
11
11
  - For CPU/GPU-heavy local inference, isolate concurrency and memory limits.
12
12
  - Use fakes for provider tests and mark live provider tests explicitly.
13
13
 
14
+ ## Streaming And Session Mechanics
15
+
16
+ - Reject a new turn while a prior turn for the same session/conversation is in-flight: check-and-claim the session at the boundary (lock or idempotency marker) and return a typed busy error; do not silently interleave two generations into one conversation state.
17
+ - Persist partial state on a timer during long streams (partial transcript, token counts) so a crash mid-stream can resume or at least account for cost; the cadence is a product decision, the mechanism belongs here.
18
+ - Close stream readers deterministically on every exit path — client disconnect, deadline, terminal error — via async context managers or cancellation handlers, or provider connections leak until pool exhaustion.
19
+ - Record token/cost usage after completion, or on terminal failure with the partial count, never only at request start; tie the usage record to the same request/session id the logs carry.
20
+ - Mark terminal vs recoverable stream states explicitly (completed / cancelled / provider-error / resumable); a consumer that cannot distinguish them retries unresumable streams.
21
+
14
22
  ## Do Not
15
23
 
16
24
  - Encode prompt policy, retrieval strategy, evaluation rubric, or model routing here; route those decisions to `llm-inference-integration`.
@@ -0,0 +1,29 @@
1
+ # Audit And History Patterns
2
+
3
+ Use this when implementing audit logs, operation records, resource history, or change tracking.
4
+
5
+ Sibling note: `go-microservice-dev/references/audit-history-patterns.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Audit Record Shape
8
+
9
+ - Capture actor, service identity, resource scope, resource id, operation name, operation type, request id, trace/log id, timestamp, result, and canonical error code.
10
+ - Store safe before/after summaries or field-level diffs when needed.
11
+ - Never store raw secrets, tokens, passwords, signatures, or credentials.
12
+ - Serialize params and data through redaction helpers; do not dump whole request or response objects.
13
+ - Use stable operation names, not function names that change during refactors.
14
+
15
+ ## Write Path
16
+
17
+ - Audit writes for critical operations belong in the same transaction or outbox as the state change when correctness depends on them (`sqlalchemy-and-migrations-patterns.md` Outbox section).
18
+ - Best-effort audit is acceptable only when explicitly non-critical; catch and log write failures with a failure metric — audit write errors must be observable even when they do not fail the main request.
19
+ - Async audit writers follow the detached side-path posture in `async-and-worker-patterns.md` (whitelisted correlation only, accept boundary at durable commit, per-attempt reservation for non-rollbackable external effects, bounded queue, serialized accept-vs-drain shutdown, loss policy matched to data class). For billing/ledger/usage records with no tested reconstruction source, use a same-transaction outbox or backpressure instead of drop — that posture is canonical there; do not restate it here.
20
+
21
+ ## Query Path
22
+
23
+ - Audit query APIs need resource-scope authorization, pagination, time-range filters, and redaction on read.
24
+ - Default ordering is deterministic — usually newest first with a stable tie-breaker.
25
+ - Large audit tables need partitioning, retention, or archive strategy before high-volume launch.
26
+
27
+ ## Tests
28
+
29
+ - Test redaction, critical-write rollback behavior, best-effort write failure visibility, exception recovery, pagination, scope filtering, and time-range boundaries.
@@ -11,6 +11,22 @@ Use this for Celery, RQ, arq, APScheduler, queue consumers, and scheduled tasks.
11
11
  - Store status for user-visible jobs.
12
12
  - Protect singleton jobs with locks or scheduler guarantees.
13
13
 
14
+ ## Tool-Specific Caveats
15
+
16
+ Defaults shift across major versions; verify against the installed tool version's docs before relying on any default named here.
17
+
18
+ - Celery ack semantics: the default early ack loses a task on worker crash; `acks_late=True` moves the ack to after completion, so a crash mid-task causes redelivery — pair it with idempotent task bodies, and decide `task_reject_on_worker_lost` deliberately rather than by default.
19
+ - Celery visibility timeout (Redis/SQS-style brokers): must exceed the longest task runtime plus retry backoff, or the broker redelivers a still-running task and it executes concurrently with itself.
20
+ - Celery prefetch and worker lifecycle: set `worker_prefetch_multiplier=1` for long tasks (default prefetch head-of-line blocks the queue behind one slow task); use `worker_max_tasks_per_child` to recycle leaky workers.
21
+ - Celery beat is a single point of scheduling: run exactly one beat instance, or guard schedule dispatch with a distributed lock; two beats double-fire every schedule.
22
+ - RQ: a queued job silently expires when its `ttl` passes before a worker picks it up, and `job_timeout` kills execution past the budget; failed jobs land in the failed registry and requeueing is an explicit operation, not automatic.
23
+ - arq: async-native — `max_tries` bounds retries, `job_timeout` bounds execution, `defer_by`/`defer_until` schedule, and results live in Redis only for the `keep_result` TTL; treat result reads after that window as misses, not errors.
24
+
25
+ ## Transactional Enqueue Boundary
26
+
27
+ - Enqueueing from inside an open DB transaction is a dual write: the broker publish does not roll back with the transaction. Either enqueue through an outbox row committed with the business write (see `python-service-architecture/references/event-driven-architecture.md` and the outbox section of `sqlalchemy-and-migrations-patterns.md`), or enqueue after commit and accept the crash window between commit and enqueue with a documented reconciliation path.
28
+ - The inverse ordering — enqueue first, then commit — hands the worker a job for state that may never commit; workers must re-read durable state, not trust the enqueue payload as proof the write happened.
29
+
14
30
  ## Request Boundary
15
31
 
16
32
  - Do not hide long work behind a synchronous request unless the timeout budget proves it is safe.
@@ -1,6 +1,6 @@
1
1
  # Batch And Artifact Patterns
2
2
 
3
- Use this for import/export scripts, backfills, reports, generated files, CSV/XLSX/PDF artifacts, and repair tools.
3
+ Use this for import/export scripts, backfills, reports, generated files, CSV/XLSX/PDF artifacts, and repair tools. For durable job status transitions, duplicate delivery, retry, and cancellation, `state-machine-task-patterns.md` is the canonical guide; this file covers row/file handling, execution, artifacts, and reports.
4
4
 
5
5
  ## Implementation
6
6
 
@@ -11,3 +11,27 @@ Use this for import/export scripts, backfills, reports, generated files, CSV/XLS
11
11
  - For object or artifact migration, record success, error, skipped, and conflict rows in replayable output so reruns can resume or audit decisions without re-discovering every item.
12
12
  - Make output artifact paths, object storage keys, retention, and download permissions explicit.
13
13
  - Test parsing, validation, edge rows, and retry/resume behavior.
14
+
15
+ ## Import Pipeline
16
+
17
+ - Represent each input row as a typed row object with row number, source name (sheet/tab/file part), raw values, normalized values, and an error list.
18
+ - Validate file type, size, sheet/partition count, header shape, start row, and column-count bounds before row parsing; normalize cells at the boundary (trim, pad missing optional columns, parse accepted list separators, reject unsupported encodings).
19
+ - Static row validation accumulates row errors; ordinary row errors do not fail the whole file. Cross-row validation detects duplicates and conflicts, then annotates every affected row.
20
+ - Enrichment resolves external names/codes to IDs through typed lookup caches; mapping misses become row errors unless the workflow is explicitly fail-fast.
21
+ - Commit only valid rows, chunk large batches, and preserve per-row outcome. Persist the job record (status, progress, counts, error-file pointer, actor, source file, idempotency key) per `state-machine-task-patterns.md`.
22
+
23
+ ## Error Reports And Export
24
+
25
+ - Error reports preserve original row order with source row, source name, and failure reasons; generate with streaming writers; upload to object storage with bounded retention and store only the object key or signed reference. Represent an empty report explicitly rather than failing generation. No secrets or sensitive raw payloads in reports.
26
+ - Exports use cursor or id-window pagination (never offset for large datasets), streaming writers, and the same authorization/resource-scope filters as API reads; long exports run as jobs with progress and a downloadable artifact status.
27
+
28
+ ## Cursor-Based Data Migration
29
+
30
+ - Read source rows by monotonic cursor or id window; discover min/max before splitting work and persist chunk boundaries so failed chunks are replayable.
31
+ - Insert into the target with idempotent create/upsert semantics before deleting from the source; delete only the cursor window that was successfully written, and record a replay/reconciliation path when source and target do not share one transaction boundary.
32
+ - Emit per-worker progress (current cursor, chunk range, migrated count, speed, terminal error); progress is best-effort — final migration state must be durable or replayable, and a closed progress channel is not proof all chunks succeeded.
33
+
34
+ ## Tests
35
+
36
+ - Test empty file, hidden/empty sheets, malformed headers, short/long rows, duplicate rows, mapping misses, partial success, error-report generation, slice idempotency, cancellation, and retry after crash.
37
+ - For cursor migrations, test empty range, cursor boundary inclusiveness, idempotent insert, delete-after-insert ordering, worker error propagation, and replaying a partial chunk.
@@ -0,0 +1,40 @@
1
+ # Notification Patterns
2
+
3
+ Use this when implementing notification clients, webhooks, operator alerts, outbound callbacks, or delivery jobs. For durable delivery state, retries, and terminal-state behavior, also apply `state-machine-task-patterns.md`. (Inbound callback *verification* is `public-api-security-patterns.md`; this file owns the outbound side.)
4
+
5
+ Sibling note: `go-microservice-dev/references/notification-patterns.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Client
8
+
9
+ - Wrap each delivery provider behind a `typing.Protocol` with a `send(message) -> Result` surface (`dependency-client-patterns.md` owns the Protocol-vs-ABC rule).
10
+ - Build requests with deadlines, content type, status-code validation, and response body size limits.
11
+ - Treat empty message batches as no-op.
12
+ - Parse provider responses into canonical success, retryable error, and permanent error.
13
+ - **Outbound URLs are an SSRF surface** when endpoints are tenant- or operator-configurable: enforce approved schemes/ports; resolve the hostname once, validate the resolved IP against private/link-local/loopback/cloud-metadata ranges, and **connect to that same validated IP** (connect-by-IP with Host/SNI set from the hostname, or a resolver-pinning client hook) — validate-then-let-the-client-re-resolve is bypassed by DNS rebinding returning a public IP to the validator and a private one to the connection; re-run the resolve-validate-pin cycle on every redirect hop, or disable redirects; prefer routing deliveries through a constrained egress proxy.
14
+
15
+ ## Message Shape
16
+
17
+ - Include notification type, recipient or endpoint reference, template id/version, actor or service identity, trace/log id, safe summary, and dedupe key.
18
+ - Render templates through typed variables, not string concatenation of raw objects.
19
+ - Redact secrets, credentials, tokens, signatures, and large payloads.
20
+
21
+ ## Async Delivery
22
+
23
+ - Critical delivery uses an outbox table or durable queue with retry count, next retry time, terminal status, and last error.
24
+ - Best-effort alerts may run asynchronously but must recover exceptions and emit logs/metrics on failure (detached-spawn firewall per `async-and-worker-patterns.md`).
25
+ - Use bounded concurrency and backoff; group or throttle repeated alerts by stable fingerprint.
26
+ - Flush or drain delivery workers during graceful shutdown when messages are critical.
27
+
28
+ ## Realtime Client Channels
29
+
30
+ - For WebSocket/SSE realtime channels, authenticate and resolve app, tenant, and user context before accepting the connection.
31
+ - Keep connection identity in a concurrency-safe registry keyed by the smallest delivery scope, and remove the client on disconnect via `finally`/context-manager cleanup.
32
+ - Rebuild initial client state from the durable store on first connect or cache miss; send a snapshot after connect, then typed/versioned deltas.
33
+ - Maintain cached unread/count state with explicit TTL refresh and atomic increments; a missing realtime cache must not create a durable-count side effect.
34
+ - Queue consumers treat disconnected clients as a no-op delivery outcome after durable storage succeeds.
35
+
36
+ ## Tests
37
+
38
+ - Test empty batch, timeout, non-2xx response, malformed response, retryable vs permanent classification, redaction, template rendering, dedupe key, throttling, and shutdown drain.
39
+ - Test the SSRF boundary: loopback/private/link-local/metadata-range targets rejected, disallowed scheme/port rejected, redirect to a blocked range rejected, and the pin exercised — assert the connection is made to the validated IP (fake resolver returning different answers on first and second resolution must not reach the second answer).
40
+ - Test realtime connect, duplicate connection, disconnect cleanup, cache-miss bootstrap, disconnected-delivery no-op, and write failure on a closed socket.
@@ -19,7 +19,7 @@ Use this for implementing Python public APIs, partner integrations, signed callb
19
19
 
20
20
  ## Safer Composition (Python 3.14+)
21
21
 
22
- - **Python 3.14 (released 7 October 2025) introduced t-strings via PEP 750** — template literals using `t"..."` syntax that evaluate to `string.templatelib.Template` objects rather than `str`, giving a consuming function access to interpolated values BEFORE they are combined into a string. The standard library exposes `string.templatelib.Template`. **t-strings are inert by themselves — they are NOT automatic injection protection by syntax**. A bare t-string only carries the raw values plus their surrounding template; safety arrives only when a trusted consumer library is t-string-aware and validates/escapes each interpolation according to its target language (SQL, shell, HTML). Writing `t"SELECT * FROM users WHERE id = {user_id}"` and passing it to an ORM/driver that has not added Template support gains nothing over an f-string — and passing it to a function expecting `str` triggers `Template.__str__()` which raises by default, surfacing the misuse rather than silently producing the wrong result. As of 2026-Q1: most popular template engines (Jinja2, Django templates) and most DB drivers do NOT yet consume `Template` directly — verify the specific library's Template support before relying on this rule. **PEP 787 (safer subprocess via t-strings) is currently Deferred to Python 3.15** per peps.python.org — PEP authors are pursuing experimental t-string subprocess work outside the stdlib through 3.14 beta before re-proposing for 3.15. Do NOT assume `subprocess.run(t"...")` works safely in 3.14; for shell/subprocess composition on 3.14, keep using `shlex.quote()` + argument-list form (`subprocess.run(["cmd", arg])`) until PEP 787 or its equivalent lands. For Python ≤3.13 targets, t-strings are unavailable — keep `shlex.quote()` / parameterized DB queries / framework-native HTML escaping; t-strings are a 3.14-and-later opt-in, not a backport.
22
+ - **Python 3.14 (released 7 October 2025) introduced t-strings via PEP 750** — template literals using `t"..."` syntax that evaluate to `string.templatelib.Template` objects rather than `str`, giving a consuming function access to interpolated values BEFORE they are combined into a string. The standard library exposes `string.templatelib.Template`. **t-strings are inert by themselves — they are NOT automatic injection protection by syntax**. A bare t-string only carries the raw values plus their surrounding template; safety arrives only when a trusted consumer library is t-string-aware and validates/escapes each interpolation according to its target language (SQL, shell, HTML). Writing `t"SELECT * FROM users WHERE id = {user_id}"` and passing it to an ORM/driver that has not added Template support gains nothing over an f-string — and passing it to a function expecting `str` triggers `Template.__str__()` which raises by default, surfacing the misuse rather than silently producing the wrong result. As of 2026-Q1: most popular template engines (Jinja2, Django templates) and most DB drivers do NOT yet consume `Template` directly — verify the specific library's Template support before relying on this rule. **PEP 787 (safer subprocess via t-strings) remains a Draft targeting Python 3.15** (postponed from 3.14; per peps.python.org, re-verified 2026-08 — not yet accepted) — PEP authors are pursuing experimental t-string subprocess work outside the stdlib before re-proposing for 3.15. Do NOT assume `subprocess.run(t"...")` works safely in 3.14; for shell/subprocess composition on 3.14, keep using `shlex.quote()` + argument-list form (`subprocess.run(["cmd", arg])`) until PEP 787 or its equivalent lands. For Python ≤3.13 targets, t-strings are unavailable — keep `shlex.quote()` / parameterized DB queries / framework-native HTML escaping; t-strings are a 3.14-and-later opt-in, not a backport.
23
23
 
24
24
  ## Signature And Replay Verification
25
25
 
@@ -0,0 +1,30 @@
1
+ # Replay Comparison Patterns
2
+
3
+ Use this when implementing replay jobs, shadow execution, response comparison, migration verification, or diff reports. Reuse `state-machine-task-patterns.md` for job transitions and terminal-state behavior.
4
+
5
+ Sibling note: `go-microservice-dev/references/replay-comparison-patterns.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Replay Jobs
8
+
9
+ - Persist replay jobs with status, target, config version, total/processed/success/failed/diff counts, timeout, concurrency, rate limit, and retention.
10
+ - Use bounded workers with deadlines (`async-and-worker-patterns.md`) for replay execution.
11
+ - Preserve only allowlisted headers and metadata; never replay credentials blindly.
12
+ - Redact captured requests and responses before storage; bound captured payload size.
13
+ - Replay targets must not commit side effects unless isolated by environment, lane, or explicit dry-run mode. Transaction rollback isolates only the local database write — a replayed handler can still send webhooks, publish messages, call payment providers, or write other datastores while its DB transaction rolls back; those adapters need environment/lane isolation or dry-run stubs of their own before rollback counts as isolation.
14
+
15
+ ## Comparator Design
16
+
17
+ - Define a comparator `typing.Protocol` such as `compare(original, replay) -> Result`; select comparators by method, route, content type, or schema version.
18
+ - Generic JSON comparators normalize strings containing JSON, Pydantic models, dicts, lists, and scalars before comparing.
19
+ - Support ignored fields by exact path/field name only when documented, loaded from config or comparator registration — never hard-coded in the generic comparator.
20
+ - Support custom field comparators for tolerances, unordered collections, timestamps, generated IDs, and approximate numerics.
21
+
22
+ ## Diff Result Shape
23
+
24
+ - Include field name, field path, original value summary, replay value summary, diff type, diff score, ignored flag, and metadata; bound value summaries so diff records do not store huge payloads.
25
+ - Record total compared fields, diff count, similarity, threshold, comparator name, and comparator version.
26
+ - Distinguish missing value, added value, type mismatch, length mismatch, value mismatch, and custom comparison.
27
+
28
+ ## Tests
29
+
30
+ - Test None/None, None/value, value/None, JSON-string normalization, model-to-dict conversion, array length mismatch, missing keys, numeric tolerance, ignored fields, custom comparator, threshold behavior, redaction, storage pagination, and replay cancellation.
@@ -0,0 +1,48 @@
1
+ # State Machine And Task Patterns
2
+
3
+ Use this when implementing durable tasks, async workflows, scheduled jobs, imports, exports, and retryable processors. This is the canonical guide for durable state transitions and terminal-state behavior; generic worker/async mechanics stay in `async-and-worker-patterns.md` and `background-job-patterns.md`.
4
+
5
+ Sibling note: `go-microservice-dev/references/state-machine-task-patterns.md` carries the Go rendering of the same discipline; content is adapted per stack and kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## State Model
8
+
9
+ - Define states as a domain-owned `Enum` (per the skill entrypoint's finite-value rule) with documented terminal states.
10
+ - Define allowed transitions in one table or policy function; do not scatter status checks across handlers.
11
+ - Terminal states reject ordinary processing and failure transitions.
12
+ - Include retry count, last error, start time, finish time, progress, result pointer, and idempotency key when tasks are inspectable.
13
+ - Store progress separately from state; progress is best-effort while state transitions are durable.
14
+ - Create the durable task row before launching background work, queueing expensive processing, or generating external artifacts — a crash must leave an inspectable, repairable record.
15
+
16
+ ## Transition Implementation
17
+
18
+ - Guard transitions with compare-and-update (`UPDATE … SET status=:new WHERE id=:id AND status=:expected` and check rowcount), `SELECT … FOR UPDATE`, or a Redis lock (`redis-cache-lock-patterns.md`) before processing.
19
+ - Lease or lock owners must be unique at the runtime-instance level, not only the service level: include pod/host identity plus process id or a random instance id so replicas cannot claim or complete each other's work.
20
+ - A lock alone does not fence: a paused worker whose lease expired can resume and commit after a new owner took over — unique owner names do not prevent this. Every durable write under a lease carries a fencing check at the store it writes to: a monotonic fencing token compared there, or owner/version compare-and-update (`… WHERE id=:id AND owner=:me AND version=:v`), so a stale holder's write fails instead of overwriting the new owner's state.
21
+ - A store-side check fences only that store: a stale worker can pass the DB compare-and-update, pause, lose its lease, and then fire a NON-transactional external effect (a charge, a send) after the new owner completed. External side effects under a lease go through intent-then-execute (`sqlalchemy-and-migrations-patterns.md` outbox / the intent pattern in `python-service-architecture/references/event-driven-architecture.md`): the durable intent row is claimed with the fencing check, and the provider call carries an idempotency key bound to that intent. This fully fences **only when the provider actually enforces the key** (rejects or deduplicates a replay). For providers that cannot — SMTP, plain webhooks, any at-least-once send with no key support — no fencing eliminates the stale-execution window: shrink it with a claim re-check immediately before the call, then classify the effect honestly as at-least-once with a duplicate-visible reconciliation path, and record the residual duplicate window instead of claiming exactly-once.
22
+ - Re-read task state under the lock before side effects.
23
+ - Make duplicate delivery normal: already-successful is success or no-op; already-terminal is no-op or a typed conflict depending on the caller's contract.
24
+ - Start transitions move pending work to processing before expensive work begins.
25
+ - Failure transitions capture canonical error code, safe message, retryable flag, retry count, and last trace/log id.
26
+ - Success transitions persist the result pointer or summary before publishing completion events; completion events are idempotent.
27
+
28
+ ## Async Processing
29
+
30
+ - Prefer queue/worker execution (Celery/RQ/arq per `background-job-patterns.md`) for work that must survive process restart.
31
+ - Delayed events and queue messages re-check current task state when consumed.
32
+ - Retry thresholds and backoff live in config or state policy, not inline literals.
33
+ - A broad `except Exception` at the task boundary converts the task to failed or retryable-failed per the retry policy (never swallow `asyncio.CancelledError` — re-raise it after cleanup, per `async-and-worker-patterns.md`).
34
+ - Work that produces a file, report, or media artifact persists the object key or result pointer before the success transition and exposes a retryable failure state when upload/finalization fails.
35
+
36
+ ## Scheduled Repair Jobs
37
+
38
+ - Scheduled checkers that repair stuck tasks need a job name, interval, distributed lock key, lock lease, per-run max batch size, and max execution time.
39
+ - A job-level lock prevents duplicate scans; a task-level lock or compare-and-update prevents duplicate repair of one record.
40
+ - Time-window selection is explicit, bounded, and based on durable timestamps (`updated_at`), not process memory.
41
+ - Repair actions re-read current task state before changing status or emitting side effects.
42
+ - Emit metrics for scan count, claimed count, repaired count, skipped count, lock conflict, failure, and duration.
43
+
44
+ ## Tests
45
+
46
+ - Test illegal transition, duplicate event, concurrent processing (two claimers, one winner), already-terminal, retry threshold, cancellation, exception-to-failed conversion, and completion-event idempotency.
47
+ - Test lease expiry with a stale worker resuming after a new owner claimed the task: the stale worker's durable write and side effect must be rejected by the fencing check.
48
+ - Test scheduled repair lock conflict, stale-window selection, per-task duplicate prevention, and repeated-run idempotency.
@@ -14,6 +14,15 @@ Use this for pytest, async tests, fixtures, fakes, ruff, mypy/pyright, and CI qu
14
14
  - Use markers to separate unit, integration, API, contract, E2E, live-infra, failure-mode, and drill tests.
15
15
  - Use `pytest-asyncio` for async code according to repo configuration.
16
16
  - Avoid long sleeps and live credentials in default tests.
17
+
18
+ ## TC Traceability And Deprecation Cascade
19
+
20
+ - **TC traceability**: link tests via `@pytest.mark.tc("TC-XX-NNN")` marker. Registers at collection time so `@skip` / `@skipif` / fixture failures still map to Bitable status. Needs the `tc` plugin from `test-artifact-management/references/tc_helpers/tc.py` loaded via `addopts = -p tc`. See `test-artifact-management/references/tc-marker-conventions.md`. Before adding tests, `grep -rn 'pytest\.mark\.tc' tests/` plus the sidecar `test/results/tc-map.jsonl` to check for existing coverage — extend rather than duplicate. When a TC is marked 废弃, grep both source and sidecar; follow deprecation cascade in `testing-strategy`. Tests without any TC link: prompt user only when the underlying code is also removed.
21
+ - **废弃级联:业务代码是否仍在用** — grep 只找出"测试函数引用了什么 import"是第一步;判断"该 import 是否还有其他 caller"才能定生死。Python 顺序:
22
+ 1. 看测试体导入的模块:`grep -E "^(from |import )" tests/test_<x>.py`
23
+ 2. 对每个产品模块(非 stdlib / 非测试 helper),找全仓库 caller:`grep -rEn "from <pkg>\.<mod>|import <pkg>\.<mod>" --include='*.py' --exclude-dir=tests`
24
+ 3. 零产品 caller → 同 commit 删该模块 + 测试;有产品 caller → 测试目标仍在用,不删测试(若 TC 已废弃但代码活,先确认产品决策)
25
+ 4. 边界:动态 import(`importlib.import_module("...")`)grep 抓不到;含 reflection 的代码人工确认;DI/插件注册(`@register` 装饰器)的产品代码需查注册表而非 import
17
26
  - For scenario tests, use `testing-strategy` to build the scenario matrix first; in Python services, map scenarios to unit tests, API/contract tests, integration tests, workflow tests, or marked E2E/live-infra tests instead of putting every case into a slow end-to-end suite.
18
27
 
19
28
  ## Test Target Split
@@ -58,4 +67,4 @@ Use this for pytest, async tests, fixtures, fakes, ruff, mypy/pyright, and CI qu
58
67
  - Keep generated or vendored code excluded according to repo policy.
59
68
  - Use coverage gates when the repo already enforces them or the change is high risk.
60
69
  - **`ruff` (Astral) is the current default-recommendation single tool for lint + format + import sort** — per Astral docs, replaces Flake8 (and dozens of plugins), Black, isort, pydocstyle, pyupgrade, autoflake with one Rust binary; ~900 lint rules (vs Pylint ~409, with ~209 rule overlap); ~tens to hundreds of times faster than the tools it replaces (e.g., the FAQ cites 250k-LOC codebase 2.5min on Pylint vs 0.4s on Ruff). Auto-fix for most violations. Use ruff for new projects by default; existing projects can migrate piecewise (linter and formatter are independent — adopt one without the other). **Pylint coverage gap to audit before retiring it**: ruff does NOT replicate Pylint's deeper semantic / data-flow analysis (Pylint's inference engine can catch e.g. argument-count mismatches across complex call chains, unreachable-after-mutation patterns, type-confusion in dynamically typed code that doesn't reach the type checker), Pylint's broader dead-code detection (beyond ruff's import-level checks — unused-method, unused-attribute, unused-private-member with project-aware reasoning), design smells (cyclomatic-complexity bands, too-many-arguments / too-many-locals / too-many-branches thresholds), duplicate-code detection (`similarities` checker), and any custom in-project Pylint plugin or rule. Strategy: ruff + a type checker (mypy / pyright / ty) replaces Pylint for ~80% of teams; teams relying on the specific Pylint capabilities above should keep Pylint as a slower secondary gate or migrate the equivalents (e.g., use `radon` for complexity, `vulture` for dead-code, a type-checker-level config for design smells) before retiring.
61
- - **`ty` (Astral, Beta announced 2025-12-16 per Astral's own announcement) is the current state of Rust-based type-checker work** — per the Astral announcement, on full check "consistently between 10x and 60x faster than mypy and Pyright" depending on workload, with editor incremental recompute on the order of ~80x faster than Pyright on a touched load-bearing file in benchmark workloads. Read these as Astral-published numbers on Astral-chosen benchmarks; actual speedup on a given codebase varies. Notable features: first-class intersection types, advanced narrowing, reachability analysis. Status: **Beta, not GA** — appropriate for piloting on a separate CI job alongside the existing mypy / pyright / basedpyright gate, NOT for replacing the production type-check gate on its own yet. **Dual-gate operational cost to size before adopting**: running ty plus mypy/pyright produces a "triage tax" — divergent type narrowing (one tool reports, the other doesn't), different stub-package assumptions (typeshed version skew), duplicated CI latency, and a "fix one, break the other" churn pattern that slows MRs. Make the pilot a non-blocking informational job and budget time for periodic finding-diff review rather than treating both as equal blocking gates. Watch for ty GA before flipping defaults. basedpyright remains a viable strict-mypy-like alternative if Pyright's strictness gaps matter and ty isn't ready.
70
+ - **`ty` (Astral, Beta announced 2025-12-16 per Astral's own announcement) is the current state of Rust-based type-checker work** — per the Astral announcement, on full check "consistently between 10x and 60x faster than mypy and Pyright" depending on workload, with editor incremental recompute on the order of ~80x faster than Pyright on a touched load-bearing file in benchmark workloads. Read these as Astral-published numbers on Astral-chosen benchmarks; actual speedup on a given codebase varies. Notable features: first-class intersection types, advanced narrowing, reachability analysis. Status: **Beta, not GA** (re-verified 2026-08: still beta; Astral projects a stable release within 2026, with the beta→stable gap focused on stability, typing-spec completeness, and first-class Pydantic/Django support) — appropriate for piloting on a separate CI job alongside the existing mypy / pyright / basedpyright gate, NOT for replacing the production type-check gate on its own yet. **Dual-gate operational cost to size before adopting**: running ty plus mypy/pyright produces a "triage tax" — divergent type narrowing (one tool reports, the other doesn't), different stub-package assumptions (typeshed version skew), duplicated CI latency, and a "fix one, break the other" churn pattern that slows MRs. Make the pilot a non-blocking informational job and budget time for periodic finding-diff review rather than treating both as equal blocking gates. Watch for ty GA before flipping defaults. basedpyright remains a viable strict-mypy-like alternative if Pyright's strictness gaps matter and ty isn't ready.
@@ -61,6 +61,8 @@ This skill coordinates gates; it does **not** itself authorize merge, tag push,
61
61
  | Reset dev/test-like branches | Yes | Target/env refs, before SHAs, dry-run/plan, force-with-lease semantics |
62
62
  | Post-merge cleanup of the merged temp feature branch (worktree/local/remote) | No — covered by the user's merge authorization (`worktree-isolation` 收尾) | The authorized MR/PR read back as merged at the current head SHA and target; the live remote source ref is absent (already cleaned by the platform) or still equals the merged MR source head (moved → preserve and ask, remote path only — eligible local cleanup proceeds per `worktree-isolation`); no other open or plan-declared MR/PR still consumes the source branch; source branch is a temp feature branch (unclear role → preserve and ask); mechanics/safety rails per `worktree-isolation` |
63
63
 
64
+ **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work.
65
+
64
66
  ## Minimal checklist
65
67
 
66
68
  - [ ] Intended scope, base/head refs, and production target identified.
@@ -34,6 +34,51 @@ Its fix is therefore not "read the real artifact" but **walk a seed list of chan
34
34
 
35
35
  **Failure shape:** an agent researching one organization from public sources declares "the public channels are mined out" five separate times; each time the user pushes back, a **previously unconsidered channel class** lands major evidence (an archive route around a bot-block, a regulator's bulk dataset, a mandatory-disclosure regime, another jurisdiction's statutory filings, a filing the organization voluntarily publishes on its own site). No artifact was hiding these — they were never on a list. The domain-side channel taxonomy lives with the research owner (`skills/multi-perspective-research/references/public-disclosure-channels.md`, repo-root relative); what belongs here is the recurrence memory that an exhaustion claim over a *channel set* needs a walked class list, not more effort inside the known classes.
36
36
 
37
+ ## Variant (d) — the OTHER agent line's store (the corpus you cannot see from where you stand)
38
+
39
+ More than one agent runs against the same repository, and each keeps its lessons in its own host's store. Those stores do not see each other. An extraction that enumerates "the lesson corpus" from the host it happens to be running on has enumerated **one line's half of it** and will report exhaustion over a corpus it never touched — the core trap, with host boundary as the thing doing the masking.
40
+
41
+ Measured in round 059: one line held hundreds of `symptom -> cause -> fix` records while another held a few dozen distilled entries, in different trees under different roots with no reference between them. The figures are deliberately not pinned here — these stores are LIVE and grow while you work, so a count written into a durable document is a timestamped measurement that starts rotting the moment it lands; a re-count during this round's own self-check already disagreed with every figure it had recorded hours earlier. State DEPTH, which is stable, not SIZE, which is not. Neither line's store contained a pointer to the other. The consequence is not abstract: one class recurred **eight times across independent sessions** on the line that lacked the abstraction, while the line that had already generalized it — same repository, same weeks — carried the remedy the whole time.
42
+
43
+ The gate: before any complete / exhausted / no-gap claim over a *lesson or failure* corpus, the landing that makes the claim carries, **for each agent line, a verdict** — `covered` / `digest-only` / `unavailable` / `excluded` — and under it one row per store backing that verdict. A per-store list without a per-line verdict is the failure this gate exists to stop: a reader cannot tell from four mixed rows whether the line was actually covered. A line whose store you cannot read is `unavailable` with the remediation attempted, never an omission; a line you deliberately leave out is `excluded` with the reason. Locate the stores by enumerating the hosts that actually work the repository, not by assuming the layout your own host uses — the other line's may be one large per-session file where yours is a directory of entries, and a path-shaped guess finds nothing and reads as absence (the same "keyed by something other than what you are asking about" failure the blocked-source ladder names).
44
+
45
+ **Where these rows live — the split that makes the gate checkable.** Two different things get confused here. The *provenance* — the corpus itself, and the project, repository, branch, ticket and contributor identifiers inside it — is what the extraction lifecycle keeps in per-host scratch, and it stays there. An agent product name and its own config path (`~/.codex/memories/`) are neither: they identify a tool this repository already names throughout, carry nothing about any project, and a leakage audit over them comes back clean. The *coverage rows* are store location, status, depth and exclusion reason; they carry no corpus content, so nothing requires them to be private, and holding them in scratch makes the gate unverifiable by the only people who will ever read the claim. So: coverage rows go **where the claim is made**, alongside it and in the same artifact; provenance stays in scratch and the coverage rows point at it by role, never by real path. If a line's store cannot be described without naming something private, that is a sanitization problem to solve in the row, not a reason to move the row. The round that introduced this variant records its own, so the rule is demonstrated rather than asserted:
46
+
47
+ **Codex line — verdict: `covered`.** One store was read to the bottom, and what was left out is left out for a stated reason, not by omission.
48
+
49
+ | store | status | depth actually read |
50
+ | --- | --- | --- |
51
+ | `~/.codex/memories/MEMORY.md` | `deep-read` | every session block expanded; every failure bullet read individually, not sampled |
52
+ | `~/.codex/skills/.extraction-work/` | `read` | all artifacts present at the time, charter/closeout fields |
53
+ | `~/.codex/sessions/` | `excluded` | not read — the distilled memory file above is already their `symptom -> cause -> fix` reduction |
54
+ | `~/.codex/memories/raw_memories.md` | `excluded` | not read — same source as the distilled file, which is structured |
55
+
56
+ **Claude Code line — verdict: `digest-only`.** Deliberately weaker, and the weakness is the point: nothing in this round may rest on this line alone.
57
+
58
+ | store | status | depth actually read |
59
+ | --- | --- | --- |
60
+ | `~/.claude/projects/<proj>/memory/` | `digest-only` | index lines only; entry bodies not expanded |
61
+ | `~/.claude/skills/.extraction-work/` | `read` | all artifacts present at the time, charter/closeout fields |
62
+
63
+ **Enumerate the line set before you fill the rows, and record how.** A table of two lines cannot reveal a third that was never listed — the gate false-greens on omission exactly the way the parent trap does on an unenumerated corpus. So the line set is *derived from something observable*, never recalled: list the agent home directories present on the machine (`ls -d ~/.*/ ` and pick the ones carrying agent state), then add any host named in this repository's own tooling and routing that has no directory yet. Record the enumeration you ran next to the verdicts, so a reader checks the SET first and the rows second. This is the channel-taxonomy failure of variant (c) applied to agent lines: nothing masks the missing line, it was simply never in the list.
64
+
65
+ The worked example below was enumerated that way, and the enumeration immediately paid for itself: run against the machine that produced this round it returned **five** agent directories, not the two the round had been working with. Three had been invisible to an author who was recalling rather than listing. Two of them (`~/.gemini`, `~/.copilot`) hold configuration and installed skills only — no session or lesson state — and drop out on evidence. The third (`~/.cursor`) is a real line carrying substantial chat state, and it earns an `excluded` row with a reason rather than silence. That is the whole point of the gate: without the enumeration the round would have claimed cross-line coverage over a set it had never established.
66
+
67
+ **Name the lines and their stores.** An anonymised table (`line A`, `line B`, "a memory file") cannot be audited: a reader cannot tell which agents were enumerated, cannot spot a third line that was never listed, and cannot check a store themselves — the gate false-greens on its own example. The private part is the corpus and the project identifiers in it, not the name of the agent that produced it. Name them, and let the sanitization rules do their job on the corpus rows instead of blanking the index.
68
+
69
+ **Cursor line — verdict: `excluded`, with a reason.**
70
+
71
+ | store | status | depth actually read |
72
+ | --- | --- | --- |
73
+ | `~/.cursor/chats/` | `excluded` | not read — raw conversation state with no distilled lesson artifact anywhere under the tree, the same class as the raw transcripts excluded on the Codex line |
74
+ | `~/.cursor/agents/` | `excluded` | empty |
75
+
76
+ **`~/.gemini`, `~/.copilot` — not lines for this purpose.** Configuration and installed skills only; no session or lesson state. Recorded here because "it turned out to hold nothing" is a finding, and leaving them off the list is how the next round re-discovers them.
77
+
78
+ Two things that table makes visible and prose did not. The lines are **asymmetric by store, not by sampling**: one keeps its lessons in the memory file and the other in the extraction-work directory, and the directory named the same on both holds nine times more on one line than the other. And a same-named path means different things per line, so a lookup shaped by your own layout returns nothing and reads as absence.
79
+
80
+ Do NOT resolve this by building a sync or mirror between the stores. Two independently-owned stores with a one-time cross-distillation is the shape that has held here; a mirror adds a mechanism to maintain and drifts silently the first time it is not run. The obligation is coverage at extraction time, not continuous replication.
81
+
37
82
  ## Variant (a0) — produced-artifact + next-run-delta (relocated gate detail)
38
83
 
39
84
  **Firing point: for a task/session retrospective, this variant fires at CHARTER time — the produced-artifact row (i) enters the charter's evidence plan as its FIRST source class (see the Evidence plan field in `source-to-skill-extraction.md`), while the next-run-delta row (ii) is a separate required charter/closeout record, not a source class — not only when an exhausted/complete claim is about to be made** (the original anchor, kept as backstop). Observed failure shape of the late anchor: a 7-session program retro that never claimed "exhausted" walked past this gate entirely, took only correction turns as evidence, and landed zero method/craft lessons while r-series reports and a benchmark corpus sat unread in the project.
@@ -146,6 +146,40 @@ instance in the same session shows the same shape at reporting altitude: an empt
146
146
  output file plus a stale status snapshot were reported as "the commit did not land" instead
147
147
  of being re-read.
148
148
 
149
+ ### The claim/evidence pair table — the single most-recorded failure class
150
+
151
+ **Invariant: the proof you hold establishes a DIFFERENT proposition than the one you are about to assert.** Not a weaker proof of the same claim — a sound proof of an adjacent claim. That is why it survives an honest self-check: the agent did verify something, and it was real.
152
+
153
+ This is the largest class in the round-059 corpus by a wide margin. The counts below are the instances that could be attributed to a specific pair on a re-read — **71 of the 400 failure records read in that pass**, across 14 pairs. (The denominator is the size of that one READ, which is fixed and re-countable from the extraction artifact; it is not the store's current size, which grows.) A coarser class-level pass over the same corpus put the shape higher still, but that figure is not reproducible from this table and is deliberately not quoted here: a table about asserting propositions your evidence does not establish must not open with one. Read 71 as a floor. Every pair is the same sentence with different nouns, which is why patching them one at a time never converged: each fix taught the next agent about `merge` versus `release` and nothing about the shape.
154
+
155
+ | You are about to claim | What your evidence actually establishes | corpus |
156
+ | --- | --- | --- |
157
+ | the reviewer approved it | the review lane returned no verdict (timeout, quota, auth failure, invalid output, exhausted budget) | 16 |
158
+ | the product or the code is defective | YOUR INVOCATION of it failed — missing runner or binary, container runtime down, sandbox denial, unwritable cache, expired credential, toolchain drift | 22 |
159
+ | released / deployed | merged | 8 |
160
+ | runtime behavior is correct | static, contract, compile, or lint evidence passed | 5 |
161
+ | the content is correct | the command exited 0 | 4 |
162
+ | this produced a product effect | CI is green / the package published | 4 |
163
+ | deletion is authorized | merging was authorized | 3 |
164
+ | the data is physically erased | refs are clean and a fresh clone looks right | 2 |
165
+ | there is a live incident | the code path is reachable | 2 |
166
+ | the item is resolved | a reply was posted | 1 |
167
+ | the capability executes | it is registered or configured | 1 |
168
+ | the caller can read it | the caller is a member | 1 |
169
+ | it is implemented | the plan validated | 1 |
170
+ | the application is authenticated | the user identity is authenticated | 1 |
171
+
172
+ **How to use it.** Not as a checklist — as a recognition aid at ONE moment: when you are about to write `done` / `complete` / `verified` / `ready` / `passed` / `covered`. Say out loud the proposition your evidence establishes, then say the proposition you are about to assert. If they are not the same sentence, report the one you have and name the one you do not. `merged; the release pipeline has not run` costs one clause and is true.
173
+
174
+ **What a reader here can and cannot check.** The rows sum to the stated total and that is verifiable in this file. The corpus behind them is NOT in this repository and cannot be: it is per-host agent session history carrying business identifiers, and it stays in private scratch under the extraction lifecycle rules. So the counts are **provenance-bound** — reproducible by whoever holds that corpus, opaque to everyone else. Treat them as what motivated the table, never as a measurement you can audit from here, and do not build a further claim on the exact number. The table earns its keep by whether the shape is recognizable when you next write `done`, which every reader can judge without the corpus.
175
+
176
+ **Why the table is a table and not a rule per row.** The rows are evidence that the invariant is real and recurrent; they are not the specification. A pair absent from this table is still the same defect — the table earns its place by making the shape recognizable, not by enumerating it. Do not extend it every time a new pair appears in the wild; extend it only when a pair recurs and the invariant alone did not catch it.
177
+
178
+ **Two families collapse into this one.** The environment row above was first classified as its own class ("an environment-layer failure reported as a product finding") and the whole-document-overwrite family as another ("the write succeeded" from "the command returned success"). Both are this invariant with different nouns, and `defect-diagnosis` already owns the substantive half of the first — its red-CI cause classification and its prove-it-from-the-tool-that-owns-the-state rule. Recording them as rows rather than as new rules is the point: the count is evidence of the shape, and three parallel rules would have taught three vocabularies instead of one invariant.
179
+
180
+ **Boundary.** This is about the PROPOSITION, orthogonal to `testing-strategy`'s strong/medium/weak evidence *quality* axis: a strong test can perfectly establish the wrong proposition, and that is the failure recorded here. Where a pair has an owner, the substantive rule lives there — release-versus-merge semantics with `release-coordination`, review verdicts in this gate, runtime-versus-static with `testing-strategy` — and this table only makes the class visible at the moment of claiming.
181
+
182
+
149
183
  ## When dual-track is mandatory
150
184
 
151
185
  | Extraction type | Review | Challenge |
@@ -202,6 +236,20 @@ Rules:
202
236
 
203
237
  ## Running the review pass
204
238
 
239
+ ### Freeze the packet against a base you have proven current
240
+
241
+ A stale base puts the upstream's newer fixes into the packet **reversed**, so the reviewer raises findings against code that is already correct — findings that read as real until someone re-checks history, and that cost a full round each. Freezing and currency are **different properties and the packet owes both**: pinning to an immutable commit stops the base moving mid-read, but says nothing about *which* commit. **Five properties, mutually independent — none implies another, so walk them as a list rather than holding them as a sentence.** Every observed failure of this check satisfied four and missed the fifth.
242
+
243
+ 1. **Target-derived, not caller-chosen.** The landing lane's base is the landing target, derived from trusted configuration rather than an arbitrary caller-supplied `--base`. There is **no declared-intent exemption**: a declaration is author-controlled and cannot distinguish a considered pin from a stale one. A review deliberately scoped to a historical base — auditing what some past release shipped — is a separate lane whose result cannot satisfy landing review.
244
+ 2. **Confirmed against the remote that owns the landing target, not against your fetch having run.** `git ls-remote <that remote> refs/heads/<branch>` is the authority — and **`origin` is not automatically it**: on a fork, `origin` is the fork and the landing target lives on the upstream, so querying `origin` confirms a SHA from the wrong authority. Record the **remote, ref, SHA, and the moment confirmed** together; a bare SHA does not preserve which authority supplied it and cannot be audited afterwards. The output must be **non-empty** before it is compared — the command exits `0` with no output for a branch that does not exist on the remote, so a naive comparison turns "the branch is gone" into a silent pass. A fetch that fails, or that succeeds without covering that branch under the configured refspec, leaves the old `origin/<branch>` resolving to a stale commit — a remote-tracking ref is a local mirror and cannot attest to its own freshness.
245
+ 3. **Pinned to an immutable object.** Resolve that tip to a SHA, pass the SHA, and record it next to the packet hash. A ref *name* is not enough: it is mutable, so a sibling agent or background fetch between your resolve and packet construction moves it and the packet is built on a base neither recorded nor checked. Re-pin after any rebase or upstream advance mid-round.
246
+ 4. **Independent of the candidate.** The base comes from the landing-target authority, outside candidate-controlled inputs. Immutability is not independence — a candidate can commit a tuned fixture and cite its perfectly immutable SHA.
247
+ 5. **Contained in HEAD.** `git rev-list --count HEAD..<pinned SHA>` must be `0`, and nothing above implies it: the count passes while a wrapper builds from a stale local `dev` or an old SHA you supplied (a property of HEAD, not of the packet), and a correctly pinned current tip still shows the target's newer commits as *reversals* once the candidate branch has diverged. Count against the **branch tip** — not the derived base commit or a `merge-base HEAD <upstream>` (ancestors of HEAD by construction, so `0` for free), not the *local* branch (the fetch did not advance it), not a *different* branch than the real base. A non-zero count is not a note-and-continue: integrate the target first (`worktree-isolation` owns that sequence and its stale-overwrite hazard), then re-pin and rebuild. **Scoping the packet to the merge base instead does not discharge this** — a three-dot range hides the target's newer commits rather than reversing them, which fixes the reviewer-facing artifact and leaves the real gap: the candidate was never exercised against the code it will land on top of. That is the right bound for reviewing what an author wrote; it is the wrong bound for a landing candidate, whose artifact is the merge.
248
+
249
+ Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, the review is not landing evidence until the base is re-pinned and the round re-run. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
250
+
251
+ **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks any of the five (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section passes a stale base and gets a verdict that looks exactly like a good one. Make that difference visible rather than silent: **the review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is recorded `base-unattested`, which is not landing evidence.** A missing attestation is then a detectable state instead of an indistinguishable one; that is the containment available to prose, and it is weaker than a gate. Mechanising the list into the wrapper is the durable form and is registered as follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
252
+
205
253
  > **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
206
254
  >
207
255
  > 1. **Resolve the organization `code-review` gate first.** It owns the installed-client checks, local `CODE_REVIEW_CLIENT_ORDER`, same-family exclusion, frozen packet, timeout, egress approval, tool boundary, and verdict parsing for Claude, Kimi, OpenCode, and Codex. Do not preselect a client from `command -v` output.