@ccoalm/ccl-skills 0.6.2 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (139) hide show
  1. package/README.md +2 -2
  2. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +11 -0
  3. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-unverified-cli-flag.sh +309 -0
  4. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_unverified_cli_flag.sh +483 -0
  5. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +5 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +10 -8
  7. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +1 -1
  8. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +16 -17
  9. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +1 -1
  10. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +195 -7
  11. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +3 -3
  12. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +13 -5
  13. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +9 -3
  14. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +9 -3
  15. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/normalize_review_timeout.sh +22 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +9 -3
  17. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +1540 -129
  18. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +8 -3
  19. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +76 -1
  20. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +1858 -3
  21. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_update_review_plan_intent.sh +789 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/update_review_plan_intent.py +513 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +1 -1
  24. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +2 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +1 -1
  26. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +14 -11
  27. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +2 -2
  28. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +5 -1
  29. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +2 -1
  30. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +13 -11
  31. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/SKILL.md +64 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/agents/openai.yaml +4 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/async-lifecycle-and-performance.md +72 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/runtime-and-project-contract.md +58 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/source-map.md +41 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/skills/nodejs-service-dev/references/verification-diagnostics-and-security.md +63 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +1 -1
  38. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +25 -9
  39. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +1 -0
  40. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +1 -1
  41. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +16 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +7 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +11 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +8 -10
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +10 -14
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +1 -1
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +135 -86
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +66 -80
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/delivery-contract.md +275 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +88 -214
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +2 -2
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +10 -8
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +4 -5
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +112 -95
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +30 -21
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +22 -3
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +20 -17
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +7 -9
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +14 -10
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +2 -0
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +1 -1
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +9 -6
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +3 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +37 -10
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +7 -1
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +8 -5
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +16 -5
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +4 -2
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +5 -1
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +1 -1
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/audit-history-architecture.md +31 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +1 -1
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +7 -4
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +2 -2
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/notification-architecture.md +28 -0
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +1 -1
  77. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/replay-comparison-architecture.md +28 -0
  78. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/workflow-state-architecture.md +39 -0
  79. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +10 -7
  80. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +8 -0
  81. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/audit-history-patterns.md +29 -0
  82. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +16 -0
  83. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +25 -1
  84. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/notification-patterns.md +40 -0
  85. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +1 -1
  86. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/replay-comparison-patterns.md +30 -0
  87. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/state-machine-task-patterns.md +48 -0
  88. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +10 -1
  89. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +2 -0
  90. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +4 -4
  91. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
  92. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +142 -4
  93. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +21 -2
  94. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +11 -9
  95. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +8 -0
  96. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +5 -4
  97. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +102 -0
  98. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +69 -0
  99. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +10 -0
  100. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +6 -6
  101. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +4 -3
  102. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +93 -2
  103. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-parallel-stack-parity.sh +119 -0
  104. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/extraction_review_gate.sh +22 -0
  105. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +49 -4
  106. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/obligation-ledger.py +2748 -0
  107. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +20 -5
  108. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/shared_git_surface_gate.py +1142 -0
  109. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_parallel_stack_parity.sh +183 -0
  110. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +19 -0
  111. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +41 -4
  112. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ci_checkout_ref_binding.sh +120 -0
  113. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_entrypoint_domain_scan_terms.sh +82 -8
  114. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_extraction_review_gate.sh +336 -0
  115. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_self_adjudication.sh +82 -10
  116. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger.sh +1416 -0
  117. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_obligation_ledger_repo_audit.sh +57 -0
  118. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +141 -4
  119. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +3 -1
  120. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_shared_git_surface_gate.sh +1696 -0
  121. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_delivery_contract.sh +2117 -0
  122. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_uiux_loading_budget.sh +316 -0
  123. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_extraction_review_state.sh +1176 -0
  124. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +31 -1
  125. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +9 -4
  126. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate_extraction_review_state.py +980 -0
  127. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +9 -6
  128. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +11 -11
  129. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +10 -2
  130. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +16 -0
  131. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +1 -1
  132. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +16 -5
  133. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +5 -3
  134. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +16 -6
  135. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/self-benchmark-baseline.md +37 -0
  136. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +7 -5
  137. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +1 -1
  138. package/dist/assets/release.json +275 -105
  139. package/package.json +1 -1
@@ -0,0 +1,40 @@
1
+ # Notification Patterns
2
+
3
+ Use this when implementing notification clients, webhooks, operator alerts, outbound callbacks, or delivery jobs. For durable delivery state, retries, and terminal-state behavior, also apply `state-machine-task-patterns.md`. (Inbound callback *verification* is `public-api-security-patterns.md`; this file owns the outbound side.)
4
+
5
+ Sibling note: `go-microservice-dev/references/notification-patterns.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Client
8
+
9
+ - Wrap each delivery provider behind a `typing.Protocol` with a `send(message) -> Result` surface (`dependency-client-patterns.md` owns the Protocol-vs-ABC rule).
10
+ - Build requests with deadlines, content type, status-code validation, and response body size limits.
11
+ - Treat empty message batches as no-op.
12
+ - Parse provider responses into canonical success, retryable error, and permanent error.
13
+ - **Outbound URLs are an SSRF surface** when endpoints are tenant- or operator-configurable: enforce approved schemes/ports; resolve the hostname once, validate the resolved IP against private/link-local/loopback/cloud-metadata ranges, and **connect to that same validated IP** (connect-by-IP with Host/SNI set from the hostname, or a resolver-pinning client hook) — validate-then-let-the-client-re-resolve is bypassed by DNS rebinding returning a public IP to the validator and a private one to the connection; re-run the resolve-validate-pin cycle on every redirect hop, or disable redirects; prefer routing deliveries through a constrained egress proxy.
14
+
15
+ ## Message Shape
16
+
17
+ - Include notification type, recipient or endpoint reference, template id/version, actor or service identity, trace/log id, safe summary, and dedupe key.
18
+ - Render templates through typed variables, not string concatenation of raw objects.
19
+ - Redact secrets, credentials, tokens, signatures, and large payloads.
20
+
21
+ ## Async Delivery
22
+
23
+ - Critical delivery uses an outbox table or durable queue with retry count, next retry time, terminal status, and last error.
24
+ - Best-effort alerts may run asynchronously but must recover exceptions and emit logs/metrics on failure (detached-spawn firewall per `async-and-worker-patterns.md`).
25
+ - Use bounded concurrency and backoff; group or throttle repeated alerts by stable fingerprint.
26
+ - Flush or drain delivery workers during graceful shutdown when messages are critical.
27
+
28
+ ## Realtime Client Channels
29
+
30
+ - For WebSocket/SSE realtime channels, authenticate and resolve app, tenant, and user context before accepting the connection.
31
+ - Keep connection identity in a concurrency-safe registry keyed by the smallest delivery scope, and remove the client on disconnect via `finally`/context-manager cleanup.
32
+ - Rebuild initial client state from the durable store on first connect or cache miss; send a snapshot after connect, then typed/versioned deltas.
33
+ - Maintain cached unread/count state with explicit TTL refresh and atomic increments; a missing realtime cache must not create a durable-count side effect.
34
+ - Queue consumers treat disconnected clients as a no-op delivery outcome after durable storage succeeds.
35
+
36
+ ## Tests
37
+
38
+ - Test empty batch, timeout, non-2xx response, malformed response, retryable vs permanent classification, redaction, template rendering, dedupe key, throttling, and shutdown drain.
39
+ - Test the SSRF boundary: loopback/private/link-local/metadata-range targets rejected, disallowed scheme/port rejected, redirect to a blocked range rejected, and the pin exercised — assert the connection is made to the validated IP (fake resolver returning different answers on first and second resolution must not reach the second answer).
40
+ - Test realtime connect, duplicate connection, disconnect cleanup, cache-miss bootstrap, disconnected-delivery no-op, and write failure on a closed socket.
@@ -19,7 +19,7 @@ Use this for implementing Python public APIs, partner integrations, signed callb
19
19
 
20
20
  ## Safer Composition (Python 3.14+)
21
21
 
22
- - **Python 3.14 (released 7 October 2025) introduced t-strings via PEP 750** — template literals using `t"..."` syntax that evaluate to `string.templatelib.Template` objects rather than `str`, giving a consuming function access to interpolated values BEFORE they are combined into a string. The standard library exposes `string.templatelib.Template`. **t-strings are inert by themselves — they are NOT automatic injection protection by syntax**. A bare t-string only carries the raw values plus their surrounding template; safety arrives only when a trusted consumer library is t-string-aware and validates/escapes each interpolation according to its target language (SQL, shell, HTML). Writing `t"SELECT * FROM users WHERE id = {user_id}"` and passing it to an ORM/driver that has not added Template support gains nothing over an f-string — and passing it to a function expecting `str` triggers `Template.__str__()` which raises by default, surfacing the misuse rather than silently producing the wrong result. As of 2026-Q1: most popular template engines (Jinja2, Django templates) and most DB drivers do NOT yet consume `Template` directly — verify the specific library's Template support before relying on this rule. **PEP 787 (safer subprocess via t-strings) is currently Deferred to Python 3.15** per peps.python.org — PEP authors are pursuing experimental t-string subprocess work outside the stdlib through 3.14 beta before re-proposing for 3.15. Do NOT assume `subprocess.run(t"...")` works safely in 3.14; for shell/subprocess composition on 3.14, keep using `shlex.quote()` + argument-list form (`subprocess.run(["cmd", arg])`) until PEP 787 or its equivalent lands. For Python ≤3.13 targets, t-strings are unavailable — keep `shlex.quote()` / parameterized DB queries / framework-native HTML escaping; t-strings are a 3.14-and-later opt-in, not a backport.
22
+ - **Python 3.14 (released 7 October 2025) introduced t-strings via PEP 750** — template literals using `t"..."` syntax that evaluate to `string.templatelib.Template` objects rather than `str`, giving a consuming function access to interpolated values BEFORE they are combined into a string. The standard library exposes `string.templatelib.Template`. **t-strings are inert by themselves — they are NOT automatic injection protection by syntax**. A bare t-string only carries the raw values plus their surrounding template; safety arrives only when a trusted consumer library is t-string-aware and validates/escapes each interpolation according to its target language (SQL, shell, HTML). Writing `t"SELECT * FROM users WHERE id = {user_id}"` and passing it to an ORM/driver that has not added Template support gains nothing over an f-string — and passing it to a function expecting `str` triggers `Template.__str__()` which raises by default, surfacing the misuse rather than silently producing the wrong result. As of 2026-Q1: most popular template engines (Jinja2, Django templates) and most DB drivers do NOT yet consume `Template` directly — verify the specific library's Template support before relying on this rule. **PEP 787 (safer subprocess via t-strings) remains a Draft targeting Python 3.15** (postponed from 3.14; per peps.python.org, re-verified 2026-08 — not yet accepted) — PEP authors are pursuing experimental t-string subprocess work outside the stdlib before re-proposing for 3.15. Do NOT assume `subprocess.run(t"...")` works safely in 3.14; for shell/subprocess composition on 3.14, keep using `shlex.quote()` + argument-list form (`subprocess.run(["cmd", arg])`) until PEP 787 or its equivalent lands. For Python ≤3.13 targets, t-strings are unavailable — keep `shlex.quote()` / parameterized DB queries / framework-native HTML escaping; t-strings are a 3.14-and-later opt-in, not a backport.
23
23
 
24
24
  ## Signature And Replay Verification
25
25
 
@@ -0,0 +1,30 @@
1
+ # Replay Comparison Patterns
2
+
3
+ Use this when implementing replay jobs, shadow execution, response comparison, migration verification, or diff reports. Reuse `state-machine-task-patterns.md` for job transitions and terminal-state behavior.
4
+
5
+ Sibling note: `go-microservice-dev/references/replay-comparison-patterns.md` carries the Go rendering; adapted per stack, kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## Replay Jobs
8
+
9
+ - Persist replay jobs with status, target, config version, total/processed/success/failed/diff counts, timeout, concurrency, rate limit, and retention.
10
+ - Use bounded workers with deadlines (`async-and-worker-patterns.md`) for replay execution.
11
+ - Preserve only allowlisted headers and metadata; never replay credentials blindly.
12
+ - Redact captured requests and responses before storage; bound captured payload size.
13
+ - Replay targets must not commit side effects unless isolated by environment, lane, or explicit dry-run mode. Transaction rollback isolates only the local database write — a replayed handler can still send webhooks, publish messages, call payment providers, or write other datastores while its DB transaction rolls back; those adapters need environment/lane isolation or dry-run stubs of their own before rollback counts as isolation.
14
+
15
+ ## Comparator Design
16
+
17
+ - Define a comparator `typing.Protocol` such as `compare(original, replay) -> Result`; select comparators by method, route, content type, or schema version.
18
+ - Generic JSON comparators normalize strings containing JSON, Pydantic models, dicts, lists, and scalars before comparing.
19
+ - Support ignored fields by exact path/field name only when documented, loaded from config or comparator registration — never hard-coded in the generic comparator.
20
+ - Support custom field comparators for tolerances, unordered collections, timestamps, generated IDs, and approximate numerics.
21
+
22
+ ## Diff Result Shape
23
+
24
+ - Include field name, field path, original value summary, replay value summary, diff type, diff score, ignored flag, and metadata; bound value summaries so diff records do not store huge payloads.
25
+ - Record total compared fields, diff count, similarity, threshold, comparator name, and comparator version.
26
+ - Distinguish missing value, added value, type mismatch, length mismatch, value mismatch, and custom comparison.
27
+
28
+ ## Tests
29
+
30
+ - Test None/None, None/value, value/None, JSON-string normalization, model-to-dict conversion, array length mismatch, missing keys, numeric tolerance, ignored fields, custom comparator, threshold behavior, redaction, storage pagination, and replay cancellation.
@@ -0,0 +1,48 @@
1
+ # State Machine And Task Patterns
2
+
3
+ Use this when implementing durable tasks, async workflows, scheduled jobs, imports, exports, and retryable processors. This is the canonical guide for durable state transitions and terminal-state behavior; generic worker/async mechanics stay in `async-and-worker-patterns.md` and `background-job-patterns.md`.
4
+
5
+ Sibling note: `go-microservice-dev/references/state-machine-task-patterns.md` carries the Go rendering of the same discipline; content is adapted per stack and kept in sync by review (not under the parallel-stack parity gate).
6
+
7
+ ## State Model
8
+
9
+ - Define states as a domain-owned `Enum` (per the skill entrypoint's finite-value rule) with documented terminal states.
10
+ - Define allowed transitions in one table or policy function; do not scatter status checks across handlers.
11
+ - Terminal states reject ordinary processing and failure transitions.
12
+ - Include retry count, last error, start time, finish time, progress, result pointer, and idempotency key when tasks are inspectable.
13
+ - Store progress separately from state; progress is best-effort while state transitions are durable.
14
+ - Create the durable task row before launching background work, queueing expensive processing, or generating external artifacts — a crash must leave an inspectable, repairable record.
15
+
16
+ ## Transition Implementation
17
+
18
+ - Guard transitions with compare-and-update (`UPDATE … SET status=:new WHERE id=:id AND status=:expected` and check rowcount), `SELECT … FOR UPDATE`, or a Redis lock (`redis-cache-lock-patterns.md`) before processing.
19
+ - Lease or lock owners must be unique at the runtime-instance level, not only the service level: include pod/host identity plus process id or a random instance id so replicas cannot claim or complete each other's work.
20
+ - A lock alone does not fence: a paused worker whose lease expired can resume and commit after a new owner took over — unique owner names do not prevent this. Every durable write under a lease carries a fencing check at the store it writes to: a monotonic fencing token compared there, or owner/version compare-and-update (`… WHERE id=:id AND owner=:me AND version=:v`), so a stale holder's write fails instead of overwriting the new owner's state.
21
+ - A store-side check fences only that store: a stale worker can pass the DB compare-and-update, pause, lose its lease, and then fire a NON-transactional external effect (a charge, a send) after the new owner completed. External side effects under a lease go through intent-then-execute (`sqlalchemy-and-migrations-patterns.md` outbox / the intent pattern in `python-service-architecture/references/event-driven-architecture.md`): the durable intent row is claimed with the fencing check, and the provider call carries an idempotency key bound to that intent. This fully fences **only when the provider actually enforces the key** (rejects or deduplicates a replay). For providers that cannot — SMTP, plain webhooks, any at-least-once send with no key support — no fencing eliminates the stale-execution window: shrink it with a claim re-check immediately before the call, then classify the effect honestly as at-least-once with a duplicate-visible reconciliation path, and record the residual duplicate window instead of claiming exactly-once.
22
+ - Re-read task state under the lock before side effects.
23
+ - Make duplicate delivery normal: already-successful is success or no-op; already-terminal is no-op or a typed conflict depending on the caller's contract.
24
+ - Start transitions move pending work to processing before expensive work begins.
25
+ - Failure transitions capture canonical error code, safe message, retryable flag, retry count, and last trace/log id.
26
+ - Success transitions persist the result pointer or summary before publishing completion events; completion events are idempotent.
27
+
28
+ ## Async Processing
29
+
30
+ - Prefer queue/worker execution (Celery/RQ/arq per `background-job-patterns.md`) for work that must survive process restart.
31
+ - Delayed events and queue messages re-check current task state when consumed.
32
+ - Retry thresholds and backoff live in config or state policy, not inline literals.
33
+ - A broad `except Exception` at the task boundary converts the task to failed or retryable-failed per the retry policy (never swallow `asyncio.CancelledError` — re-raise it after cleanup, per `async-and-worker-patterns.md`).
34
+ - Work that produces a file, report, or media artifact persists the object key or result pointer before the success transition and exposes a retryable failure state when upload/finalization fails.
35
+
36
+ ## Scheduled Repair Jobs
37
+
38
+ - Scheduled checkers that repair stuck tasks need a job name, interval, distributed lock key, lock lease, per-run max batch size, and max execution time.
39
+ - A job-level lock prevents duplicate scans; a task-level lock or compare-and-update prevents duplicate repair of one record.
40
+ - Time-window selection is explicit, bounded, and based on durable timestamps (`updated_at`), not process memory.
41
+ - Repair actions re-read current task state before changing status or emitting side effects.
42
+ - Emit metrics for scan count, claimed count, repaired count, skipped count, lock conflict, failure, and duration.
43
+
44
+ ## Tests
45
+
46
+ - Test illegal transition, duplicate event, concurrent processing (two claimers, one winner), already-terminal, retry threshold, cancellation, exception-to-failed conversion, and completion-event idempotency.
47
+ - Test lease expiry with a stale worker resuming after a new owner claimed the task: the stale worker's durable write and side effect must be rejected by the fencing check.
48
+ - Test scheduled repair lock conflict, stale-window selection, per-task duplicate prevention, and repeated-run idempotency.
@@ -14,6 +14,15 @@ Use this for pytest, async tests, fixtures, fakes, ruff, mypy/pyright, and CI qu
14
14
  - Use markers to separate unit, integration, API, contract, E2E, live-infra, failure-mode, and drill tests.
15
15
  - Use `pytest-asyncio` for async code according to repo configuration.
16
16
  - Avoid long sleeps and live credentials in default tests.
17
+
18
+ ## TC Traceability And Deprecation Cascade
19
+
20
+ - **TC traceability**: link tests via `@pytest.mark.tc("TC-XX-NNN")` marker. Registers at collection time so `@skip` / `@skipif` / fixture failures still map to Bitable status. Needs the `tc` plugin from `test-artifact-management/references/tc_helpers/tc.py` loaded via `addopts = -p tc`. See `test-artifact-management/references/tc-marker-conventions.md`. Before adding tests, `grep -rn 'pytest\.mark\.tc' tests/` plus the sidecar `test/results/tc-map.jsonl` to check for existing coverage — extend rather than duplicate. When a TC is marked 废弃, grep both source and sidecar; follow deprecation cascade in `testing-strategy`. Tests without any TC link: prompt user only when the underlying code is also removed.
21
+ - **废弃级联:业务代码是否仍在用** — grep 只找出"测试函数引用了什么 import"是第一步;判断"该 import 是否还有其他 caller"才能定生死。Python 顺序:
22
+ 1. 看测试体导入的模块:`grep -E "^(from |import )" tests/test_<x>.py`
23
+ 2. 对每个产品模块(非 stdlib / 非测试 helper),找全仓库 caller:`grep -rEn "from <pkg>\.<mod>|import <pkg>\.<mod>" --include='*.py' --exclude-dir=tests`
24
+ 3. 零产品 caller → 同 commit 删该模块 + 测试;有产品 caller → 测试目标仍在用,不删测试(若 TC 已废弃但代码活,先确认产品决策)
25
+ 4. 边界:动态 import(`importlib.import_module("...")`)grep 抓不到;含 reflection 的代码人工确认;DI/插件注册(`@register` 装饰器)的产品代码需查注册表而非 import
17
26
  - For scenario tests, use `testing-strategy` to build the scenario matrix first; in Python services, map scenarios to unit tests, API/contract tests, integration tests, workflow tests, or marked E2E/live-infra tests instead of putting every case into a slow end-to-end suite.
18
27
 
19
28
  ## Test Target Split
@@ -58,4 +67,4 @@ Use this for pytest, async tests, fixtures, fakes, ruff, mypy/pyright, and CI qu
58
67
  - Keep generated or vendored code excluded according to repo policy.
59
68
  - Use coverage gates when the repo already enforces them or the change is high risk.
60
69
  - **`ruff` (Astral) is the current default-recommendation single tool for lint + format + import sort** — per Astral docs, replaces Flake8 (and dozens of plugins), Black, isort, pydocstyle, pyupgrade, autoflake with one Rust binary; ~900 lint rules (vs Pylint ~409, with ~209 rule overlap); ~tens to hundreds of times faster than the tools it replaces (e.g., the FAQ cites 250k-LOC codebase 2.5min on Pylint vs 0.4s on Ruff). Auto-fix for most violations. Use ruff for new projects by default; existing projects can migrate piecewise (linter and formatter are independent — adopt one without the other). **Pylint coverage gap to audit before retiring it**: ruff does NOT replicate Pylint's deeper semantic / data-flow analysis (Pylint's inference engine can catch e.g. argument-count mismatches across complex call chains, unreachable-after-mutation patterns, type-confusion in dynamically typed code that doesn't reach the type checker), Pylint's broader dead-code detection (beyond ruff's import-level checks — unused-method, unused-attribute, unused-private-member with project-aware reasoning), design smells (cyclomatic-complexity bands, too-many-arguments / too-many-locals / too-many-branches thresholds), duplicate-code detection (`similarities` checker), and any custom in-project Pylint plugin or rule. Strategy: ruff + a type checker (mypy / pyright / ty) replaces Pylint for ~80% of teams; teams relying on the specific Pylint capabilities above should keep Pylint as a slower secondary gate or migrate the equivalents (e.g., use `radon` for complexity, `vulture` for dead-code, a type-checker-level config for design smells) before retiring.
61
- - **`ty` (Astral, Beta announced 2025-12-16 per Astral's own announcement) is the current state of Rust-based type-checker work** — per the Astral announcement, on full check "consistently between 10x and 60x faster than mypy and Pyright" depending on workload, with editor incremental recompute on the order of ~80x faster than Pyright on a touched load-bearing file in benchmark workloads. Read these as Astral-published numbers on Astral-chosen benchmarks; actual speedup on a given codebase varies. Notable features: first-class intersection types, advanced narrowing, reachability analysis. Status: **Beta, not GA** — appropriate for piloting on a separate CI job alongside the existing mypy / pyright / basedpyright gate, NOT for replacing the production type-check gate on its own yet. **Dual-gate operational cost to size before adopting**: running ty plus mypy/pyright produces a "triage tax" — divergent type narrowing (one tool reports, the other doesn't), different stub-package assumptions (typeshed version skew), duplicated CI latency, and a "fix one, break the other" churn pattern that slows MRs. Make the pilot a non-blocking informational job and budget time for periodic finding-diff review rather than treating both as equal blocking gates. Watch for ty GA before flipping defaults. basedpyright remains a viable strict-mypy-like alternative if Pyright's strictness gaps matter and ty isn't ready.
70
+ - **`ty` (Astral, Beta announced 2025-12-16 per Astral's own announcement) is the current state of Rust-based type-checker work** — per the Astral announcement, on full check "consistently between 10x and 60x faster than mypy and Pyright" depending on workload, with editor incremental recompute on the order of ~80x faster than Pyright on a touched load-bearing file in benchmark workloads. Read these as Astral-published numbers on Astral-chosen benchmarks; actual speedup on a given codebase varies. Notable features: first-class intersection types, advanced narrowing, reachability analysis. Status: **Beta, not GA** (re-verified 2026-08: still beta; Astral projects a stable release within 2026, with the beta→stable gap focused on stability, typing-spec completeness, and first-class Pydantic/Django support) — appropriate for piloting on a separate CI job alongside the existing mypy / pyright / basedpyright gate, NOT for replacing the production type-check gate on its own yet. **Dual-gate operational cost to size before adopting**: running ty plus mypy/pyright produces a "triage tax" — divergent type narrowing (one tool reports, the other doesn't), different stub-package assumptions (typeshed version skew), duplicated CI latency, and a "fix one, break the other" churn pattern that slows MRs. Make the pilot a non-blocking informational job and budget time for periodic finding-diff review rather than treating both as equal blocking gates. Watch for ty GA before flipping defaults. basedpyright remains a viable strict-mypy-like alternative if Pyright's strictness gaps matter and ty isn't ready.
@@ -61,6 +61,8 @@ This skill coordinates gates; it does **not** itself authorize merge, tag push,
61
61
  | Reset dev/test-like branches | Yes | Target/env refs, before SHAs, dry-run/plan, force-with-lease semantics |
62
62
  | Post-merge cleanup of the merged temp feature branch (worktree/local/remote) | No — covered by the user's merge authorization (`worktree-isolation` 收尾) | The authorized MR/PR read back as merged at the current head SHA and target; the live remote source ref is absent (already cleaned by the platform) or still equals the merged MR source head (moved → preserve and ask, remote path only — eligible local cleanup proceeds per `worktree-isolation`); no other open or plan-declared MR/PR still consumes the source branch; source branch is a temp feature branch (unclear role → preserve and ask); mechanics/safety rails per `worktree-isolation` |
63
63
 
64
+ **The matrix is a ceiling, not a floor.** A `Yes` row scopes authorization to that action and to the reversible mechanical prerequisites *inside* it — those are not re-asked. Inheritance stops there: it never covers a retry of a consumed authorization (`worktree-isolation` 合并执行协议), a follow-up action, or a prerequisite that is itself gated — that one keeps its own row, so "X needs Y" cannot launder Y's gate. Post-merge cleanup is not an instance of this inheritance; it is the separate narrow carve-out that the boundary above and its own row define. An action absent from this matrix does not acquire a gate by analogy with a listed one — route it to its owner's rules. **Absence is not permission**: anything irreversible, destructive, production-affecting, or of unclear authority preserves state and asks even with no row of its own. Only a clearly reversible, ungated action is ordinary work.
65
+
64
66
  ## Minimal checklist
65
67
 
66
68
  - [ ] Intended scope, base/head refs, and production target identified.
@@ -111,10 +111,10 @@ Use this skill to turn observed experience into durable agent skills without cop
111
111
  - Method: scan the current session's available-skills list before building the owner-generalization map; for each lifecycle stage, name the external-skill candidate alongside the ccl-skill candidate; route to the external skill when it owns the operational recipe and keep the CCL skill as the gate-keeper / cross-cutting rule layer.
112
112
  - When external packages are absent in a teammate's environment, the CCL skill's principle wording must stand alone (no broken `superpowers:*` / `gstack-*` references in executable guidance) — name them as "if installed, route to X; otherwise apply the principle inline".
113
113
  - When a user correction or self-check exposes one missed extraction dimension, sibling owner, or lifecycle stage, scan the immediate neighbors on the same axis before landing the fix. Reuse existing machinery: target-output map for lifecycle, sibling-generalization mini-map for stack/owner, and the source type's dimension enumeration for judgment axes. Do not invent new axes per task and do not walk beyond immediate neighbors. Land only the smallest needed updates and record one line per neighbor as `update`, `unchanged`, or `routed`. The trigger is a discovered miss, not every extraction.
114
- - **Class-wide routing/trigger changes require COMPLETE-set coverage in one pass — the immediate-neighbor scan above does NOT apply.** The neighbor scan covers a *single discovered miss*; a change whose rule is "every skill of class C should advertise / scope / carry X" (e.g. "every stack `*-dev`/`*-architecture` should advertise a localized-refactor trigger") is out of its scope.
114
+ - **Class-wide routing/trigger changes require COMPLETE-set coverage in one pass — the immediate-neighbor scan above does NOT apply.** The neighbor scan covers a *single discovered miss*; a change whose rule is "every skill of class C should advertise / scope / carry X" is out of its scope.
115
115
  - For such a change, first write an **explicit, narrow, source-backed class predicate** (from how the user/source phrased the class — not an expansive inference from examples; if the predicate is ambiguous or huge, downscope or ask, and record non-member exclusions).
116
116
  - Then the COMPLETE set matching that predicate is the required coverage: enumerate the installed members from the session available-skills list PLUS any referenced repo-present CCL members (absent ones get `install-drift: pending` per the next rule, never silent omission), and update or explicitly mark each `unchanged`/`routed` in ONE landing before claiming the class closed.
117
- - Landing one member-pair (e.g. Python+Go) and waiting for the user to surface each remaining member across rounds is the defect this prevents; a repeated same-class miss across rounds is the signal.
117
+ - Landing one member-pair and waiting for the user to surface each remaining member across rounds is the defect this prevents; a repeated same-class miss across rounds is the signal; member-JOIN duty: `references/source-to-skill-extraction.md#member-join-inheritance`.
118
118
  - Validation gate: the closeout map lists every predicate-matching member with a status, or the class is not closed.
119
119
  - (Ordinary single-skill edits still use owner/sibling checks, not a full-class sweep.)
120
120
  - **A referenced ccl-owned/vendored skill not installed in a host is install-drift, not a valid `not-applicable`.** When enumeration reaches a skill that exists in the canonical CCL repo (or is vendored here) and is referenced by the tree but not installed in a host, do NOT mark it `not-applicable: not installed` and move on — that records a symptom as a reason and the skill silently never fires there. Record `install-drift: pending`, surface the exact remediation, and install it ONLY via an approved user/maintainer instruction or the managed install script — do NOT silently create host symlinks or install unprompted (a host mutation changes future routing globally and can point at the wrong checkout). Until installed, that member's coverage stays interim, not omitted. This applies ONLY to ccl-owned/vendored skills; an absent *external/system* package routes to maintainer/upstream, never local install/edit.
@@ -178,9 +178,9 @@ Use this skill to turn observed experience into durable agent skills without cop
178
178
  - So the row records, per protected predicate, the removal that was **applied** and observed to turn the suite RED **for the right reason** — a bare non-zero exit does not qualify (a mutant that breaks syntax or fixture setup also exits non-zero and would bank a broken build as proof of sensitivity); the failure must be attributable to the named protected assertion, and attribution is **differential** (the owning assertion passes in the unmutated control and fails under the mutant, with no non-owning assertion failing) rather than a substring match on aggregate output. An unapplied "this mutation would fail it" is a hypothesis. `testing-strategy` owns the encoded form of that walk (route, don't copy) — for a destructive artifact the walk belongs inside the suite so a later fixture change cannot silently re-blind it.
179
179
  - A challenge skip row is allowed only for wording-only changes; trivial scope does NOT exempt a non-wording change, and ANY skill `description`/frontmatter edit — including a pure typo fix — is NOT wording-only (it changes the routing surface) — both require the full gate.
180
180
  - If a human explicitly asks to skip independent review or challenge, record the review state and residual risk honestly instead of fabricating a pass. Chat or candidate-local text may authorize an in-scope preparation/commit action, but CI authority comes from the protected platform. Distinguish a narrow exact-candidate `review_waiver` (only the review lane becomes non-blocking) from an exact-candidate `merge_authorization` (the human's final merge decision: every CI lane remains visible but none may block that merge). Neither state rewrites failures as passed.
181
- - The Agent-autonomous review budget is the initial review plus at most four challenges, five external rounds total; it never limits deep self-review, implementation, tests, or an authenticated human request. `self_review_gate` mechanically fires before external review, after findings/candidate/scope changes, at the post-budget checkpoint, and before an Agent completion claim. Its `blocks` are narrow: another external review and/or the Agent completion claim, not productive work or an authenticated human merge authorization. Candidate-controlled input cannot assert either human decision.
181
+ - `code-review` permits one review plus four challenges. **Non-wording** Agent-autonomous work MUST use `scripts/extraction_review_gate.sh` at `challenge_budget=2` (three rounds); proven wording-only work keeps the single-review path and no terminal ledger. Later human-requested review is separately attributed outside the non-wording chain, never relabelled round 4/5. Neither budget limits deep self-review, implementation, tests, or authenticated human action. `self_review_gate` fires before external review, after findings/candidate/scope changes, at post-budget, and before an Agent completion claim; it blocks only another review/completion claim, not productive work or human merge authority. Candidate input cannot assert human authority.
182
182
  - Agents cannot self-authorize skipping independent review for any shared-skill change, a challenge skip for non-wording shared-skill changes, or skipping the behavioral-evidence row / true baseline comparison for any change that alters behavior or routing.
183
- - If a required review/challenge row or its required pass is absent, skipped, inconclusive, or unavailable — a recorded `review unavailable after remediation` (or equivalent non-success) row does NOT satisfy the gate — block the Agent's completion/commit claim, run/remediate the review lane when Agent budget remains, or use an approved alternate independent reviewer with the same bounded-scope, attribution, timeout, and output-validity requirements. When Agent budget is exhausted or every lane remains inconclusive, report an `interim` checkpoint, continue self-review/implementation/tests and independent runnable work, and park only decision-dependent work. Only an authenticated human may waive the review-process gate or stop the overall iteration; neither action is inferred from a local file or Agent statement.
183
+ - Missing, skipped, inconclusive, or unavailable required review blocks Agent completion/commit; remediate within wrapper budget or use an approved alternate under the same scope, attribution, timeout, and output checks. If all lanes stay inconclusive, report `interim`. At round 3 validate the v3 receipt-bound ledger: it derives state from the supplied controller chain but cannot prove omitted history or live remote currency, so retain history and run the reference's landing recheck. Report non-success, continue independent work, and park only dependent work. Only an authenticated human may waive review or stop iteration.
184
184
  - A skill is not done until it is validated for discovery, YAML, generic wording, reference links, and at least one non-static evidence row for any non-wording extraction: source reopen, task-shape replay, runtime/rendered/device check, target-owner behavior proof, or an explicit unavailable-with-remediation record. Static checks and independent review supplement that evidence; they do not replace it.
185
185
  - A pressure scenario is not a real extraction test unless it reopens at least one relevant source artifact, reruns the observation -> judgment -> rule -> acceptance path, and either lands a discovered gap or records that no new rule was found. Re-reading only the changed skill text is a static review, not a pressure test.
186
186
  - If the primary independent reviewer hangs, returns no output, hits auth/quota/rate-limit, cannot prove tool posture, cannot show its read covered a large candidate file's middle (a fired read tool-call proves access, not content-fidelity — see the read-coverage check in `references/dual-track-review-gate.md`), or expands beyond the intended scope, do not count it as review evidence. First use the owning wrapper's documented remediation path; for Claude review/challenge, use `../code-review/SKILL.md` (`claude_review.sh`, host/direct recovery, structured output validation). The legacy bounded packet in `references/validation-and-landing.md` is debugging/advisory context only — it is never itself the gate-valid review path; the wrapper (or an approved alternate under the same lane, packet, attribution, timeout, and output-validity rules) is. If wrapper remediation still cannot produce a valid result and an approved alternate reviewer is available, switch tools while preserving the same lane (review vs challenge), bounded diff/file packet, attribution, timeout, and output-validity requirements. A free-form or hanging alternate run is still inconclusive; it does not satisfy the row.
@@ -34,6 +34,51 @@ Its fix is therefore not "read the real artifact" but **walk a seed list of chan
34
34
 
35
35
  **Failure shape:** an agent researching one organization from public sources declares "the public channels are mined out" five separate times; each time the user pushes back, a **previously unconsidered channel class** lands major evidence (an archive route around a bot-block, a regulator's bulk dataset, a mandatory-disclosure regime, another jurisdiction's statutory filings, a filing the organization voluntarily publishes on its own site). No artifact was hiding these — they were never on a list. The domain-side channel taxonomy lives with the research owner (`skills/multi-perspective-research/references/public-disclosure-channels.md`, repo-root relative); what belongs here is the recurrence memory that an exhaustion claim over a *channel set* needs a walked class list, not more effort inside the known classes.
36
36
 
37
+ ## Variant (d) — the OTHER agent line's store (the corpus you cannot see from where you stand)
38
+
39
+ More than one agent runs against the same repository, and each keeps its lessons in its own host's store. Those stores do not see each other. An extraction that enumerates "the lesson corpus" from the host it happens to be running on has enumerated **one line's half of it** and will report exhaustion over a corpus it never touched — the core trap, with host boundary as the thing doing the masking.
40
+
41
+ Measured in round 059: one line held hundreds of `symptom -> cause -> fix` records while another held a few dozen distilled entries, in different trees under different roots with no reference between them. The figures are deliberately not pinned here — these stores are LIVE and grow while you work, so a count written into a durable document is a timestamped measurement that starts rotting the moment it lands; a re-count during this round's own self-check already disagreed with every figure it had recorded hours earlier. State DEPTH, which is stable, not SIZE, which is not. Neither line's store contained a pointer to the other. The consequence is not abstract: one class recurred **eight times across independent sessions** on the line that lacked the abstraction, while the line that had already generalized it — same repository, same weeks — carried the remedy the whole time.
42
+
43
+ The gate: before any complete / exhausted / no-gap claim over a *lesson or failure* corpus, the landing that makes the claim carries, **for each agent line, a verdict** — `covered` / `digest-only` / `unavailable` / `excluded` — and under it one row per store backing that verdict. A per-store list without a per-line verdict is the failure this gate exists to stop: a reader cannot tell from four mixed rows whether the line was actually covered. A line whose store you cannot read is `unavailable` with the remediation attempted, never an omission; a line you deliberately leave out is `excluded` with the reason. Locate the stores by enumerating the hosts that actually work the repository, not by assuming the layout your own host uses — the other line's may be one large per-session file where yours is a directory of entries, and a path-shaped guess finds nothing and reads as absence (the same "keyed by something other than what you are asking about" failure the blocked-source ladder names).
44
+
45
+ **Where these rows live — the split that makes the gate checkable.** Two different things get confused here. The *provenance* — the corpus itself, and the project, repository, branch, ticket and contributor identifiers inside it — is what the extraction lifecycle keeps in per-host scratch, and it stays there. An agent product name and its own config path (`~/.codex/memories/`) are neither: they identify a tool this repository already names throughout, carry nothing about any project, and a leakage audit over them comes back clean. The *coverage rows* are store location, status, depth and exclusion reason; they carry no corpus content, so nothing requires them to be private, and holding them in scratch makes the gate unverifiable by the only people who will ever read the claim. So: coverage rows go **where the claim is made**, alongside it and in the same artifact; provenance stays in scratch and the coverage rows point at it by role, never by real path. If a line's store cannot be described without naming something private, that is a sanitization problem to solve in the row, not a reason to move the row. The round that introduced this variant records its own, so the rule is demonstrated rather than asserted:
46
+
47
+ **Codex line — verdict: `covered`.** One store was read to the bottom, and what was left out is left out for a stated reason, not by omission.
48
+
49
+ | store | status | depth actually read |
50
+ | --- | --- | --- |
51
+ | `~/.codex/memories/MEMORY.md` | `deep-read` | every session block expanded; every failure bullet read individually, not sampled |
52
+ | `~/.codex/skills/.extraction-work/` | `read` | all artifacts present at the time, charter/closeout fields |
53
+ | `~/.codex/sessions/` | `excluded` | not read — the distilled memory file above is already their `symptom -> cause -> fix` reduction |
54
+ | `~/.codex/memories/raw_memories.md` | `excluded` | not read — same source as the distilled file, which is structured |
55
+
56
+ **Claude Code line — verdict: `digest-only`.** Deliberately weaker, and the weakness is the point: nothing in this round may rest on this line alone.
57
+
58
+ | store | status | depth actually read |
59
+ | --- | --- | --- |
60
+ | `~/.claude/projects/<proj>/memory/` | `digest-only` | index lines only; entry bodies not expanded |
61
+ | `~/.claude/skills/.extraction-work/` | `read` | all artifacts present at the time, charter/closeout fields |
62
+
63
+ **Enumerate the line set before you fill the rows, and record how.** A table of two lines cannot reveal a third that was never listed — the gate false-greens on omission exactly the way the parent trap does on an unenumerated corpus. So the line set is *derived from something observable*, never recalled: list the agent home directories present on the machine (`ls -d ~/.*/ ` and pick the ones carrying agent state), then add any host named in this repository's own tooling and routing that has no directory yet. Record the enumeration you ran next to the verdicts, so a reader checks the SET first and the rows second. This is the channel-taxonomy failure of variant (c) applied to agent lines: nothing masks the missing line, it was simply never in the list.
64
+
65
+ The worked example below was enumerated that way, and the enumeration immediately paid for itself: run against the machine that produced this round it returned **five** agent directories, not the two the round had been working with. Three had been invisible to an author who was recalling rather than listing. Two of them (`~/.gemini`, `~/.copilot`) hold configuration and installed skills only — no session or lesson state — and drop out on evidence. The third (`~/.cursor`) is a real line carrying substantial chat state, and it earns an `excluded` row with a reason rather than silence. That is the whole point of the gate: without the enumeration the round would have claimed cross-line coverage over a set it had never established.
66
+
67
+ **Name the lines and their stores.** An anonymised table (`line A`, `line B`, "a memory file") cannot be audited: a reader cannot tell which agents were enumerated, cannot spot a third line that was never listed, and cannot check a store themselves — the gate false-greens on its own example. The private part is the corpus and the project identifiers in it, not the name of the agent that produced it. Name them, and let the sanitization rules do their job on the corpus rows instead of blanking the index.
68
+
69
+ **Cursor line — verdict: `excluded`, with a reason.**
70
+
71
+ | store | status | depth actually read |
72
+ | --- | --- | --- |
73
+ | `~/.cursor/chats/` | `excluded` | not read — raw conversation state with no distilled lesson artifact anywhere under the tree, the same class as the raw transcripts excluded on the Codex line |
74
+ | `~/.cursor/agents/` | `excluded` | empty |
75
+
76
+ **`~/.gemini`, `~/.copilot` — not lines for this purpose.** Configuration and installed skills only; no session or lesson state. Recorded here because "it turned out to hold nothing" is a finding, and leaving them off the list is how the next round re-discovers them.
77
+
78
+ Two things that table makes visible and prose did not. The lines are **asymmetric by store, not by sampling**: one keeps its lessons in the memory file and the other in the extraction-work directory, and the directory named the same on both holds nine times more on one line than the other. And a same-named path means different things per line, so a lookup shaped by your own layout returns nothing and reads as absence.
79
+
80
+ Do NOT resolve this by building a sync or mirror between the stores. Two independently-owned stores with a one-time cross-distillation is the shape that has held here; a mirror adds a mechanism to maintain and drifts silently the first time it is not run. The obligation is coverage at extraction time, not continuous replication.
81
+
37
82
  ## Variant (a0) — produced-artifact + next-run-delta (relocated gate detail)
38
83
 
39
84
  **Firing point: for a task/session retrospective, this variant fires at CHARTER time — the produced-artifact row (i) enters the charter's evidence plan as its FIRST source class (see the Evidence plan field in `source-to-skill-extraction.md`), while the next-run-delta row (ii) is a separate required charter/closeout record, not a source class — not only when an exhausted/complete claim is about to be made** (the original anchor, kept as backstop). Observed failure shape of the late anchor: a 7-session program retro that never claimed "exhausted" walked past this gate entirely, took only correction turns as evidence, and landed zero method/craft lessons while r-series reports and a benchmark corpus sat unread in the project.
@@ -59,6 +59,7 @@ The recurring failure: the agent declares done/covered/converged, and the *user*
59
59
  - **Independent oracle.** Where the property has no executable test — rule text, a register row, a doc — the enumeration is discharged only by an **independent oracle**: name the concrete observation that would contradict the property, say where that observation lives (the owner file and line, the primary source, the command whose output would differ), and go look. **Validate the oracle before trusting its verdict**: a check that returns "clean" because it looked in the wrong place, matched case-sensitively, used too narrow a pattern, or swallowed an error is indistinguishable from a passing property, and it fails in the dangerous direction. Before accepting a clean result you must PROVE THE CHECK CAN FAIL — point it at something you know is broken and watch it report that. Confirming it enumerated the inputs you meant is a necessary extra step, never a substitute: correct inputs say nothing about whether the predicate detects a mismatch or whether a non-zero exit was swallowed, so a check that can only ever say clean passes that weaker test. An unvalidated oracle is not weaker evidence than an imagined mutation; it is the same thing wearing a command prompt.
60
60
  - **Dimension walk.** Adding cases inside an axis you already had buys nothing against one you did not: the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering) before values — `testing-strategy` owns that list and the precision-row obligation that goes with it. A walk whose rows are all imagined mutations is exhortation wearing a checklist's clothes.
61
61
  - **Re-owe after fixes.** Whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration — that newly-added mechanism is the most dangerous line in the diff, because it has no test yet and you wrote it with your attention on the defect it repairs. "The whole enumeration" includes the **pre-cover axes sweep** (concurrency & lifecycle above all): remediation text written mid-round re-owes the draft-time axes BEFORE the candidate goes back to the reviewer, because a fix written with attention on one defect systematically re-opens the same blind-spot axes the original draft missed. And the loop has an escalation point: when the same blind-spot axis or finding class supplies findings in a **third** round, stop the per-finding loop and run one full-matrix implementer self-enumeration (the artifact's own states × failure points × orderings × residues × cross-references) on the current candidate before any further external round — letting the reviewer surface one hole per round is the reviewer-as-defect-finder failure at its most expensive (observed shape: a multi-round program burned twenty-plus single-finding rounds on one axis family; the one full-lifecycle enumeration, run at the maintainer's correction, found the remaining holes in a single batch).
62
+ - **Classify before fixing.** Persist one transition row per finding class: a stable semantic class key, its root-cause predicate, affected surface, and one disposition per occurrence. Each occurrence names the SHA-256 of its controller receipt plus the canonical JSON SHA-256 of a finding that actually appears in that receipt; the closeout must classify every controller finding exactly once. New wording or a new file is not a new class when the predicate is the same. Root-cause predicates must be unique across classes after case-folding and whitespace collapse; **only that exact normalization is mechanical**. It catches cosmetic case/whitespace key splits but does not decide whether differently worded predicates are semantically equivalent, which remains reviewer-contestable. Resolution is per occurrence, not "the last disposition wins": every `fixed`, `accepted_tradeoff`, `pre_existing_out_of_scope`, or `source_refuted` occurrence names one same-directory structured disposition-evidence file plus its SHA-256. That file binds schema version, exact current candidate, its controller receipt and finding, the disposition, a non-empty evidence list, and the ordered current/prior class occurrences it resolves. It must resolve its own occurrence; a later closing disposition that lists only itself leaves an earlier `open` occurrence unresolved until later evidence explicitly includes that exact receipt/finding pair. A `needs_human_decision` occurrence cannot be resolved by any candidate-local disposition evidence and stays unresolved until an authenticated external human/platform decision takes the separately authorized path. A controller finding disproved by first-hand source or failure-path evidence records `source_refuted`; this is a classification of an invalid finding, not a fourth disposition for a valid P0/P1. The validator binds every JSON input using duplicate-key rejection, plus the evidence file, digest, candidate, occurrence, disposition, and transition links; it does not judge the evidence text or authenticate tradeoff/scope acceptance. The reviewer or human decision-maker still owns those semantic and authority verdicts, and `ready_for_human_decision` is not approval or merge authority. On the class's third appearance, stop patching individual instances and enumerate the complete authoritative surface against that predicate. The sweep names a same-directory manifest plus its SHA-256; its candidate and exact ordered `searched_set` must match the row, and its unmatched list supplies the recorded count. `ready_for_human_decision` requires zero unresolved occurrences and zero unmatched instances; `continuation_authorization_required` and `baseline_race` retain non-zero unmatched evidence instead of lying about closure. Classification is reviewer-contestable evidence, not an author-controlled escape hatch; splitting one predicate into cosmetic sub-classes does not reset the count. `scripts/validate_extraction_review_state.py` enforces these bindings within the referenced receipt/evidence set.
62
63
  - **Graded verdict shape.** When the assessed reality is multi-dimensional or partial (capability, coverage, feasibility, quality, completion), collapsing it into one binary verdict — "done/not-done", "possible/impossible", "all correct/all wrong" — is the over-broad-absolute axis applied to your own claim layer: the swing to whichever pole feels safest to assert misrepresents a distribution, and the opposite-pole absolute ("structurally impossible", "nothing works") is the SAME defect as an unearned "done", not a humbler one. Report per-dimension status — what's strong, what's weak, what wasn't checked — with the confidence each part actually earned; and where a binary gate genuinely applies (a pass/fail check, a blocked/allowed decision), still give the clear top-line verdict after the per-dimension basis — calibration is not hedged mush. A user correcting your answers as too absolute ("每次都很绝对") is this defect's recurrence signal, same escalation as the `SKILL.md` rule states.
63
64
  - **Honesty (descriptive, not permissive).** This is recognition-dependent salience, not a mechanical gate — an agent that doesn't notice it is done-claiming cannot self-fire it; the mechanical backstops remain the closeout `interim` gates + user-signal escalation. "I didn't notice I was claiming done" does NOT waive the rule — any non-trivial completion/coverage/convergence wording must carry clean-pass evidence or an explicit interim/downscope disposition *before* you emit it. The rule targets completion/coverage/convergence assertions on work whose failure a check could catch, and never narrows the mandatory dual-track challenge (it is the always-on generalization of *self-audit to convergence*, not a replacement for the gate).
64
65
 
@@ -146,6 +147,40 @@ instance in the same session shows the same shape at reporting altitude: an empt
146
147
  output file plus a stale status snapshot were reported as "the commit did not land" instead
147
148
  of being re-read.
148
149
 
150
+ ### The claim/evidence pair table — the single most-recorded failure class
151
+
152
+ **Invariant: the proof you hold establishes a DIFFERENT proposition than the one you are about to assert.** Not a weaker proof of the same claim — a sound proof of an adjacent claim. That is why it survives an honest self-check: the agent did verify something, and it was real.
153
+
154
+ This is the largest class in the round-059 corpus by a wide margin. The counts below are the instances that could be attributed to a specific pair on a re-read — **71 of the 400 failure records read in that pass**, across 14 pairs. (The denominator is the size of that one READ, which is fixed and re-countable from the extraction artifact; it is not the store's current size, which grows.) A coarser class-level pass over the same corpus put the shape higher still, but that figure is not reproducible from this table and is deliberately not quoted here: a table about asserting propositions your evidence does not establish must not open with one. Read 71 as a floor. Every pair is the same sentence with different nouns, which is why patching them one at a time never converged: each fix taught the next agent about `merge` versus `release` and nothing about the shape.
155
+
156
+ | You are about to claim | What your evidence actually establishes | corpus |
157
+ | --- | --- | --- |
158
+ | the reviewer approved it | the review lane returned no verdict (timeout, quota, auth failure, invalid output, exhausted budget) | 16 |
159
+ | the product or the code is defective | YOUR INVOCATION of it failed — missing runner or binary, container runtime down, sandbox denial, unwritable cache, expired credential, toolchain drift | 22 |
160
+ | released / deployed | merged | 8 |
161
+ | runtime behavior is correct | static, contract, compile, or lint evidence passed | 5 |
162
+ | the content is correct | the command exited 0 | 4 |
163
+ | this produced a product effect | CI is green / the package published | 4 |
164
+ | deletion is authorized | merging was authorized | 3 |
165
+ | the data is physically erased | refs are clean and a fresh clone looks right | 2 |
166
+ | there is a live incident | the code path is reachable | 2 |
167
+ | the item is resolved | a reply was posted | 1 |
168
+ | the capability executes | it is registered or configured | 1 |
169
+ | the caller can read it | the caller is a member | 1 |
170
+ | it is implemented | the plan validated | 1 |
171
+ | the application is authenticated | the user identity is authenticated | 1 |
172
+
173
+ **How to use it.** Not as a checklist — as a recognition aid at ONE moment: when you are about to write `done` / `complete` / `verified` / `ready` / `passed` / `covered`. Say out loud the proposition your evidence establishes, then say the proposition you are about to assert. If they are not the same sentence, report the one you have and name the one you do not. `merged; the release pipeline has not run` costs one clause and is true.
174
+
175
+ **What a reader here can and cannot check.** The rows sum to the stated total and that is verifiable in this file. The corpus behind them is NOT in this repository and cannot be: it is per-host agent session history carrying business identifiers, and it stays in private scratch under the extraction lifecycle rules. So the counts are **provenance-bound** — reproducible by whoever holds that corpus, opaque to everyone else. Treat them as what motivated the table, never as a measurement you can audit from here, and do not build a further claim on the exact number. The table earns its keep by whether the shape is recognizable when you next write `done`, which every reader can judge without the corpus.
176
+
177
+ **Why the table is a table and not a rule per row.** The rows are evidence that the invariant is real and recurrent; they are not the specification. A pair absent from this table is still the same defect — the table earns its place by making the shape recognizable, not by enumerating it. Do not extend it every time a new pair appears in the wild; extend it only when a pair recurs and the invariant alone did not catch it.
178
+
179
+ **Two families collapse into this one.** The environment row above was first classified as its own class ("an environment-layer failure reported as a product finding") and the whole-document-overwrite family as another ("the write succeeded" from "the command returned success"). Both are this invariant with different nouns, and `defect-diagnosis` already owns the substantive half of the first — its red-CI cause classification and its prove-it-from-the-tool-that-owns-the-state rule. Recording them as rows rather than as new rules is the point: the count is evidence of the shape, and three parallel rules would have taught three vocabularies instead of one invariant.
180
+
181
+ **Boundary.** This is about the PROPOSITION, orthogonal to `testing-strategy`'s strong/medium/weak evidence *quality* axis: a strong test can perfectly establish the wrong proposition, and that is the failure recorded here. Where a pair has an owner, the substantive rule lives there — release-versus-merge semantics with `release-coordination`, review verdicts in this gate, runtime-versus-static with `testing-strategy` — and this table only makes the class visible at the moment of claiming.
182
+
183
+
149
184
  ## When dual-track is mandatory
150
185
 
151
186
  | Extraction type | Review | Challenge |
@@ -170,6 +205,16 @@ Only the challenge pass may be skipped, and only when this table marks challenge
170
205
  - **(a) Bounded change class.** The edit changes ONLY typo, grammar, formatting, or a meaning-preserving synonym, and changes NO trigger, scope, routing, validation, acceptance, rule/threshold/boundary text, or any other meaning. Reference/body prose that states a rule, threshold, boundary, rubric, or applies/does-not-apply line IS a semantic surface — editing it is NOT wording-only unless the change is purely typo/grammar/formatting with no meaning change. A synonym substitution usually cannot satisfy (b)'s deterministic evidence bar unless the proof avoids intent/meaning judgment. Description / frontmatter is never wording-only (see below).
171
206
  - **(b) Deterministic scope check + independent review.** Dropping challenge for the edit requires a **deterministic scope check** — recorded controller-side or diff-based scope evidence that is decidable without judging intent or meaning, such as: touched files are formatting-only by formatter output; hunks are only meaning-inert whitespace / punctuation / markdown table alignment; or token-level changes are limited to a named typo correction while the surrounding rule sentence is byte-identical. If the proof depends on a human or LLM deciding whether revised prose changes a rule, threshold, boundary, applicability, or acceptance meaning, it is NOT deterministic and challenge stays required. The deterministic evidence **AND** an independent review row confirming the same must both be present. Either piece missing, or any reviewer-flagged / unconfirmed meaning / scope / trigger / routing / validation / acceptance change, **re-arms challenge + the behavioral-evidence row** (never demoted to "recommended").
172
207
 
208
+ The executable controller proof is intentionally narrower than every edit a
209
+ human might call wording-only. It accepts only a canonical full-context
210
+ Markdown diff inside one existing skill and recomputes either punctuation-only
211
+ changed lines or one named whole-token replacement with an exact count. Use the
212
+ schema and command in `code-review/references/staged-review-contract.md`; the
213
+ result must carry both `wording_only_scope.status=passed` and the independently
214
+ reviewed `wording_only_boundary` concern. Other grammar/synonym edits, custom or
215
+ context-augmented packets, multiple skills, and any unconfirmed meaning change
216
+ take the normal challenge path.
217
+
173
218
  **If no such deterministic scope check can be formed for the change, challenge stays required.** This is a hard gate, not a default the author may waive: an LLM independent review is hypothesis-grade, so review alone never downgrades a non-wording change. Every shared-skill change still requires the independent review row regardless of class.
174
219
 
175
220
  If either required pass times out, returns empty output, is rate-limited, cannot access auth, exits nonzero, emits malformed or truncated output, fails JSON/shape parsing when structured output was requested, lacks evidence that the pass inspected the target diff/files, shows a prompt/tool-scope mismatch, or returns any `inconclusive` status, the dual-track gate has not passed **for that lane via that reviewer**. A *recoverable-lane* failure — auth, quota, rate limit, timeout, local cache/db failure, or missing capability — is NOT a terminal stop and is NOT "the gate is unrunnable on this host": before you record `blocked`, you MUST walk the **Primary reviewer failure** remediation ladder below and route to an approved independent third-party reviewer (preferably a different model family) — either your runtime's native multi-model subagent (e.g. an OpenCode `Task`/council subagent on a separate model) or a shell wrapper such as `opencode_review.sh --model <provider/model> --implementer-family <author-family> --mode review|challenge`. A primary CLI lacking auth is a routing trigger to that fallback lane, never a license to declare the gate unrunnable. Only after that ladder is exhausted — every approved independent reviewer probed and unavailable — do you record the row as `pending` or `blocked`, include the remediation attempted (which ladder steps were tried) and the next unblock action, and do not describe the skill change as solved, complete, landed, or fully closed. (A large reviewer INPUT can also be silently middle-truncated, not just the reviewer's OUTPUT — see **Read-coverage of large inputs** under *Sanity checks the gate must enforce*.)
@@ -202,6 +247,43 @@ Rules:
202
247
 
203
248
  ## Running the review pass
204
249
 
250
+ ### Freeze the packet against a base you have proven current
251
+
252
+ A stale base puts the upstream's newer fixes into the packet **reversed**, so the reviewer raises findings against code that is already correct — findings that read as real until someone re-checks history, and that cost a full round each. Freezing and currency are **different properties and the packet owes both**: pinning to an immutable commit stops the base moving mid-read, but says nothing about *which* commit. **Five properties, mutually independent — none implies another, so walk them as a list rather than holding them as a sentence.** Every observed failure of this check satisfied four and missed the fifth.
253
+
254
+ 1. **Target-derived, not caller-chosen.** The landing lane's base is the landing target, derived from trusted configuration rather than an arbitrary caller-supplied `--base`. There is **no declared-intent exemption**: a declaration is author-controlled and cannot distinguish a considered pin from a stale one. A review deliberately scoped to a historical base — auditing what some past release shipped — is a separate lane whose result cannot satisfy landing review.
255
+ 2. **Confirmed against the remote that owns the landing target, not against your fetch having run.** `git ls-remote <that remote> refs/heads/<branch>` is the authority — and **`origin` is not automatically it**: on a fork, `origin` is the fork and the landing target lives on the upstream, so querying `origin` confirms a SHA from the wrong authority. Record the **remote, ref, SHA, and the moment confirmed** together; a bare SHA does not preserve which authority supplied it and cannot be audited afterwards. The output must be **non-empty** before it is compared — the command exits `0` with no output for a branch that does not exist on the remote, so a naive comparison turns "the branch is gone" into a silent pass. A fetch that fails, or that succeeds without covering that branch under the configured refspec, leaves the old `origin/<branch>` resolving to a stale commit — a remote-tracking ref is a local mirror and cannot attest to its own freshness.
256
+ 3. **Pinned to an immutable object.** Resolve that tip to a SHA, pass the SHA, and record it next to the packet hash. A ref *name* is not enough: it is mutable, so a sibling agent or background fetch between your resolve and packet construction moves it and the packet is built on a base neither recorded nor checked. Re-pin after any rebase or upstream advance mid-round.
257
+ 4. **Independent of the candidate.** The base comes from the landing-target authority, outside candidate-controlled inputs. Immutability is not independence — a candidate can commit a tuned fixture and cite its perfectly immutable SHA.
258
+ 5. **Contained in HEAD.** `git rev-list --count HEAD..<pinned SHA>` must be `0`, and nothing above implies it: the count passes while a wrapper builds from a stale local `dev` or an old SHA you supplied (a property of HEAD, not of the packet), and a correctly pinned current tip still shows the target's newer commits as *reversals* once the candidate branch has diverged. Count against the **branch tip** — not the derived base commit or a `merge-base HEAD <upstream>` (ancestors of HEAD by construction, so `0` for free), not the *local* branch (the fetch did not advance it), not a *different* branch than the real base. A non-zero count is not a note-and-continue: integrate the target first (`worktree-isolation` owns that sequence and its stale-overwrite hazard), then re-pin and rebuild. **Scoping the packet to the merge base instead does not discharge this** — a three-dot range hides the target's newer commits rather than reversing them, which fixes the reviewer-facing artifact and leaves the real gap: the candidate was never exercised against the code it will land on top of. That is the right bound for reviewing what an author wrote; it is the wrong bound for a landing candidate, whose artifact is the merge.
259
+
260
+ Walking all five is a **point-in-time attestation, not a lock**: the target can advance after you confirm and pin, while the packet is being built or reviewed. That is what item 2's recorded confirmation moment is for — the verdict covers that base only. So **re-query the authority at landing time, immediately before acting on the verdict**: if its tip is no longer the pinned SHA, the review is not landing evidence until the base is re-pinned and the round re-run. Stating the consequence is not enough without that second query — nothing else would ever detect the movement. A force-push or branch deletion is the sharp case: the new tip need not contain the old one, so "it can only have moved forward" is not an assumption available to you. Do not paper over the window by re-checking harder; bound it, and say when it closed.
261
+
262
+ When that check detects drift, rebuild instead of improvising: stop the reviewer
263
+ lane; preserve the named candidate manifest/patch and the caller-owned ordered evidence rows;
264
+ integrate the newly attested target in an isolated worktree; reapply only the
265
+ named candidate paths/rows; regenerate derived artifacts; compare the resulting
266
+ path set with the manifest; rerun the selected tests; then re-pin and rebuild the
267
+ packet. Never use a broad reset plus `add -A`, which can silently absorb ambient
268
+ work. Keep every base attestation in one ledger scoped from the first packet
269
+ until landing or a human scheduling decision; repinning, retrying, or rebuilding
270
+ does not reset it. Each row uses one remote/ref, a contiguous sequence, a strictly
271
+ increasing RFC3339 confirmation time, the corresponding controller-receipt hash
272
+ when a round consumed it, and a same-directory hash-bound file containing the
273
+ canonical raw `ls-remote` line. A second ordered SHA change (A→B→C or A→B→A) is
274
+ the second drift and must terminate the lane as `baseline_race`, with the
275
+ unreviewed delta; the drift row and every later row must not map another
276
+ controller receipt. For every non-race closeout — `ready_for_human_decision`
277
+ and `continuation_authorization_required` alike — the final controller receipt
278
+ must consume the latest attested SHA; later same-SHA live rechecks are allowed,
279
+ but an unconsumed newer SHA is not reviewed evidence (a post-final-round drift
280
+ belongs in the next round's ledger, not appended unconsumed to this one). Do not open another
281
+ automatic rebuild after the second drift. The state validator counts these
282
+ transitions and receipt/base associations inside the complete referenced row
283
+ set. Keep independent work moving while a human chooses a landing window.
284
+
285
+ **Until a wrapper enforces these, the caller owns them — and an unenforced obligation must at least be a recorded one.** No wrapper checks the five live-Git properties (observed 2026-08: `review_gate.sh` freezes whatever base it is handed, and neither wrapper fetches), so an agent that never loads this section can still pass a stale base. The v3 closeout validator makes referenced evidence tampering and broken round association detectable, but it cannot prove that a caller supplied every historical attestation or that the recorded `ls-remote` output is still current; candidate-local receipts are consistency evidence, not remote authority or an append-only log. Therefore **each review/challenge row carries the base attestation — remote, ref, SHA, confirmation moment — and a row without one is `base-unattested`, not landing evidence**, the caller retains the complete history across rebuilds, and item 2's live authority query is still repeated immediately before landing. Mechanising the live checks and history retention into a trusted wrapper/platform is follow-up work. This is the review-lane instance of the baseline rule in `external-practice-controls.md#designing-a-behavioral-evidence-measurement`: the packet is a measurement, and its base must be one the candidate has not moved.
286
+
205
287
  > **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
206
288
  >
207
289
  > 1. **Resolve the organization `code-review` gate first.** It owns the installed-client checks, local `CODE_REVIEW_CLIENT_ORDER`, same-family exclusion, frozen packet, timeout, egress approval, tool boundary, and verdict parsing for Claude, Kimi, OpenCode, and Codex. Do not preselect a client from `command -v` output.
@@ -214,7 +296,7 @@ Rules:
214
296
 
215
297
  Primary reviewer failure is a remediation branch only when the owning gate classifies it as candidate-local. Use this ladder separately for the review lane and challenge lane:
216
298
 
217
- 1. Persist the owner-guided self-review, pass it through the required `--review-plan-file`, and run `review_gate.sh` once for the lane. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
299
+ 1. Persist the owner-guided self-review and pass it through the required `--review-plan-file`. For non-wording extraction work, run `scripts/extraction_review_gate.sh` from round 1; a strictly proven wording-only lane uses the proof-bound generic `code-review` single-review recipe in `code-review/references/staged-review-contract.md`, without chain or completion flags. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
218
300
  2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
219
301
  3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
220
302
  4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
@@ -222,11 +304,16 @@ Primary reviewer failure is a remediation branch only when the owning gate class
222
304
 
223
305
  Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
224
306
 
225
- - **Open the Agent chain on the FIRST review — it cannot be retrofitted.** An extraction's required review and challenge are one tracked multi-round run, so the round budget is decided before round 1, not after reading the review; a run that starts untracked is thrown away and restarted. Every trigger, default, flag, index, prior-result, and advisory rule behind that obligation is owned by `code-review/references/staged-review-contract.md` (Agent review chain), with the runnable pair in `code-review/SKILL.md` — take the command from there and never reconstruct it from this bullet.
307
+ - **Open the non-wording Agent chain on the FIRST review — it cannot be retrofitted.** A non-wording extraction's required review and challenge are one tracked multi-round run through `scripts/extraction_review_gate.sh`, so the round budget is decided before round 1, not after reading the review; a run that starts untracked or through the generic controller is thrown away and restarted. A strictly proven wording-only change instead uses the proof-bound single-review exception and opens no challenge chain or `complete` checkpoint. Every trigger, proof, index, prior-result and advisory rule behind those invocation shapes is owned by `code-review/references/staged-review-contract.md`, with controller options in `code-review/SKILL.md`; do not reconstruct them from this bullet.
226
308
 
227
309
  - **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
228
310
 
229
- Standard `codex review` against the target diff:
311
+ The following raw CLI shapes are **debugging diagnostics only**. They may help
312
+ isolate an owner-wrapper failure, but neither output is review/challenge evidence
313
+ and neither may replace `scripts/extraction_review_gate.sh` for a non-wording
314
+ lane.
315
+
316
+ Debug a wrapper with raw `codex review` against the target diff:
230
317
 
231
318
  ```bash
232
319
  cd <skills-repo>
@@ -262,7 +349,7 @@ Output: one finding per line with severity (P0/P1/P2), file:line, scenario, fix.
262
349
 
263
350
  ## Running the challenge pass
264
351
 
265
- `codex exec` with adversarial prompt (read-only):
352
+ Debug a wrapper with raw `codex exec` and an adversarial prompt (read-only):
266
353
 
267
354
  ```bash
268
355
  timeout 540 codex exec "<adversarial-prompt>" </dev/null \
@@ -436,6 +523,31 @@ Do NOT iterate to zero *findings* — some are intentional design tradeoffs the
436
523
 
437
524
  The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
438
525
 
526
+ Five is the generic `code-review` transport ceiling, not this extraction lane's
527
+ spend. Non-wording Agent-autonomous extraction calls go through
528
+ `scripts/extraction_review_gate.sh`, which fixes `challenge_budget=2`: the
529
+ initial review plus at most two challenges. At round 3 the autonomous lane ends.
530
+ An authenticated human may request later review, but that is separately
531
+ attributed human-requested evidence outside this chain/budget, not an Agent
532
+ round 4 or 5. Unused generic capacity never authorizes automatic continuation.
533
+ The v3 closeout validator rejects referenced receipts whose recorded budget is
534
+ not 2 and checks budget and ordering consistency within the caller-supplied
535
+ set. It cannot authenticate that the wrapper produced those receipts or that
536
+ the caller retained every earlier chain or receipt. The wrapper does not mint or
537
+ persist
538
+ `review_chain_id` or `autonomous_review_index`: the caller still supplies both,
539
+ and could start a fresh-looking chain after round 3. The validator detects bad
540
+ order inside the referenced set but cannot detect a prior chain the caller
541
+ omitted, so complete caller-owned ledger retention—and treating an Agent reset
542
+ as a contract violation—remains part of the boundary rather than a property the
543
+ local scripts prove.
544
+
545
+ A strictly proven wording-only change has no convergence loop: it uses one
546
+ generic `code-review` pass, records the independent-review row and the
547
+ challenge-not-required proof, and does not create a schema-v3 multi-round
548
+ terminal ledger. This exception does not apply to frontmatter, routing,
549
+ validation, acceptance, example, owner or behavior changes.
550
+
439
551
  This budget limits only automatic reviewer invocation. It does **not** stop implementation, tests, debugging, or deep self-review, and it does not limit an authenticated human:
440
552
 
441
553
  - A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
@@ -457,6 +569,32 @@ At the third Agent-autonomous round, do not start a fourth automatically. If fin
457
569
  - mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
458
570
  - enter `awaiting_human` only when no independent runnable work remains. This is a scheduling state, not task failure and not a human merge prohibition.
459
571
 
572
+ The terminal checkpoint is an extraction closeout record, not a state emitted by
573
+ `review_gate.py`, and its schema-v3 state is derived from evidence rather than
574
+ trusted as an author assertion. Schema-v2 closeout ledgers are rejected rather
575
+ than silently reinterpreted under the breaking occurrence/evidence shape. The
576
+ ledger and every referenced controller, completion, base, and sweep file live
577
+ in one directory and carry SHA-256s. The validator walks the ordered schema-v3
578
+ controller chain (same chain and scope,
579
+ review then contiguous challenges, packet=candidate, complete prior-result hash
580
+ prefix, fixed `challenge_budget=2`) and binds every closeout candidate to its
581
+ last receipt. Ready requires at least review + challenge; a second base drift may
582
+ stop as race immediately after round 1 rather than spending an illegal challenge
583
+ after the terminal predicate already fired.
584
+ It ends in exactly one state:
585
+
586
+ - `ready_for_human_decision`: a real `complete` receipt is `passed / self_reviewed`, binds the final external receipt and exact current candidate, and there is no unresolved finding occurrence, unreviewed delta, or unmatched sweep instance.
587
+ - `continuation_authorization_required`: round 3 itself returned `findings / post_review_budget`; a passed/unknown/inconclusive state cannot be relabelled continuation.
588
+ - `baseline_race`: the referenced ordered base rows contain a second SHA change, including A→B→A; there is no completion receipt and the unreviewed delta is non-empty. Open findings and unmatched sweep instances remain visible and do not prevent this stop state.
589
+
590
+ Run `scripts/validate_extraction_review_state.py <closeout.json>` before reporting
591
+ the state. This proves internal consistency and coverage of the files the ledger
592
+ references. It does **not** authenticate that no earlier receipt/attestation was
593
+ omitted and does not replace the live remote recheck above; the caller still owns
594
+ complete-history retention until a trusted platform owns it. An exhausted budget,
595
+ stale review, omitted evidence, or unknown lane state is never represented as
596
+ convergence.
597
+
460
598
  ### Concrete cadence
461
599
 
462
600
  For a focused single-skill change: