@ccoalm/ccl-skills 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (561) hide show
  1. package/LICENSE +201 -0
  2. package/README.md +49 -0
  3. package/dist/assets/marketplace/.agents/plugins/marketplace.json +12 -0
  4. package/dist/assets/marketplace/.claude-plugin/marketplace.json +13 -0
  5. package/dist/assets/marketplace/marketplace-manifest.json +12 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/marketplace.json +16 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/plugin.json +5 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/.codex-plugin/plugin.json +5 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/.worktree-only +3 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +45 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/subagent-start.md +12 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/hooks/AGENTS.md +19 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-delegation-owner.sh +125 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-edit-isolation.sh +102 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +1156 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +131 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/hooks/merge-authorization-prompt.sh +142 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-guard.sh +12 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-stop.sh +13 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +144 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-context.sh +87 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-start.sh +86 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-extraction-gate-stop.sh +69 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/hooks/subagent-start.sh +26 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_delegation_owner.sh +329 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_edit_isolation.sh +322 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +902 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +178 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +121 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_session_start.sh +170 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/AGENTS.md +17 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +564 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +14 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-update-skills.md +44 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-verify-skills.md +109 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-worktree-check.md +36 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/AGENTS.md +28 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/README.md +276 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.example.json +10 -0
  40. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +1307 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +941 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/SKILL.md +45 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/agents/openai.yaml +4 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +188 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/agents/openai.yaml +4 -0
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/android-dev.md +92 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/flutter-dev.md +80 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/ios-dev.md +72 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/kotlin-multiplatform.md +93 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-platform-boundaries.md +77 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +77 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/source-evidence-map.md +64 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +353 -0
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/agents/openai.yaml +4 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +419 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +126 -0
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +197 -0
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +179 -0
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +98 -0
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +93 -0
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_timeout_exit.sh +15 -0
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +1438 -0
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +324 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/concern_excerpt.py +295 -0
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/egress_schema.py +214 -0
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +642 -0
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +181 -0
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +1165 -0
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +1190 -0
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +946 -0
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_opencode_review.py +474 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +1899 -0
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_review_json.py +200 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +2845 -0
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.sh +6 -0
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/run_claude_capture.py +71 -0
  77. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +53 -0
  78. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +68 -0
  79. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +2311 -0
  80. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +1832 -0
  81. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_code_review_identity.sh +73 -0
  82. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_concern_excerpt.sh +245 -0
  83. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_egress_schema.sh +177 -0
  84. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +272 -0
  85. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +195 -0
  86. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_concurrency.sh +120 -0
  87. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_retry.sh +1005 -0
  88. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_opencode_review.sh +258 -0
  89. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +574 -0
  90. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +349 -0
  91. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +434 -0
  92. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +264 -0
  93. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +2412 -0
  94. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/verify_native_skill_binding.py +123 -0
  95. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +153 -0
  96. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/agents/openai.yaml +4 -0
  97. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +54 -0
  98. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/prevention-routing.md +36 -0
  99. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +69 -0
  100. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/agents/openai.yaml +4 -0
  101. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/references/security-review-gate.md +41 -0
  102. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +165 -0
  103. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/agents/openai.yaml +4 -0
  104. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/api-security-boundaries.md +47 -0
  105. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +160 -0
  106. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/artifact-generation-architecture.md +37 -0
  107. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/audit-history-architecture.md +29 -0
  108. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/bulk-workflow-architecture.md +33 -0
  109. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/config-rule-routing-architecture.md +34 -0
  110. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +72 -0
  111. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-modeling-and-migrations.md +79 -0
  112. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +210 -0
  113. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/dependency-platform.md +105 -0
  114. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/developer-tooling-architecture.md +38 -0
  115. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/error-contract-architecture.md +36 -0
  116. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +260 -0
  117. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/http-gateway-architecture.md +74 -0
  118. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/mq-consumer-architecture.md +38 -0
  119. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +275 -0
  120. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/notification-architecture.md +25 -0
  121. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/ops-checklist.md +57 -0
  122. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/performance-capacity-architecture.md +38 -0
  123. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/protobuf-contract-architecture.md +119 -0
  124. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/redis-cache-coordination.md +93 -0
  125. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/release-runtime-readiness.md +65 -0
  126. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/replay-comparison-architecture.md +26 -0
  127. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/runtime-observability.md +94 -0
  128. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/service-scaffold.md +76 -0
  129. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/source-evidence-map.md +55 -0
  130. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/workflow-state-architecture.md +38 -0
  131. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +159 -0
  132. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/agents/openai.yaml +4 -0
  133. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/artifact-generation-patterns.md +37 -0
  134. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/audit-history-patterns.md +28 -0
  135. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/bulk-import-export-patterns.md +56 -0
  136. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/config-rule-routing-patterns.md +38 -0
  137. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/data-access-patterns.md +55 -0
  138. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/db-schema-and-dal-patterns.md +109 -0
  139. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/dependency-client-patterns.md +130 -0
  140. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/developer-tooling-patterns.md +70 -0
  141. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/domain-feature-patterns.md +78 -0
  142. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/engineering-patterns.md +119 -0
  143. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/error-contract-patterns.md +55 -0
  144. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/feature-playbook.md +61 -0
  145. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/http-gateway-client-patterns.md +76 -0
  146. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/mq-consumer-patterns.md +55 -0
  147. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/notification-patterns.md +42 -0
  148. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/observability-implementation-patterns.md +101 -0
  149. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/performance-capacity-patterns.md +44 -0
  150. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/protobuf-contract-patterns.md +72 -0
  151. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/public-api-integration-patterns.md +56 -0
  152. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/quality-and-testing-patterns.md +91 -0
  153. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/redis-cache-lock-patterns.md +123 -0
  154. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/release-ops-patterns.md +112 -0
  155. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/reliability-patterns.md +83 -0
  156. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/replay-comparison-patterns.md +32 -0
  157. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/scaffold-and-codegen.md +86 -0
  158. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/source-evidence-map.md +54 -0
  159. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +45 -0
  160. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +80 -0
  161. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/agents/openai.yaml +4 -0
  162. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +117 -0
  163. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/agents/openai.yaml +4 -0
  164. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-approval-auto-reviewer.md +106 -0
  165. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +441 -0
  166. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-context-freshness.md +47 -0
  167. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-credentials-auth.md +13 -0
  168. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-extensions-skills.md +13 -0
  169. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-file-edit-protocol.md +129 -0
  170. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-ide-integration.md +5 -0
  171. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-input-ingestion.md +13 -0
  172. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-instruction-composition.md +13 -0
  173. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-lifecycle-hooks.md +92 -0
  174. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-messaging.md +5 -0
  175. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-runtime-bootstrap.md +5 -0
  176. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +448 -0
  177. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-task-orchestration.md +13 -0
  178. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +123 -0
  179. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-turn-lifecycle.md +131 -0
  180. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +162 -0
  181. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +156 -0
  182. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +146 -0
  183. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/retrieval-agent-safety.md +273 -0
  184. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +202 -0
  185. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/agents/openai.yaml +4 -0
  186. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +62 -0
  187. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/cross-stack-alignment.md +94 -0
  188. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/framework-choice.md +76 -0
  189. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/online-practice-uptake.md +56 -0
  190. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/platform-capabilities.md +91 -0
  191. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/product-page-checklist.md +40 -0
  192. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/qa-release.md +72 -0
  193. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/source-evidence-map.md +82 -0
  194. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +103 -0
  195. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/agents/openai.yaml +5 -0
  196. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +100 -0
  197. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +70 -0
  198. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/agents/openai.yaml +4 -0
  199. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +549 -0
  200. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +97 -0
  201. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/research-prompts.md +66 -0
  202. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/AGENTS.md +32 -0
  203. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/test-public-data-acquisition-recipes.sh +379 -0
  204. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +244 -0
  205. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/agents/openai.yaml +4 -0
  206. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +76 -0
  207. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/framework-middleware-checklist.md +142 -0
  208. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/infra-component-deployment.md +268 -0
  209. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-correlation-recipe.md +124 -0
  210. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-schema-canonical.md +208 -0
  211. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +105 -0
  212. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/obs-stack-architecture.md +107 -0
  213. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +95 -0
  214. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +11 -0
  215. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +303 -0
  216. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/agents/openai.yaml +4 -0
  217. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +163 -0
  218. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/config-center-via-etcd.md +245 -0
  219. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/custom-control-plane-boundary.md +298 -0
  220. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-cli-concrete-recipe.md +312 -0
  221. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-pipeline.md +165 -0
  222. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/env-and-lane-matrix.md +126 -0
  223. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/lane-orchestration-control-plane.md +383 -0
  224. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/multi-region-and-cluster.md +135 -0
  225. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +149 -0
  226. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/python-package-registry-release.md +462 -0
  227. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/rollback-playbook.md +123 -0
  228. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +231 -0
  229. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/version-authority-and-deprecation.md +21 -0
  230. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +276 -0
  231. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/agents/openai.yaml +4 -0
  232. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +127 -0
  233. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/framework-middleware.md +143 -0
  234. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +90 -0
  235. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/http-response-envelope-contract.md +24 -0
  236. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +127 -0
  237. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/multi-env-routing.md +192 -0
  238. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/protobuf-http-contract-signals.md +64 -0
  239. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +124 -0
  240. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/rpc-framework-recipe.md +494 -0
  241. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-choice.md +113 -0
  242. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-migration-playbook.md +231 -0
  243. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +131 -0
  244. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +235 -0
  245. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/agents/openai.yaml +4 -0
  246. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/adr-convention.md +146 -0
  247. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-checklist.md +30 -0
  248. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-evaluation-report-template.md +25 -0
  249. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-execution-spec.md +108 -0
  250. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-sop.md +457 -0
  251. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-templates.md +24 -0
  252. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/artifact-egress-confidentiality.md +58 -0
  253. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +86 -0
  254. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/cross-repo-coordination.md +46 -0
  255. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +192 -0
  256. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +62 -0
  257. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +45 -0
  258. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/diagnostic-spec-match-gate.md +36 -0
  259. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +35 -0
  260. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dormant-code-activation.md +47 -0
  261. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/existing-project-assessment-report.md +223 -0
  262. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/external-skill-augmentation.md +46 -0
  263. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/feature-deprecation-cascade.md +15 -0
  264. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/high-risk-resilience-gates.md +73 -0
  265. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-completeness-and-minimality.md +120 -0
  266. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-entry-reentry-gate.md +122 -0
  267. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/modular-monolith-heuristic.md +105 -0
  268. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +115 -0
  269. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +62 -0
  270. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-attributes.md +112 -0
  271. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-remediation-program.md +88 -0
  272. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +27 -0
  273. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +52 -0
  274. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/review-reception.md +34 -0
  275. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/shared-gate-artifact-classification.md +76 -0
  276. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/source-evidence-map.md +31 -0
  277. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/status-tracker-sync.md +77 -0
  278. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/sync-spec-repo-contract.md +25 -0
  279. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +34 -0
  280. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/worktree-mechanics.md +55 -0
  281. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/AGENTS.md +18 -0
  282. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/check-agent-contract-coverage.sh +213 -0
  283. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +136 -0
  284. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/agents/openai.yaml +9 -0
  285. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/analytics-visualization-interactions.md +206 -0
  286. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +108 -0
  287. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/complex-creation-interactions.md +194 -0
  288. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +214 -0
  289. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +53 -0
  290. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +129 -0
  291. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +97 -0
  292. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +79 -0
  293. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +63 -0
  294. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +146 -0
  295. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +250 -0
  296. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +237 -0
  297. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +65 -0
  298. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +237 -0
  299. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +324 -0
  300. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-web-desktop-patterns.md +456 -0
  301. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +114 -0
  302. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +79 -0
  303. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/resource-management-interactions.md +113 -0
  304. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/scenario-community-patterns.md +133 -0
  305. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +130 -0
  306. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +47 -0
  307. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/trust-sensitive-ai-and-data-patterns.md +96 -0
  308. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +106 -0
  309. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +176 -0
  310. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +111 -0
  311. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +157 -0
  312. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/agents/openai.yaml +4 -0
  313. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/ai-service-integration-boundaries.md +57 -0
  314. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-contract-and-schema.md +62 -0
  315. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-security-boundaries.md +39 -0
  316. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +46 -0
  317. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/async-execution-model.md +24 -0
  318. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/background-jobs-and-scheduling.md +18 -0
  319. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/batch-and-pipeline-architecture.md +11 -0
  320. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/config-secrets-runtime.md +22 -0
  321. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-modeling-and-migrations.md +64 -0
  322. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +211 -0
  323. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +263 -0
  324. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +281 -0
  325. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/observability-and-ops.md +26 -0
  326. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +20 -0
  327. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/redis-cache-coordination.md +41 -0
  328. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/reliability-and-error-contract.md +17 -0
  329. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/source-evidence-map.md +55 -0
  330. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/web-framework-boundaries.md +26 -0
  331. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +143 -0
  332. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/agents/openai.yaml +4 -0
  333. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +16 -0
  334. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/async-and-worker-patterns.md +24 -0
  335. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +18 -0
  336. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +13 -0
  337. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/dependency-client-patterns.md +39 -0
  338. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/error-handling-patterns.md +26 -0
  339. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/feature-playbook.md +43 -0
  340. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/observability-implementation-patterns.md +31 -0
  341. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/project-structure-and-tooling.md +24 -0
  342. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +52 -0
  343. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/redis-cache-lock-patterns.md +78 -0
  344. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/schema-and-validation-patterns.md +23 -0
  345. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/source-evidence-map.md +56 -0
  346. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/sqlalchemy-and-migrations-patterns.md +99 -0
  347. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +61 -0
  348. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/web-framework-patterns.md +35 -0
  349. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +91 -0
  350. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/agents/openai.yaml +4 -0
  351. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/config-runtime-readback.md +20 -0
  352. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/mr-merge-authorization.md +31 -0
  353. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/post-release-env-reset.md +31 -0
  354. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-closeout-evidence.md +20 -0
  355. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-scope-confirmation.md +21 -0
  356. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +20 -0
  357. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/test-scope-prompt.md +24 -0
  358. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/watcher-discipline.md +14 -0
  359. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/SKILL.md +64 -0
  360. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/agents/openai.yaml +4 -0
  361. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/comment-safe-release-doc.md +19 -0
  362. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-evidence-workflow.md +23 -0
  363. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-testing-scope-section.md +15 -0
  364. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +87 -0
  365. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/agents/openai.yaml +4 -0
  366. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +130 -0
  367. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/agents/openai.yaml +4 -0
  368. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/prd-composition-contract.md +35 -0
  369. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/requirement-closure-contract.md +86 -0
  370. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/security-four-questions.md +38 -0
  371. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/SKILL.md +91 -0
  372. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/agents/openai.yaml +4 -0
  373. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +88 -0
  374. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/agents/openai.yaml +4 -0
  375. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +337 -0
  376. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/agents/openai.yaml +4 -0
  377. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/analysis-parse-fix-test-challenge-replay.md +47 -0
  378. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attribution-verification.md +69 -0
  379. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/bootstrap-slim-c3-obligation-table.md +112 -0
  380. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
  381. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +162 -0
  382. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +507 -0
  383. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +86 -0
  384. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/evidence-card-template.md +51 -0
  385. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/example-domain-preselect.md +79 -0
  386. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +57 -0
  387. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-lifecycle-handoff.md +65 -0
  388. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +194 -0
  389. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +75 -0
  390. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +286 -0
  391. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +190 -0
  392. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/l0-l1-l2-routing.md +114 -0
  393. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/online-skill-review.md +47 -0
  394. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +164 -0
  395. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +90 -0
  396. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +320 -0
  397. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +16 -0
  398. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-feedback-mining.md +33 -0
  399. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-finding-standards.md +57 -0
  400. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-rubric.md +40 -0
  401. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +118 -0
  402. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/skill-listing-budget.md +19 -0
  403. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +254 -0
  404. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +658 -0
  405. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/two-source-extraction-pattern.md +167 -0
  406. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +179 -0
  407. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-routing-map.md +51 -0
  408. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +180 -0
  409. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/AGENTS.md +18 -0
  410. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +1452 -0
  411. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-evidence-card-leak.sh +491 -0
  412. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-mr-target-freshness.sh +173 -0
  413. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +488 -0
  414. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +419 -0
  415. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +197 -0
  416. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +327 -0
  417. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +401 -0
  418. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing.rb +248 -0
  419. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/generic-r0-leak-scan.sh +282 -0
  420. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/governing-chain-diff.py +321 -0
  421. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +964 -0
  422. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +708 -0
  423. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +540 -0
  424. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-lifecycle.rb +51 -0
  425. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-pending-status.rb +55 -0
  426. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +829 -0
  427. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +1203 -0
  428. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_r0_status.sh +75 -0
  429. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_register_pending_exclusion.sh +137 -0
  430. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +173 -0
  431. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_route_drift.sh +377 -0
  432. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +833 -0
  433. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +491 -0
  434. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh +114 -0
  435. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_mr_target_freshness.sh +261 -0
  436. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +538 -0
  437. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +154 -0
  438. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +190 -0
  439. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_surface_binding.sh +178 -0
  440. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_prose_target.sh +86 -0
  441. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_generic_r0_leak_scan.sh +131 -0
  442. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_git_identity_predicate_gate.sh +243 -0
  443. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_governing_chain_diff.sh +419 -0
  444. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +120 -0
  445. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +724 -0
  446. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +414 -0
  447. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_registration.sh +34 -0
  448. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +205 -0
  449. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +194 -0
  450. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_credential_cwd.sh +61 -0
  451. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +111 -0
  452. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +53 -0
  453. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +257 -0
  454. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +98 -0
  455. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/agents/openai.yaml +4 -0
  456. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/input-state-machines.md +36 -0
  457. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/streaming-rich-output.md +130 -0
  458. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/terminal-side-channels.md +96 -0
  459. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +408 -0
  460. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/agents/openai.yaml +4 -0
  461. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/AGENTS.md +18 -0
  462. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/bitable-setup.md +573 -0
  463. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/README.md +120 -0
  464. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/github-actions.yml +119 -0
  465. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/gitlab-ci.yml +76 -0
  466. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/jenkins.Jenkinsfile +106 -0
  467. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +279 -0
  468. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/gen_report.py +2807 -0
  469. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/makefile-template.md +200 -0
  470. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/report-config-schema.md +272 -0
  471. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/run_pytestless.py +475 -0
  472. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/source-to-case-workflows.md +258 -0
  473. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-marker-conventions.md +316 -0
  474. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +145 -0
  475. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/AGENTS.md +16 -0
  476. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.dart +129 -0
  477. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.go +197 -0
  478. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.py +135 -0
  479. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.ts +285 -0
  480. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/test_gen_report.py +2144 -0
  481. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +62 -0
  482. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +212 -0
  483. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/agents/openai.yaml +4 -0
  484. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +75 -0
  485. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +50 -0
  486. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/data-and-workflow-testing.md +34 -0
  487. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/design-closed-contract-oracles.md +31 -0
  488. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +71 -0
  489. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +240 -0
  490. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +235 -0
  491. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/non-functional-specialized-scenarios.md +296 -0
  492. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/rd-testing-standard-template.md +126 -0
  493. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/run-killing-mutation-walk.md +43 -0
  494. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +136 -0
  495. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/source-evidence-map.md +59 -0
  496. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/structured-tc-input-translation.md +67 -0
  497. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +392 -0
  498. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-data-and-determinism.md +39 -0
  499. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +92 -0
  500. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/unit-testing.md +46 -0
  501. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/vendored-contract-drift-checklist.md +64 -0
  502. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/verify-enforcement-mechanisms.md +18 -0
  503. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/AGENTS.md +17 -0
  504. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.py +140 -0
  505. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.test.sh +75 -0
  506. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.py +170 -0
  507. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.test.sh +87 -0
  508. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.go +198 -0
  509. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.test.sh +109 -0
  510. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/test_mutation_backup_recipe.sh +237 -0
  511. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +184 -0
  512. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/agents/openai.yaml +4 -0
  513. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/comment-safe-feishu.md +93 -0
  514. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/cross-model-co-review.md +3 -0
  515. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +60 -0
  516. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +17 -0
  517. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/session-vantage-leakage.md +58 -0
  518. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +126 -0
  519. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/agents/openai.yaml +4 -0
  520. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +47 -0
  521. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/embedded-h5-in-host.md +87 -0
  522. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +194 -0
  523. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/source-evidence-map.md +60 -0
  524. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +190 -0
  525. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +83 -0
  526. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/SKILL.md +179 -0
  527. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/agents/openai.yaml +4 -0
  528. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/shared-branch-rebase.md +25 -0
  529. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/AGENTS.md +23 -0
  530. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_status.sh +207 -0
  531. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_sweep.sh +481 -0
  532. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-status.sh +325 -0
  533. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-sweep.sh +245 -0
  534. package/dist/assets/release.json +2797 -0
  535. package/dist/claude-adapter.d.ts +9 -0
  536. package/dist/claude-adapter.js +240 -0
  537. package/dist/cli-worker.d.ts +1 -0
  538. package/dist/cli-worker.js +32 -0
  539. package/dist/cli.d.ts +22 -0
  540. package/dist/cli.js +214 -0
  541. package/dist/codex-host.d.ts +30 -0
  542. package/dist/codex-host.js +162 -0
  543. package/dist/fs-safe.d.ts +21 -0
  544. package/dist/fs-safe.js +241 -0
  545. package/dist/index.d.ts +2 -0
  546. package/dist/index.js +1 -0
  547. package/dist/manifest.d.ts +8 -0
  548. package/dist/manifest.js +135 -0
  549. package/dist/opencode-adapter.d.ts +10 -0
  550. package/dist/opencode-adapter.js +416 -0
  551. package/dist/operations.d.ts +3 -0
  552. package/dist/operations.js +956 -0
  553. package/dist/paths.d.ts +20 -0
  554. package/dist/paths.js +4 -0
  555. package/dist/types.d.ts +58 -0
  556. package/dist/types.js +1 -0
  557. package/dist/unified.d.ts +4 -0
  558. package/dist/unified.js +64 -0
  559. package/dist/version.d.ts +2 -0
  560. package/dist/version.js +5 -0
  561. package/package.json +35 -0
@@ -0,0 +1,254 @@
1
+ # Source Register
2
+
3
+ Keep source-specific provenance outside the distributed repository. Public skills must remain usable without private repositories, internal documents, local paths, contributor identities, or organization-specific examples.
4
+
5
+ ## Register Template
6
+
7
+ | Field | Required answer |
8
+ | --- | --- |
9
+ | Task or extraction name | Name the reusable workflow or source family with a source-neutral label. |
10
+ | Purpose | State the future failure or quality bar this extraction addresses. |
11
+ | Scope | List included source classes, target skills, sibling boundaries, and exclusions. |
12
+ | Depth | Record wording cleanup, targeted check, file refresh, artifact inventory, full workflow extraction, or tooling change. |
13
+ | Failure mode analysis | State the bad output, leakage, shallow rule, or overclaim the work prevents. |
14
+ | Lifecycle impact | Name the affected intake, design, implementation, testing, launch, iteration, onboarding, and documentation stages. |
15
+ | Evidence plan | List source categories and how each is inspected, routed, excluded, or marked unavailable. |
16
+ | Completion standard | Name the scenario, command, review, and source-map evidence required to finish. |
17
+
18
+ Target-output map:
19
+
20
+ | Target | Owner role | Expected decision | Source mechanism | Actual diff or reason |
21
+ | --- | --- | --- | --- | --- |
22
+
23
+ Required for upstream-owner skill changes:
24
+
25
+ | Upstream rule | Downstream owner | Expected executable behavior | Status (updated, unchanged, routed, or not-applicable) | Evidence |
26
+ | --- | --- | --- | --- | --- |
27
+ | Requirement records stay with the product workflow while generic Wiki/Base mechanics remain resource-specific | `lark-wiki` / `lark-base` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#Create or reuse only the requested resource | routed | `product-rd-workflow/SKILL.md`; `bootstrap.md` |
28
+ | Structured testcase delivery includes its testcase Base lifecycle | `test-artifact-management` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/test-artifact-management/SKILL.md#but it is optional and must not trigger creation | updated | `test-artifact-management/SKILL.md`; `test-artifact-management/references/bitable-setup.md` |
29
+ | Review-client compatibility is capability-based and leaves model selection to host configuration | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/kimi_review.sh | updated | `code-review/SKILL.md`; `code-review/scripts/kimi_review.sh`; `code-review/scripts/opencode_review.sh`; `code-review/scripts/test_review_client_compat.py` |
30
+ | A formal reviewer packet that exceeds the safe inline argv bound remains reviewable only when the alternate transport preserves exact packet identity and exposes one pathless, hash-bound packet reader without promoting candidate bytes to system-prompt authority; local file-descriptor exhaustion remains a client-runtime failure rather than model or authentication evidence | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/kimi_review.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; implementation in `code-review/scripts/kimi_review.sh` and `code-review/scripts/kimi_packet_mcp.py`, parsing in `code-review/scripts/parse_cli_review.py`, regression coverage in `code-review/scripts/test_cli_review_wrappers.sh`, `code-review/scripts/test_kimi_packet_mcp.py`, and `code-review/scripts/test_review_client_compat.py`, contract in `code-review/references/client-routing.md`, and round plan in `specs/024-kimi-packet-review/plan.md`. RED covered the old oversize refusal, the missing pathless server/parser boundary, a physical line larger than one bounded MCP result, the 16-64 KiB template-bearing argv exception, an undocumented `${HOME}` placeholder that the old live-variable allowlist treated as safe, a non-inline prompt that revealed the tail receipt before the model read the packet, a template-free candidate embedded at explicit-agent system-prompt authority, an inline argv exposure window widened from 120 to 600 seconds, and MCP delivery that skipped the forbidden-tool canary. GREEN keeps the 16 KB inline ceiling and 120-second inline timeout, runs the Read/Glob/Grep canary before both delivery modes, uses pathless UTF-8-safe byte chunks with per-call packet hash verification and parser-side exact byte matching for every over-inline packet, and withholds both candidate bytes and the receipt from non-inline prompts. MCP review retains the controller-granted timeout up to 600 seconds, overlong single-line packets remain bounded, and capability-probe/formal-run EMFILE stays a local client failure. Earlier live reviews traversed the repaired transport and supplied RED cases; the current exact-candidate Kimi run timed out, so its final verdict remains pending rather than passed. |
31
+ | A packet-only reviewer must spend its bounded lane on the frozen packet, not on workspace discovery, and a process timeout cannot erase a complete bound verdict that was already emitted and exported; wrapper timeouts inherit the controller budget, while structured provider authentication and billing errors must not be lost when stderr is empty or later events differ | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/opencode_review.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; implementation in `code-review/scripts/opencode_review.sh` and `code-review/scripts/parse_opencode_review.py`, schema in `code-review/scripts/egress_schema.py`, regression coverage in `code-review/scripts/test_opencode_review_retry.sh`, `code-review/scripts/test_opencode_review_concurrency.sh`, `code-review/scripts/test_parse_opencode_review.sh`, and `code-review/scripts/test_egress_schema.sh`, contract in `code-review/references/client-routing.md`, and plan in `specs/024-kimi-packet-review/plan.md`. RED fixtures reproduced the brittle 30-second boundary cap, the opposite failure where an uncapped hung probe consumed a long lane, completed-export timeout loss, export-only timeout recovery that could accept a partial run, incomplete idless and invalid timeout exports becoming terminal parser errors, structured HTTP 402 with empty stderr, earlier 401/402 events hidden by a later status, a wrapper that used to require `skill=true` for zero-owner runs, and a parser that later still accepted it. The generated agent now wildcard-denies all capabilities, permits only controller-selected skill names, requires `skill` exactly when the controller selected an owner, caps the local boundary probe at 60 seconds without shrinking the formal run timeout, requires matching terminal-stop evidence from both the event stream and export before timeout recovery, preserves incomplete exports as fallback-eligible timeouts while keeping explicit different-session exports terminal, and classifies the entire bounded event stream. Live small-diff review returned a formal passed verdict. The later exact-candidate run returned DeepSeek `Insufficient Balance`; the wrapper now maps that structured event to fallback-eligible quota instead of terminal transport failure. Claude was separately probed on the same packet and returned an external weekly-quota response correctly classified as fallback-eligible quota. |
32
+ | Exact packet coverage is a byte-range invariant, not a ban on EOF confirmation: after full contiguous coverage, one exact-end request may return a hash-bound empty chunk, while altered bodies, gaps, and beyond-end requests remain non-passing | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/parse_cli_review.py | updated | RED/GREEN server, parser, and wrapper fixtures in `code-review/scripts/test_kimi_packet_mcp.py`, `code-review/scripts/test_review_client_compat.py`, and `code-review/scripts/test_cli_review_wrappers.sh`; implementation in `code-review/scripts/kimi_packet_mcp.py` and `code-review/scripts/parse_cli_review.py`; contract in `specs/024-kimi-packet-review/plan.md`. |
33
+ | A compatibility probe cannot silently enlarge a controller-granted reviewer deadline: its elapsed time and validation consume the same lane budget as the formal invocation | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/claude_review.sh | updated | A delayed-baseline fixture in `code-review/scripts/test_claude_review_probe.sh` fails when the formal timeout remains unchanged and passes only when the baseline elapsed time is deducted; contract in `code-review/references/client-routing.md` and `specs/024-kimi-packet-review/plan.md`. |
34
+ | Generated reviewer instructions may interpolate only controller names that independently satisfy the package grammar, and transport peers share one bound constant instead of drifting compatible values | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/kimi_review.sh | updated | Kimi rejects an instruction-shaped `--review-skill` before inference in `code-review/scripts/test_cli_review_wrappers.sh`; `code-review/scripts/parse_cli_review.py` imports `MAX_CHUNK_BYTES` from the pathless server rather than repeating its literal. |
35
+ | A timed-out transport tail is complete only when both terminal evidence and the frozen semantic contract are complete; for profile-bound OpenCode recovery, the final export must carry exactly every required concern conclusion, while a legacy call with no frozen set cannot recover | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/parse_opencode_review.py | updated | `code-review/scripts/opencode_review.sh` passes the frozen concern IDs to the parser; `code-review/scripts/test_parse_opencode_review.sh` proves legacy, missing-coverage, and incomplete terminal evidence remain timeout while exact profile-bound coverage may recover. |
36
+ | Host vocabulary and host capability need different upgrade policies: a same-version safe-mode baseline may authorize new bare commands, but a baseline-only skill remains drift evidence and cannot become callable formal authority until pinned or controller-selected | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/parse_probe_result.py | updated | Formal review showed that a future safe-mode regression could leak a user skill while preserving help text. `code-review/scripts/init_policy_matrix.py` now expects baseline-only new skills to fall back, its mutation re-authorizes them and must fail, and `code-review/scripts/test_claude_review_probe.sh` drives the wrapper path. Command adaptation remains dynamic and same-version-bound. |
37
+ | A fail-closed reviewer stop must carry the evidence that justifies it, and its relay is bounded by construction rather than by a secret denylist | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/concern_excerpt.py | updated | `code-review/SKILL.md`; `code-review/scripts/concern_excerpt.py`; `code-review/scripts/parse_cli_review.py`; `code-review/scripts/claude_review.sh`; `code-review/scripts/AGENTS.md`; `code-review/scripts/test_concern_excerpt.sh` |
38
+ | A client-diagnostic allowlist predicated on an upstream CLI's wording re-breaks on every release, so it matches the invariant claim by shape | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/concern_excerpt.py | updated | `code-review/SKILL.md`; `code-review/scripts/parse_cli_review.py`; `code-review/scripts/test_concern_excerpt.sh` |
39
+ | Relocating the always-on injection layer is a move, not a new every-session injection surface | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-size-budget.sh | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/check-size-budget.sh`; `skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh` |
40
+ | A budget gate offers no way to nominate its own baseline: it measures the same path across revisions, and a relocation blocks for a human decision rather than being recognised automatically — supersedes the rename-following row above | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-size-budget.sh | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/check-size-budget.sh`; `skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh` |
41
+ | A hand-maintained skill catalog is only allowed to exist when the same landing makes it mechanically undriftable: the catalog and the always-on routing layer gate each other, and a set-difference on names is not that gate | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-ccl-skills.sh | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/check-ccl-skills.sh`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `skill-extraction-workflow/references/dual-track-review-gate.md`; `docs/SKILLS.md` |
42
+ | Claiming a skill must live in the always-on injection layer is a falsifiable claim, not an argument — that layer has a zero-net-growth budget, so measure the A/B benefit before spending bytes | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/eval-routing-bank.rb | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/eval-routing-bank.rb`; `skill-extraction-workflow/references/dual-track-review-gate.md`; `docs/SKILLS.md` |
43
+ | A premise about the system's own topology — who calls this, what ships where, which states are reachable — is verified before a design rests on it, and repeated non-convergence is read as a wrong premise before it is escalated as a design tradeoff | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#must not rest on a premise | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/references/dual-track-review-gate.md` |
44
+ | The requirement-front layer keeps one skill per deliverable and is named as one family, so the router names four requirement-* owners whose boundaries are decidable from the description alone | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#route clarification to | updated | `product-rd-workflow/SKILL.md`; `eval/routing-tasks.jsonl`; `docs/SKILLS.md` |
45
+ | Retargeting a pointer at a renamed skill carries no obligation of its own: the owner's bytes are reproduced by rewriting the identifier, so there is no delta to declare | `defect-diagnosis` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `defect-diagnosis/SKILL.md` pointer retarget only |
46
+ | Same retarget, risk-router side | `feature-risk-router` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `feature-risk-router/SKILL.md` pointer retarget only |
47
+ | Same retarget, Go architecture side | `go-microservice-architecture` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `go-microservice-architecture/SKILL.md` pointer retarget only |
48
+ | Same retarget, inference side | `llm-inference-integration` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `llm-inference-integration/SKILL.md` pointer retarget only |
49
+ | Same retarget, observability side | `platform-observability` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `platform-observability/SKILL.md` pointer retarget only |
50
+ | Same retarget, connectivity side | `platform-service-connectivity` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `platform-service-connectivity/SKILL.md` pointer retarget only |
51
+ | Same retarget, Python architecture side | `python-service-architecture` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `python-service-architecture/SKILL.md` pointer retarget only |
52
+ | Same retarget, doc-finalization side | `tighten-doc` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `tighten-doc/SKILL.md` pointer retarget only |
53
+ | A skill's invocation identity is behavior: renaming it retargets the delegation-owner hook's recognition token, and leaving that token behind makes the hook stop recognising a legitimate invocation and prompt forever | `multi-agent-delegation` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/multi-agent-delegation/SKILL.md#Before a worker prompt includes a delivery spec | updated | `multi-agent-delegation/SKILL.md`; reverting the hook token turns `hooks/test_guard_delegation_owner.sh` RED on the warm-dispatch assertion, control green |
54
+ | The same identity rule on the release side: the frontmatter name is what the host resolves, so a rename that does not carry it is caught by the name-matches-directory check | `platform-release-engineering` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/platform-release-engineering/SKILL.md#Language-specific library/module tagging mechanics | updated | `platform-release-engineering/SKILL.md`; reverting the frontmatter name yields `frontmatter_name_description_failed`, control ok |
55
+ | A renamed test-asset owner resolves under its new name only: the frontmatter name is what hosts bind, so the rename is a real identity change and cannot ride on a historical row that the same rename rewrote | `test-artifact-management` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/test-artifact-management/SKILL.md#Do not collapse the two axes into one enum | updated | `test-artifact-management/SKILL.md`; reverting its frontmatter name yields `frontmatter_name_description_failed`, control ok |
56
+ | A Skip clause predicated on state the utterance cannot expose never fires, so the over-claiming neighbour keeps winning: a routing judgement is written as a request SHAPE, not as a prior state | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#description | updated | `testing-strategy/SKILL.md`; the `once the layer is already chosen` wording measures 4/10, the request-shape wording 10/10 (10 rounds each, claude-haiku-4-5) |
57
+ | A coordinator that never declares its own handoff artefact loses the request to whichever neighbour owns the word: a test-scope prompt derived from confirmed release scope is a release handoff | `release-coordination` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/eval-routing-bank.rb | updated | `release-coordination/SKILL.md`; 0/10 before, 10/10 after (10 rounds each) |
58
+ | A homograph collision decides routing when nobody claims the request: the audit sense of a word and the coverage sense of the same word send an unclaimed utterance to the literally-nearest gate skill | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/eval-routing-bank.rb | updated | `skill-extraction-workflow/SKILL.md`; 3/10 before, 10/10 after, all pre-change losses at confidence 0.15–0.45 |
59
+ | A mechanism skill whose trigger is phrased as a delivery entry intercepts entry requests at high confidence; the trigger names the mechanical act, not the lifecycle moment | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/eval-routing-bank.rb | updated | `worktree-isolation/SKILL.md`; narrowing alone moves 3/10 to 6/10 — that is the ceiling of the executor-side fix |
60
+ | A reviewer-dispatch skill does not adjudicate delivery value; without that Skip it becomes the catch-all for any utterance mentioning a commit | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/SKILL.md#description | updated | `code-review/SKILL.md`; 1/10 before, 10/10 after, pre-change losses at confidence 0.55–0.85 |
61
+ | Bounding a relayed reviewer stop means removing the free-text path, not capping it — a length cap still rests on the denylist the shape summary exists to avoid; extends the relay-bounded-by-construction row above to the remaining lane, whose parser keeps the whole text for internal audit while only the egress is bounded | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/opencode_review.sh | updated | `code-review/SKILL.md`; `code-review/scripts/opencode_review.sh`; `code-review/scripts/test_opencode_review_retry.sh`; `code-review/scripts/test_review_gate.sh`; `code-review/references/client-routing.md` |
62
+
63
+ **Un-landed lesson note (reviewer-relay round).** That round surfaced a process
64
+ lesson this ledger deliberately does NOT carry as a row: *extending an invariant to
65
+ a new call site re-opens the mechanism choice, so the owning module's recorded
66
+ REJECTED alternatives are read before picking the helper.* It was self-caught
67
+ in-round with no gate firing — three review rounds each defeated a different
68
+ copy-detector before the design converged on building the verdict from an
69
+ allowlist, which is the shape the owning module had already recorded as the
70
+ resolution of the same class one level down.
71
+
72
+ It gets a note, not a row, because the table's grammar assumes a landed
73
+ upstream→downstream propagation with owner-scoped firing evidence, and this lesson
74
+ has none: no mechanism makes the next agent read the rejection. Writing a row would
75
+ have required inventing a firing path for something that does not fire — the same
76
+ record-masquerading-as-mechanism failure the lesson is about. Landing it needs a
77
+ real gate in the owning skill, which is its own shared-skill round.
78
+
79
+ **Superseded-row note (routing round).** The retarget row this ledger carried for
80
+ `testing-strategy` was true of the round that added it, and stopped being true of the
81
+ accumulated candidate: the same owner later gained a description change, so its package
82
+ no longer reproduces under the rename pairs and the class it claimed no longer holds. The
83
+ row was removed rather than left asserting something false; its accurate replacement is the
84
+ routing-surface row for the same owner, which covers both the retarget and the description.
85
+ This is the second documented bend of the append-only contract, and it has the same shape as
86
+ the first: the row described a state the tree no longer has. The general problem — rows are
87
+ written per round while the gate judges the accumulated diff — was recorded here as a
88
+ follow-up and is closed by the Round-consolidation rule below.
89
+
90
+ **Round-consolidation rule (append-once).** The fix is an operating rule on WRITE TIMING,
91
+ not a relaxation of append-only: in a multi-round program, draft candidate rows in the round
92
+ plan or scratch while rounds iterate, and APPEND each row to this ledger exactly once, in
93
+ final form, in the same squashed landing commit as the changes it declares — so the
94
+ impact-chain gate's round partition holds the rows and their declared changes together, and
95
+ no landed row can describe a state the accumulated candidate no longer has. A wrong or stale
96
+ row caught before push is corrected by re-pressing the unpushed landing commit (the row was
97
+ never landed), never by editing the ledger in a follow-up commit; after push, append-only
98
+ holds in full and the correction is a superseding row or a dated note pointing at the
99
+ replacement — never an edit or a silent deletion. Both documented bends above predate this
100
+ rule and share the shape it removes.
101
+
102
+ **Path-rewrite note (tier 3 rename, 2026-08-11).** Rows above that cite
103
+ `test-artifact-management/SKILL.md`, `platform-release-engineering/SKILL.md`, or
104
+ `multi-agent-delegation/SKILL.md` and predate that tier were written against the
105
+ paths those owners had before it: `testcase-writer/`, `platform-release-and-rollout/`,
106
+ and `agentic-execution/` respectively. The paths were retargeted mechanically so the
107
+ evidence check can still resolve them; the rows' own content and order are unchanged.
108
+ This is the one place the ledger's append-only contract bends, and only for a path
109
+ that no longer exists — read an older row's owner through this mapping.
110
+
111
+ | The firing-path anchor uses a changed line's SHAPE as a proxy for a changed obligation, so a diff shape carrying no anchorable rule line has no way to declare evidence. A frontmatter-description-only change is such a shape — and unlike the two no-behaviour classes it DOES carry behaviour, so the answer is a canonical field locator naming the description, never an exemption | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `scripts/impact-chain-gate.rb`; `scripts/test_check_ccl_impact_chain_refscripts.sh`; `references/external-practice-controls.md`. Applied mutations, each attributed differentially with the unmutated suite green: forcing the predicate true lets a body edit, a sibling-reference edit and a second frontmatter key each ride along; dropping the YAML-string check admits a nested mapping; reverting the entry boundary to indentation breaks a folded description containing a blank line; accepting any substring as the locator admits one that merely survived the edit |
112
+ | A gate change that LOOSENS a verdict removes evidence instead of adding a red, so it must clear the design-time operability check in both directions — and before a human is offered options, not only before landing; a `recommended` label asserts that the per-option evidence-removal comparison was made | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/SKILL.md#a loosening must be checked hardest | updated | `SKILL.md` operability-check scope; `references/dual-track-review-gate.md` leg (e). Observed failure: this round's first design proposed a third `not-required` class by analogy to two existing ones, and the adversarial review's first finding refuted it — the exempted class was the one carrying the most behaviour |
113
+
114
+ | A no-behaviour precondition counts behaviour, not files: a package whose non-description changes are all a proven rename retarget still carries exactly one behaviour, so the two machine-verified classes compose under one normalizer instead of the narrower one refusing what the wider one already cleared | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `scripts/impact-chain-gate.rb`; `scripts/test_check_ccl_impact_chain_refscripts.sh`; `references/external-practice-controls.md`. Applied mutations, differentially attributed with the unmutated suite green: forcing the reproduction true lets a real sibling edit ride along; degrading the normalizer to identity refuses the retarget shape it exists for. Real-input check: on the accumulated integration candidate the owner whose references were byte-proven pure retargets goes from refused to accepted, while an owner whose body carries real rule changes stays refused |
115
+
116
+ | A multi-round review program carries four process controls: a broken chain binding after an owner-file fix is a by-design dead-end recovered by an interim checkpoint naming each lane's un-run remainder plus a human continuation authorization that extends rounds without waiving any lane; convergence and closure declarations are written falsifiably with named axes and named open items; remediation text re-owes the pre-cover axes before returning to the reviewer, with a third same-class round escalating to one full-matrix self-enumeration; and ledger rows land append-once, final-form, in the same squashed round partition as the changes they declare | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; `references/dual-track-review-gate.md` (three merged clauses); `references/source-register.md` (Round-consolidation rule, closing the follow-up recorded in the routing-round note); `scripts/test_ai_coding_implementation_gates.sh` family 9. Observed failures: a tracked chain's content binding broke by design after an owner-file fix and the recovery pattern was improvised; a self-audit's "full lifecycle" claim was caught one axis short by the final challenge; a program burned twenty-plus single-finding rounds before one full-matrix enumeration closed the class; two documented append-only bends shared the rows-written-per-round shape. RED-baseline: 40 family-9 assertions (one per obligation sentence or sub-clause, per the reverse-coverage rule landed alongside; clause pins bound to their owning rule-line, anchors section-bound, ledger pins paragraph-bound), shown red under 81 applied mutations in a throwaway copy — one deletion and one relocation per pin, two cross-moves between the convergence and continuation rule-lines, one same-bullet pointer removal — first-failing label equal to the owning assertion, unmutated controls green before and after |
117
+ | For a contract or prose artifact pinned by a fixture family, pin coverage runs artifact-to-pin — every obligation sentence names the pin that reds when it is deleted, because the walk proves only the pins that exist — and a self-contained walk owes the relocation, reachability, tree-isolation, and parser-completeness probes | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/testing-strategy/SKILL.md#coverage runs in the reverse direction | updated | `testing-strategy/SKILL.md` (reverse-coverage sub-bullet with dual-side pointer); `testing-strategy/references/run-killing-mutation-walk.md` (probe classes). Observed failure: a reviewer found six silently-deletable unpinned obligations after multiple green walks; the sentence-by-sentence enumeration then closed all six at once and was recorded only in a round spec until this landing. RED-baseline: the reverse-coverage and probe pins are in family 9's applied-mutation walk above, including the same-bullet pointer-removal mutation red on the reachability assertion |
118
+
119
+
120
+ Gate: a changed upstream owner requires a non-empty row with an allowed terminal status and evidence naming that owner's SKILL.md. A header-only table is absent. Keep project-specific provenance in a private task artifact, never in this distributed register.
121
+ | An evidence row is judged against the round it landed in, never the accumulating range; classification and presence both narrow to that round while the owner-level RED floor stays cumulative | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/impact-chain-gate.rb`; `skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh`; `skill-extraction-workflow/references/external-practice-controls.md` |
122
+ | Retargeting a pointer at the renamed release coordinator carries no obligation of its own: the owner's bytes are reproduced by rewriting the identifier, so there is no delta to declare | `platform-release-engineering` | behavioral-evidence: not-required identifier-rename; observed-failure: no | updated | `platform-release-engineering/SKILL.md` pointer retarget only |
123
+ | A retarget inside a check's human-readable LABEL is not a retarget of what the check asserts: the pointer-integrity checks take the asserted needle as one argument and the label as another, so renaming the slug in the label leaves the control's verdict invariant and the owner declares a stable control rather than a reproduced package | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is itself unchanged this round; the retarget lands in the label argument of one check in `skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh`. Applied mutation, control green: reverting that label to the old slug leaves the suite GREEN (`routing_pointer_integrity_ok`), which is the pairing evidence that the change moves no verdict — an `identifier-rename` class was refused here because the same owner package carries real rule changes from an earlier round in this range, so the package does not reproduce byte-for-byte under the rename pairs |
124
+ | Re-resolving references into a deleted or renamed section extends past the document's own internal pointers to inbound ones from other artifacts, including non-Markdown files no doc linter opens; the mechanical step is a repo-wide fixed-string grep of the old heading AND its anchor slug, with the exit code read rather than the output eyeballed | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/SKILL.md#on any heading you delete or rename, run | unchanged | `tighten-doc/SKILL.md` is the changed owner key. Observed failure: `docs/ARCHITECTURE.md` and `scripts/install.sh` both pointed readers at a README section that no longer existed, and every gate stayed green because the prior rule scoped to intra-document references and `scripts/check-markdown-links.py` drops the `#fragment`. Applied checks against this repo after landing the recipe: the removed heading exits 1 (clean), the current heading exits 0 (hits), the slug pass returns the example inside the rule itself; the pre-fix recipe was RED on three counts the challenge lane applied — a `-`-leading heading parsed as an option, exit 1 misread as failure, and untracked referrers plus anchor-slug links both missed |
125
+ | A deterministic gate's terminal REASON CODES are the coverage unit — not its assertion count, and not its exit codes, since many distinct fail-closed outcomes share one nonzero exit and auditing by exit code reads them all as covered: the review controller carried 111 assertions while two of its terminal reason codes — including the fail-closed floor that fires when the configured client order holds no cross-family reviewer — had no assertion anywhere in the repository, so a regression in either was undetectable by the suite. What the mutations establish is that detection gap and its closure, not that any lane was in fact recorded clean. Enumerate a gate's terminal outcomes and check each has an owning assertion, rather than reading a large suite as coverage | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_review_gate.sh | updated | `code-review/SKILL.md` is the owner key and is itself unchanged this round; the cases land in `code-review/scripts/test_review_gate.sh`. Applied mutations, each attributed differentially. Control accounting, stated exactly because the two walks ran in different copies: the first walk's disposable copy omitted the owner's `references/`, so one unrelated check that greps a file there failed identically in its control and in every mutant, leaving attribution differential but its control not green; the second walk's copy was complete and its control was green; the final candidate in the worktree runs 138 ok / 0 FAIL. The RED here is mutation-induced, not pre-change: the added cases pass against the unmodified controller by construction, because they assert outcomes it already produces correctly, so the evidence is applied-mutation sensitivity rather than a failing baseline that the change then fixes. Attribution is per-assertion, not per-suite: each owning case passes in its control and fails under its mutant. Rewriting the initial `last_reason_code` literal fails the no-cross-family-reviewer case and nothing else; deleting the empty-packet guard fails the empty-candidate case and nothing else. The two mutations that remove the same-family `continue` also fail a recorded collateral set — 27 and 24 total FAIL lines — because that `continue` carries the whole fallback path, so their evidence is the owning flip, not containment: turning it into `break` flips the precision case that proves the skip does not kill the lane, and dropping it while exiting the fake wrapper before it appends `client_sequence` flips the no-cross-family case while leaving the empty-candidate case GREEN, which is what isolates the added earlier-write clause — that path raises before the client loop exists, so the clause was added to the first case and deliberately not to the second |
126
+ | A flag whose legal range depends on the mode cannot take one static default: the review controller defaulted `--challenge-index` to `0` while challenge mode requires `1..budget`, so the default was illegal in the only mode the flag serves and every challenge invocation that did not pass it by hand failed `invalid_input` before any provider ran. Where each invocation shape already admits exactly one legal value, derive the omitted value from the invocation instead of defaulting it, and key that derivation off the same predicate the rest of the code uses for the distinction | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/review_gate.py | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/review_gate.py` with cases in `code-review/scripts/test_review_gate.sh` and the plan in `specs/012-challenge-index-default/plan.md`. Observed failure: `docs/reviewer-lane-bootstrap-hijack.md` recorded the defect as open, and it was reproduced against the unmodified gate — challenge mode with a budget and no index returned exit 2 `invalid_input`, the same invocation with `--challenge-index 1` returned exit 0. This RED is pre-change, not mutation-induced. Three mutations were applied on the fix, each restored and byte-compared against pristine: restoring the constant default, making non-challenge resolution return 1, and relaxing the tracked range check to accept 0. Per-mutation results and suite counts live in the plan's executed-trace table rather than being restated here, because a count copied into a second place drifts from the candidate as soon as a case is added — which is exactly what a review round caught between two versions of this row. The independent review and the challenge lane both objected that the derivation keyed trackedness off `--autonomous-review-index` rather than `--review-chain-id` (chain `challenge-index-default-r2`, one finding per lane, results retained in this checkout's local review-evidence directory rather than in the tree); the concrete scenario they gave is unreachable today because an orphan index is rejected earlier with `review_chain_invalid`, but the predicate was wrong and now matches `review_chain_tracked`, with a case pinning that rejection so a later guard move cannot silently re-open it. A later round moved the resolution into the parser subclass so a direct `build_parser().parse_args(...)` cannot observe the unresolved value either |
127
+ | A reference whose target NEVER existed is invisible to every control this repo has, and naming it in a landed spec's out-of-scope list does not track it: the delete/rename inbound-reference recipe keys off a rename event that never happened here, and `check-markdown-links.py` resolves only Markdown inline-link destinations in tracked `*.md` files — so a backticked path in prose is unseen even inside Markdown, and any reference in a non-Markdown file is unseen twice over. Repair the pointer to what the repo actually does rather than back-filling the missing target, and when a slice names an out-of-scope defect, give it an owner that outlives the spec | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/scripts/AGENTS.md#A recipe step here must never cite a repo path that does not exist | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/AGENTS.md` (recipe step 4 now points at the slice's own `<NNN>-<slug>/plan.md` under specs/) and `code-review/scripts/test_init_policy_matrix.sh` (its header states the mutation check inline instead of citing the absent file), with the disposition recorded in `specs/010-review-concern-excerpt/plan.md`. Observed failure: `010`'s Scope named the `scripts/AGENTS.md` pointer as a stale out-of-scope item, the spec landed, and the pointer survived — and only one of its two copies was ever found, because the shell-script copy is outside every doc linter's file set. The RED is pre-change and re-computable rather than mutation-induced, and it is the instruction's executability, not a suite: before, step (4) directs a contributor to write into a validation log under a specs/009-claude-review-tool-boundary directory that does not exist (`test -e specs/009-claude-review-tool-boundary/` exits 1); after, it names the slice's own `<NNN>-<slug>/plan.md` under specs/ and cites an exemplar that resolves (`test -e specs/012-challenge-index-default/plan.md` exits 0). No suite moves, and none should — `test_init_policy_matrix.sh` still runs 164 cases / 0 mismatches, because its edit is a header comment; the behavioral delta being claimed is the contract step's, not the matrix's. A first attempt declared this row `not-applicable: docs-only`, which is not a value the gate accepts, and `impact-chain-gate.rb` rejected the commit — recorded because the rejection is the evidence that the declaration is checked rather than trusted. **A mechanical control was designed and rejected on evidence, not skipped**: a repo-wide existence check over backticked path-shaped tokens was prototyped and dogfooded first — 1167 such tokens, 139 occurrences across 75 distinct paths do not resolve, and the dominant class is CORRECT content (a product-agnostic skill naming a path in the consuming product repo: `test/.report-config.json`, `.feishu/project.yaml`, `ci/agent-gates.gitlab-ci.yml`) plus deliberate test fixtures. The REPOSITORY-WIDE version of that check fails the author-dogfood and marginal-cost legs of the design-time operability check, so landing it would have bought noise that trains readers to ignore the gate. **Superseded within the same landing:** narrowing the same idea to `specs/` — a directory only this repo has, whose contents are enumerable — passes both legs, and `scripts/check-spec-references.py` now runs in `make test` and CI (plan: specs/014-spec-reference-existence-gate/plan.md). The residual risk recorded one sentence earlier is therefore RETIRED for `specs/` citations; what stays review-only is the wider class this row first described — a citation into a path the repo does not own, and a prose pointer that names a section rather than a path. The generalizable half — an out-of-scope item named in a spec needs an owner or it rots regardless of the "named rather than silently dropped" wording — is **routed to `product-rd-workflow` (plan authoring), pending its own dual-track round; not landed here** |
128
+ | A reviewer-isolation gate must derive host-owned vocabulary from the exact bounded client it invokes, not from a repository snapshot: otherwise each upstream built-in-name release either disables that lane or pressures the parser to weaken its real tool and authority checks | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/claude_review.sh | updated | `code-review/SKILL.md` remains the owner key. The original `/import` failure proved that treating an unknown bare host name as a terminal customization could stop the entire reviewer chain; the interim recoverable classification prevented total chain outage but still left Claude unavailable. Claude Code 2.1.233 then reproduced the remaining defect with `auto-mode-setup`, `autocompact`, `list-agents`, and non-empty `terminal_slash_commands`: a synthetic successful owner-aware stream failed before formal verdict parsing. `claude_review.sh` now captures a same-executable safe-mode baseline from an independent empty cwd with tools, plugins, MCP, settings, workspace instructions, custom agents, and user skills disabled; it trusts that skill isolation only while the installed CLI's own `--safe-mode` help contract still states that skills are disabled, so an upstream semantic-contract change refuses and cascades rather than laundering personal vocabulary. `parse_probe_result.py` permits unique whole-string commands/skills from that same-version baseline when each is either bare or an already pinned built-in with a non-bare spelling; an unknown namespaced or path-shaped baseline entry is still a proven customization and remains terminal. Tool, plugin, MCP, permission, terminal-command, and other namespaced-customization breaches remain terminal; version mismatch and unbaselined bare host vocabulary refuse but cascade as unverified capability drift. RED/GREEN coverage is in `init_policy_matrix.py`, `test_init_policy_matrix.sh`, `test_parse_probe_result.sh`, and `test_claude_review_probe.sh`. A live 2.1.233 safe-mode baseline exposed no tools, MCP, plugins, or user-skill overlap, and later exact-candidate calls produced formal verdicts after an earlier quota response was classified separately. |
129
+ | Discharging the vocabulary-predicate class means moving the predicate onto a property the control OWNS, not enlarging the vocabulary it borrows: the reviewer-isolation check now classifies a disallowed customization entry by the SHAPE of its identifier — bare (no namespace, no path) is host vocabulary this repo cannot adjudicate, so it reports as unverifiable and cascades, while namespaced, path-shaped, duplicated and unparseable entries stay proven customizations and stay terminal. The reclassification is safe because it moves only the NEXT ACTION, never the verdict: every state set including the new one still refuses at the acceptance gate, so no input that was refused becomes accepted and the softer class buys a reviewer-client switch rather than an isolation claim | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/parse_probe_result.py | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/parse_probe_result.py` (the class and its bare-identifier predicate), `code-review/scripts/claude_review.sh` (a routing arm ahead of the terminal arm, because the wrapper routes on reason TEXT and a distinct class needs a distinct phrase — reusing the schema-drift phrase would describe an unrecognised FIELD when what is unrecognised is an IDENTIFIER, and relying on the late `*"init"*` catch-all would make routing depend on case-arm order), and the two suites, with the plan in specs/015-reviewer-lane-host-vocabulary-class/plan.md. This discharges the resolution the prior stopgap row tracked; the `import` entry itself stays, now documented as a preference optimization rather than the capability's load-bearing predicate. Scope was set by measurement against the real CLI at version 2.1.220 rather than by the one name that broke: the review-skill invocation reports 46 slash commands and 16 skills, ALL of them host built-ins, so `skills` carried the identical defect and is in scope, while `plugins` (user-installed by definition) and `mcp_servers` (pinned empty by the strict MCP flags) are excluded — recorded as a set-diff against `REQUIRED_EMPTY_INIT_FIELDS` rather than as a taste call. Pre-change RED, measured not predicted: the new policy rows fail against the unmodified parser, and the matrix is clean after. Applied mutations are attributed differentially, each in a disposable copy with the real parser byte-compared unchanged after the walk — including one that drops the class from the main-invocation predicate while leaving the probe path correct, which is the divergence shape a single-path oracle cannot see. **Exact counts live only in the plan's executed-evidence section and are deliberately not repeated here**: a count kept in two places drifts from the candidate the moment a row is added, which is what a review round caught between two versions of an earlier row — and caught again on this row, whose first version recorded a case count that three later rows made stale. Verified against the real captured init: as-captured is still accepted, a synthetic new built-in command or skill cascades, a namespaced foreign entry stays terminal. The independent review then closed the one residual the plan had tried to accept, and the disposition is the transferable part: a STRUCTURED entry was being classified on its bare `name` while a sibling key could carry path-shaped proof of a real customization, and "no path reaches acceptance" was too weak a defence — this class exists for entries whose status CANNOT BE SHOWN, and there the evidence was present and merely unread. **The same class then came back in the other branch, which is the instructive part**: the first fix guarded only the disallowed path, so a structured entry whose `name` was an ALLOWED built-in still cleared the allowlist outright and reached ACCEPTED with isolation reported verified — reproduced first-hand, and strictly worse than the finding that prompted the first fix. Per the same-class-recurrence rule the answer was not a third patch but one shape gate placed BEFORE the allowlist, so both branches inherit it; the field that legitimately carries dicts is pinned TOLERATED so the gate cannot spread to it. Whether that restriction cost anything was measured before accepting it, not assumed: the real CLI emits both host-vocabulary fields as plain strings and uses dicts only for the already-excluded `plugins`. What stays residual is only the irreducible one — a hostile CLI can steer WHICH client serves the review, already reachable through the existing drift classes and never reaching acceptance |
130
+ | A self-audit oracle proves nothing about the branch it never invokes, and a green count is not coverage: the reviewer init oracle crossed 82 cases over both parse paths while passing no expected-native-skills, which is the ONE invocation shape whose customization lists are empty in every real run — so the branch where the outage actually happened had no row at all, and 164 green cases coexisted with a total review outage. Enumerate a gate's INVOCATION SHAPES the way its terminal reason codes are enumerated, and cross each shape over every parse path that implements it | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_init_policy_matrix.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/init_policy_matrix.py` (policy clause G, a review-skill base, and two new paths carrying the native-skill flags the wrapper really passes — one selecting each parse implementation) and `code-review/scripts/test_init_policy_matrix.sh`. Observed failure: the shape gap is why the prior round's outage passed a green matrix, and it was measured here rather than argued — the pre-change matrix has no row whose expected verdict depends on the native-skill branch. The fixture had to use a selected skill whose name is NOT also a built-in skill name, because the ambiguous-selected-owner guard otherwise fires on every row and masks the verdict under test; that was hit first-hand while building the RED table. **The un-landed half of the prior round is also discharged here**: that round left the oracle's own mutation-sensitivity walk as hand-maintained scores in a module docstring, with a header note saying to re-run it by hand when the policy changes — so every added row silently invalidated the recorded numbers and nothing failed when they went stale. This round changes the policy, which is exactly when that note fires, so the walk is now EXECUTED by the suite: each mutation is applied to a disposable copy, must be found exactly once (a mutation whose anchor moved is a hard failure, not a skip, because a silently-unapplied mutation disarms the check), must make the oracle report at least one mismatch, and the real parser's digest is compared before and after so the walk cannot mutate the tree it is auditing. The prose scores are replaced by the executed walk rather than updated, since a hand-maintained count in a second place drifts from the candidate as soon as a row is added |
131
+ | A control that rejects a LOSSY reading of its input must not perform one itself, and the check that decides trust must consume the value the caller actually sent: the whole-value gate added to stop a truncated identifier from being read as host vocabulary called `.strip()` before comparing, so a wrapped allowlisted name (`"import "`, `" import"`, a tab-prefixed variant) reduced to the same token, was declared whole, and the built-in allowlist ACCEPTED it with isolation reported verified. Surrounding whitespace is part of the value; normalization inside a trust decision is limited to what the identifier derivation itself performs, and anything else disqualifies the entry | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/parse_probe_result.py | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/parse_probe_result.py`, with rows and a re-introducing mutant in `code-review/scripts/init_policy_matrix.py` and `code-review/scripts/test_init_policy_matrix.sh`. Observed failure reproduced first-hand before the change across both host-vocabulary fields and three whitespace shapes; after it, wrapped names are terminal while a clean unknown name still cascades and the baseline plus legitimate namespaced entries stay accepted. **The transferable half is the recurrence shape, not the whitespace**: this was the FOURTH appearance of one class — evidence present in the entry but never read (a dict's sibling key, a dict under an allowed name, a whitespace-hidden suffix, and now surrounding whitespace) — and the first three were each fixed where they were found. Only the fourth forced the predicate onto the whole value, which is where the class actually lives. When a class recurs, the instance a reviewer names is the symptom; the shared premise is the finding, and a fix that inherits a helper built for DIAGNOSTICS will keep re-admitting it. The bypass also predated this slice, since the allowlist always consumed the stripped token |
132
+
133
+ | A shared-gate slice can land with one lane of a two-lane gate and nothing notices, because the only record of which lanes ran is the plan's own author-written rounds table: a mandatory fail-closed gate landed with nine review-mode rounds recorded, no challenge lane run, and a landing state still reading `plan drafted`; run afterwards against the landed diff, the challenge found a one-character bypass of that gate in one round. The closeout question is therefore about LANE NAMES, not round count — read the slice's own gate section, confirm each lane it names has a recorded outcome, and confirm a chain exists for THIS slice when lanes leave evidence in a local store. Neither a high round count nor clean deterministic gates substitutes for a lane that never ran | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/dual-track-review-gate.md#Count lane names, never rounds | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is itself unchanged this round — the entrypoint is severe size debt and its own anti-monotonic-growth gate blocked an appended bullet, which is the gate working, so the rule lands in `skill-extraction-workflow/references/dual-track-review-gate.md`, the reference the Step 6 closeout row already routes to; the gate defect it describes is fixed in specs/016-spec-citation-template-exemption/plan.md with the checker change in `scripts/check-spec-references.py`. Observed failure, established from evidence rather than inferred: no chain in the local review-evidence store corresponds to that slice and no packet in it touches the checker, while the slice's own plan required both lanes. RED baseline: the bypass reproduced first-hand against the landed checker — a dead spec citation carrying a bracketed fragment, a trailing unmatched bracket, a bracketed line locator, or a bracketed note all returned exit 0 from a gate whose entire job is to reject dead citations, while the same paths without a bracket returned exit 1. Applied mutations on the fix, each attributed differentially: reverting to the raw-token test flips all five bypass cases, dropping the "no bracket survives" clause flips exactly its owning case, and accepting any segment merely containing a bracket flips two. **Deliberately NOT landed this round: a mechanical check that a shared-gate plan records both lanes.** It is the right shape of control, and inventing it here would skip its own design-time operability check — author-dogfood against existing landed specs, marginal cost per routine slice, and what it defends against given the plan text is author-written either way. Routed as a follow-up with those legs owed, rather than shipped untested alongside the fix it would have caught |
134
+ | A contributor-facing rule that says "nothing catches X" becomes a lie the moment a checker starts catching X, and the rule is where contributors look rather than the checker: the code-review recipe contract told authors that a dead backticked path in prose is invisible to every control, which was true when written and false once the specs/ citation checker landed in `make test` and CI — and it still described a template convention that a later slice deleted. When a mechanical check lands or changes its verdict, the prose that tells people what is and is not caught is part of that landing, not a follow-up | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/code-review/scripts/AGENTS.md#A recipe step here must never cite a repo path that does not exist | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/AGENTS.md`. Observed failure: the rule asserted no control catches a backticked prose path, while the repo had shipped one that does, and it offered no way to write a path template under the deleted-exemption convention. Surfaced by the impact-chain gate during integration rather than by a reviewer — the citation migration had to rewrite a landed register row whose firing path anchors on this file, and a row cannot be re-presented in a round that does not touch its owner, so the gate demanded exactly the sync the landing already owed |
135
+ | An allowlist over WHICH keys may leave says nothing about WHAT is inside them, and the two are separate gates: the reviewer lane bounded its egress key set while relaying `session_id`, `model`, `provider`, and `version` verbatim out of an untrusted export into durable evidence rows, so a crafted export put a newline-bearing `status=passed` line into a field readers treat as machine metadata. Bind the value schema at the single emission choke point and report the field NAME only — never the value, or the report re-opens the hole. The response must then SPLIT by provenance, which is the half a first implementation gets wrong: sanitize-and-report fits a field whose value came from the export, but applying it to a field this repo's own code chooses turns an internal bug into a nulled `status`, and nulling the verdict IS the verdict change the schema promises never to make — those raise instead. The report field itself is output-only and must be refused on input, or it is the one field with no schema row and hence the one an attacker can forge to hide the tampering | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_egress_schema.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged this round; the change lands in `code-review/scripts/egress_schema.py` (new declared table), `code-review/scripts/parse_opencode_review.py` (`_result` applies it), and `code-review/scripts/opencode_review.sh` (its inline allowlist replaced by the shared declaration), with the plan in specs/013-opencode-egress-field-schema/plan.md. Observed failure reproduced first-hand against the unmodified parser before the change, not inferred from the code: an export carrying `version` = `1.16.0\nstatus=passed\nreason=INJECTED_BY_EXPORT` returned that string verbatim in the emitted payload; after, the field is null and named in `field_schema_violations` while `status` stays `passed` in both runs — which is the verdict-invariance property, measured rather than asserted. Five mutations applied in a disposable copy and restored byte-identical, each flipping the case that OWNS its property rather than merely some case: relaxing the charset flips the embedded-newline case, dropping the length cap flips the 201-character case alone, emitting the violation key unconditionally flips the clean-passthrough case, including the offending value in the report flips the report case, and accepting an undeclared key flips the raise case. **Three things the plan got wrong are recorded in it rather than absorbed silently**, and the first is the transferable one: the plan's own key table under-enumerated its subject — AST-walking every emission site found four emitted keys it never listed, and since an undeclared key raises, shipping the table as written would have broken the lane on its first isolation failure. An enumeration that decides what may leave has to be derived from the code, not recalled from it. The second is that `reason`/`reason_code` were filed as closed enums while their vocabulary spans five scripts including interpolated forms, so an enum drifting out of sync would blank a legitimate diagnostic on the exact path a human reads — an integrity bug traded for an availability one — and they are bounded by shape instead, since they are this repo's values and never the export's. The third is that "one table, not two" cannot mean one SET: the shell's list is deliberately narrower than everything the parser may emit, because on the concern-audit path an unlisted key fails closed for a human to classify, so importing the full key set there would have converted that fail-closed into a pass-through for five keys — a silent widening of the very boundary this slice tightens. It is now one module holding two named sets with an import-time subset assertion and a test that the narrow one stays strict and holds no model content. **The review lane returned two P1 findings and both were accepted, and the second is the transferable one**: this plan STATED the invariant "a violation never changes the verdict", the implementation nulled `status` in the same generic branch as every other field, and the test I wrote asserted the nulling — so the suite pinned the implementation's behaviour against the plan's own claim and every gate stayed green. A written invariant is not evidence; the case that fails when it is violated is, and a test written by reading the code cannot check the code against the spec. The first finding was the mirror image of the key-set lesson above: the undeclared-key check subtracted the set that INCLUDES the output-only report field, so the one field with no schema row was the one field never validated, and a supplied value could forge the record of its own tampering. Both fixes carry mutants that fail if undone. **A third defect was caught by a different gate and is worth its own note**: the review lane's egress scanner blocked codex, kimi AND opencode — with claude already excluded as same-family, the slice had NO available reviewer — because a fixture used a credential-SHAPED literal to prove that an offending value never reaches the report. The value was a documentation example and harmless; its shape was not. `--allow-fallback-egress` would have cleared it in one flag, which is the wrong move: the gate was working, and the property under test never needed that shape, so the fixture became a non-credential-shaped canary instead |
136
+ | A plan that names known work out of scope — a defect observed in passing, a debt or hardening item, follow-up work — owes each such item a disposition that outlives the plan: an owner locator or a terminal disposition with the reason, because a landed plan is an archive with no reader and naming is visibility, not tracking; scope decisions and non-goals are decisions, not work, and owe nothing | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#Naming known work as out of scope | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged this round — its step 3 already routes plan depth and fields to `skills/product-rd-workflow/references/delivery-lifecycle.md` §Plan Authoring, where the rule lands as a required slot keyed to the known-work predicate plus one slot in the assessment-plan field list ("the disposition of findings named but not selected"). RED is pre-change and re-computable rather than mutation-induced: the base revision of that reference carries zero out-of-scope rule text, while specs/010-review-concern-excerpt/plan.md names three such items and specs/014-spec-reference-existence-gate/plan.md records the outcome — "prose in a spec that then landed. Nothing carried it forward, and only one of its two copies was ever found"; all three items advanced only when later slices re-audited the same files, and each closure contradicted the archived wording (a 16 KB ceiling recorded as 47 KB, a drop-the-field fix recorded as cap-and-redact, a close-by-deletion recorded as a missing file to restore), which is the decay the disposition requirement prevents. Downstream instances unchanged and pointed-to, not restated: the closeout locator rule in `product-rd-workflow`'s implementation-completeness reference and the MR accepting-owner row in its code-review checklist — this row adds the FIRST firing point, where the naming happens. A mechanical checker over plans' out-of-scope lists was declined on the design-time operability legs (free-prose lists would take a structured-format tax on every slice; the list is author-written either way), recorded with the pre-cover axes and dual-track rounds in specs/018-plan-oos-item-disposition/plan.md. This row lands the half the 010 row above left "routed to `product-rd-workflow` (plan authoring), pending its own dual-track round" — that pending clause is superseded here, by pointer, not by editing that row |
137
+ | An out-of-scope disposition rule drafted with a non-goals exemption sentence and an already-owned-with-reason terminal form re-opens the failure it exists to close: the exemption clause is a relabeling hole — a deferred item recast as a "non-goal" satisfies the letter while evading the duty — and an unfindable ownership assertion is the rot with extra steps, so the positive known-work predicate must delimit alone, already-owned items must cite the existing owner's durable locator, and reason-only terminal disposition is reserved for declined | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#never an ownership assertion | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged in this round; the round applies review round 1 of specs/018-plan-oos-item-disposition/plan.md (codex lane, 3 P1 findings, all accepted) to the bullet the previous row landed: the exemption sentence deleted, already-owned folded into the locator branch of the same bullet, the recomputed RED recorded in that plan. The RED is a constructed bypass run against the unchanged baseline rather than an occurred incident: the pre-fix candidate text (revision 201cda9) carries "Scope decisions and non-goals are decisions, not work, and carry no owner duty", under which a deferred item relabeled as a non-goal passes; the post-fix text has no clause to relabel into, and the ownership-assertion path now dead-ends at "never an ownership assertion". This row supersedes, by pointer, the wording of the previous row's rule cell (its "scope decisions and non-goals … owe nothing" clause and its generic "terminal disposition with the reason") — that row records round 1 as landed and stays unedited, per the ledger's append-only contract, whose machine half rejected the first attempt to edit it in place |
138
+ | An owner locator that names only a generic destination — "the owning skill", a team, a gate — is an ownership assertion wearing a path: the destination holds no item-specific record, so the plan stays the only description of the work and the visibility-without-tracking failure returns through the enumerated form itself; a locator counts only when it resolves to an item-specific durable entry that names the work. And a firing point that paraphrases its sibling instances' operative requirements re-creates the drift it exists to prevent: a cross-reference names the instance, never restates its requirement | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#resolving to an item-specific durable entry | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged in this round; the round applies challenge round 2 of specs/018-plan-oos-item-disposition/plan.md (codex lane, 3 P1): the locator enumeration in the same bullet now requires an item-specific durable entry in every form, and the two downstream cross-references are reduced to pure pointers that name the instance only. The RED is a constructed bypass against the unchanged baseline: under the pre-fix text (revision e567986) a debt item dispositioned as "owned by the owning skill" satisfies the enumerated form while no artifact anywhere names the item, which is the exact failure the rule exists to close; the post-fix text dead-ends that path at "never an ownership assertion, and never a bare destination". The round's third finding — bind controller-captured checker output into the review packet — is declined with reason in that plan's rounds record: it re-proposes the digest-bound evidence apparatus this gate evaluated and removed, and the checker claims are recomputed mechanically on the exact candidate by the merge-time gate. This row supersedes, by pointer, the locator-enumeration wording of the two rows above; both stay unedited per the ledger's append-only contract |
139
+ | A successor-slice locator is the ownership-assertion class in one more costume: a plan that dispositions an item to a successor slice that does not yet exist has named a reader who may never arrive, so the successor form counts only when the slice exists and carries the item as in-scope work — the general shape is that every locator form must survive the question "if the named destination is deleted from the future, does any existing artifact still own this item" | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#an existing successor slice that carries it | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged in this round; the round applies challenge round 3 of specs/018-plan-oos-item-disposition/plan.md (codex lane, 4 P1: three accepted, one declined with reason recorded there): the successor-slice locator form narrowed to an existing slice carrying the item as in-scope work, the slice plan's own out-list corrected (a performed sweep had been mislabeled "closed this round", a disposition the rule does not permit — moved to in-scope completed work), and the Landing state rewritten with transcribed gate output. The RED is a constructed bypass against the unchanged baseline: under the pre-fix text (revision b34fc58) an item dispositioned to a successor slice that was never created satisfies the enumerated form while no existing artifact owns the item; the post-fix text requires the slice to exist and take the work in scope. The declined finding (entrypoint-routing evidence in the packet, or an entrypoint edit) is refuted in-repo: the entrypoint already routes to §Plan Authoring at L103/L106/L130, and single-rule entrypoint growth is the anti-monotonic-growth failure. This row supersedes, by pointer, the successor-slice wording in the rows above; all stay unedited per the append-only contract |
140
+ | An owner locator that resolves to an entry nobody accepted is a second archive, and a plan author's unilateral "declined, with the reason" is termination without authority — both let the item rot exactly as before, one hop away; so the durable entry must name the work AND its accepting owner, and a terminal decline must name the authority that owns the decision or its durable decision record, which is the same authority bar the closeout instance already sets for out/deferred coverage points | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#names the work and its accepting owner | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged in this round; the round applies challenge round 5 of specs/018-plan-oos-item-disposition/plan.md (codex lane, 1 P1, accepted): the locator head now requires the entry to name the work and its accepting owner, and the terminal form names the deciding authority or its durable decision record. The RED is a constructed bypass against the unchanged baseline: under the pre-fix text (revision 034c5a4) a defect parked in an ownerless status-document entry, or unilaterally declined by the plan author, satisfies the enumerated forms while nobody accepted or was authorized to terminate the work; the post-fix text dead-ends both at "never an entry nobody accepted" and "the plan author alone cannot terminate known work". Review budget note recorded honestly: this fix lands AFTER the fifth and final Agent round, so the post-fix candidate has had no fresh full challenge — the slice is interim pending a human decision (grant a fresh challenge round, or accept the candidate), per the exhausted-budget checkpoint rule. This row supersedes, by pointer, the disposition-form wording of the rows above; all stay unedited per the append-only contract |
141
+ | Naming is not accepting: an entry that names an accepting owner who never accepted, or a decline that names an authority who never decided, is the ownership-assertion class laundered through the rule's own required fields — so the durable entry must RECORD the accepting owner's acceptance and a terminal decline must CITE the decision's durable record, the same made-decision bar the closeout instance already sets ("explicitly made that decision; record the decision reference") | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/product-rd-workflow/references/delivery-lifecycle.md#records that owner's acceptance | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged in this round; the round applies the human-granted extension chain of specs/018-plan-oos-item-disposition/plan.md (h2: review 3 P1 + 1 P2 — two accepted, two declined with recorded reasons; challenge 1 P1 accepted in both halves). The RED is a constructed bypass against the unchanged baseline: under the pre-fix text (revision 05ae061) an out-of-scope defect parked in an entry that merely NAMES an owner, or declined with a bare authority name, satisfies every enumerated field while nobody accepted or decided anything — the plan half of the same finding demonstrated it live, in this very slice's plan, whose ratification-by-merge paragraph named a future merge instead of a made decision and was replaced by proposed-declines-pending-explicit-decision. The post-fix text dead-ends both at "records that owner's acceptance" and "a bare authority name is not a decision"; the wording keeps the previous row's anchor alive as a substring, so the ledger's earlier firing paths still resolve. This row supersedes, by pointer, the accepting-owner and decline-form wording of the rows above; all stay unedited per the append-only contract |
142
+ | A routing-surface fix for a stable eval-bank failure is measurement-bound end to end: a 10-round baseline separates stable failure from jitter before any edit, the after-numbers bind to the exact final wording (an intermediate wording's clean run is void the moment the wording changes), and the repo's own gates constrain the vocabulary — the leak-scan domain pattern vetoes any literal token carrying a `code.`-plus-alphanumeric substring inside skill files (the OpenCode project-config filename is such a token; name the tool, not the file), and a severe entrypoint's zero-net-growth budget means every description addition must be offset by moving detail into its canonical reference, not by deleting phrases a mechanical gate pins | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/eval-routing.md#中间稿的通过数在措辞再变的那一刻作废 | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is changed this round: the description gains the 本仓(ccl-skills 等共享技能仓)OpenCode 项目配置·命令治理 trigger, and the UI/UX judgment-dimension nine-item enumeration moves from the entrypoint into `references/uiux-judgment-extraction.md` as the size offset (zero-loss: the list survives verbatim there under "The judgment-dimension axis"; the entrypoint keeps the pointer and the adjacency-scan instruction). Observed failure reproduced first-hand before the change: eval case route-opencode-project-config at 2 PASS of 10 valid observations against the base description (misroutes scattered over terminal-cli-dev, product-rd-workflow, requirement-intent, worktree-isolation; grader-timeout rounds re-run at a longer timeout), consistent with the n=3 record that queued it in docs/skill-taxonomy-optimization-plan.md. Post-change on the final wording: the same case at 10 of 10 (confidence median 0.92), the 8-case neighbor set 3/3 per case before and after (including the sibling snapshot-audit case and the reviewer-lane case that must keep their owners), Tier-1 eval-routing zero findings on both sides. The RED is pre-change and re-computable — the bank runner against the base description reproduces it — and the firing path is that same runner, which now guards the regression. The full round record, both gate collisions, and the byte accounting live in docs/skill-taxonomy-optimization-plan.md under the 基线补测轮 section |
143
+ | A concentrated high-confidence misroute is the opposite fingerprint of the scattered low-confidence one, and its fix is a lexical-anchor question, not an ownership question: when the expected owner's body already claims the domain but its description advertises it in the wrong language for the utterance class, a single neighbor with a same-language trigger wall steals the whole class — the repair is to give the owner the missing-language anchor plus an explicit carve-out for the adjacent domain the anchor would otherwise annex, and every wording iteration voids and re-owes the full final-arm measurement, with each final-arm artifact machine-bound to the wording it was graded on | `platform-observability` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/platform-observability/SKILL.md#Own on-call SOP design end to end | updated | `platform-observability/SKILL.md` is the owner key and is changed this round: the description gains(值班/排班 SOP、P0/P1 打断;发布值班/回滚除外)after on-call routing, offset by trimming "internal " from the boilerplate tail (799/800 chars). Observed failure reproduced first-hand before the change: eval case mem-oncall-sop at 1 PASS of 10 valid observations against the base description, thief concentrated — product-rd-workflow 9 of 9 misroutes at confidence median 0.85 — the same-language trigger-wall fingerprint, consistent with the two-arm-unstable record that queued it in docs/skill-taxonomy-optimization-plan.md. Post-change on the final wording, machine-bound per artifact (22 of 22 binding sidecars valid, single description-line sha): the case at 10 of 10 (confidence median 0.95), the 8-case neighbor set clean except one inherent-jitter compound case disclosed in the plan doc, the challenger's mixed release-watch decoy 3 of 3 back to platform-release-engineering after the carve-out, Tier-1 eval-routing zero findings both sides. The RED is pre-change and re-computable — the bank runner against the base description reproduces it — and the firing path is the frozen bank case itself, which the runner guards. Dual-track chains routing-oncall-sop-r1/r2/r3 (codex both lanes) dispositioned in the plan doc; the Agent review budget is exhausted at five rounds, so the round lands interim pending the maintainer's decision on the post-r3 delta, per the exhausted-budget checkpoint rule. The full round record lives in docs/skill-taxonomy-optimization-plan.md under the mem-oncall-sop 独立一轮 section |
144
+
145
+ **Path-rewrite note (release-coordination rename, 2026-08-11).** The test-scope-prompt
146
+ handoff row above now cites `release-coordination`; it was written against that owner's
147
+ earlier slug, `prod-release-workflow/`. The path was retargeted mechanically so the evidence
148
+ check can still resolve it; the row's own content and order are unchanged. Same bend and
149
+ same reason as the tier-3 note above — the cited path no longer exists. The rename itself
150
+ was a naming decision, not a routing repair: `prod` collided with the `product-*` family
151
+ and inverted the sense (production, not product), and the coordinator now shares the
152
+ `release-*` family with the document owner while mechanism design stays under `platform-*`.
153
+ | The impact-chain evidence check reads each source-register row against the row's own round head, which scopes the renamed-away excuse to the rename round alone: a pre-rename round that substantively changed an owner still owes — and can validly write — a row under the old name, and the first-parent round partition is fixture-proven load-bearing rather than comment-asserted | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key; this round changes its scripts only — scripts/impact-chain-gate.rb (evidence existence checked per round head via the git blob at the row scope's head; subject selection follows the git-derived rename lineage across rounds — the excused source and every transient intermediate hop are demanded and recognized wherever their SKILL.md exists at a round head, with the row-recognition mapping and the RED-floor validation walk extended to match; deleted owners were never excused, fail-closed unchanged) and scripts/test_check_ccl_impact_chain_refscripts.sh (nine fixtures). Observed failures reproduced first-hand on the unmodified gate before the fix: the escape shape case-round-scope-renamed-away-pre-rename-round returned rc 0 where 1 is owed, and the legitimate pre-rename row shape case-round-scope-renamed-away-row-at-round-head returned rc 1 where 0 is owed — both recorded as differential RED and both green after the fix with the full 76-fixture harness passing. Both dual-track lanes (chain gate-round-head-presence-r1, codex) then independently converged on the transient-hop laundering variant — X to Y to Z with an undeclared substantive Y round, invisible to cumulative-endpoint selection — reproduced RED first-hand on the first-cut candidate by case-round-scope-transient-rename-hop-undeclared (rc 0 where 1 is owed) and closed by the ordered lineage propagation, with the declared counterpart case-round-scope-transient-rename-hop-declared pinning that an honest fully-declared chain passes. The second chain (gate-round-head-presence-r2) surfaced one more P1 per lane, both verified and fixed with their own differential REDs on the lineage candidate: a cross-round delete-then-recreate lookalike read as a cumulative rename and escaped the deletion fail-closed (case-round-scope-delete-then-recreate-lookalike, rc 0 where 1 is owed — the excuse now additionally requires some single round's own pairs to contain the source, so only an atomic in-round rename qualifies), and a fully-declared rename cycle A to B to A was falsely rejected because an absent upstream name was demanded unconditionally (case-round-scope-declared-rename-cycle, rc 1 where 0 is owed — the demand is now suppressed exactly when that round's own pairs rename the name to a selected successor, keeping deletions and renames to unselected slugs demanded and fail-closed, pinned by case-round-scope-undeclared-rename-cycle and the existing curated-escape fixture). The maintainer-granted extension chain (gate-round-head-presence-h1) then converged: review passed with zero findings, and the challenge found one more P1 this round itself had introduced — git show succeeds for a TREE, so the raw-blob existence read let a directory masquerading as SKILL.md count as a present entrypoint and a nested-file anchor vouch for an effectively deleted owner; fixed with a regular-blob mode predicate (ls-tree, 100644/100755 only) at all three existence sites and pinned by case-round-scope-skillmd-directory-masquerade, differential RED on the pre-extension candidate (rc 0 where 1 is owed) and the full 77-fixture harness green after. The first-parent discriminator case-round-scope-merged-row-before-work passes on the current gate and fails at its own named assertion on a mutant whose only change removes the first-parent flag from the round walk, proving the flag load-bearing. Fixture vocabulary reuses the harness's established neutral fixture domains (platform-observability to platform-signal-evidence); no source scenario domain involved. Round record: docs/skill-taxonomy-optimization-plan.md 残留表 two rows settled; plan artifact specs/019-impact-chain-round-head-presence/plan.md |
154
+ | A measurement round's own artifact must identify the surface it graded: the routing-bank runner embeds a routing_surface block (wrapper-compatible descriptions hash, per-skill description-line hashes, graded catalog text hash, bank hash) in every report so artifact-to-wording attribution is machine-checkable by independent recomputation instead of operator assertion, and the frozen bank carries its measured boundary probes as regression guards while a characterized unstable composite sentinel is annotated in the data rather than reworded to pass | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/eval-routing-bank.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the round changes its scripts only — scripts/eval-routing-bank.rb (routing_surface self-identification, grading semantics untouched) plus the new deterministic differential test scripts/test_eval_routing_bank_surface_binding.sh registered in the fast suite, with scripts/test_check_ccl_regressions.sh updated accordingly. Observed failure: both oncall-round lanes independently found (r1 challenge subitem b, r3 challenge) that per-round artifacts carried no machine binding to the graded description wording — attribution was operator assertion, exactly the class the h1b2 audit downgraded to operator-asserted. RED reproduced first-hand: the new test fails on the pre-change runner (report has no routing_surface object) and passes after. The frozen bank gains six verbatim promotions of already-measured probes (the oncall release-watch mixed decoy 3/3 post-carve-out; the opencode combined-implementation-intent challenger sentence 5/5; four h1b evidence-only variants 3/3 each), the p3-log-plus-test sentinel keeps its wording with a machine-readable stability annotation (rebalance rejected as Goodhart per the measurement discipline), and the evidence wrapper for this round cross-checks BOTH the embedded raw-surface hash and the embedded graded-catalog hash against its own independent recomputations (7-case stub regression incl. missing/wrong-embed fail-closed shapes; the catalog leg was added when both r1 dual-track lanes independently found the multiline-YAML-description hole — a continuation-line edit changes the graded catalog but not the raw first line — and the runner test gained the multiline differential fixture proving the catalog binding load-bearing). Bound measurement, fully regenerated under the final wrapper: promoted-cases bank 3 rounds 18/18; full 136-case bank single smoke round 133/136 with all three fails dispositioned against ledger history, none touching this round's changed surface, and the one no-history case (new-tighten-vs-owner, two same-direction flips across this round's independent smoke runs) recorded as a future-round baseline target candidate. The second chain's challenge added two accepted fixes: the wrapper now also requires the report's result ids to match the bank ids exactly with self-consistent totals (a silently case-skipping evaluator can no longer bind), and the superseded generations are retained in-tree for audit. The maintainer-granted extension then ran as two partitioned chains (the 276KB whole-candidate packet exceeds the 200KB cap): the decision-surface lanes found and fixed a runner TOCTOU (tasks and hashes now derive from one immutable bank snapshot), a destroy-on-rerun wrapper hole (pre-existing round artifacts are refused up front and the round file lands only after every validity leg passes), and a non-reproducible verbatim claim (test-promoted-case-provenance.sh is the deterministic in-tree check, latest run ALL VERBATIM); the data-artifact lanes converged on regenerating the authoritative rounds under the final wrapper so every sidecar carries all four binding legs including results_match_bank — final numbers 18/18 promoted-cases and 133/136 full-bank smoke, all three fails carrying prior ledger flip history, and the one no-history case (new-tighten-vs-owner, two same-direction flips across three independent smokes) recorded as a future-round baseline target candidate. Round record: docs/skill-taxonomy-optimization-plan.md bank/runner co-change 轮落地记录; plan artifact specs/020-bank-runner-round/plan.md; evidence eval/evidence/routing-bank-runner-round-2026-08-14/ |
155
+ | A stable unidirectional routing flip is an authorization gap on the routing surface, not grader noise: when a request class the taxonomy assigns to the lifecycle owner (substance correction of an existing spec/PRD re-enters the delivery lifecycle as change management) is advertised by NO description while a neighbor advertises the lexically nearest tokens, the cheap grader lands on that neighbor consistently — the fix is the measured dual signal (positive claim on the owner + negative carve-out on the flip-target neighbor), never editing the frozen bank case; the disposition bar (stable-failure vs intrinsic-instability vs middle-band-to-maintainer) and the post-fix success bars are preregistered in the round charter before the corresponding arm runs, and growth on a severe-debt entrypoint is offset in place by removing instruction-grade description prose whose obligation the body canonically owns | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#description | updated | `product-rd-workflow/SKILL.md` changes its frontmatter description only: the claim spec·PRD·需求文档实质内容写错要改对(substance 修正) is added and the instruction-grade tail parenthetical (re-entry guidance canonically owned by the body's Implementation entry / re-entry gate; zero lexical overlap with any of the 136 bank utterances — the only 4 grep hits sit in grader-invisible why_expected/source metadata) is removed for a net -26 byte delta on the 77KB severe-debt entrypoint. The companion carve-out lands in requirement-intent (description Skip + body boundary line); the always-on bootstrap agent-context/session-start.md is deliberately unchanged (curated digest, not a description mirror; no contradicting claim; always-on growth is size-gated; the description surface is the routing authority). Observed failure: bank case new-tighten-vs-owner (这个产品 spec 里写的流程是错的,帮我改对 → expected product-rd-workflow) flipped to requirement-intent twice in the 020 round's three independent smokes (n=3, no ledger history) and this round's bound 10-round baseline reproduced it as RED first-hand: 3/10 PASS on unmodified dev c0561c7, all 7 fails unidirectional to requirement-intent at conf 0.65-0.85, binding_valid 10/10, single HEAD, single surface hash. Post-fix, final wording v2, fully bound at 411367a: target 10/10 (every round conf 0.95 vs pre-change PASS median 0.85), 11-case neighbor set before/after 3 rounds each with all 10 stable cases 3/3 in both arms (the only non-3/3 is p3-spec-then-tc, a ledger-documented pre-existing flipper with zero lexical overlap with the changed tokens: 1/3 before, 2/3 after, same flip family), must_not tighten-doc 0 hits in all 20 target rounds, T1 analyzer blocking=0 advisory=0, frozen bank byte-identical to dev with both extracts proven verbatim by test-bank-case-provenance.sh (ALL VERBATIM); the superseded v1 wording's rounds are void per discipline rule 2 and retained in superseded-wording-v1/. External grounding: Anthropic skill-authoring best practices (description is the selection surface, what+when, 1024-char cap) and standard multi-intent routing practice (positive/negative triggers; compound utterances go to accept-set grading or orchestrator fallback, which is why the maintainer's p3-log-plus-test question is answered NO for a same-style fix and routed as an accept-set/compound-annotation proposal to the next bank/runner co-change round, converging with the accepted-alternatives schema need already recorded verbatim inside the frozen bank's miss-refactor-python-unqualified why_expected). Round record: docs/skill-taxonomy-optimization-plan.md new-tighten-vs-owner 独立一轮落地记录; evidence eval/evidence/routing-new-tighten-vs-owner-2026-08-15/ |
156
+ | A deterministic gate that evaluates a stdlib constant it never requires inherits its verdict from the host's transitive-load behavior — psych loads date lazily (<=5.1), never (5.2.0-5.2.5), or eagerly (5.2.6+) — and a broad rescue StandardError folds the resulting infra NameError into the data-error verdict, so the same commit at the same base judges green on one host and red on another with no diagnostic naming the real cause; the gate declares every constant it evaluates, and the suite pins the property with a dateless-host shim (block the transitive require, allow explicit ones) whose require-stripped mutant leg must fail for the guarded reason so a broken shim can never false-green | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/impact-chain-gate.rb | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the round changes its scripts only — scripts/impact-chain-gate.rb (+require "date" with a version-window rationale comment), scripts/test_check_ccl_impact_chain_refscripts.sh (heavy-lane case-routing-surface-dateless-host: dateless leg rc 0, require-stripped mutant leg rc 1 asserting impact_chain_firing_path_missing and the refused owner by name, dated-sibling dateless leg rc 0 proving safe_load really instantiates the constant; gate_runs 77→80 pinned), and scripts/test_impact_chain_gate_dateless_host.sh registered into the fast lane (merge-time containment on a minimal synthetic repo — the dual-track chains' converged same-class P1 was that all three deep legs sat in the uninvoked heavy lane, so a later require-date regression would land green; accepted and fixed this way per the existing behaviour-in-fast/wiring-in-heavy precedent). Observed failure first-hand: PR #5 CI (ubuntu-24.04 apt ruby 3.2 / psych 5.0.1) red at the aggregated promotion diff was reproduced byte-identical locally under ruby 3.2.11 at the PR test merge commit, the RED-first suite run failed exactly at the new dateless assertion on the unfixed gate, and the one-line require turned both the suite (ruby 4.0.5 and 3.2.11) and the PR-merge-form gate run (ruby 3.2.11, exit 0) green; upstream psych source matrix v5.0.0/v5.1.2/v5.2.0/v5.2.6 verified for the load-behavior window. The rescue-narrowing follow-up (fail-loud on infra errors) is a failure-semantics change deliberately routed out of this round. Round record specs/021-impact-chain-gate-date-require/plan.md |
157
+ | A test fixture that pins vocabulary its suite does not own — a README heading literal as a seeded anchor — breaks the moment the artifact evolves, and a suite lane with no CI execution surface gives that breakage zero feedback, so the two defects compound into a silently-rotting gate; derive the fixture's anchor from the artifact at runtime under fail-fast constraints instead of pinning a literal, give every registered-but-never-executed lane a merge-time execution surface, and correct an executed-count guard that drifted while the suite could not run to the machine-counted truth rather than trusting the stale literal | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged this round; the round changes its scripts only — scripts/test_register_firing_path_wiring.sh (the seeded anchor now lives in a note file the fixture writes and commits itself, so every byte of the anchor vocabulary is suite-owned; an interim runtime-derivation-from-README design was tried and then deleted by an explicit replace ruling after fence-parsing correctness produced same-class adversarial findings across two rounds — fenced decoy, mixed-marker fence, close-suffix recognition — each RED-proven against the pre-fix logic before its fix: parsing an artifact the suite does not own WAS the defect class, deletion beats hardening; a static-inventory-equals-guard-equals-executed count check remains, final guard 14 = 13 original cases + 1 count self-test) — plus .github/workflows/ci.yml gaining the parallel regression-heavy job running the umbrella --full so the heavy suites finally execute at merge time (failure propagation proven by a real measured negative control: the prior round's --full exited non-zero on the then-broken wiring suite). Observed failure first-hand: the suite was RED on dev (the pinned heading absent from README, grep=0) and identically RED at the pre-change baseline (the 021 round's differential proof), the interim fix re-REDed at the tail guard with expected-16-saw-13 (the guard's bite observed live; squashed history makes the stale literal's provenance unknowable), and the final suite passes 14/14 with the suite-owned anchor on the same repository whose README evolution broke the original pin. Round record specs/022-heavy-lane-wiring-and-ci/plan.md |
158
+ | Verification evidence for a change is selected by the surface the diff touches and stays honest about what ran: local pre-commit/push evidence is the narrowest set that would fail for the regression (behavior → focused tests; model/user-visible output → its snapshot; docs → doc gates; published/build paths → build + built-artifact smoke; real provider → real e2e), CI owns the exhaustive matrix only where verified blocking jobs cover the affected surface, only commands actually run are reported, and a green is reusable only for an unchanged recorded tree; evidence about an agent's own output is a read-only observation of authoritative durable state outside the subject's write boundary, never a probe over its self-report; a built/installed deliverable gets one smoke through its published entry path; a test double sits at the expensive/nondeterministic/unsafe/privileged/unavailable boundary with the rest real on isolated test-owned resources; an uncovered line is first a delete candidate | `testing-strategy` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/testing-strategy/SKILL.md#observes the world, not its self-report | updated | `testing-strategy/SKILL.md` (five core rules extended in place as one-line pointers, with the entrypoint held net-negative by condensing already-referenced rules); repaired fail-closed body-as-prompt differential, arms pinned before every call (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; runner, digests, status-checked raw answers under `specs/023-agent-native-repo-borrowing/evidence/`) on the F23-shaped task "agent reports it deployed and wrote a success log — design acceptance assertions": **base 0/4, head 0/4**. This provider did not reproduce the earlier Claude delta; the no-delta result replaces the stale score instead of preserving the stronger number. Detail lands in `testing-strategy/references/ci-fixtures-and-flake-control.md` (Local Evidence Selection), `testing-strategy/references/e2e-real-flow-testing.md` (Verify The World, Published Entry Path), `testing-strategy/references/test-code-authoring-patterns.md` (§5 doubles, §6 coverage); advisory judgment fixtures `eval/behavior-fixtures.jsonl` F22/F23. Source class: an external agent-native product repository's development discipline (testing policy + pre-push checks), read as an evolving portfolio → borrowed as one industry shape, not stated as standard; dual-track chains 023-I..023-I-r11 (spec 023) |
159
+ | Live-model evidence has an honest tier: keyless/replay proves plumbing, a with-key smoke of the assembled product (isolated tenant, least-privilege credentials, reversible-tool allowlist) supplies integration/routability/deployment confidence, correctness stays on deterministic evals; a skipped with-key test is not-run and cannot back an acceptance/release claim, which needs an executed credentialed job bound to the exact candidate and its provider/model/config/deployment/env generations; model-facing text is written from the model's perspective (implementation-only concepts out, public tool-contract discriminators in, reconciled with the diagnostics redaction rule) and its wording is behavior — snapshot as drift aid under the keep test, behavioral claim carried by replay/eval | `llm-inference-integration` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/SKILL.md#skipped with-key test is | updated | `llm-inference-integration/SKILL.md` (Core Workflow step 5, one bullet) and `llm-inference-integration/references/agent-instruction-composition.md` (new section + one clarifying clause in the diagnostics sentence). Repaired fail-closed, commit-pinned differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; runner, input digests and status-checked raw outputs under `specs/023-agent-native-repo-borrowing/evidence/`) measured the with-key-skip probe at **base 4/4 vs head 4/4** and the model-facing-wording review probe at **base 2/4 vs head 4/4**. The first is explicitly no-delta; only the wording probe observed a delta. Covered without change after full read: `llm-inference-integration/references/agent-tool-dispatch.md` (exposure vs authorization, cached manifests never bypass authorization) and `llm-inference-integration/references/agent-command-sandbox.md` (profiles, network as separate grant, explicit degrade, scrubbed env/closed FDs/temp dirs/symlink canonicalization). Same source class and evidence chains as the testing-strategy row above |
160
+ | Any user- or business-visible incident keeps the unconditional formal ceremony; for an escaped defect with NO such impact, the three-part threshold decides — subtle (non-obvious mechanism a careful engineer would re-derive the hard way), systemic (it passed through a gap in tests/tooling/conventions rather than a one-off typo), and costly to rediscover — and that threshold exempts only the write-up, never the simple-vs-complex analysis depth; the write-up opens with a thirty-second executive summary (what broke, root cause in plain terms, why it escaped, the durable lesson) and links the guardrails it produced, and it stays distinct from a decision record — backward-looking failure analysis, not the design decision it triggers | `defect-diagnosis` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/defect-diagnosis/SKILL.md#must get a written postmortem when all three hold | updated | `defect-diagnosis/SKILL.md` (the postmortem-ceremony sentence in Phase C). Repaired fail-closed body-as-prompt differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; runner, input digests and status-checked raw answers under `specs/023-agent-native-repo-borrowing/evidence/`) on "green unit tests but the product crashes on start — is this worth a formal postmortem, by what test, and what must it contain?": **base 1/4, head 4/4**. Source class: an external agent-native product repository's development discipline, read as an evolving portfolio → landed as one industry shape, not a standard |
161
+ | Prose written into a durable document whose vantage is the authoring session rather than the repository is a distinct deletion class: dead design-session citations, PR/stack vantage, change narration and version stamps, review choreography, reviewer-addressed justification, derivation transcripts, hedges, and working-language slips; the test is whether a reader at HEAD with no session, thread, or draft access can resolve every reference and verify every claim, the fix restates the surviving facts before deleting the transcript around them, and the over-correction traps (issue references, suppression rationales, counterfactual-present regression pins, measured provenance) are explicitly out of scope | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/tighten-doc/SKILL.md#会话视角泄漏:句子立足于写作会话而非文档 | updated | `tighten-doc/SKILL.md` DELETE #12 pointer plus the new `tighten-doc/references/session-vantage-leakage.md`; the cross-model co-review caveat moved to `tighten-doc/references/cross-model-co-review.md` so the entrypoint stays under the severe-debt line. Repaired fail-closed differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; raw answers committed) on a README paragraph carrying five of the eight classes measured **base 4/4, head 4/4**. This provider observed no delta, and the row records that instead of retaining the earlier weak improvement. Same source class and evidence chains as the rows above |
162
+ | A decision record's alternatives section is mandatory and recorded-not-invented (an alternative never actually considered is not backfilled to look complete; an old record whose alternatives cannot be reconstructed is marked not-recorded), retention is judged by future decision value rather than word count, age, or quota, deletion of an existing surface is preceded by classifying its consumers as production / non-production / ambiguous, enforcement is reviewed at the operation that performs the side effect rather than at the schema/prompt/facade that hides it, and model-visible wording is reviewed as behavior | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/references/code-review-checklist.md#Enforcement lives in the operation that performs the side effect | updated | `product-rd-workflow/SKILL.md` is the owner key and is unchanged this round; the round changes its references only — `product-rd-workflow/references/adr-convention.md` (alternatives-considered as a required field + retention by future decision value), `product-rd-workflow/references/implementation-completeness-and-minimality.md` (consumer-corpus classification before deletion, tagged TODO for tiny simplifications), `product-rd-workflow/references/code-review-checklist.md` (enforcement-at-the-executor merged into the existing Security bullet, session-vantage into the existing comments bullet, wording-is-behavior into Testing And Evidence). Repaired fail-closed, commit-pinned differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; status-checked raw answers committed) measured the enforcement probe **4/4 on both arms**, the two-scenario alternatives probe **base 0/4, head 4/4**, and consumer-corpus classification **base 1/4, head 1/4**. Only the alternatives probe observed a delta; both no-delta results remain explicit rather than being dressed up |
163
+ | A published branch rewrite is safe only when the exact fetched remote OID is the single object carried through ancestry validation and the explicit lease, and the rewrite invalidates review-thread, approval, mergeability, CI, commit-hash, and inline-comment evidence until each is refreshed | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/worktree-isolation/SKILL.md#已推送 / 挂着 MR / 别人可能在上面工作的分支 | updated | `worktree-isolation/SKILL.md` rewrites the published-branch rebase rule into one ordered fetch → literal `FETCH_HEAD` OID → ancestry check → rebase → explicit `--force-with-lease=<branch>:<oid>` sequence, followed by review-state revalidation. Repaired fail-closed differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; status-checked raw answers and runner under `specs/023-agent-native-repo-borrowing/evidence/`) measured **base 0/4, head 4/4** |
164
+ | Timeout diagnostics separate a portable summary from sensitive raw evidence. A stalled reviewer lane may keep bounded counts, types, and booleans needed to classify the stall in its ordinary payload, but raw event, log, prompt, diff, session-id, credential, or model text requires an explicitly requested private diagnostic directory with restrictive permissions and no extended ACL, a sensitivity marker, caller-owned retention, and a test proving credential bindings are excluded. A requested owner profile alone does not prove a native-skill stall: require positive structured stream-part evidence or keep the generic timeout classification. A yielded execution handle is progress, not empty output — resume it instead of starting a second reviewer, and when the handle is lost the lane is infrastructure-inconclusive with no replacement or fallback credited | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the change lands in `code-review/scripts/opencode_review.sh`, `code-review/scripts/test_opencode_review_retry.sh`, `code-review/scripts/test_review_client_compat.py`, `code-review/scripts/test_review_gate.sh`, `code-review/scripts/AGENTS.md`, and the new `code-review/references/timeout-auth-and-capabilities.md`, with the round record in `specs/023-opencode-timeout-evidence/plan.md`. Observed failure: a timed-out owner-aware OpenCode review lost its evidence and could be read as an empty reviewer result. RED baseline: with the new assertions and the unmodified wrapper the focused suite reported four failures — precise native-stream reason, bounded summary, explicit artifact retention, and parsing of the new option; after the fix `test_opencode_review_retry.sh` ends `opencode_review_retry_tests_ok`, including empty and invalid stream negatives and success-path cleanup. Two review findings were dispositioned in that round: any JSON object could overclassify a timeout as native-skill stream progress, fixed by requiring a session id plus a non-empty structured stream part type; and the credential-exclusion assertion checked only the diagnostic root, fixed with recursive credential-file and symlink absence assertions. Declaration provenance: this row was absent when the round landed and the impact-chain gate reported the owner as changed-but-undeclared; it was inserted into this round by a later corrective rewrite of the unpushed branch, which changed no code, reference, or spec bytes of the round |
165
+ | Human review feedback on merged changes is its own extraction source class with a diff-fact evidence bar — only human-authored items count, adoption is proven by comparing the feedback-time patch with the landed patch (merge status, a resolved thread, an author's "fixed" reply, or a same-file edit are context, never proof), an unreconstructable baseline fails closed to unclear, classification against the current skill makes overlapping windows idempotent, a singleton may qualify, "no candidate" is the common correct outcome, and model-drafted skill text is never committed verbatim; a rule written for a temporary condition carries its own retirement trigger; and benchmark source enumeration includes agent-native product repositories' discipline layers, not only published skill packs | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/skill-extraction-workflow/references/review-feedback-mining.md#Adoption is a diff fact, not a thread state | updated | new `skill-extraction-workflow/references/review-feedback-mining.md` (+ its Reference Loading pointer in `skill-extraction-workflow/SKILL.md`, which is the owner key for this row, offset by condensing sibling pointer parentheticals), `skill-extraction-workflow/references/rule-consolidation.md` (temporary-rule retirement trigger; keep load-bearing invariants inline when condensing toward a reference — the failure this round's own dual-track produced), `skill-extraction-workflow/references/coverage-exhaustion-traps.md` variant (b) (agent-native product repositories as a benchmark source class), `skill-extraction-workflow/references/harness-patterns-and-eval.md` §7 (plugin-harness shape recorded as one industry form, explicitly carrying no logging/persistence guidance — that axis is deferred to a security-owner design). Repaired fail-closed body-as-prompt differential (base `759dc603` vs head `fb41e11b`; provider `codex`, model `gpt-5.6-luna`, 4 rounds/arm; runner, digests and status-checked raw answers under `specs/023-agent-native-repo-borrowing/evidence/red-baseline-023-III.json`) on "I'll mine last month's merged PRs, treat resolved threads and 'fixed' replies as adoption, and commit the model-drafted SKILL.md" measured **base 4/4, head 4/4**. This provider observed no delta; the row replaces the earlier stronger score accordingly. Same-class reviewer recurrence recorded and decided rather than patched again: five rounds asked for the new judgment fixtures to become an executed blocking gate, which `eval/AGENTS.md` forbids by construction (advisory only, human-scored, never a merge gate). Decision `narrow`: the fixture corpus states that contract in-band (record `F0-contract`), each new fixture names its owning rule and, where one exists, the independent base/head probe that measured it; making this corpus blocking would require changing that contract first, which is out of this round's scope |
166
+ | A changed `SKILL.md` body has one deterministic growth budget without retroactively reddening unrelated historical debt: YAML frontmatter is excluded from the body metric, Unicode letter/number runs and individual Han characters are counted without a segmenter, invalid UTF-8 bytes remain visible as replacement units while raw-byte sizing stays intact, new or within-limit entrypoints stop above 5000 units, historical over-limit entrypoints may stay level or shrink but not grow, rename credit is path-paired and non-growing, and an unavailable comparison fails closed without an ok token | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/check-size-budget.sh | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the behavior lands in `skill-extraction-workflow/scripts/check-size-budget.sh`, with deterministic coverage in `skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh` and the decision table in `specs/023-agent-native-repo-borrowing/plan.md`. RED first-hand on the pre-change script: the focused suite stopped at the missing `changed_entrypoint_word_delta`, and a new 5001-word entrypoint was not blocked. Differential mutations made the frontmatter fixture count YAML and disabled historical-growth blocking; each mutant turned the focused suite red while the restored candidate passed. The final suite additionally pins no-space Han under the C locale, invalid UTF-8 without disabling raw-byte sizing, an unchanged historical over-limit sibling beside a small changed entrypoint, rename baseline transfer and rename-plus-growth, landed shrink resetting the allowance, unknown-base fail-closed behavior, and wrapper rc propagation. The post-merge canonical run originally caught this row itself missing; adding this declaration makes the impact-chain obligation explicit rather than relying on unrelated register edits in an accumulated base. |
167
+ | Entrypoint capacity is an INPUT to the target-output map rather than an obstacle discovered after drafting, so a landing form forced by a full entrypoint is recorded in the map row as a constraint instead of being reported in the rationale as owner judgment | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/rule-consolidation.md#Entrypoint capacity is an INPUT to the target-output map | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the rule lands in `skill-extraction-workflow/references/rule-consolidation.md`, already reachable from the entrypoint's Reference Loading list, so the round adds no entrypoint bytes — which is the capacity rule applied to its own landing. Baseline provenance is a **recorded incident**: this round's own correction RCA (per-host extraction scratch, round 025 target-output map, Step 1) records the failure as it occurred under the unchanged text — a routing-condition rule was drafted into a `severe_debt` router entrypoint, refused by `scripts/check-size-budget.sh`, relocated to a downstream owner, and reported as the better owner until the user asked whether that owner was actually right. A constructed absence check on the base blob confirms the unchanged text permitted it: `capacity`, `check-size-budget` and `severe_debt` occur nowhere in the pre-change `SKILL.md` Core Rules or `rule-consolidation.md`, while the same greps hit the post-change file, so the check can distinguish rather than only ever reading clean. With-change compliance is observed on this round's own map, which carries an entrypoint-capacity column, records the router as `not-applicable: capacity`, and states in the row that capacity forced the landing form — the shape the base text did not require. No base/head model differential was run for this clause; the evidence is the recorded incident plus the absence check, and is declared as such rather than scored. Anchor coverage is narrower than the rule: the locator binds the opening obligation, so gutting the bullet's capacity field or its forced-form disclosure while keeping the anchored sentence would still resolve — this row does not claim the gate protects the whole rule, consistent with the honest-but-fallible trust model in `references/external-practice-controls.md`. The cited incident record itself stays per-host, because the extraction-lifecycle rule keeps provenance out of the shared tree; a bounded reviewer therefore cannot open it, which is that trust model working as designed rather than a missing artifact. `observed-failure: yes` accordingly rests on author honesty plus the independent review and challenge lanes, not on a packet-checkable object — the same basis every prose row in this ledger has. |
168
+ | The self-detect firing point has an output-shape sibling: a deliverable that has become a side-by-side comparison of named candidates, or an adopt/reject/replace verdict on a named candidate, re-opens the entry-owner question even when the entry phrasing was a narrow how-to — where a candidate is a component, service, product, library, vendor, or replacement actually up for selection and never a configuration value or technique — and its trigger shapes are walked one at a time rather than held as a conjunction, keyed to the deliverable's shape rather than to recommendation wording | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/firing-point-placement.md#re-judge the entry owner before writing further | updated | `skill-extraction-workflow/SKILL.md` is the owner key and is unchanged in this round; the rule lands in `skill-extraction-workflow/references/firing-point-placement.md`, already reachable from the entrypoint's Reference Loading list. Recorded incident: an entry judged as a configuration how-to produced several rounds of candidate comparison and a recommendation while the current-state evidence discipline and the research-routing rule both stayed dormant, surfacing only when the user challenged what the conclusion rested on. A local base/head observation was run (provider `codex`, `codex-cli 0.147.0`, local default model, 4 rounds per arm, arms differing only in whether this section is present; scenario a neutral background-job-runner retry question whose draft has grown into a three-candidate comparison recommending one on commit activity; prompts, both arm texts and the eight raw answers kept in the per-host round-025 probe directory). Oracle: does the stated next action re-open which owner governs the deliverable before continuing? **base 0/4, head 4/4.** Four declared limits, none of them argued away: (a) this is NOT a frozen measurement under `specs/*/evidence/AGENTS.md` — the head arm was read from a dirty worktree rather than an immutable commit, no arm SHA binding or per-prompt sha256 was recorded, so it is advisory local evidence and must not be cited as a round-023-class RED probe; (b) the head arm's text contains the phrasing the oracle grades, so echo cannot be excluded and the honest claim is a behavioral delta, not an independent derivation; (c) the base arm's behavior — discard the comparison, answer the narrow question — is itself a SAFE recovery, so the delta shows the rule changes HOW the divergence resolves (re-own the selection work vs. drop it), not that the base is unsafe; (d) neither arm's raw output is in any review packet, so a bounded reviewer cannot check arm equivalence or grading. The load-bearing evidence for this row is therefore the recorded incident; the observation supports it and is reported as measured without being promoted past what it can carry. Anchor coverage is likewise narrower than the rule: the locator binds the final action bullet, so deleting the trigger shapes above it would still resolve, and this row does not claim the gate protects them. A fourth wording-keyed trigger was removed rather than narrowed a third time, after it over-fired on ordinary how-to advice in two consecutive challenge rounds and was found to add no coverage over the adopt-verdict shape — recorded here as a `delete` decision under the same-class-recurrence rule, not as a silent edit; the surviving triggers additionally scope "candidate" to a component, service, product, library, vendor, or replacement under selection, closing the same over-fire semantically rather than only removing the phrase. As with the row above, this row's incident record stays per-host under the extraction-lifecycle rule and is not openable from a bounded packet by design. **Open reach limitation, decided by the risk owner rather than patched:** the detector lives in a reference reachable only once this workflow is already loaded, while the incident it addresses is precisely a session where no owner was ever selected — so it narrows the gap for an agent already inside the workflow and does not close it for one that never enters. Closing it would mean spending always-on injection bytes, which `references/dual-track-review-gate.md` forbids before an `eval-routing-bank --with-bootstrap` A/B shows a real delta; that measurement was not run this round, so the gap is recorded here rather than papered over, and the section keeps its explicit recognition-dependent framing. |
169
+ | A citation frozen into an append-only ledger line has no legal repair, because the ordinary fix — respell the token — is the one edit the ledger forbids; a repo-wide fail-closed citation gate therefore needs a row-bound waiver pinned to that line, never a grammar that exempts a SHAPE. Bind it by identity, not by resemblance, and note that CONTENT IS NOT IDENTITY: file, exact raw token, LINE NUMBER, and the SHA-256 of the line AS STORED — terminator included — must all match, so the same token one line, one byte, or one file away is a finding like any other, a verbatim COPY of the pinned row is a row nobody reviewed rather than a second instance of the waived one, and an LF-to-CRLF rewrite or a dropped final newline changes the stored bytes and must break the pin even though `splitlines()` would report the same text; key on the raw token rather than the resolved target so a locator or fragment variant inherits nothing; waive non-existence only, leaving containment escapes unwaivable; print every applied waiver; and fail the gate when a pin matches no line in a scanned waived file, since on an append-only ledger that means the pinned row was edited, reflowed, or deleted | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/check-spec-references.py | updated | No skill owner package changes in this round, so no owner key is cited: the repair is repo-root gate scope, and the ledger row that triggered it is unchanged and stays byte-stable, which is the point — the repair lands entirely in `scripts/check-spec-references.py`, its focused suite `scripts/test_check_spec_references.py`, and the plan in `specs/025-spec-reference-correction/plan.md`. Observed failure, pre-change and re-computable rather than mutation-induced: at base `2146df2` the mandatory gate exits 1 on register row 152, which writes the per-spec evidence contract as a `*/evidence/AGENTS.md` glob under specs/ while meaning the contract class, not a file; `make test` runs that gate, so the whole suite was blocked on a line nobody was permitted to fix. This is the third instance of a mechanism this repository already runs against this same ledger file, and it adopts the recorded reasoning of the first two verbatim: digests over occurrence counts, because a count says "one row" but never WHICH row, so it transfers silently when rows move. It is NOT the alternative round 016 rejected — that one allowlisted template STRINGS, predicating the gate on a vocabulary with reach over every file and every line, whereas this reach is one line and there is no grammar to evade. `CITATION_RE`, `resolve_target`, and the containment model are untouched. Thirteen mutations were applied, each attributed differentially with the unmutated suite green: dropping the digest pin lets an appended row and a one-byte reflow ride along; removing the file binding — which is enforced twice, in the key and by digesting only waived paths, so a mutation of the property has to remove both — lets an identical line in a second file ride along; waiving the whole pinned line instead of the one token lets an unrelated dead citation sharing that line ride along; keying on the resolved target lets a locator variant inherit the waiver; making a stale pin non-fatal removes the tightening that shipped alongside the loosening; dropping the not-escaped condition makes containment waivable; dropping the corpus-scope guard on staleness reds every unrelated repository the checker is pointed at, including this suite's own fixtures; comparing pins by digest alone re-waives a verbatim duplicate and a moved row; normalizing the terminator before digesting re-waives a CRLF rewrite and a missing final newline; matching the waiver token against raw text rather than parsed citations re-marks a covering-nothing entry as live, including when the token is only a substring of a longer citation or sits unbackticked in prose; removing the spend key lets one waiver cover the same dead token twice on its pinned line; forcing the waived-path presence guard off restores the bypass by omission, where deleting or untracking the waived file voided the guard instead of satisfying it; and making that guard unconditional breaks the foreign-corpus and explicit-table compatibility positives — the two are a deliberate PAIR, because the guard has two conditions that must fail in opposite directions and one mutation could only prove half of it. The focused suite goes 43 to 69 cases; 26 are added and 25 of those are RED at the frozen base, the exception being one compatibility positive that passes at base by construction because the base has no waiver table at all. The loosening cleared the design-time operability check in all four legs, recorded in the plan. Recorded and NOT closed here: nothing runs the repo gate between authoring a register row and appending it, which is the detection gap that let the defect land — a pre-append hook decision with its own blast radius, routed to its own round rather than smuggled into a correction slice Independent review defeated the first version on exactly the two halves of that identity and both fixes carry their own applied mutation: comparing pins by digest alone re-waives a verbatim duplicate and a moved row, and normalizing the terminator before digesting re-waives a CRLF rewrite and a missing final newline. The terminator mutant is deliberately the NORMALIZING one, because hashing the stripped line would also break the positive case and so would prove the digest is used without isolating terminator sensitivity. A waiver is additionally SPENT BY ONE CITATION and is only considered live when the token it names is an actual parsed citation on its pinned line: without the first, the predicate runs per regex match and one reviewed waiver covers a second unresolved citation repeated on that line; without the second, a mistyped or obsolete entry sits there covering nothing while its pin still matches, which is the covering-nothing hole the staleness check exists to close. Decide that liveness on PARSED citations, never on raw text — a token that is merely a substring of a longer citation, or that appears unbackticked in prose, is not a subject — by parsing the line once and feeding the same token set to both the liveness check and the scan, so the two cannot disagree. |
170
+ | A capability predicate that reads an external CLI's help TEXT — not just its flag names — silently outdates every hand-written fake of that CLI, and the failure is invisible: under `set -euo pipefail` the suite dies on the wrapper's command substitution before any assertion prints, so the operator sees a bare non-zero exit with empty stdout AND empty stderr. When one commit adds such a predicate and sweeps its sibling fixtures, the fixture it misses fails closed forever; and a fixture stuck at an EARLY gate makes every later assertion in its file pass for the wrong reason, so repairing it is what makes those assertions discriminate at all. Read the wrapper's own returned reason before theorising about the local toolchain, and confirm which side drifted from the history of the predicate rather than from the shape of the failure | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_parse_review_json.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the repair lands in `code-review/scripts/test_parse_review_json.sh`, with the diagnosis recorded in `specs/025-spec-reference-correction/plan.md`. Observed failure, pre-change and deterministic rather than mutation-induced: the suite exits 2 with empty stdout and stderr, reproducing identically at base `2146df2` in a clean archive copy, so `make test` has been RED on the integration branch since the predicate landed. Proven cause: `claude_review.sh` gates on a `--safe-mode` DESCRIPTION proving that safe mode disables inherited skills, while all three fakes in this file emitted their flags as one flat space-separated line carrying no descriptions; the sibling name-only check is field-based and passed, which is why only this gate failed. Two hypotheses were rejected on evidence first — that the machine's real CLI leaked in, and that the probe invocation's arguments had drifted — both killed by an argv-logging shim showing the only invocation was the help call served by the fixture on PATH. History fixes which side drifted: the single commit that introduced the predicate touched ten sibling test files, and exactly two files in that package carry a fake help surface at all: it updated the one and not the other, whose previous touch was the repository's init commit, so the wrapper's predicate is left untouched and the fixture is repaired. The repaired fixture is itself the regression control, because the description is what the predicate reads. Applied mutation: collapsing the description back to a bare flag name in one fake restores the original symptom — recorded with the first attempt's invalidity, since running a COPY of the suite from a temporary directory also exits 2 merely because it cannot resolve its sibling parser, which is a non-zero exit for the wrong reason and is not evidence. Exposed and fixed in passing: the tool-enabled case asserted an inconclusive exit while never reaching the tool-boundary check it exists to exercise, and now fails on `tool_boundary_violation` with the declared and invoked tool sets named. Recorded and NOT closed here: nothing mechanically couples the wrapper's help contract to the fakes that must satisfy it, routed to the test-layer owner; and the control that should have caught this DID fire — the suite was red — so the open question is why a red suite persisted on the integration branch, which is a maintainer decision and not the agent's to answer |
171
+ | An assertion over a value derived from ELAPSED WALL-CLOCK TIME is bounded, never pinned: the exact literal reads as the more precise choice and is the flaky one, because the quantity is integer seconds and one second of scheduler jitter flips it. Bound it on both sides so each real failure mode stays excluded — an upper bound proving the reserve was subtracted, a lower bound proving a CAP was not charged in place of the elapsed time — and keep any genuinely deterministic cap in the same case an exact match. Diagnose it by making the intermittent case deterministic rather than by re-running it, and do not accept the loudest nearby stderr line as the mechanism | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the change lands in one case of `code-review/scripts/test_opencode_review_retry.sh`, with the diagnosis in `specs/025-spec-reference-correction/plan.md`. Observed failure: one case failed under concurrent load and passed idle. Proven cause: the wrapper computes the formal budget as the remaining lane time minus an export reserve clamped to at most ten seconds, which at a 180-second lane is 170 minus whole-second elapsed, while the case grepped for the literal 170 — true only while the boundary probe and its surroundings finish inside one integer second. Made deterministic instead of re-run: forcing one second of stub boundary delay fails exactly and only that case, which is the same single failure as the intermittent one. A tempting nearby cause was rejected on evidence — the run's stderr carries a bad-file-descriptor write error from the wrapper's final stdout write, and the identical message appears on runs that exit 0, so it is an unrelated exercised path. The fix keeps the deterministic 60-second probe cap an exact match and bounds only the budget, reusing the extraction the sibling case already uses; no assertion is removed, nothing is skipped, and the wrapper is untouched. Applied mutation, evaluating the shipped predicate's own extracted bytes against synthetic recorded budgets: the new predicate passes at 170, 169, 165 and 160 and fails at 159, 155, 120, 110 and 180, so both bounds are load-bearing, while the old predicate passes only at 170 and fails at 169, which demonstrates the one-second cliff rather than arguing it. A whole-suite fifteen-second delay probe was started and abandoned at the risk owner's direction as needlessly expensive for the same conclusion. Prevention is the pattern rather than this line and is routed to the test-layer owner, not landed here |
172
+ | SUPERSEDES the immediately preceding timing-assertion row: a timing-derived value must not be repaired by replacing one guessed number with a guessed interval. Control time or observe the producer's actual operands and assert the semantic relation. A range is valid only when every value inside it satisfies the contract and applied producer mutations cannot land inside it; testing only the range's edges proves that the predicate executes, not that the producer is correct | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The preceding row is append-only and remains as the auditable first disposition, but its claimed `160..170` discrimination is false. An exact-candidate reviewer supplied a producer mutation that subtracts the ten-second export reserve twice; it emits 160 and the old range passes. The repaired case enables Bash xtrace only for its synthetic wrapper invocation, captures that invocation's actual remaining lane budget and final clamped reserve, and requires the timeout passed to `opencode run` to equal their difference. Scheduler delay is therefore represented in the observed remaining operand instead of tolerated by a broad interval, and the production wrapper remains untouched. The unmutated control records remaining 180, reserve 10 and formal 170. The applied double-reserve mutant reports exactly one failure, the named budget case, while the complete unmutated verifier suite exits 0 with `opencode_review_retry_tests_ok`. This prevents a verifier false green without adding production surface or a slow delay loop, improving both delivery correctness and feedback efficiency |
173
+ | SUPERSEDES the immediately preceding semantic-relation row: an oracle does not become independent merely because it compares two runtime values. If every operand comes from the producer under test, one producer mutation can move them together and keep the relation self-consistent. Anchor at least one relation to a controller-owned input or independently controlled observation, assert each intermediate transformation in order, and apply a producer mutation at the earliest transformation; a downstream-only mutation proves less because an upstream false green can remain hidden | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h3 challenge against candidate `04b8b47` found that the timeout verifier read both the remaining lane budget and export reserve from wrapper xtrace. An applied producer mutant subtracted the reserve inside `remaining_lane_timeout`; the old test still passed with a self-consistent remaining/reserve/final triple. The repaired test anchors the remaining budget to the externally supplied test timeout minus the wrapper's observed elapsed value, then separately checks final timeout equals remaining minus reserve. On the same mutant exactly the named budget case turns RED, while the unmutated complete verifier suite exits 0 with `opencode_review_retry_tests_ok`. The production wrapper is unchanged. `testing-strategy` already owns the general same-source-expectation and killing-mutation rules; this package test is the mechanical firing point |
174
+ | A shipped safety-table entry keeps its shipped semantics when the caller supplies that table explicitly: optional-argument omission is invocation shape, not entry identity. Under the owning repository, classify each applied entry by exact shipped key and value; the shipped object and an equal copy must retain the same presence guard, while non-shipped custom entries and foreign corpora remain inert. Prove both directions so tightening does not turn a root-wide obligation into a partial-corpus outage | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/check-spec-references.py | updated | The h3 challenge found that passing `LEDGER_CITATION_WAIVERS` explicitly applied the waiver but skipped the waived-path presence guard because the checker keyed safety on `waivers is None`. A focused RED test removes the shipped waived file and calls the helper with both the shipped object and an equal copy; the pre-fix checker records no stale pin. `scripts/check-spec-references.py` now derives required presence keys from exact shipped entry content only under its owning repository, and `scripts/test_check_spec_references.py` proves both forms go stale. Removing that classification makes exactly the new assertion fail; the unmutated focused suite and repository checker pass. Explicit non-shipped tables, explicit empty tables, and foreign corpora keep their compatibility behavior |
175
+ | SUPERSEDES the earlier independent-oracle row for the timeout verifier: operands selected from one trace must also belong to the same producer INVOCATION. Independent occurrence scans can pair the first value from an earlier boundary call with the first downstream assignment from a later call and reintroduce scheduler flakiness. Pair by trace order or an explicit call identity, and make the multi-call distinction fire in the control; a relation that passes only while both calls happen in the same clock tick has not removed the timing oracle | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h4b review found that the verifier selected the first `remaining_lane_timeout` elapsed value and the first `run_opencode` timeout assignment even though the boundary probe calls the budget function earlier. A disposable exact-test copy added a one-second delay only to that boundary invocation; exactly the named budget case turned RED on the correct production wrapper. The permanent control now includes that delay, asserts the run-paired elapsed value exceeds the first boundary elapsed value, and selects the latest elapsed observation preceding the first run-timeout assignment before asserting timeout minus elapsed and remaining minus reserve. The unmutated complete verifier suite exits 0, while the previously applied producer mutant still fails exactly the named case. Production `opencode_review.sh` remains unchanged; the added one-second control replaces probabilistic same-tick success with a deterministic multi-call distinction |
176
+ | SUPERSEDES the immediately preceding trace-order row: proximity in one producer trace is not call identity, and a producer-owned elapsed value is not an independent oracle even when paired correctly. Observe stable external entrypoints instead. Here the deterministic stub records monotonic time at the boundary and formal-run entrypoints, and the verifier compares the timeout's charged elapsed seconds with that independently measured gap, allowing only integer-second quantization. No internal call count, trace order, or producer elapsed value participates | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h6 challenge applied both counterexamples: an extra budget consultation can break positional pairing on a correct wrapper, while inflating the producer's elapsed value moves all traced relations together and keeps the old oracle green. The trace is removed from the test. The complete production-wrapper control exits 0. In a disposable copy, a mutant adds ten elapsed seconds only for this case's unique `TIMEOUT=180`; exactly the named budget assertion fails and every other case remains green. An earlier unconditional mutant caused 24 failures and is explicitly excluded as non-differential evidence. Production `opencode_review.sh` remains unchanged |
177
+ | SUPERSEDES the shipped-entry content row above: safety identity follows the shipped KEY, not mutable explanatory metadata. Once a caller presents a waiver for a shipped key under the owning repository, changing its reason text or pin tuple must not downgrade the path-presence obligation to custom-table semantics; validation of the supplied value remains a separate concern | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/check-spec-references.py | updated | The h6 challenge changed only the reason text in a copied shipped entry, removed its waived file, and reproduced findings=0, waived=0, stale=0 because application used the key while presence classification required exact key-and-value equality. The focused near-copy test was RED before the checker change: 70 cases ran and only the new assertion failed. Required presence keys now use membership in `LEDGER_CITATION_WAIVERS`; the shipped object, equal copy, and changed-reason near-copy all go stale on absence, while custom keys and foreign corpora retain their existing inertness |
178
+ | SUPERSEDES the external-entrypoint timing row above: one observed interval cannot supply both bounds when the budget starts before that interval's first endpoint. Bracket the producer with two controller-owned intervals instead — the nearest stable entrypoints give the lower bound for work that MUST be charged, while a controller timestamp before launch gives the upper bound for everything that MAY be charged. Integer-second quantization belongs outside those intervals as a narrow tolerance, never as hidden setup time | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h7 review found that lane accounting begins immediately before the boundary timeout is prepared, while the stub's boundary timestamp is recorded only after that preparation and scheduling. A correct wrapper can therefore charge more than boundary-entry to run-entry by over one second under load. The verifier now requires charged elapsed to be at least boundary-entry to run-entry minus one second and at most controller-before-wrapper-launch to run-entry plus one second. The production control exits 0; the `TIMEOUT=180` elapsed-plus-ten mutant still fails exactly the named case; no producer trace value or internal call position participates |
179
+ | SUPERSEDES only row 153's numeric suite total for the accumulated 025 slice, not its per-round evidence: that row remains correct at its introducing commit `04b8b47`, where a clean checkout runs 69 cases and exits zero. Later correction rounds add two more focused cases, so the cumulative progression is 43 to 71 cases, 28 added. Twenty-five historical RED additions plus the changed-reason near-copy and tracked-symlink diagnostic cases make 27 observed RED additions across their respective pre-fix baselines; the compatibility positive remains non-RED by construction | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/test_check_spec_references.py | updated | The h8 review compared a historical row with the later accumulated suite. Direct evidence separates them: local clean commit `04b8b47` runs 69 tests and exits 0; the changed-reason case then makes 70 and is RED before its key-membership fix; the tracked-symlink case makes 71 and is RED before scan-state classification. This appended row preserves immutable round history while giving a later reader the cumulative total explicitly |
180
+ | A path being yielded by `git ls-files` and a path being successfully SCANNED are different states. A fail-closed gate may reject both a deleted path and a tracked-but-unscannable path, but it must not diagnose both as deletion: operators otherwise restore or re-add a file that is already tracked while the real symlink, file-type, or read failure remains. Track encounter and scan completion separately, keep both non-passing, and give the difference its own stable reason suffix | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/check-spec-references.py | updated | The h8b review found the waived-path presence guard populated `scanned_waived_paths` only after a successful read, so a tracked symlink and an absent path both emitted `absent-from-tracked-corpus`. The focused test replaces the shipped waived ledger with a tracked symlink: before the fix, 71 cases run and only this case fails because `present-but-unscanned` is absent. The checker now records the waived path immediately when `git ls-files` yields it, records scan completion separately, and emits `present-but-unscanned` for their difference. Deletion and untracking still emit `absent-from-tracked-corpus`; both states remain stale and exit nonzero. NUL-bearing bytes are not part of this failure because surrogateescape scans them losslessly |
181
+ | SUPERSEDES the two-interval timing row above: the upper endpoint must begin where the producer's budget begins, not merely somewhere before it. A controller-before-process timestamp includes uncharged setup and turns that slack into a false-green allowance. Instrument the exact budget-start COMMAND from outside the producer, then bracket the first downstream action between that lane-start interval and the nearest required-work interval. Validate every observation before arithmetic, append repeated event timestamps, and assert the expected cardinality so missing or extra calls become named failures instead of shell exits or positional mispairing | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h9 challenge found three coupled gaps. The permanent fixture now uses a test-only `BASH_ENV` DEBUG hook to record monotonic time immediately before `LANE_BUDGET_STARTED`, while boundary/run stubs remain external observations. Charged elapsed must lie between lane-start-to-first-run with one-second integer quantization and the boundary-delay lower bound. Run timestamps append, the first is selected, and exactly one run is required. Missing/non-numeric observations set a guard false before arithmetic. The production control exits 0; a `TIMEOUT=180` plus-two-second elapsed mutant and a missing-run-timestamp mutant each make exactly the named case fail, with no silent shell error. Production `opencode_review.sh` remains unchanged |
182
+ | SUPERSEDES the immediately preceding lane-clock row: when the producer uses an integer shell clock, observe that SAME clock at the exact producer commands instead of comparing it with a second clock and guessing a quantization interval. DEBUG traps must be inherited into functions, every matching observation stream must append, and each expected command/event needs a cardinality assertion before arithmetic. Exact integer-clock equality then distinguishes a one-second producer mutation without timing slack | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h11 challenge found that the monotonic interval still allowed a one-second inflation and that first/last overwrites could mispair calls. The test-only `BASH_ENV` hook enables `set -T`, appends the wrapper shell's `SECONDS` at the sole lane assignment and at the elapsed command whose caller is `run_opencode`, and requires one numeric lane, elapsed, boundary, run, and formal-timeout observation before computing. Charged elapsed must equal the exact shell-clock difference. Production control exits 0; scoped plus-one-second, missing-run-clock, and duplicate-boundary mutants each make exactly the named case fail. The wrapper remains unchanged |
183
+ | Synthetic repository fixtures must copy the VCS CORPUS, not the ambient directory tree. A recursive filesystem copy silently imports untracked and ignored scratch files, so a control can fail on local debris while mutants keep their expected nonzero exits and appear attributed. Enumerate tracked paths from the source repository, copy their current worktree bytes, preserve tracked symlinks, and then initialize the fixture repository | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h11 challenge found `clone_candidate` used `shutil.copytree(source_root / "specs")`, which includes untracked and ignored files. The fixture now consumes `git ls-files -z -- specs`, copies only non-empty tracked names, preserves symlinks, and therefore still exercises current tracked candidate bytes without importing scratch. The 71-case current control exits zero. This is a hermeticity hardening discovered by challenge; no real scratch file was added to the candidate worktree, so its evidence is mutation/reasoning rather than a pre-existing repository failure |
184
+ | SUPERSEDES the exact-pre-sample shell-clock row above: even two observations from the producer's own integer clock can straddle a second boundary between the external DEBUG sample and the command's actual clock read. Bracket each producer read with a pre sample and the next DEBUG event's post sample, require exactly one numeric observation in every stream, and compare the charged result with the interval implied by those brackets. This is measured uncertainty at the read boundary, not a guessed scheduler tolerance | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h12 review identified a false RED when either pre sample crosses a `SECONDS` boundary before the producer command reads the clock. The fixture's inherited DEBUG state machine now appends lane-pre and elapsed-pre at the two uniquely identified commands, then appends lane-post and elapsed-post at their next DEBUG events. The accepted charged interval is `elapsed_pre - lane_post` through `elapsed_post - lane_pre`; six streams and the formal timeout each have cardinality and numeric guards before arithmetic. The complete production control exits zero. Scoped plus-one-second, missing-run-clock, and duplicate-boundary mutants each make exactly the named budget case fail and no other case. Production `opencode_review.sh` remains unchanged |
185
+ | SUPERSEDES the immediately preceding pre/post-bracket row: measured ambiguity is not permission to widen a safety oracle. If a one-second bracket would also admit the smallest producer regression, reject that observation, resample only the isolated probe under a strict attempt bound, and fail the named assertion on exhaustion. Make the ambiguous first attempt deterministic so retry behavior is exercised rather than merely present | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h13b review showed that the h12 interval can admit a plus-one elapsed mutation whenever a pre/post pair crosses a second boundary. The fixture now accepts only zero-width lane and elapsed pairs, with at most three isolated attempts. Its controller deliberately widens the first elapsed-post observation; the correct wrapper resamples and passes on the stable second attempt, while the scoped plus-one producer mutant fails exactly the named budget case. Missing-run-clock and duplicate-boundary mutants also fail only that case. Exhausted ambiguity is a named RED, never a false green. Production `opencode_review.sh` remains unchanged |
186
+ | SUPERSEDES only the preceding row's assumption that the second attempt is the successful one: a bounded retry proves eventual acceptance of a valid observation, not a fixed attempt index. Assert that the deterministic ambiguity was consumed and that a retry occurred; let the zero-width and attempt-bound guards decide whether attempt two or a later allowed attempt succeeded | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h14 review found that the forced first ambiguity makes attempt two usual but not guaranteed: attempt two may naturally cross a second boundary and a stable attempt three is still valid. The test now requires the forced-ambiguity marker and `budget_attempt >= 2` instead of equality with two. The existing zero-width check still rejects ambiguity, the three-attempt loop still fails closed on exhaustion, and the exact elapsed oracle is unchanged. This closes a load-dependent false RED without widening any accepted budget |
187
+ | A differential mutant is attributed to a FAILURE CLASS, not merely to a shared test case. When observation validity, retry exhaustion, and semantic arithmetic all feed one assertion label, one RED count cannot show which mechanism killed the mutant. Make prerequisite checks mutually exclusive and let later semantic checks become non-applicable rather than collateral failures when an earlier prerequisite is absent | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h15b challenge found every timing defect reported through one budget label. The fixture now separates complete numeric observations, stable sampling within the bound, and exact elapsed arithmetic. Four scoped disposable runs each produce exactly one attributable failure: plus-one elapsed names arithmetic; missing run and duplicate boundary name observation completeness; forced ambiguity on all three attempts names stable sampling. The unmutated suite exits zero. Production `opencode_review.sh` remains unchanged |
188
+ | A synthetic VCS-corpus fixture must represent tracked WORKTREE deletions as absence, not crash while copying a `git ls-files` name that no longer exists. Let the copied gate diagnose that candidate, keep the positive control bound to its explicit success token, and pin any owner-root-only safety rule to the exact mandatory invocation rather than expanding it to foreign or partial scans | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h15b challenge identified the tracked-deletion copy exception and questioned the root equality boundary. The fixture now skips tracked names absent from the worktree, conditionally copies the waived ledger outside `specs/`, preserves symlinks, and requires `spec_reference_check_ok` in the control. The root expansion is rejected under the established whole-repository contract; a new case instead pins both mandatory recipes, Makefile and CI, to `python3 scripts/check-spec-references.py .` and asserts the checker's owning root. The focused suite runs 72 tests and exits zero. An unstaged unrelated citation cannot masquerade as attributed because the positive control already requires exit zero |
189
+ | SUPERSEDES the failure-class row's remaining compound semantic label: once prerequisites have their own mutually exclusive checks, the arithmetic assertion must contain only the arithmetic relation. Controller exit, retry marker, cap, ordering, and minimum charged work are a separate class even when one producer mutation cannot currently flip them | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h16 challenge found the purported arithmetic label still conjoined controller evidence. The permanent test now separates that evidence from exact charged-elapsed equality. Five scoped runs are fully discriminating: the unmutated wrapper passes; plus-one elapsed fails only `formal timeout equals independently observed elapsed budget`; missing run and duplicate boundary fail only complete observations; three forced ambiguous attempts fail only stable sampling. Production `opencode_review.sh` remains unchanged |
190
+ | A root-scoped gate contract is pinned by executing its exact argv from its resolved working directory, not only by finding that command as text. Pair that execution with the success token, and pair every stale-only negative with explicit absence of the same token so stdout cannot contradict the exit status. Do not repair a synthetic VCS corpus by importing ambient untracked files: a missing dependency must make the positive control RED | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h16 challenge exposed two evidence gaps and proposed one incompatible expansion. The focused suite now executes `python3 scripts/check-spec-references.py .` with cwd set to the resolved owning repository after checking the Makefile and CI recipes, and requires zero exit plus `spec_reference_check_ok`. Its stale-only CLI case requires the token absent. The production main already returned before printing success when stale, so this is a proof assertion rather than a behavior fix. Copying an unstaged plan is rejected because h11 established a tracked-only candidate corpus; any missing dependency already fails the positive control. The suite remains 72 tests and exits zero |
191
+ | SUPERSEDES the controller-evidence half of the preceding arithmetic-attribution row: an independent check cannot use the same derived charged value merely with a different inequality. Put the required-work floor on the independent boundary-to-run monotonic interval, and mutate elapsed in both directions so either sign can only fail the pure equality check | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h17 challenge found `charged_elapsed >= 3` still coupled the controller label to producer arithmetic. The controller now requires at least three billion independently observed monotonic nanoseconds between boundary and run entries. Scoped plus-one and minus-one elapsed mutants each fail only `formal timeout equals independently observed elapsed budget`; missing run, duplicate boundary, and three-attempt exhaustion keep their distinct one-failure labels; the unmutated suite exits zero. Production `opencode_review.sh` remains unchanged |
192
+ | A mandatory-invocation test must bind a command to its EXECUTED CONFIGURATION, not find a free-floating substring and then run a test-chosen equivalent. Isolate the owning Make target and CI job/step, reject effective working-directory overrides, and execute the extracted exact argv from the resolved owner root. This pins the supported in-repository checker without granting its shipped table to an external installed copy | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h17 challenge showed a bare substring could survive after the recipe moved. The focused case now narrows to Makefile `test`, the CI `repository-gates` job and named Spec reference gate step, asserts no workflow/job `working-directory`, then executes exact `python3 scripts/check-spec-references.py .` from the resolved owner root and requires success. Moving `check-spec-references.py` to an unrelated tools directory is rejected as outside the established owner-root contract; partial and foreign scans remain inert. The 72-case suite exits zero |
193
+ | SUPERSEDES only the Make-target locator in the preceding mandatory-invocation row: a target name is a complete line prefix, not a substring. Find the exact `test:` line and consume only its following tab-indented recipe; otherwise an earlier `fast-test:` or comment can donate a matching command that `make test` never runs | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h18 review found the first implementation used `split("test:", 1)`. It now enumerates Makefile lines, selects `line.startswith("test:")`, collects the contiguous tab-indented recipe, and requires the exact root-scanning command there. CI step scoping and the executed owner-root argv remain paired. The 72-case suite exits zero |
194
+ | A mandatory CI command is bound by its effective CWD and ENABLEMENT, not only its job and step text. YAML top-level defaults are order-independent, so inspect every top-level defaults block wherever it appears; inspect the owning job separately; and reject disabling `if:` keys on both job and step. Scope the cwd rule to the affected job so unrelated package jobs may keep legitimate working directories | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/test_check_spec_references.py | updated | The h19 challenge showed defaults written after `jobs:` and `if: false` could leave the command text intact while preventing the owner-root gate. The focused test now scans all top-level defaults blocks regardless of order, the entire `repository-gates` job for working-directory, and its job header plus named step for `if:`. A deliberately overbroad first fix failed on the existing npm job's legitimate package cwd and was narrowed to effective scope. The exact command still executes from resolved owner root and the 72-case suite exits zero |
195
+ | A waiver pin is live only for a citation the waiver COULD apply to. Seeing the same token and bytes is insufficient when the target escapes containment: marking that pin seen leaves a permanent finding beside a permanently healthy-looking waiver that covers nothing. Apply the containment boundary to liveness as well as application | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:scripts/check-spec-references.py | updated | The h19 challenge found `seen_pins` was populated before the escape check. The checker now resolves each pinned waiver token and marks it seen only when it stays inside root. The existing containment case was RED on the new stale assertion before the change: the citation remained a finding but stale was empty. After the fix it is both unwaived and stale, while the 72-case suite and repository gate exit zero |
196
+ | SUPERSEDES the CI-enablement scope in the earlier effective-CWD row: job keys are order-independent, so inspect the full owning job for a job-level `if:` rather than only its header. A mandatory gate also fails open under `continue-on-error`; reject that key at job indentation and on the named step, without forbidding controls on unrelated steps or jobs | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h20 review found a job-level `if: false` placed after `steps` escaped the header slice, and neither job nor step pinned `continue-on-error`. The focused test now scans the full `repository-gates` slice for four-space job keys and the named Spec reference gate block for eight-space step keys. The exact command, effective cwd, and executed owner-root proof remain paired. The 72-case suite exits zero |
197
+ | SUPERSEDES only the h17 row's claim that every timeout mutant has a distinct label: diagnostic attribution is per FAILURE CLASS, so missing-run and duplicate-boundary intentionally share the complete-observation label while arithmetic drift and retry exhaustion retain their own labels. Do not describe one label per mutant when several mutants exercise the same prerequisite class | `code-review` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:skills/code-review/scripts/test_opencode_review_retry.sh | updated | The h21b challenge compared the register wording with the five mutant outputs. The implementation is already consistent with the h15b failure-class decision: plus/minus elapsed fail only exact arithmetic; missing-run and duplicate-boundary each fail only `OpenCode budget probe captures one complete numeric observation set`; three-attempt ambiguity fails only stable sampling. This row corrects the evidence claim without splitting one failure class into artificial mutant-specific assertions. Production `opencode_review.sh` remains unchanged |
198
+ | Liveness and waiver application must reject the same unresolved-target domain. Require a non-empty resolved target before a pin can be live, even when the current citation grammar cannot emit an empty target; that keeps a future grammar change from reviving a covering-nothing waiver | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h21b challenge proposed locator-only and separator-only tokens. The current `specs/` grammar makes those examples unreachable, but the liveness guard now states the full contract with `bool(waived_target)` rather than a vacuous None check. A widened-parser regression case makes `/` resolve empty and requires the pin to remain stale. The production 73-case suite exits zero |
199
+ | A mandatory CI gate depends on workflow entry and job dependency state as well as command, cwd, `if`, and `continue-on-error`. Pin the accepted trigger contract and reject `needs` on the owning job, so the command cannot remain textually perfect while its workflow or prerequisite prevents execution | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h21b challenge identified top-level trigger narrowing and skipped dependencies. The focused test now requires an unconditional `pull_request` trigger, the established `push`-to-main trigger with no path filter, and no job-level `needs` in `repository-gates`. Disposable workflow mutants for `pull_request.paths`, workflow-dispatch-only, and job `needs` each fail the named invocation case; unrelated jobs remain outside the assertion. The exact owner-root command still executes successfully |
200
+ | SUPERSEDES the widened-parser fixture in the preceding liveness row: a shared invariant should be pinned at the shared predicate, not through one mocked parser route. Give waiver application and pin liveness one helper over resolved-target emptiness and containment, then test that helper directly for empty, escaping, and eligible targets | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h22 challenge showed the mocked citation grammar proved only one hypothetical widening. `waiver_can_apply_to_target` is now called by both the spend and seen paths. Its direct case rejects an empty target and an escape and accepts a non-empty contained target, independent of how a parser or resolver produced the value. Existing exact-token, containment, stale, and application cases remain intact |
201
+ | SUPERSEDES the exact-text trigger shape in the preceding CI-entry row: validate trigger SEMANTICS, not one YAML spelling. Accept quoted keys, comments, inline event lists, legitimate event types, and branch filters that still include main; reject missing required events, path filters, main exclusion, and pull-request types that omit opened or synchronize, with an explicit diagnostic instead of a parser exception | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h22 challenge found an empty pull-request block assertion false-red on valid configuration and `list.index("on:")` failed opaquely on equivalent YAML. The test now parses with PyYAML's BaseLoader, which preserves the `on` key, and runs positive and negative trigger-contract tables. The production workflow, 75-case suite, exact owner-root command, and job/step enablement checks all pass |
202
+ | A mandatory test's import-time dependencies must be declared once for both contributors and CI. Centralize the existing pytest and PyYAML requirements rather than relying on runner-only install lines or letting an undeclared module failure look like a regression in the gate under test | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h23 review found the semantic workflow test made PyYAML a hard import without a repository test-dependency manifest. `requirements-test.txt` now declares both Python packages already installed by CI; both CI jobs consume that file, and the contribution guide gives the same installation command. The workflow assertion and all 75 focused cases remain active rather than skipping when setup is incomplete |
203
+ | SUPERSEDES the preceding dependency row's unproven firing path: dependency declaration is executable only when the test reads the manifest, checks every owning CI job consumes it after checkout from the owner root, and checks the contributor command. Pair those static mutations with an ignore-installed resolver dry-run; an already-satisfied local environment is not clean-runner evidence | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h24b challenge found the 75 cases did not cover any dependency bytes and the local dry-run reused installed packages. A new case requires exactly pytest and pyyaml, validates checkout precedes the named install step in both jobs, rejects install/job cwd overrides, and checks CONTRIBUTING. The 76-case suite exits zero. `pip install --dry-run --ignore-installed -r requirements-test.txt` independently resolves all packages and reports the would-install set without mutating the environment |
204
+ | A configuration gate should locate semantics without freezing benign labels or tool versions, and a missing semantic anchor should fail with the anchor's name rather than a bare iterator exception. Match checkout by action family, match dependency installation by command, and give each absence an explicit diagnostic | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h25 review found exact checkout v4 and install-step-name lookups could raise unlabeled StopIteration. The test now accepts checkout v5 and an arbitrary install label while still ordering the actual manifest command after checkout. Disposable probes show that benign pair stays green, while missing checkout and missing manifest install each fail with their own message. The 76-case production suite exits zero |
205
+ | SUPERSEDES command-substring detection in the preceding semantic-anchor row: a textual mention is not execution. Recognize only a non-comment run line whose stripped bytes exactly equal the required command, coerce non-string run values to non-matches, and reject workflow-level disabling controls on the matched step. Keep echo, comment, null, and list decoys in the permanent suite | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h26b challenge showed substring search accepted echo/comment decoys and could raise TypeError on null run values. `step_runs_exact_command` now owns executable-line recognition; the dependency-wiring case also rejects `if` and `continue-on-error`. A new permanent decoy table makes all four false matches RED while the real multiline install remains green. The focused suite now runs 77 tests and exits zero |
206
+ | SUPERSEDES the preceding row's line-within-script model: make the manifest install its own single-command step, then validate the whole execution hierarchy. Root/job run defaults, job `if`/`continue-on-error`/`needs`, checkout controls and alternate checkout paths, install ordering, and install step controls must all preserve owner-root execution | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h27 challenge found step-only controls left job and root bypasses. CI now separates apt tools from the exact Python dependency step. `dependency_wiring_violation` validates both required jobs across root, job, checkout, and install layers; a permanent workflow table covers every disabling/relocation class plus echo decoy, while the checkout-v5/arbitrary-label control stays green. The focused suite runs 78 tests and exits zero |
207
+ | A claimed exhaustive mutation matrix must exercise every implemented rejection branch, not only one representative per layer. Keep checkout continue-on-error and sparse checkout, install cwd and continue-on-error, and reversed checkout/install order beside the existing root/job/step mutants | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h28 review found five validator branches had no permanent mutant despite the preceding row's exhaustive wording. Those five workflows now join the rejected table and each asserts its specific violation. The production validator is unchanged; the 78-test suite and every subcase exit zero |
208
+ | SUPERSEDES completed-exhaustiveness wording in the preceding matrix rows: the durable rule is prospective—EVERY NEW rejection branch must ship with a permanent mutant that selects its branch-specific diagnostic. Current evidence names the covered branches and candidate; it does not promise that a future expanded validator remains exhaustive without expanding the table | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h29 challenge found an append-only claim of permanent exhaustiveness would become false as soon as a later branch landed alone. Candidate cc27d86 covers the currently named root/job/checkout/install branches, including the five h28 additions; future changes must extend the table in the same change. The 78-test suite exits zero |
209
+ | SUPERSEDES only the preceding row's completed-coverage evidence sentence: retain the prospective obligation without asserting that today's table is permanently or absolutely exhaustive. Evidence is an enumerated, candidate-bound list; future branches extend it rather than inheriting a blanket coverage claim | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: command:scripts/test_check_spec_references.py | updated | The h30 challenge used the complete source to identify remaining structural/cardinality branches. The matrix now adds invalid YAML/document/defaults/jobs/steps/options, missing and duplicate checkout/install, and a default two-job run whose second job is disabled. These are cc27d86-successor evidence only; the durable rule remains that later branches add their own mutants |
210
+ | Session-vantage leakage remains an entrypoint deletion class while its exhaustive categories and repair detail live in the canonical reference; rehosting that detail must preserve the stable firing-path anchor, HEAD-reader test, mandatory reference load, fact-first order, and inline no-delete guards for stable versions, issue references, suppression reasons, counterfactual pins, and measurements | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/tighten-doc/SKILL.md#会话视角泄漏:句子立足于写作会话而非文档 | updated | `tighten-doc/SKILL.md` is the changed owner key. DELETE #12 keeps the original anchor, all eight walked class labels including unanchored-referent and working-language triggers, mandatory reference load, and irreversible-deletion guards. Canonical safe-mode r9 (base `653c985`, head `744b505`; provider `claude`, model `claude-haiku-4-5`, 2 rounds/arm) used the same committed runner-time grader for all arms and measured **base 2/2, head 2/2, model-visible pointer/reference-removal mutant 0/2**. This proves preservation and carrier sensitivity, not base-to-head improvement; removing DELETE #12 also removes the only runtime load path to the canonical reference, so the mutant is an assembled-packet mutation rather than an independently hidden-reference arm. Plain `make test` automatically enforces the clean actual-`HEAD` subtree/body binding whenever this C3 repair or its evidence changes relative to fixed base discovery (`dev`, then `origin/dev`); if neither ref exists it binds fail-closed, and `--candidate` can force the same non-overridable check. Later work that does not change this C3 landing surface is not permanently pinned to historical owner trees. The measured tighten subtree is `1cc734c`, and no revision override is accepted. The regrader verifies every recorded answer hash and rejects unknown task IDs and absent arm declarations; the contract forbids canonical pass-count drift and unallowlisted arm status. Current raw answers and scores are `specs/023-agent-native-repo-borrowing/evidence/red-baseline-023-c3-tighten-current-r9.json` and its `-regrade.json` sibling. r5 is manifest audit-only; mixed r6 is also top-level audit-only. All audit files are digest-pinned and each task names r9 or r8 as its canonical superseder. r9 and r8 measured different heads; no joint model run covers both owners, so cross-owner interaction is unmeasured. Trigger, scope, routing, validation, and acceptance semantics are unchanged; C3 is the budget constraint, not the behavior claim. |
211
+ | A shared branch still defaults to merge or platform update, and any exceptional rebase must refresh the target before observing the rewritten branch, carry that branch's single fetched OID through ancestry validation and an explicit lease, and refresh invalidated review and CI state after push | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/worktree-isolation/SKILL.md#已推送 / 挂着 MR / 别人可能在上面工作的分支 | updated | `worktree-isolation/SKILL.md` is the changed owner key. Its entrypoint remains the sole exceptional-rebase action/stop sequence and keeps the default no-blind-rebase decision, explicitly alternative refspec variant, literal fetched OID, exact topology command and stop, explicit lease, force bans, all six invalidated states, and combined-tool fallback. Independent Claude review found that intermediate compact forms wrongly banned fetching the target, counted a legitimate target fetch as a second branch fetch, and rendered two alternative branch-fetch commands as a sequence; the fixes now make the target refresh literal, mark branch fetch forms as alternatives, count one pre-push fetch of rewritten branch `mybr`, and forbid only a second fetch of that branch. Canonical safe-mode r8 record (exact head `cfade5b`; provider `claude`, model `claude-haiku-4-5`, 2 rounds/arm) measured **head 2/2; combined action/reference-carrier-removal mutant 0/2**. That proves the assembled packet depends on the two carriers together, not that either one alone caused the delta; no valid base-to-head failure is claimed. Plain `make test` automatically enforces the clean actual-`HEAD` subtree/body binding whenever this C3 repair or its evidence changes relative to fixed base discovery (`dev`, then `origin/dev`); if neither ref exists it binds fail-closed, and `--candidate` can force the same non-overridable check. Later work that does not change this C3 landing surface is not permanently pinned to historical owner trees. The measured worktree subtree is `c5c478b`, the assembled skill-plus-reference hash is recomputed, and no revision override is accepted. The historical base arm is marked advisory in raw and regrade because the pre-repair runner rendered its absent reference as a named empty block; it cannot support a base-to-head claim and is the sole allowlisted canonical arm status. The repaired runner aborts required blob failures and verifies the one declared base absence before omitting that block. Raw answers and answer-hash scores are `specs/023-agent-native-repo-borrowing/evidence/red-baseline-023-c3-worktree-safe-r8.json` and its `-regrade.json` sibling; every recorded answer hash is checked and unknown task IDs fail closed. The grader accepts safe target-fetch spellings but still requires exact branch-vs-remote topology, one pre-push branch fetch, explicit OID lease, and post-push six-state revalidation. r4 and mixed r6 artifacts are audit-only, digest-pinned, and require r8 or r9 as canonical superseders. r9 and r8 measured different heads; no joint model run covers both owners, so cross-owner interaction is unmeasured. The zero-loss map documents carriers for review; C3 is the budget constraint, not a behavior claim. |
212
+ | A catalog-contract fixture must compare its worktree mutations against a test-owned source-HEAD base, not inherit the source branch's accumulated impact-chain ancestry; otherwise its pristine control can fail before the catalog assertion runs | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | `skill-extraction-workflow/SKILL.md` is the unchanged owner key. `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh` fetches the exact source HEAD into every clone, creates `ccl-test-base` at `FETCH_HEAD`, and scopes `CCL_SKILL_BASE_REF` to that ref inside `run_case`. The unpinned C3 candidate reproduced `FAIL[c6]`; unchanged `dev` passed the fast suite; the pinned exact candidate passed all eleven catalog cases. Production `check-ccl-skills.sh` base selection is unchanged. |
213
+ | SUPERSEDES only any implication that canonical r8 raw answers used the current detail grader: preserve runner-time pass/of, but compare every round's `missed` array and allow only the three named historical r8 differences; ordinary owner-only worktree edits remain pending while committed/candidate/strict drift stays fail-closed | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/worktree-isolation/SKILL.md#已推送 / 挂着 MR / 别人可能在上面工作的分支 | updated | `worktree-isolation/SKILL.md` is the changed owner key. Claude h5 found base-r1 and mutant-r1/r2 detail drift despite unchanged scores. `test_red_baseline_023_c3_regrade.rb` exact-allowlists those raw/current arrays and rejects any fourth difference. It always verifies committed `HEAD`; owner-only dirty state emits pending on an unchanged fixed base, while `--candidate`, committed C3 changes, and `CCL_C3_STRICT_EVIDENCE=1` block with recovery guidance. |
214
+ | SUPERSEDES the preceding row's completion-boundary wording: direct contract use may report owner-only dirty work as pending, but the repository `make test` gate itself MUST select `--candidate`; pending and clean runs must have different terminal tokens so a green diagnostic cannot be cited as clean preservation evidence | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/worktree-isolation/SKILL.md#已推送 / 挂着 MR / 别人可能在上面工作的分支 | updated | `worktree-isolation/SKILL.md` remains the changed owner key. Claude h6 challenge showed the permissive direct mode was also the Makefile path. `Makefile` now passes `--candidate`; dirty owners block with recovery guidance. Direct in-progress use emits `c3_owner_worktree_dirty_evidence_pending` and terminates with `c3_regrade_contract_tests_pending_dirty_owners`, while clean candidate runs alone emit `c3_regrade_contract_tests_ok_unattested`. |
215
+ | SUPERSEDES the preceding row's token-only pending boundary: every non-clean evidence state must also be nonzero so an exit-code-only direct caller cannot convert pending into pass | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#c3_regrade_contract_tests_pending_dirty_owners | updated | Claude h15 challenge and a frozen RED probe showed direct dirty-owner diagnostics ended with the pending token but rc 0. The direct path now completes deterministic diagnostics, prints `c3_regrade_contract_tests_pending_dirty_owners`, and exits 2. Candidate/strict dirty carriers still block earlier; only `c3_regrade_contract_tests_ok_unattested` exits zero. |
216
+ | SUPERSEDES the earlier whole-owner subtree and dirty completion boundary: preservation freshness must pin only the files actually assembled into the model-visible evidence, so an unrelated file elsewhere in the same skill package neither becomes stale nor requires a provider-dependent refresh | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:skills/worktree-isolation/SKILL.md#已推送 / 挂着 MR / 别人可能在上面工作的分支 | updated | Claude h7 found the subtree OID and package-root dirty check over-constrained `make test`. `test_red_baseline_023_c3_regrade.rb` now names exactly four carriers: both `SKILL.md` files, `session-vantage-leakage.md`, and `shared-branch-rebase.md`. Their committed assembled-body hashes remain unconditional; measured dirty state still blocks under `--candidate` and strict mode, while an unrelated same-package path is excluded by a permanent truth-table assertion and exact-candidate dirty/committed probes. Ruby is not a new dependency because the pre-existing first `make test` gate already requires and invokes it. |
217
+ | SUPERSEDES any claim that body-hash equality alone classifies and pins all C3 evidence: canonical raw/regrade bytes must be digest-visible before their expected body hash is trusted, every non-manifest JSON in the evidence directory must be canonical or audit-listed, and the preservation claim must explicitly stop at the assembled body rather than silently include frontmatter/routing | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: no; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#CANONICAL_FILE_SHA256 | updated | Claude h8 challenge found an expected-hash self-edit path and a filename-prefix classification gap, plus ambiguous frontmatter wording. `CANONICAL_FILE_SHA256` pins all four canonical files and the real verifier rejects same-size first/middle/last-byte mutants; classification now scans every `*.json` except the manifest. r8/r9 still hash `SKILL.md` only after frontmatter removal plus the named reference: that is the declared body-preservation boundary, while name/description/routing changes stay with routing gates. Digest constants expose candidate changes for exact-tree review; they are not provider identity or an immutable trust root. |
218
+ | SUPERSEDES the preceding row's all-JSON classification scope: the round-023 evidence directory is shared, so completeness must cover the declared C3 namespaces without reclassifying sibling I/II/IIb/IIc/III/strict artifacts | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#C3_EVIDENCE_JSON_PREFIXES | updated | The literal all-`*.json` implementation immediately made the unchanged exact candidate red because twelve pre-existing sibling artifacts are not C3 records. `C3_EVIDENCE_JSON_PREFIXES` now owns `red-baseline-023-c3-*` and `c3-*`; permanent assertions require the challenge example `c3-extra-evidence.json` to be classified and keep `red-baseline-023-II.json` outside. An arbitrarily unmarked filename cannot be semantically inferred as C3 evidence and remains an explicit naming-contract limitation. |
219
+ | SUPERSEDES the preceding C3 completion-boundary rows: a clean verdict must not combine carrier blobs from committed `HEAD` with raw, regrade, manifest, or contract bytes that exist only in the worktree; the Makefile, contract, and both declared evidence namespaces must also remain inside the C3 trigger/control surface | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#C3_CONTROL_WORKTREE_PATHS | updated | Claude h12 reproduced a clean token after committing a carrier change while leaving the matching raw, regenerated regrade, and both digest-pin edits dirty. The contract now blocks any dirty C3 plan, contract, manifest, `red-baseline-023-c3-*`, or `c3-*.json` control path; committed changes to those paths and `Makefile` enter candidate binding. Both HEAD and worktree Makefile text must retain the exact `test` recipe invocation with `--candidate`, without blocking unrelated Makefile edits. The repeated catalog finding remains rejected because eight permanent mutation cases already require nonzero, branch-specific failures under the source-HEAD base. |
220
+ | SUPERSEDES only the Makefile-cardinality implication in the preceding row: locating one valid `test:` recipe is insufficient when a later duplicate target can replace its effective commands; require exactly one target and retain the command under that unique recipe | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#make_test_candidate_invocation? | updated | Implementer h13 self-review appended a second `test:` recipe containing only `true`: the old scanner emitted the clean token while `make -n test` contained zero C3 invocations. `make_test_candidate_invocation?` now rejects duplicate targets, and its permanent mutant joins the removed-`--candidate`, moved-command, and unrelated-target controls. |
221
+ | SUPERSEDES the literal-target spelling in the preceding row: Make allows `test` inside a multi-target rule, so cardinality must count rule target lists rather than lines starting with `test:` | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#make_test_candidate_invocation? | updated | Claude h14 showed `foo test:` replaced the effective recipe while the prefix scanner still emitted clean. The helper now scans non-recipe rule lines, splits the colon-left target list, and requires exactly one declaration containing the exact `test` token. Permanent literal-duplicate and multi-target mutants both block; the unrelated-target control stays green. |
222
+ | SUPERSEDES syntax-only Makefile validation in the preceding rows: use the text parser for permanent mutation controls, but let GNU Make resolve the effective recipe and require its dry-run to contain the exact C3 candidate command once | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#effective_c3_invocations | updated | The duplicate and multi-target findings show why effective Make semantics, not a growing list of spellings, own the final boundary. The contract still rejects removed, moved, duplicate, and multi-target text mutants, then runs `make --no-print-directory -n test` and requires exactly one exact C3 `--candidate` command. This also covers variable expansion and included rule overrides on the current candidate without executing the test recipe. |
223
+ | SUPERSEDES only the inherited-environment assumption in the effective Make probe: a nested dry run must evaluate the repository recipe independently of parent Make control flags, and the production probe itself must survive a hostile inherited question-mode regression | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#line-sha256:5172-b00d-c562-745a-83e9-d49b-2881-464c-cc65-0836-8782-6c54-1ccb-5466-124d-40d0,file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#line-sha256:7f8e-93c0-5652-31f3-d4c0-f1cf-8f70-399e-11c9-512c-10a6-20ac-14ac-f9f4-e533-9c51 | updated | Claude h16's `-s`, `-i`, and `-k` examples did not reproduce, but exact h16 under `MAKEFLAGS=-q` exited 1 and `-t` suppressed the expected command. The contract clears `MAKEFLAGS`, `MFLAGS`, and `GNUMAKEFLAGS` for `make --no-print-directory -n test`; on every run it poisons those same parent variables with `-q`, then independently asserts that all three hostile values are active immediately before `Open3.capture3`. The firing paths bind both complete trimmed code lines. Moving only the poison below the call now fails the runtime assertion; removing or changing either line makes the whole-ledger resolver red. Moving both is a coordinated exact-tree edit, not a one-line bypass. |
224
+ | SUPERSEDES the shared-oracle weakness in the preceding Make-environment row: poison the parent and delete the child Make flags from independently spelled literals, so erosion of a behavior-critical deletion key makes the same production probe red without relying on a textual guard for a guard | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:specs/023-agent-native-repo-borrowing/evidence/test_red_baseline_023_c3_regrade.rb#line-sha256:cfb5-50a0-c72b-db21-0fd8-16db-4798-7e01-57bd-07ad-497a-f268-ab70-c06a-a7b0-9ae0 | updated | Claude h17 found the inherited `MAKEFLAGS=-q/-t` false-red; h20 then proved a textual inventory assertion could be replaced by a comment decoy while both contract and register stayed green. The final contract removes that self-adjudicating assertion: one literal list injects hostile parent flags and a separately spelled hash deletes the child flags. Removing only `MAKEFLAGS` from the deletion map leaves its poison active and makes the unchanged candidate fail through the real effective-recipe probe. Removing a poison key alone does not weaken production neutralization. Coordinated edits to both independent spellings remain visible exact-tree changes for the required external review. |
225
+ | SUPERSEDES literal-substring firing paths for guard implementation lines: when a comment or inert string decoy would preserve the same text, bind the SHA-256 of the complete trimmed source line and fail closed on malformed or blank-line digests | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb | updated | Claude h20 proved that a literal diagnostic left in a comment satisfied the old substring resolver. Claude h21-r2 then found the Make poison row anchored a different live line, allowing a green poison deletion followed by green neutralization deletion. `line-sha256:` now hashes each complete non-empty trimmed source line and accepts four-character hex groups so public sanitization does not mistake an opaque digest for business data. Ruby targets are hashed and lexed from unmodified file bytes and require a Ripper code token outside comments, embedded docs, heredocs, ordinary string content, and word-list separators on that line; other languages fail closed until they have an equivalent classifier. The exact poison line resolves while deletion, rewording, comment-prefix replacement, exact heredoc or `%w[...]` content copies, and a fence-bracketed heredoc transform report `line-sha256 source line absent from target`; malformed and ubiquitous empty-line digests are independently RED. Fence stripping remains limited to generic literal and heading-slug reader-prose locators. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`; `skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh`. |
226
+ | Registering a wall-clock rule and ROUTING its prevention to a test-layer owner does not close the instances that already sit beside the fix. The landing round bounded one case and left a pinned literal in a sibling case of the same file, so the identical one-second cliff re-reported later from a different line under load. When a bounded-versus-pinned repair lands, enumerate every case in the same owner that reads the same derived value and convert them in that round, or record which ones were inspected and deliberately left. A pattern row plus a routed prevention is a rule, not an inventory | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_claude_review_probe.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the change lands in one case of `code-review/scripts/test_claude_review_probe.sh`. Observed failure: the runtime main-timeout case failed in two of three full-suite runs and passed three of three standalone, with payload reason `Claude invocation timed out after 4 seconds` against an assertion pinned to the literal `after 5 seconds`. Proven cause: `claude_review.sh` derives the enforced budget as the requested timeout minus whole-second wrapper elapsed, so any integer second the preamble crosses before the main invocation lowers the reported number below the requested 5. Made deterministic instead of re-run: a disposable exact-test copy added a one-second delay to only that invocation's help probe and reproduced the identical traceback and payload; with the repaired predicate the same delayed copy returns `claude_review_runtime_tests_ok`, so the injected delay breaks that assertion and nothing else in the file. The repair reuses the bounded extraction the adjacent baseline-budget case already applies, with the upper bound at 5 because this case has no baseline consumption. Applied mutation, evaluating the shipped predicate's own extracted bytes against synthetic recorded budgets: the new predicate passes at 5, 4, 3, 2 and 1 and fails at 0, 6, 10, 600 and on a reason carrying no number, so both bounds are load-bearing, while the old predicate passes only at 5 and fails at 4, which demonstrates the one-second cliff rather than arguing it. No assertion is removed, nothing is skipped, and `claude_review.sh` is untouched |
227
+ | SUPERSEDES the bounded-versus-pinned repair's first bound: a bound wide enough to survive load is not automatically a bound that still proves what the assertion existed to prove. Replacing an exact literal with `1..requested` removed the wall-clock cliff but also stopped proving that the requested budget was plumbed through at all, so a wrapper silently enforcing a small default would pass. Choose the lower bound from the largest preamble the mechanism can actually consume, not from the widest value that cannot fail, and reject a reviewer's suggestion to widen further when the repository already holds a rule that measured ambiguity is not permission to widen a safety oracle | `code-review` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/code-review/scripts/test_claude_review_probe.sh | updated | `code-review/SKILL.md` is the owner key and is unchanged in this round; the change lands in one case of `code-review/scripts/test_claude_review_probe.sh`, with the round record in `specs/023-agent-native-repo-borrowing/c3-release-gate-repair.md`. Observed failure: an independent release review returned three findings on the same asserted line, one asking to tighten to 4..5, one to 3..5, and one to widen to 0..5. Proven cause: the reported budget is the requested timeout minus whole-second wrapper elapsed, so the bound has to exclude two different failure modes at once — a wall-clock cliff above and a silently defaulted budget below. Resolution: `3 <= n <= 5`, which allows twice the one second the deterministically reproduced flake ever consumed while still failing a 1- or 2-second default. The widening suggestion is rejected on this repository's own recorded h13b rule. Predicate mutation on the shipped bytes: passes at 5, 4 and 3; fails at 2, 1, 0, 6, 10, 600 and on a reason carrying no number, so both bounds stay load-bearing |
228
+ | A model-graded preservation record is evidence about the exact BYTES it measured, so repairing the measured carrier retires the record rather than merely dating it. When a review finds a real obligation the shrunken entrypoint dropped, expect the fix to make the owner body stale, the gate to fail closed, and the honest recovery to be a re-measurement whose numbers may be WORSE than the record it replaces — publish those numbers and withdraw the claim they no longer support, instead of keeping the flattering superseded record canonical. A demotion whose cause is a changed carrier is its own class and must not be filed under grader or rubric drift | `worktree-isolation` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/worktree-isolation/SKILL.md#右侧(远端独有)非 | updated | `worktree-isolation/SKILL.md` is the owner key and is the changed carrier in this round. Observed failure: the C3 entrypoint shrink had deleted the `右侧(远端独有)` gloss, leaving `右侧非 0 即并入、禁 rebase`; because `git rev-list --left-right --count` column order is not self-evident, a reader who guesses it inverts the stop condition and rebases over remote-only commits. The proof was inside the round's own evidence: both r8 head answers that the grader scored 2/2 with zero missed labels state the inverted reading. Restoring the two-word gloss changed the measured body and the C3 contract failed closed on stale owner body exactly as designed. It also grew a historical over-limit entrypoint by four word units, which the size gate then blocked, so the rationale-only parenthetical moved into the reference and the entrypoint ended one word below its pre-repair value. Both carrier edits had to land before re-measuring; refreshing after the first one cost a second provider round for nothing, which is the operational lesson: batch every carrier edit, then measure once. r11 is canonical at base 1/2, head 1/2, mutant 0/2 — level with base and above the mutant — against r8's head 2/2 and r10's base 0/2. r11 needs neither of r8's advisory-arm nor missed-drift allowances, both of which are now empty; r8 and r10 are audit-only under a new `superseded-stale-owner-body` class, and the worktree preservation claim is withdrawn to carrier sensitivity only |
229
+ | SUPERSEDES every earlier row's present-tense claim that `red-baseline-023-c3-worktree-safe-r8.json` is the current raw answer and that its advisory base arm is the sole allowlisted canonical arm status. On an append-only ledger a row states its OWN round, so any sentence written in the present tense about a mutable pin becomes a false claim the moment that pin moves — and it cannot be edited afterwards. Write pins and canonical sets as of-that-round facts, and when they move, supersede rather than leave the reader to discover the contradiction against the shipped contract | `skill-extraction-workflow` | behavioral-evidence: not-required wording-only; observed-failure: no | routed | No owner package changed for this correction, so no owner key is cited: the round's skill-extraction-workflow edit is the ledger itself, which this contract excludes from the owner surface. This row corrects ledger prose only and changes no executable behavior. As of this round the shipped contract pins `CANONICAL_CASES` and `CANONICAL_FILE_SHA256` to the r11 raw/regrade pair for worktree-isolation and to the r9 pair for tighten-doc; `CANONICAL_ARM_STATUS_ALLOWLIST` and `CANONICAL_MISSED_DRIFT_ALLOWLIST` are both empty; r8 and r10 sit in `AUDIT_FILE_SHA256` and in `audit-only-evidence.json` under `superseded-stale-owner-body`. Two r5 findings are rejected here on the reasoning already recorded for the same class in the repair record: the demoted raw records still carry the runner-emitted `\"canonical\": true` field because classification is manifest-driven and this contract forbids overwriting historical raw records, and r8's own raw-versus-regrade missed drift is untouched because an audit-only record is outside the canonical drift check that its now-empty allowlist governed. One r5 finding is recorded and NOT closed here: `line-sha256` accepts a code token on an unreachable line, so a guard wrapped in `if false ... end` keeps its digest green — a resolver-design gap that belongs to its own round rather than to this repair slice |
230
+ | SUPERSEDES only the open `line-sha256` reachability finding in the preceding row: exact bytes plus a Ruby code token prove that a line exists as code, not that its body can run. Parse the unmodified Ruby source and reject a matching line when it sits under a literal-dead `if`, `elsif`, `unless`, or precondition `while` / `until` branch, or inside a lambda with no live direct `.call`; a positioned Ripper builder plus code-line span filling keeps tokenless sexp forms such as `return0`, `zsuper`, `yield0`, an empty array, and multiline statement starts in the dead-line set, while `begin ... end while/until` bodies remain eligible because Ruby executes them once. A call or reassignment under a dead branch cannot change that verdict; a call inside the same lambda, inside an uncalled method, or on an expression that merely contains the lambda cannot revive it. A direct or parenthesized-inline lambda call remains eligible, as does a lambda called inside a top-level method reached through finite direct bare method calls from a later live top-level bare call. This closes the registered literal-dead finding, not every Ruby reachability form: an ordinary line inside an uncalled method body remains eligible | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh | updated | `skill-extraction-workflow/SKILL.md` is the unchanged owner key. Before the repair, both a guard under `if false` and the same guard inside an assigned but uncalled lambda returned exit 0. The focused suite now retains live top-level code, literal postcondition-loop bodies, a literal-dead reassignment followed by a live call, direct and parenthesized-inline lambda calls, and a lambda reached through live direct top-level method calls. It rejects those two bypasses, literal-dead equivalents including `elsif false`, four valid tokenless sexp statements, and a multiline statement start, a lambda self-call, a call inside an uncalled method, a call under `if false`, and a lambda merely nested in another receiver expression. It reports `line-sha256 source line is not statically reachable` for inert code, `line-sha256 target Ruby did not parse` for matching Ruby with a nil AST or an error-bearing partial AST, and the existing absent-line reason when no matching code line exists. Seventeen applied mutations independently erase parser-event positions, erase dead-body span filling, disable literal truthiness, erase `elsif` classification, admit every lambda, count a dead call, treat every method as reachable, erase live-method reachability, erase transitive method-call propagation, ignore reassignment, count a dead reassignment, erase the postcondition exception, erase the unreachable diagnostic, recursively admit a nested receiver, erase parenthesized receiver unwrapping, erase the parse diagnostic, or ignore the partial-parse error state; each makes the focused suite exit 1 at its owning assertion. Evidence: `skill-extraction-workflow/scripts/register-firing-path-resolution.rb`; `skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh`; `specs/026-register-firing-reachability/plan.md`. The classifier is deliberately static and local: acceptance means no proven-dead carrier was found, not that production executed the line or that arbitrary dynamic dispatch was solved |
231
+ | SUPERSEDES only the catalog-fixture disposition in the earlier C3 completion-boundary row: branch-specific catalog mutation failures under an explicit source-HEAD base do not cover `check-ccl-skills.sh` default base discovery. Keep the existing catalog cases pinned to their exact synthetic base, and add a separate case whose local branch tracks a synthetic base commit B, whose `origin/main` decoy points at changed HEAD H, and whose checker process has `CCL_SKILL_BASE_REF` unset. The negative arm must suppress the changed-entrypoint-only warning under `origin/main=H`; the default arm must emit it through `@{upstream}=B`. This covers the default resolver without making the result depend on the caller repository's ancestry | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | Before the repair, a focused probe exited 1 because the suite had no checker invocation with `CCL_SKILL_BASE_REF` unset. Case c12 now asserts both synthetic merge bases, proves the `origin/main` decoy hides the committed `SKILL.md` delta, unsets the override, and requires `entrypoint_body_above_recommended_for_changed_file` from the real checker. The focused suite exited 0 in 75.4 seconds with `test_check_ccl_skill_catalog: ok`. Evidence: `skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
232
+ | EXTENDS the preceding catalog default-base row to cover the resolver's second branch: after proving `@{upstream}=B` wins over the `origin/main=H` decoy, remove the fixture branch's upstream, repoint `origin/main` to B, assert that `@{upstream}` is absent, and require the same changed-entrypoint warning with `CCL_SKILL_BASE_REF` still unset. Generate deterministic over-threshold fixture content inside the clone so neither arm depends on the current skill body's size | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | The current focused suite exited 0 with `test_check_ccl_skill_catalog: ok`. The case now has an explicit wrong-ref negative control plus upstream-present and upstream-absent positive arms, all on refs and ancestry created inside the case clone. Evidence: `skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
233
+ | SUPERSEDES only the fallback-ref uniqueness claim in the preceding catalog row: removing the configured upstream does not make `origin/main` the only discoverable branch at B because the clone still carries local and remote refs at the same object. Before the fallback arm, delete every ref that points at B except `refs/remotes/origin/main`, assert that exact singleton with `for-each-ref --points-at`, then run the unset-override checker | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | Tracked Claude review r2 found that a resolver scanning another B-valued ref could make the fallback arm falsely green. The repair removes `catalog-fixture-base`, `ccl-test-base`, clone remote refs, and any other B-valued ref through a ref inventory instead of deleting only the named example. The current focused suite exited 0 with `test_check_ccl_skill_catalog: ok`; aggregate and dual-track reruns remain required because the candidate hash changed. Evidence: `skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
234
+ | SUPERSEDES the one-arm ref cleanup in the preceding catalog row: the upstream-present and upstream-absent positives share the same false-green class, so a single fixture helper must retain exactly the expected B-valued ref before either checker run. Preserve only `refs/heads/catalog-fixture-base` for the upstream arm and only `refs/remotes/origin/main` for the fallback arm, deleting every other local, remote, tag, or symbolic ref at B and asserting the singleton each time | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | Tracked Claude review r3 found that the upstream arm retained the same alternate-B false green that r2 found in the fallback arm. `retain_only_ref_at_oid` now owns both purges and assertions, preventing another arm-specific patch. The current focused suite exited 0 with `test_check_ccl_skill_catalog: ok`; aggregate and dual-track reruns remain required because the candidate hash changed. Evidence: `skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
235
+ | SUPERSEDES the positive-only default-base oracle and growth-based fixture in the preceding catalog rows: exercise upstream and `origin/main` fallback as a 2x2 empty/changed matrix with `CCL_SKILL_BASE_REF` unset in all four arms; require checker exit 0 before inspecting the warning; and make negative warning assertions safe under `set -e`. Generate a deterministic over-threshold body in a non-upstream skill while shrinking its baseline bytes, so the resolver signal is not coupled to current entrypoint size, impact-chain obligations, or size-debt growth | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | Tracked Claude review r4 found that positive-only arms could false-green when no base was resolved and everything was treated as changed. The first 2x2 run then exited 1 without diagnostics because the expected negative `grep -q` miss triggered `set -e`; wrapping it in `if` exposed a second false-green, where changed arms carried the target warning but the real checker was nonzero on owner impact-chain and size-growth gates. The final fixture synthesizes `agents-file-coverage-gate` above 5 KB but below its baseline size, and both warning helpers reject nonzero checker status. The focused suite now exits 0 with `test_check_ccl_skill_catalog: ok`. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
236
+ | SUPERSEDES only the warning-output transport in the preceding catalog row: under `set -o pipefail`, do not feed a long captured checker log to `grep -q` through a pipe, because an early match can close the reader and turn the still-writing `printf` into status 141. Search `CASE_OUTPUT` through a here-string in both the presence and absence helpers; preserve the 2x2 resolver oracle and its rc=0 requirement unchanged | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | Implementer closeout found the new positive helper had copied a pipeline shape that this repository already documents as SIGPIPE-sensitive. The exact candidate had passed, so this is preventive rather than a reproduced failure; the replacement removes the writer process from the oracle and keeps the warning predicate identical. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
237
+ | SUPERSEDES the annotated-ref handling and evidence classification in the preceding catalog rows: drive both ref deletion and its assertion from the same peeled `for-each-ref --points-at B` set, and create annotated plus nested tag aliases in the fixture so that path cannot regress unseen. Build a dedicated cataloged fixture skill in a synthetic baseline commit B, then change one same-length marker in H; no unrelated real skill's future size or existence may decide the oracle. Also correct the r2 and r3 alias rows: those reviews identified plausible false-greens but no mutation probe was run, so their evidence class is `semantic-control; observed-failure: no`, not RED-baseline; append this correction rather than rewriting historical rows | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r5 challenge reported that `%(objectname)` sees an annotated tag object while `--points-at B` peels it, and a first-hand temporary Git repo reproduced the mismatch for annotated and nested tags. The same challenge correctly identified the fixed 420-line rewrite's dependency on the live `agents-file-coverage-gate` body and an unchanged-source evidence gap. c12 now commits its own valid leaf skill plus catalog row at B, applies a byte-neutral state change at H, seeds both tag classes, and the plan quotes the exact unchanged upstream-before-origin/main resolver block. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
238
+ | EXTENDS the preceding catalog oracle's control independence: recreate named annotated and nested tag aliases before every one of the four arms, delete them through the peeled ref set, and assert their exact refs are absent outside the helper so a shared enumeration regression cannot self-certify. Insert the synthetic catalog row at one asserted existing leaf anchor instead of inheriting the meaning of the file tail, and pass the changed skill explicitly into each warning helper instead of reading mutable global state | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r6 challenge correctly identified three preventive hardening opportunities. The pre-r6 tag control already made a one-sided unpeeled deletion red in the first empty arm, so no missed failure was reproduced; per-arm seeding plus independent `show-ref --verify` assertions now makes the intended proof explicit in both empty and singleton-retention states. The dedicated catalog row is inserted immediately before the unique `agents-file-coverage-gate` leaf row, and both warning helpers take `label` and `skill` positional arguments under `set -u`. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
239
+ | EXTENDS the final catalog fixture's diagnostic separation: snapshot the peeled ref list before deleting so the fixture never mutates the ref store under a streaming `for-each-ref` producer, and calibrate the over-threshold warning once through an explicit B base before exercising default discovery. Keep the fixture bulk as ordinary body lines, not an HTML comment. Reject the r7 pseudoref finding for this concrete graph: B is created after the source fetch, same-commit `checkout -B` was reproduced without `ORIG_HEAD`, `FETCH_HEAD` still names source S rather than B, and the 2x2 negative arms already reject a resolver fixed on B | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r7 challenge identified two preventive diagnostic improvements and one incorrect concrete pseudoref premise. A temporary Git repository showed all named pseudorefs absent after same-commit `checkout -B`; Git also refused `update-ref -d FETCH_HEAD`, so speculative pseudoref deletion would add non-portable behavior without closing a present alias. The plan now states public gate success and interim R0 separately instead of using an unqualified green status. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
240
+ | SUPERSEDES the fixture's real-skill catalog anchor and inherited-Git-environment assumptions, and corrects the evidence class of the first catalog default-base row. Insert before the first row matching the catalog's generic leaf-row grammar instead of naming an unrelated skill; fail only when the catalog has no leaf row at all, which is already an invalid catalog class. Unset repository-routing Git variables inside every checker subprocess, and run c12's calibration plus four default arms under a deliberately hostile outer `GIT_DIR`, `GIT_WORK_TREE`, and `GIT_INDEX_FILE`. The original no-unset-invocation probe proved a coverage gap, not that the oracle rejected a mutated resolver, so read that first row as `semantic-control; observed-failure: no`; preserve it as history and append this correction | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r8 challenge correctly found the literal `agents-file-coverage-gate` catalog anchor and inherited Git environment as avoidable fixture couplings, and correctly separated uncovered-path evidence from a mutation RED. The current case now makes environment cleanup load-bearing on all five checker invocations. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
241
+ | SUPERSEDES only the hostile-Git wrapper's function-call assignment semantics: never rely on `VAR=value shell_function` to remain temporary across the supported Bash range when a leak would redirect later `git -C` ref mutations into the caller repository. Save whether each Git routing variable was set and its prior value, explicitly export the hostile values, invoke the checker helper through a conditional status capture, then restore or unset every variable before any subsequent fixture ref operation | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r9 release-depth review raised a P1 portability risk for Bash before 5.1. The current host did not reproduce a leak, but the consequence would cross the throwaway-clone boundary, so the wrapper now uses explicit save/export/call/restore semantics compatible with Bash 3.2 instead of trusting version-dependent temporary assignment behavior. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
242
+ | SUPERSEDES only the hostile index-path construction and rejects the r10 catalog-position premise. Derive the index as `<absolute-git-dir>/index` instead of requiring Git's newer `rev-parse --path-format=absolute --git-path` option. Keep the generic first-leaf insertion: `skill_bootstrap_leaf_in_region` is computed from names routed inside `agent-context/session-start.md`'s marked entry-routing region intersected with catalog names marked leaf; the row's byte position or heading inside `docs/SKILLS.md` is not an input to that verdict | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r10 review correctly found a Git 2.31 compatibility dependency in `--path-format=absolute`. Its other P2 assumed a bootstrap region exists in the catalog document, but the current checker source proves routing membership comes exclusively from `agent-context/session-start.md`; the plan now quotes that contract so future packet-only review can verify the rejection. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
243
+ | SUPERSEDES the hostile-context target and fixture-wide Git-routing assumptions in the preceding catalog rows. Clear caller routing variables at script entry before repository discovery or any fixture `git -C` operation, then reintroduce hostile values only around c12 checker calls. Build `GIT_DIR`, `GIT_WORK_TREE`, and `GIT_INDEX_FILE` entirely under `TEST_ROOT`; a checker regression may fail against scratch state but must never target the developer repository or its live index | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r11 review correctly found that the previous hostile variables named the developer repository and that only checker subprocesses neutralized inherited routing, leaving fixture setup and later ref mutations exposed to caller exports. The script now clears all supported repository/object routing variables before computing `REPO_ROOT`; c12 creates a disposable bare repo, worktree, and index under the existing trapped temp root for the load-bearing hostile wrapper. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
244
+ | SUPERSEDES only the helper-status implication in the earlier hostile-wrapper row. Capturing a shell function's return code is not checker-status containment when the function ends on successful `set -e`: store the checker rc in `CASE_STATUS`, explicitly return that same value, let the hostile wrapper restore Git routing before propagating it, and suppress only the caller's immediate `errexit` so the existing assertion can print captured output and adjudicate the same status | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r12 review correctly traced both c12 helper functions and found their old implicit return was always zero. The checker result was already preserved and asserted through `CASE_STATUS`; the repair aligns the wrapper return path with that source of truth without removing the diagnostic. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
245
+ | SUPERSEDES only the inherited ref-environment inventory in the catalog fixture: clear `GIT_NAMESPACE` at script entry and in every checker subshell because it changes ref visibility. Do not expand this narrow ancestry-isolation case into global/system Git configuration ownership. The hostile wrapper remains load-bearing because deleting the helper unset exposes its scratch `GIT_DIR`, worktree, and index; its restore-to-unset branch is the current entry-clean contract. Pseudorefs remain outside the production resolver, which names only `@{upstream}` and `origin/main` | `skill-extraction-workflow` | behavioral-evidence: semantic-control; observed-failure: no; firing-path: command:skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh | updated | r14 challenge's `GIT_NAMESPACE` subfinding is accepted. Its first P2 is rejected because the hostile values deliberately test the helper unset boundary: removing that cleanup redirects the checker to scratch Git state and makes the case red. Its second P2 repeats r7's rejected premise: the production resolver never scans arbitrary refs or pseudorefs, and the concrete graph has no B-valued pseudoref. `GIT_CONFIG_*` ownership is unchanged from c1-c11 and outside this round's ref-routing contract. Evidence: `skill-extraction-workflow/SKILL.md`; `skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh`; `specs/027-catalog-default-base-coverage/plan.md`. |
246
+ | SUPERSEDES the standing present-tense claim in the C3 landing rows that plain `make test` enforces the measured-carrier evidence binding: as of this round that enforcement is retired, because a mechanical gate must not outlive the claim it enforces. When a round withdraws or downgrades the measurement a gate pins — the record demoted to audit-only, its property relabelled `UNMEASURED`, or review showing the grader cannot detect the obligation being pinned — the SAME landing re-bases the gate onto a claim that survives or retires it; leaving the binding running collects a standing tax on future work for evidence the repository no longer asserts. Judge the ORACLE, not the subject: a check that matches a surface form, or a comparison with too few samples to separate its arms, measures nothing that a better subject would repair | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:Makefile#test | updated | No owner skill package changed this round, so no owner key is cited: the behavior change is `Makefile`'s `test` recipe, from which the C3 `--candidate` invocation is removed. Observed failure: the round-023 repair record had already withdrawn the tighten claim to a rubric property and the worktree claim to carrier sensitivity with preservation `UNMEASURED`, yet `OWNER_BODY_BINDINGS` kept recomputing the assembled `body_sha256` of four measured carriers at committed `HEAD` unconditionally and aborting on drift, so every committed edit to `tighten-doc/SKILL.md`, `worktree-isolation/SKILL.md`, `tighten-doc/references/session-vantage-leakage.md`, or `worktree-isolation/references/shared-branch-rebase.md` turned `make test` and `ci.yml`'s job red until an operator ran a provider refresh and relanded five pinned constants. Oracle evidence: grading is a hand-written Ruby regex contract, not a model — `claude-haiku-4-5` is the answering subject, `red-baseline-023-c3-preservation.rb` binds `grader:` to `tighten_grade`/`worktree_grade`; `worktree_grade`'s `right_zero_stop` matches the surface string and never reads the gloss that follows, so both r8 head rounds scored `pass=true, missed=[]` while glossing `mybr...$remote_oid`'s right column in mutually contradictory and both-incorrect ways. Power evidence: r9 ceiling base 2/2 = head 2/2, r10 floor base 0/2, r11 base 1/2 = head 1/2 — three arms at two rounds separate nothing at any model tier. What is NOT retired: `specs/023-agent-native-repo-borrowing/evidence/` stays frozen under its own agent contract, every raw record, regrade, manifest, and runner committed as of-that-round facts. Consequence stated plainly: with the recipe line gone, `make_test_candidate_invocation?` can no longer find its required invocation, so the contract is retired rather than merely unwired and is kept as the readable record of what was enforced. Stop-and-think friction on skill entrypoints survives without model dependency through `check-ccl-skills.sh`'s changed-entrypoint scan. Evidence: `specs/028-c3-preservation-gate-retirement/plan.md`; `make --no-print-directory -n test` contains no C3 invocation and the remaining recipe is unchanged |
247
+ | SUPERSEDES the tighten-doc C3 row's assertion that its canonical r9 record "proves preservation and carrier sensitivity". It does not, and the repair record shipped in the same round already said so: r9's head body is byte-identical to the head arms of r5 and r6, which scored differently under their own graders, and r9's grader was introduced after those results with no pre-registration of grader or task before measurement, so the 2/2 is a property of the r9 rubric and not independent proof that the shrunken body preserved its obligations. The worktree side of this pair was withdrawn to carrier sensitivity in its own round; the tighten side was left asserting preservation, so on an append-only ledger the contradiction against the shipped repair record stood uncorrected until now | `skill-extraction-workflow` | behavioral-evidence: not-required wording-only; observed-failure: no | routed | No owner package changed for this correction: the round's edit to this ledger is the ledger itself, which the impact-chain contract excludes from the owner surface. This row corrects ledger prose only and changes no executable behavior. As of this round both C3 preservation claims stand at carrier sensitivity under each run's own rubric, and preservation is `UNMEASURED` for both owners. The recurrence class is the one this ledger already named: a sentence written in the present tense about a mutable pin becomes a false claim the moment the pin moves, and on an append-only ledger it cannot be edited afterwards — so a withdrawal must supersede every row that carried the old claim, not only the row whose own arm was re-measured |
248
+ | Enforcement must not outlive its evidentiary claim, and the check that says so must fire on whoever KILLS the claim rather than on the machinery's author: the technical design gate's mechanism-operability check gains a second firing point, so a landing that withdraws a claim, restates it as unmeasured, or narrows it below what a mechanism enforces re-bases every mechanism predicated on it onto a claim that still holds, or retires it, in that same landing — enumerated from the claim's own record and from what actually invokes the mechanism; a retirement additionally walks the obligations the mechanism carried that never depended on the dead claim, since those do not die with it | `product-rd-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/product-rd-workflow/SKILL.md#and again when a landing withdraws or downgrades the evidentiary claim | updated | Owner key `product-rd-workflow/SKILL.md` (design-gate firing registry) with the rule body in `skills/product-rd-workflow/references/design-review-gate-mechanics.md`'s `Mechanism-operability check`. Observed failure: round 023 downgraded the C3 preservation claim to `UNMEASURED`, round 028 had to retire the gate two rounds later, and in between every edit to four unrelated skill carriers paid a provider refresh — the operability check existed and was well-written, but its only trigger read `whenever the design proposes new mechanical enforcement`, which the claim-withdrawing actor never matches. The preceding C3 retirement row states this rule as ledger prose; provenance is not executable guidance, which is why round 028 recorded the clause `pending` rather than landed. RED-baseline, applied mutations in a throwaway copy (never the live tree), each observed to red on its OWNING assertion with the unmutated copy green as control: deleting the re-base-or-retire sentence reds `claim liveness (destination rule)`; deleting the retirement-walk sentence reds `claim liveness (destination retirement walk)`; deleting the second trigger from the entrypoint bullet reds `claim liveness (entry signal+pointer)` as an absent firing phrase; and re-homing that trigger into its own pointerless bullet reds the SAME assertion with `its own bullet does not carry the load pointer` while the pre-existing `mechanism gate (entry signal+pointer)` assertion, which runs first, stays green — the differential that proves the new pin is not redundant with it. Entrypoint budget: the second trigger is funded by collapsing the leg-name restatement the reference already owned, so `check-size-budget.sh` reports base_body_words=10245 head_body_words=10244 with `entrypoint_word_budget_blocking_ok` and `entrypoint_size_blocking_ok` — no net growth on a historically over-limit entrypoint. Evidence: `specs/029-claim-liveness-operability-leg/plan.md` |
249
+ | The shared-skill instantiation of the operability check reaches the claim-liveness rule by pointer rather than restatement, and both halves of that rule are pinned on both sides by the deterministic fixture that already pins the relocated-rule family — destination body, retirement walk, entrypoint firing signal co-located with its load pointer, and the shared-skill pointer itself | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; changed surfaces under it are `references/dual-track-review-gate.md`'s design-time operability check (a pointer to the fifth firing point, whose actor is not the gate's author) and `scripts/test_ai_coding_implementation_gates.sh` (added assertions in the existing dual-side pin family, one per obligation the rule imposes; the set is described rather than counted, because a count in an append-only row goes stale on the next assertion and cannot be edited afterwards). Observed failure: the round-028 gate that outlived its claim WAS a shared-skill gate, so an agent who loads only the four-leg shared-skill instantiation would still not reach the rule; routing rather than duplicating keeps `product-rd-workflow` the single owner. Why the pins are split in two: the design-time assertions cannot notice the new rule's loss — all of them stay green with the claim-liveness text deleted — and `retire the gate` without the obligation walk is the exact failure that walk prevents, so a reworded retirement sentence would keep the first pin green while dropping the second. RED-baseline, applied mutation with differential attribution: rewording the sibling pointer reds `claim liveness (shared-skill instantiation pointer)` and nothing else, against a green unmutated control of the same copy; the copy's own green control also proves the harness read the mutated tree rather than the live one. Evidence: `specs/029-claim-liveness-operability-leg/plan.md` |
250
+ | Agent session persistence gains the model-visible accounting invariant in its honest form: every item in the sealed client-dispatched envelope carries exactly one durable accounting classification — reconstructable raw, reconstructable by immutable versioned reference with a digest and a policy-aligned store, or explicitly non-reconstructable with a reason — while the invariant's assertions audit record completeness and classification validity and never force raw persistence; credentials stay reference-resolved so a leaked secret is recorded as a typed exposure event, never the value and never a plain digest of it; and a digest that is the only residue of an item is keyed against a retained keyring, a plain digest allowed only beside raw content. The digest trigger is deliberately unconditional — keyed by whether raw content sits beside the digest, with no protected-class list to fall out of date (the round-029 enumeration lesson) | `llm-inference-integration` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/llm-inference-integration/references/agent-session-persistence.md#client-dispatched envelope must carry | updated | Owner key `llm-inference-integration/SKILL.md` with the rule body in `skills/llm-inference-integration/references/agent-session-persistence.md` (a persistence-policy block plus one Non-negotiables bullet; entrypoint untouched per the approved landing surface). Observed failure: three same-class review P1s on earlier drafts of this rule, recorded in `specs/023-agent-native-repo-borrowing/plan.md`'s dual-track chains — forced credential persistence (r4), mutable reference targets making "reconstructable" a false promise (r5), and enumerable plain digests of low-entropy values plus an unnamed security owner (r6) — which moved D1 out of Batch I under the pre-registered no-more-in-place-patching rule. This landing is the dedicated round that deferral required: the named security owner approved the five decision points on 2026-08-20 before implementation, and each historical failure maps to a named sentence now pinned. RED-baseline, applied mutations in a throwaway copy (never the live tree), unmutated control green before and after, each mutant observed to red on its OWNING assertion with differential attribution (the fixture exits at first failure, so every earlier assertion passed under the mutant): deleting or rewording the invariant line, the mutable-target sentence, the audits-never-forces sentence, the typed-exposure-event clause, the no-plain-digest-of-a-leaked-secret clause, the plain-digest-only-beside-raw clause, the keyed-residue clause, and the Non-negotiables bullet each red their own pin; a relocation of the bullet into another section also reds its section-bound pin. Evidence: `specs/030-d1-model-visible-accounting/plan.md` |
251
+ | The model-visible accounting clause is pinned by the deterministic fixture that already pins the relocated-rule and claim-liveness families — one section-bound assertion per obligation the clause imposes (the single-classification invariant, the immutable-reference criterion, audits-never-forces-persistence, the typed exposure event for a leaked secret, no plain digest of a leaked secret, plain-digest-only-beside-raw, keyed residue digests, and the Non-negotiables bullet), with the same stated limits as the claim-liveness family: the pins catch deletion, rewording-away, and relocation, not a weakening sentence added beside them — catching a deliberate weakening stays the dual-track review's job | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; changed surface under it is `scripts/test_ai_coding_implementation_gates.sh` (a new pin family in the existing fixture, one assertion per obligation; the set is described rather than counted, because a count in an append-only row goes stale on the next assertion and cannot be edited afterwards). Observed failure: same three review P1s as the row above — the failure class lived in shared-skill rule text, so its prevention pins belong in the shared fixture that already guards relocated rules. RED-baseline shared with the row above: the same applied-mutation set in the same throwaway copy reds each owning assertion against a green unmutated control, and the copy's own green control proves the harness read the mutated tree rather than the live one. Evidence: `specs/030-d1-model-visible-accounting/plan.md` |
252
+ | Blocked-verification remediation gains the sandbox-denial triage and controlled-escalation rule in its safe form: a failing required command is first classified from observed denial evidence — never a bare non-zero exit — as sandbox denial or genuine failure, with indeterminate and mixed/conflicting evidence failing closed as genuine and the denial required to be the sole proximate cause preventing completion; escalation is never a default action and never bypasses a genuine failure or a product sandbox under test; a controlled escalated re-run needs every condition to hold — the command reviewed-trusted this session by the operator side with trust bound to recorded content identity and the re-run executing the reviewed snapshot or verified-then-executed bytes, failing closed on mismatch (a repository under extraction is untrusted input, its entrypoints qualifying only after in-session content review), the narrowest blocked capability only with the command re-run unchanged once per approval, and host policy plus per-escalation user approval bound to this command, this capability, this session, never cached or standing — otherwise the item stays blocked with the normal remediation record. The order deliberately inverts the source rule's retry-before-diagnosing default, whose premise (a repo's own trusted commands) does not hold for a skill set facing arbitrary repositories | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/skill-extraction-workflow/references/source-to-skill-extraction.md#Sandbox-denial triage precedes any escalation | updated | Owner key `skill-extraction-workflow/SKILL.md` with the rule body in `references/source-to-skill-extraction.md` (two bullets in Blocked Verification And Source-Read Remediation; entrypoint untouched per approved landing surface D-6). Observed failure: the 023 review P1 (chain `023-plan-r10`) — the source rule generalized as written lets a repo-controlled command use a sandbox denial as a pretext to obtain host credentials or network — which moved E5 out of the landing batches under the pre-registered dedicated-spec rule. This landing is that dedicated round: the named security owner approved all six decision points in-session on 2026-08-21 before implementation. Review round 1 (organization gate, chain `031-e5-r1`, codex) P1 accepted — trust bound to the command string is a TOCTOU hole (the blocked first run or a concurrent actor rewrites what the unchanged command resolves to); amendment A1 binds trust to content identity; review round 2 (user-authorized after the autonomous budget checkpoint, chain `031-e5-ua1`, codex) found two further P1s, both accepted — verify-then-execute on a live path left a swap window, closed by amendment A2 (the re-run executes the reviewed snapshot or the exact verified bytes), and a pending final-wording approval must block the landing, so the amended wording's re-approval is decision row D-7 and gates push/MR; review round 3 (user-authorized, chain `031-e5-ua2`, codex) found the atomicity requirement bound only to the entrypoint, closed by amendment A3 (chain-wide: every repository-controlled executable or loadable code unit executes from the reviewed snapshot or is verified atomically at its point of use); review round 4 (user-authorized, chain `031-e5-ua3`, codex) falsified round 3's closure claim — sealing keyed to code left configuration, data, environment files, and symlink targets swappable — closed by amendment A4, predicated on effect rather than file kind: every repository-controlled input that can affect privileged behavior is sealed, an unsealable input refusing the escalation back to blocked (A1–A3 are its special cases); review round 5 (user-authorized, chain `031-e5-ua4`, codex) found "sealed" was a label while non-code consumption mechanics stayed code-only, closed by amendment A5: every behavior-affecting input consumes from the reviewed snapshot or from bytes opened, verified, and held stable through use; review round 6 (user-authorized, chain `031-e5-ua5`, codex) widened the origin — content fetched over the granted network/IPC/host channel is equally mutable — closed by amendment A6: the quantifier is every untrusted mutable input that can affect privileged behavior, unsealable remote inputs refusing the escalation; review round 7 (user-authorized, chain `031-e5-ua6`, codex) found approval bound only to command/capability/session could recycle a stale grant onto content re-reviewed after a mismatch, closed by amendment A7: approval also binds the reviewed content identity, a mismatch invalidates it, and re-reviewed content needs fresh user approval; review round 8 (user-authorized, chain `031-e5-ua7`, codex) found "verified" unanchored for granted-channel bytes, closed by amendment A8: verification anchors to an immutable expected identity reviewed before escalation and bound into the approval, no establishable pre-grant identity stays blocked, and the clause states the net invariant — privileged behavior under the grant derives only from operator-reviewed, approval-bound content and trusted host state; adversarial challenge round 1 (user-authorized, chain `031-e5-ua10`, codex) found mixed evidence could launder a genuine failure behind a deliberately-triggered denial, closed by amendment A9: the denial must be the sole proximate cause preventing completion, mixed or conflicting evidence classifying as genuine; adversarial challenge round 2 (user-authorized, chain `031-e5-ua11`, codex) found "observed denial" carried no provenance requirement so command-controlled stderr could fake one, closed by amendment A10: denial evidence comes from the host sandbox's own trusted enforcement or telemetry channel bound to the exact invocation and denied capability, command output alone never qualifying; adversarial challenge round 3 (user-authorized, chain `031-e5-ua12`, codex) found the denial record unbound to content identity so replaced content could ride the prior denial, closed by amendment A11: qualifying denial evidence comes from an unprivileged run of the same reviewed content identity the escalated re-run executes, a mismatch invalidating the denial record along with the approval; adversarial challenge round 4 (user-authorized, chain `031-e5-ua13`, codex) found approval bound a re-resolvable capability name, closed by amendment A12: the grant binds the canonical resolved capability identity under host policy, re-resolved and verified atomically at use, any resolution change invalidating approval and denial record alike; adversarial challenge round 5 (user-authorized, chain `031-e5-ua16`, codex) found the denied first run's completed side effects could be duplicated by the unchanged re-run, closed by amendment A13: conjunctive condition 5 requires prior external effects proven absent, rolled back, or contained before any re-run, unaccountable effects staying blocked; review round 16 (user-authorized, chain `031-e5-ua17`, codex) found containment alone left the duplicate to materialize at commit/export, closed by amendment A14: contained state is discarded or reset to its pre-run snapshot before the re-run or covered by a demonstrated idempotency or deduplication guarantee; adversarial challenge round 6 (user-authorized, chain `031-e5-ua18`, codex) found the effect-accounting evidence itself unprovenance, closed by amendment A15: accounting evidence carries the denial-evidence provenance discipline, bound to the invocation, content identity, and affected targets, command output alone never proving absence, rollback, or deduplication. RED-baseline: applied mutations in a throwaway copy (never the live tree), unmutated control green before and after, each mutant observed to red on its OWNING assertion with differential attribution (the fixture exits at first failure, so every earlier assertion passed under the mutant); a tree-isolation probe mutated the copy while the live tree's fixture stayed green, proving the harness read the mutated tree. Evidence: `specs/031-e5-controlled-privilege-escalation/plan.md` |
253
+ | The controlled-escalation clause is pinned by the deterministic fixture that already pins the relocated-rule, claim-liveness, and model-visible accounting families — one section-bound assertion per obligation the clause imposes (triage before escalation, observed-denial evidence with host-channel provenance and the command-output-never-qualifies rule, fail-closed indeterminate and mixed evidence with the sole-proximate-cause requirement, no bypass of genuine failures, the conjunctive condition gate, in-session command trust, untrusted repo entrypoints, grant minimality, no general sandbox disable, the grant binding the canonical resolved capability identity with re-resolution-at-use mismatch handling, the single unchanged re-run per approval, content-identity binding with re-verification at the escalated re-run and fail-closed mismatch handling, snapshot-or-verified-bytes execution closing the verify-to-execute swap window chain-wide for every unit in the runtime chain, effect-keyed sealing of every untrusted mutable input that can affect privileged behavior — granted-channel content included, verification anchored to a pre-reviewed approval-bound expected identity, and the net derives-only-from-reviewed-content invariant — with unsealable inputs refusing the escalation, approval binding including the reviewed content identity with mismatch invalidation and fresh approval after re-review, no cached or standing approval, the untouchable product sandbox under test, denial evidence on every outcome, and the condition-5 accounting of the first run's external effects before any re-run), with the same stated limits as those families: the pins catch deletion, rewording-away, and relocation, not a weakening sentence added beside them — catching a deliberate weakening stays the dual-track review's job | `skill-extraction-workflow` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: command:skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh | updated | Owner key `skill-extraction-workflow/SKILL.md`; changed surface under it is `scripts/test_ai_coding_implementation_gates.sh` (pin family 8 in the existing fixture, one section-bound assertion per obligation; the set is described rather than counted, because a count in an append-only row goes stale on the next assertion and cannot be edited afterwards). Observed failure: same 023 review P1 as the row above — the failure class lives in shared-skill rule text, so its prevention pins belong in the shared fixture that already guards relocated rules. RED-baseline shared with the row above: the same applied-mutation set in the same throwaway copy reds each owning assertion against a green unmutated control, and the copy-vs-live isolation probe proves the harness read the mutated tree rather than the live one. Evidence: `specs/031-e5-controlled-privilege-escalation/plan.md` |
254
+ | A no-owner deliverable doc is drafted from a locked charter plus user- or session-supplied substance — the drafting mode routes to an owner while substance is unsettled and never invents substantive decisions, mid-stream asks classify against the charter instead of appending reactively, and a loaded platform/tool skill is not the deliverable owner | `tighten-doc` | behavioral-evidence: RED-baseline; observed-failure: yes; firing-path: file:skills/tighten-doc/SKILL.md#Route to an owner when substance is unsettled | updated | `tighten-doc/SKILL.md` is the owner key (description no-owner-draft trigger + Draft-mode contract); charter method in the doc-charter-first reference of the same package; the tool-skill-masking sibling variant landed in this round's extraction-workflow reference, whose owner is covered by the process-controls rows above. Observed failure: a multi-round deliverable-doc effort where the finalization owner never fired until the user asked and the doc accreted reactively without a charter. RED-baseline: the cold-start routing miss frozen as an eval fixture in `eval/routing-tasks.jsonl`, exercised by the Tier-1 routing analyzer on the landing candidate. Corrective insertion: the round that landed this change merged with the impact-chain gate red and no row; this row was inserted by a corrective rewrite of the integrating merge per the round-023 precedent, authored post-merge from that round's recorded PR evidence and diff. |