@ccoalm/ccl-skills 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (561) hide show
  1. package/LICENSE +201 -0
  2. package/README.md +49 -0
  3. package/dist/assets/marketplace/.agents/plugins/marketplace.json +12 -0
  4. package/dist/assets/marketplace/.claude-plugin/marketplace.json +13 -0
  5. package/dist/assets/marketplace/marketplace-manifest.json +12 -0
  6. package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/marketplace.json +16 -0
  7. package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/plugin.json +5 -0
  8. package/dist/assets/marketplace/plugins/ccl-skills/.codex-plugin/plugin.json +5 -0
  9. package/dist/assets/marketplace/plugins/ccl-skills/.worktree-only +3 -0
  10. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +45 -0
  11. package/dist/assets/marketplace/plugins/ccl-skills/agent-context/subagent-start.md +12 -0
  12. package/dist/assets/marketplace/plugins/ccl-skills/hooks/AGENTS.md +19 -0
  13. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-delegation-owner.sh +125 -0
  14. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-edit-isolation.sh +102 -0
  15. package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +1156 -0
  16. package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +131 -0
  17. package/dist/assets/marketplace/plugins/ccl-skills/hooks/merge-authorization-prompt.sh +142 -0
  18. package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-guard.sh +12 -0
  19. package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-stop.sh +13 -0
  20. package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +144 -0
  21. package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-context.sh +87 -0
  22. package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-start.sh +86 -0
  23. package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-extraction-gate-stop.sh +69 -0
  24. package/dist/assets/marketplace/plugins/ccl-skills/hooks/subagent-start.sh +26 -0
  25. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_delegation_owner.sh +329 -0
  26. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_edit_isolation.sh +322 -0
  27. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +902 -0
  28. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +178 -0
  29. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +121 -0
  30. package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_session_start.sh +170 -0
  31. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/AGENTS.md +17 -0
  32. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +564 -0
  33. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +14 -0
  34. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-update-skills.md +44 -0
  35. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-verify-skills.md +109 -0
  36. package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-worktree-check.md +36 -0
  37. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/AGENTS.md +28 -0
  38. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/README.md +276 -0
  39. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.example.json +10 -0
  40. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +1307 -0
  41. package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +941 -0
  42. package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/SKILL.md +45 -0
  43. package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/agents/openai.yaml +4 -0
  44. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +188 -0
  45. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/agents/openai.yaml +4 -0
  46. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/android-dev.md +92 -0
  47. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/flutter-dev.md +80 -0
  48. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/ios-dev.md +72 -0
  49. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/kotlin-multiplatform.md +93 -0
  50. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-platform-boundaries.md +77 -0
  51. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +77 -0
  52. package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/source-evidence-map.md +64 -0
  53. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +353 -0
  54. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/agents/openai.yaml +4 -0
  55. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +419 -0
  56. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +126 -0
  57. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +197 -0
  58. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +179 -0
  59. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +98 -0
  60. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +93 -0
  61. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_timeout_exit.sh +15 -0
  62. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +1438 -0
  63. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +324 -0
  64. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/concern_excerpt.py +295 -0
  65. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/egress_schema.py +214 -0
  66. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +642 -0
  67. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +181 -0
  68. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +1165 -0
  69. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +1190 -0
  70. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +946 -0
  71. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_opencode_review.py +474 -0
  72. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +1899 -0
  73. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_review_json.py +200 -0
  74. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +2845 -0
  75. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.sh +6 -0
  76. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/run_claude_capture.py +71 -0
  77. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +53 -0
  78. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +68 -0
  79. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +2311 -0
  80. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +1832 -0
  81. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_code_review_identity.sh +73 -0
  82. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_concern_excerpt.sh +245 -0
  83. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_egress_schema.sh +177 -0
  84. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +272 -0
  85. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +195 -0
  86. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_concurrency.sh +120 -0
  87. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_retry.sh +1005 -0
  88. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_opencode_review.sh +258 -0
  89. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +574 -0
  90. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +349 -0
  91. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +434 -0
  92. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +264 -0
  93. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +2412 -0
  94. package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/verify_native_skill_binding.py +123 -0
  95. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +153 -0
  96. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/agents/openai.yaml +4 -0
  97. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +54 -0
  98. package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/prevention-routing.md +36 -0
  99. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +69 -0
  100. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/agents/openai.yaml +4 -0
  101. package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/references/security-review-gate.md +41 -0
  102. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +165 -0
  103. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/agents/openai.yaml +4 -0
  104. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/api-security-boundaries.md +47 -0
  105. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +160 -0
  106. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/artifact-generation-architecture.md +37 -0
  107. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/audit-history-architecture.md +29 -0
  108. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/bulk-workflow-architecture.md +33 -0
  109. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/config-rule-routing-architecture.md +34 -0
  110. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +72 -0
  111. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-modeling-and-migrations.md +79 -0
  112. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +210 -0
  113. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/dependency-platform.md +105 -0
  114. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/developer-tooling-architecture.md +38 -0
  115. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/error-contract-architecture.md +36 -0
  116. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +260 -0
  117. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/http-gateway-architecture.md +74 -0
  118. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/mq-consumer-architecture.md +38 -0
  119. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +275 -0
  120. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/notification-architecture.md +25 -0
  121. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/ops-checklist.md +57 -0
  122. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/performance-capacity-architecture.md +38 -0
  123. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/protobuf-contract-architecture.md +119 -0
  124. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/redis-cache-coordination.md +93 -0
  125. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/release-runtime-readiness.md +65 -0
  126. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/replay-comparison-architecture.md +26 -0
  127. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/runtime-observability.md +94 -0
  128. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/service-scaffold.md +76 -0
  129. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/source-evidence-map.md +55 -0
  130. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/workflow-state-architecture.md +38 -0
  131. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +159 -0
  132. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/agents/openai.yaml +4 -0
  133. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/artifact-generation-patterns.md +37 -0
  134. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/audit-history-patterns.md +28 -0
  135. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/bulk-import-export-patterns.md +56 -0
  136. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/config-rule-routing-patterns.md +38 -0
  137. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/data-access-patterns.md +55 -0
  138. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/db-schema-and-dal-patterns.md +109 -0
  139. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/dependency-client-patterns.md +130 -0
  140. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/developer-tooling-patterns.md +70 -0
  141. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/domain-feature-patterns.md +78 -0
  142. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/engineering-patterns.md +119 -0
  143. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/error-contract-patterns.md +55 -0
  144. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/feature-playbook.md +61 -0
  145. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/http-gateway-client-patterns.md +76 -0
  146. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/mq-consumer-patterns.md +55 -0
  147. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/notification-patterns.md +42 -0
  148. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/observability-implementation-patterns.md +101 -0
  149. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/performance-capacity-patterns.md +44 -0
  150. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/protobuf-contract-patterns.md +72 -0
  151. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/public-api-integration-patterns.md +56 -0
  152. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/quality-and-testing-patterns.md +91 -0
  153. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/redis-cache-lock-patterns.md +123 -0
  154. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/release-ops-patterns.md +112 -0
  155. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/reliability-patterns.md +83 -0
  156. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/replay-comparison-patterns.md +32 -0
  157. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/scaffold-and-codegen.md +86 -0
  158. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/source-evidence-map.md +54 -0
  159. package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +45 -0
  160. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +80 -0
  161. package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/agents/openai.yaml +4 -0
  162. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +117 -0
  163. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/agents/openai.yaml +4 -0
  164. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-approval-auto-reviewer.md +106 -0
  165. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +441 -0
  166. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-context-freshness.md +47 -0
  167. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-credentials-auth.md +13 -0
  168. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-extensions-skills.md +13 -0
  169. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-file-edit-protocol.md +129 -0
  170. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-ide-integration.md +5 -0
  171. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-input-ingestion.md +13 -0
  172. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-instruction-composition.md +13 -0
  173. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-lifecycle-hooks.md +92 -0
  174. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-messaging.md +5 -0
  175. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-runtime-bootstrap.md +5 -0
  176. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +448 -0
  177. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-task-orchestration.md +13 -0
  178. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +123 -0
  179. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-turn-lifecycle.md +131 -0
  180. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +162 -0
  181. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +156 -0
  182. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +146 -0
  183. package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/retrieval-agent-safety.md +273 -0
  184. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +202 -0
  185. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/agents/openai.yaml +4 -0
  186. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +62 -0
  187. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/cross-stack-alignment.md +94 -0
  188. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/framework-choice.md +76 -0
  189. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/online-practice-uptake.md +56 -0
  190. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/platform-capabilities.md +91 -0
  191. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/product-page-checklist.md +40 -0
  192. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/qa-release.md +72 -0
  193. package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/source-evidence-map.md +82 -0
  194. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +103 -0
  195. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/agents/openai.yaml +5 -0
  196. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +100 -0
  197. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +70 -0
  198. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/agents/openai.yaml +4 -0
  199. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +549 -0
  200. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +97 -0
  201. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/research-prompts.md +66 -0
  202. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/AGENTS.md +32 -0
  203. package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/test-public-data-acquisition-recipes.sh +379 -0
  204. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +244 -0
  205. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/agents/openai.yaml +4 -0
  206. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +76 -0
  207. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/framework-middleware-checklist.md +142 -0
  208. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/infra-component-deployment.md +268 -0
  209. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-correlation-recipe.md +124 -0
  210. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-schema-canonical.md +208 -0
  211. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +105 -0
  212. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/obs-stack-architecture.md +107 -0
  213. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +95 -0
  214. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +11 -0
  215. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +303 -0
  216. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/agents/openai.yaml +4 -0
  217. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +163 -0
  218. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/config-center-via-etcd.md +245 -0
  219. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/custom-control-plane-boundary.md +298 -0
  220. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-cli-concrete-recipe.md +312 -0
  221. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-pipeline.md +165 -0
  222. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/env-and-lane-matrix.md +126 -0
  223. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/lane-orchestration-control-plane.md +383 -0
  224. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/multi-region-and-cluster.md +135 -0
  225. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +149 -0
  226. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/python-package-registry-release.md +462 -0
  227. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/rollback-playbook.md +123 -0
  228. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +231 -0
  229. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/version-authority-and-deprecation.md +21 -0
  230. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +276 -0
  231. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/agents/openai.yaml +4 -0
  232. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +127 -0
  233. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/framework-middleware.md +143 -0
  234. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +90 -0
  235. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/http-response-envelope-contract.md +24 -0
  236. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +127 -0
  237. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/multi-env-routing.md +192 -0
  238. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/protobuf-http-contract-signals.md +64 -0
  239. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +124 -0
  240. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/rpc-framework-recipe.md +494 -0
  241. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-choice.md +113 -0
  242. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-migration-playbook.md +231 -0
  243. package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +131 -0
  244. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +235 -0
  245. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/agents/openai.yaml +4 -0
  246. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/adr-convention.md +146 -0
  247. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-checklist.md +30 -0
  248. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-evaluation-report-template.md +25 -0
  249. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-execution-spec.md +108 -0
  250. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-sop.md +457 -0
  251. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-templates.md +24 -0
  252. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/artifact-egress-confidentiality.md +58 -0
  253. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +86 -0
  254. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/cross-repo-coordination.md +46 -0
  255. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +192 -0
  256. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +62 -0
  257. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +45 -0
  258. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/diagnostic-spec-match-gate.md +36 -0
  259. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +35 -0
  260. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dormant-code-activation.md +47 -0
  261. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/existing-project-assessment-report.md +223 -0
  262. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/external-skill-augmentation.md +46 -0
  263. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/feature-deprecation-cascade.md +15 -0
  264. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/high-risk-resilience-gates.md +73 -0
  265. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-completeness-and-minimality.md +120 -0
  266. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-entry-reentry-gate.md +122 -0
  267. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/modular-monolith-heuristic.md +105 -0
  268. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +115 -0
  269. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +62 -0
  270. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-attributes.md +112 -0
  271. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-remediation-program.md +88 -0
  272. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +27 -0
  273. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +52 -0
  274. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/review-reception.md +34 -0
  275. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/shared-gate-artifact-classification.md +76 -0
  276. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/source-evidence-map.md +31 -0
  277. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/status-tracker-sync.md +77 -0
  278. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/sync-spec-repo-contract.md +25 -0
  279. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +34 -0
  280. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/worktree-mechanics.md +55 -0
  281. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/AGENTS.md +18 -0
  282. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/check-agent-contract-coverage.sh +213 -0
  283. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +136 -0
  284. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/agents/openai.yaml +9 -0
  285. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/analytics-visualization-interactions.md +206 -0
  286. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +108 -0
  287. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/complex-creation-interactions.md +194 -0
  288. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +214 -0
  289. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +53 -0
  290. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +129 -0
  291. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +97 -0
  292. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +79 -0
  293. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +63 -0
  294. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +146 -0
  295. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +250 -0
  296. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +237 -0
  297. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +65 -0
  298. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +237 -0
  299. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +324 -0
  300. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-web-desktop-patterns.md +456 -0
  301. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +114 -0
  302. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +79 -0
  303. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/resource-management-interactions.md +113 -0
  304. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/scenario-community-patterns.md +133 -0
  305. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +130 -0
  306. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +47 -0
  307. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/trust-sensitive-ai-and-data-patterns.md +96 -0
  308. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +106 -0
  309. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +176 -0
  310. package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +111 -0
  311. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +157 -0
  312. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/agents/openai.yaml +4 -0
  313. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/ai-service-integration-boundaries.md +57 -0
  314. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-contract-and-schema.md +62 -0
  315. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-security-boundaries.md +39 -0
  316. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +46 -0
  317. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/async-execution-model.md +24 -0
  318. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/background-jobs-and-scheduling.md +18 -0
  319. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/batch-and-pipeline-architecture.md +11 -0
  320. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/config-secrets-runtime.md +22 -0
  321. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-modeling-and-migrations.md +64 -0
  322. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +211 -0
  323. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +263 -0
  324. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +281 -0
  325. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/observability-and-ops.md +26 -0
  326. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +20 -0
  327. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/redis-cache-coordination.md +41 -0
  328. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/reliability-and-error-contract.md +17 -0
  329. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/source-evidence-map.md +55 -0
  330. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/web-framework-boundaries.md +26 -0
  331. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +143 -0
  332. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/agents/openai.yaml +4 -0
  333. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +16 -0
  334. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/async-and-worker-patterns.md +24 -0
  335. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +18 -0
  336. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +13 -0
  337. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/dependency-client-patterns.md +39 -0
  338. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/error-handling-patterns.md +26 -0
  339. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/feature-playbook.md +43 -0
  340. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/observability-implementation-patterns.md +31 -0
  341. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/project-structure-and-tooling.md +24 -0
  342. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +52 -0
  343. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/redis-cache-lock-patterns.md +78 -0
  344. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/schema-and-validation-patterns.md +23 -0
  345. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/source-evidence-map.md +56 -0
  346. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/sqlalchemy-and-migrations-patterns.md +99 -0
  347. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +61 -0
  348. package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/web-framework-patterns.md +35 -0
  349. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +91 -0
  350. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/agents/openai.yaml +4 -0
  351. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/config-runtime-readback.md +20 -0
  352. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/mr-merge-authorization.md +31 -0
  353. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/post-release-env-reset.md +31 -0
  354. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-closeout-evidence.md +20 -0
  355. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-scope-confirmation.md +21 -0
  356. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +20 -0
  357. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/test-scope-prompt.md +24 -0
  358. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/watcher-discipline.md +14 -0
  359. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/SKILL.md +64 -0
  360. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/agents/openai.yaml +4 -0
  361. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/comment-safe-release-doc.md +19 -0
  362. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-evidence-workflow.md +23 -0
  363. package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-testing-scope-section.md +15 -0
  364. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +87 -0
  365. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/agents/openai.yaml +4 -0
  366. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +130 -0
  367. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/agents/openai.yaml +4 -0
  368. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/prd-composition-contract.md +35 -0
  369. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/requirement-closure-contract.md +86 -0
  370. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/security-four-questions.md +38 -0
  371. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/SKILL.md +91 -0
  372. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/agents/openai.yaml +4 -0
  373. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +88 -0
  374. package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/agents/openai.yaml +4 -0
  375. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +337 -0
  376. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/agents/openai.yaml +4 -0
  377. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/analysis-parse-fix-test-challenge-replay.md +47 -0
  378. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attribution-verification.md +69 -0
  379. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/bootstrap-slim-c3-obligation-table.md +112 -0
  380. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
  381. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +162 -0
  382. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +507 -0
  383. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +86 -0
  384. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/evidence-card-template.md +51 -0
  385. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/example-domain-preselect.md +79 -0
  386. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +57 -0
  387. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-lifecycle-handoff.md +65 -0
  388. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +194 -0
  389. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +75 -0
  390. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +286 -0
  391. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +190 -0
  392. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/l0-l1-l2-routing.md +114 -0
  393. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/online-skill-review.md +47 -0
  394. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +164 -0
  395. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +90 -0
  396. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +320 -0
  397. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +16 -0
  398. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-feedback-mining.md +33 -0
  399. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-finding-standards.md +57 -0
  400. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-rubric.md +40 -0
  401. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +118 -0
  402. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/skill-listing-budget.md +19 -0
  403. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +254 -0
  404. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +658 -0
  405. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/two-source-extraction-pattern.md +167 -0
  406. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +179 -0
  407. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-routing-map.md +51 -0
  408. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +180 -0
  409. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/AGENTS.md +18 -0
  410. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +1452 -0
  411. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-evidence-card-leak.sh +491 -0
  412. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-mr-target-freshness.sh +173 -0
  413. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +488 -0
  414. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +419 -0
  415. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +197 -0
  416. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +327 -0
  417. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +401 -0
  418. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing.rb +248 -0
  419. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/generic-r0-leak-scan.sh +282 -0
  420. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/governing-chain-diff.py +321 -0
  421. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +964 -0
  422. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +708 -0
  423. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +540 -0
  424. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-lifecycle.rb +51 -0
  425. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-pending-status.rb +55 -0
  426. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +829 -0
  427. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +1203 -0
  428. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_r0_status.sh +75 -0
  429. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_register_pending_exclusion.sh +137 -0
  430. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +173 -0
  431. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_route_drift.sh +377 -0
  432. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +833 -0
  433. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +491 -0
  434. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh +114 -0
  435. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_mr_target_freshness.sh +261 -0
  436. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +538 -0
  437. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +154 -0
  438. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +190 -0
  439. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_surface_binding.sh +178 -0
  440. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_prose_target.sh +86 -0
  441. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_generic_r0_leak_scan.sh +131 -0
  442. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_git_identity_predicate_gate.sh +243 -0
  443. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_governing_chain_diff.sh +419 -0
  444. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +120 -0
  445. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +724 -0
  446. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +414 -0
  447. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_registration.sh +34 -0
  448. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +205 -0
  449. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +194 -0
  450. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_credential_cwd.sh +61 -0
  451. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +111 -0
  452. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +53 -0
  453. package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +257 -0
  454. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +98 -0
  455. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/agents/openai.yaml +4 -0
  456. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/input-state-machines.md +36 -0
  457. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/streaming-rich-output.md +130 -0
  458. package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/terminal-side-channels.md +96 -0
  459. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +408 -0
  460. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/agents/openai.yaml +4 -0
  461. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/AGENTS.md +18 -0
  462. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/bitable-setup.md +573 -0
  463. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/README.md +120 -0
  464. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/github-actions.yml +119 -0
  465. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/gitlab-ci.yml +76 -0
  466. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/jenkins.Jenkinsfile +106 -0
  467. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +279 -0
  468. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/gen_report.py +2807 -0
  469. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/makefile-template.md +200 -0
  470. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/report-config-schema.md +272 -0
  471. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/run_pytestless.py +475 -0
  472. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/source-to-case-workflows.md +258 -0
  473. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-marker-conventions.md +316 -0
  474. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +145 -0
  475. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/AGENTS.md +16 -0
  476. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.dart +129 -0
  477. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.go +197 -0
  478. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.py +135 -0
  479. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.ts +285 -0
  480. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/test_gen_report.py +2144 -0
  481. package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +62 -0
  482. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +212 -0
  483. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/agents/openai.yaml +4 -0
  484. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +75 -0
  485. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +50 -0
  486. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/data-and-workflow-testing.md +34 -0
  487. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/design-closed-contract-oracles.md +31 -0
  488. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +71 -0
  489. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +240 -0
  490. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +235 -0
  491. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/non-functional-specialized-scenarios.md +296 -0
  492. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/rd-testing-standard-template.md +126 -0
  493. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/run-killing-mutation-walk.md +43 -0
  494. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +136 -0
  495. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/source-evidence-map.md +59 -0
  496. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/structured-tc-input-translation.md +67 -0
  497. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +392 -0
  498. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-data-and-determinism.md +39 -0
  499. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +92 -0
  500. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/unit-testing.md +46 -0
  501. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/vendored-contract-drift-checklist.md +64 -0
  502. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/verify-enforcement-mechanisms.md +18 -0
  503. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/AGENTS.md +17 -0
  504. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.py +140 -0
  505. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.test.sh +75 -0
  506. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.py +170 -0
  507. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.test.sh +87 -0
  508. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.go +198 -0
  509. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.test.sh +109 -0
  510. package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/test_mutation_backup_recipe.sh +237 -0
  511. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +184 -0
  512. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/agents/openai.yaml +4 -0
  513. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/comment-safe-feishu.md +93 -0
  514. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/cross-model-co-review.md +3 -0
  515. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +60 -0
  516. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +17 -0
  517. package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/session-vantage-leakage.md +58 -0
  518. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +126 -0
  519. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/agents/openai.yaml +4 -0
  520. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +47 -0
  521. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/embedded-h5-in-host.md +87 -0
  522. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +194 -0
  523. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/source-evidence-map.md +60 -0
  524. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +190 -0
  525. package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +83 -0
  526. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/SKILL.md +179 -0
  527. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/agents/openai.yaml +4 -0
  528. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/shared-branch-rebase.md +25 -0
  529. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/AGENTS.md +23 -0
  530. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_status.sh +207 -0
  531. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_sweep.sh +481 -0
  532. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-status.sh +325 -0
  533. package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-sweep.sh +245 -0
  534. package/dist/assets/release.json +2797 -0
  535. package/dist/claude-adapter.d.ts +9 -0
  536. package/dist/claude-adapter.js +240 -0
  537. package/dist/cli-worker.d.ts +1 -0
  538. package/dist/cli-worker.js +32 -0
  539. package/dist/cli.d.ts +22 -0
  540. package/dist/cli.js +214 -0
  541. package/dist/codex-host.d.ts +30 -0
  542. package/dist/codex-host.js +162 -0
  543. package/dist/fs-safe.d.ts +21 -0
  544. package/dist/fs-safe.js +241 -0
  545. package/dist/index.d.ts +2 -0
  546. package/dist/index.js +1 -0
  547. package/dist/manifest.d.ts +8 -0
  548. package/dist/manifest.js +135 -0
  549. package/dist/opencode-adapter.d.ts +10 -0
  550. package/dist/opencode-adapter.js +416 -0
  551. package/dist/operations.d.ts +3 -0
  552. package/dist/operations.js +956 -0
  553. package/dist/paths.d.ts +20 -0
  554. package/dist/paths.js +4 -0
  555. package/dist/types.d.ts +58 -0
  556. package/dist/types.js +1 -0
  557. package/dist/unified.d.ts +4 -0
  558. package/dist/unified.js +64 -0
  559. package/dist/version.d.ts +2 -0
  560. package/dist/version.js +5 -0
  561. package/package.json +35 -0
@@ -0,0 +1,156 @@
1
+ # LLM Client Gateway
2
+
3
+ ## Boundary Shape
4
+
5
+ - Put all LLM calls behind a small gateway/client, not inside handlers or domain services.
6
+ - Accept typed messages, model selection, response mode, timeout, trace/request id, and optional tool schema.
7
+ - Return a structured response with content, reasoning/think content when available, tool calls, token usage, finish reason, raw provider metadata when allowed, and classified error.
8
+ - Keep provider credentials, URLs, and headers inside provider adapters or secret-backed config.
9
+
10
+ ## Provider Adapter Rules
11
+
12
+ - Normalize each provider behind one internal request/response contract.
13
+ - Key adapters by wire protocol, not by brand: providers speaking the same OpenAI-compatible protocol are configuration variants (base URL, model list, credentials) of one adapter, not new adapters; add an adapter only for a genuinely different wire protocol (e.g. Anthropic-native). Otherwise adapter count grows linearly with every onboarded brand. Shared transport does NOT mean shared policy: each provider still gets its own policy profile — error-class mapping locked with per-provider tests, capability/tool/safety flags, usage-accounting quirks — and "compatible" is a claim to verify per provider, not assume; when a provider's capability, auth, safety, streaming, or accounting semantics genuinely diverge, split the adapter rather than stretching config.
14
+ - Do not mix incompatible API formats in the same fallback chain unless an adapter converts payloads and response semantics explicitly.
15
+ - Maintain a per-provider/model parameter policy:
16
+ - remove unsupported streaming options;
17
+ - remove unsupported sampling parameters for reasoning models;
18
+ - preserve max output tokens or equivalent caps;
19
+ - validate response-format support before using strict JSON mode.
20
+ - Treat silent fallback as dangerous: record the failed provider, error class, fallback provider, and final outcome.
21
+
22
+ ## Streaming
23
+
24
+ - Use a single stream event envelope for content deltas, reasoning deltas, tool-call deltas, usage summary, completion, and error.
25
+ - Parse SSE defensively: ignore empty lines, handle `[DONE]`, tolerate non-JSON keepalives, and surface malformed chunks as telemetry.
26
+ - **A streaming response needs an inter-event idle timeout, separate from the total/request timeout.** The total deadline (below) bounds the *whole* call, but a stream can connect, emit a few tokens, then silently stall with the socket open and no further bytes — a wedged upstream or a dropped network path that never sends a close. The total timeout does not catch this *promptly* (it may be minutes long to allow legitimate long generations), and "stream closed before the terminal event" handling never fires because it never closes; without an inter-event idle timeout the turn is pinned until the total deadline — or forever if no total deadline is enforced. Wrap each read for the next event in an idle deadline; on expiry, cancel and treat it as a retryable stream failure (per the gateway's retry classification), not a hang. Get the details right or it misfires:
27
+ - **Reset on meaningful progress, not on any byte.** If you reset the idle timer on every SSE line, periodic comments/keepalives/heartbeats keep it alive forever while no actual content or terminal event arrives — a dead generation masked by liveness pings. Reset only on real progress events (content deltas, tool-call deltas, the terminal event); treat keepalives as "connection alive" but track a separate no-content-progress budget.
28
+ - **Size it well above normal token cadence, and make it configurable per model/operation.** Reasoning- or tool-heavy models legitimately pause tens of seconds (sometimes >60s) between tokens; a tight fixed idle timeout aborts valid work. Default generously and allow per-model/per-method override, separate from any time-to-first-token budget.
29
+ - **Terminal success wins the race atomically.** The terminal event can arrive as the idle timer fires; serialize the read-vs-timer completion so a completed stream is never cancelled-and-retried after success was already in hand (a false retry that duplicates the call).
30
+ - **One idempotent cancellation path; cancel the stream, not the connection; await cleanup before retry.** The idle timeout, the caller/total deadline, and a consumer disconnect can all abort the same stream concurrently. Route every abort through a single idempotent cancel that awaits reader cleanup before any retry, so two cancel paths don't race or leak a half-closed reader. Cancel the specific **request/stream context** and close *that* response body with a deadline — do not unconditionally *drain* it on an idle timeout, since draining a wedged stream can block until the same stalled upstream produces bytes (defeating the timeout); drain only when bounded and worth it for connection reuse. On a multiplexed transport (HTTP/2) cancel only the stream, not the shared socket, or you kill unrelated healthy in-flight streams; on a non-multiplexed connection (HTTP/1.1 keep-alive) there are no sibling streams, so closing the underlying socket is the normal way to abort a stuck read.
31
+ - **Back off on idle-timeout retries.** A stalling upstream (overload) makes *every* client hit the idle timeout at once; immediate retry multiplies the load into a storm. Count idle-timeout failures against the same bounded retry budget with jittered backoff, and honor retry-after / circuit-breaker signals — a stall is a load signal, not just a transport blip.
32
+ - **Measure idle only while waiting on the upstream read.** If the read loop is blocked pushing a delta to a slow downstream consumer (backpressure), that elapsed time is not upstream stall — counting it cancels a perfectly healthy stream. Start the idle clock when you begin awaiting the next upstream event and pause it while blocked on client/consumer delivery.
33
+ - **An idle timeout after partial progress is not freely retryable — it inherits the partial-stream rule.** Once the stream has emitted client-visible deltas or dispatched tool calls, a blind retry replays/duplicates them (the same double-dispatch hazard as any mid-stream failure). Apply the turn loop's uncommitted-partial discipline: only auto-retry an idle-timed-out stream that produced no committed progress/side effect, otherwise require a resumable or idempotent continuation rather than re-running the call. (See `agent-turn-lifecycle.md` — buffer uncommitted until the terminal event; dedupe dispatch by response + tool-call id.)
34
+ - **Stop TTFT/idle timers on every termination branch.** A timer stopped only on the success path — but not on error, cancel, fallback, or early-EOF branches — fires after the stream already ended and reports a false timeout against a finished call; audit every exit path when adding a termination branch. Stopping alone is not enough against a concurrent fire: the timer's callback must re-check terminal/generation state under the same lock or CAS that guards its side effect (a bare flag read can still lose the race; for channel timers, drain the fired channel before reuse) so a stale expiry racing the terminal event cannot cancel or re-dispatch a stream that already finished.
35
+ - **Only an *upstream* stall enters retry — a downstream/caller cancellation is terminal.** The same cancel path fires for "upstream went silent" (retryable per above) and "the consumer disconnected / caller deadline elapsed" (the client is gone — see the client-disconnect-cancels-upstream rule above). Classify the cancellation cause: replaying a generation and re-running tool work for a consumer that no longer exists is pure waste and duplicated side effects. Downstream/caller cancellation is terminal non-retryable; only upstream idle/transport failure feeds the retry policy.
36
+ - If usage only appears in the final chunk, emit a final usage event and record it once.
37
+ - Timeouts and client disconnects must cancel upstream requests; verify the disconnect→upstream-cancel path over the real transport, not only in unit tests with fake streams.
38
+ - **When evolving the stream protocol, answer identity/lifecycle questions before designing event fields.** "What fields does the new event type carry" and "which existing identity/idempotency/billing model do these events bind to" are two different questions, and the second must be answered first: which answer/revision object the events belong to, what key their usage is accounted under, and whether a retry hides a provider-side billing fork. Event designs that skip this get reworked at review or implementation time.
39
+ - **Answer-replacing stream events need explicit commit-point semantics — per plane, not one global commit.** For *answer visibility*, bind the revision's commit point to its first visible content delta: when cancel or degrade lands before commit, the previous answer must be preserved while the failure still surfaces to the caller, and caller-visible mutations (staged references/metadata) plus caller-side effects (tool dispatch) stay buffered until commit so a never-committed revision leaves no executed side effects to duplicate on replay. *Accounting is a separate plane*: provider-attempt usage/ledger records book per attempt under their own keys (per the usage rules below) regardless of whether the revision ever commits — an uncommitted revision still consumed provider tokens, and deferring or dropping its usage record is a billing hole, not tidiness. Post-commit failure semantics must also be an explicit spec decision, not an accident: when a revision commits (content already visible) and the stream later fails, either keep the partial replacement visible with the failure surfaced (progressive replacement) or roll back to the prior answer at terminal — name the choice and test it; leaving it implicit is how a partial revision silently clobbers a good answer. Define the semantics explicitly for streams with no content plane (tool-only, refusal, error-terminal): name which event commits them and how their outcome surfaces. This boundary is what chain-level replay tests force out — for stateful stream protocols, write chain-level (multi-event-sequence) tests at spec time, not after implementation.
40
+ - **OpenAI Responses API supersedes Chat Completions as the recommended surface for new integrations** per `platform.openai.com/docs/guides/migrate-to-responses` and `developers.openai.com/blog/responses-api`. Two load-bearing differences: (a) **reasoning state preservation across turns** — Responses keeps the model's reasoning context alive between API calls, whereas Chat Completions drops it; OpenAI's published evals show ~3% SWE-bench Verified improvement vs Chat Completions with identical prompts on reasoning models, and ~40-80% better cache utilization. (b) **agentic-by-default**: a single API request can call multiple built-in tools (`web_search`, `image_generation`, `file_search`, `code_interpreter`, custom functions) in one round-trip — the orchestration that previously required client-side tool-loop scaffolding moves provider-side. **CRITICAL P0**: provider-side built-in tools bypass the team's own auth / audit / rate-limit / data-egress boundaries — calls to `web_search` happen at OpenAI's edge with OpenAI's network, calls to `file_search` index data the team uploaded to OpenAI's storage, calls to `code_interpreter` execute code in OpenAI's sandbox. Required: explicitly DISABLE the built-in tools in the Responses request unless a specific tool is approved for the use case (default-on is the wrong stance for any team with compliance / data-residency / audit obligations); when enabled, require post-response reconciliation that emits one audit event per built-in tool invocation (which tool, with what arguments, returning what footprint), enforce per-tool rate limits the team owns rather than relying on provider defaults, and document the data-egress policy (what user content can the team's prompts contain when web_search / file_search is on). Migration mechanics: structured-output config moved from `response_format` to `text.format`; `reasoning_effort` default is model-specific and drifts across model versions (see `model-prompt-evaluation.md`) — do not hardcode a flat default; read the current model's documented default per route. **When to migrate**: new integrations on GPT-5.x reasoning models default to Responses; existing Chat Completions integrations stay until reasoning-state preservation OR multi-tool orchestration is load-bearing — do not migrate as drive-by during feature work. Cross-provider abstraction layers (LangChain, LlamaIndex, an in-house adapter) need to surface the API choice explicitly because SDK shapes differ; verify the adapter speaks both forms before flipping callers.
41
+ - **Prompt caching is a 2024-2025 industry-standard cost lever, not a niche optimization** — all three majors support it with different ergonomics. **Anthropic**: explicit `cache_control` blocks on prompt segments per `docs.anthropic.com/en/docs/build-with-claude/prompt-caching`; cache-read tokens are billed at ~10% of standard input price (NOT free — the 90% savings is the discount on those read tokens, not zero cost); cache-write is 1.25x base for 5-minute default TTL, 2x base for 1-hour TTL; the 1-hour TTL pays back vs uncached at roughly the third cache read once the write premium is amortized. Anthropic claims up to 90% cost savings on cache hits, but a workload that writes once and reads zero is a net loss. **OpenAI**: automatic caching on prompts ≥1024 tokens with no API change; Responses API improves cache utilization 40-80% vs Chat Completions per OpenAI's own announcement. **Google Gemini**: **implicit caching enabled by default for all Gemini 2.5+ models** per `developers.googleblog.com/en/gemini-2-5-models-now-support-implicit-caching/` — no API change needed; minimum cache-eligible request 1024 tokens (2.5 Flash) / 2048 tokens (2.5 Pro). **Architecture impact**: structure prompts so the cacheable prefix is stable (system prompt + tool/function defs + few-shot examples on top; per-request user content at the bottom). **Tenant-isolation P0**: cache keys / cacheable prefixes MUST be tenant-isolated — NEVER place tenant id, tenant-private context, tenant-specific tool definitions, or tenant-scoped few-shot examples inside the shared stable prefix unless the provider's account-isolation contract is verified AND tested (Anthropic and OpenAI account-scope caches per organization, but the moment two tenants share one API account / organization, a shared prefix that includes one tenant's context is a cross-tenant leak vector). Default position: tenant-scoped content goes BELOW the cache boundary; the stable prefix carries only tenant-neutral system prompts / tool defs / generic examples. Verify with a per-tenant cache-hit-rate audit — if tenant A's content hash ever produces a cache hit on tenant B's first request, the isolation is broken. **Footgun (Anthropic-specific)**: per Anthropic's docs, extended-thinking blocks interact with prompt caching in ways that invalidate cache entries more aggressively when the thinking state changes between turns; pin per-route extended-thinking decisions to avoid silent cost drift. For OpenAI / Google equivalents, this drop is plausible but not documented as a generalized rule at writing — measure cache-hit-rate per route before assuming the same dynamic. Treat cache-hit-rate as a per-route SLI emitted into the observability stack.
42
+ - **Keeping the prefix stable means making cost-only prefix churn cache-aware — and never deferring an authority/correctness change to save a cache write.** Any mid-conversation change to the cached prefix invalidates the cache from that point and pays the full uncached prefill on the next turn (plus, on providers with explicit cache-write pricing, a fresh cache-write charge), silently every turn after if it keeps mutating. Split such changes into two classes and handle them differently:
43
+ - **Authority-neutral, cost-only churn** (re-ordering stable prefix content, refreshing an unchanged or non-authority system-prompt section, adding authority-neutral formatting/few-shot guidance that grants no new capability) — here the recommended pattern is **deferred invalidation**: don't rebuild the prefix as a side effect of an unrelated command; instead record the change as a durable `pending` mutation bound to principal/session/config-generation, apply it **before the next model request** (not at some far-off boundary, and never silently dropped — if applying it fails, fail closed to a degraded/visible state rather than continuing on stale-but-uncommitted intent), and offer an explicit opt-in (e.g. a `--now` / "apply immediately" path) to rebuild the prefix this turn.
44
+ - **Authority/correctness/capability-surface changes** — ANY change to the callable/visible tool or capability surface (add, remove, revoke, schema/trust/exposure change), plus policy/authorization changes, memory deletion, model/behavior changes, and any privacy/principal/tenant/auth change — must invalidate / fence / re-authorize **NOW, never deferred for cost**. These follow the fail-closed snapshot/re-verify/re-authorize rules above and in `references/retrieval-agent-safety.md` / `references/agent-tool-dispatch.md`, which take precedence over cost-driven deferral. Deferring a revocation that leaves a revoked tool callable for the rest of a long turn — or deferring a newly-added tool so the model plans around a surface that isn't really there — is a correctness/security bug, not a saving. The only other routinely-acceptable mid-conversation prefix change is the deliberate context compaction already governed above (which invalidates and re-writes by design).
45
+ - **Structured output strict mode is supported by all three majors but with non-equivalent guarantees**. Anthropic: tool-use with JSON-schema definitions on `input_schema`. The provider enforces the tool-call SHAPE the model emits matches the declared `input_schema` at the wire level, but application-side validation against the schema is still required — output-token truncation can produce partial JSON, `oneOf`-style unions need explicit workaround, and edge-case constraint enforcement (`pattern` / `format` / nested `$ref`) varies; do not trust the provider as the only validator. OpenAI: `text.format` (Responses API) / `response_format` (Chat Completions) with `json_schema` strict mode per `developers.openai.com/api/docs/guides/structured-outputs` — strict mode guarantees the model output adheres to the schema, but only a subset of JSON Schema features is enforced (verify the docs page for the specific feature: `pattern` / `format` / `oneOf` / nested `$ref` support varies). Google Gemini: `responseMimeType: "application/json"` + `responseSchema` for JSON mode. **Per-feature audit before relying on strict mode**: write the schema, send a request with deliberately ambiguous prompt, verify the response shape matches AND the values are coherent (strict mode prevents shape violations, NOT semantic correctness — a model can still hallucinate field values that satisfy the schema). For high-risk extraction (billing / identity / permissions / money / quota), validate the output AGAIN with a second schema-aware parser at the application boundary; do not trust strict-mode output as the only validation layer. **Schema evolution discipline**: version the schema (carry a `schema_id` in both the prompt's schema declaration AND the model's expected output structure when the surface allows); treat adding a required field as a BREAKING change for downstream consumers including any cached responses or replay datasets; reject outputs whose `finish_reason` indicates `max_tokens` truncation as parse failures (truncated JSON looks valid until a key is cut mid-string) and retry-or-repair only against the raw captured response, not against a partially-parsed object that already dropped the truncation signal.
46
+ - **Truncation on free-text / long-form output is a continue-or-raise decision, distinct from the structured-output case above.** When a *non-structured* response stops at the output-token cap the body is a valid prefix cut at the limit; do NOT return it as if complete. **The truncation signal is per-provider AND per-API — enumerate it for every surface you support**, since a detector hard-coded to one field silently accepts a capped result from another API as complete: OpenAI Chat Completions → `finish_reason: "length"`; OpenAI Responses API → `status: "incomplete"` with `incomplete_details.reason: "max_output_tokens"` (NOT a `finish_reason`); Anthropic → `stop_reason: "max_tokens"`; Google Gemini → `candidates[].finishReason: "MAX_TOKENS"`. On a streamed response read this signal from the *terminal* event (final chunk / `message_delta` / completion), never an intermediate delta — mid-stream the field is absent, so checking an early delta sees no truncation flag and returns the capped stream as complete. Recoveries (`platform.claude.com/docs/en/build-with-claude/handling-stop-reasons` maps `max_tokens` to "raise `max_tokens` or continue the response"): **(a) raise the output-cap and re-run** when the shortfall is small and the input leaves context headroom — the simplest fix, and the one to prefer because (b)'s seam is lossy; but a *full* rerun is safe only when the capped response committed no client-visible delta or dispatched tool call — after committed streaming progress it replays/duplicates them (per the partial-stream rule above), so there use continuation (b) or surface incompleteness instead; **(b) continue-and-concatenate** for genuinely long output — re-send the *complete original request state* — EITHER via the provider's response handle (e.g. OpenAI Responses `previous_response_id`, which chains prior input/output items so you do NOT replay the transcript — but it does NOT carry the request's `instructions`, so re-send the system/developer `instructions` alongside the handle) OR by manual replay of system/developer instructions, tool definitions, and prior turns — with the assistant partial appended and a "continue from where you left off" turn, then append each continuation to an accumulator; combining the handle AND manual replay duplicates context and cost, and the bare `[user, assistant-partial, continue]` triple is only the minimal single-turn illustration — dropping the original instructions/tools lets the continuation violate the original constraints or drift to a different answer. This is the **opposite** of the JSON rule above: a truncated *structured* body is a parse failure to reject/repair, never to splice across continuations (concatenating across a cut key produces invalid JSON); the same applies when the cap is hit mid-**tool-call** — a truncated tool-call / arguments block is structured output to repair or re-request, never fed through this free-text append loop (continuation is for text content only). Guardrails the naive loop omits: **the continuation loop only decides truncated-vs-not — it does NOT own the success predicate.** Continue only while the stop signal is a *truncation* one (the per-surface signals above); on any other stop, exit the loop and hand the response to the gateway's normal terminal-outcome classification (success vs refusal vs `tool_use` / `pause_turn` vs safety/recitation vs error — owned elsewhere in this skill), which must inspect the actual output content, not just the top-level finish/status field. Never equate "not truncated" with "complete successful answer": an OpenAI refusal arrives as `finish_reason: "stop"` + `message.refusal` (Chat) or `status: "completed"` with a `refusal` content item (Responses) — both pass a naive natural-finish check yet carry no answer text, so returning them as complete ships an empty accumulator and skips refusal/fallback/telemetry. A configured `stop_sequence` end is a genuine complete answer; refusal / safety / `tool_use` / `pause_turn` stops are terminal-but-not-a-final-answer and must route to their own handling, never be returned as the finished text; **bound resources, not just attempts** — a fixed max-attempts (e.g. 3) stops an infinite loop, but also hard-cap total input+output tokens / cost / wall-time and stop on a no-progress turn (an empty or near-empty continuation), because every continuation re-sends the whole growing transcript so input cost climbs super-linearly and can itself approach the window; **distinguish the output cap from the context wall** — Anthropic surfaces the window limit as a *separate* `model_context_window_exceeded` stop reason (model/version-dependent — default on recent models, beta header on older, and an input already over the window is an HTTP 400, not this reason), whereas OpenAI reports both the output cap and the window limit under the same `incomplete_details.reason: "max_output_tokens"`, so gate continuation on token-usage-vs-window rather than the reason string alone: once the transcript fills the window, continuing just re-hits the wall — stop and surface it / raise the window (a reasoning-model cap can also return truncated with *no visible output* because reasoning tokens consumed the budget, which a blind loop would append as an empty chunk); **the seam is semantic, not byte-exact** — "continue" can re-emit or skip tokens at the boundary, so validate the join for duplication/gaps rather than trusting a raw append, and do not summarize the carried partial (it removes the exact suffix the continuation must follow). When you stop without a complete answer, surface incompleteness **out of band** (a status flag / metadata / UI marker), NOT by concatenating a notice into the accumulated output — downstream persistence, evaluation, or a further model turn will otherwise treat that notice as model-generated text.
47
+ - **Inline content-type markers within `content` deltas are a parser contract, not free-form text**: when the product needs the model to emit mixed-modality output inside one `content` stream — markdown plus structured renderable blocks (e.g. a card type, a citation block, a chart spec, a code-editable cell) — the markers MUST be a typed contract (well-named XML-style tags such as `<card-type-N>...</card-type-N>`, fenced code blocks with a typed language tag, JSON-mode segments delimited by a sentinel). The parser MUST be an incremental state machine over arbitrary chunk boundaries: SSE can split the opening delimiter, the tag name, attributes, a code fence, a JSON sentinel, or the closing delimiter across multiple chunks (e.g. one chunk ends `<car` and the next begins `d-type-1>`). A naive implementation that only buffers after a complete opening tag will emit the prefix as plain text and silently lose the structured block. Specify bounded buffering, partial-state carry-over per chunk, and explicit EOF/error behavior when a started block never closes. Free-form delimiters chosen ad-hoc per feature (one feature picks a short tag, another picks a bracket sentinel, a third nests JSON in markdown) fork the renderer per feature and break feature-cross-cutting concerns (copy, regenerate, history replay). Land the marker schema in the same contract artifact as the `tool-call` schema. **Unknown markers MUST surface to the renderer as escaped plain text / a text event, never as raw HTML or a renderable node, and surface telemetry with the payload redacted** — a permissive Markdown or `dangerouslySetInnerHTML` path turns an unknown custom tag with attributes into an XSS sink. Reasoning content is a separate channel: provider adapters MUST translate provider-specific reasoning tags (e.g. `<think>` blocks, `reasoning_content` fields) into reasoning stream events upstream of this content-marker parser; the renderable content-marker contract NEVER uses a reasoning-tag name as a content block, otherwise reasoning may leak into user-visible content or be persisted as user content by replay/history.
48
+
49
+ ## Reliability
50
+
51
+ - Every request needs context/deadline, total timeout, per-attempt timeout, retry budget, and fallback budget. The per-attempt timeout must exist explicitly in the orchestration layer that runs the retry/fallback chain — an adapter's internal HTTP-client timeout is a private implementation detail, not an orchestration guarantee, and swapping the adapter must not silently remove the bound.
52
+ - Retry only retryable network, timeout, rate-limit, and transient server failures.
53
+ - Do not retry non-idempotent tool execution without an idempotency key.
54
+ - Use bounded concurrency per route/model/provider and reject or queue when saturated.
55
+ - Separate user-visible failure from best-effort telemetry persistence; failed call-record writes should not fail successful inference.
56
+
57
+ ## Observability
58
+
59
+ Record at least:
60
+
61
+ - caller or feature route;
62
+ - request id / trace id;
63
+ - model and model version;
64
+ - prompt key and prompt version;
65
+ - latency;
66
+ - prompt/completion/reasoning/total tokens;
67
+ - finish reason;
68
+ - success/failure and error class;
69
+ - fallback path;
70
+ - tool calls requested/executed;
71
+ - redacted payload and response where policy permits.
72
+
73
+ Never log secrets, raw private user content, credentials, or unredacted files unless the product has an explicit data-retention policy and access control.
74
+
75
+ ## Generated Content Lifecycle
76
+
77
+ When model output can affect users, publishing, permissions, billing, support, or downstream automation, treat the output as a lifecycle object rather than a string response.
78
+
79
+ | State | Meaning | Allowed operations | Required metadata |
80
+ | --- | --- | --- | --- |
81
+ | candidate | Generated but not yet trusted or user accepted | preview, inspect sources, regenerate, edit, reject | model, prompt/policy version, request id, latency, token/cost, source/context summary, safety/error class |
82
+ | reviewed | Human, rule, or evaluator has checked the candidate | accept, revise, compare, reject | reviewer/evaluator, review time, decision reason, changed fields |
83
+ | accepted | User or workflow owner chose this output | publish, save, enqueue downstream action, rollback before publish when possible | accepting actor, accepted version, source candidate id |
84
+ | published | Output is visible or has affected another system | view, audit, retract/rollback if supported | publication target, trace id, published version, rollback/retraction path |
85
+ | rejected | Output must not proceed | regenerate from a new candidate or close | rejection reason, unsafe/low-quality class where relevant |
86
+
87
+ Do not silently move from generated text to published final state on high-impact routes. The release gate should prove the candidate/review/accept/publish transitions, metadata visibility, rejection path, and audit or rollback evidence.
88
+
89
+ ## Inference-As-RPC (Service-To-Service Inference Calls)
90
+
91
+ LLM clients are one kind of inference client; the broader pattern is service-to-service inference where one backend service (often Go or Java) calls an inference service (often Python over HTTP, sometimes a Triton native gRPC client). The rules below apply when the inference target is internal infrastructure, not a SaaS provider.
92
+
93
+ - **Transport**: HTTP POST + JSON body is the common cross-language transport even when the inference server can speak gRPC or Triton-native; the trade-off is universal client support and easy debugging vs the protocol-level efficiency of native binary. Document the choice per service.
94
+ - **URL construction**: derive the inference endpoint from a method-name → PSM (service identity) mapping plus a `predict` path helper, not by hard-coding the URL in every caller. A `GeneratePredictPath(method, version)` helper centralizes routing so a model migration is one change, not N.
95
+ - **Hard-coded inference IPs (`<private-ip>:<port>`) in business code are a finding**, not a shortcut. Every inference target must be resolved through service discovery (NACOS/Consul/k8s service DNS); the discovery layer owns load balancing, health filtering, and the lane/canary route.
96
+ - **Method-level timeout per model SLA**: each inference method has its own timeout (image feature extract ~90 s, image match correction ~120 s, OCR ~90 s, classification ~10 s). Choose the per-request timeout as `min(method-timeout, config-override, context-deadline)` so the caller never holds longer than the strictest applicable budget.
97
+ - **Per-attempt context** must be a sub-context of the request context; the client cancels the in-flight HTTP call when the caller's deadline elapses. Long-running inference holding the socket past the caller's cancellation is a leak.
98
+ - **Large payload routing**: when the inference call carries an image, document, or multi-MB binary, the client converts the in-memory object to a signed object-storage URL and sends the URL plus metadata, not the inline bytes. For backward compatibility, support base64 inline as a fallback with a clear size threshold. When the inference response is itself a binary (corrected image, mask, rendered overlay), the server side should return a signed URL when possible; if the contract requires inline bytes, do not strip them from the response (that breaks callers) — apply size limits at the logger / persistence layer so log bloat is bounded without changing the wire contract.
99
+ - **Signed URL trust model**: a signed URL is a bearer credential plus a server-side fetch target — it is not a "free" payload mechanism. Architecture defines: (1) allowed bucket/host allowlist (the inference server refuses to fetch from a URL outside it — SSRF defense); (2) allowed object namespace / prefix per caller; (3) maximum expiry window (short — minutes, not hours); (4) HTTP method scope (GET only for fetch URLs, PUT only for upload URLs); (5) content-length and content-type caps validated server-side before processing; (6) checksum or `If-Match` digest validation against an expected hash when the registry supports it; (7) signed query parameters redacted from logs and trace attributes. A URL that fails any of these is rejected before the model is invoked.
100
+ - **Authentication and authorization across internal inference calls**: "internal infrastructure" is not a security boundary. The inference RPC contract must specify caller identity (mTLS service identity, signed JWT, or platform-issued workload identity), caller authorization (allowlist of services / accounts permitted to call this model), tenant / resource authorization (the call carries a tenant id and the inference service confirms the caller is allowed to use that tenant's data), and quota enforcement (rate limit per caller / per model / per tenant). A model endpoint reachable by anyone who can discover it via service discovery is a finding.
101
+ - **Discovery cache miss policy is per-route, not portfolio default**: read-only idempotent inference calls (classification, OCR on already-uploaded content) MAY fall back to a documented baseline lane on cache miss. **Canary / stress / shadow / tenant-sensitive / write-with-side-effects routes MUST fail closed** when the requested lane has no healthy instance — silent fallback to production can leak canary traffic, stress traffic, or one tenant's processing into another lane. Architecture records the per-route fallback policy explicitly; the default for unmarked routes is fail-closed.
102
+ - **Trace propagation**: every outbound inference request carries W3C `traceparent` + `tracestate` (and `baggage` when needed) via the OpenTelemetry propagator's `inject`; the inference server reads them via `extract` so spans become children of the parent trace. In addition, the platform's request-id / log-id header (`X-Request-ID` / `X-Log-Id` or equivalent) travels alongside for log-level correlation. **Edge trust rules**: at the public trust boundary, the server regenerates `X-Request-ID` if the caller supplied one (or validates against a strict shape) so callers cannot spoof log correlation; internal hop-to-hop, preserve the id without rewriting. For `asyncio` / Ray / thread-pool boundaries, capture the OTel context at task creation via `opentelemetry.context.get_current()` and reattach inside the worker (`context.attach` / `context.detach`); without this, child-task spans become orphans of the parent trace.
103
+ - **Cross-stack ctx key discipline**: define context-key constants in one shared package (e.g. `TrafficRequestIdCtxKey`, `BusinessShardCtxKey`, `ExperimentRouterCtxKey`); business code reads through typed accessors, not raw string keys. The set is part of the inference contract.
104
+ - **Failure mapping**: HTTP 2xx + envelope `code != success` is a domain error; HTTP non-2xx is a transport error; both must be distinguishable in the caller's error type. A flattened generic `ServerError` that erases the inference reason makes the call chain undebuggable.
105
+ - **Inference-specific error codes**: reserve a numeric range (or a typed error class) for inference-side failures (model not found, GPU OOM, input shape mismatch, timeout on inner model call) so business retry and circuit-breaker rules can act on them differently from generic transport errors.
106
+ - **Bounded retry without circuit breaker is incomplete**: short retry loops (2-3 attempts with backoff) handle transient blips; sustained failure needs a circuit breaker per upstream / per model so the service can fail fast and shed load instead of amplifying the upstream incident.
107
+
108
+ ## Experiment And Traffic Routing For Inference
109
+
110
+ Production inference often runs more than one model variant simultaneously. Encode the routing as data, not code:
111
+
112
+ - **Experiment strategy types**: A/B test (fixed ratio), Canary (small percentage rolling up), Orthogonal (independent dimensions multiplied), Shadow (replay candidate alongside live without affecting users). Pick by goal; reuse one engine across all four.
113
+ - **ExperimentRule shape**: traffic ratio + matcher condition (tenant / route / lane / user cohort) + target model/version. Rules sum to 1.0 within a route or fail validation; silent ratio drift is a release-time bug.
114
+ - **Per-request experiment metadata**: record which experiment id, which variant, and which model version served each request. Evaluation joins to this metadata; without it, comparing model A vs B is guesswork.
115
+ - **Quality + cost metrics** (accuracy, precision, recall, F1, custom domain score, token cost, latency p95) are first-class fields on the experiment result, not afterthoughts. `confidence_threshold` and `max_retry` per experiment let downstream services react to low-confidence outputs deterministically.
116
+ - **Shadow inference** runs the candidate without returning its output to the user; persist both candidate and production output and diff offline. Shadow is the safest activation gate for high-impact changes — **but it doubles the GPU / object-storage / queue load**, so it needs its own budget. Architecture defines: (1) shadow traffic ratio (start at low single-digit %, ramp deliberately); (2) separate GPU / replica budget for shadow so it cannot starve production; (3) sampled rollout (not 100% mirror — pick representative traffic); (4) automatic cut-off when shadow error rate, latency p99, or cost-per-request exceeds the configured threshold; (5) explicit termination when the comparison concludes. A shadow that runs forever is a hidden production dependency.
117
+
118
+ ## Service Discovery, Lane Isolation, And Lifecycle
119
+
120
+ The inference fleet often spans many service instances across many environments. The discovery layer is the integration boundary; the service-side discipline determines whether it works.
121
+
122
+ - **Three-tier isolation**: service identity (PSM, project.service.kind format like `project.service.api` or `project.service.rpc`) → lane (online / canary / offline / stress) → instance (healthy / draining / down). A request resolves through all three before hitting an inference replica.
123
+ - **PSM format is strict**: a three-segment dot-delimited identity (`project.service.kind`) parsed at client construction; mistyped PSMs fail at startup, not at first request.
124
+ - **Lane derivation from environment**: read `LANE` (or the project's equivalent) at startup and tag every outbound and inbound request; in-cluster discovery filters by lane metadata so a canary call cannot fall back to production by accident.
125
+ - **Three-stage instance lifecycle**: `register → heartbeat → graceful_shutdown`. Register only after readiness probe passes (model loaded, health endpoint green). Heartbeat at a bounded cadence with explicit failure handling. On SIGTERM, deregister first, wait for in-flight requests to drain (bounded), then exit. Hidden infinite loops, unbounded waits, or process-kill exits are anti-patterns.
126
+ - **Discovery client behavior**: cache the resolved instance list with a short refresh interval (typically 5-15 s); filter to healthy + matching lane; load-balance across the survivors (round-robin or random); on cache miss fall back to a documented baseline (cross-cluster k8s service, baseline lane) before failing the call.
127
+ - **Internal-SDK as a vehicle**: the discovery + transport + retry + tracing concerns live in one internal Python WHEEL (or Go module) that every inference caller imports. Each business service does not re-implement; new patterns land in the SDK and propagate through dependency updates.
128
+
129
+ ## Multi-Language SDK Parity
130
+
131
+ When the same client/gateway SDK is mirrored across languages:
132
+
133
+ - Use one language-agnostic case manifest (case id + description) as the parity contract; adding a shared behavior forces a linked edit in every language. Assert with relational invariants (manifest ⊆ covered AND covered ⊆ manifest), not count snapshots, so drift is named, not just detected.
134
+ - Align parity guards to strength, not intent: a same-named guard that self-verifies in one language (AST scan of real tests) but is a hand-maintained string table in another looks equivalent while protecting unequally. At design time ask "does this guard fail the same way in each language?"
135
+ - Mirror protocol semantics and boundary behavior, not language idioms: error-model or reserved-field implementations may diverge per language convention, but each divergence carries a comment explaining it.
136
+ - Audit error-model field completeness across languages: one SDK exposing a raw-response escape hatch while another silently drops `retryable`/provider error code is a parity break, not an idiom difference.
137
+
138
+ ## Inference failure classification before retry
139
+
140
+ Classify inference failures before retrying: transport blips, stale auth, retryable server capacity, explicit rate limits with reset time, subscription or quota exhaustion, entitlement denial, context-window overflow, user abort, and malformed request are different outcomes. Build this as a closed taxonomy with explicit retryable semantics, and lock each provider's mapping with tests against real captured error responses, before building any fallback/cooldown policy on top: the same terminal condition surfaces under different wire shapes per provider (context-window overflow as 413 or as 400 plus an overflow error code; balance/quota exhaustion as 402), and misclassifying such terminal client errors as retryable server errors is the most common gateway defect class — retries that can never succeed, burning quota and latency. Honor retry-after/reset directives with capped waits, bound total retry amplification across SDK, gateway, queue/job, and user-visible retry layers, and let background or non-user-visible calls fail fast during capacity cascades unless their result is required for safety or permission correctness. Subscription or quota exhaustion is terminal until explicit reset or renewed-quota evidence exists; entitlement denial is terminal until fresh entitlement authorization evidence exists; user abort is terminal unless a fresh user action restarts the request; malformed requests must not retry unchanged. After stale auth or stale connection, retry only if the full original tuple still matches: principal, account or tenant, workspace, privacy state, route/adapter, provider credential/client generation, provider account/project or API-key scope, quota bucket, billing namespace, region or data-residency endpoint, entitlement state/version, authorization or policy version, prompt or policy version, model generation/params, tool-schema or capability generation, session or job id, session incarnation, and canonical rendered request digest that includes policy/prompt version, model generation/params, and tool-schema or capability generation. On any tuple drift, cancel the pending retry or rebuild, re-authorize, and re-render from scratch. Persistent unattended retry needs abort handling, periodic progress/keep-alive, and a maximum reset horizon.
141
+
142
+ ## Conversation compaction and context reconstruction
143
+
144
+ Treat conversation compaction, summary replacement, tool-result deletion, and session-memory compaction as model-visible state transformations, not token housekeeping. Every compaction path needs a durable boundary marker, preserved-segment relink metadata when raw messages survive, an explicit lossy-deletion policy, and post-compact cache invalidation for file reads, nested memory, prompt sections, classifier/speculative approvals, micro-compact state, and session-message caches. Summaries must preserve current user intent, pending tasks, unresolved errors, tool/action finality, permission and policy decisions, source-scope labels, files or skills needed for continuation, and known uncertainty; media, large tool outputs, or old turns may be replaced only with typed markers or bounded summaries that cannot be mistaken for fresh evidence. Prompt-too-long retry may drop old API-round groups only after keeping at least one summarizable group, inserting a synthetic continuation marker when needed, preserving tool-use/tool-result pairing and assistant thinking blocks, preserving or revalidating required authorization, privacy, policy, source ACL/provenance, current-intent, pending-task, and tool/action finality evidence, and recording that the summary is lossy; fail closed if the dropped groups contained the only copy of required evidence or unresolved finality that cannot be reconstructed. Session-memory compaction must wait briefly for in-flight extraction, fence every extraction by attempt id and context generation, reject late or stale extraction writes after boundary drift, session incarnation change, privacy opt-out, deletion, ignore-memory instruction, account/workspace switch, or capability/policy generation change, fall back when the summarized boundary is unknown or stale, keep enough recent text and token context under a max cap, avoid splitting tool pairs or merged assistant messages, remove stale compact boundaries from the kept segment, and fail over to ordinary compaction if the resulting context still exceeds threshold. Post-compact rehydration of files, plans, invoked skills, deferred tools, agent listings, external instructions, and hooks must be size-bounded, treated as untrusted context until revalidated, and bound to the full current authorization/source tuple before becoming model-visible: principal, account or tenant, workspace or repository identity, privacy or data-residency state, session id and incarnation, authorization or policy version, tool/source trust identity, capability or tool-schema generation, file and memory snapshot generation, and canonical source digest where available.
145
+
146
+ ## Fallback, cooldown, and degraded modes
147
+
148
+ Fallback, cooldown, and degraded modes must preserve user and product semantics. Short retry windows may preserve cache locality; long or unknown waits should switch to an approved degraded path, visible refusal with a sanitized reason code, cooldown, or standard-capacity path instead of silently changing model class, quality, speed, or cost. Any fallback/cooldown/degraded transition that changes provider, route/adapter, model class or params, region or data-residency endpoint, quota bucket, billing namespace, provider credential/client generation, provider account/project or API-key scope, entitlement state/version, authorization or policy version, prompt or policy version, tool-schema or capability generation, or canonical rendered request digest must become a new authorized and rendered attempt with separate usage attribution, not a silent continuation. Fail open only for non-critical advisory policy fetches with a safe stale cache that matches the full authorization/targeting tuple, maximum stale age, and monotonic policy/revocation epoch or fresh deny/opt-out watermark; stale advisory data must never override a newer deny, revocation, privacy opt-out, logout, account switch, or entitlement change. Fail closed for safety, entitlement, privacy, permission, destructive-action, or policy gates when no fresh-enough allow exists.
149
+
150
+ ## Usage, latency, and call-record accounting
151
+
152
+ Record token usage, latency, request success, finish reason, tool calls, retry attempt, retry-inclusive and retry-exclusive duration, fallback/cooldown/degraded-mode decision, unknown-cost markers, and caller/request id. Persist usage/cost only to the same full tuple used for retry/re-render, including principal, account or tenant, workspace, privacy/data-residency scope, route/adapter, provider credential/client generation, provider account/project or API-key scope, quota bucket, billing namespace, region or data-residency endpoint, entitlement state/version, authorization or policy version, prompt or policy version, model generation/params, tool-schema or capability generation, session or job id, session incarnation, and canonical rendered request digest; purge or invalidate restored counters after any member of that tuple changes. Deduplicate and finalize usage by request id, provider request id where available, idempotency key or usage-event id, and attempt id so retry, stream reconnect, and late provider completion paths cannot double-count or lose a terminal usage event. Use terminal ledger semantics for late, duplicate, retried, reconnected, streamed, and provider-completion usage events. Keep per-model/per-route accounting separate enough to explain quota, billing, or incident questions without leaking prompts or business data.
153
+
154
+ ## Prompt cache miss attribution
155
+
156
+ Treat prompt cache miss detection as a diagnostic control plane, not as ordinary usage telemetry. A detector must take a pre-call snapshot of the rendered prompt/cache tuple before comparing post-call cache-read tokens: system/instruction digest, tool-schema digest, cache-control scope or TTL class, route/model/effort/output parameters, mode or feature toggles that affect the provider cache key, extra request-body digest, query source class, session or agent generation, principal/workspace/privacy tuple, and capability/tool generation. Fence the pre-call snapshot, post-call comparison, and baseline mutation by immutable request/attempt identity, prompt or message generation, abort generation, provider-response identity where available, session incarnation, and rendered-request digest; reject late, retried, reconnected, or aborted completions when any identity or generation no longer matches. Bound tracked sources and evict stale entries so background agents or short-lived sessions cannot grow unbounded memory or cross-contaminate attribution. Classify cache-read drops with explicit precedence and multi-cause support: expected drops from first call, known TTL windows, intentional cache-edit deletion, compaction or context reset, and baseline reset must suppress incidents; likely server-side or routing causes must not be over-attributed to prompt drift; rendered prompt/tool/parameter drift may be claimed only when the bound pre/post tuple proves it and higher-precedence expected-drop classes are excluded; otherwise emit a bounded `unknown` or `ambiguous` cause. Diagnostic events may expose only booleans, counts, bounded deltas, enum-like cause classes, and sanitized fixed-vocabulary tool/category labels; user-configured tool names, connector names, prompts, schemas, request bodies, local paths, credentials, raw request or session ids, and free-form errors must not enter telemetry. Any local diff or support artifact that compares rendered prompts, instructions, tool descriptions, schemas, or request bodies is a privileged support artifact: write it only under an approved local diagnostic directory with bounded size and retention, never upload it automatically, gate sharing on explicit support policy, label it as raw-sensitive, and keep only a sanitized pointer or artifact class in logs.
@@ -0,0 +1,146 @@
1
+ # Model, Prompt, And Evaluation Control
2
+
3
+ ## Model Registry
4
+
5
+ Track models as controlled runtime assets:
6
+
7
+ - model id, display name, type, provider adapter, base model, active version;
8
+ - version id, parent version, status, changelog, activation time;
9
+ - default parameters and parameter policy;
10
+ - allowed use cases or routes;
11
+ - deprecation and rollback policy.
12
+
13
+ Avoid hard-coded model strings scattered through product code. Business logic should ask the registry/router for an active model by capability or route.
14
+
15
+ ## Model Version Baseline 2025-2026
16
+
17
+ - **Three major-provider model families are the credible production options at 2026-Q2**; the registry/router MUST verify exact current model strings against the live docs page before pinning, because provider naming cadence has accelerated (Anthropic ships ~quarterly minor revs, OpenAI ships sub-version reasoning-effort variants, Google ships Pro/Flash/Flash-Lite tiers separately). Authoritative-doc sources for live verification: `docs.anthropic.com/en/docs/about-claude/models` (Anthropic) / `developers.openai.com/api/docs/changelog` + OpenAI Help Center model release notes / `ai.google.dev/gemini-api/docs/changelog` (Google). General shape: Anthropic ships Opus (frontier reasoning) + Sonnet (workhorse, 200K default + 1M beta context per `anthropic.com/news/1m-context`) + Haiku (cost/speed); OpenAI ships GPT-5.x with a `reasoning_effort` parameter (default value and available effort tiers are model-version-specific and drift — do not hardcode; verify per target model per the next bullet) plus reasoning summaries; Google ships Gemini 2.5 Pro + Flash + Flash-Lite with explicit thinking-budget control. Per-route choice is a registry decision recorded in the project's model registry, not hardcoded. **Pinning + eval-baseline discipline (P0)**: pin the exact model version string in the registry config (NOT just the family name like `claude-sonnet`); block ENV-var-only flips between model versions in production (flipping `MODEL=...` to a new minor rev silently invalidates every prompt eval result against the old version); require evaluation replay or shadow comparison for ANY model-string or default-param change before flipping production traffic; emit a metric on per-route model-string drift so an unintended autoupgrade is observable, not just discoverable at the next eval. The recurring failure is: team upgrades minor rev for cost/quality reason, eval suite still passes BECAUSE eval prompts didn't exercise the regressed code path, production sees the regression a week later.
18
+ - **Extended thinking / reasoning effort is a per-request product decision, not an account-level default**. Anthropic's extended thinking (off by default) trades latency + token cost for accuracy on multi-step reasoning per `docs.anthropic.com/en/docs/build-with-claude/extended-thinking`; impacts prompt-caching efficiency (cache invalidation more aggressive with thinking on). OpenAI's `reasoning_effort` tunes the same axis per `developers.openai.com/api/docs/guides/reasoning`, but the **default differs by model version and is not monotonic** — defaults have varied across GPT-5.x revs (different revs ship different defaults; do not assume a trend) and new effort tiers (e.g. `xhigh`) appear in some revs. Do NOT hardcode a version-specific default here; confirm the default for the exact model you target against its live docs page before relying on "default behavior" — drift in this default is the most common silent cost/latency change across an OpenAI minor-version bump. `revalidate-when: OpenAI ships a new GPT-5.x reasoning rev`. Gemini exposes a thinking budget knob per `developers.googleblog.com/en/gemini-2-5-thinking-model-updates/`. **Pattern**: per-route declare reasoning intent (off / low / medium / high) based on the task's complexity AND record the decision in the prompt/route config so latency and cost regressions are traceable to an explicit knob, not provider-default drift. **Production-safety contract for reasoning enablement (P0)**: declaring reasoning intent is not enough — the per-route config MUST also carry (a) p95/p99 **latency budget** with downstream timeout propagation (reasoning can add seconds-to-tens-of-seconds to single-turn latency; downstream HTTP timeouts that worked at non-reasoning latency will fail), (b) **reasoning-token cap** to prevent runaway thinking burning 10-50× the non-reasoning token budget on a hard prompt, (c) **cost guardrail** as a per-request token + per-user-session aggregate ceiling, (d) **rollback path** to flip reasoning off at the route level under incident. Mid-route enablement of reasoning (turning thinking on after launch) is a load-bearing change that requires fresh load/eval evidence and SRE sign-off, NOT a config-only flip. Reasoning content is a separate channel from user-visible content — never persist reasoning into product history or display it as final output (see streaming rules in `llm-client-gateway.md`).
19
+
20
+ ## Prompt Registry
21
+
22
+ Track prompts as versioned assets:
23
+
24
+ - prompt key with optional system/type namespace;
25
+ - version id and parent version;
26
+ - status: draft, active, deprecated, archived;
27
+ - content, variables, defaults, tags, author, changelog;
28
+ - activation and rollback operation;
29
+ - local file or DB-backed storage depending on product maturity.
30
+
31
+ Handlers should render a named prompt version, not concatenate ad hoc prompt strings inline.
32
+
33
+ ### Prompt evolution as a first-class engineering doc
34
+
35
+ A prompt registry tracks the **current** asset set; without a parallel evolution surface, "why does this prompt look the way it does" becomes git archaeology that decays as branches merge, rebase, or get force-pushed:
36
+
37
+ - **Per-prompt evolution doc, keyed by stable IDs**: each materially-iterated prompt asset has a doc that lists, per significant version, `version-id → content-hash → landed-change-id → change-reason`. The primary key is a **stable monotonic prompt version id** (`v1, v2, ...`) — git commit hash and branch are NOT stable across rebase / squash / cherry-pick / branch-rename and must be treated as informational pointers, not identity. A content-hash (SHA-256 of the canonical prompt body) is the secondary identity that survives history rewrites; the landed-change id (PR number / merge-request id, NOT commit hash) links to the durable review record. Reasons are operational (the actual design change — a scoring-tier split, a new output field, a finer-grained output unit) not commit-message-restating. New columns are added when a new dimension matters (active branch, eval-result delta, rollback candidate).
38
+ - **Plain-text version snapshots, with separate escape hatches for size vs sensitivity**: when a registry entry would otherwise drop the prior content on activation, persist the prior version under a stable path (`<repo>/docs/prompt_versions/<prompt-key>/<version-id>__<short-descriptor>.txt`). The filename combines the stable version id with a descriptor (never just one or the other; descriptor-only collides on similar changes, id-only is unreadable in PR review). Two distinct escape hatches:
39
+ - **Size-only** (large few-shot corpora, generated examples, prompts > ~50 KB, DB-backed prompt registries with already-versioned bodies): Git LFS pointer or the registry's own versioned backend is acceptable; the access audience is unchanged from the repo. Keep a `<version-id>__<short-descriptor>.ref` pointer file in `docs/prompt_versions/` that resolves the external location.
40
+ - **Sensitivity** (prompts referencing customer data, trade-secret rubric, partner-confidential instructions): the storage backend must enforce **strictly narrower access controls than the repo** — an artifact/registry store with per-identity authorization, audit log, and short-lived access tokens. Git LFS in the common setup grants any repo-cloner access to the LFS object and is NOT a sensitivity hatch; using LFS for a sensitive prompt moves the body out of text review while keeping the same audience, which is worse than the in-repo snapshot it replaces. Keep only a sanitized title/version-id in the repo pointer.
41
+
42
+ Repo-direct snapshots are for small reviewable non-sensitive prompts; the size hatch keeps the discipline workable as the asset grows; the sensitivity hatch is non-negotiable when the body crosses a confidentiality boundary.
43
+ - **Branch fork tracking is explicit**: when one experiment branch keeps iterating a prompt while another stops at an earlier point, the evolution doc records "branch X stopped at commit Y" so a future engineer can pick the correct version when rebasing or cherry-picking. Multi-branch prompt forks are not theoretical — fix-branch + feat-branch each touching the same prompt is common, and silent fork is a real regression source.
44
+ - **Evolution doc is updated in the same PR that ships the prompt change**: a prompt change without an evolution-doc append is incomplete; the doc append is part of the code review surface so the rationale is captured at the time the engineer remembers it, not reconstructed later from commit messages.
45
+ - **Anti-pattern**: keeping prompt history only in git log + reading `git blame` to recover rationale. `git log -p consts/<prompt>.go` becomes unreadable after the first few iterations; rebases and squashes lose the intermediate states; and the change reason — the operationally important field — was never in git in the first place.
46
+
47
+ ### Automated prompt/artifact optimization — optimizer output is an untrusted candidate behind an adoption gate
48
+
49
+ Beyond hand-iterating prompts, the prompts / skill files / tool descriptions / few-shot sets can be **treated as candidate text artifacts and optimized programmatically** — wrap the artifact in an optimizer + eval harness, let the optimizer propose variants and select against a metric. This can run at the prompt/instruction level with no model-weight training (API-cost only): prompt-optimizing frameworks (e.g. DSPy's prompt optimizers; trace-reflective evolutionary optimizers such as GEPA, which read the *execution traces* of failures to propose targeted mutations rather than mutating blindly). It is a real lever, but an optimizer maximizes its *metric*, not your intent — so an evolved variant is an **untrusted candidate** until it clears an adoption gate, and is never auto-deployed:
50
+
51
+ - **Improve on a held-out split, and don't burn the holdout across runs.** Split eval data into train / val / **holdout**; the optimizer sees only train/val; report the gain on the untouched holdout. A gain only on the set the optimizer searched against is metric-overfitting (Goodhart) — the holdout must never enter the optimization loop. Repeated *adoption attempts* against the same holdout also consume it: run 30 variants and report only the winner's holdout gain and you have silently turned the holdout into validation (multiple-comparison). Track run/version counts and rotate or refresh the holdout; this is the anti-contamination rule below applied to optimization, not just to measurement.
52
+ - **Pass a behavioral gate, not just the score or the code tests.** The variant must pass the owning repo's full test/eval suite AND a prompt-**behavioral** regression suite — code/schema/parser tests do not exercise prompt behavior, so a candidate can be green while it changes refusal behavior, tool-choice propensity, verbosity, escalation, or data handling. Include negative and security cases and tool-use traces. Plus size/token **and growth** limits (an optimizer will happily 3× the prompt for +1%) and structural integrity (valid frontmatter; tool schema — param names/types — unchanged). Failed gate = reject regardless of fitness.
53
+ - **Gate on broad regression benchmarks the optimizer is NOT optimizing against — benchmarks are gates, not the fitness.** The fitness metric is task-specific (did *this* artifact do *its own* job better?); an optimizer maximizing it can silently regress unrelated capabilities or multi-turn coherence (Goodhart, one level up from the held-out point). Boundary so these don't collapse into "one eval suite": held-out catches same-task overfit, the behavioral suite catches *this* artifact's own contracts, the broad benchmark catches *other* capabilities. Run a broad capability/regression benchmark and a long-horizon / multi-turn coherence check as **separate veto gates** — with care:
54
+ - **Significance, not raw percent.** "Regressed −Y%" is a decision only with predeclared statistical handling (paired inputs, minimum N, CI/bootstrap, per-slice floors); a small delta is *inconclusive → rerun / enlarge sample*, not an automatic pass or reject (per Eval Reliability below).
55
+ - **Don't let the cheap tier become a second fitness.** Tier for cost — a cheap fast subset on every candidate — but use it as an early veto / stratified screen only, never as the finalist *ranking* objective (the optimizer will Goodhart the subset and breed cheap-set specialists). Pick finalists on primary fitness + slice coverage/diversity + predeclared constraints, and promote some diverse/random challengers so robust candidates aren't filtered out before the real gate; run the full benchmark + coherence check on those finalists.
56
+ - **The coherence gate needs a run contract, or it gets skipped.** Give the long-horizon check an owner, budget, timeout, sample shape, and evidence artifact; if it can't run, the variant is `blocked` / `deferred` with residual risk — never silently "passed" because the cheaper gates ran (per the blocked-verification discipline).
57
+ - **Benchmark, cheap subset, and coherence suite are eval assets** — version / use-count / rotate them (the cheap subset, run on every candidate, contaminates fastest); the anti-contamination rule below applies to them too.
58
+
59
+ A variant that improves fitness +X% but *significantly* regresses a broad gate is rejected — net-negative is not an improvement, and held-out-on-the-same-task does not catch a cross-capability regression. (This is the optimization-loop form of the layered-acceptance / "launch follows the product baseline, not the best component metric" rule — see `retrieval-agent-safety.md` composite-chain acceptance and `testing-strategy`'s composite-pipeline rule.)
60
+ - **Security review is a mandatory gate, not an afterthought — the evolved text is a durable injection surface.** An optimizer maximizing a helpfulness metric will, unprompted, weaken a refusal boundary, add a prompt-injection affordance, ask the model to reveal hidden/system instructions, widen a data egress, or soften a safety caveat — and that text passes score + similarity + tests, then persists in a shipped artifact. Every candidate diff must clear an adversarial security review / scan for instruction-hierarchy violations, jailbreak/override language, secret-exfiltration patterns, and safety/data-boundary weakening (owned by `references/retrieval-agent-safety.md` safety + output rules); never let fitness substitute for it.
61
+ - **Trace-derived eval data must be sanitized before it enters the optimizer, and the candidate diff audited for leakage.** When the eval/optimization set is built from real session traces, an optimizer (especially a trace-reflective one) can copy a customer name, token, internal URL, proprietary rubric, or prompt fragment straight into the "better" artifact — a durable PII/secret leak, worse than ordinary eval leakage. Only minimized/redacted traces may enter optimization; the candidate diff must pass a secret/PII/provenance-leakage audit before the PR (memory/trace hygiene per `references/retrieval-agent-safety.md`; trajectory-dataset hygiene per `inference-capacity-operations.md`).
62
+ - **Contract preservation, not just semantic similarity.** Similarity against the original is auxiliary, not sufficient — it misses small critical edits (removing a negation, "do not expose secrets" → "expose secrets when an admin asks", narrowing a safety caveat, adding a higher-priority exception). Clause-diff the must/shall/never rules, safety caveats, data boundaries, tool preconditions, and escalation rules against the original; "better at a different job" or "same words, flipped a critical clause" is a regression. For tool descriptions this is its own gate: the schema may be fixed but the *description* still drives when the model calls the tool, what it sends, whether it honors caveats, and whether it overuses a privileged tool — preserve invocation boundaries, allowed/disallowed use cases, safety caveats, and required-confirmation language.
63
+ - **Cache-stable adoption — but not for an authority/safety change.** An authority-neutral, cost-only evolved artifact (a clearer instruction, a better few-shot) deploys as a new version effective on the **next fresh session**, never hot-swapped mid-conversation — it lives in the cached prefix. But when the evolved change tightens a safety/refusal boundary, narrows a privilege, changes a tool's invocation boundary, or fixes a privacy/authorization surface, deferring it to "next session" leaves an active conversation running the stale, less-safe surface: those follow the immediate **invalidate / fence / re-authorize now** path, never deferred for cache stability (see the prefix-mutation rule in `llm-client-gateway.md`, which already splits authority-neutral churn from authority/safety/capability changes). If it can't be applied safely mid-conversation, hold or disable the affected route.
64
+ - **Human PR with the right reviewers, never direct commit.** Land via a PR carrying before/after on train/val/**holdout**, the diff, the run cost, and the constraint violations caught-and-rejected during the run; for high-risk artifacts (system prompts, safety/refusal text, privileged tool descriptions) the security/privacy/eval owner must be a required reviewer — a human rubber-stamping the fitness delta is not the gate. Risk-gate per `feature-risk-router` / `product-rd-workflow`. The optimizer proposes; a human merges. When the evolved artifact is a **CCL skill/reference** (this section lists skill files as optimizable), this gate owns only the optimizer/eval/adoption mechanics — *landing* the change still goes through `skill-extraction-workflow`'s shared-skill gates (charter, R0 leakage audit, dual-track review + challenge, owner map); an optimizer's adoption gate does not replace them.
65
+
66
+ The fitness function may be a scalar metric, unit tests / validators / schema checks, or LLM-as-judge (see Eval below); when it is LLM-as-judge it inherits all those reliability caveats (rubric drift, judge bias, position effects). A metric the optimizer can game is worse than no optimization.
67
+
68
+ ## Evaluation
69
+
70
+ Every material model or prompt change should define:
71
+
72
+ - dataset or replay source;
73
+ - expected output, rubric, or comparator;
74
+ - quality metrics such as accuracy, recall, consistency, parse success, or groundedness;
75
+ - runtime metrics such as latency, token usage, success rate, fallback rate, and cost;
76
+ - regression examples and human review notes when judgment is subjective.
77
+
78
+ Store evaluation reports with the model/prompt versions compared so future changes can reproduce the decision.
79
+
80
+ ## Eval Reliability
81
+
82
+ - Keep a stable golden set for regression checks and a rotating challenge set for newly discovered failures.
83
+ - **Separate the fixed benchmark from the regression bad-case set, and govern both.** Mapping to the two sets named above: the **stable golden set IS the fixed benchmark** (answers "is the new version at least as good" under one stable comparison), and the **rotating challenge set is where the regression bad-case set accumulates** (real failures that must never recur). Track each set's version, source, time window, sample count, slice distribution, labeler, and labeling rule.
84
+ - **Anti-contamination is the load-bearing rule: a set that has been used to train, fine-tune, tune prompts, or develop hand rules is contaminated and can no longer measure generalization — it must NOT serve as the fixed benchmark.** Benchmark/eval contamination is a top cause of inflated, misleading scores in LLM evaluation; reported gains evaporate in production because the model memorized the test. When reusing public or prior data, screen for leakage (exact-string / n-gram overlap / hash / embedding-similarity checks) before trusting a set as held-out, and prefer a benchmark refreshed/rotated over time so it does not ossify into a memorized target. Contamination is not only training: repeated inspection, manual error analysis, prompt iteration, threshold tuning, hand-rule writing, relabeling, or cherry-picking against a set all contaminate it for benchmark use ("we only looked at it" is still leakage). A contaminated set is **reclassified, retained, and versioned** as a dev / training / challenge asset — not deleted — and a fresh held-out benchmark is cut. Few-shot examples are allowed, but only sourced from non-benchmark / dev data, never from the fixed benchmark.
85
+ - **Role separation for high-risk evals**: the party optimizing the model/prompt should not also own or silently edit the benchmark set — otherwise the bar quietly bends to pass. Benchmark changes (add/remove/relabel) go through a recorded change with reason, author, and reviewer, kept auditable; keep the optimizer and the benchmark-owner roles distinct where the launch decision is high-stakes. Small-team fallback: one person may both optimize and maintain the benchmark only if benchmark edits are append-only or asynchronously reviewed by another accountable person, and high-risk launch decisions still carry a recorded independent review.
86
+ - **Continuous re-injection (the eval set is a living asset, not a one-time artifact)**: online failures, user-flagged bad outputs, and human-review findings feed back into the regression bad-case set, which runs every release. Tier the run so the gate stays affordable and honest: deterministic frozen cases run in the blocking release gate; cases that need live model calls / real retrieval index / integration run in the release / pre-ramp gate with an explicit marker, owner, and timeout; human-review-only cases produce release evidence and are NOT mislabeled as automated tests. Any **material change** re-runs the core eval before re-ramp — a passing eval from before the change does not transfer. Material = any change that can affect input distribution, retrieved context, model behavior, output schema/parser, tool availability, fallback/safety policy, or dependency/provider behavior (model version, prompt, retrieval, index, knowledge base, tool set, third-party service, or a provider default-behavior shift all qualify). If unsure, treat it as material; record any non-material classification with owner, reason, and affected surface.
87
+ - Track sample size, dataset slice, evaluator version, and run timestamp with every report.
88
+ - For LLM-as-judge, version the judge prompt/model, calibrate against human-reviewed examples, and watch for judge drift.
89
+ - Prefer paired comparisons on the same inputs when comparing model or prompt versions.
90
+ - Report both aggregate scores and concrete regression examples; aggregate-only evals hide product risk.
91
+ - Treat small score changes as inconclusive unless variance and sample size justify the decision.
92
+
93
+ ## Production Output Review (Human-In-The-Loop Sampling)
94
+
95
+ Offline eval proves a version at launch; it does not catch drift after launch. Continuous full-coverage human review does not scale, so sample deliberately — but sampling strategy must be risk-weighted, not uniform.
96
+
97
+ - **Baseline**: review a small random fraction of production outputs (a common starting range is on the order of ~1–5%, or a fixed quota such as N traces/week — tune to volume and risk) for an unbiased quality view. Random-only is insufficient — it under-samples rare, novel, and high-impact failures.
98
+ - **Priority + stratified on top of random**: actively pull outputs that users flagged negative, that an LLM-as-judge scored low, or that are novel/uncertain; stratify so each trace type, user segment, and risk class is covered, not just the high-volume happy path.
99
+ - **High-risk output classes are not random-sampled.** For classes where a single wrong output is severe — financial-fact error, mistranslation, safety/red-line answer, an irreversible user action triggered by model output — define a dedicated review rate (up to 100% for the top tier), a processing time-limit, and a named owner. An LLM-as-judge may **prioritize, annotate, or route** these items — it must NOT exclude a high-risk item from the human-eligible queue, or the worst false negatives never reach a human. Keep an independent random/stratified human sample drawn from ALL high-risk outputs (not only judge-flagged ones), audit the count of judge-dropped cases, and run periodic false-negative spot checks. For these classes the human is the final authority on legal / brand / reputational / safety calls.
100
+ - **Review budget + degradation path** (so "up to 100%" does not become an ignored ideal at volume): set a queue cap, an SLA, a sampling floor, and an escalation owner. If full review of the top tier is infeasible at the current volume, the honest responses are to reduce launch scope or obtain explicit product/compliance approval for the residual risk — never silently drop the high-risk review rate to zero.
101
+ - **Close the loop**: every confirmed bad case from this review feeds the regression bad-case set in Eval Reliability above, so the same failure is caught automatically next release. Review without re-injection is wasted signal.
102
+ - The launch/quality gate (which risk classes need this review, and the pass bar) is owned by `testing-strategy` (high-risk resilience) and `product-rd-workflow` (launch gate); this section owns the inference-side sampling + re-injection mechanics.
103
+
104
+ ## Replay And Shadow
105
+
106
+ Use replay when existing request logs can safely represent expected traffic. Use shadow when candidate output should be generated alongside production without affecting users.
107
+
108
+ For replay/shadow comparison:
109
+
110
+ - filter sensitive or disallowed traffic;
111
+ - preserve enough request context to reproduce behavior;
112
+ - define comparator thresholds before running;
113
+ - store diff examples, not only aggregate scores;
114
+ - require explicit activation after comparison passes.
115
+
116
+ Do not copy business-specific comparator labels into new products. Reuse the mechanics: input capture, candidate execution, scoring, diff retention, and rollout gate.
117
+
118
+ ### Freeze Comparison Criteria Before The Run (No Backfill)
119
+
120
+ - Define the comparator, success thresholds, sample scope, observation window, and stop rules **before** executing replay/shadow/A-B/canary. Reading the online or candidate result first and then choosing the success criterion the result happens to clear is result-fitting, not evaluation — it invalidates the comparison. This is the same discipline as the anti-contamination rule above, applied to the success bar rather than the dataset.
121
+ - **The freeze blocks moving the bar to PASS; it does not suppress a newly-discovered failure.** If a run surfaces a severe failure class the frozen rubric did not anticipate, that finding is not ignored because it was not predeclared — it becomes a launch blocker through a documented amendment plus a re-run on a fresh or re-scored sample. Forbidden: relaxing a predeclared criterion the result just missed. Required: adding a newly-found severe-failure blocker and re-validating.
122
+ - A/B experiment and shadow/canary rollout answer different questions and their conclusions are not interchangeable. A/B measures *effect difference* (is the new version better on the business/quality metric) and needs a pre-declared randomization, sample size, and stop rule. Shadow/canary controls *rollout risk* (does the new version break under real traffic) and needs a pre-declared ramp step, pause condition, rollback condition, and fallback version. Do not present a risk-control ramp as proof of effect, or an effect experiment as proof of operational safety. A combined design that measures effect while guarding rollout is valid only when randomization, power/stop rules, ramp gates, rollback, and safety monitoring are all predeclared together.
123
+ - **A continuously-monitored *statistical* comparison needs a sequential design, not repeated peeking at a fixed-horizon test.** This applies to inferential effect/quality comparisons (is the new version significantly better/worse). "Freeze before run" does not mean "look only once" — but checking a fixed-sample-size test repeatedly and stopping as soon as it looks significant (the peeking problem) inflates the false-positive rate far above the nominal level. So for an effect/quality test that is watched continuously, the stopping rule must be a pre-declared sequential method (alpha-spending or an anytime-valid confidence sequence) sized for repeated looks — not a fixed-horizon p-value glanced at every hour. Declare up front whether the comparison is fixed-horizon (analyze once at N) or sequential (continuous, alpha-adjusted); do not run fixed-horizon and then peek. This is NOT required for pure operational/SLO safety gates (error rate, latency, cost) — those use predeclared thresholds with minimum volume, persistence, and hysteresis (see `inference-capacity-operations.md`), not a hypothesis test.
124
+ - The A/B-vs-canary rollout-strategy choice is owned by `product-rd-workflow` and `platform-release-engineering`; this section owns the freeze-before-run / no-backfill comparison mechanic.
125
+
126
+ ## Offline Evaluation Pipelines For Multi-Stage Inference
127
+
128
+ When a production inference system spans multiple stages (e.g. OCR → matching → scoring) and evaluation needs end-to-end + per-stage signals:
129
+
130
+ - **Pipeline as ordered scripts**, not just one notebook: `1_download_inputs.py` → `2_run_candidate.py` → `3_evaluate_metrics.py` → `4_emit_badcase_report.py`. Each script is idempotent on its own checkpoint; rerunning step 3 does not require redownloading.
131
+ - **Database-backed inventory**: store the dataset row, the candidate output, and the evaluation metric in a relational store. The badcase report is a SQL projection, not a runtime artifact. Crashes mid-pipeline resume from the row's last-completed column.
132
+ - **Skip-on-cache + strict-mode override**: the download/run scripts should detect prior completion via a stored marker column (`regulated_local_path`, `candidate_response_hash`) and skip; a `--strict` or `--force` flag re-runs from scratch when validating reproducibility.
133
+ - **Badcase HTML report** is the human review surface; it should bundle input image / text, ground truth, candidate output, diff highlight, and a link back to the row id so reviewers can update the dataset. **Treat all model-emitted text as untrusted when rendering**: escape HTML entities on every model output before embedding into the report (`html.escape` in Python, or a templating engine with auto-escape on); render images via `<img>` with `Content-Security-Policy` or a sanitized origin so an OCR output containing `<script>` does not become stored XSS in the reviewer's browser. Reviewers open badcase reports from many sources; a sanitized report is the safe default.
134
+ - **Scheduling**: cron or manual `python pipeline/N_step.py` is acceptable for evaluation pipelines (not user-facing); event-driven triggering is over-engineering when the input is a frozen dataset. Document the cadence. **Add a run lease so cron + manual cannot overlap**: each pipeline step acquires a per-step + per-dataset lease (database row lock, Redis `SET NX EX`, or filesystem flock) before doing work and releases it on terminal state. Two concurrent runs of the same step on the same dataset must result in one waiting / skipping, not double-running. Per-row idempotency lives below the lease: each dataset row carries a `processed_at` / `result_hash` so a partial-completion replay only retries the unfinished tail.
135
+ - **Pipeline as evaluation artifact**: every model/prompt change cites the pipeline run id and links to the badcase report. Aggregate-only "F1 went up" reports without badcase examples hide regression risk.
136
+
137
+ ## Inference Quality Metrics Beyond LLM Defaults
138
+
139
+ Multi-modal inference systems (OCR, detection, structure recognition, matching, marking) need metrics beyond the LLM-defaults of accuracy / token cost:
140
+
141
+ - **Per-stage quality**: per-stage precision, recall, and F1 — pipeline-level F1 can mask a 30% recall drop in stage 3.
142
+ - **Per-class quality** when the model returns class labels: aggregate F1 is meaningless if one important class regressed.
143
+ - **Calibration metrics**: confidence score vs actual correctness — over-confident wrong outputs hurt downstream decisions even when accuracy is acceptable.
144
+ - **Throughput-quality trade-off**: per-batch latency vs per-batch quality; some optimizations (smaller batch, lower precision) gain throughput at quality cost.
145
+ - **Domain-specific rule pass-rate**: rule-based post-validators (output shape / range / cross-field constraint) catch issues that statistical metrics miss.
146
+ - **Cost per successful output** (not just per request): a model that retries to success costs more than its base latency suggests.