@ccoalm/ccl-skills 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -0
- package/README.md +49 -0
- package/dist/assets/marketplace/.agents/plugins/marketplace.json +12 -0
- package/dist/assets/marketplace/.claude-plugin/marketplace.json +13 -0
- package/dist/assets/marketplace/marketplace-manifest.json +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/marketplace.json +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/plugin.json +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.codex-plugin/plugin.json +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.worktree-only +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/subagent-start.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/AGENTS.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-delegation-owner.sh +125 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-edit-isolation.sh +102 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +1156 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/merge-authorization-prompt.sh +142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-guard.sh +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-stop.sh +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +144 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-context.sh +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-start.sh +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-extraction-gate-stop.sh +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/subagent-start.sh +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_delegation_owner.sh +329 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_edit_isolation.sh +322 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +902 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +178 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +121 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_session_start.sh +170 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/AGENTS.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +564 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-update-skills.md +44 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-verify-skills.md +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-worktree-check.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/AGENTS.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/README.md +276 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.example.json +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +1307 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +941 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/SKILL.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +188 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/android-dev.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/flutter-dev.md +80 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/ios-dev.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/kotlin-multiplatform.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-platform-boundaries.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/source-evidence-map.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +353 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +98 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_timeout_exit.sh +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +1438 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +324 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/concern_excerpt.py +295 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/egress_schema.py +214 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +642 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +181 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +1165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +1190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +946 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_opencode_review.py +474 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +1899 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_review_json.py +200 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +2845 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.sh +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/run_claude_capture.py +71 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +68 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +2311 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +1832 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_code_review_identity.sh +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_concern_excerpt.sh +245 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_egress_schema.sh +177 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +272 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +195 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_concurrency.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_retry.sh +1005 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_opencode_review.sh +258 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +574 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +349 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +434 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +264 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +2412 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/verify_native_skill_binding.py +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +153 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +54 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/prevention-routing.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/references/security-review-gate.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/api-security-boundaries.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +160 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/artifact-generation-architecture.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/audit-history-architecture.md +29 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/bulk-workflow-architecture.md +33 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/config-rule-routing-architecture.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-modeling-and-migrations.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +210 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/dependency-platform.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/developer-tooling-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/error-contract-architecture.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +260 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/http-gateway-architecture.md +74 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/mq-consumer-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +275 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/notification-architecture.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/ops-checklist.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/performance-capacity-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/protobuf-contract-architecture.md +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/redis-cache-coordination.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/release-runtime-readiness.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/replay-comparison-architecture.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/runtime-observability.md +94 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/service-scaffold.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/source-evidence-map.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/workflow-state-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +159 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/artifact-generation-patterns.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/audit-history-patterns.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/bulk-import-export-patterns.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/config-rule-routing-patterns.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/data-access-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/db-schema-and-dal-patterns.md +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/dependency-client-patterns.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/developer-tooling-patterns.md +70 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/domain-feature-patterns.md +78 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/engineering-patterns.md +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/error-contract-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/feature-playbook.md +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/http-gateway-client-patterns.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/mq-consumer-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/notification-patterns.md +42 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/observability-implementation-patterns.md +101 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/performance-capacity-patterns.md +44 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/protobuf-contract-patterns.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/public-api-integration-patterns.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/quality-and-testing-patterns.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/redis-cache-lock-patterns.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/release-ops-patterns.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/reliability-patterns.md +83 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/replay-comparison-patterns.md +32 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/scaffold-and-codegen.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/source-evidence-map.md +54 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +80 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +117 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-approval-auto-reviewer.md +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +441 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-context-freshness.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-credentials-auth.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-extensions-skills.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-file-edit-protocol.md +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-ide-integration.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-input-ingestion.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-instruction-composition.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-lifecycle-hooks.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-messaging.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-runtime-bootstrap.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +448 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-task-orchestration.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-turn-lifecycle.md +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +162 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +156 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/retrieval-agent-safety.md +273 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +202 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/cross-stack-alignment.md +94 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/framework-choice.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/online-practice-uptake.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/platform-capabilities.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/product-page-checklist.md +40 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/qa-release.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/source-evidence-map.md +82 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +103 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/agents/openai.yaml +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +100 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +70 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +549 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +97 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/research-prompts.md +66 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/AGENTS.md +32 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/test-public-data-acquisition-recipes.sh +379 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +244 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/framework-middleware-checklist.md +142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/infra-component-deployment.md +268 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-correlation-recipe.md +124 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-schema-canonical.md +208 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/obs-stack-architecture.md +107 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +95 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +303 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +163 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/config-center-via-etcd.md +245 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/custom-control-plane-boundary.md +298 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-cli-concrete-recipe.md +312 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-pipeline.md +165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/env-and-lane-matrix.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/lane-orchestration-control-plane.md +383 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/multi-region-and-cluster.md +135 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +149 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/python-package-registry-release.md +462 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/rollback-playbook.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +231 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/version-authority-and-deprecation.md +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +276 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +127 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/framework-middleware.md +143 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +90 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/http-response-envelope-contract.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +127 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/multi-env-routing.md +192 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/protobuf-http-contract-signals.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +124 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/rpc-framework-recipe.md +494 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-choice.md +113 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-migration-playbook.md +231 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +235 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/adr-convention.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-checklist.md +30 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-evaluation-report-template.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-execution-spec.md +108 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-sop.md +457 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-templates.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/artifact-egress-confidentiality.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/cross-repo-coordination.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +192 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/diagnostic-spec-match-gate.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dormant-code-activation.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/existing-project-assessment-report.md +223 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/external-skill-augmentation.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/feature-deprecation-cascade.md +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/high-risk-resilience-gates.md +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-completeness-and-minimality.md +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-entry-reentry-gate.md +122 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/modular-monolith-heuristic.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +115 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-attributes.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-remediation-program.md +88 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +27 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +52 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/review-reception.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/shared-gate-artifact-classification.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/source-evidence-map.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/status-tracker-sync.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/sync-spec-repo-contract.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/worktree-mechanics.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/check-agent-contract-coverage.sh +213 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +136 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/agents/openai.yaml +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/analytics-visualization-interactions.md +206 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +108 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/complex-creation-interactions.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +214 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +97 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +63 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +250 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +324 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-web-desktop-patterns.md +456 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/resource-management-interactions.md +113 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/scenario-community-patterns.md +133 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/trust-sensitive-ai-and-data-patterns.md +96 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +111 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +157 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/ai-service-integration-boundaries.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-contract-and-schema.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-security-boundaries.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/async-execution-model.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/background-jobs-and-scheduling.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/batch-and-pipeline-architecture.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/config-secrets-runtime.md +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-modeling-and-migrations.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +211 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +263 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +281 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/observability-and-ops.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/redis-cache-coordination.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/reliability-and-error-contract.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/source-evidence-map.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/web-framework-boundaries.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +143 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/async-and-worker-patterns.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/dependency-client-patterns.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/error-handling-patterns.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/feature-playbook.md +43 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/observability-implementation-patterns.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/project-structure-and-tooling.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +52 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/redis-cache-lock-patterns.md +78 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/schema-and-validation-patterns.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/source-evidence-map.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/sqlalchemy-and-migrations-patterns.md +99 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/web-framework-patterns.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/config-runtime-readback.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/mr-merge-authorization.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/post-release-env-reset.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-closeout-evidence.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-scope-confirmation.md +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/test-scope-prompt.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/watcher-discipline.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/SKILL.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/comment-safe-release-doc.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-evidence-workflow.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-testing-scope-section.md +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/prd-composition-contract.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/requirement-closure-contract.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/security-four-questions.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/SKILL.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +88 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +337 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/analysis-parse-fix-test-challenge-replay.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attribution-verification.md +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/bootstrap-slim-c3-obligation-table.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +162 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +507 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/evidence-card-template.md +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/example-domain-preselect.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-lifecycle-handoff.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +286 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/l0-l1-l2-routing.md +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/online-skill-review.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +164 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +90 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +320 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-feedback-mining.md +33 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-finding-standards.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-rubric.md +40 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +118 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/skill-listing-budget.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +254 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +658 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/two-source-extraction-pattern.md +167 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-routing-map.md +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +180 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +1452 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-evidence-card-leak.sh +491 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-mr-target-freshness.sh +173 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +488 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +327 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +401 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing.rb +248 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/generic-r0-leak-scan.sh +282 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/governing-chain-diff.py +321 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +964 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +708 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +540 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-lifecycle.rb +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-pending-status.rb +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +829 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +1203 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_r0_status.sh +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_register_pending_exclusion.sh +137 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +173 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_route_drift.sh +377 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +833 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +491 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_mr_target_freshness.sh +261 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +538 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +154 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_surface_binding.sh +178 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_prose_target.sh +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_generic_r0_leak_scan.sh +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_git_identity_predicate_gate.sh +243 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_governing_chain_diff.sh +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +724 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +414 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_registration.sh +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +205 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_credential_cwd.sh +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +111 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +257 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +98 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/input-state-machines.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/streaming-rich-output.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/terminal-side-channels.md +96 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +408 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/bitable-setup.md +573 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/README.md +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/github-actions.yml +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/gitlab-ci.yml +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/jenkins.Jenkinsfile +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +279 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/gen_report.py +2807 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/makefile-template.md +200 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/report-config-schema.md +272 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/run_pytestless.py +475 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/source-to-case-workflows.md +258 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-marker-conventions.md +316 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +145 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/AGENTS.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.dart +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.go +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.py +135 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.ts +285 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/test_gen_report.py +2144 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +212 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +50 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/data-and-workflow-testing.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/design-closed-contract-oracles.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +71 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +240 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +235 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/non-functional-specialized-scenarios.md +296 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/rd-testing-standard-template.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/run-killing-mutation-walk.md +43 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +136 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/source-evidence-map.md +59 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/structured-tc-input-translation.md +67 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +392 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-data-and-determinism.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/unit-testing.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/vendored-contract-drift-checklist.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/verify-enforcement-mechanisms.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/AGENTS.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.py +140 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.test.sh +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.py +170 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.test.sh +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.go +198 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.test.sh +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/test_mutation_backup_recipe.sh +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +184 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/comment-safe-feishu.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/cross-model-co-review.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +60 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/session-vantage-leakage.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/embedded-h5-in-host.md +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/source-evidence-map.md +60 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +83 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/SKILL.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/shared-branch-rebase.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/AGENTS.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_status.sh +207 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_sweep.sh +481 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-status.sh +325 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-sweep.sh +245 -0
- package/dist/assets/release.json +2797 -0
- package/dist/claude-adapter.d.ts +9 -0
- package/dist/claude-adapter.js +240 -0
- package/dist/cli-worker.d.ts +1 -0
- package/dist/cli-worker.js +32 -0
- package/dist/cli.d.ts +22 -0
- package/dist/cli.js +214 -0
- package/dist/codex-host.d.ts +30 -0
- package/dist/codex-host.js +162 -0
- package/dist/fs-safe.d.ts +21 -0
- package/dist/fs-safe.js +241 -0
- package/dist/index.d.ts +2 -0
- package/dist/index.js +1 -0
- package/dist/manifest.d.ts +8 -0
- package/dist/manifest.js +135 -0
- package/dist/opencode-adapter.d.ts +10 -0
- package/dist/opencode-adapter.js +416 -0
- package/dist/operations.d.ts +3 -0
- package/dist/operations.js +956 -0
- package/dist/paths.d.ts +20 -0
- package/dist/paths.js +4 -0
- package/dist/types.d.ts +58 -0
- package/dist/types.js +1 -0
- package/dist/unified.d.ts +4 -0
- package/dist/unified.js +64 -0
- package/dist/version.d.ts +2 -0
- package/dist/version.js +5 -0
- package/package.json +35 -0
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Agent tool-dispatch framework
|
|
2
|
+
|
|
3
|
+
Reusable mechanics for the **tool layer** of an agent runtime: how a model's requested tool calls are registered, discovered, routed to handlers, executed (including in parallel), and shaped back into model-visible results. This is the registry/router/execution plumbing — distinct from the policy engine that decides *whether* a given execution is allowed.
|
|
4
|
+
|
|
5
|
+
Use this when building or reviewing an agent's tool system. It sits beside:
|
|
6
|
+
- `agent-turn-lifecycle.md` — owns the loop that *calls* the dispatcher and consumes results.
|
|
7
|
+
- `agent-command-sandbox.md` — owns approval → sandbox → escalation for an individual command/exec tool. The dispatcher delegates the *enforcement* of a tool call to that engine; do not re-implement it here.
|
|
8
|
+
- `agent-lifecycle-hooks.md` — owns the pre/post-tool hook points the dispatcher fires around each call.
|
|
9
|
+
|
|
10
|
+
## Registry and the tool contract
|
|
11
|
+
|
|
12
|
+
Define one runtime contract every locally-executed tool implements. Keep it small and uniform:
|
|
13
|
+
|
|
14
|
+
- **execute** — run the call, return a typed result.
|
|
15
|
+
- **kind match** — which payload shapes this tool accepts (plain function call, search-surfaced call, etc.).
|
|
16
|
+
- optional metadata the runtime needs: telemetry tags, post-tool-hook payload shaping, tool-search descriptor, argument-diff consumer.
|
|
17
|
+
|
|
18
|
+
A uniform contract is what lets the router treat first-party tools, MCP/remote tools, and extension tools identically. Normalize tool **names** to one canonical form at the boundary (flatten namespacing) so dispatch, telemetry, and hook-matching all key off the same string.
|
|
19
|
+
|
|
20
|
+
## Tool discovery: static vs dynamic (searchable) tools
|
|
21
|
+
|
|
22
|
+
Not every tool should be in the model's context at once — large tool sets blow the context window and degrade selection. Split the active set:
|
|
23
|
+
|
|
24
|
+
- **Static** tools: always present in the tool spec.
|
|
25
|
+
- **Dynamic** tools: surfaced on demand via a **tool-search** entry. The model is given a lightweight search/discovery tool; matching tools are then injected into the active set only when needed. Track each tool's **origin** (static vs dynamically surfaced) so you can reason about why a tool is callable, and define its **eviction policy explicitly** (e.g. stays for the rest of the turn, LRU cap on the dynamic set, or explicit removal) rather than leaving lifetime implicit.
|
|
26
|
+
|
|
27
|
+
This is the standard answer to "I have hundreds of tools": expose a searchable index, load schemas lazily. The dispatcher must accept a call to a tool that entered the set dynamically exactly as it would a static one. Keep the searchable index **curated/trusted**, not built from user- or content-supplied free text — a poisoned index could surface a malicious tool to the model.
|
|
28
|
+
|
|
29
|
+
## Routing
|
|
30
|
+
|
|
31
|
+
The router maps an incoming tool call (name + call id + arguments payload) to the registered handler:
|
|
32
|
+
|
|
33
|
+
- Resolve by canonical name; a truly unknown tool is a typed `function_call_error`, never a panic — return it to the model so it can correct, don't crash the turn. **Distinguish "unknown" from "known but currently inactive"**: a tool that was dynamically evicted mid-turn is a race, not a model mistake — re-surface/re-activate it and dispatch, rather than telling the model it doesn't exist.
|
|
34
|
+
- Carry the **call id** through end to end; results must be correlated back to the specific call (critical for parallel execution and for tool-call↔result pairing in history). **Detect duplicate call ids within a turn** (a model can reuse an id across parallel calls); reject or disambiguate rather than mispairing results.
|
|
35
|
+
- Validate arguments against the tool's schema at the boundary; malformed arguments are a typed error result, not an exception.
|
|
36
|
+
- Distinguish a tool *execution error* (tool ran, failed) from a *dispatch error* (no such tool, bad args) — they read differently to the model.
|
|
37
|
+
|
|
38
|
+
## Parallel execution
|
|
39
|
+
|
|
40
|
+
When the model requests multiple tool calls in one response and the model/runtime supports it, execute them concurrently:
|
|
41
|
+
|
|
42
|
+
- Each in-flight call gets a child of the turn's cancellation token, so a turn abort cancels all of them. On abort, explicitly cancel **and join** each call's work (where the runtime cancels on drop, dropping the handle suffices; otherwise cancel-then-await) — a cancelled tool must stop, not detach and keep running.
|
|
43
|
+
- Share a single **turn-diff tracker** (or equivalent accumulator) across the parallel calls so concurrent effects are observed coherently, with synchronization on that shared state. **The tracker is not a substitute for resource-level safety**: two calls editing the *same* file/row concurrently corrupt or lose writes regardless of tracker locking. Run tools in parallel only across **disjoint resources**; serialize calls that touch the same resource (per-path/per-key lock) or detect and reject write conflicts.
|
|
44
|
+
- Results must be paired to their call ids and recorded in history in the **model's requested call order** (the order the calls appeared in the response), not completion order — completion order is nondeterministic and would make history non-reproducible.
|
|
45
|
+
- **Bound the parallelism** — don't fan out unboundedly across many tool calls (resource exhaustion). Respect the model's parallel-tool-calls capability flag; serialize when it's off or when tools declare ordering dependencies.
|
|
46
|
+
- Tools with side effects that may be interrupted/retried need the idempotency discipline from `agent-turn-lifecycle.md` (a cancelled-then-retried parallel batch must not double-apply).
|
|
47
|
+
|
|
48
|
+
## Result shaping and post-processing
|
|
49
|
+
|
|
50
|
+
- Convert each tool result into the model-visible output payload, and separately into the **post-tool-hook payload** (tool name, tool-use id, input, response) for observers.
|
|
51
|
+
- Apply post-tool hooks; let them inject additional context, treated as untrusted (see `agent-lifecycle-hooks.md`).
|
|
52
|
+
- Emit per-call telemetry (tool name, decision source, duration, outcome) keyed by call id.
|
|
53
|
+
- Bound result size fed back to the model; truncate large tool output **visibly and structure-aware** — truncating in the middle of structured (JSON) output yields unparseable downstream content. Use structure-aware markers or summarize, don't byte-chop.
|
|
54
|
+
|
|
55
|
+
## Code-mode: tools as a code/exec runtime (optional advanced pattern)
|
|
56
|
+
|
|
57
|
+
An alternative to many discrete tool calls: expose a single **exec** tool that runs model-authored code in a controlled runtime, where the code calls the other tools as nested in-process functions. Observed mechanics worth reusing:
|
|
58
|
+
|
|
59
|
+
- **Render each tool's JSON schema into the runtime's language types** (e.g. JSON-schema → TypeScript signatures) so the model writes correctly-typed calls against a familiar surface instead of emitting raw tool-call JSON.
|
|
60
|
+
- **Nested tool calls** from inside the exec runtime route back through the same router/registry — one dispatch path. Nested calls must still pass the **same pre/post-tool hooks and approval/sandbox gates** as top-level calls; an in-process nested call that skips them is a policy bypass.
|
|
61
|
+
- **Forbid or depth-bound re-entrant exec.** If the exec tool is itself callable from inside exec (`exec(exec(...))`), unbounded recursion blows the stack/process. Either exclude the exec tool from the nested tool set or cap recursion depth.
|
|
62
|
+
- **Yield/wait semantics**: long-running exec yields control (with a yield timeout) and can be resumed via a wait call, so the agent loop isn't blocked on a single long execution.
|
|
63
|
+
- **Per-exec output-token budget**: cap the tokens a single exec call can emit back to the model; large stdout must be truncated/bounded like any tool output.
|
|
64
|
+
- **Guard the host↔runtime numeric boundary for precision-critical integers.** When the runtime language has a narrower exact-integer range than the host (a JS/TypeScript runtime is exact only up to 2^53−1), precision-sensitive integers crossing the boundary — schema-projected id/count parameters, runtime-config fields like timeouts/token caps, and integer tool results — must be range-validated (on both inbound args and outbound tool results, before they enter the narrower runtime) or carried as a canonical decimal string or a typed int64/bigint where the transport and runtime support it (note bigint is not JSON-serializable; a free-form string invites later unsafe interpolation, so pin a canonical/branded form). A host integer past the runtime's safe range silently loses precision (and may wrap/truncate in bindings that coerce to narrower integer types), so the code executes against a *different value than the model intended*; reject out-of-range numeric inputs at the boundary rather than coercing them. (Ordinary floats like a temperature/probability map naturally to the runtime's double and don't need this — it's the exact-integer cases that bite.)
|
|
65
|
+
- **Don't assume fire-and-forget async survives a short-lived exec teardown.** If the exec runtime is short-lived and the host tears it down (or cancels the isolate) when top-level evaluation returns, any not-yet-awaited promise / pending task / timer may be cancelled or dropped — often with no model-visible error. Model code that fires off work without awaiting it (a write, a nested tool call, a flush) can lose that work. Document the teardown contract, and for effects that must complete, require the code to await them — or keep the cell alive (yield/wait) until they settle, but only *within the exec deadline / cancellation budget* (an await on a never-settling nested call must still be cut off by the section's wallclock/cancellation limits, not block the exec slot forever). (Longer-lived runtimes may instead keep the loop alive or surface an unhandled-rejection warning; the failure mode depends on the teardown contract, so state it explicitly.)
|
|
66
|
+
|
|
67
|
+
Code-mode runs **model-authored code** — it is not "a tool call that needs sandboxing", it is arbitrary code execution and demands the full treatment: the OS sandbox of `agent-command-sandbox.md` *plus* resource limits (CPU/memory/wallclock), egress control, and a per-nested-call policy check. Model code can otherwise loop forever, read secrets, or exfiltrate via a nested network tool. It trades more powerful composition (loops, conditionals, data flow between tools without a model round-trip) for that much larger attack/runtime surface — adopt it only when tool-call chaining is a real bottleneck.
|
|
68
|
+
|
|
69
|
+
## Anti-patterns
|
|
70
|
+
|
|
71
|
+
- Panicking on an unknown tool or malformed arguments instead of returning a typed error to the model (crashes the turn; model can't self-correct).
|
|
72
|
+
- Putting every tool in the model's context at once (context bloat, worse tool selection). Use static/dynamic split with tool search.
|
|
73
|
+
- Losing the call id across dispatch (results can't be paired; parallel execution corrupts history).
|
|
74
|
+
- Parallel tool futures that aren't cancelled+joined on abort (cancelled tools keep running; resource leak).
|
|
75
|
+
- Running tools that touch the *same* resource in parallel (lost writes/corruption; a shared diff tracker does not prevent this). Parallelize across disjoint resources only; lock or conflict-detect same-resource calls.
|
|
76
|
+
- Recording parallel results in completion order instead of the model's requested call order (non-reproducible history).
|
|
77
|
+
- Unbounded parallel fan-out, or ignoring the model's parallel-capability flag (resource exhaustion; calls the model can't handle).
|
|
78
|
+
- Sharing mutable cross-call state (diff tracker, history) without synchronization in the parallel path.
|
|
79
|
+
- Duplicate call ids within a turn going undetected (mispaired results).
|
|
80
|
+
- Conflating dispatch errors (no such tool / bad args) with execution errors (tool ran and failed) — the model needs to tell them apart.
|
|
81
|
+
- Re-implementing approval/sandbox logic in the router instead of delegating to the command-sandbox engine (drift between two enforcement paths).
|
|
82
|
+
- Adopting code-mode for its own sake when discrete tool calls suffice (extra runtime + sandbox surface for no real composition need).
|
|
83
|
+
- Treating code-mode as "just sandbox it" rather than full arbitrary-code-execution defense (resource limits, egress control, per-nested-call policy).
|
|
84
|
+
- Re-entrant exec with no depth bound (recursion blowup).
|
|
85
|
+
- Nested code-mode calls that skip the hooks/approval applied to top-level calls (policy bypass).
|
|
86
|
+
- Passing precision-critical host integers outside the runtime/transport safe-integer range without validation (silent precision loss → code runs on the wrong value).
|
|
87
|
+
- Assuming fire-and-forget async survives a short-lived exec teardown (it may be cancelled/dropped with no model-visible error; await it or hold the cell open).
|
|
88
|
+
- Building the dynamic tool-search index from untrusted/free-text content (poisoned tool surfaced to the model).
|
|
89
|
+
- Byte-truncating structured tool output (unparseable downstream). Truncate structure-aware or summarize.
|
|
90
|
+
- Feeding unbounded tool/exec output back into context (overflow; cost). Bound and mark truncation.
|
|
91
|
+
- Hardcoding a **cross-tool reference to a maybe-absent *concrete* tool in a tool's static schema description** — one tool's description naming another concrete callable, e.g. "prefer `<other_tool>` for X". When `<other_tool>` is absent in this deployment (disabled toolset, missing credential/API key, or simply not in the active / dynamically-surfaced set), the description steers the model to call a tool that isn't callable → a hallucinated call. Even with typed dispatch errors returned to the model for self-correction (per Routing above), it wastes a turn, degrades selection, and can leak through to the user. Fix: keep static descriptions self-contained, and add a concrete cross-tool hint only into the **rendered tool-definition snapshot for the turn, bound to the same tool-set / capability generation dispatch keys off** (not by mutating a description ad hoc between otherwise-identical turns — that churns the prompt cache; see `llm-client-gateway.md`), and only when the referenced tool is in the active set. Fine (not this anti-pattern): naming an always-present core tool, an atomically-bundled tool pair, or pointing at the tool-search/discovery capability for adjacent tools (a capability-level handoff, not a concrete absent callable). (The skill-design analog — a skill description referencing a maybe-uninstalled skill — is owned by `skill-extraction-workflow`'s "if installed, route to X; otherwise apply the principle inline" rule.)
|
|
92
|
+
|
|
93
|
+
## Persisted tool-output artifacts
|
|
94
|
+
|
|
95
|
+
Treat persisted tool-output artifacts as model-visible evidence substitutions, not ordinary temporary files or proof that the full result was read. When a tool, connector, resource reader, fetcher, or per-message budget path replaces raw output with a preview, saved artifact reference, typed marker, or read instructions, bind the artifact to principal/session/workspace, tool-call id, tool/source trust identity, schema or format label, content type class, original-size estimate, preview size and truncation state, output snapshot or digest where feasible, transcript replacement record, privacy/source-scope labels, permission/policy generation, and cleanup or retention policy; high-risk or mutable artifacts need an immutable digest, snapshot, or locked generation before replay, resume, or readback claims can rely on them. Previews, saved-path messages, and binary artifact notices are evidence of availability only; the agent must not summarize, analyze, or claim absence from the full output until it has read the needed chunks or explicitly states the unread portion. Large structured output needs a format/schema label and chunk/search strategy; binary output needs a conservative content-type to extension/viewer dispatch and a fallback when persistence, decoding, or viewer support fails. Empty tool results need a typed completion marker so model turns do not infer hidden content or stop ambiguity. Replacement decisions must be stable across resume, replay, prompt-cache reuse, and aggregate-budget enforcement: already-replaced results reapply the same model-visible marker, already-unreplaced results are not silently replaced later, and failed persistence falls back to a clearly labeled truncated or unavailable state rather than fabricating a full artifact. Persisted artifacts inherit the source tool's authorization, source labels, and untrusted-data status; converting output into a local file must not widen read permissions, erase connector/resource provenance, bypass compaction floors, or make raw output eligible for logs, telemetry, prompts, or user-visible summaries without redaction.
|
|
96
|
+
|
|
97
|
+
## Tool exposure vs tool authorization
|
|
98
|
+
|
|
99
|
+
Review tool exposure separately from tool authorization: the model-visible tool pool, deferred/searchable tools, external protocol tools, deny/allow/ask rules, and human or policy approval must be independently inspectable.
|
|
100
|
+
|
|
101
|
+
## Dynamic tool catalogs and external connectors
|
|
102
|
+
|
|
103
|
+
For dynamic tool catalogs or external connectors, verify prompt-cache partitioning by principal/session, tool source/server identity, manifest version, authorization policy version, and connector trust level; handle tool-name collisions, server/source trust boundaries, and capability-change reauthorization explicitly.
|
|
104
|
+
|
|
105
|
+
## External tool servers and connector channels
|
|
106
|
+
|
|
107
|
+
External tool servers and connector channels need a separate auth and consent contract before they become model-visible or approval-capable. Validate auth metadata schema and secure transport before use. Bind stored credentials and discovery state to server identity plus a canonical config digest, audience/resource identity, principal, account, tenant, organization, workspace, connector source, and auth-policy generation; when a product lacks one of those scopes, bind an explicit none marker instead of omitting it; never reuse credentials after server identity, config digest, URL, headers, source, audience, principal, account, tenant, organization, workspace, or auth-policy generation drift. Dynamic header helpers, environment expansion, channel bridges, and connector-provided prompts are untrusted executable or control inputs: gate workspace/local helpers on trust, bound runtime and output schema, redact secrets before logs or prompts, and fail closed for missing required credentials. Permission relay over a connector requires an active connection, explicit allowlist or policy grant, declared capability support for both conversation relay and permission relay, a structured pending request id, one-shot delete-before-resolve semantics, duplicate/unknown reply rejection, and final allow/deny binding to the original tool-call tuple including connector identity, server identity, canonical config digest, server URL, header/config digest, connector source, audience/resource identity, principal, account, tenant, organization, workspace, policy version, tool-call id, tool schema version, and normalized invocation digest. User elicitations from a connector need abort handling, hook result validation, completion notification binding to server plus elicitation id, and cancellation on malformed, late, or untrusted completion.
|
|
108
|
+
|
|
109
|
+
## Cached tool manifests and planned tool calls
|
|
110
|
+
|
|
111
|
+
Cached prompts, tool manifests, model outputs, or planned tool calls must never bypass fresh execution-time authorization.
|
|
112
|
+
|
|
113
|
+
## Remote permission/control responses as untrusted input
|
|
114
|
+
|
|
115
|
+
Treat remote permission/control responses as untrusted input: validate the discriminant and payload shape, bind each allow/deny/cancel response to the pending request id plus principal/account, session or transport incarnation, session generation, tool-call id, action/resource scope, policy version, tool source/server identity, manifest/tool schema version, and exact normalized invocation/argument digest, and reject stale responses after cancellation, account switch, reconnect generation change, policy change, or capability change.
|
|
116
|
+
|
|
117
|
+
## Tool execution side-effect authorization
|
|
118
|
+
|
|
119
|
+
Treat every tool execution as a side-effect authorization boundary. Bind allow, deny, ask, hook, classifier, and cached approval results to the exact pending tool-use id, normalized arguments, invocation digest, working directory, permission mode, policy version, principal/session identity, capability generation, tool source/server identity, and tool schema version. Re-run authorization after any hook-supplied input change, user edit, permission update, mode change, working-directory change, abort/cancel/retry, capability reload, or policy refresh; stale decisions must fail closed and must not be reused across a different invocation. The approving user or policy engine must see the final normalized invocation and resource set after all input transforms. Persisted approvals need narrowest-reusable scope, explicit revocation handling, and expiration or revalidation when policy, principal, capability, workspace, or tool schema changes.
|
|
120
|
+
|
|
121
|
+
## Search, glob, file-suggestion, and symbol discovery
|
|
122
|
+
|
|
123
|
+
Treat search, glob, file-suggestion, symbol, reference, definition, hover, and other code-discovery results as model-visible evidence control planes, not ordinary text output. Bind every discovery request to query or pattern digest, requested path or scope, type/mode/filter flags, principal/session/workspace, workspace or root-set generation, privacy/source-scope policy, tool-source identity, and capability generation before execution; run read permission and canonical containment checks before searching, indexing, or asking a code-intelligence server, and fail closed on network-path or unverifiable-containment inputs that could leak credentials or widen scope. Returned evidence needs source labels, normalized workspace-relative paths or bounded opaque labels, snapshot or index generation, freshness or partial-index state, worktree/source state where applicable, local/VCS/global/policy ignore-rule policy, hidden/binary/generated/vendor/plugin-cache exclusion policy, result mode, count, truncation, pagination offset, timeout/abort state, and whether deleted or inaccessible files were skipped during scan/stat reconciliation. A truncated, paginated, partial, timed-out, aborted, permission-filtered, stale-index, stale-server, or ignore-filtered discovery result must be labeled as incomplete and cannot support "not found", "only", "all", or absence claims without a complete fresh scan for the same tuple. File-suggestion indexes and cached discovery lists must invalidate or generation-fence after workspace/root, privacy, policy, ignore-file, source-control index, tracked/untracked file, settings, capability, or plugin/source change; background merges of untracked or newly indexed files may update suggestions only when the original cache generation still matches, otherwise discard as stale. Code-intelligence results require file snapshot/open-state proof, server initialization and generation binding, path/URI and line/column normalization, maximum file-size handling, ignored-result filtering for location queries, malformed-location rejection, and stale-server rejection after file, workspace, policy, or server-generation drift. Diagnostics may expose bounded categories, counts, incomplete-state markers, and non-reversible digests, but not raw local paths, sensitive project structure, query text, matched content, filenames, branch names, credentials, command arguments, or free-form server errors.
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
# Agent turn lifecycle & task supervision
|
|
2
|
+
|
|
3
|
+
Reusable mechanics for the layer **above** a single model call and **below** durable session storage: how an agent runtime supervises one in-flight unit of work (a "turn"/"task"), drives the sample→act→sample loop until the model stops asking for tools, and handles interruption, mid-flight steering, error taxonomy, and ordered finalization.
|
|
4
|
+
|
|
5
|
+
Use this when building or reviewing an agent runtime's turn driver. It is orthogonal to:
|
|
6
|
+
- `agent-session-persistence.md` — owns the append-only event log (the *what was recorded*); this file owns the *turn/task lifecycle* (the *how a turn runs and ends*).
|
|
7
|
+
- `agent-command-sandbox.md` — owns approval→sandbox→escalation for one tool execution; this file treats a tool call as an opaque step in the loop.
|
|
8
|
+
- `agent-tool-dispatch.md` — owns the tool registry/router/parallel execution; this file owns when the loop dispatches and how it consumes results.
|
|
9
|
+
- `agent-lifecycle-hooks.md` — owns the hook event surface; this file owns where in the lifecycle each hook is consulted.
|
|
10
|
+
|
|
11
|
+
## The turn/task model
|
|
12
|
+
|
|
13
|
+
Model the active unit of work as a supervised async task with a small, explicit contract:
|
|
14
|
+
|
|
15
|
+
- **identify** — a task kind for telemetry/UI (regular turn, review, compaction, user-shell, etc.).
|
|
16
|
+
- **run** — drives the turn to completion or cancellation; streams protocol events; returns the optional final assistant message.
|
|
17
|
+
- **abort** — cleanup hook invoked after cancellation (default no-op).
|
|
18
|
+
|
|
19
|
+
Keep the contract this small. Variant behaviors (review mode, plan mode, compaction-as-a-task) are different task *implementations*, not branches inside one mega-function.
|
|
20
|
+
|
|
21
|
+
### Single active turn invariant
|
|
22
|
+
|
|
23
|
+
Enforce at most one active turn per session. **Preempt-vs-queue is a policy choice**: a new turn can either cancel the in-flight one (a "replaced" preemption) or be serialized behind it. Whichever you choose, the invariant is that only one turn mutates history at a time — never two racing over shared history. Codex-style preemption (cancel-then-start with a distinct "replaced" reason) makes "user sends a new message while the agent is working" deterministic; a queue makes it lossless-but-delayed. Pick one deliberately.
|
|
24
|
+
|
|
25
|
+
- Hold the active-turn slot behind a lock; install a new task only via an atomic compare-and-set on that slot (check-then-act across two separate reads can double-start a turn under concurrent submitters / mailbox writers).
|
|
26
|
+
- Track per-turn state (token usage at turn start, tool-call count, citation flags) separately from the task handle so finalization can compute per-turn deltas.
|
|
27
|
+
|
|
28
|
+
### Cancellation as a token tree
|
|
29
|
+
|
|
30
|
+
Use a hierarchical cancellation primitive: session token → child per task → child per sampling request. Cancelling the parent cancels everything below; cancelling one sampling request leaves the task able to continue. Every long await inside the loop must be cancellation-aware (race the work against the token) so an interrupt takes effect within milliseconds, not at the next natural boundary.
|
|
31
|
+
|
|
32
|
+
## The sample→act loop
|
|
33
|
+
|
|
34
|
+
The core driver is a loop around a single model sampling request. Per iteration:
|
|
35
|
+
|
|
36
|
+
1. **Drain pending input** into history before building the next request. The turn's *originating* input must already be recorded in history before the first sample — that is the supervisor's job at turn start, not the drain's. This per-iteration drain is only for input that queued *during* the turn (mid-turn steering). *Defer* the drain in two cases so it doesn't fire when there's nothing new or it would split a sequence: (a) the first iteration, since the originating input is already in history and no mid-turn steering can exist yet; (b) immediately after a mid-turn compaction, so the model/tool continuation resumes before new steering is injected. Get this ordering wrong and you either re-inject the first message at the wrong point or interleave steering into a half-finished tool sequence.
|
|
37
|
+
2. **Build the request** from current history projected for the model's input modalities.
|
|
38
|
+
3. **Run one sampling request** (its own retry sub-loop, see below).
|
|
39
|
+
4. On success, decide **follow-up**: `needs_follow_up = model_requested_tools_or_continuation OR there_is_pending_substantive_input`. Pending input matters — user steering that arrived mid-turn forces another sampling pass even if the model thought it was done. Exclude pure *control* inputs (cancel/stop/interrupt) from this condition: a queued stop must end the turn, not trigger a fresh model call.
|
|
40
|
+
5. If a context-token limit is reached *and* follow-up is needed, run **mid-turn compaction** before continuing — but only at a protocol-safe boundary, never across an open tool-call transaction (a model message that requested tool calls whose results are not yet all recorded). Compacting mid-transaction can drop or reorder the tool-call/tool-result pairing and corrupt the next request. Termination relies on compaction dropping well below the limit; enforce it, don't assert it. Cap compaction retries within a turn, and if a single irreducible message (one user/tool item larger than the floor) still exceeds the limit after the cap, fail the turn with a typed context-overflow error rather than looping. See `agent-session-persistence.md` compaction rules.
|
|
41
|
+
6. If no follow-up: run **stop hooks**. A stop hook may *block* completion (inject a continuation prompt and loop again — guard against a block with no prompt) or *allow* it. Then run any legacy after-turn hook, then break.
|
|
42
|
+
7. Otherwise `continue` to the next sampling request.
|
|
43
|
+
|
|
44
|
+
**Bound the outer loop.** A stop hook that always blocks, or a source that continuously re-queues pending input, can spin the loop forever. Cap total turn iterations (and, separately, consecutive stop-hook continuations); on cap, break with a typed "max iterations" outcome. Compaction making real headroom is necessary for termination but not sufficient — the iteration cap is the backstop.
|
|
45
|
+
|
|
46
|
+
Return the last assistant message to the supervisor for uniform finalization.
|
|
47
|
+
|
|
48
|
+
### Sampling-request retry sub-loop
|
|
49
|
+
|
|
50
|
+
The inner request is itself a bounded retry loop (a runtime-configured max-retry count by error class — providers rarely advertise a runtime retry budget) over the streaming response. Classify stream errors before retrying — a cancelled token is a clean abort, a stream that closed before the terminal "completed" event is retryable, a malformed request is terminal. A turn-scoped client session may be reused across retries to preserve sticky routing / connection state, but **rebuild it when the failure implicates the pinned route or connection itself** (connection-level errors, a route returning persistent failures) — otherwise retries re-pin a poisoned upstream and burn the whole retry budget on a dead path. See `llm-client-gateway.md` for the full failure-classification taxonomy this defers to.
|
|
51
|
+
|
|
52
|
+
**Do not commit or act on partial stream output before retrying.** If a stream closed mid-response after already emitting assistant text or — worse — tool calls, a naive retry re-runs the request and can double-dispatch those tool calls (duplicated side effects) or duplicate assistant text in history. Buffer streamed deltas as *uncommitted* until the terminal "completed" event; only then record assistant items and dispatch tool calls. On a retryable mid-stream failure, discard the uncommitted partial before retrying. If your runtime must dispatch tools from a stream that can still fail, dedupe dispatch by provider response id + tool-call id so a retry cannot execute the same call twice. This is the loop-level counterpart to the idempotency rule for external effects below.
|
|
53
|
+
|
|
54
|
+
## Error taxonomy at the loop level
|
|
55
|
+
|
|
56
|
+
Distinct outcomes need distinct handling — do not collapse them into one error path:
|
|
57
|
+
|
|
58
|
+
- **Aborted** → break silently; the abort is reported via its own lifecycle event, not as an error.
|
|
59
|
+
- **Poisoned input** (e.g. an invalid/oversized inline image the provider rejects) → sanitize the offending content *in history* (replace with a placeholder) and retry under a small bounded attempt count; only surface a user-facing error if sanitization can't recover. This prevents one bad attachment from wedging the turn permanently, while the bound prevents a sanitize-retry loop if the provider keeps rejecting.
|
|
60
|
+
- **Usage/quota limit** → emit the error lifecycle event *and* propagate a typed "limit reached" signal to any higher-level goal/budget runtime; do not silently swallow.
|
|
61
|
+
- **Other errors** → emit the error event but leave the session resumable (let the user continue the conversation); do not tear down the session.
|
|
62
|
+
|
|
63
|
+
## Interruption & ordered finalization
|
|
64
|
+
|
|
65
|
+
There are two distinct terminations; keep their paths separate. **Cancellation** (user interrupt, or a "replaced" preemption) tears the task down mid-flight and is reported via the abort lifecycle event. **Terminal error** (the error arms above) ends the turn but leaves the session resumable and is reported via the error event. They share the flush-before-finality rule but differ in cleanup: only cancellation hard-aborts an in-flight task and records an interrupted-turn marker; a terminal error lets the current step finish unwinding normally. Do not route an error through the cancellation hard-abort path (you'd kill a task that was already returning) or a cancellation through the error path (you'd emit an error event for a deliberate stop).
|
|
66
|
+
|
|
67
|
+
The cancellation path:
|
|
68
|
+
|
|
69
|
+
1. Cancel the task's token.
|
|
70
|
+
2. Wait a **bounded grace window** for the task to observe cancellation and unwind cleanly, *then* hard-abort the handle if it hasn't. The grace step exists so an in-flight approval/tool wait doesn't surface as a spurious model-visible rejection after the abort. Tune the window per runtime (a small fixed value on the order of ~100 ms is typical); it is a tradeoff between interrupt latency and clean unwind, not a universal constant.
|
|
71
|
+
3. **Make the hard-abort safe against mid-write data loss.** A hard abort can land while a tool is mid-write to durable history or mid external side-effect, leaving a half-written record with no terminator. Do not rely on the timer alone: make history writes transactional/atomic (a torn write is discarded or completes, never persists partial), or only hard-abort at cancel-safe checkpoints, and snapshot/flush any pending in-flight history items the aborted task owned *before* the hard-abort — otherwise step 4's marker flushes after those items were already lost.
|
|
72
|
+
4. Run the task's `abort` cleanup. Keep this cleanup **bounded and cancellation-independent**: it must not await a lock the hard-aborted task may still hold, nor block on the killed task/tool, or the supervisor deadlocks and the turn never finalizes. Give cleanup its own timeout. Note that grace-then-hard-abort only protects *durable history* writes; non-transactional **external** side-effects (a network POST, a remote mutation) can be left half-applied or double-applied across abort/retry. Make such effects carry idempotency keys (at-least-once safe) so an interrupt or retry can't double-apply them — the turn loop cannot make a non-idempotent external call safe.
|
|
73
|
+
5. For a cancellation, record an **interrupted-turn marker** in history that carries the *reason* (user-interrupt vs replaced/superseded) *before* emitting the abort event, so the next turn's model — and a forked/branched continuation — sees that the previous turn was cut off and why.
|
|
74
|
+
6. **Commit all terminal-relevant records in one ordered durable flush before emitting finality.** This includes the final assistant message, the interrupted/error marker, and any last history items — not just the completion event. Some clients re-read the log synchronously on receipt of the finality event; if the event is flushed but a preceding record was only buffered, the client sees a terminal event with a truncated transcript. Order: durably commit the record set, confirm the flush landed, *then* emit finality. Non-negotiable for any terminator (complete, abort, error).
|
|
75
|
+
|
|
76
|
+
### Finalization on success
|
|
77
|
+
|
|
78
|
+
Finalize uniformly from the supervisor (the spawn site), not scattered across task implementations, so every task kind shares one lifecycle. **Process queued input and decide the idle-wake under the same session lock that clears the active-turn slot** — do not clear the slot, release the lock, then process leftover input. If you release between, an idle-wake (or a concurrent submission) can claim the freed slot and start a new turn *before* the prior turn's queued input is appended, so that input lands in the wrong turn or races history. Steps, all within the finalization critical section:
|
|
79
|
+
|
|
80
|
+
- Detach/clear the active-turn slot via compare-and-set on the stored turn-state identity (a stale finalizer must not clear a newer turn).
|
|
81
|
+
- Process any input that queued during the turn (re-run input hooks on it) before the slot becomes claimable.
|
|
82
|
+
- Compute **per-turn token usage** as `total_at_end − total_at_start`, clamped non-negative per field for display; emit it to telemetry and analytics keyed by turn id + thread id. Per-turn deltas (not cumulative totals) are what make cost/latency dashboards answer "which turn was expensive". The single-active-turn invariant is what makes the snapshot-difference correct; if your runtime ever allows concurrent turns or background usage on the same counter, accumulate usage from each response stream instead, since a shared global would mis-attribute concurrent consumption. Clamp only the displayed value — when `end < start` (a provider counter reset or lost accounting), also emit an anomaly signal rather than silently recording zero, or accounting bugs hide forever.
|
|
83
|
+
- Emit the terminal completion event with completed-at, duration, and time-to-first-token.
|
|
84
|
+
- If the session is now idle, run idle-lifecycle hooks and consider waking for queued background work.
|
|
85
|
+
|
|
86
|
+
### Idle wake / mailbox
|
|
87
|
+
|
|
88
|
+
Let an idle session start a fresh turn when background work is queued (a mailbox item marked "trigger a turn"). Gate it: only when the session is genuinely idle and only when a trigger item exists — but make the idle-check-and-claim a single atomic compare-and-set on the active-turn slot, not two separate reads. A check-then-act here double-starts turns when several mailbox writes or wake signals arrive concurrently. **Atomically claim (lease or state-transition) the trigger item before starting the turn**, so the item is consumed exactly once; otherwise a started turn that produces no new trigger leaves the same item visible and the session wakes itself again in a tight empty-turn loop. This is how an agent resumes autonomously after an interrupt drained pending work, without spinning when there's nothing to do.
|
|
89
|
+
|
|
90
|
+
## Anti-patterns
|
|
91
|
+
|
|
92
|
+
- Driving multiple concurrent turns over one shared history (interleaved writes, nondeterministic order). Enforce the single-active-turn invariant.
|
|
93
|
+
- Emitting the completion/abort event before the durable flush completes (lost-transcript race).
|
|
94
|
+
- Treating "model returned no tool calls" as turn-complete while pending user input exists (drops mid-turn steering).
|
|
95
|
+
- One catch-all error arm (abort, quota-limit, poisoned-input, and transient stream errors need different handling and different user-visible outcomes).
|
|
96
|
+
- A mid-turn compaction floor that doesn't actually reduce tokens (hot loop). Termination depends on compaction making real headroom.
|
|
97
|
+
- Per-task-implementation finalization (drift in what events/metrics each task emits). Finalize once at the supervisor.
|
|
98
|
+
- Non-cancellation-aware awaits inside the loop (interrupts that only take effect at the next natural boundary, so "stop" feels broken).
|
|
99
|
+
- Recording the interrupted-turn marker *after* the abort event, or not at all (next turn's model has no signal the prior turn was cut off).
|
|
100
|
+
- Hard-aborting on the grace timer while a history/side-effect write is in flight, with non-atomic writes (half-written record, no terminator). Make writes transactional or abort only at safe checkpoints.
|
|
101
|
+
- No outer-loop iteration cap (an always-blocking stop hook or a steering source that re-queues forever spins indefinitely).
|
|
102
|
+
- Check-then-act on the active-turn slot for either preemption or idle-wake (double-started turns under concurrent submitters). Use compare-and-set.
|
|
103
|
+
- Committing/dispatching partial stream output before the terminal completion event (a retry double-executes tool calls or duplicates history). Buffer uncommitted, or dedupe dispatch by response + tool-call id.
|
|
104
|
+
- Compacting across an open tool-call transaction (drops/reorders tool-call↔result pairing). Compact only at protocol-safe boundaries.
|
|
105
|
+
- Clearing the active-turn slot, releasing the lock, then processing queued input or deciding idle-wake (a new turn claims the slot first). Keep finalization's slot-clear + queued-input + idle-wake in one critical section.
|
|
106
|
+
- `abort` cleanup that awaits a lock the hard-aborted task holds, or blocks on the killed task (supervisor deadlock). Keep cleanup bounded and cancellation-independent.
|
|
107
|
+
- Routing a terminal error through the cancellation hard-abort path or vice versa (wrong event, killed-but-returning task).
|
|
108
|
+
|
|
109
|
+
## Per-turn model loop finality
|
|
110
|
+
|
|
111
|
+
Treat each per-turn model loop as a finality state machine, not just a streaming request. Bind every loop attempt to session and incarnation, message snapshot, prompt/context generation, model and tool-schema generation, permission mode and policy generation, queued input or notification snapshot, abort signal and reason, transcript-write generation, usage ledger, and max-turn or task-budget state. Persist accepted user inputs, queued inputs, streamed assistant blocks, tool results, progress/control attachments, and compact or recovery boundaries before depending on resume or terminal-result delivery; flush buffered writes before reporting a final result when the host may terminate immediately. Streaming fallback, retry, or abort must preserve assistant-message and tool-use identity: discard or tombstone orphan partial messages, pair every emitted tool call with exactly one terminal tool result, synthesize bounded error results for queued or in-progress tools when needed, and never leave an orphan tool call in transcript, cache, or protocol output. Queue and notification drains need a stable snapshot, target agent/session scope, consume-once removal only after attachment or delivery, stale/undeliverable recovery as visible uncertainty, and no cross-agent prompt leakage. If tool execution returns updated context, permissions, or tool lists, refresh schemas and re-authorize before the next model call. Distinguish user interrupt, submit interrupt, streaming abort, tool abort, hook-stopped continuation, max-turn, budget exhaustion, malformed runtime state, and model/provider error as different terminal or continuation reasons; record usage and terminal events exactly once. Diagnostics may expose bounded reason categories and counts, but not raw prompts, queued bodies, tool inputs or outputs, local paths, credentials, provider payloads, or free-form runtime errors.
|
|
112
|
+
|
|
113
|
+
## Foreground-query to background-session handoff
|
|
114
|
+
|
|
115
|
+
Treat foreground-query to background-session handoff as a session and query incarnation transition, not a UI minimize. The foreground loop must stop with a typed handoff reason that is distinct from user cancel or failure, then transfer ownership to the background owner with a fresh or explicitly rebound session and query generation, message snapshot, prompt/context generation, permission/policy/tool-schema generation, source-scope labels, and privacy state. Drain queued task notifications under a stable snapshot before any queue processor can start a replacement foreground turn; forward them to the background owner exactly once, and consume/remove them only after attachment, delivery, or typed degraded evidence. Deduplicate notifications already emitted into the foreground transcript or message list without using raw prompt bodies as durable identifiers unless the body is locally scoped, redacted, and never logged. Failed attachment construction, stale queue snapshots, session drift, or owner drift must leave visible uncertainty rather than silently dropping or replaying notifications. Foregrounding or resuming a backgrounded query must revalidate session incarnation, root or workspace, privacy and data-residency, policy, permission mode, tool/capability generation, prompt generation, transcript high-watermark, and queued-notification generation before accepting input or inheriting approvals; stale queued input, task notifications, live credentials, or cached permissions must be rejected or rebound. A transferred notification must not by itself start a new foreground query, complete a task, or authorize a tool. Diagnostics may expose bounded handoff reasons, transfer counts, duplicate/drop categories, and stale-reject categories, but not notification bodies, prompt text, queued content, local paths, task or session identifiers, credentials, or free-form runtime errors.
|
|
116
|
+
|
|
117
|
+
## Progress attachments
|
|
118
|
+
|
|
119
|
+
Treat progress attachments as model-visible status records with explicit retention semantics. High-frequency ephemeral status ticks for the same parent operation may replace the prior tick instead of appending when they carry only the current display state; stateful progress trails that carry nested messages, hook/task history, tool history, or user-relevant milestones must append or merge by a typed state machine so UI recovery, transcript replay, and resume do not lose evidence. Classify each progress type before persistence, bind it to parent operation id, session/incarnation, tool or task identity, transcript-write generation, terminal-result generation, abort signal, and privacy/source-scope labels, and reject late progress after completion, cancellation, compaction boundary, session/workspace/root/privacy/policy drift, or parent-operation mismatch. Replacement must be stable across replay and must not create false absence, completion, or work-history claims; preserved stateful progress must stay bounded by retention, compaction, and redaction policy rather than unbounded logs. Terminal progress indicators and OS-level notifications are display side channels owned by `terminal-cli-dev`; they must not be treated as transcript evidence or completion proof. Diagnostics may expose bounded progress categories, counts, replacement decisions, and stale-drop reasons, but not raw progress bodies, nested message content, tool arguments, local paths, credentials, terminal control bytes, or free-form errors.
|
|
120
|
+
|
|
121
|
+
## Plan-to-execute handoff
|
|
122
|
+
|
|
123
|
+
Treat plan-to-execute handoff as a runtime state transition, not a UI preference. Entering planning must suppress write-capable execution except explicitly scoped plan artifacts; exiting needs explicit approval provenance, the approved plan content, contained plan artifact identity, immutable plan digest or snapshot generation, approval request id, approver/source identity, principal/account/tenant/organization scope with explicit none markers when a scope is absent, session/incarnation, workspace, privacy/data-residency/source-scope authorization, permission mode and effective permission state before planning, target execution mode, policy/capability generation, and whether the user edited the plan. If the approved plan is edited in an external surface, write it back, re-snapshot it, and make implementation use the edited version only. Execution must start from the approved digest or immutable snapshot; any post-approval content, path, containment, privacy/source scope, or snapshot drift requires re-approval. Restore the prior effective permission state only when the current gate still permits every mode and rule, and when each cached or persisted approval still matches its original principal/account/tenant/organization/workspace/session/incarnation, working directory, tool source/schema/capability generation, resource scope, revocation floor, and policy generation tuple; deny/revocation precedence remains authoritative, and stale, broad, stripped, or dangerous approvals stay stripped unless both the target mode and current policy explicitly permit them. If an automatic or bypass-like mode is disabled, circuit-broken, revoked, or policy-drifted during planning, fall back to the safe default and notify the user. Planning state and plan artifacts must survive compaction, resume, and remote/local handoff only with authorized privacy/data-residency/source-scope labels, minimized or redacted diagnostics, and size bounds; forked sessions must copy or re-snapshot the plan under the child session identity and require re-approval unless the approval explicitly covers the child tuple. Stale, missing, or unauthorized plan artifacts require recovery from an authorized snapshot or message history, not silent implementation. Teammate or remote approval channels must accept replies only from the owning approver or control channel, bind them to the pending approval request, full principal/account/tenant/organization scope with explicit none markers, plan artifact identity, plan digest or snapshot generation, session/incarnation, workspace, target mode, policy/capability generation, and privacy/source-scope authorization, reject forged/late/duplicate/stale replies after any tuple drift, preserve rejection counts, and treat ambiguous event-stream finality as pending or failed closed rather than approved.
|
|
124
|
+
|
|
125
|
+
## Progress reporting as a runtime contract
|
|
126
|
+
|
|
127
|
+
Progress reporting is part of the runtime contract. Track counters and recent activity from durable events, not from UI text; classify progress snapshots as summary-only unless they are backed by durable output. Detect stalled interactive/background work separately from slow work, surface a sanitized actionable prompt, and avoid treating a progress ping as completion. Task summaries, usage, todo state, task-list storage metadata, and worktree/session metadata must be sanitized, bounded, and correlated to the exact task; do not leak raw prompts, credentials, sensitive paths, task ids beyond the intended audience, teammate or team names, task-list storage keys or paths, assignment payloads, prompt-submit bodies, raw diagnostic strings, raw tool arguments/results/bodies, or unrelated worker output. Sidecar inference that generates task progress labels or completed-tool labels is a read-only auxiliary path: deny tools and side effects even when preserving cache-key compatibility, partition or validate cache use by sidecar purpose, effective tool policy, source snapshot digest or generation, and prompt/model version, skip or isolate transcript writes, read a fresh bounded task/tool snapshot for each run, filter incomplete action/result pairs, treat the snapshot as untrusted prompt data with injection-resistant wrapping, escaping, and size caps, bind the label to the source snapshot and task or tool-batch identity, discard late labels unless the current task/tool-batch identity and source snapshot digest or generation still match at apply time, prevent overlapping runs with timer or singleflight cleanup, abort on stop, treat API errors or empty labels as absent/degraded progress rather than completion, and never let the label replace durable task, tool, permission, or side-effect finality.
|
|
128
|
+
|
|
129
|
+
## User-away or idle-return recap generation
|
|
130
|
+
|
|
131
|
+
Treat user-away or idle-return recap generation as a model-visible meta-message boundary, not a cosmetic notification or conversation compaction. Trigger only from explicit local activity, focus, or idle signals under current product, setting, policy, privacy, and supported-mode eligibility; a timer that fires during an active turn may record pending intent only until a safe turn boundary, and focus return, new user input, session/workspace/root switch, settings or policy drift, privacy change, or abort must cancel or discard stale work. The recap source set must be bounded to recent session context and authorized session memory, labeled as a recap of prior context rather than fresh evidence, wrapped as lower-precedence untrusted data, and generated through a read-only path with no tools, agents, connectors, external side effects, task completion, permission approval, or authority changes. Bind every generation to session id and incarnation, principal/workspace/privacy tuple, last real user-turn identity, message snapshot or digest, input generation, loading or turn-status generation, focus or idle-state generation, feature/settings/policy generation, memory snapshot generation, model/prompt generation, and abort signal; insert the typed meta/session message only if the same tuple still matches at apply time. Suppress duplicates at least until a new real user turn occurs, reject empty or failed generations without transcript pollution, and make any visible degraded state unobtrusive and non-authoritative. Diagnostics may expose bounded categories, counts, and degraded-state reason codes, but not raw prompts, transcript snippets, recap bodies, memory bodies, local paths, session identifiers, exact private-activity timings, credentials, or free-form model/runtime errors. Route visible return-context acceptance to `product-rd-workflow`, dialog or visual hierarchy to `product-ui-ux-design`, terminal focus mechanics to `terminal-cli-dev`, default or remote-eligibility rollout to `platform-release-engineering`, signal schema to `platform-observability`, stale/noisy/misleading recap incidents to `defect-diagnosis`, and security/adversarial review when recap input can carry private transcript, memory, or workspace context.
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
# Inference Capacity Operations
|
|
2
|
+
|
|
3
|
+
## Source-Backed Runtime Lessons
|
|
4
|
+
|
|
5
|
+
- Register or advertise an inference service only after readiness is proven. Readiness should cover the actual serving mode, such as model/application loaded, health endpoint passing, or deployment status running; a process start alone is not enough.
|
|
6
|
+
- For Ray Serve or similar deployment-graph hosts, treat `serve build`/`serve run` or a bound ingress object as deployment wiring evidence, not readiness evidence by itself. Confirm the deployment status, route availability, and loaded handler/model before registration or traffic shift.
|
|
7
|
+
- Service metadata should expose routing and diagnosis fields such as environment, partition, active version, startup timestamp, and build/commit version where available. Keep concrete registry/provider names out of generic code.
|
|
8
|
+
- Heartbeat, unregister, shutdown, and polling loops must be bounded or intentionally supervised. Avoid unbounded waits, hidden infinite loops, and process-kill behavior unless the platform contract explicitly owns it.
|
|
9
|
+
- Version and traffic routing should be explicit. If a caller requests a version, route deterministically; otherwise route by configured ratios that validate to `1.0` or fail without silently changing traffic.
|
|
10
|
+
- Inference handlers should preserve request/log ids through HTTP, gRPC, internal calls, and tests so failed predictions and support reports can be traced.
|
|
11
|
+
- Keep service host, SDK/registration, experiment/benchmark, generated artifact, model package, and runtime output folders separate. Do not convert model weights, generated configs, debug output, or benchmark scripts into default product-service architecture rules.
|
|
12
|
+
- Streaming responses need terminal state handling: timeout, context cancellation, provider error, EOF/success, partial-content persistence, reader close, and usage/cost availability after the stream is consumed.
|
|
13
|
+
|
|
14
|
+
## Capacity Controls
|
|
15
|
+
|
|
16
|
+
- Bound concurrent inference per process, route, model, provider, and tenant or user scope when needed.
|
|
17
|
+
- For batch inference, define max batch size, batch wait timeout, max ongoing requests, and backpressure behavior.
|
|
18
|
+
- For async jobs, persist task state and include lease, retry count, timeout threshold, terminal failure, and repair path.
|
|
19
|
+
- For hosted models, define warmup, health checks, model-load failure behavior, GPU/CPU resource requests, autoscaling target, and max replicas.
|
|
20
|
+
|
|
21
|
+
## Batch Serving
|
|
22
|
+
|
|
23
|
+
Batch serving should make latency/throughput tradeoffs explicit:
|
|
24
|
+
|
|
25
|
+
- `max_batch_size` caps memory and tail latency;
|
|
26
|
+
- `batch_wait_timeout` controls how long requests wait for aggregation;
|
|
27
|
+
- `max_ongoing_requests` protects the replica;
|
|
28
|
+
- per-item output ordering and error mapping must be deterministic.
|
|
29
|
+
|
|
30
|
+
Keep preprocessing and postprocessing deterministic and cheap. Expensive transformations should be measured separately from model inference.
|
|
31
|
+
|
|
32
|
+
## Load And Regression Checks
|
|
33
|
+
|
|
34
|
+
Before rollout, run a bounded capacity check for:
|
|
35
|
+
|
|
36
|
+
- QPS/concurrency saturation;
|
|
37
|
+
- p50/p95/p99 latency;
|
|
38
|
+
- timeout and retry rate;
|
|
39
|
+
- queue depth or pending job age;
|
|
40
|
+
- memory/GPU pressure;
|
|
41
|
+
- token cost per successful output;
|
|
42
|
+
- streaming first-token latency and final-token latency when applicable.
|
|
43
|
+
|
|
44
|
+
Use dry-run or report-only modes for migration/backfill/batch jobs whenever possible.
|
|
45
|
+
|
|
46
|
+
## Fine-Tuning And Local Models
|
|
47
|
+
|
|
48
|
+
- Treat fine-tuned models as registry versions with parent base model, training data lineage, training job id, parameter recipe, eval gate, and rollback path.
|
|
49
|
+
- Do not promote a fine-tuned model without comparing it against the base model and current production model on the same eval set.
|
|
50
|
+
- For local/self-hosted inference, document runtime engine, quantization, context length, batching policy, KV-cache behavior, GPU/CPU memory budget, and warmup path.
|
|
51
|
+
- Keep training data governance separate from prompt logs: consent, retention, redaction, deduplication, and evaluation leakage must be explicit.
|
|
52
|
+
- **Compressing agent/conversation trajectories into a training or eval dataset must preserve signal and avoid label/leakage bugs — not just fit a token budget.** When a recorded trajectory exceeds target length, do NOT head/tail-truncate or uniformly shrink. Protect the **head** (system prompt + first user turn + first action/tool turn = task setup) and **tail** (final actions + conclusion = outcome). The **middle is the learnable signal** for tool-use imitation / process-supervision / trajectory evals (tool choices, observation handling, correction loops, failed attempts) — compress it only after preserving the salient action/observation spans; never collapse the taught/evaluated behavior into one summary (same destruction as blind truncation). Three invariants make this safe:
|
|
53
|
+
- **Group / split / dedup / leakage-screen by original-trajectory lineage BEFORE compressing, and re-run overlap checks after.** Compress-then-split leaks: a compressed example and its near-raw sibling can straddle the train/eval boundary, or an LLM-written summary can absorb eval content (contamination — see the anti-contamination rule in `model-prompt-evaluation.md`).
|
|
54
|
+
- **An LLM-generated compression/summary note is NOT an original turn.** Keep it in metadata outside the supervised message/action stream, or explicitly mask it from loss / eval scoring; never serialize it as an assistant/tool turn — otherwise you teach the model to emit curator summaries and turn observations into synthetic labels.
|
|
55
|
+
- **Govern "salient" before selecting.** Define salience criteria up front, preserve required negative/failed/correction spans, record selector + version + reason, and audit the slice distribution before/after — post-hoc "keep the interesting parts" cherry-picks a non-representative set (worse for evals, where it inflates or deflates scores).
|
|
56
|
+
This is the offline-dataset counterpart to — and distinct from — runtime context compaction (loop mechanics in `agent-turn-lifecycle.md`, fidelity floors in `agent-session-persistence.md` / `llm-client-gateway.md`), which preserves *current intent / pending tasks* for continuation; here the goal is *training/eval signal*.
|
|
57
|
+
- Capacity tests should include cold start, model load failure, concurrent requests, batch saturation, and memory pressure.
|
|
58
|
+
|
|
59
|
+
## Operational Failure Modes
|
|
60
|
+
|
|
61
|
+
Plan for:
|
|
62
|
+
|
|
63
|
+
- provider timeout or rate limit;
|
|
64
|
+
- fallback provider incompatibility;
|
|
65
|
+
- malformed JSON/tool calls;
|
|
66
|
+
- streaming disconnect;
|
|
67
|
+
- partial batch failure;
|
|
68
|
+
- model load failure;
|
|
69
|
+
- prompt activation regression;
|
|
70
|
+
- runaway token cost;
|
|
71
|
+
- stuck async task.
|
|
72
|
+
|
|
73
|
+
Each failure mode should map to an owner-visible metric, log, trace, or report entry.
|
|
74
|
+
|
|
75
|
+
## Post-Launch Drift Monitoring And Ramp Gating
|
|
76
|
+
|
|
77
|
+
A launch eval proves a version at one point in time; inference quality drifts as input distribution, retrieved context, model behavior, or dependencies change. Define drift monitoring before ramp and wire it to ramp control, not only to a passive alert.
|
|
78
|
+
|
|
79
|
+
Distinguish two drift types because they are observable at different times. **Data drift** is a shift in the *inputs* (query mix, length, language, retrieved-context distribution, embedding distribution) — it is a *leading* signal, measurable immediately without ground truth, and is an early warning that quality may degrade. **Concept drift** is a change in the right *answer* for the same input (the input→correct-output relationship moved) — it is a *lagging* signal that usually cannot be measured directly in an LLM system with no immediate ground truth; detect it through labels, human-review/feedback, or proxy quality metrics, which arrive later. Do not treat a clean data-drift dashboard as proof quality is fine; concept/quality drift can be real while inputs look stable, and only the lagging signals will show it. "Lagging" is not "wait passively for labels": where the answer depends on fast-changing facts (prices, policy, legal status, knowledge-base content), add leading proxies for concept drift — content/source freshness and version-skew monitors, synthetic recency probes, change hooks on the upstream source of truth, and a high-risk human-review queue — so stale-wrong answers surface before the label/feedback signal arrives.
|
|
80
|
+
|
|
81
|
+
- **Drift signal taxonomy — monitor every applicable signal; mark the inapplicable ones N/A with a reason.** Input-distribution shift (new query types, length, language mix); empty/no-answer/refusal-rate shift; output-error-type distribution shift (new failure classes, not just total error count); latency (first-token and full-response); token/compute cost per successful output; timeout and failure rate; third-party/provider dependency anomaly (provider latency, error, rate-limit, or a silent default-behavior change). Not every signal applies to every product (an embedding or classifier service has no refusal rate; a batch job has no first-token latency) — record N/A explicitly so a real blind spot is distinguishable from an inapplicable signal.
|
|
82
|
+
- **Each signal has a pre-defined threshold, response time-limit, and owner before ramp.** "We monitor X" without a trigger value and an owner is not monitoring. Define the absolute or relative trigger, how long the condition must persist to fire, who is paged, and the target time to a disposition decision.
|
|
83
|
+
- **Attribute before you gate — a raw metric shift is not model drift.** A no-answer/refusal/latency/cost shift can come from seasonality, a product campaign, an abuse spike, or an upstream traffic-mix change, not the new version. Use slice-aware baselines, compare against a control / current-prod cohort, require a minimum sample/traffic volume, and exclude known seasonality and classified abuse/incident traffic before attributing a shift to the candidate.
|
|
84
|
+
- **Tag each threshold as warning or gate; only a gate-severity crossing pauses the ramp.** Warning thresholds page the owner and inform the disposition decision; gate thresholds halt expansion and, per the rollback contract, can downgrade or roll back. An alert that does not gate ramp lets a regression ride to full traffic — but auto-pausing on every warning causes flapping and, across many co-gated services, a synchronized failover storm. Before any automated pause/rollback require hysteresis, a debounce window, the minimum sample size above, and a per-route blast-radius limit; large or cross-service pauses escalate to a human rather than firing globally at once.
|
|
85
|
+
- **Define behavior under stale or missing telemetry separately.** When the gating signals are delayed or absent, do not treat "no breach observed" as healthy: high-risk routes fail closed (hold or roll back); other routes hold the current ramp step and page the owner rather than continuing to expand blind.
|
|
86
|
+
- **Keep a live-traffic observation set distinct from the frozen benchmark and the regression bad-case set.** Sample real post-launch traffic to test whether the new version fits the *current* input distribution; the frozen benchmark answers "as good as before" but cannot detect distribution drift. This sampling is subject to the same privacy discipline as eval/replay records (see `retrieval-agent-safety.md` Safety And Security): privacy-approved sampling, redaction/minimization, a retention TTL, and access control; use synthetic or aggregated substitutes where raw traffic capture is prohibited. Confirmed bad cases from this set feed the regression bad-case set (see `model-prompt-evaluation.md`).
|
|
87
|
+
- Metric pipeline, alert routing, and dashboard mechanics are owned by `platform-observability`; ramp/pause/rollback authority and time-limits are owned by `platform-release-engineering` and the launch gate in `product-rd-workflow`. This section owns which inference drift signals to watch and the ramp-gating contract.
|
|
88
|
+
|
|
89
|
+
## Triton-Class Multi-Model Serving
|
|
90
|
+
|
|
91
|
+
When inference is served through Triton (NVIDIA), Triton-compatible engines (PaddleX HPS, KServe), or any multi-model server with config-driven model loading:
|
|
92
|
+
|
|
93
|
+
- One server instance can host many models (commonly 10-30+). Each model is a directory under the model repository with a `config.pbtxt` (or equivalent) plus versioned weights. The server is the runtime; the configs are the contract.
|
|
94
|
+
- Use **ensemble scheduling** when an inference call needs `preprocess → model infer → postprocess` chained: declare the ensemble as a synthetic model so the client sees one logical call. Intermediate tensor names are part of the contract and must not collide.
|
|
95
|
+
- **Backend choice per model**: Python backend for pre/post and custom logic, TensorRT/CUDA for GPU model inference, OpenVINO IR for CPU-optimized inference, framework-specific engines (PaddleX HPS) for vendor stacks. Mixing backends in one server is normal; the server config is where the choice is recorded.
|
|
96
|
+
- **Heterogeneous instance scaling**: CPU pre/post stages declare `kind: KIND_CPU` with higher instance count (6-8) because they parallelize cheaply; GPU model stages declare `kind: KIND_GPU` with low instance count (1-2) bound by GPU memory **AND** GPU compute share. The two scale independently.
|
|
97
|
+
- **GPU compute oversubscription is a separate failure mode from GPU memory**: when multiple `instance_group { kind: KIND_GPU }` entries (across all models loaded on the same physical device) share one GPU, raising `instance_group.count` does NOT linearly increase throughput — past the point where concurrent kernels fully utilize the GPU's SM / memory-bandwidth capacity, additional instances queue inside the CUDA driver / GPU scheduler. Triton's per-model **queue duration** (`nv_inference_queue_duration_us`) measures only the time spent in Triton's scheduling queue BEFORE dispatch and does NOT include GPU-scheduler contention after dispatch; that contention shows up in **compute duration** (`nv_inference_compute_infer_duration_us`) and in end-to-end latency. Real failure mode: memory fits comfortably, the Triton config validates, the model loads, queue duration looks healthy, end-to-end P50 doubles under load.
|
|
98
|
+
- **Measure don't model**: there is no clean formula to predict the cliff. `instance_group.count` is NOT a static compute-share allocator; MPS limits are process/context-level not per-model-instance; vendor partitioning (MIG) is only meaningful when explicitly configured. Treat capacity as an **empirical profiling** problem: use Triton Model Analyzer or measure end-to-end P50 / P95 / throughput at target concurrency BEFORE and AFTER any count change. Never raise the count on intuition that "more instances = more throughput".
|
|
99
|
+
- **Choose by SLA thresholds at equal offered load, not per-metric dominance**: a higher count can win on P50-of-completed-requests while losing on P95 (long tail under contention), losing on throughput (more queueing externally), or losing on throughput stability (autoscaler thrash). The decision criterion is whether the configuration meets the **declared per-metric SLA thresholds** (P50/P95/P99 budgets, sustained throughput target, tail-variance ceiling) at target concurrency. A config that violates any SLA threshold is unsafe and rolled back, regardless of which alternative is "better on P50"; a config that meets all thresholds wins, regardless of whether the alternative beats it on a non-SLA metric. Document the chosen SLA thresholds next to the config so the next engineer's "which is better" question has a single answer.
|
|
100
|
+
- **When explicit partitioning IS available, sum one partition level at a time, not across nested levels**: each partition mechanism scopes "100%" to its own parent. MIG slices each get a fraction of a physical GPU; their allocations sum to ≤ 100% of that physical GPU and that constraint is enforced by MIG itself. MPS clients each get `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` of their MPS server's view; in the nested case where MPS runs INSIDE a MIG slice (MPS-on-MIG), the MPS server's "100%" IS the MIG slice — not the physical GPU. The sum-check is layered: at each level, the children sum to ≤ that level's budget. Don't flatten the levels — summing every MPS-client percentage across every MIG slice on a physical GPU against a single 100% double-counts and rejects valid nested configurations. Concretely: physical GPU vs MIG slices: MIG-allocator handles the sum. Multiple MPS clients on one MPS server (whether that server is on a bare GPU or inside one MIG slice): sum the clients against the MPS server's parent budget (the bare GPU's 100%, or the MIG slice's 100%). Multiple MPS servers sharing the SAME parent (multiple MPS servers on one bare GPU, or multiple MPS servers inside one MIG slice — uncommon but possible): aggregate ALL clients across ALL sibling MPS servers against the shared parent's budget — two MPS servers each running at "100%" of the same bare GPU is 200% of one GPU and oversubscribes. **MPS limits do NOT scope per-`instance_group` inside one Triton process**: `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` is per MPS client (the whole Triton process), not per model-instance — claiming you can sum it across N `instance_group` entries inside one Triton server is wrong and either rejects valid configs or grants false isolation confidence. Without partitioning at the right boundary, the rule is empirical profiling above.
|
|
101
|
+
- **Dual config files per model** (`config_min.pbtxt`, `config_max.pbtxt`) plus an env-variable selector at startup is the common pattern for switching between resource profiles (lab vs prod, dev vs canary, CPU-only vs GPU). Document which profile each environment uses.
|
|
102
|
+
- **Dynamic batching**: in Triton two settings work together. `max_batch_size > 1` at the model level **permits** batch shapes (the model can receive a tensor whose first dimension is the batch). The `dynamic_batching { ... }` scheduler block selects the dynamic batcher — present and empty means "enable with default knobs". When the block is absent, the effective behavior depends on the backend: some backends (TensorRT, ONNX Runtime in certain configs) and `auto-generated` model configs may enable dynamic batching with defaults when `max_batch_size > 1`; other backends (Python backend, custom backends) do not. **Audit by the effective generated config and observed batching metrics, not by source-`pbtxt` grep alone**: load the model and inspect Triton's `/v2/models/<name>/config` endpoint or the model status logs, then verify with batching metrics under load. `preferred_batch_size` and `max_queue_delay_microseconds` inside the block are tuning knobs that shape latency-throughput trade-off; set them explicitly when latency is sensitive (typical: `preferred_batch_size` close to expected concurrency, `max_queue_delay_microseconds` matching the SLA budget). For variable-length token inputs, add `allow_ragged_batch: true`. A config with `max_batch_size = 1` cannot batch regardless of the block.
|
|
103
|
+
- **Model warmup** belongs at startup, not at first-request latency. A warmup sample per model amortizes JIT, kernel selection, and KV-cache allocation cost so the first production request is not a cold call.
|
|
104
|
+
- **Version policy**: `model_version: -1` (latest/all) is convenient but loses reproducibility. For services with version-sensitive behavior, pin the version explicitly and rotate through a versioned route or canary policy.
|
|
105
|
+
- **Per-model GPU affinity** is necessary when GPU memory or compute is tight. Sharing one GPU device id across all models works only while combined memory and concurrent kernels fit; under contention, declare `instance_group { ... gpus: [n] }` per model.
|
|
106
|
+
- Standard ports (HTTP 8000, gRPC 8001, metrics 8002) are conventions; expose them through the platform's discovery layer rather than hard-coding client URLs.
|
|
107
|
+
|
|
108
|
+
## ML Pipeline Visualization Service
|
|
109
|
+
|
|
110
|
+
When a multi-stage inference pipeline (detection → classification → matching → structured output) is in production, debugging output drift in any single stage requires a side-by-side comparison of the stage's input, intermediate tensors, and final output. A separate visualization service is the durable answer; ad-hoc Jupyter notebooks per investigation rot fast and silently drift from production code:
|
|
111
|
+
|
|
112
|
+
- **Per-stage handler classes share the production input contract**: the visualization service declares one handler per pipeline stage (`<stage>_visualize.py`) that consumes the same input shape the production service consumes for that stage. When the production stage changes its input contract, the visualization handler is updated in the same PR — if it is not, the visualization renders stale and silently misleads debug sessions.
|
|
113
|
+
- **Common base handler for cross-stage concerns**: a `BaseHandler` owns request id, log context, image / tensor download from object storage, structured logging, and the notification webhook for shareable links. Stage handlers inherit and override only the stage-specific render logic. Without the base, every stage handler re-implements the same boilerplate and drifts.
|
|
114
|
+
- **Gradio (or equivalent) is the UI seam, not the inference seam**: the visualization service exposes Gradio components (image viewer, text panel, JSON tree) wired to the stage handlers; inference itself reuses the production stage implementation — NOT a parallel reimplementation. Concretely: (a) when the production caller SDK exposes the per-stage outputs the viz needs (often just the final stage output), the viz tool calls that SDK; (b) when the viz needs raw intermediate tensors that the public SDK does not expose, the viz tool reuses the production *stage modules* directly (the same Python classes / Go packages production runs) rather than re-implementing the stage. Both paths satisfy the rule — what is forbidden is a second copy of the stage logic that drifts from production. If neither path works (production-only Triton ensemble, no Python-level stage entry), add a sanctioned trace/debug interface on the production service (gated by internal auth) and have the viz tool consume that — never resort to a parallel implementation.
|
|
115
|
+
- **Ray Serve / equivalent for multi-tenant viz**: when multiple engineers use the tool concurrently and the visualization itself is non-trivial (image processing, large response rendering), deploy the viz service via Ray Serve so concurrent users do not block each other on the Python GIL. For low-volume internal tools, a single-process Gradio is fine; document the scale assumption.
|
|
116
|
+
- **Output artifacts go to object storage with a shareable URL — and every access path along the chain has its own controls**: every visualization render writes the rendered image / annotated tensor to object storage with a request-id-prefixed key; the response includes the URL so the engineer can share the result. The artifact path is a separate access surface from the ingress: (a) bucket ACL restricts which identities can read the prefix (internal-employee identity provider, not "anyone with the URL"); (b) signed URLs carry a short TTL (minutes for routine debug, never weeks); (c) webhook payloads / chat previews that auto-expand the URL are themselves access paths — disable link-unfurl in the channel or post hashes-only when the artifact contains sensitive customer data; (d) audit log captures who fetched which artifact, for post-incident review. In-browser rendering only is fine for one-off looks; for any incident-class debug, the artifact persists for post-mortem reference.
|
|
117
|
+
- **Internal-only network exposure with auth — ingress is the first gate, not the only one**: the viz tool exposes intermediate model outputs that are normally hidden from end users (raw output strings, confidence scores, intermediate features). Bind the ingress to an internal-only network; gate by SSO / VPN. The ingress alone does NOT cover the artifact-storage path (separate controls above), the webhook / chat surface where URLs land, or developer machines that download and cache artifacts locally — each is its own access path with its own controls. The same redaction rules that apply to logs apply to viz output for any sensitive customer data.
|
|
118
|
+
- **Mode / env selector via a single env var**: a flag like `VIZ_MODE=full | extract-only | classifier-only` selects which subset of stage handlers is active for a given deployment. Avoid one deployment trying to be all things; route a slim viz to one URL and the full pipeline viz to another when they have different access policies.
|
|
119
|
+
|
|
120
|
+
## Inference API Layer Above Triton
|
|
121
|
+
|
|
122
|
+
When a Python service wraps Triton (or directly hosts models) as the request-facing layer:
|
|
123
|
+
|
|
124
|
+
- **Three deployment shapes** are common: FastAPI + Ray Serve (multi-model orchestration with `@serve.deployment` and `autoscaling_config`), FastAPI + uvicorn lifespan (single-model CPU service with eager singleton load), and llama.cpp `llama-server` (GGUF self-contained binary for VL / GGUF-quantized LLM). Pick by model and traffic shape; document the choice per service.
|
|
125
|
+
- **Ray Serve specifics**: `autoscaling_config(min_replicas, max_replicas)` controls scale. **Backpressure has two distinct knobs**: `max_ongoing_requests` is the per-replica in-flight cap that drives autoscaling and queue-vs-route decisions (it is **not** a rejection threshold); `max_queued_requests` (at the deployment / HTTP proxy layer) is the cap on requests waiting in the router queue and **is** where rejection happens (excess returns 503 / back-pressure). Pick both intentionally — `max_ongoing_requests` follows model-replica capacity (commonly single-digit to low tens for GPU models, higher for I/O-bound work), `max_queued_requests` follows the SLA budget on queue wait. `max_batch_size` lives on the actor / handler, not at Ray level. Build the deployment via `serve build` / `serve run server_config.yaml`.
|
|
126
|
+
- **Model load lifecycle**: handlers load models on `initialize()` (Ray Serve) or lifespan startup (FastAPI) as singletons per actor or per worker. Lazy first-request load is acceptable only when readiness reflects the load state.
|
|
127
|
+
- **Per-handler timeouts** match each model's measured SLA and queue budget. One global timeout cannot cover models that span two orders of magnitude in latency; record example values only as service-local tuning, not as a generic skill default.
|
|
128
|
+
- **Per-model batching**: declare `max_batch_size` per model based on the model's actual memory profile (text embedding 24, orientation 8, layout detection 4 are typical). The orchestrator (Ray Serve actor pool) handles scheduling.
|
|
129
|
+
- **Image / large-payload handling**: accept the image as a protobuf message field (e.g. `ImageDetectReq` deserialized from JSON body) rather than a separate multipart upload; for very large payloads, accept a pre-signed object-storage URL and let the server fetch. When the inference output is itself a binary (corrected image, mask, rendered overlay), prefer returning a signed object-storage URL over inline base64 so callers do not pay the wire-cost on every response. If the API contract requires inline bytes, do not strip them from the response (that breaks callers); apply size limits at the logger / persistence layer so log bloat is bounded without changing the response contract.
|
|
130
|
+
- **GGUF + llama.cpp serving**: use `-ngl <n>` to push all layers to GPU (`-ngl 99` for full offload), `-b` / `-ub` for batch / micro-batch sizes, `-np` for parallel contexts. For OCR-style use, set `--temp 0` (greedy decode) so output is deterministic. Expose the server's OpenAI-compatible HTTP API; do not invent a new wire format on top.
|
|
131
|
+
- **multi-version handler coexistence** (v1 and v4 in the same service) is the pragmatic pattern when a model rolls out incrementally. Route by request field (`version`) or by `handler_flow_ratio` for canary traffic; do not silently swap.
|
|
132
|
+
|
|
133
|
+
## Model Artifact Integrity
|
|
134
|
+
|
|
135
|
+
Self-hosted inference depends on weights, configs, and tokenizers loaded into a process with GPU and network access. Treat artifact loading as a supply-chain step, not a file copy.
|
|
136
|
+
|
|
137
|
+
- **Immutable digests**: every model version is addressed by content digest (sha256 / model registry hash), not by mutable path or "latest" symlink. The runtime resolves digest → object-storage URI at load time and refuses to load if the resolved bytes do not match the expected digest.
|
|
138
|
+
- **Signed manifests**: where the registry supports it, sign the model manifest (weights + config + tokenizer + metadata bundle) and verify the signature before load. A model not signed by an approved key path fails closed; do not load with a warning.
|
|
139
|
+
- **Config + weight pairing**: refuse to load when `config.json` / `tokenizer.json` / `model.safetensors` come from different versions — pin all artifacts of one model to one digest.
|
|
140
|
+
- **Safe deserialization**: prefer `safetensors` (no executable code) over Python `pickle` / `torch.load(..., weights_only=False)` for untrusted-origin models. If a `pickle`-based format is unavoidable, only load from an approved internal registry with manifest signature verification; never load from a user-supplied URL or attachment.
|
|
141
|
+
- **Build-time vs runtime**: bake artifact integrity checks into the model-load path so they run in production, not only at build time. A build-time-only check is bypassable by runtime symlink / mount substitution.
|
|
142
|
+
|
|
143
|
+
## Failure Domain Isolation Across Co-Hosted Models
|
|
144
|
+
|
|
145
|
+
One Triton or Ray Serve instance commonly hosts 10-30+ models. Without isolation, one bad model (OOM, hang, infinite loop, malformed weights) takes the rest down with it.
|
|
146
|
+
|
|
147
|
+
- **Separate processes by trust and criticality**: high-trust / high-criticality models (auth-gating classifiers, billing-affecting evaluators) run in their own server process, not co-hosted with experimental / large / unstable models. Don't host an experimental VL model in the same Triton process as the production OCR critical path.
|
|
148
|
+
- **GPU isolation**: where the GPU supports it, partition with NVIDIA MIG (Multi-Instance GPU) so one model cannot monopolize compute or memory of another. Where MIG is unavailable, use `CUDA_VISIBLE_DEVICES` to pin per-process GPUs and accept the lower utilization in exchange for isolation.
|
|
149
|
+
- **CPU / memory cgroup limits**: containerize each server with explicit CPU and memory limits matched to the model's resident footprint plus headroom; a Python-backend model that leaks memory should hit its container limit and OOM-kill itself, not exhaust the host.
|
|
150
|
+
- **Python backend isolation**: when Triton's Python backend hosts user-supplied or experimental logic, run it in a separate Triton instance from compiled backend models so a Python-side crash does not cascade.
|
|
151
|
+
- **Blast-radius review**: before adding a new model to an existing multi-model server, audit what else lives there and what fails if the new model OOMs or hangs. If the answer is "the critical path", co-hosting is the wrong decision.
|
|
152
|
+
|
|
153
|
+
## Inference Service Operational Hygiene
|
|
154
|
+
|
|
155
|
+
Recurring anti-patterns observed across production inference services:
|
|
156
|
+
|
|
157
|
+
- **`print` instead of structured logging**: every inference log line should carry `request_id`, `log_id`, model name, model version, and the standard stage timing fields. `print` calls drop on container restart and cannot be aggregated.
|
|
158
|
+
- **GPU OOM not caught**: `torch.cuda.OutOfMemoryError` and equivalents must be caught, surfaced as a typed error (not 500 Internal), and surface in metrics so capacity planning can react. Silent OOM crashes look like flaky network errors to the caller.
|
|
159
|
+
- **`uvicorn --workers 1`** without justification is a single-worker bottleneck even when the host has many CPUs. Pick the worker count consciously (often 1 when the model holds a single GPU, more when CPU-bound).
|
|
160
|
+
- **Disabled framework logging** (`llama-server --log-disable` or equivalent) makes triage impossible. Keep at least warn-level logging in production and redirect to a file or sink the platform aggregates.
|
|
161
|
+
- **No `/health` / `/ready` endpoint**: readiness must reflect model-loaded state, not process-running state. Without an explicit endpoint, orchestrators and discovery layers cannot distinguish "process up" from "model ready to serve".
|
|
162
|
+
- **Mismatched runtime declarations**: a service whose `config.properties` describes one runtime (e.g. TorchServe) but whose start script launches a different runtime (e.g. Ray Serve) is a maintenance trap. Keep one canonical declaration and delete or clearly mark legacy files.
|