@ccoalm/ccl-skills 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -0
- package/README.md +49 -0
- package/dist/assets/marketplace/.agents/plugins/marketplace.json +12 -0
- package/dist/assets/marketplace/.claude-plugin/marketplace.json +13 -0
- package/dist/assets/marketplace/marketplace-manifest.json +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/marketplace.json +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.claude-plugin/plugin.json +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.codex-plugin/plugin.json +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/.worktree-only +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/session-start.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/agent-context/subagent-start.md +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/AGENTS.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-delegation-owner.sh +125 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-edit-isolation.sh +102 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/guard-merge-authorization.sh +1156 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/hooks.json +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/merge-authorization-prompt.sh +142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-guard.sh +12 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/owner-dispatch-stop.sh +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/remind-post-merge-cleanup.sh +144 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-context.sh +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/session-start.sh +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/skill-extraction-gate-stop.sh +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/subagent-start.sh +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_delegation_owner.sh +329 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_edit_isolation.sh +322 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_guard_merge_authorization.sh +902 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_merge_authorization_prompt.sh +178 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_remind_post_merge_cleanup.sh +121 -0
- package/dist/assets/marketplace/plugins/ccl-skills/hooks/test_session_start.sh +170 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/AGENTS.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/ccl-skills.ts +564 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-install-skills.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-update-skills.md +44 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-verify-skills.md +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/packages/opencode-plugin/commands/ccl-worktree-check.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/AGENTS.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/README.md +276 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.example.json +10 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/owner-dispatch.sh +1307 -0
- package/dist/assets/marketplace/plugins/ccl-skills/scripts/owner-dispatch/test.sh +941 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/SKILL.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/agents-file-coverage-gate/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/SKILL.md +188 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/android-dev.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/flutter-dev.md +80 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/ios-dev.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/kotlin-multiplatform.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-platform-boundaries.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/mobile-quality-release.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/app-cross-platform-dev/references/source-evidence-map.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/SKILL.md +353 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/client-routing.md +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/manual-invocation-and-prompts.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/staged-review-contract.md +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/references/timeout-auth-and-capabilities.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/AGENTS.md +98 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_envelope.py +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/classify_timeout_exit.sh +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/claude_review.sh +1438 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/codex_review.sh +324 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/concern_excerpt.py +295 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/egress_schema.py +214 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/init_policy_matrix.py +642 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_packet_mcp.py +181 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/kimi_review.sh +1165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/opencode_review.sh +1190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_cli_review.py +946 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_opencode_review.py +474 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_probe_result.py +1899 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/parse_review_json.py +200 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.py +2845 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/review_gate.sh +6 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/run_claude_capture.py +71 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/runtime-surface-verification-design.md +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_classify_envelope.sh +68 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_claude_review_probe.sh +2311 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_cli_review_wrappers.sh +1832 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_code_review_identity.sh +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_concern_excerpt.sh +245 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_egress_schema.sh +177 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_init_policy_matrix.sh +272 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_kimi_packet_mcp.py +195 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_concurrency.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_opencode_review_retry.sh +1005 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_opencode_review.sh +258 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_probe_result.sh +574 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_parse_review_json.sh +349 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_compat.py +434 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_client_order.sh +264 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/test_review_gate.sh +2412 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/code-review/scripts/verify_native_skill_binding.py +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/SKILL.md +153 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/diagnosis-playbook.md +54 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/defect-diagnosis/references/prevention-routing.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/SKILL.md +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/feature-risk-router/references/security-review-gate.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/SKILL.md +165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/api-security-boundaries.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/architecture-playbook.md +160 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/artifact-generation-architecture.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/audit-history-architecture.md +29 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/bulk-workflow-architecture.md +33 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/config-rule-routing-architecture.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/cross-cutting-concerns.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-modeling-and-migrations.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/data-platform-architecture.md +210 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/dependency-platform.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/developer-tooling-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/error-contract-architecture.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/event-driven-architecture.md +260 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/http-gateway-architecture.md +74 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/mq-consumer-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/multi-tenant-isolation.md +275 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/notification-architecture.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/ops-checklist.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/performance-capacity-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/protobuf-contract-architecture.md +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/redis-cache-coordination.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/release-runtime-readiness.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/replay-comparison-architecture.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/runtime-observability.md +94 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/service-scaffold.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/source-evidence-map.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-architecture/references/workflow-state-architecture.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/SKILL.md +159 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/artifact-generation-patterns.md +37 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/audit-history-patterns.md +28 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/bulk-import-export-patterns.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/config-rule-routing-patterns.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/data-access-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/db-schema-and-dal-patterns.md +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/dependency-client-patterns.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/developer-tooling-patterns.md +70 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/domain-feature-patterns.md +78 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/engineering-patterns.md +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/error-contract-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/feature-playbook.md +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/http-gateway-client-patterns.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/mq-consumer-patterns.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/notification-patterns.md +42 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/observability-implementation-patterns.md +101 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/performance-capacity-patterns.md +44 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/protobuf-contract-patterns.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/public-api-integration-patterns.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/quality-and-testing-patterns.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/redis-cache-lock-patterns.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/release-ops-patterns.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/reliability-patterns.md +83 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/replay-comparison-patterns.md +32 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/scaffold-and-codegen.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/source-evidence-map.md +54 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/go-microservice-dev/references/state-machine-task-patterns.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/SKILL.md +80 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/grill-me/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/SKILL.md +117 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-approval-auto-reviewer.md +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-command-sandbox.md +441 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-context-freshness.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-credentials-auth.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-extensions-skills.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-file-edit-protocol.md +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-ide-integration.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-input-ingestion.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-instruction-composition.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-lifecycle-hooks.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-messaging.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-runtime-bootstrap.md +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-session-persistence.md +448 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-task-orchestration.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-tool-dispatch.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/agent-turn-lifecycle.md +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/inference-capacity-operations.md +162 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/llm-client-gateway.md +156 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/model-prompt-evaluation.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/llm-inference-integration/references/retrieval-agent-safety.md +273 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/SKILL.md +202 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/contracts-and-state.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/cross-stack-alignment.md +94 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/framework-choice.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/online-practice-uptake.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/platform-capabilities.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/product-page-checklist.md +40 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/qa-release.md +72 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/miniapp-product-dev/references/source-evidence-map.md +82 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/SKILL.md +103 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/agents/openai.yaml +5 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-agent-delegation/references/multi-agent-delegation-playbook.md +100 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/SKILL.md +70 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-data-acquisition.md +549 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/public-disclosure-channels.md +97 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/references/research-prompts.md +66 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/AGENTS.md +32 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/multi-perspective-research/scripts/test-public-data-acquisition-recipes.sh +379 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/SKILL.md +244 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/alerting-and-on-call.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/framework-middleware-checklist.md +142 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/infra-component-deployment.md +268 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-correlation-recipe.md +124 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/log-schema-canonical.md +208 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/metrics-conventions.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/obs-stack-architecture.md +107 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/sli-slo-design.md +95 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-observability/references/source-register.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/SKILL.md +303 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/canary-and-rollout-strategy.md +163 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/config-center-via-etcd.md +245 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/custom-control-plane-boundary.md +298 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-cli-concrete-recipe.md +312 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/deploy-pipeline.md +165 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/env-and-lane-matrix.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/lane-orchestration-control-plane.md +383 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/multi-region-and-cluster.md +135 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/promotion-gate-and-review.md +149 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/python-package-registry-release.md +462 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/rollback-playbook.md +123 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/secret-and-config-management.md +231 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-release-engineering/references/version-authority-and-deprecation.md +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/SKILL.md +276 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/dual-sidecar-and-traffic-config-center.md +127 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/framework-middleware.md +143 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/grpc-authority-workaround.md +90 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/http-response-envelope-contract.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/mesh-architecture.md +127 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/multi-env-routing.md +192 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/protobuf-http-contract-signals.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/retry-timeout-circuit-breaker.md +124 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/rpc-framework-recipe.md +494 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-choice.md +113 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-migration-playbook.md +231 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/platform-service-connectivity/references/service-discovery-recipe.md +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/SKILL.md +235 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/adr-convention.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-checklist.md +30 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-evaluation-report-template.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-execution-spec.md +108 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-sop.md +457 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/algorithm-launch-templates.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/artifact-egress-confidentiality.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/code-review-checklist.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/cross-repo-coordination.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/delivery-lifecycle.md +192 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-review-gate-mechanics.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/design-routing-and-readiness.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/diagnostic-spec-match-gate.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dispatch-owner-skills.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/dormant-code-activation.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/existing-project-assessment-report.md +223 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/external-skill-augmentation.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/feature-deprecation-cascade.md +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/high-risk-resilience-gates.md +73 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-completeness-and-minimality.md +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/implementation-entry-reentry-gate.md +122 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/modular-monolith-heuristic.md +105 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/pre-final-continuation-gate.md +115 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/problem-resolution-and-learning.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-attributes.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/quality-remediation-program.md +88 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/rd-standards-doc-family-checklist.md +27 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/refactoring-discipline.md +52 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/review-reception.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/shared-gate-artifact-classification.md +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/source-evidence-map.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/status-tracker-sync.md +77 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/sync-spec-repo-contract.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/verify-developer-experience.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/references/worktree-mechanics.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-rd-workflow/scripts/check-agent-contract-coverage.sh +213 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/SKILL.md +136 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/agents/openai.yaml +9 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/analytics-visualization-interactions.md +206 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/behavioral-aesthetic-logic.md +108 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/complex-creation-interactions.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-execution-checklist.md +214 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-impl-naming-and-versioning.md +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-intake-and-acceptance.md +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/design-system-source-of-truth.md +97 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/external-ui-ux-quality-benchmarks.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/frontend-code-evidence-map.md +63 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/interaction-design-patterns.md +146 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/layout-recipes-and-screenshot-acceptance.md +250 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-project-token-consistency.md +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/multi-stack-strategy.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/operational-processing-workflows.md +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-mobile-patterns.md +324 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/platform-web-desktop-patterns.md +456 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-lifecycle-acceptance-and-iteration.md +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/product-surface-patterns.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/resource-management-interactions.md +113 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/scenario-community-patterns.md +133 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/source-map.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/tokens-and-components.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/trust-sensitive-ai-and-data-patterns.md +96 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-audit.md +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/ui-ux-design-development.md +176 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/product-ui-ux-design/references/visual-craft.md +111 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/SKILL.md +157 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/ai-service-integration-boundaries.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-contract-and-schema.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/api-security-boundaries.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/architecture-playbook.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/async-execution-model.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/background-jobs-and-scheduling.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/batch-and-pipeline-architecture.md +11 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/config-secrets-runtime.md +22 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-modeling-and-migrations.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/data-platform-architecture.md +211 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/event-driven-architecture.md +263 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/multi-tenant-isolation.md +281 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/observability-and-ops.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/packaging-runtime-readiness.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/redis-cache-coordination.md +41 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/reliability-and-error-contract.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/source-evidence-map.md +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-architecture/references/web-framework-boundaries.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/SKILL.md +143 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/ai-service-wiring-patterns.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/async-and-worker-patterns.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/background-job-patterns.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/batch-and-artifact-patterns.md +13 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/dependency-client-patterns.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/error-handling-patterns.md +26 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/feature-playbook.md +43 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/observability-implementation-patterns.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/project-structure-and-tooling.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/public-api-security-patterns.md +52 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/redis-cache-lock-patterns.md +78 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/schema-and-validation-patterns.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/source-evidence-map.md +56 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/sqlalchemy-and-migrations-patterns.md +99 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/testing-and-quality-patterns.md +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/python-service-dev/references/web-framework-patterns.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/SKILL.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/config-runtime-readback.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/mr-merge-authorization.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/post-release-env-reset.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-closeout-evidence.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/release-scope-confirmation.md +21 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/tag-and-prod-pipeline-gate.md +20 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/test-scope-prompt.md +24 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-coordination/references/watcher-discipline.md +14 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/SKILL.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/comment-safe-release-doc.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-evidence-workflow.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/release-doc-writer/references/release-testing-scope-section.md +15 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/SKILL.md +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-baseline/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/SKILL.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/prd-composition-contract.md +35 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/requirement-closure-contract.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-doc-writer/references/security-four-questions.md +38 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/SKILL.md +91 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-intent/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/SKILL.md +88 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/requirement-scope/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/SKILL.md +337 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/analysis-parse-fix-test-challenge-replay.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/attribution-verification.md +69 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/bootstrap-slim-c3-obligation-table.md +112 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/coverage-exhaustion-traps.md +45 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/description-authoring.md +162 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/dual-track-review-gate.md +507 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/eval-routing.md +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/evidence-card-template.md +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/example-domain-preselect.md +79 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/external-practice-controls.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-lifecycle-handoff.md +65 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/extraction-quickstart.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/firing-point-placement.md +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/harness-patterns-and-eval.md +286 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/incident-postmortem-extraction.md +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/l0-l1-l2-routing.md +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/online-skill-review.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/parallel-stack-references-pattern.md +164 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/r0-leakage-audit.md +90 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/recurring-anti-patterns-checklist.md +320 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/resume-paused-delivery.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-feedback-mining.md +33 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-finding-standards.md +57 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/review-rubric.md +40 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/rule-consolidation.md +118 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/skill-listing-budget.md +19 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-register.md +254 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/source-to-skill-extraction.md +658 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/two-source-extraction-pattern.md +167 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-judgment-extraction.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/uiux-routing-map.md +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/references/validation-and-landing.md +180 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-ccl-skills.sh +1452 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-evidence-card-leak.sh +491 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-mr-target-freshness.sh +173 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-size-budget.sh +488 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/check-sync-pointers.sh +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-golden-trace.rb +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-health.rb +327 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing-bank.rb +401 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/eval-routing.rb +248 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/generic-r0-leak-scan.sh +282 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/governing-chain-diff.py +321 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/impact-chain-gate.rb +964 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/register-firing-path-resolution.rb +708 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/skill-behavior-eval.py +540 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-lifecycle.rb +51 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/source-register-pending-status.rb +55 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_ai_coding_implementation_gates.sh +829 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_impact_chain_refscripts.sh +1203 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_r0_status.sh +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_register_pending_exclusion.sh +137 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_regressions.sh +173 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_route_drift.sh +377 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_size_budget.sh +833 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_skill_catalog.sh +491 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_ccl_source_register_lifecycle.sh +114 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_mr_target_freshness.sh +261 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_check_sync_pointers.sh +538 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_controlled_escalation_pins.sh +154 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_grader_diagnostics.sh +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_bank_surface_binding.sh +178 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_eval_routing_prose_target.sh +86 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_generic_r0_leak_scan.sh +131 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_git_identity_predicate_gate.sh +243 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_governing_chain_diff.sh +419 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_impact_chain_gate_dateless_host.sh +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_resolution.sh +724 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_register_firing_path_wiring.sh +414 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_regression_runner_registration.sh +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_bank_integrity.sh +205 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_routing_pointer_integrity.sh +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_credential_cwd.sh +61 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_cross_refs.sh +111 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/test_validate_skill_root_depth.sh +53 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/skill-extraction-workflow/scripts/validate-skill.sh +257 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/SKILL.md +98 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/input-state-machines.md +36 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/streaming-rich-output.md +130 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/terminal-cli-dev/references/terminal-side-channels.md +96 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/SKILL.md +408 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/AGENTS.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/bitable-setup.md +573 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/README.md +120 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/github-actions.yml +119 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/gitlab-ci.yml +76 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/ci_templates/jenkins.Jenkinsfile +106 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/classical-test-design-techniques.md +279 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/gen_report.py +2807 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/makefile-template.md +200 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/report-config-schema.md +272 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/run_pytestless.py +475 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/source-to-case-workflows.md +258 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-marker-conventions.md +316 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc-review-and-prioritization.md +145 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/AGENTS.md +16 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.dart +129 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.go +197 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.py +135 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/tc_helpers/tc.ts +285 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/test_gen_report.py +2144 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/test-artifact-management/references/update-lifecycle.md +62 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/SKILL.md +212 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/ci-fixtures-and-flake-control.md +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/client-runtime-test-matrices.md +50 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/data-and-workflow-testing.md +34 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/design-closed-contract-oracles.md +31 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/e2e-real-flow-testing.md +71 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/fitness-functions.md +240 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/integration-contract-testing.md +235 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/non-functional-specialized-scenarios.md +296 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/rd-testing-standard-template.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/run-killing-mutation-walk.md +43 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/scenario-testing.md +136 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/source-evidence-map.md +59 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/structured-tc-input-translation.md +67 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-code-authoring-patterns.md +392 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-data-and-determinism.md +39 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/test-topology-and-commands.md +92 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/unit-testing.md +46 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/vendored-contract-drift-checklist.md +64 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/references/verify-enforcement-mechanisms.md +18 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/AGENTS.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.py +140 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/client-terminal-ansi-check.test.sh +75 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.py +170 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-ast-check.test.sh +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.go +198 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/lang-basics-go-check.test.sh +109 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/testing-strategy/scripts/test_mutation_backup_recipe.sh +237 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/SKILL.md +184 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/comment-safe-feishu.md +93 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/cross-model-co-review.md +3 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/delivery-face-closeout.md +60 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/doc-charter-first.md +17 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/tighten-doc/references/session-vantage-leakage.md +58 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/SKILL.md +126 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/complex-workspace-patterns.md +47 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/embedded-h5-in-host.md +87 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/react-architecture.md +194 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/source-evidence-map.md +60 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-quality-release.md +190 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/web-react-dev/references/web-ui-quality.md +83 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/SKILL.md +179 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/agents/openai.yaml +4 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/references/shared-branch-rebase.md +25 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/AGENTS.md +23 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_status.sh +207 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/test_worktree_sweep.sh +481 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-status.sh +325 -0
- package/dist/assets/marketplace/plugins/ccl-skills/skills/worktree-isolation/scripts/worktree-sweep.sh +245 -0
- package/dist/assets/release.json +2797 -0
- package/dist/claude-adapter.d.ts +9 -0
- package/dist/claude-adapter.js +240 -0
- package/dist/cli-worker.d.ts +1 -0
- package/dist/cli-worker.js +32 -0
- package/dist/cli.d.ts +22 -0
- package/dist/cli.js +214 -0
- package/dist/codex-host.d.ts +30 -0
- package/dist/codex-host.js +162 -0
- package/dist/fs-safe.d.ts +21 -0
- package/dist/fs-safe.js +241 -0
- package/dist/index.d.ts +2 -0
- package/dist/index.js +1 -0
- package/dist/manifest.d.ts +8 -0
- package/dist/manifest.js +135 -0
- package/dist/opencode-adapter.d.ts +10 -0
- package/dist/opencode-adapter.js +416 -0
- package/dist/operations.d.ts +3 -0
- package/dist/operations.js +956 -0
- package/dist/paths.d.ts +20 -0
- package/dist/paths.js +4 -0
- package/dist/types.d.ts +58 -0
- package/dist/types.js +1 -0
- package/dist/unified.d.ts +4 -0
- package/dist/unified.js +64 -0
- package/dist/version.d.ts +2 -0
- package/dist/version.js +5 -0
- package/package.json +35 -0
|
@@ -0,0 +1,507 @@
|
|
|
1
|
+
# Dual-Track Review Gate
|
|
2
|
+
|
|
3
|
+
Independent-review gate for shared-skill changes, deep extraction, and any skill change that ships operational, architectural, or security-sensitive rules. `Shared-skill change` means any change under a shared skill package, including `SKILL.md`, references, scripts, validators, templates, generated outputs, metadata, and examples — **plus the repo's own plugin-shipped command / behavior surfaces outside `skills/`**: any executable, command, server, monitor, hook, settings, or behavior-activation surface the plugin ships and a host runs — today root `hooks/*` (SessionStart / PreToolUse shell on every session and tool call), `scripts/install.sh` (teammates run it at install), plugin `bin/` (added to the Bash `PATH` while enabled); the same applies to `.mcp.json`, `.lsp.json`, `monitors/**`, `settings.json`, and any manifest path that declares or redirects those. Those are the install-time and runtime code-execution supply-chain surface every install runs, so a change to them is a shared-skill change — independent review always, adversarial challenge for any non-wording change — and is never `not-applicable: docs-only`. Every shared-skill change requires an independent review before commit. Non-wording shared-skill changes also require an adversarial challenge before commit.
|
|
4
|
+
|
|
5
|
+
This gate owns *when* two passes run, *what each catches*, and what a finding state does to a landing (block / accept / defer / rerun). How findings are severity-calibrated (P0/P1/P2), what makes a finding actionable vs noise, and what makes a `findings` / `no-findings` / `inconclusive` result *valid* are owned by `review-finding-standards.md` — apply that standard to every finding either pass produces, so convergence is measured against a shared bar rather than each reviewer's private sense of "P1" or a rubber-stamped empty result.
|
|
6
|
+
|
|
7
|
+
## Why two passes
|
|
8
|
+
|
|
9
|
+
| Pass | What it catches | What it misses |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| **Fact/consistency review** (`codex review` or equivalent) | Technical inaccuracies (semantic claims that are wrong), contradictions across references, sanitization gaps (residual ccl-specific names/IPs/hostnames), over-prescription (saying MUST when industry has multiple acceptable patterns), API/path inconsistencies | Production-safety chaos modes: race conditions, data-loss paths, security holes, algorithm flaws, operational footguns |
|
|
12
|
+
| **Adversarial challenge** (`codex exec` with adversarial prompt) | Production-down scenarios under chaos, attacker-style abuse paths, edge cases that break the documented happy path, sequencing/timing issues, broken algorithms at small/large scale | The flat-text consistency layer (handled by the first pass) |
|
|
13
|
+
|
|
14
|
+
### Closeout reads the LANE NAMES, not the round count
|
|
15
|
+
|
|
16
|
+
A slice can record many rounds and still have run only one lane, because the only
|
|
17
|
+
record of which lanes ran is the landing artifact's own author-written rounds
|
|
18
|
+
table. At closeout:
|
|
19
|
+
|
|
20
|
+
- **Count lane names, never rounds**: enumerate the lanes the slice's own gate section requires, and confirm each one has a recorded outcome.
|
|
21
|
+
- **A rounds table naming one lane while the gate requires two is `interim`, never landed** — no round count and no green deterministic gate substitutes for a lane that never ran.
|
|
22
|
+
- **Where lanes leave evidence in a local store, a chain must exist for THIS slice** — check what the stored packets actually covered rather than assuming one is there.
|
|
23
|
+
|
|
24
|
+
Observed: a mandatory, fail-closed repository gate landed with nine review-mode
|
|
25
|
+
rounds recorded, no challenge lane run, and a landing state still reading
|
|
26
|
+
`plan drafted`. Run afterwards against the landed diff, the challenge found a
|
|
27
|
+
one-character bypass of that gate in a single round — appending one bracket
|
|
28
|
+
anywhere in a citation skipped resolution entirely, so the gate reported success
|
|
29
|
+
on exactly the input it exists to reject. The defect lived in "how would this
|
|
30
|
+
break", which is the question only the missing lane asks.
|
|
31
|
+
|
|
32
|
+
Skipping the challenge is how P0/P1 issues survive into shared skills. Real example from one multi-batch real-project extraction:
|
|
33
|
+
- Review pass found 8 issues (P2/P3): all factual/consistency.
|
|
34
|
+
- Challenge pass found 19 issues (3 P0 + 14 P1 + 2 P2): production-down + security + algorithm + footgun.
|
|
35
|
+
|
|
36
|
+
The 11 challenge-only findings included: `DELETE /lane` cascade without safety gates, mesh reconcile without serialization, snapshot DR corrupting live quorum, log-error-diff comparing counts not rates, `docker login -p` leaking password to process args, chat callback without signature/replay protection, FQDN fallback bypassing lane filtering. None of these were caught by review.
|
|
37
|
+
|
|
38
|
+
## Self-audit to convergence BEFORE the gate — the passes are an adversarial backstop, not your defect-finder
|
|
39
|
+
|
|
40
|
+
**The two passes are a final *independent adversarial backstop*, not your defect-discovery loop.** The recurring failure: an agent ships unconverged work into review/challenge, then treats the third-party model as the primary mechanism that finds what's wrong and uses its findings as the task list — review/challenge → fix → re-review/re-challenge → fix. It burns rounds, offloads the implementer's own diligence, and (because each fix can add a fresh P0/P1 — see *Iterating*) routinely fails to converge.
|
|
41
|
+
|
|
42
|
+
Before invoking either pass, self-audit the candidate to the point where *you* expect it to pass: for any non-wording semantic rule / gate / status / verdict / code-mechanism change, first **state its load-bearing invariants** — the properties that must hold on *every* path ("exactly one terminalizer fires", "no per-object state outlives its owner's close", "a deadline only shrinks across hops", "the required schema fields always survive truncation", "input is bounded before any hot-path step touches it") — make each a checkable assertion, and **record them under the implementer self-review row's edge/failure-paths element** (the row `product-rd`/`extraction-quickstart` already require — not a new field): each applicable invariant → its derived failure mode → the concrete check or a residual-risk disposition. Invariants absent, shallow, or vacuous make that element **inconclusive, not a checked box**; if the change genuinely has none, record `invariant-pass: not-applicable` with a reason held to the same gate-fireability standard (not self-waived — "it's just prose, not a mechanism" is the decorative-gate dodge). Many failure modes are just an unstated invariant broken, so derive them from the invariants too — an **additive lens**, not the universal root cause (some P0/P1s instead come from rollout ordering, authority, or a factual-contract miss that no single invariant captures; those are the axis list below). Doing the invariant pass is a large part of what separates a happy-path draft from senior-grade code. Then independently enumerate every path/branch/state that affects the outcome and pull each into the test matrix, run the complete checks/suite the owner/risk gate selects (risk-matched, not blanket), and re-check acceptance criteria, edge/failure paths, and the **recurring first-draft blind-spot axes** yourself (security/authority/data-loss, concurrency & lifecycle, resource bounds, rollout/migration ordering, over-broad absolutes, and enumeration-completeness — under-listing a set the rule itself defines — the set the challenge most reliably supplies, so pre-cover each applicable one; the full enumeration is the draft-time corollary in `SKILL.md`). Good shape: *"the only new runtime path is the no-jq fallback; I enumerated every merge-affecting branch into the matrix, asserted the jq and no-jq paths give identical exit codes on the same fixture, ran the full repo suite — now I hand it to the challenge as a last adversarial backstop."*
|
|
43
|
+
|
|
44
|
+
**Self-audit never narrows the gate.** Converging your own work first does NOT shorten, soften, rescope, or pre-bias the required challenge: a self-audited candidate still gets a fresh full *adversarial* pass under the convergence rules below (`Do not bias the re-challenge`), never a "confirm my work" pass. The self-audit conclusion is *your* controller-side evidence — never feed it to the challenger as framing that asks it to confirm your work; challenge prompts stay adversarial and diff-scoped. This is the same ordering `product-rd-workflow`'s technical-design gate and `extraction-quickstart.md` already mandate; a *valid* persisted self-review row (validity rules in `extraction-quickstart.md`) is **ordering evidence only** — proof the self-audit ran *first*, not proof its claims are *true*. Test/verification claims still need command/artifact evidence, and a backfilled or thin row is invalid. Don't restate the row's field list here.
|
|
45
|
+
|
|
46
|
+
"Self-review does not count as dual-track" (*What does NOT count*, below) means self-review is not sufficient *evidence* of independence — it does NOT make the self-audit-to-convergence preparation skippable. And if review/challenge is the first place a basic scope/contract/test/security issue surfaces, that is a process defect in the self-audit loop — but the finding is still a normal gate finding: resolve or disposition it under this gate (fix, or documented accept/defer/out-of-scope), never waive or downgrade it as "should have been caught earlier." The process-defect repair (close the self-audit gap, then rerun) is additional, not a substitute.
|
|
47
|
+
|
|
48
|
+
**A design pivot resets the self-audit obligation.** When the candidate is substantially redesigned mid-gate (a mechanism replaced, a capability torn down, a rewrite beyond the findings being fixed), the earlier self-audit covered the OLD candidate: redo the closure self-audit on the NEW candidate before re-entering the challenge. Sliding from a pivot straight back into challenge → fix → re-challenge is the exact loop this section forbids, and it recurs precisely at pivots because the prior audit feels "already done."
|
|
49
|
+
|
|
50
|
+
**Partition findings before fixing: mechanical fixes vs design decisions.** When a round returns findings, classify each before starting fix work: a *mechanical* finding (bug, missing check, wrong value) goes on the fix list; a *design-level* finding — one that questions a mechanism's cost, operability, trust-model fit, or existence (the remove-the-capability signal in `SKILL.md`), OR one whose remediation would expand scope into an explicitly-deferred concern (the scope-direction / controller-cut-scope signal in `SKILL.md`), recognized on its FIRST appearance (recurrence across rounds is only the reviewer-lane stop/reframe escalation, never the point at which you first classify) — is a **risk-owner decision item**: present keep / delete / narrow / replace — or, only after the current-phase-impact test and compound split (see `Findings, autonomous budget, and human authority`) leave a residual with no current-phase impact, cut-scope-to-phase-boundary — to the user/maintainer BEFORE investing hardening rounds in the questioned mechanism. Executing "fix the findings" by hardening a mechanism whose design finding was never decided pays the hardening cost twice — once to build, once to tear down.
|
|
51
|
+
|
|
52
|
+
**A finding-fix that widens or hardens a validator needs a false-positive sweep against an explicit accept-set oracle.** Before landing a fix that strengthens a check in response to a finding (advisory → blocking, a narrower accept-set, a new rejection class), set-diff the strengthened check against an authoritative universe of legitimate content — a finite schema, the surface's own inventory of shapes, or the existing corpus the check will scan; when no finite oracle exists, name the representative legitimate classes checked, the sampling boundary, and the uncovered residual as explicit risk. Listing the salient examples that came to mind is not a sweep — the fix must state its precision exposure, not only close the recall gap. Failure shape: a "block malformed ledger rows" fix that would have false-positived on the register's other legitimate table shapes, reverted one round later.
|
|
53
|
+
|
|
54
|
+
### The self-adversary enumeration — method detail (relocated from `SKILL.md`)
|
|
55
|
+
|
|
56
|
+
The recurring failure: the agent declares done/covered/converged, and the *user* has to push — "keep going", "did you verify", "that's not actually covered", "深度分析了么" — before the agent runs the loop that would have caught the gap. The `SKILL.md` rule carries the red lines (clean-fresh-result-or-`interim`, `unverified` labelling, risk-owner acceptance); this section carries the method detail.
|
|
57
|
+
|
|
58
|
+
- **Mutation enumeration.** List every property the candidate states — in the rule text, the commit message, a register row, a doc, or a test name — and for each one name the mutation that would make it RED. Where the property has an executable test, apply the mutation and watch it go RED; bound the blast radius — never disable an authorization, idempotency, or deletion guard and then exercise it against a shared or live dependency, where buying a RED can destroy real data or perform a real unauthorized action; mutate against isolated dependencies or at the lowest layer that avoids them, and where neither is possible record the property `unverified` — naming a plausible-sounding one and moving on is the same self-certification this rule exists to stop.
|
|
59
|
+
- **Independent oracle.** Where the property has no executable test — rule text, a register row, a doc — the enumeration is discharged only by an **independent oracle**: name the concrete observation that would contradict the property, say where that observation lives (the owner file and line, the primary source, the command whose output would differ), and go look. **Validate the oracle before trusting its verdict**: a check that returns "clean" because it looked in the wrong place, matched case-sensitively, used too narrow a pattern, or swallowed an error is indistinguishable from a passing property, and it fails in the dangerous direction. Before accepting a clean result you must PROVE THE CHECK CAN FAIL — point it at something you know is broken and watch it report that. Confirming it enumerated the inputs you meant is a necessary extra step, never a substitute: correct inputs say nothing about whether the predicate detects a mismatch or whether a non-zero exit was swallowed, so a check that can only ever say clean passes that weaker test. An unvalidated oracle is not weaker evidence than an imagined mutation; it is the same thing wearing a command prompt.
|
|
60
|
+
- **Dimension walk.** Adding cases inside an axis you already had buys nothing against one you did not: the enumeration walks dimensions (shape / provenance-and-trust / cardinality / semantics / ordering) before values — `testing-strategy` owns that list and the precision-row obligation that goes with it. A walk whose rows are all imagined mutations is exhortation wearing a checklist's clothes.
|
|
61
|
+
- **Re-owe after fixes.** Whatever you produced while fixing a previous round's findings is part of the current candidate and re-owes the whole enumeration — that newly-added mechanism is the most dangerous line in the diff, because it has no test yet and you wrote it with your attention on the defect it repairs. "The whole enumeration" includes the **pre-cover axes sweep** (concurrency & lifecycle above all): remediation text written mid-round re-owes the draft-time axes BEFORE the candidate goes back to the reviewer, because a fix written with attention on one defect systematically re-opens the same blind-spot axes the original draft missed. And the loop has an escalation point: when the same blind-spot axis or finding class supplies findings in a **third** round, stop the per-finding loop and run one full-matrix implementer self-enumeration (the artifact's own states × failure points × orderings × residues × cross-references) on the current candidate before any further external round — letting the reviewer surface one hole per round is the reviewer-as-defect-finder failure at its most expensive (observed shape: a multi-round program burned twenty-plus single-finding rounds on one axis family; the one full-lifecycle enumeration, run at the maintainer's correction, found the remaining holes in a single batch).
|
|
62
|
+
- **Graded verdict shape.** When the assessed reality is multi-dimensional or partial (capability, coverage, feasibility, quality, completion), collapsing it into one binary verdict — "done/not-done", "possible/impossible", "all correct/all wrong" — is the over-broad-absolute axis applied to your own claim layer: the swing to whichever pole feels safest to assert misrepresents a distribution, and the opposite-pole absolute ("structurally impossible", "nothing works") is the SAME defect as an unearned "done", not a humbler one. Report per-dimension status — what's strong, what's weak, what wasn't checked — with the confidence each part actually earned; and where a binary gate genuinely applies (a pass/fail check, a blocked/allowed decision), still give the clear top-line verdict after the per-dimension basis — calibration is not hedged mush. A user correcting your answers as too absolute ("每次都很绝对") is this defect's recurrence signal, same escalation as the `SKILL.md` rule states.
|
|
63
|
+
- **Honesty (descriptive, not permissive).** This is recognition-dependent salience, not a mechanical gate — an agent that doesn't notice it is done-claiming cannot self-fire it; the mechanical backstops remain the closeout `interim` gates + user-signal escalation. "I didn't notice I was claiming done" does NOT waive the rule — any non-trivial completion/coverage/convergence wording must carry clean-pass evidence or an explicit interim/downscope disposition *before* you emit it. The rule targets completion/coverage/convergence assertions on work whose failure a check could catch, and never narrows the mandatory dual-track challenge (it is the always-on generalization of *self-audit to convergence*, not a replacement for the gate).
|
|
64
|
+
|
|
65
|
+
### Pre-cover axis detail (relocated from `SKILL.md`)
|
|
66
|
+
|
|
67
|
+
Across a long operational-rule/code extraction series the challenge supplies *the same handful of axes* as the recurring P0/P1 — first drafts systematically nail the functional / cost / happy-path and omit a predictable set. The per-axis instance lists for the six first-draft blind-spot axes:
|
|
68
|
+
|
|
69
|
+
- **(1) security / privacy / authority / data-loss** — weakened safety/refusal/authorization, secret/PII into a durable artifact, non-restorable delete vs archive, lost/orphaned/duplicated work, untrusted input treated as authority, runs-against-prod/live-creds instead of a sandbox.
|
|
70
|
+
- **(2) concurrency & lifecycle** — races, deadlock (e.g. holding a lock through a drain/callback), use-after-close/free, resurrection after delete, double-free/double-close, cleanup ordering, at-most-once/fires-once.
|
|
71
|
+
- **(3) resource bounds** — an unbounded default/timeout/buffer/retry, a leaked registry/goroutine/task entry, a missing max backstop.
|
|
72
|
+
- **(4) rollout / migration ordering** — a step that breaks not-yet-upgraded consumers, or abandons a live bug to do the clean refactor first.
|
|
73
|
+
- **(5) over-broad absolute** — an "always/never" rule or impl that breaks a legitimate case and needs a scoped exception.
|
|
74
|
+
- **(6) enumeration-completeness (the mirror of (5) — under-listing, not over-listing)** — when the rule itself DEFINES or LISTS a set (allowed values/schemes, risk tags, change-shapes, state/error classes, reconciliation outcomes, OS/runtime/transport variants, the cases a guard must cover), the first draft lists the *salient* members and silently omits siblings; the check here is a **set-diff, not a negative case** — diff your enumerated set against its authoritative complete source (the same file's own table/enum, the primary-source spec, or an explicitly declared variant matrix in the changed artifact — not a taste call), not just the members that came to mind.
|
|
75
|
+
|
|
76
|
+
That the challenge reliably catches them does NOT make pre-covering redundant: the challenge is the safety net, not the first line, and a teammate whose challenge is weaker or (against the gate) skipped otherwise ships the blind draft.
|
|
77
|
+
|
|
78
|
+
Three per-edit instances of the axes above that recur because their canonical rules live AWAY from the editing surface (observed shape: three externally-caught defects in one batch, each covered by a rule the author had read but that had no firing point at the edit site): **(axis 1/6, new fillable fields)** every record field/slot/status value the diff INTRODUCES gets the record-field forgery question before handoff — what makes a false fill fail? — a field without a validation bar is the `SKILL.md` record-field corollary shipping a new forgery surface; **(axis 6, renames/recounts)** a change that renames, re-counts, or re-labels anything runs the family-wide stale-term scan to zero (tighten-doc closeout owns the rule; run it on skill text too — skill text IS a reader-facing doc) before review, not after the reviewer greps it for you; **(axis 5, gate-satisfying substitutions)** when a mechanical gate rejects your first choice, check the substitute against the gate's INTENT, not only its predicate — a predicate-passing substitute that misses the intent is self-goodharting under gate pressure (the circular-anchor shape: a firing-path anchor that points at prose ABOUT firing paths instead of the decision point that fires this change).
|
|
79
|
+
|
|
80
|
+
### Design-time operability check (relocated from `SKILL.md`)
|
|
81
|
+
|
|
82
|
+
For any new mechanical gate, validator, or evidence apparatus — **and for any change that makes an existing one's verdict stricter** — run the four legs at design time, not after challenge rounds force them. These four legs all fire at design time; the check has a second firing point they do not cover, because its actor is not the gate's author: when a landing **withdraws or downgrades the evidentiary claim an existing gate rests on**, that gate is re-based or retired in the same landing — see the claim-liveness rule in `product-rd-workflow/references/design-review-gate-mechanics.md`, which owns it, including the obligation walk a retirement owes.
|
|
83
|
+
|
|
84
|
+
- **(a) author dogfood, scaled to the gate's statefulness** — for a gate that is base-relative, stateful, or evidence-regenerating, the intended authoring workflow (multi-commit development, a rebase, one routine follow-up edit) must pass it end-to-end under the SAME base resolution CI uses, before the gate lands (a gate whose own author's branch fails it ships a broken contract); a trivial stateless check needs only a proportional smoke run (this leg stays risk-matched — it never demands synthetic multi-commit ceremony for a one-shot grep).
|
|
85
|
+
- **(b) marginal-cost statement** — record what the cheapest routine change costs under the gate (recompute/regenerate/rerun burden); a gate whose per-iteration cost defeats normal development gets lightened or redesigned at design time.
|
|
86
|
+
- **(c) trust-model fit** — name what the mechanism defends against under its DECLARED trust model; machinery that only defends against adversaries the trust model already excludes (e.g. content digests where the author can regenerate every hash) buys redundant detection at full complexity cost — prefer the lighter mechanism that keeps the enforceable core.
|
|
87
|
+
- **(e) loosening check — an exemption must name the class of change that stops owing evidence, and that class must be one with no behaviour to evidence.** Fires whenever a change makes a gate accept what it used to reject: a new exemption class, a waived requirement, a widened accept set. A loosening is easier to get wrong than a tightening and shows up later, because it produces no red for anyone to notice — the gate simply stops asking. Two obligations, both outcomes rather than procedures. **First, state the exempted class in behavioural terms and check it against the repo's own definition of that class**: an exemption named `not-required` asserts *no behaviour*, so if any rule in the tree already classifies that same diff shape as behaviour-changing, the exemption contradicts it and the answer is to fix the *anchor/evidence form* for that shape, never to drop the evidence. **Second, if the class does carry behaviour, the exemption must be replaced by a way to SUPPLY the evidence** — widen where the anchor may land, add an evidence form the shape can satisfy — because the class with the most behaviour is exactly the one an exemption hurts most. **A precedent of the same shape is not a justification**: reaching for an existing exemption class because the root cause rhymes with an earlier one transfers the solution without checking the disanalogy, and the disanalogy is usually the load-bearing part. Failure shape: a gate anchor that structurally cannot bind to a frontmatter-only change was answered with a third `not-required` class by analogy to two existing ones, even though the same repository elsewhere states that any frontmatter edit is a routing-surface change and the author had just measured its routing delta; the adversarial review caught it on the first finding, and the correct fix was to let the anchor bind to the changed description instead.
|
|
88
|
+
- **Proposal-time corollary.** This check fires **before** a design option list reaches a human, not only before landing. An option list must carry, per option, what evidence that option removes and from which class of change; a `recommended` label asserts that comparison was made. A recommendation resting on *consistency with precedent* rather than on that comparison is unearned — and when a rejected option's stated downside reads back as a true statement about the domain, that option is probably the right one.
|
|
89
|
+
- **(d) premise check for a verdict-tightening change — a clean run on the CURRENT corpus is not evidence.** Fires whenever the change makes the verdict stricter for inputs that previously passed (a reporting check turned blocking, a warning turned error, a widened reject set, a new blocking gate landing over an existing corpus). Today's zero-violation count is the sum of inputs that pass because they are *registered or correct* and inputs that pass by an **incidental property of how they happen to be written**; the tightening silently converts the second group into reds with no legal exit — and the author never sees it, because it fires on whoever edits those inputs next. So verify the premise, not the state: drive the gate's own input toward the worst case it will legitimately see — saturate the incidental property across the whole scanned set, or apply the cheapest edit to the input nearest the threshold — in a **throwaway copy**, never the live tree, and re-run. Then **classify what fails against what the tightening was SUPPOSED to reject** — "it failed" alone is a false green, because on a gate whose whole point is to start rejecting something, the worst-case input fails by design and you learn nothing. An **intended** rejection is the change working. The finding is the other kind: an input that was meant to stay valid and fails only because of an incidental property of how it is written. **The leg is satisfied by an outcome, not by having run the procedure: no such input may be able to hit the tightened verdict with no way out.** So the disposition menu for one of these is closed, with exactly three exits: give it a **legal exit** before landing (register/exempt it, or narrow the rule); keep the tightened behavior **non-blocking** until it has one; or obtain an **explicit risk-owner deferral** carrying a repair path the affected inputs can actually take **at or before the moment the blocking verdict turns on** — a repair that only becomes usable later is the non-blocking exit, not this one, because until it lands the next ordinary edit still hard-fails with nowhere to go. All three name the affected inputs, never an unenumerated blanket. "I ran the procedure" and a bare "accepted cost" are not exits — accepting a residual that still hard-fails the next ordinary edit *is* the outcome this leg exists to prevent. Run the perturb-and-classify exercise **once per newly-tightened predicate and per materially distinct class of affected input** — a change tightening two predicates is not covered by exercising the louder one, and the class you never perturbed is exactly where the next contributor's ordinary edit lands; representative sampling is allowed only when you record why the sampled classes cover the omitted ones. Keep it proportional — scope the saturation to the gate's own input set, it is a one-time design-time run, and a reasoned narrowing is fine while silent skipping is not; where the worst case cannot be constructed cheaply, the premise is `unverified` — and `unverified` is **not a landing state for the blocking form** (an un-run check is not a passed check, the same semantics this repo already gives its `*_unevaluated` size-gate token) — route it through the non-blocking or authorized-deferral exit above, per affected input class.
|
|
90
|
+
|
|
91
|
+
Failure shape (a)–(c): the digest-bound evidence apparatus evaluated and removed in `external-practice-controls.md` (behavioral-evidence-and-attestation) — its full-suite double-arm regeneration cost broke its own author's multi-commit branch, and it survived review as a fix-list item until the maintainer's cost challenge tore it down to a static gate.
|
|
92
|
+
|
|
93
|
+
Failure shape (d): a route-existence check flipped from warn-only to blocking on a corpus measured at zero violations. The zero was partly accidental — a canonical vocabulary token that the check's own exemption tables never registered sat on several lines that happened to carry no trigger word, so the first ordinary wording edit to any of them would have reddened CI with a diagnostic telling the author to restore something that never existed. Legs (a)–(c) all passed: the author's own branch was green, the marginal cost of a typical edit was zero, and the trust model was sound. Six external review rounds missed it too — they read the diff, and the evidence was a statistic over the corpus *outside* it. One saturation run surfaced every affected line at once.
|
|
94
|
+
|
|
95
|
+
### Green verdicts that were structurally incapable of being red (relocated from `SKILL.md` §Self-audit)
|
|
96
|
+
|
|
97
|
+
The self-audit rule says to prove the oracle can fail before trusting its clean verdict. Two shapes make that proof skippable because the verdict *looks* like it came from the thing you meant to run:
|
|
98
|
+
|
|
99
|
+
- **A pipeline reports its LAST stage's status.** `make test | tail -20` exits with `tail`'s status, so a failing `make` reads as exit 0; the same holds for `| grep`, `| head`, and any `$?` read after a pipe. Redirect and read the command's own status (`cmd > log 2>&1; echo $?`) or set `pipefail`. Observed: a full-suite failure was reported as passing, and the mistake was caught only because the summary line the suite prints on success was missing from the captured tail.
|
|
100
|
+
- **A check run against the wrong tree passes silently.** When the work lives in a git worktree, an ambient `cwd` the harness may reset between tool calls sends relative-path commands to the primary checkout, which is clean — so the gate evaluates a tree that does not contain the change. Resolve the target by absolute path (`git -C <abs>`, `bash <abs>/script.sh <abs>`); `worktree-isolation` owns that discipline. Observed: a catalog test's pristine-tree case passed against the primary checkout while the real candidate was blocked.
|
|
101
|
+
|
|
102
|
+
Both are the same defect as a check that can only ever say clean: if you cannot produce the red path on demand, the green is not evidence.
|
|
103
|
+
|
|
104
|
+
### A new gate owes the negative half (relocated from `SKILL.md` §behavioral-evidence)
|
|
105
|
+
|
|
106
|
+
A `RED-baseline` row proves the gate fires on what it must catch. That is one side. The other side — proof it does **not** fire on untouched, legacy, or out-of-scope inputs — is the half authors skip, because those cases feel uninteresting and the suite is already green without them. A gate that also blocks what it should spare is a regression wearing a green suite, and it lands where it hurts most: on repositories and fixtures that never opted into the change.
|
|
107
|
+
|
|
108
|
+
So a new or tightened gate records spare-case rows beside its block-case rows. Observed: a catalog gate shipped with five block-case rows, all green; nothing checked what it did to a repository whose catalog predated the format, and a nested test fixture found the hard failure instead of the author.
|
|
109
|
+
|
|
110
|
+
### Residency in the always-on layer is a measurable claim (relocated from `SKILL.md` §behavioral-evidence)
|
|
111
|
+
|
|
112
|
+
The always-on injection layer (`agent-context/session-start.md`) enters every session on every host and carries a zero-net-growth budget enforced by `scripts/check-size-budget.sh`. "This skill/rule has to be resident" therefore spends a scarce, contested resource — and before there was a way to measure it, the claim could only be argued.
|
|
113
|
+
|
|
114
|
+
Measure the A/B first: `scripts/eval-routing-bank.rb --with-bootstrap` runs the task bank with and without the injected entry-routing region, so the delta the layer buys is observable. Spend bytes only on a delta that shows up. When the measured delta is zero the honest landing is to leave the content out — not to argue the budget down, and not to trade an existing rule away to fund it. Observed: seven skills a plan asserted were routing gaps showed 19/22 → 19/22 with the full skill listing and 20/21 → 21/22 with descriptions truncated to 250 chars; they stayed out, and the layer ended smaller than it started.
|
|
115
|
+
|
|
116
|
+
### Verify a load-bearing premise before designing on it; non-convergence is a premise smell
|
|
117
|
+
|
|
118
|
+
The self-audit rules above catch a claim you make about your own change. This one catches
|
|
119
|
+
the claim you never examined because it felt like background fact: who calls this script,
|
|
120
|
+
what the packaging actually ships, what a field in a gate's output means, which states are
|
|
121
|
+
reachable at all. Those are premises about the system's own topology, and a design built on
|
|
122
|
+
a wrong one cannot converge no matter how many rounds it runs.
|
|
123
|
+
|
|
124
|
+
**The trigger is a design decision, scope cut, or escalation whose correctness depends on
|
|
125
|
+
such a claim.** Before the design work — not after the third round — falsify it with the
|
|
126
|
+
cheap local command that settles it: grep the callers, read the gate source at the line that
|
|
127
|
+
defines the field, list the packaging roots. The predictor of these failures is not that
|
|
128
|
+
reasoning was unavailable; it is that verification was cheap and skipped.
|
|
129
|
+
|
|
130
|
+
**Escalation corollary — the dangerous one, because it looks responsible.** When successive
|
|
131
|
+
review rounds keep surfacing new instances of the same class, the documented reading is a
|
|
132
|
+
design smell (question whether the capability should exist). Check the premise first: if a
|
|
133
|
+
premise is wrong, rounds will not converge and the pattern is indistinguishable from a
|
|
134
|
+
genuine design dilemma. Handing that non-convergence to a human as a decision *feels* like
|
|
135
|
+
responsible escalation while the missing input is one command, not a judgment. Escalate only
|
|
136
|
+
after the premises the tradeoff rests on have been falsified or confirmed.
|
|
137
|
+
|
|
138
|
+
- **A design decision, scope cut, or escalation must not rest on a premise about the system own topology that a cheap local command would settle** — grep the callers, read the gate source at the line defining the field, list the packaging roots — **and non-convergence across review rounds must never be escalated as a design tradeoff before those premises are checked.**
|
|
139
|
+
|
|
140
|
+
Observed, both inside one session: a shared gate was redesigned to stay compatible with
|
|
141
|
+
downstream repositories that a one-line grep showed it is never installed into; and three
|
|
142
|
+
rounds of design were spent on a state ("a checkout without the always-on layer") that the
|
|
143
|
+
caller list showed only test fixtures can reach — the third round was escalated to the
|
|
144
|
+
maintainer as a design tradeoff, and their one question about the premise closed it. A third
|
|
145
|
+
instance in the same session shows the same shape at reporting altitude: an empty background
|
|
146
|
+
output file plus a stale status snapshot were reported as "the commit did not land" instead
|
|
147
|
+
of being re-read.
|
|
148
|
+
|
|
149
|
+
## When dual-track is mandatory
|
|
150
|
+
|
|
151
|
+
| Extraction type | Review | Challenge |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| New shared skill | required | required |
|
|
154
|
+
| Deep extraction (multi-batch over multiple sessions) | required | required |
|
|
155
|
+
| Multi-skill landing (≥ 2 skills changed) | required | required |
|
|
156
|
+
| Skill that ships operational rules (deploy, rollback, mesh, secret, audit) | required | required |
|
|
157
|
+
| Skill that ships architectural rules (boundary, lifecycle, retention) | required | required |
|
|
158
|
+
| Skill that ships security-sensitive rules (CORS, auth, secrets, callbacks) | required | required |
|
|
159
|
+
| Skill description / frontmatter rewrite (changes to triggers, Proactively-invoke, Skip-when, or capability statement) | required | required |
|
|
160
|
+
| Wording-only shared-skill edit (typo, grammar, formatting — NOT triggers, NOT routing, NOT description content) | required | not required |
|
|
161
|
+
| Single trivial shared-skill update with any non-wording change | required | required |
|
|
162
|
+
| Single trivial non-shared-skill update with no operational/security claims | optional | not required |
|
|
163
|
+
| Generator-owned shared skill regenerated through its tool | required independent review; generator validation is additional, not a substitute | required for any non-wording shared-skill change |
|
|
164
|
+
| Generator-owned non-shared skill regenerated through its tool | required (the generator's own) | required if rules changed |
|
|
165
|
+
|
|
166
|
+
Only the challenge pass may be skipped, and only when this table marks challenge not required; record an explicit `challenge: not-required, reason: ...` row in the validation log. Independent review for shared-skill changes has no skip row. A required review or challenge that is missing, inconclusive, or skipped blocks commit and landing; the work may only be reported as an uncommitted interim checkpoint until the required pass succeeds.
|
|
167
|
+
|
|
168
|
+
**Canonical wording-only criterion + the deterministic scope check (challenge-skip gate).** This is the single source of truth the L0/L1/L2 risk view (`l0-l1-l2-routing.md`) defers to; do not redefine it elsewhere. An edit qualifies as `wording-only` — and so may take the challenge-not-required row above — **only when BOTH** of the following hold, never on the author's say-so:
|
|
169
|
+
|
|
170
|
+
- **(a) Bounded change class.** The edit changes ONLY typo, grammar, formatting, or a meaning-preserving synonym, and changes NO trigger, scope, routing, validation, acceptance, rule/threshold/boundary text, or any other meaning. Reference/body prose that states a rule, threshold, boundary, rubric, or applies/does-not-apply line IS a semantic surface — editing it is NOT wording-only unless the change is purely typo/grammar/formatting with no meaning change. A synonym substitution usually cannot satisfy (b)'s deterministic evidence bar unless the proof avoids intent/meaning judgment. Description / frontmatter is never wording-only (see below).
|
|
171
|
+
- **(b) Deterministic scope check + independent review.** Dropping challenge for the edit requires a **deterministic scope check** — recorded controller-side or diff-based scope evidence that is decidable without judging intent or meaning, such as: touched files are formatting-only by formatter output; hunks are only meaning-inert whitespace / punctuation / markdown table alignment; or token-level changes are limited to a named typo correction while the surrounding rule sentence is byte-identical. If the proof depends on a human or LLM deciding whether revised prose changes a rule, threshold, boundary, applicability, or acceptance meaning, it is NOT deterministic and challenge stays required. The deterministic evidence **AND** an independent review row confirming the same must both be present. Either piece missing, or any reviewer-flagged / unconfirmed meaning / scope / trigger / routing / validation / acceptance change, **re-arms challenge + the behavioral-evidence row** (never demoted to "recommended").
|
|
172
|
+
|
|
173
|
+
**If no such deterministic scope check can be formed for the change, challenge stays required.** This is a hard gate, not a default the author may waive: an LLM independent review is hypothesis-grade, so review alone never downgrades a non-wording change. Every shared-skill change still requires the independent review row regardless of class.
|
|
174
|
+
|
|
175
|
+
If either required pass times out, returns empty output, is rate-limited, cannot access auth, exits nonzero, emits malformed or truncated output, fails JSON/shape parsing when structured output was requested, lacks evidence that the pass inspected the target diff/files, shows a prompt/tool-scope mismatch, or returns any `inconclusive` status, the dual-track gate has not passed **for that lane via that reviewer**. A *recoverable-lane* failure — auth, quota, rate limit, timeout, local cache/db failure, or missing capability — is NOT a terminal stop and is NOT "the gate is unrunnable on this host": before you record `blocked`, you MUST walk the **Primary reviewer failure** remediation ladder below and route to an approved independent third-party reviewer (preferably a different model family) — either your runtime's native multi-model subagent (e.g. an OpenCode `Task`/council subagent on a separate model) or a shell wrapper such as `opencode_review.sh --model <provider/model> --implementer-family <author-family> --mode review|challenge`. A primary CLI lacking auth is a routing trigger to that fallback lane, never a license to declare the gate unrunnable. Only after that ladder is exhausted — every approved independent reviewer probed and unavailable — do you record the row as `pending` or `blocked`, include the remediation attempted (which ladder steps were tried) and the next unblock action, and do not describe the skill change as solved, complete, landed, or fully closed. (A large reviewer INPUT can also be silently middle-truncated, not just the reviewer's OUTPUT — see **Read-coverage of large inputs** under *Sanity checks the gate must enforce*.)
|
|
176
|
+
|
|
177
|
+
**Description / frontmatter rewrites are NOT wording-only.** The `description` field is the routing-surface that AI clients match user requests against to decide which skill to auto-invoke; changing it changes which asks the skill catches and which sibling skills it collides with. Treat every description rewrite as a multi-skill landing for the purposes of this gate, even when only one file changed. See `references/description-authoring.md` for the structure and validation checklist.
|
|
178
|
+
|
|
179
|
+
## Behavioral-evidence row (required for every non-wording shared-skill change)
|
|
180
|
+
|
|
181
|
+
Review + challenge check the change for defects; neither checks whether it actually moves agent behavior the intended way. Record one behavioral-evidence row per non-wording shared-skill change, in the same validation file. The building blocks already exist — pressure scenarios (`references/validation-and-landing.md`) and the test-case-first hard-check in `check-ccl-skills.sh`; this row makes recording one of them mandatory and uniform.
|
|
182
|
+
|
|
183
|
+
**Primary status** (pick one):
|
|
184
|
+
|
|
185
|
+
| Status | Use when |
|
|
186
|
+
|---|---|
|
|
187
|
+
| `RED-baseline` | the change alters behavior or routing (trigger / scope / routing / validation / acceptance). Run the scenario WITHOUT the change first (baseline failure), then WITH it (compliance) |
|
|
188
|
+
| `semantic-control` | a non-wording but semantic-preserving mechanical refactor (e.g. mega-bullet split per the B0 checklist). The **reviewer** confirms NO change to trigger / scope / routing / validation / acceptance, and an existing scenario or control still behaves identically. NOT for pure formatting — that is wording-only and needs no row |
|
|
189
|
+
| `not-applicable: docs-only` | the change touches NO file under `skills/**` and no skill-loaded guidance — i.e. `README` / `ARCHITECTURE` / `CONTRIBUTING` / `docs/**` only. Forbidden for `SKILL.md`, `references/**`, validators, templates, examples, and the plugin-shipped command/behavior surfaces (`hooks/*`, `scripts/install.sh`, `bin/`, `.mcp.json`, `.lsp.json`, `monitors/**`, `settings.json`): those are behavioral or executable source even when they read like prose |
|
|
190
|
+
|
|
191
|
+
For a `RED-baseline`, the evidence form can be a before-after task diff, a golden trace, or a pressure scenario — these are *how* you show baseline→compliance, not standalone substitutes for it. For a routing-surface / hub-skill change you can run the golden-trace form as a REAL headless-agent run via the F4 Tier-3 harness (optional, higher-fidelity than a recorded scenario — see `validation-and-landing.md` Behavioral Validation); a recorded scenario is not by itself a failing baseline.
|
|
192
|
+
|
|
193
|
+
**A change that ships a DESTRUCTIVE or irreversible operation** (a script/recipe that deletes/overwrites/prunes worktrees, branches, files, records; bulk or `--force`-class mutation) must EXECUTE its protected/NEGATIVE cases as part of the evidence — run it (in a throwaway repo/dir) and assert it does NOT touch what it must keep (unmerged commits, dirty/untracked/ignored files, the protected/default targets), not only that it acts on the intended input. A positive-only test ("it removed the merged one") does not establish safety; the data-loss footguns hide in the must-NOT-touch set, and the adversarial challenge will hunt exactly there (expect ~3–4 rounds for a stateful destructive recipe). Dry-run-default + a no-`--force` second net are design safeguards, not a substitute for executing the negative cases. **Executing the negative cases is necessary but NOT sufficient — the probes must be SENSITIVE, and this is the half that ships blind.** A negative probe that would still pass with the protecting predicate removed is evidence of nothing, and a suite of them reads as thorough coverage: the recurring shape is a degenerate fixture that makes every probe short-circuit on an unrelated conservative branch, so the safety predicate is never reached and a green suite certifies a hole. So the evidence row for a destructive change records, per protected predicate, that its removal was **applied** and observed to turn the suite RED — an unapplied "this mutation would fail it" is a hypothesis — and for the encoded form of that walk `testing-strategy` owns the rule (route, don't copy). Failure shape: a 9-probe negative suite for a worktree-pruning script passed fully green with its "only remove an actually-merged branch" predicate stubbed to always-true, because the fixture put the integration ref at the default branch's tip.
|
|
194
|
+
|
|
195
|
+
**Skill-TYPE refines the evidence FORM, never the status** (advisory; only when creating or substantially authoring a *whole* skill, not every small extraction). FIRST pick the primary status from the table above based ONLY on the diff — a whole new skill almost always adds a routing/discovery surface plus acceptance/use behavior, so it is a `RED-baseline` even when its CONTENT is Reference material; skill type never downgrades that status and never invents a fourth one. THEN use the skill type (classify per `superpowers:writing-skills` if installed, else inline — **Technique** = a method with steps; **Pattern** = a way of thinking; **Reference** = API/syntax/lookup material; plus **discipline-enforcing** = a rule/gate, which in `writing-skills` is a 4th *test category*, NOT a 4th Skill Type and NOT a 4th status) only to choose the evidence ARTIFACT within the already-chosen status: Technique → a RED-GREEN behavior trace; Pattern → applied-to-a-real-scenario evidence; Reference → source fact/coverage PLUS a lightweight retrieval + application + gap check (per `writing-skills`: can an agent find the right entry, apply it correctly, and are common cases covered — "correct docs, unusable skill" is the failure to catch), not link-correctness alone; discipline-enforcing rule/gate (TDD-like — and most of THIS workflow's own rules, e.g. worktree-isolation and the dual-track gate, are this kind) → its `RED-baseline` artifact is a **pressure scenario** (does the agent comply under combined stress — time / sunk-cost / exhaustion — not merely understand the rule academically). Type does NOT widen `not-applicable: docs-only`: a Reference-type skill under `skills/**` is still a shared-skill change needing review/challenge and a recorded evidence row.
|
|
196
|
+
|
|
197
|
+
Rules:
|
|
198
|
+
|
|
199
|
+
- `RED-baseline` is required whenever the change alters behavior or routing. `semantic-control` is valid ONLY with reviewer confirmation that none of trigger/scope/routing/validation/acceptance changed — an author cannot self-assert it.
|
|
200
|
+
- The row must give a concrete locator + evidence shape, not a bare status: artifact path / commit / transcript / command, the exact prompt or scenario, and expected-vs-actual. For `RED-baseline`, record BOTH the without-change (baseline failure) and with-change (compliance) results, and name the baseline's **provenance type**: a *recorded incident* (cite where the failure is actually recorded — transcript, note, issue; the cited record must describe a failure that occurred, not prescribe a method) or a *constructed scenario* (run against BOTH the unchanged baseline and the changed rule, with an openable artifact for each run — a scenario "run" only mentally, only against the patched text, or only as reviewer discussion does not count). A prescriptive source — a method-bar or best-practice note with no failure recorded — cannot be cited as an occurred failure and is not by itself a valid `RED-baseline`; it may seed the constructed scenario's design or support the rule's rationale, but the RED evidence is the run artifact. Narrating what "would have" failed as if it happened is a fabricated evidence row, the same defect class as fabricated verification output. A status word with no openable artifact is not a valid row, same as a missing review row.
|
|
201
|
+
- A missing or unreviewable behavioral-evidence row blocks landing the same way a missing review row does; until it exists the work is an uncommitted interim checkpoint.
|
|
202
|
+
|
|
203
|
+
## Running the review pass
|
|
204
|
+
|
|
205
|
+
> **Pick the reviewer — route through the owning wrapper before hand-rolling.** The gate needs an independent reviewer; it is **tool-agnostic**. Invariants regardless of tool: prefer a model from a **different family than the author** (cross-model catches shared blind spots — it is symmetric: Claude-authored → OpenAI-family reviews, OpenAI-authored → Claude/Moonshot reviews), and treat any sign-off as **hypothesis-grade** (verify load-bearing claims against primary sources). The agent running the gate, before reviewing:
|
|
206
|
+
>
|
|
207
|
+
> 1. **Resolve the organization `code-review` gate first.** It owns the installed-client checks, local `CODE_REVIEW_CLIENT_ORDER`, same-family exclusion, frozen packet, timeout, egress approval, tool boundary, and verdict parsing for Claude, Kimi, OpenCode, and Codex. Do not preselect a client from `command -v` output.
|
|
208
|
+
> 2. **Pass the actual implementer family.** A known same-family Claude, Moonshot/Kimi, or OpenAI/Codex client is excluded before inference; OpenCode is excluded after its exported session binds the actual provider/model family. An unmapped family is inconclusive, not permission to guess.
|
|
209
|
+
> 3. **Do not add a separate model behavior probe.** Claude, Kimi, and Codex validate the formal invocation stream; OpenCode additionally uses only the non-inference `debug agent ccl-review` structural check before its formal run. Availability/help checks are local and are not review evidence.
|
|
210
|
+
> 4. **Use another approved wrapper or runtime-native independent lane only when the organization gate itself is absent or cannot start.** Preserve the same bounded packet, read-only/no-exec posture, independent family, reviewer-emitted verdict, and review-vs-challenge separation. Raw CLI commands are debugging/reference shapes, not a shortcut around a terminal ccl-gate result.
|
|
211
|
+
> 5. **Run review and challenge separately.** Each result must bind its own mode, packet, selected reviewer, family, and verdict. A non-empty finding remains a finding; the controller never rewrites it into a pass.
|
|
212
|
+
|
|
213
|
+
### Primary Reviewer Unavailable
|
|
214
|
+
|
|
215
|
+
Primary reviewer failure is a remediation branch only when the owning gate classifies it as candidate-local. Use this ladder separately for the review lane and challenge lane:
|
|
216
|
+
|
|
217
|
+
1. Persist the owner-guided self-review, pass it through the required `--review-plan-file`, and run `review_gate.sh` once for the lane. A host-returned live `session_id`/execution handle is still that same run: poll it to terminal exit and do not start a replacement or fallback from its empty current output. Record the handle type, opaque host transcript/tool-call reference and terminal exit. If the handle is lost, the lane is infrastructure-inconclusive/manual-review-required and no replacement or fallback may be started or credited; process-tree and wrapper artifacts are diagnostic only. This is a procedural host obligation because the inner gate cannot observe the outer handle; machine enforcement requires a trusted host adapter. Never persist a credential-like raw handle in shared evidence. If Claude then returns `auth_path_unavailable`, perform the one documented host rerun with the same frozen candidate, plan, stage, and `--host-remediation-attempted`.
|
|
218
|
+
2. Let the gate continue only for its allowlisted candidate-local classes: missing client/provider, bounded auth failure, quota/rate limit, timeout, missing capability, or malformed model output. It records every skipped/attempted client.
|
|
219
|
+
3. Packet/input/binding/tool-boundary, egress, same-family, mode, and unknown failures are not manually bypassed. A terminal result stops that lane.
|
|
220
|
+
4. If the organization gate could not start at all, an approved alternate wrapper or runtime-native lane may be used only with the same bounded packet, independent family, no-write/no-exec boundary, attribution, and parseable verdict. Do not invent a one-off provider/model chain.
|
|
221
|
+
5. A fallback result satisfies only the exact lane it ran. If no candidate returns a conclusive verdict, keep the work `interim`.
|
|
222
|
+
|
|
223
|
+
Do not call a manual ad-hoc run "fallback review" unless it meets the same evidence bar. Correcting a CLI argument mistake and rerunning is remediation; waiting forever, killing the process, or accepting partial stdout is not evidence.
|
|
224
|
+
|
|
225
|
+
- **Open the Agent chain on the FIRST review — it cannot be retrofitted.** An extraction's required review and challenge are one tracked multi-round run, so the round budget is decided before round 1, not after reading the review; a run that starts untracked is thrown away and restarted. Every trigger, default, flag, index, prior-result, and advisory rule behind that obligation is owned by `code-review/references/staged-review-contract.md` (Agent review chain), with the runnable pair in `code-review/SKILL.md` — take the command from there and never reconstruct it from this bullet.
|
|
226
|
+
|
|
227
|
+
- **Compose the packet — it is the mechanism that decides which finding classes are reachable at all.** The lever and its constraints are owned by `code-review`'s `SKILL.md` (packet-bounded reviewer; `--paths` only narrows; `--diff-file` supplies a packet you assembled), including the rule that an "input insufficient to judge" finding is an input defect rather than a candidate defect. What this workflow adds is the shared-skill inclusion list: alongside the diff, carry (a) the canonical rule or contract text the changed clause must not contradict, (b) the sibling clauses in the same section or file, (c) the derived carriers that restate the change — commit message, MR body, register row, `description` surface, (d) the actual output of any gate or script the change touches. Measured over one 11-round gate on a prose-rule change, a diff-only packet surfaced only defects in the tail of the just-edited sentence; the composed packet is what surfaced cross-clause contradiction, cross-carrier drift, and silent weakening of the canonical wording. Reach for packet composition before inventing another prose rule or a wording-level grep for the same defect class, and pick between candidate mechanisms by their hit rate over the round's actual findings, not by whether they feel in scope.
|
|
228
|
+
|
|
229
|
+
Standard `codex review` against the target diff:
|
|
230
|
+
|
|
231
|
+
```bash
|
|
232
|
+
cd <skills-repo>
|
|
233
|
+
# Current codex: a custom PROMPT is mutually exclusive with --base / --uncommitted / --commit.
|
|
234
|
+
# Give a prompt (default scope = working-tree uncommitted changes):
|
|
235
|
+
timeout 540 codex review "<prompt>" \
|
|
236
|
+
-c 'model_reasoning_effort="high"' \
|
|
237
|
+
-c notify='[]' \
|
|
238
|
+
--enable web_search_cached
|
|
239
|
+
# …or a scope flag with NO custom prompt: codex review --base <base-sha-or-branch> -c '…'
|
|
240
|
+
# Flag arity is version-sensitive — run `codex review --help` if either form errors.
|
|
241
|
+
# (codex review reads stdin only when PROMPT is `-`; the </dev/null gotcha below is codex-exec-specific.)
|
|
242
|
+
```
|
|
243
|
+
|
|
244
|
+
Prompt template (adapt as needed):
|
|
245
|
+
|
|
246
|
+
```
|
|
247
|
+
IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/.
|
|
248
|
+
Stay focused on the diff.
|
|
249
|
+
|
|
250
|
+
Focus: <one-paragraph context — what was added/changed, what the references document>.
|
|
251
|
+
|
|
252
|
+
Specifically check:
|
|
253
|
+
1. Internal consistency across the touched references (any contradictions?).
|
|
254
|
+
2. Sanitization gaps (any residual ccl-specific names, IPs, hostnames, namespace names, vendor terms that look like real internal artifacts).
|
|
255
|
+
3. Technical accuracy of generic claims (etcd Watch semantics, Istio VirtualService routing, framework patterns, OTel/Prom/VM sizing, gRPC quirks, k8s control-plane behavior — whatever the diff touches).
|
|
256
|
+
4. Over-prescription: is anything written as MUST when industry has multiple acceptable patterns?
|
|
257
|
+
5. Cross-skill handoff coherence: boundaries clear, no overlap or gap between the touched skills.
|
|
258
|
+
6. Anti-patterns: is anything in 'Common Pitfalls' misclassified?
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
Output: one finding per line with severity (P0/P1/P2), file:line, scenario, fix. Apply each finding before the second pass.
|
|
262
|
+
|
|
263
|
+
## Running the challenge pass
|
|
264
|
+
|
|
265
|
+
`codex exec` with adversarial prompt (read-only):
|
|
266
|
+
|
|
267
|
+
```bash
|
|
268
|
+
timeout 540 codex exec "<adversarial-prompt>" </dev/null \
|
|
269
|
+
-C <skills-repo> \
|
|
270
|
+
-s read-only \
|
|
271
|
+
-c 'model_reasoning_effort="high"' \
|
|
272
|
+
-c notify='[]' \
|
|
273
|
+
--enable web_search_cached \
|
|
274
|
+
--json
|
|
275
|
+
# timeout + -c notify='[]' inline on purpose (see notes below): a stuck round-end
|
|
276
|
+
# notifier otherwise hangs a finished run; a timeout-kill (exit 124) = inconclusive lane, not a pass.
|
|
277
|
+
# BUT only disable a NON-policy notifier: if the host's notify is an audit/DLP/transcript control,
|
|
278
|
+
# drop -c notify='[]' and use a bounded non-blocking wrapper instead (see notes) — don't skip a required control.
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
Redirect stdin from `/dev/null` (**`codex exec` specifically**): `codex exec` reads instructions from stdin, so when the prompt is large, or it is wrapped in `timeout` / run in the background, it otherwise blocks on `Reading additional input from stdin...` and returns no output — it looks like a hang but is just waiting on stdin. `</dev/null` makes it read only the prompt argument. If a `codex exec` call sits at zero output, suspect a missing `</dev/null` before killing it. (`codex review` reads stdin only when PROMPT is `-`, so it does not have this gotcha — a quiet `codex review` has another cause.)
|
|
282
|
+
|
|
283
|
+
A **second, distinct** hang cause is the round-**end** `notify` hook, not the start-of-turn stdin read above: when the host's codex config sets a `notify` program (run on turn completion), codex waits for that program to return before exiting, so a `notify` command that blocks makes an otherwise-finished `codex exec`/`codex review` hang at the *end* with the full answer already produced. Neutralize it per-invocation with `-c notify='[]'` (override to no notifier — per-invocation only, so a notifier the host configured for normal runs is untouched). Only blanket-disable a **non-policy desktop/notification** hook; if the configured notifier is an audit / DLP / transcript-capture control, preserve it with a bounded non-blocking wrapper instead of disabling it, or you produce a clean-looking reviewer lane that skipped a required control. Always wrap the call in `timeout` so a stuck notifier bounds to a failed lane instead of an indefinite hang. A `timeout`-killed call (exit 124) or any partial/nonzero result is an **inconclusive lane per the failure rule above — never recorded as a clean "no findings" pass**; the point is to fail the lane fast and remediate, not to accept a truncated review. Verify the override for the installed version before relying on it — `codex --help` / `codex exec --help`, or a strict-config probe (`codex --strict-config -c notify='[]' <subcmd>` errors out on an unrecognized key, stays clean when `notify` is valid) — codex config keys drift across versions (there is no `codex config` subcommand to consult).
|
|
284
|
+
|
|
285
|
+
Prompt template:
|
|
286
|
+
|
|
287
|
+
```
|
|
288
|
+
IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/.
|
|
289
|
+
|
|
290
|
+
ADVERSARIAL CHALLENGE mode. Review the new/changed reference files at:
|
|
291
|
+
<list the files explicitly>.
|
|
292
|
+
|
|
293
|
+
Your job: find ways an engineer who follows these references LITERALLY will produce
|
|
294
|
+
a production failure, security incident, data loss, or operational disaster. Be
|
|
295
|
+
brutal. Think like a chaos engineer + attacker + paranoid SRE + auditor.
|
|
296
|
+
|
|
297
|
+
Specifically hunt for:
|
|
298
|
+
1. Race conditions in described workflows.
|
|
299
|
+
2. Data loss / corruption paths.
|
|
300
|
+
3. Security holes (auth bypass, replay, credential leakage, CORS, mesh bypass).
|
|
301
|
+
4. Failure cascades (control plane down → no rollback; partition; upgrade ordering;
|
|
302
|
+
trust-domain misconfig).
|
|
303
|
+
5. Algorithm flaws (thresholds at edge values, sample-size assumptions, retry budgets,
|
|
304
|
+
timeout chains).
|
|
305
|
+
6. Operational footguns (broad-glob cleanups, single-instance "long-term" stores,
|
|
306
|
+
HPA thrashing, override paths too easy).
|
|
307
|
+
7. Subtle inconsistencies between references.
|
|
308
|
+
8. Things readers would do WORSE than ad-hoc scripts if they followed the reference
|
|
309
|
+
verbatim.
|
|
310
|
+
9. Gate-fireability / bypass-by-omission (REQUIRED when the change adds or edits a
|
|
311
|
+
rule/gate/status/verdict — "when X, do/block Y"). Can Y be reached WITHOUT ever
|
|
312
|
+
producing X? Is the trigger condition X a mandatorily-recorded field/output of an
|
|
313
|
+
upstream step, or an optional signal a reader can simply never emit (no verdict
|
|
314
|
+
recorded = never "rejected" = ships)? Who is allowed to set X — can the gated party
|
|
315
|
+
self-adjudicate it? A gate whose trigger is never mandatorily produced is decorative.
|
|
316
|
+
List every literal path that reaches the gated outcome while skipping the gate.
|
|
317
|
+
|
|
318
|
+
Output one finding per issue with: severity (P0 production-down / P1 high / P2
|
|
319
|
+
medium), file:line, the failure scenario in concrete steps, and the fix. Be terse.
|
|
320
|
+
No compliments.
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
For a rule/gate/status change specifically, item 9 is not optional padding: a generic
|
|
324
|
+
challenge (items 1–8) will pass a decorative gate because none of those classes asks
|
|
325
|
+
"is the trigger ever produced." The failure shape this prevents: a gate keyed on a
|
|
326
|
+
`rejected`/approved/verified verdict that nothing requires anyone to record, so the
|
|
327
|
+
gated outcome ships by simply never emitting the verdict — and a challenge run without
|
|
328
|
+
item 9 signs off on it.
|
|
329
|
+
|
|
330
|
+
Use `--json` to capture reasoning traces and tool calls cleanly. Parse the JSONL stream with a small Python or jq script as documented in `gstack-codex` skill.
|
|
331
|
+
|
|
332
|
+
## Recording findings + fixes
|
|
333
|
+
|
|
334
|
+
For each pass, record in the extraction's working file (e.g. `<project>-extraction-summary.md`):
|
|
335
|
+
|
|
336
|
+
```
|
|
337
|
+
## Review pass (codex review)
|
|
338
|
+
- Findings: N total (a P0 / b P1 / c P2 / d P3)
|
|
339
|
+
- R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
|
|
340
|
+
- Applied: M fixes (commit: <sha>)
|
|
341
|
+
- Deferred: <list with reason>
|
|
342
|
+
|
|
343
|
+
## Challenge pass (codex exec adversarial)
|
|
344
|
+
- Findings: N total (a P0 / b P1 / c P2)
|
|
345
|
+
- R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
|
|
346
|
+
- Gate-fireability applicability: <yes — change adds/edits a semantic rule/gate/status/verdict | no — valid ONLY when the diff is wording-only or adds/edits no semantic rule/gate/status/verdict>
|
|
347
|
+
- Item 9 exercised: <locator to the captured prompt/transcript/JSONL showing the bypass-by-omission probe actually ran (not a pasted self-assertion) | n/a per line above>
|
|
348
|
+
- Applied: M fixes (commit: <sha>)
|
|
349
|
+
- Deferred: <list with reason>
|
|
350
|
+
|
|
351
|
+
## Behavioral evidence
|
|
352
|
+
- Status: <RED-baseline | semantic-control | not-applicable: docs-only>
|
|
353
|
+
- Artifact + outcome: <locator: path/commit/transcript/command> · <exact prompt/scenario> · <expected vs actual> (RED-baseline: include both without-change and with-change results; semantic-control: name the reviewer who confirmed no behavior change)
|
|
354
|
+
```
|
|
355
|
+
|
|
356
|
+
Each fix lands as its own commit when meaningful; small fixes can batch. The commit message lists the issues addressed (by review-pass severity + short description).
|
|
357
|
+
|
|
358
|
+
## Sanity checks the gate must enforce
|
|
359
|
+
|
|
360
|
+
- Both passes ran against the actual target diff, not an older snapshot.
|
|
361
|
+
- All P0 findings have an applied fix OR an explicit acceptance with documented mitigation in a deferred-fix note + a tracking ticket.
|
|
362
|
+
- All P1 findings have an applied fix OR an explicit deferred row with reason.
|
|
363
|
+
- P2/P3 findings may be deferred more freely; track for the next pass.
|
|
364
|
+
- Re-run sanitization after fixes (fixes can introduce new leakage).
|
|
365
|
+
- Re-run `check-ccl-skills.sh` after fixes.
|
|
366
|
+
- **Inspect the validator output for `alias_audit_unavailable`.** `check-ccl-skills.sh` exits 0 even when the private R0 alias audit never ran (`ALIAS_AUDIT_CMD` unset); in that case its final token is `ccl_skill_check_interim_ok` (with `r0_status=public-fallback`), not the clean token. When unset it runs the public fallback (`generic-r0-leak-scan.sh`) and prints `generic_r0_leak_scan_ok` + `alias_audit_unavailable`; that fallback is diff-scoped public evidence, not the private audit. Only a clean private audit ends on `ccl_skill_check_clean_ok` (with `r0_status=private-ok`); a deprecated `ccl_skill_check_ok` line is still printed for old consumers but is no longer the last line or a sufficient signal. Review/challenge evidence MUST read the full output and treat any `alias_audit_unavailable` / `ccl_skill_check_interim_ok` as **private R0 not run** — a no-findings review that ignores it (or that accepts the clean generic fallback as if it were the private audit) is **inconclusive** for landing readiness, not a clean pass. The R0 lane is satisfied only when the output shows `alias_audit_ok` / `ccl_skill_check_clean_ok` or the record cites a named private-profile result (see `r0-leakage-audit.md`); neither the deprecated `ccl_skill_check_ok`, the interim token, nor `generic_r0_leak_scan_ok` alone closes it.
|
|
367
|
+
- A behavioral-evidence row exists with an allowed status and a referenced artifact; `RED-baseline` is used for any change that alters behavior or routing.
|
|
368
|
+
- **Read-coverage of large inputs**: if any touched file, required reference, or supplied diff exceeds ~200 lines or ~8 KiB (`wc -lc`, objective — not the author's to wave off as "not relied on"), the pass is **inconclusive** unless its transcript shows chunked reads (each under **both** limits) covering the changed hunks plus their owning/referenced sections; naming the exact unread line ranges only **downscopes** the verdict to what was covered — any unread changed hunk or owning/referenced section keeps the pass **inconclusive**, it does not make it conclusive — a single whole-file read drops the middle (codex keeps only head+tail past 256 lines / 10 KiB, [openai/codex#6426](https://github.com/openai/codex/issues/6426)). Tool-call presence alone (an `nl -ba`/`sed -n` line) does NOT satisfy this — a single oversized `nl -ba` still drops the middle. This is the review-lane application of the always-on read-in-chunks rule (`agent-context/session-start.md`; detail in `references/source-to-skill-extraction.md#read-in-chunks-large-reads-lose-the-middle`). The 256-line / 10-KiB cutoff is version-drift-prone: record `codex --version` and probe locally when the exact cutoff is load-bearing. (Separately, `project_doc_max_bytes` is a host config budget for project-doc ingestion — verify the version's default — and does NOT affect tool-output truncation.)
|
|
369
|
+
- **For a NON-WORDING change that adds/edits a semantic rule/gate/status/verdict, the challenge row's `R0 evidence`, `Gate-fireability applicability`, and `Item 9 exercised` fields must be present, and the latter two must be affirmative** (a wording-only edit inside a rule sentence — with NO trigger/scope/routing/validation/acceptance meaning change, per the strict line-31 table — is out of scope; mark applicability `no` with that reason). A generic challenge that reused an older prompt without item 9 passes a decorative gate, and a row with no `R0 evidence` field can hide `alias_audit_unavailable` by omission, so a row missing those fields is **inconclusive** for the gate change regardless of finding count, and must be re-run with item 9 plus R0 evidence recorded. This is the canonical validity criterion; the no-findings section and the per-round log reference it rather than restating it.
|
|
370
|
+
- **Tool-boundary capability review heuristic (advisory):** when a wrapper claims `no tools`, `read-only tools`, or an exact tool set, reviewers should distinguish permission allow-rules from availability restrictions and inspect the installed tool's preferred live-help branch, effective argv/config, inherited extension surfaces, and runtime-init data when exposed. A fake CLI that omits the production-preferred flag proves only its fallback, and prompt wording or directory-add flags are not filesystem sandboxes. This is a review lens, not a new landing field/status gate; executable enforcement belongs in the wrapper's capability probes, runtime checks, and regression tests.
|
|
371
|
+
|
|
372
|
+
## When the challenge pass returns "no findings"
|
|
373
|
+
|
|
374
|
+
A genuine no-finding result from a high-reasoning-effort adversarial pass is rare but valid. Validate by:
|
|
375
|
+
- Confirming the prompt actually asked for adversarial finding (not a summary).
|
|
376
|
+
- Confirming the pass read the actual files (look for tool-call lines like `nl -ba <file>` or `sed -n`). For inputs over ~200 lines / ~8 KiB this is necessary but NOT sufficient — a single oversized `nl -ba` still drops the middle; satisfy the read-coverage check above (chunked ranges covering the changed hunks + owning/referenced sections).
|
|
377
|
+
- Confirming the model had time to run (high reasoning + bounded by timeout).
|
|
378
|
+
- For a non-wording rule/gate/status change, confirming the challenge row's `Gate-fireability applicability` and `Item 9 exercised` fields are both present and affirmative (per the sanity-check validity criterion above) — a no-findings pass from a prompt that never probed bypass-by-omission is inconclusive for the gate, not a clean pass.
|
|
379
|
+
|
|
380
|
+
If the pass returned no findings within seconds, treat it as failed (likely a prompt error or auth issue) and rerun.
|
|
381
|
+
|
|
382
|
+
## Iterating: challenge → fix → re-challenge
|
|
383
|
+
|
|
384
|
+
A single challenge pass is not always enough. Fix-ups can introduce new bugs, and challenge passes have stochastic depth — what one pass missed, the next may surface. Plan for multiple rounds when the change is non-trivial.
|
|
385
|
+
|
|
386
|
+
### The pattern
|
|
387
|
+
|
|
388
|
+
Each round produces three classes of finding:
|
|
389
|
+
|
|
390
|
+
1. **Genuine issues from the original change** — what challenge was meant to catch.
|
|
391
|
+
2. **New issues introduced by the previous round's fixes** — e.g. a fix narrows a regex but the narrowed version misses a real case; a fix moves a routing pointer but the new location creates a different collision.
|
|
392
|
+
3. **Pre-existing issues codex notices on second look** — often P2/P3, often deferrable, but worth recording.
|
|
393
|
+
|
|
394
|
+
After applying fixes from round N, re-run the challenge. The next round should produce strictly fewer findings AND no new P0/P1 from the round-N fixes themselves. If round N+1 surfaces a P0 the round-N fix introduced, the fix was wrong — revert or redesign before continuing.
|
|
395
|
+
|
|
396
|
+
### Do not bias the re-challenge (gate integrity)
|
|
397
|
+
|
|
398
|
+
Each round — first and every re-challenge — must give the reviewer the artifact and an open "find any remaining P0/P1" instruction. The one thing to withhold is your own **fix-claim**: any list, changelog, or self-assessed conclusion stating which issues you fixed or that prior findings are resolved ("I fixed X, Y, Z", "all prior findings addressed"). A fix-claim primes the reviewer to confirm those specific items and skim the rest, so a "CONVERGED" verdict certifies a gate that was never actually re-tested. This holds even when every claimed fix is real: findings from the first pass, or newly introduced by the patch, go unexamined. A round run with a fix-claim through *any* channel does not count — re-run it clean before claiming convergence or landing.
|
|
399
|
+
|
|
400
|
+
Run the reviewer in a **fresh, isolated context**. A re-challenge in the same chat/session that already contains "I fixed X/Y" from earlier turns is already primed even if the new prompt is clean — the fix-claim leaked through conversation history and tool logs. Use a separate reviewer invocation (a fresh `codex exec`, a new session) whose only inputs are what you deliberately supply; if you cannot isolate, explicitly audit that no fix-claim sits anywhere in the reviewer-visible context, not just in the prompt you typed.
|
|
401
|
+
|
|
402
|
+
Withhold the fix-claim through every channel that reaches the reviewer, including commit messages in a `git show` / PR-patch / log view, and fix-claims **embedded in the artifact itself** — a changelog entry, PR-template line, or code comment that says "fixed the P1 cancellation race". Redact or omit such review metadata for the gate run (strip or mask in-artifact fix-claim lines and commit messages), not a disclaimer that the text is unverified — a disclaimer still primes. The sole exception: when the fix-claim text *is* the product surface under review (e.g. you are reviewing the changelog's accuracy itself), it stays, because then judging it is the task.
|
|
403
|
+
|
|
404
|
+
Removing the fix-claim must not shrink or distort the artifact. Give the reviewer the **complete, exact landing candidate** — the full file list and full diff for the whole change set, never a path-filtered subset or a hand-assembled patch. A subset is itself a bias: it hides a finding the author didn't think to include (a P1 introduced in a teardown file outside the "interesting" path). Three precision requirements so this can't be gamed:
|
|
405
|
+
|
|
406
|
+
- **Right SHAs.** Diff the verified base (the target merge-base with the branch you'll land into) against the *exact* head that will land (the PR head / landing-candidate HEAD), not an intermediate commit or a stale base. A "complete" diff of the wrong revision range certifies nothing.
|
|
407
|
+
- **Everything that will land.** If staged-but-uncommitted, untracked, or generated files are part of the landing artifact, include them (commit to a verified candidate tree, or require a clean working tree first). `git diff <base> <head>` silently omits them.
|
|
408
|
+
- **Redaction preserves content.** When you strip an embedded fix-claim, neutralize only the narration with a placeholder and keep the file, hunk, and surrounding lines intact. If the line also carries behavior or user-facing content, you cannot mask it without changing what's reviewed — fall back to the product-surface exception and leave it in.
|
|
409
|
+
|
|
410
|
+
Withholding is only about the *fix-claim narration*, never about the *scope or fidelity of code shown*.
|
|
411
|
+
|
|
412
|
+
Neutral, non-leading context is encouraged, not withheld (starving the reviewer causes the opposite failure — a false "no findings"). Safe to supply: the artifact/diff, the original requirement / acceptance criteria, reproduction steps, and the threat model. A scope statement is allowed but must be **additive, never exclusive** — "review the entire artifact for any P0/P1, especially surfaces X/Y/Z", never "focus only on X". An exclusive scope is path-filtering by instruction: the full diff is attached but the reviewer is steered past a P1 introduced elsewhere (a teardown file outside the named surface). Two cases need care:
|
|
413
|
+
|
|
414
|
+
- **Prior accepted/deferred risks**: supply them so settled tradeoffs aren't re-litigated, but only with their acceptance evidence (who accepted, when, why) attached, and tell the reviewer it may re-raise any of them if this change alters their severity. Without verifiable acceptance evidence, present the item as open — an author must not silently relabel a live P0/P1 as "accepted" to suppress it.
|
|
415
|
+
- **Prior reviewer findings**: may be resurfaced as raw open items in the reviewer's own words (never "I fixed X"), but only as the *complete* prior set or a stable reference the reviewer can actually open — never an author-picked subset, since choosing which to resurface re-biases scope exactly like a fix-claim. Any resurfaced material (inline or behind a reference) must be sanitized to reviewer findings plus acceptance evidence only — a linked thread or doc that still contains author fix-claims re-primes through the back door and is not a valid "complete history". If the full history is too large to inline, attach at minimum every prior P0/P1 verbatim plus an accessible pointer to the remainder, and tell the reviewer the omitted lower-severity items still need a severity re-check under this change.
|
|
416
|
+
|
|
417
|
+
(Self-extracted: an agent-runtime reference in this tree was reported CONVERGED off a fix-list-primed re-challenge; an unbiased re-run surfaced real remaining P1s — partial-stream double-dispatch, cancellation-vs-error path split, finalization-vs-idle-wake race, abort-cleanup deadlock.)
|
|
418
|
+
|
|
419
|
+
### Findings, autonomous budget, and human authority
|
|
420
|
+
|
|
421
|
+
A candidate may be claimed review-ready only when **every** remaining P0/P1 finding has a disposition — there are exactly three, each evidence-backed, not the author's word. Missing disposition blocks the readiness claim and the next external review; it does not stop implementation or unrelated work:
|
|
422
|
+
|
|
423
|
+
- **fixed** — the patch resolves it (and the re-challenge that confirms this is unprimed, per the gate-integrity rule above);
|
|
424
|
+
- **accepted** — recorded as `deferred: <reason>` with acceptance evidence (who/when/why);
|
|
425
|
+
- **pre-existing & out-of-scope** — recorded with evidence it existed at the base SHA **and** that this change does not increase its reachability or severity. A latent deadlock the old code never reached but the new code now can is in-scope, not pre-existing. An unaudited "that was already broken" is not a valid disposition.
|
|
426
|
+
|
|
427
|
+
A **scope-cut / out-of-phase** finding (the scope-direction signal in `SKILL.md`) is not a new disposition or severity. On first appearance, before implementing, it is a **classification checkpoint, not a cut**: **(1)** test current-phase impact — any current execution path, data/authority boundary, contract commitment, or diff-added reachability **or severity**; a remediation that merely looks deferred never proves its absence (a live defect whose repair happens to need future machinery is current-phase); **(2)** split a compound finding. These signals can co-occur: if the recurrence is *also* same-risk-class, run the delete-capability existence evaluation (`keep / delete / narrow / replace`) on it too — scope-cut applies only to the no-current-impact residual and never waives that evaluation. The current-phase core takes its normal disposition above at its real severity — P0/P1 via the three dispositions, P2/P3 under normal non-blocking handling — and is **never cut**. Only a residual with **no** current-phase impact — *proven by recorded negative evidence* (the specific execution paths / data-authority boundaries / contracts / reachability actually checked), never by assertion — is a cut candidate: the controller proposes that residual and **a risk owner distinct from that proposing controller reviews the original finding, confirms the impact classification and split, and decides before anything is reverted** (a self-approval fails the `accepted` who/when/why bar; with no distinct risk owner available, do not cut at any severity — not-cutting governs only the scope-cut operation and never downgrades blocking, so the finding then follows its normal disposition, P0/P1 blocking and P2/P3 non-blocking). An approved cut lands as one of the three dispositions above plus a tracked follow-up **recorded now in a durable tracker that the deferred phase OWNS and its own entry/acceptance gate CONSUMES on entry** (a reverse-consumed link, not just a forward reference), with an owner and a reopen trigger, so that phase cannot claim done until the finding is re-dispositioned under the then-current gate. You need not build that gate now, but the tracker must be one the phase's start is bound to read; if no such phase-owned, gate-consumed tracker exists (the reverse link can't be established), do not cut — keep the P0/P1 blocking. The cut thus defers rather than discards, and the controller never self-accepts convergence. Recurrence across 2+ rounds that stays **undispositioned** — or was implemented despite a prior out-of-phase classification — is the reviewer-lane stop/reframe escalation; an already evidence-backed accepted/deferred finding that a later unbiased challenge legitimately repeats is not, **but only after it is re-verified against the exact current candidate** (fresh recorded evidence that its current-phase paths/boundaries/reachability/severity have not changed since acceptance — the same freshness bar as `pre-existing & out-of-scope`; any delta invalidates the old disposition and reruns classification, split, and risk-owner approval). Throughout, the severity that routes a finding is the **reviewer's assigned severity** (per `references/review-finding-standards.md`), never re-rated by the proposing controller, and the scope-cut safeguards (recorded no-impact evidence, distinct-owner ratification, phase-owned consumed tracker) attach to the cut operation **at any severity** — a finding can never be discarded by self-rating it P2/P3.
|
|
428
|
+
|
|
429
|
+
- When the recurring surface is a **self-adjudication clause** — decidable test: the clause's classification verb has NO named test whose output produces the classification, so the receiver/author judges it — the `keep / delete / narrow / replace` decision must first ask whether an existing mechanical or semi-mechanical test can carry the adjudication: name the test whose output settles the classification, **and the mapping from its output to the classes** — which output means which class — plus the actual result or the evidence contract that will produce it. Naming a test is not routing to it: a test whose output cannot discriminate the classes, or one named with no recorded output-to-class mapping, leaves the adjudication exactly where it was and does not satisfy this rule. A `keep` that retains self-adjudication prose, or a `narrow`/`replace` that adds more prose bindings, is landed only when the decision record (the same-class rule's recorded decision, in the commit body or register row) names the reason no existing test could carry it — a reason left in chat does not count; two challenge rounds attacking the same self-adjudication surface are the signal that prose is the wrong layer — a clause routed to an existing test inherits that test's evidence bar instead of the adjudicator's say-so (worked instance: a review-reception clause that left "is this finding scope-adding" to the receiver survived two rounds of attacks on that self-classification until the classification was routed to the existing structural-minimality test). The same question fires at drafting time for any new reception/discipline-style clause that would grant a self-adjudication.
|
|
430
|
+
|
|
431
|
+
"No *new* P0/P1 this round" and "findings stabilized into the same categories" are necessary but **not sufficient** — a finding repeated unchanged across rounds is still unresolved and still blocks landing until it gets one of the three dispositions. Convergence means *no undispositioned P0/P1 remains*, not *no new P0/P1 appeared*.
|
|
432
|
+
|
|
433
|
+
**A convergence or closure declaration must be written falsifiably.** Name the exact candidate identity it covers, each lane's terminal evidence, the axes/dimensions the closing self-audit actually crossed, and every standing open item by name (e.g. "the final challenge's own fix has not itself been re-challenged") — an aggregate "converged / all axes closed" whose axes are unnamed cannot be checked false and is inconclusive, and any "full X" adjective is scoped to the named axes, never wider. The named enumeration is what lets a fresh challenge falsify the claim by pointing at an un-crossed axis (observed both ways in one program: a self-audit that named its five walked axes was caught exactly one axis short by the final challenge — the naming is why the gap was findable — and the honest handoff that named its open item let the human choose between one fresh pass and explicit risk acceptance instead of inheriting a false "done").
|
|
434
|
+
|
|
435
|
+
Do NOT iterate to zero *findings* — some are intentional design tradeoffs the user already rejected the alternative for, some are genuinely pre-existing. Forcing the finding count to zero either over-corrects or scope-creeps. The bar is zero *undispositioned P0/P1*, which is different from zero findings.
|
|
436
|
+
|
|
437
|
+
The initial independent review plus Agent-initiated challenges share one **Agent-autonomous external-review budget of at most five rounds**. The initial review consumes round 1, so `challenge_budget` is `0..4`. Candidate edits, commits, rebases, amended plans, renamed slices, or a fresh controller invocation do not create more Agent authority. A stateless local controller cannot prove omitted history against a caller that controls its files, so the consuming workflow must preserve the complete review ledger and treat an Agent-created reset as a contract violation.
|
|
438
|
+
|
|
439
|
+
This budget limits only automatic reviewer invocation. It does **not** stop implementation, tests, debugging, or deep self-review, and it does not limit an authenticated human:
|
|
440
|
+
|
|
441
|
+
- A human may request another review or self-review, stop a live review or the overall iteration, commit, or merge. Record human-requested review separately from Agent-autonomous rounds.
|
|
442
|
+
- A human merge/risk decision must come from platform-authenticated authority outside the candidate diff, such as a protected maintainer approval. A repository file, branch flag, CLI argument, environment variable, model statement, or Agent-written note is not human authentication.
|
|
443
|
+
- A narrow authenticated `review_waiver` clears only the review-process gate for the exact candidate and records decision-maker, time, reason, residual findings, and accepted risk.
|
|
444
|
+
- A distinct authenticated `merge_authorization` is the human's final decision for the exact candidate. CI still runs and reports review/build/test/security/compliance failures, but none remains merge-blocking after that decision. Report `merge_authorized_by_human` / `failed_but_human_overridden`; never rewrite any underlying result as `passed` or discard residual findings.
|
|
445
|
+
- A distinct authenticated **`continuation_authorization`** is the third human state, for a budget that is exhausted or has dead-ended: it waives nothing and decides no merge — both lanes stay intact and blocking — the human only authorizes further external rounds toward convergence, each recorded as human-authorized (never counted as Agent-autonomous) and run as a fresh chain bound to the current candidate — a fresh chain restarts the candidate binding, never the history: it carries forward the complete review ledger and every prior round's focuses and dispositions, per the Agent-review-chain fields of `code-review`'s staged review contract. The grant itself is scope-bound, not reusable: it names the granting session and either one exact candidate or, explicitly, this program's rounds to convergence in that session — a candidate or session outside the named scope requires a fresh authorization, so recording rounds as human-authorized can never launder an expired or broader-than-granted continuation. One dead-end is **by design, not an error**: a finding's fix that edits the owner package's own files breaks the review chain's content binding, so the tracker rightly refuses both another autonomous round and a challenge bound to the stale prior result. The recovery is always the same shape — an `interim` checkpoint that names each lane's terminal state and the exact un-run remainder ("challenge not yet run against any candidate", "the final fix is pinned but not re-challenged"), then the human's continuation authorization or their explicit risk acceptance with the record as the disposition trail. Never Agent self-authorization, and never a lane waiver inferred from the human's silence or from the authorization to continue.
|
|
446
|
+
|
|
447
|
+
When a round returns findings, hand them to the implementer before another autonomous review. The implementer verifies each failure path, classifies it as a local fix, false positive, deferred risk, or human decision, and records targeted self-review plus tests. Do not blindly apply every suggestion and do not use the reviewer as the primary defect finder.
|
|
448
|
+
|
|
449
|
+
The mechanical reminder is `self_review_gate`, not prose alone. It records outstanding and satisfied triggers, the narrow blocked actions, and the productive actions that remain allowed. It fires before external review, after findings, after a tracked candidate change, on risk/scope escalation, at the post-budget checkpoint, and before a completion claim. A final passed review stays `completion_gated=true` until the exact-candidate local `complete` checkpoint validates the new deep-self-review plan; this checkpoint invokes no reviewer and grants no human authority.
|
|
450
|
+
|
|
451
|
+
In this gate, `stop`, `terminal`, `abort`, or `revert` applies to the current reviewer lane, readiness claim, or defective dependent slice unless an authenticated human explicitly stops the overall iteration. Repeated root cause, two no-progress attempts, or recurring findings trigger a method change, narrower reproduction, redesign, validation switch, or parked decision item; they never auto-stop unrelated runnable work.
|
|
452
|
+
|
|
453
|
+
At the third Agent-autonomous round, do not start a fourth automatically. If findings remain:
|
|
454
|
+
|
|
455
|
+
- keep fixing local bugs, testing, and self-reviewing under `post_review_budget / human_decision_required`;
|
|
456
|
+
- record the last externally reviewed candidate and every later candidate delta; stale review evidence never certifies changed content;
|
|
457
|
+
- mark findings that need product/design/risk authority as `needs_human_decision`, freeze only dependent work, and continue independent runnable slices;
|
|
458
|
+
- enter `awaiting_human` only when no independent runnable work remains. This is a scheduling state, not task failure and not a human merge prohibition.
|
|
459
|
+
|
|
460
|
+
### Concrete cadence
|
|
461
|
+
|
|
462
|
+
For a focused single-skill change:
|
|
463
|
+
- **Round 1 — independent review**: inspect the self-reviewed candidate broadly.
|
|
464
|
+
- **Round 2 — challenge 1**: after implementer triage, attack the highest-risk unresolved surface with an unprimed prompt.
|
|
465
|
+
- **Round 3 — challenge 2**: verify remaining/new attack paths. This is the final Agent-initiated external round; findings feed the post-budget checkpoint rather than an automatic round 4.
|
|
466
|
+
|
|
467
|
+
Broad extractions use the same three-round Agent budget. Continue their implementation in smaller independent slices after budget exhaustion; a human may explicitly request further review when useful.
|
|
468
|
+
|
|
469
|
+
### Anti-patterns
|
|
470
|
+
|
|
471
|
+
- **Single-round challenge → done**. The round-1 fix-up itself may introduce bugs. Always do at least one re-challenge after a non-trivial fix-up.
|
|
472
|
+
- **Iterating external review until zero findings**. Stop Agent reviewer calls at the configured budget. Stabilized or repeated findings are recorded, triaged, and may cause a method/design change or a parked dependent slice; implementation and independent work continue.
|
|
473
|
+
- **Treating "no new high-severity findings" as "ready to ship" without recording the deferred items**. Deferred findings still need a written reason in the validation log.
|
|
474
|
+
- **Treating every tiny edit as an automatic new external round**. Re-run deep self-review at the required checkpoint; consume another Agent review round only when the retained chain and risk call for it, or when a human explicitly requests one.
|
|
475
|
+
- **Re-running with a softer prompt after fixes**. Use the same adversarial framing every round; weakening the prompt to make later rounds "pass" defeats the purpose.
|
|
476
|
+
|
|
477
|
+
### Recording the loop
|
|
478
|
+
|
|
479
|
+
Add one row per round to the validation log:
|
|
480
|
+
|
|
481
|
+
```
|
|
482
|
+
## Challenge pass — round N (codex exec adversarial)
|
|
483
|
+
- Diff scope: <files / commit range / sha>
|
|
484
|
+
- Findings: N total (a P0 / b P1 / c P2)
|
|
485
|
+
- R0 evidence: <alias_audit_ok | named private-profile result: project-alias/process-retro/both | alias_audit_unavailable or generic_r0_leak_scan_ok => private R0 not run / interim, not landing-clean>
|
|
486
|
+
- Gate-fireability applicability: <yes | no — reason>; Item 9 exercised: <captured prompt/transcript/JSONL locator per the single-pass field above, not a pasted self-assertion | n/a>
|
|
487
|
+
- New since prior round: <count> (subset of above; flag round-introduced bugs)
|
|
488
|
+
- Stabilized: <list of findings carried over without change>
|
|
489
|
+
- Applied: M fixes (commit: <sha>)
|
|
490
|
+
- Deferred: <list with reason>
|
|
491
|
+
- Decision: continue implementation / park dependent slice / await human / human stop, because <reason>
|
|
492
|
+
```
|
|
493
|
+
|
|
494
|
+
A complete dual-track-validated change names every round explicitly. Skipping rounds without recording the decision is the same as not running them.
|
|
495
|
+
|
|
496
|
+
## What does NOT count as dual-track
|
|
497
|
+
|
|
498
|
+
- Two reviews of the same kind (e.g. two consistency reviews from different reviewers — both still miss chaos modes).
|
|
499
|
+
- Self-review by the extractor before submitting (catches obvious things but not adversarial scenarios) — insufficient as independent *evidence*, but NOT skippable: the self-audit-to-convergence preparation above is still required before the gate runs.
|
|
500
|
+
- Static validation script (covers YAML/links/sanitization, not chaos modes).
|
|
501
|
+
- Human PR review without an explicit challenge framing (humans default to consistency review unless prompted).
|
|
502
|
+
|
|
503
|
+
The challenge pass is **structurally different** from review — it must be invoked with an adversarial prompt. The same model can do both passes, but each pass needs its own prompt and its own output.
|
|
504
|
+
|
|
505
|
+
## Cost note
|
|
506
|
+
|
|
507
|
+
Challenge pass at high reasoning typically costs 2-5× review pass in tokens. For a ~3 kLOC reference diff, expect ~250-500k tokens on challenge vs ~50-100k on review. The value of one P0 finding caught before landing dwarfs the cost difference; do not skip on cost.
|