codex-orchestrator 2.0.2 → 2.0.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (221) hide show
  1. package/CHANGELOG.md +44 -426
  2. package/README.md +135 -34
  3. package/dist/src/index.d.ts +11 -1
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/index.js +5 -0
  6. package/dist/src/index.js.map +1 -1
  7. package/dist/src/v2/acceptance-proof.d.ts +3 -0
  8. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  9. package/dist/src/v2/acceptance-proof.js +2 -8
  10. package/dist/src/v2/acceptance-proof.js.map +1 -1
  11. package/dist/src/v2/adapters/gh-issue-adapter.d.ts +5 -3
  12. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  13. package/dist/src/v2/adapters/gh-issue-adapter.js +67 -12
  14. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  15. package/dist/src/v2/adapters/issues.d.ts +16 -2
  16. package/dist/src/v2/adapters/issues.d.ts.map +1 -1
  17. package/dist/src/v2/adapters/issues.js +15 -5
  18. package/dist/src/v2/adapters/issues.js.map +1 -1
  19. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts +1 -0
  20. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts.map +1 -1
  21. package/dist/src/v2/adapters/mission-coordinator-lock.js +5 -1
  22. package/dist/src/v2/adapters/mission-coordinator-lock.js.map +1 -1
  23. package/dist/src/v2/cli-contract.d.ts +3 -3
  24. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  25. package/dist/src/v2/cli-contract.js +9 -1
  26. package/dist/src/v2/cli-contract.js.map +1 -1
  27. package/dist/src/v2/cli.d.ts +24 -0
  28. package/dist/src/v2/cli.d.ts.map +1 -0
  29. package/dist/src/v2/{candidate-cli.js → cli.js} +30 -26
  30. package/dist/src/v2/cli.js.map +1 -0
  31. package/dist/src/v2/code-review-report.d.ts +66 -0
  32. package/dist/src/v2/code-review-report.d.ts.map +1 -0
  33. package/dist/src/v2/code-review-report.js +259 -0
  34. package/dist/src/v2/code-review-report.js.map +1 -0
  35. package/dist/src/v2/codex-process.d.ts +8 -1
  36. package/dist/src/v2/codex-process.d.ts.map +1 -1
  37. package/dist/src/v2/codex-process.js +22 -0
  38. package/dist/src/v2/codex-process.js.map +1 -1
  39. package/dist/src/v2/config.d.ts +4 -3
  40. package/dist/src/v2/config.d.ts.map +1 -1
  41. package/dist/src/v2/config.js +8 -3
  42. package/dist/src/v2/config.js.map +1 -1
  43. package/dist/src/v2/contained-report-operation.d.ts +100 -0
  44. package/dist/src/v2/contained-report-operation.d.ts.map +1 -0
  45. package/dist/src/v2/contained-report-operation.js +200 -0
  46. package/dist/src/v2/contained-report-operation.js.map +1 -0
  47. package/dist/src/v2/containment.d.ts +10 -0
  48. package/dist/src/v2/containment.d.ts.map +1 -1
  49. package/dist/src/v2/containment.js +49 -1
  50. package/dist/src/v2/containment.js.map +1 -1
  51. package/dist/src/v2/direct-delivery.d.ts +96 -0
  52. package/dist/src/v2/direct-delivery.d.ts.map +1 -0
  53. package/dist/src/v2/direct-delivery.js +482 -0
  54. package/dist/src/v2/direct-delivery.js.map +1 -0
  55. package/dist/src/v2/immutable-workflow-publisher.d.ts +40 -0
  56. package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -0
  57. package/dist/src/v2/immutable-workflow-publisher.js +218 -0
  58. package/dist/src/v2/immutable-workflow-publisher.js.map +1 -0
  59. package/dist/src/v2/implementation-reviewer.d.ts +81 -0
  60. package/dist/src/v2/implementation-reviewer.d.ts.map +1 -0
  61. package/dist/src/v2/implementation-reviewer.js +157 -0
  62. package/dist/src/v2/implementation-reviewer.js.map +1 -0
  63. package/dist/src/v2/owner-control-lock.d.ts +41 -0
  64. package/dist/src/v2/owner-control-lock.d.ts.map +1 -0
  65. package/dist/src/v2/owner-control-lock.js +174 -0
  66. package/dist/src/v2/owner-control-lock.js.map +1 -0
  67. package/dist/src/v2/proof-report.d.ts.map +1 -1
  68. package/dist/src/v2/proof-report.js +55 -29
  69. package/dist/src/v2/proof-report.js.map +1 -1
  70. package/dist/src/v2/route-continuations.d.ts +32 -0
  71. package/dist/src/v2/route-continuations.d.ts.map +1 -0
  72. package/dist/src/v2/route-continuations.js +2 -0
  73. package/dist/src/v2/route-continuations.js.map +1 -0
  74. package/dist/src/v2/route-coordinator.d.ts +77 -0
  75. package/dist/src/v2/route-coordinator.d.ts.map +1 -0
  76. package/dist/src/v2/route-coordinator.js +370 -0
  77. package/dist/src/v2/route-coordinator.js.map +1 -0
  78. package/dist/src/v2/route-decision.d.ts +129 -0
  79. package/dist/src/v2/route-decision.d.ts.map +1 -0
  80. package/dist/src/v2/route-decision.js +400 -0
  81. package/dist/src/v2/route-decision.js.map +1 -0
  82. package/dist/src/v2/run-issue.d.ts +64 -6
  83. package/dist/src/v2/run-issue.d.ts.map +1 -1
  84. package/dist/src/v2/run-issue.js +962 -92
  85. package/dist/src/v2/run-issue.js.map +1 -1
  86. package/dist/src/v2/run-store.d.ts +25 -1
  87. package/dist/src/v2/run-store.d.ts.map +1 -1
  88. package/dist/src/v2/run-store.js +129 -4
  89. package/dist/src/v2/run-store.js.map +1 -1
  90. package/dist/src/v2/runtime-assets.d.ts +15 -13
  91. package/dist/src/v2/runtime-assets.d.ts.map +1 -1
  92. package/dist/src/v2/runtime-assets.js +263 -416
  93. package/dist/src/v2/runtime-assets.js.map +1 -1
  94. package/dist/src/v2/runtime.d.ts +17 -9
  95. package/dist/src/v2/runtime.d.ts.map +1 -1
  96. package/dist/src/v2/runtime.js +564 -64
  97. package/dist/src/v2/runtime.js.map +1 -1
  98. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  99. package/dist/src/v2/setup-cli.js +4 -10
  100. package/dist/src/v2/setup-cli.js.map +1 -1
  101. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  102. package/dist/src/v2/setup-runtime.js +19 -131
  103. package/dist/src/v2/setup-runtime.js.map +1 -1
  104. package/dist/src/v2/setup-store.d.ts +0 -5
  105. package/dist/src/v2/setup-store.d.ts.map +1 -1
  106. package/dist/src/v2/setup-store.js +3 -106
  107. package/dist/src/v2/setup-store.js.map +1 -1
  108. package/dist/src/v2/setup.d.ts +6 -43
  109. package/dist/src/v2/setup.d.ts.map +1 -1
  110. package/dist/src/v2/setup.js +13 -192
  111. package/dist/src/v2/setup.js.map +1 -1
  112. package/dist/src/v2/spec-coordinator.d.ts +85 -0
  113. package/dist/src/v2/spec-coordinator.d.ts.map +1 -0
  114. package/dist/src/v2/spec-coordinator.js +88 -0
  115. package/dist/src/v2/spec-coordinator.js.map +1 -0
  116. package/dist/src/v2/spec-delivery.d.ts +143 -0
  117. package/dist/src/v2/spec-delivery.d.ts.map +1 -0
  118. package/dist/src/v2/spec-delivery.js +401 -0
  119. package/dist/src/v2/spec-delivery.js.map +1 -0
  120. package/dist/src/v2/triage-route.d.ts +68 -0
  121. package/dist/src/v2/triage-route.d.ts.map +1 -0
  122. package/dist/src/v2/triage-route.js +223 -0
  123. package/dist/src/v2/triage-route.js.map +1 -0
  124. package/dist/src/v2/waiting-human-coordinator.d.ts +49 -0
  125. package/dist/src/v2/waiting-human-coordinator.d.ts.map +1 -0
  126. package/dist/src/v2/waiting-human-coordinator.js +509 -0
  127. package/dist/src/v2/waiting-human-coordinator.js.map +1 -0
  128. package/dist/src/v2/waiting-human.d.ts +143 -0
  129. package/dist/src/v2/waiting-human.d.ts.map +1 -0
  130. package/dist/src/v2/waiting-human.js +408 -0
  131. package/dist/src/v2/waiting-human.js.map +1 -0
  132. package/dist/src/v2/workflow-assets.d.ts +98 -0
  133. package/dist/src/v2/workflow-assets.d.ts.map +1 -0
  134. package/dist/src/v2/workflow-assets.js +646 -0
  135. package/dist/src/v2/workflow-assets.js.map +1 -0
  136. package/docs/deep-dive.md +275 -52
  137. package/internal-workflow/docs/agents/bug-workflow-routing.md +24 -0
  138. package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
  139. package/internal-workflow/docs/agents/coding-skill-routing.md +123 -0
  140. package/internal-workflow/docs/agents/confidence-rubric.md +65 -0
  141. package/internal-workflow/docs/agents/contract-test-ledger.md +60 -0
  142. package/internal-workflow/docs/agents/review-gates.md +42 -0
  143. package/internal-workflow/docs/agents/review-protocol.md +98 -0
  144. package/internal-workflow/docs/agents/tool-usage.md +88 -0
  145. package/internal-workflow/evals/coding-skill-evals.json +66 -0
  146. package/internal-workflow/manifest.json +1 -0
  147. package/internal-workflow/operations/acceptance-proof/SKILL.md +9 -0
  148. package/internal-workflow/operations/ambiguity-review/SKILL.md +5 -0
  149. package/internal-workflow/operations/code-review/SKILL.md +23 -0
  150. package/internal-workflow/operations/implementation/SKILL.md +24 -0
  151. package/internal-workflow/operations/spec-author/SKILL.md +12 -0
  152. package/internal-workflow/operations/spec-review/SKILL.md +12 -0
  153. package/internal-workflow/operations/triage/SKILL.md +12 -0
  154. package/internal-workflow/profiles/analyst_deep.toml +9 -0
  155. package/internal-workflow/profiles/implementer_standard.toml +9 -0
  156. package/internal-workflow/profiles/proof_agent.toml +8 -0
  157. package/internal-workflow/profiles/reviewer_deep.toml +9 -0
  158. package/internal-workflow/profiles/reviewer_standard.toml +9 -0
  159. package/internal-workflow/schemas/ambiguity-review-v1.json +1 -0
  160. package/internal-workflow/schemas/code-review-v1.json +1 -0
  161. package/internal-workflow/schemas/implementation-report-v1.json +1 -0
  162. package/internal-workflow/schemas/proof-report-v1.json +1 -0
  163. package/internal-workflow/schemas/spec-author-v1.json +1 -0
  164. package/internal-workflow/schemas/spec-review-v1.json +30 -0
  165. package/internal-workflow/schemas/triage-route-v1.json +1 -0
  166. package/internal-workflow/skills/acceptance-proof/agents/openai.yaml +6 -0
  167. package/{internal-skills → internal-workflow/skills}/agent-auto/SKILL.md +6 -1
  168. package/internal-workflow/skills/agent-auto/agents/openai.yaml +6 -0
  169. package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
  170. package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
  171. package/internal-workflow/skills/code-review/SKILL.md +279 -0
  172. package/internal-workflow/skills/code-review/agents/openai.yaml +4 -0
  173. package/internal-workflow/skills/code-review/references/bug-classes.md +56 -0
  174. package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
  175. package/internal-workflow/skills/code-review/references/framework-lenses.md +34 -0
  176. package/internal-workflow/skills/code-review/references/targeted-recipes.md +49 -0
  177. package/internal-workflow/skills/diagnosing-bugs/SKILL.md +138 -0
  178. package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +6 -0
  179. package/internal-workflow/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +41 -0
  180. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +102 -0
  181. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +6 -0
  182. package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +31 -0
  183. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +146 -0
  184. package/internal-workflow/skills/implementation-spec-review/SKILL.md +115 -0
  185. package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +6 -0
  186. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
  187. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
  188. package/internal-workflow/skills/small-task-implementer/SKILL.md +104 -0
  189. package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +6 -0
  190. package/internal-workflow/skills/spec-implementer/SKILL.md +126 -0
  191. package/internal-workflow/skills/spec-implementer/agents/openai.yaml +6 -0
  192. package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
  193. package/internal-workflow/skills/spec-implementer/references/review-loop.md +94 -0
  194. package/internal-workflow/skills/tdd/SKILL.md +72 -0
  195. package/internal-workflow/skills/tdd/agents/openai.yaml +6 -0
  196. package/internal-workflow/skills/tdd/interface-design.md +31 -0
  197. package/internal-workflow/skills/tdd/mocking.md +59 -0
  198. package/internal-workflow/skills/tdd/refactoring.md +10 -0
  199. package/internal-workflow/skills/tdd/tests.md +77 -0
  200. package/internal-workflow/skills/triage/AGENT-BRIEF.md +192 -0
  201. package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +101 -0
  202. package/internal-workflow/skills/triage/SKILL.md +134 -0
  203. package/internal-workflow/skills/triage/agents/openai.yaml +6 -0
  204. package/package.json +14 -8
  205. package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
  206. package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
  207. package/dist/src/v2/adapters/target-activity-fence.js +0 -249
  208. package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
  209. package/dist/src/v2/candidate-cli.d.ts +0 -22
  210. package/dist/src/v2/candidate-cli.d.ts.map +0 -1
  211. package/dist/src/v2/candidate-cli.js.map +0 -1
  212. package/dist/src/v2/legacy-cutover.d.ts +0 -52
  213. package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
  214. package/dist/src/v2/legacy-cutover.js +0 -87
  215. package/dist/src/v2/legacy-cutover.js.map +0 -1
  216. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/SKILL.md +0 -0
  217. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/android.md +0 -0
  218. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/browser.md +0 -0
  219. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/ios.md +0 -0
  220. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/android-lease.mjs +0 -0
  221. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/ios-lease.mjs +0 -0
@@ -0,0 +1,279 @@
1
+ ---
2
+ name: "code-review"
3
+ description: "Evidence-first review of code, PRs, commits, regressions, or review-and-fix work using correctness and standards/cleanup lenses. Auto-fix only qualifying high-confidence severe issues."
4
+ ---
5
+
6
+ # Code Review
7
+
8
+ This skill performs evidence-based code review. It is not a style pass and not a summary. Treat the change as potentially wrong until independent review tracks fail to break it.
9
+
10
+ The review always covers two lenses:
11
+
12
+ - **Correctness reviewer**: bugs, regressions, runtime behavior, security, contracts, caches, concurrency, framework rules, and failure paths.
13
+ - **Spec & standards reviewer**: requested behavior, documented repo standards, architecture fit, duplication, cleanup, and workaround-shaped implementation.
14
+
15
+ The main agent is the coordinator. It pins the review target, assigns both
16
+ lenses to one reviewer for `simple` and `medium`, and splits them across two
17
+ independent reviewers only for `high`. It verifies the strongest findings,
18
+ applies only safe fixes, and returns a concise findings-first report.
19
+
20
+ Full review is bounded to the settled diff, its authority, changed owners, and
21
+ callers or contracts with plausible fan-out. `Full` means complete coverage of
22
+ that assigned scope once; it does not mean a repository-wide audit. Do not load
23
+ unrelated modules or activate optional lenses without a diff signal or mandatory
24
+ Review Focus.
25
+
26
+ ## When To Use
27
+
28
+ Use this skill when the user asks for:
29
+
30
+ - code review, PR review, commit audit, regression scan, or bug hunt
31
+ - review since a branch, commit, tag, merge-base, or working tree state
32
+ - review and fix of critical or high-severity high-confidence defects
33
+ - framework-focused review such as `NestJS`, `Next.js`, `Flutter`, or `Dart`
34
+
35
+ For every implementation profile, the spec/standards lens includes bounded
36
+ cleanup for duplication, obsolete paths, workaround branches, and unjustified
37
+ abstractions. High-risk work assigns that lens to its own reviewer in the same
38
+ parallel final wave. A concrete evidenced simplification risk named by the
39
+ user, approved source, or repo policy amplifies this lens inside the same review
40
+ activation; it never creates a separate cleanup gate.
41
+
42
+ Exception for approved spec execution: follow
43
+ `../spec-implementer/references/review-loop.md`. Intermediate code-review
44
+ checkpoints activate only their assigned Review Focus; final cleanup coverage
45
+ uses the durable Review Plan and canonical Defect Ledger.
46
+
47
+ ## Implementation Review Adapter
48
+
49
+ When this skill is called from `$spec-implementer`:
50
+
51
+ - read `../spec-implementer/references/review-loop.md` and the persisted
52
+ `## Implementation Review State`
53
+ - accept the scheduled mode, session, revision, and lenses from that state
54
+ - pin the target and give reviewers the owner-defined capsule
55
+ - return the usable result and stable defect updates to the executor
56
+ - keep cleanup findings in the spec/standards lineage and canonical Defect Ledger
57
+
58
+ Do not infer a fresh review loop, choose another mode, or make the owner's
59
+ terminal decision inside this Adapter.
60
+
61
+ ## Progressive References
62
+
63
+ Read only the references the current review needs:
64
+
65
+ - Detailed bug classes: `references/bug-classes.md`
66
+ - Cleanup lens method: `references/cleanup-lens.md`
67
+ - Framework lenses for Next.js, NestJS, Flutter, and Dart: `references/framework-lenses.md`
68
+ - Targeted recipes for recurring diff shapes: `references/targeted-recipes.md`
69
+ - Contract test ledger: `../../docs/agents/contract-test-ledger.md`
70
+ - Shared confidence rubric: `../../docs/agents/confidence-rubric.md`
71
+
72
+ Load `references/framework-lenses.md` when the user names a framework or files/configs strongly imply one. Load `references/targeted-recipes.md` and `../../docs/agents/contract-test-ledger.md` when the diff shape matches new fields, retries, DTO/schema/runtime contracts, caches, state merge precedence, ordering, evidence/snapshots, determinism, or aggregation summaries. Load `references/bug-classes.md` for substantial reviews or broad bug hunts.
73
+ Load `references/cleanup-lens.md` when the spec/standards lens is assigned. Use
74
+ its bounded method by default and its amplified method only for a concrete
75
+ evidenced simplification risk supplied as mandatory Review Focus.
76
+
77
+ ## Coordinator Workflow
78
+
79
+ ### 1. Pin The Review Target
80
+
81
+ Identify exactly what is being reviewed.
82
+
83
+ - If the user gave a fixed point, use it directly: branch, commit SHA, tag, `main`, `HEAD~5`, etc.
84
+ - If they did not, infer from context:
85
+ - current uncommitted work: `git diff` plus staged diff if relevant
86
+ - branch review: `git diff <base>...HEAD`
87
+ - commit review: `git show <commit>`
88
+ - If there is no safe inference, ask one short question: "Review against which branch or commit?"
89
+
90
+ Capture:
91
+
92
+ - `git status --short`
93
+ - diff stat
94
+ - the exact diff command used
95
+ - commit list when reviewing a branch range
96
+ - changed files and nearest related tests/docs
97
+
98
+ Use three-dot diff for branch/base reviews: `git diff <fixed-point>...HEAD`.
99
+
100
+ ### 2. Discover Spec And Standards
101
+
102
+ Do this before reviewer tracks so both tracks receive bounded inputs.
103
+
104
+ Spec sources, in priority order:
105
+
106
+ 1. Issue or PR references in commit messages, branch names, PR metadata, or user prompt.
107
+ 2. A spec/PRD/plan path supplied by the user.
108
+ 3. Matching files under `docs/`, `specs/`, `.scratch/`, or local issue folders.
109
+ 4. If none exists, continue and mark the spec axis as "no spec found" instead of inventing requirements.
110
+
111
+ Standards sources:
112
+
113
+ - `AGENTS.md`, `CLAUDE.md`, `CONTRIBUTING.md`
114
+ - `CONTEXT.md`, context maps, domain docs, ADRs
115
+ - `STYLE.md`, `STANDARDS.md`, style guides, review checklists
116
+ - `.editorconfig`, ESLint, Biome, Prettier, TypeScript, analyzer, or framework configs
117
+ - relevant test files and existing examples in the touched area
118
+
119
+ Machine-enforced config matters as context, but do not spend review findings on issues a required formatter/linter would already catch unless the tool is absent or failing.
120
+
121
+ ### 3. Activate Lenses
122
+
123
+ Before deep review, decide which lenses apply.
124
+
125
+ - Always activate the general correctness and spec/standards lenses.
126
+ - Always apply bounded cleanup inside spec/standards; amplify it only for a
127
+ concrete evidenced Review Focus, never from size or risk labels alone.
128
+ - Add framework lenses when explicit or strongly implied by files/configs.
129
+ - Add targeted recipes when the diff shape matches them.
130
+ - If the user, plan, or implementation spec provides `Review Focus`, treat each listed lens, targeted recipe, invariant, and risk as mandatory. Do not replace it with a generic review; report any focus item that cannot be verified as a verification gap.
131
+ - If inference is uncertain, say so and continue with the best-supported review instead of pretending certainty.
132
+
133
+ Mandatory delta lenses:
134
+
135
+ - **Contract Delta Review:** If the diff expands a shared contract, enum/status, schema, DTO, permission, capability, token, or scope, search old consumers and report `changed contract -> searched call sites -> risky fallback/default branches -> tests or fix`.
136
+ - **Backend Trust Boundary Review:** If the diff adds a credential, OAuth scope, permission, role-sensitive write, or external capability, verify the backend-side guard. UI controls, query/state flags, and caller intent are not authorization.
137
+
138
+ Common framework signals:
139
+
140
+ - **Next.js**: `next.config.*`, `app/**`, `pages/**`, route handlers, server actions, `use server`, `use client`, revalidation APIs.
141
+ - **NestJS**: `@nestjs/*`, controllers, providers, modules, guards, pipes, interceptors, DTOs, `nest-cli.json`.
142
+ - **Flutter**: `pubspec.yaml`, `lib/**`, widgets, navigation, bloc/provider/riverpod/notifiers.
143
+ - **Dart**: `.dart`, `Future`, `Stream`, isolates, generated serializers, null-safety constructs.
144
+
145
+ ### 4. Run Profile-Selected Reviewers
146
+
147
+ For `simple`, run one `reviewer_fast`; for `medium`, run one
148
+ `reviewer_standard`. Assign that child both correctness and spec/standards
149
+ lenses in one bounded Full review. For `high`, run two `reviewer_deep` children
150
+ in parallel with one disjoint lens each. Invoking `$code-review` authorizes
151
+ these reviewers. Preserve every fulfilled launch handle if a parallel peer
152
+ fails, then close all launched children. Tell every reviewer not to edit or
153
+ revert unrelated work.
154
+
155
+ For spec-driven checkpoints, obey the track assignment in the persisted Review
156
+ Plan instead of automatically launching both default tracks.
157
+
158
+ Medium is the normal review profile. API, persistence, multiple files, or a
159
+ shared-looking name do not select `high` unless evidence proves both a material
160
+ failure consequence and an uncertainty amplifier.
161
+
162
+ When already inside an assigned reviewer child, execute its assigned lens set
163
+ inline and return it to root; do not spawn a grandchild.
164
+ If the user forbids delegation, report the independent review gate as waived or
165
+ unavailable according to the parent workflow; root must not self-review inline.
166
+
167
+ #### Correctness Reviewer Brief
168
+
169
+ Include the exact diff command or commit under review, changed file list, commit list, active references, any `Review Focus`, and an instruction to read surrounding execution paths.
170
+
171
+ > Review runtime correctness adversarially. Hunt concrete bugs in control flow, state, async/concurrency, security, auth, contracts, schemas, caching, persistence, UI/API integration, and framework-specific behavior. If a Review Focus is provided, explicitly test the named risks first, such as duplicate side effects, retry/idempotency, ordering, source-of-truth ownership, partial failure, DTO/schema drift, or false user-facing state. For each finding, provide file/line, trigger path, impact, why guards do not prevent it, severity, and confidence. Do not report style nits or generic "needs tests" comments.
172
+ > For bugfixes, check claim boundaries: what changed tests prove, what they do not prove, and whether sibling execution paths can still violate the claimed invariant.
173
+ > When Contract Delta Review applies, follow new values through old consumers and default/fallback branches before trusting local tests. When Backend Trust Boundary Review applies, verify the server-side authorization predicate at the write/callback boundary.
174
+
175
+ #### Spec & Standards Reviewer Brief
176
+
177
+ Include the exact diff command or commit under review, spec source path/content or "no spec found", standards source list, changed file list, commit list, any `Review Focus`, and an instruction to cite the spec or standard behind each finding.
178
+
179
+ > Review the change against the requested work and repo standards. Check missing requirements, partial behavior, scope creep, undocumented contract changes, architecture drift, duplicate source-of-truth logic, dead/legacy branches, and workaround-shaped implementation. If a Review Focus is provided, explicitly verify each named ownership, scope, validation, and invariant risk against the spec. Cite the spec or standard when available. If there is no spec, skip requirement claims and focus on documented standards and architecture evidence.
180
+ > Apply `references/cleanup-lens.md`: use bounded cleanup by default, or the amplified method when a concrete evidenced simplification risk is mandatory Review Focus. Keep cleanup findings in this review's normal Defect Ledger. Do not create a cleanup-only verdict or separate pass.
181
+
182
+ ### 5. Aggregate And Verify
183
+
184
+ The coordinator must not blindly relay reviewer output.
185
+
186
+ 1. Deduplicate findings across tracks.
187
+ 2. Re-read the relevant code for the strongest findings.
188
+ 3. Drop findings that lack a concrete trigger path.
189
+ 4. Reclassify severity/confidence using `../../docs/agents/confidence-rubric.md` if evidence does not support the label.
190
+ 5. For real contract defects, identify the missing or inadequate Contract Test Ledger invariant when TDD/spec evidence is available.
191
+ 6. Confirm every mandatory `Review Focus` item and mandatory delta lens was reviewed; if not, report the unverified item as a verification gap.
192
+ 7. Decide whether auto-fix is allowed.
193
+ 8. Run the narrowest meaningful verification after any fix.
194
+
195
+ Keep the two axes visible in your own notes, but present the final report by severity unless the user explicitly asked for side-by-side Standards/Spec output.
196
+
197
+ After repairs, verify medium or low findings through the coordinator's direct
198
+ failure-path check plus affected validation. Launch Closure only for the
199
+ severity, protected-contract, or invalidated-coverage triggers owned by
200
+ `review-protocol.md`. For a scheduled Closure, send the bounded capsule to the
201
+ selected reviewer lineage and return its defect updates.
202
+
203
+ ## Evidence Standard
204
+
205
+ A valid finding explains:
206
+
207
+ - what breaks
208
+ - why it breaks
209
+ - the input, sequence, role, tenant, environment, or timing that triggers it
210
+ - where the defect lives
211
+ - why existing guards do not prevent it
212
+ - severity and confidence
213
+
214
+ Do not file:
215
+
216
+ - style nits disguised as correctness issues
217
+ - speculative races without a shared-state path
218
+ - generic "needs tests" comments without a concrete regression risk
219
+ - architecture discomfort without wrong ownership, duplication, leakage, or a regression path
220
+ - performance comments without a hot path or failure mode
221
+
222
+ ## Auto-Fix Policy
223
+
224
+ Automatically fix only when all are true:
225
+
226
+ - severity is critical or high
227
+ - confidence is high under `../../docs/agents/confidence-rubric.md`
228
+ - root cause is clear
229
+ - correct fix is narrow and low-risk
230
+ - fix matches local project patterns
231
+ - verification is available, or the edit is obviously safe and syntax-checkable
232
+
233
+ When auto-fixing:
234
+
235
+ - patch only the bug
236
+ - add/update behavior tests when regression risk is meaningful and the codebase supports it; for contract defects, add the missing ledger invariant first when a ledger exists or is being created
237
+ - rerun relevant verification
238
+ - never revert unrelated user changes
239
+
240
+ Do not auto-fix ambiguous semantics, product decisions, broad refactors, or low-confidence concerns. Report them with evidence.
241
+
242
+ ## Output Contract
243
+
244
+ For review-only tasks:
245
+
246
+ 1. Findings first, ordered by severity.
247
+ 2. File and line references for each finding.
248
+ 3. Trigger path, impact, severity, confidence, and evidence.
249
+ 4. Open questions, assumptions, or verification gaps.
250
+ 5. Short summary only after findings.
251
+
252
+ For review-and-fix tasks:
253
+
254
+ 1. State which critical or high-severity high-confidence issues were fixed.
255
+ 2. Report remaining findings that were not fixed.
256
+ 3. Give one short verification note and any blocked checks.
257
+ 4. Keep the closing summary user-facing and outcome-based.
258
+
259
+ If there are no findings, say so clearly and mention residual test or verification gaps.
260
+
261
+ For spec-driven review checkpoints or final review gates, include a compact review handoff that can feed the executor's Final Risk Handoff: reviewed target, Review Focus status, high/critical findings fixed or remaining, skipped checks, and residual verification gaps. Keep findings first.
262
+
263
+ If inline review comments are requested, emit one `::code-comment{...}` directive per actionable finding.
264
+
265
+ ## Tooling Defaults
266
+
267
+ - Use `rg`/`rg --files` for code search.
268
+ - Prefer parallel reads for status, diff, changed files, related modules, tests, repo docs, and standards.
269
+ - Use official docs or Context7 only when a finding depends on version-sensitive framework/library behavior.
270
+ - Run the narrowest meaningful tests first, then required lint/build/analyzer checks for touched areas.
271
+ - Keep raw command output out of the final response unless the user asks for it.
272
+
273
+ ## Decision Defaults
274
+
275
+ - If the user says "review", default to findings-first review.
276
+ - If the user says "review and fix", auto-fix only critical or high-severity high-confidence issues that satisfy the auto-fix policy.
277
+ - If the user provides a framework focus, explicitly use it.
278
+ - If no focus is provided, infer active lenses from the diff and say when they materially affected findings.
279
+ - Ask clarification only when the correct review target or fix would otherwise be risky or ambiguous.
@@ -0,0 +1,4 @@
1
+ interface:
2
+ display_name: "Code Review"
3
+ short_description: "Profile-routed correctness and standards review"
4
+ default_prompt: "Use $code-review for one final profile-selected wave; high risk uses two parallel reviewers and the spec/standards lens includes bounded cleanup."
@@ -0,0 +1,56 @@
1
+ # Code Review Bug Classes
2
+
3
+ Use this reference for substantial reviews and broad bug hunts.
4
+
5
+ ## Review Mindset
6
+
7
+ - Findings first; summary second.
8
+ - Prefer a few high-signal findings over a long speculative list.
9
+ - Read enough surrounding code to understand the real execution path.
10
+ - Include pre-existing bugs only when they materially affect the reviewed path.
11
+ - Treat workaround-shaped code as a review target.
12
+ - Treat duplicated business rules, builders, cleanup rules, normalization, cache keys, restart paths, and persistence math as likely drift vectors.
13
+ - Treat "tests pass" as insufficient when contracts, ownership, or source-of-truth logic moved.
14
+ - Do not stop after the edited hunk unless the change is truly trivial.
15
+
16
+ ## Non-Negotiable Passes
17
+
18
+ Every substantial review covers:
19
+
20
+ 1. **Scope**: exact diff/commit/branch/files under review.
21
+ 2. **Execution path**: entrypoints, callers, callees, side effects, state writes, network/DB boundaries, cleanup.
22
+ 3. **Architecture fit**: correct owning layer, no leaky abstractions, no duplicate source of truth, no symptom patch in the wrong layer.
23
+ 4. **Invariants**: validation, authorization, ordering, idempotency, data shape, permissions, cache visibility, and lifecycle rules.
24
+ 5. **Failure modes**: empty, null, duplicate, stale, delayed, retried, timeout, cancellation, partial failure, concurrent execution.
25
+ 6. **Blast radius**: DTOs, schemas, cache keys, feature flags, config, metrics, tests, consumers, migrations, and backward compatibility.
26
+ 7. **Verification**: narrow tests/lint/build/analyzer for touched areas, or a clear note when a check cannot run.
27
+
28
+ ## Bug Classes To Hunt
29
+
30
+ - Control-flow mistakes: wrong branch, inverted predicate, missing return, incorrect default, off-by-one, pagination errors.
31
+ - State/lifecycle bugs: stale state, forgotten reset, double writes, orphaned cleanup, leaks, inconsistent derived state.
32
+ - Async/concurrency bugs: missing `await`, race windows, non-idempotent retries, shared mutable state, closure mutation in retryable callbacks.
33
+ - Partial-failure bugs: one side effect succeeds while durable follow-up fails.
34
+ - Contract drift: DTO/schema/type mismatch, nullable changes, enum drift, serialization differences, `ObjectId`/`Date` runtime mismatch.
35
+ - Cache bugs: stale reads, bad cache keys, cross-tenant leakage, missed invalidation, TTL mismatch.
36
+ - Auth regressions: missing actor checks, wrong tenant scope, optional-param bypass, client-only enforcement.
37
+ - Data correctness: timezones, DST, rounding, units, locale parsing, duplicate filtering, sort instability.
38
+ - UI/API regressions: optimistic state not rolled back, loading/error/empty states broken, incompatible response handling.
39
+ - Observability gaps: critical failures suppressed without logs, metrics, retries, or surfaced errors.
40
+ - Workarounds: magic delays, one-off guards, forced ordering, duplicated normalization, cross-layer patches.
41
+ - Architecture drift: business logic in the wrong layer, source-of-truth splits, unnecessary coupling.
42
+ - Deep-module drift: shallow pass-through modules, hypothetical one-adapter seams, lost locality, weak leverage, or tests reaching inside the implementation instead of crossing the Module Interface.
43
+
44
+ ## Mandatory Questions
45
+
46
+ Ask the relevant subset for each meaningful path:
47
+
48
+ - What happens with empty, null, duplicated, delayed, retried, stale, or out-of-order input?
49
+ - What happens when a dependency throws, times out, returns partial data, or returns stale data?
50
+ - Can two executions interleave and corrupt state, leak data, or duplicate work?
51
+ - Are validation and authorization enforced before side effects?
52
+ - Does the change rely on a type guarantee that is weaker at runtime?
53
+ - Do all readers and writers agree on field names, units, nullability, and semantics?
54
+ - If state is rebuilt, cloned, merged, normalized, serialized, or persisted, are all required fields preserved?
55
+ - Can feature flags, optional params, defaults, or fallback paths bypass intended behavior?
56
+ - If failure happens halfway through, what durable state remains and who repairs it?
@@ -0,0 +1,52 @@
1
+ # Cleanup Lens
2
+
3
+ Use this method inside the code-review spec/standards lens. Its job is to reduce
4
+ maintenance surface without changing required observable behavior. It is never
5
+ a standalone gate, activation, verdict, or coverage class.
6
+
7
+ ## Depth
8
+
9
+ - **Bounded:** inspect material additions and replacements for obvious duplicate
10
+ owners, obsolete paths, workaround branches, dead code, and unjustified
11
+ abstractions.
12
+ - **Amplified:** when mandatory Review Focus names a concrete evidenced
13
+ simplification risk, inventory every material addition, replacement,
14
+ compatibility path, and runtime owner related to that risk. Do not amplify
15
+ from file count, implementation size, or review profile alone.
16
+
17
+ ## In Scope
18
+
19
+ - duplicated logic, sources of truth, registrations, or old/new paths kept in parallel
20
+ - dead helpers, flags, adapters, branches, comments, tests, or documentation left by the change
21
+ - workaround-shaped conditionals, magic ordering, symptom patches, and unnecessary state
22
+ - compatibility or fallback behavior without current repository, source-authority, or production evidence
23
+ - new services, events/listeners, adapters, or indirection with one current consumer and only speculative reuse
24
+ - ownership placement when moving or deleting code restores an existing owner without redesigning the system
25
+
26
+ Do not turn functional correctness, security, performance, product decisions,
27
+ missing regression proof, broad architecture redesign, or style preferences
28
+ into cleanup findings. Route them to the applicable code-review lens.
29
+
30
+ ## Method
31
+
32
+ Classify each material complexity decision:
33
+
34
+ - `KEEP`: a current invariant requires it and evidence or a boundary proof supports it.
35
+ - `SIMPLIFY`: required behavior can use a smaller existing seam or fewer states or branches.
36
+ - `REMOVE`: no current behavior, authority, consumer, or compatibility evidence requires it.
37
+
38
+ A one-producer/one-consumer abstraction defaults to `SIMPLIFY` unless a concrete
39
+ lifecycle, transaction, dependency-direction, or fanout invariant requires the
40
+ boundary. Do not request a new abstraction unless it reduces current
41
+ duplication or restores an existing owner now.
42
+
43
+ Report a finding only with exact evidence, concrete maintenance cost, and a
44
+ behavior-preserving fix. Uncertain removal is a non-blocking follow-up. In
45
+ amplified mode, include concise `KEEP | SIMPLIFY | REMOVE` decisions in the
46
+ spec/standards handoff; bounded mode needs decisions only when they explain a
47
+ finding or protected invariant.
48
+
49
+ Cleanup repairs retain the same defect IDs. Medium or low cleanup repairs use coordinator verification plus affected validation. Launch affected-lens Closure
50
+ only when the shared review protocol requires it for severity, protected
51
+ contract impact, or invalidated mandatory coverage; add correctness whenever
52
+ that Closure repair may alter observable behavior.
@@ -0,0 +1,34 @@
1
+ # Code Review Framework Lenses
2
+
3
+ Load this reference when a framework is explicit or strongly implied by files/configs.
4
+
5
+ ## Next.js
6
+
7
+ - Server/client component boundaries, hydration assumptions, browser APIs on server.
8
+ - `fetch` cache modes, route segment caching, `revalidatePath`/`revalidateTag`, stale data behavior.
9
+ - Route handlers/server actions: input validation, auth, serialization of dates/errors/nullables.
10
+ - Loading, empty, error, navigation, optimistic update, rollback, back/refresh behavior.
11
+ - SSR/client mismatch from locale, feature flags, sessions, params, and search params.
12
+
13
+ ## NestJS
14
+
15
+ - Request flow through controller, guard, pipe, interceptor, service, repository, events/queues.
16
+ - DTO validation versus runtime payload, especially transforms, enums, optionals, nested objects.
17
+ - Server-side auth and tenant scope before side effects.
18
+ - Transaction boundaries, partial writes, outbox/queue ordering, retry safety.
19
+ - Exception mapping, HTTP status behavior, cache decorators, request scope, singleton mutable state, cron/queue concurrency.
20
+
21
+ ## Flutter
22
+
23
+ - Widget lifecycle: `init`, `build`, async callbacks, `dispose`, state updates after unmount.
24
+ - Navigation/back stack, dialog/sheet dismissal, duplicate submit from rapid taps.
25
+ - State management transitions, stale emissions, missed listeners, dropped fields.
26
+ - Loading, offline, empty, error, retry states.
27
+ - Platform permissions, keyboard insets, small-screen overflow, animation/controller/stream cleanup.
28
+
29
+ ## Dart
30
+
31
+ - Null-safety assumptions, casts, `late`, `!`, JSON parsing, generated serializers.
32
+ - `Future`, `Stream`, timer, isolate ordering, cancellation, uncaught errors.
33
+ - Equality, copy semantics, immutable updates, mutation during iteration.
34
+ - Date/duration/timezone handling, numeric precision, parsing, generic runtime assumptions.
@@ -0,0 +1,49 @@
1
+ # Code Review Targeted Recipes
2
+
3
+ Load this reference when the diff shape matches one of these recurring risks.
4
+
5
+ ## Activation
6
+
7
+ - new field or contract: run **New Field Threading**
8
+ - transaction/retry/queue/background job: run **Retryable Persistence**
9
+ - DTO/schema/type/persistence change: run **Runtime Contract Alignment**
10
+ - cache or invalidation change: run **Cache Coherence**
11
+ - server/AI/cache/fallback state merged into client state: run **State Merge Precedence**
12
+ - preview/trace/summary/score/winner field: run **Aggregation Cardinality**
13
+
14
+ ## New Field Threading
15
+
16
+ 1. Search for the field name.
17
+ 2. Search nearby `build`, `apply`, `adjust`, `replan`, `normalize`, `merge`, `compile`, `serialize`, `clone`, and fallback helpers.
18
+ 3. Confirm the field survives primary construction, correction/replan paths, persistence reload, response serialization, and externally visible tests.
19
+
20
+ ## Retryable Persistence
21
+
22
+ 1. Read the whole transaction/retry callback.
23
+ 2. List outer-scope variables referenced inside it.
24
+ 3. Flag mutation of outer arrays, counters, iterators, derived inputs, or one-shot streams that changes behavior on retry.
25
+ 4. Compare transactional and non-transactional paths for parity.
26
+
27
+ ## Runtime Contract Alignment
28
+
29
+ 1. Compare DTO validation, internal type, persistence schema, normalization, API response shape, and frontend/state consumers.
30
+ 2. Treat `string` versus `ObjectId`, `Date` versus string, enum drift, and nested optional mismatches as review targets.
31
+ 3. Prefer boundary validation when malformed input should be rejected.
32
+
33
+ ## Cache Coherence
34
+
35
+ 1. Identify cache key inputs, tenant/user scope, feature flags, locale, and permissions.
36
+ 2. Confirm invalidation covers every write path and revalidation timing matches user-visible expectations.
37
+ 3. Check stale reads, cross-tenant leakage, and optimistic UI/cache rollback.
38
+
39
+ ## State Merge Precedence
40
+
41
+ 1. Identify precedence between user-entered state, server state, AI-generated state, fallback state, cache state, and defaults.
42
+ 2. Verify omitted fields, empty objects, `false`, `0`, or `null` cannot erase stronger user choices unless intended.
43
+ 3. Check both backend merge logic and frontend state update logic.
44
+
45
+ ## Aggregation Cardinality
46
+
47
+ 1. Determine whether source data is global, per-group, per-item, per-step, or per-tenant.
48
+ 2. Verify top-level summary/trace/score/winner fields do not collapse multiple meaningful entities into one misleading value.
49
+ 3. Treat `.find()`, first-selected, `sort()[0]`, and flat-array winner logic as suspicious when multiple groups can each have a result.
@@ -0,0 +1,138 @@
1
+ ---
2
+ name: diagnosing-bugs
3
+ description: Debug hard, flaky, unclear, or performance bugs through reproduce, minimize, hypothesize, instrument, fix, and regression-test. Trigger for nondeterministic failures, unclear breakage, or performance regressions.
4
+ ---
5
+
6
+ # Diagnosing Bugs
7
+
8
+ A discipline for hard bugs. Skip phases only when explicitly justified.
9
+
10
+ Routing precedence: use `$CODEX_ORCHESTRATOR_WORKFLOW_ROOT/docs/agents/bug-workflow-routing.md`. This skill owns the feedback loop; after the loop proves the bug, return to the original intent: diagnosis-only output or implementation through `code-debugger`.
11
+
12
+ Use `$CODEX_ORCHESTRATOR_WORKFLOW_ROOT/docs/agents/confidence-rubric.md` when deciding whether a hypothesis, root cause, or fix is high-confidence enough to act on. Low-confidence concerns are questions or verification gaps, not proven causes.
13
+
14
+ When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
15
+
16
+ ## Phase 1 — Build a feedback loop
17
+
18
+ **This is the skill.** Everything else is mechanical. If you have a **tight** pass/fail signal for the bug — one that goes red on _this_ bug — you will find the cause; bisection, hypothesis-testing, and instrumentation all just consume it. If you don't have one, no amount of staring at code will save you.
19
+
20
+ Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
21
+
22
+ ### Ways to construct one — try them in roughly this order
23
+
24
+ 1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
25
+ 2. **Curl / HTTP script** against a running dev server.
26
+ 3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
27
+ 4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
28
+ 5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
29
+ 6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
30
+ 7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
31
+ 8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
32
+ 9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
33
+ 10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
34
+
35
+ Build the right feedback loop, and the bug is 90% fixed.
36
+
37
+ ### Tighten the loop
38
+
39
+ Treat the loop as a product. Once you have _a_ loop, **tighten** it:
40
+
41
+ - Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
42
+ - Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
43
+ - Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
44
+
45
+ A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight — a debugging superpower.
46
+
47
+ ### Non-deterministic bugs
48
+
49
+ The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
50
+
51
+ ### When you genuinely cannot build a loop
52
+
53
+ Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
54
+
55
+ ### Completion criterion — a tight loop that goes red
56
+
57
+ Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command** — a script path, a test invocation, a curl — that you have **already run at least once** (paste the invocation and its output), and that is:
58
+
59
+ - [ ] **Red-capable** — it drives the actual bug code path and asserts the **user's exact symptom**, so it can go red on this bug and green once fixed. Not "runs without erroring" — it must be able to _catch this specific bug_.
60
+ - [ ] **Deterministic** — same verdict every run (flaky bugs: a pinned, high reproduction rate, per above).
61
+ - [ ] **Fast** — seconds, not minutes.
62
+ - [ ] **Agent-runnable** — you can run it unattended; a human in the loop only via `scripts/hitl-loop.template.sh`.
63
+
64
+ If you catch yourself reading code to build a theory before this command exists, **stop — jumping straight to a hypothesis is the exact failure this skill prevents.** No red-capable command, no Phase 2.
65
+
66
+ ## Phase 2 — Reproduce + minimise
67
+
68
+ Run the loop. Watch it go red — the bug appears.
69
+
70
+ Confirm:
71
+
72
+ - [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
73
+ - [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
74
+ - [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
75
+
76
+ ### Minimise
77
+
78
+ Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut — keep only what's load-bearing for the failure.
79
+
80
+ Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer moving parts left to suspect) and becomes the clean regression test in Phase 5.
81
+
82
+ Done when **every remaining element is load-bearing** — removing any one of them makes the loop go green.
83
+
84
+ Do not proceed until you have reproduced **and** minimised.
85
+
86
+ ## Phase 3 — Hypothesise
87
+
88
+ Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
89
+
90
+ Each hypothesis must be **falsifiable**: state the prediction it makes.
91
+
92
+ > Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
93
+
94
+ If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
95
+
96
+ **Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
97
+
98
+ ## Phase 4 — Instrument
99
+
100
+ Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
101
+
102
+ Tool preference:
103
+
104
+ 1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
105
+ 2. **Targeted logs** at the boundaries that distinguish hypotheses.
106
+ 3. Never "log everything and grep".
107
+
108
+ **Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
109
+
110
+ **Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
111
+
112
+ ## Phase 5 — Fix + regression test
113
+
114
+ Write the regression test **before the fix** — but only if there is a **correct seam** for it.
115
+
116
+ A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
117
+
118
+ **If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
119
+
120
+ If a correct seam exists:
121
+
122
+ 1. Turn the minimised repro into a failing test at that seam.
123
+ 2. Watch it fail.
124
+ 3. Apply the fix.
125
+ 4. Watch it pass.
126
+ 5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
127
+
128
+ ## Phase 6 — Cleanup + post-mortem
129
+
130
+ Required before declaring done:
131
+
132
+ - [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
133
+ - [ ] Regression test passes (or absence of seam is documented)
134
+ - [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
135
+ - [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
136
+ - [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
137
+
138
+ **Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `$improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
@@ -0,0 +1,6 @@
1
+ interface:
2
+ display_name: "Diagnosing Bugs"
3
+ short_description: "Run a disciplined hard-bug investigation loop"
4
+ default_prompt: "Use $diagnosing-bugs to reproduce, minimise, diagnose, and verify this difficult bug."
5
+ policy:
6
+ allow_implicit_invocation: true